Commit Graph
1726 Commits
Author SHA1 Message Date
remco f496d48cee Add renovate.json
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / tests-ui (pull_request) Skipped
CI / preflight (pull_request) Skipped
CI / check (pull_request) Failing after 21s
CI / tests-unit (pull_request) Skipped
CI / tests-integration (pull_request) Skipped
CI / deploy (pull_request) Skipped
2026-10-11 17:00:08 +00:00
openhands 7d8874c8cd fix: biome lint suppressions + admin-badges.ts formatting
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Failing after 23s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / preflight (push) Skipped
CI / tests-ui (push) Skipped
CI / deploy (push) Skipped
2026-10-11 18:48:26 +02:00
openhands a06db72c65 Update: admin panel fixes, badge icons, studio improvements, gitea runner v5
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
2026-10-11 18:44:02 +02:00
openhands afc8909d3f fix(studio): the furniture list was rendered at opacity 0 and never painted
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 2m2s
CI / tests-unit (push) Successful in 2m3s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m40s
The Studio furniture pane was present in the DOM the whole time and still
looked empty. Every row existed, had real dimensions, and its image loaded
with a 200 — and none of it was visible.

The pane was wrapped in a `motion/react-m` element declared with
`initial={{ opacity: 0 }}` and `animate={{ opacity: 1 }}`. That minimal entry
renders the element but never runs the animation, so the inline style stayed at
opacity 0 and the content was painted transparently forever. Measured on the
live release: the table sat at opacity 1 directly inside a wrapper pinned at
`style="opacity: 0"`.

It went unnoticed because playwright.ui.config.ts sets reducedMotion to
"reduce", under which the animation is skipped and the element lands straight
on its final value. The existing Studio specs therefore passed while the real
browser showed nothing. That also means the perf win from the minimal entry
was never actually delivering a working fade — it only hid the breakage.

The decorative 150ms fade is now a CSS animation (.studio-list-fade-in). A CSS
animation cannot strand content this way: if it never runs, the element is
simply opaque. motion/react-m had exactly one usage in the app and is gone;
the four files that use the full motion/react are untouched and unaffected.

Adds e2e/ui/studio-visibility.spec.ts, which opts out of reduced motion and
asserts no ancestor of a furniture row is faded below 0.9. Verified it fails
on the old code with `Received: 0` and passes on the new, so this cannot
regress silently again.
2026-10-11 18:05:13 +02:00
openhands 990ebdb158 fix(pwa): retire the service worker instead of just fixing it
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 35s
CI / tests-integration (push) Successful in 1m58s
CI / tests-unit (push) Successful in 2m22s
CI / tests-ui (push) Successful in 2m50s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m50s
The previous commit stopped the worker from caching build output, which
removed the cause of the blank Catalog Studio but left the worker itself
serving a purpose worth one offline fallback: three public endpoints
(/api/home, /api/online, /api/radio/config). Behind Cloudflare, for a site
whose visitors are online, that fallback rarely fires and goes stale exactly
where it is most likely to matter.

So the worker goes away rather than staying as a mostly-inert layer that every
future release still has to keep correct.

public/sw.js is not deleted, because deleting it would leave every existing
registration alive and still in control of the page, caching as it did before.
It becomes the opposite of what it was: on activate it deletes every cache,
unregisters itself and reloads open clients so they stop being controlled.
Browsers that already installed a worker therefore uninstall it on their next
visit; browsers that never had one are unaffected.

src/components/pwa-register.tsx is removed and unmounted from the root
layout, so nothing registers a worker any more and the kill switch only ever
runs for the visitors who need it.

The manifest is deliberately untouched. src/app/manifest.ts is an ordinary
Next.js route, independent of any worker, and Chrome installs from a manifest
alone, so the site stays installable on a phone.

Verified nothing depended on it: no other navigator.serviceWorker or caches.*
reference exists in the app, the theme runs from the plain /scripts/
theme-init.js script, and no Studio, catalog or import module touches a
worker. Typecheck and lint clean, 2326 unit tests and 72 UI tests green,
including every Studio and catalog spec.
2026-10-11 17:31:15 +02:00
openhands 1a4888df50 fix(pwa): stop the service worker replaying chunks from a previous release
Gitea Actions Runner Test / test-job (push) Successful in 3s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 2m11s
CI / tests-unit (push) Successful in 2m16s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m54s
CI / deploy (push) Successful in 2m47s
The Catalog Studio rendered as a blank white page after a deploy, while the
server was serving it correctly: /admin/studio/furni answered 200 with 110KB
of HTML and every API it calls answered 200 with data. No script chunk 404ed
during the session and no JavaScript threw, so the failure was entirely in
what the browser chose to execute.

The service worker wrapped /assets/ and /_next/static/ in a Cache Storage
entry served cache-first, under a hardcoded name (atom-v3) with no build id
and no revalidation. A release therefore could not invalidate it: the browser
kept replaying the previous release's chunks against the new HTML, and React
never hydrated. The chunk branch also had no .catch(), unlike the API branch
below it, so a rejected fetch silently dropped the <script> instead of
surfacing a network error the browser could retry.

That cache also bought nothing. nginx already sends /_next/static/ and
/assets/ as `max-age=31536000, immutable`, and those filenames are
content-hashed, so the HTTP cache is both sufficient and safe — a changed
file gets a new name. The worker was the only layer able to go stale across a
deploy, and it was the one doing it.

Both prefixes now fall through to the network and are left to the HTTP cache.
The API offline fallback and network-first navigations are unchanged, and the
cache names move to v4 so an existing worker is replaced. SW_VERSION moves
with them: it is part of the registration URL, so without the bump the browser
never refetches sw.js and the old worker keeps running.

Unrelated but confirmed while tracing this: /assets/images/themes/arctic-ice.png
is requested by the browser and 404s, but that string exists nowhere in the
code, the database or the build. It was a stale reference served out of this
same cache, not a missing asset.
2026-10-11 17:21:53 +02:00
openhands 2978483ea9 fix(assets): serve furniture icons from the gamedata tree, add stack-height SQL generator
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 2m3s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m40s
CI / deploy (push) Successful in 3m47s
Commit the two changes that were live in the working tree but never
recorded.

The nginx change adds a location for /swf/dcr/hof_furni/icons/ that reads
/var/www/Gamedata/icons/ off disk and falls back to the CMS bundle for
anything missing. The CMS built its icon URLs from that path, but the icons
actually live in the gamedata tree: of the 16,263 classnames in items_base,
15,338 are present there under the safe name, against 11,267 under
public/swf/dcr/hof_furni/icons. The swf tree also keeps colour variants
behind a literal "*" in the filename, so every request using the safe name
missed, and hof.furni.url already points the Nitro client at the gamedata
tree. A miss is a no-store 404, never a cached one — same reasoning as
@gamedata_missing, since add_header without "always" says nothing about a
404 and Cloudflare would otherwise apply the zone TTL.

scripts/generate-stack-height-sql.ts emits a SQL file that repairs
items_base dimensions and interaction columns from the logic inside each
bundle, which is where they actually live rather than in furnidata. It
writes a file and touches no database, so the result is reviewed before it
is run.
2026-10-11 17:01:53 +02:00
openhands 43742e8d99 fix(catalog): decode WebP bundle textures so furniture icons resolve again
Every write path normalises a bundle's texture to WebP Lossless, so the
catalog icon was being read back with a PNG-only decoder. decodePng throws
on anything that is not a PNG, the callers caught that and returned null,
and the user-visible result was "no icon (not in source or bundle)" for
every furniture whose source does not serve a standalone icon.

Measured against the production asset tree, all 18,505 bundles were WebP;
icon extraction succeeded on 0 of them. src/lib/services/imager/
decode-texture.ts keeps PNG on the dependency-free decoder and routes WebP
through sharp, which is already a dependency and already encodes these
textures. extractFurniIconPng and getPetIconPng become async; the five
call sites (upload, clone import, furni import, icon repair and both icon
routes) already awaited their surrounding work.

Same root cause, second bug: the spritesheet frame key. Converters disagree
on packing — some keep a trailing ".png", and some lowercase the whole key
while leaving the bundle name mixed-case, so "LTD_fashionistaf" looks up
frame "LTD_fashionistaf_LTD_fashionistaf_icon_a" and never finds
"ltd_fashionistaf_ltd_fashionistaf_icon_a". Any mixed-case classname could
therefore never match, which is most of the catalogue. findFrame tries the
two exact spellings, then falls back to one case-insensitive pass.
Extraction now succeeds on 18,483 of 18,505 bundles (99.88%); the 22
remainder are data, not code — 9 bundles ship no icon asset, 11 do not
parse.

Third: three catalogue icons exist only as .gif while catalogueIconUrl
hardcoded .png, so the picker offered icons that could only ever 404, and
291 icons that ship as both formats were listed twice. The API now dedupes
per id and the two renderers retry with .gif before falling back to the
placeholder, matching what catalog-image-picker already did. Verified live:
/gamedata/.../icon_1542.png returns 404 while icon_1542.gif returns 200.

Separately, close the last hole in the memory cap. Every script in
package.json routes through scripts/with-memory-cap.sh, but invoking the
builder directly — from a terminal, an IDE or an agent — skipped the
wrapper and ran unbounded, on a host with no swap where the OOM killer
picks its victim across the whole machine. next.config.ts now refuses a
production build that the wrapper has not marked, before anything
allocates. next dev and next start are deliberately unaffected.

README gains a Memory-capped commands section covering the per-script
ceilings, the backends and the ulimit -v trap, and its stale version and
script tables are corrected.
2026-10-11 17:01:37 +02:00
openhands 6793f77733 fix(deploy): stub git status in the simulation harness
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m37s
CI / tests-unit (push) Successful in 1m39s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m20s
CI / deploy (push) Successful in 3m6s
The clean-tree guard in ci-deploy.sh aborts the release when
`git status --porcelain` prints anything. The harness stubs `git` as a
shell function that only special-cases `rev-parse`; every other
subcommand fell through to its `ls-remote`-shaped printf, so `git status`
emitted a fake refs/heads/main line and the guard failed on every
scenario.

Introduced in fdb7af7e, which added the guard without teaching the
harness about it, so all 21 deploy tests have been failing since. This
stubs `status` to report a clean tree, and adds a scenario that asserts
the guard actually stops a release before it builds, migrates or starts
anything — the guard itself had no coverage, which is how it could break
silently in the first place.
2026-10-10 17:26:12 +02:00
openhands adffac7360 feat(catalog): store furniture bundles as .hab instead of .nitro
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-unit (push) Failing after 1m37s
CI / tests-integration (push) Successful in 1m37s
CI / tests-ui (push) Successful in 2m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Every bundle the CMS writes — upload, clone, sync, repair and the pet /
effect / figure importers — now lands as `<classname>.hab`, the extension
this deployment's renderer asks for. `.hab` and `.nitro` are the same
container, so an upload of either extension is accepted.

Resolution goes through one module, src/lib/furni/bundle-file.ts, so
nothing has to know the extension twice. Every existence check probes
`.hab` first and falls back to `.nitro`: the on-disk asset set is still
predominantly `.nitro`, and without the fallback Studio would report every
imported item as missing and the cleanup scan would classify 18k live
bundles as fake leftovers. Downloads are unchanged — Habbo's CDN and every
configured clone source still serve `.nitro`, so the conversion happens on
write, not on request.

Deliberately unchanged: the staged-attachment store in furni-attachment.ts
keys on a UUID and never reaches the client, so renaming it would break
in-flight recovery jobs.

Adds scripts/migrate-nitro-to-hab.ts to rename the existing asset set. It
refuses to run without --dry-run or --yes, never overwrites an existing
.hab, never deletes, and is idempotent.

Note: renderer-config.json lives outside this repo and was patched to
.hab separately; that file is served with a 30-day max-age, so returning
clients need a cms-client cache purge to pick the change up.
2026-10-10 17:09:36 +02:00
openhands fdb7af7ef5 fix(deploy): unblock every rebuild on BuildKit's host-network refusal
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-unit (push) Failing after 2m4s
CI / tests-integration (push) Successful in 2m6s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Docker 29.1.3 ships BuildKit v0.26, which refuses to grant a build host
networking unless each caller passes --allow=network.host. All three rebuild
paths asked for it, and `docker compose build` has no flag to grant it, so a
rebuild failed immediately with "additional privileges requested". The live
container was never replaced, which is exactly the reported symptom: the site
kept serving the previous release after a rebuild.

Nothing in the build actually needs host networking. It uses the network only
for apk, pnpm and next/font/google — all outbound internet, which the default
bridge provides. Verified by building both the full runner image and the
migrations stage with --no-cache after dropping the flag.

Runtime `network_mode: host` stays: blue/green needs per-release host ports
(3002/3003) and nginx reaches each slot over 127.0.0.1.

The second gap is how a rebuild could still ship the wrong code. ci-deploy.sh
stamped every image with HEAD's revision label, and verify-deployed-release.mjs
only re-checks that same label, so a dirty working tree produced an image that
claimed to be release $sha while containing uncommitted code. docker-update.sh
already refused this; ci-deploy.sh now does too, before any build work.
2026-10-10 13:07:13 +02:00
openhands 48291ab641 fix: correct the ACL revoke migration and update the login redirect e2e
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m56s
CI / tests-unit (push) Successful in 2m8s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m44s
CI / deploy (push) Successful in 2m51s
The deploy gate found two defects in the previous commits, both mine.

0034_acl_midrank_revoke.sql never applied: it joined `acl_roles` on
`ar.model_type`, a column that table does not have (only
`acl_model_permissions` does). It now joins on the id and keeps the
`model_type` check where it belongs.

That hid a second, worse bug. The rank was extracted with
`SUBSTRING(slug, 7)`, but MySQL's SUBSTRING is 1-based and the digits start at
position 6, right after `rank_`. rank_10 therefore parsed as 0 and rank_7 as an
empty string, so every rank >= 7 would have lost exactly the grants the
migration exists to preserve — the ACL repair would have made things worse, not
better. Now reads from position 6.

Verified against a real MariaDB with a fixture covering rank_1, rank_6, rank_7,
rank_9, rank_10 and a non-rank slug: only the sub-7 roles lose their non-view
admin.* grants, the multi-digit and higher ranks keep everything, and the
non-rank slug is untouched. The mail index was checked the same way — it
applies idempotently and EXPLAIN confirms `users_mail_index` with rows: 1.

news.spec.ts expected to land on /me after signing in. That expectation predates
the `?from=` honouring added in 39332149, which lands a bounced admin back where
they were heading. The same step navigates to /admin/articles/new explicitly a
few lines later, so nothing depended on it; the assertion now covers the redirect
target instead.
2026-10-09 18:21:09 +02:00
openhands 6b34293cd3 fix: apply the 44px button target only to coarse pointers
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 27s
CI / tests-integration (push) Successful in 1m35s
CI / tests-unit (push) Successful in 1m38s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m20s
CI / deploy (push) Failing after 2m55s
The UI screenshot suite caught a regression from the previous commit: the admin
article form grew 3px and its baseline no longer matched.

The cause was the blanket `.btn` min-height bump from 40px to 44px. That was
the wrong way round — 40px already clears WCAG 2.2 AA, which asks for 24px, and
44px is a touch-target guideline. Growing every button for mouse users only made
each admin dialog and table 4px taller for no benefit.

The 44px floor now sits behind `@media (pointer: coarse)`, so finger input gets
the comfortable target and desktop keeps its density. A hybrid laptop still
uses the desktop metrics for its trackpad, which is the behaviour the previous
attempt got wrong in both directions.
2026-10-09 18:04:21 +02:00
openhands 5cb42c8dfb fix: stop the commandocentrum audit call leaking an unhandled rejection
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 27s
CI / tests-unit (push) Successful in 1m35s
CI / tests-integration (push) Successful in 1m37s
CI / preflight (push) Skipped
CI / tests-ui (push) Failing after 2m25s
CI / deploy (push) Skipped
The full suite passed but exited non-zero, which fails CI: three unhandled
rejections came out of commandocentrum's fire-and-forget audit call.

The cause is a real defect, not a test artefact. `auditAction` wrapped
`logAudit(...)` in a try/catch to honour "auditing must never fail the command
it describes", but logAudit is async, so the catch can never see its rejection.
A failing audit insert therefore surfaced as an unhandled rejection instead of
being swallowed — in production that is a request taking down over a logging
failure. The catch is now on the promise itself.

The test now mocks the audit service explicitly instead of leaning on the fake
db lacking `insert`, and asserts both that an entry is logged and that a
rejecting audit still lets the command succeed.
2026-10-09 17:57:28 +02:00
openhands 759ae91745 perf: cache search and news archive, drop motion/react from public pages
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 26s
CI / tests-unit (push) Failing after 1m37s
CI / tests-integration (push) Successful in 1m38s
CI / preflight (push) Skipped
CI / tests-ui (push) Failing after 2m24s
CI / deploy (push) Skipped
Closes the four remaining LOW items.

Search and news archive caching
- A leading-wildcard LIKE cannot use an index, so every /search section cost a
  COUNT(*) scan plus an ordered page fetch, and /news did the same for its
  archive. Both now cache: search per section for 30s, the archive for 60s
  under the existing news revision so publishing an article drops it at once.
- Sections are cached independently, so one slow query cannot hold up the rest
  and a failure is not cached as a result.
- The cached value is passed through cacheSafe() so the Redis path and the
  in-process path return the same types; without it a cache hit would hand the
  events grid a string where a miss hands it a Date, and it calls toISOString()
  on that field. Dates are revived on the way out so the public signatures of
  loadNewsArchive and loadPublicSearch are unchanged.
- Archive entries are keyed on the REQUESTED page rather than the clamped one,
  so two requests that clamp onto the same page cannot alias each other.

motion/react out of the public bundle
- Converted the six public-facing users: the radio player, the typewriter text
  (a motion.span with no animation props at all), the photo lightbox, the
  animated counter, the footer CMS-info popup and the scroll reveal. That was
  the actual entry points — the counter and the popup reach the public home page
  and footer through static imports, so removing only the three originally named
  would have left the library in the bundle anyway.
- Each animation moved to a CSS class, and the two that animate on exit now hold
  the element for the length of the fade, which is what AnimatePresence used to
  do.
- motion/react now only ships with /admin and the two already-lazy nav panels.
- Two safety fixes came out of this: the scroll reveal starts at opacity 0, so
  it is forced visible under prefers-reduced-motion and via a <noscript> rule in
  the root layout; and it now emits the .motion-reveal class, which the theme
  panel's "Scroll Reveal" toggle selects and which previously matched nothing.
- The CMS-info backdrop became a real button in a pointer-transparent layer
  instead of a handler on a static element, so click-outside-to-dismiss is
  reachable by keyboard.

Fewer duplicate router refreshes
- Next.js re-renders the current route as part of a server action's own response
  when that action revalidates, and applies it with a seeded navigation; the
  router only skips its own update when the action did NOT revalidate. So the
  refresh after such an action fetched the same tree twice.
- useServerAction takes an opt-in `revalidated` flag that skips it. It is opt-in
  per call rather than derived from an action name, since a rename would
  silently change behaviour. Applied to the two user-facing call sites whose
  actions were verified to revalidate their own route.

Touch targets
- .btn was the one shared control at 40px; it and the lightbox and CMS-info
  close buttons are now 44px, as is the password toggle (the auth input already
  reserved 44px for it). The remaining 32px icon buttons pass WCAG 2.2 AA, which
  only asks for 24px; enlarging those inside inputs and overlays was left alone
  because it risks visual breakage that cannot be checked from here.
2026-10-09 17:35:38 +02:00
openhands 179484642f feat: per-account login lockout, mail index, resend captcha, i18n scoping
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Failing after 1m45s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m29s
CI / deploy (push) Skipped
Closes the four HIGH/MEDIUM items left open after the previous pass.

Login lockout
- The only login limits were keyed on the client IP, so a distributed attempt
  could grind on one account indefinitely. Added a per-account lockout with a
  budget of 8 failures per 15 minutes.
- The bucket is keyed on the RESOLVED account id, not on the submitted string:
  users may sign in with either username or e-mail and neither the lookup nor
  the input normaliser folds case, so an input-keyed bucket would hand out a
  fresh budget per spelling of the same account.
- precheckLogin and NextAuth's authorize share the bucket, so the pre-check
  cannot be used to buy extra attempts and a client that skips it entirely is
  still bounded. Both check the lockout BEFORE verifying the password: the
  success path clears the counter, which would otherwise walk a locked account
  straight back in on the right password.
- A successful login clears the failures, which needs two new primitives in
  rate-limit.ts: peekRateLimit (read-only, does not consume a unit) and
  clearRateLimit.
- Fixed a latent inconsistency while doing so: the in-process bucket capped its
  counter at the limit while Redis' INCR kept climbing, so the two backends
  disagreed about how far over the limit a key was. Both now track the true
  count.

Mail lookup index
- Added an index on users.mail (0035). Password reset, e-mail verification and
  the resend cooldown all resolve a single account from a submitted address and
  were full table scans of `users`. Deliberately non-unique: legacy rows can
  hold the same address more than once, so a unique index would fail to apply.

Resend captcha
- /verify's resend form triggers real outbound mail and was reachable with only
  a cooldown. It now runs the configured captcha before the account lookup and
  before any send.

Client message payload
- The root layout serialised the whole catalogue into every page. pages.admin
  and admin are ~177 KB of the ~235 KB and are unreachable from the public route
  group, so that layout now installs its own provider with the staff namespaces
  removed. Nested providers replace rather than merge, which is why this has to
  live in the segment layout. /admin, /mod, /client and /admin-next keep the
  full set; a guard test fails if a public page ever references a staff
  namespace.
2026-10-09 17:12:50 +02:00
openhands 6cc45d7413 feat: harden atoms-nexst against review findings (37 items)
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Failing after 1m49s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m31s
CI / deploy (push) Skipped
Second review pass covering security, performance, admin tooling and the
public/room flows. All HIGH and MEDIUM findings from the audit are resolved;
nothing in this commit changes the visible feature set.

Authentication & session security
- CSP is now set on the request headers in the proxy, which is what Next.js
  uses to derive the render nonce, so the nonce is effective.
- 2FA: an already-enabled user cannot re-enroll, the setup endpoint is
  rate-limited per account, and confirmed codes are persisted so the second
  secret no longer silently never applies.
- Password reset revokes the ticket, authTicket and all personal access
  tokens, and bumps the token version so existing sessions die. The same
  revocation is now wired into the staff-side password reset.
- /reset and /verify return a stable error code instead of raw text; the
  mail lookups are ordered by id so duplicates cannot vary between runs.
- Resending the verification mail gets a per-address cooldown on top of the
  per-user limit.
- Issue API tokens with the narrower radio/ticket ability set instead of "*".

Authorization & input handling
- Mid-rank staff can no longer keep dynamically granted non-view admin.*
  permissions: existing grants are revoked by migration and the grant lookup
  is restricted to "%.view". Rank guards use the dynamic super-admin check.
- Alerting a user is permission-checked and audited like the other tools.
- Material mutations (giveCredits/giveDuckets/giveDiamonds, the admin user
  actions route, bulk user actions) are capped and rank-guarded, and bulk
  ids are bounded.
- updateRoom / updateRoomItem write through a field allowlist, and items
  may only be edited through their own room.
- Classnames reaching the filesystem are validated before use so a crafted
  value cannot escape the asset directories.
- The word filter now also covers offline mails, guild forum threads and
  replies, and user mottos.
- Media uploads are validated by magic bytes, /api/media requires the page
  edit permission, APP_URL must be configured once mail is enabled, and the
  diagnostics error route checks the fetch site header.

Admin tooling
- Secret settings render masked and cannot be overwritten with a blank or
  an arbitrary raw key; radio credentials are new password inputs.
- Commandocentrum balance changes are audited.
- Admin list pagination reads the caller's per-page instead of the max, and
  the log exporter caps offset and search length.

Performance
- Catalog translations are cached per module, with a cheap revision hash;
  the public online count uses a stale window instead of hammering the DB.
- The cache warmup now primes the payload the home route actually reads.
- TopHeader batches its queries into one round trip, and LCP avatars load
  eagerly.
- motion/react and sonner are no longer part of the root layout; the nav
  dropdown and mobile nav panels are lazy client chunks. Anonymous visitors
  again get the navigation chrome, and public pages get an edge cacheable
  response.

Accessibility
- Nested <main> elements in phase pages became <section>; the page entrance
  and route progress animations are pure CSS that respect reduced motion.
2026-10-09 16:19:48 +02:00
openhands 3933214953 feat(auth): implement all 16 homepage/login/register review items
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m47s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m30s
CI / deploy (push) Failing after 2m56s
- add countArticles() (published-only, mirrors news-list) and warm total_articles
- localize homepage metadata; bind articleCount to both stats; unique photo alts
- drop duplicate news date and the mascot preload priorities
- extract shared AuthPageFrame/AuthUsersCards used by /login and /register
- login: localized noindex metadata, session redirect via safeRedirectPath,
  ?from passthrough from proxy, unified auth roster cache keys, registered notice
- register: localized metadata, session redirect to /me, unified cache keys
- add resend-verification flow on /verify with rate-limited non-enumerable action
- add safeRedirectPath() with unit tests
- register form: live requirements checklist + password mismatch guard
- login form: unverified state with resend-link CTA
- honour prefers-reduced-motion in TypewriterText
- add 6 translations across all 25 locales
2026-10-08 18:49:27 +02:00
openhands 8561c3f85e fix(ci): accept any runtime the engines range supports
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 42s
CI / tests-unit (push) Successful in 1m42s
CI / tests-integration (push) Successful in 1m47s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m28s
CI / deploy (push) Successful in 3m12s
The active-runtime assertion required process.versions.node to equal
.nvmrc exactly, so any Node.js patch release broke `toolchain:check` and
the act CI run even though package.json engines (>=26.10.0 <27) supports
the newer runtime.

Keep .nvmrc and the Docker base image exactly pinned for reproducibility
(both still asserted), but validate the running runtime against the
engines range instead. Verified: toolchain:check, lint, typecheck,
i18n:check, hk:matrix:check and all 3391 tests pass.
2026-10-08 17:28:56 +02:00
openhands d2350a6427 fix(ci): install bash in the Docker build stage
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Failing after 17s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The build script now runs every heavy command through
scripts/with-memory-cap.sh, which is bash (arrays, BASH_REMATCH). Alpine's
node image ships busybox ash, not bash, so the builder stage failed with
`sh: bash: not found` (exit 127): ci-deploy.sh could not build the image.

Verified: full docker build passes and compiles all 279 routes.
2026-10-08 17:19:06 +02:00
openhands 0845f80768 fix(ops): run heavy commands under a hard memory cap to stop host OOM kills
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 35s
CI / tests-integration (push) Successful in 2m5s
CI / tests-unit (push) Successful in 2m30s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 3m3s
CI / deploy (push) Failing after 44s
The host runs with vm.overcommit_memory=0 and no swap, so a process that
grows past free memory makes the kernel OOM-kill across the whole machine
-- the Turbopack build (commit 3d828a61) could take out the database,
nginx or the live release.

Add scripts/with-memory-cap.sh: it moves a command into its own systemd
scope with MemoryMax, so only that cgroup gets OOM-killed (verified: a
Turbopack build died at its 6GB cap, host untouched). Build/analyze/dev/
test*/typecheck now run under explicit caps; ulimit -v is only an explicit
opt-in because it bounds virtual address space per process and 10g/20g both
break V8-based builds. Docker and GitLab builds run in their own isolated
containers with a read-only cgroupfs and opt out explicitly (webpack +
--max-old-space-size stay their bound).

Measured: webpack build peaks ~6.5GB RSS, so 10GB leaves headroom within
the 23.5GB host.
2026-10-07 20:17:00 +02:00
openhands 3265c149da style(landing): add micro-interactions and polish public pages
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m48s
CI / tests-integration (push) Successful in 1m52s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m35s
CI / deploy (push) Successful in 3m7s
2026-10-06 22:59:03 +02:00
openhands 57710a7fe3 style(landing): polish the public index, login and register screens
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 26s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / preflight (push) Skipped
CI / tests-ui (push) Skipped
CI / deploy (push) Skipped
- add theme-aware helpers: btn-brand, btn-glass-dark, auth-input, auth-label,
  auth-alert, explore-pill, avatar-tile, aurora blobs and a hero scroll cue
- reorder backdrop-filter declarations so the glass blur survives the
  production CSS optimizer in modern Chromium
- hero: aurora glow, frosted recent-users chip, premium CTA buttons and cue
- explore nav: icon pills with hover arrows; features: gradient icon tiles
- stats: single glass panel with column dividers; join CTA: aurora + ring
- auth forms: visible labels, icon inputs, eye/password toggle, gradient
  submit and pill footer links
- auth pages: gradient card frame, aurora accents on the intro panel and
  hover-lift avatar tiles; unified pill-shaped top bar
- drop the hard-coded register banner image in favor of the framed card
2026-10-06 22:00:02 +02:00
openhands 5b2eb91c5c fix(ops): stop a compose replica from blocking the blue/green release
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m57s
CI / tests-unit (push) Successful in 2m3s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m49s
CI / deploy (push) Successful in 2m48s
The deploy failed after the build, the migrations and the browser gate:
"Port 3002 is already in use". The holder was `epicnext-cms`, a compose
replica of release 6bffc537 that the daily scripts/docker-update.sh cron
had recreated at 03:30 with restart=unless-stopped. nginx serves the green
slot on 3003, so that replica was squatting the blue slot the next
candidate needed, and live traffic never noticed.

It got there because the updater's CI-ownership guard only tested
epicnext-cms-app. After a cutover to the green slot that container is
stopped, renamed and deleted, so the guard stopped firing while the host
stayed CI-managed.

- scripts/docker-update.sh: refuse a compose deployment on a CI host by
  checking both slot containers and the nginx upstream, which is the only
  thing that still marks the host as blue/green while a slot is idle.
- scripts/ci-deploy.sh: retire a compose replica of this checkout from
  the candidate port before starting the candidate, so a stray replica
  can never block a release again. Never a slot container, never the port
  nginx serves; anything else still fails loudly in assert_port_free.
- Tests cover both directions: a squatting replica is removed and the
  release lands, a replica on the live port is left alone.
2026-10-05 21:13:49 +02:00
openhands 11ad6d4376 fix(ci): measure route bundles from webpack manifests
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 1m41s
CI / tests-unit (push) Successful in 1m44s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m31s
CI / deploy (push) Failing after 2m25s
The performance report measured nothing. It only read `entryJSFiles` from
each route's client-reference manifest, a field Turbopack emits and webpack
does not. When the build moved to webpack (3d828a61) every route fell
through to the "unavailable" branch, and because the report is informational
and exits 0 on an unavailable metric, nothing failed and the budgets quietly
stopped being enforced.

Derive the envelope from clientModules[*].chunks when entryJSFiles is
absent, which is the same source Next's own static-routes-info uses for
webpack builds. Webpack interleaves numeric chunk ids with file names in
those arrays, so ids are skipped by shape while a malformed chunk path still
throws — otherwise a broken manifest would quietly under-report a route.
entryJSFiles still wins when present, since it is per-segment and therefore
the tighter envelope, and the per-chunk origin label is shared rather than
the absolute node_modules path webpack records, which would otherwise bloat
report.json.

Six tests cover the webpack layout: id filtering, deduplication of a chunk
reached by several client modules, the origin label, the malformed-path
rejection, the no-chunks-at-all case, and entryJSFiles taking precedence.

Re-measured on the current build, all six routes are inside their budgets
again. Note /admin/studio/furni now sits at ~98% of its gzip limit, so one
more dependency on that route will trip it; docs/performance-budgets.md
records the webpack baseline numbers and how to recalibrate.

Verified: 3385 tests, typecheck and biome clean, and the report now emits
measured rows instead of six unavailable ones.
2026-10-05 20:49:53 +02:00
openhands 6c3d81920e fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m33s
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.

jobs-worker never ran

`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.

The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.

/api/health answered 200 with the database down

The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.

The runtime had no V8 heap cap

NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.

Storage ownership was only repaired for one path

ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.

nginx: robots.txt was a guaranteed 404, and TLS never resumed

`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.

Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
2026-10-05 20:25:22 +02:00
openhands 108c6ce03d fix(ci): make the lint gate fail for real and stop byparr leaking disk
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 2m3s
CI / tests-unit (push) Successful in 2m18s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 3m6s
CI / deploy (push) Failing after 3m14s
The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.

Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
  (key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
  function-declaration handlers, with the reasoning that dropping them broke
  the tree and save-on-Ctrl+S once already (704e3363)
- scope the remaining noArrayIndexKey / useExhaustiveDependencies exemptions to
  the three files that need them, in biome.json instead of scattered comments

Storage, on a host that had grown to 81% disk:
- byparr starts a Firefox per request and never removes the profile it leaves in
  the container's writable layer. With no volume mounted, nothing else reclaimed
  it: 716 profiles / 6.8 GB in two days, ~1.7 GB/day. docker-prune.sh now removes
  orphaned profiles, identifying live ones by the open fd in /proc/<pid>/fd rather
  than by age, because browsers stay warm for ~27 hours here — longer than the
  leak window, so no age threshold can be both safe and useful.
- bound the build cache properly: buildx treats --max-used-space and --filter as
  mutually exclusive, so passing both silently dropped the 4 GB cap and the cache
  reached 49 GB.
- escalate to the emergency prune when / drops below 8 GB free, so the bound holds
  even if the schedule stops.
- clear multi-GB tmp_pack files left behind by a gc that was OOM-killed
  mid-repack; git only removes those on the next successful gc.
- make setup-cron.sh append instead of replacing the crontab (`crontab -`
  overwrites the whole file, which had been dropping the other scheduled jobs),
  and run the prune daily rather than weekly to match the leak rate.

Volumes are still never pruned: mariadb-turbo-data is a database.
2026-10-05 17:24:12 +02:00
openhands 6bffc53779 refactor(auth): merge the duplicate login form and localize the auth screens
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 1m8s
CI / tests-integration (push) Successful in 1m53s
CI / tests-unit (push) Successful in 1m59s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m42s
CI / deploy (push) Successful in 4m33s
`home-login-form.tsx` and `login-form.tsx` were two ~240-line near-identical
components. Delete the former and give `LoginForm` a `variant` prop:

- `variant="page"`   sr-only labels plus the register/forgot footer (/login)
- `variant="compact"` visible labels, no footer (homepage sidebar)

Field ids now come from `useId()`, so the two usages can never collide, and the
hardcoded "Show"/"Hide"/"Loading" strings are translated.

Localization of the login and register screens:

- `home-login-form.tsx` was entirely hardcoded English.
- `passwordStrength()` returned hardcoded "Weak"/"Fair"/"Good"/"Strong".
- `register.ts` returned only English strings. It now returns a
  locale-independent `code` next to the message, and the form renders
  `t(code)` with the English string as a fallback.
- Backfilled the new keys across all 25 locales, plus the login/register
  strings that were still English in most of them. `ar`, `fi` and `ja` had
  their entire login/register namespace in English and are now filled in.
  Locale parity stays at 0 missing keys, as `i18n:check` requires.

Copy that did not match the enforced rules: the UI advertised "min 8 chars"
(EN) / "min 6 tekens" (NL) while registration requires 12 characters plus an
uppercase, a lowercase, a digit and a special character. Corrected in every
locale. `password-reset.ts` enforced only 6 characters and is raised to 12 to
match registration.

Accessibility: `login-form.tsx` had no `<label>`, no `id` and no `required` on
any field. All three are now present, and error banners are announced with
`role="alert"`.

Adds `src/i18n/auth-messages.test.ts`, which asserts every `RegisterErrorCode`
resolves to a non-empty message in all 25 locales; verified it fails when a key
is removed. The existing register tests now also assert the error `code`.
2026-10-04 18:50:23 +02:00
openhands 3d828a61ab fix(build): build with webpack because the Turbopack build is OOM-killed
`next build` on Turbopack never completes on this app. The compiler is a
single native process whose RSS grows monotonically with no plateau:

    0.9G -> 1.6G -> 2.8G -> 5.0G -> 5.5G -> 6.2G -> killed

It still dies with 4GB of swap attached, at 12GB RSS. The build workers are
only 0.17GB each, so `experimental.cpus` is not the lever either.

A `--max-old-space-size` cap cannot help: measured with a 2GB cap, RSS still
reached 8GB, because the memory is native Turbopack (Rust) memory rather than
the V8 heap. The cap added in 1c9ddcd4 was therefore inert and only created
false confidence, so it is dropped from the build script.

Webpack builds the same 329 routes in ~95s with a ~6GB peak.

Ruled out by measurement: the 25 bundled locale files (stubbing 24 of them
from 7.1MB down to 276KB still peaked at 11GB), worker count, and the
flatten/unflatten message pipeline (600 iterations cost 6.4s and settle at
39MB of heap).

Verified: `pnpm run build` exits 0, TypeScript passes, 279/279 static pages are
generated, and the standalone output boots and serves /, /login and /register.
2026-10-04 18:50:11 +02:00
openhands 7f07c111ac perf(studio): load motion's minimal entry instead of the full component library
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Successful in 1m52s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m44s
CI / deploy (push) Successful in 2m34s
/admin/studio/furni sat at 94.9% of its initial-JS budget (427436 of
450560 gzip bytes), so the next feature would have broken the build. Of the
98228 gzip bytes unique to that route, a large part is framer-motion.

This file uses motion twice, for one thing: a 150ms opacity fade on the result
pane when viewMode changes. Importing `motion/react` to get it pulls in
framer-motion's complete component library — 73 internal modules — plus its
render components, drag/gesture and projection code, none of which is
rendered here.

`motion/react-m` ships only the element factories: 2 internal modules, and the
same initial/animate/transition props, so the fade is unchanged. It exports the
elements flat rather than under a `motion.` namespace, so the import becomes
`div as Mdiv` and the two JSX tags are renamed to match.

I could not measure the resulting bundle here: the local build is OOM-killed
(exit 137) with the running containers on the host, so the actual saving is
unverified. The CI build reports it in build-reports, and the number in this
commit message should be read as a hypothesis, not a measurement.

Verified: typecheck clean, lint clean, and the 10 studio UI tests pass —
including the pane and navigation specs that exercise the view switch.
2026-10-03 19:21:02 +02:00
openhands 8218039c64 test(live): stop the live suites inheriting the production database
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m57s
CI / tests-integration (push) Successful in 2m1s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m51s
CI / deploy (push) Successful in 1m42s
Seven suites read .env with a bare `process.env[key] = value`, which
overwrites whatever the shell already set. That made the DATABASE_URL from
the production .env authoritative, so a single environment variable was
enough to aim them at the live hotel database:

  RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts

Three of those suites then repair the catalog in place: catalog-audit-repair-live
and catalog-repair-direct-live rewrite catalog_items and delete duplicate
classnames, and clone-bulk-import-live bulk-imports every cloneable item. None
of that is undoable, and nothing in their output said the target was
production rather than a sandbox.

Added src/test/live-env.ts with one shared loader, and pointed all seven suites
at it:

- Values already in the real environment win, so an explicit DATABASE_URL on
  the command line is always respected.
- DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the
  production one from .env.
- Anything that is not loopback is treated as production and redirected.
- Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning
  saying the suite repairs the catalog.

Tests in src/test/live-env.test.ts run the loader against a temporary .env so
the real project file is never read, and cover the redirect, the shell
override, non-loopback detection, the opt-in and quote stripping. A second
block asserts each of the seven suites no longer contains an inline
`process.env[...] =` assignment. Verified four of them fail against the old
loader.

This does not enable the suites; they stay gated behind their RUN_* flags.
It only removes the possibility of them silently hitting production.

Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean.
2026-10-03 19:03:28 +02:00
openhands f705c67fc7 revert(docker-compose): keep the cms services the contract tests require
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m44s
CI / tests-ui (push) Successful in 2m34s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
The previous commit removed the `cms` and `cms-green` services from
docker-compose.yml. That was overreach and it broke two tests:

- src/lib/docker-build-contract.test.ts asserts the compose build passes
  NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}, so a compose-built image
  carries its release id.
- scripts/proxy-config.test.mjs resolves `docker compose config` and asserts
  the `cms` service's host networking, volumes, healthcheck and image tag.

Both encode that docker-compose.yml is a maintained deployment surface, not a
leftover. Removing it was not my call to make while fixing a deploy.

Restored verbatim. The stray container that actually blocked port 3002 is
already gone, and nothing recreates it: there is no systemd unit or pm2
ecosystem that runs `docker compose up`, and `restart: unless-stopped` only
applies to a container that still exists. So the blocker is resolved by the
container removal alone, and compose stays intact for manual and reviewed use.
2026-10-03 18:51:06 +02:00
openhands c8b3054527 fix(deploy): free port 3002 and stop compose from competing for the slots
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy could not start its candidate because port 3002 was held by
`epicnext-cms`, a `docker compose up` replica built from the `local` image and
serving no traffic. Everything else in the pipeline was healthy: the image
built, the news browser gate passed and migrations were current.

The container was unusable for this pipeline for two reasons. It ran a
different image than any release, and its name did not match the slot the
deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms`
while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to
agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never
could. deploy.sh already documents that compose "never managed the release
that actually ran", so the service was stale by its own account.

Removed the stray container and dropped the `cms` and `cms-green` services (plus
the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot
resurrect a replica that permanently occupies a blue/green slot. byparr is
untouched.

Also fixed the diagnostic from the previous commit, which blamed every running
container. `docker ps --filter publish=` returns nothing for --net=host
containers, so the fallback listed all of them and buried the real holder
among seven innocent ones. It now resolves the listening PID from `ss` back to
its container through /proc/<pid>/cgroup and names only that one, with the
exact `docker rm -f` command to run.

Verified: port 3002 free, live release on 3003 still serving
(status ok, database and redis true), deploy simulation 26 passed, typecheck.
2026-10-03 18:45:51 +02:00
openhands 8ec3df541e fix(deploy): name the container blocking a port and silence phantom cleanup
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m33s
CI / deploy (push) Failing after 1m30s
Two follow-ups from the blocked deploy.

The rollback path called `docker logs` and `docker rm -f` on the candidate
unconditionally. When the port check refuses to start it, the container was
never created, so both printed "No such container: epicnext-cms-app" — noise
that looked like a second, unrelated failure and buried the real message.
Both calls are now guarded by `docker inspect`.

assert_port_free() now reports which container holds the port and flags it when
it is not a blue/green slot this script manages. The previous output listed
every container and said only "port already in use", which is a dead end: on
this host the holder is `epicnext-cms` (a `docker compose up` replica on port
3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ
because docker-compose.yml pins `container_name: epicnext-cms` for the `cms`
service; slot B happens to match, which is why 3003 deploys fine and 3002 never
can. The message now names the squatter, explains that live traffic is
unaffected, and gives the next action.

Deploy simulation: 26 passed.
2026-10-03 18:37:12 +02:00
openhands 8ee144745a fix(deploy): stub ss in the deploy harness and cover the port-conflict path
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m47s
CI / tests-unit (push) Successful in 1m54s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m40s
CI / deploy (push) Failing after 1m41s
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.

The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.

Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.

Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.

Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
2026-10-03 18:30:00 +02:00
openhands 64ad9baf39 fix(deploy): trust the nginx upstream when picking the live slot
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Failing after 1m54s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m46s
CI / deploy (push) Skipped
The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.

read_active_port() now orders its sources by how well they describe reality:

1. The nginx upstream file. It is the only source that says where public
   traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.

answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.

start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.

Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
2026-10-03 18:22:40 +02:00
openhands 704e33638f fix: restore six useEffect dependencies removed while silencing lint
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m38s
CI / tests-integration (push) Successful in 1m40s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m24s
CI / deploy (push) Failing after 2m58s
The previous commit dropped biome-ignore comments to clear
useExhaustiveDependencies diagnostics and, in doing so, also deleted the
dependencies themselves. Six components were left with effects that no longer
react to the state they read. Every one of these is a real behaviour
regression, not a lint preference:

- health-check-client: checkEmulator is a function declaration, so it gets a
  fresh identity each render. As an effect dependency that re-fires the effect
  after every setState, polling /api/admin/devops/health in a loop. Wrapped in
  useCallback so the identity is stable.
- article-recovery: reload restarts the autosave timer for the "Retry recovery"
  button. Without it in the deps that button is a no-op. The counter had been
  renamed to _reload to satisfy the unused-variable rule.
- catalog-integrity-panel: same pattern; refresh starts a new read-only scan,
  so the rescan control did nothing.
- catalog-search: refreshKey re-runs the query after a bulk edit, so results
  were not refreshed after catalog edits. The selection-reset effect also lost
  catalogType, so switching catalog no longer cleared the selection.
- catalog-image-picker: dropped debounced (the search term) and name (the
  error reset), so image search and error state no longer reacted to input.
- icon-picker: dropped iconImage, so a failed load left the placeholder on the
  next icon too.

Each restored dependency carries a biome-ignore with the reason it is
load-bearing, so the diagnostic can be re-derived instead of silently
disappearing again.

Verified: typecheck, lint clean on all six, unit 3315 passed, integration 20
passed, UI 72 passed / 2 skipped.
2026-10-03 18:07:56 +02:00
openhands eddb7edea4 fix: make all CI jobs pass (integration, ui) and restore prefix dialog reset
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m34s
CI / tests-unit (push) Successful in 1m35s
CI / tests-ui (push) Successful in 2m20s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m37s
Three failing test suites blocked CI. All three were test defects, not
application bugs.

Integration tests (integration/database.test.ts)
------------------------------------------------
The suite set NODE_ENV=test, which makes cache.cached() short-circuit both
its Redis read (src/lib/cache.ts:226) and its write (:249). A suite whose
stated purpose is exercising the real Redis path therefore never touched
Redis. Switched to NODE_ENV=development, the only non-production value
src/env.ts accepts, so the shared-cache code paths are genuinely covered.

Three assertions then needed correcting for real Redis semantics:

- `await cache.cached(...)` followed by `.resolves` can never hold: await
  yields a value, not a Promise. Assert the value directly.
- A cached negative result is stored as the JSON encoding of null, so
  `redis.get(key)` returns "null", not null.
- The news negative-cache key does not exist at all, so `ttl()` returned -2.
  Now that the write path is live the key is created and the TTL assertion
  holds as originally written.

UI tests (src/app/admin/prefixes/prefix-dialog.tsx)
---------------------------------------------------
The form-reset effect had `isOpen` removed from its dependency array. The
component returns null when closed, so the effect only ever ran on mount:
reopening the dialog no longer cleared the fields and a dismissed-but-
unsaved edit reappeared. Two tests in e2e/ui/unsaved-changes.spec.ts caught
this. Restored the dependency and documented why it is load-bearing.

The remaining edits in this branch drop stale biome-ignore comments that
suppressed useExhaustiveDependencies and noArrayIndexKey diagnostics. Where
the suppression had been load-bearing for behaviour, the underlying
dependency is now listed explicitly rather than silenced.

Verified: check (toolchain, audit, lint, i18n, typecheck), unit 3315
passed, integration 20 passed, UI 72 passed / 2 skipped.
2026-10-03 17:02:49 +02:00
openhands 1c9ddcd48a fix(cms): increase docker mem limit to 6gb and enforce node max-old-space-size to prevent OOM killer crashes
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m30s
CI / tests-integration (push) Failing after 1m34s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m17s
CI / deploy (push) Skipped
2026-10-02 22:25:16 +02:00
openhands 30ff970c38 chore: upgrade to pnpm v12, update dependencies, and fix msw v3 typescript types
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / preflight (push) Skipped
CI / tests-ui (push) Skipped
CI / deploy (push) Skipped
2026-10-02 21:59:39 +02:00
openhands f99980052b perf: optimize cache layer for speed and stability
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 34s
CI / tests-integration (push) Failing after 33m57s
CI / tests-unit (push) Failing after 33m57s
CI / tests-ui (push) Failing after 33m56s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- Remove random TTL jitter to prevent unpredictable cache drops
- Add deterministic LRU eviction with proper entry cleanup
- Improve cache deduplication to prevent duplicate computations
- Skip Redis I/O during tests for faster, more stable execution
- Optimize depth calculation in catalog tree nodes
- Maintain backward compatibility and full test coverage (3331 passed)
2026-10-02 17:16:03 +02:00
openhands f181cd6af4 chore: ignore local runtime snapshots under backups/
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m41s
CI / tests-unit (push) Successful in 1m46s
CI / tests-ui (push) Successful in 2m32s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m56s
backups/catalog-integrity/pre-fix.json is a one-off database snapshot taken
during an incident, not source. Kept on disk for reference, out of git.
2026-10-01 18:07:56 +02:00
openhands 8a6d92afd8 fix(cloudflare): cache the gamedata tree at the edge with respect_origin
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m44s
CI / tests-integration (push) Successful in 1m44s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m31s
CI / deploy (push) Successful in 2m16s
Icons are plain .png, a cacheable extension by default, so the zone's
"Browser Cache TTL = 1 year" pinned them to max-age=31536000 regardless of
the 300/3600/604800 that nginx sends per class. Extend the edge rule to
/gamedata/ and keep respect_origin, so the nginx header wins and a 404
(notably no-store from the gamedata 404 handler) is never pinned.
2026-10-01 18:00:48 +02:00
openhands dbaccd7cfc fix(gamedata): never cache a missing gamedata file
A missing gamedata file got no Cache-Control at all, because add_header
without `always` only applies to 2xx/3xx. Cloudflare then fell back to the
zone setting "Browser Cache TTL = 1 year", so the 404 came back as
`max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's
browser and at the edge. An icon requested while its import was still
running stayed a 404 for the rest of the year, even after the file existed.
That was the "some icons load, some don't" report.

Give every gamedata location a named 404 handler that sends no-store, and
split icons/ out as its own cache class: those files are rewritten under
the same name (repair-icons, reimport), so an hourly must-revalidate keeps
a repaired icon visible within the hour instead of days later.
2026-10-01 18:00:36 +02:00
openhands 4e036b08d5 fix(proxy): drop request rate limiting from gamedata entirely
/gamedata/* is served straight from disk by nginx; no request hits the CMS
backend or a database, so a request-rate limit protects nothing while
costing players their icons. A room load fires hundreds of these files in
one burst, which every limit turned into visible 503s.

Removed the static zone from the gamedata locations. Traefik's
epicnabbo-gamedata router likewise carries no rateLimit middleware.
/client/ and /nitro-client/ keep theirs, and the main route keeps the
30r/s page budget plus the server-wide connection limit.

Measured: 1000 icon requests fired fully in parallel now all return 200,
while 200 parallel requests on / are still rejected.
2026-10-01 17:56:48 +02:00
openhands f0dcf440a7 fix(proxy): split rate limiting into page and static zones
The single server-scope limit_req (30r/s) treated a page load and a room
load as the same thing. Loading a Nitro room fires several hundred gamedata
icons in one burst, which that zone answered with 503s, so icons showed up
late in the client.

Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata
and client asset locations, and apply the page-rate zone explicitly on the
main route instead of at server scope. Connection limit stays server-wide.

Measured: 900 icon requests in burst now all return 200, while 200 parallel
requests on / are still rejected.
2026-10-01 17:49:41 +02:00
openhands 4a1211a931 feat(proxy): add per-IP rate and connection limits
The edge had no limit_req/limit_conn at all, so a single client could
flood the Next.js backend and the Nitro client with unbounded parallel
requests. Traefik's logs already showed this: bursts of gamedata icon
requests answered with 429.

Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed
on the real client IP, applied at server scope so both cached assets and
proxied API routes share one budget. The burst is deliberately generous
because the Nitro client fetches gamedata and icons in bursts when
loading a room.
2026-10-01 17:25:49 +02:00
openhands e3c010f383 fix(proxy): raise nginx worker rlimit above worker_connections
nginx inherited systemd's soft LimitNOFILE of 1024, so every start logged
"2048 worker_connections exceed open file resource limit: 1024" and the
worker_connections value could not actually be reached.

Set worker_rlimit_nofile to 65536. Bounded from above by a systemd drop-in
at /etc/systemd/system/nginx.service.d/override.conf (LimitNOFILE=65536),
since the master's hard limit caps what workers may request.
2026-10-01 17:11:04 +02:00
openhands a6cc3cafa9 fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.

The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.

The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.

- `writeFurniData` now purges the gamedata edge tag itself. One place covers
  import, batch, resync, regen, nitro-editor, translate and dedupe. It is
  fire-and-forget and swallowed at every level: a stale edge copy is bounded
  by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
  `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
  and the content-addressed trees (c_images, album*, clothes) keep the long
  TTL. `must-revalidate` is the point — the client now revalidates instead of
  replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
  purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
  Cloudflare was unreachable at write time.
2026-10-01 15:16:48 +02:00
openhands cede541813 fix(catalog): route item-table writes to the catalog they belong to
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m16s
The items table is shared between both catalogs, but its four mutating
actions were normal-only: moving, reordering, creating and updating a
Builder Club offer wrote to catalog_items, so a BC edit either landed in
the wrong catalog or hit an unknown column.

Pass the catalog from the table through the actions and let the server
resolve it. BC rows have no price, points or currency column, so the BC
commands strip those fields instead of rejecting them. Moving and
reordering now share one command that locks the category and writes the
table for the same catalog, and BC writes revalidate the BC route.
2026-09-30 20:06:48 +02:00