Docker 29.1.3 ships BuildKit v0.26, which refuses to grant a build host
networking unless each caller passes --allow=network.host. All three rebuild
paths asked for it, and `docker compose build` has no flag to grant it, so a
rebuild failed immediately with "additional privileges requested". The live
container was never replaced, which is exactly the reported symptom: the site
kept serving the previous release after a rebuild.
Nothing in the build actually needs host networking. It uses the network only
for apk, pnpm and next/font/google — all outbound internet, which the default
bridge provides. Verified by building both the full runner image and the
migrations stage with --no-cache after dropping the flag.
Runtime `network_mode: host` stays: blue/green needs per-release host ports
(3002/3003) and nginx reaches each slot over 127.0.0.1.
The second gap is how a rebuild could still ship the wrong code. ci-deploy.sh
stamped every image with HEAD's revision label, and verify-deployed-release.mjs
only re-checks that same label, so a dirty working tree produced an image that
claimed to be release $sha while containing uncommitted code. docker-update.sh
already refused this; ci-deploy.sh now does too, before any build work.
The deploy gate found two defects in the previous commits, both mine.
0034_acl_midrank_revoke.sql never applied: it joined `acl_roles` on
`ar.model_type`, a column that table does not have (only
`acl_model_permissions` does). It now joins on the id and keeps the
`model_type` check where it belongs.
That hid a second, worse bug. The rank was extracted with
`SUBSTRING(slug, 7)`, but MySQL's SUBSTRING is 1-based and the digits start at
position 6, right after `rank_`. rank_10 therefore parsed as 0 and rank_7 as an
empty string, so every rank >= 7 would have lost exactly the grants the
migration exists to preserve — the ACL repair would have made things worse, not
better. Now reads from position 6.
Verified against a real MariaDB with a fixture covering rank_1, rank_6, rank_7,
rank_9, rank_10 and a non-rank slug: only the sub-7 roles lose their non-view
admin.* grants, the multi-digit and higher ranks keep everything, and the
non-rank slug is untouched. The mail index was checked the same way — it
applies idempotently and EXPLAIN confirms `users_mail_index` with rows: 1.
news.spec.ts expected to land on /me after signing in. That expectation predates
the `?from=` honouring added in 39332149, which lands a bounced admin back where
they were heading. The same step navigates to /admin/articles/new explicitly a
few lines later, so nothing depended on it; the assertion now covers the redirect
target instead.
The UI screenshot suite caught a regression from the previous commit: the admin
article form grew 3px and its baseline no longer matched.
The cause was the blanket `.btn` min-height bump from 40px to 44px. That was
the wrong way round — 40px already clears WCAG 2.2 AA, which asks for 24px, and
44px is a touch-target guideline. Growing every button for mouse users only made
each admin dialog and table 4px taller for no benefit.
The 44px floor now sits behind `@media (pointer: coarse)`, so finger input gets
the comfortable target and desktop keeps its density. A hybrid laptop still
uses the desktop metrics for its trackpad, which is the behaviour the previous
attempt got wrong in both directions.
The full suite passed but exited non-zero, which fails CI: three unhandled
rejections came out of commandocentrum's fire-and-forget audit call.
The cause is a real defect, not a test artefact. `auditAction` wrapped
`logAudit(...)` in a try/catch to honour "auditing must never fail the command
it describes", but logAudit is async, so the catch can never see its rejection.
A failing audit insert therefore surfaced as an unhandled rejection instead of
being swallowed — in production that is a request taking down over a logging
failure. The catch is now on the promise itself.
The test now mocks the audit service explicitly instead of leaning on the fake
db lacking `insert`, and asserts both that an entry is logged and that a
rejecting audit still lets the command succeed.
Closes the four remaining LOW items.
Search and news archive caching
- A leading-wildcard LIKE cannot use an index, so every /search section cost a
COUNT(*) scan plus an ordered page fetch, and /news did the same for its
archive. Both now cache: search per section for 30s, the archive for 60s
under the existing news revision so publishing an article drops it at once.
- Sections are cached independently, so one slow query cannot hold up the rest
and a failure is not cached as a result.
- The cached value is passed through cacheSafe() so the Redis path and the
in-process path return the same types; without it a cache hit would hand the
events grid a string where a miss hands it a Date, and it calls toISOString()
on that field. Dates are revived on the way out so the public signatures of
loadNewsArchive and loadPublicSearch are unchanged.
- Archive entries are keyed on the REQUESTED page rather than the clamped one,
so two requests that clamp onto the same page cannot alias each other.
motion/react out of the public bundle
- Converted the six public-facing users: the radio player, the typewriter text
(a motion.span with no animation props at all), the photo lightbox, the
animated counter, the footer CMS-info popup and the scroll reveal. That was
the actual entry points — the counter and the popup reach the public home page
and footer through static imports, so removing only the three originally named
would have left the library in the bundle anyway.
- Each animation moved to a CSS class, and the two that animate on exit now hold
the element for the length of the fade, which is what AnimatePresence used to
do.
- motion/react now only ships with /admin and the two already-lazy nav panels.
- Two safety fixes came out of this: the scroll reveal starts at opacity 0, so
it is forced visible under prefers-reduced-motion and via a <noscript> rule in
the root layout; and it now emits the .motion-reveal class, which the theme
panel's "Scroll Reveal" toggle selects and which previously matched nothing.
- The CMS-info backdrop became a real button in a pointer-transparent layer
instead of a handler on a static element, so click-outside-to-dismiss is
reachable by keyboard.
Fewer duplicate router refreshes
- Next.js re-renders the current route as part of a server action's own response
when that action revalidates, and applies it with a seeded navigation; the
router only skips its own update when the action did NOT revalidate. So the
refresh after such an action fetched the same tree twice.
- useServerAction takes an opt-in `revalidated` flag that skips it. It is opt-in
per call rather than derived from an action name, since a rename would
silently change behaviour. Applied to the two user-facing call sites whose
actions were verified to revalidate their own route.
Touch targets
- .btn was the one shared control at 40px; it and the lightbox and CMS-info
close buttons are now 44px, as is the password toggle (the auth input already
reserved 44px for it). The remaining 32px icon buttons pass WCAG 2.2 AA, which
only asks for 24px; enlarging those inside inputs and overlays was left alone
because it risks visual breakage that cannot be checked from here.
Closes the four HIGH/MEDIUM items left open after the previous pass.
Login lockout
- The only login limits were keyed on the client IP, so a distributed attempt
could grind on one account indefinitely. Added a per-account lockout with a
budget of 8 failures per 15 minutes.
- The bucket is keyed on the RESOLVED account id, not on the submitted string:
users may sign in with either username or e-mail and neither the lookup nor
the input normaliser folds case, so an input-keyed bucket would hand out a
fresh budget per spelling of the same account.
- precheckLogin and NextAuth's authorize share the bucket, so the pre-check
cannot be used to buy extra attempts and a client that skips it entirely is
still bounded. Both check the lockout BEFORE verifying the password: the
success path clears the counter, which would otherwise walk a locked account
straight back in on the right password.
- A successful login clears the failures, which needs two new primitives in
rate-limit.ts: peekRateLimit (read-only, does not consume a unit) and
clearRateLimit.
- Fixed a latent inconsistency while doing so: the in-process bucket capped its
counter at the limit while Redis' INCR kept climbing, so the two backends
disagreed about how far over the limit a key was. Both now track the true
count.
Mail lookup index
- Added an index on users.mail (0035). Password reset, e-mail verification and
the resend cooldown all resolve a single account from a submitted address and
were full table scans of `users`. Deliberately non-unique: legacy rows can
hold the same address more than once, so a unique index would fail to apply.
Resend captcha
- /verify's resend form triggers real outbound mail and was reachable with only
a cooldown. It now runs the configured captcha before the account lookup and
before any send.
Client message payload
- The root layout serialised the whole catalogue into every page. pages.admin
and admin are ~177 KB of the ~235 KB and are unreachable from the public route
group, so that layout now installs its own provider with the staff namespaces
removed. Nested providers replace rather than merge, which is why this has to
live in the segment layout. /admin, /mod, /client and /admin-next keep the
full set; a guard test fails if a public page ever references a staff
namespace.
Second review pass covering security, performance, admin tooling and the
public/room flows. All HIGH and MEDIUM findings from the audit are resolved;
nothing in this commit changes the visible feature set.
Authentication & session security
- CSP is now set on the request headers in the proxy, which is what Next.js
uses to derive the render nonce, so the nonce is effective.
- 2FA: an already-enabled user cannot re-enroll, the setup endpoint is
rate-limited per account, and confirmed codes are persisted so the second
secret no longer silently never applies.
- Password reset revokes the ticket, authTicket and all personal access
tokens, and bumps the token version so existing sessions die. The same
revocation is now wired into the staff-side password reset.
- /reset and /verify return a stable error code instead of raw text; the
mail lookups are ordered by id so duplicates cannot vary between runs.
- Resending the verification mail gets a per-address cooldown on top of the
per-user limit.
- Issue API tokens with the narrower radio/ticket ability set instead of "*".
Authorization & input handling
- Mid-rank staff can no longer keep dynamically granted non-view admin.*
permissions: existing grants are revoked by migration and the grant lookup
is restricted to "%.view". Rank guards use the dynamic super-admin check.
- Alerting a user is permission-checked and audited like the other tools.
- Material mutations (giveCredits/giveDuckets/giveDiamonds, the admin user
actions route, bulk user actions) are capped and rank-guarded, and bulk
ids are bounded.
- updateRoom / updateRoomItem write through a field allowlist, and items
may only be edited through their own room.
- Classnames reaching the filesystem are validated before use so a crafted
value cannot escape the asset directories.
- The word filter now also covers offline mails, guild forum threads and
replies, and user mottos.
- Media uploads are validated by magic bytes, /api/media requires the page
edit permission, APP_URL must be configured once mail is enabled, and the
diagnostics error route checks the fetch site header.
Admin tooling
- Secret settings render masked and cannot be overwritten with a blank or
an arbitrary raw key; radio credentials are new password inputs.
- Commandocentrum balance changes are audited.
- Admin list pagination reads the caller's per-page instead of the max, and
the log exporter caps offset and search length.
Performance
- Catalog translations are cached per module, with a cheap revision hash;
the public online count uses a stale window instead of hammering the DB.
- The cache warmup now primes the payload the home route actually reads.
- TopHeader batches its queries into one round trip, and LCP avatars load
eagerly.
- motion/react and sonner are no longer part of the root layout; the nav
dropdown and mobile nav panels are lazy client chunks. Anonymous visitors
again get the navigation chrome, and public pages get an edge cacheable
response.
Accessibility
- Nested <main> elements in phase pages became <section>; the page entrance
and route progress animations are pure CSS that respect reduced motion.
The active-runtime assertion required process.versions.node to equal
.nvmrc exactly, so any Node.js patch release broke `toolchain:check` and
the act CI run even though package.json engines (>=26.10.0 <27) supports
the newer runtime.
Keep .nvmrc and the Docker base image exactly pinned for reproducibility
(both still asserted), but validate the running runtime against the
engines range instead. Verified: toolchain:check, lint, typecheck,
i18n:check, hk:matrix:check and all 3391 tests pass.
The build script now runs every heavy command through
scripts/with-memory-cap.sh, which is bash (arrays, BASH_REMATCH). Alpine's
node image ships busybox ash, not bash, so the builder stage failed with
`sh: bash: not found` (exit 127): ci-deploy.sh could not build the image.
Verified: full docker build passes and compiles all 279 routes.
The host runs with vm.overcommit_memory=0 and no swap, so a process that
grows past free memory makes the kernel OOM-kill across the whole machine
-- the Turbopack build (commit 3d828a61) could take out the database,
nginx or the live release.
Add scripts/with-memory-cap.sh: it moves a command into its own systemd
scope with MemoryMax, so only that cgroup gets OOM-killed (verified: a
Turbopack build died at its 6GB cap, host untouched). Build/analyze/dev/
test*/typecheck now run under explicit caps; ulimit -v is only an explicit
opt-in because it bounds virtual address space per process and 10g/20g both
break V8-based builds. Docker and GitLab builds run in their own isolated
containers with a read-only cgroupfs and opt out explicitly (webpack +
--max-old-space-size stay their bound).
Measured: webpack build peaks ~6.5GB RSS, so 10GB leaves headroom within
the 23.5GB host.
- add theme-aware helpers: btn-brand, btn-glass-dark, auth-input, auth-label,
auth-alert, explore-pill, avatar-tile, aurora blobs and a hero scroll cue
- reorder backdrop-filter declarations so the glass blur survives the
production CSS optimizer in modern Chromium
- hero: aurora glow, frosted recent-users chip, premium CTA buttons and cue
- explore nav: icon pills with hover arrows; features: gradient icon tiles
- stats: single glass panel with column dividers; join CTA: aurora + ring
- auth forms: visible labels, icon inputs, eye/password toggle, gradient
submit and pill footer links
- auth pages: gradient card frame, aurora accents on the intro panel and
hover-lift avatar tiles; unified pill-shaped top bar
- drop the hard-coded register banner image in favor of the framed card
The deploy failed after the build, the migrations and the browser gate:
"Port 3002 is already in use". The holder was `epicnext-cms`, a compose
replica of release 6bffc537 that the daily scripts/docker-update.sh cron
had recreated at 03:30 with restart=unless-stopped. nginx serves the green
slot on 3003, so that replica was squatting the blue slot the next
candidate needed, and live traffic never noticed.
It got there because the updater's CI-ownership guard only tested
epicnext-cms-app. After a cutover to the green slot that container is
stopped, renamed and deleted, so the guard stopped firing while the host
stayed CI-managed.
- scripts/docker-update.sh: refuse a compose deployment on a CI host by
checking both slot containers and the nginx upstream, which is the only
thing that still marks the host as blue/green while a slot is idle.
- scripts/ci-deploy.sh: retire a compose replica of this checkout from
the candidate port before starting the candidate, so a stray replica
can never block a release again. Never a slot container, never the port
nginx serves; anything else still fails loudly in assert_port_free.
- Tests cover both directions: a squatting replica is removed and the
release lands, a replica on the live port is left alone.
The performance report measured nothing. It only read `entryJSFiles` from
each route's client-reference manifest, a field Turbopack emits and webpack
does not. When the build moved to webpack (3d828a61) every route fell
through to the "unavailable" branch, and because the report is informational
and exits 0 on an unavailable metric, nothing failed and the budgets quietly
stopped being enforced.
Derive the envelope from clientModules[*].chunks when entryJSFiles is
absent, which is the same source Next's own static-routes-info uses for
webpack builds. Webpack interleaves numeric chunk ids with file names in
those arrays, so ids are skipped by shape while a malformed chunk path still
throws — otherwise a broken manifest would quietly under-report a route.
entryJSFiles still wins when present, since it is per-segment and therefore
the tighter envelope, and the per-chunk origin label is shared rather than
the absolute node_modules path webpack records, which would otherwise bloat
report.json.
Six tests cover the webpack layout: id filtering, deduplication of a chunk
reached by several client modules, the origin label, the malformed-path
rejection, the no-chunks-at-all case, and entryJSFiles taking precedence.
Re-measured on the current build, all six routes are inside their budgets
again. Note /admin/studio/furni now sits at ~98% of its gzip limit, so one
more dependency on that route will trip it; docs/performance-budgets.md
records the webpack baseline numbers and how to recalibrate.
Verified: 3385 tests, typecheck and biome clean, and the report now emits
measured rows instead of six unavailable ones.
Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.
jobs-worker never ran
`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.
The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.
/api/health answered 200 with the database down
The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.
The runtime had no V8 heap cap
NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.
Storage ownership was only repaired for one path
ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.
nginx: robots.txt was a guaranteed 404, and TLS never resumed
`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.
Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.
Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
(key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
function-declaration handlers, with the reasoning that dropping them broke
the tree and save-on-Ctrl+S once already (704e3363)
- scope the remaining noArrayIndexKey / useExhaustiveDependencies exemptions to
the three files that need them, in biome.json instead of scattered comments
Storage, on a host that had grown to 81% disk:
- byparr starts a Firefox per request and never removes the profile it leaves in
the container's writable layer. With no volume mounted, nothing else reclaimed
it: 716 profiles / 6.8 GB in two days, ~1.7 GB/day. docker-prune.sh now removes
orphaned profiles, identifying live ones by the open fd in /proc/<pid>/fd rather
than by age, because browsers stay warm for ~27 hours here — longer than the
leak window, so no age threshold can be both safe and useful.
- bound the build cache properly: buildx treats --max-used-space and --filter as
mutually exclusive, so passing both silently dropped the 4 GB cap and the cache
reached 49 GB.
- escalate to the emergency prune when / drops below 8 GB free, so the bound holds
even if the schedule stops.
- clear multi-GB tmp_pack files left behind by a gc that was OOM-killed
mid-repack; git only removes those on the next successful gc.
- make setup-cron.sh append instead of replacing the crontab (`crontab -`
overwrites the whole file, which had been dropping the other scheduled jobs),
and run the prune daily rather than weekly to match the leak rate.
Volumes are still never pruned: mariadb-turbo-data is a database.
`home-login-form.tsx` and `login-form.tsx` were two ~240-line near-identical
components. Delete the former and give `LoginForm` a `variant` prop:
- `variant="page"` sr-only labels plus the register/forgot footer (/login)
- `variant="compact"` visible labels, no footer (homepage sidebar)
Field ids now come from `useId()`, so the two usages can never collide, and the
hardcoded "Show"/"Hide"/"Loading" strings are translated.
Localization of the login and register screens:
- `home-login-form.tsx` was entirely hardcoded English.
- `passwordStrength()` returned hardcoded "Weak"/"Fair"/"Good"/"Strong".
- `register.ts` returned only English strings. It now returns a
locale-independent `code` next to the message, and the form renders
`t(code)` with the English string as a fallback.
- Backfilled the new keys across all 25 locales, plus the login/register
strings that were still English in most of them. `ar`, `fi` and `ja` had
their entire login/register namespace in English and are now filled in.
Locale parity stays at 0 missing keys, as `i18n:check` requires.
Copy that did not match the enforced rules: the UI advertised "min 8 chars"
(EN) / "min 6 tekens" (NL) while registration requires 12 characters plus an
uppercase, a lowercase, a digit and a special character. Corrected in every
locale. `password-reset.ts` enforced only 6 characters and is raised to 12 to
match registration.
Accessibility: `login-form.tsx` had no `<label>`, no `id` and no `required` on
any field. All three are now present, and error banners are announced with
`role="alert"`.
Adds `src/i18n/auth-messages.test.ts`, which asserts every `RegisterErrorCode`
resolves to a non-empty message in all 25 locales; verified it fails when a key
is removed. The existing register tests now also assert the error `code`.
`next build` on Turbopack never completes on this app. The compiler is a
single native process whose RSS grows monotonically with no plateau:
0.9G -> 1.6G -> 2.8G -> 5.0G -> 5.5G -> 6.2G -> killed
It still dies with 4GB of swap attached, at 12GB RSS. The build workers are
only 0.17GB each, so `experimental.cpus` is not the lever either.
A `--max-old-space-size` cap cannot help: measured with a 2GB cap, RSS still
reached 8GB, because the memory is native Turbopack (Rust) memory rather than
the V8 heap. The cap added in 1c9ddcd4 was therefore inert and only created
false confidence, so it is dropped from the build script.
Webpack builds the same 329 routes in ~95s with a ~6GB peak.
Ruled out by measurement: the 25 bundled locale files (stubbing 24 of them
from 7.1MB down to 276KB still peaked at 11GB), worker count, and the
flatten/unflatten message pipeline (600 iterations cost 6.4s and settle at
39MB of heap).
Verified: `pnpm run build` exits 0, TypeScript passes, 279/279 static pages are
generated, and the standalone output boots and serves /, /login and /register.
/admin/studio/furni sat at 94.9% of its initial-JS budget (427436 of
450560 gzip bytes), so the next feature would have broken the build. Of the
98228 gzip bytes unique to that route, a large part is framer-motion.
This file uses motion twice, for one thing: a 150ms opacity fade on the result
pane when viewMode changes. Importing `motion/react` to get it pulls in
framer-motion's complete component library — 73 internal modules — plus its
render components, drag/gesture and projection code, none of which is
rendered here.
`motion/react-m` ships only the element factories: 2 internal modules, and the
same initial/animate/transition props, so the fade is unchanged. It exports the
elements flat rather than under a `motion.` namespace, so the import becomes
`div as Mdiv` and the two JSX tags are renamed to match.
I could not measure the resulting bundle here: the local build is OOM-killed
(exit 137) with the running containers on the host, so the actual saving is
unverified. The CI build reports it in build-reports, and the number in this
commit message should be read as a hypothesis, not a measurement.
Verified: typecheck clean, lint clean, and the 10 studio UI tests pass —
including the pane and navigation specs that exercise the view switch.
Seven suites read .env with a bare `process.env[key] = value`, which
overwrites whatever the shell already set. That made the DATABASE_URL from
the production .env authoritative, so a single environment variable was
enough to aim them at the live hotel database:
RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts
Three of those suites then repair the catalog in place: catalog-audit-repair-live
and catalog-repair-direct-live rewrite catalog_items and delete duplicate
classnames, and clone-bulk-import-live bulk-imports every cloneable item. None
of that is undoable, and nothing in their output said the target was
production rather than a sandbox.
Added src/test/live-env.ts with one shared loader, and pointed all seven suites
at it:
- Values already in the real environment win, so an explicit DATABASE_URL on
the command line is always respected.
- DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the
production one from .env.
- Anything that is not loopback is treated as production and redirected.
- Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning
saying the suite repairs the catalog.
Tests in src/test/live-env.test.ts run the loader against a temporary .env so
the real project file is never read, and cover the redirect, the shell
override, non-loopback detection, the opt-in and quote stripping. A second
block asserts each of the seven suites no longer contains an inline
`process.env[...] =` assignment. Verified four of them fail against the old
loader.
This does not enable the suites; they stay gated behind their RUN_* flags.
It only removes the possibility of them silently hitting production.
Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean.
The previous commit removed the `cms` and `cms-green` services from
docker-compose.yml. That was overreach and it broke two tests:
- src/lib/docker-build-contract.test.ts asserts the compose build passes
NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}, so a compose-built image
carries its release id.
- scripts/proxy-config.test.mjs resolves `docker compose config` and asserts
the `cms` service's host networking, volumes, healthcheck and image tag.
Both encode that docker-compose.yml is a maintained deployment surface, not a
leftover. Removing it was not my call to make while fixing a deploy.
Restored verbatim. The stray container that actually blocked port 3002 is
already gone, and nothing recreates it: there is no systemd unit or pm2
ecosystem that runs `docker compose up`, and `restart: unless-stopped` only
applies to a container that still exists. So the blocker is resolved by the
container removal alone, and compose stays intact for manual and reviewed use.
The deploy could not start its candidate because port 3002 was held by
`epicnext-cms`, a `docker compose up` replica built from the `local` image and
serving no traffic. Everything else in the pipeline was healthy: the image
built, the news browser gate passed and migrations were current.
The container was unusable for this pipeline for two reasons. It ran a
different image than any release, and its name did not match the slot the
deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms`
while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to
agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never
could. deploy.sh already documents that compose "never managed the release
that actually ran", so the service was stale by its own account.
Removed the stray container and dropped the `cms` and `cms-green` services (plus
the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot
resurrect a replica that permanently occupies a blue/green slot. byparr is
untouched.
Also fixed the diagnostic from the previous commit, which blamed every running
container. `docker ps --filter publish=` returns nothing for --net=host
containers, so the fallback listed all of them and buried the real holder
among seven innocent ones. It now resolves the listening PID from `ss` back to
its container through /proc/<pid>/cgroup and names only that one, with the
exact `docker rm -f` command to run.
Verified: port 3002 free, live release on 3003 still serving
(status ok, database and redis true), deploy simulation 26 passed, typecheck.
Two follow-ups from the blocked deploy.
The rollback path called `docker logs` and `docker rm -f` on the candidate
unconditionally. When the port check refuses to start it, the container was
never created, so both printed "No such container: epicnext-cms-app" — noise
that looked like a second, unrelated failure and buried the real message.
Both calls are now guarded by `docker inspect`.
assert_port_free() now reports which container holds the port and flags it when
it is not a blue/green slot this script manages. The previous output listed
every container and said only "port already in use", which is a dead end: on
this host the holder is `epicnext-cms` (a `docker compose up` replica on port
3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ
because docker-compose.yml pins `container_name: epicnext-cms` for the `cms`
service; slot B happens to match, which is why 3003 deploys fine and 3002 never
can. The message now names the squatter, explains that live traffic is
unaffected, and gives the next action.
Deploy simulation: 26 passed.
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.
The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.
Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.
Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.
Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.
read_active_port() now orders its sources by how well they describe reality:
1. The nginx upstream file. It is the only source that says where public
traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.
answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.
start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.
Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
The previous commit dropped biome-ignore comments to clear
useExhaustiveDependencies diagnostics and, in doing so, also deleted the
dependencies themselves. Six components were left with effects that no longer
react to the state they read. Every one of these is a real behaviour
regression, not a lint preference:
- health-check-client: checkEmulator is a function declaration, so it gets a
fresh identity each render. As an effect dependency that re-fires the effect
after every setState, polling /api/admin/devops/health in a loop. Wrapped in
useCallback so the identity is stable.
- article-recovery: reload restarts the autosave timer for the "Retry recovery"
button. Without it in the deps that button is a no-op. The counter had been
renamed to _reload to satisfy the unused-variable rule.
- catalog-integrity-panel: same pattern; refresh starts a new read-only scan,
so the rescan control did nothing.
- catalog-search: refreshKey re-runs the query after a bulk edit, so results
were not refreshed after catalog edits. The selection-reset effect also lost
catalogType, so switching catalog no longer cleared the selection.
- catalog-image-picker: dropped debounced (the search term) and name (the
error reset), so image search and error state no longer reacted to input.
- icon-picker: dropped iconImage, so a failed load left the placeholder on the
next icon too.
Each restored dependency carries a biome-ignore with the reason it is
load-bearing, so the diagnostic can be re-derived instead of silently
disappearing again.
Verified: typecheck, lint clean on all six, unit 3315 passed, integration 20
passed, UI 72 passed / 2 skipped.
Three failing test suites blocked CI. All three were test defects, not
application bugs.
Integration tests (integration/database.test.ts)
------------------------------------------------
The suite set NODE_ENV=test, which makes cache.cached() short-circuit both
its Redis read (src/lib/cache.ts:226) and its write (:249). A suite whose
stated purpose is exercising the real Redis path therefore never touched
Redis. Switched to NODE_ENV=development, the only non-production value
src/env.ts accepts, so the shared-cache code paths are genuinely covered.
Three assertions then needed correcting for real Redis semantics:
- `await cache.cached(...)` followed by `.resolves` can never hold: await
yields a value, not a Promise. Assert the value directly.
- A cached negative result is stored as the JSON encoding of null, so
`redis.get(key)` returns "null", not null.
- The news negative-cache key does not exist at all, so `ttl()` returned -2.
Now that the write path is live the key is created and the TTL assertion
holds as originally written.
UI tests (src/app/admin/prefixes/prefix-dialog.tsx)
---------------------------------------------------
The form-reset effect had `isOpen` removed from its dependency array. The
component returns null when closed, so the effect only ever ran on mount:
reopening the dialog no longer cleared the fields and a dismissed-but-
unsaved edit reappeared. Two tests in e2e/ui/unsaved-changes.spec.ts caught
this. Restored the dependency and documented why it is load-bearing.
The remaining edits in this branch drop stale biome-ignore comments that
suppressed useExhaustiveDependencies and noArrayIndexKey diagnostics. Where
the suppression had been load-bearing for behaviour, the underlying
dependency is now listed explicitly rather than silenced.
Verified: check (toolchain, audit, lint, i18n, typecheck), unit 3315
passed, integration 20 passed, UI 72 passed / 2 skipped.
- Remove random TTL jitter to prevent unpredictable cache drops
- Add deterministic LRU eviction with proper entry cleanup
- Improve cache deduplication to prevent duplicate computations
- Skip Redis I/O during tests for faster, more stable execution
- Optimize depth calculation in catalog tree nodes
- Maintain backward compatibility and full test coverage (3331 passed)
Icons are plain .png, a cacheable extension by default, so the zone's
"Browser Cache TTL = 1 year" pinned them to max-age=31536000 regardless of
the 300/3600/604800 that nginx sends per class. Extend the edge rule to
/gamedata/ and keep respect_origin, so the nginx header wins and a 404
(notably no-store from the gamedata 404 handler) is never pinned.
A missing gamedata file got no Cache-Control at all, because add_header
without `always` only applies to 2xx/3xx. Cloudflare then fell back to the
zone setting "Browser Cache TTL = 1 year", so the 404 came back as
`max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's
browser and at the edge. An icon requested while its import was still
running stayed a 404 for the rest of the year, even after the file existed.
That was the "some icons load, some don't" report.
Give every gamedata location a named 404 handler that sends no-store, and
split icons/ out as its own cache class: those files are rewritten under
the same name (repair-icons, reimport), so an hourly must-revalidate keeps
a repaired icon visible within the hour instead of days later.
/gamedata/* is served straight from disk by nginx; no request hits the CMS
backend or a database, so a request-rate limit protects nothing while
costing players their icons. A room load fires hundreds of these files in
one burst, which every limit turned into visible 503s.
Removed the static zone from the gamedata locations. Traefik's
epicnabbo-gamedata router likewise carries no rateLimit middleware.
/client/ and /nitro-client/ keep theirs, and the main route keeps the
30r/s page budget plus the server-wide connection limit.
Measured: 1000 icon requests fired fully in parallel now all return 200,
while 200 parallel requests on / are still rejected.
The single server-scope limit_req (30r/s) treated a page load and a room
load as the same thing. Loading a Nitro room fires several hundred gamedata
icons in one burst, which that zone answered with 503s, so icons showed up
late in the client.
Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata
and client asset locations, and apply the page-rate zone explicitly on the
main route instead of at server scope. Connection limit stays server-wide.
Measured: 900 icon requests in burst now all return 200, while 200 parallel
requests on / are still rejected.
The edge had no limit_req/limit_conn at all, so a single client could
flood the Next.js backend and the Nitro client with unbounded parallel
requests. Traefik's logs already showed this: bursts of gamedata icon
requests answered with 429.
Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed
on the real client IP, applied at server scope so both cached assets and
proxied API routes share one budget. The burst is deliberately generous
because the Nitro client fetches gamedata and icons in bursts when
loading a room.
nginx inherited systemd's soft LimitNOFILE of 1024, so every start logged
"2048 worker_connections exceed open file resource limit: 1024" and the
worker_connections value could not actually be reached.
Set worker_rlimit_nofile to 65536. Bounded from above by a systemd drop-in
at /etc/systemd/system/nginx.service.d/override.conf (LimitNOFILE=65536),
since the master's hard limit caps what workers may request.
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.
The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.
The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.
- `writeFurniData` now purges the gamedata edge tag itself. One place covers
import, batch, resync, regen, nitro-editor, translate and dedupe. It is
fire-and-forget and swallowed at every level: a stale edge copy is bounded
by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
`max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
and the content-addressed trees (c_images, album*, clothes) keep the long
TTL. `must-revalidate` is the point — the client now revalidates instead of
replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
Cloudflare was unreachable at write time.
The items table is shared between both catalogs, but its four mutating
actions were normal-only: moving, reordering, creating and updating a
Builder Club offer wrote to catalog_items, so a BC edit either landed in
the wrong catalog or hit an unknown column.
Pass the catalog from the table through the actions and let the server
resolve it. BC rows have no price, points or currency column, so the BC
commands strip those fields instead of rejecting them. Moving and
reordering now share one command that locks the category and writes the
table for the same catalog, and BC writes revalidate the BC route.
The previous commit taught bulk editing and delete-with-restore about the BC
catalog. Neither actually worked, and one of them was destructive.
`catalog_items_bc` has six columns: id, item_ids, page_id, catalog_name,
order_number, extradata. There is no price, points, currency, offer_id, limit or
membership column on it. The bulk path read and wrote columns that do not
exist, and the UPDATE was aimed at catalog_items while the SELECT came from
catalog_items_bc — so a BC category move wrote into the normal catalog. Two
tests now pin that pairing: reads and writes have to stay in the same table.
Underneath it the BC table was never being read at all. The inline editor
fetched `/api/admin/catalog/items?pageId=N` without the catalog, so opening a BC
category showed the normal catalog's offers, and the route selected BC rows
directly instead of going through the loader, skipping the furni enrichment the
table needs to render anything but a bare caption. Both catalogs now take the
same path, and the catalog is in the fetch callback's dependencies — without
that, a switch keeps reading the previous catalog's rows through a stale
closure.
Because a BC offer has no price, the editor no longer offers one. The server
refuses price, points and currency changes with a readable message instead of
letting them reach the database as an unknown-column error, and a BC bulk edit
is what it can actually be: a category move.
BC deletions also went through a bare DELETE, which made them the one catalog
mutation with no way back. They now keep their rows and hand back a restoreId
like the normal ones. The catalog is recorded in the audit target rather than
in the payload, so a restore can never put a BC row into the normal offers
table.
A bulk offer edit is the catalog mutation that rewrites hundreds of rows at
once, and it was the only one writing nothing to the staff activity log: 22 of
the 43 catalog actions logged, this one did not. The entry it now writes says
what changed, not just that something did, because the log has no undo of its
own and "bulk updated 200 offers" cannot answer the question it exists for.
Deleting offers had no inverse at all. Every removed row is now kept at delete
time and the caller gets a restoreId back, so an accidental multi-select is a
click rather than a hand-edit of the table. The undo toast covers the common
case; a RecentDeletionsPanel holds the same records so a delete noticed later is
still reachable. Three refusals guard it: an id that another offer has since
taken, a category that no longer exists (which would leave an offer that sells
nowhere and shows under no page), and a delete whose restore record cannot be
written — that one rolls back rather than deleting without a way back. Reading
the audit row FOR UPDATE is also what stops two restores of one deletion from
both inserting.
sendCatalogUpdate() overwrote hotel-status.json on every write, so "which
imports reached the hotel" was answerable for the last attempt only, and a
failure two imports ago was gone by the time anyone looked. That file is now
also appended to as a bounded 50-entry tail.
The tree route carried four copies of the same page-select-plus-counts
shaping, of which the BC branches had already drifted: one counted offers
through the VARCHAR-tolerant helper, the other inline and swallowing errors.
All of it is one readPages() now, and readFullTree sends both catalogs through
one depth computation instead of delegating normal to getTreeFlat while
computing BC here — a split that left two implementations behind one function
name. getTreeFlat is gone. The BC ancestor walk also went from 20 levels to 50,
matching getAncestors, so a deeply nested catalog no longer loses its
breadcrumb.
Bulk editing reaches the BC catalog, which previously had no way to edit or
duplicate offers in bulk. The catalog is part of the operation identity now, so
replaying one request key against the other catalog is not mistaken for the
same work.
Integration tests failed to import: the next/cache mock supplied only
revalidatePath, and catalog-totals calls unstable_cache at module scope.
The previous commit made imports update the Studio without a reload, but the
guarantee only held inside the tab that started the import and only as long as
every read succeeded. Four holes were left, and this closes them.
A session that mounted the tree before an import kept the pre-import tree for
the rest of its life, because ensureCatalogTreeLoaded() was a once-per-session
no-op. It now asks the server whether what it holds is still current. The answer
is a revision: sendCatalogUpdate() already runs after every catalog write, so it
bumps one, and clients read it on mount, on focus, on a 20s poll and from other
tabs over a BroadcastChannel. An import that finishes in another tab, another
browser or the job worker now lands here too.
A failed read used to be swallowed, which is the worst outcome available: the
rail kept showing pre-import counts as if they were current and nothing said so.
The snapshot now carries the error, the rail shows it with a retry, and the
previous tree stays on screen because stale beats empty.
Every settled import pulled the entire flat tree, which is the one payload that
grows with the size of the catalog. The revision doubles as the ETag on
mode=full, so an unchanged catalog answers 304 and the poll costs a file read.
An import could also report success for an offer the hotel will never sell: a
hidden or disabled page, an item_ids that misses the furni id, a zero amount.
importSingleFurni reads its own row back and reports each of those as a warning,
where the import report already is, instead of leaving it to surface as "the
import did not work" in the client.
Finally, the catalog items table no longer falls back to router.refresh() —
onRefresh is now required, so every mutation ends in a refresh of the caller's
own data instead of a route re-render that threw away editor state and scroll
position. useServerAction keeps its default, because 47 callers across the app
depend on it. The 750-line CatalogTree in catalog-tree.tsx was dead code that
kept its own stale tree and three more router.refresh() calls; only CatalogIcon
and LAYOUT_COLORS are still imported, so the rest is gone.
Tests: the store now covers revisions, 304s, probe failures and error recovery;
a jsdom test mounts a consumer and asserts the tree updates in place with no
navigation; the old organize-imports e2e asserted nothing about the endpoints
the code actually calls, and is replaced by one that asserts a cross-tab write
lands in the mounted categories without a reload.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The live catalog store only covered part of the import surface. A durable
job settled, a sync queue drained, a .nitro upload or a clone run left the
Studio rail and the stats bar showing pre-import numbers until the page was
reloaded, and the Catalog Manager kept a second tree that never saw writes
made elsewhere in the session.
Every one of those paths now pulls the tree again, and the refresh carries
the totals with it: importing writes catalog rows server-side, so the counts
the store holds were stale for the rest of the session.
- refreshCatalogTree shares one request between concurrent callers and queues
a single follow-up read when a write lands mid-flight, so a burst of edits
costs at most one extra read.
- useFurnitureJobs treats its first payload as a baseline, so a page load no
longer replays every past import as "just settled", and hands the settled
jobs to the callback.
- The Catalog Manager pushes its own mutations into the store and re-reads its
active tab when the store changes.
- The 30s unstable_cache on the admin totals is now tagged and invalidated from
every catalog write, including the import worker, so it no longer survives an
import even across a hard reload.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The resync route rebuilt every entry from items_base plus the *official
Habbo* furnidata. A classname that only exists on a retro hotel (leet.ws
and friends) is absent from the official set, so lookupOfficialHabboFurni
returned null and the entry was written with revision 0, category
"unknown" and an empty description — even though the clone import had
those exact values available at import time from the source's own
furnidata.
- resync/route.ts: when a classname is genuinely missing from official
Habbo, fall back to the configured clone sources. Their furnidata is
indexed by normalized classname and each entry is coerced into the
OfficialHabboFurniEntry shape, which is the same JSON shape, so it drives
the existing buildFurniEntry fallbacks for revision, category, name,
description, defaultdir, partcolors, specialtype, furniline, environment,
rare and bc. items_base stays authoritative for id, spriteId and dims,
and public_name still wins over the source name, matching the import.
The index is memoized per request, not at module scope: a module-level
cache would pin the source list for the life of the process and a source
added later would never be picked up. fetchSourceFurnidata already caches
per URL, so this costs one parse rather than a network round-trip.
Disabled sources and sources that fail to respond are skipped, so an
unreachable hotel degrades to the previous items_base-only behaviour
instead of failing the run. Official Habbo still wins whenever it has the
classname, so existing behaviour is unchanged for everything but the
retro-only case.
Applies to every resync mode, so the pre-existing ?missing=1 sweep picks
this up too.
"Missing furnidata" was only a filter in the Studio status dropdown, so
imported items whose classname was absent from FurnitureData.json could
be found but not fixed from that screen. Only the Catalog Audit page could
repair them, and only globally.
Adds the same shape of quick action that "no nitro" already had:
- studio-client.tsx: a "N no furnidata" shortcut next to the "N no nitro"
button that sets the missingFurnidata status filter, and a bulk "Add
missing furnidata (N)" button for the selected rows. Both only appear
when there is something to act on. Rows that come back repaired flip
to hasFurnidata: true so the badges and counts update in place; rows
the server reported in errors keep their state.
- resync/route.ts: accepts an optional { classnames: string[] } body to
target exactly the selected rows. classnames are resolved through the
same normalized local index the listing uses to decide hasFurnidata, so
the rows written are the rows flagged as missing. The upsert is already
idempotent, and RCON updateCatalog + updateItems run afterwards so the
emulator picks the new entries up.
Also clears the Studio furnidata cache after a write, which this route
never did: without it the listing kept serving a stale hasFurnidata for
up to the 30s cache TTL, so a repair looked like it had done nothing.
PERMS is now imported from permission-slugs (identical re-export) so the
route no longer pulls next-auth into tests.
- studio-filters.test.ts: pins the missingFurnidata branch, in particular
that an unchecked item (hasFurnidata undefined) is not treated as missing.
The existing ?days / ?missing / ?broken / ?all modes are unchanged; the
body is only consulted when it carries a classnames array.
The emulator complained "Sibling order 2 is used more than once" 74 times
in production, covering 19 pages across 12 parents. The cause was that
nearly every page-creation path passed orderNum 0, so each new page
collided with whatever sibling already sat at 0 or 1, and the Park
placeholder pages all shared the sentinel values 99 and 999.
The production rows themselves are repaired out of band (renumbered 1..N
per affected parent, ordered by order_num then id so the existing visual
order is preserved, plus one dangling catalog_items row whose
items_base no longer existed removed). This commit stops it recurring.
- hierarchy.ts: add nextFreeSiblingOrder(), which ignores -1 and 0 as
the same root set, treats a missing orderNum as 0, and returns an order
strictly above the highest sibling in use. An explicit order is still
honoured whenever it is free, so callers that genuinely want a position
keep it.
- page-commands.ts: createPageCommand resolves the real order through
nextFreeSiblingOrder instead of writing the requested 0 straight through.
Note that furni-import.ts and upload-import.ts still take their order
from the furnidata catInfo.order, so two categories carrying the same
furnidata order can still collide. That is caught by the emulator audit
and repaired by fixEmulatorIssues(), but it is not prevented here.
Organising imports, the Studio furni batch, the catalog totals and the
"import from a source" stats all used to need a full page reload, or at
best a router.refresh() that re-rendered the whole admin route, before
anything on screen reflected what the import had just written.
- live-catalog-merge.ts (new): pure tree and total arithmetic. Applies a
delta of created pages, added offers and moved offers, recomputes depth
for the touched subtree, bumps parent child counts and the item totals.
Returns the input untouched when a delta is empty, so subscribers can
bail out instead of re-rendering. Depth resolution tolerates a parent
cycle in a dirty DB and still terminates, matching getTreeFlat.
- use-live-catalog.ts (new): one module-level store exposed through
useSyncExternalStore, so every consumer shares a single instance without
threading a provider through the admin layout. Deltas only apply to the
"normal" catalog, so public and public_handlers trees stay separate.
seedCatalogTotals() takes the first server value per mode and never
overwrites it afterwards, so a later hard render cannot make the header
totals jump backwards.
- actions/catalog.ts: organizeImportFurni now reports each group through
the new OrganizedPageChange, carrying parentId, pageLayout, the icon,
isNew and the per-source movedFrom counts, so the client can fold the
result into the tree without reading the page back.
- organize-imports-dialog.tsx: drops useRouter and router.refresh(); the
response is applied as a delta the moment the run finishes.
- studio-client.tsx: reads the tree from the store instead of freezing it
with useState(initialTree), loads it on mount when empty, and refreshes
it once a batch import settles. The batch is server-side and derives its
import pages from furnidata, so that one path re-reads the tree via
GET /api/admin/catalog/tree?mode=full rather than trusting the delta.
- studio/furni/page.tsx: stops calling getTreeFlat() and no longer passes
initialTree; the store is the single source of truth for the rail.
- import-clone-client.tsx: tracks which items are already present, so
present and clonable update per cloned row instead of only at the end.
- catalog-manager-dialog.tsx: seeds the totals once and renders the live
values, so the header reflects an import that just ran.
- e2e/ui/fixtures/entry.tsx: drops the removed initialTree prop.
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
picks the one that really answers; the upstream file only serves as a
fallback when zero or both slots respond. A stray 'docker compose up' (or a
clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
only seed it when missing, never overwrite what a deploy wrote. This is the
root cause of tonight's 502: a nginx-sync run reset the snippet (written to
green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
incoming CF-Connecting-IP, 403 any other peer that presents one
(spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.
- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
blue/green upstream snippet; config backed by scripts/nginx-sync.sh
(idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
no-op unless Cloudflare is configured; scripts/cf-purge.sh and
cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
The previous dependency upgrade rewrote the file with two-space indentation, which violated the Biome formatter setting (indentStyle: tab) and broke `biome check .`.
Patch release, bug fixes only, no breaking changes. Two entries are
relevant to this repository: a stack overflow when spying on
Set.prototype.add, and the hanging-process reporter switching to its ESM
entrypoint, which drops why-is-node-running 2.3.0, siginfo and stackback
in favour of why-is-node-running 3.2.2.
Verified with the CI unit command: 3302 tests pass under --maxWorkers=4,
the coverage run clears its thresholds, pnpm deps:audit reports no known
vulnerabilities, and the lockfile stays consistent under
--frozen-lockfile.
The timeout constant made the first it() line exceed the line width, so
`biome check .` failed with a format error. Reformat and confirm the
three publication tests still pass.
These drive real git processes against a local bare remote, so their cost
is process spawns competing with every other Vitest worker. Measured on
CI they take 23-30s each, and the 30s override was crossed by 37ms, so the
run failed on wall-clock rather than on behaviour.
Replace the three hand-picked 30_000 values with one documented constant
at 120_000, which keeps a genuine hang visible while clearing the observed
spread. The global default stays at 10s so nothing else is loosened.
Dependency updates are handled manually, so the scheduled Renovate job
only cost a daily privileged Docker run on the deploy host. The empty
cache directory it maintained is gone too, and the operations note now
records that updates are manual instead of describing bot behaviour.
load-env.ts reads .env.local before .env and gives it precedence, so a
developer override file can hold real secrets. Only .env was ignored, so
that file was one 'git add .' away from being committed.
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy
directory's .env. When that variable was missing the deploy had already
built an image and run the browser gate before pnpm db:migrate aborted on
an empty value, so a release was paid for in full and then thrown away.
Check for the variable right after the .env is copied, before the build,
and say plainly that the live release was not touched. The deploy test
fixture gains a DATABASE_URL so it mirrors a working deploy directory
instead of the broken one.
The admin.studio.nitroCleanup section (102 keys) only existed in en and
nl, so 23 locales fell back to English for the entire Nitro Cleanup
panel. The referrals and dailyRewards keys were missing from the same
23 locales, and en itself was missing 6 keys that nl had.
Add the missing keys to every locale with translations, so all 25
locales now carry the same 6063 keys.
Add a gamedata cleanup to the Nitro Cleanup panel: a read-only preview
plus an apply run that dedupes FurnitureData classnames and removes
rows and figure entries that reference nothing.
Three passes run in a fixed order, because cleanFigureMap has to precede
cleanFigureData: dropping the part that points at a set is what makes
that set unreferenced.
A pass refuses to write when it would delete more than maxRemovals rows
(default 500) and reports the reason, a wrong asset directory otherwise
turns every row into an orphan and one call would empty the file. Passes
that would act on empty input (no libraries, no sets) treat that as a
missing file rather than as a reason to delete everything. Every write
copies the file to a timestamped backup first, so a pass that turns out
to be wrong can be undone by hand.
The plan reads FurnitureData once and hands the parsed copy to both
furniture passes; the file is tens of megabytes in a real deployment.
A bundle whose texture is already VP8L was returned untouched, so a
stale spritesheet.meta.image survived the normalisation and the client
could not find the texture member. Rebuild the archive in that case and
reuse the existing VP8L bytes instead of decoding them again.
Two more .nitro entry points in the main import path still wrote the
supplied buffer verbatim: an attached `providedNitro` and a bundle pulled
back by `resolveMissingNitro`. Both are real furniture imports, so they
could still land a PNG texture while the SWF, clone and upload paths
produced WebP.
Route both through the same normalisation, falling back to the original
bytes with a warning if the texture cannot be decoded.
Importing from a hotel wrote the downloaded .nitro to disk untouched, so
official PNG textures stayed PNG and only SWF imports ended up as WebP.
Every Studio import should produce the same format regardless of where the
bytes came from, so both clone and upload paths now run the bundle through
toWebpLosslessBundle.
The helper decodes the texture and re-encodes it with the same VP8L options
the SWF importer uses, so the artwork round-trips bit-for-bit, and lets
createNitroBundle relabel the member and repair the meta.image pointer. A
bundle that is already lossless WebP is returned untouched, making the
operation idempotent and safe to run on re-import. A colour variant that
shares a library keeps the member base name it arrived with.
A texture that cannot be decoded keeps its original format with a warning
instead of failing the import: the bundle is valid, and losing a furniture
item over a codec edge case is worse than a slightly larger texture.
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a
third-party tool that ships a WebP member while still pointing
spritesheet.meta.image at a .png was accepted and stored as-is. The client
resolves the spritesheet through that pointer, so the result was a file
that validates fine and then renders nothing.
Re-write the bundle through createNitroBundle on import, which labels the
member from the actual bytes and repairs the pointer. No texture is
re-encoded, so the bytes stay identical, and the member keeps the base
name it arrived with so `chair*2` colour variants that share the `chair`
library are not renamed.
Newly converted .nitro bundles now store their spritesheet as WebP VP8L
instead of PNG, so imports land much smaller without changing a single
pixel. The texture member and spritesheet.meta.image are both labelled
from the actual bytes, never from a caller's assumption.
- encode through sharp with lossless and exact, so colour hidden under
alpha 0 survives; this mirrors ImageSharp's TransparentColorMode.Preserve
- detect PNG/WebP by magic bytes and reject anything the client cannot
render, on create, download and upload paths
- keep the source format when deriving size-32 sheets, scaling composites
and editing metadata, so existing bundles are never silently rewritten
- report fidelity in the studio: the compression panel re-encodes with the
same options the importer uses, so it cannot drift and invent false
warnings, and shows PNG/WebP size estimates
convertSwfToNitro and buildSpritesheet are now async, so the worker, the
main-thread fallback and every import call site await them. PNG stays
supported for existing bundles and icon sidecars are untouched.
`next build` had never completed on this host, so three real defects were
sitting in the tree untested. All three are now fixed and the build is green.
- The build was not memory-bound the way it looked. Turbopack's builder reached
20.5GB RSS and died, and raising `--max-old-space-size` could never have
helped: that flag caps the V8 heap, while the 20GB sat in Turbopack's own Rust
allocator. The first symptom was misleading because the process doing the
allocating is a grandchild of `npx`, so watching the direct child shows a
95MB shim the whole time. Building with `--webpack` puts the build back under
the JS heap, where the flag actually applies: peak 5.9GB, 150s, exit 0.
- withAdmin's second parameter was typed `{ params?: ... }` and given a `= {}`
default, which made it optional and `RouteContext | undefined`. Next's
generated route types assert that argument against `ParamCheck<RouteContext>`
and reject it, across 113 route files. `tsc --noEmit` cannot see this, because
Next only adds `.next/types` to the project during a production build — so the
type check that everyone runs locally was structurally incapable of catching
the only type error that blocks a deploy. `params` is now required, which is
also what the code already assumed: it is awaited with no guard. The 35 test
call sites that invoked a handler with one argument now pass a real context,
and the await got a guard so a direct internal call cannot turn a missing
context into a 500.
- `src/app/api/admin/import/furni/route.ts` re-exported `ensureDirectories` and
`importSingleFurni` for "backward compatibility" that nothing used; the batch
route imports from `@/lib/services/furni-import` directly. Next rejects any
value export from a route module that is not an HTTP verb or config, so this
had been breaking the build for as long as it existed. Removed.
- `isomorphic-dompurify` builds its server-side DOM through jsdom. Bundled, that
pulls jsdom's `browser/default-stylesheet.css` into the server chunk, where the
path no longer resolves, and page-data collection dies with ENOENT on every
page that sanitizes HTML. Marked external so Node resolves it from
node_modules and the standalone tracer includes it.
The remaining build warning is a pre-existing circular dependency between
chunks that share the webpack runtime. It costs hash reuse, not correctness, and
is left alone rather than churned here.
Verified: build exit 0, 276 static pages generated, 3223 tests pass, tsc and
biome clean.
Follow-up to f81b114b, addressing the three ways that commit could make things
worse rather than better. All three were verified against the real database or
by breaking the test and watching it fail.
- The grace window is now capped at 120s. A window is a cushion for the TTL
boundary, not a second TTL, but the call sites treated it as the latter: the
5 min values/staff routes and the 10 min teams route asked for a window as
long as or longer than their own TTL, so a single large staleMs silently
doubled how far behind a value could be served. Nothing marked those as
unsafe, because nothing looked wrong. The cap lives in the cache rather than
at the call sites so no future route can reintroduce it. Routes that asked
for less than 120s (the 10s online poll, the 20s news cache) are unchanged,
so their intended cushion still does its job.
- A request no longer queues behind an arbitrarily old render. Sharing a render
is what collapses a cold-cache stampede into one render, but a hung render
used to hold up everyone who arrived after it. A newcomer past 2s now serves
the placeholder instead of waiting, reusing the ImagerUnavailableError path
that "both upstreams down" already takes. The caller that actually started
the render keeps waiting, which is correct: it is the one whose image this
is. When the join window is removed the new test hangs for the full 10s it
was meant to prevent, which is the tail this bounds.
- The information_schema row-count estimate is gone; the counters are exact
again. Running it against the live database: users 165, rooms 92, camera_web
0, and the estimate was 0.00% off on all three. At 165 rows an index scan is
cheaper than the extra round trip the estimate needed, so the optimisation
bought nothing and traded a guaranteed-correct member count for an
approximation that InnoDB would only make less accurate as the table grows.
The exactness is now pinned by tests: a real zero stays zero, a database
error propagates instead of becoming a number, and each counter counts the
table it claims to. The module stays, because the homepage and the boot
warm-up writing different values to the same cache key is its own bug.
The module comment records the measured numbers, because "COUNT(*) is too slow"
sounds true in the abstract and is false here.
3223 tests pass.
Three separate things that were each costing more than they needed to on the
hot path.
- Single-flight avatar renders. The disk cache was checked first and a miss
went straight to the upstream, with nothing shared between callers, so a page
requesting dozens of avatars at once turned N concurrent requests for one
figure into N renders. A render is the most expensive operation this app
does, and the duplication happened exactly when the cache had nothing to
offer. Eight concurrent requests now cause one render instead of eight. The
map lives on globalThis because Next can evaluate the module more than once
per process, and two copies would each start their own render.
- Let public read-only routes be cached by a shared cache. Every JSON response
was `cache-control: no-store`, so a CDN in front of the app could not answer
any of it and every request reached the origin. publicCacheControl() opts a
route in with s-maxage and stale-while-revalidate, using the same TTL as the
server-side cache so the two layers cannot disagree. The default stays
no-store: most routes here are personalised, admin-only or auth-dependent.
/api/badges/leaderboard is deliberately left alone because it returns
per-viewer rank entries to signed-in callers.
Note this only takes effect once a cache rule exists for /api/* at the CDN, or
the explicit `cache: "no-store"` is dropped from the client fetches (24 files
do that today, including the /api/online poll). The headers alone are inert
until one of those happens.
- Take the homepage row counts from the storage engine estimate instead of
COUNT(*), which walks an index and gets slower as the tables grow. A missing
or zero estimate falls back to the exact count rather than ever showing a
wrong zero. The online count stays exact: it is an indexed read over a small
subset and a few seconds of drift reads as broken rather than approximate.
The counters move into one module because the homepage and the boot warm-up
populate the same cache keys, so two implementations would race to write
different values into the same entry.
3223 tests pass.
The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.
- Evict least-recently-used instead, and raise the default budget to 2000
(CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
lives on the entry, so one call site opting in protects every reader of that
key. A failed background refresh keeps serving the last good value instead of
falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
Redis key and publishes a signal, so a value written by one process is no
longer served stale by the others for the rest of its TTL. A failed Redis
delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
call, with a pub/sub signal to drop the local copy when it rotates. A Redis
outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
the old cap counted files and never removed anything while entries were
fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
runtime imaging cache.
Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.
3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
A fallback render drops the requested effect and is only a degraded
stand-in, so writing it to the 30 day disk cache kept serving the worse
image long after the local renderer recovered. Cache primary renders only
and let the next request pick up the real render.
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.
Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
Add an alerting/stats layer over the existing CrowdSec integration:
- New crowdsec-alerts.ts: cooldown-gated ops alerts (Redis NX lock, TTL from
HEALTH_ALERT_COOLDOWN_MIN) fanning out through the app's sendAlert service.
Raised for daily quota exhaustion, block bursts (5-min window past
CROWDSEC_ALERT_BLOCK_BURST), and signal-push failures.
- New crowdsec-stats.ts: daily counters (lookups/blocks/reports/report_fail)
in Redis with a 14-day reader for the admin panel.
- Shared 403/429 backoff: the pause marker now lives in Redis
(crowdsec:backoff-until) so every instance honours it, not just the process
that hit the limit.
- Atomic quota reservation: INCR-before-call with self-rollback on overshoot,
so concurrent instances can never slip calls past the daily ceiling.
- Admin anti-DDoS page gains a last-14-days activity table next to the quota bar.
- Bound the in-process verdict cache (FIFO eviction at 2000 entries) so a
flood of distinct bucket-tripping IPs cannot grow it without limit.
- Record block metadata (reputation, score, behaviors, category, TTL) in
antiddos:block:meta:{ip}, surfaced as the reason in the admin block list;
unban now also clears the metadata and report locks.
- Track daily CTI enrichment usage in Redis (crowdsec:usage:{date}); warn
once at 80% and pause lookups until tomorrow at CROWDSEC_CTI_DAILY_QUOTA
(default 10000, 0 = unlimited) so a via-spread DDoS cannot burn the plan.
- Add opt-in signal push to the CrowdSec community (CAPI watcher): stable
auto-generated 48-char machine_id/password pair persisted in Redis (or via
env), one-time registration, cached JWT login, optional Console enrollment,
and POST /v3/signals with a ban decision, deduped per IP. Never throws and
reports last status to the admin panel with a verify action.
- Admin page: quota usage bar, reporting status/verify channel, and CrowdSec
block reasons in the active-blocks list.
Align the active runtime with the pinned version across .nvmrc, package.json
engines and the Dockerfile base images, so scripts/check-node-toolchain.mjs
passes on the CI host running Node 26.10.0.
- new crowdsec-api lib: CTI lookup (GET /smoke/{ip}, freemium x-api-key), verdict parser with false-positive veto, 1h Redis + in-memory verdict cache, NX lock dedupe, 403/429 backoff; writes only the shared antiddos:block:{ip} key (value "crowdsec") and never touches Cloudflare
- gate fires it fire-and-forget for IPs that already tripped a rate bucket, so known-bad IPs are hard-blocked before the local maxViolations threshold
- runtime config: crowdsecAutoBlock toggle, score threshold (0-5, default 4), block TTL (default 24h); boot defaults CROWDSEC_AUTO_BLOCK_ENABLED / CROWDSEC_BLOCK_SCORE / CROWDSEC_BLOCK_TTL_SECONDS
- admin panel: CrowdSec stat card, verify-connection action, score/TTL settings, CrowdSec source badge in the blocked-IPs list
- credentials live in env only (CROWDSEC_API_KEY); block is enforced per-request via proxy on the resolved X-Forwarded-For / CF-Connecting-IP
- tests: crowdsec-api unit suite + ddos-guard integration suite (early-block, threshold, cache dedupe, backoff)
cloudflare-api unit tests drove the real Redis connection when REDIS_URL was set (CI), causing cross-test bleed. Mock @/lib/redis with an in-memory fake identical to the gate integration test.
- gate creates a zone IP Access Rule (block) for proxied offenders that hit the block threshold, deduped until the tiered block expires
- cloudflare-api lib: verified endpoints, create/delete/verify/list helpers, Redis-backed tracking + 30s TTL sweep (instrumentation worker + admin render)
- runtime toggle cloudflareAutoBlock in antiddos config; boot default CLOUDFLARE_AUTO_BLOCK_ENABLED
- admin panel: Cloudflare edge-blocks card with verify + remove-rule actions; unban also lifts the edge block
- credentials live in env only (CLOUDFLARE_API_TOKEN / CLOUDFLARE_ZONE_ID)
Bump the framework to the latest 16.3.6 patch release. Typecheck passes and
the homepage renders (HTTP 200) on the dev server with Next 16.3.6 under
Turbopack.
Split the heavy test suites out of the check job so coverage, MariaDB/Redis
integration and Playwright UI tests run concurrently on the host runner
(capacity raised to 4) instead of back-to-back (~2min wall-time saving).
Deploy and preflight now gate on all three test jobs.
Point PLAYWRIGHT_BROWSERS_PATH at the persistent /opt/ms-playwright dir on
the host runner so 'playwright install chromium' is an instant no-op after
the first run (was ~100s CDN download per job).
- Switch test-runner and renovate workflows from ubuntu-latest to
self-hosted now that a native host runner is running as a systemd service
- Replace remaining hardcoded color utilities in the homepage with theme
tokens and inline rgba styles to satisfy the no-hardcoded-colors contract
- Restore dual UserAvatarThumbnail usage on the homepage (hero avatar stack
plus community grid) to satisfy the public avatar presentation contract
- Create src/middleware.ts for per-request nonce-based CSP
- Integrate src/lib/csp.ts to build the CSP header dynamically
- Add src/middleware.test.ts to verify CSP header is set with nonce
- Biome lint and TypeScript checks pass
- Update txSelect and db.select mocks in draw-badge.test.ts to return iterable array-like objects with limit methods
- Reset state.price in beforeEach
- Fix test assertions for unsafe character stripping test
- All 3,066 tests now pass cleanly
The palette lives in plain CSS (:root/ThemeVars/admin remap), so Tailwind
never generated bg-primary, bg-destructive, text-foreground and similar
utilities. Destructive buttons rendered as invisible white text on light
surfaces (e.g. the Nitro cleanup delete button). Re-declare the color tokens
as @theme inline so utilities resolve through var() and runtime theme
overrides keep working.
- Add persistent disk cache for rendered avatars/badges (storage/imaging)
so repeats never touch the flaky local renderer and cached renders
survive upstream downtime
- Serve cache-first with stale-on-error; cut primary/fallback timeouts
from 10s/6s to 4s/4s so failing images cannot stall pages
- Avatar proxy now returns a graceful 200 silhouette instead of 502 when
no renderer can produce a figure, so no broken-image glyphs appear
- Badge endpoint becomes a caching proxy trying configured CDN, public
Habbo CDN and local /swf copy in order, and drops the fragile IP rate
limit that could blank badge streams
- Route all site badge images (profile, me, badges, apply pages) through
the cached proxy instead of hot-linking images.habbo.com
- Track referral attribution at registration via ?ref code with
same-IP and duplicate-pair guards
- Add daily login rewards with streak tracking, claim flow and
sendCurrency payout backed by RCON with DB fallback
- Add admin pages for referral settings and the daily reward schedule
- Add migration 0033 with tables, seed schedule, settings and ACL grants
- Add admin.referrals.* and admin.dailyrewards.* permission slugs
- Localize new copy in en, nl and it
Header and hero/stats counters each opened their own EventSource to the
online-count stream; a shared subscriber now opens a single socket and
multicasts to every mounted counter. The entrance count-up animation skips
its requestAnimationFrame loop when the user prefers reduced motion.
Keep every animation transform/opacity-only so frames never repaint:
- hero ring and loading glow pulse via opacity instead of background-position / box-shadow
- button shine sweeps with transform, not left
- floating glass chips drop animated backdrop-filter (it re-samples every frame)
- promote continuously animated layers (particles, halo, float) with will-change
- drop the negligible blur on moving clouds and remove the unused gradient-shift
Add background_effect (aurora/particles), background_overlay tint and
opacity to the theme manager, rendered site-wide by ThemeVars on every
public page. Polish the home and register pages (hero mascot, live stat
pulse, date pills, photo strip, CTA band, theme-aware register intro,
i18n for home/register section).
Extract FurniThumb, LayoutPreview, and GroupItemList into a dedicated
mall-helpers module alongside OrganizeImportsDialog. Preserves all
virtualization, drag-and-drop, and preview behavior while reducing
the main dialog component size.
Split the Add Item dialog form fields into its own module,
reducing the main table component size while preserving all
form fields, validation and handler logic.
Replace raw db.execute tuple casts with queryRows/rowsFrom/execResult/
affectedRows helpers from lib/db, drop redundant mysql2 casts on typed
query builders, and centralize per-test fakeForm into test/fake-form.
Update db mocks in tests so helpers resolve against mocked execute.
- password.ts: derive plain and salted digest detection from one DIGEST_SCHEMES
table instead of parallel hardcoded lists, so adding a family is one row.
- auth.ts: move 2FA challenge verification into twofactor-verification.ts and
the website login-log insert into website-login-log.ts, slimming the
NextAuth provider to orchestration only.
- deps: bump @formatjs/icu-messageformat-parser, @tanstack/react-query, jszip,
lucide-react, motion (patch/minor only). @types/react stay pinned per
pnpm-workspace.yaml; next-auth is already at the newest available (v5 beta).
Expand checkLogin to auto-detect and migrate every common retro CMS password
format to bcrypt on login:
- combined digests: md5(md5(pass)), md5(sha1(pass)), sha1(md5(pass)),
double sha1/sha256/sha512 and md5<->sha256/sha512 combinations
- salted digests of all families (md5/sha1/sha256/sha512) with embedded
salt using : $ @ _ separators, verifying both salt+pass and pass+salt
- plaintext fallback stays as the final catch-all
All formats verified on login and rewritten to bcrypt, so accounts work
whenever they come from any legacy CMS.
Legacy md5/argon2id hashes are now always upgraded to bcrypt on login, so
the CONVERT_PASSWORDS flag is no longer used. Drop it from env schema,
.env.example, the docker installer, and test mocks.
checkLogin now verifies and migrates all known password formats without
configuration: bcrypt, argon2id/argon2i/argon2d, unsalted md5/sha1/sha256/
sha512, double-md5 (UberCMS/Butterfly), salted md5 with embedded salt
(hash:salt, salt:hash, hash$salt), and a guarded plaintext fallback.
Every successful legacy login rewrites the stored hash to bcrypt, so the
CONVERT_PASSWORDS flag is no longer required (kept for deploy compatibility).
The deps update bumped react to 19.3.0 but left the lockfile resolving
@types/react to 19.3.0 while package.json and pnpm-workspace.yaml pin
19.2.18/19.2.7, breaking pnpm install --frozen-lockfile with
ERR_PNPM_OUTDATED_LOCKFILE. Re-resolve the two type packages against the
pinned specifiers (react 19.3.0 unchanged).
Theme Manager under /admin-next/hotel/theme-manager lets the owner save, apply, rename, delete, import, and export custom themes, plus set a custom site background by URL or upload. Themes are stored in WebsiteSetting/custom_themes JSON so they survive CMS updates.
The cleanup scan used to read every .nitro bundle in full and decompress
the large PNG texture just to confirm the file is structurally valid. On
directories with hundreds of thousands of bundles this took minutes, the
reverse proxy cut the request at its 30s timeoutable with an HTML 504, and
the panel then crashed with "Unexpected token '<'".
Validate bundles with a cheap header-only read (a few KB, no decompression)
that mirrors parseNitroBundle's byte layout; only files whose header looks
suspicious get the expensive full parse. Robust against downloads that
landed as an HTML error page, truncated or zero-filled files. The scan
drops from minutes to seconds on large nitro directories.
Also guard the panel against non-JSON (proxy error page / HTML) responses
so it reports a clear error message instead of a JSON parse failure.
Scan distinguishes fake, broken, and orphaned SWF/icon assets with age
metadata, deletes per asset kind, re-downloads broken nitro bundles from
configured sources, auto-cleans old fake leftovers, and exports a JSON
manifest. Adds rebuild and auto-clean API endpoints with audit coverage
and a housekeeping preview route under the hotel domain.
Verified: full vitest suite (2213 tests), typecheck, and biome all pass.
AvatarImage is used on server pages via UserAvatarThumbnail but lacked
'use client', so the onError handler on its <img> could not cross the
RSC boundary. /login rendered the error page, hanging the news e2e
journey until the 240s test timeout.
- Mark AvatarImage as a client component like ProfileImage
- Replace inline <img onError> on mod/users server pages with the
client AvatarImage component
- Restore build_attempted=1 in ci-preflight.sh so the exit trap
removes the temporary image tag
- Remove publish-container.test.ts and its harness (publication
workflow and script were removed in fff284aa)
- Update deploy-workflow-contract and docker-build-contract tests
to assert that publication has been removed
- Native HTML5 drag & drop between categories with drop-target highlight
- Virtualized item lists via @tanstack/react-virtual (fixed 34px rows)
- Auto-batching of large groups (500 items / 50 groups per run)
- Duplicate detection against the destination page with badge + summary
- Per-category layout preview grid
- Undo history for item moves (single + batch, tracks source groups)
- Days-range selector to load older imports (30/60/90/180)
- LocalStorage persistence of user settings (mode, destination, price, days)
- Added translation keys across all 25 locales
Make /admin/catalog a full-screen catalog studio that replaces the old
listing plus separate [id]/builder-club detail pages:
- Embed CatalogManagerWorkspace on /admin/catalog with a Normal/Builder
Club toggle, Catalog Sync status, packages (normal), Organize imports
and a diagnostics link to /admin/studio/maintenance.
- Manage BC items directly in the studio Items tab (new BcItemsEditor,
CRUD via existing bc actions; /api/admin/catalog/items now serves BC).
- Inline editor: add pageTextTeaser field for both catalogs and remove
the legacy full-editor links.
- Remove the 'Open full editor' context action from the tree.
- Move catalog-items-table (dir + barrel) and catalog-translate-tab out
of the app route into src/components/admin/catalog and update all
importers.
- Keep /admin/catalog/[id], builder-club/[id] and /admin/catalog/maintenance
as redirects into the new studio; consolidate maintenance panels into
/admin/studio/maintenance and point the nav item there.
- Delete the old listing/table/tabs/forms and the standalone bc-manager.
Catalog Studio:
- Cross-parent drag & drop now uses optimistic updates with rollback
on failure (no more full tree reload / visible delay)
- Subpage creation adds the node optimistically then refreshes parent
only (was full tree reload)
- Single page deletion refreshes only the affected parent (was full
tree reload)
- Root page creation replaces native prompt() with an inline input
in the root tab bar
- Escape key no longer closes the dialog when an input field is focused
- TreeNodeUpdate type now supports parentId and orderNum for
optimistic structural changes
Docker:
- docker-prune.sh default mode now aggressively cleans all unreferenced
build cache, images >1h old, and stopped containers >1h old
(was 72h/7d/24h which let cache grow past 80% on every push)
Add superRefine rule in src/env.ts ensuring that if one PayPal credential (PAYPAL_CLIENT_ID or PAYPAL_SECRET) is set in production, the other is also required, catching configuration drift at startup.
Introduce getCachedAdminCount to cache un-filtered table count(*) queries in Redis for admin lists (starting with UsersPage), avoiding heavy full table scans on every request while keeping exact counts for search/filtered queries.
Add rateLimit protection to /api/paypal/create, /api/paypal/capture, /api/tokens, /api/radio/shouts, and /api/articles/[slug]/comment to prevent abuse and spamming.
The 5-minute disk probe now reclaims storage automatically: from 85% it runs the gentle age-windowed Docker prune, from 90% it drops the age windows (docker-prune.sh --force: all unused build cache and unreferenced images, all stopped containers) so a mount can never silently max out. Alerts still fire at 85/90/95% and their hint now points at non-Docker growth when reclaiming is not enough. Force mode is reserved for the worker; deploys keep the gentle mode. Volumes are off-limits in every path.
Add a pure df parser (disk-usage.ts) with 85/90/95% threshold classification, a diskPressure() alert (Discord/email/alert_logs, severity escalates with fill), and a 5-minute host-side probe in jobs-worker.ts that raises one alert per crossing mount, cooldown-gated per mount+level. Real mounts only: overlay/tmpfs pseudo filesystems are ignored.
Add scripts/docker-prune.sh (build cache >72h capped at 4g, unreferenced images >7d, stopped containers >24h; never volumes), run it after every CI deploy and compose update, and schedule a nightly prune from the host-side jobs-worker. Tighten the deployment contract tests to assert the scoped-prune boundaries.
The interactive batch ignored the client's Translate option, so translated names never landed in the language files during a bulk run. Patch each successful item through patchLocalizedFurniDataEntries (mutex-guarded, best-effort) when translation is requested, surfacing failures as item warnings instead of failing the import.
Mirror interactive batch runs into the import-job store so interrupted imports (restart, time-out, disconnect) can be resumed from Import History. Items are checkpointed as they settle (coalesced, serialized saves) and the mirror starts 'running' so the boot-time worker marks it 'interrupted' instead of double-importing; done items are never re-imported. Add bounded backoff retry for transient download/connection failures before marking an item failed, and point the client's time-out/network toasts at Import History.
Highlight the active search term in names/classnames (grid + table), add zebra striping and a left accent bar on selected table rows, fade the results list when switching grid/table or on first load, and swap the broken-icon fallback for a cleaner placeholder.
Raise batch concurrency (furni 3->12, clone 10->12) with a Speed control next to Translate. Skip the SWF download when a .nitro bundle already exists on disk (color variants share the base nitro), and stop flagging that as a failed download.
- Render the table view through a virtualizer too, using a shared grid
template so the sticky header and rows keep perfect column alignment
- Keep semantic table/row/cell elements while virtualizing
- Cap the batch item-details list to the latest 60 rows (newest first)
- Coalesce per-item progress events server-side (120ms throttle) in both
the exact-import and clone SSE batch runners
- Add TTL-based caching for local index lookup, furnidata classnames,
catalog id set, nitro file presence and import stats
- Invalidate caches after furnace single/batch/clone imports
- Rewrite batch progress with elapsed time, rate and verification chips
- Virtualize the grid with @tanstack/react-virtual and replace the
Load more button with infinite scroll via an IntersectionObserver
The "Create" CTA used totalSelected (the sum of furni items across
approved groups) as its plural count, so with many imported items it
claimed to create thousands of pages. One approved group creates exactly
one page, so the label now counts approved groups instead.
- hero-ring: continuously shifting gradient hairline around the hero
frame (mask-drawn, reduced-motion safe).
- glass-chip: frosted floating pills under the CTAs that bob on a
staggered loop, live online counter keeps ticking for the Online chip.
- Taller hero for more presence on desktop.
The first cloud pass was to subtle: only five, up to 190s per crossing,
and they froze off-screen under prefers-reduced-motion. Rework:
- Eight clouds across the top 60% of the viewport with visible drift
(28s–66s loops, staggered by negative delays so they are always mid
scene on load).
- Higher opacity/steeper size contrast in light mode; dark mode dims
them slightly.
- Reduced motion now freezes a static, evenly-spread cloud field across
the width instead of pushing the clouds off-screen.
- New src/lib/site-icons.ts: base64 data URIs for the 13 tiny classic
icons actually referenced in markup (100–2000 bytes), replacing extra
requests with inline payloads. home.png, dynamic flags and currency
sets stay on the filesystem.
- Default favicon is now served server-side as a base64 SVG data URI
(memoized), while a DB-configured custom favicon still takes priority.
- Icons render through <Image unoptimized>, so data URIs pass through
untouched on all affected pages (home, login, register, settings,
navigation, auth top bar, client loading).
Add a fixed, decorative layer of soft clouds that slowly float across
the Habbo sky behind the content:
- Five clouds at staggered sizes, heights, opacities and loop timings
(60s–190s) so the drift feels organic.
- Dimmed further in dark mode; frozen by prefers-reduced-motion.
- Pure decoration: aria-hidden, pointer-events: none, no color
utilities in markup (cloud shapes live in globals.css).
- Slow Ken Burns drift on the hero artwork for cinematic depth.
- Gentle breathing pulse on the brand glows behind Frank and the hero.
- Silkier reveal easing (cubic-bezier .22/1/.36/1, 0.55s) site-wide.
- Eased, longer hover transitions on the hero CTAs.
- Smooth page scrolling, all guarded by prefers-reduced-motion.
Give the home page a clear, professional information hierarchy:
- Add eyebrow labels + headings for Features, Live stats and Community
sections using reusable, theme-aware .eyebrow / .section-title styles.
- Loosen the vertical rhythm (gap-8/10) so each block breathes.
- New messages resolve via the existing English fallback for all locales.
Professional tidy-up of the landing experience:
- SurfaceCard: unified rounded-xl radius for a crisper, consistent look.
- Hero: matches the new card radius and gains a dual-direction title
shadow so the headline stays readable over the header artwork.
- Login/register: drop the duplicated "no account / have an account"
paragraphs — the forms already ship an inline footer, so one clear CTA
cluster remains and the side column is cleaner.
Refine the public UI for a cleaner, more professional and scannable
landing experience without leaving the classic Habbo style:
- SurfaceCard: softer layered shadow, gradient accent hairline on the top
edge and a bolder header title across all public cards.
- Home: gradient hotel-name in the hero headline, shine effect on the
primary CTA, hairline on the top bar.
- Login/register: consistent avatar tiles with rounded corners, subtle
borders and a gentle hover lift; uniform username sizing.
- Add reusable theme-aware .card-hairline and .gradient-text utilities.
Replace the premium dark-gaming redesign of home, login and register with
the original classic landing (Habbo sky background, AuthTopBar, SurfaceCard
layout). Keeps the theme background visible again and adds a subtle
theme-aware brand halo behind the hero/Frank plus a soft primary glow on
card hover.
Also upgrades dependencies: next 16.3.5, vite 8.3.0 (typescript 7.0.2 was
already latest). Temporarily lowers pnpm minimumReleaseAge to 60 min so the
fresh 16.3.5 release can be installed; restore to 1440 once it is 24h old.
Replaced the classic Habbo landing style on the home, login and register pages with a modern premium dark-gaming look: always-dark hero canvas with brand glows, grid overlay and ambient orbs, glass panels, gradient text and floating art. Adds shared LandingTopBar, AuthShell and BrandFrank components plus reusable premium CSS utilities. Build, typecheck and lint pass.
resolveNitroFrame now matches spritesheet frame keys that carry a .png
suffix or namespaced naming, and isScale/scaleName preserve that suffix.
Broken source sprites (missing frames or references to icon artwork) are
skipped and reported instead of aborting the whole generation, and the
studio UI surfaces the skipped count.
Adds a second mode to the organize-imports dialog: instead of creating
one new page per approved group (which could produce dozens of tiny
pages), the user can pick an existing destination page and have every
approved item moved into it. No catalog page is created in this mode.
- organizeImportFurni: groups accept destinationPageId; when set, the
existing page is reused, new offers append after its current highest
order, moved offers keep their original name, and the real page
caption is used for logging and results
- OrganizeImportsDialog: mode toggle (create pages / move into page),
searchable destination picker via /api/admin/catalog/tree?search=,
name/icon/layout editors hidden in move mode, button shows a move
count, and the success toast reports moved/added instead of pages
- en + nl translations for the new mode, destination, and move keys
The organize-imports route defaulted to a 500-item limit, silently
hiding offers beyond the first batch of imported pages. Remove the
effective cap (limit now means 'all', guarded only by a 50k lint cap)
so every offer already sitting in the import tree is returned and
grouped.
The original GET route used a correlated NOT EXISTS / FIND_IN_SET
subquery over the entire catalog_items table for every recent import
audit entry, causing server timeouts when the audit log or catalog
grew large. The per-item host-page validation inside the create
action also issued one SELECT + one UPDATE per moved offer.
Changes:
- GET /api/admin/import/organize: replace the correlated subquery
with a bounded candidate list and a JS-side placed-set check, then
resolve all needed base items in a single indexed SELECT. This
bounds the query cost regardless of catalog or audit log size.
- organizeImportFurni action: validate mover ids in one SELECT, then
batch every move per group into a single UPDATE with a CASE
expression instead of one UPDATE per item.
- OrganizeImportsDialog: add a 45-second abort timeout on the fetch
and a distinct load-error state so the UI never silently hangs.
- Add 'loadError' translation key (en + nl).
Adds a Studio 'Organize imports' dialog that groups recently imported
furniture and furniture already sitting in the auto-created import
pages into suggested catalog categories. Each group is presented with
its suggested name, icon, and layout which can be overridden before
approval; approved groups are turned into real catalog pages in a
single atomic export run. Offers already inside the import subtree are
moved to the new pages; brand-new furniture gets a fresh offer.
- groupSuggestedCategories: generic pure helper reusing the same
per-item label heuristic that drives suggestCategoryName; items with
no label land in a 'Other Furni' remainder bucket
- GET /api/admin/import/organize: returns items from the imported
furniture tree (catalog_items JOIN items_base via page id set from
the imported-furniture root) union audit-logged but not-yet-placed
recent imports; marks alreadyPlaced vs new
- organizeImportFurni server action: validates import-page membership
before moving any offer, creates pages + inserts/moves items in one
withCatalogExport snapshot, logs activity
- OrganizeImportsDialog: full-featured Studio dialog with price panel
(applies to new offers only), parent select, per-group approval,
editable name/icon/layout with suggestion reset chips, source tags
- import-pages.ts server helper: locates the imported-furniture tree
- Studio nav: 'Organize imports' button with FolderTree icon,
gated on CATALOG_EDIT, wired next to the catalog manager
- en + nl translations for organizeImports.* keys with ICU plurals
- groupSuggestedCategories unit tests (deterministic grouping,
remainder handling, size+alpha ordering, label consistency)
The catalog manager (with auto-category wizard) now lives inside the Studio
navigation for teams that manage furniture inline. The button is gated on
CATALOG_EDIT; only users with that permission see the launcher.
- Split server/client Studio layout to derive permissions server-side
- Add optional triggerLabel prop to CatalogManagerDialog for custom labels
- Wire the Catalog manager button in the Studio nav right cluster
- Keep existing /admin/catalog entry points unchanged
- Add AutoCategoryDialog: pick furni (or empty page), suggest caption/icon/layout, live preview
- Add createAutoCategory server action (page + offers in one export) with CatalogKind
- Extend furni search API with interactionType
- Share ShopTile and refactor inline-editor/items-shop-preview to use it
- Add suggestion heuristics (suggestCategoryName/dIcon/layout) with tests
- Add autoCategory translations (en/nl)
- Update all DragonflyDB references to Valkey in README and docker-compose.yml
- Update install instructions to use Valkey package repository and .deb download
- Update configuration paths from /etc/dragonfly/ to /etc/valkey/
- Update version requirement to Valkey 8.x+ (successor to Redis OSS)
The Arcturus errors "page hierarchy contains a cycle page 354 and 357" and
"sibling order 1 is used more than once (111 problems)" come from
catalog_pages, not catalog_items: pages 354/357 point at themselves, and
many parents have child pages sharing the same order_num. Extend the
emulator catalog scan + fix to detect both: pages whose parent chain loops
back get detached (parent_id = 0 on the highest cycle member) and every
affected parent's children are renumbered sequentially, preserving their
current relative order.
Add scripts/diag-emulator.ts to inspect the live catalog state.
Block invalid catalog_items writes at the API level (points currency
allowlist, non-negative prices, positive amount, limited stack >= sold
count, unique sibling order numbers) and auto-assign unique order numbers
on bulk create. Add a catalog-maintenance scan + transactional repair that
fixes pre-existing rows: resets unsupported points_type, clamps negative
costs, sets amount to 1, raises limited_stack, renumbers duplicate orders
and deletes offers with missing page/item references. Surface the issue
count and a fix button in the admin maintenance panel.
Also: add enabled/retired flag to clone sources, classify poster and
currency furniture in item-kind, and remove the obsolete update-Nitrov3.sh.
translateCatalogItems previously only patched the master FurnitureData.json
via patchFurniEntryNames but never updated the per-language files
(FurnitureData_nl.json, etc.). Custom/imported furniture translated through
the catalog Translate tab was therefore invisible in localized builds.
After patching the master file, the action now also calls
patchLocalizedFurniDataEntries so LibreTranslate translates the English
names into all 13 supported languages.
Exclude generated drizzle-kit snapshot artifacts from formatting checks (drizzle/drafts/meta), which made biome scan a 360KB generated JSON for 23s. Fix the pre-existing lint errors in error-monitor, article-form and the admin-search-permissions mock so pnpm biome:lint is green in CI.
Run vitest without coverage by default (pnpm test) and add pnpm test:coverage which enforces the coverage thresholds. CI keeps using the coverage run so thresholds are still enforced on every push.
Drop the unused knip dead-code check and the husky+lint-staged pre-commit hook pipeline. All checks remain covered by the CI workflow (lint, typecheck, i18n, tests).
Drop the leftover Playwright browser install and e2e smoke test from the deployment script, and update the deployment contract tests to cover the verify-deployed-release smoke check instead.
- Move the pnpm store, apk and .next caches into --mount=type=cache so
dependencies are shared across builds instead of duplicated in fresh
image layers (was the source of unbounded disk growth).
- Replace the deprecated --keep-storage prune flag in ci-deploy.sh with
the working --max-used-space=4g (buildx v0.37 renamed the flag). The
deprecated flag silently did nothing, so the BuildKit cache kept
growing unbounded (was 15.86GB); it is now capped at 4GB after every
deploy.
Runs the full catalog audit (runCatalogAudit) against the live hotel DB with
no repair/sql options — a pure read pass — and cross-checks the emitted
summary against independent DB + asset-dir measurements: row totals, duplicate
classname groups, orphaned catalog references and missing catalog / nitro /
icon counts must match exactly, with zero error events and a final
'batch_complete'.
Guarded by RUN_CATALOG_AUDIT_LIVE=1 so CI never runs it; loads the real .env
because vitest fakes DATABASE_URL.
generate-drizzle-schema.mjs imported 'dotenv/config' but dotenv is not a
dependency, failing knip and crashing 'pnpm db:schema:generate'. Load .env
with Node's built-in process.loadEnvFile like the other scripts do
(scripts/load-env.ts), never overriding vars already set in the shell.
- Create the runtime write targets (/app/storage, /app/public/nitro-assets,
/app/public/swf, /var/www/Gamedata) owned by UID/GID 33 in the runner image
so running without the bound volumes no longer hits ENOENT.
- Bake a HEALTHCHECK into the image so `docker run` (ci-deploy.sh) also reports
Docker-level health; compose can still override it with its own probe.
- Add the dockerfile:1 syntax pragma and ignore non-pnpm lockfiles so a stray
package-lock.json/yarn.lock can never taint the build context.
next.config.ts falls back to `git rev-parse HEAD`, but the build context has
no .git (excluded by .dockerignore), so every CI build printed fatal git
errors and stamped the release as "unknown". Pass the deploy commit sha as a
NEXT_DEPLOYMENT_ID build-arg so git is never invoked and the actual commit
reaches NEXT_PUBLIC_CMS_RELEASE and deploymentId.
The committed lockfile records overrides from pnpm-workspace.yaml. With the
narrower manifest COPY, pnpm install --frozen-lockfile ran without the
workspace file and failed with ERR_PNPM_LOCKFILE_CONFIG_MISMATCH in CI.
Copy package.json, pnpm-lock, pnpm-workspace and .npmrc together so the
override config present in the lockfile is also supplied at install time.
- Use floating node:alpine that tracks the latest supported LTS; pnpm
bootstrap follows package.json's packageManager pin.
- Drop corepack (removed from node:26), install pnpm via npm global.
- Add pnpm fetch + offline install for stable dependency-layer caching.
- Run as non-root nextjs (UID/GID 33 = host www-data) with tini as PID 1
for correct signal handling.
- Open node engines to >=20.9.0 so patches/minors float automatically.
- Add docker-preflight.sh (per-VPS checks incl. --fix) and gate docker-update.sh
so Node major upgrades require explicit review while patches deploy silently.
The site-settings loader kept an in-process map forever after a Redis miss
and promoted DEFAULTS (no logo/theme) to Redis on any DB error, so a build
that started before the DB was reachable stuck the site on the preset logo
and default theme until a manual reload or full restart.
- Redis miss now reloads from the database instead of the stale in-process map
- a DB failure returns defaults only as an in-process last resort and never
writes them to Redis, so the shared cache can't be poisoned by a transient
error at startup
- regression tests: DB re-read on Redis miss after cache expiry, defaults never
promoted to Redis, recovery from transient DB failure
- fix(studio): stop markBatchDone infinite recursion so batches complete
- refactor(api): merge api-response into api and drop the duplicate module
- refactor(media): extract shared media loader and URL validator
- refactor(ui): extract shared LoadingSpinner for site and admin groups
- refactor(dates): consolidate raw date formatting into formatDate util
- refactor(logs): share a single generic log-list loader across tables
- refactor(theme): merge both ColorField components and reuse contrast helpers
- refactor(import): extract shared ImportErrorBanner and SearchInput
- chore(): remove dead theme-editor-tabs after inlining tab components
- New buildDutchLanguage() function translates only 'nl' language
- Proper error handling with specific messages for network errors
- Safe template literal usage, guarded .find() for Dutch language
- Consistent with existing buildAllLanguages pattern
Remove unused admin-helpers.test.ts (dead code)
- Add retries for furnidata fetch (3 attempts, 1s delay)
- Better error messages for user (distinguish 502/400/other)
- Client-side: show detailed error from SSE stream when available
- Remove unused admin-helpers.test.ts
- Add e2e job (needs deploy, main/master only) that installs the
Playwright browser and smoke-tests the live container on :3002
- Keep deploy-job contract slice from bleeding into the e2e job
- Cover the e2e job in the CI workflow contract test
- Ignore Playwright output dirs (test-results, playwright-report,
blob-report)
Zonder dit loopt nieuwe code tegen een oud schema aan zodra een push
migraties bevat. Idempotent (toegepaste migraties worden overgeslagen),
draait op de host met de productie-.env, vóór de container-replace.
docker run --env-file behoudt letterlijke quotes (bewezen test),
waardoor DATABASE_URL ongeldig was en de container crashte. Nu wordt
.env gesourced en elke sleutel met -e doorgegeven: exact dezelfde
waarden als bij de build. Contract-test verbiedt --env-file.
- Deploy kopieert de productie-.env van de host in de build-context:
Next.js bakt NEXT_PUBLIC_* in en valideert DATABASE_URL/HOTEL_NAME
(SKIP_ENV_VALIDATION is verboden voor productie, zie src/env.ts).
- Dockerfile builder installeert git (next.config.ts deploymentId).
- Deploy-container krijgt --env-file + dezelfde volumes als compose,
stopt ook de oude compose-container (poort 3002) en rolt terug via
compose bij een falende health check.
De host heeft Docker iptables uitgeschakeld, dus build-containers op
bridge hebben geen outbound internet; 'npm install -g pnpm' in de
builder-stage hing daardoor. Met --network=host krijgt de build wel
registry-toegang. Contract-test vergrendelt de flag.
pnpm 11 has no --jobs option for install; the stray positional '1'
flipped the command into 'add' mode, which then rejected both
--frozen-lockfile and --jobs ('Unknown options', 'pnpm help add').
Reproduced locally, fixed, verified install succeeds. Add contract
regression test.
Root cause: de job-container (docker mode) heeft GEEN outbound
internet naar GitHub, waardoor actions/checkout@v4 faalde met
'Unable to clone ... i/o timeout'.
Oplossing: zowel check als deploy draaien nu op self-hosted (host)
waar Node 26.8.1 + pnpm 11.25.0 geïnstalleerd zijn en internet
beschikbaar is. Dit is de enige betrouwbare setup in deze omgeving.
Runner is tevens hernoemd naar 'Epic runner'.
- Runner hernoemd naar 'Epic runner'
- Env variabelen als global env (niet per step)
- fetch-depth: 1 voor snellere checkout
- Knip verwijderd (traag, niet kritiek)
- Docker BuildKit caching
- Health check opgeschoond
- Show arrow (›) appears when toolbar is hidden, regardless of pos state
- Null-safe pos top/left (?. ?? 8) to prevent TS errors
- Arrow positioned fixed top-right instead of depending on balk-positie
- Redesign in-game client toolbar with refined glass styling and buttons
- Live online count + emulator status streamed over SSE multicast
- Who's-online tooltip throttled to reduce repeated requests
- Fix show button so it always returns the hidden toolbar
- Use theme CSS variables instead of hardcoded colors
- Add toolbar translations across all 25 locales
AnimatedCounter restarted from 0 on every online-count cache refresh (~10s)
because the effect depended on value. Rewrote it as a single requestAnimationFrame
animation with ease-out; changes to value after the animation only update the
display statically.
NITRO_REF="$(git ls-remote https://github.com/duckietm/Nitro-V3.git main 2>/dev/null | awk '{print $1}')"
RENDER_REF="$(git ls-remote https://github.com/duckietm/Nitro_Render_V3.git main 2>/dev/null | awk '{print $1}')"
EMU_REF="$(git ls-remote https://github.com/duckietm/Polaris-Emulator.git main 2>/dev/null | awk '{print $1}')"
{
echo "# EpicNext-CMS ${VERSION}"
echo ""
echo "> Modern, high-performance CMS for Habbo hotel emulators — built on Next.js 16, React 19 and Drizzle ORM. Integrates with Polaris / Arcturus Morningstar databases."
echo ""
echo "## Menu"
echo "- [What is EpicNext-CMS?](#what-is-epicnext-cms)"
echo "EpicNext-CMS is a full public-facing hotel website plus an administrative panel. It features NextAuth authentication (bcrypt with MD5 upgrade), real-time RCON communication with the emulator, Server-Sent Events for live radio, smooth page transitions and extensive extensibility. Full documentation: https://gitlab.epicnabbo.nl/remco/EpicNext-Cms/src/branch/main/README.md"
echo ""
echo '<a id="system-requirements"></a>'
echo "## System Requirements"
echo ""
echo "What you need to install before running the CMS:"
echo ""
echo "| Component | Version | Notes |"
echo "| --------- | ------- | ----- |"
echo "| Node.js | 26.7.0 | Current release pinned in .nvmrc |"
echo "The CMS shares the emulator database. Import the Polaris/Arcturus database first, then create the CMS schema:"
echo '```sql'
echo "CREATE DATABASE IF NOT EXISTS epicnext_cms CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;"
echo '```'
echo ""
echo "### 3. Configure Environment"
echo '```bash'
echo "cp .env.example .env"
echo '```'
echo ""
echo "Edit .env with at minimum: DATABASE_URL, AUTH_SECRET, HOTEL_NAME and APP_URL. See .env.example for RCON, email, Redis, OAuth and PayPal options."
echo ""
echo "### 4. Run CMS Migrations"
echo '```bash'
echo "pnpm db:migrate"
echo '```'
echo ""
echo "Creates all CMS-owned tables (website_*, radio_*, acl_*, admin_audit_log). Emulator tables are never touched. Check status with pnpm db:migrate:status. Runtime types come from the committed Drizzle schema (src/db/schema.ts) via '@/lib/db'."
echo ""
echo "### 5. Polaris Emulator"
echo ""
echo "Clone and build the emulator (requires Java 17+ and Maven 3.9+):"
echo "Place the built Habbo-*-jar-with-dependencies.jar next to **config.ini** (see [setup/emulator/config.ini](https://gitlab.epicnabbo.nl/remco/EpicNext-Cms/src/branch/main/setup/emulator/config.ini)), then create a systemd unit from [setup/emulator/emulator.service](https://gitlab.epicnabbo.nl/remco/EpicNext-Cms/src/branch/main/setup/emulator/emulator.service) with the [emulator](https://gitlab.epicnabbo.nl/remco/EpicNext-Cms/src/branch/main/setup/emulator/emulator) launcher so it starts on boot. The bundled update-Nitrov3.sh in this repo automates cloning, building and updating the emulator and Nitro — run it any time to pull the latest commits and rebuild:"
echo '```bash'
echo "./update-Nitrov3.sh"
echo '```'
echo ""
echo "### 6. Nitro V3 & Renderer"
echo ""
echo "Clone both Nitro repos and build the client:"
echo "Copy the reference configs from [setup/nitro/](https://gitlab.epicnabbo.nl/remco/EpicNext-Cms/src/branch/main/setup/nitro) into /var/www/Nitro-V3/public/configuration, keep them as *.json, and replace **MY_DOMAIN** with your domain, API URL and gamedata paths (see the Full setup guide, NitroV3_And_Emulator.md)."
echo ""
echo "### 7. Catalogus (catalog & gamedata)"
echo ""
echo "Catalogus holds the daily-updated catalog/gamedata. Clone the Beta-3 branch alongside the other components:"
The CMS editor at `/admin/translations/cms` now uses the same bundled catalogs as the request-time translator. Runtime changes are stored separately in `storage/cms-translations/<locale>.json`, inside the existing persistent `/app/storage` mount. No source-file write, environment-variable change or rebuild is required to apply an edit.
Only overrides are saved. Unchanged messages continue to receive updates from Git. Saves use a file lock, an atomic replacement and a revision check; a stale editor cannot overwrite another operator's changes. ICU syntax, argument names, rich-text tags and allowed keys are checked server-side. Invalid or obsolete overrides are excluded when reading a new release. Storage read errors are reported to the CMS error monitor and public pages fall back to bundled text.
The page starts in the operator's language. It shows all English reference keys, including missing translations, and supports search by key, translated text or English source. Filters separate missing, identical and modified text. Drafts survive language switches and failed saves. Users without `SETTINGS_EDIT` can review and export but cannot save.
## Audit outcome, 6 September 2026
- Scanned 25 JSON catalogs; 22 languages are selectable. The small Arabic, Finnish and Japanese catalogs are legacy files and remain outside the supported locale list.
- Repaired 66 malformed ICU messages across the 22 active catalogs, including HTML fragments and unescaped JSON examples.
- Added 199 missing English reference keys used by page components, with Italian translations, and fixed the incorrect navigation namespace in the admin error page.
- Completed the 31 previously missing Italian reference keys.
- Repaired missing `count` and `preset` variables in other locales.
- Final checks: zero malformed messages, zero argument/tag mismatches and zero missing references among the statically resolved translation calls.
- English contains 3,423 reference keys. Italian covers all of them; 629 values match English. Dutch is missing 524 reference keys and has 618 identical values. Matching English can be intentional for names and technical labels; this is not proof of translation quality.
- Found 1,577 literal JSX text candidates outside translation calls. These include labels, technical strings and names; they are an editorial inventory, not 1,577 confirmed bugs. The largest concentrations are the catalog item table (118), Studio main component (79), import audit (53), sound management (50) and permission editor (42).
## Repeatable checks
-`pnpm i18n:check`: fails on malformed messages, incompatible variables/tags, empty messages, source parsing errors or missing statically referenced keys. Runs in Gitea CI.
-`pnpm i18n:audit`: prints coverage and findings.
-`node scripts/audit-cms-translations.mjs --json`: full machine-readable inventory, including file and line references for literal JSX candidates.
Static analysis resolves literal translator namespaces and literal message keys. Dynamic key construction, prose embedded in arbitrary JavaScript strings, and the linguistic accuracy of all 22 translations still require targeted review. English fallback remains explicit; copying English into other catalogs would hide untranslated entries and is intentionally avoided.
## Validation
Automated tests exercise message syntax, actual translator output, persistent overrides, invalid-message rejection, revision conflicts and reset behavior. A browser fixture mounts the real editor and verifies missing-key editing, validation, failed-save preservation, language switching, successful saves, read-only access and mobile layout. Server actions are simulated in that fixture; authenticated production editing requires a staff session.
Data: 6 settembre 2026. Base verificata: main, commit 9b0ea2fb. Documento di proposta, non implementazione approvata. Il sito /admin/catalog reindirizza al login senza sessione staff: nessuna prova delle mutazioni sul database di produzione. Le criticità indicate sono percorsi verificati nel codice; gli effetti concorrenti richiedono riproduzione controllata.
## Obiettivo e perimetro
Un catalogo HK con un solo ambiente di lavoro, regole coerenti e operazioni recuperabili. Conservare stile HK, icone reali, salvataggio diretto, catalogo normale e Builder Club. Nessun ritorno dei Preferiti.
Inclusi: albero, categorie, offerte, prezzi, bundle, disponibilità, proprietà condivise dei furni, traduzioni, ricerca, anteprima, operazioni massive, manutenzione, collegamenti a Catalog Studio e sincronizzazione Git/hotel.
Catalog Studio conserva la responsabilità di importare, convertire e riparare .nitro, icone e furnidata. Il refactor collega questi strumenti alla selezione del catalogo; non comporta riscrivere il convertitore, cambiare protocollo dell'emulatore o sostituire il sistema Git/Gitea già esistente.
## Inventario verificato
Sono già presenti virtualizzazione dell'albero, trascinamento, multiselezione, griglia/tabella, editor dei prezzi, anteprima negozio, import massivo, traduzioni, manutenzione e coda di export Git. Vanno riutilizzati.
| File | Righe attuali | Responsabilità da separare |
Le dimensioni aiutano a trovare i punti di intervento: l'obiettivo non è un limite arbitrario di righe, ma responsabilità verificabili e riutilizzabili.
## Problemi e interventi
| Priorità | Evidenza | Intervento |
|---|---|---|
| P0 | sortable-tree.tsx invia il riordino dei fratelli con Promise.allSettled, una action per riga; catalog.ts aggiorna RCON per ciascuna | Un comando batch con lista completa, validazione, transazione e un solo evento di aggiornamento |
| P0 | catalog-items.ts riordina le offerte con update sequenziali fuori transazione | Stesso contratto atomico per l'ordine delle offerte |
| P0 | updateCatalogPage accetta parentId direttamente; il controllo cicli è separato in movePage | Validazione comune per creazione, form, spostamento e API; controllare destinazioni inesistenti e concorrenza |
| P0 | deletePage normale sposta figli, elimina offerte e pagina separatamente | Transazione, analisi dell'impatto e snapshot ripristinabile |
| P0 | updateCatalogItem modifica pagina, offerta e items_base con scritture separate | Transazione e verifica che il furno appartenga all'offerta; scope distinto per proprietà condivise |
| P1 | cascadeDelete non mantiene un insieme di nodi visitati; il calcolo profondità BC è ricorsivo senza guardia ai cicli | Lettura tollerante di dati incoerenti, diagnostica e arresto sicuro delle traversate |
| P1 | handleEditTab modifica lo stato prima della conferma; chiusura X bypassa la protezione | Un unico controllo delle modifiche per cambio pagina, offerta, scheda, uscita e navigazione |
| P1 | loadPage/loadItemsData non annullano o identificano la richiesta precedente | AbortController e identità della selezione; solo la risposta corrente può aggiornare l'editor |
| P1 | Il salvataggio ignora il booleano restituito da RCON; Git opera in coda | Distinguere DB salvato, invio hotel riuscito/fallito e stato Git; retry senza risalvare i dati |
| P1 | loadCatalogItemsData usa Number(value) || fallback per order_number, offer_id e amount | Definire semantica di zero/null per campo e testare il round trip prima di cambiare i fallback |
| P2 | Tutte le offerte e metadati sono caricati insieme; filtro con CAST(page_id AS CHAR) | Misurare query/payload, separare elenco e dettagli, paginazione e adapter compatibile INT/VARCHAR |
| P2 | Editor normale/BC e form condividono solo parte delle regole; testi anche letterali | Contratti comuni, differenze BC esplicite, traduzioni e permessi coerenti |
1.**Pulizia dei file mantenendo tutti gli editor:** rischio iniziale basso, ma conserva duplicazioni e differenze operative. Utile solo come passaggio iniziale.
2.**Refactor progressivo con un editor principale — consigliato:** servizi comuni prima, poi promozione del Visual Manager a pagina. Permette piccoli rilasci e confronti tra vecchio e nuovo percorso.
3.**Riscrittura completa:** libertà maggiore, ma più rischio di perdere casi speciali, compatibilità DB e funzioni già presenti. Non giustificata dall'inventario attuale.
## Architettura proposta
Modulo src/features/catalog con confini chiari:
- domain/: tipi Page, Offer, FurnitureReference, CatalogKind; validazione gerarchie, prezzi, bundle e disponibilità; nessuna dipendenza React/DB.
Le route e le action attuali rimangono inizialmente adapter sottili. Un unico risultato di operazione include ID operazione, revisione, elementi modificati, eventuali errori di campo e stato sincronizzazione. Non introdurre nuove librerie prima di verificare i limiti degli strumenti già installati.
Flusso di scrittura: permesso → validazione → verifica revisione → transazione DB con audit → risposta di salvataggio → aggiornamento hotel/export Git. La durabilità del passaggio DB→coda va garantita con un evento persistito nella transazione o meccanismo equivalente verificato. Un fallimento Git/RCON non deve far ripetere una creazione già committata. Riutilizzare il worker e la coda esistenti, aggiungendo idempotenza dove manca.
## UX proposta
Pagina /admin/catalog con barra: Normale/BC, ricerca, nuova categoria, aggiungi furni, stato operazioni. Sotto: categorie a sinistra, offerte al centro, dettagli a destra. Il pannello dettagli si richiude; su schermi piccoli diventa una vista dedicata. Un solo scorrimento per ciascuna area, azioni di salvataggio sempre raggiungibili.
La URL conserva catalogo, categoria, offerta, vista e ricerca; i campi non salvati restano nello stato locale. Indietro/avanti e ricaricamento devono riaprire il contesto corretto. I vecchi URL dei dettagli continuano a funzionare.
Tre oggetti riconoscibili:
- Categoria: percorso, titolo, icona, layout, visibilità e requisiti.
- Offerta: prezzo, valuta, quantità, componenti bundle, disponibilità e ordine.
- Furno condiviso: classname, sprite, dimensioni e interazioni; mostrare quante offerte lo referenziano prima di una modifica globale.
Idee operative:
- Ricerca trasversale per nome, classname, ID pagina/offerta/furno e sprite ID, con percorso nei risultati.
- Selettore visuale di categoria e layout; proprietà tecniche nelle Avanzate.
- Prezzi con icone reali delle valute; mostrare il prima/dopo delle operazioni massive, arrotondamenti ed elementi esclusi.
- Multiselezione con riepilogo di spostamento/eliminazione; dopo un errore mantenere selezionati i falliti.
- Anteprima del negozio già esistente integrata nel contesto; non presentarla come prova completa del comportamento del client hotel.
- Diagnostica su richiesta: offerta, SQL, furnidata, Nitro e icona separati. Collegamento a Catalog Studio sul furno esatto; nessuna scansione pesante a ogni apertura.
- Storico di chi/cosa/quando con differenze e ripristino. Il ripristino controlla revisioni successive: non sovrascrive in silenzio modifiche di altri operatori e non annulla acquisti già avvenuti.
- Riepilogo visibile: salvato, invio hotel, Git. Gli errori hanno riferimento al monitor CMS.
- Stati vuoti, errori, caricamento e sola lettura distinti; traduzioni complete e uso da tastiera.
## Programma di lavoro e criteri di uscita
| Lotto | Consegna | Criterio per proseguire |
|---|---|---|
| 1. Baseline | Matrice funzioni/route/permessi normale e BC; fixture con bundle, LTD, offerte speciali, zeri/null, alberi incoerenti; misure query e rete | Tutti i flussi esistenti hanno una destinazione nel piano, senza omissioni |
| 2. Integrità | Validatori, transazioni di riordino/spostamento/eliminazione, gerarchie sicure, revisioni | Un fallimento intermedio non lascia dati parziali; due operatori non si sovrascrivono |
| 3. Servizi condivisi | Query/command/repository e risultato comune; vecchie route come adapter | Vecchie UI superano le stesse prove con il nuovo backend |
| 4. Stato editor | Unica gestione delle modifiche, richieste annullabili, risposta coerente con selezione | Annullare l'uscita conserva tutto; cambi rapidi mostrano sempre l'ultima selezione |
| 5. Pagina unificata | Visual Manager nella pagina, griglia/tabella condivise, URL, layout adattivo | Parità normale/BC e vecchi link conservati; niente perdita di scroll o azioni nascoste |
| 6. Operazioni avanzate | Ricerca, editor bundle, prezzi massivi con differenze, storico e diagnosi contestuale | Gli effetti sono spiegati prima dell'applicazione; retry applica solo ciò che manca |
| 7. Sincronizzazione | Stato DB/hotel/Git, operazioni persistenti, retry/idempotenza | Guasto dopo commit e riavvio worker non duplicano né perdono l'operazione |
| 8. Prestazioni e rimozione duplicati | Paginazione, caricamento progressivo, accessibilità, eliminazione vecchi componenti | Confronto misurato e prove finali; nessuna route o funzione rimasta senza equivalente |
I lotti 2 e 7 condividono il contratto delle operazioni: progettare subito evento persistente e idempotenza, anche se la UI di stato arriva dopo. Nessuna stima in giorni finché non sono note dimensioni reali del catalogo, varianti DB e casi speciali attivi. Ogni lotto può richiedere più PR piccole; niente sostituzione monolitica.
## Verifica e rilascio
Test unitari delle regole; integrazione su MariaDB per rollback, concorrenza e varianti INT/VARCHAR; browser con permessi lettura/modifica, desktop e schermo ridotto. Simulare doppio submit, timeout, risposta fuori ordine, fallimento RCON, Git non raggiungibile, riavvio dopo commit. Conservare test e componenti esistenti finché la parità non è dimostrata.
Registrare baseline e risultati per categorie grandi/piccole: richieste per riordino, tempo DB, payload, tempo fino a editor utilizzabile, risposte fallite. Non promettere percentuali senza dati.
Rilascio progressivo con selezione reversibile del nuovo editor. Il ritorno alla UI precedente deve usare gli stessi servizi corretti. Migrazioni additive e compatibili; il rollback dell'app non deve richiedere la cancellazione di dati. Eliminare le vecchie UI solo dopo parità verificata. Confermare CI, deploy, health e prove staff prima di dichiarare risolto il flusso live.
## Primo passo consigliato
Lotti 1 e 2: inventario di compatibilità e correzione delle operazioni a rischio, mantenendo inizialmente l'aspetto corrente. Poi estrarre i servizi e unificare l'editor. È la sequenza che permette di migliorare UX senza portare avanti gli stessi difetti dentro una nuova schermata.
| Release | Deployment checks HTTP health, release identity and Chromium pages before marking the running image verified. Registry publication reuses that verified digest on the shared runner; independent hosts build and verify their own image. |
| Shared HK | Dialogs scroll within the available viewport. Table column preferences persist per page in the current browser session; filters already remain in the URL. No favorites added. |
| Catalog Studio | Dedicated detail drawer and five separate completeness states: SQL, offers, furnidata, icon and Nitro. Sprite/type conflicts are consistently excluded from import and linked to audit. Missing source files are never represented as available. |
| Operations | Command center includes permission-filtered error groups, personal import failures, open support tickets and news drafts. A failed source is shown separately from an empty source. |
| Error center | Retained occurrence counts, first/last times, identified users, release counts, self-assignment and recognized local links. Assignment requires edit permission and is audited. |
| News | Private server-backed autosave, recover/discard/retry, prior saved revisions restored as a draft, and concurrent-edit detection. New status/search filters and pagination make older drafts reachable. |
| Jobs | Cancellation finishes the current item and stops pending items. Completed work stays completed. Uncertain interrupted mutations are excluded from retry. Live lease checks stop known stale worker writes. |
| Installation | `/admin/devops/installation` shows release, DB latency, Redis, emulator, storage permissions, migration history and worker heartbeat. Renderer defaults are recognized. Registry access is explicitly unverified from the web process. |
| Public dashboard | Current/next published event, clearer unread-message action, useful empty/error states and mobile layout refinements. |
| Audit | Exact actor/action and UTC date filters, readable recorded before/after values and permission-protected CSV of the filtered page, capped at 100 rows. |
| Performance | Active import polling remains 5 seconds; idle polling is 30 seconds and pauses in hidden tabs. History returns at most 30 owned jobs and initially renders 50 items per job. Health probes are deduplicated within a render; diagnostics report observed probe duration. |
| Text | New messages are translated in English, Italian and Dutch. Other locales have explicit English fallback strings. Existing translation debt is not reported as resolved. |
## Deployment requirements
- Apply migration `0026_article_editor_recovery.sql` through `pnpm db:migrate` before enabling the new news editor. It adds private drafts and revision tables without modifying existing articles.
- Keep `storage` persistent and writable. Error assignments and import cancellation markers use the existing shared storage.
- Run the existing `pnpm jobs:worker` process with the installation configuration. It now publishes a heartbeat to Redis every minute. A web process alone does not establish that scheduled-news jobs are running.
- The deploy runner installs Chromium before cutover. Browser checks visit only public login/news/staff pages; they do not create production content or authenticate staff. The host must satisfy Chromium system-library requirements.
- Use the existing Gitea registry secrets. The CMS does not read or display those credentials.
## Boundaries
- Operations is a bounded operational summary: errors use at most 1,000 retained events and imports use the latest 30 owned jobs. Empty checked records do not prove that all historical work is resolved.
- Import history bounds payloads and concurrent file reads. Directory metadata scanning and worker enumeration still scale with stored history.
- Redis lease checks and file saves are separate operations. They reduce stale writes but do not provide atomic fencing across Redis, SQL and filesystem operations. An import interrupted after SQL may require local-data inspection.
- Revision history displays the latest 20 saved versions; revisions are retained in the database. New-article recovery has one private slot per staff account.
- The database connection probe is a measurement, not a performance benchmark. No throughput or latency improvement is claimed without production measurements.
- A rollback restores an application image; it does not reverse database migrations. The new tables are additive.
## Verification
Local verification on 2026-09-09:
- Production build succeeded with fixture configuration and an intentionally unavailable database. Build-time fallback logs are expected in this check.
- Full Vitest run: 254 files passed, 4 skipped; 1,466 tests passed, 6 skipped. Coverage thresholds passed (19.89% lines); this does not imply exhaustive coverage.
- Global Biome rules/import checks passed across 1,328 files; modified source/locales were formatted separately to avoid unrelated Windows line-ending changes.
- Translation audit: no invalid ICU messages, variable mismatches or missing static references. Existing locale gaps and 1,560 hardcoded-text candidates still need editorial work; they are not silently marked translated.
- Headless Edge verification of actual shared components with compiled CSS and the CMS theme at widths 1,280 and 390 pixels: the switch changes state and thumb position, the dialog stays within the 720px viewport, the final button is reachable, and the background page does not scroll. This isolated fixture does not establish full authenticated-page parity.
- Deployment rollback, verified-digest publication, ownership/cancellation and concurrent news edit behavior have focused regression tests.
Docker runtime, database migration execution and authenticated browser flows still require an integration environment. Check the Gitea pipeline and live release identifier after publication. A successful local build is not production verification.
## Catalog packages
The normal catalog toolbar now opens a dedicated Catalog packages dialog.
Create a named draft from selected categories and their descendants, either to
update those categories or copy them under a chosen parent. Drafts are shared
with authorized staff; saving a draft does not modify the live catalog.
Edit category metadata and offer prices, or review bulk price changes before
applying them to the draft. The catalog preview supports category navigation,
search, real local furniture icons and rank/Club/VIP access simulation. It does
not render the game client or evaluate ancestors outside the selected package;
special layouts and that access limitation are disclosed in the preview.
Publication requires a saved draft and a fresh review. Concurrent source changes
block publication; version checks prevent one editor overwriting another.
Copying preserves the underlying category and offer fields and remaps internal
references while retaining furniture IDs and assets. The catalog writes and
published result are committed together; retrying the same published package
does not copy it again. Failures in hotel notifications, audit or Git export
scheduling after commit are reported as warnings rather than failed publication.
Apply additive migration `0027_catalog_packages.sql` before opening this tool.
Existing migration automation discovers the file. Limits are 200 categories,
500 offers and 8 MB of package data. Package source checks inspect at most 20,000
catalog categories. Publication briefly locks category rows while validating and
writing changes, so large live catalogs should be checked under realistic load.
English, Italian and Dutch copy is provided; other locales use the new English
strings pending translation. No new dependency is required.
Validation: 1,506 tests passed (six skipped), type checking, lint, translation
contracts and a production build with fixture configuration. A browser fixture
verified the real dialog at 1280 and 390 pixels; server actions were simulated.
The new database migration and package publication have not run in production.
## Housekeeping search and user overview
The existing global search now includes furniture, normal/Club catalog categories
and both ticket sources. Results respect module permissions, accept single-digit
IDs, and remain usable when one source fails. Keyboard navigation and cancellation
prevent stale search responses from replacing newer results.
User details open on an operational overview with up to five records from each
authorized source: bans, active mute, support tickets, help tickets, reports,
payments, catalog purchases and audit activity. Failed sources are distinguished
from empty results. User detail and edit pages also apply log permissions before
loading activity and exposing counters.
Italian and Dutch navigation, user management, news and support labels were
reviewed. The new search and overview copy has English, Italian and Dutch text;
other locales receive English fallback strings. This is a focused editorial pass,
not a full translation of every CMS page. No database migration or dependency
change is required. Browser checks use real components with simulated data at
1280 and 390 pixels; production database behavior still needs deployment validation.
## Operational reliability and recovery
- Audit history now records category settings (normal/Club), individual/bulk offer prices, update-mode package publication and news edits inside the write transaction. The audit screen previews and restores individual changes after locking and comparing the current recorded fields. Deleted records, hierarchy changes, LTD counters, imports and historical entries without complete snapshots cannot be restored. Restores create their own history entry. Apply migration `0028_history_snapshots.sql` before running this version: it widens audit snapshots to MEDIUMTEXT without deleting existing data.
-`/admin/operations` reuses durable import jobs, stable history pagination and owner-scoped failed-item retries. A deterministic child ID prevents duplicate retries. The shared Git export queue displays actual pending/running state and latest result; synchronous/SSE synchronization remains linked rather than represented as a durable job history.
- Catalog maintenance includes a read-only integrity report for normal/Club categories, offers, furniture references and local icons. Known sentinel IDs are preserved. Only categories pointing to a missing positive parent have an automated repair: preview lists every affected category and apply compares the locked graph before reattaching those categories at the root. No records are deleted. Other issues require an explicit manual edit. Reports display up to 200 issues with complete counts; unavailable icon storage is distinguished from missing assets.
- CMS errors support exact release and time filters plus frequency sorting. Counts refer to retained matching events, while group resolution remains current across releases.
- User/settings forms now protect unsaved edits and preserve failed submissions. User/news validation errors appear at the affected fields; settings show returned validation errors inline. Existing submission locking is retained and tested in a browser.
-`/admin/permissions/preview` shows one role's section access and known CMS grants using live ACL and the existing highest-rank policy. It never changes sessions. Additional user roles, navigation customization and record-specific authorization remain explicit limits of the preview.
New UI copy is supplied in English, Italian and Dutch; other locales receive English fallback strings. No new runtime dependencies. Browser fixtures use real UI components with simulated server responses; the database migration and production behavior have not been exercised on the live hotel.
Validation for this increment: 1,589 tests passed, six skipped; TypeScript, Biome, translation contracts and fixture production build passed. Browser checks covered user/settings/news forms, permission preview, CMS errors, integrity preview and history restore at 1280 and 390 pixels. Double submission, stale preview, blocked navigation, field focus and horizontal overflow were checked with simulated server actions. No live database writes or deployment were performed.
## September 11: public pages and staff workflows
- Docker stages now follow the exact `.nvmrc` release, enforced by the toolchain check.
-`/news` supports search, ordering by effective publication date and real pagination.
-`/events` supports upcoming/ongoing/completed filters, explicit UTC week windows, local displayed times and personal registrations.
-`/search` searches users, open rooms, published news and events with independent pagination and partial failure states.
-`/me` shows support replies, incoming friend requests, the next registered event and available referral rewards. Reply availability does not claim unread status.
- Profile privacy is managed in `/settings`. Wallet values are private by default; visitors do not receive hidden sections in HTML. Photo galleries initially show six photos and can expand to the loaded limit of 24.
- Ticket desks support waiting-for-staff and assignment filters, with elapsed time since the latest reply.
- HK table views save filters, order and visible columns per account and table (maximum 20). Existing session column preferences remain available until a named view is applied.
- Official and clone synchronization run through the existing durable import queue. Reloading restores history; interrupted uncertain writes still require inspection before repair. Successful items are not repeated.
- Publication preflight validates URL syntax/protocols and schedules, shows affected page links and keeps existing article previews. It does not claim remote URLs are reachable. Drafts remain savable. Partial event updates preserve omitted fields.
- Admin APIs return `x-operation-id`; server errors, staff audit records and import jobs share correlation context. Error and audit screens link to each other. Older records without this context remain readable.
- Public reads distinguish unavailability from empty results and real 404s, preserving independently available sections on home, dashboard, staff, photos, rankings and groups/forums.
### Data and verification
Additive migrations `0029_admin_table_views.sql` and `0030_profile_privacy.sql` run through the existing deployment migration runner. They create CMS-owned tables and do not change emulator user settings. Keep the existing shared storage volume and background jobs worker for durable imports.
Browser verification used real components with controlled data fixtures at 1280 and 390 pixels, including failure and partial-result cases. Production compilation and full lint were checked locally. No production content was created during those checks; real authenticated content and external source availability remain environment-dependent.
## Original furniture bundle recovery and progress
Catalog Studio queued imports and repairs now search other enabled Nitro sources when their initial downloads/conversion produce no local bundle. Recovery checks an exact classname, floor/wall type and positive revision against the source furnidata, then validates the bundle filename, internal name and PNG texture before writing it. It does not copy the alternative source's prices, IDs or descriptive metadata.
The recovery pass checks at most eight eligible sources, excludes the selected source, and has a 20-second network budget with four-second request limits. Catalog downloads are capped at 20 MiB; bundle downloads and attachment decompression are capped at 50 MiB. A bounded catalog cache avoids downloading full furnidata for every item. Blocked or incompatible sources can still require the original bundle to be attached manually.
Import history displays the current phase and elapsed time, the last phase on failure/interruption, and the source/revision of a recovered bundle. Phases are persisted under the existing worker lease; a retry clears old phase/provenance fields. Synchronization jobs also report their existing importer phases, but their clone-specific asset strategy is unchanged.
No additional secrets or environment variables are needed. Source definitions remain managed through the existing source configuration.
## Catalog workflow and public diagnostics
- Bulk offer edits retain the existing preview and now record complete price/category snapshots. The success notification offers an atomic batch Undo for 15 seconds; individual changes remain restorable from audit history afterward. Undo rejects changed records or missing categories rather than overwriting newer edits. No additional migration is needed beyond the existing history snapshot migration.
- Catalog Studio preserves the existing in-place search/filter/selection flow and now restores list scroll and focus after closing furniture details. Escape closes details before clearing selection. Obsolete list responses cannot replace a newer search; switching source invalidates pending review preparation.
- Import review groups furniture needing completion, conflicting records and unverified components. Each row lists missing or unknown components and suggests the next action. These are inspection recommendations; final import validation remains authoritative.
- DevOps performance has separate HK API and public-page groups, each limited to 200 recent samples for one hour. Public instrumentation measures root server page invocations on /me, news, profiles, events and search, including measured database/external work. It excludes metadata, separately rendered children, cached responses that do not invoke the page, network transfer and browser rendering. Static build invocations are excluded. Routes use fixed labels without usernames or search terms; component outcomes are distinct from HTTP status codes.
- Shared unsaved-change protection now coordinates dirty forms and protects global-search navigation. Prefix, badge and room-furniture editors protect explicit dismissal and remain open after failed saves. Browser Back is intercepted when the cancellable Navigation API is available; reload/close and links retain their existing protection. Prefix saving now waits for the real server action result before closing.
This opt-in operator tool creates one MariaDB logical dump and copies explicitly selected persistent files. It never runs during install, update or CI deployment. The only restore operation is a **disposable drill**: there is no production restore command, database target, destination directory or overwrite option.
## Scope and prerequisites
Use the existing project Node toolchain and Docker CLI/Engine on a Linux host. Creation uses a short-lived `mariadb:11.4.5` client on the Docker host network, so `127.0.0.1` means that host. Use a local Docker Engine/context with the same filesystem; remote Docker daemons and Docker Desktop are not supported for creation. The image must already be available or downloadable through the operator's normal image policy. No packages are installed by this tool.
Choose the single application database explicitly. Its tables must use InnoDB; empty databases, system schemas and unsupported engines are rejected. The backup account needs access to every application table, view, trigger, routine and event being exported. Account/grant provisioning belongs to the operator; the tool does not change privileges. Restore compatibility is checked with the pinned MariaDB image, not guaranteed across arbitrary server versions, plugins, collations or external schema dependencies.
Before starting, pause **every database/file writer** and schema changer for the whole creation command: CMS requests that write, workers, schedulers, emulator processes, MariaDB events, import jobs and other tools using these resources. Keep them paused until the command exits. `--writers-quiesced` records your acknowledgement; it does not stop services or prove that they are stopped. No DDL may run while dumping. InnoDB's transaction snapshot alone cannot make independently copied files consistent with rows or coordinate other services. The before/after table inventory and second source-file hash pass detect many concurrent changes, but cannot replace quiescing.
The dump explicitly uses `--single-transaction --quick --skip-lock-tables --routines --events --triggers --hex-blob --tz-utc`. It retains schema/data and named SQL objects while streaming rows. See [MariaDB dump snapshot and object options](https://mariadb.com/docs/server/clients-and-utilities/backup-restore-and-import-clients/mariadb-dump). This is a logical application backup, not point-in-time recovery: binlogs, server accounts/grants, server configuration, Redis state and unrelated databases are outside its scope.
## Private configuration
Create a private JSON file **outside the clone and every selected source directory**, for example `/etc/epicnext/backup.json`. Use a directory accessible only to the operator and file mode `0600`. Edit it with the host's normal private configuration workflow; do not put a password in a shell command or commit the file.
```json
{
"database":{
"host":"127.0.0.1",
"port":3306,
"user":"REPLACE_WITH_BACKUP_ACCOUNT",
"password":"REPLACE_PRIVATELY",
"database":"REPLACE_WITH_APPLICATION_DATABASE"
},
"roots":{
"storage":"/srv/epicnext/storage",
"nitro":"/srv/epicnext/public/nitro-assets",
"swf":"/srv/epicnext/public/swf",
"gamedata":"/var/www/Gamedata"
}
}
```
Replace the example clone path and database settings. `storage`, `nitro` and `swf` are required existing directories, including when empty. `gamedata` is optional; omit its key only when those files are independently backed up or not used. All paths must be absolute, distinct, non-overlapping and free of `..` and symlink/junction components. Links, hardlinked files, special files and secret configuration filenames such as `.env`, `.docker-install` and `persistent.path` inside a selected root cause rejection. The configuration itself may not be inside any source root. Keep writers and directory ownership controlled for the duration; this is not a filesystem snapshot resistant to hostile concurrent renames.
The tool reads only this explicit JSON. It does not source `.env`, inspect a running application's environment or inherit its secrets into Docker. The MariaDB password is written to a random private temporary directory/file (`0700`/`0600`) and read through a read-only container mount with `--defaults-file` as the first client option. It is absent from process arguments and tool logs; source configuration and absolute source paths are absent from the manifest. Temporary credentials are removed on ordinary success/failure. See [MariaDB option-file handling](https://mariadb.com/docs/server/clients-and-utilities/backup-restore-and-import-clients/mariadb-dump#defaults-file-name).
Only Docker connection/runtime environment variables are passed to child processes. Do not point `DOCKER_HOST`/`DOCKER_CONTEXT` at another host or enable shell tracing around private configuration work.
## Create and verify
Provision a private backup parent directory, with adequate free space, outside all source roots. Choose a **new** absolute artifact directory for each run. After pausing the writers described above, run from the clone root:
The date is an example; use a new name for the actual run. Existing directories are refused, including earlier incomplete attempts. A successful artifact contains:
```text
manifest.json
database.sql
files/storage/...
files/nitro/...
files/swf/...
files/gamedata/... (only when selected)
```
`manifest.json` is written last and marks the format complete. It records each relative file path, byte count and SHA-256, empty directories, selected logical roots, creation time and database inventory. The inventory includes table row counts and MariaDB extended table checksums plus names/types of views, triggers, routines and events. These checksums validate restored table contents; they are not cryptographic signatures or a substitute for application-level checks. See [MariaDB CHECKSUM TABLE semantics and version limits](https://mariadb.com/docs/server/reference/sql-statements/table-statements/checksum-table).
On a caught failure the newly created artifact is removed; an interrupted process can leave an incomplete directory, which verification refuses. Source roots and existing backup directories are never overwritten. Files are copied and hashed as streams, then source hashes are checked again across the dump interval. This performs multiple full reads of the assets and table data; allow sufficient time and disk capacity during the maintenance window.
After creation you may resume writers. Copy the completed artifact to the designated protected backup location according to the operator's retention/encryption policy. SQL and uploaded files contain application data and may themselves contain sensitive values. Hashes detect corruption against the manifest, not an attacker who can replace both. Keep the manifest and artifact under trusted access control. Secrets, TLS material and deployment configuration excluded from this artifact need their separate recovery procedure.
## Isolated restore drill
Run the drill against a trusted completed artifact:
The drill first rejects missing, extra, changed or unsafe paths/files. It copies the persistent files into a new private temporary directory and verifies their contents. It creates a randomly named MariaDB container with **no network and no published ports**, a fresh password supplied by file, and a disposable database volume. SQL is imported through the container's standard input; there is no connection to the source database. The event scheduler stays off. The resulting table counts/checksums and object inventory must match the manifest. This proves dump importability and the recorded contents, not a full CMS/emulator startup or external-service recovery.
On ordinary completion/failure it removes the container, its anonymous volume and temporary files. Cleanup failure makes the drill fail. A host crash or forced process termination can interrupt cleanup; inspect only resources labeled `cms.backup-drill=true` and private `cms-backup-private-*` temporary directories from that run, and review their ownership before manual removal. Never substitute an existing database/container into this procedure.
A real production recovery remains a separate, reviewed procedure with its own deployment configuration, credentials, downtime and application checks. This tool deliberately cannot perform it.
## Repository verification
```sh
pnpm exec vitest run --coverage.enabled=false scripts/backup
pnpm exec vitest run --config vitest.integration.config.ts integration/backup.test.ts
```
The local suite exercises real file copies, empty directories, SQL/file tampering, manifest traversal, links, secret-file rejection, overwrites, cleanup and changes during backup. The integration suite requires Docker: failure to start MariaDB fails the suite. It uses a real database with Unicode text, large unsigned identifiers, a foreign key, view, trigger, procedure and event; it creates an artifact, restores it to a second disposable server and checks both byte corruption and a SQL content change with a recomputed file hash. A passing local file suite alone is not evidence that the MariaDB drill ran.
Push work to a `codex/**` branch and open a pull request targeting `main` or `master`. CI runs the existing `check` job first. After it passes, the new `preflight` job builds the production Dockerfile and runs the isolated news browser suite against that exact image. Review both results before merging; repository branch protection can require `check` and `preflight` for pull requests.
Each execution uses `epicnext-cms:preflight-<commit>-<random suffix>`, including retries and separate push/PR runs of the same commit. The full checked-out commit is passed as `NEXT_DEPLOYMENT_ID` and `NEWS_E2E_RELEASE`; `NEWS_E2E_IMAGE` identifies that execution's image. Failures in dependency/browser setup, Docker build, news tests or cleanup fail the job. News browser artifacts are uploaded even when the gate fails.
The script installs dependencies with the frozen lockfile and installs Chromium on the CI runner. The real Docker build uses the existing Dockerfile's fixture build settings. It never copies or sources a deployment `.env`, connects to a VPS, runs live migrations, updates live containers, publishes a registry image or changes release tags. The existing isolated news runner owns its disposable MariaDB, Redis and application containers. Cleanup removes only the preflight tag and its empty private temporary directory; it does not prune Docker resources.
The `deploy` and `publish-container` conditions remain restricted to pushes on `main`/`master`. Deployment still runs its own news gate before live migrations/cutover. A successful branch preflight supplies earlier evidence; the deployed commit is independently checked again.
On a Linux development or CI host with the project toolchain, Docker Engine and normal browser prerequisites, the same gate can be run from a clean checkout:
```sh
bash scripts/ci-preflight.sh
```
Shell orchestration is covered by `pnpm exec vitest run --coverage.enabled=false src/lib/ci-preflight.test.ts`. Those tests execute the real shell script with external command boundaries simulated; they prove ordering, failure propagation, unique tags and cleanup scope. They do not build an image or run the news browser suite. The branch/PR CI job provides that Docker/browser evidence.
Dependency updates are manual: Renovate has been removed, so there is no bot
opening upgrade pull requests. Node/pnpm upgrades remain coordinated with
Docker and the runner toolchain.
Obsolete global overrides were removed; a scoped esbuild override remains because
Drizzle Kit's loader still resolves a vulnerable legacy development-server build.
The two deprecated esbuild-kit packages remain upstream dependencies of Drizzle
Kit; replacing the ORM is not warranted for this tooling issue.
## HK request performance
Open **DevOps → Request performance** (`/admin/devops/performance`). Access requires `DEVOPS_VIEW`; collection only retains requests that passed authentication and permission checks through `withAdmin`.
The view shows recent requests, their server duration, completed database query count/time/errors, monitored curl download count/time/errors, HTTP status and the existing operation ID. Search by route or operation ID, sort by duration or recency, and filter requests taking at least one second.
### Measurement boundaries
- Duration ends when the handler returns its response. Streaming completion, background jobs, browser rendering and public pages are not measured.
- Database spans wrap the shared mysql2 promise pool and transaction connection query/execute calls. Query parameters and SQL are never collected. Explicit transaction connection acquisition and transaction control methods are not separate spans.
- External spans currently cover `curlFetchText` and `curlDownload`; ordinary fetch calls and other integrations are not covered.
- Concurrent dependency durations can exceed wall-clock request duration. Do not subtract the sums to infer application CPU time.
- Dynamic route values and query strings are removed. No request bodies, headers, usernames, download URLs or SQL text are stored.
### Storage and operation
No new environment variables or packages are required. Redis stores at most 200 recent samples under `cms:performance:v1`, with one-hour retention. Writes are best effort, restricted to an already-ready connection and at most four pending batches; the request never waits for persistence. The view reads at most 200 entries and falls back after one second if shared storage is unavailable.
An in-process buffer preserves up to 200 samples during outages. The page explicitly labels this local mode; it is per instance and disappears on restart. Filtering applies to the retained samples, not complete traffic history. Under load or during storage outages, some shared samples may be omitted.
Use a Linux host with Docker Engine, the Compose plugin, Git and `flock`. This Compose file uses host networking and port 3002. Provide an existing compatible Habbo MariaDB database, reachable Redis, persistent storage and an HTTP(S) reverse proxy. The installer does not provision the emulator, a database, TLS, or a registry account.
Clone this repository from the Gitea URL supplied by your administrator, enter the clone, then run:
```sh
bash cms install
```
The wizard asks for the public URL and delivery mode. **Source** (the default for a new installation) builds the locked source locally and only needs repository access plus access to public build dependencies. **Prebuilt** downloads the application and migrations from the repository's `docker-image.txt`; a private package needs a separate registry login with package-read permission. Git access alone may not grant package access. Existing saved mode and `.env` are preserved. `bash cms install --configure-only` saves configuration without starting containers.
Credentials are entered on the host, never in the dashboard. Do not paste `.env` into support tickets. Installation prepares `public/nitro-assets`, `public/swf`, `storage`, and `/var/www/Gamedata` for UID/GID 33 without recursively taking ownership of existing files. Existing nested assets may still need operator permission repair. The updater tests access to those mounts as the candidate container user before migrations; this checks directory access, not every nested file or filesystem capacity.
After startup, visit `/admin/devops/installation` with the appropriate permission. Verify database, Redis, storage and migration status. Worker heartbeat is a separate runtime signal: a healthy HTTP endpoint does not prove an import worker is processing jobs. Check the worker status and investigate a missing/stale heartbeat before scheduling imports. The dashboard is read-only and cannot start Docker, upgrade the host or grant registry access.
## Client IP trust at the reverse proxy
The application validates and normalizes client addresses from `cf-connecting-ip`, the first `x-forwarded-for` entry, then `x-real-ip`. It never accepts `x-real-client-ip`; that legacy derived header is also stripped by the Next.js proxy. Missing or invalid addresses resolve to `0.0.0.0` for rate limits and audit records. API routes use the same resolver even though they do not run through the Next.js proxy.
These headers are trustworthy only when the ingress sanitizes them. Configure the reverse proxy to discard client-supplied forwarding/derived headers and replace the accepted address from a verified connection or a specifically trusted upstream proxy. Do not append an untrusted incoming `x-forwarded-for` chain and then treat its first entry as authoritative. Forward `cf-connecting-ip` only after verifying that it came through your trusted CDN path; otherwise remove it.
Restrict direct access to the application port so requests must pass through that ingress. The provided Compose file uses host networking with `HOSTNAME=0.0.0.0`; it does not enforce this restriction or provision nginx/Traefik trust rules. Verify the host firewall and actual reverse-proxy configuration before relying on client IPs for blocking, auditing or abuse limits. Repository tests prove rejection of the derived-header bypass and malformed addresses; they do not certify the deployed forwarding trust chain.
## Opt-in profile: Nginx on the same host
Use this profile for a new Linux Compose clone whose public HTTPS endpoint is Nginx on that same host. It leaves `docker-compose.yml` and CI-managed production deployments unchanged. Nginx must include `ngx_http_realip_module`; check `nginx -V` before using the templates. The CMS remains on the existing host network for database/Redis connectivity, but its HTTP process listens on `127.0.0.1:3002`. Host networking shares the host network namespace; a `ports:` mapping would not provide the restriction. See [Docker host networking](https://docs.docker.com/engine/network/drivers/host/).
1. Run `bash cms install --configure-only` and choose the intended public HTTPS URL. Obtain a valid certificate for that hostname using the host's existing certificate-management process. The templates do not issue certificates or configure renewal.
2. In the clone's existing `.env`, add or replace this single setting, preserving all other values:
Keep `APP_URL`, `AUTH_URL` and the saved installer public URL on the same canonical `https://` hostname. The profile pins the existing port 3002 as well as the loopback address. Do not place `COMPOSE_FILE` in `.docker-install`; that file accepts only `MODE` and `PUBLIC_URL`.
3. Copy [nginx-direct.example.conf](../../deployment/proxy/nginx-direct.example.conf) into the host's Nginx configuration directory, outside this Git clone. Replace **every** `hotel.example` and both certificate paths. Load it at `http` scope, for example through `/etc/nginx/conf.d/cms.conf`. Review any existing virtual host for that hostname to avoid two competing configurations. The example handles only CMS HTTP traffic; emulator WebSocket and other hotel services need their own reviewed ingress.
4. Validate the effective Nginx configuration with `nginx -t`, then enable/reload it through the host's normal service-management procedure. Prepare this endpoint before starting the installer, because installation verifies the saved public URL. It can return an upstream-unavailable response until the CMS starts.
5. From the clone root, remove conflicting Compose overrides from the deployment shell and validate the model:
These `unset` commands remove shell overrides; they do not remove the `COMPOSE_FILE` line in `.env`. Do not set `COMPOSE_DISABLE_ENV_FILE=1`, use an alternate `--env-file`, or supply a competing shell `COMPOSE_FILE` for this workflow. Environment values can override `.env` selection. Do not print or paste the full rendered Compose configuration, because it contains runtime credentials. See [Compose predefined variables and precedence](https://docs.docker.com/compose/how-tos/environment-variables/envvars/).
6. Confirm the running configuration without dumping the environment:
The listener must be `127.0.0.1:3002`, not `0.0.0.0:3002` or `[::]:3002`. From another machine, the host's public IP on port 3002 must be unreachable. Check both address families when the host has IPv6. Loopback isolation covers the CMS process only: other containers and host services, including Byparr, retain their existing bindings.
The direct template uses the original socket peer (`$realip_remote_addr`), overwrites the two accepted forwarding headers, and removes incoming `CF-Connecting-IP`, `X-Real-Client-IP` and `Forwarded`. It fixes forwarded host/protocol to the configured HTTPS origin. An unrelated inherited real-IP rule cannot turn a caller-supplied header into the forwarded client address in this mode. See [Nginx original-peer variables](https://nginx.org/en/docs/http/ngx_http_realip_module.html) and [header replacement/removal](https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_set_header).
Response buffering is disabled for progress streams, and this example adds no proxy response cache or CORS policy. Its 64 MiB ingress body cap accommodates the existing 52 MiB Studio attachment limit; route and Server Action limits remain authoritative and may be lower. TLS 1.2/1.3 are configured explicitly. See [Nginx buffering](https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_buffering) and [TLS protocols](https://nginx.org/en/docs/http/ngx_http_ssl_module.html#ssl_protocols).
### Why the profile survives an update
The installer preserves an existing `.env`. The installer invokes `docker-update.sh`, and the updater's config, build, candidate run, cutover and rollback paths all call **bare `docker compose` from the clone root**, without `-f`. Compose therefore reads the saved file list on each invocation. The base file appears first so its relative mount/build paths remain rooted in the clone; the later profile overrides only the CMS environment and healthcheck. No installer/updater patch is required for this selection. See [Compose merge order and relative paths](https://docs.docker.com/compose/how-tos/multiple-compose-files/merge/).
Keep the profile selected for every routine `bash cms update` and selected-release update. A one-off `docker compose -f ... up` does not persist selection for later installer/updater commands. Do not add an untracked `docker-compose.override.yml`: it makes the clone dirty unless separately excluded and is bypassed when `COMPOSE_FILE` selects an explicit list. Before selecting an older release, verify that `deployment/proxy/compose.loopback.yml` exists in that release; a missing selected file fails configuration rather than silently using the public bind.
This profile does not change `scripts/ci-deploy.sh`, which owns a separate `docker run` deployment and currently sets its own public binding. Do not use the clone installer to take over a host managed by CI. Restricting that deployment's listener requires a separate compatibility review and rollout.
### Separate mode: verified remote edge or Cloudflare
A remote proxy cannot connect to the CMS loopback listener directly. The supported topology here is **remote edge → TLS → Nginx on the CMS host → loopback CMS**, retaining the same Compose profile. Use [nginx-trusted-proxy.example.conf](../../deployment/proxy/nginx-trusted-proxy.example.conf) instead of the direct template. Do not enable both templates for one hostname.
For your own remote edge, replace `203.0.113.10/32` in both the `geo` peer allowlist and `set_real_ip_from` with the exact approved connection source addresses. The reserved example address deliberately permits no real edge. Require the edge to overwrite `X-Forwarded-For` with one verified client IP, use the configured hostname, enforce public HTTPS, and validate this origin's TLS certificate. The origin checks the original socket peer before accepting the rewritten address. Never use `0.0.0.0/0`, `::/0` or arbitrary client networks as trusted proxies. See [Nginx real-IP trust configuration](https://nginx.org/en/docs/http/ngx_http_realip_module.html).
For Cloudflare, use that same restricted-edge mode and replace both peer lists with the **current verified IPv4 and IPv6 Cloudflare ranges**, maintained by the operator; use `real_ip_header CF-Connecting-IP` instead of `X-Forwarded-For`. Configure Full (strict) TLS and review authenticated origin pulls. Obtain ranges from [Cloudflare's official IP list](https://www.cloudflare.com/ips/) and follow its [visitor-IP restoration guidance](https://developers.cloudflare.com/support/troubleshooting/restoring-visitor-ips/restoring-original-visitor-ips/). Do not copy a historical range list from a support ticket or trust `CF-Connecting-IP` merely because it is present. This setup still removes the CF header before the CMS and forwards only Nginx's normalized result in XFF/X-Real-IP.
The direct template intentionally records the CDN/edge socket address when placed behind an unconfigured CDN; it does not silently trust an upstream header. After configuring the restricted-edge mode, verify with requests from an allowed edge and a disallowed direct client, including forged CF/XFF headers, and check the recorded client address. Neither the repository nor the installer changes the host firewall or certifies another proxy's header behavior.
### Local verification and deployment boundary
`pnpm exec vitest run --coverage.enabled=false scripts/proxy-config.test.mjs` runs the installed Docker Compose CLI against a disposable clone configuration, without contacting Docker Engine, building images or reading the real `.env`. It verifies selection through `.env`, loopback/port override precedence, the IPv4 healthcheck and preservation of mounts, release selection and host networking. It explicitly skips when Compose is unavailable. This is not a container-start or network-isolation test.
`pnpm test:integration` additionally starts disposable Nginx containers from the actual templates, supplies a temporary test certificate, and sends real HTTPS requests with forged identity headers. It checks direct-mode replacement even with an inherited real-IP rule, rejection of untrusted peers, and acceptance through an explicitly trusted peer. This requires Docker Engine and the OpenSSL CLI and does not read deployment credentials. The templates must still pass `nginx -t` on the intended host after its hostname/certificate substitution, then the listener and trusted-header checks above; the disposable fixture cannot certify that host or its firewall.
## Routine and selected-release updates
```sh
bash cms update
```
This requires a clean clone and fast-forwards its configured Git upstream, then builds/pulls artifacts for that exact commit. It validates the application and migration revision labels, validates runtime configuration and storage, runs migrations and verifies the recreated CMS locally and through the saved public URL.
To install a specific published version, fetch it and select its commit on the host first:
```sh
git fetch origin
git switch --detach <reviewed-commit-or-tag>
bash cms update --skip-pull
```
`--skip-pull` deliberately uses the checked-out commit, including detached HEAD. For routine updates again, switch back to your tracked deployment branch. The script does not invent version-to-schema compatibility or automatically change branches.
In prebuilt mode you can additionally require immutable artifacts from the configured repository:
Replace both placeholders with publisher-provided digests. Both are mandatory together. Revision labels must match the checked-out commit; a digest alone does not establish schema compatibility. Older migration images without a revision label must be republished from matching source, or the same checkout can be installed in source mode. No migration runs when the artifact or candidate validation fails.
## Compatibility and recovery
Before upgrading, read the selected release's migration changes and take a database backup with a tested restore procedure. Preserve `.env`, persistent assets and the previous release identifier. The current tooling does not provide a validated matrix of supported source/target database versions. Pinning an old commit is therefore not a supported database downgrade procedure.
On a failure after container replacement, the updater attempts to restore the previous image and checks it locally. **Image rollback does not reverse database migrations.** A migration may partially apply or make the old application incompatible, including when migration fails before container replacement. Recover the database only through the reviewed backup/restore procedure and coordinate downtime; do not assume restarting the old image recovers it. First installation has no prior image to restore. Host logs remain in `logs/docker-update.log` and may contain application/database diagnostics; restrict access.
Local Git Bash tests cover selection validation and mocked failure paths. They do not establish Linux container startup, runtime filesystem permissions, registry availability or real MariaDB upgrade compatibility.
Public API bearer authentication accepts only tokens owned by the exact `App\Models\User` model. The owner ID must be a positive, safely representable user ID, and the token must satisfy its existing expiration check. Both plaintext tokens and the existing `{id}|{plaintext}` request format remain supported; only the SHA-256 hash is looked up in the database.
The `abilities` column must contain a non-empty JSON array of non-empty strings. Null, malformed JSON, non-array JSON, empty arrays, non-string entries, and entries with surrounding whitespace are rejected. A valid `"*"` entry grants access to all existing bearer-protected endpoints. Other permissions match exactly: there is no `tickets:*` expansion, implicit read/write inheritance, or fallback to unrestricted access.
| `badges:read` | Personal viewer data in `GET /api/badges/leaderboard` |
For example, `["tickets:read","radio:read"]` allows reading the owner's tickets and radio points. It cannot create tickets, send replies, post article comments, or send radio shouts. Endpoint ownership checks and rate limits still apply after scope authorization.
Required-token endpoints return the existing generic `401 Unauthorized` response when authorization fails. The badge leaderboard remains public: a denied bearer token receives the anonymous view, without personal viewer data. When an Authorization header is present, this endpoint does not use a session cookie to bypass a denied token. Session-only requests continue to personalize the leaderboard normally.
## Compatibility and maintenance
Existing valid wildcard tokens remain compatible. The existing session-authenticated `POST /api/tokens` endpoint continues issuing `["*"]`; this change does not add token-creation options or alter stored tokens. Legacy null, malformed, empty, differently cased model names, and unrelated model tokens are intentionally denied. Review and replace affected tokens with explicit intended scopes, or reissue through the existing token endpoint when full access is appropriate.
Every new bearer-authenticated endpoint must pass its required abilities to `bearerUserId`. Multiple required abilities use AND semantics. Omitting the requirements, or passing an empty list, requires a wildcard token rather than granting arbitrary scoped tokens access.
No plaintext token or stored hash is added to error responses or logs by these checks. The existing issuance endpoint returns plaintext once by design.
The supplied review describes commit `baeb54ae` plus a separate port for another hotel. Its “Fixed” labels were not evidence that the changes existed in EpicNext-Cms. This verification inspected canonical `main` at `52f6d149` and the corrective changes prepared here. No exploit or authenticated mutation was performed against production.
| Supplied finding | Verified state in baseline | Correction / remaining boundary |
| --- | --- | --- |
| 1. Logo authorization/upload | Confirmed missing action permission and per-file validation | Require settings edit before input or storage access; bounded decoded raster uploads |
| 3. Active uploaded SVG | Confirmed SVG served inline without route CSP | Route CSP sandbox and nosniff on success/errors; existing SVG served as attachment |
| 4. Client IP spoofing | Confirmed direct trust in caller-controlled `x-real-client-ip` | Shared validated resolver ignores that header. Forwarded headers still require trusted ingress that overwrites them and prevents direct public origin access |
| 5. Email token action exports | Confirmed token helpers in a `use server` module | Move token creation/validation and delivery to a server-only module. Registration and verification call it internally |
| 6. Locale cookie | Confirmed missing allowlist | Supported locales only, validate before reading/writing cookies |
| 7. Email header injection | Confirmed unsanitized values in sendmail headers | Reject control characters before any mail transport or file fallback; includes configured sender |
| 8. Gateway CORS | Supplied gateway path is outside this repository | Read-only GET to our `/api/health` with an unrelated Origin returned a fixed `https://epicnabbo.nl` allow-origin and no allow-credentials. This does not reproduce the report on that route, nor certify every host/route |
| 9. Token abilities | Confirmed abilities and owner type not checked by bearer authentication | Enforce User owner type and explicit endpoint abilities; existing wildcard user tokens remain supported |
| 10. Broad script CDN | Confirmed unrestricted jsDelivr script source, without a source-code consumer | Remove the broad script source; retain required captcha/analytics sources and nonce |
The additional `withNitroStaff` code and its tests mentioned in the supplied port do not exist in this checkout; they were not assumed to have been reviewed or imported.
## Evidence and limits
- Regression tests exercise authorization before I/O, actual file decoding, SVG/error response headers, token-boundary exports, token abilities, forged derived-IP headers across consumers, locale values, and mail header control characters.
- An updated `pnpm audit --json` reported zero known advisories. This is a dependency database result, not proof that application code has no vulnerabilities.
- Next.js treats exported Server Actions as public endpoints; unused actions can also be removed by the compiler. The email refactor removes the action boundary entirely instead of relying on whether a specific build exports an unused helper. See [Next.js data security](https://nextjs.org/docs/app/guides/data-security).
- The framework also has its own Server Action body limit. The logo defect was absence of application-level file validation, not evidence of literally unlimited bytes through every deployment layer.
- No live database, user accounts, uploaded files, or gateway configuration were modified during verification. These changes do not constitute a penetration test or an audit of the emulator, host, or all CMS endpoints.
- No nginx/Traefik ingress configuration is versioned here. The deployment guide records the forwarding-header trust requirement. That external boundary remains unverified.
## Follow-up identified during verification
The separate comment review is now implemented: both the website form and REST API use one submission service, require a published article whose publication time is due, apply the same moderation, and share a five-attempt/30-second per-user quota. The publication check locks the current article in the insertion transaction. Regression tests cover both entrypoints; real MariaDB/Redis coverage includes publication eligibility, word filtering and alternating submissions. Moderation retains its existing fail-open behavior on service outages. Form input beyond 255 characters is now rejected instead of truncated, and temporary API storage failures return 503. Real integration execution remains a required CI check.
The command writes `report.json` and `report.md` and prints the Markdown report. An optional `PERFORMANCE_COMMIT_SHA` environment variable records the commit declared by the build caller; the script does not infer that an existing build matches the current checkout. JSON also records `BUILD_ID`, Node/zlib versions, manifest provenance, exact file paths, sizes and source entries.
## What is measured
For each configured App Router route, resolve its exact app path using `app-path-routes-manifest.json` and `server/app-paths-manifest.json`. Read its generated `page_client-reference-manifest.js` as a JSON assignment **without executing JavaScript**. Use its sibling `page/build-manifest.json`, falling back to the root build manifest only if that sibling is absent.
The **initial entry envelope** is the union of route bootstrap `rootMainFilesTree[appPath]` (or `rootMainFiles`) and every client chunk that route's client-reference manifest lists. This includes layout, page and boundary/loading entries.
The manifest exposes those chunks differently per bundler. Turbopack emits an explicit per-segment `entryJSFiles` map; webpack emits no such field and records chunks only per client module, as `clientModules[*].chunks`, in `[chunkId, fileName, chunkId, fileName, …]` order. The report reads `entryJSFiles` when present and otherwise derives the same envelope from `clientModules`, which is the source Next's own `static-routes-info` uses. Numeric chunk ids are skipped; a malformed chunk *path* still fails rather than being dropped, so a broken manifest cannot quietly under-report a route.
> The build runs webpack (`next build --webpack`), so the `clientModules` path is the live one. An earlier revision only read `entryJSFiles`, and after the switch to webpack every route reported `unavailable` while the command still exited 0 — the budgets were silently not being measured. When a bundler switch changes the manifest layout again, re-check this section rather than trusting a clean exit.
The definition follows the data exposed by the installed Next 16.3.8 build and the `getLinkAndScriptTags` / `getRequiredScripts` renderer helpers; it is deliberately a build-artifact envelope, not a browser network trace. Conditional rendering, redirects, streaming and browser caches can change actual requests.
- Raw bytes are filesystem byte lengths of unique JavaScript assets in that envelope.
- Gzip bytes are the **sum of independent gzip level 9 compressions** of those files using the recorded Node/zlib runtime. They are not gzip of concatenated source, nor observed CDN transfer sizes.
- Deployment query strings and `/_next/` prefixes are normalized before deduplication. Shared files count once per route; each route is measured independently, with no misleading cross-route total.
- Legacy `nomodule` polyfills are measured separately, outside the modern initial budget. CSS, source maps, images, external scripts, HTML/RSC payloads and async-only chunks absent from the manifest's chunk lists are excluded.
- This report makes no claims about execution cost, LCP, hydration time or real-user performance.
## Initial limits
The first limits are **baseline bytes × 1.15, rounded upward to the next 10 KiB (10,240 bytes)** independently for raw and gzip. They are provisional size alerts, not validated speed targets. Baseline: local production build `build-TfctsWXpff2fKS`, Next 16.3.4 **Turbopack**; its source commit was not inferred.
The production build now runs webpack, so the numbers it reports are not directly comparable to the baseline below. Re-measured on the current webpack build the routes land at `/me` 786138/247738, `/news` 781601/245672, `/events` 782011/245923, `/search` 783262/246578, `/admin/catalog` 1172089/370843, `/admin/studio/furni` 1374128/440546 (raw/gzip). All remain inside the limits below, but `/admin/studio/furni` sits at ~98% of its gzip limit, so the next dependency added to that route will trip it. Recalibrate the table and `scripts/performance-budgets.json` together if the intent is to reset the baseline on webpack.
| Route | Baseline raw bytes | Baseline gzip bytes | Raw limit | Gzip limit |
Configured limits are positive integer bytes; `null` explicitly means observe-only. `scripts/performance-budgets.json` remains `mode: informational`. Exceeding a limit produces `over-budget` and a warning, with exit code 0. Missing production `BUILD_ID`, unsupported manifests, missing routes or missing referenced assets produce `unavailable` with a reason and **no partial/zero total**, also exit code 0. Malformed budget configuration or an unwritable output directory fails the command. This keeps initial CI reporting non-blocking while preventing invalid configuration from quietly disabling limits.
Synthetic tests cover shared-chunk deduplication, exact byte/gzip calculations, route bootstrap selection, missing data, safe parsing and CLI exit behavior. Run `pnpm exec vitest run --coverage.enabled=false scripts/performance-report.test.mjs`.
User approved bulk editing and complete category duplication, keeping Visual Manager in its button-opened modal.
Use subagent-driven-development for the independent duplication task; root owns bulk operations and final review.
- [x] Duplicate category subtree and offers atomically for normal/BC; preview counts, destination/name, fresh IDs, preserve furniture references; fail on stale source, invalid destination, cycles and insert failure. Shared permissions, single export/RCON outcome. UI in Visual Manager.
- [x] Extend selected-offer operations to preview and atomically apply prices, currency, and destination. Visibility belongs to categories, so expose it separately with exact affected categories; no fictitious per-offer flag. Preserve unselected and unrelated fields; conflicts prevent overwriting concurrent changes.
- [x] Meaningful domain/transaction/action tests, browser fixtures for modal preview/confirm, typecheck/Biome/i18n/Knip; review limits and commit exact scope.
No production data mutations, added dependencies or schema migrations. Full category duplication shares existing items_base definitions and asset files. Repeated confirmations must be disabled while saving. Preview should be revalidated under locks before write. Existing direct save and permissions remain.
## Implementation notes
Bulk editing now works from selected normal-catalog offers and from selection across global searches. Only explicitly chosen prices/currency/destination fields change. Prices support set/add/percentage with integer rounding and range validation; 500 selected offers per transaction. Preview binds rows, changes and category names; concurrent changes reject confirmation. Category visibility remains on the existing category controls, not a fabricated offer property. BC has no offer pricing and does not expose that editor.
Duplication copies up to 500 categories and 5000 offers, preserving bundle and asset references. New root starts disabled and hidden. Includes and internal offer IDs are remapped. Explicit normal offer IDs use the existing allocator and do not require AUTO_INCREMENT; database collisions roll back the whole copy. Preview binds destination siblings as well as source data so confirmed preview replay is rejected. Dirty editor state blocks copying saved data until drafts are saved/reset.
No database migration or dependency added. Database transaction tests use mocks and do not prove live MariaDB locking behavior. Browser fixtures use real components with mocked actions, not production writes. Existing allocator is process-local; competing external imports may cause a safe rollback requiring a fresh preview. No automatic retry of uncertain copy commits.
Verification complete: 1339 unit tests passed, 5 skipped. Typecheck/Knip/i18n and changed-file Biome checked. Real-component browser fixtures passed bulk selection, preview, conflicts, pending guards, mobile bounds and table refresh without phantom drafts; duplication nested inside actual manager verified menus, picker, preview and dirty-state guards. Fixed manager portal layering and invalid menu-label nesting uncovered by those tests. Browser actions are mocked; no live data written or deployment claimed.
> For agentic workers: use superpowers:subagent-driven-development for independent changes and review each deliverable.
Goal: deliver the approved progressive catalog refactor while preserving current features.
Architecture: shared domain validation and transactional commands behind existing action contracts; unified editor reuses current views and tools.
Stack: Next, React, Drizzle/MariaDB, Zod, Vitest, Playwright; no added dependencies.
Spec: docs/CATALOG_REFACTOR_PLAN.md
Constraints: preserve normal/BC differences, direct save, real icons, permissions, existing routes, no favorites. No production data mutations for testing.
- [x] Characterize compatibility and command rules with tests; capture baseline source map.
- [x] Domain hierarchy/reorder validation and transactional page commands (normal/BC), deterministic locking, common updates/deletion.
- [ ] Verify unit tests, database integration if runtime available, browser, typecheck, lint/i18n, full suite; review final scope and remaining environment-only checks.
Execution notes: changes are progressive, no destructive migration. Existing UI adapters stay until functional parity is tested. Whole-program completion must not be claimed from the first deliverable.
## First implementation delivery (2026-09-06)
Completed: shared normal/BC page and offer commands; transactional page reorder/delete/move and offer create/update/reorder; hierarchy validation and optimistic page-save checks; embedded default manager with classic views retained; request cancellation, editor unsaved-change guards and URL selection; bounded global category/offer/furniture search; separate hotel/Git status and hotel retry; export-finalization failures no longer mask committed server actions. Existing permissions and direct-save behavior retained. No added dependency or DB migration.
Verification: full unit suite 1308 passed / 5 skipped before final retry-classification regression; final focused catalog suite 96 passed. TypeScript, i18n static validation, Knip and Biome on all 54 changed source files pass. Browser fixtures cover editor loading races, retry failures, unsaved changes, URL history, read-only, normal/BC, search and 375px layout. Database calls and mutations are mocked in those fixtures. Full-repository formatter check reports pre-existing Windows CRLF formatting differences; unrelated files were not reformatted.
Remaining roadmap: per-operation durable dispatch/outbox, category/offer snapshot history and conflict-aware undo, contextual maintenance diagnosis, pagination/performance measurement against representative real data, further decomposition of retained legacy views. Existing coarse audit/export infrastructure remains; the new status file is not a durable outbox and RCON socket success does not prove client application. External furni importer mutation internals remain outside shared editor commands.
Environment checks still required: real MariaDB locking/rollback integration, staff-authenticated end-to-end smoke test and pipeline/deployment health. No production mutations performed; no production deployment claimed.
# Security and operational reliability implementation plan
Goal: finish the five approved follow-ups with independently verified commits.
Architecture: share comment policy between session and bearer entrypoints; opt-in same-host proxy configuration; run real news browser checks against the already-built candidate in disposable services; correlate existing diagnostics with deliveries; verify database and persistent-file backup restoration in isolation.
Stack: existing Next, MariaDB, Redis, Playwright, Testcontainers and Docker; no new dependencies.
Design: user-approved numbered proposal in this task, 2026-09-13.
Global constraints: preserve current public/HK UX and ACL; no production test content or proxy/firewall changes; no credentials in output; root owns Git on canonical main. Complete each block's checks before an exact-file commit and push. Confirm final CI, container publication and live release.
1. Comments — src/actions/article-comments.ts, API comment route and shared policy/tests. Add regression cases for hidden/future articles, moderation, cross-channel limit and safe failures; reproduce them, implement, run focused and integration checks. Publicly available article predicate is checked on both entrypoints.
2. Proxy — deployment/proxy templates and installation guide/tests. Override must survive installer/update/rollback, force loopback and replace incoming identity headers. Validate the merged Compose config and Nginx syntax; do not apply to the host.
3. Real news — e2e/news-real runner/fixture plus ci-deploy gate and harness tests. Start only disposable MariaDB/Redis and the local candidate image; real staff login, draft, preview, publish and anonymous read. Fail before live migration/cutover on any error, clean all fixture resources. Require successful CI execution.
4. Diagnostics — carry persisted operation/delivery identifiers into error records; link filtered deliveries and diagnostics with permission checks. Preserve request correlation separately. Tests cover exact matching, hostile IDs, permissions and retry outcomes.
5. Recovery — backup creation and isolated restore drill for database plus explicit persistent directories. Keep credentials off argv/logs, reject unsafe paths and incomplete/tampered artifacts. Test real database restore and file checksums with disposable data, record limits for cross-service consistency. Never overwrite production during a drill.
Status: all five blocks implemented and locally checked. Required final gates: CI real database/proxy/backup suites, candidate news browser journey, deployment/container completion and live release verification. Extra scheduler deadlock discovered in the real concurrency test is fixed with bounded transaction retries. Evidence and boundaries accompany each delivered block.
Run `pnpm test:integration` on a Docker-capable host. Missing Docker or failed container setup fails the suite. CI runs this check before deployment.
Sixteen database tests exercise the production database commands, news actions, public article query and delivery worker against MariaDB 11.4.5 and Redis 7.4.2:
- Migration CLI replay/status, committed catalog bulk edits and undo history, complete rollback after an audit insert fails, and competing catalog previews.
- Real Redis expiry metadata and cache-key isolation, concurrent request idempotency, operation/outbox rollback, and exclusive delivery claims.
- Concurrent duplicate draft creation and publication produce one article, one revision, one audit update and one effect per operation. The public query changes from cached absence to the full published content, including Unicode and a body larger than a TEXT column.
- A database trigger rejects the publication effect after the article, revision and audit writes. The transaction restores all preceding state; the identical request can then retry successfully without duplicate history.
- A trigger rejects the second scheduled-publication effect. Both article updates and both operations roll back, including the first queued effect.
- Competing scheduler ticks publish each due article once, preserve future articles and drafts, retain an unsigned bigint ID beyond JavaScript's safe integer range, and attribute a legacy authorless article to the system actor.
- A real Redis client disconnect leaves a publication committed and readable directly from MariaDB. The worker records a pending failed attempt. Reconnecting retains the stale cached absence until a successful outbox retry rotates the cache revision; the public query then returns the published article. The test advances only the queued retry timestamp to avoid sleeping.
Each execution starts disposable containers with random exposed ports and generated passwords. No production URLs, volumes or credentials are used. The real migration CLI is copied beneath a temporary isolated fixture root so its environment loader cannot read the checkout environment file. Cleanup attempts all connections and containers even if a previous cleanup fails. The fixture inspection connection reads timestamps as UTC; the application keeps its production connection settings and host timezone. Each test receives a fresh news-cache revision, and the disconnect test reconnects its client in a `finally` block.
Only authorization/session lookup, translation lookup, Next.js revalidation/redirects, and the external publication webhook are mocked for the news actions. MariaDB, Drizzle, transactions, article revisions, history, operation deduplication, outbox claims/retries, Redis caching, the publication scheduler, public article lookup and the news delivery handler use their production implementations.
The fixture models the emulator's catalog/audit baseline plus the legacy article columns. Real migrations 0025-0029 and 0031 supply publication, editorial recovery, catalog packages, history and operations tables. This does not certify every historical emulator schema or migration.
These are application-service integration tests, not browser or HTTP end-to-end tests: they do not start a Next.js server, render the news page, verify login/ACL behavior, or send external webhooks. The cache outage test closes and restores the actual application Redis connection; it does not stop the Redis server or model a multi-host network partition. Passing TypeScript or unit tests without Docker is not a passing result for this suite.
Public comment pagination and reaction aggregation are covered by unit tests using the production Drizzle query builder with a substituted database transport, plus server-rendered page tests with substituted service results. This integration fixture does not create users or article-reactions tables and does not execute the public pagination and aggregation queries against MariaDB; its article-comments table is used for submission tests. Comment submission now additionally uses real MariaDB and Redis to verify publication eligibility, filtering and the shared quota across the actual form/API handlers, with authentication replaced at the boundary. The article-cache upgrade test verifies that legacy cached absence is ignored under the versioned key; it is a unit test, separate from the real Redis publication and delivery checks above.
Run `pnpm test:news:real` with `NEWS_E2E_IMAGE=epicnext-cms:<candidate-tag>`. Docker CLI, a working Docker daemon, the OpenSSL CLI with `-addext` support, installed project dependencies and Playwright Chromium are required. The runner deliberately fails when the image or Docker is unavailable. It never silently skips the browser journey.
`NEWS_E2E_RELEASE=<full-commit-sha>` additionally verifies the image revision label. The runner resolves the local image to its immutable ID before starting it. It does not build or pull the application image. MariaDB 11.4.5 and Redis 7.4.2 Alpine are disposable Testcontainers dependencies and may be pulled when absent.
The deployment script runs this gate after building the candidate and extracting its performance report, before live database migrations and container cutover. The general Playwright suite excludes `e2e/news-real`; this suite has its own configuration and requires the runner.
## What the browser proves
One Chromium journey exercises the candidate's normal Docker entrypoint and standalone Next server:
1. An anonymous request to the staff editor reaches the login page.
2. The browser signs in through the actual login form, credential precheck and Auth.js handler. The test verifies the real Secure/HttpOnly session cookie and database login record.
3. A rank 7 editor, below an occupied rank 9 owner, accesses news through three explicit ACL grants: `admin.dashboard`, `admin.news.view` and `admin.news.edit`.
4. The browser types Unicode content into the bundled TinyMCE editor and saves a draft through the real server action. The persisted HTML must equal what the form submitted.
5. An independent anonymous browser sees the not-found screen; the public articles API returns an empty list. Redis contains the cached null article. Next can stream a not-found screen with HTTP 200, so this check verifies the rendered 404 screen as well as absence from the API.
6. The editor reopens and previews the saved draft in the real preview iframe. Its title, summary and rendered body match; it remains a draft with no article revision created by previewing.
7. Publishing through the editor produces one article, a publication timestamp, one previous-draft revision, a before/after audit entry, two completed operation results and two news-refresh outbox entries.
8. The anonymous page and API immediately show the article, using a new shared Redis cache revision. The browser remains signed out.
No authentication, application HTTP responses, mutations, database calls or cache calls are mocked. The browser blocks resources outside the local fixture origin, such as external avatars. This does not replace any application response. A local HTTPS edge passes requests to Next and sets trusted forwarding headers, allowing the production Secure-cookie behavior to run normally.
The fresh browser contexts retain normal application service-worker registration. The application worker does not cache article or authentication responses. Playwright's worker-blocking injection is avoided because it throws inside the sandboxed preview; browser errors are still checked without filtering. The preview is closed through its Close button, since keyboard events inside its sandbox do not reach the parent dialog.
## Isolation and cleanup
- Each run creates a random Docker network and fresh MariaDB, Redis and candidate containers. All mapped ports are allocated dynamically. The HTTPS listener binds only to `127.0.0.1`.
- No production container, database, volume or credentials are reused. Container configuration is explicit. Child processes receive a small environment whitelist; checkout `.env` and installation service credentials are not forwarded.
- The database uses the current ORM column definitions for 29 tables needed by login, site/admin layouts and news. Unique constraints and composite primary keys are preserved. The emulator-owned `permission_ranks` fixture provides the rank authority columns this flow queries. Real CMS migrations 0026 and 0031 create the recovery/revision and operation/outbox tables. This is a focused fixture, not a replacement for the emulator's complete schema or migration coverage.
- Passwords, Redis credentials and the Auth.js secret are generated per run. The owner has an unknown random password and is never used to bypass ACL checks. Email verification is enabled with a verified fixture staff account. CAPTCHA and forced staff 2FA use their ordinary disabled installation settings; their challenges are outside this flow.
- OpenSSL generates a fresh localhost certificate and private key in the OS temporary directory for each run. The certificate lasts one day and covers `127.0.0.1` and `localhost`. Playwright trusts this self-signed endpoint only in this suite. No certificate or key is committed or copied into build artifacts; cleanup removes both with the temporary credentials.
- Credentials used by the test worker are stored in an OS temporary directory with a mode-0600 file, then removed. Cleanup stops only containers and the network created by this runner; Testcontainers also registers them with its resource reaper.
- Failure artifacts are under `test-results/news-real` and `playwright-report/news-real`. Server logs redact generated passwords/secrets. Playwright traces can contain the short-lived fixture login/session data; all associated services are destroyed after the run.
This browser gate checks the synchronous editor/publication path and durable delivery intent. Scheduler concurrency, rollback, duplicate requests, worker delivery and Redis outage/recovery remain covered by `pnpm test:integration`. It does not claim to test emulator connectivity, external notifications, CAPTCHA/2FA challenges or production data.
The local workstation currently has no working Docker daemon. Type checks, lint and fixture/bootstrap checks can run there; a passing real browser result must come from the Docker-capable CI gate.
Run `pnpm test:ui` after `pnpm install --frozen-lockfile` and `pnpm exec playwright install chromium`.
The fixture server bundles real CMS components with local HTTP/server-action fixtures. It does not use the database or production credentials. Production smoke tests remain in the separate `pnpm test:e2e` command.
The suite covers accessible settings switches, news editor/preview, support-table filters, import history, attachment recovery and news/event recovery. Axe checks WCAG A/AA rules. The sandboxed preview cannot execute Axe; its accessible name and keyboard exit are checked separately. Automated checks do not replace a full accessibility audit.
Visual references are committed under `__screenshots__/<platform>/<viewport>`. Normal runs compare references and must not update them. For an intentional UI change, use `pnpm test:ui:update`, inspect every changed PNG and commit only accepted references. Keep the Playwright browser version aligned with the lockfile. Linux references must be captured on the Linux runner; Windows references are not interchangeable.
Fixtures contain synthetic data only. Do not add production requests, environment secrets, ignored local documents or database access to tests. CI retains test reports and failures as the `ui-results` artifact.
Loaded 100 of 1412 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.