Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.
jobs-worker never ran
`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.
The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.
/api/health answered 200 with the database down
The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.
The runtime had no V8 heap cap
NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.
Storage ownership was only repaired for one path
ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.
nginx: robots.txt was a guaranteed 404, and TLS never resumed
`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.
Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.
The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.
The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.
- `writeFurniData` now purges the gamedata edge tag itself. One place covers
import, batch, resync, regen, nitro-editor, translate and dedupe. It is
fire-and-forget and swallowed at every level: a stale edge copy is bounded
by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
`max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
and the content-addressed trees (c_images, album*, clothes) keep the long
TTL. `must-revalidate` is the point — the client now revalidates instead of
replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
Cloudflare was unreachable at write time.
The 5-minute disk probe now reclaims storage automatically: from 85% it runs the gentle age-windowed Docker prune, from 90% it drops the age windows (docker-prune.sh --force: all unused build cache and unreferenced images, all stopped containers) so a mount can never silently max out. Alerts still fire at 85/90/95% and their hint now points at non-Docker growth when reclaiming is not enough. Force mode is reserved for the worker; deploys keep the gentle mode. Volumes are off-limits in every path.
Add a pure df parser (disk-usage.ts) with 85/90/95% threshold classification, a diskPressure() alert (Discord/email/alert_logs, severity escalates with fill), and a 5-minute host-side probe in jobs-worker.ts that raises one alert per crossing mount, cooldown-gated per mount+level. Real mounts only: overlay/tmpfs pseudo filesystems are ignored.
Add scripts/docker-prune.sh (build cache >72h capped at 4g, unreferenced images >7d, stopped containers >24h; never volumes), run it after every CI deploy and compose update, and schedule a nightly prune from the host-side jobs-worker. Tighten the deployment contract tests to assert the scoped-prune boundaries.
Align onlyBuiltDependencies with the workspace, fail fast without AUTH_SECRET in production, and delete catalog_items via VARCHAR-safe SQL so page deletes do not leave orphans.
Co-authored-by: Cursor <[email protected]>
Sentry is opt-in via DSN env vars; logger uses structured pino JSON in prod; badge uploads are normalized to GIF with sharp.
Co-authored-by: Cursor <[email protected]>
Beyond parity — the web-feasible versions of the "host-only" items plus
extras AtomCMS doesn't have:
- App-level abuse/DDoS guard (src/lib/services/abuse-guard.ts): counts
requests per IP and auto-adds flooders to website_ip_blacklist (enforced
by the access guard) + fires ddosDetected(). OFF by default, tunable via
settings. The iptables layer stays host-only; this is the real app-tier
mitigation. Access guard now also enforces the IP blacklist (cached).
- PWA: a themeable web manifest (src/app/manifest.ts) + a service worker
(public/sw.js, cache-first assets / network-first pages) registered after
hydration — the hotel is now installable.
- /api/health: DB + emulator(RCON) + runtime status probe.
- /developers: a public API documentation page covering every REST endpoint
with its method, path and auth requirement.
- jobs-worker: daily emulator JAR backup (runs host-side in the worker, like
AtomCMS's backup command) — copies + prunes; no-ops unless EMULATOR_JAR_PATH
+ EMULATOR_BACKUP_DIR are set.
Verified live (prod, amx_test): /api/health ok, manifest + sw served, docs
page renders, normal pages unaffected by the guard. tsc 0, vitest 49/49,
next build 0.
Security (launch blockers):
- src/middleware.ts (edge): forwards x-pathname + real client IP.
- access-guard.ts (Node, from root layout): routes non-staff to /maintenance
when maintenance mode is on, banned users to /banned. New /banned + /maintenance
pages (the consumers the admin toggle was missing). Admin layout enforces
force_staff_2fa before /admin.
- staff-activity.ts audit log wired into ban/lift/give-currency/set-rank actions.
Infra (parallel agents): alert service (alert_logs + Discord embed + email),
PayPal top-up (create/capture API routes + /shop/topup), cron worker
(scripts/jobs-worker.ts via croner: emulator-ping->alert, maintenance-check,
bans-cleanup), social connections page, admin radio settings/banners/ranks.
Public radio subsystem: /radio (+schedule, shouts+post, contests, giveaways,
apply, leaderboard) and /apply/staff + /apply/team submission forms. Radio nav
link added. .env.example documents the new optional vars.
(radio song-requests dropped: its table is a stub in AtomCMS — columns added by
un-modeled alter-migrations.)
Verified: tsc exit 0, vitest 48/48, next build exit 0 (82 page routes).