The deploy failed after the build, the migrations and the browser gate:
"Port 3002 is already in use". The holder was `epicnext-cms`, a compose
replica of release 6bffc537 that the daily scripts/docker-update.sh cron
had recreated at 03:30 with restart=unless-stopped. nginx serves the green
slot on 3003, so that replica was squatting the blue slot the next
candidate needed, and live traffic never noticed.
It got there because the updater's CI-ownership guard only tested
epicnext-cms-app. After a cutover to the green slot that container is
stopped, renamed and deleted, so the guard stopped firing while the host
stayed CI-managed.
- scripts/docker-update.sh: refuse a compose deployment on a CI host by
checking both slot containers and the nginx upstream, which is the only
thing that still marks the host as blue/green while a slot is idle.
- scripts/ci-deploy.sh: retire a compose replica of this checkout from
the candidate port before starting the candidate, so a stray replica
can never block a release again. Never a slot container, never the port
nginx serves; anything else still fails loudly in assert_port_free.
- Tests cover both directions: a squatting replica is removed and the
release lands, a replica on the live port is left alone.
Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.
jobs-worker never ran
`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.
The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.
/api/health answered 200 with the database down
The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.
The runtime had no V8 heap cap
NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.
Storage ownership was only repaired for one path
ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.
nginx: robots.txt was a guaranteed 404, and TLS never resumed
`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.
Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
The deploy could not start its candidate because port 3002 was held by
`epicnext-cms`, a `docker compose up` replica built from the `local` image and
serving no traffic. Everything else in the pipeline was healthy: the image
built, the news browser gate passed and migrations were current.
The container was unusable for this pipeline for two reasons. It ran a
different image than any release, and its name did not match the slot the
deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms`
while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to
agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never
could. deploy.sh already documents that compose "never managed the release
that actually ran", so the service was stale by its own account.
Removed the stray container and dropped the `cms` and `cms-green` services (plus
the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot
resurrect a replica that permanently occupies a blue/green slot. byparr is
untouched.
Also fixed the diagnostic from the previous commit, which blamed every running
container. `docker ps --filter publish=` returns nothing for --net=host
containers, so the fallback listed all of them and buried the real holder
among seven innocent ones. It now resolves the listening PID from `ss` back to
its container through /proc/<pid>/cgroup and names only that one, with the
exact `docker rm -f` command to run.
Verified: port 3002 free, live release on 3003 still serving
(status ok, database and redis true), deploy simulation 26 passed, typecheck.
Two follow-ups from the blocked deploy.
The rollback path called `docker logs` and `docker rm -f` on the candidate
unconditionally. When the port check refuses to start it, the container was
never created, so both printed "No such container: epicnext-cms-app" — noise
that looked like a second, unrelated failure and buried the real message.
Both calls are now guarded by `docker inspect`.
assert_port_free() now reports which container holds the port and flags it when
it is not a blue/green slot this script manages. The previous output listed
every container and said only "port already in use", which is a dead end: on
this host the holder is `epicnext-cms` (a `docker compose up` replica on port
3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ
because docker-compose.yml pins `container_name: epicnext-cms` for the `cms`
service; slot B happens to match, which is why 3003 deploys fine and 3002 never
can. The message now names the squatter, explains that live traffic is
unaffected, and gives the next action.
Deploy simulation: 26 passed.
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.
The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.
Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.
Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.
Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.
read_active_port() now orders its sources by how well they describe reality:
1. The nginx upstream file. It is the only source that says where public
traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.
answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.
start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.
Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
picks the one that really answers; the upstream file only serves as a
fallback when zero or both slots respond. A stray 'docker compose up' (or a
clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
only seed it when missing, never overwrite what a deploy wrote. This is the
root cause of tonight's 502: a nginx-sync run reset the snippet (written to
green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.
- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
blue/green upstream snippet; config backed by scripts/nginx-sync.sh
(idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
no-op unless Cloudflare is configured; scripts/cf-purge.sh and
cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy
directory's .env. When that variable was missing the deploy had already
built an image and run the browser gate before pnpm db:migrate aborted on
an empty value, so a release was paid for in full and then thrown away.
Check for the variable right after the .env is copied, before the build,
and say plainly that the live release was not touched. The deploy test
fixture gains a DATABASE_URL so it mirrors a working deploy directory
instead of the broken one.
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.
Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
Add scripts/docker-prune.sh (build cache >72h capped at 4g, unreferenced images >7d, stopped containers >24h; never volumes), run it after every CI deploy and compose update, and schedule a nightly prune from the host-side jobs-worker. Tighten the deployment contract tests to assert the scoped-prune boundaries.
Drop the leftover Playwright browser install and e2e smoke test from the deployment script, and update the deployment contract tests to cover the verify-deployed-release smoke check instead.
- Move the pnpm store, apk and .next caches into --mount=type=cache so
dependencies are shared across builds instead of duplicated in fresh
image layers (was the source of unbounded disk growth).
- Replace the deprecated --keep-storage prune flag in ci-deploy.sh with
the working --max-used-space=4g (buildx v0.37 renamed the flag). The
deprecated flag silently did nothing, so the BuildKit cache kept
growing unbounded (was 15.86GB); it is now capped at 4GB after every
deploy.
next.config.ts falls back to `git rev-parse HEAD`, but the build context has
no .git (excluded by .dockerignore), so every CI build printed fatal git
errors and stamped the release as "unknown". Pass the deploy commit sha as a
NEXT_DEPLOYMENT_ID build-arg so git is never invoked and the actual commit
reaches NEXT_PUBLIC_CMS_RELEASE and deploymentId.