Two follow-ups from the blocked deploy.
The rollback path called `docker logs` and `docker rm -f` on the candidate
unconditionally. When the port check refuses to start it, the container was
never created, so both printed "No such container: epicnext-cms-app" — noise
that looked like a second, unrelated failure and buried the real message.
Both calls are now guarded by `docker inspect`.
assert_port_free() now reports which container holds the port and flags it when
it is not a blue/green slot this script manages. The previous output listed
every container and said only "port already in use", which is a dead end: on
this host the holder is `epicnext-cms` (a `docker compose up` replica on port
3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ
because docker-compose.yml pins `container_name: epicnext-cms` for the `cms`
service; slot B happens to match, which is why 3003 deploys fine and 3002 never
can. The message now names the squatter, explains that live traffic is
unaffected, and gives the next action.
Deploy simulation: 26 passed.
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.
The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.
Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.
Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.
Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.
read_active_port() now orders its sources by how well they describe reality:
1. The nginx upstream file. It is the only source that says where public
traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.
answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.
start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.
Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
- Remove random TTL jitter to prevent unpredictable cache drops
- Add deterministic LRU eviction with proper entry cleanup
- Improve cache deduplication to prevent duplicate computations
- Skip Redis I/O during tests for faster, more stable execution
- Optimize depth calculation in catalog tree nodes
- Maintain backward compatibility and full test coverage (3331 passed)
Icons are plain .png, a cacheable extension by default, so the zone's
"Browser Cache TTL = 1 year" pinned them to max-age=31536000 regardless of
the 300/3600/604800 that nginx sends per class. Extend the edge rule to
/gamedata/ and keep respect_origin, so the nginx header wins and a 404
(notably no-store from the gamedata 404 handler) is never pinned.
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.
The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.
The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.
- `writeFurniData` now purges the gamedata edge tag itself. One place covers
import, batch, resync, regen, nitro-editor, translate and dedupe. It is
fire-and-forget and swallowed at every level: a stale edge copy is bounded
by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
`max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
and the content-addressed trees (c_images, album*, clothes) keep the long
TTL. `must-revalidate` is the point — the client now revalidates instead of
replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
Cloudflare was unreachable at write time.
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
picks the one that really answers; the upstream file only serves as a
fallback when zero or both slots respond. A stray 'docker compose up' (or a
clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
only seed it when missing, never overwrite what a deploy wrote. This is the
root cause of tonight's 502: a nginx-sync run reset the snippet (written to
green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
incoming CF-Connecting-IP, 403 any other peer that presents one
(spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.
- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
blue/green upstream snippet; config backed by scripts/nginx-sync.sh
(idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
no-op unless Cloudflare is configured; scripts/cf-purge.sh and
cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy
directory's .env. When that variable was missing the deploy had already
built an image and run the browser gate before pnpm db:migrate aborted on
an empty value, so a release was paid for in full and then thrown away.
Check for the variable right after the .env is copied, before the build,
and say plainly that the live release was not touched. The deploy test
fixture gains a DATABASE_URL so it mirrors a working deploy directory
instead of the broken one.
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a
third-party tool that ships a WebP member while still pointing
spritesheet.meta.image at a .png was accepted and stored as-is. The client
resolves the spritesheet through that pointer, so the result was a file
that validates fine and then renders nothing.
Re-write the bundle through createNitroBundle on import, which labels the
member from the actual bytes and repairs the pointer. No texture is
re-encoded, so the bytes stay identical, and the member keeps the base
name it arrived with so `chair*2` colour variants that share the `chair`
library are not renamed.
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.
Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
Legacy md5/argon2id hashes are now always upgraded to bcrypt on login, so
the CONVERT_PASSWORDS flag is no longer used. Drop it from env schema,
.env.example, the docker installer, and test mocks.
- Restore build_attempted=1 in ci-preflight.sh so the exit trap
removes the temporary image tag
- Remove publish-container.test.ts and its harness (publication
workflow and script were removed in fff284aa)
- Update deploy-workflow-contract and docker-build-contract tests
to assert that publication has been removed
Catalog Studio:
- Cross-parent drag & drop now uses optimistic updates with rollback
on failure (no more full tree reload / visible delay)
- Subpage creation adds the node optimistically then refreshes parent
only (was full tree reload)
- Single page deletion refreshes only the affected parent (was full
tree reload)
- Root page creation replaces native prompt() with an inline input
in the root tab bar
- Escape key no longer closes the dialog when an input field is focused
- TreeNodeUpdate type now supports parentId and orderNum for
optimistic structural changes
Docker:
- docker-prune.sh default mode now aggressively cleans all unreferenced
build cache, images >1h old, and stopped containers >1h old
(was 72h/7d/24h which let cache grow past 80% on every push)
The 5-minute disk probe now reclaims storage automatically: from 85% it runs the gentle age-windowed Docker prune, from 90% it drops the age windows (docker-prune.sh --force: all unused build cache and unreferenced images, all stopped containers) so a mount can never silently max out. Alerts still fire at 85/90/95% and their hint now points at non-Docker growth when reclaiming is not enough. Force mode is reserved for the worker; deploys keep the gentle mode. Volumes are off-limits in every path.
Add a pure df parser (disk-usage.ts) with 85/90/95% threshold classification, a diskPressure() alert (Discord/email/alert_logs, severity escalates with fill), and a 5-minute host-side probe in jobs-worker.ts that raises one alert per crossing mount, cooldown-gated per mount+level. Real mounts only: overlay/tmpfs pseudo filesystems are ignored.
Add scripts/docker-prune.sh (build cache >72h capped at 4g, unreferenced images >7d, stopped containers >24h; never volumes), run it after every CI deploy and compose update, and schedule a nightly prune from the host-side jobs-worker. Tighten the deployment contract tests to assert the scoped-prune boundaries.