Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.
The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.
The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.
- `writeFurniData` now purges the gamedata edge tag itself. One place covers
import, batch, resync, regen, nitro-editor, translate and dedupe. It is
fire-and-forget and swallowed at every level: a stale edge copy is bounded
by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
`max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
and the content-addressed trees (c_images, album*, clothes) keep the long
TTL. `must-revalidate` is the point — the client now revalidates instead of
replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
Cloudflare was unreachable at write time.
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
picks the one that really answers; the upstream file only serves as a
fallback when zero or both slots respond. A stray 'docker compose up' (or a
clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
only seed it when missing, never overwrite what a deploy wrote. This is the
root cause of tonight's 502: a nginx-sync run reset the snippet (written to
green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
incoming CF-Connecting-IP, 403 any other peer that presents one
(spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.
- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
blue/green upstream snippet; config backed by scripts/nginx-sync.sh
(idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
no-op unless Cloudflare is configured; scripts/cf-purge.sh and
cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy
directory's .env. When that variable was missing the deploy had already
built an image and run the browser gate before pnpm db:migrate aborted on
an empty value, so a release was paid for in full and then thrown away.
Check for the variable right after the .env is copied, before the build,
and say plainly that the live release was not touched. The deploy test
fixture gains a DATABASE_URL so it mirrors a working deploy directory
instead of the broken one.
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a
third-party tool that ships a WebP member while still pointing
spritesheet.meta.image at a .png was accepted and stored as-is. The client
resolves the spritesheet through that pointer, so the result was a file
that validates fine and then renders nothing.
Re-write the bundle through createNitroBundle on import, which labels the
member from the actual bytes and repairs the pointer. No texture is
re-encoded, so the bytes stay identical, and the member keeps the base
name it arrived with so `chair*2` colour variants that share the `chair`
library are not renamed.
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.
Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
Legacy md5/argon2id hashes are now always upgraded to bcrypt on login, so
the CONVERT_PASSWORDS flag is no longer used. Drop it from env schema,
.env.example, the docker installer, and test mocks.
- Restore build_attempted=1 in ci-preflight.sh so the exit trap
removes the temporary image tag
- Remove publish-container.test.ts and its harness (publication
workflow and script were removed in fff284aa)
- Update deploy-workflow-contract and docker-build-contract tests
to assert that publication has been removed
Catalog Studio:
- Cross-parent drag & drop now uses optimistic updates with rollback
on failure (no more full tree reload / visible delay)
- Subpage creation adds the node optimistically then refreshes parent
only (was full tree reload)
- Single page deletion refreshes only the affected parent (was full
tree reload)
- Root page creation replaces native prompt() with an inline input
in the root tab bar
- Escape key no longer closes the dialog when an input field is focused
- TreeNodeUpdate type now supports parentId and orderNum for
optimistic structural changes
Docker:
- docker-prune.sh default mode now aggressively cleans all unreferenced
build cache, images >1h old, and stopped containers >1h old
(was 72h/7d/24h which let cache grow past 80% on every push)
The 5-minute disk probe now reclaims storage automatically: from 85% it runs the gentle age-windowed Docker prune, from 90% it drops the age windows (docker-prune.sh --force: all unused build cache and unreferenced images, all stopped containers) so a mount can never silently max out. Alerts still fire at 85/90/95% and their hint now points at non-Docker growth when reclaiming is not enough. Force mode is reserved for the worker; deploys keep the gentle mode. Volumes are off-limits in every path.
Add a pure df parser (disk-usage.ts) with 85/90/95% threshold classification, a diskPressure() alert (Discord/email/alert_logs, severity escalates with fill), and a 5-minute host-side probe in jobs-worker.ts that raises one alert per crossing mount, cooldown-gated per mount+level. Real mounts only: overlay/tmpfs pseudo filesystems are ignored.
Add scripts/docker-prune.sh (build cache >72h capped at 4g, unreferenced images >7d, stopped containers >24h; never volumes), run it after every CI deploy and compose update, and schedule a nightly prune from the host-side jobs-worker. Tighten the deployment contract tests to assert the scoped-prune boundaries.
The Arcturus errors "page hierarchy contains a cycle page 354 and 357" and
"sibling order 1 is used more than once (111 problems)" come from
catalog_pages, not catalog_items: pages 354/357 point at themselves, and
many parents have child pages sharing the same order_num. Extend the
emulator catalog scan + fix to detect both: pages whose parent chain loops
back get detached (parent_id = 0 on the highest cycle member) and every
affected parent's children are renumbered sequentially, preserving their
current relative order.
Add scripts/diag-emulator.ts to inspect the live catalog state.