A missing gamedata file got no Cache-Control at all, because add_header
without `always` only applies to 2xx/3xx. Cloudflare then fell back to the
zone setting "Browser Cache TTL = 1 year", so the 404 came back as
`max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's
browser and at the edge. An icon requested while its import was still
running stayed a 404 for the rest of the year, even after the file existed.
That was the "some icons load, some don't" report.
Give every gamedata location a named 404 handler that sends no-store, and
split icons/ out as its own cache class: those files are rewritten under
the same name (repair-icons, reimport), so an hourly must-revalidate keeps
a repaired icon visible within the hour instead of days later.
/gamedata/* is served straight from disk by nginx; no request hits the CMS
backend or a database, so a request-rate limit protects nothing while
costing players their icons. A room load fires hundreds of these files in
one burst, which every limit turned into visible 503s.
Removed the static zone from the gamedata locations. Traefik's
epicnabbo-gamedata router likewise carries no rateLimit middleware.
/client/ and /nitro-client/ keep theirs, and the main route keeps the
30r/s page budget plus the server-wide connection limit.
Measured: 1000 icon requests fired fully in parallel now all return 200,
while 200 parallel requests on / are still rejected.
The single server-scope limit_req (30r/s) treated a page load and a room
load as the same thing. Loading a Nitro room fires several hundred gamedata
icons in one burst, which that zone answered with 503s, so icons showed up
late in the client.
Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata
and client asset locations, and apply the page-rate zone explicitly on the
main route instead of at server scope. Connection limit stays server-wide.
Measured: 900 icon requests in burst now all return 200, while 200 parallel
requests on / are still rejected.
The edge had no limit_req/limit_conn at all, so a single client could
flood the Next.js backend and the Nitro client with unbounded parallel
requests. Traefik's logs already showed this: bursts of gamedata icon
requests answered with 429.
Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed
on the real client IP, applied at server scope so both cached assets and
proxied API routes share one budget. The burst is deliberately generous
because the Nitro client fetches gamedata and icons in bursts when
loading a room.
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.
The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.
The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.
- `writeFurniData` now purges the gamedata edge tag itself. One place covers
import, batch, resync, regen, nitro-editor, translate and dedupe. It is
fire-and-forget and swallowed at every level: a stale edge copy is bounded
by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
`max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
and the content-addressed trees (c_images, album*, clothes) keep the long
TTL. `must-revalidate` is the point — the client now revalidates instead of
replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
Cloudflare was unreachable at write time.
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
incoming CF-Connecting-IP, 403 any other peer that presents one
(spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.
- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
blue/green upstream snippet; config backed by scripts/nginx-sync.sh
(idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
no-op unless Cloudflare is configured; scripts/cf-purge.sh and
cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.