Commit Graph
10 Commits
Author SHA1 Message Date
openhands dbaccd7cfc fix(gamedata): never cache a missing gamedata file
A missing gamedata file got no Cache-Control at all, because add_header
without `always` only applies to 2xx/3xx. Cloudflare then fell back to the
zone setting "Browser Cache TTL = 1 year", so the 404 came back as
`max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's
browser and at the edge. An icon requested while its import was still
running stayed a 404 for the rest of the year, even after the file existed.
That was the "some icons load, some don't" report.

Give every gamedata location a named 404 handler that sends no-store, and
split icons/ out as its own cache class: those files are rewritten under
the same name (repair-icons, reimport), so an hourly must-revalidate keeps
a repaired icon visible within the hour instead of days later.
2026-10-01 18:00:36 +02:00
openhands 4e036b08d5 fix(proxy): drop request rate limiting from gamedata entirely
/gamedata/* is served straight from disk by nginx; no request hits the CMS
backend or a database, so a request-rate limit protects nothing while
costing players their icons. A room load fires hundreds of these files in
one burst, which every limit turned into visible 503s.

Removed the static zone from the gamedata locations. Traefik's
epicnabbo-gamedata router likewise carries no rateLimit middleware.
/client/ and /nitro-client/ keep theirs, and the main route keeps the
30r/s page budget plus the server-wide connection limit.

Measured: 1000 icon requests fired fully in parallel now all return 200,
while 200 parallel requests on / are still rejected.
2026-10-01 17:56:48 +02:00
openhands f0dcf440a7 fix(proxy): split rate limiting into page and static zones
The single server-scope limit_req (30r/s) treated a page load and a room
load as the same thing. Loading a Nitro room fires several hundred gamedata
icons in one burst, which that zone answered with 503s, so icons showed up
late in the client.

Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata
and client asset locations, and apply the page-rate zone explicitly on the
main route instead of at server scope. Connection limit stays server-wide.

Measured: 900 icon requests in burst now all return 200, while 200 parallel
requests on / are still rejected.
2026-10-01 17:49:41 +02:00
openhands 4a1211a931 feat(proxy): add per-IP rate and connection limits
The edge had no limit_req/limit_conn at all, so a single client could
flood the Next.js backend and the Nitro client with unbounded parallel
requests. Traefik's logs already showed this: bursts of gamedata icon
requests answered with 429.

Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed
on the real client IP, applied at server scope so both cached assets and
proxied API routes share one budget. The burst is deliberately generous
because the Nitro client fetches gamedata and icons in bursts when
loading a room.
2026-10-01 17:25:49 +02:00
openhands e3c010f383 fix(proxy): raise nginx worker rlimit above worker_connections
nginx inherited systemd's soft LimitNOFILE of 1024, so every start logged
"2048 worker_connections exceed open file resource limit: 1024" and the
worker_connections value could not actually be reached.

Set worker_rlimit_nofile to 65536. Bounded from above by a systemd drop-in
at /etc/systemd/system/nginx.service.d/override.conf (LimitNOFILE=65536),
since the master's hard limit caps what workers may request.
2026-10-01 17:11:04 +02:00
openhands a6cc3cafa9 fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.

The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.

The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.

- `writeFurniData` now purges the gamedata edge tag itself. One place covers
  import, batch, resync, regen, nitro-editor, translate and dedupe. It is
  fire-and-forget and swallowed at every level: a stale edge copy is bounded
  by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
  `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
  and the content-addressed trees (c_images, album*, clothes) keep the long
  TTL. `must-revalidate` is the point — the client now revalidates instead of
  replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
  purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
  Cloudflare was unreachable at write time.
2026-10-01 15:16:48 +02:00
openhands 7507c3b55c fix(deploy): detect the actually-live blue/green slot, stop nginx-sync clobbering the upstream
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m40s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m28s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m5s
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
  picks the one that really answers; the upstream file only serves as a
  fallback when zero or both slots respond. A stray 'docker compose up' (or a
  clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
  only seed it when missing, never overwrite what a deploy wrote. This is the
  root cause of tonight's 502: a nginx-sync run reset the snippet (written to
  green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
2026-09-28 23:38:22 +02:00
openhands 90b65c92a2 feat(proxy): sync Cloudflare ranges at nginx+Traefik, block IP spoofing
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m51s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 20s
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
  live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
  incoming CF-Connecting-IP, 403 any other peer that presents one
  (spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
  nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
2026-09-28 23:35:00 +02:00
openhands 7697728d07 feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.

- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
  blue/green upstream snippet; config backed by scripts/nginx-sync.sh
  (idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
  gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
  no-op unless Cloudflare is configured; scripts/cf-purge.sh and
  cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
  guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
2026-09-28 21:55:18 +02:00
Simo 9a2d73a6f6 feat(docker): provide verified-header proxy profiles for self-hosted clones
CI / check (push) Failing after 1m45s
CI / deploy (push) Skipped
CI / publish-container (push) Skipped
2026-09-13 20:15:10 +02:00