9bc600579a3f4030c36a3b0c236618f1af6cb92a
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2978483ea9 |
fix(assets): serve furniture icons from the gamedata tree, add stack-height SQL generator
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 2m3s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m47s
Commit the two changes that were live in the working tree but never recorded. The nginx change adds a location for /swf/dcr/hof_furni/icons/ that reads /var/www/Gamedata/icons/ off disk and falls back to the CMS bundle for anything missing. The CMS built its icon URLs from that path, but the icons actually live in the gamedata tree: of the 16,263 classnames in items_base, 15,338 are present there under the safe name, against 11,267 under public/swf/dcr/hof_furni/icons. The swf tree also keeps colour variants behind a literal "*" in the filename, so every request using the safe name missed, and hof.furni.url already points the Nitro client at the gamedata tree. A miss is a no-store 404, never a cached one — same reasoning as @gamedata_missing, since add_header without "always" says nothing about a 404 and Cloudflare would otherwise apply the zone TTL. scripts/generate-stack-height-sql.ts emits a SQL file that repairs items_base dimensions and interaction columns from the logic inside each bundle, which is where they actually live rather than in furnidata. It writes a file and touches no database, so the result is reviewed before it is run. |
||
|
|
6c3d81920e |
fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than the code. Each one had a signature that looked like a network or permissions problem and was actually a configuration or ordering bug. jobs-worker never ran `import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM evaluates a module's imports in source order and the first import reaches `@/env`, which validates process.env at import time. The ZodError on DATABASE_URL therefore fired before load-env ever executed, so the worker could only start from a shell that had already exported the configuration. Nothing supervised it either, so scheduled articles, catalog export, JAR and database backups, disk alerts and the ops health probe have all been dead; `cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and added deployment/systemd/cms-jobs-worker.service with Restart=always. The JAR backup additionally pointed at './emulator/Arcturus.jar', which does not exist and would go stale on the next emulator upgrade. resolveEmulatorJar now accepts a file, a directory or a wildcard and picks the newest JAR, the same way emulator.service picks its build, and reports an unresolvable path once instead of logging an opaque copyFile ENOENT every night. /api/health answered 200 with the database down The route documented this as intentional, and ci-deploy.sh worked around it by grepping the body for '"database":true'. The container healthcheck did not, so Docker reported containers healthy while every page 500'd. The status is now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis and the emulator deliberately do not fail the container — both have in-process fallbacks, so failing them would trade a slow site for an outage. The runtime had no V8 heap cap NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its heap from host memory (23.5 GB) while the container was limited to 4 GB, so the kernel OOM-killed the process mid-request — the same failure mode as the 14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit (v2 with a v1 fallback) and sets 70% of it, respecting an explicit override. Storage ownership was only repaired for one path ci-deploy.sh chowned storage/imaging and nothing else, so storage/catalog-git/hotel-status.json kept coming back root:root and /api/admin/catalog/status kept throwing EACCES. All eight writable storage paths are repaired now. The silent-failure mode is the reason this mattered: these writes sit inside try/catch, so a wrong owner looks like a slow page rather than an error. nginx: robots.txt was a guaranteed 404, and TLS never resumed `index index.html` without a `root` left every try_files resolving against /etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES on each stat and nginx logs a failed stat at crit, which is where 149 crit lines per scan came from. robots.txt answered from that same broken location, so crawlers were pointed at a file they could never read while sitemap.xml kept advertising it. Added `root`, proxied robots.txt to the CMS, added ssl_session_cache (there was no session resumption at all), and set Restart=on-failure in a systemd override, since the packaged unit ships Restart=no and nginx is the only thing serving the site. Verified against the running host: 3379 tests, typecheck and biome clean, nginx -t passes, health returns 200 with every check green, and the worker has run for hours at NRestarts=0 with a heartbeat refreshing each minute. |
||
|
|
dbaccd7cfc |
fix(gamedata): never cache a missing gamedata file
A missing gamedata file got no Cache-Control at all, because add_header without `always` only applies to 2xx/3xx. Cloudflare then fell back to the zone setting "Browser Cache TTL = 1 year", so the 404 came back as `max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's browser and at the edge. An icon requested while its import was still running stayed a 404 for the rest of the year, even after the file existed. That was the "some icons load, some don't" report. Give every gamedata location a named 404 handler that sends no-store, and split icons/ out as its own cache class: those files are rewritten under the same name (repair-icons, reimport), so an hourly must-revalidate keeps a repaired icon visible within the hour instead of days later. |
||
|
|
4e036b08d5 |
fix(proxy): drop request rate limiting from gamedata entirely
/gamedata/* is served straight from disk by nginx; no request hits the CMS backend or a database, so a request-rate limit protects nothing while costing players their icons. A room load fires hundreds of these files in one burst, which every limit turned into visible 503s. Removed the static zone from the gamedata locations. Traefik's epicnabbo-gamedata router likewise carries no rateLimit middleware. /client/ and /nitro-client/ keep theirs, and the main route keeps the 30r/s page budget plus the server-wide connection limit. Measured: 1000 icon requests fired fully in parallel now all return 200, while 200 parallel requests on / are still rejected. |
||
|
|
f0dcf440a7 |
fix(proxy): split rate limiting into page and static zones
The single server-scope limit_req (30r/s) treated a page load and a room load as the same thing. Loading a Nitro room fires several hundred gamedata icons in one burst, which that zone answered with 503s, so icons showed up late in the client. Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata and client asset locations, and apply the page-rate zone explicitly on the main route instead of at server scope. Connection limit stays server-wide. Measured: 900 icon requests in burst now all return 200, while 200 parallel requests on / are still rejected. |
||
|
|
4a1211a931 |
feat(proxy): add per-IP rate and connection limits
The edge had no limit_req/limit_conn at all, so a single client could flood the Next.js backend and the Nitro client with unbounded parallel requests. Traefik's logs already showed this: bursts of gamedata icon requests answered with 429. Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed on the real client IP, applied at server scope so both cached assets and proxied API routes share one budget. The burst is deliberately generous because the Nitro client fetches gamedata and icons in bursts when loading a room. |
||
|
|
a6cc3cafa9 |
fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached twice and nobody could reach the client. The catalog items loader kept its own 30s TTL copy of FurnitureData.json next to the mtime-validated cache in `furni-data.ts`. An import cleared only the second one, so the catalog table kept serving pre-import furnidata — empty descriptions and revisions — until the TTL ran out. The loader now reads through `readFurniData`, which revalidates on mtime+size and is reset by every write, so there is exactly one cache and it cannot go stale on its own. `invalidateFurniDataCache` and its single call site are gone with it. The client was worse: nginx served all of /gamedata/ with `max-age=604800`, and the `cms-gamedata` purge that would have fixed it hung off the catalog Git export, which is disabled in production. A freshly imported item was invisible in the client for up to seven days no matter how often you imported. - `writeFurniData` now purges the gamedata edge tag itself. One place covers import, batch, resync, regen, nitro-editor, translate and dedupe. It is fire-and-forget and swallowed at every level: a stale edge copy is bounded by the edge TTL, so a failed purge must never fail an import. - nginx splits /gamedata/ by how mutable the content is: config/ gets `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`, and the content-addressed trees (c_images, album*, clothes) keep the long TTL. `must-revalidate` is the point — the client now revalidates instead of replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the purge still reaches them. - A 30-minute safety-net purge in the jobs worker covers the case where Cloudflare was unreachable at write time. |
||
|
|
90b65c92a2 |
feat(proxy): sync Cloudflare ranges at nginx+Traefik, block IP spoofing
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m51s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 20s
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from live CF IPv4/IPv6 ranges plus Traefik bridge and loopback - nginx-cms.conf: forward real client IP only from trusted peers, strip incoming CF-Connecting-IP, 403 any other peer that presents one (spoof gate); direct game clients on :9443 stay unaffected - cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs - nginx-sync.sh: install the cloudflare-ips.conf snippet - cms_upstream_servers.conf: point default at the live green slot 3003 |
||
|
|
7697728d07 |
feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single Cache-Control owner per route: the app stays the source, nginx only manages headers, and Cloudflare stores the public API allowlist at the edge. - deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the blue/green upstream snippet; config backed by scripts/nginx-sync.sh (idempotent install + reload, --check/--force). - nginx serves Cache-Tag headers on the public allowlist (cms-public), gamedata, client and camera responses so the edge and purge stay in sync. - src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that no-op unless Cloudflare is configured; scripts/cf-purge.sh and cf-setup-cache.sh create and purge the cache rule. - src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls. - Purge hooks after catalog exports (public + gamedata) and on shop, team, guild, photo and rare-values edits; ci-deploy purges after each release. - src/proxy.ts excludes the imaging/images docs from the middleware matcher. |