Commit Graph
4 Commits
Author SHA1 Message Date
openhands 6c3d81920e fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.

jobs-worker never ran

`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.

The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.

/api/health answered 200 with the database down

The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.

The runtime had no V8 heap cap

NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.

Storage ownership was only repaired for one path

ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.

nginx: robots.txt was a guaranteed 404, and TLS never resumed

`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.

Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
2026-10-05 20:25:22 +02:00
openhands 7507c3b55c fix(deploy): detect the actually-live blue/green slot, stop nginx-sync clobbering the upstream
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m40s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m28s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m5s
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and
  picks the one that really answers; the upstream file only serves as a
  fallback when zero or both slots respond. A stray 'docker compose up' (or a
  clobbered snippet) can no longer derail the next deploy's cutover.
- nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh;
  only seed it when missing, never overwrite what a deploy wrote. This is the
  root cause of tonight's 502: a nginx-sync run reset the snippet (written to
  green:3003 by the last cutover) back to the dead slot A:3002.
- cms_upstream_servers.conf: restore the fresh-host seed default to slot A.
2026-09-28 23:38:22 +02:00
openhands 90b65c92a2 feat(proxy): sync Cloudflare ranges at nginx+Traefik, block IP spoofing
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m51s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 20s
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from
  live CF IPv4/IPv6 ranges plus Traefik bridge and loopback
- nginx-cms.conf: forward real client IP only from trusted peers, strip
  incoming CF-Connecting-IP, 403 any other peer that presents one
  (spoof gate); direct game clients on :9443 stay unaffected
- cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the
  nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs
- nginx-sync.sh: install the cloudflare-ips.conf snippet
- cms_upstream_servers.conf: point default at the live green slot 3003
2026-09-28 23:35:00 +02:00
openhands 7697728d07 feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.

- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
  blue/green upstream snippet; config backed by scripts/nginx-sync.sh
  (idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
  gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
  no-op unless Cloudflare is configured; scripts/cf-purge.sh and
  cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
  guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
2026-09-28 21:55:18 +02:00