chore: upgrade to pnpm v12, vitest v5 and resolve deprecated subdependencies
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped

This commit is contained in:
openhands committed 2026-09-27 19:42:34 +02:00
1 parent f187cd70a9
commit e4f83a8036
18 files changed
+1692 -1121

No files matched your search

+149 -11
View File
@@ -685,6 +685,69 @@ For bulk-written tables (game-data JSON, imports) a containerized **`mariadb-tur
The CMS caches through `cached()` / `cachedQuery()` (`src/lib/cache.ts`): an in-memory fast path with a Redis (Valkey-compatible) fallback and single-flight protection against cache stampedes. Public API routes (home, leaderboard, online count, radio, photos, …) fill Redis keys per route; verify with `redis-cli DBSIZE`. With Cloudflare in front, static and `/imaging/*` responses are additionally cached edge-side (`cf-cache-status`), cutting origin load further.
#### nginx micro-cache (edge of the origin)
A short-lived nginx cache sits in front of the origin for an explicit whitelist of
public API reads. It is configured in `/etc/nginx/sites-enabled/cms.conf` and its
backends come from `/etc/nginx/snippets/cms_upstream_servers.conf`.
Why nginx and not the app: Next.js emits
`private, no-cache, no-store, max-age=0, must-revalidate` for dynamic route
handlers, and its response helper then drops the `Cache-Control` set by
`publicCacheControl()`. Every one of the whitelisted routes was verified to be
free of `auth()`, `cookies()`, `headers()` and request data, so the edge may
cache them on their behalf.
| Class | Endpoints | TTL |
| --- | --- | --- |
| 1 | `/api/online`, `/api/online/count` | 10s |
| 2 | `/api/photos`, `/api/leaderboard`, `/api/radio/current-dj`, `/api/radio/points/leaderboard` | 60s |
| 3 | `/api/staff`, `/api/teams`, `/api/guilds`, `/api/shop`, `/api/shop/categories`, `/api/values`, `/api/values/categories`, `/api/values/[0-9]+` | 300s |
`/api/radio/shouts` is intentionally excluded: it is live chat.
Three details are easy to get wrong and are worth keeping intact:
- **`proxy_ignore_headers "Cache-Control" "Vary";` is what makes the cache work.**
`proxy_hide_header` only strips the header from the response to the client; the
cache module has already read the upstream's `no-store` by then and stores
nothing. Ignore the upstream directive and let nginx emit a single
`Cache-Control` of its own.
- **`$upstream_status` is empty on a cache HIT** — it is only populated on a MISS.
The header map therefore matches both `|2` and the empty-status form, so HITs
get the same class as the MISS that stored them. A status-gated map alone sends
every HIT down the `default` branch and quietly returns `private`.
- **A session must never be advertised as `public`.** With a session cookie nginx
already refuses to store the response (`proxy_no_cache`), but without the skip
flag in the header map the response would still claim `public, s-maxage=...` and
Cloudflare could share it. The map keys on
`"$cms_cc_class|$cms_skip_cache|$upstream_status"`.
HTML is deliberately **not** cached: `src/app/(site)/page.tsx` performs `auth()`
with a redirect to `/me` and reads `headers()` for a per-request CSP nonce.
### Files changed for stability and caching
| File | Change |
| --- | --- |
| `/etc/nginx/nginx.conf` | `worker_connections 4096`, `multi_accept on`, TLS 1.2/1.3 only, `proxy_cache_path` for the micro-cache |
| `/etc/nginx/sites-enabled/cms.conf` | single correct `Cache-Control` per path, micro-cache on the whitelist, session/Authorization bypass, SSE block with buffering off, `upstream cms_app` + `proxy_next_upstream` |
| `/etc/nginx/snippets/cms_upstream_servers.conf` | backend list, rewritten by `ci-deploy.sh` during a cutover |
| `src/app/api/online/count/stream/route.ts` | `X-Accel-Buffering: no` |
| `src/app/api/radio/stream/route.ts` | `X-Accel-Buffering: no` |
| `src/app/api/admin/import/{audit,furni/route,furni/batch,furni/batch-regen,furni/repair-icons}/route.ts` | `X-Accel-Buffering: no` |
| `src/app/api/admin/studio/nitro-cleanup{,/rebuild}/route.ts` | `X-Accel-Buffering: no` |
| `scripts/ci-deploy.sh` | resource limits, blue/green cutover, `switch_upstream`, rollback split by cutover state |
| `deploy.sh` | reduced to a wrapper around `ci-deploy.sh` |
| `docker-compose.yml` | two replica services from a shared anchor, resource limits, port-aware healthcheck, dead `mariadb-turbo` removed |
| `src/lib/services/cache-warmup.ts` | corrected a comment that described delays the code never had |
| `src/lib/ci-deploy.test.ts`, `src/test/ci-deploy-harness.sh` | blue/green coverage; nginx paths are injectable so the in-place path stays tested |
| `ecosystem.config.cjs` | removed (PM2 supervised no process, and nothing referenced it) |
> The nginx files live on the host only, not in this repository. Rebuilding the
> host loses the cache configuration — keep a copy under version control if that
> becomes a real risk.
---
## Cloudflare & Anti-DDoS Protection
@@ -902,23 +965,98 @@ expire.
---
## Production Deployment (PM2)
## Production Deployment (blue/green)
```bash
pnpm build
pm2 start pnpm --name "next" -- start
pm2 save
PM2 is **not** used in production. It is neither started nor supervised here; the
CMS runs as a single Docker container per replica, started by
`scripts/ci-deploy.sh`.
```
git push main
└─ Gitea Actions "deploy" job
└─ bash scripts/ci-deploy.sh
├─ build image epicnext-cms:<sha>
├─ pnpm db:migrate
├─ start candidate on the idle port (live replica keeps serving)
├─ health gate + release-hash check + Playwright e2e
├─ switch the nginx upstream, reload (cutover)
└─ stop and retire the previous replica
```
Restart after updates:
### Why host ports instead of `--scale`
```bash
git pull
pnpm install
pnpm build
pm2 restart next
The containers run with `network_mode: host`, so every replica binds a **host**
port rather than sharing a published docker port. `docker compose up --scale
cms=2` therefore cannot work here: the replicas would collide on 3002. Instead
there are two fixed slots, and each release lands in the idle one:
| Slot | Host port | Container name |
| --- | --- | --- |
| A | 3002 | `epicnext-cms-app` |
| B | 3003 | `epicnext-cms-green` |
`docker-compose.yml` defines both services (`cms`, and `cms-green` behind the
`green` profile) from one shared YAML anchor so they cannot drift apart. Note
that the CI deploy path deliberately uses `docker run` rather than compose, so
the resource limits are declared in **both** places — limits that only existed
in compose would never apply to a real release.
### The nginx upstream is the switch
nginx does not know about container names; it reads a plain list of backends from
`/etc/nginx/snippets/cms_upstream_servers.conf`, which `ci-deploy.sh` rewrites
during the cutover:
```nginx
upstream cms_app {
least_conn;
keepalive 32;
include /etc/nginx/snippets/cms_upstream_servers.conf;
}
```
```
server 127.0.0.1:3002 max_fails=2 fail_timeout=10s;
```
`switch_upstream()` writes the new backend, runs `nginx -t`, and only then
reloads. If the test fails the previous file is restored untouched, so a typo can
never take the site down.
### Failover between replicas
`max_fails`/`fail_timeout` mark a dead replica as unavailable, but only
`proxy_next_upstream` actually retries the other one:
```nginx
proxy_next_upstream error timeout invalid_header http_502 http_503 http_504;
proxy_next_upstream_tries 2;
proxy_next_upstream_timeout 10s;
```
`non_idempotent` is deliberately **absent**, so `POST` is never replayed against
the second replica — a retried import or write would otherwise apply twice. A
stream that already emitted events is likewise not resumed elsewhere.
### Resource limits
`--memory=4g --memory-swap=5g --cpus=2 --pids-limit=512` per replica. The limits
were previously removed ("Next may run unrestricted"); on a shared host that
means a single leak can starve MariaDB, nginx and Traefik.
> **Watch this on the first heavy import.** The 4 GiB ceiling is not yet
> exercised by a real bulk import. If the container is OOM-killed during a large
> furniture import, raise `mem_limit` in `docker-compose.yml` *and* `mem_limit` /
> `mem_swap_limit` in `scripts/ci-deploy.sh` together.
### Manual deploys
`./deploy.sh` is a thin wrapper around `scripts/ci-deploy.sh`, so a manual
release and a CI release follow exactly the same path. The previous version
killed whatever held port 3002 with `fuser -k` and ran `docker compose down`
before anything new existed — guaranteed downtime on every failed build. The
branch guard inside `ci-deploy.sh` only accepts `main`/`master`.
---
## Scripts