chore: upgrade to pnpm v12, vitest v5 and resolve deprecated subdependencies
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
This commit is contained in:
1 parent
f187cd70a9
commit
e4f83a8036
18 files changed
+1692
-1121
No files matched your search
@@ -685,6 +685,69 @@ For bulk-written tables (game-data JSON, imports) a containerized **`mariadb-tur
|
||||
|
||||
The CMS caches through `cached()` / `cachedQuery()` (`src/lib/cache.ts`): an in-memory fast path with a Redis (Valkey-compatible) fallback and single-flight protection against cache stampedes. Public API routes (home, leaderboard, online count, radio, photos, …) fill Redis keys per route; verify with `redis-cli DBSIZE`. With Cloudflare in front, static and `/imaging/*` responses are additionally cached edge-side (`cf-cache-status`), cutting origin load further.
|
||||
|
||||
#### nginx micro-cache (edge of the origin)
|
||||
|
||||
A short-lived nginx cache sits in front of the origin for an explicit whitelist of
|
||||
public API reads. It is configured in `/etc/nginx/sites-enabled/cms.conf` and its
|
||||
backends come from `/etc/nginx/snippets/cms_upstream_servers.conf`.
|
||||
|
||||
Why nginx and not the app: Next.js emits
|
||||
`private, no-cache, no-store, max-age=0, must-revalidate` for dynamic route
|
||||
handlers, and its response helper then drops the `Cache-Control` set by
|
||||
`publicCacheControl()`. Every one of the whitelisted routes was verified to be
|
||||
free of `auth()`, `cookies()`, `headers()` and request data, so the edge may
|
||||
cache them on their behalf.
|
||||
|
||||
| Class | Endpoints | TTL |
|
||||
| --- | --- | --- |
|
||||
| 1 | `/api/online`, `/api/online/count` | 10s |
|
||||
| 2 | `/api/photos`, `/api/leaderboard`, `/api/radio/current-dj`, `/api/radio/points/leaderboard` | 60s |
|
||||
| 3 | `/api/staff`, `/api/teams`, `/api/guilds`, `/api/shop`, `/api/shop/categories`, `/api/values`, `/api/values/categories`, `/api/values/[0-9]+` | 300s |
|
||||
|
||||
`/api/radio/shouts` is intentionally excluded: it is live chat.
|
||||
|
||||
Three details are easy to get wrong and are worth keeping intact:
|
||||
|
||||
- **`proxy_ignore_headers "Cache-Control" "Vary";` is what makes the cache work.**
|
||||
`proxy_hide_header` only strips the header from the response to the client; the
|
||||
cache module has already read the upstream's `no-store` by then and stores
|
||||
nothing. Ignore the upstream directive and let nginx emit a single
|
||||
`Cache-Control` of its own.
|
||||
- **`$upstream_status` is empty on a cache HIT** — it is only populated on a MISS.
|
||||
The header map therefore matches both `|2` and the empty-status form, so HITs
|
||||
get the same class as the MISS that stored them. A status-gated map alone sends
|
||||
every HIT down the `default` branch and quietly returns `private`.
|
||||
- **A session must never be advertised as `public`.** With a session cookie nginx
|
||||
already refuses to store the response (`proxy_no_cache`), but without the skip
|
||||
flag in the header map the response would still claim `public, s-maxage=...` and
|
||||
Cloudflare could share it. The map keys on
|
||||
`"$cms_cc_class|$cms_skip_cache|$upstream_status"`.
|
||||
|
||||
HTML is deliberately **not** cached: `src/app/(site)/page.tsx` performs `auth()`
|
||||
with a redirect to `/me` and reads `headers()` for a per-request CSP nonce.
|
||||
|
||||
### Files changed for stability and caching
|
||||
|
||||
| File | Change |
|
||||
| --- | --- |
|
||||
| `/etc/nginx/nginx.conf` | `worker_connections 4096`, `multi_accept on`, TLS 1.2/1.3 only, `proxy_cache_path` for the micro-cache |
|
||||
| `/etc/nginx/sites-enabled/cms.conf` | single correct `Cache-Control` per path, micro-cache on the whitelist, session/Authorization bypass, SSE block with buffering off, `upstream cms_app` + `proxy_next_upstream` |
|
||||
| `/etc/nginx/snippets/cms_upstream_servers.conf` | backend list, rewritten by `ci-deploy.sh` during a cutover |
|
||||
| `src/app/api/online/count/stream/route.ts` | `X-Accel-Buffering: no` |
|
||||
| `src/app/api/radio/stream/route.ts` | `X-Accel-Buffering: no` |
|
||||
| `src/app/api/admin/import/{audit,furni/route,furni/batch,furni/batch-regen,furni/repair-icons}/route.ts` | `X-Accel-Buffering: no` |
|
||||
| `src/app/api/admin/studio/nitro-cleanup{,/rebuild}/route.ts` | `X-Accel-Buffering: no` |
|
||||
| `scripts/ci-deploy.sh` | resource limits, blue/green cutover, `switch_upstream`, rollback split by cutover state |
|
||||
| `deploy.sh` | reduced to a wrapper around `ci-deploy.sh` |
|
||||
| `docker-compose.yml` | two replica services from a shared anchor, resource limits, port-aware healthcheck, dead `mariadb-turbo` removed |
|
||||
| `src/lib/services/cache-warmup.ts` | corrected a comment that described delays the code never had |
|
||||
| `src/lib/ci-deploy.test.ts`, `src/test/ci-deploy-harness.sh` | blue/green coverage; nginx paths are injectable so the in-place path stays tested |
|
||||
| `ecosystem.config.cjs` | removed (PM2 supervised no process, and nothing referenced it) |
|
||||
|
||||
> The nginx files live on the host only, not in this repository. Rebuilding the
|
||||
> host loses the cache configuration — keep a copy under version control if that
|
||||
> becomes a real risk.
|
||||
|
||||
---
|
||||
|
||||
## Cloudflare & Anti-DDoS Protection
|
||||
@@ -902,23 +965,98 @@ expire.
|
||||
|
||||
---
|
||||
|
||||
## Production Deployment (PM2)
|
||||
## Production Deployment (blue/green)
|
||||
|
||||
```bash
|
||||
pnpm build
|
||||
pm2 start pnpm --name "next" -- start
|
||||
pm2 save
|
||||
PM2 is **not** used in production. It is neither started nor supervised here; the
|
||||
CMS runs as a single Docker container per replica, started by
|
||||
`scripts/ci-deploy.sh`.
|
||||
|
||||
```
|
||||
git push main
|
||||
└─ Gitea Actions "deploy" job
|
||||
└─ bash scripts/ci-deploy.sh
|
||||
├─ build image epicnext-cms:<sha>
|
||||
├─ pnpm db:migrate
|
||||
├─ start candidate on the idle port (live replica keeps serving)
|
||||
├─ health gate + release-hash check + Playwright e2e
|
||||
├─ switch the nginx upstream, reload (cutover)
|
||||
└─ stop and retire the previous replica
|
||||
```
|
||||
|
||||
Restart after updates:
|
||||
### Why host ports instead of `--scale`
|
||||
|
||||
```bash
|
||||
git pull
|
||||
pnpm install
|
||||
pnpm build
|
||||
pm2 restart next
|
||||
The containers run with `network_mode: host`, so every replica binds a **host**
|
||||
port rather than sharing a published docker port. `docker compose up --scale
|
||||
cms=2` therefore cannot work here: the replicas would collide on 3002. Instead
|
||||
there are two fixed slots, and each release lands in the idle one:
|
||||
|
||||
| Slot | Host port | Container name |
|
||||
| --- | --- | --- |
|
||||
| A | 3002 | `epicnext-cms-app` |
|
||||
| B | 3003 | `epicnext-cms-green` |
|
||||
|
||||
`docker-compose.yml` defines both services (`cms`, and `cms-green` behind the
|
||||
`green` profile) from one shared YAML anchor so they cannot drift apart. Note
|
||||
that the CI deploy path deliberately uses `docker run` rather than compose, so
|
||||
the resource limits are declared in **both** places — limits that only existed
|
||||
in compose would never apply to a real release.
|
||||
|
||||
### The nginx upstream is the switch
|
||||
|
||||
nginx does not know about container names; it reads a plain list of backends from
|
||||
`/etc/nginx/snippets/cms_upstream_servers.conf`, which `ci-deploy.sh` rewrites
|
||||
during the cutover:
|
||||
|
||||
```nginx
|
||||
upstream cms_app {
|
||||
least_conn;
|
||||
keepalive 32;
|
||||
include /etc/nginx/snippets/cms_upstream_servers.conf;
|
||||
}
|
||||
```
|
||||
|
||||
```
|
||||
server 127.0.0.1:3002 max_fails=2 fail_timeout=10s;
|
||||
```
|
||||
|
||||
`switch_upstream()` writes the new backend, runs `nginx -t`, and only then
|
||||
reloads. If the test fails the previous file is restored untouched, so a typo can
|
||||
never take the site down.
|
||||
|
||||
### Failover between replicas
|
||||
|
||||
`max_fails`/`fail_timeout` mark a dead replica as unavailable, but only
|
||||
`proxy_next_upstream` actually retries the other one:
|
||||
|
||||
```nginx
|
||||
proxy_next_upstream error timeout invalid_header http_502 http_503 http_504;
|
||||
proxy_next_upstream_tries 2;
|
||||
proxy_next_upstream_timeout 10s;
|
||||
```
|
||||
|
||||
`non_idempotent` is deliberately **absent**, so `POST` is never replayed against
|
||||
the second replica — a retried import or write would otherwise apply twice. A
|
||||
stream that already emitted events is likewise not resumed elsewhere.
|
||||
|
||||
### Resource limits
|
||||
|
||||
`--memory=4g --memory-swap=5g --cpus=2 --pids-limit=512` per replica. The limits
|
||||
were previously removed ("Next may run unrestricted"); on a shared host that
|
||||
means a single leak can starve MariaDB, nginx and Traefik.
|
||||
|
||||
> **Watch this on the first heavy import.** The 4 GiB ceiling is not yet
|
||||
> exercised by a real bulk import. If the container is OOM-killed during a large
|
||||
> furniture import, raise `mem_limit` in `docker-compose.yml` *and* `mem_limit` /
|
||||
> `mem_swap_limit` in `scripts/ci-deploy.sh` together.
|
||||
|
||||
### Manual deploys
|
||||
|
||||
`./deploy.sh` is a thin wrapper around `scripts/ci-deploy.sh`, so a manual
|
||||
release and a CI release follow exactly the same path. The previous version
|
||||
killed whatever held port 3002 with `fuser -k` and ran `docker compose down`
|
||||
before anything new existed — guaranteed downtime on every failed build. The
|
||||
branch guard inside `ci-deploy.sh` only accepts `main`/`master`.
|
||||
|
||||
---
|
||||
|
||||
## Scripts
|
||||
|
||||
Reference in new issue
Block a user