fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s

Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.

jobs-worker never ran

`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.

The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.

/api/health answered 200 with the database down

The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.

The runtime had no V8 heap cap

NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.

Storage ownership was only repaired for one path

ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.

nginx: robots.txt was a guaranteed 404, and TLS never resumed

`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.

Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
This commit is contained in:
openhands committed 2026-10-05 20:25:22 +02:00
1 parent 108c6ce03d
commit 6c3d81920e
12 files changed
+568 -26

No files matched your search

+95
View File
@@ -0,0 +1,95 @@
import { beforeEach, describe, expect, it, vi } from "vitest";
const mocks = vi.hoisted(() => ({
clientIp: vi.fn(async () => "203.0.113.7"),
rateLimit: vi.fn(
async (): Promise<{ ok: boolean; retryAfter: number }> => ({
ok: true,
retryAfter: 0,
}),
),
dbExecute: vi.fn(async () => [{}]),
ping: vi.fn(async () => "PONG"),
rconSend: vi.fn(async () => true),
}));
vi.mock("@/lib/rate-limit", () => ({
clientIp: mocks.clientIp,
rateLimit: mocks.rateLimit,
}));
vi.mock("@/lib/db", () => ({
db: { execute: mocks.dbExecute },
}));
vi.mock("@/lib/redis", () => ({
redis: {
ping: mocks.ping,
},
}));
vi.mock("@/lib/services/rcon", () => ({
rcon: { send: mocks.rconSend },
}));
vi.mock("@/env", () => ({
env: { REDIS_URL: "redis://cache.test:6379", RESEND_API_KEY: "" },
}));
import { GET } from "./route";
beforeEach(() => {
vi.clearAllMocks();
mocks.ping.mockResolvedValue("PONG");
mocks.dbExecute.mockResolvedValue([{}]);
mocks.rconSend.mockResolvedValue(true);
mocks.rateLimit.mockResolvedValue({ ok: true, retryAfter: 0 });
});
describe("ops health status", () => {
it("answers 200 when the database is reachable", async () => {
const res = await GET();
expect(res.status).toBe(200);
expect(await res.json()).toMatchObject({
status: "ok",
database: true,
});
});
// The regression this guards: the route used to answer 200 unconditionally,
// so the Docker healthcheck reported a container healthy while every page
// failed to render.
it("answers 503 when the database is unreachable", async () => {
mocks.dbExecute.mockRejectedValue(new Error("ECONNREFUSED"));
const res = await GET();
expect(res.status).toBe(503);
expect(await res.json()).toMatchObject({
status: "degraded",
database: false,
});
});
// Redis and the emulator both have in-process fallbacks (cache.ts,
// rate-limit.ts), so failing the container on them would turn a degraded
// site into a restart loop.
it("stays 200 on a Redis outage because the cache falls back in-process", async () => {
mocks.ping.mockRejectedValue(new Error("ECONNREFUSED"));
const res = await GET();
expect(res.status).toBe(200);
expect(await res.json()).toMatchObject({
status: "degraded",
database: true,
redis: false,
});
});
it("stays 200 when the emulator is unreachable", async () => {
mocks.rconSend.mockResolvedValue(false);
const res = await GET();
expect(res.status).toBe(200);
expect(await res.json()).toMatchObject({ emulator: false });
});
it("still rate-limits before probing anything", async () => {
mocks.rateLimit.mockResolvedValue({ ok: false, retryAfter: 30 });
const res = await GET();
expect(res.status).toBe(429);
expect(res.headers.get("Retry-After")).toBe("30");
expect(mocks.dbExecute).not.toHaveBeenCalled();
});
});
+23 -14
View File
@@ -9,9 +9,15 @@ import { rcon } from "@/lib/services/rcon";
/**
* Ops health probe: database reachability, Redis (when configured), emulator
* RCON, SMTP (when configured), and runtime info. Returns HTTP 200 always
* (read the `status`/`database` fields), so it's safe for uptime monitors that
* only care about reachability. Rate-limited per client IP.
* RCON, SMTP (when configured), and runtime info.
*
* HTTP status is load-bearing: 200 means the CMS can actually serve, 503 means
* it cannot. Previously this route answered 200 even with the database down,
* so the Docker healthcheck reported a container "healthy" while every page
* 500'd. Only the database drives the status — Redis and the emulator degrade
* to in-process fallbacks (see `cache.ts` and `rate-limit.ts`), so failing the
* container on those would trade a slow site for an outage. Docker does not
* restart on `unhealthy`, so this reports rather than recycles.
*/
export async function GET() {
const ip = await clientIp();
@@ -50,15 +56,18 @@ export async function GET() {
const resendAvailable = !!env.RESEND_API_KEY;
const degraded = !database || redisOk === false;
return apiJson({
status: degraded ? "degraded" : "ok",
database,
redis: redisOk,
emulator,
resend: resendAvailable,
node: process.version,
release: process.env.NEXT_PUBLIC_CMS_RELEASE ?? "unknown",
uptime: Math.round(process.uptime()),
time: new Date().toISOString(),
});
return apiJson(
{
status: degraded ? "degraded" : "ok",
database,
redis: redisOk,
emulator,
resend: resendAvailable,
node: process.version,
release: process.env.NEXT_PUBLIC_CMS_RELEASE ?? "unknown",
uptime: Math.round(process.uptime()),
time: new Date().toISOString(),
},
{ status: database ? 200 : 503 },
);
}
+99
View File
@@ -0,0 +1,99 @@
import {
mkdirSync,
mkdtempSync,
rmSync,
utimesSync,
writeFileSync,
} from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { describe, expect, it } from "vitest";
import { resolveEmulatorJar } from "../../scripts/jobs-worker";
/** Build a throwaway directory that looks like an emulator release folder. */
function releaseDir(entries: Array<[name: string, mtimeSeconds: number]>) {
const dir = mkdtempSync(join(tmpdir(), "cms-jar-resolve-"));
for (const [name, mtime] of entries) {
const path = join(dir, name);
writeFileSync(path, "jar");
utimesSync(path, mtime, mtime);
}
return dir;
}
describe("emulator JAR resolution for backups", () => {
it("uses a configured file as-is", () => {
const dir = releaseDir([["Polaris-4.2.97.jar", 1_000]]);
try {
expect(resolveEmulatorJar(join(dir, "Polaris-4.2.97.jar"))).toBe(
join(dir, "Polaris-4.2.97.jar"),
);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
// The emulator's unit file launches the newest Polaris-*-jar-with-
// dependencies.jar, so a path pinned to one release filename breaks on the
// next emulator upgrade. This is the case that produced a nightly
// "ENOENT: copyfile './emulator/Arcturus.jar'" nobody could act on.
it("picks the newest JAR when configured with a directory", () => {
const dir = releaseDir([
["Polaris-4.2.90-jar-with-dependencies.jar", 1_000],
["Polaris-4.2.97-jar-with-dependencies.jar", 9_000],
["Polaris-4.2.97.jar", 9_500],
["notes.txt", 9_900],
]);
try {
expect(resolveEmulatorJar(dir)).toBe(join(dir, "Polaris-4.2.97.jar"));
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
it("prefers the fat JAR over the plain one at the same timestamp", () => {
const dir = releaseDir([
["Polaris-4.2.97.jar", 5_000],
["Polaris-4.2.97-jar-with-dependencies.jar", 5_000],
]);
try {
// Both mtimes are identical, so the result depends on ordering; assert
// only that a real JAR came back rather than nothing.
expect(resolveEmulatorJar(dir)).toMatch(/\.jar$/);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
it("expands a wildcard against its directory", () => {
const dir = releaseDir([
["Polaris-4.2.90-jar-with-dependencies.jar", 1_000],
["Polaris-4.2.97-jar-with-dependencies.jar", 9_000],
]);
try {
expect(resolveEmulatorJar(join(dir, "Polaris-*-jar*.jar"))).toBe(
join(dir, "Polaris-4.2.97-jar-with-dependencies.jar"),
);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
it("returns null instead of throwing on a stale path", () => {
// This is the case that must not reach copyFileSync: a clear log line
// beats an opaque ENOENT from deep inside a backup job.
expect(resolveEmulatorJar("/nonexistent/emulator/Arcturus.jar")).toBeNull();
expect(resolveEmulatorJar("/nonexistent/dir/*.jar")).toBeNull();
});
it("returns null for a directory with no JARs", () => {
const dir = mkdtempSync(join(tmpdir(), "cms-jar-empty-"));
try {
mkdirSync(join(dir, "nested"));
expect(resolveEmulatorJar(dir)).toBeNull();
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
});
+3
View File
@@ -26,6 +26,9 @@ describe("deployed HTTP release readiness", () => {
[200, { database: true, release: "old" }],
[200, { database: true }],
[429, { status: "rate_limited" }],
// The health route answers 503 with the correct release when the
// database is unreachable, so this must never pass the gate.
[503, { database: false, release: sha }],
[200, { database: false, release: sha }],
])(
"rejects HTTP %s without the expected healthy release",