fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s
The in-process cache was a FIFO of 500 entries that was never touched on a read, so a key polled on every request could be evicted by an unrelated burst of dynamic keys. That looked exactly like the cache being cleared at random, and it is what made the site fall back to the database unpredictably. - Evict least-recently-used instead, and raise the default budget to 2000 (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key only leaves when a hotter one takes its place. - Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window lives on the entry, so one call site opting in protects every reader of that key. A failed background refresh keeps serving the last good value instead of falling through to the origin, and is reported once rather than per read. - Invalidate across processes. invalidateKey() now clears memory, deletes the Redis key and publishes a signal, so a value written by one process is no longer served stale by the others for the rest of its TTL. A failed Redis delete no longer skips the broadcast. - Guard against a refresh that started before an invalidation writing its outdated result back into the cache. - Read the news revision at most once a second per process instead of on every call, with a pub/sub signal to drop the local copy when it rotates. A Redis outage now degrades to the in-process cache rather than to no cache at all. - Warm the hot public keys on boot, so the first visitors after a deploy do not each pay for a miss. - Count hits, misses, stale serves, errors and evictions per key, exposed at GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and a dead origin all look identical from the outside. - Enforce the imaging cache budget for real: records are .img/.json pairs, so the old cap counted files and never removed anything while entries were fresh. Sweeps are throttled per directory and prune to a low-water mark. - Cap the JWT version map, and stop per-test scratch roots from littering the runtime imaging cache. Public read-only endpoints get grace windows; admin, account and auth data deliberately stays fresh. Redis TTLs get a little jitter so keys written together no longer expire together. 3209 tests pass. next build could not be verified on this host: the optimized build is OOM-killed before prerender, so this has not run in a real Next runtime yet.
This commit is contained in:
1 parent
f490fcc9da
commit
203399aab7
43 files changed
+1737
-212
No files matched your search
+81
-8
@@ -10,7 +10,7 @@ import {
|
||||
unlink,
|
||||
writeFile,
|
||||
} from "node:fs/promises";
|
||||
import { join } from "node:path";
|
||||
import { extname, join } from "node:path";
|
||||
|
||||
/**
|
||||
* Persistent disk cache for imaging renders (avatars, badges).
|
||||
@@ -25,11 +25,30 @@ import { join } from "node:path";
|
||||
* the directory cannot grow without bound.
|
||||
*/
|
||||
|
||||
const MAX_ENTRIES = 20_000;
|
||||
function maxEntries(): number {
|
||||
const raw = (process.env.IMAGING_CACHE_MAX_ENTRIES || "").trim();
|
||||
const parsed = Number.parseInt(raw, 10);
|
||||
return Number.isSafeInteger(parsed) && parsed > 0 ? parsed : 20_000;
|
||||
}
|
||||
|
||||
const MAX_AGE_MS = 30 * 24 * 60 * 60 * 1000;
|
||||
|
||||
// Pruning reads and stats every file in the directory, so running it on every
|
||||
// write costs thousands of syscalls per rendered avatar. A write only needs the
|
||||
// sweep to have happened "recently enough" — the budget is enforced over hours,
|
||||
// not per request, so the exact moment does not matter.
|
||||
const PRUNE_INTERVAL_MS = 5 * 60 * 1000;
|
||||
|
||||
// Prune down to this fraction of the budget, so the sweep only runs again after
|
||||
// a meaningful amount of new renders rather than on the very next write.
|
||||
const PRUNE_TARGET_RATIO = 0.9;
|
||||
|
||||
const IMG_ROOT_DEFAULT = join(process.cwd(), "storage", "imaging");
|
||||
|
||||
// Keyed per directory so throttling the avatars sweep does not also throttle the
|
||||
// badges one.
|
||||
const lastPruneAt = new Map<string, number>();
|
||||
|
||||
function imagingCacheDir(kind: "avatars" | "badges"): string {
|
||||
// Unit tests point this at a scratch root so they never read or pollute
|
||||
// the runtime cache.
|
||||
@@ -98,21 +117,75 @@ export async function writeImagingCache(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Keep the cache inside its budget.
|
||||
*
|
||||
* Records are stored as an `.img`/`.json` pair, so the file count is twice the
|
||||
* render count. Age alone is not enough to bound the directory: a burst of
|
||||
* popular figures can push it far past the budget while every file is still
|
||||
* fresh, and then nothing is ever removed. So expired files go first, and if that
|
||||
* is not enough the oldest remaining files go too, down to a low-water mark so
|
||||
* the next sweep is not immediately due again.
|
||||
*/
|
||||
async function pruneImagingCache(directory: string): Promise<void> {
|
||||
try {
|
||||
const now = Date.now();
|
||||
const lastRun = lastPruneAt.get(directory) ?? 0;
|
||||
if (now - lastRun < PRUNE_INTERVAL_MS) return;
|
||||
lastPruneAt.set(directory, now);
|
||||
|
||||
const names = await readdir(directory);
|
||||
if (names.length <= MAX_ENTRIES) return;
|
||||
const cutoffMs = Date.now() - MAX_AGE_MS;
|
||||
// Group the two halves of each record so a pair is never left orphaned.
|
||||
const records = new Map<string, string[]>();
|
||||
for (const name of names) {
|
||||
const base = name.slice(0, name.length - extname(name).length);
|
||||
const files = records.get(base);
|
||||
if (files) files.push(name);
|
||||
else records.set(base, [name]);
|
||||
}
|
||||
const budget = maxEntries();
|
||||
if (records.size <= budget) return;
|
||||
|
||||
const cutoffMs = now - MAX_AGE_MS;
|
||||
const survivors: { files: string[]; mtimeMs: number }[] = [];
|
||||
let live = 0;
|
||||
for (const files of records.values()) {
|
||||
try {
|
||||
const file = join(directory, name);
|
||||
const info = await stat(file);
|
||||
if (info.mtimeMs < cutoffMs) await unlink(file);
|
||||
const info = await stat(join(directory, files[0] as string));
|
||||
if (info.mtimeMs < cutoffMs) {
|
||||
await removeRecord(directory, files);
|
||||
continue;
|
||||
}
|
||||
survivors.push({ files, mtimeMs: info.mtimeMs });
|
||||
live++;
|
||||
} catch {
|
||||
// Skip entries that disappeared between listing and unlink.
|
||||
// Skip records that disappeared between listing and stat.
|
||||
}
|
||||
}
|
||||
|
||||
const target = Math.floor(budget * PRUNE_TARGET_RATIO);
|
||||
if (live <= target) return;
|
||||
survivors.sort((a, b) => a.mtimeMs - b.mtimeMs);
|
||||
for (const record of survivors) {
|
||||
if (live <= target) break;
|
||||
try {
|
||||
await removeRecord(directory, record.files);
|
||||
live--;
|
||||
} catch {
|
||||
// Already gone.
|
||||
}
|
||||
}
|
||||
} catch {
|
||||
// Nothing to prune or directory missing.
|
||||
}
|
||||
}
|
||||
|
||||
async function removeRecord(directory: string, files: string[]): Promise<void> {
|
||||
for (const name of files) {
|
||||
try {
|
||||
await unlink(join(directory, name));
|
||||
} catch {
|
||||
// Already gone.
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in new issue
Block a user