fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s

The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.

- Evict least-recently-used instead, and raise the default budget to 2000
  (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
  only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
  lives on the entry, so one call site opting in protects every reader of that
  key. A failed background refresh keeps serving the last good value instead of
  falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
  Redis key and publishes a signal, so a value written by one process is no
  longer served stale by the others for the rest of its TTL. A failed Redis
  delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
  outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
  call, with a pub/sub signal to drop the local copy when it rotates. A Redis
  outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
  each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
  GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
  a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
  the old cap counted files and never removed anything while entries were
  fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
  runtime imaging cache.

Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.

3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
This commit is contained in:
openhands committed 2026-09-25 18:26:45 +02:00
1 parent f490fcc9da
commit 203399aab7
43 files changed
+1737 -212

No files matched your search

+203 -29
View File
@@ -1,46 +1,143 @@
import "server-only";
import { onCacheInvalidated } from "@/lib/cache-invalidation";
import { recordCacheOutcome } from "@/lib/cache-stats";
import { logger } from "@/lib/logger";
import { redis } from "@/lib/redis";
type CacheEntry<T> = { data: T; expiresAt: number };
const memory = new Map<string, CacheEntry<unknown>>();
type CacheEntry = {
data: unknown;
/** Fresh until this timestamp. */
expiresAt: number;
/** Stale values stay servable until this timestamp (stale-while-revalidate). */
staleUntil: number;
};
// Cap the in-process map so dynamic keys (leaderboard currencies, article
// slugs, …) can never grow it without bound. Oldest entries are evicted.
const MAX_MEMORY_ENTRIES = 500;
export interface CachedOptions {
/**
* How long past the TTL the last value may still be served while it is
* refreshed in the background. `0` (the default) recomputes synchronously on
* expiry, which is what callers that need an immediately-fresh value want.
* A small grace turns a TTL boundary from a blocking origin call into a
* non-blocking one, so a mass expiry can never stall requests.
*
* The window is stored on the entry, not on the caller, so a key only needs
* *one* call site to opt in: every other reader of that key benefits from the
* grace window too. That matters for hot shared keys like `online_count`,
* which are read from a dozen places — opt in once, centrally.
*/
staleMs?: number;
}
const memory = new Map<string, CacheEntry>();
function readPositiveInt(name: string, fallback: number): number {
const raw = (process.env[name] || "").trim();
if (!raw) return fallback;
const parsed = Number.parseInt(raw, 10);
return Number.isSafeInteger(parsed) && parsed > 0 ? parsed : fallback;
}
// In-process budget. Entries are evicted least-recently-*used* (see getMemory),
// so a hot key only leaves when a hotter one takes its place. The default is
// generous because the layer exists to keep origin calls off the hot path.
const MAX_MEMORY_ENTRIES = readPositiveInt("CACHE_MEMORY_MAX_ENTRIES", 2_000);
// Single-flight: a key being (re)computed is awaited by concurrent callers
// instead of each starting its own `fn()` (cache-stampede protection).
const inFlight = new Map<string, Promise<unknown>>();
// Bumped by every invalidation. A refresh that started before the bump must not
// write its (now outdated) result back into the cache, or an invalidation
// would be silently undone by a refresh that was already in flight.
const generations = new Map<string, number>();
const MAX_GENERATION_ENTRIES = MAX_MEMORY_ENTRIES * 2;
// Keys whose most recent background refresh failed, so a persistently broken
// origin reports once instead of on every read of a hot key.
const reportedFailures = new Set<string>();
/** Drop a key from the in-process cache (used when an upstream value changes). */
export function invalidateMemory(key: string): void {
generations.set(key, (generations.get(key) ?? 0) + 1);
memory.delete(key);
inFlight.delete(key);
reportedFailures.delete(key);
// A generation only matters to a refresh that is still running; once nothing
// is in flight and the value is gone, the counter is dead weight.
if (generations.size > MAX_GENERATION_ENTRIES) {
for (const key of generations.keys()) {
if (!memory.has(key) && !inFlight.has(key)) generations.delete(key);
if (generations.size <= MAX_GENERATION_ENTRIES) break;
}
}
}
function pruneExpired(now: number): void {
for (const [key, entry] of memory) {
if (entry.expiresAt <= now) memory.delete(key);
if (entry.staleUntil <= now) memory.delete(key);
}
}
function setMemory<T>(key: string, entry: CacheEntry<T>): void {
function setMemory(key: string, entry: CacheEntry): void {
if (memory.has(key)) {
memory.set(key, entry);
return;
}
// Reclaim before evicting so a full map of long-lived entries never has to
// drop a live one to make room.
if (memory.size >= MAX_MEMORY_ENTRIES) {
// Drop expired entries first, then evict the oldest (insertion order).
pruneExpired(Date.now());
if (memory.size >= MAX_MEMORY_ENTRIES) {
const oldest = memory.keys().next().value;
if (oldest !== undefined) memory.delete(oldest);
// Map iteration order is insertion order and getMemory() re-inserts
// on every read, so the first key is the least recently used.
const lru = memory.keys().next().value;
if (lru !== undefined) {
memory.delete(lru);
recordCacheOutcome(lru, "evicted");
}
}
}
memory.set(key, entry);
}
/** Current in-process budget usage, for the admin cache report. */
export function cacheBudget(): { entries: number; limit: number } {
return { entries: memory.size, limit: MAX_MEMORY_ENTRIES };
}
/**
* Read a key and mark it as most recently used. Re-inserting is what makes the
* eviction above true LRU: without it, a key that is read on every single
* request still sits at the front of the insertion order and gets evicted by an
* unrelated burst of dynamic keys, which looks exactly like random clearing.
*/
function getMemory(key: string): CacheEntry | undefined {
const entry = memory.get(key);
if (entry === undefined) return undefined;
memory.delete(key);
memory.set(key, entry);
return entry;
}
/** `setex` with a little jitter so keys written together do not expire together. */
function redisTtlSeconds(ttlSec: number): number {
const jitter = Math.min(5, Math.floor(ttlSec * 0.1));
return ttlSec + Math.floor(Math.random() * (jitter + 1));
}
// A wrong REDIS_URL or a Redis outage does not fail visibly: the cache keeps
// answering from memory and every instance quietly stops sharing. Say so once.
let warnedSharedCacheDown = false;
function warnIfSharedCacheDown(): void {
if (redis && redis.status !== "end") return;
if (warnedSharedCacheDown || process.env.NODE_ENV !== "production") return;
warnedSharedCacheDown = true;
logger.error(
"[cache] Shared (Redis) caching is unavailable. Entries are per-process only, so they are dropped on restart and not shared between instances. Check REDIS_URL and that Redis is reachable.",
);
}
/**
* Redis-first cached query with an in-memory fallback.
* Use for read-heavy endpoints polled by the browser (online count, etc.).
@@ -49,47 +146,103 @@ export async function cached<T>(
key: string,
ttlMs: number,
fn: () => Promise<T>,
options: CachedOptions = {},
): Promise<T> {
const ttlSec = Math.ceil(ttlMs / 1000);
// In-memory fast path — served first so repeated reads within a TTL window
// don't each pay a Redis round-trip (Redis is still the shared fallback).
const staleMs = Math.max(0, options.staleMs ?? 0);
// `existing &&` short-circuits so Date.now() is never evaluated during
// prerender when the map is empty (keeps `next build` prerendering clean).
const existing = memory.get(key);
if (existing && existing.expiresAt > Date.now()) {
return existing.data as T;
const existing = getMemory(key);
if (existing) {
const now = Date.now();
if (existing.expiresAt > now) {
recordCacheOutcome(key, "hit");
return existing.data as T;
}
if (existing.staleUntil > now) {
recordCacheOutcome(key, "stale");
// Serve the last good value and refresh behind it: the caller never
// waits on the origin, and concurrent readers share one refresh.
if (!inFlight.has(key)) {
void refresh(key, ttlMs, staleMs, fn).catch((error: unknown) => {
// A failed refresh keeps the stale entry until its grace window
// runs out, then the next read recomputes synchronously. Never
// let it surface as an unhandled rejection, and never log it per
// read: a broken origin would otherwise flood the log from a
// single hot key.
if (reportedFailures.has(key)) return;
reportedFailures.add(key);
logger.warn("[cache] background refresh failed", {
key,
error: String(error),
});
});
}
return existing.data as T;
}
memory.delete(key);
}
// Single-flight: a concurrent request already recomputing this key.
const pending = inFlight.get(key);
if (pending) return (await pending) as T;
if (pending) {
recordCacheOutcome(key, "miss");
return (await pending) as T;
}
recordCacheOutcome(key, "miss");
return refresh(key, ttlMs, staleMs, fn);
}
/**
* Recompute `key` and publish the result. Runs behind a single-flight guard, so
* concurrent misses (and background refreshes) share one origin call.
*/
function refresh<T>(
key: string,
ttlMs: number,
staleMs: number,
fn: () => Promise<T>,
): Promise<T> {
const generation = generations.get(key) ?? 0;
const compute = (async (): Promise<T> => {
// Redis path (shared across instances).
if (redis && redis.status !== "end") {
try {
const cached = await redis.get(key);
if (cached !== null && cached !== undefined) {
const data = JSON.parse(cached) as T;
setMemory(key, { data, expiresAt: Date.now() + ttlMs });
const raw = await redis.get(key);
if (raw !== null && raw !== undefined) {
const data = JSON.parse(raw) as T;
commit(key, data, ttlMs, staleMs, generation);
return data;
}
} catch {
/* fall through to fn */
}
} else {
warnIfSharedCacheDown();
}
const data = await fn();
const data = await fn().catch((error: unknown) => {
recordCacheOutcome(key, "error");
throw error;
});
if (redis && redis.status !== "end") {
try {
await redis.setex(key, ttlSec, JSON.stringify(data));
} catch {
/* non-critical: memory cache still works */
// An invalidation landed while `fn()` was running: keep the value out of
// the cache so the next read recomputes instead of resurrecting stale data.
if ((generations.get(key) ?? 0) === generation) {
if (redis && redis.status !== "end") {
try {
await redis.setex(
key,
redisTtlSeconds(Math.ceil(ttlMs / 1000)),
JSON.stringify(data),
);
} catch {
/* non-critical: memory cache still works */
}
}
reportedFailures.delete(key);
commit(key, data, ttlMs, staleMs, generation);
}
setMemory(key, { data, expiresAt: Date.now() + ttlMs });
return data;
})().finally(() => {
@@ -99,3 +252,24 @@ export async function cached<T>(
inFlight.set(key, compute);
return compute;
}
function commit(
key: string,
data: unknown,
ttlMs: number,
staleMs: number,
generation: number,
): void {
if ((generations.get(key) ?? 0) !== generation) return;
const now = Date.now();
setMemory(key, {
data,
expiresAt: now + ttlMs,
staleUntil: now + ttlMs + staleMs,
});
}
// Another process (the jobs worker, a second instance) invalidating a key must
// drop this process's copy too, otherwise a value stays stale here for the full
// TTL after everyone else already saw the new one.
onCacheInvalidated(invalidateMemory);