fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s

The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.

- Evict least-recently-used instead, and raise the default budget to 2000
  (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
  only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
  lives on the entry, so one call site opting in protects every reader of that
  key. A failed background refresh keeps serving the last good value instead of
  falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
  Redis key and publishes a signal, so a value written by one process is no
  longer served stale by the others for the rest of its TTL. A failed Redis
  delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
  outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
  call, with a pub/sub signal to drop the local copy when it rotates. A Redis
  outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
  each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
  GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
  a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
  the old cap counted files and never removed anything while entries were
  fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
  runtime imaging cache.

Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.

3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
This commit is contained in:
openhands committed 2026-09-25 18:26:45 +02:00
1 parent f490fcc9da
commit 203399aab7
43 files changed
+1737 -212

No files matched your search

+119 -6
View File
@@ -55,13 +55,126 @@ describe("cached (memory-only, no Redis)", () => {
expect(fn).toHaveBeenCalledTimes(2);
});
it("evicts the oldest entry when the in-memory cache is full", async () => {
it("evicts the least recently used entry when the cache is full", async () => {
const fn = vi.fn(async () => 1);
for (let i = 0; i < 600; i++) {
await cached(`bulk-${i}`, 10_000, fn);
// Fill well past the budget (default 2 000 entries).
for (let i = 0; i < 2_100; i++) {
await cached(`bulk-${i}`, 60_000, fn);
}
// Key "bulk-0" was evicted (insertion order), so it must recompute.
await cached("bulk-0", 10_000, fn);
expect(fn).toHaveBeenCalledTimes(601);
// "bulk-0" is the least recently used, so it must have been evicted.
await cached("bulk-0", 60_000, fn);
expect(fn).toHaveBeenCalledTimes(2_101);
});
it("keeps a hot key alive while colder keys churn through the cache", async () => {
const hot = vi.fn(async () => "hot");
const cold = vi.fn(async () => "cold");
const hotKey = `hot-${Math.random()}`;
expect(await cached(hotKey, 60_000, hot)).toBe("hot");
// Every iteration reads the hot key first, then floods the cache with
// fresh one-off keys. Reading must count as using the entry, so the hot
// key survives even though it was inserted first by a long way.
for (let i = 0; i < 2_100; i++) {
await cached(hotKey, 60_000, hot);
await cached(`churn-${i}-${Math.random()}`, 60_000, cold);
}
expect(hot).toHaveBeenCalledTimes(1);
});
});
describe("cached (stale-while-revalidate)", () => {
it("serves the stale value and refreshes behind it", async () => {
let count = 0;
const fn = vi.fn(async () => ++count);
const key = `swr-${Math.random()}`;
const ttl = 20;
expect(await cached(key, ttl, fn, { staleMs: 10_000 })).toBe(1);
await new Promise((r) => setTimeout(r, 40));
// The expired entry is still served, so the caller never waits on the
// origin, and the refresh happens behind the response.
expect(await cached(key, ttl, fn, { staleMs: 10_000 })).toBe(1);
await vi.waitFor(() => expect(fn).toHaveBeenCalledTimes(2));
expect(await cached(key, ttl, fn, { staleMs: 10_000 })).toBe(2);
});
it("recomputes synchronously once the grace window has passed", async () => {
let count = 0;
const fn = vi.fn(async () => ++count);
const key = `swr-expiry-${Math.random()}`;
const ttl = 20;
expect(await cached(key, ttl, fn, { staleMs: 20 })).toBe(1);
await new Promise((r) => setTimeout(r, 80));
expect(await cached(key, ttl, fn, { staleMs: 20 })).toBe(2);
expect(fn).toHaveBeenCalledTimes(2);
});
it("keeps serving the stale value when a background refresh fails", async () => {
const key = `swr-fail-${Math.random()}`;
const fn = vi
.fn()
.mockResolvedValueOnce("first")
.mockRejectedValue(new Error("origin down"));
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
await new Promise((r) => setTimeout(r, 40));
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
await vi.waitFor(() => expect(fn).toHaveBeenCalledTimes(2));
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
});
it("does not cache the result of a refresh invalidated mid-flight", async () => {
const key = `swr-invalidate-${Math.random()}`;
let resolveSlow: (value: string) => void = () => {};
const slow = vi.fn(
() =>
new Promise<string>((resolve) => {
resolveSlow = resolve;
}),
);
const pending = cached(key, 10_000, slow, { staleMs: 10_000 });
// The refresh is in flight and an invalidation lands before it resolves.
invalidateMemory(key);
resolveSlow("computed-before-invalidation");
await pending;
// The outdated value must not have been written back, so the next read
// recomputes instead of serving what the invalidation just discarded.
const fresh = vi.fn(async () => "fresh");
expect(await cached(key, 10_000, fresh)).toBe("fresh");
expect(await cached(key, 10_000, fresh)).toBe("fresh");
expect(fresh).toHaveBeenCalledTimes(1);
});
it("recovers once the origin comes back after failed background refreshes", async () => {
const key = `swr-recover-${Math.random()}`;
const fn = vi
.fn()
.mockResolvedValueOnce("first")
.mockRejectedValueOnce(new Error("origin down"))
.mockRejectedValueOnce(new Error("origin down"))
.mockResolvedValue("recovered");
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
await new Promise((r) => setTimeout(r, 40));
// Each stale read kicks off one background refresh; while the origin keeps
// failing the caller keeps getting the last good value.
for (const expectedCalls of [2, 3]) {
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
await vi.waitFor(() => expect(fn).toHaveBeenCalledTimes(expectedCalls));
// Let the failed refresh settle so the next read starts a new one
// instead of joining the still-registered in-flight promise.
await new Promise((r) => setTimeout(r, 10));
}
// Once a refresh succeeds, the fresh value replaces the stale one.
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("first");
await vi.waitFor(() => expect(fn).toHaveBeenCalledTimes(4));
await vi.waitFor(async () =>
expect(await cached(key, 20, fn, { staleMs: 10_000 })).toBe("recovered"),
);
});
});