Commit Graph
2 Commits
Author SHA1 Message Date
openhands f81b114b69 perf(cache): single-flight avatar renders, cacheable public reads, cheap row counts
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m30s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m7s
Three separate things that were each costing more than they needed to on the
hot path.

- Single-flight avatar renders. The disk cache was checked first and a miss
  went straight to the upstream, with nothing shared between callers, so a page
  requesting dozens of avatars at once turned N concurrent requests for one
  figure into N renders. A render is the most expensive operation this app
  does, and the duplication happened exactly when the cache had nothing to
  offer. Eight concurrent requests now cause one render instead of eight. The
  map lives on globalThis because Next can evaluate the module more than once
  per process, and two copies would each start their own render.

- Let public read-only routes be cached by a shared cache. Every JSON response
  was `cache-control: no-store`, so a CDN in front of the app could not answer
  any of it and every request reached the origin. publicCacheControl() opts a
  route in with s-maxage and stale-while-revalidate, using the same TTL as the
  server-side cache so the two layers cannot disagree. The default stays
  no-store: most routes here are personalised, admin-only or auth-dependent.
  /api/badges/leaderboard is deliberately left alone because it returns
  per-viewer rank entries to signed-in callers.

  Note this only takes effect once a cache rule exists for /api/* at the CDN, or
  the explicit `cache: "no-store"` is dropped from the client fetches (24 files
  do that today, including the /api/online poll). The headers alone are inert
  until one of those happens.

- Take the homepage row counts from the storage engine estimate instead of
  COUNT(*), which walks an index and gets slower as the tables grow. A missing
  or zero estimate falls back to the exact count rather than ever showing a
  wrong zero. The online count stays exact: it is an indexed read over a small
  subset and a few seconds of drift reads as broken rather than approximate.

The counters move into one module because the homepage and the boot warm-up
populate the same cache keys, so two implementations would race to write
different values into the same entry.

3223 tests pass.
2026-09-25 18:42:23 +02:00
openhands 203399aab7 fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s
The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.

- Evict least-recently-used instead, and raise the default budget to 2000
  (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
  only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
  lives on the entry, so one call site opting in protects every reader of that
  key. A failed background refresh keeps serving the last good value instead of
  falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
  Redis key and publishes a signal, so a value written by one process is no
  longer served stale by the others for the rest of its TTL. A failed Redis
  delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
  outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
  call, with a pub/sub signal to drop the local copy when it rotates. A Redis
  outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
  each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
  GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
  a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
  the old cap counted files and never removed anything while entries were
  fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
  runtime imaging cache.

Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.

3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
2026-09-25 18:26:45 +02:00