Commit Graph
2 Commits
Author SHA1 Message Date
openhands 155bf750c3 fix(cache): bound grace windows, cap render queues, and drop the useless estimate
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m55s
CI / tests-unit (push) Failing after 2m14s
CI / tests-ui (push) Failing after 36m38s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Follow-up to f81b114b, addressing the three ways that commit could make things
worse rather than better. All three were verified against the real database or
by breaking the test and watching it fail.

- The grace window is now capped at 120s. A window is a cushion for the TTL
  boundary, not a second TTL, but the call sites treated it as the latter: the
  5 min values/staff routes and the 10 min teams route asked for a window as
  long as or longer than their own TTL, so a single large staleMs silently
  doubled how far behind a value could be served. Nothing marked those as
  unsafe, because nothing looked wrong. The cap lives in the cache rather than
  at the call sites so no future route can reintroduce it. Routes that asked
  for less than 120s (the 10s online poll, the 20s news cache) are unchanged,
  so their intended cushion still does its job.

- A request no longer queues behind an arbitrarily old render. Sharing a render
  is what collapses a cold-cache stampede into one render, but a hung render
  used to hold up everyone who arrived after it. A newcomer past 2s now serves
  the placeholder instead of waiting, reusing the ImagerUnavailableError path
  that "both upstreams down" already takes. The caller that actually started
  the render keeps waiting, which is correct: it is the one whose image this
  is. When the join window is removed the new test hangs for the full 10s it
  was meant to prevent, which is the tail this bounds.

- The information_schema row-count estimate is gone; the counters are exact
  again. Running it against the live database: users 165, rooms 92, camera_web
  0, and the estimate was 0.00% off on all three. At 165 rows an index scan is
  cheaper than the extra round trip the estimate needed, so the optimisation
  bought nothing and traded a guaranteed-correct member count for an
  approximation that InnoDB would only make less accurate as the table grows.
  The exactness is now pinned by tests: a real zero stays zero, a database
  error propagates instead of becoming a number, and each counter counts the
  table it claims to. The module stays, because the homepage and the boot
  warm-up writing different values to the same cache key is its own bug.

The module comment records the measured numbers, because "COUNT(*) is too slow"
sounds true in the abstract and is false here.

3223 tests pass.
2026-09-25 18:55:33 +02:00
openhands f81b114b69 perf(cache): single-flight avatar renders, cacheable public reads, cheap row counts
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m30s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m7s
Three separate things that were each costing more than they needed to on the
hot path.

- Single-flight avatar renders. The disk cache was checked first and a miss
  went straight to the upstream, with nothing shared between callers, so a page
  requesting dozens of avatars at once turned N concurrent requests for one
  figure into N renders. A render is the most expensive operation this app
  does, and the duplication happened exactly when the cache had nothing to
  offer. Eight concurrent requests now cause one render instead of eight. The
  map lives on globalThis because Next can evaluate the module more than once
  per process, and two copies would each start their own render.

- Let public read-only routes be cached by a shared cache. Every JSON response
  was `cache-control: no-store`, so a CDN in front of the app could not answer
  any of it and every request reached the origin. publicCacheControl() opts a
  route in with s-maxage and stale-while-revalidate, using the same TTL as the
  server-side cache so the two layers cannot disagree. The default stays
  no-store: most routes here are personalised, admin-only or auth-dependent.
  /api/badges/leaderboard is deliberately left alone because it returns
  per-viewer rank entries to signed-in callers.

  Note this only takes effect once a cache rule exists for /api/* at the CDN, or
  the explicit `cache: "no-store"` is dropped from the client fetches (24 files
  do that today, including the /api/online poll). The headers alone are inert
  until one of those happens.

- Take the homepage row counts from the storage engine estimate instead of
  COUNT(*), which walks an index and gets slower as the tables grow. A missing
  or zero estimate falls back to the exact count rather than ever showing a
  wrong zero. The online count stays exact: it is an indexed read over a small
  subset and a few seconds of drift reads as broken rather than approximate.

The counters move into one module because the homepage and the boot warm-up
populate the same cache keys, so two implementations would race to write
different values into the same entry.

3223 tests pass.
2026-09-25 18:42:23 +02:00