Follow-up to f81b114b, addressing the three ways that commit could make things
worse rather than better. All three were verified against the real database or
by breaking the test and watching it fail.
- The grace window is now capped at 120s. A window is a cushion for the TTL
boundary, not a second TTL, but the call sites treated it as the latter: the
5 min values/staff routes and the 10 min teams route asked for a window as
long as or longer than their own TTL, so a single large staleMs silently
doubled how far behind a value could be served. Nothing marked those as
unsafe, because nothing looked wrong. The cap lives in the cache rather than
at the call sites so no future route can reintroduce it. Routes that asked
for less than 120s (the 10s online poll, the 20s news cache) are unchanged,
so their intended cushion still does its job.
- A request no longer queues behind an arbitrarily old render. Sharing a render
is what collapses a cold-cache stampede into one render, but a hung render
used to hold up everyone who arrived after it. A newcomer past 2s now serves
the placeholder instead of waiting, reusing the ImagerUnavailableError path
that "both upstreams down" already takes. The caller that actually started
the render keeps waiting, which is correct: it is the one whose image this
is. When the join window is removed the new test hangs for the full 10s it
was meant to prevent, which is the tail this bounds.
- The information_schema row-count estimate is gone; the counters are exact
again. Running it against the live database: users 165, rooms 92, camera_web
0, and the estimate was 0.00% off on all three. At 165 rows an index scan is
cheaper than the extra round trip the estimate needed, so the optimisation
bought nothing and traded a guaranteed-correct member count for an
approximation that InnoDB would only make less accurate as the table grows.
The exactness is now pinned by tests: a real zero stays zero, a database
error propagates instead of becoming a number, and each counter counts the
table it claims to. The module stays, because the homepage and the boot
warm-up writing different values to the same cache key is its own bug.
The module comment records the measured numbers, because "COUNT(*) is too slow"
sounds true in the abstract and is false here.
3223 tests pass.
The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.
- Evict least-recently-used instead, and raise the default budget to 2000
(CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
lives on the entry, so one call site opting in protects every reader of that
key. A failed background refresh keeps serving the last good value instead of
falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
Redis key and publishes a signal, so a value written by one process is no
longer served stale by the others for the rest of its TTL. A failed Redis
delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
call, with a pub/sub signal to drop the local copy when it rotates. A Redis
outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
the old cap counted files and never removed anything while entries were
fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
runtime imaging cache.
Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.
3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
- Cache read-heavy public API routes via redisCache (leaderboard, values,
shop, articles, photos, guilds, teams, staff, users, home, radio, badges)
- Add single-flight and bounded-memory cache layer with unit tests
- Parallelize independent DB queries on search, rares, shop, staff, polls
and profile pages
- Push radio points leaderboard aggregation to SQL with a LIMIT
- Split studio-client and import-furni-client into focused modules
- Clean up next.config.ts
Database:
- Add missing indexes (users.credits, users_currency(type,amount),
users_settings.respects_received, camera_web.timestamp,
messenger_offline.user_id) via migrations 0020/0021
- Use partial .select() everywhere instead of SELECT * (tickets, users,
rooms, audit logs, catalog tree, polls, radio, password reset)
- Add queryPrepared/queryPreparedOne (server-side prepared statements)
and switch the login check to a prepared statement; drop dead
cache options from the pool config
- Raise total_users/total_rooms COUNT(*) cache TTL to 5m
Caching:
- Consolidate the three cache helpers (cached, redisCache, cachedQuery)
into a single memory-first implementation backed by Redis
- invalidateKey now clears the in-process cache as well as Redis
- Cache homepage sections, news list, and leaderboard tabs; share one
news_list cache key between homepage and news archive
- siteSettings: in-process cache with TTL so repeated getters no longer
pay a Redis round-trip per call
- Share a 10s poll cache across all radio SSE connections
- Normalize timestamps after cache reads (Redis JSON round-trip)
Assets:
- Enable AVIF/WebP via images.formats and remove unoptimized from news
covers and the homepage hero (149KB jpg) with proper sizes/priority
- Support ?format=webp|avif|png in the /imaging proxy via sharp
Other:
- Fix pnpm supply-chain minimumReleaseAge failures by excluding the
freshly-published packages (next 16.3.1, hookform resolvers 5.8.0,
resend 6.20.0)
- Remove unused before/after fields from housekeeping AuditEntry