Follow-up to f81b114b, addressing the three ways that commit could make things
worse rather than better. All three were verified against the real database or
by breaking the test and watching it fail.
- The grace window is now capped at 120s. A window is a cushion for the TTL
boundary, not a second TTL, but the call sites treated it as the latter: the
5 min values/staff routes and the 10 min teams route asked for a window as
long as or longer than their own TTL, so a single large staleMs silently
doubled how far behind a value could be served. Nothing marked those as
unsafe, because nothing looked wrong. The cap lives in the cache rather than
at the call sites so no future route can reintroduce it. Routes that asked
for less than 120s (the 10s online poll, the 20s news cache) are unchanged,
so their intended cushion still does its job.
- A request no longer queues behind an arbitrarily old render. Sharing a render
is what collapses a cold-cache stampede into one render, but a hung render
used to hold up everyone who arrived after it. A newcomer past 2s now serves
the placeholder instead of waiting, reusing the ImagerUnavailableError path
that "both upstreams down" already takes. The caller that actually started
the render keeps waiting, which is correct: it is the one whose image this
is. When the join window is removed the new test hangs for the full 10s it
was meant to prevent, which is the tail this bounds.
- The information_schema row-count estimate is gone; the counters are exact
again. Running it against the live database: users 165, rooms 92, camera_web
0, and the estimate was 0.00% off on all three. At 165 rows an index scan is
cheaper than the extra round trip the estimate needed, so the optimisation
bought nothing and traded a guaranteed-correct member count for an
approximation that InnoDB would only make less accurate as the table grows.
The exactness is now pinned by tests: a real zero stays zero, a database
error propagates instead of becoming a number, and each counter counts the
table it claims to. The module stays, because the homepage and the boot
warm-up writing different values to the same cache key is its own bug.
The module comment records the measured numbers, because "COUNT(*) is too slow"
sounds true in the abstract and is false here.
3223 tests pass.
Three separate things that were each costing more than they needed to on the
hot path.
- Single-flight avatar renders. The disk cache was checked first and a miss
went straight to the upstream, with nothing shared between callers, so a page
requesting dozens of avatars at once turned N concurrent requests for one
figure into N renders. A render is the most expensive operation this app
does, and the duplication happened exactly when the cache had nothing to
offer. Eight concurrent requests now cause one render instead of eight. The
map lives on globalThis because Next can evaluate the module more than once
per process, and two copies would each start their own render.
- Let public read-only routes be cached by a shared cache. Every JSON response
was `cache-control: no-store`, so a CDN in front of the app could not answer
any of it and every request reached the origin. publicCacheControl() opts a
route in with s-maxage and stale-while-revalidate, using the same TTL as the
server-side cache so the two layers cannot disagree. The default stays
no-store: most routes here are personalised, admin-only or auth-dependent.
/api/badges/leaderboard is deliberately left alone because it returns
per-viewer rank entries to signed-in callers.
Note this only takes effect once a cache rule exists for /api/* at the CDN, or
the explicit `cache: "no-store"` is dropped from the client fetches (24 files
do that today, including the /api/online poll). The headers alone are inert
until one of those happens.
- Take the homepage row counts from the storage engine estimate instead of
COUNT(*), which walks an index and gets slower as the tables grow. A missing
or zero estimate falls back to the exact count rather than ever showing a
wrong zero. The online count stays exact: it is an indexed read over a small
subset and a few seconds of drift reads as broken rather than approximate.
The counters move into one module because the homepage and the boot warm-up
populate the same cache keys, so two implementations would race to write
different values into the same entry.
3223 tests pass.
The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.
- Evict least-recently-used instead, and raise the default budget to 2000
(CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
lives on the entry, so one call site opting in protects every reader of that
key. A failed background refresh keeps serving the last good value instead of
falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
Redis key and publishes a signal, so a value written by one process is no
longer served stale by the others for the rest of its TTL. A failed Redis
delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
call, with a pub/sub signal to drop the local copy when it rotates. A Redis
outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
the old cap counted files and never removed anything while entries were
fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
runtime imaging cache.
Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.
3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
A fallback render drops the requested effect and is only a degraded
stand-in, so writing it to the 30 day disk cache kept serving the worse
image long after the local renderer recovered. Cache primary renders only
and let the next request pick up the real render.
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.
Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
Add an alerting/stats layer over the existing CrowdSec integration:
- New crowdsec-alerts.ts: cooldown-gated ops alerts (Redis NX lock, TTL from
HEALTH_ALERT_COOLDOWN_MIN) fanning out through the app's sendAlert service.
Raised for daily quota exhaustion, block bursts (5-min window past
CROWDSEC_ALERT_BLOCK_BURST), and signal-push failures.
- New crowdsec-stats.ts: daily counters (lookups/blocks/reports/report_fail)
in Redis with a 14-day reader for the admin panel.
- Shared 403/429 backoff: the pause marker now lives in Redis
(crowdsec:backoff-until) so every instance honours it, not just the process
that hit the limit.
- Atomic quota reservation: INCR-before-call with self-rollback on overshoot,
so concurrent instances can never slip calls past the daily ceiling.
- Admin anti-DDoS page gains a last-14-days activity table next to the quota bar.
- Bound the in-process verdict cache (FIFO eviction at 2000 entries) so a
flood of distinct bucket-tripping IPs cannot grow it without limit.
- Record block metadata (reputation, score, behaviors, category, TTL) in
antiddos:block:meta:{ip}, surfaced as the reason in the admin block list;
unban now also clears the metadata and report locks.
- Track daily CTI enrichment usage in Redis (crowdsec:usage:{date}); warn
once at 80% and pause lookups until tomorrow at CROWDSEC_CTI_DAILY_QUOTA
(default 10000, 0 = unlimited) so a via-spread DDoS cannot burn the plan.
- Add opt-in signal push to the CrowdSec community (CAPI watcher): stable
auto-generated 48-char machine_id/password pair persisted in Redis (or via
env), one-time registration, cached JWT login, optional Console enrollment,
and POST /v3/signals with a ban decision, deduped per IP. Never throws and
reports last status to the admin panel with a verify action.
- Admin page: quota usage bar, reporting status/verify channel, and CrowdSec
block reasons in the active-blocks list.
Align the active runtime with the pinned version across .nvmrc, package.json
engines and the Dockerfile base images, so scripts/check-node-toolchain.mjs
passes on the CI host running Node 26.10.0.
- new crowdsec-api lib: CTI lookup (GET /smoke/{ip}, freemium x-api-key), verdict parser with false-positive veto, 1h Redis + in-memory verdict cache, NX lock dedupe, 403/429 backoff; writes only the shared antiddos:block:{ip} key (value "crowdsec") and never touches Cloudflare
- gate fires it fire-and-forget for IPs that already tripped a rate bucket, so known-bad IPs are hard-blocked before the local maxViolations threshold
- runtime config: crowdsecAutoBlock toggle, score threshold (0-5, default 4), block TTL (default 24h); boot defaults CROWDSEC_AUTO_BLOCK_ENABLED / CROWDSEC_BLOCK_SCORE / CROWDSEC_BLOCK_TTL_SECONDS
- admin panel: CrowdSec stat card, verify-connection action, score/TTL settings, CrowdSec source badge in the blocked-IPs list
- credentials live in env only (CROWDSEC_API_KEY); block is enforced per-request via proxy on the resolved X-Forwarded-For / CF-Connecting-IP
- tests: crowdsec-api unit suite + ddos-guard integration suite (early-block, threshold, cache dedupe, backoff)
cloudflare-api unit tests drove the real Redis connection when REDIS_URL was set (CI), causing cross-test bleed. Mock @/lib/redis with an in-memory fake identical to the gate integration test.
- gate creates a zone IP Access Rule (block) for proxied offenders that hit the block threshold, deduped until the tiered block expires
- cloudflare-api lib: verified endpoints, create/delete/verify/list helpers, Redis-backed tracking + 30s TTL sweep (instrumentation worker + admin render)
- runtime toggle cloudflareAutoBlock in antiddos config; boot default CLOUDFLARE_AUTO_BLOCK_ENABLED
- admin panel: Cloudflare edge-blocks card with verify + remove-rule actions; unban also lifts the edge block
- credentials live in env only (CLOUDFLARE_API_TOKEN / CLOUDFLARE_ZONE_ID)
Bump the framework to the latest 16.3.6 patch release. Typecheck passes and
the homepage renders (HTTP 200) on the dev server with Next 16.3.6 under
Turbopack.
Split the heavy test suites out of the check job so coverage, MariaDB/Redis
integration and Playwright UI tests run concurrently on the host runner
(capacity raised to 4) instead of back-to-back (~2min wall-time saving).
Deploy and preflight now gate on all three test jobs.
Point PLAYWRIGHT_BROWSERS_PATH at the persistent /opt/ms-playwright dir on
the host runner so 'playwright install chromium' is an instant no-op after
the first run (was ~100s CDN download per job).
- Switch test-runner and renovate workflows from ubuntu-latest to
self-hosted now that a native host runner is running as a systemd service
- Replace remaining hardcoded color utilities in the homepage with theme
tokens and inline rgba styles to satisfy the no-hardcoded-colors contract
- Restore dual UserAvatarThumbnail usage on the homepage (hero avatar stack
plus community grid) to satisfy the public avatar presentation contract
- Create src/middleware.ts for per-request nonce-based CSP
- Integrate src/lib/csp.ts to build the CSP header dynamically
- Add src/middleware.test.ts to verify CSP header is set with nonce
- Biome lint and TypeScript checks pass
- Update txSelect and db.select mocks in draw-badge.test.ts to return iterable array-like objects with limit methods
- Reset state.price in beforeEach
- Fix test assertions for unsafe character stripping test
- All 3,066 tests now pass cleanly
The palette lives in plain CSS (:root/ThemeVars/admin remap), so Tailwind
never generated bg-primary, bg-destructive, text-foreground and similar
utilities. Destructive buttons rendered as invisible white text on light
surfaces (e.g. the Nitro cleanup delete button). Re-declare the color tokens
as @theme inline so utilities resolve through var() and runtime theme
overrides keep working.
- Add persistent disk cache for rendered avatars/badges (storage/imaging)
so repeats never touch the flaky local renderer and cached renders
survive upstream downtime
- Serve cache-first with stale-on-error; cut primary/fallback timeouts
from 10s/6s to 4s/4s so failing images cannot stall pages
- Avatar proxy now returns a graceful 200 silhouette instead of 502 when
no renderer can produce a figure, so no broken-image glyphs appear
- Badge endpoint becomes a caching proxy trying configured CDN, public
Habbo CDN and local /swf copy in order, and drops the fragile IP rate
limit that could blank badge streams
- Route all site badge images (profile, me, badges, apply pages) through
the cached proxy instead of hot-linking images.habbo.com
- Track referral attribution at registration via ?ref code with
same-IP and duplicate-pair guards
- Add daily login rewards with streak tracking, claim flow and
sendCurrency payout backed by RCON with DB fallback
- Add admin pages for referral settings and the daily reward schedule
- Add migration 0033 with tables, seed schedule, settings and ACL grants
- Add admin.referrals.* and admin.dailyrewards.* permission slugs
- Localize new copy in en, nl and it