9bc600579a3f4030c36a3b0c236618f1af6cb92a
640
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
43742e8d99 |
fix(catalog): decode WebP bundle textures so furniture icons resolve again
Every write path normalises a bundle's texture to WebP Lossless, so the catalog icon was being read back with a PNG-only decoder. decodePng throws on anything that is not a PNG, the callers caught that and returned null, and the user-visible result was "no icon (not in source or bundle)" for every furniture whose source does not serve a standalone icon. Measured against the production asset tree, all 18,505 bundles were WebP; icon extraction succeeded on 0 of them. src/lib/services/imager/ decode-texture.ts keeps PNG on the dependency-free decoder and routes WebP through sharp, which is already a dependency and already encodes these textures. extractFurniIconPng and getPetIconPng become async; the five call sites (upload, clone import, furni import, icon repair and both icon routes) already awaited their surrounding work. Same root cause, second bug: the spritesheet frame key. Converters disagree on packing — some keep a trailing ".png", and some lowercase the whole key while leaving the bundle name mixed-case, so "LTD_fashionistaf" looks up frame "LTD_fashionistaf_LTD_fashionistaf_icon_a" and never finds "ltd_fashionistaf_ltd_fashionistaf_icon_a". Any mixed-case classname could therefore never match, which is most of the catalogue. findFrame tries the two exact spellings, then falls back to one case-insensitive pass. Extraction now succeeds on 18,483 of 18,505 bundles (99.88%); the 22 remainder are data, not code — 9 bundles ship no icon asset, 11 do not parse. Third: three catalogue icons exist only as .gif while catalogueIconUrl hardcoded .png, so the picker offered icons that could only ever 404, and 291 icons that ship as both formats were listed twice. The API now dedupes per id and the two renderers retry with .gif before falling back to the placeholder, matching what catalog-image-picker already did. Verified live: /gamedata/.../icon_1542.png returns 404 while icon_1542.gif returns 200. Separately, close the last hole in the memory cap. Every script in package.json routes through scripts/with-memory-cap.sh, but invoking the builder directly — from a terminal, an IDE or an agent — skipped the wrapper and ran unbounded, on a host with no swap where the OOM killer picks its victim across the whole machine. next.config.ts now refuses a production build that the wrapper has not marked, before anything allocates. next dev and next start are deliberately unaffected. README gains a Memory-capped commands section covering the per-script ceilings, the backends and the ulimit -v trap, and its stale version and script tables are corrected. |
||
|
|
6793f77733 |
fix(deploy): stub git status in the simulation harness
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m37s
CI / tests-unit (push) Successful in 1m39s
CI / tests-ui (push) Successful in 2m20s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m6s
The clean-tree guard in ci-deploy.sh aborts the release when
`git status --porcelain` prints anything. The harness stubs `git` as a
shell function that only special-cases `rev-parse`; every other
subcommand fell through to its `ls-remote`-shaped printf, so `git status`
emitted a fake refs/heads/main line and the guard failed on every
scenario.
Introduced in
|
||
|
|
adffac7360 |
feat(catalog): store furniture bundles as .hab instead of .nitro
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-unit (push) Failing after 1m37s
CI / tests-integration (push) Successful in 1m37s
CI / tests-ui (push) Successful in 2m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Every bundle the CMS writes — upload, clone, sync, repair and the pet / effect / figure importers — now lands as `<classname>.hab`, the extension this deployment's renderer asks for. `.hab` and `.nitro` are the same container, so an upload of either extension is accepted. Resolution goes through one module, src/lib/furni/bundle-file.ts, so nothing has to know the extension twice. Every existence check probes `.hab` first and falls back to `.nitro`: the on-disk asset set is still predominantly `.nitro`, and without the fallback Studio would report every imported item as missing and the cleanup scan would classify 18k live bundles as fake leftovers. Downloads are unchanged — Habbo's CDN and every configured clone source still serve `.nitro`, so the conversion happens on write, not on request. Deliberately unchanged: the staged-attachment store in furni-attachment.ts keys on a UUID and never reaches the client, so renaming it would break in-flight recovery jobs. Adds scripts/migrate-nitro-to-hab.ts to rename the existing asset set. It refuses to run without --dry-run or --yes, never overwrites an existing .hab, never deletes, and is idempotent. Note: renderer-config.json lives outside this repo and was patched to .hab separately; that file is served with a 30-day max-age, so returning clients need a cms-client cache purge to pick the change up. |
||
|
|
fdb7af7ef5 |
fix(deploy): unblock every rebuild on BuildKit's host-network refusal
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-unit (push) Failing after 2m4s
CI / tests-integration (push) Successful in 2m6s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Docker 29.1.3 ships BuildKit v0.26, which refuses to grant a build host networking unless each caller passes --allow=network.host. All three rebuild paths asked for it, and `docker compose build` has no flag to grant it, so a rebuild failed immediately with "additional privileges requested". The live container was never replaced, which is exactly the reported symptom: the site kept serving the previous release after a rebuild. Nothing in the build actually needs host networking. It uses the network only for apk, pnpm and next/font/google — all outbound internet, which the default bridge provides. Verified by building both the full runner image and the migrations stage with --no-cache after dropping the flag. Runtime `network_mode: host` stays: blue/green needs per-release host ports (3002/3003) and nginx reaches each slot over 127.0.0.1. The second gap is how a rebuild could still ship the wrong code. ci-deploy.sh stamped every image with HEAD's revision label, and verify-deployed-release.mjs only re-checks that same label, so a dirty working tree produced an image that claimed to be release $sha while containing uncommitted code. docker-update.sh already refused this; ci-deploy.sh now does too, before any build work. |
||
|
|
759ae91745 |
perf: cache search and news archive, drop motion/react from public pages
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 26s
CI / tests-unit (push) Failing after 1m37s
CI / tests-integration (push) Successful in 1m38s
CI / tests-ui (push) Failing after 2m24s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Closes the four remaining LOW items. Search and news archive caching - A leading-wildcard LIKE cannot use an index, so every /search section cost a COUNT(*) scan plus an ordered page fetch, and /news did the same for its archive. Both now cache: search per section for 30s, the archive for 60s under the existing news revision so publishing an article drops it at once. - Sections are cached independently, so one slow query cannot hold up the rest and a failure is not cached as a result. - The cached value is passed through cacheSafe() so the Redis path and the in-process path return the same types; without it a cache hit would hand the events grid a string where a miss hands it a Date, and it calls toISOString() on that field. Dates are revived on the way out so the public signatures of loadNewsArchive and loadPublicSearch are unchanged. - Archive entries are keyed on the REQUESTED page rather than the clamped one, so two requests that clamp onto the same page cannot alias each other. motion/react out of the public bundle - Converted the six public-facing users: the radio player, the typewriter text (a motion.span with no animation props at all), the photo lightbox, the animated counter, the footer CMS-info popup and the scroll reveal. That was the actual entry points — the counter and the popup reach the public home page and footer through static imports, so removing only the three originally named would have left the library in the bundle anyway. - Each animation moved to a CSS class, and the two that animate on exit now hold the element for the length of the fade, which is what AnimatePresence used to do. - motion/react now only ships with /admin and the two already-lazy nav panels. - Two safety fixes came out of this: the scroll reveal starts at opacity 0, so it is forced visible under prefers-reduced-motion and via a <noscript> rule in the root layout; and it now emits the .motion-reveal class, which the theme panel's "Scroll Reveal" toggle selects and which previously matched nothing. - The CMS-info backdrop became a real button in a pointer-transparent layer instead of a handler on a static element, so click-outside-to-dismiss is reachable by keyboard. Fewer duplicate router refreshes - Next.js re-renders the current route as part of a server action's own response when that action revalidates, and applies it with a seeded navigation; the router only skips its own update when the action did NOT revalidate. So the refresh after such an action fetched the same tree twice. - useServerAction takes an opt-in `revalidated` flag that skips it. It is opt-in per call rather than derived from an action name, since a rename would silently change behaviour. Applied to the two user-facing call sites whose actions were verified to revalidate their own route. Touch targets - .btn was the one shared control at 40px; it and the lightbox and CMS-info close buttons are now 44px, as is the password toggle (the auth input already reserved 44px for it). The remaining 32px icon buttons pass WCAG 2.2 AA, which only asks for 24px; enlarging those inside inputs and overlays was left alone because it risks visual breakage that cannot be checked from here. |
||
|
|
179484642f |
feat: per-account login lockout, mail index, resend captcha, i18n scoping
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Failing after 1m45s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m29s
CI / deploy (push) Skipped
Closes the four HIGH/MEDIUM items left open after the previous pass. Login lockout - The only login limits were keyed on the client IP, so a distributed attempt could grind on one account indefinitely. Added a per-account lockout with a budget of 8 failures per 15 minutes. - The bucket is keyed on the RESOLVED account id, not on the submitted string: users may sign in with either username or e-mail and neither the lookup nor the input normaliser folds case, so an input-keyed bucket would hand out a fresh budget per spelling of the same account. - precheckLogin and NextAuth's authorize share the bucket, so the pre-check cannot be used to buy extra attempts and a client that skips it entirely is still bounded. Both check the lockout BEFORE verifying the password: the success path clears the counter, which would otherwise walk a locked account straight back in on the right password. - A successful login clears the failures, which needs two new primitives in rate-limit.ts: peekRateLimit (read-only, does not consume a unit) and clearRateLimit. - Fixed a latent inconsistency while doing so: the in-process bucket capped its counter at the limit while Redis' INCR kept climbing, so the two backends disagreed about how far over the limit a key was. Both now track the true count. Mail lookup index - Added an index on users.mail (0035). Password reset, e-mail verification and the resend cooldown all resolve a single account from a submitted address and were full table scans of `users`. Deliberately non-unique: legacy rows can hold the same address more than once, so a unique index would fail to apply. Resend captcha - /verify's resend form triggers real outbound mail and was reachable with only a cooldown. It now runs the configured captcha before the account lookup and before any send. Client message payload - The root layout serialised the whole catalogue into every page. pages.admin and admin are ~177 KB of the ~235 KB and are unreachable from the public route group, so that layout now installs its own provider with the staff namespaces removed. Nested providers replace rather than merge, which is why this has to live in the segment layout. /admin, /mod, /client and /admin-next keep the full set; a guard test fails if a public page ever references a staff namespace. |
||
|
|
6cc45d7413 |
feat: harden atoms-nexst against review findings (37 items)
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m31s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Second review pass covering security, performance, admin tooling and the public/room flows. All HIGH and MEDIUM findings from the audit are resolved; nothing in this commit changes the visible feature set. Authentication & session security - CSP is now set on the request headers in the proxy, which is what Next.js uses to derive the render nonce, so the nonce is effective. - 2FA: an already-enabled user cannot re-enroll, the setup endpoint is rate-limited per account, and confirmed codes are persisted so the second secret no longer silently never applies. - Password reset revokes the ticket, authTicket and all personal access tokens, and bumps the token version so existing sessions die. The same revocation is now wired into the staff-side password reset. - /reset and /verify return a stable error code instead of raw text; the mail lookups are ordered by id so duplicates cannot vary between runs. - Resending the verification mail gets a per-address cooldown on top of the per-user limit. - Issue API tokens with the narrower radio/ticket ability set instead of "*". Authorization & input handling - Mid-rank staff can no longer keep dynamically granted non-view admin.* permissions: existing grants are revoked by migration and the grant lookup is restricted to "%.view". Rank guards use the dynamic super-admin check. - Alerting a user is permission-checked and audited like the other tools. - Material mutations (giveCredits/giveDuckets/giveDiamonds, the admin user actions route, bulk user actions) are capped and rank-guarded, and bulk ids are bounded. - updateRoom / updateRoomItem write through a field allowlist, and items may only be edited through their own room. - Classnames reaching the filesystem are validated before use so a crafted value cannot escape the asset directories. - The word filter now also covers offline mails, guild forum threads and replies, and user mottos. - Media uploads are validated by magic bytes, /api/media requires the page edit permission, APP_URL must be configured once mail is enabled, and the diagnostics error route checks the fetch site header. Admin tooling - Secret settings render masked and cannot be overwritten with a blank or an arbitrary raw key; radio credentials are new password inputs. - Commandocentrum balance changes are audited. - Admin list pagination reads the caller's per-page instead of the max, and the log exporter caps offset and search length. Performance - Catalog translations are cached per module, with a cheap revision hash; the public online count uses a stale window instead of hammering the DB. - The cache warmup now primes the payload the home route actually reads. - TopHeader batches its queries into one round trip, and LCP avatars load eagerly. - motion/react and sonner are no longer part of the root layout; the nav dropdown and mobile nav panels are lazy client chunks. Anonymous visitors again get the navigation chrome, and public pages get an edge cacheable response. Accessibility - Nested <main> elements in phase pages became <section>; the page entrance and route progress animations are pure CSS that respect reduced motion. |
||
|
|
3933214953 |
feat(auth): implement all 16 homepage/login/register review items
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m47s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m30s
CI / deploy (push) Failing after 2m56s
- add countArticles() (published-only, mirrors news-list) and warm total_articles - localize homepage metadata; bind articleCount to both stats; unique photo alts - drop duplicate news date and the mascot preload priorities - extract shared AuthPageFrame/AuthUsersCards used by /login and /register - login: localized noindex metadata, session redirect via safeRedirectPath, ?from passthrough from proxy, unified auth roster cache keys, registered notice - register: localized metadata, session redirect to /me, unified cache keys - add resend-verification flow on /verify with rate-limited non-enumerable action - add safeRedirectPath() with unit tests - register form: live requirements checklist + password mismatch guard - login form: unverified state with resend-link CTA - honour prefers-reduced-motion in TypewriterText - add 6 translations across all 25 locales |
||
|
|
5b2eb91c5c |
fix(ops): stop a compose replica from blocking the blue/green release
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m57s
CI / tests-unit (push) Successful in 2m3s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m49s
CI / deploy (push) Successful in 2m48s
The deploy failed after the build, the migrations and the browser gate:
"Port 3002 is already in use". The holder was `epicnext-cms`, a compose
replica of release
|
||
|
|
6c3d81920e |
fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-integration (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than the code. Each one had a signature that looked like a network or permissions problem and was actually a configuration or ordering bug. jobs-worker never ran `import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM evaluates a module's imports in source order and the first import reaches `@/env`, which validates process.env at import time. The ZodError on DATABASE_URL therefore fired before load-env ever executed, so the worker could only start from a shell that had already exported the configuration. Nothing supervised it either, so scheduled articles, catalog export, JAR and database backups, disk alerts and the ops health probe have all been dead; `cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and added deployment/systemd/cms-jobs-worker.service with Restart=always. The JAR backup additionally pointed at './emulator/Arcturus.jar', which does not exist and would go stale on the next emulator upgrade. resolveEmulatorJar now accepts a file, a directory or a wildcard and picks the newest JAR, the same way emulator.service picks its build, and reports an unresolvable path once instead of logging an opaque copyFile ENOENT every night. /api/health answered 200 with the database down The route documented this as intentional, and ci-deploy.sh worked around it by grepping the body for '"database":true'. The container healthcheck did not, so Docker reported containers healthy while every page 500'd. The status is now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis and the emulator deliberately do not fail the container — both have in-process fallbacks, so failing them would trade a slow site for an outage. The runtime had no V8 heap cap NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its heap from host memory (23.5 GB) while the container was limited to 4 GB, so the kernel OOM-killed the process mid-request — the same failure mode as the 14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit (v2 with a v1 fallback) and sets 70% of it, respecting an explicit override. Storage ownership was only repaired for one path ci-deploy.sh chowned storage/imaging and nothing else, so storage/catalog-git/hotel-status.json kept coming back root:root and /api/admin/catalog/status kept throwing EACCES. All eight writable storage paths are repaired now. The silent-failure mode is the reason this mattered: these writes sit inside try/catch, so a wrong owner looks like a slow page rather than an error. nginx: robots.txt was a guaranteed 404, and TLS never resumed `index index.html` without a `root` left every try_files resolving against /etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES on each stat and nginx logs a failed stat at crit, which is where 149 crit lines per scan came from. robots.txt answered from that same broken location, so crawlers were pointed at a file they could never read while sitemap.xml kept advertising it. Added `root`, proxied robots.txt to the CMS, added ssl_session_cache (there was no session resumption at all), and set Restart=on-failure in a systemd override, since the packaged unit ships Restart=no and nginx is the only thing serving the site. Verified against the running host: 3379 tests, typecheck and biome clean, nginx -t passes, health returns 200 with every check green, and the worker has run for hours at NRestarts=0 with a heartbeat refreshing each minute. |
||
|
|
108c6ce03d |
fix(ci): make the lint gate fail for real and stop byparr leaking disk
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 2m3s
CI / tests-unit (push) Successful in 2m18s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 3m6s
CI / deploy (push) Failing after 3m14s
The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.
Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
(key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
function-declaration handlers, with the reasoning that dropping them broke
the tree and save-on-Ctrl+S once already (
|
||
|
|
8218039c64 |
test(live): stop the live suites inheriting the production database
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m57s
CI / tests-integration (push) Successful in 2m1s
CI / tests-ui (push) Successful in 2m51s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m42s
Seven suites read .env with a bare `process.env[key] = value`, which overwrites whatever the shell already set. That made the DATABASE_URL from the production .env authoritative, so a single environment variable was enough to aim them at the live hotel database: RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts Three of those suites then repair the catalog in place: catalog-audit-repair-live and catalog-repair-direct-live rewrite catalog_items and delete duplicate classnames, and clone-bulk-import-live bulk-imports every cloneable item. None of that is undoable, and nothing in their output said the target was production rather than a sandbox. Added src/test/live-env.ts with one shared loader, and pointed all seven suites at it: - Values already in the real environment win, so an explicit DATABASE_URL on the command line is always respected. - DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the production one from .env. - Anything that is not loopback is treated as production and redirected. - Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning saying the suite repairs the catalog. Tests in src/test/live-env.test.ts run the loader against a temporary .env so the real project file is never read, and cover the redirect, the shell override, non-loopback detection, the opt-in and quote stripping. A second block asserts each of the seven suites no longer contains an inline `process.env[...] =` assignment. Verified four of them fail against the old loader. This does not enable the suites; they stay gated behind their RUN_* flags. It only removes the possibility of them silently hitting production. Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean. |
||
|
|
8ee144745a |
fix(deploy): stub ss in the deploy harness and cover the port-conflict path
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m47s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m41s
The port-conflict guard added in the previous commit made the six existing blue/green deployment simulation tests fail. assert_port_free() shells out to ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node — but not ss. Because the runner is self-hosted and the containers use --net=host, the simulation saw the production CMS containers holding 3002 and 3003 and refused to start its own candidate. The harness now stubs ss. It reports no listener for every scenario except 'port-taken', which reserves whichever port the script asks about, so the simulation stays independent of the host it runs on. Also switched the ss probe from `command -v ss` to `type ss`. The stub is a shell function delivered through BASH_ENV; `command -v` happens to find it, but `type` is the reliable test for "is this resolvable", and the two differ across shells. Added a regression test for the guard itself: with the candidate port already occupied, the deploy must fail, must not have run `docker run`, and must leave the nginx upstream untouched on the old port — no half-finished cutover. Verified it fails when the assert_port_free call is removed. Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped. |
||
|
|
30ff970c38 |
chore: upgrade to pnpm v12, update dependencies, and fix msw v3 typescript types
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / preflight (push) Skipped
CI / tests-ui (push) Skipped
CI / deploy (push) Skipped
|
||
|
|
f99980052b |
perf: optimize cache layer for speed and stability
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 34s
CI / tests-integration (push) Failing after 33m57s
CI / tests-unit (push) Failing after 33m57s
CI / tests-ui (push) Failing after 33m56s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- Remove random TTL jitter to prevent unpredictable cache drops - Add deterministic LRU eviction with proper entry cleanup - Improve cache deduplication to prevent duplicate computations - Skip Redis I/O during tests for faster, more stable execution - Optimize depth calculation in catalog tree nodes - Maintain backward compatibility and full test coverage (3331 passed) |
||
|
|
a6cc3cafa9 |
fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached twice and nobody could reach the client. The catalog items loader kept its own 30s TTL copy of FurnitureData.json next to the mtime-validated cache in `furni-data.ts`. An import cleared only the second one, so the catalog table kept serving pre-import furnidata — empty descriptions and revisions — until the TTL ran out. The loader now reads through `readFurniData`, which revalidates on mtime+size and is reset by every write, so there is exactly one cache and it cannot go stale on its own. `invalidateFurniDataCache` and its single call site are gone with it. The client was worse: nginx served all of /gamedata/ with `max-age=604800`, and the `cms-gamedata` purge that would have fixed it hung off the catalog Git export, which is disabled in production. A freshly imported item was invisible in the client for up to seven days no matter how often you imported. - `writeFurniData` now purges the gamedata edge tag itself. One place covers import, batch, resync, regen, nitro-editor, translate and dedupe. It is fire-and-forget and swallowed at every level: a stale edge copy is bounded by the edge TTL, so a failed purge must never fail an import. - nginx splits /gamedata/ by how mutable the content is: config/ gets `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`, and the content-addressed trees (c_images, album*, clothes) keep the long TTL. `must-revalidate` is the point — the client now revalidates instead of replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the purge still reaches them. - A 30-minute safety-net purge in the jobs worker covers the case where Cloudflare was unreachable at write time. |
||
|
|
e0efbef30d |
fix(catalog): make the Builder Club catalog read and write its own offers
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m23s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m6s
The previous commit taught bulk editing and delete-with-restore about the BC catalog. Neither actually worked, and one of them was destructive. `catalog_items_bc` has six columns: id, item_ids, page_id, catalog_name, order_number, extradata. There is no price, points, currency, offer_id, limit or membership column on it. The bulk path read and wrote columns that do not exist, and the UPDATE was aimed at catalog_items while the SELECT came from catalog_items_bc — so a BC category move wrote into the normal catalog. Two tests now pin that pairing: reads and writes have to stay in the same table. Underneath it the BC table was never being read at all. The inline editor fetched `/api/admin/catalog/items?pageId=N` without the catalog, so opening a BC category showed the normal catalog's offers, and the route selected BC rows directly instead of going through the loader, skipping the furni enrichment the table needs to render anything but a bare caption. Both catalogs now take the same path, and the catalog is in the fetch callback's dependencies — without that, a switch keeps reading the previous catalog's rows through a stale closure. Because a BC offer has no price, the editor no longer offers one. The server refuses price, points and currency changes with a readable message instead of letting them reach the database as an unknown-column error, and a BC bulk edit is what it can actually be: a category move. BC deletions also went through a bare DELETE, which made them the one catalog mutation with no way back. They now keep their rows and hand back a restoreId like the normal ones. The catalog is recorded in the audit target rather than in the payload, so a restore can never put a BC row into the normal offers table. |
||
|
|
cebcf440c5 |
feat(catalog): record bulk edits, make deletions reversible, unify the tree read
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m23s
A bulk offer edit is the catalog mutation that rewrites hundreds of rows at once, and it was the only one writing nothing to the staff activity log: 22 of the 43 catalog actions logged, this one did not. The entry it now writes says what changed, not just that something did, because the log has no undo of its own and "bulk updated 200 offers" cannot answer the question it exists for. Deleting offers had no inverse at all. Every removed row is now kept at delete time and the caller gets a restoreId back, so an accidental multi-select is a click rather than a hand-edit of the table. The undo toast covers the common case; a RecentDeletionsPanel holds the same records so a delete noticed later is still reachable. Three refusals guard it: an id that another offer has since taken, a category that no longer exists (which would leave an offer that sells nowhere and shows under no page), and a delete whose restore record cannot be written — that one rolls back rather than deleting without a way back. Reading the audit row FOR UPDATE is also what stops two restores of one deletion from both inserting. sendCatalogUpdate() overwrote hotel-status.json on every write, so "which imports reached the hotel" was answerable for the last attempt only, and a failure two imports ago was gone by the time anyone looked. That file is now also appended to as a bounded 50-entry tail. The tree route carried four copies of the same page-select-plus-counts shaping, of which the BC branches had already drifted: one counted offers through the VARCHAR-tolerant helper, the other inline and swallowing errors. All of it is one readPages() now, and readFullTree sends both catalogs through one depth computation instead of delegating normal to getTreeFlat while computing BC here — a split that left two implementations behind one function name. getTreeFlat is gone. The BC ancestor walk also went from 20 levels to 50, matching getAncestors, so a deeply nested catalog no longer loses its breadcrumb. Bulk editing reaches the BC catalog, which previously had no way to edit or duplicate offers in bulk. The catalog is part of the operation identity now, so replaying one request key against the other catalog is not mistaken for the same work. Integration tests failed to import: the next/cache mock supplied only revalidatePath, and catalog-totals calls unstable_cache at module scope. |
||
|
|
9550b3d66f |
feat(catalog): make the live catalog self-correcting and honest about failure
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m34s
CI / tests-integration (push) Failing after 1m34s
CI / tests-ui (push) Successful in 2m19s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The previous commit made imports update the Studio without a reload, but the guarantee only held inside the tab that started the import and only as long as every read succeeded. Four holes were left, and this closes them. A session that mounted the tree before an import kept the pre-import tree for the rest of its life, because ensureCatalogTreeLoaded() was a once-per-session no-op. It now asks the server whether what it holds is still current. The answer is a revision: sendCatalogUpdate() already runs after every catalog write, so it bumps one, and clients read it on mount, on focus, on a 20s poll and from other tabs over a BroadcastChannel. An import that finishes in another tab, another browser or the job worker now lands here too. A failed read used to be swallowed, which is the worst outcome available: the rail kept showing pre-import counts as if they were current and nothing said so. The snapshot now carries the error, the rail shows it with a retry, and the previous tree stays on screen because stale beats empty. Every settled import pulled the entire flat tree, which is the one payload that grows with the size of the catalog. The revision doubles as the ETag on mode=full, so an unchanged catalog answers 304 and the poll costs a file read. An import could also report success for an offer the hotel will never sell: a hidden or disabled page, an item_ids that misses the furni id, a zero amount. importSingleFurni reads its own row back and reports each of those as a warning, where the import report already is, instead of leaving it to surface as "the import did not work" in the client. Finally, the catalog items table no longer falls back to router.refresh() — onRefresh is now required, so every mutation ends in a refresh of the caller's own data instead of a route re-render that threw away editor state and scroll position. useServerAction keeps its default, because 47 callers across the app depend on it. The 750-line CatalogTree in catalog-tree.tsx was dead code that kept its own stale tree and three more router.refresh() calls; only CatalogIcon and LAYOUT_COLORS are still imported, so the rest is gone. Tests: the store now covers revisions, 304s, probe failures and error recovery; a jsdom test mounts a consumer and asserts the tree updates in place with no navigation; the old organize-imports e2e asserted nothing about the endpoints the code actually calls, and is replaced by one that asserts a cross-tab write lands in the mounted categories without a reload. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
28ce0f911c |
fix(catalog): keep the live catalog truthful after every import path
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Failing after 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m33s
CI / deploy (push) Skipped
The live catalog store only covered part of the import surface. A durable job settled, a sync queue drained, a .nitro upload or a clone run left the Studio rail and the stats bar showing pre-import numbers until the page was reloaded, and the Catalog Manager kept a second tree that never saw writes made elsewhere in the session. Every one of those paths now pulls the tree again, and the refresh carries the totals with it: importing writes catalog rows server-side, so the counts the store holds were stale for the rest of the session. - refreshCatalogTree shares one request between concurrent callers and queues a single follow-up read when a write lands mid-flight, so a burst of edits costs at most one extra read. - useFurnitureJobs treats its first payload as a baseline, so a page load no longer replays every past import as "just settled", and hands the settled jobs to the callback. - The Catalog Manager pushes its own mutations into the store and re-reads its active tab when the store changes. - The 30s unstable_cache on the admin totals is now tagged and invalidated from every catalog write, including the import worker, so it no longer survives an import even across a hard reload. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
7697728d07 |
feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single Cache-Control owner per route: the app stays the source, nginx only manages headers, and Cloudflare stores the public API allowlist at the edge. - deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the blue/green upstream snippet; config backed by scripts/nginx-sync.sh (idempotent install + reload, --check/--force). - nginx serves Cache-Tag headers on the public allowlist (cms-public), gamedata, client and camera responses so the edge and purge stay in sync. - src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that no-op unless Cloudflare is configured; scripts/cf-purge.sh and cf-setup-cache.sh create and purge the cache rule. - src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls. - Purge hooks after catalog exports (public + gamedata) and on shop, team, guild, photo and rare-values edits; ci-deploy purges after each release. - src/proxy.ts excludes the imaging/images docs from the middleware matcher. |
||
|
|
e4f83a8036 |
chore: upgrade to pnpm v12, vitest v5 and resolve deprecated subdependencies
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
|
||
|
|
d2d01141f1 |
style: apply Biome formatting to the catalog release test
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m49s
CI / tests-unit (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 20s
The timeout constant made the first it() line exceed the line width, so `biome check .` failed with a format error. Reformat and confirm the three publication tests still pass. |
||
|
|
c2981bd970 |
test(catalog): give the Git publication tests room to finish
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
These drive real git processes against a local bare remote, so their cost is process spawns competing with every other Vitest worker. Measured on CI they take 23-30s each, and the 30s override was crossed by 37ms, so the run failed on wall-clock rather than on behaviour. Replace the three hand-picked 30_000 values with one documented constant at 120_000, which keeps a genuine hang visible while clearing the observed spread. The global default stays at 10s so nothing else is loosened. |
||
|
|
80d7ae14ba |
fix(ci): fail fast when the deploy dir has no DATABASE_URL
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 34s
CI / tests-unit (push) Successful in 1m42s
CI / tests-integration (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m37s
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy directory's .env. When that variable was missing the deploy had already built an image and run the browser gate before pnpm db:migrate aborted on an empty value, so a release was paid for in full and then thrown away. Check for the variable right after the .env is copied, before the build, and say plainly that the live release was not touched. The deploy test fixture gains a DATABASE_URL so it mirrors a working deploy directory instead of the broken one. |
||
|
|
944527e078 |
feat(nitro-cleanup): dedupe FurnitureData and clean dangling figure entries
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m58s
CI / tests-unit (push) Successful in 2m16s
CI / tests-ui (push) Successful in 3m8s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m30s
Add a gamedata cleanup to the Nitro Cleanup panel: a read-only preview plus an apply run that dedupes FurnitureData classnames and removes rows and figure entries that reference nothing. Three passes run in a fixed order, because cleanFigureMap has to precede cleanFigureData: dropping the part that points at a set is what makes that set unreferenced. A pass refuses to write when it would delete more than maxRemovals rows (default 500) and reports the reason, a wrong asset directory otherwise turns every row into an orphan and one call would empty the file. Passes that would act on empty input (no libraries, no sets) treat that as a missing file rather than as a reason to delete everything. Every write copies the file to a timestamped backup first, so a pass that turns out to be wrong can be undone by hand. The plan reads FurnitureData once and hands the parsed copy to both furniture passes; the file is tens of megabytes in a real deployment. |
||
|
|
d176fad4da |
fix(nitro): repair stale meta.image in bundles that are already lossless
A bundle whose texture is already VP8L was returned untouched, so a stale spritesheet.meta.image survived the normalisation and the client could not find the texture member. Rebuild the archive in that case and reuse the existing VP8L bytes instead of decoding them again. |
||
|
|
9ee22db8ba |
fix(nitro): normalise attached and recovered .nitro bundles too
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 40s
CI / tests-unit (push) Successful in 1m45s
CI / tests-integration (push) Successful in 1m55s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m53s
Two more .nitro entry points in the main import path still wrote the supplied buffer verbatim: an attached `providedNitro` and a bundle pulled back by `resolveMissingNitro`. Both are real furniture imports, so they could still land a PNG texture while the SWF, clone and upload paths produced WebP. Route both through the same normalisation, falling back to the original bytes with a warning if the texture cannot be decoded. |
||
|
|
b26e2e0de4 |
feat(nitro): normalise hotel and uploaded bundles to WebP Lossless
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 2m17s
CI / tests-unit (push) Failing after 2m29s
CI / tests-ui (push) Successful in 3m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Importing from a hotel wrote the downloaded .nitro to disk untouched, so official PNG textures stayed PNG and only SWF imports ended up as WebP. Every Studio import should produce the same format regardless of where the bytes came from, so both clone and upload paths now run the bundle through toWebpLosslessBundle. The helper decodes the texture and re-encodes it with the same VP8L options the SWF importer uses, so the artwork round-trips bit-for-bit, and lets createNitroBundle relabel the member and repair the meta.image pointer. A bundle that is already lossless WebP is returned untouched, making the operation idempotent and safe to run on re-import. A colour variant that shares a library keeps the member base name it arrived with. A texture that cannot be decoded keeps its original format with a warning instead of failing the import: the bundle is valid, and losing a furniture item over a codec edge case is worse than a slightly larger texture. |
||
|
|
306e209e29 |
fix(nitro): normalise uploaded bundles so meta.image matches the texture
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 39s
CI / tests-unit (push) Successful in 2m3s
CI / tests-integration (push) Successful in 2m7s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m50s
CI / deploy (push) Failing after 1m45s
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a third-party tool that ships a WebP member while still pointing spritesheet.meta.image at a .png was accepted and stored as-is. The client resolves the spritesheet through that pointer, so the result was a file that validates fine and then renders nothing. Re-write the bundle through createNitroBundle on import, which labels the member from the actual bytes and repairs the pointer. No texture is re-encoded, so the bytes stay identical, and the member keeps the base name it arrived with so `chair*2` colour variants that share the `chair` library are not renamed. |
||
|
|
17de94d984 |
feat(nitro): convert imported SWF bundles to WebP Lossless
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m59s
CI / tests-integration (push) Successful in 2m19s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m56s
CI / deploy (push) Successful in 2m18s
Newly converted .nitro bundles now store their spritesheet as WebP VP8L instead of PNG, so imports land much smaller without changing a single pixel. The texture member and spritesheet.meta.image are both labelled from the actual bytes, never from a caller's assumption. - encode through sharp with lossless and exact, so colour hidden under alpha 0 survives; this mirrors ImageSharp's TransparentColorMode.Preserve - detect PNG/WebP by magic bytes and reject anything the client cannot render, on create, download and upload paths - keep the source format when deriving size-32 sheets, scaling composites and editing metadata, so existing bundles are never silently rewritten - report fidelity in the studio: the compression panel re-encodes with the same options the importer uses, so it cannot drift and invent false warnings, and shows PNG/WebP size estimates convertSwfToNitro and buildSpritesheet are now async, so the worker, the main-thread fallback and every import call site await them. PNG stays supported for existing bundles and icon sidecars are untouched. |
||
|
|
420210ffa0 |
fix(build): make the production build pass, and stop it eating 20GB
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m39s
CI / tests-unit (push) Successful in 1m43s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m31s
CI / deploy (push) Successful in 2m17s
`next build` had never completed on this host, so three real defects were
sitting in the tree untested. All three are now fixed and the build is green.
- The build was not memory-bound the way it looked. Turbopack's builder reached
20.5GB RSS and died, and raising `--max-old-space-size` could never have
helped: that flag caps the V8 heap, while the 20GB sat in Turbopack's own Rust
allocator. The first symptom was misleading because the process doing the
allocating is a grandchild of `npx`, so watching the direct child shows a
95MB shim the whole time. Building with `--webpack` puts the build back under
the JS heap, where the flag actually applies: peak 5.9GB, 150s, exit 0.
- withAdmin's second parameter was typed `{ params?: ... }` and given a `= {}`
default, which made it optional and `RouteContext | undefined`. Next's
generated route types assert that argument against `ParamCheck<RouteContext>`
and reject it, across 113 route files. `tsc --noEmit` cannot see this, because
Next only adds `.next/types` to the project during a production build — so the
type check that everyone runs locally was structurally incapable of catching
the only type error that blocks a deploy. `params` is now required, which is
also what the code already assumed: it is awaited with no guard. The 35 test
call sites that invoked a handler with one argument now pass a real context,
and the await got a guard so a direct internal call cannot turn a missing
context into a 500.
- `src/app/api/admin/import/furni/route.ts` re-exported `ensureDirectories` and
`importSingleFurni` for "backward compatibility" that nothing used; the batch
route imports from `@/lib/services/furni-import` directly. Next rejects any
value export from a route module that is not an HTTP verb or config, so this
had been breaking the build for as long as it existed. Removed.
- `isomorphic-dompurify` builds its server-side DOM through jsdom. Bundled, that
pulls jsdom's `browser/default-stylesheet.css` into the server chunk, where the
path no longer resolves, and page-data collection dies with ENOENT on every
page that sanitizes HTML. Marked external so Node resolves it from
node_modules and the standalone tracer includes it.
The remaining build warning is a pre-existing circular dependency between
chunks that share the webpack runtime. It costs hash reuse, not correctness, and
is left alone rather than churned here.
Verified: build exit 0, 276 static pages generated, 3223 tests pass, tsc and
biome clean.
|
||
|
|
155bf750c3 |
fix(cache): bound grace windows, cap render queues, and drop the useless estimate
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m55s
CI / tests-unit (push) Failing after 2m14s
CI / preflight (push) Skipped
CI / tests-ui (push) Failing after 36m38s
CI / deploy (push) Skipped
Follow-up to
|
||
|
|
f81b114b69 |
perf(cache): single-flight avatar renders, cacheable public reads, cheap row counts
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m30s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m7s
Three separate things that were each costing more than they needed to on the hot path. - Single-flight avatar renders. The disk cache was checked first and a miss went straight to the upstream, with nothing shared between callers, so a page requesting dozens of avatars at once turned N concurrent requests for one figure into N renders. A render is the most expensive operation this app does, and the duplication happened exactly when the cache had nothing to offer. Eight concurrent requests now cause one render instead of eight. The map lives on globalThis because Next can evaluate the module more than once per process, and two copies would each start their own render. - Let public read-only routes be cached by a shared cache. Every JSON response was `cache-control: no-store`, so a CDN in front of the app could not answer any of it and every request reached the origin. publicCacheControl() opts a route in with s-maxage and stale-while-revalidate, using the same TTL as the server-side cache so the two layers cannot disagree. The default stays no-store: most routes here are personalised, admin-only or auth-dependent. /api/badges/leaderboard is deliberately left alone because it returns per-viewer rank entries to signed-in callers. Note this only takes effect once a cache rule exists for /api/* at the CDN, or the explicit `cache: "no-store"` is dropped from the client fetches (24 files do that today, including the /api/online poll). The headers alone are inert until one of those happens. - Take the homepage row counts from the storage engine estimate instead of COUNT(*), which walks an index and gets slower as the tables grow. A missing or zero estimate falls back to the exact count rather than ever showing a wrong zero. The online count stays exact: it is an indexed read over a small subset and a few seconds of drift reads as broken rather than approximate. The counters move into one module because the homepage and the boot warm-up populate the same cache keys, so two implementations would race to write different values into the same entry. 3223 tests pass. |
||
|
|
203399aab7 |
fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m33s
CI / deploy (push) Successful in 2m43s
The in-process cache was a FIFO of 500 entries that was never touched on a read, so a key polled on every request could be evicted by an unrelated burst of dynamic keys. That looked exactly like the cache being cleared at random, and it is what made the site fall back to the database unpredictably. - Evict least-recently-used instead, and raise the default budget to 2000 (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key only leaves when a hotter one takes its place. - Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window lives on the entry, so one call site opting in protects every reader of that key. A failed background refresh keeps serving the last good value instead of falling through to the origin, and is reported once rather than per read. - Invalidate across processes. invalidateKey() now clears memory, deletes the Redis key and publishes a signal, so a value written by one process is no longer served stale by the others for the rest of its TTL. A failed Redis delete no longer skips the broadcast. - Guard against a refresh that started before an invalidation writing its outdated result back into the cache. - Read the news revision at most once a second per process instead of on every call, with a pub/sub signal to drop the local copy when it rotates. A Redis outage now degrades to the in-process cache rather than to no cache at all. - Warm the hot public keys on boot, so the first visitors after a deploy do not each pay for a miss. - Count hits, misses, stale serves, errors and evictions per key, exposed at GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and a dead origin all look identical from the outside. - Enforce the imaging cache budget for real: records are .img/.json pairs, so the old cap counted files and never removed anything while entries were fresh. Sweeps are throttled per directory and prune to a low-water mark. - Cap the JWT version map, and stop per-test scratch roots from littering the runtime imaging cache. Public read-only endpoints get grace windows; admin, account and auth data deliberately stays fresh. Redis TTLs get a little jitter so keys written together no longer expire together. 3209 tests pass. next build could not be verified on this host: the optimized build is OOM-killed before prerender, so this has not run in a real Next runtime yet. |
||
|
|
f490fcc9da |
fix(imaging): stop caching fallback renders
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m40s
CI / tests-unit (push) Successful in 1m53s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m59s
A fallback render drops the requested effect and is only a degraded stand-in, so writing it to the 30 day disk cache kept serving the worse image long after the local renderer recovered. Cache primary renders only and let the next request pick up the real render. |
||
|
|
fe5a7a6185 |
fix(imaging): keep avatars rendering, cacheable and reliably timed
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m40s
CI / tests-unit (push) Successful in 1m51s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m10s
Effect renders need a little over 4s, which the 4s primary timeout cut off, so every avatar with the default effect fell through to an unreachable public fallback and rendered as a placeholder. Raise the primary budget above the observed render cost and shorten the fallback budget. Also stop the proxy from stamping no-store over the avatar and media responses, so browsers keep the long-lived Cache-Control the route already sends, and recreate the imaging cache directories with the container user on every deploy, since root ownership made those cache writes fail silently. |
||
|
|
7f6febf906 |
fix(security): drop URLhaus feed, validate CIDR ranges, pass unknown client IPs
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m46s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m40s
|
||
|
|
84d53139a9 |
feat(security): opt-in local CrowdSec LAPI bouncer on the Docker engine
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m37s
CI / tests-integration (push) Successful in 1m55s
CI / tests-ui (push) Successful in 2m23s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m38s
|
||
|
|
3e1a3f92c8 |
feat(security): recovery alerts, gate-block sharing, rolling-window burst and admin breakdown for CrowdSec
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m42s
CI / deploy (push) Successful in 2m3s
|
||
|
|
301edd2c9a |
feat(security): ops alerts, shared backoff, atomic quota and daily stats for CrowdSec
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m36s
CI / tests-unit (push) Successful in 1m40s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m28s
CI / deploy (push) Successful in 2m3s
Add an alerting/stats layer over the existing CrowdSec integration: - New crowdsec-alerts.ts: cooldown-gated ops alerts (Redis NX lock, TTL from HEALTH_ALERT_COOLDOWN_MIN) fanning out through the app's sendAlert service. Raised for daily quota exhaustion, block bursts (5-min window past CROWDSEC_ALERT_BLOCK_BURST), and signal-push failures. - New crowdsec-stats.ts: daily counters (lookups/blocks/reports/report_fail) in Redis with a 14-day reader for the admin panel. - Shared 403/429 backoff: the pause marker now lives in Redis (crowdsec:backoff-until) so every instance honours it, not just the process that hit the limit. - Atomic quota reservation: INCR-before-call with self-rollback on overshoot, so concurrent instances can never slip calls past the daily ceiling. - Admin anti-DDoS page gains a last-14-days activity table next to the quota bar. |
||
|
|
5e4fc9ab59 |
feat(security): give back to CrowdSec and harden the CTI budget
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m34s
CI / tests-unit (push) Successful in 1m36s
CI / tests-ui (push) Successful in 2m22s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m53s
- Bound the in-process verdict cache (FIFO eviction at 2000 entries) so a
flood of distinct bucket-tripping IPs cannot grow it without limit.
- Record block metadata (reputation, score, behaviors, category, TTL) in
antiddos:block:meta:{ip}, surfaced as the reason in the admin block list;
unban now also clears the metadata and report locks.
- Track daily CTI enrichment usage in Redis (crowdsec:usage:{date}); warn
once at 80% and pause lookups until tomorrow at CROWDSEC_CTI_DAILY_QUOTA
(default 10000, 0 = unlimited) so a via-spread DDoS cannot burn the plan.
- Add opt-in signal push to the CrowdSec community (CAPI watcher): stable
auto-generated 48-char machine_id/password pair persisted in Redis (or via
env), one-time registration, cached JWT login, optional Console enrollment,
and POST /v3/signals with a ban decision, deduped per IP. Never throws and
reports last status to the admin panel with a verify action.
- Admin page: quota usage bar, reporting status/verify channel, and CrowdSec
block reasons in the active-blocks list.
|
||
|
|
f32a6dadd0 |
feat(security): auto-block repeat offenders via CrowdSec community reputation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Failing after 17s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- new crowdsec-api lib: CTI lookup (GET /smoke/{ip}, freemium x-api-key), verdict parser with false-positive veto, 1h Redis + in-memory verdict cache, NX lock dedupe, 403/429 backoff; writes only the shared antiddos:block:{ip} key (value "crowdsec") and never touches Cloudflare
- gate fires it fire-and-forget for IPs that already tripped a rate bucket, so known-bad IPs are hard-blocked before the local maxViolations threshold
- runtime config: crowdsecAutoBlock toggle, score threshold (0-5, default 4), block TTL (default 24h); boot defaults CROWDSEC_AUTO_BLOCK_ENABLED / CROWDSEC_BLOCK_SCORE / CROWDSEC_BLOCK_TTL_SECONDS
- admin panel: CrowdSec stat card, verify-connection action, score/TTL settings, CrowdSec source badge in the blocked-IPs list
- credentials live in env only (CROWDSEC_API_KEY); block is enforced per-request via proxy on the resolved X-Forwarded-For / CF-Connecting-IP
- tests: crowdsec-api unit suite + ddos-guard integration suite (early-block, threshold, cache dedupe, backoff)
|
||
|
|
6264f9fb20 |
test(security): make Cloudflare block tests deterministic under CI Redis
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m39s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m31s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m1s
cloudflare-api unit tests drove the real Redis connection when REDIS_URL was set (CI), causing cross-test bleed. Mock @/lib/redis with an in-memory fake identical to the gate integration test. |
||
|
|
4479753160 |
feat(security): mirror anti-DDoS blocks to Cloudflare edge via API
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Failing after 1m40s
CI / preflight (push) Skipped
CI / tests-ui (push) Successful in 2m28s
CI / deploy (push) Skipped
- gate creates a zone IP Access Rule (block) for proxied offenders that hit the block threshold, deduped until the tiered block expires - cloudflare-api lib: verified endpoints, create/delete/verify/list helpers, Redis-backed tracking + 30s TTL sweep (instrumentation worker + admin render) - runtime toggle cloudflareAutoBlock in antiddos config; boot default CLOUDFLARE_AUTO_BLOCK_ENABLED - admin panel: Cloudflare edge-blocks card with verify + remove-rule actions; unban also lifts the edge block - credentials live in env only (CLOUDFLARE_API_TOKEN / CLOUDFLARE_ZONE_ID) |
||
|
|
f0c27eb815 |
feat(security): Cloudflare-aware IP trust and admin-tunable anti-DDoS
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m50s
CI / tests-unit (push) Successful in 1m52s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m32s
- resolveClientIp: trust cf-connecting-ip only behind cf-ray/cdn-loop, use nginx x-real-ip otherwise (anti-spoof) - antiddos-config: Redis-backed live config (antiddos:config) with 30s cache, 13 ANTI_DDOS_* env vars - ddos-guard: consume tunable rates/tiers via getAntiddosConfig - admin panel at /admin/devops/antiddos (save/reset/unban actions, PERMS.SETTINGS_VIEW) - register new admin page in housekeeping migration matrix (146 -> 147) |
||
|
|
fd4d0fa1cb |
feat(security): harden anti-DDoS gate with scanner triage, tiered blocks and in-process global halt
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m0s
|
||
|
|
98a184953a |
feat(security): add Redis-backed app-layer anti-DDoS rate limiting to proxy
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
|
||
|
|
b13b3a50ff |
security: switch default hashing to Argon2id, fix tests
- hashPassword now uses Argon2id (memory-hard, GPU-resistant) via hash-wasm - verifyPassword checks both Argon2id and bcrypt - Legacy hashes (bcrypt, argon2, md5, sha1, sha256, sha512, combined, salted) auto-migrate to Argon2id on successful login - Updated all password tests to expect Argon2id format - Register validation: min 12 chars, max 128, upper+lower+digit+special required - Username restricted to [A-Za-z0-9_-], reserved names blocked - Disposable email domains blocked - Fixed parameter names for hash-wasm argon2id API (memorySize, iterations, parallelism, hashLength) |
||
|
|
ac60a867d9 |
security: harden authentication (register/login)
- Username: restrict to [A-Za-z0-9_-], block reserved names (admin, mod, root, etc.), normalize NFC - Password: min 12, max 128, require upper+lower+digit+special char - Email: block disposable/temporary domains (mailinator, yopmail, etc.) - Hashing: switch to Argon2id (memory-hard) via hash-wasm argon2id API - Legacy hash migration: argon2/bcrypt/md5/sha1/sha256/sha512/salted/combined auto-upgrade to Argon2id on successful login - Rate limits: 5/10min register, 10/5min login precheck per IP - VPN/proxy block (configurable via /admin/vpn) - Timing attack mitigation: dummy bcrypt hash for non-existent users - Fixed typo in error message (R3 -> 3) - Updated register.test.ts to match new validation rules |