Commit Graph
1169 Commits
Author SHA1 Message Date
openhands 5b2eb91c5c fix(ops): stop a compose replica from blocking the blue/green release
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m57s
CI / tests-unit (push) Successful in 2m3s
CI / tests-ui (push) Successful in 2m49s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m48s
The deploy failed after the build, the migrations and the browser gate:
"Port 3002 is already in use". The holder was `epicnext-cms`, a compose
replica of release 6bffc537 that the daily scripts/docker-update.sh cron
had recreated at 03:30 with restart=unless-stopped. nginx serves the green
slot on 3003, so that replica was squatting the blue slot the next
candidate needed, and live traffic never noticed.

It got there because the updater's CI-ownership guard only tested
epicnext-cms-app. After a cutover to the green slot that container is
stopped, renamed and deleted, so the guard stopped firing while the host
stayed CI-managed.

- scripts/docker-update.sh: refuse a compose deployment on a CI host by
  checking both slot containers and the nginx upstream, which is the only
  thing that still marks the host as blue/green while a slot is idle.
- scripts/ci-deploy.sh: retire a compose replica of this checkout from
  the candidate port before starting the candidate, so a stray replica
  can never block a release again. Never a slot container, never the port
  nginx serves; anything else still fails loudly in assert_port_free.
- Tests cover both directions: a squatting replica is removed and the
  release lands, a replica on the live port is left alone.
2026-10-05 21:13:49 +02:00
openhands 6c3d81920e fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than
the code. Each one had a signature that looked like a network or permissions
problem and was actually a configuration or ordering bug.

jobs-worker never ran

`import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM
evaluates a module's imports in source order and the first import reaches
`@/env`, which validates process.env at import time. The ZodError on
DATABASE_URL therefore fired before load-env ever executed, so the worker
could only start from a shell that had already exported the configuration.
Nothing supervised it either, so scheduled articles, catalog export, JAR and
database backups, disk alerts and the ops health probe have all been dead;
`cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and
added deployment/systemd/cms-jobs-worker.service with Restart=always.

The JAR backup additionally pointed at './emulator/Arcturus.jar', which does
not exist and would go stale on the next emulator upgrade. resolveEmulatorJar
now accepts a file, a directory or a wildcard and picks the newest JAR, the
same way emulator.service picks its build, and reports an unresolvable path
once instead of logging an opaque copyFile ENOENT every night.

/api/health answered 200 with the database down

The route documented this as intentional, and ci-deploy.sh worked around it
by grepping the body for '"database":true'. The container healthcheck did not,
so Docker reported containers healthy while every page 500'd. The status is
now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis
and the emulator deliberately do not fail the container — both have in-process
fallbacks, so failing them would trade a slow site for an outage.

The runtime had no V8 heap cap

NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its
heap from host memory (23.5 GB) while the container was limited to 4 GB, so
the kernel OOM-killed the process mid-request — the same failure mode as the
14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit
(v2 with a v1 fallback) and sets 70% of it, respecting an explicit override.

Storage ownership was only repaired for one path

ci-deploy.sh chowned storage/imaging and nothing else, so
storage/catalog-git/hotel-status.json kept coming back root:root and
/api/admin/catalog/status kept throwing EACCES. All eight writable storage
paths are repaired now. The silent-failure mode is the reason this mattered:
these writes sit inside try/catch, so a wrong owner looks like a slow page
rather than an error.

nginx: robots.txt was a guaranteed 404, and TLS never resumed

`index index.html` without a `root` left every try_files resolving against
/etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES
on each stat and nginx logs a failed stat at crit, which is where 149 crit
lines per scan came from. robots.txt answered from that same broken location,
so crawlers were pointed at a file they could never read while sitemap.xml
kept advertising it. Added `root`, proxied robots.txt to the CMS, added
ssl_session_cache (there was no session resumption at all), and set
Restart=on-failure in a systemd override, since the packaged unit ships
Restart=no and nginx is the only thing serving the site.

Verified against the running host: 3379 tests, typecheck and biome clean,
nginx -t passes, health returns 200 with every check green, and the worker has
run for hours at NRestarts=0 with a heartbeat refreshing each minute.
2026-10-05 20:25:22 +02:00
openhands 108c6ce03d fix(ci): make the lint gate fail for real and stop byparr leaking disk
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 2m3s
CI / tests-unit (push) Successful in 2m18s
CI / tests-ui (push) Successful in 3m6s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 3m14s
The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.

Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
  (key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
  function-declaration handlers, with the reasoning that dropping them broke
  the tree and save-on-Ctrl+S once already (704e3363)
- scope the remaining noArrayIndexKey / useExhaustiveDependencies exemptions to
  the three files that need them, in biome.json instead of scattered comments

Storage, on a host that had grown to 81% disk:
- byparr starts a Firefox per request and never removes the profile it leaves in
  the container's writable layer. With no volume mounted, nothing else reclaimed
  it: 716 profiles / 6.8 GB in two days, ~1.7 GB/day. docker-prune.sh now removes
  orphaned profiles, identifying live ones by the open fd in /proc/<pid>/fd rather
  than by age, because browsers stay warm for ~27 hours here — longer than the
  leak window, so no age threshold can be both safe and useful.
- bound the build cache properly: buildx treats --max-used-space and --filter as
  mutually exclusive, so passing both silently dropped the 4 GB cap and the cache
  reached 49 GB.
- escalate to the emergency prune when / drops below 8 GB free, so the bound holds
  even if the schedule stops.
- clear multi-GB tmp_pack files left behind by a gc that was OOM-killed
  mid-repack; git only removes those on the next successful gc.
- make setup-cron.sh append instead of replacing the crontab (`crontab -`
  overwrites the whole file, which had been dropping the other scheduled jobs),
  and run the prune daily rather than weekly to match the leak rate.

Volumes are still never pruned: mariadb-turbo-data is a database.
2026-10-05 17:24:12 +02:00
openhands 6bffc53779 refactor(auth): merge the duplicate login form and localize the auth screens
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 1m8s
CI / tests-integration (push) Successful in 1m53s
CI / tests-unit (push) Successful in 1m59s
CI / tests-ui (push) Successful in 2m42s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 4m33s
`home-login-form.tsx` and `login-form.tsx` were two ~240-line near-identical
components. Delete the former and give `LoginForm` a `variant` prop:

- `variant="page"`   sr-only labels plus the register/forgot footer (/login)
- `variant="compact"` visible labels, no footer (homepage sidebar)

Field ids now come from `useId()`, so the two usages can never collide, and the
hardcoded "Show"/"Hide"/"Loading" strings are translated.

Localization of the login and register screens:

- `home-login-form.tsx` was entirely hardcoded English.
- `passwordStrength()` returned hardcoded "Weak"/"Fair"/"Good"/"Strong".
- `register.ts` returned only English strings. It now returns a
  locale-independent `code` next to the message, and the form renders
  `t(code)` with the English string as a fallback.
- Backfilled the new keys across all 25 locales, plus the login/register
  strings that were still English in most of them. `ar`, `fi` and `ja` had
  their entire login/register namespace in English and are now filled in.
  Locale parity stays at 0 missing keys, as `i18n:check` requires.

Copy that did not match the enforced rules: the UI advertised "min 8 chars"
(EN) / "min 6 tekens" (NL) while registration requires 12 characters plus an
uppercase, a lowercase, a digit and a special character. Corrected in every
locale. `password-reset.ts` enforced only 6 characters and is raised to 12 to
match registration.

Accessibility: `login-form.tsx` had no `<label>`, no `id` and no `required` on
any field. All three are now present, and error banners are announced with
`role="alert"`.

Adds `src/i18n/auth-messages.test.ts`, which asserts every `RegisterErrorCode`
resolves to a non-empty message in all 25 locales; verified it fails when a key
is removed. The existing register tests now also assert the error `code`.
2026-10-04 18:50:23 +02:00
openhands 7f07c111ac perf(studio): load motion's minimal entry instead of the full component library
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Successful in 1m52s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m34s
/admin/studio/furni sat at 94.9% of its initial-JS budget (427436 of
450560 gzip bytes), so the next feature would have broken the build. Of the
98228 gzip bytes unique to that route, a large part is framer-motion.

This file uses motion twice, for one thing: a 150ms opacity fade on the result
pane when viewMode changes. Importing `motion/react` to get it pulls in
framer-motion's complete component library — 73 internal modules — plus its
render components, drag/gesture and projection code, none of which is
rendered here.

`motion/react-m` ships only the element factories: 2 internal modules, and the
same initial/animate/transition props, so the fade is unchanged. It exports the
elements flat rather than under a `motion.` namespace, so the import becomes
`div as Mdiv` and the two JSX tags are renamed to match.

I could not measure the resulting bundle here: the local build is OOM-killed
(exit 137) with the running containers on the host, so the actual saving is
unverified. The CI build reports it in build-reports, and the number in this
commit message should be read as a hypothesis, not a measurement.

Verified: typecheck clean, lint clean, and the 10 studio UI tests pass —
including the pane and navigation specs that exercise the view switch.
2026-10-03 19:21:02 +02:00
openhands 8218039c64 test(live): stop the live suites inheriting the production database
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m57s
CI / tests-integration (push) Successful in 2m1s
CI / tests-ui (push) Successful in 2m51s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m42s
Seven suites read .env with a bare `process.env[key] = value`, which
overwrites whatever the shell already set. That made the DATABASE_URL from
the production .env authoritative, so a single environment variable was
enough to aim them at the live hotel database:

  RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts

Three of those suites then repair the catalog in place: catalog-audit-repair-live
and catalog-repair-direct-live rewrite catalog_items and delete duplicate
classnames, and clone-bulk-import-live bulk-imports every cloneable item. None
of that is undoable, and nothing in their output said the target was
production rather than a sandbox.

Added src/test/live-env.ts with one shared loader, and pointed all seven suites
at it:

- Values already in the real environment win, so an explicit DATABASE_URL on
  the command line is always respected.
- DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the
  production one from .env.
- Anything that is not loopback is treated as production and redirected.
- Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning
  saying the suite repairs the catalog.

Tests in src/test/live-env.test.ts run the loader against a temporary .env so
the real project file is never read, and cover the redirect, the shell
override, non-loopback detection, the opt-in and quote stripping. A second
block asserts each of the seven suites no longer contains an inline
`process.env[...] =` assignment. Verified four of them fail against the old
loader.

This does not enable the suites; they stay gated behind their RUN_* flags.
It only removes the possibility of them silently hitting production.

Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean.
2026-10-03 19:03:28 +02:00
openhands 8ee144745a fix(deploy): stub ss in the deploy harness and cover the port-conflict path
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m47s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m41s
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.

The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.

Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.

Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.

Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
2026-10-03 18:30:00 +02:00
openhands 704e33638f fix: restore six useEffect dependencies removed while silencing lint
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m38s
CI / tests-integration (push) Successful in 1m40s
CI / tests-ui (push) Successful in 2m24s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m58s
The previous commit dropped biome-ignore comments to clear
useExhaustiveDependencies diagnostics and, in doing so, also deleted the
dependencies themselves. Six components were left with effects that no longer
react to the state they read. Every one of these is a real behaviour
regression, not a lint preference:

- health-check-client: checkEmulator is a function declaration, so it gets a
  fresh identity each render. As an effect dependency that re-fires the effect
  after every setState, polling /api/admin/devops/health in a loop. Wrapped in
  useCallback so the identity is stable.
- article-recovery: reload restarts the autosave timer for the "Retry recovery"
  button. Without it in the deps that button is a no-op. The counter had been
  renamed to _reload to satisfy the unused-variable rule.
- catalog-integrity-panel: same pattern; refresh starts a new read-only scan,
  so the rescan control did nothing.
- catalog-search: refreshKey re-runs the query after a bulk edit, so results
  were not refreshed after catalog edits. The selection-reset effect also lost
  catalogType, so switching catalog no longer cleared the selection.
- catalog-image-picker: dropped debounced (the search term) and name (the
  error reset), so image search and error state no longer reacted to input.
- icon-picker: dropped iconImage, so a failed load left the placeholder on the
  next icon too.

Each restored dependency carries a biome-ignore with the reason it is
load-bearing, so the diagnostic can be re-derived instead of silently
disappearing again.

Verified: typecheck, lint clean on all six, unit 3315 passed, integration 20
passed, UI 72 passed / 2 skipped.
2026-10-03 18:07:56 +02:00
openhands eddb7edea4 fix: make all CI jobs pass (integration, ui) and restore prefix dialog reset
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m34s
CI / tests-unit (push) Successful in 1m35s
CI / tests-ui (push) Successful in 2m20s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m37s
Three failing test suites blocked CI. All three were test defects, not
application bugs.

Integration tests (integration/database.test.ts)
------------------------------------------------
The suite set NODE_ENV=test, which makes cache.cached() short-circuit both
its Redis read (src/lib/cache.ts:226) and its write (:249). A suite whose
stated purpose is exercising the real Redis path therefore never touched
Redis. Switched to NODE_ENV=development, the only non-production value
src/env.ts accepts, so the shared-cache code paths are genuinely covered.

Three assertions then needed correcting for real Redis semantics:

- `await cache.cached(...)` followed by `.resolves` can never hold: await
  yields a value, not a Promise. Assert the value directly.
- A cached negative result is stored as the JSON encoding of null, so
  `redis.get(key)` returns "null", not null.
- The news negative-cache key does not exist at all, so `ttl()` returned -2.
  Now that the write path is live the key is created and the TTL assertion
  holds as originally written.

UI tests (src/app/admin/prefixes/prefix-dialog.tsx)
---------------------------------------------------
The form-reset effect had `isOpen` removed from its dependency array. The
component returns null when closed, so the effect only ever ran on mount:
reopening the dialog no longer cleared the fields and a dismissed-but-
unsaved edit reappeared. Two tests in e2e/ui/unsaved-changes.spec.ts caught
this. Restored the dependency and documented why it is load-bearing.

The remaining edits in this branch drop stale biome-ignore comments that
suppressed useExhaustiveDependencies and noArrayIndexKey diagnostics. Where
the suppression had been load-bearing for behaviour, the underlying
dependency is now listed explicitly rather than silenced.

Verified: check (toolchain, audit, lint, i18n, typecheck), unit 3315
passed, integration 20 passed, UI 72 passed / 2 skipped.
2026-10-03 17:02:49 +02:00
openhands 30ff970c38 chore: upgrade to pnpm v12, update dependencies, and fix msw v3 typescript types
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
2026-10-02 21:59:39 +02:00
openhands f99980052b perf: optimize cache layer for speed and stability
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 34s
CI / tests-ui (push) Failing after 33m56s
CI / tests-integration (push) Failing after 33m57s
CI / tests-unit (push) Failing after 33m57s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- Remove random TTL jitter to prevent unpredictable cache drops
- Add deterministic LRU eviction with proper entry cleanup
- Improve cache deduplication to prevent duplicate computations
- Skip Redis I/O during tests for faster, more stable execution
- Optimize depth calculation in catalog tree nodes
- Maintain backward compatibility and full test coverage (3331 passed)
2026-10-02 17:16:03 +02:00
openhands a6cc3cafa9 fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached
twice and nobody could reach the client.

The catalog items loader kept its own 30s TTL copy of FurnitureData.json next
to the mtime-validated cache in `furni-data.ts`. An import cleared only the
second one, so the catalog table kept serving pre-import furnidata — empty
descriptions and revisions — until the TTL ran out. The loader now reads
through `readFurniData`, which revalidates on mtime+size and is reset by
every write, so there is exactly one cache and it cannot go stale on its own.
`invalidateFurniDataCache` and its single call site are gone with it.

The client was worse: nginx served all of /gamedata/ with `max-age=604800`,
and the `cms-gamedata` purge that would have fixed it hung off the catalog Git
export, which is disabled in production. A freshly imported item was invisible
in the client for up to seven days no matter how often you imported.

- `writeFurniData` now purges the gamedata edge tag itself. One place covers
  import, batch, resync, regen, nitro-editor, translate and dedupe. It is
  fire-and-forget and swallowed at every level: a stale edge copy is bounded
  by the edge TTL, so a failed purge must never fail an import.
- nginx splits /gamedata/ by how mutable the content is: config/ gets
  `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`,
  and the content-addressed trees (c_images, album*, clothes) keep the long
  TTL. `must-revalidate` is the point — the client now revalidates instead of
  replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the
  purge still reaches them.
- A 30-minute safety-net purge in the jobs worker covers the case where
  Cloudflare was unreachable at write time.
2026-10-01 15:16:48 +02:00
openhands cede541813 fix(catalog): route item-table writes to the catalog they belong to
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m16s
The items table is shared between both catalogs, but its four mutating
actions were normal-only: moving, reordering, creating and updating a
Builder Club offer wrote to catalog_items, so a BC edit either landed in
the wrong catalog or hit an unknown column.

Pass the catalog from the table through the actions and let the server
resolve it. BC rows have no price, points or currency column, so the BC
commands strip those fields instead of rejecting them. Moving and
reordering now share one command that locks the category and writes the
table for the same catalog, and BC writes revalidate the BC route.
2026-09-30 20:06:48 +02:00
openhands e0efbef30d fix(catalog): make the Builder Club catalog read and write its own offers
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m23s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m6s
The previous commit taught bulk editing and delete-with-restore about the BC
catalog. Neither actually worked, and one of them was destructive.

`catalog_items_bc` has six columns: id, item_ids, page_id, catalog_name,
order_number, extradata. There is no price, points, currency, offer_id, limit or
membership column on it. The bulk path read and wrote columns that do not
exist, and the UPDATE was aimed at catalog_items while the SELECT came from
catalog_items_bc — so a BC category move wrote into the normal catalog. Two
tests now pin that pairing: reads and writes have to stay in the same table.

Underneath it the BC table was never being read at all. The inline editor
fetched `/api/admin/catalog/items?pageId=N` without the catalog, so opening a BC
category showed the normal catalog's offers, and the route selected BC rows
directly instead of going through the loader, skipping the furni enrichment the
table needs to render anything but a bare caption. Both catalogs now take the
same path, and the catalog is in the fetch callback's dependencies — without
that, a switch keeps reading the previous catalog's rows through a stale
closure.

Because a BC offer has no price, the editor no longer offers one. The server
refuses price, points and currency changes with a readable message instead of
letting them reach the database as an unknown-column error, and a BC bulk edit
is what it can actually be: a category move.

BC deletions also went through a bare DELETE, which made them the one catalog
mutation with no way back. They now keep their rows and hand back a restoreId
like the normal ones. The catalog is recorded in the audit target rather than
in the payload, so a restore can never put a BC row into the normal offers
table.
2026-09-30 19:17:20 +02:00
openhands cebcf440c5 feat(catalog): record bulk edits, make deletions reversible, unify the tree read
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m23s
A bulk offer edit is the catalog mutation that rewrites hundreds of rows at
once, and it was the only one writing nothing to the staff activity log: 22 of
the 43 catalog actions logged, this one did not. The entry it now writes says
what changed, not just that something did, because the log has no undo of its
own and "bulk updated 200 offers" cannot answer the question it exists for.

Deleting offers had no inverse at all. Every removed row is now kept at delete
time and the caller gets a restoreId back, so an accidental multi-select is a
click rather than a hand-edit of the table. The undo toast covers the common
case; a RecentDeletionsPanel holds the same records so a delete noticed later is
still reachable. Three refusals guard it: an id that another offer has since
taken, a category that no longer exists (which would leave an offer that sells
nowhere and shows under no page), and a delete whose restore record cannot be
written — that one rolls back rather than deleting without a way back. Reading
the audit row FOR UPDATE is also what stops two restores of one deletion from
both inserting.

sendCatalogUpdate() overwrote hotel-status.json on every write, so "which
imports reached the hotel" was answerable for the last attempt only, and a
failure two imports ago was gone by the time anyone looked. That file is now
also appended to as a bounded 50-entry tail.

The tree route carried four copies of the same page-select-plus-counts
shaping, of which the BC branches had already drifted: one counted offers
through the VARCHAR-tolerant helper, the other inline and swallowing errors.
All of it is one readPages() now, and readFullTree sends both catalogs through
one depth computation instead of delegating normal to getTreeFlat while
computing BC here — a split that left two implementations behind one function
name. getTreeFlat is gone. The BC ancestor walk also went from 20 levels to 50,
matching getAncestors, so a deeply nested catalog no longer loses its
breadcrumb.

Bulk editing reaches the BC catalog, which previously had no way to edit or
duplicate offers in bulk. The catalog is part of the operation identity now, so
replaying one request key against the other catalog is not mistaken for the
same work.

Integration tests failed to import: the next/cache mock supplied only
revalidatePath, and catalog-totals calls unstable_cache at module scope.
2026-09-30 18:19:36 +02:00
openhandsandClaude Opus 4.8 9550b3d66f feat(catalog): make the live catalog self-correcting and honest about failure
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m34s
CI / tests-integration (push) Failing after 1m34s
CI / tests-ui (push) Successful in 2m19s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The previous commit made imports update the Studio without a reload, but the
guarantee only held inside the tab that started the import and only as long as
every read succeeded. Four holes were left, and this closes them.

A session that mounted the tree before an import kept the pre-import tree for
the rest of its life, because ensureCatalogTreeLoaded() was a once-per-session
no-op. It now asks the server whether what it holds is still current. The answer
is a revision: sendCatalogUpdate() already runs after every catalog write, so it
bumps one, and clients read it on mount, on focus, on a 20s poll and from other
tabs over a BroadcastChannel. An import that finishes in another tab, another
browser or the job worker now lands here too.

A failed read used to be swallowed, which is the worst outcome available: the
rail kept showing pre-import counts as if they were current and nothing said so.
The snapshot now carries the error, the rail shows it with a retry, and the
previous tree stays on screen because stale beats empty.

Every settled import pulled the entire flat tree, which is the one payload that
grows with the size of the catalog. The revision doubles as the ETag on
mode=full, so an unchanged catalog answers 304 and the poll costs a file read.

An import could also report success for an offer the hotel will never sell: a
hidden or disabled page, an item_ids that misses the furni id, a zero amount.
importSingleFurni reads its own row back and reports each of those as a warning,
where the import report already is, instead of leaving it to surface as "the
import did not work" in the client.

Finally, the catalog items table no longer falls back to router.refresh() —
onRefresh is now required, so every mutation ends in a refresh of the caller's
own data instead of a route re-render that threw away editor state and scroll
position. useServerAction keeps its default, because 47 callers across the app
depend on it. The 750-line CatalogTree in catalog-tree.tsx was dead code that
kept its own stale tree and three more router.refresh() calls; only CatalogIcon
and LAYOUT_COLORS are still imported, so the rest is gone.

Tests: the store now covers revisions, 304s, probe failures and error recovery;
a jsdom test mounts a consumer and asserts the tree updates in place with no
navigation; the old organize-imports e2e asserted nothing about the endpoints
the code actually calls, and is replaced by one that asserts a cross-tab write
lands in the mounted categories without a reload.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-09-30 15:15:34 +02:00
openhandsandClaude Opus 4.8 28ce0f911c fix(catalog): keep the live catalog truthful after every import path
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Failing after 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The live catalog store only covered part of the import surface. A durable
job settled, a sync queue drained, a .nitro upload or a clone run left the
Studio rail and the stats bar showing pre-import numbers until the page was
reloaded, and the Catalog Manager kept a second tree that never saw writes
made elsewhere in the session.

Every one of those paths now pulls the tree again, and the refresh carries
the totals with it: importing writes catalog rows server-side, so the counts
the store holds were stale for the rest of the session.

- refreshCatalogTree shares one request between concurrent callers and queues
  a single follow-up read when a write lands mid-flight, so a burst of edits
  costs at most one extra read.
- useFurnitureJobs treats its first payload as a baseline, so a page load no
  longer replays every past import as "just settled", and hands the settled
  jobs to the callback.
- The Catalog Manager pushes its own mutations into the store and re-reads its
  active tab when the store changes.
- The 30s unstable_cache on the admin totals is now tagged and invalidated from
  every catalog write, including the import worker, so it no longer survives an
  import even across a hard reload.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-09-30 14:45:26 +02:00
openhands d73baf1458 fix(catalog): take furnidata values from the clone source for retro items
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m33s
CI / tests-unit (push) Successful in 1m33s
CI / tests-ui (push) Successful in 2m21s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m53s
The resync route rebuilt every entry from items_base plus the *official
Habbo* furnidata. A classname that only exists on a retro hotel (leet.ws
and friends) is absent from the official set, so lookupOfficialHabboFurni
returned null and the entry was written with revision 0, category
"unknown" and an empty description — even though the clone import had
those exact values available at import time from the source's own
furnidata.

- resync/route.ts: when a classname is genuinely missing from official
  Habbo, fall back to the configured clone sources. Their furnidata is
  indexed by normalized classname and each entry is coerced into the
  OfficialHabboFurniEntry shape, which is the same JSON shape, so it drives
  the existing buildFurniEntry fallbacks for revision, category, name,
  description, defaultdir, partcolors, specialtype, furniline, environment,
  rare and bc. items_base stays authoritative for id, spriteId and dims,
  and public_name still wins over the source name, matching the import.

  The index is memoized per request, not at module scope: a module-level
  cache would pin the source list for the life of the process and a source
  added later would never be picked up. fetchSourceFurnidata already caches
  per URL, so this costs one parse rather than a network round-trip.

  Disabled sources and sources that fail to respond are skipped, so an
  unreachable hotel degrades to the previous items_base-only behaviour
  instead of failing the run. Official Habbo still wins whenever it has the
  classname, so existing behaviour is unchanged for everything but the
  retro-only case.

Applies to every resync mode, so the pre-existing ?missing=1 sweep picks
this up too.
2026-09-29 16:02:58 +02:00
openhands 2f7e557d5e feat(catalog): add a Studio button to fix missing furnidata entries
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m33s
CI / tests-unit (push) Successful in 1m38s
CI / tests-ui (push) Successful in 2m26s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m1s
"Missing furnidata" was only a filter in the Studio status dropdown, so
imported items whose classname was absent from FurnitureData.json could
be found but not fixed from that screen. Only the Catalog Audit page could
repair them, and only globally.

Adds the same shape of quick action that "no nitro" already had:

- studio-client.tsx: a "N no furnidata" shortcut next to the "N no nitro"
  button that sets the missingFurnidata status filter, and a bulk "Add
  missing furnidata (N)" button for the selected rows. Both only appear
  when there is something to act on. Rows that come back repaired flip
  to hasFurnidata: true so the badges and counts update in place; rows
  the server reported in errors keep their state.
- resync/route.ts: accepts an optional { classnames: string[] } body to
  target exactly the selected rows. classnames are resolved through the
  same normalized local index the listing uses to decide hasFurnidata, so
  the rows written are the rows flagged as missing. The upsert is already
  idempotent, and RCON updateCatalog + updateItems run afterwards so the
  emulator picks the new entries up.
  Also clears the Studio furnidata cache after a write, which this route
  never did: without it the listing kept serving a stale hasFurnidata for
  up to the 30s cache TTL, so a repair looked like it had done nothing.
  PERMS is now imported from permission-slugs (identical re-export) so the
  route no longer pulls next-auth into tests.
- studio-filters.test.ts: pins the missingFurnidata branch, in particular
  that an unchecked item (hasFurnidata undefined) is not treated as missing.

The existing ?days / ?missing / ?broken / ?all modes are unchanged; the
body is only consulted when it carries a classnames array.
2026-09-29 15:48:32 +02:00
openhands 4be7eaed59 fix(catalog): never create a page that reuses a sibling's order number
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 41s
CI / tests-unit (push) Successful in 1m53s
CI / tests-integration (push) Successful in 2m19s
CI / tests-ui (push) Successful in 2m46s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m9s
The emulator complained "Sibling order 2 is used more than once" 74 times
in production, covering 19 pages across 12 parents. The cause was that
nearly every page-creation path passed orderNum 0, so each new page
collided with whatever sibling already sat at 0 or 1, and the Park
placeholder pages all shared the sentinel values 99 and 999.

The production rows themselves are repaired out of band (renumbered 1..N
per affected parent, ordered by order_num then id so the existing visual
order is preserved, plus one dangling catalog_items row whose
items_base no longer existed removed). This commit stops it recurring.

- hierarchy.ts: add nextFreeSiblingOrder(), which ignores -1 and 0 as
  the same root set, treats a missing orderNum as 0, and returns an order
  strictly above the highest sibling in use. An explicit order is still
  honoured whenever it is free, so callers that genuinely want a position
  keep it.
- page-commands.ts: createPageCommand resolves the real order through
  nextFreeSiblingOrder instead of writing the requested 0 straight through.

Note that furni-import.ts and upload-import.ts still take their order
from the furnidata catInfo.order, so two categories carrying the same
furnidata order can still collide. That is caught by the emulator audit
and repaired by fixEmulatorIssues(), but it is not prevented here.
2026-09-29 15:34:44 +02:00
openhands 9cc57cddfc feat(catalog): update the catalog live after an import, no page refresh
Organising imports, the Studio furni batch, the catalog totals and the
"import from a source" stats all used to need a full page reload, or at
best a router.refresh() that re-rendered the whole admin route, before
anything on screen reflected what the import had just written.

- live-catalog-merge.ts (new): pure tree and total arithmetic. Applies a
  delta of created pages, added offers and moved offers, recomputes depth
  for the touched subtree, bumps parent child counts and the item totals.
  Returns the input untouched when a delta is empty, so subscribers can
  bail out instead of re-rendering. Depth resolution tolerates a parent
  cycle in a dirty DB and still terminates, matching getTreeFlat.
- use-live-catalog.ts (new): one module-level store exposed through
  useSyncExternalStore, so every consumer shares a single instance without
  threading a provider through the admin layout. Deltas only apply to the
  "normal" catalog, so public and public_handlers trees stay separate.
  seedCatalogTotals() takes the first server value per mode and never
  overwrites it afterwards, so a later hard render cannot make the header
  totals jump backwards.
- actions/catalog.ts: organizeImportFurni now reports each group through
  the new OrganizedPageChange, carrying parentId, pageLayout, the icon,
  isNew and the per-source movedFrom counts, so the client can fold the
  result into the tree without reading the page back.
- organize-imports-dialog.tsx: drops useRouter and router.refresh(); the
  response is applied as a delta the moment the run finishes.
- studio-client.tsx: reads the tree from the store instead of freezing it
  with useState(initialTree), loads it on mount when empty, and refreshes
  it once a batch import settles. The batch is server-side and derives its
  import pages from furnidata, so that one path re-reads the tree via
  GET /api/admin/catalog/tree?mode=full rather than trusting the delta.
- studio/furni/page.tsx: stops calling getTreeFlat() and no longer passes
  initialTree; the store is the single source of truth for the rail.
- import-clone-client.tsx: tracks which items are already present, so
  present and clonable update per cloned row instead of only at the end.
- catalog-manager-dialog.tsx: seeds the totals once and renders the live
  values, so the header reflects an import that just ran.
- e2e/ui/fixtures/entry.tsx: drops the removed initialTree prop.
2026-09-29 15:34:36 +02:00
openhands 7697728d07 feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single
Cache-Control owner per route: the app stays the source, nginx only manages
headers, and Cloudflare stores the public API allowlist at the edge.

- deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the
  blue/green upstream snippet; config backed by scripts/nginx-sync.sh
  (idempotent install + reload, --check/--force).
- nginx serves Cache-Tag headers on the public allowlist (cms-public),
  gamedata, client and camera responses so the edge and purge stay in sync.
- src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that
  no-op unless Cloudflare is configured; scripts/cf-purge.sh and
  cf-setup-cache.sh create and purge the cache rule.
- src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls.
- Purge hooks after catalog exports (public + gamedata) and on shop, team,
  guild, photo and rare-values edits; ci-deploy purges after each release.
- src/proxy.ts excludes the imaging/images docs from the middleware matcher.
2026-09-28 21:55:18 +02:00
openhands e4f83a8036 chore: upgrade to pnpm v12, vitest v5 and resolve deprecated subdependencies
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
2026-09-27 19:42:34 +02:00
openhands d2d01141f1 style: apply Biome formatting to the catalog release test
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m49s
CI / tests-unit (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 20s
The timeout constant made the first it() line exceed the line width, so
`biome check .` failed with a format error. Reformat and confirm the
three publication tests still pass.
2026-09-27 19:26:08 +02:00
openhands c2981bd970 test(catalog): give the Git publication tests room to finish
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
These drive real git processes against a local bare remote, so their cost
is process spawns competing with every other Vitest worker. Measured on
CI they take 23-30s each, and the 30s override was crossed by 37ms, so the
run failed on wall-clock rather than on behaviour.

Replace the three hand-picked 30_000 values with one documented constant
at 120_000, which keeps a genuine hang visible while clearing the observed
spread. The global default stays at 10s so nothing else is loosened.
2026-09-27 19:20:44 +02:00
openhands 80d7ae14ba fix(ci): fail fast when the deploy dir has no DATABASE_URL
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 34s
CI / tests-unit (push) Successful in 1m42s
CI / tests-integration (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m37s
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy
directory's .env. When that variable was missing the deploy had already
built an image and run the browser gate before pnpm db:migrate aborted on
an empty value, so a release was paid for in full and then thrown away.

Check for the variable right after the .env is copied, before the build,
and say plainly that the live release was not touched. The deploy test
fixture gains a DATABASE_URL so it mirrors a working deploy directory
instead of the broken one.
2026-09-27 18:57:10 +02:00
openhands bc00ecf08c feat(i18n): complete message parity across all 25 locales
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 37s
CI / tests-unit (push) Successful in 1m49s
CI / tests-integration (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m29s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m37s
The admin.studio.nitroCleanup section (102 keys) only existed in en and
nl, so 23 locales fell back to English for the entire Nitro Cleanup
panel. The referrals and dailyRewards keys were missing from the same
23 locales, and en itself was missing 6 keys that nl had.

Add the missing keys to every locale with translations, so all 25
locales now carry the same 6063 keys.
2026-09-27 18:38:58 +02:00
openhands 944527e078 feat(nitro-cleanup): dedupe FurnitureData and clean dangling figure entries
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m58s
CI / tests-unit (push) Successful in 2m16s
CI / tests-ui (push) Successful in 3m8s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m30s
Add a gamedata cleanup to the Nitro Cleanup panel: a read-only preview
plus an apply run that dedupes FurnitureData classnames and removes
rows and figure entries that reference nothing.

Three passes run in a fixed order, because cleanFigureMap has to precede
cleanFigureData: dropping the part that points at a set is what makes
that set unreferenced.

A pass refuses to write when it would delete more than maxRemovals rows
(default 500) and reports the reason, a wrong asset directory otherwise
turns every row into an orphan and one call would empty the file. Passes
that would act on empty input (no libraries, no sets) treat that as a
missing file rather than as a reason to delete everything. Every write
copies the file to a timestamped backup first, so a pass that turns out
to be wrong can be undone by hand.

The plan reads FurnitureData once and hands the parsed copy to both
furniture passes; the file is tens of megabytes in a real deployment.
2026-09-27 17:14:30 +02:00
openhands d176fad4da fix(nitro): repair stale meta.image in bundles that are already lossless
A bundle whose texture is already VP8L was returned untouched, so a
stale spritesheet.meta.image survived the normalisation and the client
could not find the texture member. Rebuild the archive in that case and
reuse the existing VP8L bytes instead of decoding them again.
2026-09-27 17:13:52 +02:00
openhands 9ee22db8ba fix(nitro): normalise attached and recovered .nitro bundles too
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 40s
CI / tests-unit (push) Successful in 1m45s
CI / tests-integration (push) Successful in 1m55s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m53s
Two more .nitro entry points in the main import path still wrote the
supplied buffer verbatim: an attached `providedNitro` and a bundle pulled
back by `resolveMissingNitro`. Both are real furniture imports, so they
could still land a PNG texture while the SWF, clone and upload paths
produced WebP.

Route both through the same normalisation, falling back to the original
bytes with a warning if the texture cannot be decoded.
2026-09-27 16:00:44 +02:00
openhands b26e2e0de4 feat(nitro): normalise hotel and uploaded bundles to WebP Lossless
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 2m17s
CI / tests-unit (push) Failing after 2m29s
CI / tests-ui (push) Successful in 3m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Importing from a hotel wrote the downloaded .nitro to disk untouched, so
official PNG textures stayed PNG and only SWF imports ended up as WebP.
Every Studio import should produce the same format regardless of where the
bytes came from, so both clone and upload paths now run the bundle through
toWebpLosslessBundle.

The helper decodes the texture and re-encodes it with the same VP8L options
the SWF importer uses, so the artwork round-trips bit-for-bit, and lets
createNitroBundle relabel the member and repair the meta.image pointer. A
bundle that is already lossless WebP is returned untouched, making the
operation idempotent and safe to run on re-import. A colour variant that
shares a library keeps the member base name it arrived with.

A texture that cannot be decoded keeps its original format with a warning
instead of failing the import: the bundle is valid, and losing a furniture
item over a codec edge case is worse than a slightly larger texture.
2026-09-27 15:55:17 +02:00
openhands 306e209e29 fix(nitro): normalise uploaded bundles so meta.image matches the texture
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 39s
CI / tests-unit (push) Successful in 2m3s
CI / tests-integration (push) Successful in 2m7s
CI / tests-ui (push) Successful in 2m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m45s
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a
third-party tool that ships a WebP member while still pointing
spritesheet.meta.image at a .png was accepted and stored as-is. The client
resolves the spritesheet through that pointer, so the result was a file
that validates fine and then renders nothing.

Re-write the bundle through createNitroBundle on import, which labels the
member from the actual bytes and repairs the pointer. No texture is
re-encoded, so the bytes stay identical, and the member keeps the base
name it arrived with so `chair*2` colour variants that share the `chair`
library are not renamed.
2026-09-27 15:47:44 +02:00
openhands 17de94d984 feat(nitro): convert imported SWF bundles to WebP Lossless
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m59s
CI / tests-integration (push) Successful in 2m19s
CI / tests-ui (push) Successful in 2m56s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m18s
Newly converted .nitro bundles now store their spritesheet as WebP VP8L
instead of PNG, so imports land much smaller without changing a single
pixel. The texture member and spritesheet.meta.image are both labelled
from the actual bytes, never from a caller's assumption.

- encode through sharp with lossless and exact, so colour hidden under
  alpha 0 survives; this mirrors ImageSharp's TransparentColorMode.Preserve
- detect PNG/WebP by magic bytes and reject anything the client cannot
  render, on create, download and upload paths
- keep the source format when deriving size-32 sheets, scaling composites
  and editing metadata, so existing bundles are never silently rewritten
- report fidelity in the studio: the compression panel re-encodes with the
  same options the importer uses, so it cannot drift and invent false
  warnings, and shows PNG/WebP size estimates

convertSwfToNitro and buildSpritesheet are now async, so the worker, the
main-thread fallback and every import call site await them. PNG stays
supported for existing bundles and icon sidecars are untouched.
2026-09-27 15:42:21 +02:00
openhands 420210ffa0 fix(build): make the production build pass, and stop it eating 20GB
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m39s
CI / tests-unit (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m31s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m17s
`next build` had never completed on this host, so three real defects were
sitting in the tree untested. All three are now fixed and the build is green.

- The build was not memory-bound the way it looked. Turbopack's builder reached
  20.5GB RSS and died, and raising `--max-old-space-size` could never have
  helped: that flag caps the V8 heap, while the 20GB sat in Turbopack's own Rust
  allocator. The first symptom was misleading because the process doing the
  allocating is a grandchild of `npx`, so watching the direct child shows a
  95MB shim the whole time. Building with `--webpack` puts the build back under
  the JS heap, where the flag actually applies: peak 5.9GB, 150s, exit 0.

- withAdmin's second parameter was typed `{ params?: ... }` and given a `= {}`
  default, which made it optional and `RouteContext | undefined`. Next's
  generated route types assert that argument against `ParamCheck<RouteContext>`
  and reject it, across 113 route files. `tsc --noEmit` cannot see this, because
  Next only adds `.next/types` to the project during a production build — so the
  type check that everyone runs locally was structurally incapable of catching
  the only type error that blocks a deploy. `params` is now required, which is
  also what the code already assumed: it is awaited with no guard. The 35 test
  call sites that invoked a handler with one argument now pass a real context,
  and the await got a guard so a direct internal call cannot turn a missing
  context into a 500.

- `src/app/api/admin/import/furni/route.ts` re-exported `ensureDirectories` and
  `importSingleFurni` for "backward compatibility" that nothing used; the batch
  route imports from `@/lib/services/furni-import` directly. Next rejects any
  value export from a route module that is not an HTTP verb or config, so this
  had been breaking the build for as long as it existed. Removed.

- `isomorphic-dompurify` builds its server-side DOM through jsdom. Bundled, that
  pulls jsdom's `browser/default-stylesheet.css` into the server chunk, where the
  path no longer resolves, and page-data collection dies with ENOENT on every
  page that sanitizes HTML. Marked external so Node resolves it from
  node_modules and the standalone tracer includes it.

The remaining build warning is a pre-existing circular dependency between
chunks that share the webpack runtime. It costs hash reuse, not correctness, and
is left alone rather than churned here.

Verified: build exit 0, 276 static pages generated, 3223 tests pass, tsc and
biome clean.
2026-09-25 19:47:10 +02:00
openhands 155bf750c3 fix(cache): bound grace windows, cap render queues, and drop the useless estimate
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m55s
CI / tests-unit (push) Failing after 2m14s
CI / tests-ui (push) Failing after 36m38s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Follow-up to f81b114b, addressing the three ways that commit could make things
worse rather than better. All three were verified against the real database or
by breaking the test and watching it fail.

- The grace window is now capped at 120s. A window is a cushion for the TTL
  boundary, not a second TTL, but the call sites treated it as the latter: the
  5 min values/staff routes and the 10 min teams route asked for a window as
  long as or longer than their own TTL, so a single large staleMs silently
  doubled how far behind a value could be served. Nothing marked those as
  unsafe, because nothing looked wrong. The cap lives in the cache rather than
  at the call sites so no future route can reintroduce it. Routes that asked
  for less than 120s (the 10s online poll, the 20s news cache) are unchanged,
  so their intended cushion still does its job.

- A request no longer queues behind an arbitrarily old render. Sharing a render
  is what collapses a cold-cache stampede into one render, but a hung render
  used to hold up everyone who arrived after it. A newcomer past 2s now serves
  the placeholder instead of waiting, reusing the ImagerUnavailableError path
  that "both upstreams down" already takes. The caller that actually started
  the render keeps waiting, which is correct: it is the one whose image this
  is. When the join window is removed the new test hangs for the full 10s it
  was meant to prevent, which is the tail this bounds.

- The information_schema row-count estimate is gone; the counters are exact
  again. Running it against the live database: users 165, rooms 92, camera_web
  0, and the estimate was 0.00% off on all three. At 165 rows an index scan is
  cheaper than the extra round trip the estimate needed, so the optimisation
  bought nothing and traded a guaranteed-correct member count for an
  approximation that InnoDB would only make less accurate as the table grows.
  The exactness is now pinned by tests: a real zero stays zero, a database
  error propagates instead of becoming a number, and each counter counts the
  table it claims to. The module stays, because the homepage and the boot
  warm-up writing different values to the same cache key is its own bug.

The module comment records the measured numbers, because "COUNT(*) is too slow"
sounds true in the abstract and is false here.

3223 tests pass.
2026-09-25 18:55:33 +02:00
openhands f81b114b69 perf(cache): single-flight avatar renders, cacheable public reads, cheap row counts
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m30s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m7s
Three separate things that were each costing more than they needed to on the
hot path.

- Single-flight avatar renders. The disk cache was checked first and a miss
  went straight to the upstream, with nothing shared between callers, so a page
  requesting dozens of avatars at once turned N concurrent requests for one
  figure into N renders. A render is the most expensive operation this app
  does, and the duplication happened exactly when the cache had nothing to
  offer. Eight concurrent requests now cause one render instead of eight. The
  map lives on globalThis because Next can evaluate the module more than once
  per process, and two copies would each start their own render.

- Let public read-only routes be cached by a shared cache. Every JSON response
  was `cache-control: no-store`, so a CDN in front of the app could not answer
  any of it and every request reached the origin. publicCacheControl() opts a
  route in with s-maxage and stale-while-revalidate, using the same TTL as the
  server-side cache so the two layers cannot disagree. The default stays
  no-store: most routes here are personalised, admin-only or auth-dependent.
  /api/badges/leaderboard is deliberately left alone because it returns
  per-viewer rank entries to signed-in callers.

  Note this only takes effect once a cache rule exists for /api/* at the CDN, or
  the explicit `cache: "no-store"` is dropped from the client fetches (24 files
  do that today, including the /api/online poll). The headers alone are inert
  until one of those happens.

- Take the homepage row counts from the storage engine estimate instead of
  COUNT(*), which walks an index and gets slower as the tables grow. A missing
  or zero estimate falls back to the exact count rather than ever showing a
  wrong zero. The online count stays exact: it is an indexed read over a small
  subset and a few seconds of drift reads as broken rather than approximate.

The counters move into one module because the homepage and the boot warm-up
populate the same cache keys, so two implementations would race to write
different values into the same entry.

3223 tests pass.
2026-09-25 18:42:23 +02:00
openhands 203399aab7 fix(cache): true LRU, stale-while-revalidate and cross-process invalidation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m43s
The in-process cache was a FIFO of 500 entries that was never touched on a
read, so a key polled on every request could be evicted by an unrelated burst
of dynamic keys. That looked exactly like the cache being cleared at random,
and it is what made the site fall back to the database unpredictably.

- Evict least-recently-used instead, and raise the default budget to 2000
  (CACHE_MEMORY_MAX_ENTRIES). Reading a key now marks it as used, so a hot key
  only leaves when a hotter one takes its place.
- Add opt-in stale-while-revalidate (CachedOptions.staleMs). The grace window
  lives on the entry, so one call site opting in protects every reader of that
  key. A failed background refresh keeps serving the last good value instead of
  falling through to the origin, and is reported once rather than per read.
- Invalidate across processes. invalidateKey() now clears memory, deletes the
  Redis key and publishes a signal, so a value written by one process is no
  longer served stale by the others for the rest of its TTL. A failed Redis
  delete no longer skips the broadcast.
- Guard against a refresh that started before an invalidation writing its
  outdated result back into the cache.
- Read the news revision at most once a second per process instead of on every
  call, with a pub/sub signal to drop the local copy when it rotates. A Redis
  outage now degrades to the in-process cache rather than to no cache at all.
- Warm the hot public keys on boot, so the first visitors after a deploy do not
  each pay for a miss.
- Count hits, misses, stale serves, errors and evictions per key, exposed at
  GET /api/admin/devops/cache. Without it a wrong REDIS_URL, a full budget and
  a dead origin all look identical from the outside.
- Enforce the imaging cache budget for real: records are .img/.json pairs, so
  the old cap counted files and never removed anything while entries were
  fresh. Sweeps are throttled per directory and prune to a low-water mark.
- Cap the JWT version map, and stop per-test scratch roots from littering the
  runtime imaging cache.

Public read-only endpoints get grace windows; admin, account and auth data
deliberately stays fresh. Redis TTLs get a little jitter so keys written
together no longer expire together.

3209 tests pass. next build could not be verified on this host: the optimized
build is OOM-killed before prerender, so this has not run in a real Next
runtime yet.
2026-09-25 18:26:45 +02:00
openhands f490fcc9da fix(imaging): stop caching fallback renders
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m40s
CI / tests-unit (push) Successful in 1m53s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m59s
A fallback render drops the requested effect and is only a degraded
stand-in, so writing it to the 30 day disk cache kept serving the worse
image long after the local renderer recovered. Cache primary renders only
and let the next request pick up the real render.
2026-09-24 23:34:54 +02:00
openhands fe5a7a6185 fix(imaging): keep avatars rendering, cacheable and reliably timed
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m40s
CI / tests-unit (push) Successful in 1m51s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m10s
Effect renders need a little over 4s, which the 4s primary timeout cut off,
so every avatar with the default effect fell through to an unreachable
public fallback and rendered as a placeholder. Raise the primary budget
above the observed render cost and shorten the fallback budget.

Also stop the proxy from stamping no-store over the avatar and media
responses, so browsers keep the long-lived Cache-Control the route already
sends, and recreate the imaging cache directories with the container user
on every deploy, since root ownership made those cache writes fail
silently.
2026-09-24 23:22:28 +02:00
openhands 7f6febf906 fix(security): drop URLhaus feed, validate CIDR ranges, pass unknown client IPs
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m46s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m40s
2026-09-24 19:19:22 +02:00
openhands 84d53139a9 feat(security): opt-in local CrowdSec LAPI bouncer on the Docker engine
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m37s
CI / tests-integration (push) Successful in 1m55s
CI / tests-ui (push) Successful in 2m23s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m38s
2026-09-24 18:08:18 +02:00
openhands 3e1a3f92c8 feat(security): recovery alerts, gate-block sharing, rolling-window burst and admin breakdown for CrowdSec
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m42s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
2026-09-23 15:06:16 +02:00
openhands 301edd2c9a feat(security): ops alerts, shared backoff, atomic quota and daily stats for CrowdSec
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m36s
CI / tests-unit (push) Successful in 1m40s
CI / tests-ui (push) Successful in 2m28s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
Add an alerting/stats layer over the existing CrowdSec integration:

- New crowdsec-alerts.ts: cooldown-gated ops alerts (Redis NX lock, TTL from
  HEALTH_ALERT_COOLDOWN_MIN) fanning out through the app's sendAlert service.
  Raised for daily quota exhaustion, block bursts (5-min window past
  CROWDSEC_ALERT_BLOCK_BURST), and signal-push failures.
- New crowdsec-stats.ts: daily counters (lookups/blocks/reports/report_fail)
  in Redis with a 14-day reader for the admin panel.
- Shared 403/429 backoff: the pause marker now lives in Redis
  (crowdsec:backoff-until) so every instance honours it, not just the process
  that hit the limit.
- Atomic quota reservation: INCR-before-call with self-rollback on overshoot,
  so concurrent instances can never slip calls past the daily ceiling.
- Admin anti-DDoS page gains a last-14-days activity table next to the quota bar.
2026-09-23 14:45:35 +02:00
openhands 5e4fc9ab59 feat(security): give back to CrowdSec and harden the CTI budget
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m34s
CI / tests-unit (push) Successful in 1m36s
CI / tests-ui (push) Successful in 2m22s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m53s
- Bound the in-process verdict cache (FIFO eviction at 2000 entries) so a
  flood of distinct bucket-tripping IPs cannot grow it without limit.
- Record block metadata (reputation, score, behaviors, category, TTL) in
  antiddos:block:meta:{ip}, surfaced as the reason in the admin block list;
  unban now also clears the metadata and report locks.
- Track daily CTI enrichment usage in Redis (crowdsec:usage:{date}); warn
  once at 80% and pause lookups until tomorrow at CROWDSEC_CTI_DAILY_QUOTA
  (default 10000, 0 = unlimited) so a via-spread DDoS cannot burn the plan.
- Add opt-in signal push to the CrowdSec community (CAPI watcher): stable
  auto-generated 48-char machine_id/password pair persisted in Redis (or via
  env), one-time registration, cached JWT login, optional Console enrollment,
  and POST /v3/signals with a ban decision, deduped per IP. Never throws and
  reports last status to the admin panel with a verify action.
- Admin page: quota usage bar, reporting status/verify channel, and CrowdSec
  block reasons in the active-blocks list.
2026-09-23 14:24:44 +02:00
openhands f32a6dadd0 feat(security): auto-block repeat offenders via CrowdSec community reputation
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Failing after 17s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- new crowdsec-api lib: CTI lookup (GET /smoke/{ip}, freemium x-api-key), verdict parser with false-positive veto, 1h Redis + in-memory verdict cache, NX lock dedupe, 403/429 backoff; writes only the shared antiddos:block:{ip} key (value "crowdsec") and never touches Cloudflare
- gate fires it fire-and-forget for IPs that already tripped a rate bucket, so known-bad IPs are hard-blocked before the local maxViolations threshold
- runtime config: crowdsecAutoBlock toggle, score threshold (0-5, default 4), block TTL (default 24h); boot defaults CROWDSEC_AUTO_BLOCK_ENABLED / CROWDSEC_BLOCK_SCORE / CROWDSEC_BLOCK_TTL_SECONDS
- admin panel: CrowdSec stat card, verify-connection action, score/TTL settings, CrowdSec source badge in the blocked-IPs list
- credentials live in env only (CROWDSEC_API_KEY); block is enforced per-request via proxy on the resolved X-Forwarded-For / CF-Connecting-IP
- tests: crowdsec-api unit suite + ddos-guard integration suite (early-block, threshold, cache dedupe, backoff)
2026-09-23 13:03:19 +02:00
openhands 6264f9fb20 test(security): make Cloudflare block tests deterministic under CI Redis
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m39s
CI / tests-unit (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m31s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m1s
cloudflare-api unit tests drove the real Redis connection when REDIS_URL was set (CI), causing cross-test bleed. Mock @/lib/redis with an in-memory fake identical to the gate integration test.
2026-09-22 23:31:41 +02:00
openhands 4479753160 feat(security): mirror anti-DDoS blocks to Cloudflare edge via API
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m38s
CI / tests-unit (push) Failing after 1m40s
CI / tests-ui (push) Successful in 2m28s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- gate creates a zone IP Access Rule (block) for proxied offenders that hit the block threshold, deduped until the tiered block expires
- cloudflare-api lib: verified endpoints, create/delete/verify/list helpers, Redis-backed tracking + 30s TTL sweep (instrumentation worker + admin render)
- runtime toggle cloudflareAutoBlock in antiddos config; boot default CLOUDFLARE_AUTO_BLOCK_ENABLED
- admin panel: Cloudflare edge-blocks card with verify + remove-rule actions; unban also lifts the edge block
- credentials live in env only (CLOUDFLARE_API_TOKEN / CLOUDFLARE_ZONE_ID)
2026-09-22 23:27:15 +02:00
openhands f0c27eb815 feat(security): Cloudflare-aware IP trust and admin-tunable anti-DDoS
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m50s
CI / tests-unit (push) Successful in 1m52s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m32s
- resolveClientIp: trust cf-connecting-ip only behind cf-ray/cdn-loop, use nginx x-real-ip otherwise (anti-spoof)
- antiddos-config: Redis-backed live config (antiddos:config) with 30s cache, 13 ANTI_DDOS_* env vars
- ddos-guard: consume tunable rates/tiers via getAntiddosConfig
- admin panel at /admin/devops/antiddos (save/reset/unban actions, PERMS.SETTINGS_VIEW)
- register new admin page in housekeeping migration matrix (146 -> 147)
2026-09-22 22:22:51 +02:00
openhands fd4d0fa1cb feat(security): harden anti-DDoS gate with scanner triage, tiered blocks and in-process global halt
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m0s
2026-09-22 21:57:09 +02:00
openhands 98a184953a feat(security): add Redis-backed app-layer anti-DDoS rate limiting to proxy
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 1m42s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
2026-09-22 21:48:40 +02:00