6c3d81920efbc222cbe9bd5de86fe272a3ee11e1
1701
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6c3d81920e |
fix(ops): supervise the job worker and stop the health probe from lying
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than the code. Each one had a signature that looked like a network or permissions problem and was actually a configuration or ordering bug. jobs-worker never ran `import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM evaluates a module's imports in source order and the first import reaches `@/env`, which validates process.env at import time. The ZodError on DATABASE_URL therefore fired before load-env ever executed, so the worker could only start from a shell that had already exported the configuration. Nothing supervised it either, so scheduled articles, catalog export, JAR and database backups, disk alerts and the ops health probe have all been dead; `cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and added deployment/systemd/cms-jobs-worker.service with Restart=always. The JAR backup additionally pointed at './emulator/Arcturus.jar', which does not exist and would go stale on the next emulator upgrade. resolveEmulatorJar now accepts a file, a directory or a wildcard and picks the newest JAR, the same way emulator.service picks its build, and reports an unresolvable path once instead of logging an opaque copyFile ENOENT every night. /api/health answered 200 with the database down The route documented this as intentional, and ci-deploy.sh worked around it by grepping the body for '"database":true'. The container healthcheck did not, so Docker reported containers healthy while every page 500'd. The status is now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis and the emulator deliberately do not fail the container — both have in-process fallbacks, so failing them would trade a slow site for an outage. The runtime had no V8 heap cap NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its heap from host memory (23.5 GB) while the container was limited to 4 GB, so the kernel OOM-killed the process mid-request — the same failure mode as the 14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit (v2 with a v1 fallback) and sets 70% of it, respecting an explicit override. Storage ownership was only repaired for one path ci-deploy.sh chowned storage/imaging and nothing else, so storage/catalog-git/hotel-status.json kept coming back root:root and /api/admin/catalog/status kept throwing EACCES. All eight writable storage paths are repaired now. The silent-failure mode is the reason this mattered: these writes sit inside try/catch, so a wrong owner looks like a slow page rather than an error. nginx: robots.txt was a guaranteed 404, and TLS never resumed `index index.html` without a `root` left every try_files resolving against /etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES on each stat and nginx logs a failed stat at crit, which is where 149 crit lines per scan came from. robots.txt answered from that same broken location, so crawlers were pointed at a file they could never read while sitemap.xml kept advertising it. Added `root`, proxied robots.txt to the CMS, added ssl_session_cache (there was no session resumption at all), and set Restart=on-failure in a systemd override, since the packaged unit ships Restart=no and nginx is the only thing serving the site. Verified against the running host: 3379 tests, typecheck and biome clean, nginx -t passes, health returns 200 with every check green, and the worker has run for hours at NRestarts=0 with a heartbeat refreshing each minute. |
||
|
|
108c6ce03d |
fix(ci): make the lint gate fail for real and stop byparr leaking disk
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 2m3s
CI / tests-unit (push) Successful in 2m18s
CI / tests-ui (push) Successful in 3m6s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 3m14s
The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.
Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
(key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
function-declaration handlers, with the reasoning that dropping them broke
the tree and save-on-Ctrl+S once already (
|
||
|
|
6bffc53779 |
refactor(auth): merge the duplicate login form and localize the auth screens
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 1m8s
CI / tests-integration (push) Successful in 1m53s
CI / tests-unit (push) Successful in 1m59s
CI / tests-ui (push) Successful in 2m42s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 4m33s
`home-login-form.tsx` and `login-form.tsx` were two ~240-line near-identical components. Delete the former and give `LoginForm` a `variant` prop: - `variant="page"` sr-only labels plus the register/forgot footer (/login) - `variant="compact"` visible labels, no footer (homepage sidebar) Field ids now come from `useId()`, so the two usages can never collide, and the hardcoded "Show"/"Hide"/"Loading" strings are translated. Localization of the login and register screens: - `home-login-form.tsx` was entirely hardcoded English. - `passwordStrength()` returned hardcoded "Weak"/"Fair"/"Good"/"Strong". - `register.ts` returned only English strings. It now returns a locale-independent `code` next to the message, and the form renders `t(code)` with the English string as a fallback. - Backfilled the new keys across all 25 locales, plus the login/register strings that were still English in most of them. `ar`, `fi` and `ja` had their entire login/register namespace in English and are now filled in. Locale parity stays at 0 missing keys, as `i18n:check` requires. Copy that did not match the enforced rules: the UI advertised "min 8 chars" (EN) / "min 6 tekens" (NL) while registration requires 12 characters plus an uppercase, a lowercase, a digit and a special character. Corrected in every locale. `password-reset.ts` enforced only 6 characters and is raised to 12 to match registration. Accessibility: `login-form.tsx` had no `<label>`, no `id` and no `required` on any field. All three are now present, and error banners are announced with `role="alert"`. Adds `src/i18n/auth-messages.test.ts`, which asserts every `RegisterErrorCode` resolves to a non-empty message in all 25 locales; verified it fails when a key is removed. The existing register tests now also assert the error `code`. |
||
|
|
3d828a61ab |
fix(build): build with webpack because the Turbopack build is OOM-killed
`next build` on Turbopack never completes on this app. The compiler is a
single native process whose RSS grows monotonically with no plateau:
0.9G -> 1.6G -> 2.8G -> 5.0G -> 5.5G -> 6.2G -> killed
It still dies with 4GB of swap attached, at 12GB RSS. The build workers are
only 0.17GB each, so `experimental.cpus` is not the lever either.
A `--max-old-space-size` cap cannot help: measured with a 2GB cap, RSS still
reached 8GB, because the memory is native Turbopack (Rust) memory rather than
the V8 heap. The cap added in
|
||
|
|
7f07c111ac |
perf(studio): load motion's minimal entry instead of the full component library
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Successful in 1m52s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m34s
/admin/studio/furni sat at 94.9% of its initial-JS budget (427436 of 450560 gzip bytes), so the next feature would have broken the build. Of the 98228 gzip bytes unique to that route, a large part is framer-motion. This file uses motion twice, for one thing: a 150ms opacity fade on the result pane when viewMode changes. Importing `motion/react` to get it pulls in framer-motion's complete component library — 73 internal modules — plus its render components, drag/gesture and projection code, none of which is rendered here. `motion/react-m` ships only the element factories: 2 internal modules, and the same initial/animate/transition props, so the fade is unchanged. It exports the elements flat rather than under a `motion.` namespace, so the import becomes `div as Mdiv` and the two JSX tags are renamed to match. I could not measure the resulting bundle here: the local build is OOM-killed (exit 137) with the running containers on the host, so the actual saving is unverified. The CI build reports it in build-reports, and the number in this commit message should be read as a hypothesis, not a measurement. Verified: typecheck clean, lint clean, and the 10 studio UI tests pass — including the pane and navigation specs that exercise the view switch. |
||
|
|
8218039c64 |
test(live): stop the live suites inheriting the production database
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m57s
CI / tests-integration (push) Successful in 2m1s
CI / tests-ui (push) Successful in 2m51s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m42s
Seven suites read .env with a bare `process.env[key] = value`, which overwrites whatever the shell already set. That made the DATABASE_URL from the production .env authoritative, so a single environment variable was enough to aim them at the live hotel database: RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts Three of those suites then repair the catalog in place: catalog-audit-repair-live and catalog-repair-direct-live rewrite catalog_items and delete duplicate classnames, and clone-bulk-import-live bulk-imports every cloneable item. None of that is undoable, and nothing in their output said the target was production rather than a sandbox. Added src/test/live-env.ts with one shared loader, and pointed all seven suites at it: - Values already in the real environment win, so an explicit DATABASE_URL on the command line is always respected. - DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the production one from .env. - Anything that is not loopback is treated as production and redirected. - Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning saying the suite repairs the catalog. Tests in src/test/live-env.test.ts run the loader against a temporary .env so the real project file is never read, and cover the redirect, the shell override, non-loopback detection, the opt-in and quote stripping. A second block asserts each of the seven suites no longer contains an inline `process.env[...] =` assignment. Verified four of them fail against the old loader. This does not enable the suites; they stay gated behind their RUN_* flags. It only removes the possibility of them silently hitting production. Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean. |
||
|
|
f705c67fc7 |
revert(docker-compose): keep the cms services the contract tests require
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m44s
CI / tests-ui (push) Successful in 2m34s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
The previous commit removed the `cms` and `cms-green` services from
docker-compose.yml. That was overreach and it broke two tests:
- src/lib/docker-build-contract.test.ts asserts the compose build passes
NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}, so a compose-built image
carries its release id.
- scripts/proxy-config.test.mjs resolves `docker compose config` and asserts
the `cms` service's host networking, volumes, healthcheck and image tag.
Both encode that docker-compose.yml is a maintained deployment surface, not a
leftover. Removing it was not my call to make while fixing a deploy.
Restored verbatim. The stray container that actually blocked port 3002 is
already gone, and nothing recreates it: there is no systemd unit or pm2
ecosystem that runs `docker compose up`, and `restart: unless-stopped` only
applies to a container that still exists. So the blocker is resolved by the
container removal alone, and compose stays intact for manual and reviewed use.
|
||
|
|
c8b3054527 |
fix(deploy): free port 3002 and stop compose from competing for the slots
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy could not start its candidate because port 3002 was held by `epicnext-cms`, a `docker compose up` replica built from the `local` image and serving no traffic. Everything else in the pipeline was healthy: the image built, the news browser gate passed and migrations were current. The container was unusable for this pipeline for two reasons. It ran a different image than any release, and its name did not match the slot the deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms` while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never could. deploy.sh already documents that compose "never managed the release that actually ran", so the service was stale by its own account. Removed the stray container and dropped the `cms` and `cms-green` services (plus the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot resurrect a replica that permanently occupies a blue/green slot. byparr is untouched. Also fixed the diagnostic from the previous commit, which blamed every running container. `docker ps --filter publish=` returns nothing for --net=host containers, so the fallback listed all of them and buried the real holder among seven innocent ones. It now resolves the listening PID from `ss` back to its container through /proc/<pid>/cgroup and names only that one, with the exact `docker rm -f` command to run. Verified: port 3002 free, live release on 3003 still serving (status ok, database and redis true), deploy simulation 26 passed, typecheck. |
||
|
|
8ec3df541e |
fix(deploy): name the container blocking a port and silence phantom cleanup
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m30s
Two follow-ups from the blocked deploy. The rollback path called `docker logs` and `docker rm -f` on the candidate unconditionally. When the port check refuses to start it, the container was never created, so both printed "No such container: epicnext-cms-app" — noise that looked like a second, unrelated failure and buried the real message. Both calls are now guarded by `docker inspect`. assert_port_free() now reports which container holds the port and flags it when it is not a blue/green slot this script manages. The previous output listed every container and said only "port already in use", which is a dead end: on this host the holder is `epicnext-cms` (a `docker compose up` replica on port 3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ because docker-compose.yml pins `container_name: epicnext-cms` for the `cms` service; slot B happens to match, which is why 3003 deploys fine and 3002 never can. The message now names the squatter, explains that live traffic is unaffected, and gives the next action. Deploy simulation: 26 passed. |
||
|
|
8ee144745a |
fix(deploy): stub ss in the deploy harness and cover the port-conflict path
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m47s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m41s
The port-conflict guard added in the previous commit made the six existing blue/green deployment simulation tests fail. assert_port_free() shells out to ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node — but not ss. Because the runner is self-hosted and the containers use --net=host, the simulation saw the production CMS containers holding 3002 and 3003 and refused to start its own candidate. The harness now stubs ss. It reports no listener for every scenario except 'port-taken', which reserves whichever port the script asks about, so the simulation stays independent of the host it runs on. Also switched the ss probe from `command -v ss` to `type ss`. The stub is a shell function delivered through BASH_ENV; `command -v` happens to find it, but `type` is the reliable test for "is this resolvable", and the two differ across shells. Added a regression test for the guard itself: with the candidate port already occupied, the deploy must fail, must not have run `docker run`, and must leave the nginx upstream untouched on the old port — no half-finished cutover. Verified it fails when the assert_port_free call is removed. Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped. |
||
|
|
64ad9baf39 |
fix(deploy): trust the nginx upstream when picking the live slot
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Failing after 1m54s
CI / tests-ui (push) Successful in 2m46s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy failed with "Expected release never became healthy" after 30 attempts. Root cause: read_active_port() counted the slots answering /api/health and only consulted the nginx upstream when the count was not exactly one. On this host both slots were healthy, so it fell back to the upstream file, but a leftover epicnext-cms:local replica was holding slot A (3002). The candidate was assigned that occupied port, docker run died with EADDRINUSE, and the health probe then answered from the pre-existing container on that port. That container reports release "unknown" because it was built without NEXT_DEPLOYMENT_ID, so the release comparison could never match and the deploy timed out blaming a release that was never serving. read_active_port() now orders its sources by how well they describe reality: 1. The nginx upstream file. It is the only source that says where public traffic actually enters; everything below it is a consequence. 2. A healthy slot matching that pointer. 3. The other slot when the pointer names a dead port. 4. The pointer itself when nothing answers, so rollback still has a target. 5. Slot A when no upstream file exists at all. answers_health() was added as a retry-free sibling of healthy(); port detection should not spend 90 seconds per slot on a process that is either running now or never will. start_candidate() now calls assert_port_free() before docker run, so an occupied port fails immediately and names the listener and the containers involved, instead of surfacing later as a misleading health-check timeout. Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from the real script rather than copying them, and covers the regression: with both slots healthy and nginx serving slot B, the result must not be slot A. Verified the test fails against the old logic and passes against the new. Wired into the check job so this is caught before an image is built. |
||
|
|
704e33638f |
fix: restore six useEffect dependencies removed while silencing lint
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m38s
CI / tests-integration (push) Successful in 1m40s
CI / tests-ui (push) Successful in 2m24s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m58s
The previous commit dropped biome-ignore comments to clear useExhaustiveDependencies diagnostics and, in doing so, also deleted the dependencies themselves. Six components were left with effects that no longer react to the state they read. Every one of these is a real behaviour regression, not a lint preference: - health-check-client: checkEmulator is a function declaration, so it gets a fresh identity each render. As an effect dependency that re-fires the effect after every setState, polling /api/admin/devops/health in a loop. Wrapped in useCallback so the identity is stable. - article-recovery: reload restarts the autosave timer for the "Retry recovery" button. Without it in the deps that button is a no-op. The counter had been renamed to _reload to satisfy the unused-variable rule. - catalog-integrity-panel: same pattern; refresh starts a new read-only scan, so the rescan control did nothing. - catalog-search: refreshKey re-runs the query after a bulk edit, so results were not refreshed after catalog edits. The selection-reset effect also lost catalogType, so switching catalog no longer cleared the selection. - catalog-image-picker: dropped debounced (the search term) and name (the error reset), so image search and error state no longer reacted to input. - icon-picker: dropped iconImage, so a failed load left the placeholder on the next icon too. Each restored dependency carries a biome-ignore with the reason it is load-bearing, so the diagnostic can be re-derived instead of silently disappearing again. Verified: typecheck, lint clean on all six, unit 3315 passed, integration 20 passed, UI 72 passed / 2 skipped. |
||
|
|
eddb7edea4 |
fix: make all CI jobs pass (integration, ui) and restore prefix dialog reset
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m34s
CI / tests-unit (push) Successful in 1m35s
CI / tests-ui (push) Successful in 2m20s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m37s
Three failing test suites blocked CI. All three were test defects, not application bugs. Integration tests (integration/database.test.ts) ------------------------------------------------ The suite set NODE_ENV=test, which makes cache.cached() short-circuit both its Redis read (src/lib/cache.ts:226) and its write (:249). A suite whose stated purpose is exercising the real Redis path therefore never touched Redis. Switched to NODE_ENV=development, the only non-production value src/env.ts accepts, so the shared-cache code paths are genuinely covered. Three assertions then needed correcting for real Redis semantics: - `await cache.cached(...)` followed by `.resolves` can never hold: await yields a value, not a Promise. Assert the value directly. - A cached negative result is stored as the JSON encoding of null, so `redis.get(key)` returns "null", not null. - The news negative-cache key does not exist at all, so `ttl()` returned -2. Now that the write path is live the key is created and the TTL assertion holds as originally written. UI tests (src/app/admin/prefixes/prefix-dialog.tsx) --------------------------------------------------- The form-reset effect had `isOpen` removed from its dependency array. The component returns null when closed, so the effect only ever ran on mount: reopening the dialog no longer cleared the fields and a dismissed-but- unsaved edit reappeared. Two tests in e2e/ui/unsaved-changes.spec.ts caught this. Restored the dependency and documented why it is load-bearing. The remaining edits in this branch drop stale biome-ignore comments that suppressed useExhaustiveDependencies and noArrayIndexKey diagnostics. Where the suppression had been load-bearing for behaviour, the underlying dependency is now listed explicitly rather than silenced. Verified: check (toolchain, audit, lint, i18n, typecheck), unit 3315 passed, integration 20 passed, UI 72 passed / 2 skipped. |
||
|
|
1c9ddcd48a |
fix(cms): increase docker mem limit to 6gb and enforce node max-old-space-size to prevent OOM killer crashes
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m30s
CI / tests-integration (push) Failing after 1m34s
CI / tests-ui (push) Successful in 2m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
|
||
|
|
30ff970c38 |
chore: upgrade to pnpm v12, update dependencies, and fix msw v3 typescript types
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
|
||
|
|
f99980052b |
perf: optimize cache layer for speed and stability
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 34s
CI / tests-ui (push) Failing after 33m56s
CI / tests-integration (push) Failing after 33m57s
CI / tests-unit (push) Failing after 33m57s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
- Remove random TTL jitter to prevent unpredictable cache drops - Add deterministic LRU eviction with proper entry cleanup - Improve cache deduplication to prevent duplicate computations - Skip Redis I/O during tests for faster, more stable execution - Optimize depth calculation in catalog tree nodes - Maintain backward compatibility and full test coverage (3331 passed) |
||
|
|
f181cd6af4 |
chore: ignore local runtime snapshots under backups/
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m41s
CI / tests-unit (push) Successful in 1m46s
CI / tests-ui (push) Successful in 2m32s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m56s
backups/catalog-integrity/pre-fix.json is a one-off database snapshot taken during an incident, not source. Kept on disk for reference, out of git. |
||
|
|
8a6d92afd8 |
fix(cloudflare): cache the gamedata tree at the edge with respect_origin
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m44s
CI / tests-integration (push) Successful in 1m44s
CI / tests-ui (push) Successful in 2m31s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m16s
Icons are plain .png, a cacheable extension by default, so the zone's "Browser Cache TTL = 1 year" pinned them to max-age=31536000 regardless of the 300/3600/604800 that nginx sends per class. Extend the edge rule to /gamedata/ and keep respect_origin, so the nginx header wins and a 404 (notably no-store from the gamedata 404 handler) is never pinned. |
||
|
|
dbaccd7cfc |
fix(gamedata): never cache a missing gamedata file
A missing gamedata file got no Cache-Control at all, because add_header without `always` only applies to 2xx/3xx. Cloudflare then fell back to the zone setting "Browser Cache TTL = 1 year", so the 404 came back as `max-age=31536000` with `cf-cache-status: HIT` — pinned in the visitor's browser and at the edge. An icon requested while its import was still running stayed a 404 for the rest of the year, even after the file existed. That was the "some icons load, some don't" report. Give every gamedata location a named 404 handler that sends no-store, and split icons/ out as its own cache class: those files are rewritten under the same name (repair-icons, reimport), so an hourly must-revalidate keeps a repaired icon visible within the hour instead of days later. |
||
|
|
4e036b08d5 |
fix(proxy): drop request rate limiting from gamedata entirely
/gamedata/* is served straight from disk by nginx; no request hits the CMS backend or a database, so a request-rate limit protects nothing while costing players their icons. A room load fires hundreds of these files in one burst, which every limit turned into visible 503s. Removed the static zone from the gamedata locations. Traefik's epicnabbo-gamedata router likewise carries no rateLimit middleware. /client/ and /nitro-client/ keep theirs, and the main route keeps the 30r/s page budget plus the server-wide connection limit. Measured: 1000 icon requests fired fully in parallel now all return 200, while 200 parallel requests on / are still rejected. |
||
|
|
f0dcf440a7 |
fix(proxy): split rate limiting into page and static zones
The single server-scope limit_req (30r/s) treated a page load and a room load as the same thing. Loading a Nitro room fires several hundred gamedata icons in one burst, which that zone answered with 503s, so icons showed up late in the client. Add a separate static zone (1000r/s, burst 1000, nodelay) for the gamedata and client asset locations, and apply the page-rate zone explicitly on the main route instead of at server scope. Connection limit stays server-wide. Measured: 900 icon requests in burst now all return 200, while 200 parallel requests on / are still rejected. |
||
|
|
4a1211a931 |
feat(proxy): add per-IP rate and connection limits
The edge had no limit_req/limit_conn at all, so a single client could flood the Next.js backend and the Nitro client with unbounded parallel requests. Traefik's logs already showed this: bursts of gamedata icon requests answered with 429. Add limit_req (30r/s, burst 60, nodelay) and limit_conn (30) zones keyed on the real client IP, applied at server scope so both cached assets and proxied API routes share one budget. The burst is deliberately generous because the Nitro client fetches gamedata and icons in bursts when loading a room. |
||
|
|
e3c010f383 |
fix(proxy): raise nginx worker rlimit above worker_connections
nginx inherited systemd's soft LimitNOFILE of 1024, so every start logged "2048 worker_connections exceed open file resource limit: 1024" and the worker_connections value could not actually be reached. Set worker_rlimit_nofile to 65536. Bounded from above by a systemd drop-in at /etc/systemd/system/nginx.service.d/override.conf (LimitNOFILE=65536), since the master's hard limit caps what workers may request. |
||
|
|
a6cc3cafa9 |
fix(catalog): read furnidata from one cache, purge the gamedata edge on write
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m52s
CI / tests-unit (push) Successful in 1m56s
CI / tests-ui (push) Successful in 2m48s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m27s
Furniture was not always loading completely because the same file was cached twice and nobody could reach the client. The catalog items loader kept its own 30s TTL copy of FurnitureData.json next to the mtime-validated cache in `furni-data.ts`. An import cleared only the second one, so the catalog table kept serving pre-import furnidata — empty descriptions and revisions — until the TTL ran out. The loader now reads through `readFurniData`, which revalidates on mtime+size and is reset by every write, so there is exactly one cache and it cannot go stale on its own. `invalidateFurniDataCache` and its single call site are gone with it. The client was worse: nginx served all of /gamedata/ with `max-age=604800`, and the `cms-gamedata` purge that would have fixed it hung off the catalog Git export, which is disabled in production. A freshly imported item was invisible in the client for up to seven days no matter how often you imported. - `writeFurniData` now purges the gamedata edge tag itself. One place covers import, batch, resync, regen, nitro-editor, translate and dedupe. It is fire-and-forget and swallowed at every level: a stale edge copy is bounded by the edge TTL, so a failed purge must never fail an import. - nginx splits /gamedata/ by how mutable the content is: config/ gets `max-age=300, must-revalidate`, bundled/ `max-age=3600, must-revalidate`, and the content-addressed trees (c_images, album*, clothes) keep the long TTL. `must-revalidate` is the point — the client now revalidates instead of replaying the old body. All three keep `Cache-Tag: cms-gamedata` so the purge still reaches them. - A 30-minute safety-net purge in the jobs worker covers the case where Cloudflare was unreachable at write time. |
||
|
|
cede541813 |
fix(catalog): route item-table writes to the catalog they belong to
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m16s
The items table is shared between both catalogs, but its four mutating actions were normal-only: moving, reordering, creating and updating a Builder Club offer wrote to catalog_items, so a BC edit either landed in the wrong catalog or hit an unknown column. Pass the catalog from the table through the actions and let the server resolve it. BC rows have no price, points or currency column, so the BC commands strip those fields instead of rejecting them. Moving and reordering now share one command that locks the category and writes the table for the same catalog, and BC writes revalidate the BC route. |
||
|
|
e0efbef30d |
fix(catalog): make the Builder Club catalog read and write its own offers
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m23s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m6s
The previous commit taught bulk editing and delete-with-restore about the BC catalog. Neither actually worked, and one of them was destructive. `catalog_items_bc` has six columns: id, item_ids, page_id, catalog_name, order_number, extradata. There is no price, points, currency, offer_id, limit or membership column on it. The bulk path read and wrote columns that do not exist, and the UPDATE was aimed at catalog_items while the SELECT came from catalog_items_bc — so a BC category move wrote into the normal catalog. Two tests now pin that pairing: reads and writes have to stay in the same table. Underneath it the BC table was never being read at all. The inline editor fetched `/api/admin/catalog/items?pageId=N` without the catalog, so opening a BC category showed the normal catalog's offers, and the route selected BC rows directly instead of going through the loader, skipping the furni enrichment the table needs to render anything but a bare caption. Both catalogs now take the same path, and the catalog is in the fetch callback's dependencies — without that, a switch keeps reading the previous catalog's rows through a stale closure. Because a BC offer has no price, the editor no longer offers one. The server refuses price, points and currency changes with a readable message instead of letting them reach the database as an unknown-column error, and a BC bulk edit is what it can actually be: a category move. BC deletions also went through a bare DELETE, which made them the one catalog mutation with no way back. They now keep their rows and hand back a restoreId like the normal ones. The catalog is recorded in the audit target rather than in the payload, so a restore can never put a BC row into the normal offers table. |
||
|
|
cebcf440c5 |
feat(catalog): record bulk edits, make deletions reversible, unify the tree read
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m23s
A bulk offer edit is the catalog mutation that rewrites hundreds of rows at once, and it was the only one writing nothing to the staff activity log: 22 of the 43 catalog actions logged, this one did not. The entry it now writes says what changed, not just that something did, because the log has no undo of its own and "bulk updated 200 offers" cannot answer the question it exists for. Deleting offers had no inverse at all. Every removed row is now kept at delete time and the caller gets a restoreId back, so an accidental multi-select is a click rather than a hand-edit of the table. The undo toast covers the common case; a RecentDeletionsPanel holds the same records so a delete noticed later is still reachable. Three refusals guard it: an id that another offer has since taken, a category that no longer exists (which would leave an offer that sells nowhere and shows under no page), and a delete whose restore record cannot be written — that one rolls back rather than deleting without a way back. Reading the audit row FOR UPDATE is also what stops two restores of one deletion from both inserting. sendCatalogUpdate() overwrote hotel-status.json on every write, so "which imports reached the hotel" was answerable for the last attempt only, and a failure two imports ago was gone by the time anyone looked. That file is now also appended to as a bounded 50-entry tail. The tree route carried four copies of the same page-select-plus-counts shaping, of which the BC branches had already drifted: one counted offers through the VARCHAR-tolerant helper, the other inline and swallowing errors. All of it is one readPages() now, and readFullTree sends both catalogs through one depth computation instead of delegating normal to getTreeFlat while computing BC here — a split that left two implementations behind one function name. getTreeFlat is gone. The BC ancestor walk also went from 20 levels to 50, matching getAncestors, so a deeply nested catalog no longer loses its breadcrumb. Bulk editing reaches the BC catalog, which previously had no way to edit or duplicate offers in bulk. The catalog is part of the operation identity now, so replaying one request key against the other catalog is not mistaken for the same work. Integration tests failed to import: the next/cache mock supplied only revalidatePath, and catalog-totals calls unstable_cache at module scope. |
||
|
|
9550b3d66f |
feat(catalog): make the live catalog self-correcting and honest about failure
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m34s
CI / tests-integration (push) Failing after 1m34s
CI / tests-ui (push) Successful in 2m19s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The previous commit made imports update the Studio without a reload, but the guarantee only held inside the tab that started the import and only as long as every read succeeded. Four holes were left, and this closes them. A session that mounted the tree before an import kept the pre-import tree for the rest of its life, because ensureCatalogTreeLoaded() was a once-per-session no-op. It now asks the server whether what it holds is still current. The answer is a revision: sendCatalogUpdate() already runs after every catalog write, so it bumps one, and clients read it on mount, on focus, on a 20s poll and from other tabs over a BroadcastChannel. An import that finishes in another tab, another browser or the job worker now lands here too. A failed read used to be swallowed, which is the worst outcome available: the rail kept showing pre-import counts as if they were current and nothing said so. The snapshot now carries the error, the rail shows it with a retry, and the previous tree stays on screen because stale beats empty. Every settled import pulled the entire flat tree, which is the one payload that grows with the size of the catalog. The revision doubles as the ETag on mode=full, so an unchanged catalog answers 304 and the poll costs a file read. An import could also report success for an offer the hotel will never sell: a hidden or disabled page, an item_ids that misses the furni id, a zero amount. importSingleFurni reads its own row back and reports each of those as a warning, where the import report already is, instead of leaving it to surface as "the import did not work" in the client. Finally, the catalog items table no longer falls back to router.refresh() — onRefresh is now required, so every mutation ends in a refresh of the caller's own data instead of a route re-render that threw away editor state and scroll position. useServerAction keeps its default, because 47 callers across the app depend on it. The 750-line CatalogTree in catalog-tree.tsx was dead code that kept its own stale tree and three more router.refresh() calls; only CatalogIcon and LAYOUT_COLORS are still imported, so the rest is gone. Tests: the store now covers revisions, 304s, probe failures and error recovery; a jsdom test mounts a consumer and asserts the tree updates in place with no navigation; the old organize-imports e2e asserted nothing about the endpoints the code actually calls, and is replaced by one that asserts a cross-tab write lands in the mounted categories without a reload. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
28ce0f911c |
fix(catalog): keep the live catalog truthful after every import path
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Failing after 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The live catalog store only covered part of the import surface. A durable job settled, a sync queue drained, a .nitro upload or a clone run left the Studio rail and the stats bar showing pre-import numbers until the page was reloaded, and the Catalog Manager kept a second tree that never saw writes made elsewhere in the session. Every one of those paths now pulls the tree again, and the refresh carries the totals with it: importing writes catalog rows server-side, so the counts the store holds were stale for the rest of the session. - refreshCatalogTree shares one request between concurrent callers and queues a single follow-up read when a write lands mid-flight, so a burst of edits costs at most one extra read. - useFurnitureJobs treats its first payload as a baseline, so a page load no longer replays every past import as "just settled", and hands the settled jobs to the callback. - The Catalog Manager pushes its own mutations into the store and re-reads its active tab when the store changes. - The 30s unstable_cache on the admin totals is now tagged and invalidated from every catalog write, including the import worker, so it no longer survives an import even across a hard reload. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
d73baf1458 |
fix(catalog): take furnidata values from the clone source for retro items
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 29s
CI / tests-integration (push) Successful in 1m33s
CI / tests-unit (push) Successful in 1m33s
CI / tests-ui (push) Successful in 2m21s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m53s
The resync route rebuilt every entry from items_base plus the *official Habbo* furnidata. A classname that only exists on a retro hotel (leet.ws and friends) is absent from the official set, so lookupOfficialHabboFurni returned null and the entry was written with revision 0, category "unknown" and an empty description — even though the clone import had those exact values available at import time from the source's own furnidata. - resync/route.ts: when a classname is genuinely missing from official Habbo, fall back to the configured clone sources. Their furnidata is indexed by normalized classname and each entry is coerced into the OfficialHabboFurniEntry shape, which is the same JSON shape, so it drives the existing buildFurniEntry fallbacks for revision, category, name, description, defaultdir, partcolors, specialtype, furniline, environment, rare and bc. items_base stays authoritative for id, spriteId and dims, and public_name still wins over the source name, matching the import. The index is memoized per request, not at module scope: a module-level cache would pin the source list for the life of the process and a source added later would never be picked up. fetchSourceFurnidata already caches per URL, so this costs one parse rather than a network round-trip. Disabled sources and sources that fail to respond are skipped, so an unreachable hotel degrades to the previous items_base-only behaviour instead of failing the run. Official Habbo still wins whenever it has the classname, so existing behaviour is unchanged for everything but the retro-only case. Applies to every resync mode, so the pre-existing ?missing=1 sweep picks this up too. |
||
|
|
2f7e557d5e |
feat(catalog): add a Studio button to fix missing furnidata entries
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 28s
CI / tests-integration (push) Successful in 1m33s
CI / tests-unit (push) Successful in 1m38s
CI / tests-ui (push) Successful in 2m26s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m1s
"Missing furnidata" was only a filter in the Studio status dropdown, so
imported items whose classname was absent from FurnitureData.json could
be found but not fixed from that screen. Only the Catalog Audit page could
repair them, and only globally.
Adds the same shape of quick action that "no nitro" already had:
- studio-client.tsx: a "N no furnidata" shortcut next to the "N no nitro"
button that sets the missingFurnidata status filter, and a bulk "Add
missing furnidata (N)" button for the selected rows. Both only appear
when there is something to act on. Rows that come back repaired flip
to hasFurnidata: true so the badges and counts update in place; rows
the server reported in errors keep their state.
- resync/route.ts: accepts an optional { classnames: string[] } body to
target exactly the selected rows. classnames are resolved through the
same normalized local index the listing uses to decide hasFurnidata, so
the rows written are the rows flagged as missing. The upsert is already
idempotent, and RCON updateCatalog + updateItems run afterwards so the
emulator picks the new entries up.
Also clears the Studio furnidata cache after a write, which this route
never did: without it the listing kept serving a stale hasFurnidata for
up to the 30s cache TTL, so a repair looked like it had done nothing.
PERMS is now imported from permission-slugs (identical re-export) so the
route no longer pulls next-auth into tests.
- studio-filters.test.ts: pins the missingFurnidata branch, in particular
that an unchecked item (hasFurnidata undefined) is not treated as missing.
The existing ?days / ?missing / ?broken / ?all modes are unchanged; the
body is only consulted when it carries a classnames array.
|
||
|
|
4be7eaed59 |
fix(catalog): never create a page that reuses a sibling's order number
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 41s
CI / tests-unit (push) Successful in 1m53s
CI / tests-integration (push) Successful in 2m19s
CI / tests-ui (push) Successful in 2m46s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m9s
The emulator complained "Sibling order 2 is used more than once" 74 times in production, covering 19 pages across 12 parents. The cause was that nearly every page-creation path passed orderNum 0, so each new page collided with whatever sibling already sat at 0 or 1, and the Park placeholder pages all shared the sentinel values 99 and 999. The production rows themselves are repaired out of band (renumbered 1..N per affected parent, ordered by order_num then id so the existing visual order is preserved, plus one dangling catalog_items row whose items_base no longer existed removed). This commit stops it recurring. - hierarchy.ts: add nextFreeSiblingOrder(), which ignores -1 and 0 as the same root set, treats a missing orderNum as 0, and returns an order strictly above the highest sibling in use. An explicit order is still honoured whenever it is free, so callers that genuinely want a position keep it. - page-commands.ts: createPageCommand resolves the real order through nextFreeSiblingOrder instead of writing the requested 0 straight through. Note that furni-import.ts and upload-import.ts still take their order from the furnidata catInfo.order, so two categories carrying the same furnidata order can still collide. That is caught by the emulator audit and repaired by fixEmulatorIssues(), but it is not prevented here. |
||
|
|
9cc57cddfc |
feat(catalog): update the catalog live after an import, no page refresh
Organising imports, the Studio furni batch, the catalog totals and the "import from a source" stats all used to need a full page reload, or at best a router.refresh() that re-rendered the whole admin route, before anything on screen reflected what the import had just written. - live-catalog-merge.ts (new): pure tree and total arithmetic. Applies a delta of created pages, added offers and moved offers, recomputes depth for the touched subtree, bumps parent child counts and the item totals. Returns the input untouched when a delta is empty, so subscribers can bail out instead of re-rendering. Depth resolution tolerates a parent cycle in a dirty DB and still terminates, matching getTreeFlat. - use-live-catalog.ts (new): one module-level store exposed through useSyncExternalStore, so every consumer shares a single instance without threading a provider through the admin layout. Deltas only apply to the "normal" catalog, so public and public_handlers trees stay separate. seedCatalogTotals() takes the first server value per mode and never overwrites it afterwards, so a later hard render cannot make the header totals jump backwards. - actions/catalog.ts: organizeImportFurni now reports each group through the new OrganizedPageChange, carrying parentId, pageLayout, the icon, isNew and the per-source movedFrom counts, so the client can fold the result into the tree without reading the page back. - organize-imports-dialog.tsx: drops useRouter and router.refresh(); the response is applied as a delta the moment the run finishes. - studio-client.tsx: reads the tree from the store instead of freezing it with useState(initialTree), loads it on mount when empty, and refreshes it once a batch import settles. The batch is server-side and derives its import pages from furnidata, so that one path re-reads the tree via GET /api/admin/catalog/tree?mode=full rather than trusting the delta. - studio/furni/page.tsx: stops calling getTreeFlat() and no longer passes initialTree; the store is the single source of truth for the rail. - import-clone-client.tsx: tracks which items are already present, so present and clonable update per cloned row instead of only at the end. - catalog-manager-dialog.tsx: seeds the totals once and renders the live values, so the header reflects an import that just ran. - e2e/ui/fixtures/entry.tsx: drops the removed initialTree prop. |
||
|
|
7507c3b55c |
fix(deploy): detect the actually-live blue/green slot, stop nginx-sync clobbering the upstream
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-unit (push) Successful in 1m40s
CI / tests-integration (push) Successful in 1m42s
CI / tests-ui (push) Successful in 2m28s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m5s
- ci-deploy.sh: read_active_port() now probes both slots on /api/health and picks the one that really answers; the upstream file only serves as a fallback when zero or both slots respond. A stray 'docker compose up' (or a clobbered snippet) can no longer derail the next deploy's cutover. - nginx-sync.sh: cms_upstream_servers.conf is runtime-owned by ci-deploy.sh; only seed it when missing, never overwrite what a deploy wrote. This is the root cause of tonight's 502: a nginx-sync run reset the snippet (written to green:3003 by the last cutover) back to the dead slot A:3002. - cms_upstream_servers.conf: restore the fresh-host seed default to slot A. |
||
|
|
90b65c92a2 |
feat(proxy): sync Cloudflare ranges at nginx+Traefik, block IP spoofing
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m51s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 20s
- cloudflare-ips.conf (new): geo $cms_trusted_edge + set_real_ip_from from live CF IPv4/IPv6 ranges plus Traefik bridge and loopback - nginx-cms.conf: forward real client IP only from trusted peers, strip incoming CF-Connecting-IP, 403 any other peer that presents one (spoof gate); direct game clients on :9443 stay unaffected - cf-ips-sync.sh (new): fetch cloudflare.com/ips-v4/-v6, regenerate the nginx snippet and Traefik websecure.forwardedHeaders.trustedIPs - nginx-sync.sh: install the cloudflare-ips.conf snippet - cms_upstream_servers.conf: point default at the live green slot 3003 |
||
|
|
7697728d07 |
feat(cache): single-owner caching across nginx, edge and content edits
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 28s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m41s
CI / tests-ui (push) Successful in 2m27s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 3m35s
Rebuild production nginx from the repo (deployment/proxy/*) with a single Cache-Control owner per route: the app stays the source, nginx only manages headers, and Cloudflare stores the public API allowlist at the edge. - deployment/proxy: nginx.conf, mime.types, nginx-cms.conf and the blue/green upstream snippet; config backed by scripts/nginx-sync.sh (idempotent install + reload, --check/--force). - nginx serves Cache-Tag headers on the public allowlist (cms-public), gamedata, client and camera responses so the edge and purge stay in sync. - src/lib/edge-cache.ts + tests: coalesced, fire-and-forget edge purges that no-op unless Cloudflare is configured; scripts/cf-purge.sh and cf-setup-cache.sh create and purge the cache rule. - src/lib/cloudflare-api.ts: purgeCacheByTags/purgeCacheByUrls. - Purge hooks after catalog exports (public + gamedata) and on shop, team, guild, photo and rare-values edits; ci-deploy purges after each release. - src/proxy.ts excludes the imaging/images docs from the middleware matcher. |
||
|
|
30dcecd530 |
style: restore tab indentation in package.json
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m39s
CI / tests-integration (push) Successful in 1m43s
CI / tests-ui (push) Successful in 2m26s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 18s
The previous dependency upgrade rewrote the file with two-space indentation, which violated the Biome formatter setting (indentStyle: tab) and broke `biome check .`. |
||
|
|
e4f83a8036 |
chore: upgrade to pnpm v12, vitest v5 and resolve deprecated subdependencies
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Failing after 21s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
|
||
|
|
f187cd70a9 |
chore(deps): bump vitest and @vitest/coverage-v8 to 5.0.2
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m51s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 18s
Patch release, bug fixes only, no breaking changes. Two entries are relevant to this repository: a stack overflow when spying on Set.prototype.add, and the hanging-process reporter switching to its ESM entrypoint, which drops why-is-node-running 2.3.0, siginfo and stackback in favour of why-is-node-running 3.2.2. Verified with the CI unit command: 3302 tests pass under --maxWorkers=4, the coverage run clears its thresholds, pnpm deps:audit reports no known vulnerabilities, and the lockfile stays consistent under --frozen-lockfile. |
||
|
|
d2d01141f1 |
style: apply Biome formatting to the catalog release test
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 32s
CI / tests-integration (push) Successful in 1m49s
CI / tests-unit (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 20s
The timeout constant made the first it() line exceed the line width, so `biome check .` failed with a format error. Reformat and confirm the three publication tests still pass. |
||
|
|
c2981bd970 |
test(catalog): give the Git publication tests room to finish
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Failing after 22s
CI / tests-unit (push) Skipped
CI / tests-integration (push) Skipped
CI / tests-ui (push) Skipped
CI / preflight (push) Skipped
CI / deploy (push) Skipped
These drive real git processes against a local bare remote, so their cost is process spawns competing with every other Vitest worker. Measured on CI they take 23-30s each, and the 30s override was crossed by 37ms, so the run failed on wall-clock rather than on behaviour. Replace the three hand-picked 30_000 values with one documented constant at 120_000, which keeps a genuine hang visible while clearing the observed spread. The global default stays at 10s so nothing else is loosened. |
||
|
|
d9ef7f8360 |
chore(ci): remove Renovate
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 35s
CI / tests-unit (push) Successful in 1m59s
CI / tests-integration (push) Successful in 1m58s
CI / tests-ui (push) Successful in 2m43s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 19s
Dependency updates are handled manually, so the scheduled Renovate job only cost a daily privileged Docker run on the deploy host. The empty cache directory it maintained is gone too, and the operations note now records that updates are manual instead of describing bot behaviour. |
||
|
|
1f9ad02410 |
chore: ignore .env.local and .env.*.local
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 40s
CI / tests-integration (push) Successful in 2m35s
CI / tests-unit (push) Failing after 3m24s
CI / tests-ui (push) Successful in 4m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
load-env.ts reads .env.local before .env and gives it precedence, so a developer override file can hold real secrets. Only .env was ignored, so that file was one 'git add .' away from being committed. |
||
|
|
80d7ae14ba |
fix(ci): fail fast when the deploy dir has no DATABASE_URL
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 34s
CI / tests-unit (push) Successful in 1m42s
CI / tests-integration (push) Successful in 1m48s
CI / tests-ui (push) Successful in 2m35s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m37s
pnpm db:migrate runs on the host and reads DATABASE_URL from the deploy directory's .env. When that variable was missing the deploy had already built an image and run the browser gate before pnpm db:migrate aborted on an empty value, so a release was paid for in full and then thrown away. Check for the variable right after the .env is copied, before the build, and say plainly that the live release was not touched. The deploy test fixture gains a DATABASE_URL so it mirrors a working deploy directory instead of the broken one. |
||
|
|
bc00ecf08c |
feat(i18n): complete message parity across all 25 locales
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 37s
CI / tests-unit (push) Successful in 1m49s
CI / tests-integration (push) Successful in 1m50s
CI / tests-ui (push) Successful in 2m29s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m37s
The admin.studio.nitroCleanup section (102 keys) only existed in en and nl, so 23 locales fell back to English for the entire Nitro Cleanup panel. The referrals and dailyRewards keys were missing from the same 23 locales, and en itself was missing 6 keys that nl had. Add the missing keys to every locale with translations, so all 25 locales now carry the same 6063 keys. |
||
|
|
944527e078 |
feat(nitro-cleanup): dedupe FurnitureData and clean dangling figure entries
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 1m58s
CI / tests-unit (push) Successful in 2m16s
CI / tests-ui (push) Successful in 3m8s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m30s
Add a gamedata cleanup to the Nitro Cleanup panel: a read-only preview plus an apply run that dedupes FurnitureData classnames and removes rows and figure entries that reference nothing. Three passes run in a fixed order, because cleanFigureMap has to precede cleanFigureData: dropping the part that points at a set is what makes that set unreferenced. A pass refuses to write when it would delete more than maxRemovals rows (default 500) and reports the reason, a wrong asset directory otherwise turns every row into an orphan and one call would empty the file. Passes that would act on empty input (no libraries, no sets) treat that as a missing file rather than as a reason to delete everything. Every write copies the file to a timestamped backup first, so a pass that turns out to be wrong can be undone by hand. The plan reads FurnitureData once and hands the parsed copy to both furniture passes; the file is tens of megabytes in a real deployment. |
||
|
|
d176fad4da |
fix(nitro): repair stale meta.image in bundles that are already lossless
A bundle whose texture is already VP8L was returned untouched, so a stale spritesheet.meta.image survived the normalisation and the client could not find the texture member. Rebuild the archive in that case and reuse the existing VP8L bytes instead of decoding them again. |
||
|
|
9ee22db8ba |
fix(nitro): normalise attached and recovered .nitro bundles too
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 40s
CI / tests-unit (push) Successful in 1m45s
CI / tests-integration (push) Successful in 1m55s
CI / tests-ui (push) Successful in 2m37s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m53s
Two more .nitro entry points in the main import path still wrote the supplied buffer verbatim: an attached `providedNitro` and a bundle pulled back by `resolveMissingNitro`. Both are real furniture imports, so they could still land a PNG texture while the SWF, clone and upload paths produced WebP. Route both through the same normalisation, falling back to the original bytes with a warning if the texture cannot be decoded. |
||
|
|
b26e2e0de4 |
feat(nitro): normalise hotel and uploaded bundles to WebP Lossless
Gitea Actions Runner Test / test-job (push) Successful in 0s
CI / check (push) Successful in 31s
CI / tests-integration (push) Successful in 2m17s
CI / tests-unit (push) Failing after 2m29s
CI / tests-ui (push) Successful in 3m17s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Importing from a hotel wrote the downloaded .nitro to disk untouched, so official PNG textures stayed PNG and only SWF imports ended up as WebP. Every Studio import should produce the same format regardless of where the bytes came from, so both clone and upload paths now run the bundle through toWebpLosslessBundle. The helper decodes the texture and re-encodes it with the same VP8L options the SWF importer uses, so the artwork round-trips bit-for-bit, and lets createNitroBundle relabel the member and repair the meta.image pointer. A bundle that is already lossless WebP is returned untouched, making the operation idempotent and safe to run on re-import. A colour variant that shares a library keeps the member base name it arrived with. A texture that cannot be decoded keeps its original format with a warning instead of failing the import: the bundle is valid, and losing a furniture item over a codec edge case is worse than a slightly larger texture. |
||
|
|
306e209e29 |
fix(nitro): normalise uploaded bundles so meta.image matches the texture
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 39s
CI / tests-unit (push) Successful in 2m3s
CI / tests-integration (push) Successful in 2m7s
CI / tests-ui (push) Successful in 2m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m45s
An uploaded .nitro was written to disk byte-for-byte, so a bundle from a third-party tool that ships a WebP member while still pointing spritesheet.meta.image at a .png was accepted and stored as-is. The client resolves the spritesheet through that pointer, so the result was a file that validates fine and then renders nothing. Re-write the bundle through createNitroBundle on import, which labels the member from the actual bytes and repairs the pointer. No texture is re-encoded, so the bytes stay identical, and the member keeps the base name it arrived with so `chair*2` colour variants that share the `chair` library are not renamed. |