Author SHA1 Message Date
openhands 7f07c111ac perf(studio): load motion's minimal entry instead of the full component library
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Successful in 1m52s
CI / tests-ui (push) Successful in 2m44s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m34s
/admin/studio/furni sat at 94.9% of its initial-JS budget (427436 of
450560 gzip bytes), so the next feature would have broken the build. Of the
98228 gzip bytes unique to that route, a large part is framer-motion.

This file uses motion twice, for one thing: a 150ms opacity fade on the result
pane when viewMode changes. Importing `motion/react` to get it pulls in
framer-motion's complete component library — 73 internal modules — plus its
render components, drag/gesture and projection code, none of which is
rendered here.

`motion/react-m` ships only the element factories: 2 internal modules, and the
same initial/animate/transition props, so the fade is unchanged. It exports the
elements flat rather than under a `motion.` namespace, so the import becomes
`div as Mdiv` and the two JSX tags are renamed to match.

I could not measure the resulting bundle here: the local build is OOM-killed
(exit 137) with the running containers on the host, so the actual saving is
unverified. The CI build reports it in build-reports, and the number in this
commit message should be read as a hypothesis, not a measurement.

Verified: typecheck clean, lint clean, and the 10 studio UI tests pass —
including the pane and navigation specs that exercise the view switch.
2026-10-03 19:21:02 +02:00
openhands 8218039c64 test(live): stop the live suites inheriting the production database
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 31s
CI / tests-unit (push) Successful in 1m57s
CI / tests-integration (push) Successful in 2m1s
CI / tests-ui (push) Successful in 2m51s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 1m42s
Seven suites read .env with a bare `process.env[key] = value`, which
overwrites whatever the shell already set. That made the DATABASE_URL from
the production .env authoritative, so a single environment variable was
enough to aim them at the live hotel database:

  RUN_CATALOG_AUDIT_LIVE=1 pnpm vitest run src/lib/services/catalog-audit-repair-live.test.ts

Three of those suites then repair the catalog in place: catalog-audit-repair-live
and catalog-repair-direct-live rewrite catalog_items and delete duplicate
classnames, and clone-bulk-import-live bulk-imports every cloneable item. None
of that is undoable, and nothing in their output said the target was
production rather than a sandbox.

Added src/test/live-env.ts with one shared loader, and pointed all seven suites
at it:

- Values already in the real environment win, so an explicit DATABASE_URL on
  the command line is always respected.
- DATABASE_URL defaults to the sandbox on port 3307 rather than inheriting the
  production one from .env.
- Anything that is not loopback is treated as production and redirected.
- Reaching production requires ALLOW_PRODUCTION_LIVE_DB=1 and logs a warning
  saying the suite repairs the catalog.

Tests in src/test/live-env.test.ts run the loader against a temporary .env so
the real project file is never read, and cover the redirect, the shell
override, non-loopback detection, the opt-in and quote stripping. A second
block asserts each of the seven suites no longer contains an inline
`process.env[...] =` assignment. Verified four of them fail against the old
loader.

This does not enable the suites; they stay gated behind their RUN_* flags.
It only removes the possibility of them silently hitting production.

Unit suite: 3330 passed, 12 skipped. Typecheck and lint clean.
2026-10-03 19:03:28 +02:00
openhands f705c67fc7 revert(docker-compose): keep the cms services the contract tests require
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m44s
CI / tests-ui (push) Successful in 2m34s
CI / preflight (push) Skipped
CI / deploy (push) Successful in 2m3s
The previous commit removed the `cms` and `cms-green` services from
docker-compose.yml. That was overreach and it broke two tests:

- src/lib/docker-build-contract.test.ts asserts the compose build passes
  NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}, so a compose-built image
  carries its release id.
- scripts/proxy-config.test.mjs resolves `docker compose config` and asserts
  the `cms` service's host networking, volumes, healthcheck and image tag.

Both encode that docker-compose.yml is a maintained deployment surface, not a
leftover. Removing it was not my call to make while fixing a deploy.

Restored verbatim. The stray container that actually blocked port 3002 is
already gone, and nothing recreates it: there is no systemd unit or pm2
ecosystem that runs `docker compose up`, and `restart: unless-stopped` only
applies to a container that still exists. So the blocker is resolved by the
container removal alone, and compose stays intact for manual and reviewed use.
2026-10-03 18:51:06 +02:00
openhands c8b3054527 fix(deploy): free port 3002 and stop compose from competing for the slots
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy could not start its candidate because port 3002 was held by
`epicnext-cms`, a `docker compose up` replica built from the `local` image and
serving no traffic. Everything else in the pipeline was healthy: the image
built, the news browser gate passed and migrations were current.

The container was unusable for this pipeline for two reasons. It ran a
different image than any release, and its name did not match the slot the
deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms`
while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to
agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never
could. deploy.sh already documents that compose "never managed the release
that actually ran", so the service was stale by its own account.

Removed the stray container and dropped the `cms` and `cms-green` services (plus
the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot
resurrect a replica that permanently occupies a blue/green slot. byparr is
untouched.

Also fixed the diagnostic from the previous commit, which blamed every running
container. `docker ps --filter publish=` returns nothing for --net=host
containers, so the fallback listed all of them and buried the real holder
among seven innocent ones. It now resolves the listening PID from `ss` back to
its container through /proc/<pid>/cgroup and names only that one, with the
exact `docker rm -f` command to run.

Verified: port 3002 free, live release on 3003 still serving
(status ok, database and redis true), deploy simulation 26 passed, typecheck.
2026-10-03 18:45:51 +02:00
openhands 8ec3df541e fix(deploy): name the container blocking a port and silence phantom cleanup
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m43s
CI / tests-unit (push) Successful in 1m47s
CI / tests-ui (push) Successful in 2m33s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m30s
Two follow-ups from the blocked deploy.

The rollback path called `docker logs` and `docker rm -f` on the candidate
unconditionally. When the port check refuses to start it, the container was
never created, so both printed "No such container: epicnext-cms-app" — noise
that looked like a second, unrelated failure and buried the real message.
Both calls are now guarded by `docker inspect`.

assert_port_free() now reports which container holds the port and flags it when
it is not a blue/green slot this script manages. The previous output listed
every container and said only "port already in use", which is a dead end: on
this host the holder is `epicnext-cms` (a `docker compose up` replica on port
3002), while the deploy manages slot A as `epicnext-cms-app`. The names differ
because docker-compose.yml pins `container_name: epicnext-cms` for the `cms`
service; slot B happens to match, which is why 3003 deploys fine and 3002 never
can. The message now names the squatter, explains that live traffic is
unaffected, and gives the next action.

Deploy simulation: 26 passed.
2026-10-03 18:37:12 +02:00
openhands 8ee144745a fix(deploy): stub ss in the deploy harness and cover the port-conflict path
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m47s
CI / tests-unit (push) Successful in 1m54s
CI / tests-ui (push) Successful in 2m40s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 1m41s
The port-conflict guard added in the previous commit made the six existing
blue/green deployment simulation tests fail. assert_port_free() shells out to
ss, and the simulation harness stubs git, curl, docker, nginx, pnpm and node —
but not ss. Because the runner is self-hosted and the containers use
--net=host, the simulation saw the production CMS containers holding 3002 and
3003 and refused to start its own candidate.

The harness now stubs ss. It reports no listener for every scenario except
'port-taken', which reserves whichever port the script asks about, so the
simulation stays independent of the host it runs on.

Also switched the ss probe from `command -v ss` to `type ss`. The stub is a
shell function delivered through BASH_ENV; `command -v` happens to find it,
but `type` is the reliable test for "is this resolvable", and the two differ
across shells.

Added a regression test for the guard itself: with the candidate port already
occupied, the deploy must fail, must not have run `docker run`, and must leave
the nginx upstream untouched on the old port — no half-finished cutover.
Verified it fails when the assert_port_free call is removed.

Deploy simulation: 26 passed. Full unit suite: 3316 passed, 12 skipped.
2026-10-03 18:30:00 +02:00
openhands 64ad9baf39 fix(deploy): trust the nginx upstream when picking the live slot
Gitea Actions Runner Test / test-job (push) Successful in 1s
CI / check (push) Successful in 30s
CI / tests-integration (push) Successful in 1m48s
CI / tests-unit (push) Failing after 1m54s
CI / tests-ui (push) Successful in 2m46s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.

read_active_port() now orders its sources by how well they describe reality:

1. The nginx upstream file. It is the only source that says where public
   traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.

answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.

start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.

Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
2026-10-03 18:22:40 +02:00
openhands 704e33638f fix: restore six useEffect dependencies removed while silencing lint
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 29s
CI / tests-unit (push) Successful in 1m38s
CI / tests-integration (push) Successful in 1m40s
CI / tests-ui (push) Successful in 2m24s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m58s
The previous commit dropped biome-ignore comments to clear
useExhaustiveDependencies diagnostics and, in doing so, also deleted the
dependencies themselves. Six components were left with effects that no longer
react to the state they read. Every one of these is a real behaviour
regression, not a lint preference:

- health-check-client: checkEmulator is a function declaration, so it gets a
  fresh identity each render. As an effect dependency that re-fires the effect
  after every setState, polling /api/admin/devops/health in a loop. Wrapped in
  useCallback so the identity is stable.
- article-recovery: reload restarts the autosave timer for the "Retry recovery"
  button. Without it in the deps that button is a no-op. The counter had been
  renamed to _reload to satisfy the unused-variable rule.
- catalog-integrity-panel: same pattern; refresh starts a new read-only scan,
  so the rescan control did nothing.
- catalog-search: refreshKey re-runs the query after a bulk edit, so results
  were not refreshed after catalog edits. The selection-reset effect also lost
  catalogType, so switching catalog no longer cleared the selection.
- catalog-image-picker: dropped debounced (the search term) and name (the
  error reset), so image search and error state no longer reacted to input.
- icon-picker: dropped iconImage, so a failed load left the placeholder on the
  next icon too.

Each restored dependency carries a biome-ignore with the reason it is
load-bearing, so the diagnostic can be re-derived instead of silently
disappearing again.

Verified: typecheck, lint clean on all six, unit 3315 passed, integration 20
passed, UI 72 passed / 2 skipped.
2026-10-03 18:07:56 +02:00
21 changed files with 559 additions and 158 deletions

No files matched your search

+6
View File
@@ -33,6 +33,12 @@ jobs:
- name: Toolchain check
run: node scripts/check-node-toolchain.mjs
# Port selection decides which blue/green slot stays live. Getting it
# wrong starts the candidate on an occupied port, so the regression that
# caused a failed deploy is covered here, before any image is built.
- name: Deploy port-selection tests
run: bash scripts/ci-deploy-ports.test.sh
- name: Install dependencies
run: pnpm install --frozen-lockfile
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env bash
# Tests for the port-selection and port-conflict logic in scripts/ci-deploy.sh.
#
# Background: on a host where both blue/green slots answer /api/health, the
# original read_active_port() counted healthy slots and only consulted the nginx
# upstream when the count was not exactly 1. With two healthy slots it fell back
# to the upstream file, but an operator `docker compose up` can leave an extra
# replica behind, after which the fallback picked slot A regardless of which slot
# was really live. The candidate then tried to start on an occupied port, and the
# health probe answered from the pre-existing container on that port instead of
# the candidate — producing 30 failed "expected release never became healthy"
# attempts against a release that was never serving.
#
# The functions are extracted from ci-deploy.sh rather than copied so this test
# cannot drift from the script it protects.
set -Eeuo pipefail
deploy_script="$(dirname "$0")/ci-deploy.sh"
[[ -r "$deploy_script" ]] || { echo "cannot read $deploy_script" >&2; exit 1; }
# Pull the two functions out of the real script.
extract() {
sed -n "/^$1() {/,/^}/p" "$deploy_script"
}
read_active_port_fn="$(extract read_active_port)"
assert_port_free_fn="$(extract assert_port_free)"
answers_health_fn="$(extract answers_health)"
if [ -z "$read_active_port_fn" ] || [ -z "$assert_port_free_fn" ] || [ -z "$answers_health_fn" ]; then
echo "could not extract functions from $deploy_script" >&2
exit 1
fi
slot_a_port=3002
slot_b_port=3003
fail() { echo "FAIL: $*" >&2; exit 1; }
# ── read_active_port ──────────────────────────────────────────────────────────
# $1 = upstream body ("none" for a missing file), $2..$3 = ports that answer.
run_read_active_port() {
local body="$1" a="$2" b="$3" tmp
tmp="$(mktemp)"
if [ "$body" = "none" ]; then
tmp=/tmp/ci-deploy-test-nonexistent-upstream-$$
rm -f "$tmp"
else
printf '%s\n' "$body" >"$tmp"
fi
CMS_UPSTREAM_FILE="$tmp" \
PORT_A_HEALTHY="$a" PORT_B_HEALTHY="$b" \
bash -c "
slot_a_port=$slot_a_port
slot_b_port=$slot_b_port
upstream_file=\"\$CMS_UPSTREAM_FILE\"
$read_active_port_fn
# Defined after the extracted function on purpose: answers_health is a
# collaborator here, and the test substitutes a deterministic stub for it.
answers_health() {
local p=\$1 want
case \$p in
$slot_a_port) want=\"\$PORT_A_HEALTHY\" ;;
$slot_b_port) want=\"\$PORT_B_HEALTHY\" ;;
*) want='' ;;
esac
[ \"\$want\" = yes ]
}
read_active_port
echo
" 2>/dev/null
rm -f "$tmp"
}
# nginx points at slot B and both answer -> trust the upstream file.
got="$(run_read_active_port 'server 127.0.0.1:3003 max_fails=2;' yes yes)"
[ "$got" = "$slot_b_port" ] || fail "nginx->3003 with both healthy: got '$got', want 3003"
got="$(run_read_active_port 'server 127.0.0.1:3002 max_fails=2;' yes yes)"
[ "$got" = "$slot_a_port" ] || fail "nginx->3002 with both healthy: got '$got', want 3002"
# The regression: both healthy, nginx points at B, but slot A is an unrelated
# leftover replica. The upstream file is the only thing that knows which slot is
# live, so it must win.
got="$(run_read_active_port 'server 127.0.0.1:3003 max_fails=2;' yes yes)"
[ "$got" != "$slot_a_port" ] || fail "both healthy: fell back to slot A while nginx serves 3003"
# Upstream names a dead slot: fall back to a slot that actually answers, never to
# the dead port itself.
got="$(run_read_active_port 'server 127.0.0.1:3002 max_fails=2;' no yes)"
[ "$got" = "$slot_b_port" ] || fail "nginx->3002 unhealthy, B healthy: got '$got', want 3003"
# Nothing answers at all: read_active_port still has to name a slot, otherwise the
# rollback path has no target.
got="$(run_read_active_port 'server 127.0.0.1:3003 max_fails=2;' no no)"
[ "$got" = "$slot_b_port" ] || fail "nothing healthy: got '$got', want the upstream port 3003"
# No upstream file at all: pick a slot that answers.
got="$(run_read_active_port none no yes)"
[ "$got" = "$slot_b_port" ] || fail "no upstream, B healthy: got '$got', want 3003"
got="$(run_read_active_port none yes no)"
[ "$got" = "$slot_a_port" ] || fail "no upstream, A healthy: got '$got', want 3002"
# ── assert_port_free ─────────────────────────────────────────────────────────
# Runs against real loopback ports: 3999 is intentionally unused, so the check
# must report it free.
bash -c "
$assert_port_free_fn
assert_port_free 3999 candidate >/dev/null 2>&1
" || fail "a port with no listener must be reported as free"
# On this host 3002 is held by a CMS container, so the check must fail. Skip when
# it genuinely is free, otherwise the assertion would be meaningless.
if ss -ltn 2>/dev/null | grep -qE '127\.0\.0\.1:3002|0\.0\.0\.0:3002'; then
if bash -c "
$assert_port_free_fn
assert_port_free 3002 candidate >/dev/null 2>&1
"; then
fail "an occupied port must be rejected, but assert_port_free returned success"
fi
fi
echo 'Deploy port-selection tests passed'
+124 -22
View File
@@ -76,6 +76,13 @@ healthy() {
return 1
}
# Zelfde check als `healthy`, maar zonder retries. Voor het bepalen van de
# actieve poort willen we geen 90 seconden per slot wachten: daar gaat het om
# een al draaiend proces dat nu of nooit antwoordt.
answers_health() {
curl -sf --max-time 5 "http://127.0.0.1:$1/api/health" | grep -q '"database":true'
}
# Staat er een blue/green-upstream? Zonder die bestanden blijft dit script op de
# oude, in-place cutover vallen, zodat een host met een andere nginx-indeling
# niet stilvalt op een upgrade.
@@ -85,29 +92,123 @@ detect_blue_green() {
return 0
}
# Welke poort is op dit moment ÉCHT live? Kijk niet naar het upstream-bestand
# (dat kan door een losse `docker compose up` zijn ingehaald en naar een dood
# slot wijzen), maar test welk slot werkelijk antwoordt op /api/health. Alleen
# in een dubbelzinnige situatie (geen óf beide slots gezond) valt het script
# terug op de huidige nginx-pointer; onbekend = slot A (eerste release op 3002).
# Welke poort is op dit moment ÉCHT live?
#
# Volgorde van vertrouwen:
# 1. Het nginx-upstream-bestand. Dat is de enige bron die aangeeft wáár het
# publieke verkeer daadwerkelijk binnenkomt; alles daaronder is gevolg.
# 2. Een gezond slot dat overeenkomt met die aanwijzing.
# 3. Precies één gezond slot (een verse host met geen upstream-bestand).
#
# De eerdere versie telde gezonde slots en gebruikte de fallback pas als er 0 of
# 2+ waren. Op een host waar beide slots tegelijk gezond zijn — bijvoorbeeld
# doordat een losse `docker compose up` een extra replica heeft achtergelaten —
# gaf dat een willekeurige keuze, en dan kon de kandidaat op een bezette poort
# starten (EADDRINUSE) terwijl de health-check de reeds draaiende container op
# die poort beantwoordde. De release-vergelijking faalde dan 30 keer op een
# container die toevallig een andere release draaide.
read_active_port() {
local live="" port="" result=""
local count=0
for port in "$slot_a_port" "$slot_b_port"; do
if curl -sf --max-time 3 "http://127.0.0.1:$port/api/health" | grep -q '"database":true'; then
live="$live $port"
fi
done
for port in $live; do count=$((count + 1)); result="$port"; done
if [ "$count" -eq 1 ]; then
printf '%s' "$result"
local port="" pointed=""
if [ -r "$upstream_file" ]; then
port="$(grep -oE '127\.0\.0\.1:(3002|3003)' "$upstream_file" 2>/dev/null | head -1 | cut -d: -f2 || true)"
fi
if [ -n "$port" ] && answers_health "$port"; then
printf '%s' "$port"
return 0
fi
port="$(grep -oE '127\.0\.0\.1:[0-9]+' "$upstream_file" 2>/dev/null | head -1 | cut -d: -f2 || true)"
case "$port" in
"$slot_b_port") printf '%s' "$slot_b_port" ;;
*) printf '%s' "$slot_a_port" ;;
# Het upstream-bestand wijst naar een slot dat niet antwoordt. Kies dan het
# enige andere gezonde slot, anders is er niets om op te bouwen.
for candidate in "$slot_a_port" "$slot_b_port"; do
[ "$candidate" = "$port" ] && continue
if answers_health "$candidate"; then
echo "nginx points at ${port:-unknown}, which is unhealthy; ${candidate} answers instead" >&2
printf '%s' "$candidate"
return 0
fi
done
# Geen enkel slot antwoordt. Vertrouw dan op het bestand, zodat een
# rollback-poging toch het vorige slot kan starten.
if [ -n "$port" ]; then
printf '%s' "$port"
return 0
fi
printf '%s' "$slot_a_port"
}
# Poort-bezetting controleren vóór het starten van de kandidaat.
#
# Zonder deze check zorgt `docker run` er stilzwijgend voor dat de kandidaat
# dood gaat op EADDRINUSE, terwijl de health-check ondertussen de reeds draaiende
# container op diezelfde poort beantwoordt. Dat levert een misleidende
# "expected release never became healthy" op in plaats van de echte oorzaak.
# Elke listener wordt hierboven concreet genoemd, inclusief de container die
# hem vasthoudt.
assert_port_free() {
local port="$1" name="$2"
local holders=""
# `type`, niet `command -v`: de deploy-simulatietests leveren `ss` als
# shell-functie via BASH_ENV, en `command -v` herkent die wel op Bash maar de
# functie is niet geëxporteerd naar de subshell van start_candidate. Met `type`
# blijft de stub ook daar zichtbaar, zodat de test geen echte hostpoorten
# hoeft te zien.
if type ss >/dev/null 2>&1; then
# `ss` drukt altijd een kolomkop af, ook als er geen listener is. Filter op
# LISTEN, anders zou elke vrije poort als bezet gemeld worden.
holders="$(ss -ltnp "sport = :$port" 2>/dev/null | grep -F 'LISTEN' || true)"
fi
[ -z "$holders" ] && return 0
echo "Port $port is already in use, cannot start candidate $name" >&2
printf '%s\n' "$holders" >&2
# Noem exact het container dat de poort vasthoudt.
#
# `docker ps --filter publish=` werkt niet: de app draait met --net=host en
# publiceert dus geen poorten, dus die filter levert altijd niets op. In plaats
# daarvan volgen we de luisterende PID uit `ss` terug naar de container via
# /proc/<pid>/cgroup. Een eerdere versie noemde álle draaiende containers als
# belkenners, wat de echte boosdochter (epicnext-cms) onder een zee van
# onschuldige containers begraven.
local squatter="squatter_pids"
squatter_pids="$(printf '%s\n' "$holders" | grep -oP 'pid=\K[0-9]+' | sort -u || true)"
if [ -n "$squatter_pids" ]; then
local pid cid owner=""
for pid in $squatter_pids; do
cid="$(sed -n 's#.*docker-\([0-9a-f]\{64\}\)\.scope#\1#p' "/proc/$pid/cgroup" 2>/dev/null | head -1)"
[ -n "$cid" ] || continue
owner="$(docker inspect --format '{{.Name}} ({{.Config.Image}})' "$cid" 2>/dev/null || true)"
[ -n "$owner" ] && printf 'Held by container: %s\n' "${owner#/}" >&2
done
fi
echo "" >&2
# Blauwe/groene releases beheren hun eigen slots. Een container met een andere
# naam die toevallig op een van deze poorten draait — meestal een
# `docker compose up`-replica — staat los van de pipeline en blokkeert de
# release. Live verkeer loopt via het nginx-upstream over het andere slot en is
# dus niet geraakt.
case "$owner" in
*"/$name"*|*"/$slot_b_container"*)
echo "Note: the holder looks like a managed slot container; re-check the port mapping above." >&2 ;;
*)
cat >&2 <<EOF
This port is held by a container that is not a blue/green slot, so the deploy
cannot start the candidate. Live traffic is unaffected: nginx keeps serving
the other slot until cutover.
Remove the stray container and re-run the deploy:
docker rm -f $(printf '%s' "$owner" | sed -n 's#.*/\([^ ]*\).*#\1#p')
If it comes back after a reboot, it is started by docker-compose.yml rather
than by this script; delete or disable that service.
EOF
;;
esac
return 1
}
# Zet de nginx-upstream op de nieuwe poort en herlaadt graceful.
@@ -153,6 +254,7 @@ switch_upstream() {
# gelden.
start_candidate() {
local port="$1" name="$2"
assert_port_free "$port" "$name"
(
set -a
# shellcheck disable=SC1091
@@ -183,7 +285,7 @@ finish() {
if [ "$cutover_started" -eq 0 ]; then
# De live release draait nog ongestoord; alleen de kandidaat opruimen.
echo "Deployment failed before cutover; the live release was never stopped" >&2
if [ "$candidate_attempted" -eq 1 ] && [ -n "$new_container" ]; then
if [ "$candidate_attempted" -eq 1 ] && [ -n "$new_container" ] && docker inspect "$new_container" >/dev/null 2>&1; then
docker logs "$new_container" --tail 50 >&2 || true
docker rm -f "$new_container" || true
fi
@@ -191,9 +293,9 @@ finish() {
# nginx wijst nu naar de kandidaat. Eerst het verkeer terug, dan pas de
# kandidaat weghalen, anders zou de site 502-en terwijl we terugdraaien.
echo "Deployment failed after cutover; rolling back to port $old_port" >&2
if [ -n "$new_container" ]; then docker logs "$new_container" --tail 50 >&2 || true; fi
if [ -n "$new_container" ] && docker inspect "$new_container" >/dev/null 2>&1; then docker logs "$new_container" --tail 50 >&2 || true; fi
if [ -n "$old_port" ]; then switch_upstream "$old_port" || true; fi
if [ -n "$new_container" ]; then docker rm -f "$new_container" || true; fi
if [ -n "$new_container" ] && docker inspect "$new_container" >/dev/null 2>&1; then docker rm -f "$new_container" || true; fi
if [ -n "$old_container" ] && docker start "$old_container" >/dev/null 2>&1; then
if healthy "$old_port"; then
echo "Rollback verified on port $old_port"
+6 -3
View File
@@ -1,7 +1,7 @@
"use client";
import { RefreshCw } from "lucide-react";
import { useEffect, useState } from "react";
import { useCallback, useEffect, useState } from "react";
import { Badge } from "@/components/ui/badge";
import { Button } from "@/components/ui/button";
@@ -11,7 +11,10 @@ export function HealthCheckClient() {
);
const [checking, setChecking] = useState(false);
async function checkEmulator() {
// Must be stable: the effect below depends on it. A plain function
// declaration gets a new identity on every render, so the effect would
// re-fire after each setState and poll the health endpoint in a loop.
const checkEmulator = useCallback(async () => {
setChecking(true);
setStatus("checking");
try {
@@ -27,7 +30,7 @@ export function HealthCheckClient() {
} finally {
setChecking(false);
}
}
}, []);
useEffect(() => {
checkEmulator();
+7 -2
View File
@@ -20,13 +20,14 @@ export function useArticleRecovery(
ReturnType<typeof loadArticleRecovery>
> | null>(null);
const [message, setMessage] = useState("autosaveLoading");
const [_reload, setReload] = useState(0);
const [reload, setReload] = useState(0);
const version = useRef(0);
const request = useRef<Promise<void> | null>(null);
const last = useRef("");
const dirtyRef = useRef(dirty);
dirtyRef.current = dirty;
const blocked = useRef(true);
// biome-ignore lint/correctness/useExhaustiveDependencies: reload restarts the autosave timer for the recovery retry without remounting; read via setReload
useEffect(() => {
let active = true;
blocked.current = true;
@@ -79,7 +80,11 @@ export function useArticleRecovery(
active = false;
clearInterval(timer);
};
}, [key, form, saving]);
// reload restarts the autosave timer without remounting the editor or
// touching the current form contents; it is what the "Retry recovery"
// button bumps after reconciling server versions. Removing it from the
// deps makes that button a no-op.
}, [key, form, saving, reload]);
return {
draft: recovery?.draft?.payload,
wait: () => request.current ?? Promise.resolve(),
@@ -76,6 +76,7 @@ export function CatalogImagePicker({
);
// Reset and fetch on open / search change
// biome-ignore lint/correctness/useExhaustiveDependencies: debounced is the search term; refetching on its change is intended
useEffect(() => {
if (!open) return;
setImages([]);
@@ -83,7 +84,8 @@ export function CatalogImagePicker({
setHasMore(true);
fetchImages(0, true);
if (scrollRef.current) scrollRef.current.scrollTop = 0;
}, [open, fetchImages]);
// debounced drives the search term, so a changed query must refetch.
}, [open, fetchImages, debounced]);
const handleScroll = useCallback(() => {
const el = scrollRef.current;
@@ -233,7 +235,8 @@ function CatalogImageThumb({
const [error, setError] = useState(0);
// Reset error state when name changes
useEffect(() => setError(0), []);
// biome-ignore lint/correctness/useExhaustiveDependencies: name is a prop; the reset is meant to follow it
useEffect(() => setError(0), [name]);
if (!name || error >= 2) {
return (
@@ -33,9 +33,10 @@ export function CatalogIntegrityPanel({ canRepair }: { canRepair: boolean }) {
const [busy, setBusy] = useState(false);
const [error, setError] = useState<"failed" | "stalePreview" | null>(null);
const [applied, setApplied] = useState<number | null>(null);
const [_refresh, setRefresh] = useState(0);
const [refresh, setRefresh] = useState(0);
const request = useRef(0);
const applying = useRef(false);
// biome-ignore lint/correctness/useExhaustiveDependencies: refresh starts a new read-only scan; read via setRefresh
useEffect(() => {
const version = ++request.current;
setBusy(true);
@@ -57,7 +58,9 @@ export function CatalogIntegrityPanel({ canRepair }: { canRepair: boolean }) {
return () => {
request.current++;
};
}, [catalog]);
// refresh starts a new read-only scan; without it the rescan control
// does nothing.
}, [catalog, refresh]);
async function prepare() {
setBusy(true);
setPreview(null);
+4 -1
View File
@@ -214,7 +214,10 @@ function IconPreview({
size?: number;
}) {
const [error, setError] = useState(false);
useEffect(() => setError(false), []);
// Clear the previous load failure when the icon changes, otherwise a
// placeholder sticks to the next image too.
// biome-ignore lint/correctness/useExhaustiveDependencies: iconImage is a prop; the reset is meant to follow it
useEffect(() => setError(false), [iconImage]);
if (error || iconImage <= 0) {
return (
@@ -26,7 +26,13 @@ import {
Trash2,
X,
} from "lucide-react";
import { motion } from "motion/react";
// motion/react-m is the minimal entry. The full motion/react pulls in
// framer-motion's entire component library (73 internal modules) for what this
// file needs: one fade-in on the result pane. The minimal entry ships only the
// element factories (2 modules) and exposes the same initial/animate/transition
// props, so the fade behaves identically. Only the element factory is
// imported, since that is the single one this file renders.
import { div as Mdiv } from "motion/react-m";
import {
lazy,
Suspense,
@@ -2053,7 +2059,7 @@ export function StudioClient({
</p>
</div>
) : (
<motion.div
<Mdiv
key={viewMode}
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
@@ -2497,7 +2503,7 @@ export function StudioClient({
</span>
)}
</div>
</motion.div>
</Mdiv>
)}
</div>
@@ -48,7 +48,9 @@ export function CatalogSearch({
const [selectionError, setSelectionError] = useState<string | null>(null);
const destinationRequest = useRef(0);
const destinationBusy = useRef(false);
const [_refreshKey, setRefreshKey] = useState(0);
const [refreshKey, setRefreshKey] = useState(0);
// Catalog switches invalidate the selection and destination requests.
// biome-ignore lint/correctness/useExhaustiveDependencies: a catalog switch is exactly what should reset this
useEffect(() => {
destinationRequest.current += 1;
destinationBusy.current = false;
@@ -59,7 +61,7 @@ export function CatalogSearch({
return () => {
destinationRequest.current += 1;
};
}, []);
}, [catalogType]);
async function loadDestinations() {
if (destinationBusy.current) return;
destinationBusy.current = true;
@@ -99,6 +101,7 @@ export function CatalogSearch({
});
if (!pages) void loadDestinations();
}
// biome-ignore lint/correctness/useExhaustiveDependencies: refreshKey re-runs the query after a bulk edit
useEffect(() => {
const request = requests.start();
setResults([]);
@@ -129,7 +132,9 @@ export function CatalogSearch({
clearTimeout(timer);
requests.cancel();
};
}, [query, catalogType, requests]);
// refreshKey re-runs the query after a bulk edit changed rows that this
// search would otherwise keep showing as stale.
}, [query, catalogType, requests, refreshKey]);
return (
<section
aria-label={t("globalSearch")}
+17
View File
@@ -299,6 +299,23 @@ describe("blue/green cutover", () => {
expect(r.upstream).not.toContain("127.0.0.1:3003");
});
// Regressie: een poort die al in gebruik is liet `docker run` stilletjes op
// EADDRINUSE sterven. De health-probe beantwoordde daarna vanaf de
// reeds draaiende container op diezelfde poort, waardoor de
// release-vergelijking 30 keer op een verkeerde release faalde in plaats
// van op de echte oorzaak te wijzen.
it("refuses to start the candidate on a port that is already in use", () => {
const r = simulateBlueGreen("port-taken");
expect(r.status, r.output).not.toBe(0);
expect(r.output).toContain("already in use");
// Er is geen kandidaat gestart, dus er is ook niets om te verwijderen.
expect(r.calls).not.toContain("docker run");
// De live release draait ongestoord door en nginx wijst nog steeds
// naar de oude poort: geen halve cutover.
expect(r.upstream).toContain("127.0.0.1:3002");
expect(r.upstream).not.toContain("127.0.0.1:3003");
});
it.each([
"run-failure",
"health-failure",
@@ -1,9 +1,9 @@
// @ts-nocheck
// Runs repairMissingIcons + repairMissingNitros directly (no source comparison
// phase, which is the slow part of runCatalogAudit). Logs progress to a file.
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CATALOG_AUDIT_LIVE === "1";
const LOG = "/tmp/catalog-asset-repair.log";
@@ -16,21 +16,7 @@ describe.skipIf(!runLive)("catalog asset repair (live)", () => {
let repairNitros: typeof import("@/lib/services/repair-nitros").repairMissingNitros;
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
process.env.SKIP_ENV_VALIDATION = "1";
[repairIcons, repairNitros] = await Promise.all([
+2 -17
View File
@@ -12,11 +12,10 @@
// Usage:
// RUN_CATALOG_AUDIT_LIVE=1 pnpm exec vitest run --coverage.enabled=false \
// src/lib/services/catalog-audit-live.test.ts
import { existsSync, readFileSync } from "node:fs";
import { readdir } from "node:fs/promises";
import { resolve } from "node:path";
import { sql } from "drizzle-orm";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CATALOG_AUDIT_LIVE === "1";
const KNOWN_EVENT_TYPES = new Set([
@@ -41,21 +40,7 @@ describe.skipIf(!runLive)("catalog audit live (read-only)", () => {
beforeAll(async () => {
// vitest's test.env injects a throwaway DATABASE_URL (root:root@:3306).
// Replace it with the real .env value so the audit targets the hotel DB.
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
const [dbMod, auditMod, typesMod] = await Promise.all([
import("@/lib/db"),
@@ -9,9 +9,9 @@
// Run with the REAL DATABASE_URL loaded from .env:
// RUN_CATALOG_AUDIT_LIVE=1 pnpm exec vitest run --coverage.enabled=false \
// src/lib/services/catalog-audit-repair-live.test.ts
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CATALOG_AUDIT_LIVE === "1";
const LOG = "/tmp/catalog-audit-repair.log";
@@ -27,21 +27,7 @@ describe.skipIf(!runLive)("catalog audit live repair", () => {
let runCatalogAudit: typeof import("@/lib/services/catalog-audit").runCatalogAudit;
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
const auditMod = await import("@/lib/services/catalog-audit");
runCatalogAudit = auditMod.runCatalogAudit;
@@ -1,8 +1,8 @@
// @ts-nocheck
// Read-only audit run that writes the full summary to a log file.
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CATALOG_AUDIT_LIVE === "1";
const LOG = "/tmp/catalog-audit-summary.log";
@@ -14,21 +14,7 @@ describe.skipIf(!runLive)("catalog audit read-only summary", () => {
let runCatalogAudit: typeof import("@/lib/services/catalog-audit").runCatalogAudit;
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
const auditMod = await import("@/lib/services/catalog-audit");
runCatalogAudit = auditMod.runCatalogAudit;
@@ -1,8 +1,7 @@
// @ts-nocheck
// Direct repair execution against the live DB. Skips the slow clone-source
// comparison entirely and just runs the repair functions.
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
const LOG = "/tmp/catalog-repair.log";
const log = (msg: string) => {
@@ -12,6 +11,7 @@ const log = (msg: string) => {
import { sql } from "drizzle-orm";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CATALOG_AUDIT_LIVE === "1";
@@ -20,21 +20,7 @@ describe.skipIf(!runLive)("catalog repair direct (live)", () => {
let repair: typeof import("@/lib/services/catalog-repair");
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
const [dbMod, repairMod] = await Promise.all([
import("@/lib/db"),
@@ -2,9 +2,9 @@
// Bulk-import live run: imports every cloneable item from the sources that
// reliably serve .nitro assets (SodaStudios, Hubbly). One-off operation to
// shrink missingFromSources before disabling dead sources.
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CLONE_BULK_LIVE === "1";
const LOG = "/tmp/clone-bulk.log";
@@ -15,21 +15,7 @@ const IMPORT_IDS = ["default-sodastudios", "default-hubbly"];
describe.skipIf(!runLive)("clone bulk import", () => {
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
});
it("imports clonable items from delivering sources", async () => {
@@ -2,9 +2,9 @@
// Feasibility probe: splits the audit's missing-from-sources list per clone
// source, marks which sources can deliver .nitro assets, and runs a tiny pilot
// import (N items) to measure per-item cost. Writes a report to /tmp.
import { appendFileSync, existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { appendFileSync } from "node:fs";
import { beforeAll, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "@/test/live-env";
const runLive = process.env.RUN_CLONE_FEASIBILITY_LIVE === "1";
const LOG = "/tmp/clone-feasibility.log";
@@ -12,21 +12,7 @@ const log = (msg: string) => appendFileSync(LOG, `${msg}\n`);
describe.skipIf(!runLive)("clone feasibility", () => {
beforeAll(async () => {
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const m = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!m) continue;
let value = m[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
process.env[m[1]] = value;
}
}
loadEnvForLiveTests();
});
it("reports per-source clonable counts + pilot", async () => {
+23 -1
View File
@@ -38,6 +38,28 @@ curl() {
echo '{"database":true}'
}
sleep() { :; }
# De poortconflict-check (assert_port_free in scripts/ci-deploy.sh) leest de
# echte luisterende sockets. Zonder deze stub zou de simulatie de containers van
# de productiehost zien — de test draait immers op dezelfde machine — en elke
# blauwe/groene scenario laten falen op een poort die de simulatie zelf net
# virtueel heeft toegewezen. Standaard is geen enkele poort bezet; het scenario
# 'port-taken' reserveert expliciet de kandidaatpoort.
ss() {
local requested="" arg
for arg in "$@"; do
case "$arg" in
sport=*) requested="${arg#sport=}" ;;
'sport'|'='|':') ;;
*:*) [ -z "$requested" ] && requested="${arg##*:}" ;;
esac
done
if [ "$SCENARIO" = port-taken ] && [ -n "$requested" ]; then
printf 'State Recv-Q Send-Q Local Address:Port Peer Address:Port\n'
printf 'LISTEN 0 511 0.0.0.0:%s 0.0.0.0:*\n' "$requested"
return 0
fi
printf 'State Recv-Q Send-Q Local Address:Port Peer Address:Port\n'
}
docker() {
echo "docker $*" >> "$TEST_DIR/calls"
local name="${@: -1}"
@@ -70,7 +92,7 @@ docker() {
*) return 0 ;;
esac
}
export -f git flock pnpm curl sleep docker nginx candidate_is_new
export -f git flock pnpm curl sleep ss docker nginx candidate_is_new
node() {
echo "node $*" >> "$TEST_DIR/calls"
+126
View File
@@ -0,0 +1,126 @@
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join, resolve } from "node:path";
import { afterEach, describe, expect, it } from "vitest";
import { loadEnvForLiveTests } from "./live-env";
const ORIGINAL_ENV = { ...process.env };
const SANDBOX = "mysql://[email protected]:3307/habbo_sandbox?charset=utf8mb4";
/**
* Run the helper against a temporary .env so the real project file is never
* read and the ambient environment cannot leak in.
*/
function runWithDotEnv(dotEnv: string | null, shell: Record<string, string>) {
const dir = mkdtempSync(join(tmpdir(), "live-env-test-"));
const previousCwd = process.cwd();
try {
if (dotEnv !== null) writeFileSync(join(dir, ".env"), dotEnv);
// process.chdir keeps process.cwd() inside the helper pointing at dir.
process.chdir(dir);
for (const key of Object.keys(process.env)) delete process.env[key];
Object.assign(process.env, shell);
loadEnvForLiveTests();
return { ...process.env };
} finally {
process.chdir(previousCwd);
rmSync(dir, { recursive: true, force: true });
}
}
afterEach(() => {
for (const key of Object.keys(process.env)) delete process.env[key];
Object.assign(process.env, ORIGINAL_ENV);
});
describe("loadEnvForLiveTests", () => {
// The suites these protect rewrite the catalog, delete duplicate
// classnames and bulk-import thousands of rows. A single env var used to be
// enough to aim them at the live hotel database.
it("never inherits a production database from .env", () => {
const env = runWithDotEnv(
[
"DATABASE_URL='mysql://cms:[email protected]:3306/habbo'",
"HOTEL_NAME='Test Hotel'",
].join("\n"),
{},
);
expect(env.DATABASE_URL).toBe(SANDBOX);
// Non-database settings still load, so the suites stay usable.
expect(env.HOTEL_NAME).toBe("Test Hotel");
});
it("keeps an explicit shell override", () => {
const env = runWithDotEnv(
"DATABASE_URL='mysql://cms:[email protected]:3306/habbo'",
{
DATABASE_URL: "mysql://[email protected]:3399/other?charset=utf8mb4",
},
);
expect(env.DATABASE_URL).toBe(
"mysql://[email protected]:3399/other?charset=utf8mb4",
);
});
it("treats any non-loopback database as production", () => {
const env = runWithDotEnv(null, {
DATABASE_URL: "mysql://user:[email protected]:3306/habbo",
});
expect(env.DATABASE_URL).toBe(SANDBOX);
});
it("allows production only behind an explicit opt-in", () => {
const production = "mysql://cms:[email protected]:3306/habbo";
const env = runWithDotEnv(null, {
DATABASE_URL: production,
ALLOW_PRODUCTION_LIVE_DB: "1",
});
expect(env.DATABASE_URL).toBe(production);
});
it("stays on the sandbox when the opt-in is set but the URL is safe", () => {
const env = runWithDotEnv(null, {
DATABASE_URL: "mysql://[email protected]:3307/habbo_sandbox?charset=utf8mb4",
ALLOW_PRODUCTION_LIVE_DB: "1",
});
expect(env.DATABASE_URL).toBe(SANDBOX);
});
it("supplies a sandbox when .env has no database at all", () => {
const env = runWithDotEnv("HOTEL_NAME='Test Hotel'", {});
expect(env.DATABASE_URL).toBe(SANDBOX);
});
it("strips quotes the way the previous inline loader did", () => {
const env = runWithDotEnv(
['AUTH_SECRET="quoted-secret"', "OTHER='single'"].join("\n"),
{},
);
expect(env.AUTH_SECRET).toBe("quoted-secret");
expect(env.OTHER).toBe("single");
});
});
describe("live suites no longer inline the .env loader", () => {
const suites = [
"catalog-audit-repair-live",
"catalog-audit-live",
"clone-bulk-import-live",
"clone-feasibility-live",
"catalog-repair-direct-live",
"catalog-asset-repair-live",
"catalog-audit-summary-live",
];
it.each(suites)("%s delegates to the shared loader", (name) => {
const source = readFileSync(
resolve(process.cwd(), "src/lib/services", `${name}.test.ts`),
"utf8",
);
// A bare assignment would overwrite an operator's explicit DATABASE_URL,
// which is how production used to win.
expect(source).not.toMatch(/process\.env\[m\[1\]\]\s*=/);
expect(source).toContain("loadEnvForLiveTests()");
});
});
+75
View File
@@ -0,0 +1,75 @@
import { existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
/**
* Sandbox database for the live suites. Production runs on port 3306; the
* throwaway MariaDB used by the rehearsal suites listens on 3307.
*/
export const SANDBOX_DATABASE_URL =
"mysql://[email protected]:3307/habbo_sandbox?charset=utf8mb4";
/**
* Load .env for the live suites without ever pointing them at production.
*
* Every live suite used to read .env with a bare
* `process.env[key] = value`, which overwrites whatever the shell already set.
* That made the DATABASE_URL from the production .env authoritative: setting
* RUN_CATALOG_AUDIT_LIVE=1 pointed three suites — catalog-audit-repair-live,
* catalog-repair-direct-live and clone-bulk-import-live — at the live hotel
* database, where they rewrite the catalog, delete duplicate classnames and
* bulk-import thousands of rows. The failure mode was silent: the gate is a
* single environment variable, and nothing in the suite said the target was
* production.
*
* Rules now:
* 1. Values already present in the real environment win, so an explicit
* DATABASE_URL on the command line is always respected.
* 2. DATABASE_URL defaults to the sandbox, never to production.
* 3. Targeting production requires ALLOW_PRODUCTION_LIVE_DB=1 and says so
* loudly, because it is destructive and not undoable.
*/
export function loadEnvForLiveTests(): void {
const allowProduction = process.env.ALLOW_PRODUCTION_LIVE_DB === "1";
const envFile = resolve(process.cwd(), ".env");
if (existsSync(envFile)) {
for (const line of readFileSync(envFile, "utf8").split(/\r?\n/)) {
const match = line.match(/^\s*([A-Za-z_][A-Za-z0-9_]*)\s*=\s*(.*?)\s*$/);
if (!match) continue;
let value = match[2].trim();
if (
(value.startsWith('"') && value.endsWith('"')) ||
(value.startsWith("'") && value.endsWith("'"))
) {
value = value.slice(1, -1);
}
// Already-set values win: the shell is a deliberate override.
if (process.env[match[1]] === undefined) {
process.env[match[1]] = value;
}
}
}
const databaseUrl = process.env.DATABASE_URL ?? "";
const targetsProduction =
databaseUrl.includes(":3306") ||
(databaseUrl.includes("@") && !databaseUrl.includes("127.0.0.1"));
if (allowProduction) {
if (targetsProduction) {
console.warn(
"[live-test] ALLOW_PRODUCTION_LIVE_DB=1 and DATABASE_URL points at " +
"a remote database. This suite repairs the catalog in place.",
);
}
return;
}
if (targetsProduction || !databaseUrl) {
process.env.DATABASE_URL = SANDBOX_DATABASE_URL;
console.warn(
`[live-test] redirected DATABASE_URL to the sandbox at ${SANDBOX_DATABASE_URL}. ` +
"Set ALLOW_PRODUCTION_LIVE_DB=1 to override.",
);
}
}