The deploy failed with "Expected release never became healthy" after 30
attempts. Root cause: read_active_port() counted the slots answering
/api/health and only consulted the nginx upstream when the count was not
exactly one. On this host both slots were healthy, so it fell back to the
upstream file, but a leftover epicnext-cms:local replica was holding slot A
(3002). The candidate was assigned that occupied port, docker run died with
EADDRINUSE, and the health probe then answered from the pre-existing
container on that port. That container reports release "unknown" because it
was built without NEXT_DEPLOYMENT_ID, so the release comparison could never
match and the deploy timed out blaming a release that was never serving.
read_active_port() now orders its sources by how well they describe reality:
1. The nginx upstream file. It is the only source that says where public
traffic actually enters; everything below it is a consequence.
2. A healthy slot matching that pointer.
3. The other slot when the pointer names a dead port.
4. The pointer itself when nothing answers, so rollback still has a target.
5. Slot A when no upstream file exists at all.
answers_health() was added as a retry-free sibling of healthy(); port
detection should not spend 90 seconds per slot on a process that is either
running now or never will.
start_candidate() now calls assert_port_free() before docker run, so an
occupied port fails immediately and names the listener and the containers
involved, instead of surfacing later as a misleading health-check timeout.
Added scripts/ci-deploy-ports.test.sh, which extracts the two functions from
the real script rather than copying them, and covers the regression: with
both slots healthy and nginx serving slot B, the result must not be slot A.
Verified the test fails against the old logic and passes against the new.
Wired into the check job so this is caught before an image is built.
Dependency updates are handled manually, so the scheduled Renovate job
only cost a daily privileged Docker run on the deploy host. The empty
cache directory it maintained is gone too, and the operations note now
records that updates are manual instead of describing bot behaviour.
Split the heavy test suites out of the check job so coverage, MariaDB/Redis
integration and Playwright UI tests run concurrently on the host runner
(capacity raised to 4) instead of back-to-back (~2min wall-time saving).
Deploy and preflight now gate on all three test jobs.
Point PLAYWRIGHT_BROWSERS_PATH at the persistent /opt/ms-playwright dir on
the host runner so 'playwright install chromium' is an instant no-op after
the first run (was ~100s CDN download per job).
- Switch test-runner and renovate workflows from ubuntu-latest to
self-hosted now that a native host runner is running as a systemd service
- Replace remaining hardcoded color utilities in the homepage with theme
tokens and inline rgba styles to satisfy the no-hardcoded-colors contract
- Restore dual UserAvatarThumbnail usage on the homepage (hero avatar stack
plus community grid) to satisfy the public avatar presentation contract
Run vitest without coverage by default (pnpm test) and add pnpm test:coverage which enforces the coverage thresholds. CI keeps using the coverage run so thresholds are still enforced on every push.
Drop the unused knip dead-code check and the husky+lint-staged pre-commit hook pipeline. All checks remain covered by the CI workflow (lint, typecheck, i18n, tests).
- Add e2e job (needs deploy, main/master only) that installs the
Playwright browser and smoke-tests the live container on :3002
- Keep deploy-job contract slice from bleeding into the e2e job
- Cover the e2e job in the CI workflow contract test
- Ignore Playwright output dirs (test-results, playwright-report,
blob-report)
Zonder dit loopt nieuwe code tegen een oud schema aan zodra een push
migraties bevat. Idempotent (toegepaste migraties worden overgeslagen),
draait op de host met de productie-.env, vóór de container-replace.
docker run --env-file behoudt letterlijke quotes (bewezen test),
waardoor DATABASE_URL ongeldig was en de container crashte. Nu wordt
.env gesourced en elke sleutel met -e doorgegeven: exact dezelfde
waarden als bij de build. Contract-test verbiedt --env-file.
- Deploy kopieert de productie-.env van de host in de build-context:
Next.js bakt NEXT_PUBLIC_* in en valideert DATABASE_URL/HOTEL_NAME
(SKIP_ENV_VALIDATION is verboden voor productie, zie src/env.ts).
- Dockerfile builder installeert git (next.config.ts deploymentId).
- Deploy-container krijgt --env-file + dezelfde volumes als compose,
stopt ook de oude compose-container (poort 3002) en rolt terug via
compose bij een falende health check.
De host heeft Docker iptables uitgeschakeld, dus build-containers op
bridge hebben geen outbound internet; 'npm install -g pnpm' in de
builder-stage hing daardoor. Met --network=host krijgt de build wel
registry-toegang. Contract-test vergrendelt de flag.
pnpm 11 has no --jobs option for install; the stray positional '1'
flipped the command into 'add' mode, which then rejected both
--frozen-lockfile and --jobs ('Unknown options', 'pnpm help add').
Reproduced locally, fixed, verified install succeeds. Add contract
regression test.
Root cause: de job-container (docker mode) heeft GEEN outbound
internet naar GitHub, waardoor actions/checkout@v4 faalde met
'Unable to clone ... i/o timeout'.
Oplossing: zowel check als deploy draaien nu op self-hosted (host)
waar Node 26.8.1 + pnpm 11.25.0 geïnstalleerd zijn en internet
beschikbaar is. Dit is de enige betrouwbare setup in deze omgeving.
Runner is tevens hernoemd naar 'Epic runner'.
- Runner hernoemd naar 'Epic runner'
- Env variabelen als global env (niet per step)
- fetch-depth: 1 voor snellere checkout
- Knip verwijderd (traag, niet kritiek)
- Docker BuildKit caching
- Health check opgeschoond