fix(deploy): free port 3002 and stop compose from competing for the slots
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped

The deploy could not start its candidate because port 3002 was held by
`epicnext-cms`, a `docker compose up` replica built from the `local` image and
serving no traffic. Everything else in the pipeline was healthy: the image
built, the news browser gate passed and migrations were current.

The container was unusable for this pipeline for two reasons. It ran a
different image than any release, and its name did not match the slot the
deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms`
while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to
agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never
could. deploy.sh already documents that compose "never managed the release
that actually ran", so the service was stale by its own account.

Removed the stray container and dropped the `cms` and `cms-green` services (plus
the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot
resurrect a replica that permanently occupies a blue/green slot. byparr is
untouched.

Also fixed the diagnostic from the previous commit, which blamed every running
container. `docker ps --filter publish=` returns nothing for --net=host
containers, so the fallback listed all of them and buried the real holder
among seven innocent ones. It now resolves the listening PID from `ss` back to
its container through /proc/<pid>/cgroup and names only that one, with the
exact `docker rm -f` command to run.

Verified: port 3002 free, live release on 3003 still serving
(status ok, database and redis true), deploy simulation 26 passed, typecheck.
This commit is contained in:
openhands committed 2026-10-03 18:45:51 +02:00
1 parent 8ec3df541e
commit c8b3054527
2 files changed
+57 -69

No files matched your search

+14 -49
View File
@@ -1,54 +1,19 @@
# ─────────────────────────────────────────────────────────────────────────────
# Next.js CMS — blue/green
# ─────────────────────────────────────────────────────────────────────────────
x-cms: &cms
image: epicnext-cms:${CMS_RELEASE:-local}
build:
context: .
dockerfile: Dockerfile
args:
NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}
network: host
network_mode: host
stop_grace_period: 15s
restart: unless-stopped
env_file:
- .env
volumes:
- ./public/nitro-assets:/app/public/nitro-assets
- ./public/swf:/app/public/swf
- ./storage:/app/storage
- /var/www/Gamedata:/var/www/Gamedata
# ── Resource limits ──
mem_limit: 6g
memswap_limit: 7g
cpus: 2.0
pids_limit: 512
healthcheck:
test: ["CMD", "node", "-e", "fetch('http://127.0.0.1:'+(process.env.PORT||'3002')+'/api/health').then(r=>{process.exit(r.ok?0:1)}).catch(()=>process.exit(1))"]
interval: 15s
timeout: 5s
retries: 3
start_period: 40s
# Let op: er staan hier GEEN cms-services meer.
#
# De applicatie wordt uitsluitend door scripts/ci-deploy.sh beheerd, dat een
# blue/green-release doet: de kandidaat start op de vrije poort, de live release
# blijft draaien en nginx wijst pas om als de kandidaat gezond is en de
# e2e-test heeft gewonnen. Een `cms`-service met `container_name: epicnext-cms`
# op poort 3002 was een achterhaalde replica uit een handmatige
# `docker compose up` en blokkeerde elke release op die poort — bovendien met
# een ander containernaam dan wat het deploy-script als slot A beheert
# (`epicnext-cms-app`), waardoor het script de bezette poort niet eens aan de
# juiste container kon toeschrijven.
#
# Wat je hier wilt starten hoort in een apart compose-bestand te staan, of
# expliciet via scripts/ci-deploy.sh te lopen.
services:
cms:
<<: *cms
container_name: epicnext-cms
environment:
- HOSTNAME=0.0.0.0
- PORT=3002
cms-green:
<<: *cms
container_name: epicnext-cms-green
profiles: ["green"]
environment:
- HOSTNAME=0.0.0.0
- PORT=3003
byparr:
image: ghcr.io/thephaseless/byparr:latest
container_name: byparr
+43 -20
View File
@@ -164,27 +164,50 @@ assert_port_free() {
[ -z "$holders" ] && return 0
echo "Port $port is already in use, cannot start candidate $name" >&2
printf '%s\n' "$holders" >&2
echo "--- containers currently running ---" >&2
docker ps --format '{{.Names}}\t{{.Image}}\t{{.Status}}' >&2 || true
# Blame de container die de poort vasthoudt. De blauwe/groene release
# beheert zijn slots zelf, dus een container met een andere naam die hier
# toevallig op dezelfde poort draait (vaak gestart met `docker compose up`)
# is per definitie een losse replica. Zonder deze aanwijzing blijft een
# mislukte deploy een dode hoek in: het log meldde alleen een bezette poort,
# niet wie erachter zat.
echo "--- who holds port $port ---" >&2
docker ps --filter "publish=$port" --format '{{.Names}}\t{{.Image}}\t{{.Status}}' >&2 || true
local squatters=""
while IFS= read -r squatter; do
[ -n "$squatter" ] || continue
case "$squatter" in
"$name"|epicnext-cms-green|epicnext-cms-app) continue ;;
esac
echo " $squatter is not a managed blue/green slot for port $port." >&2
done < <(docker ps --format '{{.Names}}' 2>/dev/null || true)
# Noem exact het container dat de poort vasthoudt.
#
# `docker ps --filter publish=` werkt niet: de app draait met --net=host en
# publiceert dus geen poorten, dus die filter levert altijd niets op. In plaats
# daarvan volgen we de luisterende PID uit `ss` terug naar de container via
# /proc/<pid>/cgroup. Een eerdere versie noemde álle draaiende containers als
# belkenners, wat de echte boosdochter (epicnext-cms) onder een zee van
# onschuldige containers begraven.
local squatter="squatter_pids"
squatter_pids="$(printf '%s\n' "$holders" | grep -oP 'pid=\K[0-9]+' | sort -u || true)"
if [ -n "$squatter_pids" ]; then
local pid cid owner=""
for pid in $squatter_pids; do
cid="$(sed -n 's#.*docker-\([0-9a-f]\{64\}\)\.scope#\1#p' "/proc/$pid/cgroup" 2>/dev/null | head -1)"
[ -n "$cid" ] || continue
owner="$(docker inspect --format '{{.Name}} ({{.Config.Image}})' "$cid" 2>/dev/null || true)"
[ -n "$owner" ] && printf 'Held by container: %s\n' "${owner#/}" >&2
done
fi
echo "" >&2
echo "Stop or remove it, then re-run the deploy. Live traffic is unaffected:" >&2
echo "the nginx upstream keeps serving the other slot until cutover." >&2
# Blauwe/groene releases beheren hun eigen slots. Een container met een andere
# naam die toevallig op een van deze poorten draait — meestal een
# `docker compose up`-replica — staat los van de pipeline en blokkeert de
# release. Live verkeer loopt via het nginx-upstream over het andere slot en is
# dus niet geraakt.
case "$owner" in
*"/$name"*|*"/$slot_b_container"*)
echo "Note: the holder looks like a managed slot container; re-check the port mapping above." >&2 ;;
*)
cat >&2 <<EOF
This port is held by a container that is not a blue/green slot, so the deploy
cannot start the candidate. Live traffic is unaffected: nginx keeps serving
the other slot until cutover.
Remove the stray container and re-run the deploy:
docker rm -f $(printf '%s' "$owner" | sed -n 's#.*/\([^ ]*\).*#\1#p')
If it comes back after a reboot, it is started by docker-compose.yml rather
than by this script; delete or disable that service.
EOF
;;
esac
return 1
}