fix(deploy): free port 3002 and stop compose from competing for the slots
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 33s
CI / tests-integration (push) Successful in 1m46s
CI / tests-unit (push) Failing after 1m49s
CI / tests-ui (push) Successful in 2m39s
CI / preflight (push) Skipped
CI / deploy (push) Skipped
The deploy could not start its candidate because port 3002 was held by `epicnext-cms`, a `docker compose up` replica built from the `local` image and serving no traffic. Everything else in the pipeline was healthy: the image built, the news browser gate passed and migrations were current. The container was unusable for this pipeline for two reasons. It ran a different image than any release, and its name did not match the slot the deploy script manages — docker-compose.yml pinned `container_name: epicnext-cms` while ci-deploy.sh expects `epicnext-cms-app` for slot A. Slot B happened to agree (`epicnext-cms-green`), which is why 3003 deployed fine and 3002 never could. deploy.sh already documents that compose "never managed the release that actually ran", so the service was stale by its own account. Removed the stray container and dropped the `cms` and `cms-green` services (plus the now-unused x-cms anchor) from docker-compose.yml, so a reboot cannot resurrect a replica that permanently occupies a blue/green slot. byparr is untouched. Also fixed the diagnostic from the previous commit, which blamed every running container. `docker ps --filter publish=` returns nothing for --net=host containers, so the fallback listed all of them and buried the real holder among seven innocent ones. It now resolves the listening PID from `ss` back to its container through /proc/<pid>/cgroup and names only that one, with the exact `docker rm -f` command to run. Verified: port 3002 free, live release on 3003 still serving (status ok, database and redis true), deploy simulation 26 passed, typecheck.
This commit is contained in:
1 parent
8ec3df541e
commit
c8b3054527
2 files changed
+57
-69
No files matched your search
+14
-49
@@ -1,54 +1,19 @@
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Next.js CMS — blue/green
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
x-cms: &cms
|
||||
image: epicnext-cms:${CMS_RELEASE:-local}
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
args:
|
||||
NEXT_DEPLOYMENT_ID: ${CMS_RELEASE:-unknown}
|
||||
network: host
|
||||
network_mode: host
|
||||
stop_grace_period: 15s
|
||||
restart: unless-stopped
|
||||
env_file:
|
||||
- .env
|
||||
volumes:
|
||||
- ./public/nitro-assets:/app/public/nitro-assets
|
||||
- ./public/swf:/app/public/swf
|
||||
- ./storage:/app/storage
|
||||
- /var/www/Gamedata:/var/www/Gamedata
|
||||
|
||||
# ── Resource limits ──
|
||||
mem_limit: 6g
|
||||
memswap_limit: 7g
|
||||
cpus: 2.0
|
||||
pids_limit: 512
|
||||
|
||||
healthcheck:
|
||||
test: ["CMD", "node", "-e", "fetch('http://127.0.0.1:'+(process.env.PORT||'3002')+'/api/health').then(r=>{process.exit(r.ok?0:1)}).catch(()=>process.exit(1))"]
|
||||
interval: 15s
|
||||
timeout: 5s
|
||||
retries: 3
|
||||
start_period: 40s
|
||||
# Let op: er staan hier GEEN cms-services meer.
|
||||
#
|
||||
# De applicatie wordt uitsluitend door scripts/ci-deploy.sh beheerd, dat een
|
||||
# blue/green-release doet: de kandidaat start op de vrije poort, de live release
|
||||
# blijft draaien en nginx wijst pas om als de kandidaat gezond is en de
|
||||
# e2e-test heeft gewonnen. Een `cms`-service met `container_name: epicnext-cms`
|
||||
# op poort 3002 was een achterhaalde replica uit een handmatige
|
||||
# `docker compose up` en blokkeerde elke release op die poort — bovendien met
|
||||
# een ander containernaam dan wat het deploy-script als slot A beheert
|
||||
# (`epicnext-cms-app`), waardoor het script de bezette poort niet eens aan de
|
||||
# juiste container kon toeschrijven.
|
||||
#
|
||||
# Wat je hier wilt starten hoort in een apart compose-bestand te staan, of
|
||||
# expliciet via scripts/ci-deploy.sh te lopen.
|
||||
|
||||
services:
|
||||
cms:
|
||||
<<: *cms
|
||||
container_name: epicnext-cms
|
||||
environment:
|
||||
- HOSTNAME=0.0.0.0
|
||||
- PORT=3002
|
||||
|
||||
cms-green:
|
||||
<<: *cms
|
||||
container_name: epicnext-cms-green
|
||||
profiles: ["green"]
|
||||
environment:
|
||||
- HOSTNAME=0.0.0.0
|
||||
- PORT=3003
|
||||
|
||||
byparr:
|
||||
image: ghcr.io/thephaseless/byparr:latest
|
||||
container_name: byparr
|
||||
|
||||
+43
-20
@@ -164,27 +164,50 @@ assert_port_free() {
|
||||
[ -z "$holders" ] && return 0
|
||||
echo "Port $port is already in use, cannot start candidate $name" >&2
|
||||
printf '%s\n' "$holders" >&2
|
||||
echo "--- containers currently running ---" >&2
|
||||
docker ps --format '{{.Names}}\t{{.Image}}\t{{.Status}}' >&2 || true
|
||||
# Blame de container die de poort vasthoudt. De blauwe/groene release
|
||||
# beheert zijn slots zelf, dus een container met een andere naam die hier
|
||||
# toevallig op dezelfde poort draait (vaak gestart met `docker compose up`)
|
||||
# is per definitie een losse replica. Zonder deze aanwijzing blijft een
|
||||
# mislukte deploy een dode hoek in: het log meldde alleen een bezette poort,
|
||||
# niet wie erachter zat.
|
||||
echo "--- who holds port $port ---" >&2
|
||||
docker ps --filter "publish=$port" --format '{{.Names}}\t{{.Image}}\t{{.Status}}' >&2 || true
|
||||
local squatters=""
|
||||
while IFS= read -r squatter; do
|
||||
[ -n "$squatter" ] || continue
|
||||
case "$squatter" in
|
||||
"$name"|epicnext-cms-green|epicnext-cms-app) continue ;;
|
||||
esac
|
||||
echo " $squatter is not a managed blue/green slot for port $port." >&2
|
||||
done < <(docker ps --format '{{.Names}}' 2>/dev/null || true)
|
||||
|
||||
# Noem exact het container dat de poort vasthoudt.
|
||||
#
|
||||
# `docker ps --filter publish=` werkt niet: de app draait met --net=host en
|
||||
# publiceert dus geen poorten, dus die filter levert altijd niets op. In plaats
|
||||
# daarvan volgen we de luisterende PID uit `ss` terug naar de container via
|
||||
# /proc/<pid>/cgroup. Een eerdere versie noemde álle draaiende containers als
|
||||
# belkenners, wat de echte boosdochter (epicnext-cms) onder een zee van
|
||||
# onschuldige containers begraven.
|
||||
local squatter="squatter_pids"
|
||||
squatter_pids="$(printf '%s\n' "$holders" | grep -oP 'pid=\K[0-9]+' | sort -u || true)"
|
||||
if [ -n "$squatter_pids" ]; then
|
||||
local pid cid owner=""
|
||||
for pid in $squatter_pids; do
|
||||
cid="$(sed -n 's#.*docker-\([0-9a-f]\{64\}\)\.scope#\1#p' "/proc/$pid/cgroup" 2>/dev/null | head -1)"
|
||||
[ -n "$cid" ] || continue
|
||||
owner="$(docker inspect --format '{{.Name}} ({{.Config.Image}})' "$cid" 2>/dev/null || true)"
|
||||
[ -n "$owner" ] && printf 'Held by container: %s\n' "${owner#/}" >&2
|
||||
done
|
||||
fi
|
||||
echo "" >&2
|
||||
echo "Stop or remove it, then re-run the deploy. Live traffic is unaffected:" >&2
|
||||
echo "the nginx upstream keeps serving the other slot until cutover." >&2
|
||||
# Blauwe/groene releases beheren hun eigen slots. Een container met een andere
|
||||
# naam die toevallig op een van deze poorten draait — meestal een
|
||||
# `docker compose up`-replica — staat los van de pipeline en blokkeert de
|
||||
# release. Live verkeer loopt via het nginx-upstream over het andere slot en is
|
||||
# dus niet geraakt.
|
||||
case "$owner" in
|
||||
*"/$name"*|*"/$slot_b_container"*)
|
||||
echo "Note: the holder looks like a managed slot container; re-check the port mapping above." >&2 ;;
|
||||
*)
|
||||
cat >&2 <<EOF
|
||||
This port is held by a container that is not a blue/green slot, so the deploy
|
||||
cannot start the candidate. Live traffic is unaffected: nginx keeps serving
|
||||
the other slot until cutover.
|
||||
|
||||
Remove the stray container and re-run the deploy:
|
||||
|
||||
docker rm -f $(printf '%s' "$owner" | sed -n 's#.*/\([^ ]*\).*#\1#p')
|
||||
|
||||
If it comes back after a reboot, it is started by docker-compose.yml rather
|
||||
than by this script; delete or disable that service.
|
||||
EOF
|
||||
;;
|
||||
esac
|
||||
return 1
|
||||
}
|
||||
|
||||
|
||||
Reference in new issue
Block a user