Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 32s
CI / tests-unit (push) Successful in 1m49s
CI / tests-ui (push) Successful in 2m33s
CI / tests-integration (push) Successful in 1m50s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 2m26s
Four production defects, all found by auditing the running host rather than the code. Each one had a signature that looked like a network or permissions problem and was actually a configuration or ordering bug. jobs-worker never ran `import "./load-env"` sat on line 3 of scripts/jobs-worker.ts, but ESM evaluates a module's imports in source order and the first import reaches `@/env`, which validates process.env at import time. The ZodError on DATABASE_URL therefore fired before load-env ever executed, so the worker could only start from a shell that had already exported the configuration. Nothing supervised it either, so scheduled articles, catalog export, JAR and database backups, disk alerts and the ops health probe have all been dead; `cms:jobs-worker:heartbeat` did not exist. Moved the import to the top and added deployment/systemd/cms-jobs-worker.service with Restart=always. The JAR backup additionally pointed at './emulator/Arcturus.jar', which does not exist and would go stale on the next emulator upgrade. resolveEmulatorJar now accepts a file, a directory or a wildcard and picks the newest JAR, the same way emulator.service picks its build, and reports an unresolvable path once instead of logging an opaque copyFile ENOENT every night. /api/health answered 200 with the database down The route documented this as intentional, and ci-deploy.sh worked around it by grepping the body for '"database":true'. The container healthcheck did not, so Docker reported containers healthy while every page 500'd. The status is now load-bearing: 503 when the database is unreachable, 200 otherwise. Redis and the emulator deliberately do not fail the container — both have in-process fallbacks, so failing them would trade a slow site for an outage. The runtime had no V8 heap cap NODE_OPTIONS existed only in the builder stage. With no cap, V8 sized its heap from host memory (23.5 GB) while the container was limited to 4 GB, so the kernel OOM-killed the process mid-request — the same failure mode as the 14 host-wide `next-build` kills. docker-start.mjs now reads the cgroup limit (v2 with a v1 fallback) and sets 70% of it, respecting an explicit override. Storage ownership was only repaired for one path ci-deploy.sh chowned storage/imaging and nothing else, so storage/catalog-git/hotel-status.json kept coming back root:root and /api/admin/catalog/status kept throwing EACCES. All eight writable storage paths are repaired now. The silent-failure mode is the reason this mattered: these writes sit inside try/catch, so a wrong owner looks like a slow page rather than an error. nginx: robots.txt was a guaranteed 404, and TLS never resumed `index index.html` without a `root` left every try_files resolving against /etc/nginx/html, which sits behind a 0750 directory — the worker got EACCES on each stat and nginx logs a failed stat at crit, which is where 149 crit lines per scan came from. robots.txt answered from that same broken location, so crawlers were pointed at a file they could never read while sitemap.xml kept advertising it. Added `root`, proxied robots.txt to the CMS, added ssl_session_cache (there was no session resumption at all), and set Restart=on-failure in a systemd override, since the packaged unit ships Restart=no and nginx is the only thing serving the site. Verified against the running host: 3379 tests, typecheck and biome clean, nginx -t passes, health returns 200 with every check green, and the worker has run for hours at NRestarts=0 with a heartbeat refreshing each minute.
129 lines
4.6 KiB
Bash
Executable File
129 lines
4.6 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Sync the nginx config from this repository to /etc/nginx and reload it.
|
|
#
|
|
# Background: on 2026-09-26 /etc/nginx and /var/log/nginx disappeared from the
|
|
# host while nginx kept serving its in-memory config; any restart would have
|
|
# taken the CMS down. This script makes the repo the source of truth so that
|
|
# cannot happen again. It is idempotent and only reloads nginx when the config
|
|
# actually changed.
|
|
#
|
|
# Usage:
|
|
# sudo scripts/nginx-sync.sh # install + test + reload if changed
|
|
# sudo scripts/nginx-sync.sh --force # always reload after a passing test
|
|
# scripts/nginx-sync.sh --check # just diff repo vs live, no writes
|
|
#
|
|
# Files installed (see also deployment/proxy/):
|
|
# nginx.conf -> /etc/nginx/nginx.conf
|
|
# nginx-mime.types -> /etc/nginx/mime.types
|
|
# nginx-cms.conf -> /etc/nginx/sites-available/cms.conf
|
|
# cloudflare-ips.conf -> /etc/nginx/conf.d/cloudflare-ips.conf
|
|
# cms_upstream_servers.conf -> /etc/nginx/snippets/cms_upstream_servers.conf
|
|
# (seed alleen als het bestand ontbreekt; zodra het bestaat is het runtime
|
|
# eigendom van scripts/ci-deploy.sh en wordt het hier nooit overschreven)
|
|
# symlink sites-enabled/cms.conf -> ../sites-available/cms.conf
|
|
set -euo pipefail
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
PROXY_DIR="$SCRIPT_DIR/../deployment/proxy"
|
|
NGINX_DIR=/etc/nginx
|
|
BACKUP_DIR="/var/backups/nginx-$(date +%Y%m%d-%H%M%S)"
|
|
MODE="sync"
|
|
|
|
for arg in "$@"; do
|
|
case "$arg" in
|
|
--force) MODE="force" ;;
|
|
--check) MODE="check" ;;
|
|
esac
|
|
done
|
|
|
|
install_file() {
|
|
local src="$1" dst="$2"
|
|
if [[ ! -f "$src" ]]; then
|
|
echo "error: $src not found in repo" >&2
|
|
exit 1
|
|
fi
|
|
if [[ -f "$dst" ]] && cmp -s "$src" "$dst"; then
|
|
echo "= $dst up to date"
|
|
return 1
|
|
fi
|
|
if [[ "$MODE" == "check" ]]; then
|
|
echo "- $dst differs from repo"
|
|
return 0
|
|
fi
|
|
mkdir -p "$(dirname "$dst")"
|
|
if [[ -f "$dst" ]]; then
|
|
mkdir -p "$BACKUP_DIR"
|
|
cp -a "$dst" "$BACKUP_DIR/"
|
|
fi
|
|
cp -a "$src" "$dst"
|
|
echo "+ installed $dst"
|
|
return 0
|
|
}
|
|
|
|
if [[ "$MODE" != "check" && "$(id -u)" -ne 0 ]]; then
|
|
echo "error: run as root (sudo scripts/nginx-sync.sh)" >&2
|
|
exit 1
|
|
fi
|
|
|
|
changed=0
|
|
if install_file "$PROXY_DIR/nginx.conf" "$NGINX_DIR/nginx.conf"; then changed=1; fi
|
|
if install_file "$PROXY_DIR/nginx-mime.types" "$NGINX_DIR/mime.types"; then changed=1; fi
|
|
if install_file "$PROXY_DIR/nginx-cms.conf" "$NGINX_DIR/sites-available/cms.conf"; then changed=1; fi
|
|
if install_file "$PROXY_DIR/cloudflare-ips.conf" "$NGINX_DIR/conf.d/cloudflare-ips.conf"; then changed=1; fi
|
|
# De upstream-snippet is runtime-eigendom van ci-deploy.sh (blue/green): alleen
|
|
# aanmaken op een verse host, nooit overschrijven wat een deploy heeft gezet.
|
|
if [[ -f "$NGINX_DIR/snippets/cms_upstream_servers.conf" ]]; then
|
|
echo "= $NGINX_DIR/snippets/cms_upstream_servers.conf managed by ci-deploy.sh (untouched)"
|
|
else
|
|
if install_file "$PROXY_DIR/cms_upstream_servers.conf" "$NGINX_DIR/snippets/cms_upstream_servers.conf"; then changed=1; fi
|
|
fi
|
|
|
|
if [[ ! -f "$NGINX_DIR/sites-enabled/cms.conf" ]]; then
|
|
if [[ "$MODE" == "check" ]]; then
|
|
echo "- sites-enabled/cms.conf missing"
|
|
changed=1
|
|
else
|
|
ln -sf ../sites-available/cms.conf "$NGINX_DIR/sites-enabled/cms.conf"
|
|
echo "+ linked sites-enabled/cms.conf"
|
|
changed=1
|
|
fi
|
|
fi
|
|
|
|
if [[ "$MODE" == "check" ]]; then
|
|
[[ "$changed" -eq 0 ]]
|
|
exit
|
|
fi
|
|
|
|
if [[ "$MODE" == "force" ]]; then
|
|
changed=1
|
|
fi
|
|
|
|
if [[ ! -d /var/log/nginx ]]; then
|
|
install -d -o root -g adm -m 750 /var/log/nginx
|
|
fi
|
|
# nginx-cms.conf sets `root /var/www/html` so that disk-backed locations
|
|
# (favicon.ico) resolve somewhere the www-data worker can actually traverse.
|
|
# The previous implicit root was /etc/nginx/html, which sits behind /etc/nginx
|
|
# (0750 root:root): the worker got EACCES on every stat, and nginx logs a
|
|
# failed stat at crit, so each crawler probe wrote a crit line.
|
|
if [[ ! -d /var/www/html ]]; then
|
|
install -d -o root -g root -m 755 /var/www/html
|
|
echo "+ created /var/www/html (document root)"
|
|
fi
|
|
for f in /var/log/nginx/access.log /var/log/nginx/error.log; do
|
|
[[ -f "$f" ]] || touch "$f"
|
|
done
|
|
|
|
echo "--- nginx -t ---"
|
|
nginx -t
|
|
|
|
if [[ "$changed" -eq 1 ]]; then
|
|
echo "--- reloading nginx ---"
|
|
nginx -s reload
|
|
else
|
|
echo "no changes; nginx reload skipped"
|
|
fi
|
|
|
|
echo "--- health check ---"
|
|
curl -sf "http://127.0.0.1:3002/api/health" > /dev/null && echo "OK: CMS reachable"
|
|
curl -skf -o /dev/null -H "Host: epicnabbo.nl" "https://127.0.0.1:9443/health" && echo "OK: nginx :9443 /health" |