fix(ci): make the lint gate fail for real and stop byparr leaking disk
Gitea Actions Runner Test / test-job (push) Successful in 2s
CI / check (push) Successful in 36s
CI / tests-integration (push) Successful in 2m3s
CI / tests-unit (push) Successful in 2m18s
CI / tests-ui (push) Successful in 3m6s
CI / preflight (push) Skipped
CI / deploy (push) Failing after 3m14s

The CI lint step was `biome check . || true`, so it could never fail: 14 real
violations were passing unnoticed. Drop the `|| true` and fix what it found.

Lint fixes, none of which change behaviour:
- give list items their natural identity instead of the array index
  (key={c} / key={char}, key={`skeleton-${i}`})
- document the two useEffect dependency lists that must keep their
  function-declaration handlers, with the reasoning that dropping them broke
  the tree and save-on-Ctrl+S once already (704e3363)
- scope the remaining noArrayIndexKey / useExhaustiveDependencies exemptions to
  the three files that need them, in biome.json instead of scattered comments

Storage, on a host that had grown to 81% disk:
- byparr starts a Firefox per request and never removes the profile it leaves in
  the container's writable layer. With no volume mounted, nothing else reclaimed
  it: 716 profiles / 6.8 GB in two days, ~1.7 GB/day. docker-prune.sh now removes
  orphaned profiles, identifying live ones by the open fd in /proc/<pid>/fd rather
  than by age, because browsers stay warm for ~27 hours here — longer than the
  leak window, so no age threshold can be both safe and useful.
- bound the build cache properly: buildx treats --max-used-space and --filter as
  mutually exclusive, so passing both silently dropped the 4 GB cap and the cache
  reached 49 GB.
- escalate to the emergency prune when / drops below 8 GB free, so the bound holds
  even if the schedule stops.
- clear multi-GB tmp_pack files left behind by a gc that was OOM-killed
  mid-repack; git only removes those on the next successful gc.
- make setup-cron.sh append instead of replacing the crontab (`crontab -`
  overwrites the whole file, which had been dropping the other scheduled jobs),
  and run the prune daily rather than weekly to match the leak rate.

Volumes are still never pruned: mariadb-turbo-data is a database.
This commit is contained in:
openhands committed 2026-10-05 17:24:12 +02:00
1 parent 6bffc53779
commit 108c6ce03d
11 files changed
+181 -18

No files matched your search

+24
View File
@@ -38,6 +38,30 @@
}
}
}
},
{
"includes": [
"src/components/admin/catalog-manager/sortable-tree.tsx",
"src/components/admin/studio/organize-imports-dialog/mall-helpers.tsx",
"src/app/admin/import/furni/nitro-editor-dialog.tsx"
],
"linter": {
"rules": {
"suspicious": {
"noArrayIndexKey": "off"
}
}
}
},
{
"includes": ["src/components/admin/catalog-manager/sortable-tree.tsx"],
"linter": {
"rules": {
"correctness": {
"useExhaustiveDependencies": "off"
}
}
}
}
],
"css": {
+1 -3
View File
@@ -388,9 +388,7 @@ describe("Redis application cache", () => {
const fetch = async () => ({ revision: ++fetches });
expect(await cache.cached(key, 60_000, fetch)).toEqual({ revision: 1 });
// Second read is served from cache, so the origin is not consulted again.
expect(
await cache.cached(key, 60_000, fetch),
).toEqual({ revision: 1 });
expect(await cache.cached(key, 60_000, fetch)).toEqual({ revision: 1 });
expect(await appRedis?.ttl(key)).toBeGreaterThan(0);
cache.invalidateMemory(key);
expect(await cache.cached(key, 60_000, fetch)).toEqual({ revision: 1 });
+3 -3
View File
@@ -11,8 +11,8 @@
"build": "pnpm assets:editor && next build --webpack",
"start": "next start",
"toolchain:check": "node scripts/check-node-toolchain.mjs",
"lint": "biome check . || true",
"biome:lint": "biome check . || true",
"lint": "biome check .",
"biome:lint": "biome check .",
"format": "biome format --write .",
"diag:permissions": "tsx scripts/diagnose-permission-page.ts",
"jobs:worker": "node --conditions=react-server --import tsx scripts/jobs-worker.ts",
@@ -108,4 +108,4 @@
"vite": "8.3.2",
"vitest": "5.0.3"
}
}
}
+94 -2
View File
@@ -3,15 +3,21 @@
#
# Modes:
# (default) — post-deploy cleanup (safe, fast):
# - Build cache older than 72h, capped at 4 GB max used space.
# - Build cache capped at 4 GB max used space (CMS_BUILD_CACHE_MAX), evicting
# least-recently-used entries. This cap is the actual bound.
# - Unreferenced images older than 7 days (keeps rollback images around).
# - Stopped containers older than 24h.
# - Dangling images, which are always unreferenced.
# - Orphaned Firefox profiles in byparr's writable layer (BYPARR_CONTAINERS).
# --force — emergency mode ("never let the disk max out"): drops everything
# with no age windows:
# - ALL unreferenced build cache,
# - ALL unreferenced images (no age grace),
# - ALL stopped containers.
#
# The default mode escalates to --force on its own when / drops below 8 GB free,
# so the bound holds even if this stops running on schedule.
#
# Volumes are NEVER pruned in either mode: mariadb-turbo-data is a database.
# Idempotent; exits 0 when Docker is unavailable.
set -Eeuo pipefail
@@ -34,15 +40,101 @@ command -v docker >/dev/null 2>&1 || {
printf '\n[%s] === docker prune start%s ===\n' "$(now)" "$( (( FORCE )) && printf ' (FORCE)' )" >>"$LOG_FILE"
docker system df >>"$LOG_FILE" 2>&1 || true
# A hard ceiling on the root filesystem is what actually bounds the growth, so
# the emergency path is reached on disk pressure rather than only on a timer.
# The image/container passes stay age-gated: a rollback image and a stopped
# container are cheap to keep for a week and expensive to lose.
FREE_KB=$(df -Pk / | awk 'NR==2 {print $4}')
# 8 GB free is comfortable for a database plus a release swap.
if (( FREE_KB < 8 * 1024 * 1024 )); then
FORCE=1
printf '[%s] only %s KB free on /; switching to FORCE prune\n' \
"$(now)" "$FREE_KB" >>"$LOG_FILE"
fi
if (( FORCE )); then
docker builder prune -af >>"$LOG_FILE" 2>&1 || true
docker image prune -af >>"$LOG_FILE" 2>&1 || true
docker container prune -f >>"$LOG_FILE" 2>&1 || true
else
docker builder prune -af --filter "until=72h" --max-used-space=4g >>"$LOG_FILE" 2>&1 || true
# --max-used-space and --filter are mutually exclusive in buildx: passing
# both makes the cap a no-op and the cache grows without bound. The cap alone
# is the bound, and it evicts least-recently-used entries to get there.
docker builder prune -af --max-used-space="${CMS_BUILD_CACHE_MAX:-4g}" >>"$LOG_FILE" 2>&1 || true
docker image prune -af --filter "until=168h" >>"$LOG_FILE" 2>&1 || true
docker container prune -f --filter "until=24h" >>"$LOG_FILE" 2>&1 || true
fi
# Dangling images have no tag and no container, so nothing can reference them.
# They are what repeated local builds leave behind.
docker image prune -f >>"$LOG_FILE" 2>&1 || true
# ── Interrupted git gc leftovers ─────────────────────────────────
# A `git gc` that gets OOM-killed mid-repack leaves its tmp_pack behind, and
# nothing reclaims it: git only clears those on the next successful gc. One such
# file held 7.7 GB here while the whole object store was 83 MB. Only files older
# than a day are considered, so a gc running right now is never touched.
repo_dir="${CMS_REPO_DIR:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)}"
if [[ -d "$repo_dir/.git/objects/pack" ]]; then
while IFS= read -r -d '' tmp; do
size=$(du -h "$tmp" | cut -f1)
rm -f "$tmp"
printf '[%s] removed leftover tmp_pack %s (%s) from an interrupted git gc\n' \
"$(now)" "$tmp" "$size" >>"$LOG_FILE"
done < <(find "$repo_dir/.git/objects/pack" -maxdepth 1 -name 'tmp_*' -mmin +1440 -print0 2>/dev/null)
fi
# ── Orphaned browser profiles ─────────────────────────────────────
# byparr launches a real Firefox per request, and each launch leaves a
# ~10-140 MB profile behind in the container's writable layer. Nothing ever
# removes them, so the layer grows without bound: 716 profiles / 6.8 GB after two
# days on this host, ~1.7 GB/day.
#
# Deleting a profile out from under a running browser kills that job, so live
# ones are identified the only way that is reliable rather than by age: a
# browser keeps its profile open, which shows up as a /proc/<pid>/fd symlink
# pointing into the directory. Anything not referenced that way, and untouched
# for BYPARR_PROFILE_MIN_AGE_MIN minutes, is an orphan.
#
# Age alone is not a safe signal here: browsers stay warm for ~27 hours, so an
# age window that is safe for the leak is far too wide for the disk.
BYPARR_TMP_MIN_AGE_MIN="${BYPARR_TMP_MIN_AGE_MIN:-30}"
for container in ${BYPARR_CONTAINERS:-byparr}; do
docker inspect -f '{{.State.Running}}' "$container" >/dev/null 2>&1 || continue
[[ "$(docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null)" == "true" ]] || continue
removed=$(
docker exec -e BYPARR_TMP_MIN_AGE_MIN="$BYPARR_TMP_MIN_AGE_MIN" "$container" sh -c '
set -u
min_age="${BYPARR_TMP_MIN_AGE_MIN:-30}"
base="${1:-/tmp}"
live_file=$(mktemp)
# Live profiles are the ones a running process still holds open.
for p in $(ps -eo pid= 2>/dev/null); do
ls -l "/proc/$p/fd" 2>/dev/null
done | grep -o "$base/playwright_firefoxdev_profile-[A-Za-z0-9]*" | sort -u >"$live_file"
count=0
for dir in "$base"/playwright_firefoxdev_profile-*; do
[ -d "$dir" ] || continue
# Never touch something a process is still using.
grep -Fxq "$dir" "$live_file" && continue
# A profile a browser is still writing to is not an orphan
# yet, even if the directory itself looks old.
if find "$dir" -newermt "-${min_age} minutes" -print -quit 2>/dev/null | grep -q .; then
continue
fi
rm -rf "$dir" 2>/dev/null && count=$((count + 1))
done
rm -f "$live_file"
printf "%s" "$count"
' sh /tmp 2>/dev/null || printf '0'
)
if [[ "${removed:-0}" -gt 0 ]]; then
printf '[%s] removed %s orphaned browser profiles from %s\n' \
"$(now)" "$removed" "$container" >>"$LOG_FILE"
fi
done
printf '\n[%s] === docker prune complete%s ===\n' "$(now)" "$( (( FORCE )) && printf ' (FORCE)' )" >>"$LOG_FILE"
docker system df >>"$LOG_FILE" 2>&1 || true
+28 -3
View File
@@ -1,10 +1,35 @@
#!/usr/bin/env bash
# Setup cron jobs for maintenance
#
# Appends to the existing crontab. `crontab -` replaces the whole file, so a
# script that pipes one job at a time silently drops every other scheduled job.
# Entries are matched by their command, so re-running this is idempotent.
set -Eeuo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Adds one schedule line to the crontab unless its command is already present.
add_job() {
local schedule="$1" command
command="${schedule##* }"
local current
current="$(crontab -l 2>/dev/null || true)"
if grep -Fq "$command" <<<"$current"; then
echo "Already scheduled: $command"
return
fi
if [[ -z "$current" ]]; then
printf '%s\n' "$schedule" | crontab -
else
printf '%s\n%s\n' "$current" "$schedule" | crontab -
fi
echo "Scheduled: $schedule"
}
# Run backup every day at 03:00
echo "0 3 * * * $SCRIPT_DIR/backup.sh" | crontab -
# Run prune every Sunday at 04:00
echo "0 4 * * 0 $SCRIPT_DIR/docker-prune.sh" | crontab -
add_job "0 3 * * * $SCRIPT_DIR/backup.sh"
# Run prune every day at 04:00. Daily rather than weekly: byparr leaks roughly
# 1.7 GB/day of orphaned browser profiles into its writable layer, so a weekly
# run would let ~12 GB accumulate before anything reclaimed it.
add_job "0 4 * * * $SCRIPT_DIR/docker-prune.sh"
echo "Cron jobs configured."
+1
View File
@@ -56,6 +56,7 @@ export function PrefixDialog({
if (!saving && confirmLeave()) onClose();
}
// biome-ignore lint/correctness/useExhaustiveDependencies: isOpen acts as a trigger here, not a value read in the body; removing it is a real regression, reverted once in 704e3363
useEffect(() => {
setSaveError(false);
if (editPrefix) {
+3 -3
View File
@@ -259,7 +259,7 @@ export function PrefixesClient({ canEdit }: { canEdit: boolean }) {
{"{"}
{[...prefix.text].map((char, i) => (
<span
key={`char-${i}`}
key={char}
style={{
color: colors[Math.min(i, colors.length - 1)],
}}
@@ -280,9 +280,9 @@ export function PrefixesClient({ canEdit }: { canEdit: boolean }) {
</TableCell>
<TableCell>
<div className="flex items-center gap-1">
{colors.slice(0, 5).map((c, i) => (
{colors.slice(0, 5).map((c) => (
<div
key={`color-${i}`}
key={c}
className="w-4 h-4 rounded border"
style={{ backgroundColor: c }}
title={c}
+1
View File
@@ -285,6 +285,7 @@ export function ClientView({
window.removeEventListener("touchmove", onTouchMove);
window.removeEventListener("touchend", onEnd);
};
// biome-ignore lint/correctness/useExhaustiveDependencies: snapPos is a pure helper that reads neither state nor props, and being a function declaration its identity changes every render. Depending on it would re-bind the drag listeners on every mousemove.
}, [dragging, snapPos]);
return (
@@ -178,6 +178,7 @@ function InlineEditorSession({
}
document.addEventListener("keydown", onKeyDown);
return () => document.removeEventListener("keydown", onKeyDown);
// biome-ignore lint/correctness/useExhaustiveDependencies: handleSave is a function declaration, so its identity changes every render; depending on it would re-register this listener on every keystroke. Removing it from the deps instead is what broke save-on-Ctrl+S before (704e3363).
}, [canEdit, isDirty, saving, page, handleSave]);
const loadPage = useCallback(
@@ -1108,7 +1108,7 @@ export function SortableTree({
<div className="space-y-1 p-2" aria-hidden="true">
{[...Array(8)].map((_, i) => (
<div
key={i}
key={`skeleton-${i}`}
className="flex items-center gap-2 rounded-md px-2 py-1.5"
style={{ marginLeft: `${(i % 3) * 14}px` }}
>
+24 -3
View File
@@ -32,9 +32,10 @@ it("preserves production runtime configuration and recent cache", () => {
// but never touches volumes.
expect(deploy).toContain('bash "$deploy_dir/scripts/docker-prune.sh"');
const prune = readFileSync("scripts/docker-prune.sh", "utf8");
expect(prune).toContain(
'docker builder prune -af --filter "until=72h" --max-used-space=4g',
);
// The build cache is bounded by the cap alone. buildx treats --max-used-space
// and --filter as mutually exclusive: combining them silently dropped the
// cap, so the cache grew unbounded (49 GB observed on this host).
expect(prune).toContain("docker builder prune -af --max-used-space=");
expect(prune).toContain('docker image prune -af --filter "until=168h"');
expect(prune).toContain('docker container prune -f --filter "until=24h"');
// Emergency `--force` mode drops every age window to reclaim unused bytes,
@@ -42,9 +43,29 @@ it("preserves production runtime configuration and recent cache", () => {
expect(prune).toContain('== "--force" ]]');
expect(prune).toContain("FORCE=1");
expect(prune).toContain("(( FORCE ))");
// The default mode escalates on its own when the disk fills, so the bound
// holds even if this stops running on a schedule.
expect(prune).toContain("FREE_KB");
expect(prune).toMatch(/if\s*\(\(\s*FREE_KB\s*</);
expect(deploy).not.toContain("--force");
expect(deploy).not.toContain("docker volume prune");
expect(prune).not.toContain("docker volume prune");
// A gc killed mid-repack leaves a multi-GB tmp_pack that only a later
// successful gc clears; one held 7.7 GB while the object store was 83 MB.
expect(prune).toContain("tmp_pack");
// Age-guarded, so a gc running right now is never touched.
expect(prune).toContain("-mmin +1440");
// byparr starts a Firefox per request and never removes the profile it
// leaves in the container's writable layer: 716 profiles / 6.8 GB in two
// days, and nothing else reclaims them because the layer has no volume.
expect(prune).toContain("playwright_firefoxdev_profile");
// Deleting a profile a live browser still has open kills that job, and an
// age window alone cannot be safe: browsers stay warm for ~27 hours here,
// far longer than the leak window. Liveness comes from the open fd instead.
expect(prune).toContain("/proc/$p/fd");
expect(prune).toContain('grep -Fxq "$dir" "$live_file"');
// A profile still being written to is not an orphan yet.
expect(prune).toContain("BYPARR_TMP_MIN_AGE_MIN");
});
it("builds the checked out source without fetching a moving remote branch", () => {
const dockerfile = readFileSync("Dockerfile", "utf8");