fix(agent): root deployed nix store paths so the pruner cannot collect a live service #565
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!565
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/nix-gcroot-live-services"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The incident
The agent builds with
nix build --no-link, so a deployed store path has no GC root: nothing but the systemd unit'sExecStartreferences it, and nix does not read unit files.On 2026-09-09 the hourly
forgegraph-store-prunetimer rannix-collect-garbage -dthirteen minutes after annpm-registrydeploy on hetzner-fg and deleted the path the live process was running from. The service stayedactiveand kept serving reads out of already-loaded files, then failed on the first lazily required module: everynpm publishreturned400 Cannot find module '../encodings'while anonymous GETs looked perfectly healthy. Restarting would have turned it into a hard outage, because the binary was gone.This was not specific to the registry.
hetzner-masterwas found with five units pointing at store paths that no longer exist (splat,turntable-bot,habit,habit-prod,veritas— all currently inactive, so they would have failed on next start).The fix
Two independent guards, because a node can be running an older agent:
1.
agent/internal/deploy—registerGCRootpins the path asforgegraph-<service>-livebefore the slow parts of activation (migration, health probe), and rotates the prior path to-previousso the rollback target survives the next sweep too. A failed restart or health probe re-points-liveat whatever it rolled back to. Symlinks swap atomically, so a prune running mid-deploy never observes a missing root.2.
ops/store-prune—pin_live_servicesre-derives the same root from everyforgegraph-*.serviceunit before collecting, and warns rather than rooting when a unit already points at a path that is gone. This covers hosts still on an older agent.Tests
gcroot_test.go: first deploy, previous rotation, idempotence (re-deploying the same path creates no-previous), non-store-path rejection, rollback restore, no leftover.tmplinks.ops/store-prune/test-forgegraph-store-prune.shgains a live-service pinning case; the unit dir, gcroots dir and store prefix are now overridable so it runs hermetically instead of against the real/nix.go test ./internal/deploy/,go vet,gofmtand the store-prune harness all pass.Already applied by hand
Roots were created manually on hetzner-fg and hetzner-master for the currently live services, the registry was redeployed onto a rebuilt path, and the prune timer is running again. This PR makes it durable.
🤖 Generated with Claude Code
Preview environment is live: https://pr-565-forgegraph.forgegraf.com
Deployed
0b7e2dbdwith the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.The active-build guard had no floor. On 2026-09-09 hetzner-bob sat at 99% (2.7G free) carrying 59G of pnpm store across two users while CI ran near-continuously, so every hourly sweep logged "active build detected — deferring" and reclaimed nothing. Meanwhile the api database suite failed on every run with emulate schema load failed: CREATE TYPE "repo_provider" AS ENUM(...) -> the database system is in recovery mode because the Postgres those tests spawn could not write. Three PRs in a stack went red on it and it reads exactly like a flake, which is how it survived. It is a feedback loop: the failures trigger retries, the retries keep a build active, and an always-active build keeps deferring the sweep. The node cannot climb out on its own. Below CRITICAL_FREE_GB (default 5) the guard no longer applies. The builds it protects are already failing at that point, so risking one corrupted install beats a node that stays wedged. Above the floor the guard is unchanged. The harness now covers both sides: 10G free with an active build still defers, 0G free with an active build sweeps and says so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Preview environment is live: https://pr-565-forgegraph.forgegraf.com
Deployed
0edcfd62with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.Pushed
5cff84a5f0d8: fix a pinning bug before this rolls out, plus what rollout actually takesThe registry was garbage-collected a second time (down 2026-09-13 06:51 → 2026-09-14 03:40, restored by break-glass redeploy to
df0c19a). That made this PR urgent, so I checked it against real unit files before recommending it — and found a bug.The bug
pin_live_servicesmade one<unit>-livelink per unit and rewrote it for each store path withln -sfn. A unit that references two paths kept only the last one protected — and the order came fromsort -u, i.e. hash order, so which one survived was a coin flip. The common shape is exactly that:A collected runtime kills the service on restart as surely as a collected app. hetzner-master's turntable-bot unit has this shape (its paths are already gone; the unit is disabled).
Fix: every distinct path gets a root. The first path in file order keeps
<unit>-live, the name the agent uses too. The rest become<unit>-live-extra-<hash>, and old extras are removed first so replaced builds stay collectable. New test covers two paths and the stale-extra cleanup; all 5 prune tests pass.Checked against real units: I ran the new function under POSIX
shagainst the actual unit files on hetzner-fg and hetzner-master, writing into a throwaway gcroots dir. It pins the same roots as before — 1 on hetzner-fg (the registry), 8 on hetzner-master. No live unit there has two paths today.Merging this does not protect anything by itself
ops/store-prune/forgegraph-store-prune.shfg node cp, per its README)nix-collect-garbagevia systemd's minimal PATH. vanuc has none.agent/internal/deploy/gcroot.goFG_DISABLE_SELF_UPDATE=1), and that pin is a deliberate, standing decision. A release will not reach the registry's node without someone choosing to lift it.The prune-script half doesn't need the agent: it derives live paths from the unit files on disk just before collection. Installing the script alone on hetzner-fg protects the registry across redeploys, without touching the pin. That's the recommended first rollout step.
Existing breakage (not caused here)
The dry runs found 9 units pointing at already-collected builds, all inactive and disabled leftovers, none serving:
habit,deploy-test-001,7a71b456…fg-daily-dose,fg-splat,fg-turntable-bot,habit-prod,habit,veritas, plus the two-path turntable unit🤖 Generated with Claude Code
Preview environment is live: https://pr-565-forgegraph.forgegraf.com
Deployed
5cff84a5with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.