fix(agent): root deployed nix store paths so the pruner cannot collect a live service #565

Merged
gmackie merged 3 commits from fix/nix-gcroot-live-services into main 2026-09-14 05:34:06 +00:00
Owner

The incident

The agent builds with nix build --no-link, so a deployed store path has no GC root: nothing but the systemd unit's ExecStart references it, and nix does not read unit files.

On 2026-09-09 the hourly forgegraph-store-prune timer ran nix-collect-garbage -d thirteen minutes after an npm-registry deploy on hetzner-fg and deleted the path the live process was running from. The service stayed active and kept serving reads out of already-loaded files, then failed on the first lazily required module: every npm publish returned 400 Cannot find module '../encodings' while anonymous GETs looked perfectly healthy. Restarting would have turned it into a hard outage, because the binary was gone.

This was not specific to the registry. hetzner-master was found with five units pointing at store paths that no longer exist (splat, turntable-bot, habit, habit-prod, veritas — all currently inactive, so they would have failed on next start).

The fix

Two independent guards, because a node can be running an older agent:

1. agent/internal/deploy — registerGCRoot pins the path as forgegraph-<service>-live before the slow parts of activation (migration, health probe), and rotates the prior path to -previous so the rollback target survives the next sweep too. A failed restart or health probe re-points -live at whatever it rolled back to. Symlinks swap atomically, so a prune running mid-deploy never observes a missing root.

2. ops/store-prune — pin_live_services re-derives the same root from every forgegraph-*.service unit before collecting, and warns rather than rooting when a unit already points at a path that is gone. This covers hosts still on an older agent.

Tests

  • gcroot_test.go: first deploy, previous rotation, idempotence (re-deploying the same path creates no -previous), non-store-path rejection, rollback restore, no leftover .tmp links.
  • ops/store-prune/test-forgegraph-store-prune.sh gains a live-service pinning case; the unit dir, gcroots dir and store prefix are now overridable so it runs hermetically instead of against the real /nix.

go test ./internal/deploy/, go vet, gofmt and the store-prune harness all pass.

Already applied by hand

Roots were created manually on hetzner-fg and hetzner-master for the currently live services, the registry was redeployed onto a rebuilt path, and the prune timer is running again. This PR makes it durable.

🤖 Generated with Claude Code

## The incident The agent builds with `nix build --no-link`, so a deployed store path has **no GC root**: nothing but the systemd unit's `ExecStart` references it, and nix does not read unit files. On 2026-09-09 the hourly `forgegraph-store-prune` timer ran `nix-collect-garbage -d` thirteen minutes after an `npm-registry` deploy on hetzner-fg and deleted the path the live process was running from. The service stayed `active` and kept serving reads out of already-loaded files, then failed on the first lazily required module: every `npm publish` returned `400 Cannot find module '../encodings'` while anonymous GETs looked perfectly healthy. Restarting would have turned it into a hard outage, because the binary was gone. This was not specific to the registry. `hetzner-master` was found with **five** units pointing at store paths that no longer exist (`splat`, `turntable-bot`, `habit`, `habit-prod`, `veritas` — all currently inactive, so they would have failed on next start). ## The fix Two independent guards, because a node can be running an older agent: **1. `agent/internal/deploy`** — `registerGCRoot` pins the path as `forgegraph-<service>-live` *before* the slow parts of activation (migration, health probe), and rotates the prior path to `-previous` so the rollback target survives the next sweep too. A failed restart or health probe re-points `-live` at whatever it rolled back to. Symlinks swap atomically, so a prune running mid-deploy never observes a missing root. **2. `ops/store-prune`** — `pin_live_services` re-derives the same root from every `forgegraph-*.service` unit before collecting, and *warns* rather than rooting when a unit already points at a path that is gone. This covers hosts still on an older agent. ## Tests - `gcroot_test.go`: first deploy, previous rotation, idempotence (re-deploying the same path creates no `-previous`), non-store-path rejection, rollback restore, no leftover `.tmp` links. - `ops/store-prune/test-forgegraph-store-prune.sh` gains a live-service pinning case; the unit dir, gcroots dir and store prefix are now overridable so it runs hermetically instead of against the real `/nix`. `go test ./internal/deploy/`, `go vet`, `gofmt` and the store-prune harness all pass. ## Already applied by hand Roots were created manually on hetzner-fg and hetzner-master for the currently live services, the registry was redeployed onto a rebuilt path, and the prune timer is running again. This PR makes it durable. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(agent): root deployed nix store paths so the pruner cannot collect a live service
All checks were successful
CI / gitleaks (pull_request) Successful in 9s
CI / storybook (pull_request) Successful in 1m15s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 10m40s
0b7e2dbd8e
The agent builds with `nix build --no-link`, so a deployed store path had no
GC root: nothing but the systemd unit's ExecStart referenced it, and nix does
not read unit files. On 2026-09-09 the hourly forgegraph-store-prune timer ran
`nix-collect-garbage -d` thirteen minutes after an npm-registry deploy on
hetzner-fg and deleted the path the live process was running from. The service
stayed "active" and kept serving reads out of already-loaded files, then failed
on the first lazily required module — every publish returned
400 "Cannot find module '../encodings'" while GETs looked fine. A restart would
have been an outage: the binary was gone. hetzner-master was found in the same
state with five units pointing at paths that no longer exist.

Two independent guards, because a node can be running an older agent:

- deploy: registerGCRoot pins the path as forgegraph-<service>-live before the
  slow parts of activation (migration, health probe) run, and keeps the prior
  path as -previous so the rollback target survives the next sweep too. A
  failed restart or health probe re-points -live at what it rolled back to.
  Symlinks are swapped atomically so a prune running mid-deploy never sees a
  missing root.
- store-prune: pin_live_services re-derives the same root from every
  forgegraph-*.service unit before collecting, and warns (rather than rooting)
  when a unit already points at a path that is gone.

Tests: gcroot_test.go covers first deploy, previous rotation, idempotence,
non-store-path rejection and rollback restore; the store-prune harness gains a
live-service pinning case (unit prefixes are now overridable so it runs
hermetically off /nix).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Author
Owner

Preview environment is live: https://pr-565-forgegraph.forgegraf.com

Deployed 0b7e2dbd with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-565-forgegraph.forgegraf.com Deployed `0b7e2dbd` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
fix(ops): let store-prune sweep when the disk is already critical
All checks were successful
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 2m0s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m58s
0edcfd6282
The active-build guard had no floor. On 2026-09-09 hetzner-bob sat at 99%
(2.7G free) carrying 59G of pnpm store across two users while CI ran
near-continuously, so every hourly sweep logged "active build detected —
deferring" and reclaimed nothing. Meanwhile the api database suite failed on
every run with

  emulate schema load failed:
    CREATE TYPE "repo_provider" AS ENUM(...)
    -> the database system is in recovery mode

because the Postgres those tests spawn could not write. Three PRs in a stack
went red on it and it reads exactly like a flake, which is how it survived.

It is a feedback loop: the failures trigger retries, the retries keep a build
active, and an always-active build keeps deferring the sweep. The node cannot
climb out on its own.

Below CRITICAL_FREE_GB (default 5) the guard no longer applies. The builds it
protects are already failing at that point, so risking one corrupted install
beats a node that stays wedged. Above the floor the guard is unchanged.

The harness now covers both sides: 10G free with an active build still defers,
0G free with an active build sweeps and says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Author
Owner

Preview environment is live: https://pr-565-forgegraph.forgegraf.com

Deployed 0edcfd62 with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-565-forgegraph.forgegraf.com Deployed `0edcfd62` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
Author
Owner

Pushed 5cff84a5f0d8: fix a pinning bug before this rolls out, plus what rollout actually takes

The registry was garbage-collected a second time (down 2026-09-13 06:51 → 2026-09-14 03:40, restored by break-glass redeploy to df0c19a). That made this PR urgent, so I checked it against real unit files before recommending it — and found a bug.

The bug

pin_live_services made one <unit>-live link per unit and rewrote it for each store path with ln -sfn. A unit that references two paths kept only the last one protected — and the order came from sort -u, i.e. hash order, so which one survived was a coin flip. The common shape is exactly that:

ExecStart=/nix/store/…-nodejs-22/bin/node /nix/store/…-app/server.mjs

A collected runtime kills the service on restart as surely as a collected app. hetzner-master's turntable-bot unit has this shape (its paths are already gone; the unit is disabled).

Fix: every distinct path gets a root. The first path in file order keeps <unit>-live, the name the agent uses too. The rest become <unit>-live-extra-<hash>, and old extras are removed first so replaced builds stay collectable. New test covers two paths and the stale-extra cleanup; all 5 prune tests pass.

Checked against real units: I ran the new function under POSIX sh against the actual unit files on hetzner-fg and hetzner-master, writing into a throwaway gcroots dir. It pins the same roots as before — 1 on hetzner-fg (the registry), 8 on hetzner-master. No live unit there has two paths today.

Merging this does not protect anything by itself

part how it reaches nodes status
ops/store-prune/forgegraph-store-prune.sh installed by hand (fg node cp, per its README) hetzner-fg, hetzner-worker, hetzner-bob and labnuc run main's pre-#565 script. hetzner-master runs an older copy that lacks the active-build guard and only finds nix-collect-garbage via systemd's minimal PATH. vanuc has none.
agent/internal/deploy/gcroot.go agent release hetzner-fg is pinned (FG_DISABLE_SELF_UPDATE=1), and that pin is a deliberate, standing decision. A release will not reach the registry's node without someone choosing to lift it.

The prune-script half doesn't need the agent: it derives live paths from the unit files on disk just before collection. Installing the script alone on hetzner-fg protects the registry across redeploys, without touching the pin. That's the recommended first rollout step.

Existing breakage (not caused here)

The dry runs found 9 units pointing at already-collected builds, all inactive and disabled leftovers, none serving:

  • hetzner-fg: habit, deploy-test-001, 7a71b456…
  • hetzner-master: fg-daily-dose, fg-splat, fg-turntable-bot, habit-prod, habit, veritas, plus the two-path turntable unit

🤖 Generated with Claude Code

## Pushed `5cff84a5f0d8`: fix a pinning bug before this rolls out, plus what rollout actually takes The registry was garbage-collected a **second** time (down 2026-09-13 06:51 → 2026-09-14 03:40, restored by break-glass redeploy to `df0c19a`). That made this PR urgent, so I checked it against real unit files before recommending it — and found a bug. ### The bug `pin_live_services` made one `<unit>-live` link per unit and rewrote it for each store path with `ln -sfn`. A unit that references two paths kept only the last one protected — and the order came from `sort -u`, i.e. hash order, so which one survived was a coin flip. The common shape is exactly that: ``` ExecStart=/nix/store/…-nodejs-22/bin/node /nix/store/…-app/server.mjs ``` A collected runtime kills the service on restart as surely as a collected app. hetzner-master's turntable-bot unit has this shape (its paths are already gone; the unit is disabled). **Fix:** every distinct path gets a root. The first path in file order keeps `<unit>-live`, the name the agent uses too. The rest become `<unit>-live-extra-<hash>`, and old extras are removed first so replaced builds stay collectable. New test covers two paths and the stale-extra cleanup; all 5 prune tests pass. **Checked against real units:** I ran the new function under POSIX `sh` against the actual unit files on hetzner-fg and hetzner-master, writing into a throwaway gcroots dir. It pins the same roots as before — 1 on hetzner-fg (the registry), 8 on hetzner-master. No live unit there has two paths today. ### Merging this does not protect anything by itself | part | how it reaches nodes | status | |---|---|---| | `ops/store-prune/forgegraph-store-prune.sh` | **installed by hand** (`fg node cp`, per its README) | hetzner-fg, hetzner-worker, hetzner-bob and labnuc run main's pre-#565 script. hetzner-master runs an **older** copy that lacks the active-build guard and only finds `nix-collect-garbage` via systemd's minimal PATH. vanuc has none. | | `agent/internal/deploy/gcroot.go` | agent release | **hetzner-fg is pinned (`FG_DISABLE_SELF_UPDATE=1`)**, and that pin is a deliberate, standing decision. A release will not reach the registry's node without someone choosing to lift it. | The prune-script half doesn't need the agent: it derives live paths from the unit files on disk just before collection. **Installing the script alone on hetzner-fg protects the registry across redeploys**, without touching the pin. That's the recommended first rollout step. ### Existing breakage (not caused here) The dry runs found 9 units pointing at already-collected builds, all **inactive and disabled** leftovers, none serving: - **hetzner-fg:** `habit`, `deploy-test-001`, `7a71b456…` - **hetzner-master:** `fg-daily-dose`, `fg-splat`, `fg-turntable-bot`, `habit-prod`, `habit`, `veritas`, plus the two-path turntable unit 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(store-prune): root every store path a unit references, not just one
All checks were successful
CI / gitleaks (pull_request) Successful in 8s
CI / storybook (pull_request) Successful in 3m51s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 12m18s
5cff84a5f0
pin_live_services made a single <unit>-live link per unit and wrote it
once per store path with `ln -sfn`, so a unit referencing two paths kept
only the last one protected. The common shape is a runtime plus the app
it runs:

  ExecStart=/nix/store/…-nodejs-22/bin/node /nix/store/…-app/server.mjs

The paths come out of `sort -u`, i.e. hash order, so which one survived
was a coin flip, and a collected runtime kills the service on its next
restart just as surely as a collected app. It is not hypothetical:
hetzner-master's turntable-bot unit has exactly that shape (both paths
are already gone, and the unit is disabled).

Every distinct path now gets a root. The first in file order keeps
<unit>-live, the name the agent's gcroot uses too; the rest become
<unit>-live-extra-<hash>. Extras from a previous deploy are removed
first, so replaced builds stay collectable instead of accumulating.

Verified by running the new function against the real unit files on
hetzner-fg and hetzner-master into a throwaway gcroots dir, under POSIX
sh: it pins exactly what the old version did (1 and 8 roots), since no
live unit there has two paths today. The new test covers the two-path
case and the stale-extra cleanup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Author
Owner

Preview environment is live: https://pr-565-forgegraph.forgegraf.com

Deployed 5cff84a5 with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-565-forgegraph.forgegraf.com Deployed `5cff84a5` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!565
No description provided.