fix(ci): unblock the recovery lane, and report the real CI queue depth #549
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!549
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/ci-recovery-lane-and-queue-visibility"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up to the CI investigation. Two separate problems, both of which made CI look broken while the dashboards said it was fine.
1. The recovery lane went red on every push, and its repair was working
control-plane-recovery.ymlfailed 11 of 15 runs, including the last four. Every one dies identically, and the logs show the repair had already succeeded (full binding set promoted to 100%,FG_ENCRYPTION_KEYincluded):The 404 is a
curl -fsSagainst the Worker scriptforgegraph-router.fg traffic disableremoved that script on 2026-08-30 at 20:35; the router diagnostics were added the next day in #540. Underset -euo pipefailthe 404 ends the job.The consequence is worse than the red: the 404 fires on the first diagnostic curl, so
direct_worker_node_http,direct_worker_hub_httpand the apex retry loop after it have never executed. Those are the only checks that prove the repair worked. In a real outage this lane repairs the Worker and then dies before verifying anything.Fixed by probing the router without
-fand branching on the status code. An absent router is now reported (router_present=no) rather than raised, and the verification tail runs.2. The queue could not show a backlog
buildCiOverviewcomputesqueue.depthby filtering the newest 50 rows the caller fetched (limit: 50at every call site). So the reported depth cannot exceed 50, and builds queued long enough to fall out of that window vanish from it.oldestQueuedAgeMshas the same flaw: it is the oldest in the window, not the oldest. A queue that is not draining renders as a small, calm number.buildQueue.runnerCapacitymeanwhile does a realcount(*), so the two already disagreed.New
summarizeCiQueuetakes counts fromcount(*)and timestamps frommin()over the whole table, with the staleness cutoffs evaluated against the database clock rather than the isolate's, and separates cases that need different responses:no-runners— waiting work and zero active runners. Reported ahead of depth, because it explains it; a queue of 40 with no runner is stopped, not slow.stalled— queued past 15m, or rows stuck inrunningpast 60m that no runner will finish. The stuck-running case is invisible in every "is CI busy?" view precisely because CI is not busy.backlogged— deep but young. A capacity story, kept distinct from stuck rows.flowing/idle.Surfaced four ways:
buildQueue.queueStats,GET /api/fg/ci/queue(fleet-wide, not app-scoped, since a backlog is shared-capacity), a panel on/cithat renders only when there is something to see, andfg ci queuewhich exits non-zero when stalled so it can gate a script. The queue tile on/cinow reads the real aggregate.What the queue actually looks like
Measured across all 77 apps while writing this: 1 build queued (2 minutes old), 0 running. Forgejo's runner queue was empty too, with jobs finishing in 0.1 to 12.8 minutes. The build queue is not the backlog.
The backlog is 94 open pull requests: npm-registry 25, ForgeGraph 15, StreamConductor 11, playtrek-engine 8, creator 7. Of those, 35 have failing CI and 93 have no review. Median age is 5 days, p90 is 52 days, and 12 are older than 30 days. There is no fleet-wide PR view in the web UI today (only per-repo
/repos/[id]/pulls), which is a separate gap from this PR.Verification
packages/apiandapps/webboth typecheck clean. The pre-existingstripPrefixandforgejo-apierrors turned out to be stale injected workspace copies; refreshingnode_modules/@forgegraph/{db,api}cleared all of them.summarizeCiQueuecover each verdict, the future-timestamp clamp, and an unparseable timestamp.go build,go vet, and the newformatQueueAgetest green. The Go and TypeScript duration formatters are pinned to each other by test so the CLI and the page describe the same wait identically.oxlintclean on every changed file.curl -fsScalls in the tail either target scripts that exist or sit insideifconditions whereset -edoes not apply.Not changed
The recovery lane still runs on
push: branches: [main]and still force-deploys the control-plane Worker on every merge, replacing its secrets with four and rebuilding the rest by version inheritance. That is a real risk worth a decision, but it changes your recovery posture rather than fixing a bug, so it is flagged and left alone here.No index was added for the new aggregate. It scans
buildsfiltered to in-flight rows, which is the same shape as the existingrunnerCapacitycount, and adding DDL would turn production deploys red until it is applied by hand. Worth revisiting if the table grows.🤖 Generated with Claude Code
915c18831fc799a3504bc799a3504baa47a0de14aa47a0de14d1ff4689a5d1ff4689a59f544f0a96