fix(ci): unblock the recovery lane, and report the real CI queue depth #549

Merged
gmackie merged 1 commit from fix/ci-recovery-lane-and-queue-visibility into main 2026-09-04 06:51:06 +00:00
Owner

Follow-up to the CI investigation. Two separate problems, both of which made CI look broken while the dashboards said it was fine.

1. The recovery lane went red on every push, and its repair was working

control-plane-recovery.yml failed 11 of 15 runs, including the last four. Every one dies identically, and the logs show the repair had already succeeded (full binding set promoted to 100%, FG_ENCRYPTION_KEY included):

direct_worker_url=https://forgegraf.gmac.workers.dev
curl: (22) The requested URL returned error: 404
Job 'restore' failed

The 404 is a curl -fsS against the Worker script forgegraph-router. fg traffic disable removed that script on 2026-08-30 at 20:35; the router diagnostics were added the next day in #540. Under set -euo pipefail the 404 ends the job.

The consequence is worse than the red: the 404 fires on the first diagnostic curl, so direct_worker_node_http, direct_worker_hub_http and the apex retry loop after it have never executed. Those are the only checks that prove the repair worked. In a real outage this lane repairs the Worker and then dies before verifying anything.

Fixed by probing the router without -f and branching on the status code. An absent router is now reported (router_present=no) rather than raised, and the verification tail runs.

2. The queue could not show a backlog

buildCiOverview computes queue.depth by filtering the newest 50 rows the caller fetched (limit: 50 at every call site). So the reported depth cannot exceed 50, and builds queued long enough to fall out of that window vanish from it. oldestQueuedAgeMs has the same flaw: it is the oldest in the window, not the oldest. A queue that is not draining renders as a small, calm number. buildQueue.runnerCapacity meanwhile does a real count(*), so the two already disagreed.

New summarizeCiQueue takes counts from count(*) and timestamps from min() over the whole table, with the staleness cutoffs evaluated against the database clock rather than the isolate's, and separates cases that need different responses:

  • no-runners — waiting work and zero active runners. Reported ahead of depth, because it explains it; a queue of 40 with no runner is stopped, not slow.
  • stalled — queued past 15m, or rows stuck in running past 60m that no runner will finish. The stuck-running case is invisible in every "is CI busy?" view precisely because CI is not busy.
  • backlogged — deep but young. A capacity story, kept distinct from stuck rows.
  • flowing / idle.

Surfaced four ways: buildQueue.queueStats, GET /api/fg/ci/queue (fleet-wide, not app-scoped, since a backlog is shared-capacity), a panel on /ci that renders only when there is something to see, and fg ci queue which exits non-zero when stalled so it can gate a script. The queue tile on /ci now reads the real aggregate.

What the queue actually looks like

Measured across all 77 apps while writing this: 1 build queued (2 minutes old), 0 running. Forgejo's runner queue was empty too, with jobs finishing in 0.1 to 12.8 minutes. The build queue is not the backlog.

The backlog is 94 open pull requests: npm-registry 25, ForgeGraph 15, StreamConductor 11, playtrek-engine 8, creator 7. Of those, 35 have failing CI and 93 have no review. Median age is 5 days, p90 is 52 days, and 12 are older than 30 days. There is no fleet-wide PR view in the web UI today (only per-repo /repos/[id]/pulls), which is a separate gap from this PR.

Verification

  • packages/api and apps/web both typecheck clean. The pre-existing stripPrefix and forgejo-api errors turned out to be stale injected workspace copies; refreshing node_modules/@forgegraph/{db,api} cleared all of them.
  • 12 new tests for summarizeCiQueue cover each verdict, the future-timestamp clamp, and an unparseable timestamp.
  • go build, go vet, and the new formatQueueAge test green. The Go and TypeScript duration formatters are pinned to each other by test so the CLI and the page describe the same wait identically.
  • oxlint clean on every changed file.
  • Workflow YAML validated. The remaining curl -fsS calls in the tail either target scripts that exist or sit inside if conditions where set -e does not apply.

Not changed

The recovery lane still runs on push: branches: [main] and still force-deploys the control-plane Worker on every merge, replacing its secrets with four and rebuilding the rest by version inheritance. That is a real risk worth a decision, but it changes your recovery posture rather than fixing a bug, so it is flagged and left alone here.

No index was added for the new aggregate. It scans builds filtered to in-flight rows, which is the same shape as the existing runnerCapacity count, and adding DDL would turn production deploys red until it is applied by hand. Worth revisiting if the table grows.

🤖 Generated with Claude Code

Follow-up to the CI investigation. Two separate problems, both of which made CI look broken while the dashboards said it was fine. ## 1. The recovery lane went red on every push, and its repair was working `control-plane-recovery.yml` failed 11 of 15 runs, including the last four. Every one dies identically, and the logs show the repair had already succeeded (full binding set promoted to 100%, `FG_ENCRYPTION_KEY` included): ``` direct_worker_url=https://forgegraf.gmac.workers.dev curl: (22) The requested URL returned error: 404 Job 'restore' failed ``` The 404 is a `curl -fsS` against the Worker script `forgegraph-router`. `fg traffic disable` removed that script on 2026-08-30 at 20:35; the router diagnostics were added the next day in #540. Under `set -euo pipefail` the 404 ends the job. The consequence is worse than the red: the 404 fires on the **first** diagnostic curl, so `direct_worker_node_http`, `direct_worker_hub_http` and the apex retry loop after it have never executed. Those are the only checks that prove the repair worked. In a real outage this lane repairs the Worker and then dies before verifying anything. Fixed by probing the router without `-f` and branching on the status code. An absent router is now reported (`router_present=no`) rather than raised, and the verification tail runs. ## 2. The queue could not show a backlog `buildCiOverview` computes `queue.depth` by filtering the newest 50 rows the caller fetched (`limit: 50` at every call site). So the reported depth cannot exceed 50, and builds queued long enough to fall out of that window vanish from it. `oldestQueuedAgeMs` has the same flaw: it is the oldest *in the window*, not the oldest. A queue that is not draining renders as a small, calm number. `buildQueue.runnerCapacity` meanwhile does a real `count(*)`, so the two already disagreed. New `summarizeCiQueue` takes counts from `count(*)` and timestamps from `min()` over the whole table, with the staleness cutoffs evaluated against the **database** clock rather than the isolate's, and separates cases that need different responses: - `no-runners` — waiting work and zero active runners. Reported ahead of depth, because it explains it; a queue of 40 with no runner is stopped, not slow. - `stalled` — queued past 15m, or rows stuck in `running` past 60m that no runner will finish. The stuck-running case is invisible in every "is CI busy?" view precisely because CI is not busy. - `backlogged` — deep but young. A capacity story, kept distinct from stuck rows. - `flowing` / `idle`. Surfaced four ways: `buildQueue.queueStats`, `GET /api/fg/ci/queue` (fleet-wide, not app-scoped, since a backlog is shared-capacity), a panel on `/ci` that renders only when there is something to see, and `fg ci queue` which exits non-zero when stalled so it can gate a script. The queue tile on `/ci` now reads the real aggregate. ## What the queue actually looks like Measured across all 77 apps while writing this: **1 build queued** (2 minutes old), **0 running**. Forgejo's runner queue was empty too, with jobs finishing in 0.1 to 12.8 minutes. The build queue is not the backlog. The backlog is **94 open pull requests**: npm-registry 25, ForgeGraph 15, StreamConductor 11, playtrek-engine 8, creator 7. Of those, 35 have failing CI and 93 have no review. Median age is 5 days, p90 is 52 days, and 12 are older than 30 days. There is no fleet-wide PR view in the web UI today (only per-repo `/repos/[id]/pulls`), which is a separate gap from this PR. ## Verification - `packages/api` and `apps/web` both typecheck clean. The pre-existing `stripPrefix` and `forgejo-api` errors turned out to be stale injected workspace copies; refreshing `node_modules/@forgegraph/{db,api}` cleared all of them. - 12 new tests for `summarizeCiQueue` cover each verdict, the future-timestamp clamp, and an unparseable timestamp. - `go build`, `go vet`, and the new `formatQueueAge` test green. The Go and TypeScript duration formatters are pinned to each other by test so the CLI and the page describe the same wait identically. - `oxlint` clean on every changed file. - Workflow YAML validated. The remaining `curl -fsS` calls in the tail either target scripts that exist or sit inside `if` conditions where `set -e` does not apply. ## Not changed The recovery lane still runs on `push: branches: [main]` and still force-deploys the control-plane Worker on every merge, replacing its secrets with four and rebuilding the rest by version inheritance. That is a real risk worth a decision, but it changes your recovery posture rather than fixing a bug, so it is flagged and left alone here. No index was added for the new aggregate. It scans `builds` filtered to in-flight rows, which is the same shape as the existing `runnerCapacity` count, and adding DDL would turn production deploys red until it is applied by hand. Worth revisiting if the table grows. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(ci): unblock the recovery lane, and report the real CI queue depth
Some checks failed
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 2m19s
CI / ci (pull_request) Has been cancelled
915c18831f
Two things kept CI looking broken while the numbers said otherwise.

1. control-plane-recovery.yml failed 11 of 15 runs, including the last
   four, while its repair succeeded every time. `fg traffic disable`
   deleted the forgegraph-router Worker on 2026-08-30 20:35; the router
   diagnostics added the next day (#540) curl -fsS that script, so the
   404 ended the job under `set -euo pipefail`. Everything after it --
   the node probe, the hub-token probe and the apex retry loop, i.e. the
   only checks that prove the repair worked -- has never run since.
   Probe the router without -f and branch on the status code: an absent
   router is a fact to report, not a failure to raise.

2. The queue could not show a backlog. `buildCiOverview` derives
   queue.depth by filtering the newest 50 rows the caller fetched, so it
   cannot report more than 50 and anything queued long enough to fall
   out of that window disappears -- a queue that is not draining renders
   as a small, calm number.

   New `summarizeCiQueue` takes counts from `count(*)` and timestamps
   from `min()` over the whole builds table, and separates the cases
   that need different responses: no-runners (a stopped queue, reported
   ahead of depth because it explains it), stalled (aged out, or rows
   stuck in `running` that no runner will finish -- invisible in every
   "is CI busy?" view precisely because CI is not busy), backlogged
   (deep but young) and flowing.

   Surfaced as buildQueue.queueStats, GET /api/fg/ci/queue, a panel on
   /ci that appears only when there is something to see, and
   `fg ci queue` (non-zero exit when stalled, so it can gate a script).

The one-box/apex risk in this workflow is unchanged and deliberate: it
still runs on every push to main and still force-deploys the
control-plane Worker. That is a separate call, flagged but not taken
here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gmackie force-pushed fix/ci-recovery-lane-and-queue-visibility from 915c18831f
Some checks failed
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 2m19s
CI / ci (pull_request) Has been cancelled
to c799a3504b
All checks were successful
CI / gitleaks (pull_request) Successful in 9s
CI / storybook (pull_request) Successful in 2m15s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 10m37s
2026-09-04 04:14:33 +00:00
Compare
gmackie force-pushed fix/ci-recovery-lane-and-queue-visibility from c799a3504b
All checks were successful
CI / gitleaks (pull_request) Successful in 9s
CI / storybook (pull_request) Successful in 2m15s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 10m37s
to aa47a0de14
All checks were successful
CI / gitleaks (pull_request) Successful in 9s
CI / storybook (pull_request) Successful in 1m51s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 11m3s
2026-09-04 04:41:06 +00:00
Compare
gmackie force-pushed fix/ci-recovery-lane-and-queue-visibility from aa47a0de14
All checks were successful
CI / gitleaks (pull_request) Successful in 9s
CI / storybook (pull_request) Successful in 1m51s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 11m3s
to d1ff4689a5
Some checks failed
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 1m35s
CI / ci (pull_request) Has been cancelled
2026-09-04 06:31:03 +00:00
Compare
gmackie force-pushed fix/ci-recovery-lane-and-queue-visibility from d1ff4689a5
Some checks failed
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 1m35s
CI / ci (pull_request) Has been cancelled
to 9f544f0a96
All checks were successful
CI / gitleaks (pull_request) Successful in 8s
CI / storybook (pull_request) Successful in 1m48s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m57s
2026-09-04 06:38:04 +00:00
Compare
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!549
No description provided.