perf(alerts): look changesets up by id instead of hashing the whole table #506

Merged
gmackie merged 1 commit from perf/settled-alerts-index-lookup into main 2026-08-27 22:29:38 +00:00
Owner

Follow-up to #502, before it ships. #502 is merged but has never deployed — it is one of the commits stuck behind the registry-401 deploy blackout — so this corrects it while it is still unreleased rather than after.

The problem. The resolver joined on a computed key:

alerts.dedupe_key = 'ci_failed:' || changesets.id

No index can serve that. EXPLAIN (ANALYZE) on production:

Hash Join  (actual time=1.664..1.665 rows=0)
  ->  Index Scan using alert_status_idx on alerts  (rows=11)
  ->  Seq Scan on changesets  (rows=3063)        <- every poll
Execution Time: 1.701 ms

1.7ms is not slow. But it runs on every agent poll — roughly 840 times an hour across seven nodes — and the scan grows with the changesets table rather than with the handful of firing alerts it is actually about. That is the same shape as the unindexed poll that starved this database this morning (bob re-reading session_event), so it is not a pattern to leave sitting in the agent hot path.

The fix. The dedupe key already contains the changeset id, so the join was recovering something we already had. Read the firing alerts by index, parse the ids out of the keys, look them up by primary key:

Index Scan using alert_status_idx on alerts       0.066 ms
Index Scan using changesets_pkey  (11 loops)      0.299 ms

0.37ms, and no sequential scan — cost now scales with firing alerts, not with total changesets.

Behaviour is unchanged: still resolve-only, still bounded, still leaves a genuinely open changeset alerting. 5 tests updated for the two-query shape, all green.

Caught because a peer flagged that the blackout will ship several commits at once unbisected and suggested reviewing mine before it goes out — worth doing, since this one touches the poll path.

CI note: typecheck will fail here on @preflight/runreport until #505 lands. That is the registry-401 blackout, not this change — the package cannot be installed at all right now, locally or in CI, and the failing files (run-report-provider.ts, runs/[testRunId]/page.tsx) are untouched by this diff.

🤖 Generated with Claude Code

Follow-up to #502, before it ships. #502 is merged but has never deployed — it is one of the commits stuck behind the registry-401 deploy blackout — so this corrects it while it is still unreleased rather than after. **The problem.** The resolver joined on a computed key: ```sql alerts.dedupe_key = 'ci_failed:' || changesets.id ``` No index can serve that. `EXPLAIN (ANALYZE)` on production: ``` Hash Join (actual time=1.664..1.665 rows=0) -> Index Scan using alert_status_idx on alerts (rows=11) -> Seq Scan on changesets (rows=3063) <- every poll Execution Time: 1.701 ms ``` 1.7ms is not slow. But it runs on **every agent poll — roughly 840 times an hour across seven nodes** — and the scan grows with the changesets table rather than with the handful of firing alerts it is actually about. That is the same shape as the unindexed poll that starved this database this morning (bob re-reading `session_event`), so it is not a pattern to leave sitting in the agent hot path. **The fix.** The dedupe key already contains the changeset id, so the join was recovering something we already had. Read the firing alerts by index, parse the ids out of the keys, look them up by primary key: ``` Index Scan using alert_status_idx on alerts 0.066 ms Index Scan using changesets_pkey (11 loops) 0.299 ms ``` **0.37ms, and no sequential scan** — cost now scales with firing alerts, not with total changesets. Behaviour is unchanged: still resolve-only, still bounded, still leaves a genuinely open changeset alerting. 5 tests updated for the two-query shape, all green. Caught because a peer flagged that the blackout will ship several commits at once unbisected and suggested reviewing mine before it goes out — worth doing, since this one touches the poll path. **CI note:** `typecheck` will fail here on `@preflight/runreport` until #505 lands. That is the registry-401 blackout, not this change — the package cannot be installed at all right now, locally or in CI, and the failing files (`run-report-provider.ts`, `runs/[testRunId]/page.tsx`) are untouched by this diff. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
perf(alerts): look changesets up by id instead of hashing the whole table
All checks were successful
CI / gitleaks (pull_request) Successful in 6s
CI / storybook (pull_request) Successful in 1m25s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 11m17s
d10c648383
The settled-changeset resolver joined on a computed key,
alerts.dedupe_key = 'ci_failed:' || changesets.id, which no index can serve.
EXPLAIN on production: the alert side is an index scan for 11 rows, then a
sequential scan of all 3,063 changesets to build the hash. 1.7ms, but it runs
on every agent poll -- roughly 840 times an hour across the fleet -- and the
scan grows with the changesets table rather than with the handful of firing
alerts it is actually about.

That is the same shape as the unindexed poll that starved this database this
morning, so it is not a pattern to leave in the hot path. The dedupe key
already contains the changeset id: read the firing alerts by index, parse the
ids out, and look them up by primary key.
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!506
No description provided.