fix(traffic): stop the apex route stealing forgegraf.com, and serialise the reconcile #479

Merged
gmackie merged 1 commit from fix/apex-route-and-reconcile-claim into main 2026-08-26 23:19:37 +00:00
Owner

Follow-up to #475, which fixed only half the problem. Both halves proven by production evidence.

1. The apex route declaration was the real steal

Dropping --domain forgegraf.com from the deploy command wasn't enough — apps/web/wrangler.toml also declared:

routes = [ { pattern = "forgegraf.com", custom_domain = true } ]

so every deploy re-attached the apex to the forgegraf worker regardless of the flag. #475's own deploy (21:14:43–21:18:38) did it again, and the reconcile pulled the apex back at 21:17:59, mid-deploy.

ForgeGraph owns hostname attachment: forgegraf.com is a one-box hostname bound to forgegraph-router, and this worker is reached through the router's PRIMARY service binding, so it needs no hostname of its own. traffic.disable still hands the apex back to this worker if one-box is ever turned off, so nothing is orphaned. wrangler.staging.toml keeps its beta.forgegraf.com route — that hostname is not one-box-managed.

2. The reconcile was not serialised across isolates

The drift exposed a hazard in #458's sweep. Its interval gate is module-level state, which on Workers is per-isolate — it serialises nothing. Two isolates swept together and interleaved delete-then-ensure on the same hostname, leaving forgegraf.com attached to no worker in between.

Evidence: two traffic.hostnames_reconciled events for forgegraf.com at the same second (21:17:59), the second reporting reattached (a re-create) rather than rebound (a move) — i.e. one isolate found the apex unattached because the other was mid-swap. A hostname attached to nothing does not resolve, so that window is a brief outage of the platform's own apex.

Each split is now claimed with an atomic conditional UPDATE ... WHERE updated_at < now() - interval, so exactly one isolate proceeds per split per interval. No migration needed.

59/59 lifecycle + router tests green (new: "skips a split another isolate already claimed this interval"); tsc clean.

Unrelated observation: drain > writes the weight to zero... runs 2.7–4.1s against vitest's 5s default and intermittently times out. Pre-existing (reproduced on a clean tree); the suite's retry: 2 absorbs it, but it deserves its own fix.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f

Follow-up to #475, which fixed only half the problem. Both halves proven by production evidence. ## 1. The apex route declaration was the real steal Dropping `--domain forgegraf.com` from the deploy command wasn't enough — `apps/web/wrangler.toml` also declared: ```toml routes = [ { pattern = "forgegraf.com", custom_domain = true } ] ``` so **every** deploy re-attached the apex to the `forgegraf` worker regardless of the flag. #475's own deploy (21:14:43–21:18:38) did it again, and the reconcile pulled the apex back at 21:17:59, mid-deploy. ForgeGraph owns hostname attachment: `forgegraf.com` is a one-box hostname bound to `forgegraph-router`, and this worker is reached through the router's PRIMARY service binding, so it needs no hostname of its own. `traffic.disable` still hands the apex back to this worker if one-box is ever turned off, so nothing is orphaned. `wrangler.staging.toml` keeps its `beta.forgegraf.com` route — that hostname is not one-box-managed. ## 2. The reconcile was not serialised across isolates The drift exposed a hazard in #458's sweep. Its interval gate is module-level state, which on Workers is **per-isolate** — it serialises nothing. Two isolates swept together and interleaved `delete`-then-`ensure` on the same hostname, leaving `forgegraf.com` attached to no worker in between. Evidence: two `traffic.hostnames_reconciled` events for `forgegraf.com` at the *same second* (21:17:59), the second reporting `reattached` (a re-create) rather than `rebound` (a move) — i.e. one isolate found the apex unattached because the other was mid-swap. A hostname attached to nothing does not resolve, so that window is a brief outage of the platform's own apex. Each split is now claimed with an atomic conditional `UPDATE ... WHERE updated_at < now() - interval`, so exactly one isolate proceeds per split per interval. No migration needed. 59/59 lifecycle + router tests green (new: "skips a split another isolate already claimed this interval"); tsc clean. **Unrelated observation:** `drain > writes the weight to zero...` runs 2.7–4.1s against vitest's 5s default and intermittently times out. Pre-existing (reproduced on a clean tree); the suite's `retry: 2` absorbs it, but it deserves its own fix. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
fix(traffic): stop the apex route stealing forgegraf.com, and serialise the reconcile
All checks were successful
CI / gitleaks (pull_request) Successful in 6s
CI / storybook (pull_request) Successful in 1m26s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m42s
92ae13ca1b
Removing --domain from the deploy command was not enough: apps/web/
wrangler.toml declared

  routes = [{ pattern = "forgegraf.com", custom_domain = true }]

so every deploy re-attached the apex to the forgegraf worker regardless.
The deploy at 21:14-21:18 today did it again, and the reconcile pulled it
back mid-deploy. ForgeGraph owns hostname attachment and the router
reaches this worker by service binding, so the declaration is dropped;
traffic.disable still hands the apex back if one-box is turned off.

That drift also exposed a hazard in the reconcile itself. Its interval
gate is module state, which on Workers is per-isolate and serialises
nothing: two isolates swept together and interleaved delete-then-ensure
on the same hostname, leaving forgegraf.com attached to no worker in
between (two hostnames_reconciled events in the same second, the second
a re-create rather than a move). Each split is now claimed with an
atomic conditional update, so exactly one isolate acts per interval.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
Author
Owner

Preview environment is live: https://pr-479-forgegraph.forgegraf.com

Deployed 92ae13ca with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-479-forgegraph.forgegraf.com Deployed `92ae13ca` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!479
No description provided.