fix(traffic): stop the apex route stealing forgegraf.com, and serialise the reconcile #479
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!479
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/apex-route-and-reconcile-claim"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up to #475, which fixed only half the problem. Both halves proven by production evidence.
1. The apex route declaration was the real steal
Dropping
--domain forgegraf.comfrom the deploy command wasn't enough —apps/web/wrangler.tomlalso declared:so every deploy re-attached the apex to the
forgegrafworker regardless of the flag. #475's own deploy (21:14:43–21:18:38) did it again, and the reconcile pulled the apex back at 21:17:59, mid-deploy.ForgeGraph owns hostname attachment:
forgegraf.comis a one-box hostname bound toforgegraph-router, and this worker is reached through the router's PRIMARY service binding, so it needs no hostname of its own.traffic.disablestill hands the apex back to this worker if one-box is ever turned off, so nothing is orphaned.wrangler.staging.tomlkeeps itsbeta.forgegraf.comroute — that hostname is not one-box-managed.2. The reconcile was not serialised across isolates
The drift exposed a hazard in #458's sweep. Its interval gate is module-level state, which on Workers is per-isolate — it serialises nothing. Two isolates swept together and interleaved
delete-then-ensureon the same hostname, leavingforgegraf.comattached to no worker in between.Evidence: two
traffic.hostnames_reconciledevents forforgegraf.comat the same second (21:17:59), the second reportingreattached(a re-create) rather thanrebound(a move) — i.e. one isolate found the apex unattached because the other was mid-swap. A hostname attached to nothing does not resolve, so that window is a brief outage of the platform's own apex.Each split is now claimed with an atomic conditional
UPDATE ... WHERE updated_at < now() - interval, so exactly one isolate proceeds per split per interval. No migration needed.59/59 lifecycle + router tests green (new: "skips a split another isolate already claimed this interval"); tsc clean.
Unrelated observation:
drain > writes the weight to zero...runs 2.7–4.1s against vitest's 5s default and intermittently times out. Pre-existing (reproduced on a clean tree); the suite'sretry: 2absorbs it, but it deserves its own fix.🤖 Generated with Claude Code
https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
Removing --domain from the deploy command was not enough: apps/web/ wrangler.toml declared routes = [{ pattern = "forgegraf.com", custom_domain = true }] so every deploy re-attached the apex to the forgegraf worker regardless. The deploy at 21:14-21:18 today did it again, and the reconcile pulled it back mid-deploy. ForgeGraph owns hostname attachment and the router reaches this worker by service binding, so the declaration is dropped; traffic.disable still hands the apex back if one-box is turned off. That drift also exposed a hazard in the reconcile itself. Its interval gate is module state, which on Workers is per-isolate and serialises nothing: two isolates swept together and interleaved delete-then-ensure on the same hostname, leaving forgegraf.com attached to no worker in between (two hostnames_reconciled events in the same second, the second a re-create rather than a move). Each split is now claimed with an atomic conditional update, so exactly one isolate acts per interval. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71fPreview environment is live: https://pr-479-forgegraph.forgegraf.com
Deployed
92ae13cawith the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.