fix(traffic): re-read splits before reconciling, and stop the self-deploy stealing forgegraf.com #475

Merged
gmackie merged 2 commits from fix/reconcile-stale-split into main 2026-08-26 21:11:10 +00:00
Owner

Two fixes for the same story: hostname ownership, both found by live evidence within hours of #458 deploying.

1. Re-read a split before reconciling it

reconcileEnabledSplitHostnames listed enabled splits once per sweep, then acted on that snapshot. A Cloudflare round-trip for an earlier split gives a concurrent enable/disable time to land; acting on the stale row means moving a hostname back onto a router disable has already deleted. The move detaches the hostname before re-attaching, so only restore-on-failure keeps the site serving.

Observed live 2026-08-26 while retiring a duplicate latchflow-beta app: traffic.disabled at 18:26:34, the sweep acted at 18:26:39 and 404'd attaching beta.latchflow.io to the just-deleted latchflow-beta-router. The restore held (302 throughout) — but a failed restore there is precisely the habitplay.io dark-hostname failure mode.

Each split is now re-read immediately before its Cloudflare calls and skipped if it is no longer enabled, no longer in a reconcilable state, has no hostnames, or has just changed state. The allowed-state list is hoisted to RECONCILABLE_STATES so the query and the guard cannot drift.

2. Stop our own deploy stealing forgegraf.com

deploy.yml ran opennextjs-cloudflare deploy --domain forgegraf.com, which re-attached the apex to the forgegraf worker on every self-deploy — pulling it off forgegraph-router. So between each deploy and the next reconcile sweep, ForgeGraph silently bypassed its own one-box canary, and the swap left the hostname attached to nothing in between.

Evidence: four traffic.hostnames_reconciled events for forgegraf.com on 2026-08-26 alone (09:52, 15:41, 16:31, 18:21), the last one a re-create rather than a move — i.e. the apex was attached to no worker when the sweep found it. The router reaches the worker through a service binding, so it needs no hostname of its own.

This is the production-deploy steal-back case #458 predicted, with the platform itself as the offender.

17/17 lifecycle + 41/41 router tests green (new: "skips a split that was disabled between the listing and the sweep"); tsc clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f

Two fixes for the same story: hostname ownership, both found by live evidence within hours of #458 deploying. ## 1. Re-read a split before reconciling it `reconcileEnabledSplitHostnames` listed enabled splits once per sweep, then acted on that snapshot. A Cloudflare round-trip for an earlier split gives a concurrent enable/disable time to land; acting on the stale row means moving a hostname back onto a router `disable` has already deleted. The move detaches the hostname before re-attaching, so only restore-on-failure keeps the site serving. **Observed live 2026-08-26** while retiring a duplicate `latchflow-beta` app: `traffic.disabled` at 18:26:34, the sweep acted at 18:26:39 and 404'd attaching `beta.latchflow.io` to the just-deleted `latchflow-beta-router`. The restore held (302 throughout) — but a failed restore there is precisely the habitplay.io dark-hostname failure mode. Each split is now re-read immediately before its Cloudflare calls and skipped if it is no longer `enabled`, no longer in a reconcilable state, has no hostnames, or has just changed state. The allowed-state list is hoisted to `RECONCILABLE_STATES` so the query and the guard cannot drift. ## 2. Stop our own deploy stealing forgegraf.com `deploy.yml` ran `opennextjs-cloudflare deploy --domain forgegraf.com`, which re-attached the apex to the `forgegraf` worker on **every** self-deploy — pulling it off `forgegraph-router`. So between each deploy and the next reconcile sweep, ForgeGraph silently bypassed its own one-box canary, and the swap left the hostname attached to nothing in between. Evidence: four `traffic.hostnames_reconciled` events for `forgegraf.com` on 2026-08-26 alone (09:52, 15:41, 16:31, 18:21), the last one a **re-create** rather than a move — i.e. the apex was attached to no worker when the sweep found it. The router reaches the worker through a service binding, so it needs no hostname of its own. This is the production-deploy steal-back case #458 predicted, with the platform itself as the offender. 17/17 lifecycle + 41/41 router tests green (new: "skips a split that was disabled between the listing and the sweep"); tsc clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
fix(traffic): re-read a split before reconciling its hostnames
Some checks failed
CI / gitleaks (pull_request) Has been cancelled
CI / storybook (pull_request) Has been cancelled
CI / ci (pull_request) Has been cancelled
92269874d6
The reconcile listed enabled splits once per sweep and then acted on
that snapshot. A Cloudflare round-trip for an earlier split gives a
concurrent enable/disable time to land, and acting on the stale row
means moving a hostname back onto a router that disable has already
deleted — the move detaches the hostname first, so only the
restore-on-failure path keeps the site serving.

Observed live 2026-08-26 retiring a duplicate app: disable at :34, the
sweep acted at :39 and 404'd on the deleted router. The restore held and
beta.latchflow.io stayed up, but a failed restore there is exactly how a
hostname goes dark.

Each split is now re-read immediately before its Cloudflare calls and
skipped if it is no longer enabled, no longer in a reconcilable state,
has no hostnames, or has just changed state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
fix(deploy): stop the self-deploy stealing forgegraf.com from its own router
All checks were successful
CI / gitleaks (pull_request) Successful in 6s
CI / storybook (pull_request) Successful in 1m39s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m55s
5126085d71
`--domain forgegraf.com` re-attached the apex to the `forgegraf` worker on
every ForgeGraph deploy, yanking it off forgegraph-router. Between the
deploy and the next reconcile sweep the platform silently bypassed its own
one-box canary, and the swap left the hostname attached to nothing in
between — 2026-08-26 shows four traffic.hostnames_reconciled events for
forgegraf.com in one day, one of them a re-create rather than a move.

The router reaches this worker through a service binding, so the worker
needs no hostname of its own; hostname ownership belongs to the route
table and the one-box router.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
gmackie changed title from fix(traffic): re-read a split before reconciling its hostnames to fix(traffic): re-read splits before reconciling, and stop the self-deploy stealing forgegraf.com 2026-08-26 18:35:26 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!475
No description provided.