feat(traffic): strip platform-managed hostnames from wrangler config on every deploy #482

Merged
gmackie merged 2 commits from feat/strip-managed-routes into main 2026-08-27 01:58:54 +00:00
Owner

Closes the last steal vector, found by auditing the whole fleet rather than one app.

The finding

Of the 22 enabled one-box apps, 18 declare their own one-box hostname in their wrangler config. Only controlsfoundry, streamconductor, jobs-pulse and (after #479) forgegraph are clean.

Agent 0.1.57's route strip fires only when the deploy carries a --name override, so canary and preview deploys were already safe — but the one-box primary deploy, whose config name legitimately matches the target, slipped through. Every production deploy of those 18 apps re-attached its public hostname from the router to the primary worker: the split was bypassed until the next reconcile sweep, with a brief unattached window during the swap. habit-app, latchflow, veritas, netcontrol, insure, test-dojo and calzone are among them.

The fix

The control plane sends the split's hostnames as managedHostnames in the deploy payload (only for cloudflare-workers deployments with an enabled split), and the agent removes only the route entries matching them:

  • { "pattern": "shop.example.com", "custom_domain": true } → stripped
  • { "pattern": "api.shop.example.com" } → kept (not platform-managed)
  • "shop.example.com/*" and route = "..." scalars → matched on the hostname part
  • per-env blocks handled; the app's name never touched

So an app can still declare routes of its own; it just can't claim a hostname the router owns. This makes the #458 reconcile a safety net rather than the mechanism that holds routing together.

The alternative was 18 per-repo config PRs, which fixes today's fleet but not the next app onboarded.

Agent tests + pkg/client green, go vet clean, tsc clean on apps/web. Needs an agent release to take effect; until then the reconcile continues to cover it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f

Closes the last steal vector, found by auditing the whole fleet rather than one app. ## The finding Of the 22 enabled one-box apps, **18 declare their own one-box hostname in their wrangler config**. Only `controlsfoundry`, `streamconductor`, `jobs-pulse` and (after #479) `forgegraph` are clean. Agent 0.1.57's route strip fires only when the deploy carries a `--name` override, so **canary and preview deploys were already safe** — but the one-box **primary** deploy, whose config name legitimately matches the target, slipped through. Every production deploy of those 18 apps re-attached its public hostname from the router to the primary worker: the split was bypassed until the next reconcile sweep, with a brief unattached window during the swap. `habit-app`, `latchflow`, `veritas`, `netcontrol`, `insure`, `test-dojo` and `calzone` are among them. ## The fix The control plane sends the split's hostnames as `managedHostnames` in the deploy payload (only for `cloudflare-workers` deployments with an enabled split), and the agent removes **only** the route entries matching them: - `{ "pattern": "shop.example.com", "custom_domain": true }` → stripped - `{ "pattern": "api.shop.example.com" }` → **kept** (not platform-managed) - `"shop.example.com/*"` and `route = "..."` scalars → matched on the hostname part - per-`env` blocks handled; the app's `name` never touched So an app can still declare routes of its own; it just can't claim a hostname the router owns. This makes the #458 reconcile a safety net rather than the mechanism that holds routing together. The alternative was 18 per-repo config PRs, which fixes today's fleet but not the next app onboarded. Agent tests + `pkg/client` green, `go vet` clean, `tsc` clean on apps/web. **Needs an agent release to take effect**; until then the reconcile continues to cover it. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
feat(traffic): strip platform-managed hostnames from wrangler config on every deploy
Some checks failed
CI / gitleaks (pull_request) Has been cancelled
CI / storybook (pull_request) Has been cancelled
CI / ci (pull_request) Has been cancelled
52510e25c1
An audit of the 22 enabled one-box apps found 18 declaring their own
one-box hostname in wrangler config. Agent 0.1.57's strip only fires on a
--name override, so canary and preview deploys were safe but the one-box
PRIMARY deploy — whose config name legitimately matches the target — kept
re-attaching the public hostname to the primary worker. The split was then
bypassed until the next reconcile sweep, with a brief unattached window
during the swap.

The control plane now sends the split's hostnames as managedHostnames in
the deploy payload, and the agent removes only the route entries matching
them. Routes the platform does not manage are left untouched, so an app
can still declare an API subdomain of its own. Matching is on the hostname
part of the pattern, so bare, wildcard and {pattern, custom_domain} forms
all match.

This makes the reconcile a safety net rather than the mechanism.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
Author
Owner

Preview environment is live: https://pr-482-forgegraph.forgegraf.com

Deployed 52510e25 with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-482-forgegraph.forgegraf.com Deployed `52510e25` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
test(traffic): warm the module import so drain stops flaking
All checks were successful
CI / gitleaks (pull_request) Successful in 6s
CI / storybook (pull_request) Successful in 1m45s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m31s
31ec5f84a1
Every test in this file imports traffic-lifecycle dynamically so the
vi.mock factories apply, and resolving it pulls in drizzle and the whole
schema. That 3-4s cost landed on whichever test imported first — drain —
putting it against vitest's 5s default and failing the suite
intermittently (reproduced on a clean tree: 2.7s, 3.6s, 4.1s, and two
timeouts under load). retry:2 hid it at the cost of re-running the file.

Importing once in beforeAll moves the cost outside any test's timeout:
drain goes 3209ms -> 2ms and the file runs in 1.1s instead of ~4s.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CMpzX1b6swjezEptw3T71f
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!482
No description provided.