fix(traffic): give the one-box canary its secrets #477
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!477
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/traffic-lifecycle-recovery"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Root cause of a live control-plane incident: forgegraf.com served from forgegraph-onebox, a Worker with zero secrets, so every bearer request 401d and all 7 node agents went offline for ~30 minutes.
The enable path comments that the one-box target "shares secrets" with the stage. True for a node service (env file from stage secrets each deploy); false for a Worker, where Cloudflare scopes secrets per worker name. The canary is created with none, deploying code adds none, and they cannot be copied off the primary because Worker secrets are write-only. They must be pushed from our encrypted store.
putWorkerSecretalready existed with exactly one caller, none in the one-box path.Why it is worse than a degraded canary:
verifyBearerTokenWithDBreturns invalid when FG_API_TOKEN is unset, before the DB is consulted. A canary owning the production hostname denies all machine auth platform-wide, and the deploy meant to replace it needs that same auth to finish, so it cannot self-recover.enable now pushes stage secrets to the canary and refuses to enable if any fail, naming them. Half-configured is the dangerous state.
Preview environment is live: https://pr-477-forgegraph.forgegraf.com
Deployed
f27bd9f1with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.Pushing the stage's secrets is not proof the canary is complete, and my own previous commit made exactly that mistake: it verified the pushes succeeded, not that the required secrets exist. A canary can receive every stage secret, report zero failures, and still be fatally incomplete. On this install 26 of the primary Worker's 41 secrets are not in ForgeGraph's store at all -- FG_API_TOKEN, FG_SESSION_KEY, FG_ENCRYPTION_KEY, BETTER_AUTH_SECRET, every OAuth client secret. They were set directly with wrangler secret put <key> Create or update a secret for a Worker POSITIONALS key The variable name to be accessible in the Worker [string] [required] GLOBAL FLAGS -c, --config Path to Wrangler configuration file [string] --cwd Run as if Wrangler was started in the specified directory instead of the current working directory [string] -e, --env Environment to use for operations, and for selecting .env and .dev.vars files [string] --env-file Path to an .env file to load - can be specified multiple times - values from earlier files are overridden by values in later files [array] -h, --help Show help [boolean] -v, --version Show version number [boolean] OPTIONS --name Name of the Worker. If this is not specified, it will default to the name specified in your Wrangler config file. [string], and Worker secrets are write-only, so nothing in the system can read them to copy them. The sync would have pushed 15 of 41 and declared success. Cloudflare withholds secret values but not secret *names*, so parity is checkable without ever holding plaintext -- which is the only option here, since for those 26 no plaintext exists anywhere the server can reach. enable now compares the two Workers' secret names and refuses with PRECONDITION_FAILED, naming what is missing and how to add it, rather than handing the production hostname to a Worker that cannot authenticate. That is the check that would have prevented today's incident. 331 tests pass across packages/api. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Preview environment is live: https://pr-477-forgegraph.forgegraf.com
Deployed
66180854with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.Reviewed. The diagnosis is right and this is not superseded —
git grep listWorkerSecretNames|syncSecretsToCanary|assertCanaryCanServeTrafficagainst main returns nothing, and the three traffic commits that landed since (82d29c2d#475,d2af6f91#479,849263f3#482) all address hostname stealing, a different facet of the same incident.The red CI is not this PR's fault. Task
23991failed inVet and test Go agenton a transient module-proxy error:The same log reports
test | passed | 1151 passed / 0 failed / 10 skipped. A re-run should go green.The blocker: this would stop deploys fleet-wide, including forgegraf.com
assertCanaryCanServeTrafficfires inshiftAndBake, which every one-box deploy passes through. But the only thing that ever pushes secrets to a canary issyncSecretsToCanary, called from exactly one place — theenablemutation inpackages/api/src/routers/traffic.ts:303. Nothing re-syncs on deploy.I checked the production control-plane DB:
All 22 enabled splits —
calzone,controlsfoundry,creator,crucible,daily-dose,driftport,fabforge,festigram,forgegraph,fryos,habit-app,insure,jobs-pulse,latchflow,linear-clone,netcontrol,omnidat-app,playtrek,streamconductor,test-dojo,trip,veritas— were enabled before this code existed, so none of their canaries has ever received a secret push. On merge, each one's next production deploy throws at the shift step, and re-runningenableis the only remediation.It is worse for ForgeGraph itself: the PR body notes 26 of the primary's 41 secrets were set out-of-band with
wrangler secret putand are not in the store, so even a disable/re-enable would fail the strict parity check. That is a self-blocking condition on the platform's own deploy path.Suggested shape: land the enable-time sync now, but either put the parity guard behind a flag / warn-only mode, or add a shift-time sync before the shift-time assert, so an already-enabled split self-heals instead of dead-ending. The 26 out-of-band secrets need backfilling into the store before a strict check goes live either way.
Two smaller issues
listWorkerSecretNamesfails in opposite directions. It returns[]wheneverdata.successis false and never checksresp.ok. A CF API error on the primary lookup therefore makes the guard silently pass (fails open, defeating it); the same error on the canary lookup reports every primary secret as missing and blocks the deploy (fails closed, spuriously).deploymentTargetsselect inshiftAndBakeruns before theworkerNamecheck, so every shift pays a DB roundtrip even on node-platform lanes. Cosmetic.Not merging — the guard's rollout needs a decision, not a patch.
64ceead063d975761077Preview environment is live: https://pr-477-forgegraph.forgegraf.com
Deployed
d9757610with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.