fix(traffic): close the three ways a Worker secret sync misreports itself #542
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!542
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/harden-worker-secret-sync"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Three fixes from the 2026-08-30 control-plane incident. Each one made a correct-looking system misreport its own state, and each cost real time tonight.
1. A plain
[vars]name wedgedenablepermanentlyCloudflare answers
10053 binding name already in usetosecret puton a name already bound as an env var. The Worker has the value — it just isn't a secret. ButsyncSecretsToCanaryrecorded it asfailed, andtraffic.enablethrows on anyfailedentry with nothing gating that throw. So the split could never be enabled, and no retry could clear it.Hit live on
OTEL_EXPORTER_OTLP_ENDPOINT, whichagent/cmd/agent/deploy_env.go:57injects as a var at deploy time while the stage store also carried it:Now classified
skipped. A genuine push error is stillfailed— there's a test for that, because collapsing the two would hide real breakage.2. The CLI sync could re-introduce what #541 removed
forge secret sync --target cloudflarehad no denylist, so it could pushDATABASE_URLback onto a Worker through the other door. It targetsforgegraf, which is why it hasn't bitten yet — but the apex now routes through the one-box router, so that Worker serves production.Mirrors
WORKER_FORBIDDEN_SECRET_KEYS, with a test on each side so the two lists can't drift apart silently.3. The recovery workflow retried instead of diagnosing
It probed
forgegraf.com60 times over two minutes while printing the answer on the line immediately above the loop:Direct worker healthy + domain broken is a routing fault, not a secrets fault. Retrying cannot fix a hostname pointed at a different script. Five consecutive red runs reported only
503.It now fails immediately and says so, naming the query that identifies the hostname's owner. The other branch is explicit too: if the direct probe was also unhealthy, it says the fault is the Worker, not the binding.
Tests
src/libfiles; 3 new covering the 10053 split, including that a real CF error is stillfailedgo build ./...clean, YAML parsesNote:
packages/apivitest can't run through its normal entrypoint locally (PGlite global setup fails, pre-existing), so these were run with a scoped config. That gap is exactly what let a broken test through on #541, so this time the whole blast radius was run rather than one file.🤖 Generated with Claude Code
Preview environment is live: https://pr-542-forgegraph.forgegraf.com
Deployed
d5ccfc8awith the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.