one-box worker migration deploys a renamed CF worker without migrating secrets (two production outages) #452

Open
opened 2026-08-25 08:20:09 +00:00 by gmackie · 0 comments
Owner

The one-box deployment model deploys a NEW-named Cloudflare worker (<app>-onebox / <app>-app-web) and moves the custom domain, but does not migrate the old worker's secrets. The new worker then throws at startup env validation, every route 500s, and /api/health can still report healthy (it skips validation), so monitoring misses it.

Observed twice on 2026-08-24/25:

  1. playtrek — pipeline deploys --name playtrek-onebox; worker had 0 secrets vs 26 on nextjs. Three production deploys failed with cloudflare_worker_deploy_failed (CF error 10021 in authEnv); the stored deploy journal truncates before the actual error line, which made diagnosis needlessly hard. Recovered by bulk-loading recoverable secrets onto onebox.
  2. preflight-app — preflight-app-web deployed 02:30 with 0 secrets while old preflight worker held all 19 (ASC keys, signing certs, OAuth). Whole build farm down (runner register fail-loop). Recovered by deploying current main to the OLD worker name (same-name deploys keep secrets; wrangler config re-claimed the domain).

Asks:

  • Migrate (or re-point) worker secrets as part of the one-box rename, or fail the deploy pre-flight when the target worker is missing required secrets.
  • Stop truncating journalOutput before the failing command's stderr — the wrangler error was cut out in every failed deploy record.
  • Consider a post-deploy probe beyond /api/health (e.g. / status) since env-validation failures don't show there.

Filed from the recovery session; full recipes recorded in the playtrek repo's ops learnings.

The one-box deployment model deploys a NEW-named Cloudflare worker (`<app>-onebox` / `<app>-app-web`) and moves the custom domain, but does **not** migrate the old worker's secrets. The new worker then throws at startup env validation, every route 500s, and `/api/health` can still report healthy (it skips validation), so monitoring misses it. Observed twice on 2026-08-24/25: 1. **playtrek** — pipeline deploys `--name playtrek-onebox`; worker had 0 secrets vs 26 on `nextjs`. Three production deploys failed with `cloudflare_worker_deploy_failed` (CF error 10021 in `authEnv`); the stored deploy journal truncates before the actual error line, which made diagnosis needlessly hard. Recovered by bulk-loading recoverable secrets onto onebox. 2. **preflight-app** — `preflight-app-web` deployed 02:30 with 0 secrets while old `preflight` worker held all 19 (ASC keys, signing certs, OAuth). Whole build farm down (runner register fail-loop). Recovered by deploying current main to the OLD worker name (same-name deploys keep secrets; wrangler config re-claimed the domain). Asks: - Migrate (or re-point) worker secrets as part of the one-box rename, or fail the deploy pre-flight when the target worker is missing required secrets. - Stop truncating `journalOutput` before the failing command's stderr — the wrangler error was cut out in every failed deploy record. - Consider a post-deploy probe beyond `/api/health` (e.g. `/` status) since env-validation failures don't show there. Filed from the recovery session; full recipes recorded in the playtrek repo's ops learnings.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph#452
No description provided.