fix(deploy): stop Corepack prompting on the runner's pty after a reboot #574

Merged
gmackie merged 2 commits from fix/corepack-deploy-prompt into main 2026-09-14 03:52:57 +00:00
Owner

Production has not deployed since 2026-09-09. The last successful deploy was run 17946 (df0c19ab). All eight main deploys since then were cancelled. Five hit the job's 25-minute timeout, and three were cancelled when a newer push landed behind them. So #564, #567 and #571 are merged but unshipped — their routes (/api/fg/contracts, /api/fg/alchemy/*) 404 in production while forge-health reports 200. CI stayed green throughout; only Deploy ForgeGraf / deploy shows "cancelled", which reads like concurrency and isn't.

Cause

Every timed-out run stops at the same line, straight after pnpm install:

Done in 7.2s using pnpm v10.19.0
! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-10.19.0.tgz
this step has been cancelled: ctx: context deadline exceeded
  1. The deploy's db:push runs as sudo -u postgres env HOME=/tmp … pnpm. HOME=/tmp dates from 2026-05-18, to avoid a Corepack EACCES on /root/package.json.
  2. That puts Corepack's pnpm cache in /tmp. hetzner-fg rebooted 2026-09-10 02:21, which emptied it; the first timeout was 27 minutes later.
  3. The next run has to download pnpm, and Corepack asks Do you want to continue? [Y/n] first — but only when stdin is a terminal.
  4. The runner gives steps one: act_runner is built with creack/pty, and sudoers has Defaults use_pty. Nothing answers, so the step waits until it's killed. The question has no trailing newline, so it never reaches the job log.

Verified on the runner host

stdin prompt enabled COREPACK_ENABLE_DOWNLOAD_PROMPT=0
pty (script -qec) stops at [Y/n] prints 10.19.0
/dev/null notice, then downloads prints 10.19.0
open pipe notice, then downloads prints 10.19.0

Only the pty case reproduces the hang, which is why simpler repros miss it. It also isn't the network: the runner is host-mode, outbound traffic is allowed, and registry.npmjs.org answers in ~50ms as both root and postgres.

Fix

Add COREPACK_ENABLE_DOWNLOAD_PROMPT=0 to the env list in buildLocalDrizzlePushCommand. It has to live in the command: sudo's env_reset strips any job-level variable. deploy-staging.yml runs the same script, so it's covered too.

The exact-array test is updated, and there's a regression test asserting the flag sits inside the sudo env segment. 5/5 pass.

Heads-up

While reproducing this I ran the command once with the real HOME=/tmp, which re-downloaded pnpm into that cache. The next deploy will probably pass even without this change — until the next reboot empties /tmp again. A green deploy right now doesn't prove the fix; the pty table above does.

Merging this self-deploys and should ship #564/#567/#571 along with it.

🤖 Generated with Claude Code

**Production has not deployed since 2026-09-09.** The last successful deploy was run 17946 (`df0c19ab`). All eight main deploys since then were cancelled. Five hit the job's 25-minute timeout, and three were cancelled when a newer push landed behind them. So #564, #567 and #571 are merged but unshipped — their routes (`/api/fg/contracts`, `/api/fg/alchemy/*`) 404 in production while `forge-health` reports 200. CI stayed green throughout; only `Deploy ForgeGraf / deploy` shows "cancelled", which reads like concurrency and isn't. ## Cause Every timed-out run stops at the same line, straight after `pnpm install`: ``` Done in 7.2s using pnpm v10.19.0 ! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-10.19.0.tgz this step has been cancelled: ctx: context deadline exceeded ``` 1. The deploy's `db:push` runs as `sudo -u postgres env HOME=/tmp … pnpm`. `HOME=/tmp` dates from 2026-05-18, to avoid a Corepack EACCES on `/root/package.json`. 2. That puts Corepack's pnpm cache in `/tmp`. **hetzner-fg rebooted 2026-09-10 02:21**, which emptied it; the first timeout was 27 minutes later. 3. The next run has to download pnpm, and Corepack asks `Do you want to continue? [Y/n]` first — **but only when stdin is a terminal.** 4. The runner gives steps one: `act_runner` is built with `creack/pty`, and sudoers has `Defaults use_pty`. Nothing answers, so the step waits until it's killed. The question has no trailing newline, so it never reaches the job log. ## Verified on the runner host | stdin | prompt enabled | `COREPACK_ENABLE_DOWNLOAD_PROMPT=0` | |---|---|---| | pty (`script -qec`) | **stops at `[Y/n]`** | prints `10.19.0` | | `/dev/null` | notice, then downloads | prints `10.19.0` | | open pipe | notice, then downloads | prints `10.19.0` | Only the pty case reproduces the hang, which is why simpler repros miss it. It also isn't the network: the runner is host-mode, outbound traffic is allowed, and `registry.npmjs.org` answers in ~50ms as both root and postgres. ## Fix Add `COREPACK_ENABLE_DOWNLOAD_PROMPT=0` to the env list in `buildLocalDrizzlePushCommand`. It **has to live in the command**: sudo's `env_reset` strips any job-level variable. `deploy-staging.yml` runs the same script, so it's covered too. The exact-array test is updated, and there's a regression test asserting the flag sits inside the sudo `env` segment. 5/5 pass. ## Heads-up While reproducing this I ran the command once with the real `HOME=/tmp`, which re-downloaded pnpm into that cache. **The next deploy will probably pass even without this change** — until the next reboot empties `/tmp` again. A green deploy right now doesn't prove the fix; the pty table above does. Merging this self-deploys and should ship #564/#567/#571 along with it. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(deploy): stop Corepack prompting on the runner's pty after a reboot
Some checks failed
CI / gitleaks (pull_request) Successful in 8s
forgegraph/ci CI failed
CI / ci (pull_request) Failing after 5m7s
CI / storybook (pull_request) Failing after 5m51s
edee980785
Every self-deploy since the 2026-09-10 reboot of the prod runner hung for
25 minutes and was killed, so production has served the 2026-09-09 build
while #564, #567 and #571 sat merged but unshipped. CI stayed green; only
the deploy job was "cancelled", which reads like concurrency and isn't.

The deploy's db:push runs as `sudo -u postgres env HOME=/tmp … pnpm`.
HOME=/tmp puts Corepack's pnpm cache in /tmp, and the reboot emptied it,
so the next run had to download pnpm again. Corepack asks
"Do you want to continue? [Y/n]" before downloading whenever stdin is a
terminal, and the runner gives steps one (act_runner uses creack/pty;
sudoers has use_pty). Nothing answers, so the step waited until killed.
The question has no trailing newline, which is why the job log ends at
"! Corepack is about to download …" and never shows it.

Reproduced on the runner host: under a pty the command stops at the
prompt; with COREPACK_ENABLE_DOWNLOAD_PROMPT=0 it prints 10.19.0. With
stdin as /dev/null or a pipe it never hangs, which is why simple repros
miss it.

The flag has to be in the command itself: sudo's env_reset strips any
job-level variable. deploy-staging.yml uses the same script, so it is
fixed too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gmackie scheduled this pull request to auto merge when all checks succeed 2026-09-14 02:36:50 +00:00
ci: re-run after npm.forgegraf.com was restored
All checks were successful
CI / gitleaks (pull_request) Successful in 8s
CI / storybook (pull_request) Successful in 3m33s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 11m23s
9a03811b59
The first CI run failed in pnpm install with ERR_PNPM_FETCH_502 from the
private registry, which had been down since 2026-09-13 (its store path
was garbage-collected). No code change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!574
No description provided.