fix(deploy): stop Corepack prompting on the runner's pty after a reboot #574
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!574
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/corepack-deploy-prompt"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Production has not deployed since 2026-09-09. The last successful deploy was run 17946 (
df0c19ab). All eight main deploys since then were cancelled. Five hit the job's 25-minute timeout, and three were cancelled when a newer push landed behind them. So #564, #567 and #571 are merged but unshipped — their routes (/api/fg/contracts,/api/fg/alchemy/*) 404 in production whileforge-healthreports 200. CI stayed green throughout; onlyDeploy ForgeGraf / deployshows "cancelled", which reads like concurrency and isn't.Cause
Every timed-out run stops at the same line, straight after
pnpm install:db:pushruns assudo -u postgres env HOME=/tmp … pnpm.HOME=/tmpdates from 2026-05-18, to avoid a Corepack EACCES on/root/package.json./tmp. hetzner-fg rebooted 2026-09-10 02:21, which emptied it; the first timeout was 27 minutes later.Do you want to continue? [Y/n]first — but only when stdin is a terminal.act_runneris built withcreack/pty, and sudoers hasDefaults use_pty. Nothing answers, so the step waits until it's killed. The question has no trailing newline, so it never reaches the job log.Verified on the runner host
COREPACK_ENABLE_DOWNLOAD_PROMPT=0script -qec)[Y/n]10.19.0/dev/null10.19.010.19.0Only the pty case reproduces the hang, which is why simpler repros miss it. It also isn't the network: the runner is host-mode, outbound traffic is allowed, and
registry.npmjs.organswers in ~50ms as both root and postgres.Fix
Add
COREPACK_ENABLE_DOWNLOAD_PROMPT=0to the env list inbuildLocalDrizzlePushCommand. It has to live in the command: sudo'senv_resetstrips any job-level variable.deploy-staging.ymlruns the same script, so it's covered too.The exact-array test is updated, and there's a regression test asserting the flag sits inside the sudo
envsegment. 5/5 pass.Heads-up
While reproducing this I ran the command once with the real
HOME=/tmp, which re-downloaded pnpm into that cache. The next deploy will probably pass even without this change — until the next reboot empties/tmpagain. A green deploy right now doesn't prove the fix; the pty table above does.Merging this self-deploys and should ship #564/#567/#571 along with it.
🤖 Generated with Claude Code