docs(field-reports): the self-hosted runners never restore the pnpm cache #481

Merged
gmackie merged 1 commit from docs/runner-cache-field-report into main 2026-08-27 01:34:57 +00:00
Owner

Every job on forgegraph-ci / forgegraph-worker logs pnpm cache is not found and does a full cold install over the network. When that network is flaky the install dies, and the PR fails for reasons unrelated to its contents.

#340 burned five runs on this. Against its exact failing head:

check result
pnpm install --frozen-lockfile locally succeeds
storybook build locally, this branch succeeds
storybook build locally, main succeeds
storybook job on #480 (branch off same main) passed
manifests vs main identical
regenerating the lockfile zero-line diff

Only the CI install fails.

Four of those five runs were spent looking in the wrong place — storybook config, workspace globs, hoisting — because the swallowing retry loop reported storybook: not found instead of the install failure. #480 fixed the message; this documents what the message now reveals.

The likely deadlock: setup-node saves its cache only after a successful install, so a run that cannot install never writes one, and the next run is cold too. Nothing inside the workflow breaks that cycle.

Where the boundary is: gmackie/bob uses identical setup-node + cache: pnpm config and restores a 787 MB cache fine on ubuntu-latest. Its deploy jobs, which run on forgegraph-worker, failed the same day with context canceled. Same config, different runner, different outcome.

Related same-root: bob #93 failed twice on Request timeout: /astral-sh/uv/releases/download/0.12.6/... before passing on a third attempt.

Docs only.

Every job on `forgegraph-ci` / `forgegraph-worker` logs `pnpm cache is not found` and does a **full cold install over the network**. When that network is flaky the install dies, and the PR fails for reasons unrelated to its contents. #340 burned **five runs** on this. Against its exact failing head: | check | result | |---|---| | `pnpm install --frozen-lockfile` locally | succeeds | | storybook build locally, this branch | succeeds | | storybook build locally, `main` | succeeds | | storybook job on #480 (branch off same main) | **passed** | | manifests vs main | identical | | regenerating the lockfile | zero-line diff | Only the CI install fails. **Four of those five runs were spent looking in the wrong place** — storybook config, workspace globs, hoisting — because the swallowing retry loop reported `storybook: not found` instead of the install failure. #480 fixed the message; this documents what the message now reveals. **The likely deadlock:** `setup-node` saves its cache only after a *successful* install, so a run that cannot install never writes one, and the next run is cold too. Nothing inside the workflow breaks that cycle. **Where the boundary is:** `gmackie/bob` uses identical `setup-node` + `cache: pnpm` config and restores a 787 MB cache fine on `ubuntu-latest`. Its *deploy* jobs, which run on `forgegraph-worker`, failed the same day with `context canceled`. Same config, different runner, different outcome. Related same-root: bob #93 failed twice on `Request timeout: /astral-sh/uv/releases/download/0.12.6/...` before passing on a third attempt. Docs only.
docs(field-reports): the self-hosted runners never restore the pnpm cache
All checks were successful
CI / gitleaks (pull_request) Successful in 6s
CI / storybook (pull_request) Successful in 1m40s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m18s
8dc5515d05
Every job on forgegraph-ci / forgegraph-worker logs "pnpm cache is not
found" and does a full cold install over the network. When that network is
flaky the install dies, and the PR fails for reasons unrelated to its
contents.

#340 burned five runs on this. Against its exact failing head, the install,
the storybook build, and the same build on main all succeed locally; the
storybook job passed on #480 off the same main; the manifests match main and
regenerating the lockfile is a zero-line diff. Only the CI install fails.

Four of those runs were spent looking at storybook config, workspace globs
and hoisting, because the swallowing retry loop reported "storybook: not
found" instead of the install failure. #480 fixed the message; this documents
what the message now reveals.

The likely deadlock is that setup-node saves its cache only after a
successful install, so a run that cannot install never writes one, and the
next run is cold too. Nothing inside the workflow breaks that cycle.

Boundary worth noting: gmackie/bob uses identical setup-node + cache: pnpm
config and restores a 787 MB cache fine on ubuntu-latest. Its deploy jobs, on
forgegraph-worker, failed the same day with "context canceled". Same config,
different runner.
Author
Owner

Preview environment is live: https://pr-481-forgegraph.forgegraf.com

Deployed 8dc5515d with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-481-forgegraph.forgegraf.com Deployed `8dc5515d` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
Author
Owner

Preview environment is live: https://pr-481-forgegraph.forgegraf.com

Deployed 8dc5515d with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.

Preview environment is live: https://pr-481-forgegraph.forgegraf.com Deployed `8dc5515d` with the beta stage's environment. It redeploys on every push and is destroyed when this PR closes.
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!481
No description provided.