docs(field-reports): withdraw the "cold cache blocks PRs" claim #500

Merged
gmackie merged 1 commit from docs/correct-cold-cache-report into main 2026-08-27 20:59:00 +00:00
Owner

Corrects the report merged this morning as #481.

It used PR #340's five failures as its worked example. That attribution was wrong. #340 failed for a reason of its own — it adds @preflight:registry=https://npm.forgegraf.com/ to .npmrc, that registry answers 401 unauthenticated, and ci.yml configured no token for it. After adding the token it went green on the first run, with no change to the runners.

What survives

  • The runners really do log pnpm cache is not found every run, while gmackie/bob restores a 787 MB cache fine on ubuntu-latest with identical config.
  • bob #93 really did fail twice on Request timeout: /astral-sh/uv/releases/download/... — a public GitHub release, no auth involved, so a genuine egress drop.

Both are worth fixing on cost and reliability grounds.

What does not

Any demonstrated case of the cold cache blocking a PR. The one example offered turned out to be a credential problem wearing a network problem's clothes. Title amended to match.

Why the wrong reasoning is left visible

Because how it misled is the useful part. All of these were true and all pointed the wrong way:

  • pnpm install --frozen-lockfile succeeded locally on the exact failing head
  • the storybook build succeeded locally on both the branch and main
  • the same job passed on a different PR off the same main
  • manifests matched main; regenerating the lockfile was a zero-line diff

A warm pnpm store plus a personal ~/.npmrc hides a missing CI credential perfectly. The report now says to diff .npmrc for new scoped registries before blaming the runner.

#480 is unaffected — the retry loop really did report storybook: not found instead of the install failure, which is why four of the five runs were spent looking in the wrong place.

Corrects the report merged this morning as #481. It used PR #340's five failures as its worked example. **That attribution was wrong.** #340 failed for a reason of its own — it adds `@preflight:registry=https://npm.forgegraf.com/` to `.npmrc`, that registry answers **401** unauthenticated, and `ci.yml` configured no token for it. After adding the token it went green on the first run, with no change to the runners. ## What survives - The runners really do log `pnpm cache is not found` every run, while `gmackie/bob` restores a 787 MB cache fine on `ubuntu-latest` with identical config. - bob #93 really did fail twice on `Request timeout: /astral-sh/uv/releases/download/...` — a **public** GitHub release, no auth involved, so a genuine egress drop. Both are worth fixing on cost and reliability grounds. ## What does not Any demonstrated case of the cold cache *blocking a PR*. The one example offered turned out to be a credential problem wearing a network problem's clothes. Title amended to match. ## Why the wrong reasoning is left visible Because how it misled is the useful part. All of these were true and all pointed the wrong way: - `pnpm install --frozen-lockfile` succeeded locally on the exact failing head - the storybook build succeeded locally on both the branch and `main` - the same job passed on a different PR off the same `main` - manifests matched `main`; regenerating the lockfile was a zero-line diff A warm pnpm store plus a personal `~/.npmrc` hides a missing CI credential perfectly. The report now says to diff `.npmrc` for new scoped registries before blaming the runner. #480 is unaffected — the retry loop really did report `storybook: not found` instead of the install failure, which is why four of the five runs were spent looking in the wrong place.
docs(field-reports): withdraw the "cold cache blocks PRs" claim
All checks were successful
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 1m40s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m39s
75f4c6f668
The report I merged this morning as #481 used PR #340's five failures as its
worked example. That attribution was wrong.

#340 failed for a reason of its own: it adds
@preflight:registry=https://npm.forgegraf.com/ to .npmrc, that registry answers
401 unauthenticated, and ci.yml configured no token for it. After adding the
token it went green on the first run, with no change to the runners.

What survives, and stays in the report: the runners really do log "pnpm cache
is not found" every run while bob restores a 787 MB cache fine on
ubuntu-latest, and bob #93 really did fail twice on a public GitHub release
download — no auth involved, so a real egress drop. Both are worth fixing.

What does not survive: any demonstrated case of the cold cache blocking a PR.
The one example offered turned out to be a credential problem wearing a network
problem's clothes. Title amended to match.

Kept the wrong reasoning visible rather than deleting it, because the way it
misled is the useful part: install succeeds locally, build succeeds locally,
the same job passes on another PR, manifests match, lockfile regenerates to a
zero-line diff — all true, all pointing the wrong way. A warm pnpm store plus a
personal ~/.npmrc hides a missing CI credential perfectly.

#480 is unaffected: the retry loop really did report "storybook: not found"
instead of the install failure, and that is why four of the five runs were
spent looking in the wrong place.
gmackie deleted branch docs/correct-cold-cache-report 2026-08-27 20:59:01 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!500
No description provided.