feat(alchemy): host Alchemy's state store so correlation is a fact, not a join #571

Merged
gmackie merged 3 commits from feat/alchemy-state-store into main 2026-09-12 02:48:44 +00:00
Owner

Decision A: adopting Alchemy

This is the ForgeGraph side, and it is server only — it changes no deploy path, and nothing calls it until an app opts in. Branches from main, independent of the Phase 1/2/3a work.

Plan: https://7n04n7hhuesf.postplan.dev — Phase 3b.1 + 3b.2.

Why host the store rather than let Alchemy keep state elsewhere: the correlation between a semantic node and the Cloudflare resource behind it becomes something ForgeGraph already holds, instead of a join it reconstructs by scraping. That is the entire argument for this design.

The document is stored verbatim

Alchemy owns that format, it is a beta dependency, and its own API declares the payload free-form. Storing it as-is means an Alchemy change cannot corrupt what we hold. The correlation columns beside it are derived on every write and are an index over the document, never the source of truth — which is why every one is nullable, and why correlate returns nulls instead of throwing on a document it does not understand. It runs during a deploy; refusing to record state we cannot fully parse would break the deploy it is only observing.

Rows are workspace-scoped. Alchemy's stack is its own namespace and carries no tenancy, so the bearer token is the only thing separating two workspaces' stacks — the workspace is part of the unique key.

The contract was read, not guessed

I took it off Alchemy's own client and reference server. Two details would have been wrong otherwise, and both fail in ways that look like something else:

detail what happens if you guess
a missing resource is null, not 404 the client maps s == null to "absent"; a 404 reads as a transport failure and aborts the deploy
the fqn arrives percent-encoded it is a /-joined namespace path, so an unencoded lookup silently misses

Two more, chosen deliberately: a write echoes the document the client sent rather than our stored copy, so a storage bug cannot look like a successful write; and /version is unauthenticated, because the client probes it to decide whether a store is too old to talk to and must be able to learn that without a valid token.

A bug this found

The handlers originally discriminated auth failures with instanceof NextResponse. Any other Response subclass falls through to the success path — an auth bypass, not a type error. The test caught it because it mocks a plain Response. Replaced with an explicit tagged result.

Verification

24 tests. Compatibility checked before writing anything: alchemy 2.0.0-beta.77 peers on effect >=4.0.0-beta.105, which the template's 4.0.0-rc.112 satisfies (rc sorts above beta).

⚠️ Before merging

Apply 0103_alchemy_state.sql then 0103a_alchemy_state_grants.sql to the live database, or prod deploys go red on schema drift.

Next, deliberately not here

A live contract test running Alchemy's real client against these routes, before any app is pointed at them. The unit tests pin the contract as I read it; only the real client proves I read it right.

🤖 Generated with Claude Code

## Decision A: adopting Alchemy This is the ForgeGraph side, and it is **server only** — it changes no deploy path, and nothing calls it until an app opts in. Branches from `main`, independent of the Phase 1/2/3a work. Plan: https://7n04n7hhuesf.postplan.dev — Phase 3b.1 + 3b.2. **Why host the store** rather than let Alchemy keep state elsewhere: the correlation between a semantic node and the Cloudflare resource behind it becomes something ForgeGraph *already holds*, instead of a join it reconstructs by scraping. That is the entire argument for this design. ## The document is stored verbatim Alchemy owns that format, it is a beta dependency, and its own API declares the payload free-form. Storing it as-is means an Alchemy change cannot corrupt what we hold. The correlation columns beside it are derived on every write and are an *index over* the document, never the source of truth — which is why every one is nullable, and why `correlate` returns nulls instead of throwing on a document it does not understand. It runs during a deploy; refusing to record state we cannot fully parse would break the deploy it is only observing. Rows are workspace-scoped. Alchemy's `stack` is its own namespace and carries no tenancy, so the bearer token is the only thing separating two workspaces' stacks — the workspace is part of the unique key. ## The contract was read, not guessed I took it off Alchemy's own client and reference server. Two details would have been wrong otherwise, and both fail in ways that look like something else: | detail | what happens if you guess | |---|---| | a missing resource is **`null`, not 404** | the client maps `s == null` to "absent"; a 404 reads as a transport failure and aborts the deploy | | the fqn arrives **percent-encoded** | it is a `/`-joined namespace path, so an unencoded lookup silently misses | Two more, chosen deliberately: a write echoes the document the client sent rather than our stored copy, so a storage bug cannot look like a successful write; and `/version` is unauthenticated, because the client probes it to decide whether a store is too old to talk to and must be able to learn that without a valid token. ## A bug this found The handlers originally discriminated auth failures with `instanceof NextResponse`. Any other `Response` subclass falls through to the **success path** — an auth bypass, not a type error. The test caught it because it mocks a plain `Response`. Replaced with an explicit tagged result. ## Verification 24 tests. Compatibility checked before writing anything: alchemy `2.0.0-beta.77` peers on `effect >=4.0.0-beta.105`, which the template's `4.0.0-rc.112` satisfies (rc sorts above beta). ## ⚠️ Before merging Apply `0103_alchemy_state.sql` then `0103a_alchemy_state_grants.sql` to the live database, or prod deploys go red on schema drift. ## Next, deliberately not here A live contract test running Alchemy's **real client** against these routes, before any app is pointed at them. The unit tests pin the contract as I read it; only the real client proves I read it right. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(alchemy): host Alchemy's state store so correlation is a fact, not a join
All checks were successful
CI / gitleaks (pull_request) Successful in 8s
CI / storybook (pull_request) Successful in 2m22s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 10m8s
fbf78b04e2
Adopting Alchemy (Decision A). This is the ForgeGraph side and it is server
only — it changes no deploy path and nothing calls it until an app opts in.

Hosting the state store rather than letting Alchemy keep state elsewhere is
the whole point: the correlation between a semantic node and the Cloudflare
resource behind it becomes something ForgeGraph already holds, instead of a
join it has to reconstruct by scraping.

- packages/db: `alchemy_state` stores Alchemy's document VERBATIM. Alchemy
  owns that format, it is a beta dependency, and its own API declares the
  payload free-form, so an Alchemy change cannot corrupt what we hold. The
  correlation columns beside it (resourceType, logicalId, instanceId,
  physicalId, physicalName, status) are derived on every write and are an
  index over the document, never the source of truth — which is why every one
  of them is nullable.
- Rows are workspace-scoped. Alchemy's `stack` is its own namespace and
  carries no tenancy, so the bearer token is the only thing keeping two
  workspaces' stacks apart; the workspace is part of the unique key.
- packages/api: `correlate` never throws on a document it does not understand.
  It runs on every state write during a deploy, and refusing to record state
  we cannot fully parse would break the deploy it is only observing.
- apps/web: the eleven endpoints Alchemy's client calls, with the handlers in
  one module so the wire contract cannot drift between route files.

The contract was read off Alchemy's own client and reference server rather
than guessed, and two details would have been wrong otherwise:

  - a missing resource is `null`, NOT 404 — the client maps `s == null` to
    "absent" and reads a 404 as a transport failure;
  - the fqn arrives percent-encoded, because it is a `/`-joined namespace
    path that cannot ride a URL segment raw.

Also: a write echoes the document the client sent rather than our stored copy,
so a storage bug cannot look like a successful write; `/version` is
unauthenticated because the client probes it to decide whether a store is too
old to talk to, and must be able to learn that without a valid token.

Found and fixed while testing: the handlers discriminated auth failures with
`instanceof NextResponse`. Any other Response subclass falls through to the
success path — an auth bypass, not a type error. Replaced with a tagged
result.

24 tests. Verified compatible: alchemy 2.0.0-beta.77 peers on
`effect >=4.0.0-beta.105`, which the template's 4.0.0-rc.112 satisfies.

⚠️ Apply 0103_alchemy_state.sql then 0103a_alchemy_state_grants.sql BEFORE
merge, or prod deploys go red on schema drift.

Next, and deliberately not in this PR: a live contract test running Alchemy's
real client against these routes, before any app is pointed at them.

Plan: https://7n04n7hhuesf.postplan.dev (Phase 3b.1 + 3b.2)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test(alchemy): pin the fqn encoding chain against Alchemy's real client
All checks were successful
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 1m26s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 9m6s
32dc407141
The state-store handlers were written against a reading of Alchemy's
client. Running the real client against them turned one of those
readings into a fact and corrected another.

The fqn is percent-encoded TWICE before it reaches us: Alchemy's client
encodes it by hand (it is a "/"-joined namespace path), and Effect's
HttpApiClient then encodes every path param again in compilePath. Next's
route matcher decodes exactly one layer. So decodeFqn decoding once is
correct — but looked at from the handler alone it reads like a
double-decode bug, and "fixing" it silently corrupts any fqn containing
a percent-escape: the resource resolves under the wrong key, listResources
reports the mangled name, and Alchemy's orphan sweep concludes the real
resource is gone. I nearly made that change before checking compilePath.

The new tests model all four steps of the chain rather than just our end
of it, over the fqns where the layers can disagree unnoticed.

Verified end to end with alchemy 2.0.0-beta.77's actual client driving
these actual handlers — 20 assertions, on next dev AND on workerd via
OpenNext, since production is Workers and param decoding is exactly the
behaviour that can differ. Both runtimes agree. docs/alchemy-state-store.md
records the contract, why each response shape is what it is, and how to
re-run the live test when the pinned alchemy version moves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Author
Owner

Ran Alchemy's real client against these handlers

The tests here asserted my reading of a client we don't control. I ran the actual one — alchemy@2.0.0-beta.77's makeHttpStateStore — against these actual handlers, on a real Next server. 20 assertions, all passing, covering every endpoint plus fqn round-trips.

It confirmed most of the contract and corrected one thing I had wrong, which is now the reason for the extra commit.

The fqn is encoded twice, not once

I was about to "fix" decodeFqn as a double-decode bug. Next's route matcher decodes every dynamic param (route-matcher.js:19), so decoding again in the handler looks plainly redundant. It isn't:

  1. Alchemy's client percent-encodes the fqn by hand — it's a /-joined namespace path (HttpStateStore.ts).
  2. Effect's HttpApiClient then percent-encodes every path param again while interpolating the route (compilePath, HttpApiClient.ts:724).
  3. Next decodes exactly one layer.

Two encodes, one decode by the runtime, one by us. Correct as written.

Had I removed it, any fqn containing a percent-escape (My%20Thing) would resolve under the wrong key — silently. Worse: listResources would then report the mangled name, and Alchemy's orphan sweep would conclude the real resource is gone and destroy it. A live D1 or bucket, for a one-line "cleanup".

The new tests model all four steps of the chain rather than just our end of it, over the fqns where the layers can disagree unnoticed: Root/Child/MainDb, My%20Thing, a%b, sp ace, uniçøde, q?=&#hash.

Run on workerd too, not just next dev

Production is Workers via OpenNext, and param decoding is exactly the behaviour that can differ between runtimes. I built the harness with opennextjs-cloudflare and re-ran the same 20 assertions under wrangler dev --local. Both runtimes agree.

Confirmed unchanged

  • missing resource → 200 + null (the client maps s == null to absent; a 404 gets retried 5× then fails the deploy)
  • delete → 204, empty body
  • listResources → array of fqn strings; replaced-resources → array of documents
  • /version unauthenticated, returns 5
  • deleteStack passes stage through only when narrowing

One incidental correction to a note in the handler docblock: on a write the client ignores our response body entirely (Effect.map(() => request.value)) — it keeps its own canonical object, Redacted values and all. Echoing is still right, just not load-bearing.

docs/alchemy-state-store.md records the contract, why each response shape is what it is, and how to re-run the live test when the pinned alchemy version moves.

Migration 0103 + 0103a still need applying to prod before this merges.

🤖 Generated with Claude Code

## Ran Alchemy's real client against these handlers The tests here asserted my *reading* of a client we don't control. I ran the actual one — `alchemy@2.0.0-beta.77`'s `makeHttpStateStore` — against these actual handlers, on a real Next server. 20 assertions, all passing, covering every endpoint plus fqn round-trips. It confirmed most of the contract and **corrected one thing I had wrong**, which is now the reason for the extra commit. ### The fqn is encoded twice, not once I was about to "fix" `decodeFqn` as a double-decode bug. Next's route matcher decodes every dynamic param (`route-matcher.js:19`), so decoding again in the handler looks plainly redundant. It isn't: 1. Alchemy's client percent-encodes the fqn **by hand** — it's a `/`-joined namespace path (`HttpStateStore.ts`). 2. Effect's `HttpApiClient` then percent-encodes **every path param again** while interpolating the route (`compilePath`, `HttpApiClient.ts:724`). 3. Next decodes exactly one layer. Two encodes, one decode by the runtime, one by us. Correct as written. Had I removed it, any fqn containing a percent-escape (`My%20Thing`) would resolve under the wrong key — silently. Worse: `listResources` would then report the mangled name, and Alchemy's orphan sweep would conclude the real resource is gone and destroy it. A live D1 or bucket, for a one-line "cleanup". The new tests model all four steps of the chain rather than just our end of it, over the fqns where the layers can disagree unnoticed: `Root/Child/MainDb`, `My%20Thing`, `a%b`, `sp ace`, `uniçøde`, `q?=&#hash`. ### Run on workerd too, not just `next dev` Production is Workers via OpenNext, and param decoding is exactly the behaviour that can differ between runtimes. I built the harness with `opennextjs-cloudflare` and re-ran the same 20 assertions under `wrangler dev --local`. Both runtimes agree. ### Confirmed unchanged - missing resource → `200` + `null` (the client maps `s == null` to absent; a 404 gets retried 5× then fails the deploy) - delete → `204`, empty body - `listResources` → array of **fqn strings**; `replaced-resources` → array of **documents** - `/version` unauthenticated, returns `5` - `deleteStack` passes `stage` through only when narrowing One incidental correction to a note in the handler docblock: on a write the client **ignores our response body** entirely (`Effect.map(() => request.value)`) — it keeps its own canonical object, Redacted values and all. Echoing is still right, just not load-bearing. `docs/alchemy-state-store.md` records the contract, why each response shape is what it is, and how to re-run the live test when the pinned alchemy version moves. Migration `0103` + `0103a` still need applying to prod before this merges. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
merge: bring main into feat/alchemy-state-store
All checks were successful
CI / gitleaks (pull_request) Successful in 7s
CI / storybook (pull_request) Successful in 1m50s
forgegraph/ci CI passed
CI / ci (pull_request) Successful in 11m9s
d113ad8300
#567 landed first and touched the same two registration points, so this
merges main in rather than rebasing — the branch is already pushed and a
force-push wedges Forgejo's merge status in "checking".

Both conflicts were additive and resolved as the union of the two sides:
packages/db/src/schema/index.ts re-exports alchemy-state alongside the
two contract tables, and packages/api's exports map keeps ./lib/
contract-ingest next to the alchemy entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gmackie scheduled this pull request to auto merge when all checks succeed 2026-09-12 02:12:35 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!571
No description provided.