feat(alchemy): host Alchemy's state store so correlation is a fact, not a join #571
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!571
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/alchemy-state-store"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Decision A: adopting Alchemy
This is the ForgeGraph side, and it is server only — it changes no deploy path, and nothing calls it until an app opts in. Branches from
main, independent of the Phase 1/2/3a work.Plan: https://7n04n7hhuesf.postplan.dev — Phase 3b.1 + 3b.2.
Why host the store rather than let Alchemy keep state elsewhere: the correlation between a semantic node and the Cloudflare resource behind it becomes something ForgeGraph already holds, instead of a join it reconstructs by scraping. That is the entire argument for this design.
The document is stored verbatim
Alchemy owns that format, it is a beta dependency, and its own API declares the payload free-form. Storing it as-is means an Alchemy change cannot corrupt what we hold. The correlation columns beside it are derived on every write and are an index over the document, never the source of truth — which is why every one is nullable, and why
correlatereturns nulls instead of throwing on a document it does not understand. It runs during a deploy; refusing to record state we cannot fully parse would break the deploy it is only observing.Rows are workspace-scoped. Alchemy's
stackis its own namespace and carries no tenancy, so the bearer token is the only thing separating two workspaces' stacks — the workspace is part of the unique key.The contract was read, not guessed
I took it off Alchemy's own client and reference server. Two details would have been wrong otherwise, and both fail in ways that look like something else:
null, not 404s == nullto "absent"; a 404 reads as a transport failure and aborts the deploy/-joined namespace path, so an unencoded lookup silently missesTwo more, chosen deliberately: a write echoes the document the client sent rather than our stored copy, so a storage bug cannot look like a successful write; and
/versionis unauthenticated, because the client probes it to decide whether a store is too old to talk to and must be able to learn that without a valid token.A bug this found
The handlers originally discriminated auth failures with
instanceof NextResponse. Any otherResponsesubclass falls through to the success path — an auth bypass, not a type error. The test caught it because it mocks a plainResponse. Replaced with an explicit tagged result.Verification
24 tests. Compatibility checked before writing anything: alchemy
2.0.0-beta.77peers oneffect >=4.0.0-beta.105, which the template's4.0.0-rc.112satisfies (rc sorts above beta).⚠️ Before merging
Apply
0103_alchemy_state.sqlthen0103a_alchemy_state_grants.sqlto the live database, or prod deploys go red on schema drift.Next, deliberately not here
A live contract test running Alchemy's real client against these routes, before any app is pointed at them. The unit tests pin the contract as I read it; only the real client proves I read it right.
🤖 Generated with Claude Code
Adopting Alchemy (Decision A). This is the ForgeGraph side and it is server only — it changes no deploy path and nothing calls it until an app opts in. Hosting the state store rather than letting Alchemy keep state elsewhere is the whole point: the correlation between a semantic node and the Cloudflare resource behind it becomes something ForgeGraph already holds, instead of a join it has to reconstruct by scraping. - packages/db: `alchemy_state` stores Alchemy's document VERBATIM. Alchemy owns that format, it is a beta dependency, and its own API declares the payload free-form, so an Alchemy change cannot corrupt what we hold. The correlation columns beside it (resourceType, logicalId, instanceId, physicalId, physicalName, status) are derived on every write and are an index over the document, never the source of truth — which is why every one of them is nullable. - Rows are workspace-scoped. Alchemy's `stack` is its own namespace and carries no tenancy, so the bearer token is the only thing keeping two workspaces' stacks apart; the workspace is part of the unique key. - packages/api: `correlate` never throws on a document it does not understand. It runs on every state write during a deploy, and refusing to record state we cannot fully parse would break the deploy it is only observing. - apps/web: the eleven endpoints Alchemy's client calls, with the handlers in one module so the wire contract cannot drift between route files. The contract was read off Alchemy's own client and reference server rather than guessed, and two details would have been wrong otherwise: - a missing resource is `null`, NOT 404 — the client maps `s == null` to "absent" and reads a 404 as a transport failure; - the fqn arrives percent-encoded, because it is a `/`-joined namespace path that cannot ride a URL segment raw. Also: a write echoes the document the client sent rather than our stored copy, so a storage bug cannot look like a successful write; `/version` is unauthenticated because the client probes it to decide whether a store is too old to talk to, and must be able to learn that without a valid token. Found and fixed while testing: the handlers discriminated auth failures with `instanceof NextResponse`. Any other Response subclass falls through to the success path — an auth bypass, not a type error. Replaced with a tagged result. 24 tests. Verified compatible: alchemy 2.0.0-beta.77 peers on `effect >=4.0.0-beta.105`, which the template's 4.0.0-rc.112 satisfies. ⚠️ Apply 0103_alchemy_state.sql then 0103a_alchemy_state_grants.sql BEFORE merge, or prod deploys go red on schema drift. Next, and deliberately not in this PR: a live contract test running Alchemy's real client against these routes, before any app is pointed at them. Plan: https://7n04n7hhuesf.postplan.dev (Phase 3b.1 + 3b.2) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>Ran Alchemy's real client against these handlers
The tests here asserted my reading of a client we don't control. I ran the actual one —
alchemy@2.0.0-beta.77'smakeHttpStateStore— against these actual handlers, on a real Next server. 20 assertions, all passing, covering every endpoint plus fqn round-trips.It confirmed most of the contract and corrected one thing I had wrong, which is now the reason for the extra commit.
The fqn is encoded twice, not once
I was about to "fix"
decodeFqnas a double-decode bug. Next's route matcher decodes every dynamic param (route-matcher.js:19), so decoding again in the handler looks plainly redundant. It isn't:/-joined namespace path (HttpStateStore.ts).HttpApiClientthen percent-encodes every path param again while interpolating the route (compilePath,HttpApiClient.ts:724).Two encodes, one decode by the runtime, one by us. Correct as written.
Had I removed it, any fqn containing a percent-escape (
My%20Thing) would resolve under the wrong key — silently. Worse:listResourceswould then report the mangled name, and Alchemy's orphan sweep would conclude the real resource is gone and destroy it. A live D1 or bucket, for a one-line "cleanup".The new tests model all four steps of the chain rather than just our end of it, over the fqns where the layers can disagree unnoticed:
Root/Child/MainDb,My%20Thing,a%b,sp ace,uniçøde,q?=&#hash.Run on workerd too, not just
next devProduction is Workers via OpenNext, and param decoding is exactly the behaviour that can differ between runtimes. I built the harness with
opennextjs-cloudflareand re-ran the same 20 assertions underwrangler dev --local. Both runtimes agree.Confirmed unchanged
200+null(the client mapss == nullto absent; a 404 gets retried 5× then fails the deploy)204, empty bodylistResources→ array of fqn strings;replaced-resources→ array of documents/versionunauthenticated, returns5deleteStackpassesstagethrough only when narrowingOne incidental correction to a note in the handler docblock: on a write the client ignores our response body entirely (
Effect.map(() => request.value)) — it keeps its own canonical object, Redacted values and all. Echoing is still right, just not load-bearing.docs/alchemy-state-store.mdrecords the contract, why each response shape is what it is, and how to re-run the live test when the pinned alchemy version moves.Migration
0103+0103astill need applying to prod before this merges.🤖 Generated with Claude Code