fix(agent): bound postgres metrics payload to fix large-upload stall #23
No reviewers
Labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
gmackie/ForgeGraph!23
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/pg-metrics-lean-payload"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Root cause (finally) of master's metrics timeout. Captured via
GODEBUG=http2debugon the live agent: the metrics POST body was ~95KB (28 DBs x table stats + full slow-query text). Go's HTTP client stalls mid-upload on a body that large to the Cloudflare edge — the trace shows it sending the body, receiving WINDOW_UPDATE frames, then a 15s timeout. A ~4KB db-level-only payload always succeeds (and curl handles 80KB fine), so it's a Go large-body-upload interaction, triggered purely by payload size. Earlier fixes (h2 pings, fresh conn, backoff in #22) didn't address the size, so the hang persisted.Fix: keep lightweight db-level metrics (size/connections/cache — what
GET /api/fg/databasesand the view use) for every DB each cycle, but collect the heavy per-table + slow-query detail for only a rotating 2 DBs per cycle; cap slow queries 10->5 and truncate query text to 1000 chars. Every payload now stays ~10KB, well under the h2 flow-control window; all DBs' detail still refreshes within a few minutes, and the per-app postgres-metrics UI keeps its data. nildetailDBs= full detail (on-demand CLI path).Verified on master (0.1.15): 0 report errors + 0 collection errors over 175s (was every cycle); clean 60s cadence (66 rows / 3 min = 2 cycles x 33 DBs, no retry churn);
GET /api/fg/databasesshows all 33 DBs live. New tests for detail-gating + truncate; full suite green.🤖 Generated with Claude Code