fix(agent): h2 keep-alive pings + metrics retry backoff #22

Merged
gmackie merged 1 commit from fix/agent-metrics-conn into main 2026-06-25 23:26:15 +00:00
Owner

master's postgres-metrics POST intermittently hung until the agent's 15s client timeout, so it retried every heartbeat cycle (log spam + extra psql load).

  • h2 health-check pings (ReadIdleTimeout/PingTimeout via http2.ConfigureTransports): a silently-dropped pooled HTTP/2 connection to the edge is now detected and replaced instead of hanging requests until the client timeout.
  • Retry backoff: advance the metrics throttle on every attempt, so a persistent failure retries at the 60s cadence instead of re-collecting + re-POSTing every cycle (verified: errors now 60s apart, not 30s).
  • Factor the IPv4-pinning dialer into a helper.

Honest caveat: a deeper intermittent Go-HTTP hang on master is not fully resolved — curl -4 with the identical payload is 20/20 fast, yet the agent still occasionally times out (ruled out: connection pooling, h2 reuse, IP family, payload size, proxy env, network). It needs runtime-level instrumentation (GODEBUG/tcpdump) to root-cause. Metrics still flow (server completes the insert; GET /api/fg/databases shows live data for all 33 DBs); this PR halves the churn and improves connection robustness.

🤖 Generated with Claude Code

master's postgres-metrics POST intermittently hung until the agent's 15s client timeout, so it retried every heartbeat cycle (log spam + extra psql load). - **h2 health-check pings** (`ReadIdleTimeout`/`PingTimeout` via `http2.ConfigureTransports`): a silently-dropped pooled HTTP/2 connection to the edge is now detected and replaced instead of hanging requests until the client timeout. - **Retry backoff**: advance the metrics throttle on *every* attempt, so a persistent failure retries at the 60s cadence instead of re-collecting + re-POSTing every cycle (verified: errors now 60s apart, not 30s). - Factor the IPv4-pinning dialer into a helper. **Honest caveat:** a deeper *intermittent* Go-HTTP hang on master is **not** fully resolved — curl -4 with the identical payload is 20/20 fast, yet the agent still occasionally times out (ruled out: connection pooling, h2 reuse, IP family, payload size, proxy env, network). It needs runtime-level instrumentation (GODEBUG/tcpdump) to root-cause. Metrics still flow (server completes the insert; `GET /api/fg/databases` shows live data for all 33 DBs); this PR halves the churn and improves connection robustness. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
gmackie force-pushed fix/agent-metrics-conn from 67453402de
Some checks failed
AI Code Review / review (pull_request) Failing after 8s
CI / ci (pull_request) Has been cancelled
to ad18587645
Some checks failed
AI Code Review / review (pull_request) Failing after 35s
CI / ci (pull_request) Successful in 16m54s
2026-06-25 23:05:27 +00:00
Compare
gmackie changed title from fix(agent): fresh connection + backoff for postgres metrics POST to fix(agent): h2 keep-alive pings + metrics retry backoff 2026-06-25 23:05:46 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
gmackie/ForgeGraph!22
No description provided.