DEV_GUIDE.md and the three READMEs still advertise Go 1.25.7 and
golangci-lint v2.7, but CI has since moved on:
- backend/go.mod declares go 1.26.5, and backend-ci.yml / release.yml /
security-scan.yml all resolve the toolchain via
`go-version-file: backend/go.mod` and then hard-assert
`go version | grep -q 'go1.26.5'`.
- backend-ci.yml pins golangci-lint to v2.9.
So a contributor following DEV_GUIDE.md installs a linter two minor
versions behind CI (different findings locally vs. in CI) and expects a
Go version that the workflow's own assertion step rejects.
Update all ten stale references, and note in the CI section which files
the Go version assertion lives in so future bumps don't miss one.
Docs only, no code or workflow changes.
glm-5.2 has no entry in the fallback table, so getFallbackPricing falls
through to `strings.Contains(modelLower, "glm-5")` and prices it at
GLM-5 rates ($1.00 in / $3.20 out per MTok) instead of the official
z.ai rates ($1.40 / $4.40) — roughly 27% under.
LiteLLM carries no bare `glm-5.2` key either (only provider-prefixed
`cloudflare/@cf/zai-org/glm-5.2` and `fireworks_ai/.../glm-5p2`, which
the lookup candidates never match), so the request always lands on the
fallback path and the discrepancy shows up directly in usage logs.
Add the glm-5.2 entry (same price as glm-5.1 per docs.z.ai) and match it
before the bare `glm-5` branch, with a note that dotted variants must
precede it. The existing regression test asserting the old glm-5 price
is updated accordingly.
Source: https://docs.z.ai/guides/overview/pricing
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codex/OpenAI upstream switched from a WebSocket pool to HTTP/2. The
outbound H2 transport set neither ReadIdleTimeout nor ResponseHeaderTimeout
(the OpenAI profile forces ResponseHeaderTimeout=0), so a pooled H2
connection silently killed by a proxy/NAT becomes a "dead connection":
both ends believe it is alive and a request assigned to it hangs until the
OS TCP retransmit timeout (minutes) before the first byte — observed as an
8m37s TTFT with an eventual 200. Occasional (only when a request lands on a
dead pooled conn) and across all groups (shared OpenAI transport); worse on
larger idle pools.
Explicitly configure http2 on the openai_h2 transport and enable active
PING health checks (ReadIdleTimeout=15s, PingTimeout=15s) so dead
connections are detected and evicted at the source, instead of relying on
ResponseHeaderTimeout as an after-the-fact backstop. Scoped to the
openai_h2 path only; Claude/Gemini (default) and h1 modes are untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>