For OpenAI-compatible API-key accounts, /v1/responses requests with
image-generation intent could be scheduled to accounts whose upstream
does not support the Responses API (extra.openai_responses_supported=false).
The flag was only consulted at forward time, where such accounts are
silently downgraded to a Chat-Completions path that cannot produce images,
causing upstream 4xx/5xx or canceled requests.
Fix:
- Add endpoint capability OpenAIEndpointCapabilityResponses. Its check in
SupportsOpenAIEndpointCapability excludes only OpenAI API-key accounts
probed as unsupported (mirroring the forward-time downgrade condition);
OAuth/Grok/unprobed accounts keep existing behavior, and a responses-
capable upstream must still pass the chat_completions gate. Reusing the
existing requiredCapability plumbing makes every scheduler filter path
enforce it with no scheduler signature changes.
- Request the responses capability at the HTTP Responses and
ResponsesWebSocket call sites only when imageIntent && platform==openai,
so non-image requests keep the downgrade path and Grok's own image path
is untouched.
- Normalize max_tokens -> max_output_tokens on the native responses
forward path (PlatformOpenAI), and strip prompt_cache_options alongside
prompt_cache_retention/safety_identifier.
/v1/images/generations continues to use native image capability (unchanged).
Tests: capability truth table, scheduler exclusion of unsupported accounts,
and forward-path transform behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KXpzKKvsb5jW2GvgBQqnnZ
Update the modelConflict / mappingConflict strings (en + zh) to state that
model names are matched case-insensitively, so an existing entry (e.g.
"GLM-5.2") already covers all case variants and the lowercase variant does not
need to be added. This addresses the confusion in #3394 where the rejection of
a case-only duplicate looked like inconsistent behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The "[Billing] Using fallback pricing for model: X" warn was emitted on
every request for any model not in LiteLLM but matched by getFallbackPricing
(e.g. "glm-5.2" substring-matches the "glm-5" fallback price). Via the stdlib
log bridge it is inferred as WARN and persisted to ops_system_logs, producing
tens of thousands of rows/day. It fired even when channel pricing already
overrides the price, and from several call sites (resolver, account-stats
pricing, non-channel billing, admin lookup).
Dedup the warn at the source (BillingService.GetModelPricing) via a sync.Map
keyed by the already-lowercased model name, so each model logs at most once
per process while keeping one audit line. Billing amounts are unchanged.
Also clarify the model/mapping pattern conflict error to state that names are
matched case-insensitively, so an existing entry (e.g. "GLM-5.2") already
covers all case variants and the lowercase variant need not be added — the
behavior the issue mistook for a case-sensitivity bug.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Locks in that Claude Code detection keys on the billing block prefix +
cc_entrypoint=cli, not on the cch field that the new CLI (and now our own
mimicry) no longer sends:
- BillingBlockRecognizedWithoutCCH: an identity-prose-less sub-request whose
system block is `x-anthropic-billing-header: cc_version=...; cc_entrypoint=cli;`
(no cch) is still detected as Claude Code.
- NoCCHBlockStillRequiresClaudeCodeUA: dropping cch did not loosen detection —
a non-claude-cli UA is still rejected, so ClaudeCodeOnly groups can't be spoofed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Recent Claude Code CLI versions no longer emit the cch=... signature field in
their x-anthropic-billing-header system block (issue #3358). sub2api still
injected cch=00000 when mimicking Claude Code for OAuth accounts and optionally
signed it, so mimicked requests now diverge from real CLI traffic — the opposite
of what the mimicry is for.
- buildBillingAttributionText emits the block without the cch=00000 segment;
cc_version + cc_entrypoint=cli are kept (detection and Anthropic's first-party
signal rely on the block, not on cch).
- Retire signing: remove the two enableCCH signBillingHeaderCCH call sites in
buildUpstreamRequest / buildCountTokensRequest and delete the now-dead
signBillingHeaderCCH, cchPlaceholderRe, cchSeed, xxHash64Seeded helpers.
- enable_cch_signing is now a documented no-op (kept for backward compat).
- Drop the obsolete signing tests (TestSignBillingHeaderCCH, TestXXHash64Seeded,
TestSanitizeMustBeBeforeCCHSigning_HashConsistency) and update the prompt test
to assert the injected block no longer carries cch=.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Covers the #3358 fix:
- StripsUnsupportedClaudeCodeTokens reproduces the prod 400 — the four Vertex-
rejected tokens (advisor-tool, prompt-caching-scope, redact-thinking,
thinking-token-count) plus the identity betas are stripped while whitelisted
tokens survive. Fails before the builder fix, passes after.
- DropsHeaderWhenAllUnsupported: no anthropic-beta header is sent when every
client token is filtered out.
- BodySanitizeKeysOnFinalBeta: body.context_management is stripped based on the
final beta, not the raw client value.
- BlocksViaBetaPolicy: an admin block rule on a Vertex account returns BetaBlockedError.
- TestFilterVertexBetaTokens unit-tests whitelist/drop-set/dedupe/empty.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Vertex AI's Anthropic endpoint rejects unknown anthropic-beta tokens with
HTTP 400. buildUpstreamRequestAnthropicVertex forwarded the client header
verbatim via the allowedHeaders whitelist, so recent Claude Code CLIs that
send advisor-tool-2026-03-01, prompt-caching-scope-2026-01-05,
redact-thinking-2026-02-12 and thinking-token-count-2026-05-13 broke every
Vertex service_account request, even though plain account-test requests passed.
This is the only upstream builder that bypassed beta filtering: the
OAuth/API-key path uses computeFinalAnthropicBeta and the Bedrock path uses
filterBedrockBetaTokens. Close the gap with a Vertex-specific whitelist
(vertexSupportedBetaTokens) mirroring bedrockSupportedBetaTokens, plus the
existing BetaPolicy block check:
- evaluateBetaPolicy block check (symmetric to resolveBedrockBetaTokensForRequest)
- filterVertexBetaTokens strips policy-filtered + non-whitelisted tokens
- body context_management sanitize now keys on the final beta, not the raw client value
- overwrite the anthropic-beta header after the whitelist copy loop with the final value
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The OpenAI/Codex 5h "used %" inversion that caused fresh accounts to show
~96-99% used (PR #2918, commit b65dde63) was already reverted in #2993, so the
stored value is now the correct "used %" again. This commit hardens that fix:
1. Regression test locking in direct "used %" semantics. The semantics have
flip-flopped twice (#2918 -> #2993) with no value-level guard — a fresh
account (secondary_used_percent=1, 5h window) must store
codex_5h_used_percent=1, not 99.
2. Stale-bounded self-heal in resolveOpenAIQuotaUtilization (the single
auto-pause chokepoint). An account poisoned with an inflated used% gets
excluded from scheduling, and a paused account never receives traffic to
refresh its snapshot — so it stayed stuck until the window's reset_at passed
(up to 5h/7d). When codex_usage_updated_at is older than 2h, the account is
no longer auto-paused on that snapshot; it gets one request whose response
headers refresh the snapshot and self-heal it. A missing timestamp is treated
as fresh (stays paused), and an actively-served exhausted account refreshes
the timestamp every response so it never crosses the bound — it cannot escape
auto-pause.
No change to Normalize(); no 100-x reintroduced; no new dependency wiring.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Each retry in the SetOutboxWatermark loop now gets its own 5s context.
Previously a shared context could already be expired when the second or
third attempt ran, making the retries pointless.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Fixes#1691
- pollOutbox() reused a 10s context for SetOutboxWatermark after event
processing could take much longer, causing "outbox watermark write
failed: context deadline exceeded". The watermark never advanced so
the same 200 events were reprocessed every poll cycle, spiking CPU.
Now uses an independent 5s context with up to 3 retries (200ms apart).
- When multiple Codex accounts sharing the same 21-22 groups are all
rate-limited in quick succession, each account_changed event triggered
redundant bucket rebuild attempts for the same groups. Introduce
batchSeenKey{groupID, platform} and thread a seen map through the
handler chain; rebuildBucketsForPlatform skips (group, platform) pairs
already rebuilt within the same poll batch (~80% fewer rebuild calls in
the 5-accounts-same-groups scenario).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
When multiple goroutines/workers concurrently refresh the same OAuth token,
the first succeeds but invalidates the old refresh_token (rotation). Subsequent
attempts using the stale token get invalid_grant, which was incorrectly treated
as non-retryable, permanently marking the account as ERROR.
Three complementary fixes:
1. Race-aware recovery: after invalid_grant, re-read DB to check if another
worker already refreshed (refresh_token changed) — return success instead
of error
2. In-process mutex (sync.Map of per-account locks): prevents concurrent
refreshes within the same process, complementing the Redis distributed lock
3. Increase default lock TTL from 30s to 60s to reduce TTL-expiry races
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace charset→base64url double-encoding with standard random
bytes→base64url approach to match official client behavior and avoid
risk control detection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Align OAuth scopes with upstream Claude Code client which now includes
the user:file_upload scope for file upload support.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When all failover accounts are exhausted, handleFailoverExhausted maps
the upstream status code (e.g. 403) to a client-facing code (e.g. 502)
but did not write the original code to the gin context. This caused ops
error logs to show the mapped code instead of the real upstream code.
Call SetOpsUpstreamError before mapUpstreamError in all failover-
exhausted paths so that ops_error_logger captures the true upstream
status code and message.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The `pattern="\d+\.\d+\.\d+"` on the min_claude_code_version input caused
the browser's native HTML5 form validation to silently block form submission
when the value was invalid or when the hidden gateway tab was active. This
resulted in no network request being sent when clicking Save on any tab.
Backend already validates semver format and returns a proper 400 error,
so the frontend pattern attribute is redundant.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add compile-time interface assertion for sessionWindowMockRepo
- Fix flaky fallback test by capturing time.Now() before calling UpdateSessionWindow
- Replace stale hardcoded timestamps with dynamic future values
- Add millisecond detection and bounds validation for reset header timestamp
- Use pause/resume pattern for interval in UsageProgressBar to avoid idle timers on large lists
- Fix gofmt comment alignment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 5h window reset time displayed for Setup Token accounts was inaccurate
because UpdateSessionWindow predicted the window end as "current hour + 5h"
instead of reading the actual `anthropic-ratelimit-unified-5h-reset` response
header. This caused the countdown to differ from the official Claude page.
Backend: parse the reset header (Unix timestamp) and use it as the real
window end, falling back to the hour-truncated prediction only when the
header is absent. Also correct stale predictions when a subsequent request
provides the real reset time.
Frontend: add a reactive 60s timer so the reset countdown in
UsageProgressBar ticks down in real-time instead of freezing at the
initial value.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
OAuth upstreams (ChatGPT) reject requests containing role:"system" in
the input array with HTTP 400 "System messages are not allowed". Extract
such items before forwarding and merge their content into the top-level
instructions field, prepending to any existing value.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add AdminResetQuota service method to reset daily/weekly usage windows
- Add POST /api/v1/admin/subscriptions/:id/reset-quota handler and route
- Add resetQuota API function in frontend subscriptions client
- Add reset quota button, confirmation dialog, and handlers in SubscriptionsView
- Add i18n keys for reset quota feature in zh and en locales
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>