- Responses WebSocket composite whitelist stays openai+grok: CN accounts
cannot pass the WSv2 ingress transport filter and the WS HTTP bridge
has no Responses conversion for them, so admitting CN targets only
turns a clear policy rejection into a misleading no-available-account
- widen the responses/input_tokens and messages/count_tokens composite
gates to the shared openai-compatible text whitelist so CN targets get
the same local-estimate token counting as generation
- restore the Claude default-model fallback for standalone CN groups
(admin candidates and custom models list) while keeping composite
listings limited to CN account mapping keys
- refresh the scheduler bulk-rebuild comment and bucket capacity hints
for the 8-platform canonical set
Review fixes for the profit-control feature commit:
- Image intent no longer disables the profit gate. The shared /v1/responses
handler previously skipped the pricing context (and therefore the gate)
whenever the platform-wide image intent predicate matched, which includes
Codex's passive image_gen namespace declaration: any client could disable
admission control for anthropic/gemini/antigravity groups by declaring a
namespace tool in the request body. Both /v1/responses paths now always
install the token pricing context; image intent only drives capability
routing and image billing. Mixed token+image requests stay token-gated;
only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
ungated: Grok media (billed by media multipliers; also prevents in-flight
video lookups from turning into spurious 404s), OpenAI-group count_tokens
(unbilled), and Live calls (duration billed) carry a suppress marker that
every install point honors, including the defensive scheduler-entry
install.
- The gate resolved during selection now travels back to handlers on the
AccountSelectionResult. The shared gateway installed the gate only on a
scheduler-local context, so handler-side post-slot terminal rechecks and
post-admission sticky binding were no-ops for anthropic/gemini/
antigravity/shared-grok requests (and for composite-routed member groups
on the OpenAI path). Handlers re-apply the carried gate via
ContextWithSelectionProfitGate before the terminal recheck and binding;
the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
BindStickySessionAfterProfitAdmission falls back to the official eager
bind when no gate is installed (wait paths lost their only binding point
otherwise), reads the pre-existing binding at bind time only when gated
(removes the unconditional per-request Redis read the feature added to
the shared handlers), and the legacy engine's three selection-time
binding writes are skipped under a gate so a terminally vetoed account
can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
every turn (BeforeTurn) and bill each turn with its own instant, closing
the connect-at-valley/bill-at-valley window; a turn that fails the
recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
staler snapshot object (UpdatedAt guard), and observer counters are
documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
landed upstream as 191.
New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Fixes#5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.
Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.
Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
into a ForwardResult and return it alongside the error, for both the
regular Anthropic path and the API-key passthrough path. Invariants:
UpstreamFailoverError always keeps result=nil (failover retries are
billed as the successful attempt, never twice), and zero observed
usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
(dropped_stopped) from operator-configured drop/sample overflow
drops; billing tasks now fall back to inline synchronous execution
only during the shutdown window, while explicit drop/sample overflow
semantics are preserved. Image usage keeps its mandatory fallback
for both drop kinds via the new mode.Dropped() helper.
Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
Expose the active weekly subscription window through /v1/usage, calculate offsets with the same normalized page size used by queries, and keep user-facing date ranges on the browser's local calendar date.
Constraint: Preserve existing response fields and avoid new dependencies
Rejected: Keep duplicate inline date formatters | a shared local-date utility prevents the same UTC regression in both views
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep Offset and Limit based on the same normalized page size
Tested: go test ./internal/pkg/pagination ./internal/handler; go vet ./internal/pkg/pagination ./internal/handler; frontend 923 tests; pnpm typecheck; pnpm lint:check; pnpm build
Not-tested: Live API request against a deployed subscription
Related: #4121
Move anthropic-beta header filtering from separate FilterBedrockBetaHeader
into ApplyBedrockCCCompat, so one function handles all CC compat processing
(body cleanup + header filtering). Change signature from ctx to *gin.Context
to access request headers. Remove the redundant separate call in handler.
Service layer writes a complete JSON error response then returns error.
Handler's ensureForwardErrorResponse couldn't distinguish this from
"no response written" and appended an SSE event, corrupting the body.
Use gin.Context flag: service marks MarkResponseCommitted(c) after
writing, ensureForwardErrorResponse checks IsResponseCommitted(c)
and skips. Zero function signature changes, zero error wrapping.
When a Forward implementation already wrote a complete non-SSE (JSON) error
response to the client and returned an error -- e.g. the case-400 passthrough
in GatewayService.handleErrorResponse -- the handler unconditionally called
ensureForwardErrorResponse, which detected the writer was already written and
appended a fallback `data: {"type":"error",...}` SSE frame. The client then
received a corrupted body: the upstream JSON immediately followed by a stray
`data:` line.
Add gatewayForwardErrorAlreadyCommunicated (and the OpenAI counterpart) to
detect this case -- writer size changed AND Content-Type is not
text/event-stream -- and skip the fallback. SSE streams that only flushed
keepalive pings or partial data still receive a protocol-compliant terminal
frame, so strict SDKs (Codex CLI) do not see a silent EOF.
Applied consistently across the Messages / ChatCompletions / Responses
gateway handlers and the OpenAI chat/images handlers. Added regression tests
covering JSON passthrough, mid-stream SSE 400, nil-error and no-write cases.
Keep Anthropic request body rewrites attempt-local and synchronize the accepted wire body only after upstream success so failover and retry paths do not reuse stale parsed state.
Keep large gateway payloads as raw body ranges and bind OpenAI parsed-body caches to the body bytes so failover and mapping do not reuse stale mutable state.