Add gateway.models_list_read_max_bytes with the existing 8 MiB behavior as its default, and apply it consistently to generic, Codex, and Antigravity model-list reads.
Read one sentinel byte for Codex manifests so oversized responses return an explicit bounded upstream error instead of malformed JSON.
P2-1/P2-3 from review:
- fetchCNQuota: credential-invalid now judged by StatusCode 401/403
(aligned with fetchCNBalance) instead of `!Success && !CredentialValid` —
CN quota service only sets CredentialValid=true on the success path, so
500/429/zhipu business errors were all misclassified as failed instead
of error.
- fetchCNBalance: snapshot carries new BalanceLow flag computed with the
scheduler's exact criterion (`!Available || allCNBalancesBelowThreshold`)
against Gateway.CNProviders.BalanceThreshold (ctor now takes cfg; wire
regenerated). quotaDegradedHint reports "balance low" instead of the old
`<=0` check, so an account already paused by the scheduler (balance 5 /
threshold 10) no longer shows green in the monitor.
- threshold helper falls back to viper default 0.5 for nil/<=0 config to
avoid a zero-threshold regression where balance=0 stops alerting.
Tests: CN quota status-code matrix (rewrites the test that cemented the
old behavior), balance-low matrix (below-threshold / unavailable /
multi-currency healthy), threshold-from-config; PayG stubs now set
Available explicitly (zero-value trap).
The channel cache holds a groupID -> platform map with a 10 minute TTL, and
only channel Create/Update/Delete call invalidateCache(). Changing a group's
platform through the admin API therefore leaves the cache pointing at the old
platform for up to 10 minutes.
Channel pricing, model mapping and the model whitelist are all matched per
platform, so during that window the lookups silently miss: pricing falls back
to the global LiteLLM price list, renames stop applying and the whitelist
stops restricting. Nothing is logged.
Inject a narrow ChannelCacheInvalidator into the admin service (same shape as
the existing APIKeyAuthCacheInvalidator) and call it from UpdateGroup only when
the platform actually changed. The dependency is optional -- when it is nil the
cache simply rebuilds on TTL expiry, as before.
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.
- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
(V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
218/219/220 rules (#5408).
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.
Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
resolution failure surfaces as a moderation error and never silently
falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
proxy_not_active) without leaking credentials
Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
non-blockingly; save and test payloads carry proxy_id; zh/en i18n