Commit Graph
206 Commits
Author SHA1 Message Date
shaw aa673062e2 fix(composite): keep CN rollout on fully supported paths
- Responses WebSocket composite whitelist stays openai+grok: CN accounts
  cannot pass the WSv2 ingress transport filter and the WS HTTP bridge
  has no Responses conversion for them, so admitting CN targets only
  turns a clear policy rejection into a misleading no-available-account
- widen the responses/input_tokens and messages/count_tokens composite
  gates to the shared openai-compatible text whitelist so CN targets get
  the same local-estimate token counting as generation
- restore the Claude default-model fallback for standalone CN groups
  (admin candidates and custom models list) while keeping composite
  listings limited to CN account mapping keys
- refresh the scheduler bulk-rebuild comment and bucket capacity hints
  for the 8-platform canonical set
2026-08-19 14:58:32 +08:00
wucm667 b171bb0e4a fix(composite): support CN providers
Extend composite routing, pricing, migrations, and admin options for Kimi, GLM, and DeepSeek.
2026-08-19 13:01:25 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
IanShaw027 245d069602 fix(grok): 二轮评审残留 — Search 叠加计费与 fail-closed 安全
- SearchCost 叠加 token(openai/gateway),未定价 warn
- Token URL 校验失败回落 DefaultTokenURL,禁止 Effective 旁路
- free 判定收窄 paidSignal(仅 plan/月额度),usage% 不否决 free
- web_search: mandatory 计费、uuid request_id、上游 failover 重选账号
- 调度阈值:7d/30d 不跨期 until + 48h stale 可选跳过
- SanitizeStoredCredentials 接入 create/update/bulk/SSO/ApplyOAuth
- ApplyOAuth 成功清除 grok_needs_reauth;SSO 允许 header_override_enabled
- VideoModelPrices 视为媒体定价完整;realtime 正常关闭仍计费
2026-08-07 17:40:21 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuangandshaw fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuangandshaw 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
shaw bd52e5d770 fix(gateway): record observed usage when anthropic stream is interrupted
Fixes #5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.

Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.

Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
  into a ForwardResult and return it alongside the error, for both the
  regular Anthropic path and the API-key passthrough path. Invariants:
  UpstreamFailoverError always keeps result=nil (failover retries are
  billed as the successful attempt, never twice), and zero observed
  usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
  shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
  (dropped_stopped) from operator-configured drop/sample overflow
  drops; billing tasks now fall back to inline synchronous execution
  only during the shutdown window, while explicit drop/sample overflow
  semantics are preserved. Image usage keeps its mandatory fallback
  for both drop kinds via the new mode.Dropped() helper.

Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
2026-08-01 11:29:58 +08:00
eyre 248236ce6d fix(gateway): 修复模拟响应使用 Bedrock msg_bdrk_ 格式,改为正宗 Anthropic msg_01 格式
问题:
探针拦截(suggestion mode / warmup / max_tokens=1 haiku)的模拟响应以及
Gemini/Antigravity 兼容层生成的 message ID 不符合 Anthropic 官方 API 格式,
容易被客户端识别为非正宗响应。

修复:
1. generateRealisticMsgID():msg_bdrk_ + 24字符 → msg_01 + 22位 Base62
   (与官方 API 返回的 msg_011CdS6b8gAhoKWdW9jE87Zs 格式一致)
2. 去掉固定的 msg_mock_suggestion / msg_mock_warmup,统一使用随机 ID
3. 流式响应格式对齐官方:
   - message_start 增加 stop_details/cache token 字段
   - content_block_start 字段顺序修正
   - message_delta.usage 只含 output_tokens
4. 非流式响应:增加 stop_details:null,移除非标准 total_tokens
5. Gemini Messages/ChatCompletions 兼容层:msg_ + hex → msg_01 + Base62
6. Antigravity response/stream transformer:msg_ + 12位 → msg_01 + 22位 Base62

验证方式:对照 Anthropic 官方 API 实际响应格式确认。
2026-07-27 15:20:03 +00:00
Edison42 1c0cb24c7e feat(usage): persist client session identifiers 2026-07-24 01:22:34 +08:00
Heatherm Huang ee332cee64 Fix composite model defaults for linked platforms 2026-07-23 09:20:52 +08:00
Heatherm Huang 3a683fff55 Fix composite route alias attribution 2026-07-23 09:20:52 +08:00
Heatherm Huang c8d1e2e16f Harden composite group product surfaces 2026-07-23 09:19:25 +08:00
Heatherm Huang ebc1028771 Add composite group routing 2026-07-23 09:19:24 +08:00
wucm667 addd5ef1dc [verified] fix: align sync cache billing after failover 2026-07-20 22:43:11 +08:00
cyh 2264a33085 fix(gateway): normalize Claude Code 1m model suffix 2026-07-17 21:59:57 +08:00
Wesley LiddickandGitHub 8bfbc5ca99 Merge pull request #4485 from Sub2API-Devs/dev
feat(security-audit): 新增 OpenAI 兼容提示词审计能力与安全审计控制台
2026-07-17 16:15:26 +08:00
mt21625457 d11bdb13f5 feat(security-audit): add OpenAI-compatible prompt auditing 2026-07-17 00:39:39 +08:00
Heatherm Huang 115116e8bf fix(grok): repair OAuth routing regressions 2026-07-16 16:52:25 +08:00
Tian Lee f59a6ed74c feat: 增加 API Key 计费倍率自省接口 2026-07-15 22:42:53 +08:00
shaw a0593b0bf8 fix(gateway): 客户端断开后 failover 静默终止,不再误报 502 账号耗尽
上游请求经 detachUpstreamContext(WithoutCancel) 有意脱离客户端取消(保
计费),但 failover 循环仍用原始 c.Request.Context() 重新选号:客户端
断开后上游返回 520 等可 failover 错误时,重新选号必然得到 context
canceled,被误判为账号耗尽,记录并返回通用 502。

修复:客户端已断开 ⇒ failover 静默终止。

- 新增 failoverClientGone(c):请求 ctx 已取消时先停 compact 心跳
  (建立 happens-before,对齐其它终结路径),响应未提交则标 499
  (statusClientClosedRequest,与并发槽取消路径同惯例)
- 7 个 OpenAI 内联 failover 循环(Responses/Messages/chat_completions/
  embeddings/images/grok_media/alpha_search)加双 guard:换号前 +
  选号失败分支入口;guard 位于 ReportOpenAIAccountScheduleResult(false)
  之后、RecordOpenAIAccountSwitch/池模式重试之前,账号健康副作用
  (service 层 detached ctx)不受影响
- FailoverState.HandleFailoverError/HandleSelectionExhausted 入口加
  ctx.Err() 检查返回 FailoverCanceled,取消不再改动 failover 状态;
  全部 10 个 FailoverCanceled 分支统一调用 failoverClientGone 归类 499
- 上游 detach 与计费设计不变;真实上游 520 事件仍完整落 ops
  (面板显示码 COALESCE(upstream_status_code,status_code)=520,
  错误率/告警口径不变)

测试:新增 openai_responses_failover_cancel_test.go 复现 issue 场景
(520+取消 ⇒ 不切号、499、无 502 终态)+ 在线客户端对照(正常切换、
耗尽 502);failover_loop_test.go 补入口取消用例并修正取消语义断言。

Fixes #4257
2026-07-15 09:29:05 +08:00
Wesley LiddickandGitHub b4aa3eb023 Merge pull request #4148 from feeeei/fix/retry_count
fix: 池模式同账号重试次数配置对 Anthropic/Gemini/通用转发路径生效
2026-07-13 15:32:40 +08:00
feeeei c7c933776d fix: 池模式同账号重试次数配置对 Anthropic/Gemini/通用转发路径生效
e643fc38 引入 pool_mode_retry_count 账号配置时,只覆盖了 OpenAI 族
handler 的内联重试循环;走 HandleFailoverError 的 Anthropic/
Antigravity/Gemini/通用转发路径仍硬编码同账号重试 3 次,配置不生效。

为 HandleFailoverError 增加 retryLimit 参数,由调用方传入
account.GetPoolModeRetryCount();未配置账号默认仍为 3 次,行为不变。
2026-07-13 12:57:00 +08:00
yan9651688 3605a316af Keep usage ranges consistent across API and dashboards
Expose the active weekly subscription window through /v1/usage, calculate offsets with the same normalized page size used by queries, and keep user-facing date ranges on the browser's local calendar date.

Constraint: Preserve existing response fields and avoid new dependencies
Rejected: Keep duplicate inline date formatters | a shared local-date utility prevents the same UTC regression in both views
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep Offset and Limit based on the same normalized page size
Tested: go test ./internal/pkg/pagination ./internal/handler; go vet ./internal/pkg/pagination ./internal/handler; frontend 923 tests; pnpm typecheck; pnpm lint:check; pnpm build
Not-tested: Live API request against a deployed subscription
Related: #4121
2026-07-13 11:35:46 +08:00
Wesley LiddickandGitHub 0438057c0b Merge pull request #3816 from Vibeone/fix/ops-inband-sse-error-logging
fix(ops): 记录固化 200 SSE 流上的就地错误,修复流内限流不进错误看板
2026-07-09 16:35:29 +08:00
InCerry 53a5c45bd8 fix(gateway): cap lenient json normalization
Fixes #3540
2026-07-09 11:15:52 +08:00
Eyre921 5aba53d542 fix(ops): 记录固化 200 SSE 流上的就地错误,修复流内限流不进错误看板
流式请求一旦 flush 了 keepalive ping,HTTP 状态码即固化为 200;此后
出现的错误(等待并发槽位超时后回退的限流、Wait 后二次计费校验失败、
流开始后才无可用账号等)只能就地以 SSE error 帧回传。而 ops_error_logger
以 status>=400 为采集触发条件,这类挂在 200 流上的失败此前会在错误看板里
完全隐形——客户端能收到 rate_limit_error,但网关侧没有任何错误记录可供排障。

- service: 新增 OpsStreamError 上下文 + MarkOpsStreamError/GetOpsStreamError,
  采用「首个标记生效」保留根因错误,避免被后续通用兜底帧覆盖。
- handler: handleStreamingAwareError 在 streamStarted 分支标记流内错误。
- handler: OpsErrorLoggerMiddleware 在 status<400 且无上游错误上下文时,
  据标记补记一条错误日志;分级用 IntendedStatus(如并发限流 429),
  StatusCode 仍记 wire 的 200。上游透传错误已由 upstream-context 分支落库,
  故此路径不重复记录。
- 补充单测覆盖补记、no-op、skip_monitoring 跳过与首个标记生效。
2026-07-08 02:51:51 +00:00
li 40c563c4ae fix(gateway): 记录请求体解析失败的真实原因,不再吞错
400 "Failed to parse request body" 此前丢弃底层错误,无法区分
JSON 真非法、还是 body 被截断/被中间件提前消费。

- 服务层 invalid json 错误增补 len/offset/非法字符信息
  (仅诊断元数据,不含 body 内容,可安全 wrap);
- handler 层新增 logRequestBodyParseFailure,向服务端日志输出
  底层错误 + body 长度 + 转义后的 head/tail 片段(各 256B),
  客户端响应文案保持不变;
- 接入全部 9 处入站解析点(messages/count_tokens/responses/
  chat_completions/embeddings,Anthropic 与 OpenAI 网关)。

Fixes #3715
2026-07-07 13:53:39 +08:00
wucm667 41cdd438d7 fix(gateway): honor Anthropic custom models list 2026-07-05 08:35:34 +08:00
Wesley LiddickandGitHub 2fc4fef847 Merge pull request #3310 from heathermhuang/codex/grok-subscription-support
feat: add grok subscription support
2026-06-26 15:41:52 +08:00
shaw fcd3bc1272 fix: return 404 model_not_found instead of 503 when no account supports the model 2026-06-26 15:38:06 +08:00
Heatherm Huang 39be1ec97f feat: add grok subscription support 2026-06-26 10:36:09 +08:00
Wesley LiddickandGitHub 7fb3e1aacd Merge pull request #3247 from alfadb/fix/reasoning-and-thinking-protocol
fix(gateway): 整合推理强度与思考协议处理(替代 #2155 / #2136 / #3246)
2026-06-16 20:29:21 +08:00
Wesley LiddickandGitHub 03ec90a25a Merge pull request #3223 from feitianbubu/fix/intercept-streaming-haiku-probe
fix: 修复CC Switch改为流式测试后请求失败的问题
2026-06-16 20:27:33 +08:00
alfadb a05d9e87c0 feat(billing): 国产模型 thinking-enabled 自动填充 reasoning_effort 默认值
问题:Kimi/GLM/MiniMax 等国产 LLM 协议层只有 thinking on/off 开关,没有
reasoning_effort 档位概念。客户端启用 thinking 后 usage_log.reasoning_effort
长期为 NULL,无法在用量分析里区分 'thinking 开启' 与 'thinking 关闭'。

方案:仅在 'thinking 启用 + 上游属于 passback-required 国产模型 + 客户端
未明确指定 effort' 三者同时成立时,给 usage_log.reasoning_effort 写默认值
'high'(与 DeepSeek thinking-enabled 默认 effort 一致)。

设计原则:
1. **白名单**:仅 ResolveThinkingProtocol == PassbackRequired 集合内,
   且排除原生支持 effort 的 DeepSeek(避免覆盖客户端意图)。
2. **fail-open**:客户端显式传 effort 时永远不覆盖。
3. **未来兼容**:如 Kimi 后续加入真 effort 档位,客户端开始发 effort,
   guard (3) 自动让出,本逻辑变 no-op。

实现:
- gateway_request.go: 加 DefaultEffortForThinkingEnabled (按模型白名单)
  + OpenAIBodyHasThinkingEnabled (检测 OpenAI 协议 body 里的 thinking.type)
  + ApplyThinkingEnabledFallback (包装现有 extractor 的 nil-then-default 逻辑)
- gateway_handler.go: Anthropic 路径两处(主 + retry)对称补充
- OpenAI 路径全覆盖:openai_gateway_service.go (passthrough + non-passthrough)
  + openai_gateway_chat_completions_raw.go + openai_gateway_responses_chat_fallback.go
  + openai_ws_http_bridge.go + openai_ws_forwarder.go
- 跨协议路径:gateway_forward_as_chat_completions.go (CC client → Anthropic upstream)
  + gateway_forward_as_responses.go (Responses client → Anthropic upstream)
  + gemini_chat_completions_compat_service.go (一致性保持)

未覆盖:openai_ws_v2_passthrough_adapter.go 两处。原因:该 adapter 持有的是
session-level 客户端原始 model,没有 *Account 句柄无法走 GetMappedModel。
WS v2 当前对国产模型场景不重要(pi 调用 Kimi/GLM/MiniMax 走 sync HTTP),
留待后续如果出现 WS v2 + 国产模型用例时单独处理。

测试:
- TestDefaultEffortForThinkingEnabled (14 用例):覆盖 Kimi/GLM/MiniMax 大小写、
  Qwen thinking 变体、DeepSeek 排除、Claude/GPT/Gemini 不命中。
- TestOpenAIBodyHasThinkingEnabled (8 用例):covers enabled/adaptive/disabled、
  大小写、空 body、缺字段、invalid JSON fail-safe。
- TestApplyThinkingEnabledFallback (9 用例):现有 effort 不覆盖、nil + 启用 +
  passback → high、nil + disabled → nil、nil + 启用 + 排除模型 → nil。
2026-06-16 19:37:31 +08:00
jjawandshaw b0579c4891 fix: move user wait queue accounting off hot path 2026-06-16 11:41:54 +08:00
feitianbubu b256f91141 fix(gateway): intercept max_tokens=1 haiku probes for streaming requests too 2026-06-11 20:46:47 +08:00
Wesley LiddickandGitHub dd709f5985 Merge pull request #3181 from codeQuest-fly/fix/gateway-upstream-error-double-write
fix: avoid double-writing error frame on non-stream upstream errors
2026-06-10 09:26:18 +08:00
erio 12962bab24 refactor(bedrock): merge header filtering into ApplyBedrockCCCompat
Move anthropic-beta header filtering from separate FilterBedrockBetaHeader
into ApplyBedrockCCCompat, so one function handles all CC compat processing
(body cleanup + header filtering). Change signature from ctx to *gin.Context
to access request headers. Remove the redundant separate call in handler.
2026-06-10 00:18:09 +08:00
erio 6c88631690 fix(gateway): prevent double-write on error passthrough responses
Service layer writes a complete JSON error response then returns error.
Handler's ensureForwardErrorResponse couldn't distinguish this from
"no response written" and appended an SSE event, corrupting the body.

Use gin.Context flag: service marks MarkResponseCommitted(c) after
writing, ensureForwardErrorResponse checks IsResponseCommitted(c)
and skips. Zero function signature changes, zero error wrapping.
2026-06-10 00:15:51 +08:00
dailingfei 914c059f4a fix: avoid double-writing error frame on non-stream upstream errors
When a Forward implementation already wrote a complete non-SSE (JSON) error
response to the client and returned an error -- e.g. the case-400 passthrough
in GatewayService.handleErrorResponse -- the handler unconditionally called
ensureForwardErrorResponse, which detected the writer was already written and
appended a fallback `data: {"type":"error",...}` SSE frame. The client then
received a corrupted body: the upstream JSON immediately followed by a stray
`data:` line.

Add gatewayForwardErrorAlreadyCommunicated (and the OpenAI counterpart) to
detect this case -- writer size changed AND Content-Type is not
text/event-stream -- and skip the fallback. SSE streams that only flushed
keepalive pings or partial data still receive a protocol-compliant terminal
frame, so strict SDKs (Codex CLI) do not see a silent EOF.

Applied consistently across the Messages / ChatCompletions / Responses
gateway handlers and the OpenAI chat/images handlers. Added regression tests
covering JSON passthrough, mid-stream SSE 400, nil-error and no-write cases.
2026-06-09 22:46:06 +08:00
Wesley LiddickandGitHub 5f63fe1945 Merge pull request #2927 from moonagic/main
fix antigravity gemini rate limit and account scheduling
2026-06-01 14:50:36 +08:00
moonagic a01686c637 fix antigravity gemini rate limit and account scheduling
Squash of 4 commits:
- Fix Gemini rate limit scheduling
- fix antigravity gemini rate limit scheduling
- fix antigravity gemini limited account scheduling
- fix antigravity test stubs for default lint
2026-05-31 22:00:03 +08:00
name 2caee9d884 refactor(gateway): snapshot usage worker inputs 2026-05-30 20:01:39 +08:00
name 619e5ae619 refactor(gateway): isolate anthropic body rewrites
Keep Anthropic request body rewrites attempt-local and synchronize the accepted wire body only after upstream success so failover and retry paths do not reuse stale parsed state.
2026-05-30 20:01:39 +08:00
name b1c4be4ac8 refactor(gateway): remove parsed request object graphs
Keep large gateway payloads as raw body ranges and bind OpenAI parsed-body caches to the body bytes so failover and mapping do not reuse stale mutable state.
2026-05-30 20:01:38 +08:00
name d8cbf9ab5c refactor(gateway): introduce request body refs 2026-05-30 20:01:38 +08:00
Wesley LiddickandGitHub 69e7c4db30 Merge pull request #2865 from wey-gu/feat/usage-request-context
fix(gateway): preserve usage request context
2026-05-29 16:21:59 +08:00
Wey Gu 2bd3125d0f Preserve usage request context 2026-05-28 22:44:25 +08:00
gaoren002 56e96fdd8c fix: classify concurrency acquire failures 2026-05-28 10:03:41 +00:00