Commit Graph
191 Commits
Author SHA1 Message Date
cyh 2264a33085 fix(gateway): normalize Claude Code 1m model suffix 2026-07-17 21:59:57 +08:00
Wesley LiddickandGitHub 8bfbc5ca99 Merge pull request #4485 from Sub2API-Devs/dev
feat(security-audit): 新增 OpenAI 兼容提示词审计能力与安全审计控制台
2026-07-17 16:15:26 +08:00
mt21625457 d11bdb13f5 feat(security-audit): add OpenAI-compatible prompt auditing 2026-07-17 00:39:39 +08:00
Heatherm Huang 115116e8bf fix(grok): repair OAuth routing regressions 2026-07-16 16:52:25 +08:00
Tian Lee f59a6ed74c feat: 增加 API Key 计费倍率自省接口 2026-07-15 22:42:53 +08:00
shaw a0593b0bf8 fix(gateway): 客户端断开后 failover 静默终止,不再误报 502 账号耗尽
上游请求经 detachUpstreamContext(WithoutCancel) 有意脱离客户端取消(保
计费),但 failover 循环仍用原始 c.Request.Context() 重新选号:客户端
断开后上游返回 520 等可 failover 错误时,重新选号必然得到 context
canceled,被误判为账号耗尽,记录并返回通用 502。

修复:客户端已断开 ⇒ failover 静默终止。

- 新增 failoverClientGone(c):请求 ctx 已取消时先停 compact 心跳
  (建立 happens-before,对齐其它终结路径),响应未提交则标 499
  (statusClientClosedRequest,与并发槽取消路径同惯例)
- 7 个 OpenAI 内联 failover 循环(Responses/Messages/chat_completions/
  embeddings/images/grok_media/alpha_search)加双 guard:换号前 +
  选号失败分支入口;guard 位于 ReportOpenAIAccountScheduleResult(false)
  之后、RecordOpenAIAccountSwitch/池模式重试之前,账号健康副作用
  (service 层 detached ctx)不受影响
- FailoverState.HandleFailoverError/HandleSelectionExhausted 入口加
  ctx.Err() 检查返回 FailoverCanceled,取消不再改动 failover 状态;
  全部 10 个 FailoverCanceled 分支统一调用 failoverClientGone 归类 499
- 上游 detach 与计费设计不变;真实上游 520 事件仍完整落 ops
  (面板显示码 COALESCE(upstream_status_code,status_code)=520,
  错误率/告警口径不变)

测试:新增 openai_responses_failover_cancel_test.go 复现 issue 场景
(520+取消 ⇒ 不切号、499、无 502 终态)+ 在线客户端对照(正常切换、
耗尽 502);failover_loop_test.go 补入口取消用例并修正取消语义断言。

Fixes #4257
2026-07-15 09:29:05 +08:00
Wesley LiddickandGitHub b4aa3eb023 Merge pull request #4148 from feeeei/fix/retry_count
fix: 池模式同账号重试次数配置对 Anthropic/Gemini/通用转发路径生效
2026-07-13 15:32:40 +08:00
feeeei c7c933776d fix: 池模式同账号重试次数配置对 Anthropic/Gemini/通用转发路径生效
e643fc38 引入 pool_mode_retry_count 账号配置时,只覆盖了 OpenAI 族
handler 的内联重试循环;走 HandleFailoverError 的 Anthropic/
Antigravity/Gemini/通用转发路径仍硬编码同账号重试 3 次,配置不生效。

为 HandleFailoverError 增加 retryLimit 参数,由调用方传入
account.GetPoolModeRetryCount();未配置账号默认仍为 3 次,行为不变。
2026-07-13 12:57:00 +08:00
yan9651688 3605a316af Keep usage ranges consistent across API and dashboards
Expose the active weekly subscription window through /v1/usage, calculate offsets with the same normalized page size used by queries, and keep user-facing date ranges on the browser's local calendar date.

Constraint: Preserve existing response fields and avoid new dependencies
Rejected: Keep duplicate inline date formatters | a shared local-date utility prevents the same UTC regression in both views
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep Offset and Limit based on the same normalized page size
Tested: go test ./internal/pkg/pagination ./internal/handler; go vet ./internal/pkg/pagination ./internal/handler; frontend 923 tests; pnpm typecheck; pnpm lint:check; pnpm build
Not-tested: Live API request against a deployed subscription
Related: #4121
2026-07-13 11:35:46 +08:00
Wesley LiddickandGitHub 0438057c0b Merge pull request #3816 from Vibeone/fix/ops-inband-sse-error-logging
fix(ops): 记录固化 200 SSE 流上的就地错误,修复流内限流不进错误看板
2026-07-09 16:35:29 +08:00
InCerry 53a5c45bd8 fix(gateway): cap lenient json normalization
Fixes #3540
2026-07-09 11:15:52 +08:00
Eyre921 5aba53d542 fix(ops): 记录固化 200 SSE 流上的就地错误,修复流内限流不进错误看板
流式请求一旦 flush 了 keepalive ping,HTTP 状态码即固化为 200;此后
出现的错误(等待并发槽位超时后回退的限流、Wait 后二次计费校验失败、
流开始后才无可用账号等)只能就地以 SSE error 帧回传。而 ops_error_logger
以 status>=400 为采集触发条件,这类挂在 200 流上的失败此前会在错误看板里
完全隐形——客户端能收到 rate_limit_error,但网关侧没有任何错误记录可供排障。

- service: 新增 OpsStreamError 上下文 + MarkOpsStreamError/GetOpsStreamError,
  采用「首个标记生效」保留根因错误,避免被后续通用兜底帧覆盖。
- handler: handleStreamingAwareError 在 streamStarted 分支标记流内错误。
- handler: OpsErrorLoggerMiddleware 在 status<400 且无上游错误上下文时,
  据标记补记一条错误日志;分级用 IntendedStatus(如并发限流 429),
  StatusCode 仍记 wire 的 200。上游透传错误已由 upstream-context 分支落库,
  故此路径不重复记录。
- 补充单测覆盖补记、no-op、skip_monitoring 跳过与首个标记生效。
2026-07-08 02:51:51 +00:00
li 40c563c4ae fix(gateway): 记录请求体解析失败的真实原因,不再吞错
400 "Failed to parse request body" 此前丢弃底层错误,无法区分
JSON 真非法、还是 body 被截断/被中间件提前消费。

- 服务层 invalid json 错误增补 len/offset/非法字符信息
  (仅诊断元数据,不含 body 内容,可安全 wrap);
- handler 层新增 logRequestBodyParseFailure,向服务端日志输出
  底层错误 + body 长度 + 转义后的 head/tail 片段(各 256B),
  客户端响应文案保持不变;
- 接入全部 9 处入站解析点(messages/count_tokens/responses/
  chat_completions/embeddings,Anthropic 与 OpenAI 网关)。

Fixes #3715
2026-07-07 13:53:39 +08:00
wucm667 41cdd438d7 fix(gateway): honor Anthropic custom models list 2026-07-05 08:35:34 +08:00
Wesley LiddickandGitHub 2fc4fef847 Merge pull request #3310 from heathermhuang/codex/grok-subscription-support
feat: add grok subscription support
2026-06-26 15:41:52 +08:00
shaw fcd3bc1272 fix: return 404 model_not_found instead of 503 when no account supports the model 2026-06-26 15:38:06 +08:00
Heatherm Huang 39be1ec97f feat: add grok subscription support 2026-06-26 10:36:09 +08:00
Wesley LiddickandGitHub 7fb3e1aacd Merge pull request #3247 from alfadb/fix/reasoning-and-thinking-protocol
fix(gateway): 整合推理强度与思考协议处理(替代 #2155 / #2136 / #3246)
2026-06-16 20:29:21 +08:00
Wesley LiddickandGitHub 03ec90a25a Merge pull request #3223 from feitianbubu/fix/intercept-streaming-haiku-probe
fix: 修复CC Switch改为流式测试后请求失败的问题
2026-06-16 20:27:33 +08:00
alfadb a05d9e87c0 feat(billing): 国产模型 thinking-enabled 自动填充 reasoning_effort 默认值
问题:Kimi/GLM/MiniMax 等国产 LLM 协议层只有 thinking on/off 开关,没有
reasoning_effort 档位概念。客户端启用 thinking 后 usage_log.reasoning_effort
长期为 NULL,无法在用量分析里区分 'thinking 开启' 与 'thinking 关闭'。

方案:仅在 'thinking 启用 + 上游属于 passback-required 国产模型 + 客户端
未明确指定 effort' 三者同时成立时,给 usage_log.reasoning_effort 写默认值
'high'(与 DeepSeek thinking-enabled 默认 effort 一致)。

设计原则:
1. **白名单**:仅 ResolveThinkingProtocol == PassbackRequired 集合内,
   且排除原生支持 effort 的 DeepSeek(避免覆盖客户端意图)。
2. **fail-open**:客户端显式传 effort 时永远不覆盖。
3. **未来兼容**:如 Kimi 后续加入真 effort 档位,客户端开始发 effort,
   guard (3) 自动让出,本逻辑变 no-op。

实现:
- gateway_request.go: 加 DefaultEffortForThinkingEnabled (按模型白名单)
  + OpenAIBodyHasThinkingEnabled (检测 OpenAI 协议 body 里的 thinking.type)
  + ApplyThinkingEnabledFallback (包装现有 extractor 的 nil-then-default 逻辑)
- gateway_handler.go: Anthropic 路径两处(主 + retry)对称补充
- OpenAI 路径全覆盖:openai_gateway_service.go (passthrough + non-passthrough)
  + openai_gateway_chat_completions_raw.go + openai_gateway_responses_chat_fallback.go
  + openai_ws_http_bridge.go + openai_ws_forwarder.go
- 跨协议路径:gateway_forward_as_chat_completions.go (CC client → Anthropic upstream)
  + gateway_forward_as_responses.go (Responses client → Anthropic upstream)
  + gemini_chat_completions_compat_service.go (一致性保持)

未覆盖:openai_ws_v2_passthrough_adapter.go 两处。原因:该 adapter 持有的是
session-level 客户端原始 model,没有 *Account 句柄无法走 GetMappedModel。
WS v2 当前对国产模型场景不重要(pi 调用 Kimi/GLM/MiniMax 走 sync HTTP),
留待后续如果出现 WS v2 + 国产模型用例时单独处理。

测试:
- TestDefaultEffortForThinkingEnabled (14 用例):覆盖 Kimi/GLM/MiniMax 大小写、
  Qwen thinking 变体、DeepSeek 排除、Claude/GPT/Gemini 不命中。
- TestOpenAIBodyHasThinkingEnabled (8 用例):covers enabled/adaptive/disabled、
  大小写、空 body、缺字段、invalid JSON fail-safe。
- TestApplyThinkingEnabledFallback (9 用例):现有 effort 不覆盖、nil + 启用 +
  passback → high、nil + disabled → nil、nil + 启用 + 排除模型 → nil。
2026-06-16 19:37:31 +08:00
jjawandshaw b0579c4891 fix: move user wait queue accounting off hot path 2026-06-16 11:41:54 +08:00
feitianbubu b256f91141 fix(gateway): intercept max_tokens=1 haiku probes for streaming requests too 2026-06-11 20:46:47 +08:00
Wesley LiddickandGitHub dd709f5985 Merge pull request #3181 from codeQuest-fly/fix/gateway-upstream-error-double-write
fix: avoid double-writing error frame on non-stream upstream errors
2026-06-10 09:26:18 +08:00
erio 12962bab24 refactor(bedrock): merge header filtering into ApplyBedrockCCCompat
Move anthropic-beta header filtering from separate FilterBedrockBetaHeader
into ApplyBedrockCCCompat, so one function handles all CC compat processing
(body cleanup + header filtering). Change signature from ctx to *gin.Context
to access request headers. Remove the redundant separate call in handler.
2026-06-10 00:18:09 +08:00
erio 6c88631690 fix(gateway): prevent double-write on error passthrough responses
Service layer writes a complete JSON error response then returns error.
Handler's ensureForwardErrorResponse couldn't distinguish this from
"no response written" and appended an SSE event, corrupting the body.

Use gin.Context flag: service marks MarkResponseCommitted(c) after
writing, ensureForwardErrorResponse checks IsResponseCommitted(c)
and skips. Zero function signature changes, zero error wrapping.
2026-06-10 00:15:51 +08:00
dailingfei 914c059f4a fix: avoid double-writing error frame on non-stream upstream errors
When a Forward implementation already wrote a complete non-SSE (JSON) error
response to the client and returned an error -- e.g. the case-400 passthrough
in GatewayService.handleErrorResponse -- the handler unconditionally called
ensureForwardErrorResponse, which detected the writer was already written and
appended a fallback `data: {"type":"error",...}` SSE frame. The client then
received a corrupted body: the upstream JSON immediately followed by a stray
`data:` line.

Add gatewayForwardErrorAlreadyCommunicated (and the OpenAI counterpart) to
detect this case -- writer size changed AND Content-Type is not
text/event-stream -- and skip the fallback. SSE streams that only flushed
keepalive pings or partial data still receive a protocol-compliant terminal
frame, so strict SDKs (Codex CLI) do not see a silent EOF.

Applied consistently across the Messages / ChatCompletions / Responses
gateway handlers and the OpenAI chat/images handlers. Added regression tests
covering JSON passthrough, mid-stream SSE 400, nil-error and no-write cases.
2026-06-09 22:46:06 +08:00
Wesley LiddickandGitHub 5f63fe1945 Merge pull request #2927 from moonagic/main
fix antigravity gemini rate limit and account scheduling
2026-06-01 14:50:36 +08:00
moonagic a01686c637 fix antigravity gemini rate limit and account scheduling
Squash of 4 commits:
- Fix Gemini rate limit scheduling
- fix antigravity gemini rate limit scheduling
- fix antigravity gemini limited account scheduling
- fix antigravity test stubs for default lint
2026-05-31 22:00:03 +08:00
name 2caee9d884 refactor(gateway): snapshot usage worker inputs 2026-05-30 20:01:39 +08:00
name 619e5ae619 refactor(gateway): isolate anthropic body rewrites
Keep Anthropic request body rewrites attempt-local and synchronize the accepted wire body only after upstream success so failover and retry paths do not reuse stale parsed state.
2026-05-30 20:01:39 +08:00
name b1c4be4ac8 refactor(gateway): remove parsed request object graphs
Keep large gateway payloads as raw body ranges and bind OpenAI parsed-body caches to the body bytes so failover and mapping do not reuse stale mutable state.
2026-05-30 20:01:38 +08:00
name d8cbf9ab5c refactor(gateway): introduce request body refs 2026-05-30 20:01:38 +08:00
Wesley LiddickandGitHub 69e7c4db30 Merge pull request #2865 from wey-gu/feat/usage-request-context
fix(gateway): preserve usage request context
2026-05-29 16:21:59 +08:00
Wey Gu 2bd3125d0f Preserve usage request context 2026-05-28 22:44:25 +08:00
gaoren002 56e96fdd8c fix: classify concurrency acquire failures 2026-05-28 10:03:41 +00:00
lyen1688andlyen1688 f597c1581b feat(group): 支持自定义 /v1/models 模型列表 2026-05-27 18:00:45 +08:00
benjaminandSisyphus c3e7476992 fix(gateway): mark local platform gates business-limited
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-26 17:19:23 +08:00
6b39b344d8 feat(quota): 用户 × 平台 USD 配额
为用户在 anthropic/openai/gemini/antigravity 四个平台上提供日/周/月
三个窗口的 USD 配额管控。配额语义:未设置=不限制,0=禁用,>0=美元上限。

两层模型:
- 配置层:系统默认配额,以及 email/linuxdo/oidc/wechat/github/google/
  dingtalk 七个鉴权来源的默认配额,存于 settings,以嵌套 JSON 整体读写
  (系统 1 个 key + 每个来源 1 个 key),整体替换语义。
- 运行时层:user_platform_quota 表按用户记录实际配额,与配置层解耦。

后端:新增 ent schema 与 140_user_platform_quotas.sql 迁移、repository
与 service 端口、计费链路集成、管理端与用户端读写接口。
前端:管理端设置页配额编辑、用户配额管理 Modal、用户 Dashboard 展示、
中英文案。

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 10:49:20 +08:00
Jamie WongandClaude Opus 4.7 b34cc71bee fix(openai): also emit response.failed in ensureForwardErrorResponse after Writer.Written
Case B: when a slot wait flushes SSE ping comments first (Writer.Written
becomes true), the previous ensureForwardErrorResponse short-circuited
on `c.Writer.Written()` and returned false without notifying the client.
Subsequent upstream errors (http2 timeout, stream INTERNAL_ERROR, etc.)
produced silent EOF; Codex CLI reported "stream closed before
response.completed" just like the user-slot timeout case.

Remove the Written() early return; coerce streamStarted to true when
Writer has already been written to, and let handleStreamingAwareError
walk the existing logic — which now (thanks to the previous commits)
emits a protocol-compliant response.failed for /responses paths and the
legacy `event: error` for others.

Update tests that previously asserted "do not override written response":
the new contract is to *append* an SSE terminal frame so the client sees
a clean close instead of EOF. recoverResponsesPanic inherits this fix.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 22:00:56 +08:00
Jamie WongandClaude Opus 4.7 cff2f291be fix(openai): also match bare /responses route in handleStreamingAwareError
The first revision compared GetInboundEndpoint(c) against EndpointResponses
("/v1/responses"). NormalizeInboundEndpoint only recognizes paths that
contain the literal "/v1/responses" substring, but the project actually
registers six /responses routes — three of which (top-level
r.POST("/responses", ...) and codexDirect's "/backend-api/codex/responses")
have FullPath values without the "/v1" prefix and therefore fall through
to the default branch.

Codex CLI users targeting the bare /responses route at the production
deployment (observed 2026-05-24 ~11:05 UTC, user 16) never reached the
new writeResponsesFailedSSE path: the endpoint check was false, the
legacy `event: error` frame fired, and the strict SDK kept reporting
"stream closed before response.completed".

Replace the strict equality check with inboundIsResponses(c), which
uses suffix detection on FullPath (falling back to URL.Path when
FullPath is empty in test fixtures) and covers all six route variants:

  /v1/responses[/...]
  /responses[/...]
  /backend-api/codex/responses[/...]

Add test table covering all routes plus negative cases.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 19:32:08 +08:00
Jamie WongandClaude Opus 4.7 5e5c2062bf fix(openai): emit response.failed for /v1/responses after stream started
When /v1/responses streaming hits the user/account concurrency wait, the
wait loop sends SSE ping comments to keep the connection alive, which
flushes HTTP 200 + headers. If the wait then times out (or any other
post-flush error fires), handleStreamingAwareError previously emitted a
generic `event: error` frame. Codex CLI requires the stream to end with
a Responses terminal event (response.completed/failed/incomplete/cancelled),
so it reports "stream closed before response.completed" and the user-facing
rate-limit intent is lost.

This change detects inbound = /v1/responses in both handleStreamingAwareError
implementations and emits a protocol-compliant response.failed event whose
field set mirrors apicompat.makeResponsesCompletedEvent
(id/object/model/status/output/error). The synthetic id reuses
ctxkey.RequestID so client errors can be grepped against server logs.
sequence_number is intentionally omitted to preserve monotonicity on streams
that already emitted real events.

Other inbound endpoints (/v1/chat/completions, /v1/messages) keep their
legacy formats untouched.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 10:58:29 +08:00
erio fe1c6c958b feat(bedrock): add Claude Code compatibility for AWS Bedrock
- Export ApplyBedrockCCCompat() in GatewayService, called after channel
  model mapping to ensure mapped model ID is used for Opus 4.7+ detection
- Add sanitizeBedrockCCFields(): remove service_tier/interface_geo/
  context_management, inject max_tokens/anthropic_version defaults
- Add sanitizeBedrockCCBetaTokens(): filter anthropic_beta to keep only
  Bedrock-supported tokens, reusing autoInjectBedrockBetaTokens and
  filterBedrockBetaTokens for consistent rules
- Remove unsupported beta tokens (interleaved-thinking, context-management)
  from whitelist based on AWS official docs
- Simplify IsBedrockCCCompatEnabled() to check boolean toggle directly,
  applying CC compat to all accounts regardless of platform
- Add unit tests for IsBedrockCCCompatEnabled (8 cases),
  sanitizeBedrockCCFields (8 cases), sanitizeBedrockCCBetaTokens (7 cases)
- Update bedrock beta policy tests for removed auto-injection
2026-05-21 11:46:24 +08:00
wucm667 90b2b2a757 feat(usage): 用户 API Key 用量页支持按日明细 2026-05-20 15:48:38 +08:00
name 2eb622f2f6 Remove ops retry replay storage 2026-05-19 19:37:41 +08:00
wucm667 6381f9e37d fix(openai): 识别上游静默拒绝并触发 failover 2026-05-19 15:48:36 +08:00
Wesley LiddickandGitHub 8a4ee578cb Merge pull request #2451 from wucm667/codex/issue-2237-gemini-chat-completions
fix(gateway): 修复 Gemini 组 Chat Completions 路由
2026-05-19 14:47:52 +08:00
benjaminandSisyphus 6acb46c113 fix: 标记通用网关本地调度容量错误
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-18 16:52:32 +08:00
wucm667 2ec1d331e0 fix(gateway): return Gemini models for Gemini groups 2026-05-15 11:33:26 +08:00
shaw fff4a300c6 feat(risk-control): add content moderation audit 2026-05-07 09:14:47 +08:00
shaw 733627cf9d fix: improve sticky session scheduling 2026-04-30 11:38:11 +08:00