Commit Graph
618 Commits
Author SHA1 Message Date
Wesley LiddickandGitHub fd6cd474d6 Merge pull request #5846 from lbyxiaolizi/fix/responses-chat-malformed-tool-arguments
fix(apicompat): reject malformed tool-call arguments
2026-08-22 13:35:02 +08:00
Wesley LiddickandGitHub 6244090c1c Merge pull request #5487 from an-epiphany/fix/file-part-min
fix(apicompat): chat/completions 的 file part 转换为 Responses input_file,不再静默丢弃
2026-08-22 13:34:18 +08:00
Wesley LiddickandGitHub 844b118785 Merge pull request #5938 from Hakunm/fix/google-one-model-catalog
fix(gemini): 限制 Google One OAuth 模型目录 / constrain Google One model catalog
2026-08-22 13:34:05 +08:00
Wesley LiddickandGitHub f646a1f974 Merge pull request #5632 from 3219378872/fix/apicompat-streaming-tool-name-empty
fix(apicompat): omit empty tool name on streamed arguments deltas
2026-08-21 17:49:55 +08:00
Wesley LiddickandGitHub 4eb7630ab5 Merge pull request #5625 from sweetcornna/fix/antigravity-official-daily-endpoint
fix(antigravity): use official daily endpoint
2026-08-21 17:40:29 +08:00
Hakunm f98a056f75 fix(gemini): constrain Google One model catalog 2026-08-21 02:33:00 +08:00
IanShaw 2e68b10aad 完善 Grok 内容拒绝计费与媒体兼容 2026-08-20 09:12:52 -07:00
IanShaw 39485f2e28 更新 Grok 默认模型与官方计费目录 2026-08-20 02:32:22 -07:00
IanShaw ed4207a16f 校正 Grok 模型目录计费与工具出站 2026-08-20 01:03:29 -07:00
Nai Long 7th 9ede0f7165 fix(grok): promote tool-search discoveries into callable tools 2026-08-20 12:37:39 +08:00
Nai Long 7th 5b2089c5a3 fix(grok): lower Codex tool-search discovery outputs 2026-08-20 06:55:47 +08:00
lbyxiaolizi fbc9ee626d fix(apicompat): narrow malformed tool-call handling 2026-08-19 20:49:31 +08:00
lbyxiaolizi e2d9ce0cad fix(apicompat): reject malformed tool-call arguments 2026-08-19 18:48:35 +08:00
hansnow fefd0d5145 test(apicompat): avoid unchecked tool history assertions 2026-08-19 15:31:18 +08:00
hansnow e4896c41d2 test(apicompat): satisfy client tool type assertions 2026-08-19 15:26:19 +08:00
hansnow b94e484e23 fix(openai): preserve client tools across WS bridge turns 2026-08-19 15:19:33 +08:00
Wesley LiddickandGitHub e943f817b1 Merge pull request #5729 from lbyxiaolizi/fix/responses-chat-reasoning-content-passback
fix(openai-compat): Responses→Chat 桥接按 reasoning item id 缓存回注 reasoning_content (#5520)
2026-08-19 14:03:03 +08:00
Wesley LiddickandGitHub 58ccea4eaa Merge pull request #5767 from hansnow/fix/ws-http-bridge-custom-tools
fix(openai): 补齐客户端工具终止事件恢复
2026-08-18 16:21:50 +08:00
hansnow c253bd2c72 fix(openai): restore client tools in terminal events 2026-08-18 15:48:02 +08:00
Wesley LiddickandGitHub 37732dcd34 Merge pull request #5725 from tamseno/fix/gemini-include-server-side-tool-invocations
fix(gemini): support includeServerSideToolInvocations in GeminiToolConfig
2026-08-18 15:35:01 +08:00
yaxin 1ba92449c7 fix(gemini): wire includeServerSideToolInvocations into the typed transform path
The struct field alone never reached the wire: the raw passthrough
pipeline is covered by enableMixedGeminiToolInvocations (#5711), but
TransformClaudeToGeminiWithOptions builds GeminiToolConfig from scratch
and never set the flag, so gemini-* models entering through the Claude
format gateway could still hit the upstream 400 from issue #5709.

- Set IncludeServerSideToolInvocations=true when the built tool
  declarations mix functionDeclarations with googleSearch, matching the
  raw-path injection semantics.
- Replace the marshal-roundtrip-only test with behavior tests that
  drive TransformClaudeToGeminiWithOptions: mixed tools set the flag,
  function-only and web-search-only requests leave it unset.
2026-08-18 15:05:54 +08:00
o2e 16e4f7ecc3 修复 Codex 额度探针模型兼容性 2026-08-18 13:08:28 +08:00
lbyxiaolizi 401dd43b4b fix(apicompat): 链式工具调用回放本轮 reasoning_content
codex 0.147.0 真实 resume 历史复现:DeepSeek 每轮只在开头流式产生一次
reasoning,reasoning → call A → output A → call B 的链式调用中 call B 前
没有 reasoning item,其 assistant 消息缺 reasoning_content,DeepSeek
thinking mode 400 整个历史。

buildChatMessagesFromItems 新增 lastTurnReasoning:记录本轮最近一次
reasoning,跨 tool output 存活,仅被 user 侧 item 清除;assistant 消息
(文本或 tool_calls)在 pendingReasoning 为空时回放本轮 reasoning。
2026-08-17 18:51:17 +08:00
lbyxiaolizi 612436a5a7 fix(openai-compat): Responses→Chat 桥接按 reasoning item id 缓存回注 reasoning_content
修复 #5520:Codex 经 force_chat_completions 桥接到 DeepSeek thinking 上游时,
历史中的 encrypted-only reasoning item(summary 为空 + 不透明 encrypted_content,
远程 compaction / 跨会话恢复后常见)取不出明文,后续 assistant 消息缺
reasoning_content,DeepSeek 400 "The reasoning_content in the thinking mode
must be passed back to the API",客户端仅看到通用 502。

reasoning item 的 id 一定会被客户端回传,以其为 key 做服务端缓存:

- GatewayCache 新增 Set/GetReasoningContent(Redis,默认 TTL 7 天)
- 响应侧:流式扫 response.output_item.done、非流式扫 output,把 reasoning
  全文按 item id 写缓存;客户端断连后 drain 期间用 detached ctx 仍会写完
- 请求侧:apicompat 新增 ResponsesToChatCompletionsRequestWithOptions 与
  ReasoningContentByID 钩子,encrypted-only item 查缓存补回 pendingReasoning;
  缓存 miss/出错一律 fail-open 维持原行为
- 自愈:历史里带明文 summary 的 reasoning item 顺手刷新缓存,覆盖 Redis
  flush / 跨实例漂移

测试:apicompat(命中恢复/miss 保持原样/明文优先)、service 端到端(流式
写缓存、请求侧回注+自愈)、repository miniredis 存取。
2026-08-17 16:48:36 +08:00
Tamseno 3c3bb2fa19 fix(gemini): support includeServerSideToolInvocations in GeminiToolConfig
- Add IncludeServerSideToolInvocations field to GeminiToolConfig to prevent dropping client tool settings.
- Fix HTTP 400 error when mixing built-in tools (e.g. Google Search) with function calling on Gemini 3.6/3.7 models.
- Add serialization/deserialization unit test TestGeminiToolConfig_IncludeServerSideToolInvocations.

Fixes #5709
2026-08-17 15:23:49 +08:00
lyen1688 cb7b03795d feat: 优化分组用量统计 2026-08-14 22:34:43 +08:00
风起 bafd2e293f fix(apicompat): omit empty tool name on streamed arguments deltas
The responses->chat streaming converter re-emits the tool name as an empty
string on every function_call_arguments.delta because ChatFunctionCall.Name
has no omitempty. OpenAI-compatible clients accumulate the name from the
first tool-call delta, so the trailing "name":"" overwrites it and the
call fails dispatch (unknown tool ""). Empty names also break the next
round-trip upstream (missing field 'name').

Fixes #5631
2026-08-14 06:32:38 +00:00
Cornna 21c07e8351 fix(antigravity): use official daily endpoint
Align the daily Cloud Code host with the current Antigravity client so requests do not use the legacy sandbox hostname.
2026-08-13 21:31:42 -07:00
IanShaw027 b61e4bcc45 fix: 对齐 JWT 常量 gofmt,并补全 available groups 契约字段
golangci-lint 要求 const 块按类型对齐;分组 DTO 新增
long_context_pricing_enabled 后,契约夹具未同步导致比较失败。
2026-08-13 08:55:56 +08:00
IanShaw027 c4d883b8da feat: Chat 与 Responses 往返保留 x_search,并补 sources 抽取
Chat Completions 的 {"type":"x_search"} 会被 apicompat 丢掉,独立
/x_search 又没带 include 与结构化提示,主路径容易返回空结果仍计费。

- Chat↔Responses 保留 x_search 过滤字段与 tool_choice
- declared 只注册实际存活的 x_search,web_search 选择项仍丢弃
- 上游补 include x_search_call.action.sources 与结构化输出提示
2026-08-13 08:37:57 +08:00
IanShaw027 363cc4994b fix: SuperGrokPro 用 4.5 窗口区分 Heavy,容量抖动只封单模型
JWT 的 SuperGrokPro 同时覆盖 SuperGrok 与 Heavy,账单月额度又滞后,
账号会误升/误降。multi-agent 容量抖动还会把整号提出调度。

- CanonicalGrokPlan:明确 JWT 优先;模糊档仅采新鲜的 grok-4.5 Responses 窗口
- 配额快照写入 plan_from_45_responses 时间戳,过期信号不用
- engine_overloaded 对 multi-agent 只封当前模型 0.5s–5min
2026-08-13 07:49:14 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
IanShaw027 bb9e74285e feat: 从 JWT tier 识别 Grok 订阅档位,刷新后覆盖失效订阅
Grok Build access token 带有数字 tier claim(0=free、1=supergrok、5=heavy、6=lite),
原先只解 email/sub/team_id,refresh 还会把旧档位抄回去,订阅失效后一直显示 Heavy。

- 解码 JWT tier,并归一化 display / header 别名(含 free-tier、SuperGrok Lite)
- 仅从 access token 读取档位;新 JWT 覆盖已存凭证,AT 无 claim 时保留旧值
- 用量与 free-cache 判定以当前 AT 为准,账单月额度仅作降级
2026-08-13 01:23:37 +08:00
Wesley LiddickandGitHub 6876477371 Merge pull request #5304 from wucm667/fix/issue-5302-chat-reasoning-alias
fix(apicompat): accept chat reasoning alias
2026-08-11 13:59:49 +08:00
Pengap 4d4a0be1ad fix(apicompat): chat/completions file part 不再被静默丢弃,转换为 Responses input_file
/v1/chat/completions → Responses 转换层此前只处理 text 和 image_url 两种
content part,type:"file"(PDF 附件)被静默丢弃:请求返回 200、模型正常
回答,但 prompt 里没有文件(prompt_tokens 只剩纯文字)。

现将 file part 映射为 Responses API 的 input_file(filename/file_data/
file_id 透传),与同网关 /v1/responses + input_file 实测可用的格式一致。
无 file_data 且无 file_id 的空 file part 跳过,与空 image URL 行为一致。

参考上游 PR #2497(因混入无关改动未合并)。
2026-08-10 13:59:20 +08:00
IanShaw027 7eb1310701 fix(grok): close free-by-default billing and related review blockers
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.

M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
2026-08-08 14:39:22 +08:00
IanShaw027 1f58e25ab3 Merge upstream/main into feat/grok-complete-integration
冲突集中在 chat completions / messages 两条 Responses 转发路径:
upstream 给 OpenAIForwardResult 增加了 UpstreamResponseModel 与
UpstreamResponseModelConflict(配套 beginUpstreamResponseModelObservation
观测器),本分支在同样位置把返回值改成了具名变量以便挂 Grok 原生搜索计数。
两侧不互斥,合并结果同时保留上游的响应模型观测字段与 Grok SearchCount 逻辑。

frontend/pnpm-lock.yaml 取 upstream 版本:package.json 与 upstream 完全一致,
本地差异只是 pnpm install 的重解析噪音。
2026-08-08 11:12:56 +08:00
IanShaw027 165b072908 fix(grok): gateway media/voice routing, models, and status UI polish
Align gateway Grok media/voice paths and model lists, harden upstream failure
and quota handling, clear non-Grok video generation config migration, and polish
temp-unsched/status indicators with model whitelist updates.
2026-08-08 01:07:27 +08:00
IanShaw027 6d632eec45 fix(grok): tighten OAuth SSO flow and hide password login
Require oauth state/redirect consistency, fail closed on missing proxy,
and remove password login from create/reauth UI (admin-only password path stays off by default).
2026-08-08 01:07:19 +08:00
Wesley LiddickandGitHub 155c494964 Merge pull request #5399 from fengshao1227/fix/responses-anthropic-invalid-content-blocks
fix(apicompat): Responses→Anthropic 转换不再发出上游会拒收的 content block
2026-08-07 23:21:59 +08:00
li 64090de664 fix(apicompat): Responses→Anthropic 转换不再发出上游会拒收的 content block
convertResponsesInputToAnthropic 的 default 分支把未知 item 的 content 逐字
透传成 Anthropic user 消息,Responses 专有的分片类型会原样进入上游请求体。
最典型的是工具执行后回放的 reasoning item:带 content 数组时,reasoning_text
块直接发给 Anthropic,上游回 400 Request body format invalid,而该 item 会一直
留在会话历史里,导致此后每一轮都继续失败。

同时修正两处会产出空内容消息的路径——Anthropic 拒收空内容消息与空白 text 块:
分片全部不可识别时,user 消息退化成 content:""、assistant 消息退化成单个空
text 块。

改动:
- type=reasoning 显式跳过。Anthropic 无法摄入 OpenAI reasoning:encrypted_content
  不透明,thinking 重放需要上游签发的 signature。Codex 常见形态(只带 summary +
  encrypted_content)本来就会被丢弃,这里让带 content 的形态行为一致。
- default 分支改走 convertResponsesUserToAnthropicContent 白名单转换,保留其中
  可识别的文本/图片,丢弃其余分片。
- user / assistant 分支在转换结果为空内容或纯空白文本时跳过该消息。

Fixes #5329
2026-08-07 23:09:03 +08:00
Brisbanehuang db0bff82c7 feat(usage): audit upstream response models
(cherry picked from commit 839036224f795c8ee5dc6718a2a14372a45eea44)
2026-08-07 09:40:11 -04:00
li 02fbcbe3ad fix(ratelimit): 守卫按端点来源门控,并与冷却键对齐模型口径
上一版守卫只看模型类型,不区分请求从哪个端点进来。OAuth 账号的 /v1/images/*
上游同样是 Codex Responses(openai_images_responses.go → handleOpenAIImagesErrorResponse
→ handleOpenAIAccountUpstreamError → HandleUpstreamModelNotFound),所以专用生图
端点也会命中 plan-gated 分支。账号确实不具备生图能力时跳过冷却,会让调度层失去
唯一的刹车:每个请求都完整走一遍号池,对上游形成无上界的 400 放大。

改动:
- 新增 ctxkey.OpenAIImagesEndpoint 与 WithOpenAIImagesEndpoint /
  OpenAIImagesEndpointFromContext,在 handler/openai_images.go 入口置位;
  与 OpenAIImageGenerationIntent 区分——后者在 /v1/responses 带图片模型时也会置位。
- 守卫下移到 modelKey 计算之后,抽成 shouldSkipCodexPlanGatedImageModelCooldown,
  仅在 plan-gated 分支、且非 /v1/images/* 入站时生效。
- 同时判断 requestedModel 与最终 modelKey:冷却键走 account.GetMappedModel,
  账号可以把文本别名映射到 gpt-image-*,只判请求模型会漏掉这种形态。
2026-08-07 20:49:39 +08:00
IanShaw027 72a56f862c fix(grok): free 500k + 对齐 personal-dev tier/时间窗
- soft-gate 默认额度 500k tokens / 滚动 24h / 95%(门禁 475k)
- soft-gate 仅显式 free OAuth(subscription_tier/plan_type == free)
- isKnownGrokFreeAccount 按 personal-dev(usage% 为 paid 证据、仅 credentials tier)
- 调度阈值仅 grok_sched header quota 窗,去掉 billing 7d/30d 候选
2026-08-07 18:52:53 +08:00
IanShaw027 d7c9e7167b fix(grok): 三轮评审 — 流式 Search 去重与调度/计费加固
- P0: 直播 SSE SearchCount 跨事件 call_id 去重,避免 ~2× 附加费
- SearchCount/Audio/WebSearchCalls 走 mandatory usage task
- web_search:uuid 等 forced request_id 优先于 client/local
- Sanitize 始终剥离 cookie;ApplyOAuth 清 grok_needs_reauth_at
- free 判定:paid 证据压过陈旧 free 凭据
- Gateway 列表应用 free soft-gate;token/body-read 可 failover
- web_search 重试支持 WaitPlan 获取;周 PeriodEnd 不再回填月 end
2026-08-07 17:58:23 +08:00
IanShaw027 a3aae134ee fix(grok): 加固 SSO 凭证与 Cookie 隔离 2026-08-07 16:34:47 +08:00
IanShaw027 25d2b03e90 fix(grok): 加固 OAuth 会话共享与一次性消费 2026-08-07 16:29:48 +08:00
Wesley LiddickandGitHub 32e4de7942 Merge pull request #5231 from fengshao1227/fix/upstream-dial-timeout
fix(upstream): set explicit TCP dial timeout on upstream transports
2026-08-07 16:24:53 +08:00
Wesley LiddickandGitHub 22ef761c27 Merge pull request #5309 from wucm667/fix/issue-5308-antigravity-gemini36
fix(antigravity): support Gemini 3.6 Flash models
2026-08-07 16:14:37 +08:00
IanShaw027 79df1647d4 feat(grok): 账单绝对金额、调度阈值 UI、搜索/Voice 计费与 /v1/web_search
- BillingSummary 输出 prepaid/monthly_used/on_demand 绝对金额,并映射官方 7d/30d 进度条
- 账号编辑页支持 credentials.account_scheduling_threshold 覆盖;Settings 文案细化
- groups.search_price_per_1k + 管理端/缓存全链路;SearchCount/AudioUsage 请求级计费
- Voice TTS/STT/Realtime 成功路径 RecordUsage;独立 /v1/web_search 原生 Grok 搜索
2026-08-07 15:52:55 +08:00