Commit Graph
1454 Commits
Author SHA1 Message Date
shaw bd1ccd9736 Merge remote-tracking branch 'origin/main' into fix/issue-5796-composite-new-platforms 2026-08-19 14:47:34 +08:00
Wesley LiddickandGitHub 7d9c958482 Merge pull request #5810 from Pluviobyte/codex/fix-responses-input-tokens
fix(codex): handle Responses input token preflight
2026-08-19 14:43:33 +08:00
wucm667 b171bb0e4a fix(composite): support CN providers
Extend composite routing, pricing, migrations, and admin options for Kimi, GLM, and DeepSeek.
2026-08-19 13:01:25 +08:00
Kingsley 58e147fba6 feat(composite): support Codex endpoints 2026-08-19 04:40:32 +00:00
Rain bfac49fef9 fix(codex): handle responses input token preflight 2026-08-19 11:34:18 +08:00
Wesley LiddickandGitHub c6f4fbde49 Merge pull request #5676 from Perfecto23/agent/openai-capacity-failover
fix(openai): recover message-only capacity failures before output
2026-08-19 11:11:23 +08:00
Wesley LiddickandGitHub 359fd12b2e Merge pull request #5749 from Randark-JMT/chore/remove-sora-leftovers
chore: remove leftover Sora references after platform removal
2026-08-19 09:41:25 +08:00
Wesley LiddickandGitHub e0c48a19ed Merge pull request #5761 from Randark-JMT/feat/channel-monitor-quota-mode
feat(monitor): 渠道监控配额模式——关联账号展示用量/余额(重启 #5387)
2026-08-18 17:36:27 +08:00
shaw 8f6f459835 fix(channels): support kimi/zhipu/deepseek platforms in channel pricing
- ChannelsView platformOrder now includes the three CN provider
  platforms so channel pricing can be configured for them; composite
  group expansion/attachment stays limited to the five main platforms
  (matches backend isConcreteRequestPlatform and composite-routes
  target_platform validation)
- SyncPricingModels maps kimi->moonshot, zhipu->zhipu,
  deepseek->deepseek; also fixes gemini mapping to "google" which
  matched zero catalog entries (provider key is "gemini")
- Add CN platform colors to channel pricing model tag classes
2026-08-18 17:12:01 +08:00
Randark 41344c20ff feat(monitor): wire quota fetcher & expose check_mode in handlers
- handler DTO: create/update 接收 check_mode/account_id,provider oneof 扩至 8 家,
  endpoint/api_key 改为 omitempty(条件必填下沉 service 校验);
  monitor/checkResult/historyItem 响应透传 check_mode/account_id/quota
- 用户端 latest_quota 由 channel_monitor_show_quota 控制,关闭时服务端剥离
- wire: NewChannelMonitorQuotaFetcher 以具体服务类型收参(窄接口包内保留供
  stub),ProvideChannelMonitorRunner 注入后 SetQuotaFetcher
2026-08-18 04:14:15 +00:00
Randark 6a6fd304f6 feat(settings): channel_monitor_show_quota public setting (default off)
- 新增公开设置 key(迁移 226 已插入默认 false),控制用户端监控页
  是否展示配额/余额;管理端不受影响
- 解析 fail-closed:仅字面 "true" 视为开启(对齐 available_channels_enabled
  语义,而非 hide_throughput 的 fail-open)
- 全链路贯通:domain key 常量 → setting_public(公开读取 + ChannelMonitorRuntime
  .ShowQuota + 注入 payload)→ settings_view/parse/update → handler DTO
  (admin/user 响应 + admin 更新请求)→ api_contract_test 两个 wantJSON 块
2026-08-18 04:06:18 +00:00
Randark 7e45634df9 chore: remove leftover Sora references after platform removal
PR #1463 removed the Sora platform, but some references survived:

- README/README_CN/README_JA kept the 'Sora status (temporarily
  unavailable)' sections and gateway.sora_* docs that the removal PR
  never touched.
- deploy/config.example.yaml still documented ~130 lines of sora_*
  gateway keys, the top-level sora: direct-client/storage block, and
  token_refresh.sync_linked_sora_accounts - none of which map to any
  field in the config structs anymore.
- The OIDC login PR (02a66a01c, branched off pre-removal main and
  merged 4 days after #1463) re-added the dead
  PublicSettings.SoraClientEnabled field, which no code ever sets.
- A release sync (748a84d87) re-introduced sora i18n keys that the
  later i18n split (d9e514f98) faithfully carried into
  locales/{zh,en}/admin/{overview,settings}.ts. No component references
  any of these keys.

This drops all of the above. Pure deletions, no behavior change.
2026-08-18 00:51:30 +00:00
shaw 6bf335965a merge main 并修复与 #5730 的语义冲突
main 侧 #5730 新增的 openai_gateway_cn_fixes_test.go 按旧 11 参签名调用
calculateOpenAIRecordUsageCost;本分支为该函数新增了第 12 个参数
pricingAt。文本无冲突但 test build 会失败,此处按本分支对同类测试
调用点的既有处理方式补传 time.Time{}。
2026-08-17 22:27:22 +08:00
lyen1688 9f24a55305 功能:支持渠道模型分时倍率定价 2026-08-17 19:45:07 +08:00
shaw 10c8b70203 fix(cn-providers): 修复 CN 分组五项功能缺陷(调度闸门/计费/断开漏记/count_tokens/403)
对已合并 PR #5666 + 分组入口放行后的全量功能审计发现的 P0/P1 修复,
全部先经代码与厂商文档实证再实施:

1. /v1/messages 调度闸门(P0):sanitizeGroupMessagesDispatchFields 对非
   openai 平台恒置 AllowMessagesDispatch=false,而闸门豁免名单只有 grok,
   CN 分组经正常途径创建后恒 403——原生 Anthropic 直通(Claude Code 主用例)
   完全不可达。修复:闸门对 CN 与 grok 同语义豁免;count_tokens 处的内联
   裸检查统一走同一 helper;ResolveMessagesDispatchModel 对 CN 早退,避免
   openai 专属的 gpt-5.x 默认映射发给 CN 上游。

2. 计费候选链(P0):候选链兜底含客户端原始模型名,配合 getFallbackPricing
   的 claude→Sonnet 统一兜底,映射的 CN 模型无价时 CN 流量会按 Claude 原价
   (数倍~数十倍)静默误计,且 usage 日志显示 claude-* 名无从察觉。修复:
   CN 账号的 claude-* 候选仅在显式分组/渠道定价时放行;候选全滤空时按
   ErrModelPricingUnavailable 走零成本+告警落账(顺带修复原空候选错误会
   丢弃整条 usage 记录的次生问题)。

3. 断开/中断漏记(P0,#5148 对齐,惠及 openai 平台):messages/responses/
   chat_completions 三个 handler 的错误路径此前在 err!=nil 时丢弃携带的
   部分 result——客户端断开排水后的完整 usage 被丢,payg 上游照常计费而
   平台漏记(anthropic 网关早有同修复,openai 网关缺失)。修复:错误路径
   result 非空时照常提交 usage;failover 错误恒 result=nil 无重复计费。
   Responses×anthropic 流式转换器同时改为断开后继续排水至流自然结束
   (末尾 message_delta 的 output_tokens 不再丢),finalize 帧补工具名反转
   与客户端工具还原、仅在客户端仍连接时写出。

4. count_tokens(P1,证据修正):经实证三家 Anthropic 兼容层均无
   /v1/messages/count_tokens(DeepSeek 官方文档无此端点且注明
   anthropic-version 被忽略;OpenModel 标注该端点 Anthropic only),
   anthropic 协议转发上游=常态 404,且错误处置缺模型上下文会把不计费的
   探测放大成整账号停调。修复:CN 全协议一律本地 tiktoken 估算(与 Grok
   同方案),删除上游转发死代码。

5. 403 处置(P1):CN 此前落入通用 handleAuthError,单次 HTML 403(CDN/
   代理拦截页)即永久禁用,且 403 在 failover 集里会逐账号重放连环禁用
   整组。修复:CN 与 openai 同口径——HTML 豁免 + 3 次累计 + 临时冷却。

新增回归测试 8 项:闸门豁免(含 openai 仍受控断言)、CN 调度映射空返回、
候选过滤三态、空候选零成本落账、断开排水 usage 完整性、HTML-403 零处罚、
结构化 403 首次临时停调。handler/service 全包测试通过。
2026-08-17 17:28:53 +08:00
shaw 7cdca9e495 feat(groups): 放行 kimi/zhipu/deepseek 平台分组创建入口
PR #5666 引入 CN 平台后,路由/调度/前端类型均已支持 CN 平台分组,但分组
创建入口两头缺失:后端 Create/UpdateGroupRequest 的 platform oneof 白名单
与前端 GroupsView 平台选项都没有三平台,导致 CN 账号「无可用分组」、整条
流量链路不通(composite 不能作为替代:CN 不可为 composite 路由目标)。

- group_handler.go: 两处 oneof 加 kimi/zhipu/deepseek;composite 路由目标
  白名单有意不动(DetectModelPlatform/isConcreteRequestPlatform 均无 CN 分支)
- GroupsView: platformOptions/platformFilterOptions 补三项;两处徽章配色链
  按 platformColors.ts 色系补 CN 分支
- i18n: admin.groups.platforms 补 kimi/zhipu/deepseek 键(zh/en),缺键时
  分组徽章/GroupRPM/RateMultipliers 弹窗/ChannelsView 会渲染原始 key
- GroupBadge: badgeClass/labelClass 补 CN 配色
- 新增表驱动测试:9 平台 Create/Update 全放行、非法值(别名/大小写/空格)
  全拒绝、composite target 对 CN 保持拒绝的守卫
2026-08-17 17:28:51 +08:00
Perfecto c3063e01a4 fix(openai): recover message-only capacity failures 2026-08-15 21:30:56 +08:00
Randark 901a0439f1 feat: 国产供应商一等支持(Kimi/Zhipu/DeepSeek 多协议 + 配额/余额监控)
后端:
- 协议凭证维度 credentials[api_protocol] ∈ chat_completions(默认)/anthropic/responses(deepseek)
- /v1/messages 零转换直通原生 Anthropic 端点(kimi/zhipu/deepseek),CC/Responses
  入站交叉组合走 apicompat 双向转换链(responses/chat_completions anthropic-native 转发器)
- count_tokens:anthropic 协议透传原生端点;其余 CN 协议本地 tiktoken 估算
- Coding Plan 额度探测(5h/weekly 滚动窗口)+ payg 余额探测(kimi/deepseek),
  deepseek 双币种 CNY+USD 明细,任一币种达标不停调
- 周期任务 [CNBalance] 并发探测 + 预算随工作量放大;响应式 429 冷却到最早窗口
  重置点;余额不足可恢复临时停调;智谱 CREDIT_LIMIT 不污染窗口解析
- CC→anthropic 流式客户端断开后继续排水上游保住 usage 计量

前端:
- 创建/编辑弹窗 account_mode + api_protocol + base_url 联动预设(含 watcher 竞态防护)
- 用量单元格:kimi/zhipu coding 显示 5h/weekly 窗口,kimi/deepseek payg 显示余额,
  多币种并列展示;探测失败保留快照;挂载自动探测 5min 去抖
- 调度阈值设置面板补 kimi/zhipu 平台(对齐后端 AllowedSchedulingThresholdPlatforms)
2026-08-15 10:37:51 +00:00
shaw 8ae6d8f67e fix(openai): send session-level beta features and probe native compaction v2
OpenAI sunset the legacy unary /responses/compact endpoint (404, #5598,
#5624), so the account "compact probe" in the admin UI kept failing even for
healthy accounts, and the beta-feature negotiation header was only attached
to compaction turns.

Beta features (codex-rs session/mod.rs build_model_client_beta_features_header
+ client.rs build_responses_headers): the header is a session-level constant
attached to every /responses request, the WS handshake and /responses/compact.
Enumerating FEATURES shows no Experimental feature is enabled by default, so a
default install sends exactly "remote_compaction_v2". Mirror that:

- OAuth requests without a client-declared header get the default shape, so we
  no longer produce a "header only on compaction turns" pattern real Codex
  never emits (#5586 chains that strip the header)
- a client-declared header is preserved as-is: non-empty without v2 means the
  user disabled the feature and the gateway must not rewrite that
- native v2 turns (compaction_trigger in body) always ensure v2 is present
- non-OAuth upstreams keep the compaction-turn-only behaviour
- the WS injection sits outside the client-header copy block so prewarm and
  turn handshakes cannot land in different pool compatibility buckets

Compact probe now exercises native v2 (streaming /responses +
compaction_trigger) instead of the dead endpoint. Success requires an actual
compaction output item — scanning output_item.done/added, the terminal
response.output[] and the whole-JSON fallback — so a 2xx that silently drops
the trigger is reported as unsupported (the "got 0 items" class, #5478,
#5648). Probe identity is now UUID-shaped and applies the account's
convergence, matching real traffic on the same endpoint.
2026-08-15 16:35:08 +08:00
Wesley LiddickandGitHub 1d3b9665c8 Merge pull request #5641 from InCerryGit/fix/issue-5624-remote-compaction-v2
fix(openai): preserve remote compaction v2 responses endpoint
2026-08-15 13:46:31 +08:00
lyen1688 cb7b03795d feat: 优化分组用量统计 2026-08-14 22:34:43 +08:00
InCerryGit a8b9ea22b7 fix(openai): separate native and legacy compaction routing 2026-08-14 18:18:53 +08:00
InCerryGit 9662cff2e7 fix(openai): preserve remote compaction v2 responses endpoint 2026-08-14 16:23:52 +08:00
IanShaw027 678eb22a40 fix: Realtime 仅在观察到音频后计费,并修正标志位求值顺序
原先 elapsed>0 就出账,握手失败也会扣费;随后用 audioObserved,
但又在同一个 return 里先 Load 再收 errCh,标志位恒为 false,会话全部漏计。

- 先等中继结束再读 audioObserved
- 无音频或零时长不出账;每次连接独立 request id
2026-08-13 08:38:10 +08:00
IanShaw027 c4d883b8da feat: Chat 与 Responses 往返保留 x_search,并补 sources 抽取
Chat Completions 的 {"type":"x_search"} 会被 apicompat 丢掉,独立
/x_search 又没带 include 与结构化提示,主路径容易返回空结果仍计费。

- Chat↔Responses 保留 x_search 过滤字段与 tool_choice
- declared 只注册实际存活的 x_search,web_search 选择项仍丢弃
- 上游补 include x_search_call.action.sources 与结构化输出提示
2026-08-13 08:37:57 +08:00
IanShaw027 0de6d7e9ba feat: 新增独立 /x_search,走原生 x_search 并沿用搜索计费
Grok 分组原先只有 /web_search,无法带 handle/日期过滤,也无法走
xAI 的 x_search tool。

- POST /x_search(仅 Grok 分组),复用 web_search 的审计、failover 与按次计费
- 上游 Responses 强制 x_search;计费模型记为 grok-x-search
2026-08-13 07:49:19 +08:00
IanShaw027 f3d9491071 feat: 分组支持逐模型定价,并可关闭长上下文阶梯
运营需要按分组覆盖渠道/内置价,且部分套餐不应自动吃 200k 倍率。
原先只能改渠道价卡,分组侧只剩 Voice 三列。

- groups 新增 model_pricing / long_context_pricing_enabled,解析链改为 Group → Channel → 内置
- 关闭长上下文时 token 模型只取最低档;video 按秒计费可写进同一价卡
- 回退价对齐官方卡:4.5 缓存 $0.30、4.3/imagine/audio/search 默认值一并校正
2026-08-13 07:49:08 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
Wesley LiddickandGitHub a29fce4a61 Merge pull request #5511 from wucm667/fix/pr-5234-ws-audit-logging
fix(security-audit): restore websocket audit logs
2026-08-12 09:58:24 +08:00
shaw a3bbf35cbd Merge branch 'main' into fix/issue-5029-openai-passthrough-pool-auth-retry
Resolve conflict in backend/internal/handler/openai_gateway_handler_test.go.

main and this branch each appended a passthrough upstream stub plus a test at
the same two insertion points:

  main   openAIHTTPPassthroughSSERateLimitUpstream
         TestOpenAIResponses_APIKeyPassthroughSSERateLimitUsesConfiguredPoolRetry
  branch openAIHTTPPassthroughAuthFailoverUpstream
         TestOpenAIResponses_APIKeyPassthroughPoolAuthFailureRetriesThenSwitchesToHealthyAccount

Both sides are kept verbatim; the only edit is giving each stub its own
calls() body instead of sharing the trailing one. No assertion was changed.

openai_gateway_passthrough.go and openai_oauth_passthrough_test.go merged
automatically.
2026-08-11 14:09:07 +08:00
Wesley LiddickandGitHub b918874f81 Merge pull request #5403 from cyhhao/fix/codex-capacity-exponential-backoff
fix(openai): back off capacity retries exponentially
2026-08-11 13:58:06 +08:00
wucm667 2d9920ba7d fix(security-audit): restore websocket audit logs 2026-08-11 11:56:46 +08:00
pigzwyandshaw 9096492b55 feat(billing): support safe upstream response model billing 2026-08-10 18:45:14 +08:00
Wesley LiddickandGitHub 10a4c6e3ad Merge pull request #5234 from wucm667/fix/issue-5230-deduplicate-latest-turn-audit
fix(security-audit): deduplicate websocket turn audits
2026-08-10 10:53:17 +08:00
Wesley LiddickandGitHub f3c7a1a8c4 Merge pull request #5295 from wucm667/fix/issue-5289-streaming-upstream-error
fix: emit response.failed when compact keepalive commits headers but no SSE payload
2026-08-10 10:52:49 +08:00
Wesley LiddickandGitHub 30d0405388 Merge pull request #5464 from wucm667/fix/issue-5455-api-key-input-validation
fix(api-key): validate quota and expiry inputs
2026-08-10 10:52:04 +08:00
wucm667 f5c108c836 fix(api-key): validate quota and expiry inputs 2026-08-09 22:42:51 +08:00
lyen1688andlyen1688 bbc8b6e906 完善大文件备份分卷上传与恢复 2026-08-09 20:58:07 +08:00
shaw 563a72ca73 feat: add default-off switch for email domain registration quota
PR #5423 relaxed the email suffix whitelist: once a whitelist is
configured, non-whitelisted registrable domains are each allowed to
register one account. That behavior activated unconditionally.

Add registration_email_domain_quota_enabled (default false) to gate it:

- Off (default): restore pre-#5423 strict whitelist semantics — with a
  non-empty whitelist, non-whitelisted domains are rejected with
  EMAIL_SUFFIX_NOT_ALLOWED; the register/verify views restore the
  client-side whitelist pre-check and allowed-domain hint.
- On: keep #5423 behavior — one account per non-whitelisted registrable
  domain (EMAIL_DOMAIN_REGISTRATION_LIMIT).
- Empty whitelist keeps allowing all domains in both states.

Gating lives in validateRegistrationEmailQuota and (as a race-safety
backstop) createUserWithRegistrationEmailGuard; the repository-level
domain lock + in-tx recheck is unchanged. The admin update field is
*bool (omitted = keep current) so stale full-payload saves cannot
silently flip the switch. Email binding and OAuth auto-signup keep
their strict policy, and pending-OAuth bind-login for existing
accounts is unaffected because the handler resolves existing emails
before the quota check.

Frontend adds the toggle to admin settings (zh/en copy; whitelist hint
restored to strict wording, quota wording moved to the new toggle) and
exposes the flag via public settings + SSR injection payload.

Tests: #5423 quota tests now enable the switch explicitly; new
default-off regression tests cover register/send-code/async/pending
OAuth/OIDC create-account plus both register views; API contract JSON
and the injection drift guard are updated.
2026-08-09 15:53:40 +08:00
Wesley LiddickandGitHub f2da30bcd9 Merge pull request #5423 from lyen1688/feat/email-domain-registration-quota
完善邮箱域名注册额度策略
2026-08-09 15:16:26 +08:00
shaw d92edc01be Merge origin/main into feat/channel-monitor-v2-ops-ui
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.

- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
  (V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
  GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
  save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
  and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
  218/219/220 rules (#5408).
2026-08-09 12:11:35 +08:00
lyen1688andlyen1688 4999231d61 修复邮箱域名注册额度策略 2026-08-08 21:05:20 +08:00
IanShaw027 7eb1310701 fix(grok): close free-by-default billing and related review blockers
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.

M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
2026-08-08 14:39:22 +08:00
IanShaw027 cec922d335 fix(grok): clear golangci-lint findings on complete-integration branch
Check Close/CloseNow errors, drop unused helpers and dead constants,
lowercase ST1005 error strings, and stop discarding unwrap status as an
unused assignment so CI golangci-lint passes.
2026-08-08 12:56:40 +08:00
IanShaw027 1f58e25ab3 Merge upstream/main into feat/grok-complete-integration
冲突集中在 chat completions / messages 两条 Responses 转发路径:
upstream 给 OpenAIForwardResult 增加了 UpstreamResponseModel 与
UpstreamResponseModelConflict(配套 beginUpstreamResponseModelObservation
观测器),本分支在同样位置把返回值改成了具名变量以便挂 Grok 原生搜索计数。
两侧不互斥,合并结果同时保留上游的响应模型观测字段与 Grok SearchCount 逻辑。

frontend/pnpm-lock.yaml 取 upstream 版本:package.json 与 upstream 完全一致,
本地差异只是 pnpm install 的重解析噪音。
2026-08-08 11:12:56 +08:00
IanShaw027 f3bac4619e feat(keys): expand Grok client samples and tune free soft-gate default
Ship Use Key templates that match Grok Build / Codex best practice: env
vars + multi-model config.toml with api_backend=responses, env_key over
hardcoded secrets, and clearer shell/path guidance for Claude/Codex/OpenCode.

Also set free_quota_token_limit default to 500k (24h soft-gate), clean up
personal-dev-only comments, and keep billing test fixtures aligned.
2026-08-08 10:26:16 +08:00
IanShaw027 e01ce90d47 fix(grok): harden voice request ids, video pending, and search pricing alerts
Mint durable grok_audio/grok_realtime usage ids, avoid CLI headers on api.x.ai
voice, retry video pending store and fail-closed when snapshot is missing without
status duration, align pure-video ImageCount tests, and escalate unset search
price_per_1k to error-level logs.
2026-08-08 09:45:12 +08:00
IanShaw027 12db0f906a fix(grok): drop account-test ZDR path and align media CLI headers
Remove optional upload_url / fake connectivity-only success from admin video
tests. Stamp Grok CLI headers only on the CLI proxy so OAuth media against
api.x.ai can complete and preview video like the gateway path.
2026-08-08 09:45:12 +08:00
IanShaw027 35faaa6d21 feat(grok): register custom-voices CRUD and audio download gateway routes
Forward list/get/patch/delete and reference-audio paths with safe path segment
encoding, method passthrough, and empty-body GET/DELETE handling.
2026-08-08 08:48:38 +08:00
IanShaw027 85b65284ec fix(grok): set async video duration_ms from create accept to done discovery
Store CreatedAt on pending billing at video create and use wall-clock E2E
latency when status/content first observes official done+video.url, so usage
logs no longer record only the single poll hop.
2026-08-08 08:48:38 +08:00