Commit Graph
715 Commits
Author SHA1 Message Date
shaw 6bf335965a merge main 并修复与 #5730 的语义冲突
main 侧 #5730 新增的 openai_gateway_cn_fixes_test.go 按旧 11 参签名调用
calculateOpenAIRecordUsageCost;本分支为该函数新增了第 12 个参数
pricingAt。文本无冲突但 test build 会失败,此处按本分支对同类测试
调用点的既有处理方式补传 time.Time{}。
2026-08-17 22:27:22 +08:00
lyen1688 9f24a55305 功能:支持渠道模型分时倍率定价 2026-08-17 19:45:07 +08:00
shaw 7cdca9e495 feat(groups): 放行 kimi/zhipu/deepseek 平台分组创建入口
PR #5666 引入 CN 平台后,路由/调度/前端类型均已支持 CN 平台分组,但分组
创建入口两头缺失:后端 Create/UpdateGroupRequest 的 platform oneof 白名单
与前端 GroupsView 平台选项都没有三平台,导致 CN 账号「无可用分组」、整条
流量链路不通(composite 不能作为替代:CN 不可为 composite 路由目标)。

- group_handler.go: 两处 oneof 加 kimi/zhipu/deepseek;composite 路由目标
  白名单有意不动(DetectModelPlatform/isConcreteRequestPlatform 均无 CN 分支)
- GroupsView: platformOptions/platformFilterOptions 补三项;两处徽章配色链
  按 platformColors.ts 色系补 CN 分支
- i18n: admin.groups.platforms 补 kimi/zhipu/deepseek 键(zh/en),缺键时
  分组徽章/GroupRPM/RateMultipliers 弹窗/ChannelsView 会渲染原始 key
- GroupBadge: badgeClass/labelClass 补 CN 配色
- 新增表驱动测试:9 平台 Create/Update 全放行、非法值(别名/大小写/空格)
  全拒绝、composite target 对 CN 保持拒绝的守卫
2026-08-17 17:28:51 +08:00
Randark 901a0439f1 feat: 国产供应商一等支持(Kimi/Zhipu/DeepSeek 多协议 + 配额/余额监控)
后端:
- 协议凭证维度 credentials[api_protocol] ∈ chat_completions(默认)/anthropic/responses(deepseek)
- /v1/messages 零转换直通原生 Anthropic 端点(kimi/zhipu/deepseek),CC/Responses
  入站交叉组合走 apicompat 双向转换链(responses/chat_completions anthropic-native 转发器)
- count_tokens:anthropic 协议透传原生端点;其余 CN 协议本地 tiktoken 估算
- Coding Plan 额度探测(5h/weekly 滚动窗口)+ payg 余额探测(kimi/deepseek),
  deepseek 双币种 CNY+USD 明细,任一币种达标不停调
- 周期任务 [CNBalance] 并发探测 + 预算随工作量放大;响应式 429 冷却到最早窗口
  重置点;余额不足可恢复临时停调;智谱 CREDIT_LIMIT 不污染窗口解析
- CC→anthropic 流式客户端断开后继续排水上游保住 usage 计量

前端:
- 创建/编辑弹窗 account_mode + api_protocol + base_url 联动预设(含 watcher 竞态防护)
- 用量单元格:kimi/zhipu coding 显示 5h/weekly 窗口,kimi/deepseek payg 显示余额,
  多币种并列展示;探测失败保留快照;挂载自动探测 5min 去抖
- 调度阈值设置面板补 kimi/zhipu 平台(对齐后端 AllowedSchedulingThresholdPlatforms)
2026-08-15 10:37:51 +00:00
lyen1688 cb7b03795d feat: 优化分组用量统计 2026-08-14 22:34:43 +08:00
IanShaw027 f3d9491071 feat: 分组支持逐模型定价,并可关闭长上下文阶梯
运营需要按分组覆盖渠道/内置价,且部分套餐不应自动吃 200k 倍率。
原先只能改渠道价卡,分组侧只剩 Voice 三列。

- groups 新增 model_pricing / long_context_pricing_enabled,解析链改为 Group → Channel → 内置
- 关闭长上下文时 token 模型只取最低档;video 按秒计费可写进同一价卡
- 回退价对齐官方卡:4.5 缓存 $0.30、4.3/imagine/audio/search 默认值一并校正
2026-08-13 07:49:08 +08:00
pigzwyandshaw 9096492b55 feat(billing): support safe upstream response model billing 2026-08-10 18:45:14 +08:00
lyen1688andlyen1688 bbc8b6e906 完善大文件备份分卷上传与恢复 2026-08-09 20:58:07 +08:00
shaw 563a72ca73 feat: add default-off switch for email domain registration quota
PR #5423 relaxed the email suffix whitelist: once a whitelist is
configured, non-whitelisted registrable domains are each allowed to
register one account. That behavior activated unconditionally.

Add registration_email_domain_quota_enabled (default false) to gate it:

- Off (default): restore pre-#5423 strict whitelist semantics — with a
  non-empty whitelist, non-whitelisted domains are rejected with
  EMAIL_SUFFIX_NOT_ALLOWED; the register/verify views restore the
  client-side whitelist pre-check and allowed-domain hint.
- On: keep #5423 behavior — one account per non-whitelisted registrable
  domain (EMAIL_DOMAIN_REGISTRATION_LIMIT).
- Empty whitelist keeps allowing all domains in both states.

Gating lives in validateRegistrationEmailQuota and (as a race-safety
backstop) createUserWithRegistrationEmailGuard; the repository-level
domain lock + in-tx recheck is unchanged. The admin update field is
*bool (omitted = keep current) so stale full-payload saves cannot
silently flip the switch. Email binding and OAuth auto-signup keep
their strict policy, and pending-OAuth bind-login for existing
accounts is unaffected because the handler resolves existing emails
before the quota check.

Frontend adds the toggle to admin settings (zh/en copy; whitelist hint
restored to strict wording, quota wording moved to the new toggle) and
exposes the flag via public settings + SSR injection payload.

Tests: #5423 quota tests now enable the switch explicitly; new
default-off regression tests cover register/send-code/async/pending
OAuth/OIDC create-account plus both register views; API contract JSON
and the injection drift guard are updated.
2026-08-09 15:53:40 +08:00
shaw d92edc01be Merge origin/main into feat/channel-monitor-v2-ops-ui
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.

- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
  (V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
  GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
  save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
  and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
  218/219/220 rules (#5408).
2026-08-09 12:11:35 +08:00
IanShaw027 1f58e25ab3 Merge upstream/main into feat/grok-complete-integration
冲突集中在 chat completions / messages 两条 Responses 转发路径:
upstream 给 OpenAIForwardResult 增加了 UpstreamResponseModel 与
UpstreamResponseModelConflict(配套 beginUpstreamResponseModelObservation
观测器),本分支在同样位置把返回值改成了具名变量以便挂 Grok 原生搜索计数。
两侧不互斥,合并结果同时保留上游的响应模型观测字段与 Grok SearchCount 逻辑。

frontend/pnpm-lock.yaml 取 upstream 版本:package.json 与 upstream 完全一致,
本地差异只是 pnpm install 的重解析噪音。
2026-08-08 11:12:56 +08:00
IanShaw027 12db0f906a fix(grok): drop account-test ZDR path and align media CLI headers
Remove optional upload_url / fake connectivity-only success from admin video
tests. Stamp Grok CLI headers only on the CLI proxy so OAuth media against
api.x.ai can complete and preview video like the gateway path.
2026-08-08 09:45:12 +08:00
IanShaw027 68faeac837 fix(grok): restore base URL resolution and operator settings wiring
Honor account GetGrokBaseURLOr policy for official vs custom endpoints,
and wire settings resolution used by responses/chat URL builders.
2026-08-08 01:07:19 +08:00
IanShaw027 6d632eec45 fix(grok): tighten OAuth SSO flow and hide password login
Require oauth state/redirect consistency, fail closed on missing proxy,
and remove password login from create/reauth UI (admin-only password path stays off by default).
2026-08-08 01:07:19 +08:00
IanShaw027 d0767eab9d feat(grok): admin account test modes with real media preview
Add mode-first connectivity probes for text/image/video/search/tts/stt/realtime,
standalone voice and web-search paths, media upload options, and in-browser
image/audio/video preview (including ZDR-safe b64 images and edit validation).
2026-08-08 01:07:08 +08:00
Brisbanehuang db0bff82c7 feat(usage): audit upstream response models
(cherry picked from commit 839036224f795c8ee5dc6718a2a14372a45eea44)
2026-08-07 09:40:11 -04:00
IanShaw027 d7c9e7167b fix(grok): 三轮评审 — 流式 Search 去重与调度/计费加固
- P0: 直播 SSE SearchCount 跨事件 call_id 去重,避免 ~2× 附加费
- SearchCount/Audio/WebSearchCalls 走 mandatory usage task
- web_search:uuid 等 forced request_id 优先于 client/local
- Sanitize 始终剥离 cookie;ApplyOAuth 清 grok_needs_reauth_at
- free 判定:paid 证据压过陈旧 free 凭据
- Gateway 列表应用 free soft-gate;token/body-read 可 failover
- web_search 重试支持 WaitPlan 获取;周 PeriodEnd 不再回填月 end
2026-08-07 17:58:23 +08:00
IanShaw027 245d069602 fix(grok): 二轮评审残留 — Search 叠加计费与 fail-closed 安全
- SearchCost 叠加 token(openai/gateway),未定价 warn
- Token URL 校验失败回落 DefaultTokenURL,禁止 Effective 旁路
- free 判定收窄 paidSignal(仅 plan/月额度),usage% 不否决 free
- web_search: mandatory 计费、uuid request_id、上游 failover 重选账号
- 调度阈值:7d/30d 不跨期 until + 48h stale 可选跳过
- SanitizeStoredCredentials 接入 create/update/bulk/SSO/ApplyOAuth
- ApplyOAuth 成功清除 grok_needs_reauth;SSO 允许 header_override_enabled
- VideoModelPrices 视为媒体定价完整;realtime 正常关闭仍计费
2026-08-07 17:40:21 +08:00
IanShaw027 7a81468282 fix(grok): 按 review 优先级修复计费漏扣、auth 投影与 free 门禁
P0: content 与 status 共用 claim 计费;稳定 grok-video request_id;
auth 热路径投影 video_model_prices。
P1: claim 失败释放可重试;视频绑定 TTL≥24h;SSO 凭据白名单与脱敏;
Token URL 校验;free soft-gate 默认 1M 并统一 free 判定;
Voice 预检余额;STT 抗低报;Realtime 失败不计费;search 未定价告警。
2026-08-07 17:15:43 +08:00
IanShaw027 856a96a217 test(grok): 对齐配额模型同步与 CLI 身份断言 2026-08-07 16:55:36 +08:00
IanShaw027 962da308cc fix(grok): 限制导入探测队列并去重任务 2026-08-07 16:36:12 +08:00
IanShaw027 79df1647d4 feat(grok): 账单绝对金额、调度阈值 UI、搜索/Voice 计费与 /v1/web_search
- BillingSummary 输出 prepaid/monthly_used/on_demand 绝对金额,并映射官方 7d/30d 进度条
- 账号编辑页支持 credentials.account_scheduling_threshold 覆盖;Settings 文案细化
- groups.search_price_per_1k + 管理端/缓存全链路;SearchCount/AudioUsage 请求级计费
- Voice TTS/STT/Realtime 成功路径 RecordUsage;独立 /v1/web_search 原生 Grok 搜索
2026-08-07 15:52:55 +08:00
IanShaw027 99ad01f6de feat(grok): 网关 Voice TTS/STT/Realtime 与分组音频定价
- 中继 xAI Voice HTTP(/tts /stt /custom-voices)与 Realtime WS(/realtime)
  仅 Grok 分组可用;账号选调度沿用 chat 能力,不依赖创作中心
- Voice URL 固定走 api.x.ai(CLI proxy base 自动回落)
- 分组级 audio_realtime / TTS / STT 单价:schema + migration 218 +
  admin API + GroupsView 表单与 i18n(无创作中心 UI)
2026-08-07 15:08:39 +08:00
IanShaw027 7c62382d04 feat(grok): 吸收调度阈值、配额解析、批量用量与 CLI 身份
从 personal-dev 取优并接线到 feat/grok-complete-integration:

- 调度阈值:grok_sched_* 写入、ApplyAccountSchedulingThreshold、
  account_scheduling_thresholds 设置与 Settings UI、网关过滤
- 配额头:x-rate-limit 别名、相对秒 reset、tier/entitlement 扩展
- 批量 usage:GetUsageBatch + POST /usage/batch + AccountsView 合并加载
- CLI 身份:xai.cli_identity 统一 pin/UA;service 层 applyGrokCLIHeaders
  与 transport 最终改写对齐;media eligibility 模块落地
- UsageCell:月度 cents→USD 与 30d 进度条展示

保留既有 failure taxonomy / sticky / free recovery 等 HEAD cools。
2026-08-07 14:53:18 +08:00
IanShaw027 d0930c4bdb fix(grok): 完善密码与SSO授权能力控制 2026-08-07 14:13:07 +08:00
IanShaw027 eb6c9663e7 feat(grok): 视频按模型族配置每秒单价,并补齐管理端映射设置
分组新增 video_model_prices(JSONB);计费优先模型×分辨率覆盖,
其次旧分辨率三列,最后官方按模型族默认价。同步补齐 admin 设置
中的 grok_default_text_model / 跨客户端映射开关,并保持
grok-imagine-video-1.5 请求模型 identity 不被静默改写。
2026-08-07 12:44:08 +08:00
IanShaw027 370bdcf695 feat(grok): 补齐密码登录与 SSO 校验,统一 OAuth 凭证形态
在 main 既有 SSO→Build 批量导入之上增加 sso-token 校验与账号密码授权。
密码仅用于换取 SSO 再转 Build OAuth,明文与 raw SSO 均不落库。
2026-08-07 12:38:49 +08:00
IanShaw027 a5beecb92a feat(channel-monitor-v2): 接入模式开关、路由门控与依赖注入
注册 admin/user 路由与 feature/mode 守卫,串联 Wire DI 与设置读写,
公开 channel_monitor_mode 与 hide_throughput 等运行时标志。
2026-08-07 11:05:32 +08:00
shaw 8e102b3a0f fix: 完善腾讯验证码区域适配与 CSP 白名单
修复国内站和国际站 SDK 构造、验证容器、票据重置及动态资源加载问题,并补充认证流程回归测试。
2026-08-06 20:34:37 +08:00
feeeei 26e0a89323 人机验证增加阿里云验证码 2.0
沿用腾讯天御验证码引入的多服务商模型:aliyun_captcha_enabled 作为独立
开关,与 Cloudflare Turnstile、腾讯天御三方互斥(保存校验 + 运行时
CAPTCHA_PROVIDER_CONFLICT)。后台「安全与认证」合并为单张人机验证卡片:
总开关 + 服务商单选(Turnstile / 腾讯天御 / 阿里云),选中即启用该家并
关闭其它,落库仍是三个独立开关键,由前端映射保证互斥。

阿里云侧同时支持 aliyun 中国站与国际站(alibabacloud.com):两站前端脚本、
region 取值与服务端 API 完全一致,仅账号与 AccessKey 相互独立,因此由
「服务地域」决定线路即可——中国内地走 captcha.cn-shanghai.aliyuncs.com,
非中国内地(新加坡)走 captcha.ap-southeast-1.aliyuncs.com,AccessKey
取自持有该实例的账号,无需在配置中区分站点。

- AliyunCaptchaService 对称 TencentCaptchaService:服务端校验走官方 SDK
  VerifyIntelligentCaptcha,调用异常按 fail-closed 拦截,与 Turnstile
  网络错误行为对称;保存设置时真实探测 AK/SK 有效性
- 保护面对齐腾讯扩展入口:VerifyTencentCaptchaIfEnabled 通用化为
  VerifyActionCaptchaIfEnabled,OAuth 登录启动、passkey 登录在阿里云
  启用时同样拦截;Turnstile 维持既有覆盖不扩大
- 前端 AliyunCaptchaWidget 为表单内预验证按钮(popup 模式),同时暴露
  verify() 供 OAuth 启动、passkey 等动作入口程序化弹窗;未预验证直接
  提交时弹窗兜底。SDK 按钮绑定异步完成,弹窗未出现前按 tick 重试触发,
  并轮询弹窗可见性识别用户关闭
- captchaVerifyParam 复用 turnstile_token 请求字段提交;公开设置下发
  aliyun_captcha_enabled / scene_id / prefix / region
- CSP 放行验证码 CDN:script-src/style-src 加 *.alicdn.com
2026-08-04 20:57:15 +08:00
Wesley LiddickandGitHub 8b3fe664dc Merge pull request #5261 from lyen1688/feat/tencent-captcha-gate
新增腾讯天御验证码认证门禁
2026-08-04 16:39:55 +08:00
lyen1688 e592c5f9e0 新增腾讯天御验证码认证门禁 2026-08-04 15:09:29 +08:00
zhiyu 2eb24814fe fix(codex): 强制统一出站身份并让客户端版本号跟随官方发布
上游 /backend-api/codex 在容量紧张时按客户端身份分优先级降载,被降载的请求
HTTP 200 后立刻推流内 server_is_overloaded。此前网关对配不出官方身份的客户端
整体回退到硬编码的 codex_cli_rs/0.144.1(落后官方 4 个发布),这些请求稳定
落在被优先丢弃的一侧。

- 强制统一出口:所有 OAuth 出站的 User-Agent / originator / version 一律改写
  为网关规范身份,客户端自报身份不参与构造;HTTP / 透传 / WS / alpha-search /
  探针全覆盖。compat 桥接故意删除 originator 的路径保持 no-op。
- 版本号收敛为单一来源,运行时优先级为面板覆写 → 自动同步值 → 内置常量;
  UA 与 version 头同源派生,不再各自硬编码。
- 新增 3 小时自动同步官方客户端最新稳定版,面板可关闭,无需为跟版本而发版。
- 流内 server_is_overloaded / slow_down 改为先在同账号有界重试再切号,并标记为
  请求级瞬时故障,不再据此临时封禁账号。
- 移除被取代的降载身份黑名单、浏览器 UA 兜底及其辅助函数。
2026-08-03 20:14:58 +08:00
shaw 54a2bcfd15 fix(openai): harden reset-credit refresh and account recovery
Review follow-ups on the reset-credit caching flow:

- Recover account state BEFORE (and independently of) the reset-credit
  display cache. A failed cache refresh could previously abort the run and
  leave the account rate-limited — the very reason the credit was spent
  (#3672 / #3740). The recovered account row is now returned even when the
  cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
  the panel reset call a larger timeout. A client abort no longer strands a
  consumed (non-refundable) credit with an unrecovered account, and the
  chained upstream calls can no longer exceed the client timeout and invite a
  retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
  instead of a side-effecting GET flag, so the write is covered by the audit
  middleware. A rejected snapshot write now degrades to cache_persisted=false
  instead of turning a successful upstream read into a 502 that left the card
  without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
  drop expired credits (clamping the count) when rehydrating, so a stale
  cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
  storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
  patched row also enters the auto-refresh silent window.
2026-08-03 14:40:55 +08:00
rick147 a0802f00b6 feat: cache OpenAI reset credit details 2026-08-02 21:32:31 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuangandshaw 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
Brisbanehuangandshaw b0f5007f04 feat(billing-probe): optionally sync account rate from upstream declared rate
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.

- new per-account flag upstream_billing_rate_sync_enabled stored next to
  the probe flag in account extra: enabling sync force-enables the
  probe, disabling the probe cascades sync off, and eligibility follows
  IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
  to an explicit whitelist of the five supported API-key platforms so
  future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
  (finite, within bounds, not rounded to zero at the rate_multiplier
  decimal(10,4) scale) writes back; failed/unsupported/invalid probes
  leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
  UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
  and applies it atomically with the snapshot under the same
  identity/snapshot compare-and-swap, so a probe result observed on a
  stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
  applies the form without overwriting a rate that a probe synchronized
  after the edit form was loaded (nil rateMultiplier = not edited);
  once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
  account has rate sync enabled (whole batch fails with a dedicated
  error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
  enable/disable coupling enforced in the form), synced-rate tooltip on
  the rate cell, bulk edit modal warns and blocks rate edits that hit
  sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
  repo tests for the extended CAS, real-PostgreSQL integration tests
  (rate written only for successful+enabled accounts, manual rate
  protected after sync disabled, admin edit preserved across concurrent
  probe sync), handler/API contract updates, frontend specs for modal
  coupling, bulk rejection and rate cell
2026-08-01 22:11:09 +08:00
Wesley LiddickandGitHub d9fba8fe78 Merge pull request #5101 from Tongzai123/feat/admin-select-all-filtered-results
feat(admin): 支持按筛选结果全选账号
2026-08-01 08:54:38 +08:00
shaw 948b63c9ca feat(moderation): route content moderation through configurable proxy server
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.

Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
  resolution failure surfaces as a moderation error and never silently
  falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
  config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
  0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
  proxy_not_active) without leaking credentials

Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
  non-blockingly; save and test payloads carry proxy_id; zh/en i18n
2026-07-31 23:12:22 +08:00
Litong a35ff9613e feat(admin): 为账号批量删除增加并发限制 2026-07-30 18:03:54 +08:00
wucm667 739c0ff9c5 feat(home): add compact home page preset to avoid abuse classification
- Add compact_home_enabled setting to provide a minimal landing page
- Preserves custom home_content priority over compact mode
- Renders only site identity, navigation, and login/dashboard link
- Avoids marketing copy that triggers anti-fraud systems
- Includes focused backend/frontend tests and i18n support

Fixes #5065
2026-07-30 16:28:19 +08:00
feeeei 720c405e35 feat: add model plaza with group-scoped pricing showcase
- public /model-plaza page (standalone + admin-embedded) listing groups
  with discounted effective prices alongside LiteLLM official reference
- faceted platform/group/rate filters: cross-dimension options gray out
  instead of disappearing, platform-tinted chips via accent color-mix
- paid-price columns highlighted with per-platform tint band
- OptionalJWT middleware so anonymous and signed-in users share one route
- admin settings: enable switch, require-auth switch, markdown description
2026-07-28 16:19:41 +08:00
Wesley LiddickandGitHub 2e432173f7 Merge pull request #4920 from alexj11324/feat/passkey-auth
feat: add passkey authentication
2026-07-28 14:58:37 +08:00
shaw fead4c7ec3 feat(security): add panel API rate limiting to protect DB from high-frequency requests
用户可高频刷面板接口(usage/dashboard 等重聚合查询)直接打爆数据库:
现有限流器只覆盖登录/注册等公开认证入口,登录后的全部面板端点无任何限流。

三层防护(阈值均可在后台可视化配置,panel_rate_limit_settings):

1. 认证面板接口按「用户 ID」限流,与来源 IP 无关——反向代理/NAT 共享出口
   (所有请求源地址坍缩为 127.0.0.1 等)不会互相误伤:
   - Global 档(默认 240 rpm/账号):user/auth/payment/admin 全部登录后路由
   - Heavy 档(默认 60 rpm/账号):/usage、/usage/dashboard/*、
     /user/api-keys/:id/usage/daily 等重 SQL 聚合端点叠加计数
   - 管理员默认豁免(可关闭)

2. 无认证公开接口(/api/v1/settings/*,每次请求都查 DB)按安全客户端 IP
   限流(默认 300 rpm/IP);回环/私网/链路本地地址(反代内部转发地址)
   一律跳过计数,杜绝把整条反代链路合并进同一个桶造成大面积误拦截。

3. 修复既有隐患:auth 入口限流的 IP 取值从 c.ClientIP() 切换到与审计日志/
   会话绑定/API Key ACL 同源的安全客户端 IP 解析(尊重后台「信任反代转发
   IP」开关快照)。原实现下默认反代部署(未配置 server.trusted_proxies)
   所有用户共享同一个登录限流桶,既会全员误拦也可被单人恶意占满形成登录
   DoS;开关关闭时行为与原来完全一致。

工程约束:
- 配置热路径走进程内缓存(atomic.Value + singleflight,60s TTL),
  限流中间件零 DB 访问;保存后当前节点立即生效
- 面板限流 Redis 故障 fail-open(auth 入口保持原有 fail-close)
- 429 响应携带 Retry-After;错误码 RATE_LIMITED
- 支付 webhook / 公开支付回调有意不挂限流
- 新增 GET/PUT /api/v1/admin/settings/panel-rate-limit;设置页安全 tab
  新增「面板接口限流」卡片(zh/en i18n 全量)

测试:rate_limiter/panel_rate_limit/setting_panel_rate_limit 单测全绿;
routes、handler/admin、-tags unit 契约测试通过;前端 vue-tsc/ESLint/
SettingsView spec(26/26,含新增交互用例)/i18n 守卫全部通过。
2026-07-27 15:12:51 +08:00
Wesley LiddickandGitHub a74e11c26a Merge pull request #4868 from visa2/fix/settings-partial-update-clobber
fix(settings): keep fields a settings PUT never sent at their stored value
2026-07-27 11:44:40 +08:00
Zhixuan Jiang 357c5b917b feat: add passkey sign-in settings control 2026-07-26 11:07:12 -04:00
Wey Gu 1850e00955 fix(admin): filter usage logs by request id 2026-07-26 00:53:13 +08:00
visa2andClaude Opus 5 0b5903d458 fix(settings): keep fields a settings PUT never sent at their stored value
PUT /api/v1/admin/settings is a whole-document write. The admin UI always sends
the complete document, so saving from the settings page is unaffected both
before and after this change. The bug is only reachable when an API client calls
the endpoint directly and sends just the fields it wants to change, which is the
natural assumption for a PUT on a settings resource.

Value-typed fields of UpdateSettingsRequest bind to their zero value when the
payload omits them, and buildSystemSettingsUpdates writes every key
unconditionally, so such a caller has no way to say "leave this one alone".
Omitting a field and explicitly clearing it are indistinguishable on the wire.
A caller that sends only the field it wants to change, e.g.

    {"risk_control_enabled": true}

sets that flag and clears every other unguarded field in the same request.
Measured against a fully configured store, one such call empties site_name,
site_subtitle, api_base_url, contact_info and doc_url, and turns
registration_enabled, email_verify_enabled, invitation_code_enabled and
turnstile_enabled off. turnstile_enabled alone gates the captcha on login,
register, forgot-password and both verify-code endpoints, and
email_verify_enabled is a precondition of IsPasswordResetEnabled.

The damage is easy to miss. site_name has a built-in fallback, so
getStringOrDefault renders the cleared value as the default product name and the
login page visibly changes, while the toggles just go quiet. Reopening the
settings page reads the already-cleared state back into the form, so correcting
the one visible field and saving persists the rest of the damage.

Fields that grew their own guard already survive this: the SMTP block falls back
to the previous values when smtp_host arrives empty, secret fields are written
only when non-empty, and 132 request fields are pointers whose handler merges an
omitted field with the stored value. This generalizes that pattern rather than
adding a fourth ad-hoc guard.

The handler now decodes the payload a second time as a raw field map, resolves
the setting key each absent field would have written, and hands that set to the
service, which drops those keys before SetMultiple, so the stored value is never
touched. Fields the payload does carry are written as before, giving the caller
the partial-update semantics it was already assuming. The mapping is reflected off
the request's json tags so new fields are covered without maintaining a list;
smtp_from_email is the only field whose json name differs from its setting key
and is aliased explicitly.

Only value-typed fields are filtered. Pointer fields keep whole-document
behaviour on purpose: forwarded_client_ip_headers and
api_key_acl_trust_forwarded_ip depend on being rewritten on every save to
re-normalize fail-closed state, which the malformed forwarded-client-IP header
test pins down.

An explicitly sent empty value is still a deliberate clear; only absent fields
are preserved. A partial write refreshes the in-process caches from storage
instead of from the request struct, which holds zero values for whatever the
caller omitted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 21:03:00 +08:00
alfadb b403f88f51 fix(ollama): 避免刷新候选饥饿
ListDue 在 LIMIT 前用与 service 纯函数一致的 debounce/max-wait/backoff
规则筛真正 due 组,防止有活动但未到期的组占满每轮 20 名额。
2026-07-25 15:37:06 +08:00