Commit Graph
2 Commits
Author SHA1 Message Date
alfadb 5c5283979b doc(thinking-protocol): clarify mappedModel vs originalModel semantics per call path
回应 PR #3247 Copilot review:

1. NormalizeChineseLLMThinking 文档写 'MiniMax M3 / M3.x' 但实现覆盖
   minimax-m* (M2.x 也命中)。改成 'MiniMax M-series (M2.x/M3.x)'。

2. Copilot 质疑 gemini_messages_compat 传 originalModel 不是
   mappedModel 是误用。实际是有意为之的跨协议场景:
   - 上游是 Gemini,但被剥离的 body 是 Anthropic 格式
   - 剥离逻辑要按客户端请求的 Anthropic 子协议族判定
   - 传 mappedModel (gemini-3.1-pro) 会被判为 Unknown→不剥离→retry 死循环

   未改代码逻辑,改为更新文档:
   - thinking_protocol.go ResolveThinkingProtocol 添加「调用路径语义」
     说明 Anthropic gateway 与 Gemini compat 两条路径的参数语义差异。
   - gemini_messages_compat_service.go 调用点加 inline 注释解释
     路径语义为什么该传 originalModel。
2026-06-16 19:37:31 +08:00
alfadbandClaude Opus 4.7 6baf00d784 fix(gateway): protocol-aware thinking-block filtering for Anthropic-compatible upstreams
The gateway's thinking-block handling was designed for Anthropic's strict
semantics (drop blocks with missing/invalid signature), but third-party
Anthropic-compatible upstreams have INVERTED semantics:

  * DeepSeek `/anthropic`, Kimi `/coding`, GLM, Moonshot, qwen-*-thinking
    require ALL historical thinking blocks to round-trip verbatim.
  * Stripping any of them produces:
    400 "The content[].thinking in the thinking mode must be passed back
    to the API"

Without this fix, every multi-turn request from a thinking-capable client
(Claude Code, pi, etc.) to such upstreams loses its thinking blocks and
fails. This becomes especially painful when an account's model_mapping
maps `claude-sonnet-4-6 → deepseek-v4-pro` — `reqModel` looks Anthropic
but the upstream contract is the opposite.

Approach
--------

Branch all thinking-block transforms by the *mapped* model id (after
account model_mapping is applied), classifying into three families:

  * `anthropic-strict`     claude-/opus-/sonnet-/haiku-      → existing behaviour
  * `passback-required`    deepseek-/kimi-/moonshot-/glm-/   → preserve verbatim
                           qwen-*-thinking
  * `unknown`              other models                       → conservative
                                                               (preserve, no retry)

Affected entry points (all guarded):

  * Pre-filter on outbound:   `FilterThinkingBlocks`
    Previously dropped blocks with missing/invalid signature; now skips
    entirely for non-strict families. Pre-filter is needed because the
    post-error retry path can run out of budget on long conversations
    (maxRetryElapsed = 10 s).
  * 400 retry rectifier:      `FilterThinkingBlocksForRetry`
    Disables top-level thinking and converts thinking → text. Now skips
    for passback-required (those 400s aren't signature errors and any
    transformation breaks the round-trip contract).
  * 400 retry rectifier (tools): `FilterSignatureSensitiveBlocksForRetry`
    Same family-aware short-circuit.
  * 400 detector:             `shouldRectifySignatureError`
    Returns false for passback-required, so the retry path doesn't even
    fire.

Tests
-----

  * `thinking_protocol_test.go` — classifier across all known vendor
    prefixes plus edge cases (empty, case, qwen non-thinking).
  * `thinking_protocol_filter_integration_test.go` — locks in that the
    three filter entry points return the body byte-for-byte unchanged
    when the model id is passback-required or unknown, and still strip
    invalid blocks for anthropic-strict.

This PR supersedes #1350 (which only added the pre-filter without the
upstream-family awareness, and would have made third-party upstreams
worse). Once merged, please close #1350.

Reference issues:
  - NousResearch/hermes-agent#16748 — DeepSeek /anthropic strip behaviour
  - NousResearch/hermes-agent#15700 — DeepSeek thinking:disabled requirement

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-16 19:37:31 +08:00