The Anthropic→Responses streaming converter omits two things the OpenAI
Responses wire format requires, which breaks clients that reconstruct the
response from the event stream (rather than just reading deltas).
1. response.content_part.added is never emitted.
A message item is opened with content: [], and the OpenAI SDK's
accumulating stream helper (client.responses.stream) only appends a
content part when it sees content_part.added. Without it, the next
output_text.delta indexes output.content[content_index] and raises:
File "openai/lib/streaming/responses/_responses.py", line 352,
in accumulate_event
content = output.content[event.content_index]
IndexError: list index out of range
Raw iteration (responses.create(stream=True)) does not accumulate and is
unaffected, which is why this went unnoticed.
2. response.completed carries Output: []ResponsesOutput{}.
get_final_response() and tracing integrations parse the terminal event's
response directly, so callers see an empty output_text even though the
deltas streamed correctly. This one is invisible when only watching the
stream render.
Also carries the full text on output_text.done / content_part.done (deltas
carry increments only, done events carry the whole part) and fills in
content/arguments/summary on output_item.done, which had the same
empty-payload issue.
Reproduced against a live Anthropic-platform group with the openai Python
SDK 2.46.0; Arize Phoenix's playground hits the same path. Verified before
(IndexError) and after (full text via get_final_response()).
Adds regression tests covering event ordering, done-event payloads, and the
terminal event's output for both text and tool calls.
Adversarial re-verification found one more divergence from the
double-conversion chain: an upstream 200 with empty choices (or a nil
response) left stop_reason as an empty string, while the old chain
reports end_turn. Derive the fallback from the content blocks — the
guard never fires when choices exist, since every finish_reason maps to
a non-empty stop_reason.
Also strengthen the equivalence tests: compare tool_use Content[].Input
and tool_call Function.Arguments across bridges, and add an
empty-choices stop_reason parity case.
Adversarial review of the direct bridge found divergences from the
double-conversion chain it replaces; all are now aligned and covered by
tests that fail against the previous implementation:
- flush argument fragments buffered before a deferred tool announcement,
and announce name-less tools at finalize, so no tool arguments are lost
when upstreams stream arguments before the name
- fold text-only user array content into a single string; parts form only
when an image requires it (strict chat upstreams reject array content)
- drop tool_choice pointing at undeclared tools and unknown choice types
- treat cache_write_tokens/cache_creation_tokens as alternate spellings
(prefer write), not additive
- generate a response id when the upstream omits one
- derive stop_reason from blocks for content_filter/unknown finish reasons
- emit input_json_delta "{}" when a tool block closes without argument
deltas
Three tests covering the CC→Responses→Anthropic finalize path:
- TestStreamingParallelToolUseNoGhostDelta: full end-to-end with two
parallel tool_calls, asserts every content_block_delta targets a
started block
- TestStreamingParallelToolUseSecondToolPackedArgsDone: focused test
where tool 1 streams deltas and tool 2 has packed .done args
- TestStreamingThreeParallelToolsAllPackedDone: three tools all with
packed .done, the most extreme case
All three fail on unfixed code (ghost deltas on wrong indices) and
pass after the fix.
Skip the Responses API intermediate representation on the /v1/messages
force-chat path. Previously every streaming token ran through two state
machines (CC→Responses→Anthropic); now it runs through one (CC→Anthropic).
New file backend/internal/pkg/apicompat/chatcompletions_anthropic_bridge.go:
- AnthropicToChatCompletionsRequest: request-side direct conversion
- ChatCompletionsResponseToAnthropic: non-stream response direct conversion
- ChatCompletionsChunkToAnthropicEvents + Finalize: single streaming state machine
openai_gateway_messages_chat_fallback.go rewired to use the direct bridge.
Existing 7 ForceChatCompletions end-to-end tests pass unchanged; 22 new
unit tests added including equivalence tests vs the double-conversion path.
resToAnthHandleFuncArgsDone used state.ContentBlockIndex directly
instead of looking up from OutputIndexToBlockIdx like
resToAnthHandleFuncArgsDelta does. When multiple tool_calls arrive
in parallel and arguments come as a packed .done (no prior delta),
the second+ tool would emit content_block_delta on an index that
was never content_block_start'ed, causing Claude Code to report
"Content block not found".
Fix: resolve block index from OutputIndexToBlockIdx, and skip the
delta if the block is already closed or the index doesn't match
the current open block.
Closes#4193
WSv2 egress relays upstream events verbatim without the HTTP-path
namespace restore, so flattening requests that take the WSv2 branch
would surface flattened tool names the client cannot match. Resolve
the WS transport decision before flattening and skip flattening only
when the request will actually go WSv2 (passthrough accounts return
via HTTP before the WSv2 branch and still flatten).
Also check all type assertions in responses_namespace_test.go to
satisfy golangci-lint errcheck.
Resolve conflict in openai_gateway_passthrough.go streaming path: keep
main's normalizeCompletedImageGenerationStatus normalization ahead of
this branch's namespace restore block, mirroring the established order
in openai_gateway_response_handling.go.
Preserve Chat Completions parallel_tool_calls when converting requests to the Responses API, and map Responses parallel_tool_calls back when falling back to Chat Completions upstreams.
Cover both true and explicit false values so clients can disable parallel tool calls without the field being dropped by omitempty.
When converting a Chat Completions stream into Responses events, the first
tool_call delta chunk was copied wholesale into stream state (including
function.arguments), then the same chunk's arguments were accumulated again by
the shared `+=` block. For OpenAI this is harmless because its first tool_call
chunk carries empty arguments, but upstreams that pack id+name+arguments into a
single chunk (e.g. GLM/Zhipu) end up with doubled arguments such as
{"cmd":"ls"}{"cmd":"ls"}. Codex then fails to parse the tool call with
"trailing characters", breaking every tool invocation.
Reset the copied arguments so the shared accumulator counts them exactly once,
keeping the emitted delta and the final done/arguments consistent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR #3016's merge appended a verbatim second copy of four
TestStream_Reasoning* functions into
chatcompletions_responses_stream_lifecycle_test.go, causing
'redeclared in this block' build failures that broke both the
test and golangci-lint CI jobs.
Remove the duplicate block; each test now appears once.