The Anthropic→Responses streaming converter omits two things the OpenAI
Responses wire format requires, which breaks clients that reconstruct the
response from the event stream (rather than just reading deltas).
1. response.content_part.added is never emitted.
A message item is opened with content: [], and the OpenAI SDK's
accumulating stream helper (client.responses.stream) only appends a
content part when it sees content_part.added. Without it, the next
output_text.delta indexes output.content[content_index] and raises:
File "openai/lib/streaming/responses/_responses.py", line 352,
in accumulate_event
content = output.content[event.content_index]
IndexError: list index out of range
Raw iteration (responses.create(stream=True)) does not accumulate and is
unaffected, which is why this went unnoticed.
2. response.completed carries Output: []ResponsesOutput{}.
get_final_response() and tracing integrations parse the terminal event's
response directly, so callers see an empty output_text even though the
deltas streamed correctly. This one is invisible when only watching the
stream render.
Also carries the full text on output_text.done / content_part.done (deltas
carry increments only, done events carry the whole part) and fills in
content/arguments/summary on output_item.done, which had the same
empty-payload issue.
Reproduced against a live Anthropic-platform group with the openai Python
SDK 2.46.0; Arize Phoenix's playground hits the same path. Verified before
(IndexError) and after (full text via get_final_response()).
Adds regression tests covering event ordering, done-event payloads, and the
terminal event's output for both text and tool calls.
Behind a reverse proxy (e.g. nginx with X-Real-IP), admin audit logs and
session IP/UA binding always recorded 127.0.0.1 because they hardcoded
the gin trusted_proxies chain, while API key IP restriction already
honored the "trust forwarded client IP" system setting.
- add ip.GetSecurityClientIP(c, trustForwarded) as the single source of
truth for security-sensitive client IP selection; API key auth
middlewares (main + google) refactored onto it with zero behavior change
- SessionBindingContext(cfg) now resolves the client IP via the same
toggle and injects it into the request context; token issuance,
binding enforcement and its mismatch audit record all read the
injected value, so issue/verify can never diverge
- audit log middleware and audit-log clear trace record the same
security client IP (middleware.SecurityClientIP), falling back to the
trusted proxy chain when the injection is absent
- settings UI hint (zh/en) documents the broadened toggle scope and the
one-time re-login after toggling while session binding is enabled
With the toggle off (default) behavior is byte-for-byte unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
- admin.accounts.oauth.openai.mobileRefreshTokenAuth was referenced by
OAuthAuthorizationFlow.vue since 9f8cffe88 (Mobile RT entry) but never
added to zh/en locales, rendering the raw key in the add-account wizard
- admin.accounts.oauth.openai.accessTokenAuth has the same latent issue
since 26060e702 (Sora AT import); currently hidden but fixed alongside
Co-Authored-By: Claude <noreply@anthropic.com>
Reuse applyGrokFreeMessagesFunctionToolCacheRoute on native /v1/responses
and the Grok WS HTTP bridge so Free OAuth requests with client function
tools get the same mixed-tools cache route as the Messages bridge
(append/convert web_search and x_search).
Also dedupe: Grok Build already declares function tools named web_search,
so naive append caused "Duplicate tool names: web_search". Convert those
function entries to native tool types and skip duplicates.
Only Free OAuth accounts (isKnownGrokFreeAccount); paid/unknown unchanged.
PR #4425 was authored before #4429 widened NewUserHandler with the
step-up TOTP and user services, and merged without a rebase, breaking
typecheck on main.
Reconcile the OAuth media route with the manual endpoint-switch redesign
(7f5d067af): media leaves for api.x.ai only when text traffic resolves to
the CLI gateway host; manually selected official/regional/custom endpoints
keep serving media as-is.
Admins often recreate groups with the same pricing, routing, and account membership. A server-side duplicate creates an inactive copy for review, preserves eligible account priorities, and recovers ambiguous retries without creating extra groups.
Constraint: Group has no neutral JSON metadata field for durable operation recovery
Constraint: Model routing references account IDs, so copied configuration requires matching bindings
Rejected: Rebuild from the list response | it omits configuration and account priority details
Rejected: Store operation identity in business configuration | it would pollute real group settings
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated groups inactive until an administrator reviews the copied configuration
Tested: Go unit and full tests, go vet, integration-tag compile, frontend Vitest, lint, typecheck, production build, and Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite
Admins often recreate monitors with the same endpoint, model, and request settings. A server-side duplicate keeps the stored API key out of the browser, creates a disabled copy for review, and uses stable operation identity to recover ambiguous retries without creating extra rows.
Constraint: Stored monitor API keys must never be returned to the browser
Constraint: Applying a request template must preserve internal duplicate recovery metadata
Rejected: Rebuild the monitor from list data | list responses only contain a masked API key
Rejected: Copy runtime state and history | a duplicate should start as an unverified configuration
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated monitors disabled until an administrator reviews and enables them
Tested: Go unit tests for repository, service, and admin handler; integration-tag compile; go vet; golangci-lint v2.9; frontend Vitest, ESLint, typecheck, production build; Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite
For OpenAI-compatible API-key accounts, /v1/responses requests with
image-generation intent could be scheduled to accounts whose upstream
does not support the Responses API (extra.openai_responses_supported=false).
The flag was only consulted at forward time, where such accounts are
silently downgraded to a Chat-Completions path that cannot produce images,
causing upstream 4xx/5xx or canceled requests.
Fix:
- Add endpoint capability OpenAIEndpointCapabilityResponses. Its check in
SupportsOpenAIEndpointCapability excludes only OpenAI API-key accounts
probed as unsupported (mirroring the forward-time downgrade condition);
OAuth/Grok/unprobed accounts keep existing behavior, and a responses-
capable upstream must still pass the chat_completions gate. Reusing the
existing requiredCapability plumbing makes every scheduler filter path
enforce it with no scheduler signature changes.
- Request the responses capability at the HTTP Responses and
ResponsesWebSocket call sites only when imageIntent && platform==openai,
so non-image requests keep the downgrade path and Grok's own image path
is untouched.
- Normalize max_tokens -> max_output_tokens on the native responses
forward path (PlatformOpenAI), and strip prompt_cache_options alongside
prompt_cache_retention/safety_identifier.
/v1/images/generations continues to use native image capability (unchanged).
Tests: capability truth table, scheduler exclusion of unsupported accounts,
and forward-path transform behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KXpzKKvsb5jW2GvgBQqnnZ