Scan client-injected assistant/tool/model turns, fail closed when config cannot
be trusted after startup or stale invalidation, reuse probe tokens only for the
same base URL, and restrict localhost dials to loopback addresses.
Co-authored-by: Cursor <cursoragent@cursor.com>
Reuse applyGrokFreeMessagesFunctionToolCacheRoute on native /v1/responses
and the Grok WS HTTP bridge so Free OAuth requests with client function
tools get the same mixed-tools cache route as the Messages bridge
(append/convert web_search and x_search).
Also dedupe: Grok Build already declares function tools named web_search,
so naive append caused "Duplicate tool names: web_search". Convert those
function entries to native tool types and skip duplicates.
Only Free OAuth accounts (isKnownGrokFreeAccount); paid/unknown unchanged.
PR #4425 was authored before #4429 widened NewUserHandler with the
step-up TOTP and user services, and merged without a rebase, breaking
typecheck on main.
Reconcile the OAuth media route with the manual endpoint-switch redesign
(7f5d067af): media leaves for api.x.ai only when text traffic resolves to
the CLI gateway host; manually selected official/regional/custom endpoints
keep serving media as-is.
Admins often recreate groups with the same pricing, routing, and account membership. A server-side duplicate creates an inactive copy for review, preserves eligible account priorities, and recovers ambiguous retries without creating extra groups.
Constraint: Group has no neutral JSON metadata field for durable operation recovery
Constraint: Model routing references account IDs, so copied configuration requires matching bindings
Rejected: Rebuild from the list response | it omits configuration and account priority details
Rejected: Store operation identity in business configuration | it would pollute real group settings
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated groups inactive until an administrator reviews the copied configuration
Tested: Go unit and full tests, go vet, integration-tag compile, frontend Vitest, lint, typecheck, production build, and Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite
Admins often recreate monitors with the same endpoint, model, and request settings. A server-side duplicate keeps the stored API key out of the browser, creates a disabled copy for review, and uses stable operation identity to recover ambiguous retries without creating extra rows.
Constraint: Stored monitor API keys must never be returned to the browser
Constraint: Applying a request template must preserve internal duplicate recovery metadata
Rejected: Rebuild the monitor from list data | list responses only contain a masked API key
Rejected: Copy runtime state and history | a duplicate should start as an unverified configuration
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated monitors disabled until an administrator reviews and enables them
Tested: Go unit tests for repository, service, and admin handler; integration-tag compile; go vet; golangci-lint v2.9; frontend Vitest, ESLint, typecheck, production build; Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite
For OpenAI-compatible API-key accounts, /v1/responses requests with
image-generation intent could be scheduled to accounts whose upstream
does not support the Responses API (extra.openai_responses_supported=false).
The flag was only consulted at forward time, where such accounts are
silently downgraded to a Chat-Completions path that cannot produce images,
causing upstream 4xx/5xx or canceled requests.
Fix:
- Add endpoint capability OpenAIEndpointCapabilityResponses. Its check in
SupportsOpenAIEndpointCapability excludes only OpenAI API-key accounts
probed as unsupported (mirroring the forward-time downgrade condition);
OAuth/Grok/unprobed accounts keep existing behavior, and a responses-
capable upstream must still pass the chat_completions gate. Reusing the
existing requiredCapability plumbing makes every scheduler filter path
enforce it with no scheduler signature changes.
- Request the responses capability at the HTTP Responses and
ResponsesWebSocket call sites only when imageIntent && platform==openai,
so non-image requests keep the downgrade path and Grok's own image path
is untouched.
- Normalize max_tokens -> max_output_tokens on the native responses
forward path (PlatformOpenAI), and strip prompt_cache_options alongside
prompt_cache_retention/safety_identifier.
/v1/images/generations continues to use native image capability (unchanged).
Tests: capability truth table, scheduler exclusion of unsupported accounts,
and forward-path transform behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KXpzKKvsb5jW2GvgBQqnnZ