Persist reasoning ceilings and exact mappings for OpenAI groups, enforce them across HTTP and WebSocket forwarding, and invalidate cached auth snapshots.
Resolve const-block conflict in openai_gateway_grok_cache.go:
keep #4590's client tool cache constants, drop grokFreeRolling24hTokenLimit
(moved to pkg/xai as IsGrokFreeRolling24hTokenLimit with legacy 2M support).
The in-place update ran entirely inside c.Request.Context(). Browsers and
reverse proxies commonly abort long-idle requests (axios global timeout
30s, nginx proxy_read_timeout 60s by default), which canceled the request
context mid-download and killed every slow update with
'download failed: context canceled' while the version stayed unchanged.
Users behind slow GitHub links saw the update button fail at a wall-clock
ceiling (~60s) on every attempt (#4504).
- Run PerformUpdate and RollbackToVersion on a context detached from the
request (context.WithoutCancel) and bounded by a 15-minute deadline so
the 10-minute GitHub download client owns its own timeout. A client
disconnect no longer aborts the binary swap; retries then hit the
system operation lock or report 'Already up to date'.
- Raise the frontend timeout for the update/rollback calls from the
global 30s axios default to 15 minutes so the browser can actually
wait for the result.
Fixes#4504
The Codex models manifest path returned upstream 401s straight to the
client without feeding them into the account state machinery. A revoked
or invalidated OAuth account therefore stayed active and schedulable,
kept being selected for subsequent /models requests, and produced
repeated 502s until an admin ran a manual connection test (#4544).
- Route ChatGPT-backend manifest 401s through the shared upstream-error
handling: token cache invalidation, temp-unschedulable cooldown for
refreshable OAuth accounts, permanent disable for
token_revoked/token_invalidated, plus the runtime scheduling block.
- Treat ChatGPT-backend manifest 401s as failover-eligible so the
current /models request can switch to a healthy account instead of
returning 502. Custom API key upstream 401s keep the existing
no-failover, no-disable behavior since their /models auth is not
authoritative for the account.
- Skip Agent Identity accounts: their 401s can be task-scoped and have
a dedicated recovery flow.
- Attach the selected account to the ops error-log context so /models
failures record the account_id.
Fixes#4544
Co-authored-by: Cursor <cursoragent@cursor.com>
Reconcile with #4539 (grok media account model mapping), now on main:
- handler/grok_media.go non-failover error path keeps both changes —
#4539's grokMediaScheduleModel(account, routingModel, nil) schedule
attribution and this branch's IsResponseCommitted guard
- auto-merged sections verified: routing/classify use #4539's routingModel,
video lookup owner-binding and no-failover semantics intact, ForwardGrokMedia
keeps mapping block (skipped for lookup endpoints via RequiresRequestBody),
empty-image failover, and video-status URL rewrite in order
Resolve conflicts with main:
- service/grok_media.go: keep both post-response blocks — #4497's empty
image-output failover (main) runs first for image endpoints, then this
branch's video-status content-URL rewrite; the endpoint conditions are
mutually exclusive
- handler/openai_gateway_credential_failover_loop_test.go: mark the stub
OAuth accounts media-eligible via the grok_media_eligible extra override,
because this branch moved grok media failover coverage to the generation
endpoint, which is now gated by #4540's paid-eligibility probe on main
Resolve conflicts with main:
- handler/grok_media_test.go: keep both new tests (schedule model test from
this branch, eligibility gating tests from #4540)
- service/openai_gateway_grok_test.go: keep both new tests; update the image
cases of the mapping table test to return a non-empty image payload because
#4497 (already on main) now converts empty image responses into an upstream
failover error