Files
hapi/docs/api/client-contract/errors.md
T
Junmo KimandGitHub fd2822bab5 fix(agy): allow switching existing sessions to newly available models (#1814)
* fix(agy): say what the model picker is actually waiting on

The spinner in the New Session AGY picker read "Checking Antigravity
authentication…", but nothing at that point checks authentication — the
machine is running `agy models`, and the sign-in prompt is a separate
branch below it, shown only when agy reports the failure.

Name the wait after the work: "Fetching available models…", the same
words agy prints while it fetches.

* refactor(agy): describe a probe by its outcome, not by its response

The probe function returned a finished `AgyModelsResponse`, so "agy could
not be reached" and "agy listed no models" both arrived as a successful
response carrying the hardcoded mirror, and the caller could no longer
tell which had happened. Every policy decision about that answer has to
live inside the probe as a result.

Hand back what the probe observed — a live catalog, an auth failure, or
nothing usable — and let the caller turn it into a response. Same
behaviour: the mirror still stands in for both failure modes, and the
60s cache still holds whatever came out.

* fix(agy): serve the model catalog stale-while-revalidate

The `agy models` probe is a whole agy invocation — around 3s on a good
day, 15s when it times out — and the 60s window meant the New Session
picker paid that again a minute after the last look.

Keep the last listing agy actually returned and answer from it: fresh for
ten minutes, then still answered while a probe refreshes behind it, until
the entry is a day old and stops standing in for the machine at all. A
probe that times out or loses auth leaves that entry alone, so a blip no
longer empties a working picker, and the hardcoded mirror is no longer
recorded as if the machine had reported it.

Three things fall out of that and are handled here. A machine whose
sign-in has actually gone bad would otherwise look healthy for a day, so
an auth failure rides along with the catalog it can still serve — and,
because nothing else would re-probe a catalog that is still fresh, a
warning riding on the answer is itself a reason to look again. A failed
probe is not repeated on the very next request either, or a machine where
agy hangs would spawn it once per poll.

An explicit refresh always costs a probe, and never rides one that was
already running when it was asked for.

* fix(agy): let Retry force a fresh model catalog probe

With the catalog held for ten minutes, Retry would otherwise hand back the
answer it was pressed to replace, so the intent travels to the machine:
`?refresh=true` on the machine route, an optional RPC param, and a
one-shot flag on the query so ordinary mount and focus refetches stay
cheap. Every hop is optional, so a hub and a runner on different versions
still talk — the older side ignores it and answers from its cache.

Retry also has to be reachable, and honest, in the state that needs it.
The machine now answers with both a usable catalog and the sign-in failure
behind it, so the picker keeps the list and says why it may be out of
date, with the button right there rather than only once there is nothing
left to show. Pressing it runs agy, which can take tens of seconds, so the
button says so while it does.

The client contract covers both: `agy-models` is the one catalog route
that can carry an `error` on a successful response, and the one that
takes a refresh parameter.

* fix(agy): use the machine catalog in the in-session model picker

New Session already asks the machine what `agy models` lists, but a
session that is already open offered the built-in list in
`shared/src/models.ts`. That list is a hand-maintained mirror, so a model
agy started offering after the last release could be picked for a new
session and not for the one already running.

Point the composer at the same machine catalog. The mirror stays as the
fallback for the moment before the machine answers, and a model the
session is already on is kept selectable — and readable, when it is one
of the known presets — even after the catalog moves on without it.

* fix(agy): announce a model catalog re-check that changed the answer

Serving the last known catalog answers the picker instantly, but a picker
that was already open kept showing that answer until the user closed and
reopened it — the machine had no way to say it had found something newer.

Say it on the stream that already carries machine changes. The machine
daemon — the only process that answers `<machineId>:listAgyModels` —
emits it, and the hub forwards it as `machine-agy-models-updated` with
nothing but the machineId. Namespace resolution, per-machine delivery and
reconnect replay all come from the existing path.

What counts as a change is what the route would answer, not what sits in
the cache. That distinction carries the cases: a sign-in that lapsed or
came back changes no models yet changes what the user is told; a machine
whose agy was signed out has been answering from the hardcoded mirror, and
its first real listing is the largest change there is, for every client
except the one awaiting it.

* fix(agy): re-read the announced machine's model catalog

On `machine-agy-models-updated`, cancel and refetch that one machine's
catalog query — the app's global connection is always subscribed, so an
open picker redraws wherever it is.

Cancelling first is what makes it correct rather than merely likely.
query-core cancels an in-flight fetch only when the query already holds
data, so a picker opening for the first time would otherwise join the
request already on its way and settle on the listing the announcement
replaced. The refetch is answered from the machine's cache, so it starts
no probe and cannot bounce another announcement back.

A reconnect the hub could not replay takes the resync path, which clears
the agy catalogs the same way — that path has no announcement to fall back
on, so it is the one that can least afford to join a stale request.

* fix(agy): keep a model the user picked when the catalog moves under them

The catalog can now change while the New Session form is open, and the
form dropped any selection the machine no longer advertised — including
one the user had just made.

Keep that one, and list it as no longer listed so the form does not imply
agy is still offering it. A model restored from a draft or a saved
preference is still dropped: it may never have been runnable here.

* fix(agy): announce uncached authentication changes
2026-09-11 14:19:37 +08:00

8.2 KiB
Raw Blame History

Errors

Error semantics for all /api/* endpoints. Grounded in hub/src/web/routes/guards.ts and the individual route files under hub/src/web/routes/; reference consumer web/src/api/client.ts (ApiError, parseErrorCode).

Body shape

Error responses are JSON:

{ "error": "Session is inactive", "code": "session_inactive" }
  • error — human-readable message. Never match on it: consumers i18n it and the hub may reword it (this rule is stated in guards.ts itself).
  • code — optional stable machine-readable discriminator. Clients branch on (status, code).
  • issues — present on some 400s: Zod validation details (parsed.error.issues or .flatten() output). Useful for logging, not for branching.
  • A few endpoints add context fields (e.g. 413 export adds count/limit; 422 reopen adds missing[]).

When no code is present, branch on status alone and treat the failure generically. The web reference falls back to using error as a pseudo-code when code is absent (parseErrorCode) — acceptable for logging, not for logic.

Status × code table

Status code Where (source) Meaning / client action
400 — All routes with Zod bodies ({error: 'Invalid body'}, some with issues) Client bug — fix the request; do not retry
400 — Flavor gates (sessions.ts: wrong-flavor model/mode endpoints) Hide the control for this flavor
400 scratchlist_attachment_invalid, scratchlist_entry_empty, attachment-limit codes from validateScratchlistAttachmentsForWrite sessions.ts scratchlist routes Surface validation message
401 — Middleware (middleware/auth.ts) and POST /api/auth (routes/auth.ts) See Auth → 401 bodies; middleware 401 → silent re-auth once, /api/auth 401 → re-pair
403 — guards.ts (Session access denied, Machine access denied); owner-only routes (usage.ts, storage.ts, hubSettings.ts, voice.ts) Namespace mismatch / not hub owner — hide the surface, don't retry
403 access_denied RPC-flow results mapped in sessions.ts (resume/reopen/cursor-chat-store) Same as above
404 — guards.ts (Session not found, Machine not found); permissions.ts (Request not found); scratchlist entry/attachment; events.ts visibility (Subscription not found) Stale reference — refresh the parent list
404 session_not_found, machine_not_found Coded variants from resume/reopen/restart-runner result mapping Same
409 session_inactive guards.ts requireSession(requireActive) — send/steer/abort/approve/deny/config on an inactive session Offer Reopen (the web router does exactly this on this code)
409 scratchlist_at_cap sessions.ts scratchlist create (200-entry cap) Show cap notice; do not retry
409 scratchlist_attachment_in_use sessions.ts attachment delete while still referenced Detach from entry first
409 resume_unavailable sessions.ts resume/reopen result mapping Session can't be resumed (e.g. unsupported state)
409 metadata_conflict sessions.ts reopen result mapping Refetch session, retry once at most
409 runner_upgrade_required machines.ts Agent availability Upgrade and restart the runner; disable session creation
409 — (version conflict) sessions.ts PATCH rename/summary, machines.ts PATCH rename — message mentions version/concurrently; no code Concurrent edit — refetch and reapply
409 — sessions.ts delete-while-active, archive of plain inactive row, fork/rewind refusals, remote-only config on terminal-controlled sessions (controlledByUser) Surface message; refresh session state
413 — sessions.ts upload (> 50 MB decoded), export too large ({error, count, limit}); voice.ts transcription (Audio file too large, 25 MB audio / ~26 MB body) Reduce payload
413 scratchlist_attachment_too_large sessions.ts scratchlist upload Reduce attachment
422 — sessions.ts reopen with incomplete metadata ({error, missing[]}); title-suggestion pass-through (TitleSuggestionError, statuses 422/429/502/503) Not reopenable / feature unavailable
429 — Title suggestion (provider rate limit) Back off
500 — Catch-all in most routes ({error: message}) Log; generic failure UI
502 — Title suggestion upstream failure; restart-runner unknown error Retry later
503 — guards.ts requireSyncEngine → {error: 'Not connected'} — hub subsystems not up (startup/shutdown window); also Telegram-disabled on /api/auth initData path Retry with backoff
503 no_machine_online resume/reopen/spawn-flow result mapping (sessions.ts) The machine that owns the session is offline — tell the user to start the runner
503 machine_offline machines.ts restart-runner Same
503 rpc_target_missing machines.ts pi/codex model catalogs (RPC_TARGET_MISSING_ERROR_CODE in shared/src/rpcMethods.ts — RPC handler unregistered or socket disconnected) CLI-side target gone — treat as offline

RPC-wrapped endpoints

Many endpoints do not answer from hub state — the hub relays the request over Socket.IO to the session's CLI process (or the machine's runner) and forwards the result: git/file/directory/search, generated images, uploads, model catalogs, Agent availability, slash-commands, skills, spawn, list-directory, paths/exists. (The mode/model/effort config endpoints are RPC-backed too, but map apply-failures to 409 with a message.) Their failure modes differ from plain endpoints:

  1. CLI reachable, command failed → HTTP 200 with {success: false, error} (e.g. runRpc in hub/src/web/routes/git.ts catches RPC errors, including the 30 s RPC timeout, and returns them as a JSON envelope). Clients must check the success field on every RPC-shaped response; HTTP 200 alone means nothing.
  2. CLI offline / handler missing → depends on the route: the model-catalog routes in machines.ts map RpcTargetMissingError to 503 rpc_target_missing; git.ts-style routes fold it into the 200 {success: false} envelope; resume/reopen surface 503 no_machine_online.
  3. Hub subsystems not up → 503 Not connected from requireSyncEngine (brief startup/shutdown window).

Practical rule: treat success: false, 503 rpc_target_missing, and 503 no_machine_online as the same user-facing condition — "the computer running this session is not reachable" — with the raw error string available in a details view.

One route qualifies that rule. GET /api/machines/:id/agy-models keeps serving the last catalog the machine got out of agy models while the CLI re-checks in the background, so it can answer success: true and carry an error: the list is usable, and error says why it may be stale (typically the machine's agy sign-in has lapsed). Render it beside the catalog rather than instead of it, and offer ?refresh=true as the way to ask again — a plain repeat is answered from the same cache.

When that background re-check lands a different listing, the machine says so over the existing event stream rather than making clients ask: machine-agy-models-updated (see SSE) carries the machineId and nothing else. Refetch that machine's route on it — the answer comes from the machine's cache, so it costs no agy run and produces no further event. It is emitted whenever the re-check changes what this route would answer — a different listing, or a sign-in warning that appeared or cleared — and not when it changes neither, so a failed re-check that raises a warning does announce. The machine's very first listing is never announced: whoever triggered it is already awaiting it. Requests during that window are answered from the machine's cache and do not launch agy.

Retry guidance

Class Retry?
400 / 403 / 404 / 409 / 413 / 422 No (fix input, refresh state, or hide surface)
401 (middleware) Once, after silent re-auth (Auth)
429 / 502 / 503 Yes, with backoff
200 {success: false} Manual retry only (user-initiated) — the CLI answered and said no