Use one native app-server for terminal, Web and phone clients while retaining the existing CLI and Runner lifecycle.
Synchronize native queues, permissions, question history and steering state; preserve explicit permission precedence and per-turn usage models. Resume inactive clear commands through Runner and reject independent child cold resumes.
Add shared-runtime regression tests, generated protocol fixtures and lifecycle documentation.
Bridge main-session PermissionRequest hooks without suppressing the native
terminal dialog. Reconcile replies against native results and clean up on
timeout, cancellation, mode switches, and session changes.
Keep reply IDs distinct from native tool IDs across web and native clients;
add protocol fixtures and regression tests.
Refs #1796
* fix(agy): say what the model picker is actually waiting on
The spinner in the New Session AGY picker read "Checking Antigravity
authentication…", but nothing at that point checks authentication — the
machine is running `agy models`, and the sign-in prompt is a separate
branch below it, shown only when agy reports the failure.
Name the wait after the work: "Fetching available models…", the same
words agy prints while it fetches.
* refactor(agy): describe a probe by its outcome, not by its response
The probe function returned a finished `AgyModelsResponse`, so "agy could
not be reached" and "agy listed no models" both arrived as a successful
response carrying the hardcoded mirror, and the caller could no longer
tell which had happened. Every policy decision about that answer has to
live inside the probe as a result.
Hand back what the probe observed — a live catalog, an auth failure, or
nothing usable — and let the caller turn it into a response. Same
behaviour: the mirror still stands in for both failure modes, and the
60s cache still holds whatever came out.
* fix(agy): serve the model catalog stale-while-revalidate
The `agy models` probe is a whole agy invocation — around 3s on a good
day, 15s when it times out — and the 60s window meant the New Session
picker paid that again a minute after the last look.
Keep the last listing agy actually returned and answer from it: fresh for
ten minutes, then still answered while a probe refreshes behind it, until
the entry is a day old and stops standing in for the machine at all. A
probe that times out or loses auth leaves that entry alone, so a blip no
longer empties a working picker, and the hardcoded mirror is no longer
recorded as if the machine had reported it.
Three things fall out of that and are handled here. A machine whose
sign-in has actually gone bad would otherwise look healthy for a day, so
an auth failure rides along with the catalog it can still serve — and,
because nothing else would re-probe a catalog that is still fresh, a
warning riding on the answer is itself a reason to look again. A failed
probe is not repeated on the very next request either, or a machine where
agy hangs would spawn it once per poll.
An explicit refresh always costs a probe, and never rides one that was
already running when it was asked for.
* fix(agy): let Retry force a fresh model catalog probe
With the catalog held for ten minutes, Retry would otherwise hand back the
answer it was pressed to replace, so the intent travels to the machine:
`?refresh=true` on the machine route, an optional RPC param, and a
one-shot flag on the query so ordinary mount and focus refetches stay
cheap. Every hop is optional, so a hub and a runner on different versions
still talk — the older side ignores it and answers from its cache.
Retry also has to be reachable, and honest, in the state that needs it.
The machine now answers with both a usable catalog and the sign-in failure
behind it, so the picker keeps the list and says why it may be out of
date, with the button right there rather than only once there is nothing
left to show. Pressing it runs agy, which can take tens of seconds, so the
button says so while it does.
The client contract covers both: `agy-models` is the one catalog route
that can carry an `error` on a successful response, and the one that
takes a refresh parameter.
* fix(agy): use the machine catalog in the in-session model picker
New Session already asks the machine what `agy models` lists, but a
session that is already open offered the built-in list in
`shared/src/models.ts`. That list is a hand-maintained mirror, so a model
agy started offering after the last release could be picked for a new
session and not for the one already running.
Point the composer at the same machine catalog. The mirror stays as the
fallback for the moment before the machine answers, and a model the
session is already on is kept selectable — and readable, when it is one
of the known presets — even after the catalog moves on without it.
* fix(agy): announce a model catalog re-check that changed the answer
Serving the last known catalog answers the picker instantly, but a picker
that was already open kept showing that answer until the user closed and
reopened it — the machine had no way to say it had found something newer.
Say it on the stream that already carries machine changes. The machine
daemon — the only process that answers `<machineId>:listAgyModels` —
emits it, and the hub forwards it as `machine-agy-models-updated` with
nothing but the machineId. Namespace resolution, per-machine delivery and
reconnect replay all come from the existing path.
What counts as a change is what the route would answer, not what sits in
the cache. That distinction carries the cases: a sign-in that lapsed or
came back changes no models yet changes what the user is told; a machine
whose agy was signed out has been answering from the hardcoded mirror, and
its first real listing is the largest change there is, for every client
except the one awaiting it.
* fix(agy): re-read the announced machine's model catalog
On `machine-agy-models-updated`, cancel and refetch that one machine's
catalog query — the app's global connection is always subscribed, so an
open picker redraws wherever it is.
Cancelling first is what makes it correct rather than merely likely.
query-core cancels an in-flight fetch only when the query already holds
data, so a picker opening for the first time would otherwise join the
request already on its way and settle on the listing the announcement
replaced. The refetch is answered from the machine's cache, so it starts
no probe and cannot bounce another announcement back.
A reconnect the hub could not replay takes the resync path, which clears
the agy catalogs the same way — that path has no announcement to fall back
on, so it is the one that can least afford to join a stale request.
* fix(agy): keep a model the user picked when the catalog moves under them
The catalog can now change while the New Session form is open, and the
form dropped any selection the machine no longer advertised — including
one the user had just made.
Keep that one, and list it as no longer listed so the form does not imply
agy is still offering it. A model restored from a draft or a saved
preference is still dropped: it may never have been runnable here.
* fix(agy): announce uncached authentication changes
* feat(shared): steer capability gates and live steered signal schemas
- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
agents can deliver queued messages into the active turn (pi, codex,
cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
messages-consumed live signal (never persisted by the hub)
* feat(cli): queue reservations and steered messages-consumed option
- MessageQueue2 gains takeByLocalId/restoreReservation/
beginReservationDispatch/commitReservation so an async steer can reserve
a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery
* feat(codex): mid-turn steer via app-server turn/steer (#888)
- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
reserves the queued row, validates it against the active turn (no
control commands, matching mode hash), injects via turn/steer with an
epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered
* feat(web): Steered badge and steer gating for codex sessions
- HappyUserMessage shows a ↳ Steered badge fed by the live
messages-consumed steered signal, preserved across server echoes and
refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
(upstream typecheck breakage)
* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer
- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
finished): the hub RPC acks once dispatch succeeds — never on the
concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
restores the row so the message still delivers via turn/start, and a
dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure
* fix(codex): reconcile dispatched steers before restoring; align error copy
- A dispatched turn/steer whose completion fails (disconnect / protocol
error) is now reconciled via thread/read by clientUserMessageId before
the queued row is restored — the instruction is only re-delivered by
turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
(Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
and reconcile-rejected outcomes
* fix(codex): consume the row at dispatch; drop background reconcile
- The hub RPC acks and the queue row is consumed as soon as stdin accepts
turn/steer; completion is background-only logging. A dispatched steer is
never restored, so the same localId cannot be re-delivered via turn/start
after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
failure
- steer.completed rejection is always handled (no unhandled rejection on
the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
dispatch failure restores it
* fix(codex): distinguish definite rejection from indeterminate completion
- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
a definite app-server rejection restores the row (instruction was never
accepted, so turn/start cannot duplicate it); an indeterminate outcome
leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
outcome (row stays reserved) and dispatch failure
* fix(codex): reconcile indeterminate steers instead of a permanent reservation
- After an indeterminate completion (disconnect/protocol), reconcile the
thread by clientUserMessageId immediately: accepted → commit + consumed,
provably rejected → restore, still unreadable → keep the reservation and
retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
reconciliation consumes; rejected path restores
* fix(codex): accept all thread item shapes; retry reconcile; ack through abort
- Reconcile matcher accepts userMessage/user_message with clientId/
client_id, matching the shapes the thread parser supports — an accepted
steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
reported steered on dispatch, so commit + messages-consumed must reach
it even when an abort resets the queue in between
* fix(codex): reinit reconnected app-server; keep reconcile retries alive
- thread/read after a disconnect auto-connects a fresh app-server, which
must be initialized before any request — reconcile now ensures
connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
so recovery without external traffic is eventually observed
- launcher mock gains isConnected
* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK
- Reconciliation runs on a self-rescheduling 1s timer independent of the
main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
observe app-server recovery; abort clears nothing implicitly — the ACK
path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
so ensureAppServerInitialized re-initializes a fresh process before
thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
keeps reserved, explicit rejection restores
* fix(codex): bind reconciliation to the launcher lifecycle
- runSteerReconciliation clears any armed retry timer on entry and never
installs a second one, so loop-top and timer-driven passes cannot
multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
pending map is dropped, so an unresolved steer can never respawn an
app-server after cleanup (remote-to-local switch included)
* fix(codex): report steered only after app-server acceptance
- The handler now awaits steer.completed (the inject-acceptance response):
an explicit JSON-RPC rejection surfaces as failed and restores the row
for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
reconciled' and keeps the row reserved while the timer-driven thread
reconciliation runs
- dispatch-failure path also swallows the paired completion rejection
* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait
- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
steer reservation: the hub neither deletes the row nor stamps invoked_at
(new CancelMessageResponse 'busy' status; web restores the optimistic
row); pushIsolateAndClear and reset/close share cancelReservations so
/clear-style commands cannot have a rejected steer resurrect a discarded
prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
lost response is indeterminate and funnels into thread reconciliation
instead of stranding the reservation
- tests updated for the tri-state cancel contract
* fix(codex,web): busy-aware edit flow; bound reconciliation reads
- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
never prefills the composer when the row is inside an async steer, so a
second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
connected-but-silent app-server cannot hold the reservation in-flight
indefinitely
* fix(steer): inFlight-dominated cancel acks; bounded reconciliation
- hub cancel-queued-message acks check inFlight before removed: a stale
duplicate socket reporting removed can no longer delete the durable row
while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
rejection window, a dispatched steer that the app-server never proved
(client ids dropped on restart) is committed instead of polling
thread/read forever
- pre-dispatch failures (abort before write included) never enter
reconciliation — they restore the row and report failure
* fix(steer): persist indeterminate outcomes without replay
* fix(steer): make ambiguous delivery restart-safe
* fix(steer): recover crash-held rows and preserve retry dedup
* fix(steer): ack retries and bound stdin dispatch
* fix(steer): reconcile indeterminate dispatches and serialize retries
* fix(codex): classify stdin callback failures as indeterminate
* fix(steer): recheck indeterminate cancels after ACK
* fix(steer): close retry and abort races
* fix(steer): serialize live retries and abort admission
* fix(steer): distinguish live dispatching from unknown
* fix(steer): keep ACK failures held and reconcile busy cancel
* fix(steer): distinguish held cancel from removal
* fix(store): combine schema v24 migrations
* fix(store): reserve schema v25 for steer delivery state
* fix(steer): keep held cancel state and notify requeue
* fix(steer): release explicitly cancelled unknown reservations
* fix(codex): reject cancelled reservations before native steer
* fix(codex): make reservation restore atomic with state
* fix(codex): terminate abandoned transport writes
* fix(steer): own abandoned app-server lifecycle and consume races
* fix(codex): confirm dispatch and recover abandoned turns
* test(codex): mock abandoned transport callback
* fix(codex): clear visible turn state on transport loss
* fix(steer): claim retries and cover native delivery state
* fix(native): preserve indeterminate state on Android hydration
* fix(steer): make retry claims single-winner
* fix(steer): serialize concurrent retry claims
* fix(socket): tolerate missing steer-state ACK callbacks
* fix(native): serialize retry operations
* docs(web): document unknown steer delivery and retry controls
* fix(steer): handle retry failures and abort-before-connect
* fix(steer): reinitialize after transport loss and finish iOS retry errors
* fix(steer): preserve indeterminate rows across reconnect gaps
* test(web): mock indeterminate queued recovery state
* fix(steer): recover consumed ACK tombstones
* fix(steer): expose consumed cancel tombstones