* feat(shared): steer capability gates and live steered signal schemas - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which agents can deliver queued messages into the active turn (pi, codex, cursor ACP; legacy stream-json cursor excluded) - AgentState.steeringActive, DecryptedMessage.steered and messages-consumed live signal (never persisted by the hub) * feat(cli): queue reservations and steered messages-consumed option - MessageQueue2 gains takeByLocalId/restoreReservation/ beginReservationDispatch/commitReservation so an async steer can reserve a queued row without racing the main loop's turn/start drain - emitMessagesConsumed accepts steered: true to mark mid-turn delivery * feat(codex): mid-turn steer via app-server turn/steer (#888) - CodexAppServerClient.steerTurn + TurnSteerParams/Response types - CodexRemoteLauncher registers the steer-queued-message RPC handler: reserves the queued row, validates it against the active turn (no control commands, matching mode hash), injects via turn/steer with an epoch guard that invalidates in-flight steers on abort/cleanup - steeringActive agent state tracks the active-turn window - hub syncEngine gate opens to codex; messages-consumed relays steered * feat(web): Steered badge and steer gating for codex sessions - HappyUserMessage shows a ↳ Steered badge fed by the live messages-consumed steered signal, preserved across server echoes and refetches (mergeMessages carries the optimistic marker) - SessionChat gates canSteer via isSteeringSupportedForSession instead of the pi-only check - clearStaleQueuedStatus normalizes a queued status on an invoked message - fix(web): drop duplicate showSessionSummaryInChat in markdown test (upstream typecheck breakage) * fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise codex and pi only; cursor joins when its soft-steer handler lands (#1609) - turn/steer now splits dispatch (stdin accepted) from completion (turn finished): the hub RPC acks once dispatch succeeds — never on the concurrent turn's completion, which can exceed the 30s RPC window - queue row commits only after the turn settles; a rejected/aborted steer restores the row so the message still delivers via turn/start, and a dispatched steer is never restored (no duplicate delivery) - steer carries clientUserMessageId (echoed as userMessage.clientId) so ambiguous transport failures can reconcile the thread later - client tests cover dispatch/complete split and stdin-write failure * fix(codex): reconcile dispatched steers before restoring; align error copy - A dispatched turn/steer whose completion fails (disconnect / protocol error) is now reconciled via thread/read by clientUserMessageId before the queued row is restored — the instruction is only re-delivered by turn/start when the thread never received it - Reconcile targets the pinned steer thread, not whichever turn is current when completion fails - syncEngine unsupported-flavor error now matches the capability gate (Pi and Codex only until the cursor handler lands) - launcher tests cover steer success (ack on dispatch), reconcile-accepted and reconcile-rejected outcomes * fix(codex): consume the row at dispatch; drop background reconcile - The hub RPC acks and the queue row is consumed as soon as stdin accepts turn/steer; completion is background-only logging. A dispatched steer is never restored, so the same localId cannot be re-delivered via turn/start after the caller was told the steer succeeded - Dispatch failure (stdin write error) still restores the row and reports failure - steer.completed rejection is always handled (no unhandled rejection on the dispatch-failure path) - tests updated: completion failure after dispatch keeps the row consumed; dispatch failure restores it * fix(codex): distinguish definite rejection from indeterminate completion - Transport-level failures (timeout, abort, disconnect, spawn, protocol) carry an indeterminate marker; explicit JSON-RPC error responses do not - After a dispatched steer, turn completion resolves → commit + consumed; a definite app-server rejection restores the row (instruction was never accepted, so turn/start cannot duplicate it); an indeterminate outcome leaves the row reserved so it can never be delivered twice - Completion handling registers before awaiting dispatch so the dispatch-failure path cannot leak an unhandled rejection - client/launcher tests cover explicit rejection (restore), indeterminate outcome (row stays reserved) and dispatch failure * fix(codex): reconcile indeterminate steers instead of a permanent reservation - After an indeterminate completion (disconnect/protocol), reconcile the thread by clientUserMessageId immediately: accepted → commit + consumed, provably rejected → restore, still unreadable → keep the reservation and retry from the main-loop top on later passes (post-reconnect) - A row never sits in dispatching forever: the hub cannot stamp it invoked while the instruction may never have been accepted - tests: indeterminate keeps reserved while thread unreadable; accepted reconciliation consumes; rejected path restores * fix(codex): accept all thread item shapes; retry reconcile; ack through abort - Reconcile matcher accepts userMessage/user_message with clientId/ client_id, matching the shapes the thread parser supports — an accepted steer can no longer be misclassified as rejected - A pending reconciliation schedules a wakeLoop retry, so a temporary app-server outage cannot strand the reservation behind waitForTurnOrRecovery - The success-path ACK no longer checks the steer epoch: the hub already reported steered on dispatch, so commit + messages-consumed must reach it even when an abort resets the queue in between * fix(codex): reinit reconnected app-server; keep reconcile retries alive - thread/read after a disconnect auto-connects a fresh app-server, which must be initialized before any request — reconcile now ensures connect + initialize (isConnected getter added to the client) - every still-unknown loop-top reconciliation schedules the next retry, so recovery without external traffic is eventually observed - launcher mock gains isConnected * fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK - Reconciliation runs on a self-rescheduling 1s timer independent of the main loop (wakes it too), so idle loops and waitForTurnOrRecovery still observe app-server recovery; abort clears nothing implicitly — the ACK path commits and consumes even when the reservation was cancelled - Absence of a durable client id is ambiguous: unmatched reads stay 'unknown' and keep retrying instead of restoring the row - CodexAppServerClient tracks initialized state (reset on disconnect/exit) so ensureAppServerInitialized re-initializes a fresh process before thread/read; initialize failures leave the flag false for the next retry - tests: accepted reconciliation via scheduled timer, indeterminate keeps reserved, explicit rejection restores * fix(codex): bind reconciliation to the launcher lifecycle - runSteerReconciliation clears any armed retry timer on entry and never installs a second one, so loop-top and timer-driven passes cannot multiply - shuttingDown is set when the main loop ends: timers are cleared and the pending map is dropped, so an unresolved steer can never respawn an app-server after cleanup (remote-to-local switch included) * fix(codex): report steered only after app-server acceptance - The handler now awaits steer.completed (the inject-acceptance response): an explicit JSON-RPC rejection surfaces as failed and restores the row for the normal turn/start path instead of a false steered - Transport failure after dispatch reports 'Steer outcome is being reconciled' and keeps the row reserved while the timer-driven thread reconciliation runs - dispatch-failure path also swallows the paired completion rejection * fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait - MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching steer reservation: the hub neither deletes the row nor stamps invoked_at (new CancelMessageResponse 'busy' status; web restores the optimistic row); pushIsolateAndClear and reset/close share cancelReservations so /clear-style commands cannot have a rejected steer resurrect a discarded prompt - turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a lost response is indeterminate and funnels into thread reconciliation instead of stranding the reservation - tests updated for the tri-state cancel contract * fix(codex,web): busy-aware edit flow; bound reconciliation reads - QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it never prefills the composer when the row is inside an async steer, so a second client cannot send a duplicate - reconcileSteerByClientId bounds thread/read with a 5s timeout so a connected-but-silent app-server cannot hold the reservation in-flight indefinitely * fix(steer): inFlight-dominated cancel acks; bounded reconciliation - hub cancel-queued-message acks check inFlight before removed: a stale duplicate socket reporting removed can no longer delete the durable row while another socket is dispatching the steer - reconciliation entries expire after 60s and mark delivered: after the rejection window, a dispatched steer that the app-server never proved (client ids dropped on restart) is committed instead of polling thread/read forever - pre-dispatch failures (abort before write included) never enter reconciliation — they restore the row and report failure * fix(steer): persist indeterminate outcomes without replay * fix(steer): make ambiguous delivery restart-safe * fix(steer): recover crash-held rows and preserve retry dedup * fix(steer): ack retries and bound stdin dispatch * fix(steer): reconcile indeterminate dispatches and serialize retries * fix(codex): classify stdin callback failures as indeterminate * fix(steer): recheck indeterminate cancels after ACK * fix(steer): close retry and abort races * fix(steer): serialize live retries and abort admission * fix(steer): distinguish live dispatching from unknown * fix(steer): keep ACK failures held and reconcile busy cancel * fix(steer): distinguish held cancel from removal * fix(store): combine schema v24 migrations * fix(store): reserve schema v25 for steer delivery state * fix(steer): keep held cancel state and notify requeue * fix(steer): release explicitly cancelled unknown reservations * fix(codex): reject cancelled reservations before native steer * fix(codex): make reservation restore atomic with state * fix(codex): terminate abandoned transport writes * fix(steer): own abandoned app-server lifecycle and consume races * fix(codex): confirm dispatch and recover abandoned turns * test(codex): mock abandoned transport callback * fix(codex): clear visible turn state on transport loss * fix(steer): claim retries and cover native delivery state * fix(native): preserve indeterminate state on Android hydration * fix(steer): make retry claims single-winner * fix(steer): serialize concurrent retry claims * fix(socket): tolerate missing steer-state ACK callbacks * fix(native): serialize retry operations * docs(web): document unknown steer delivery and retry controls * fix(steer): handle retry failures and abort-before-connect * fix(steer): reinitialize after transport loss and finish iOS retry errors * fix(steer): preserve indeterminate rows across reconnect gaps * test(web): mock indeterminate queued recovery state * fix(steer): recover consumed ACK tombstones * fix(steer): expose consumed cancel tombstones
14 KiB
Message pagination, windowing, and optimistic sends
Audience: Implementers of native HAPI clients (iOS / Android). This page specifies the message paging protocol (GET /api/sessions/:id/messages), the epoch reset contract, the tail-sync loop, recommended client windowing, and the optimistic-send / cancel lifecycle. Companion pages: sse, messages, rest.
Source of truth: shared/src/apiTypes.ts (MessagesQuerySchema, MessagesResponse, SendMessageRequestSchema), hub/src/web/routes/messages.ts, hub/src/sync/messageService.ts, hub/src/store/messages.ts, reference client web/src/lib/message-window-store.ts + web/src/lib/messages.ts.
Position key
Messages are ordered by a compound position, not by seq alone:
position = (at, seq) where at = invokedAt ?? createdAt
Ascending by at, ties broken by seq. Rationale: a queued user message sits at its createdAt until the agent consumes it, at which point invokedAt is stamped and the row moves forward to its invocation position. seq (per-session insert counter) alone would freeze queued rows at enqueue order. Every cursor in this protocol is therefore a (seq, at) pair — both halves are always required together.
GET /api/sessions/:id/messages
Query parameters (MessagesQuerySchema; all numbers coerced from strings):
| Param | Type | Constraint |
|---|---|---|
limit |
int | 1–200. Default 50 when omitted (the web reference always sends 200). |
beforeSeq + beforeAt |
int + int | Page strictly older than this position. Pairwise required. |
afterSeq + afterAt |
int + int | Page strictly newer than this position. Pairwise required. |
untilSeq + untilAt |
int + int | Inclusive snapshot head for a catch-up loop. Pairwise required; requires an after cursor. |
epoch |
int ≥ 0 | Client's cached epoch. Requires an after cursor. |
Validation rules (violations are 400 {"error":"Invalid query","issues":…}):
beforeAt⇄beforeSeq,afterAt⇄afterSeq,untilAt⇄untilSeqmust each be provided together.beforeandafterare mutually exclusive.untilandepochare only valid alongsideafter.
Session errors: 404 not found, 403 foreign namespace (see errors).
Response shape
type MessagesResponse = {
messages: DecryptedMessage[] // ascending display order
page: {
direction: 'latest' | 'before' | 'after'
limit: number
epoch: number // server's current epoch for this session
reset: boolean // true ⇒ discard your window, this page replaces it
nextBeforeSeq: number | null // cursor for the next OLDER page
nextBeforeAt: number | null
nextAfterSeq: number | null // cursor for the next NEWER page
nextAfterAt: number | null
snapshotHeadSeq: number | null // newest position at snapshot time
snapshotHeadAt: number | null
hasMore: boolean // more rows exist in the requested direction
}
}
latest (no cursor)
Newest limit rows by position, plus — out of band — every uninvoked local user message (queued rows, including future-scheduled ones), so a fresh client still sees the queued bar even when those rows fall outside the page. The out-of-band rows are pinned to every latest response and do not affect the cursor: nextBefore* anchors to the oldest row of the position-ordered page proper. hasMore = at least one row exists before that. If a page contains only server-side-filtered rows (see messages), the hub auto-advances to older pages until it can return something or history is exhausted.
before
Rows strictly older than the cursor. nextBefore* = oldest row of this page; hasMore = at least one row older than that. The response also carries the current epoch — compare it to your cached one (see below).
after
Rows strictly newer than the cursor, bounded by an inclusive snapshot head = min(until, currentHead) (or whichever exists). This keeps a catch-up loop from chasing messages appended while it runs. Responses:
- Client
epoch≠ server epoch ⇒ the server ignores the cursor and returns the latest page withreset: true(direction: 'latest'). - Snapshot head ≤ cursor ⇒ empty page,
hasMore: false,nextAfter*echoes the cursor. - Otherwise:
nextAfter*= last row's position,hasMore=nextAfter < snapshotHead.
Epoch
epoch is a per-session monotonic counter (message_epochs table, starts at 0) that is bumped whenever history changes in a way that invalidates composite cursors (hub/src/store/messages.ts):
- a new row lands before the current head position (out-of-order insert, e.g. transcript import with an earlier timestamp);
- a queued message is deleted (cancel);
- rewind / history replace (
replaceSessionMessagesFrom); - messages are copied/merged between sessions (both sides), fork hydration.
Client contract:
afterrequest — always send your cachedepoch. On mismatch the server answers with the latest page andreset: true; discard the entire local window and replace it with that page.beforerequest — the response'spage.epochmay differ from your cached one; if it does, your cursors are meaningless: drop cursor state, flag the window for a latest reset, and run a fresh tail sync (web:fetchOlderMessages→epoch-resetoutcome).- A structural change is also announced live via the
messages-invalidatedSSE event — on the open session, clear the window and tail-sync from scratch.
Tail-sync loop
Run after connect, after an SSE resume: 'gap' handshake, on session open, and when told to (messages-invalidated). Reference: runTailSync in web/src/lib/message-window-store.ts.
- No usable state (no newest cursor, no cached epoch, or a reset is pending):
GET …/messages?limit=200(latest), replace/merge into the window, storepage.epoch,nextBefore*(older-page cursor) andsnapshotHead*(newest cursor). Done. - Have cursor + epoch: loop
GET …/messages?afterSeq&afterAt&epoch[&untilSeq&untilAt]&limit=200, whereafterstarts at your newest cursor anduntilis thesnapshotHead*captured from the first response of the loop (fixes the target so the loop terminates).page.resetordirection: 'latest'⇒ replace the window with this page; stop.- Otherwise merge the rows, advance
after = nextAfter*, update the newest cursor tomax(current, nextAfter); stop whenhasMoreis false. - Guard: if
nextAfterdid not advance past the previous cursor, abort with an error (protocol violation, do not spin).
New live rows keep arriving via the SSE message-received event; ingest them and advance the newest cursor to max(current, incoming position). Only run one tail sync at a time per session; if events force another (e.g. a reset was flagged mid-loop), queue a trailing run.
Client windowing (normative recommendation)
Constants from the web reference (web/src/lib/message-window-store.ts):
| Constant | Value | Meaning |
|---|---|---|
PAGE_SIZE |
200 | Request size for every page fetch. |
VISIBLE_WINDOW_SIZE |
400 | Max regular rows kept in tail mode (following live bottom). |
HISTORY_WINDOW_SIZE |
600 | Max regular rows kept in history mode (user scrolled back). |
OLDER_LOAD_WINDOW_SIZE |
800 | Temporary cap while an older page is being merged (prepend). |
AGENT_RUN_WINDOW_SIZE |
800 | Separate trim bucket for codex agent-run-* rows so background-agent traces don't evict chat. |
Rules:
- Tail mode trims from the top (oldest dropped). Dropping rows ⇒ set
hasMore: trueand recompute the older-page cursor from the oldest kept row. - History mode trims from the bottom (newest dropped). Dropping newest rows means your window no longer reaches the tail ⇒ flag "latest reset required": on returning to tail mode, discard cursors and fetch a fresh latest page rather than trusting stale ones.
- Queued rows are never trimmed (user messages with
invokedAt === null, see below) — they are re-merged after every trim. - Persist the window (messages + cursors + epoch) per session for instant cold-start rendering; on re-activation with a persisted cursor, still fetch a fresh latest page first (another client may have advanced the session by many pages) and reconcile.
Optimistic sends
Send: POST /api/sessions/:id/messages (see constraints below). Reference: web/src/hooks/mutations/useSendMessage.ts, mergeMessages in web/src/lib/messages.ts.
Lifecycle:
- Generate a client-side
localIdand append an optimistic row:{id: localId, seq: null, localId, invokedAt: null, scheduledAt, createdAt: now, status:'sending', content: {role:'user', content:{type:'text', text, attachments?}, meta:{deliveryMode}}}. A row is optimistic iffid === localId. - On POST success: status →
queuedif the session is currently thinking, elsesent. On failure: drop the row and restore the composer (or keep it asfailedwith a retry affordance when attachments are involved). - Echo: the hub emits
message-receivedcarrying the stored row (serverid, realseq, samelocalId). Merging a stored row whoselocalIdmatches an optimistic row replaces the optimistic one, preserving the client-sidestatusand any already-knowninvokedAtthe server row lacks. Fallback when nolocalIdecho matches: drop an optimisticsentrow when a server user message lands within 10 s of the same position. messages-consumed {localIds, invokedAt}(SSE): stampinvokedAtand flip status tosenton matching rows (skipfailedones). This is what moves a message out of the queued bar and into the thread at its invocation position.messages-indeterminate {localIds}(SSE): the steer outcome is unknown. KeepinvokedAt: null, markdeliveryState:'indeterminate', exclude the row from automatic replay, and show explicit Retry/Cancel actions.messages-requeued {localIds}(SSE): an explicit Retry restored normal queue delivery; cleardeliveryState.message-cancelled {messageId, localId?}(SSE): remove the row (match either id).
Queued semantics: a user message is "queued" iff invokedAt === null strictly, deliveryState !== 'indeterminate', and status !== 'failed'. An indeterminate row remains visible in the unresolved-delivery bar but is not eligible for automatic delivery. undefined means already-invoked (rows from pre-V8 hubs omit the field) — only rows explicitly carrying null belong in the queued bar. Server-side, rows sent without a localId are stamped invoked at insert and can never be queued.
Queued-state recovery
After a reconnect whose handshake said resume: 'gap' (an ok resume replayed the consume/cancel events already), the consumed/cancelled events for your queued rows may have been lost. Reference: web/src/lib/queued-state-reconciliation.ts.
- Finish a tail sync.
- Collect candidate
localIds: user rows withinvokedAt === null, excluding optimistic rows stillsending/failed. POST /api/sessions/:id/messages/queued-statewith{"localIds": […]}(max 1000 per call; batch above that) →{queuedLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}.- Apply
invokedLocalMessagesexactly likemessages-consumed; drop candidates that are in neither list (deleted server-side).
Send constraints
POST /api/sessions/:id/messages body (SendMessageRequestSchema):
| Field | Type | Rules |
|---|---|---|
text |
string | Required (route also accepts empty text when attachments is non-empty). |
localId |
string | Optional but required for scheduledAt, and required in practice: without it the row is stamped invoked at insert (no queue/ack/cancel path). |
attachments |
AttachmentMetadata[] |
Optional. Not allowed with scheduledAt. |
scheduledAt |
epoch ms | Optional. Must be ≤ now + 7 days; requires localId; no attachments; never steer. |
deliveryMode |
'queue' | 'steer' |
Optional, default queue. steer is honored only for Pi-flavor sessions and never for scheduled sends — the hub silently normalizes everything else to queue, and deferred/replayed delivery (reconnect backfill, retries, scheduled release) always degrades steer to queue. |
Response {"ok": true}. Sending to an inactive session returns 409 {"error":"Session is inactive","code":"session_inactive"} — resume/reopen first, and note the resumed session may have a different id (migrate drafts and re-target, see rest).
Cancel and steer
Cancel: DELETE /api/sessions/:id/messages/:messageId — :messageId may be the server id or the localId. Response union (CancelMessageResponseSchema):
| Response | Meaning | Client action |
|---|---|---|
{"status":"cancelled","localId":string|null} |
Row deleted (or already gone). Bumps the epoch. | Remove the row. |
{"status":"invoked","message":DecryptedMessage} |
Too late — the agent consumed it before the cancel landed. | Ingest the returned message as the authoritative row (correct invokedAt, status sent); do not resurrect the queued snapshot. |
{"status":"busy","localId":string} |
A live steer is still resolving. | Restore the row as indeterminate; reconcile queued state before allowing Retry/Cancel. |
Other subscribers learn the same outcome via message-cancelled / messages-consumed SSE events.
Steer a queued message into the current turn: POST /api/sessions/:id/messages/:messageId/steer (Pi sessions) → SteerQueuedMessageResponseSchema:
| Response | Client action |
|---|---|
{"status":"steered","localId"} |
Keep the row queued-side; it is being injected into the live turn. |
{"status":"invoked","message"} |
Already consumed — ingest the message. |
{"status":"failed","error","localId":string|null} |
Surface the error; the row remains queued. |