Files
hapi/docs/api/client-contract/pagination.md
T
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00

195 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Message pagination, windowing, and optimistic sends
**Audience:** Implementers of native HAPI clients (iOS / Android). This page specifies the message paging protocol (`GET /api/sessions/:id/messages`), the epoch reset contract, the tail-sync loop, recommended client windowing, and the optimistic-send / cancel lifecycle. Companion pages: [sse](./sse.md), [messages](./messages.md), [rest](./rest.md).
Source of truth: `shared/src/apiTypes.ts` (`MessagesQuerySchema`, `MessagesResponse`, `SendMessageRequestSchema`), `hub/src/web/routes/messages.ts`, `hub/src/sync/messageService.ts`, `hub/src/store/messages.ts`, reference client `web/src/lib/message-window-store.ts` + `web/src/lib/messages.ts`.
---
## Position key
Messages are ordered by a **compound position**, not by `seq` alone:
```
position = (at, seq) where at = invokedAt ?? createdAt
```
Ascending by `at`, ties broken by `seq`. Rationale: a queued user message sits at its `createdAt` until the agent consumes it, at which point `invokedAt` is stamped and the row **moves forward** to its invocation position. `seq` (per-session insert counter) alone would freeze queued rows at enqueue order. Every cursor in this protocol is therefore a `(seq, at)` **pair** — both halves are always required together.
---
## `GET /api/sessions/:id/messages`
Query parameters (`MessagesQuerySchema`; all numbers coerced from strings):
| Param | Type | Constraint |
|---|---|---|
| `limit` | int | 1–200. **Default 50** when omitted (the web reference always sends 200). |
| `beforeSeq` + `beforeAt` | int + int | Page strictly older than this position. Pairwise required. |
| `afterSeq` + `afterAt` | int + int | Page strictly newer than this position. Pairwise required. |
| `untilSeq` + `untilAt` | int + int | Inclusive snapshot head for a catch-up loop. Pairwise required; **requires an `after` cursor**. |
| `epoch` | int ≥ 0 | Client's cached epoch. **Requires an `after` cursor.** |
Validation rules (violations are `400 {"error":"Invalid query","issues":…}`):
- `beforeAt`⇄`beforeSeq`, `afterAt`⇄`afterSeq`, `untilAt`⇄`untilSeq` must each be provided together.
- `before` and `after` are **mutually exclusive**.
- `until` and `epoch` are only valid alongside `after`.
Session errors: `404` not found, `403` foreign namespace (see [errors](./errors.md)).
### Response shape
```ts
type MessagesResponse = {
messages: DecryptedMessage[] // ascending display order
page: {
direction: 'latest' | 'before' | 'after'
limit: number
epoch: number // server's current epoch for this session
reset: boolean // true ⇒ discard your window, this page replaces it
nextBeforeSeq: number | null // cursor for the next OLDER page
nextBeforeAt: number | null
nextAfterSeq: number | null // cursor for the next NEWER page
nextAfterAt: number | null
snapshotHeadSeq: number | null // newest position at snapshot time
snapshotHeadAt: number | null
hasMore: boolean // more rows exist in the requested direction
}
}
```
### `latest` (no cursor)
Newest `limit` rows by position, **plus** — out of band — every uninvoked local user message (queued rows, including future-scheduled ones), so a fresh client still sees the queued bar even when those rows fall outside the page. The out-of-band rows are pinned to every latest response and do **not** affect the cursor: `nextBefore*` anchors to the oldest row of the position-ordered page proper. `hasMore` = at least one row exists before that. If a page contains only server-side-filtered rows (see [messages](./messages.md)), the hub auto-advances to older pages until it can return something or history is exhausted.
### `before`
Rows strictly older than the cursor. `nextBefore*` = oldest row of this page; `hasMore` = at least one row older than that. The response also carries the current `epoch` — compare it to your cached one (see below).
### `after`
Rows strictly newer than the cursor, bounded by an **inclusive snapshot head** = `min(until, currentHead)` (or whichever exists). This keeps a catch-up loop from chasing messages appended while it runs. Responses:
- Client `epoch` ≠ server epoch ⇒ the server ignores the cursor and returns the **latest page with `reset: true`** (`direction: 'latest'`).
- Snapshot head ≤ cursor ⇒ empty page, `hasMore: false`, `nextAfter*` echoes the cursor.
- Otherwise: `nextAfter*` = last row's position, `hasMore` = `nextAfter < snapshotHead`.
---
## Epoch
`epoch` is a per-session monotonic counter (`message_epochs` table, starts at 0) that is bumped whenever history changes in a way that invalidates composite cursors (`hub/src/store/messages.ts`):
- a new row lands **before** the current head position (out-of-order insert, e.g. transcript import with an earlier timestamp);
- a queued message is deleted (cancel);
- rewind / history replace (`replaceSessionMessagesFrom`);
- messages are copied/merged between sessions (both sides), fork hydration.
Client contract:
- **`after` request** — always send your cached `epoch`. On mismatch the server answers with the latest page and `reset: true`; **discard the entire local window** and replace it with that page.
- **`before` request** — the response's `page.epoch` may differ from your cached one; if it does, your cursors are meaningless: drop cursor state, flag the window for a latest reset, and run a fresh tail sync (web: `fetchOlderMessages` → `epoch-reset` outcome).
- A structural change is also announced live via the `messages-invalidated` SSE event — on the open session, clear the window and tail-sync from scratch.
---
## Tail-sync loop
Run after connect, after an SSE `resume: 'gap'` handshake, on session open, and when told to (`messages-invalidated`). Reference: `runTailSync` in `web/src/lib/message-window-store.ts`.
1. **No usable state** (no newest cursor, no cached epoch, or a reset is pending): `GET …/messages?limit=200` (latest), replace/merge into the window, store `page.epoch`, `nextBefore*` (older-page cursor) and `snapshotHead*` (newest cursor). Done.
2. **Have cursor + epoch**: loop
- `GET …/messages?afterSeq&afterAt&epoch[&untilSeq&untilAt]&limit=200`, where `after` starts at your newest cursor and `until` is the `snapshotHead*` captured from the **first** response of the loop (fixes the target so the loop terminates).
- `page.reset` or `direction: 'latest'` ⇒ replace the window with this page; stop.
- Otherwise merge the rows, advance `after = nextAfter*`, update the newest cursor to `max(current, nextAfter)`; stop when `hasMore` is false.
- Guard: if `nextAfter` did not advance past the previous cursor, abort with an error (protocol violation, do not spin).
New live rows keep arriving via the SSE `message-received` event; ingest them and advance the newest cursor to `max(current, incoming position)`. Only run one tail sync at a time per session; if events force another (e.g. a reset was flagged mid-loop), queue a trailing run.
---
## Client windowing (normative recommendation)
Constants from the web reference (`web/src/lib/message-window-store.ts`):
| Constant | Value | Meaning |
|---|---|---|
| `PAGE_SIZE` | 200 | Request size for every page fetch. |
| `VISIBLE_WINDOW_SIZE` | 400 | Max regular rows kept in **tail** mode (following live bottom). |
| `HISTORY_WINDOW_SIZE` | 600 | Max regular rows kept in **history** mode (user scrolled back). |
| `OLDER_LOAD_WINDOW_SIZE` | 800 | Temporary cap while an older page is being merged (prepend). |
| `AGENT_RUN_WINDOW_SIZE` | 800 | Separate trim bucket for codex `agent-run-*` rows so background-agent traces don't evict chat. |
Rules:
- **Tail mode** trims from the top (oldest dropped). Dropping rows ⇒ set `hasMore: true` and recompute the older-page cursor from the oldest kept row.
- **History mode** trims from the bottom (newest dropped). Dropping newest rows means your window no longer reaches the tail ⇒ flag "latest reset required": on returning to tail mode, discard cursors and fetch a fresh latest page rather than trusting stale ones.
- **Queued rows are never trimmed** (user messages with `invokedAt === null`, see below) — they are re-merged after every trim.
- Persist the window (messages + cursors + epoch) per session for instant cold-start rendering; on re-activation with a persisted cursor, still fetch a fresh latest page first (another client may have advanced the session by many pages) and reconcile.
---
## Optimistic sends
Send: `POST /api/sessions/:id/messages` (see constraints below). Reference: `web/src/hooks/mutations/useSendMessage.ts`, `mergeMessages` in `web/src/lib/messages.ts`.
Lifecycle:
1. Generate a client-side `localId` and append an **optimistic row**: `{id: localId, seq: null, localId, invokedAt: null, scheduledAt, createdAt: now, status:'sending', content: {role:'user', content:{type:'text', text, attachments?}, meta:{deliveryMode}}}`. A row is *optimistic* iff `id === localId`.
2. On POST success: status → `queued` if the session is currently thinking, else `sent`. On failure: drop the row and restore the composer (or keep it as `failed` with a retry affordance when attachments are involved).
3. **Echo**: the hub emits `message-received` carrying the stored row (server `id`, real `seq`, same `localId`). Merging a stored row whose `localId` matches an optimistic row **replaces** the optimistic one, preserving the client-side `status` and any already-known `invokedAt` the server row lacks. Fallback when no `localId` echo matches: drop an optimistic `sent` row when a server user message lands within **10 s** of the same position.
4. **`messages-consumed {localIds, invokedAt}`** (SSE): stamp `invokedAt` and flip status to `sent` on matching rows (skip `failed` ones). This is what moves a message out of the queued bar and into the thread at its invocation position.
5. **`messages-indeterminate {localIds}`** (SSE): the steer outcome is unknown. Keep `invokedAt: null`, mark `deliveryState:'indeterminate'`, exclude the row from automatic replay, and show explicit Retry/Cancel actions.
6. **`messages-requeued {localIds}`** (SSE): an explicit Retry restored normal queue delivery; clear `deliveryState`.
7. **`message-cancelled {messageId, localId?}`** (SSE): remove the row (match either id).
**Queued semantics**: a user message is "queued" iff `invokedAt === null` **strictly**, `deliveryState !== 'indeterminate'`, and `status !== 'failed'`. An indeterminate row remains visible in the unresolved-delivery bar but is not eligible for automatic delivery. `undefined` means already-invoked (rows from pre-V8 hubs omit the field) — only rows explicitly carrying `null` belong in the queued bar. Server-side, rows sent without a `localId` are stamped invoked at insert and can never be queued.
### Queued-state recovery
After a reconnect whose handshake said `resume: 'gap'` (an `ok` resume replayed the consume/cancel events already), the consumed/cancelled events for your queued rows may have been lost. Reference: `web/src/lib/queued-state-reconciliation.ts`.
1. Finish a tail sync.
2. Collect candidate `localId`s: user rows with `invokedAt === null`, excluding optimistic rows still `sending`/`failed`.
3. `POST /api/sessions/:id/messages/queued-state` with `{"localIds": […]}` (max 1000 per call; batch above that) → `{queuedLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}`.
4. Apply `invokedLocalMessages` exactly like `messages-consumed`; drop candidates that are in **neither** list (deleted server-side).
---
## Send constraints
`POST /api/sessions/:id/messages` body (`SendMessageRequestSchema`):
| Field | Type | Rules |
|---|---|---|
| `text` | string | Required (route also accepts empty text when `attachments` is non-empty). |
| `localId` | string | Optional but **required for `scheduledAt`**, and required in practice: without it the row is stamped invoked at insert (no queue/ack/cancel path). |
| `attachments` | `AttachmentMetadata[]` | Optional. Not allowed with `scheduledAt`. |
| `scheduledAt` | epoch ms | Optional. Must be ≤ now + **7 days**; requires `localId`; no attachments; never `steer`. |
| `deliveryMode` | `'queue'` \| `'steer'` | Optional, default `queue`. `steer` is honored only for **Pi-flavor** sessions and never for scheduled sends — the hub silently normalizes everything else to `queue`, and deferred/replayed delivery (reconnect backfill, retries, scheduled release) always degrades `steer` to `queue`. |
Response `{"ok": true}`. Sending to an inactive session returns `409 {"error":"Session is inactive","code":"session_inactive"}` — resume/reopen first, and note the resumed session **may have a different id** (migrate drafts and re-target, see [rest](./rest.md)).
---
## Cancel and steer
**Cancel**: `DELETE /api/sessions/:id/messages/:messageId` — `:messageId` may be the server id **or** the `localId`. Response union (`CancelMessageResponseSchema`):
| Response | Meaning | Client action |
|---|---|---|
| `{"status":"cancelled","localId":string\|null}` | Row deleted (or already gone). Bumps the epoch. | Remove the row. |
| `{"status":"invoked","message":DecryptedMessage}` | Too late — the agent consumed it before the cancel landed. | **Ingest the returned message** as the authoritative row (correct `invokedAt`, status `sent`); do not resurrect the queued snapshot. |
| `{"status":"busy","localId":string}` | A live steer is still resolving. | Restore the row as indeterminate; reconcile queued state before allowing Retry/Cancel. |
Other subscribers learn the same outcome via `message-cancelled` / `messages-consumed` SSE events.
**Steer a queued message into the current turn**: `POST /api/sessions/:id/messages/:messageId/steer` (Pi sessions) → `SteerQueuedMessageResponseSchema`:
| Response | Client action |
|---|---|
| `{"status":"steered","localId"}` | Keep the row queued-side; it is being injected into the live turn. |
| `{"status":"invoked","message"}` | Already consumed — ingest the message. |
| `{"status":"failed","error","localId":string\|null}` | Surface the error; the row remains queued. |