Files
hapi/ios/Packages/HapiKit/Sources/HapiClient/Stores/SyncEventRouting.swift
T
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00

101 lines
4.3 KiB
Swift

import Foundation
import HapiProtocol
/// Fan-out from an SSE subscription's event stream to the stores — the
/// session-list half of the event routing, mirroring the global-scope rules
/// of `web/src/hooks/useSSE.ts` via the Android port
/// (`SyncEventRouter` + `StoreSyncTargets`):
///
/// - session lifecycle events → `SessionListStoring.applySessionEvent`;
/// - `machine-updated` → `MachineListStoring.applyMachineEvent`;
/// - global-scope message-stream events refresh the session list where the
/// web invalidates it (`messages-invalidated`, `messages-consumed`,
/// `message-cancelled`, `scheduled-matured`, and a `message-received`
/// carrying `scheduledAt` — they all move the hub-computed
/// scheduled/queued fields the client cannot derive);
/// - `toast` → the injected callback;
/// - a `gap` handshake verdict triggers the full REST resync (session list +
/// cached details + machines); `ok` means the hub replays every missed
/// event right after the handshake, so the REST resync is skipped;
/// - `heartbeat`/`connection-changed`/unknown types are engine/no-op
/// territory here.
///
/// The app's `HubSession` owns one of these per hub and calls
/// ``handleHandshake(resume:)`` / ``route(_:scope:)`` from its SSE consume
/// loop. Since M2f the app's per-chat `ChatSession` reuses this router for
/// the session-scope pipe's non-message events and delivers the
/// message-stream family to the open `MessageWindowController` itself —
/// awaiting each ingest from its single consume task, which is what
/// preserves SSE arrival order into the window actor (the role the Android
/// port's `StoreSyncTargets` channel plays).
@MainActor
public struct SyncEventRouter {
private let sessions: any SessionListStoring
private let machines: any MachineListStoring
private let onToast: @MainActor (ToastPayload) -> Void
public init(
sessions: any SessionListStoring,
machines: any MachineListStoring,
onToast: @escaping @MainActor (ToastPayload) -> Void = { _ in }
) {
self.sessions = sessions
self.machines = machines
self.onToast = onToast
}
/// `resume == .ok` skips the resync (the replay that follows covers every
/// missed event); `.gap` — or, defensively, anything else — triggers the
/// full refetch. (`SSEClient` already maps an absent wire verdict to
/// `.gap`, per the contract's absence rule.)
public func handleHandshake(resume: ResumeVerdict?) {
guard resume != .ok else { return }
requestFullResync()
}
/// Full REST resync: session list + cached details, then machines.
/// Failures are swallowed — offline keeps the snapshot state and the
/// next reconnect retries.
public func requestFullResync() {
let sessions = self.sessions
let machines = self.machines
Task { @MainActor in
try? await sessions.fullResync()
try? await machines.refresh()
}
}
/// Routes one decoded `SyncEvent` from the subscription identified by
/// `scope` (the dual-subscription model: the global pipe drives the
/// list; per-chat `.session` pipes arrive with M2f and their
/// message-stream events belong to the message window, not the list).
public func route(_ event: SyncEvent, scope: SSEScope = .global) {
switch event {
case .sessionAdded, .sessionUpdated, .sessionRemoved, .sessionEnded:
sessions.applySessionEvent(event)
case .machineUpdated(_, let machineId, let data):
machines.applyMachineEvent(machineId: machineId, data: data)
case .messagesInvalidated, .messagesConsumed, .messagesIndeterminate, .messagesRequeued, .messageCancelled, .scheduledMatured:
guard scope == .global else { return }
sessions.scheduleRefresh()
case .messageReceived(_, _, let message):
guard scope == .global else { return }
if message.scheduledAt != nil {
sessions.scheduleRefresh()
}
case .toast(_, let payload):
onToast(payload)
case .heartbeat, .connectionChanged, .unknown:
// Heartbeats feed the SSE watchdog; the handshake surfaces
// through `handleHandshake`; unknown types are forward-compat
// no-ops.
break
}
}
}