* fix(web): keep streamed reasoning/text block ids stable across snapshot rows
Streaming snapshots of one stream (pi/codex reasoning and text) arrive as
separate message rows, and the window store retires older rows as newer
snapshots land. The timeline derived the block id from whichever row was
first seen, so the id (and the threadMessageId built from it) churned on
every snapshot, remounting the rendered reasoning panel mid-stream and
replaying its open animation — the panel visibly flashed/re-rendered on
every snapshot tick.
Derive the block id from the stream id when present (unique per stream,
stable across snapshot rows) so the block is updated in place and the
smooth streaming keeps appending to the previous text. Row-derived ids
remain the fallback for content without a stream id.
Also rerun gen:fixtures to refresh the two golden fixtures affected by
the new id shape.
* fix(ios,android): mirror stream-stable block ids in native chat ports
The native HapiKit (Swift) and protocol (Kotlin) chat pipelines are ports
of the web reducerTimeline and are pinned by the same golden fixtures in
shared/fixtures/chat. After the web-side change to derive streamed
reasoning/text block ids from the stream id, the ports still produced
row-derived ids, so the iOS/Android fixture conformance suites went red
on the two refreshed fixtures.
Apply the same streamId-first id derivation (row-derived fallback kept)
to both ports so all three pipelines project identical block ids.
* fix(web,ios,android): reject blank stream ids as block identity
Blank ('' or whitespace-only) stream ids are not streams per the wire
semantics in shared/src/messages.ts (readReasoningStreamId trims before
accepting). The previous nullish fallback let accepted payloads carrying
blank ids through, so every such row shared one empty block id: the
merge maps collided and assistant-ui occurrence suffixes churned with
list position, reintroducing remounts.
Normalize with a trim guard in all three pipelines (web, HapiKit,
protocol) and add a web regression test covering both empty and
whitespace-only ids.
* fix(ios): use normalized stream id for block construction identity
The blank-id guard was applied to lookup and map insertion but block
construction still read the raw optional, so accepted payloads carrying
blank/whitespace ids produced blocks sharing one blank SwiftUI identity
instead of falling back to row-derived ids (web/Android already used the
normalized local). Hoist the nonBlankStreamId result and reuse it for
lookup, block identity, and insertion in both the text and reasoning
branches.
Also add native coverage for stream identity: stream-id derivation for
text/reasoning plus blank ('' and whitespace-only) fallbacks, which the
golden fixtures do not exercise.
* fix(web): pin blank stream-id identity contract in golden fixtures
Update the two stale fixture descriptions (stream-keyed blocks are now
keyed by the stream id, not the first message) and add a generated
conformance fixture covering empty and whitespace-only codex data.id
values for both reasoning and text: blank ids are not stream identities,
so each payload keeps its own row-derived block id instead of collapsing
onto a shared blank identity. Web, iOS, and Android all run this same
golden fixture.
* feat(hub): make title provider max_tokens and timeout env-tunable
Reasoning models used as title providers (e.g. GLM thinking models) need
more than 64 completion tokens and more than the hardcoded 10s timeout to
emit a title, and the only workaround was patching the compiled binary
after every install.
Expose both knobs via HAPI_TITLE_PROVIDER_MAX_TOKENS and
HAPI_TITLE_PROVIDER_TIMEOUT_MS, following the existing
HAPI_TITLE_SUGGESTION_RATE_LIMIT pattern; defaults are unchanged.
* docs(hub): document title provider max_tokens/timeout env knobs
Add the two new HAPI_TITLE_PROVIDER_* variables to the title-provider
configuration table in the installation guide, and extend the provider
test to cover the timeout abort path (the signal fires and rejects the
in-flight request).
---------
Co-authored-by: HongChenGG <HongChenGG@users.noreply.github.com>
* fix(acp): carry the live reasoning marker on the wire payload
ACP agents stream thoughts a token at a time, so the handler coalesces
them into a buffer and re-sends the whole buffer under a stable stream
id every 250ms. The converter dropped the marker that says a payload is
one of those throttled snapshots, leaving the hub unable to tell a
replaceable snapshot from the settled message that closes the stream.
Mirrors how the text variant already forwards streamSnapshot.
* fix(hub): keep one stored message per reasoning stream
OpenCode reasoning arrives as a series of growing snapshots sharing one
stream id, and every snapshot was persisted as its own message. A 26h
session reached 48,844 rows and 63MB, and because the web budgets a
fixed number of messages, its 400-message window covered barely three
minutes of conversation — scrolling up walked through duplicate
snapshots instead of history.
Retire a stream's earlier live snapshots once their replacement is
stored. Sweeping only after the insert matters: the two statements are
separate transactions, so clearing first would leave a window where a
crash takes the whole stream. Only rows marked live are eligible and the
replacement is spared, so a stream always keeps at least one row and the
settled message that closes it is never removed.
Live rendering is unchanged: the web still receives every snapshot and
already folds them by stream id.
* fix(web): spend the message window on conversation, not repeated snapshots
The window budgets raw messages, but a reasoning stream renders as a
single folded block no matter how many snapshots it arrived in. On
sessions recorded before the hub started retiring them, those snapshots
fill the window on their own: in one 26h session the newest 400 messages
covered 202 seconds, so scrolling up paged through duplicates instead of
history.
Collapse each stream to its newest snapshot before trimming. Rendering
is unchanged — the timeline already folds them by stream id — and rows
without a stream id are never touched.
* fix(ios,android): port reasoning-snapshot compaction to the native windows
The window logic in HapiProtocol and :core:protocol is a one-to-one port
of the web store, so collapsing superseded reasoning snapshots only on
the web left the native windows budgeting raw snapshot rows. The hub
stores one row per stream now, but a client that already holds the older
snapshots still spends its window on them.
Add the same stream-id reader and compaction to both ports, in the shape
each already uses for agent-run rows, and pin the behaviour with a
pagination fixture. Both fixture suites enumerate shared/fixtures/pagination
from disk, so the ports cannot drift from the web again without CI saying
so.
* feat(shared): steer capability gates and live steered signal schemas
- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
agents can deliver queued messages into the active turn (pi, codex,
cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
messages-consumed live signal (never persisted by the hub)
* feat(cli): queue reservations and steered messages-consumed option
- MessageQueue2 gains takeByLocalId/restoreReservation/
beginReservationDispatch/commitReservation so an async steer can reserve
a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery
* feat(codex): mid-turn steer via app-server turn/steer (#888)
- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
reserves the queued row, validates it against the active turn (no
control commands, matching mode hash), injects via turn/steer with an
epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered
* feat(web): Steered badge and steer gating for codex sessions
- HappyUserMessage shows a ↳ Steered badge fed by the live
messages-consumed steered signal, preserved across server echoes and
refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
(upstream typecheck breakage)
* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer
- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
finished): the hub RPC acks once dispatch succeeds — never on the
concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
restores the row so the message still delivers via turn/start, and a
dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure
* fix(codex): reconcile dispatched steers before restoring; align error copy
- A dispatched turn/steer whose completion fails (disconnect / protocol
error) is now reconciled via thread/read by clientUserMessageId before
the queued row is restored — the instruction is only re-delivered by
turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
(Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
and reconcile-rejected outcomes
* fix(codex): consume the row at dispatch; drop background reconcile
- The hub RPC acks and the queue row is consumed as soon as stdin accepts
turn/steer; completion is background-only logging. A dispatched steer is
never restored, so the same localId cannot be re-delivered via turn/start
after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
failure
- steer.completed rejection is always handled (no unhandled rejection on
the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
dispatch failure restores it
* fix(codex): distinguish definite rejection from indeterminate completion
- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
a definite app-server rejection restores the row (instruction was never
accepted, so turn/start cannot duplicate it); an indeterminate outcome
leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
outcome (row stays reserved) and dispatch failure
* fix(codex): reconcile indeterminate steers instead of a permanent reservation
- After an indeterminate completion (disconnect/protocol), reconcile the
thread by clientUserMessageId immediately: accepted → commit + consumed,
provably rejected → restore, still unreadable → keep the reservation and
retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
reconciliation consumes; rejected path restores
* fix(codex): accept all thread item shapes; retry reconcile; ack through abort
- Reconcile matcher accepts userMessage/user_message with clientId/
client_id, matching the shapes the thread parser supports — an accepted
steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
reported steered on dispatch, so commit + messages-consumed must reach
it even when an abort resets the queue in between
* fix(codex): reinit reconnected app-server; keep reconcile retries alive
- thread/read after a disconnect auto-connects a fresh app-server, which
must be initialized before any request — reconcile now ensures
connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
so recovery without external traffic is eventually observed
- launcher mock gains isConnected
* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK
- Reconciliation runs on a self-rescheduling 1s timer independent of the
main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
observe app-server recovery; abort clears nothing implicitly — the ACK
path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
so ensureAppServerInitialized re-initializes a fresh process before
thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
keeps reserved, explicit rejection restores
* fix(codex): bind reconciliation to the launcher lifecycle
- runSteerReconciliation clears any armed retry timer on entry and never
installs a second one, so loop-top and timer-driven passes cannot
multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
pending map is dropped, so an unresolved steer can never respawn an
app-server after cleanup (remote-to-local switch included)
* fix(codex): report steered only after app-server acceptance
- The handler now awaits steer.completed (the inject-acceptance response):
an explicit JSON-RPC rejection surfaces as failed and restores the row
for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
reconciled' and keeps the row reserved while the timer-driven thread
reconciliation runs
- dispatch-failure path also swallows the paired completion rejection
* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait
- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
steer reservation: the hub neither deletes the row nor stamps invoked_at
(new CancelMessageResponse 'busy' status; web restores the optimistic
row); pushIsolateAndClear and reset/close share cancelReservations so
/clear-style commands cannot have a rejected steer resurrect a discarded
prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
lost response is indeterminate and funnels into thread reconciliation
instead of stranding the reservation
- tests updated for the tri-state cancel contract
* fix(codex,web): busy-aware edit flow; bound reconciliation reads
- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
never prefills the composer when the row is inside an async steer, so a
second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
connected-but-silent app-server cannot hold the reservation in-flight
indefinitely
* fix(steer): inFlight-dominated cancel acks; bounded reconciliation
- hub cancel-queued-message acks check inFlight before removed: a stale
duplicate socket reporting removed can no longer delete the durable row
while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
rejection window, a dispatched steer that the app-server never proved
(client ids dropped on restart) is committed instead of polling
thread/read forever
- pre-dispatch failures (abort before write included) never enter
reconciliation — they restore the row and report failure
* fix(steer): persist indeterminate outcomes without replay
* fix(steer): make ambiguous delivery restart-safe
* fix(steer): recover crash-held rows and preserve retry dedup
* fix(steer): ack retries and bound stdin dispatch
* fix(steer): reconcile indeterminate dispatches and serialize retries
* fix(codex): classify stdin callback failures as indeterminate
* fix(steer): recheck indeterminate cancels after ACK
* fix(steer): close retry and abort races
* fix(steer): serialize live retries and abort admission
* fix(steer): distinguish live dispatching from unknown
* fix(steer): keep ACK failures held and reconcile busy cancel
* fix(steer): distinguish held cancel from removal
* fix(store): combine schema v24 migrations
* fix(store): reserve schema v25 for steer delivery state
* fix(steer): keep held cancel state and notify requeue
* fix(steer): release explicitly cancelled unknown reservations
* fix(codex): reject cancelled reservations before native steer
* fix(codex): make reservation restore atomic with state
* fix(codex): terminate abandoned transport writes
* fix(steer): own abandoned app-server lifecycle and consume races
* fix(codex): confirm dispatch and recover abandoned turns
* test(codex): mock abandoned transport callback
* fix(codex): clear visible turn state on transport loss
* fix(steer): claim retries and cover native delivery state
* fix(native): preserve indeterminate state on Android hydration
* fix(steer): make retry claims single-winner
* fix(steer): serialize concurrent retry claims
* fix(socket): tolerate missing steer-state ACK callbacks
* fix(native): serialize retry operations
* docs(web): document unknown steer delivery and retry controls
* fix(steer): handle retry failures and abort-before-connect
* fix(steer): reinitialize after transport loss and finish iOS retry errors
* fix(steer): preserve indeterminate rows across reconnect gaps
* test(web): mock indeterminate queued recovery state
* fix(steer): recover consumed ACK tombstones
* fix(steer): expose consumed cancel tombstones
- decode fs.stat-derived epoch fields leniently (fractional mtimeMs from
real hubs broke machines/files decode and pinned the offline banner)
- offline banner: only a failed sessions fetch counts; machines/baseline
failures are advisory; live SSE emission clears it and seeds the unread
baseline (all-rows-unread gray dot cascade)
- session row: single weighted title (short names no longer truncated at
half width), unread dot moved beside the timestamp
- composer restyle: borderless input pill with inline mic, hand-drawn
vector glyphs replacing emoji icons, 42dp round actions, park-draft
relocated to the session overflow menu
Interaction layer turning the read-only chat into a working remote control:
- Composer: multiline input bar with optimistic sends (appendOptimistic ->
POST -> status settle), queue-by-default delivery with a long-press
Send&Steer intent while a turn is active, abort button during thinking,
tap-to-retry on failed rows, per-session drafts (DataStore, hub-scoped
keys), attachment chip seam for M4.
- session_inactive (409) recovery: one POST /resume then retry; a
superseding session id seeds the new window, migrates the draft and
emits SessionSuperseded for renavigation (web resolveSessionId parity).
- Queued bar: uninvoked sends with Cancel (DELETE; invoked-race ingests the
authoritative row as sent), Edit (cancel + composer prefill, draft-kept
guard) and Steer (POST steer; invoked answers reconcile a missed consume).
reconcileQueuedState now runs on chat open and on session-pipe gap.
- Permission actions: flavor-exact bodies mirroring PermissionFooter.tsx --
claude {} / allowTools / mode:acceptEdits, codex-family decision:
approved / approved_for_session / abort -- plus AskUserQuestion flat
answers and request_user_input nested answers forms; optimistic
Resolving/AlreadyHandled overrides settled by the agentState patch.
- Session config sheet: catalog-driven permission-mode picker, claude
static model/effort catalogs (ported to :core:protocol catalog), codex
models via GET /codex-models (new HapiApi endpoint + wire types) with
per-model reasoning efforts; optimistic detail updates rolled back to
server truth on error.
- Lifecycle: ProcessLifecycleOwner -> SseEngine.setLifecycleForeground +
POST /api/visibility per subscription (VisibilityReporter fed by the new
SyncTargets.onHandshake hook); the global SSE pipe moved from the session
list VM to HubGraph lifetime (GlobalSsePipe) so queued/consumed
bookkeeping and list badges stay fresh while a chat is open.
- Tests: VM-level interaction suite (optimistic send/failure/retry,
inactive-resume both id paths, cancel invoked-race, steer, exact
approve/deny body JSON incl. both answers formats, config optimistic +
rollback, drafts) + GlobalSsePipe tests; full gate green (554 tests,
assembleDebug).
:core:protocol window/ — pure state machine ported function-for-function from
web/src/lib/message-window-store.ts + messages.ts: merge-by-(id|localId) with
optimistic echo reconciliation, position ordering (invokedAt ?? createdAt, seq,
ASCII id tie-break), trim-preserving-queued with the codex agent-run budget,
epoch reset handling, latest-replace with request-baseline identity
preservation, consumed/cancelled/queued-reconcile transitions, tail/history
modes, hydrate/persist shapes. MessageRetention ports the null-decision tree
of normalizeDecryptedMessage (dedup with the B-M2a pipeline port flagged).
:core:data store/ — per-session MessageWindowStore (StateFlow + Mutex,
single-flight tail sync with trailing drain, fetchOlder with epoch-reset
resync, SSE ingest hooks, optimistic sends, queued-state reconciliation,
seedFrom for resume id changes) + WindowSnapshots (atomic JSON files, LRU 10;
JsonSnapshotStore dedup TODO) + MessageWindowStores registry. Minimal
MessagesApi interface (sealed MessagesQuery) extracted over the two message
endpoints, implemented by HapiApi.
Gate: PaginationFixtureTest replays all 11 shared/fixtures/pagination scripts
against the real store — requests, older-load outcomes, reconcile candidates
and the final window projection all exact — 11/11 green; plus targeted
concurrency/snapshot/seed unit tests. :app:assembleDebug green.