Group ordinary Codex commands with default tools across web, iOS and Android while preserving exploration and user-shell boundaries.
Add regression tests, generated protocol fixtures and shared-command coverage in the iOS transcript UI suite.
Render ExitPlanMode and exit_plan_mode Markdown from input.plan in iOS and Android conversations. Preserve approvals, raw source and diagnostics while hiding empty output placeholders and prewarming plan documents.
Add generated protocol fixtures and native regression coverage for long plans, live updates, recycling, themes and typography.
Use one native app-server for terminal, Web and phone clients while retaining the existing CLI and Runner lifecycle.
Synchronize native queues, permissions, question history and steering state; preserve explicit permission precedence and per-turn usage models. Resume inactive clear commands through Runner and reject independent child cold resumes.
Add shared-runtime regression tests, generated protocol fixtures and lifecycle documentation.
Bridge main-session PermissionRequest hooks without suppressing the native
terminal dialog. Reconcile replies against native results and clean up on
timeout, cancellation, mode switches, and session changes.
Keep reply IDs distinct from native tool IDs across web and native clients;
add protocol fixtures and regression tests.
Refs #1796
* fix(web): keep streamed reasoning/text block ids stable across snapshot rows
Streaming snapshots of one stream (pi/codex reasoning and text) arrive as
separate message rows, and the window store retires older rows as newer
snapshots land. The timeline derived the block id from whichever row was
first seen, so the id (and the threadMessageId built from it) churned on
every snapshot, remounting the rendered reasoning panel mid-stream and
replaying its open animation — the panel visibly flashed/re-rendered on
every snapshot tick.
Derive the block id from the stream id when present (unique per stream,
stable across snapshot rows) so the block is updated in place and the
smooth streaming keeps appending to the previous text. Row-derived ids
remain the fallback for content without a stream id.
Also rerun gen:fixtures to refresh the two golden fixtures affected by
the new id shape.
* fix(ios,android): mirror stream-stable block ids in native chat ports
The native HapiKit (Swift) and protocol (Kotlin) chat pipelines are ports
of the web reducerTimeline and are pinned by the same golden fixtures in
shared/fixtures/chat. After the web-side change to derive streamed
reasoning/text block ids from the stream id, the ports still produced
row-derived ids, so the iOS/Android fixture conformance suites went red
on the two refreshed fixtures.
Apply the same streamId-first id derivation (row-derived fallback kept)
to both ports so all three pipelines project identical block ids.
* fix(web,ios,android): reject blank stream ids as block identity
Blank ('' or whitespace-only) stream ids are not streams per the wire
semantics in shared/src/messages.ts (readReasoningStreamId trims before
accepting). The previous nullish fallback let accepted payloads carrying
blank ids through, so every such row shared one empty block id: the
merge maps collided and assistant-ui occurrence suffixes churned with
list position, reintroducing remounts.
Normalize with a trim guard in all three pipelines (web, HapiKit,
protocol) and add a web regression test covering both empty and
whitespace-only ids.
* fix(ios): use normalized stream id for block construction identity
The blank-id guard was applied to lookup and map insertion but block
construction still read the raw optional, so accepted payloads carrying
blank/whitespace ids produced blocks sharing one blank SwiftUI identity
instead of falling back to row-derived ids (web/Android already used the
normalized local). Hoist the nonBlankStreamId result and reuse it for
lookup, block identity, and insertion in both the text and reasoning
branches.
Also add native coverage for stream identity: stream-id derivation for
text/reasoning plus blank ('' and whitespace-only) fallbacks, which the
golden fixtures do not exercise.
* fix(web): pin blank stream-id identity contract in golden fixtures
Update the two stale fixture descriptions (stream-keyed blocks are now
keyed by the stream id, not the first message) and add a generated
conformance fixture covering empty and whitespace-only codex data.id
values for both reasoning and text: blank ids are not stream identities,
so each payload keeps its own row-derived block id instead of collapsing
onto a shared blank identity. Web, iOS, and Android all run this same
golden fixture.
* feat(hub): make title provider max_tokens and timeout env-tunable
Reasoning models used as title providers (e.g. GLM thinking models) need
more than 64 completion tokens and more than the hardcoded 10s timeout to
emit a title, and the only workaround was patching the compiled binary
after every install.
Expose both knobs via HAPI_TITLE_PROVIDER_MAX_TOKENS and
HAPI_TITLE_PROVIDER_TIMEOUT_MS, following the existing
HAPI_TITLE_SUGGESTION_RATE_LIMIT pattern; defaults are unchanged.
* docs(hub): document title provider max_tokens/timeout env knobs
Add the two new HAPI_TITLE_PROVIDER_* variables to the title-provider
configuration table in the installation guide, and extend the provider
test to cover the timeout abort path (the signal fires and rejects the
in-flight request).
---------
Co-authored-by: HongChenGG <HongChenGG@users.noreply.github.com>
* fix(acp): carry the live reasoning marker on the wire payload
ACP agents stream thoughts a token at a time, so the handler coalesces
them into a buffer and re-sends the whole buffer under a stable stream
id every 250ms. The converter dropped the marker that says a payload is
one of those throttled snapshots, leaving the hub unable to tell a
replaceable snapshot from the settled message that closes the stream.
Mirrors how the text variant already forwards streamSnapshot.
* fix(hub): keep one stored message per reasoning stream
OpenCode reasoning arrives as a series of growing snapshots sharing one
stream id, and every snapshot was persisted as its own message. A 26h
session reached 48,844 rows and 63MB, and because the web budgets a
fixed number of messages, its 400-message window covered barely three
minutes of conversation — scrolling up walked through duplicate
snapshots instead of history.
Retire a stream's earlier live snapshots once their replacement is
stored. Sweeping only after the insert matters: the two statements are
separate transactions, so clearing first would leave a window where a
crash takes the whole stream. Only rows marked live are eligible and the
replacement is spared, so a stream always keeps at least one row and the
settled message that closes it is never removed.
Live rendering is unchanged: the web still receives every snapshot and
already folds them by stream id.
* fix(web): spend the message window on conversation, not repeated snapshots
The window budgets raw messages, but a reasoning stream renders as a
single folded block no matter how many snapshots it arrived in. On
sessions recorded before the hub started retiring them, those snapshots
fill the window on their own: in one 26h session the newest 400 messages
covered 202 seconds, so scrolling up paged through duplicates instead of
history.
Collapse each stream to its newest snapshot before trimming. Rendering
is unchanged — the timeline already folds them by stream id — and rows
without a stream id are never touched.
* fix(ios,android): port reasoning-snapshot compaction to the native windows
The window logic in HapiProtocol and :core:protocol is a one-to-one port
of the web store, so collapsing superseded reasoning snapshots only on
the web left the native windows budgeting raw snapshot rows. The hub
stores one row per stream now, but a client that already holds the older
snapshots still spends its window on them.
Add the same stream-id reader and compaction to both ports, in the shape
each already uses for agent-run rows, and pin the behaviour with a
pagination fixture. Both fixture suites enumerate shared/fixtures/pagination
from disk, so the ports cannot drift from the web again without CI saying
so.
Two new golden-fixture suites, expectations machine-generated from the web
implementation (same source-of-truth principle as the chat suite):
- shared/fixtures/sse/ (12 cases): session-updated versioned-patch
application via applySessionDetailPatch — strict version gates for
metadata/agentState/todos/teamState, out-of-order arrival, teamState:null
clear, max-monotonic updatedAt, flat last-write-wins fields, sub-minute
activeAt keep-alive drop, scratchlistUpdatedAt trigger, and the pinned
actual behavior that activeTurnStartedAt is NOT applied by the patch path.
Inputs are stored schema-normalized and validated against SessionSchema /
strict SessionPatchSchema at generation time; per-patch applied/unchanged
verdicts are part of the contract.
- shared/fixtures/pagination/ (11 cases): op scripts driving the real
message-window store with a scripted ApiClient — latest page + SSE ingest,
before-cursor older pages, epoch-mismatch reset (window discard + recorded
internal latest request), reset:true replace preserving optimistic rows,
localId echo reconciliation, messages-consumed invokedAt stamping (no
cursor advance), message-cancelled removal, cancel-too-late invoked-row
ingest, hidden-row cursor advance, 400-row trim preserving queued rows,
and queued-state gap recovery. Documents pin the exact requests the store
issued, older-load outcomes, reconcile candidates, and a minimal window
projection (ids/order, queued/optimistic flags, hasMore, epoch, viewMode,
compound cursors).
Generator: web/scripts/fixtures/{sse,pagination}/ extend the K5 framework
(canonical serialization reused; suites pruned of stale files). Web
self-conformance tests replay every fixture against the real implementation
in bun run test:web. README documents schemas, projections, replay
contracts, and native consumption for both suites.
No time injection needed: the store's only Date.now() reads gate
notification throttling, which never reaches persisted or projected state.