mirror of
https://github.com/wu736139669/hapi.git
synced 2026-10-09 19:29:41 +00:00
1de9613df60912efa9b5f623f322ecbddc298157
731
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1de9613df6 | docs: align documentation with current implementation | ||
|
|
26fddff038 |
fix(chat): restore grouping for shared Codex commands
Group ordinary Codex commands with default tools across web, iOS and Android while preserving exploration and user-shell boundaries. Add regression tests, generated protocol fixtures and shared-command coverage in the iOS transcript UI suite. |
||
|
|
d29713aeda |
fix(native): render plan proposals inline
Render ExitPlanMode and exit_plan_mode Markdown from input.plan in iOS and Android conversations. Preserve approvals, raw source and diagnostics while hiding empty output placeholders and prewarming plan documents. Add generated protocol fixtures and native regression coverage for long plans, live updates, recycling, themes and typography. |
||
|
|
d18c01b4b6 |
revert(codex): remove Luna Reserve fallback (#1780)
Revert
|
||
|
|
0c4abcb3d1 |
feat(codex): share sessions across terminal and web
Use one native app-server for terminal, Web and phone clients while retaining the existing CLI and Runner lifecycle. Synchronize native queues, permissions, question history and steering state; preserve explicit permission precedence and per-turn usage models. Resume inactive clear commands through Runner and reject independent child cold resumes. Add shared-runtime regression tests, generated protocol fixtures and lifecycle documentation. |
||
|
|
a729456682 |
fix(claude): answer local permission prompts from web
Bridge main-session PermissionRequest hooks without suppressing the native terminal dialog. Reconcile replies against native results and clean up on timeout, cancellation, mode switches, and session changes. Keep reply IDs distinct from native tool IDs across web and native clients; add protocol fixtures and regression tests. Refs #1796 |
||
|
|
68422ed6d3 |
fix(web): use conversation content for session reference eligibility (#1800)
* fix(web): base session reference eligibility on conversation content * fix(hub): refresh replacement content after clear abort * fix(hub): refresh content on idempotent clear aborts |
||
|
|
de4af23482 | fix(web): simplify session summary setting copy (#1803) | ||
|
|
2eb11a5d2c |
fix(web): clarify round usage metadata labels (#1805)
* fix(web): clarify round usage metadata labels * test(web): cover duration boundary rounding |
||
|
|
d124c106a4 |
feat(web): file context menu with copy path / absolute path / add to composer (#1808)
Add a right-click (desktop) / long-press (mobile) menu to file entries on the Files page: git change rows, file search results, and the directory tree. Actions: - Copy path (workspace-relative) - Copy absolute path (resolved against session.metadata.path) - Add to composer (append a backticked reference to the session composer draft and navigate back to the chat) Closes #1807 |
||
|
|
fd2822bab5 |
fix(agy): allow switching existing sessions to newly available models (#1814)
* fix(agy): say what the model picker is actually waiting on The spinner in the New Session AGY picker read "Checking Antigravity authentication…", but nothing at that point checks authentication — the machine is running `agy models`, and the sign-in prompt is a separate branch below it, shown only when agy reports the failure. Name the wait after the work: "Fetching available models…", the same words agy prints while it fetches. * refactor(agy): describe a probe by its outcome, not by its response The probe function returned a finished `AgyModelsResponse`, so "agy could not be reached" and "agy listed no models" both arrived as a successful response carrying the hardcoded mirror, and the caller could no longer tell which had happened. Every policy decision about that answer has to live inside the probe as a result. Hand back what the probe observed — a live catalog, an auth failure, or nothing usable — and let the caller turn it into a response. Same behaviour: the mirror still stands in for both failure modes, and the 60s cache still holds whatever came out. * fix(agy): serve the model catalog stale-while-revalidate The `agy models` probe is a whole agy invocation — around 3s on a good day, 15s when it times out — and the 60s window meant the New Session picker paid that again a minute after the last look. Keep the last listing agy actually returned and answer from it: fresh for ten minutes, then still answered while a probe refreshes behind it, until the entry is a day old and stops standing in for the machine at all. A probe that times out or loses auth leaves that entry alone, so a blip no longer empties a working picker, and the hardcoded mirror is no longer recorded as if the machine had reported it. Three things fall out of that and are handled here. A machine whose sign-in has actually gone bad would otherwise look healthy for a day, so an auth failure rides along with the catalog it can still serve — and, because nothing else would re-probe a catalog that is still fresh, a warning riding on the answer is itself a reason to look again. A failed probe is not repeated on the very next request either, or a machine where agy hangs would spawn it once per poll. An explicit refresh always costs a probe, and never rides one that was already running when it was asked for. * fix(agy): let Retry force a fresh model catalog probe With the catalog held for ten minutes, Retry would otherwise hand back the answer it was pressed to replace, so the intent travels to the machine: `?refresh=true` on the machine route, an optional RPC param, and a one-shot flag on the query so ordinary mount and focus refetches stay cheap. Every hop is optional, so a hub and a runner on different versions still talk — the older side ignores it and answers from its cache. Retry also has to be reachable, and honest, in the state that needs it. The machine now answers with both a usable catalog and the sign-in failure behind it, so the picker keeps the list and says why it may be out of date, with the button right there rather than only once there is nothing left to show. Pressing it runs agy, which can take tens of seconds, so the button says so while it does. The client contract covers both: `agy-models` is the one catalog route that can carry an `error` on a successful response, and the one that takes a refresh parameter. * fix(agy): use the machine catalog in the in-session model picker New Session already asks the machine what `agy models` lists, but a session that is already open offered the built-in list in `shared/src/models.ts`. That list is a hand-maintained mirror, so a model agy started offering after the last release could be picked for a new session and not for the one already running. Point the composer at the same machine catalog. The mirror stays as the fallback for the moment before the machine answers, and a model the session is already on is kept selectable — and readable, when it is one of the known presets — even after the catalog moves on without it. * fix(agy): announce a model catalog re-check that changed the answer Serving the last known catalog answers the picker instantly, but a picker that was already open kept showing that answer until the user closed and reopened it — the machine had no way to say it had found something newer. Say it on the stream that already carries machine changes. The machine daemon — the only process that answers `<machineId>:listAgyModels` — emits it, and the hub forwards it as `machine-agy-models-updated` with nothing but the machineId. Namespace resolution, per-machine delivery and reconnect replay all come from the existing path. What counts as a change is what the route would answer, not what sits in the cache. That distinction carries the cases: a sign-in that lapsed or came back changes no models yet changes what the user is told; a machine whose agy was signed out has been answering from the hardcoded mirror, and its first real listing is the largest change there is, for every client except the one awaiting it. * fix(agy): re-read the announced machine's model catalog On `machine-agy-models-updated`, cancel and refetch that one machine's catalog query — the app's global connection is always subscribed, so an open picker redraws wherever it is. Cancelling first is what makes it correct rather than merely likely. query-core cancels an in-flight fetch only when the query already holds data, so a picker opening for the first time would otherwise join the request already on its way and settle on the listing the announcement replaced. The refetch is answered from the machine's cache, so it starts no probe and cannot bounce another announcement back. A reconnect the hub could not replay takes the resync path, which clears the agy catalogs the same way — that path has no announcement to fall back on, so it is the one that can least afford to join a stale request. * fix(agy): keep a model the user picked when the catalog moves under them The catalog can now change while the New Session form is open, and the form dropped any selection the machine no longer advertised — including one the user had just made. Keep that one, and list it as no longer listed so the form does not imply agy is still offering it. A model restored from a draft or a saved preference is still dropped: it may never have been runnable here. * fix(agy): announce uncached authentication changes |
||
|
|
2b402d24d9 |
fix(web): show Codex round usage metadata (#1685)
* fix(web): show Codex round usage metadata * fix(web): skip imported Codex round duration |
||
|
|
ceb9314350 |
feat(web): add scroll-to-bottom button (#1694)
* feat(web): add scroll-to-bottom button * fix(web): type counted scroll button props * fix(web): retarget smooth tail scroll * chore: retrigger pull request checks |
||
|
|
7031d60eb4 |
fix(opencode): expose model-specific reasoning effort options (#1716)
* fix(cli): discover opencode thought_level via set_config_option on model switch * fix(cli): apply and refresh opencode model switches so thought_level stays discoverable * fix(web): track opencode effort options across model switches * feat(cli,hub,web): dynamic opencode effort options in new-session form * test(cli): avoid platform-specific process event narrowing * fix(opencode): address variant discovery review findings * fix(opencode): synchronize effort options with model targets * fix(opencode): roll back rejected model targets * fix(cli): guard opencode variant probe workspace paths * fix(web): clear stale opencode effort on model switch * fix(cli): clear stale opencode effort metadata * fix(web): reset stale effort options on model switch * fix(web): reset opencode effort poll budget * test(web): enforce opencode effort poll budget |
||
|
|
f27e58741b |
feat(web,ios,android): let Claude pick a permission mode when creating a session (#1751)
* refactor(web): route create-form permission control through one native-select predicate Extract usesNativePermissionSelect(flavor) (grok || codex-family, matching the existing iOS/Android predicate of the same name) and route PermissionField's select-vs-toggle gate through it instead of an inline condition. Rename the codex-family-only state codexFamilyPermissionMode to nativePermissionMode since it now backs a shared predicate, not just the codex family. Behavior is unchanged: any stale sessionStorage draft written under the old codexFamilyPermissionMode key has no value under the new key and falls back to 'default', which only matters within a single browser tab's lifetime. * feat(web): let Claude pick a permission mode when creating a session Claude was the only create-form flavor still on the global HAPI YOLO toggle while grok and the codex family got the native permission select, so there was no way to start a session in Plan Mode without creating it first and switching the mode from the composer. usesNativePermissionSelect now gates the control for claude as well, and the spawn body carries permissionMode (including 'default') instead of yolo, which is the shape the other native-select flavors already send. The stored hapi:newSession:yolo preference is bridged into the select rather than dropped, but only for the flavors that have actually moved onto it (LEGACY_YOLO_BRIDGE_AGENTS = codex, claude). copilot, gemini, kimi and opencode moved earlier and settled on 'default'; re-enabling Yolo for them now would widen permissions rather than migrate a preference. This narrows the sessionStorage draft bridge too, which until now fired for the whole codex family with no allow-list, so their draft restores yield 'default' instead of 'yolo' — same-tab-lifetime state only. Claude and the codex family share one nativePermissionMode state and their mode sets do not overlap, so the existing agent-change reset plus the flavor filters in the draft loader and the stored launch settings are what keep a codex mode out of a Claude spawn. Adds the regression test that pins it: pick a mode under codex, switch to Claude, create, assert the payload carries 'default'. * feat(ios): let Claude pick a permission mode when creating a session Extend usesNativePermissionSelect to include claude alongside grok and the codex family, matching the web change. buildSpawnRequest now derives both yolo and permissionMode from that single predicate instead of two separate local flags, so claude sends permissionMode (including 'default') and no longer sends yolo. Unlike web, iOS carries no persistent YOLO preference across sessions to migrate — the toggle only lives in the in-memory form or a draft deleted on success — so there is no bridging logic to add here. * feat(android): let Claude pick a permission mode when creating a session Extend usesNativePermissionSelect to include claude alongside grok and the codex family, matching the web and iOS changes. buildSpawnRequest derives both yolo and permissionMode from that single predicate, so claude sends permissionMode (including 'default') and no longer sends yolo. Unlike web, Android has no persistent YOLO preference to migrate: the toggle only lives in the form draft, which is deleted once a session is created. The agent-switch test asserted claude renders the YOLO toggle; it now checks the native select for claude and keeps the toggle assertion on cursor, which still carries it. |
||
|
|
33f51fd0ab | feat(web): add configurable PWA taskbar badge (#1759) | ||
|
|
4dac00ccd9 |
feat(web): add bulk mark-all-read session action (#1773)
* feat(web): add bulk mark-all-read session action * fix(web): mark duplicate sessions as read * fix(web): count hidden duplicate unread sessions * fix(web): exclude empty session stubs from bulk read |
||
|
|
8357da0a9d |
feat(codex): support backend-authorized Luna Reserve fallback (#1780)
* feat(codex): reconcile Luna Reserve fallback and conditional usage * fix(codex): preserve queued settings and reconcile Reserve sessions |
||
|
|
66ea35963e | feat(web): add word wrap to file source previews (#1754) | ||
|
|
b14bb0fae5 |
fix(web): let the installed Android status bar follow the system (#1742)
An installed WebAPK reads manifest theme_color once at install time and uses it as a fixed toolbar color; runtime document theme-color updates never repaint it, only the icon tint follows. With theme_color set to white the status bar stayed white while dark app themes turned the icons white too, hiding the clock, signal and battery. Chrome falls back to white when no theme color is present and to black when the system is dark and no dark color is declared, so dropping theme_color gives a status bar that follows the Android system theme with icons that stay readable in both. The dark manifest override that could have supplied a per-scheme color is not an option: Chrome removed dark color parsing from the manifest parser in 128 and the mojom fields are marked obsolete. The document theme-color meta keeps driving every non-installed surface, so browser tabs are unchanged. |
||
|
|
e26312b51f | fix(web): remember file browser tab preference (#1761) | ||
|
|
c191e073b2 |
fix(web): compact recent path labels (#1757)
* fix(web): compact recent path labels * fix(web): expose full recent path to assistive tech |
||
|
|
52a76a0bc4 |
fix(web): clarify session summary setting copy (#1760)
* fix(web): clarify session summary settings * fix(web): clarify summary emission copy |
||
|
|
17ee052d9a |
fix(web): preserve the visible chat window during rewind (#1766)
* fix(web): preserve chat window during rewind * fix(web): scope rewind invalidation preservation * fix(web): clear unknown rewind boundaries * fix(web): deduplicate rewind invalidations * fix(web): retain rewind dedupe history |
||
|
|
9e954d7c1b | fix(web): center inactive session notice text (#1770) | ||
|
|
bb055ae74e | fix(web): preserve sidebar viewport across pin updates (#1776) | ||
|
|
3873e58496 |
fix(web): keep streamed reasoning/text block ids stable across snapshot rows (#1741)
* fix(web): keep streamed reasoning/text block ids stable across snapshot rows
Streaming snapshots of one stream (pi/codex reasoning and text) arrive as
separate message rows, and the window store retires older rows as newer
snapshots land. The timeline derived the block id from whichever row was
first seen, so the id (and the threadMessageId built from it) churned on
every snapshot, remounting the rendered reasoning panel mid-stream and
replaying its open animation — the panel visibly flashed/re-rendered on
every snapshot tick.
Derive the block id from the stream id when present (unique per stream,
stable across snapshot rows) so the block is updated in place and the
smooth streaming keeps appending to the previous text. Row-derived ids
remain the fallback for content without a stream id.
Also rerun gen:fixtures to refresh the two golden fixtures affected by
the new id shape.
* fix(ios,android): mirror stream-stable block ids in native chat ports
The native HapiKit (Swift) and protocol (Kotlin) chat pipelines are ports
of the web reducerTimeline and are pinned by the same golden fixtures in
shared/fixtures/chat. After the web-side change to derive streamed
reasoning/text block ids from the stream id, the ports still produced
row-derived ids, so the iOS/Android fixture conformance suites went red
on the two refreshed fixtures.
Apply the same streamId-first id derivation (row-derived fallback kept)
to both ports so all three pipelines project identical block ids.
* fix(web,ios,android): reject blank stream ids as block identity
Blank ('' or whitespace-only) stream ids are not streams per the wire
semantics in shared/src/messages.ts (readReasoningStreamId trims before
accepting). The previous nullish fallback let accepted payloads carrying
blank ids through, so every such row shared one empty block id: the
merge maps collided and assistant-ui occurrence suffixes churned with
list position, reintroducing remounts.
Normalize with a trim guard in all three pipelines (web, HapiKit,
protocol) and add a web regression test covering both empty and
whitespace-only ids.
* fix(ios): use normalized stream id for block construction identity
The blank-id guard was applied to lookup and map insertion but block
construction still read the raw optional, so accepted payloads carrying
blank/whitespace ids produced blocks sharing one blank SwiftUI identity
instead of falling back to row-derived ids (web/Android already used the
normalized local). Hoist the nonBlankStreamId result and reuse it for
lookup, block identity, and insertion in both the text and reasoning
branches.
Also add native coverage for stream identity: stream-id derivation for
text/reasoning plus blank ('' and whitespace-only) fallbacks, which the
golden fixtures do not exercise.
* fix(web): pin blank stream-id identity contract in golden fixtures
Update the two stale fixture descriptions (stream-keyed blocks are now
keyed by the stream id, not the first message) and add a generated
conformance fixture covering empty and whitespace-only codex data.id
values for both reasoning and text: blank ids are not stream identities,
so each payload keeps its own row-derived block id instead of collapsing
onto a shared blank identity. Web, iOS, and Android all run this same
golden fixture.
* feat(hub): make title provider max_tokens and timeout env-tunable
Reasoning models used as title providers (e.g. GLM thinking models) need
more than 64 completion tokens and more than the hardcoded 10s timeout to
emit a title, and the only workaround was patching the compiled binary
after every install.
Expose both knobs via HAPI_TITLE_PROVIDER_MAX_TOKENS and
HAPI_TITLE_PROVIDER_TIMEOUT_MS, following the existing
HAPI_TITLE_SUGGESTION_RATE_LIMIT pattern; defaults are unchanged.
* docs(hub): document title provider max_tokens/timeout env knobs
Add the two new HAPI_TITLE_PROVIDER_* variables to the title-provider
configuration table in the installation guide, and extend the provider
test to cover the timeout abort path (the signal fires and rejects the
in-flight request).
---------
Co-authored-by: HongChenGG <HongChenGG@users.noreply.github.com>
|
||
|
|
aa0c1dc808 |
fix(codex): fail closed on ambiguous Web Rewind boundaries (#1707)
* fix(codex): fail closed on ambiguous Web Rewind boundaries * fix(codex): reject malformed native rewind history * fix(codex): reject duplicate rewind identifiers * feat(codex): offer Fork fallback for ambiguous rewind * fix(codex): gate rewind Fork fallback on exact native boundary * fix(codex): reject unresolved rewind boundaries safely * chore: trigger PR checks * fix(codex): gate safe rewind fallback on fork support * fix(codex): require complete user ids for rewind fallback |
||
|
|
e5a8212f4a | feat(session): validate agents and browse workspace directories | ||
|
|
a9d09cb3ff |
feat(web): support composer attachment drag reordering (#1682)
* feat(web): support composer attachment drag reordering * fix(web): handle forward attachment reordering * fix(web): enlarge file attachment touch targets * fix(web): preserve staged attachment order |
||
|
|
be1ef2a2e4 |
feat(dsh): integrate DeepSeek Harness through ACP (#1632)
* feat(dsh): add DeepSeek Harness ACP flavor * fix(dsh): update mobile flavor catalogs * fix(dsh): keep mobile spawn policy managed * fix(dsh): keep managed policy and prompt retry * fix(dsh): suppress unsupported runner policy flags * fix(dsh): align native managed-policy UX |
||
|
|
661e9b4eb7 | feat(web): show Claude round usage metadata (#1655) | ||
|
|
dd2e978fe3 |
feat(web): add mark unread session action (#1649)
* feat(web): add mark unread session action * fix(web): keep explicit unread state on selected sessions * fix(web): address mark unread review feedback * fix(web): preserve newer manual unread state |
||
|
|
e3f13c4eb6 | feat(web): anchor composer settings sheet to the clicked value button's section (#1627) | ||
|
|
fdf4fdaa38 |
feat(web,hub): make large conversation export warning-only (#1648)
* feat(export): warn before large session downloads * fix(export): surface server size errors |
||
|
|
338e405442 | fix(brand): sharpen iOS and PWA icons | ||
|
|
0aebf39c78 |
fix(opencode): keep one stored message per reasoning stream (#1643)
* fix(acp): carry the live reasoning marker on the wire payload ACP agents stream thoughts a token at a time, so the handler coalesces them into a buffer and re-sends the whole buffer under a stable stream id every 250ms. The converter dropped the marker that says a payload is one of those throttled snapshots, leaving the hub unable to tell a replaceable snapshot from the settled message that closes the stream. Mirrors how the text variant already forwards streamSnapshot. * fix(hub): keep one stored message per reasoning stream OpenCode reasoning arrives as a series of growing snapshots sharing one stream id, and every snapshot was persisted as its own message. A 26h session reached 48,844 rows and 63MB, and because the web budgets a fixed number of messages, its 400-message window covered barely three minutes of conversation — scrolling up walked through duplicate snapshots instead of history. Retire a stream's earlier live snapshots once their replacement is stored. Sweeping only after the insert matters: the two statements are separate transactions, so clearing first would leave a window where a crash takes the whole stream. Only rows marked live are eligible and the replacement is spared, so a stream always keeps at least one row and the settled message that closes it is never removed. Live rendering is unchanged: the web still receives every snapshot and already folds them by stream id. * fix(web): spend the message window on conversation, not repeated snapshots The window budgets raw messages, but a reasoning stream renders as a single folded block no matter how many snapshots it arrived in. On sessions recorded before the hub started retiring them, those snapshots fill the window on their own: in one 26h session the newest 400 messages covered 202 seconds, so scrolling up paged through duplicates instead of history. Collapse each stream to its newest snapshot before trimming. Rendering is unchanged — the timeline already folds them by stream id — and rows without a stream id are never touched. * fix(ios,android): port reasoning-snapshot compaction to the native windows The window logic in HapiProtocol and :core:protocol is a one-to-one port of the web store, so collapsing superseded reasoning snapshots only on the web left the native windows budgeting raw snapshot rows. The hub stores one row per stream now, but a client that already holds the older snapshots still spends its window on them. Add the same stream-id reader and compaction to both ports, in the shape each already uses for agent-run rows, and pin the behaviour with a pagination fixture. Both fixture suites enumerate shared/fixtures/pagination from disk, so the ports cannot drift from the web again without CI saying so. |
||
|
|
9bdccf0cff | fix(web): unify message action buttons (#1645) | ||
|
|
0f7a3da68b |
feat(cursor): mid-turn Steer via concurrent ACP session/prompt (#888) (#1609)
* feat(shared): steer capability gates and live steered signal schemas - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which agents can deliver queued messages into the active turn (pi, codex, cursor ACP; legacy stream-json cursor excluded) - AgentState.steeringActive, DecryptedMessage.steered and messages-consumed live signal (never persisted by the hub) * feat(cli): queue reservations and steered messages-consumed option - MessageQueue2 gains takeByLocalId/restoreReservation/ beginReservationDispatch/commitReservation so an async steer can reserve a queued row without racing the main loop's turn/start drain - emitMessagesConsumed accepts steered: true to mark mid-turn delivery * feat(codex): mid-turn steer via app-server turn/steer (#888) - CodexAppServerClient.steerTurn + TurnSteerParams/Response types - CodexRemoteLauncher registers the steer-queued-message RPC handler: reserves the queued row, validates it against the active turn (no control commands, matching mode hash), injects via turn/steer with an epoch guard that invalidates in-flight steers on abort/cleanup - steeringActive agent state tracks the active-turn window - hub syncEngine gate opens to codex; messages-consumed relays steered * feat(web): Steered badge and steer gating for codex sessions - HappyUserMessage shows a ↳ Steered badge fed by the live messages-consumed steered signal, preserved across server echoes and refetches (mergeMessages carries the optimistic marker) - SessionChat gates canSteer via isSteeringSupportedForSession instead of the pi-only check - clearStaleQueuedStatus normalizes a queued status on an invoked message - fix(web): drop duplicate showSessionSummaryInChat in markdown test (upstream typecheck breakage) * feat(acp): split request dispatch from completion and add soft steer - AcpStdioTransport.sendRequestWithDispatch separates stdin-accepted dispatch from the JSON-RPC response, keeping sendRequest behavior unchanged - AcpSdkBackend tracks concurrent session/prompt requests with an activePromptRequests counter (main prompt + soft steers); response completion stays pending until every concurrent prompt settles - beginSoftSteerPrompt kicks off a concurrent session/prompt (Cursor GUI Send semantics — no cancel, no handler swap) returning {dispatched, completed}; softSteerPrompt awaits the full response for direct callers * feat(cursor): mid-turn soft steer via concurrent session/prompt (#888) - CursorAcpRemoteLauncher registers the steer-queued-message RPC handler: reserves the queued row, rejects control commands and mode mismatches, then soft-injects via beginSoftSteerPrompt without canceling the in-flight turn - Acks the hub once stdin accepts the inject (not on turn completion) to stay inside the 30s RPC window; the launcher stays busy until the concurrent prompt settles so handlers are not swapped mid-inject - steeringActive agent state mirrors the active-turn window; abort and cleanup reset it and invalidate pending steers - Legacy stream-json Cursor sessions register a steer handler that reports unsupported * fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise codex and pi only; cursor joins when its soft-steer handler lands (#1609) - turn/steer now splits dispatch (stdin accepted) from completion (turn finished): the hub RPC acks once dispatch succeeds — never on the concurrent turn's completion, which can exceed the 30s RPC window - queue row commits only after the turn settles; a rejected/aborted steer restores the row so the message still delivers via turn/start, and a dispatched steer is never restored (no duplicate delivery) - steer carries clientUserMessageId (echoed as userMessage.clientId) so ambiguous transport failures can reconcile the thread later - client tests cover dispatch/complete split and stdin-write failure * feat(shared): advertise cursor in the steer gate now that its handler lands Cursor ACP sessions pass the web and hub steer gates; legacy stream-json cursor sessions stay excluded. * fix(codex): reconcile dispatched steers before restoring; align error copy - A dispatched turn/steer whose completion fails (disconnect / protocol error) is now reconciled via thread/read by clientUserMessageId before the queued row is restored — the instruction is only re-delivered by turn/start when the thread never received it - Reconcile targets the pinned steer thread, not whichever turn is current when completion fails - syncEngine unsupported-flavor error now matches the capability gate (Pi and Codex only until the cursor handler lands) - launcher tests cover steer success (ack on dispatch), reconcile-accepted and reconcile-rejected outcomes * fix(codex): consume the row at dispatch; drop background reconcile - The hub RPC acks and the queue row is consumed as soon as stdin accepts turn/steer; completion is background-only logging. A dispatched steer is never restored, so the same localId cannot be re-delivered via turn/start after the caller was told the steer succeeded - Dispatch failure (stdin write error) still restores the row and reports failure - steer.completed rejection is always handled (no unhandled rejection on the dispatch-failure path) - tests updated: completion failure after dispatch keeps the row consumed; dispatch failure restores it * fix(cursor): consume the row at dispatch; keep waiters for prompt gating - The hub RPC acks and the queue row is consumed as soon as stdin accepts the concurrent session/prompt; completion is background-only. A dispatched steer is never restored (no duplicate via the next prompt) - softSteerWaiters are registered before awaiting dispatch so the main loop's finally cannot start the next prompt mid-inject; they still gate prompt handover on completion - tests updated: post-dispatch ACP rejection keeps the row consumed * fix(cursor,hub): never hang teardown on unresolved soft steer; align diagnostics - Prompt-finally waits for soft-steer completion only when not exiting; the outer finally no longer waits at all — cleanup() disconnects the ACP transport, which rejects pending requests and settles the waiters - syncEngine gate diagnostics and JSDoc name all supported flavors (Pi, Codex, Cursor ACP) - regression test: Switch with an unresolved soft-steer completion still reaches teardown * fix(codex): distinguish definite rejection from indeterminate completion - Transport-level failures (timeout, abort, disconnect, spawn, protocol) carry an indeterminate marker; explicit JSON-RPC error responses do not - After a dispatched steer, turn completion resolves → commit + consumed; a definite app-server rejection restores the row (instruction was never accepted, so turn/start cannot duplicate it); an indeterminate outcome leaves the row reserved so it can never be delivered twice - Completion handling registers before awaiting dispatch so the dispatch-failure path cannot leak an unhandled rejection - client/launcher tests cover explicit rejection (restore), indeterminate outcome (row stays reserved) and dispatch failure * fix(codex): reconcile indeterminate steers instead of a permanent reservation - After an indeterminate completion (disconnect/protocol), reconcile the thread by clientUserMessageId immediately: accepted → commit + consumed, provably rejected → restore, still unreadable → keep the reservation and retry from the main-loop top on later passes (post-reconnect) - A row never sits in dispatching forever: the hub cannot stamp it invoked while the instruction may never have been accepted - tests: indeterminate keeps reserved while thread unreadable; accepted reconciliation consumes; rejected path restores * fix(cursor): abort drops soft-steer waiters so the next prompt never blocks - Ordinary Abort (shouldExit false) now clears softSteerWaiters: the prompt finally cannot wait forever on a soft steer whose completion is unbounded; the ACP cancel rejects in-flight requests, and cleanup() settles leftovers on session end - regression test: unresolved soft-steer completion after Abort no longer blocks the next prompt * fix(codex): accept all thread item shapes; retry reconcile; ack through abort - Reconcile matcher accepts userMessage/user_message with clientId/ client_id, matching the shapes the thread parser supports — an accepted steer can no longer be misclassified as rejected - A pending reconciliation schedules a wakeLoop retry, so a temporary app-server outage cannot strand the reservation behind waitForTurnOrRecovery - The success-path ACK no longer checks the steer epoch: the hub already reported steered on dispatch, so commit + messages-consumed must reach it even when an abort resets the queue in between * fix(cursor,acp): abort force-settles soft-steer bookkeeping - AcpSdkBackend.abortSoftSteers() drops the concurrent-prompt counter and notifies response-complete so the next turn's waitForResponseComplete() cannot block on a soft steer that will never settle after abort - handleAbort calls it before clearing the waiters; the main prompt's own finishPromptRequest stays guarded by Math.max(0, ...) - unit tests cover counter release and no-op when idle * fix(codex): reinit reconnected app-server; keep reconcile retries alive - thread/read after a disconnect auto-connects a fresh app-server, which must be initialized before any request — reconcile now ensures connect + initialize (isConnected getter added to the client) - every still-unknown loop-top reconciliation schedules the next retry, so recovery without external traffic is eventually observed - launcher mock gains isConnected * test(acp): match finishPromptRequest epoch signature in whitebox test * fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK - Reconciliation runs on a self-rescheduling 1s timer independent of the main loop (wakes it too), so idle loops and waitForTurnOrRecovery still observe app-server recovery; abort clears nothing implicitly — the ACK path commits and consumes even when the reservation was cancelled - Absence of a durable client id is ambiguous: unmatched reads stay 'unknown' and keep retrying instead of restoring the row - CodexAppServerClient tracks initialized state (reset on disconnect/exit) so ensureAppServerInitialized re-initializes a fresh process before thread/read; initialize failures leave the flag false for the next retry - tests: accepted reconciliation via scheduled timer, indeterminate keeps reserved, explicit rejection restores * fix(codex): bind reconciliation to the launcher lifecycle - runSteerReconciliation clears any armed retry timer on entry and never installs a second one, so loop-top and timer-driven passes cannot multiply - shuttingDown is set when the main loop ends: timers are cleared and the pending map is dropped, so an unresolved steer can never respawn an app-server after cleanup (remote-to-local switch included) * fix(cursor): abort releases an in-progress soft-steer wait - The prompt-finally wait races Promise.allSettled against the abort signal: an Abort that clears the waiters now also releases a wait that already started, so the launcher always reaches the next queued prompt * fix(codex): report steered only after app-server acceptance - The handler now awaits steer.completed (the inject-acceptance response): an explicit JSON-RPC rejection surfaces as failed and restores the row for the normal turn/start path instead of a false steered - Transport failure after dispatch reports 'Steer outcome is being reconciled' and keeps the row reserved while the timer-driven thread reconciliation runs - dispatch-failure path also swallows the paired completion rejection * fix(cursor,acp): commit on ACP acceptance; distinguish transport failures - AcpStdioTransport marks transport-level failures (timeout, closed, stdin write) as indeterminate; explicit JSON-RPC error responses are not - The steer handler commits + consumes on completion (ACP acceptance) and restores the row on an explicit rejection; an indeterminate transport failure keeps the row reserved so a delivered instruction is never re-sent, and the ACK reaches the hub even when abort reset the queue - launcher/transport tests updated for the three outcomes * fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait - MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching steer reservation: the hub neither deletes the row nor stamps invoked_at (new CancelMessageResponse 'busy' status; web restores the optimistic row); pushIsolateAndClear and reset/close share cancelReservations so /clear-style commands cannot have a rejected steer resurrect a discarded prompt - turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a lost response is indeterminate and funnels into thread reconciliation instead of stranding the reservation - tests updated for the tri-state cancel contract * test(cursor): match tri-state cancel contract for dispatching steers * fix(cursor): drop duplicate promptInFlight declaration after upstream merge * fix(codex,web): busy-aware edit flow; bound reconciliation reads - QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it never prefills the composer when the row is inside an async steer, so a second client cannot send a duplicate - reconcileSteerByClientId bounds thread/read with a 5s timeout so a connected-but-silent app-server cannot hold the reservation in-flight indefinitely * fix(steer): inFlight-dominated cancel acks; bounded reconciliation - hub cancel-queued-message acks check inFlight before removed: a stale duplicate socket reporting removed can no longer delete the durable row while another socket is dispatching the steer - reconciliation entries expire after 60s and mark delivered: after the rejection window, a dispatched steer that the app-server never proved (client ids dropped on restart) is committed instead of polling thread/read forever - pre-dispatch failures (abort before write included) never enter reconciliation — they restore the row and report failure * fix(cursor,acp,web): indeterminate close marks, steer gating precision - AcpStdioTransport.rejectAllPending marks close/protocol failures indeterminate, so an accepted-but-close-interrupted soft steer restores nothing (no duplicate delivery) - the abort race in the soft-steer wait removes its listener in finally (no accumulation across repeated waits) - SessionChat gates the Steer button on agentState.steeringActive for codex/cursor instead of the queued-grace thinking flag, so Steer is not exposed before the launcher can accept it - codex pre-dispatch abort never enters reconciliation (merged from #1606) * fix(steer): persist indeterminate outcomes without replay * fix(cursor): hold ambiguous steers for explicit resolution * fix(steer): make ambiguous delivery restart-safe * fix(cursor): make ambiguous delivery restart-safe * fix(steer): recover crash-held rows and preserve retry dedup * fix(steer): ack retries and bound stdin dispatch * fix(cursor): reject steers when prompt generation changes * fix(steer): reconcile indeterminate dispatches and serialize retries * fix(cursor): preserve soft-steer reservations across abort * fix(codex): classify stdin callback failures as indeterminate * fix(steer): recheck indeterminate cancels after ACK * fix(steer): close retry and abort races * fix(cursor): hold ambiguous dispatch failures * fix(steer): serialize live retries and abort admission * fix(steer): distinguish live dispatching from unknown * fix(cursor): bound ACP dispatch acknowledgements * fix(steer): keep ACK failures held and reconcile busy cancel * fix(cursor): preserve state when dispatch ACK is uncertain * fix(steer): distinguish held cancel from removal * fix(store): combine schema v24 migrations * fix(cursor): distinguish held cancel from removal * fix(store): reserve schema v25 for steer delivery state * fix(cursor): suppress late ACP updates after abort * fix(steer): keep held cancel state and notify requeue * fix(cursor): isolate late updates after abort * fix(steer): release explicitly cancelled unknown reservations * test(cursor): cover explicit held cancellation * fix(codex): reject cancelled reservations before native steer * fix(cursor): reject cancelled reservations before ACP steer * fix(codex): make reservation restore atomic with state * fix(cursor): make reservation restore atomic with state * fix(codex): terminate abandoned transport writes * fix(cursor): hard-stop abandoned writes and update native queue state * fix(steer): own abandoned app-server lifecycle and consume races * fix(cursor): isolate aborts and add native retry resolution * fix(codex): confirm dispatch and recover abandoned turns * test(codex): mock abandoned transport callback * fix(native): reconcile retry responses * fix(codex): clear visible turn state on transport loss * fix(steer): claim retries and cover native delivery state * fix(native): resync busy cancel outcomes * fix(native): preserve indeterminate state on Android hydration * fix(steer): make retry claims single-winner * fix(cursor): hold restore failures for explicit resolution * fix(steer): serialize concurrent retry claims * fix(socket): tolerate missing steer-state ACK callbacks * fix(native): serialize retry operations * docs(web): document unknown steer delivery and retry controls * fix(steer): handle retry failures and abort-before-connect * fix(cursor): drain foreground prompt after soft-steer abort * fix(steer): reinitialize after transport loss and finish iOS retry errors * fix(steer): preserve indeterminate rows across reconnect gaps * test(web): mock indeterminate queued recovery state * fix(steer): recover consumed ACK tombstones * fix(steer): expose consumed cancel tombstones * fix(cursor): drain soft steers before handler replacement * fix(cursor): preserve buffered output on abort |
||
|
|
47c0768c27 | feat(brand): refresh HAPI icons across platforms | ||
|
|
f0e5ba9c0f |
feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which agents can deliver queued messages into the active turn (pi, codex, cursor ACP; legacy stream-json cursor excluded) - AgentState.steeringActive, DecryptedMessage.steered and messages-consumed live signal (never persisted by the hub) * feat(cli): queue reservations and steered messages-consumed option - MessageQueue2 gains takeByLocalId/restoreReservation/ beginReservationDispatch/commitReservation so an async steer can reserve a queued row without racing the main loop's turn/start drain - emitMessagesConsumed accepts steered: true to mark mid-turn delivery * feat(codex): mid-turn steer via app-server turn/steer (#888) - CodexAppServerClient.steerTurn + TurnSteerParams/Response types - CodexRemoteLauncher registers the steer-queued-message RPC handler: reserves the queued row, validates it against the active turn (no control commands, matching mode hash), injects via turn/steer with an epoch guard that invalidates in-flight steers on abort/cleanup - steeringActive agent state tracks the active-turn window - hub syncEngine gate opens to codex; messages-consumed relays steered * feat(web): Steered badge and steer gating for codex sessions - HappyUserMessage shows a ↳ Steered badge fed by the live messages-consumed steered signal, preserved across server echoes and refetches (mergeMessages carries the optimistic marker) - SessionChat gates canSteer via isSteeringSupportedForSession instead of the pi-only check - clearStaleQueuedStatus normalizes a queued status on an invoked message - fix(web): drop duplicate showSessionSummaryInChat in markdown test (upstream typecheck breakage) * fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer - STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise codex and pi only; cursor joins when its soft-steer handler lands (#1609) - turn/steer now splits dispatch (stdin accepted) from completion (turn finished): the hub RPC acks once dispatch succeeds — never on the concurrent turn's completion, which can exceed the 30s RPC window - queue row commits only after the turn settles; a rejected/aborted steer restores the row so the message still delivers via turn/start, and a dispatched steer is never restored (no duplicate delivery) - steer carries clientUserMessageId (echoed as userMessage.clientId) so ambiguous transport failures can reconcile the thread later - client tests cover dispatch/complete split and stdin-write failure * fix(codex): reconcile dispatched steers before restoring; align error copy - A dispatched turn/steer whose completion fails (disconnect / protocol error) is now reconciled via thread/read by clientUserMessageId before the queued row is restored — the instruction is only re-delivered by turn/start when the thread never received it - Reconcile targets the pinned steer thread, not whichever turn is current when completion fails - syncEngine unsupported-flavor error now matches the capability gate (Pi and Codex only until the cursor handler lands) - launcher tests cover steer success (ack on dispatch), reconcile-accepted and reconcile-rejected outcomes * fix(codex): consume the row at dispatch; drop background reconcile - The hub RPC acks and the queue row is consumed as soon as stdin accepts turn/steer; completion is background-only logging. A dispatched steer is never restored, so the same localId cannot be re-delivered via turn/start after the caller was told the steer succeeded - Dispatch failure (stdin write error) still restores the row and reports failure - steer.completed rejection is always handled (no unhandled rejection on the dispatch-failure path) - tests updated: completion failure after dispatch keeps the row consumed; dispatch failure restores it * fix(codex): distinguish definite rejection from indeterminate completion - Transport-level failures (timeout, abort, disconnect, spawn, protocol) carry an indeterminate marker; explicit JSON-RPC error responses do not - After a dispatched steer, turn completion resolves → commit + consumed; a definite app-server rejection restores the row (instruction was never accepted, so turn/start cannot duplicate it); an indeterminate outcome leaves the row reserved so it can never be delivered twice - Completion handling registers before awaiting dispatch so the dispatch-failure path cannot leak an unhandled rejection - client/launcher tests cover explicit rejection (restore), indeterminate outcome (row stays reserved) and dispatch failure * fix(codex): reconcile indeterminate steers instead of a permanent reservation - After an indeterminate completion (disconnect/protocol), reconcile the thread by clientUserMessageId immediately: accepted → commit + consumed, provably rejected → restore, still unreadable → keep the reservation and retry from the main-loop top on later passes (post-reconnect) - A row never sits in dispatching forever: the hub cannot stamp it invoked while the instruction may never have been accepted - tests: indeterminate keeps reserved while thread unreadable; accepted reconciliation consumes; rejected path restores * fix(codex): accept all thread item shapes; retry reconcile; ack through abort - Reconcile matcher accepts userMessage/user_message with clientId/ client_id, matching the shapes the thread parser supports — an accepted steer can no longer be misclassified as rejected - A pending reconciliation schedules a wakeLoop retry, so a temporary app-server outage cannot strand the reservation behind waitForTurnOrRecovery - The success-path ACK no longer checks the steer epoch: the hub already reported steered on dispatch, so commit + messages-consumed must reach it even when an abort resets the queue in between * fix(codex): reinit reconnected app-server; keep reconcile retries alive - thread/read after a disconnect auto-connects a fresh app-server, which must be initialized before any request — reconcile now ensures connect + initialize (isConnected getter added to the client) - every still-unknown loop-top reconciliation schedules the next retry, so recovery without external traffic is eventually observed - launcher mock gains isConnected * fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK - Reconciliation runs on a self-rescheduling 1s timer independent of the main loop (wakes it too), so idle loops and waitForTurnOrRecovery still observe app-server recovery; abort clears nothing implicitly — the ACK path commits and consumes even when the reservation was cancelled - Absence of a durable client id is ambiguous: unmatched reads stay 'unknown' and keep retrying instead of restoring the row - CodexAppServerClient tracks initialized state (reset on disconnect/exit) so ensureAppServerInitialized re-initializes a fresh process before thread/read; initialize failures leave the flag false for the next retry - tests: accepted reconciliation via scheduled timer, indeterminate keeps reserved, explicit rejection restores * fix(codex): bind reconciliation to the launcher lifecycle - runSteerReconciliation clears any armed retry timer on entry and never installs a second one, so loop-top and timer-driven passes cannot multiply - shuttingDown is set when the main loop ends: timers are cleared and the pending map is dropped, so an unresolved steer can never respawn an app-server after cleanup (remote-to-local switch included) * fix(codex): report steered only after app-server acceptance - The handler now awaits steer.completed (the inject-acceptance response): an explicit JSON-RPC rejection surfaces as failed and restores the row for the normal turn/start path instead of a false steered - Transport failure after dispatch reports 'Steer outcome is being reconciled' and keeps the row reserved while the timer-driven thread reconciliation runs - dispatch-failure path also swallows the paired completion rejection * fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait - MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching steer reservation: the hub neither deletes the row nor stamps invoked_at (new CancelMessageResponse 'busy' status; web restores the optimistic row); pushIsolateAndClear and reset/close share cancelReservations so /clear-style commands cannot have a rejected steer resurrect a discarded prompt - turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a lost response is indeterminate and funnels into thread reconciliation instead of stranding the reservation - tests updated for the tri-state cancel contract * fix(codex,web): busy-aware edit flow; bound reconciliation reads - QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it never prefills the composer when the row is inside an async steer, so a second client cannot send a duplicate - reconcileSteerByClientId bounds thread/read with a 5s timeout so a connected-but-silent app-server cannot hold the reservation in-flight indefinitely * fix(steer): inFlight-dominated cancel acks; bounded reconciliation - hub cancel-queued-message acks check inFlight before removed: a stale duplicate socket reporting removed can no longer delete the durable row while another socket is dispatching the steer - reconciliation entries expire after 60s and mark delivered: after the rejection window, a dispatched steer that the app-server never proved (client ids dropped on restart) is committed instead of polling thread/read forever - pre-dispatch failures (abort before write included) never enter reconciliation — they restore the row and report failure * fix(steer): persist indeterminate outcomes without replay * fix(steer): make ambiguous delivery restart-safe * fix(steer): recover crash-held rows and preserve retry dedup * fix(steer): ack retries and bound stdin dispatch * fix(steer): reconcile indeterminate dispatches and serialize retries * fix(codex): classify stdin callback failures as indeterminate * fix(steer): recheck indeterminate cancels after ACK * fix(steer): close retry and abort races * fix(steer): serialize live retries and abort admission * fix(steer): distinguish live dispatching from unknown * fix(steer): keep ACK failures held and reconcile busy cancel * fix(steer): distinguish held cancel from removal * fix(store): combine schema v24 migrations * fix(store): reserve schema v25 for steer delivery state * fix(steer): keep held cancel state and notify requeue * fix(steer): release explicitly cancelled unknown reservations * fix(codex): reject cancelled reservations before native steer * fix(codex): make reservation restore atomic with state * fix(codex): terminate abandoned transport writes * fix(steer): own abandoned app-server lifecycle and consume races * fix(codex): confirm dispatch and recover abandoned turns * test(codex): mock abandoned transport callback * fix(codex): clear visible turn state on transport loss * fix(steer): claim retries and cover native delivery state * fix(native): preserve indeterminate state on Android hydration * fix(steer): make retry claims single-winner * fix(steer): serialize concurrent retry claims * fix(socket): tolerate missing steer-state ACK callbacks * fix(native): serialize retry operations * docs(web): document unknown steer delivery and retry controls * fix(steer): handle retry failures and abort-before-connect * fix(steer): reinitialize after transport loss and finish iOS retry errors * fix(steer): preserve indeterminate rows across reconnect gaps * test(web): mock indeterminate queued recovery state * fix(steer): recover consumed ACK tombstones * fix(steer): expose consumed cancel tombstones |
||
|
|
2fbd98dfe4 |
fix(pi): offer model-accurate thinking levels in the create-session form (#1626)
* fix(web): filter Pi effort options by model thinkingLevelMap in new session form The create-session EffortField called getPiThinkingLevelOptions without the selected model's thinkingLevelMap, so the Pi effort select always showed the static off..high list: levels the model marks unsupported stayed visible and xhigh/max never appeared even for models that opt in. Pass the map through piSelectedModel (mirrors HappyComposer), and extend the stale-effort reset effect so a level unsupported by the newly selected model falls back to auto. * feat(cli): probe machine Pi models over RPC to carry thinkingLevelMap The machine-level Pi model probe parsed the `pi --list-models` text table, which only exposes provider/model/thinking-yes-no — thinkingLevelMap (and name/contextWindow) never reached the create-session form, so model-accurate thinking levels could not render there (xhigh/max are map-opt-in and were permanently hidden; see the companion web commit). Replace the table probe with a short-lived `pi --mode rpc` child (--no-session --no-extensions --no-skills --no-prompt-templates --no-tools) that issues get_available_models and reuses the session path's parsePiModels schema, so machine and session catalogs share one wire contract. The lightweight probe also measures faster than the table probe (~0.8s vs ~1.6-2.4s) and drops the whitespace-table parsing entirely. Cache/inflight dedupe/timeout structure is unchanged. * fix: address review findings on the Pi model probe and effort reset Three review findings on the RPC probe / effort-map change: - (high) Probe teardown killed only the direct child PID: with shell:true on Windows that is the shell, orphaning the interactive pi RPC process on every successful probe and on timeout; finish() also resolved before the process was confirmed gone. Use killProcessByChildProcess (taskkill /T on Windows, Unix tree kill, SIGKILL escalation) and settle only after the tree teardown completes. - (medium) An explicit get_available_models success:false response was discarded, so Pi's own error text was lost and the interactive child hung until the generic 15s timeout. parsePiModelsProbeLine now returns a three-way result (unrelated/models/error), also validating the response id, and the probe rejects immediately with Pi's error. - (medium) Switching an xhigh/max-capable model back to Default left the now-hidden effort in state and submitted it while the select visually fell back to auto. The reset effect now also covers model === 'auto' (undefined map) while still not resetting mid-resolve for a concrete model. Tests: probe-line failure/foreign-id cases; NewSession restore-to-Default reset and map-opt-in retention (the latter guards the former against a false green from the restore path). * fix(web): type the Pi model test mock as PiModelSummary The inline mock element type omitted thinkingLevelMap, so the new capability-driven tests failed typecheck (TS2353) and the required test job stopped before the unit tests ran. Use the shared PiModelSummary type so the mock cannot drift from the wire contract again. * fix(web): reconcile hidden Pi effort when model discovery fails A restored explicit Pi model never resolves when the machine catalog request fails, so piSelectedModel stays null for good. The reset effect required a resolved model, so it skipped reconciliation, while EffortField rendered with an undefined map and hid xhigh/max. Creation is only gated on the loading state, not on the error, so handleCreate could still forward the stale hidden level for Pi to reject or clamp. Treat a failed catalog as a settled selection (alongside Default and a resolved model) and reconcile against the undefined map; keep skipping the reset while a concrete model is still resolving without an error, so a restored xhigh/max survives until the map can prove it valid. Test asserts the spawn payload carries no effort after a failed catalog with a restored xhigh; verified it fails when the error branch is reverted. * fix(cli): honor a failed probe process-tree teardown killProcessByChildProcess reports survivors by resolving false, but the probe discarded that result and settled anyway. The Windows graceful path is taskkill /T without /F and escalates nothing on its own, so a probe child that refuses the signal would be reported as cleaned up and the catalog cached, letting interactive Pi processes accumulate across refreshes. Escalate to the forced teardown when the graceful one reports survivors, and reject (caching nothing, so the next call re-probes) when even that fails. Skip the check when the child has no pid: nothing can leak, and replacing a spawn ENOENT with a teardown error would only obscure the real failure. Adds probe lifecycle tests (spawn + teardown helper mocked) for the escalation path, the reject-and-do-not-cache path, the no-pid spawn-failure path, and the timeout path; verified they fail when the escalation is reverted. * fix(cli): keep extensions enabled for the Pi model probe --no-extensions silently dropped providers contributed through pi.registerProvider, which the old `pi --list-models` probe did list. Users with such an extension would have lost those models in the create-session form only. Verified with a project-local .pi/extensions provider in the same cwd: the probe with --no-extensions returned 29 models, while both the default run and the old table probe returned 30 including the extension's model (and its thinkingLevelMap). Dropping the flag restores parity; the extension provider now comes back with its map intact. Discovery that cannot contribute models stays disabled (--no-session, --no-skills, --no-prompt-templates, --no-tools). Cost: the probe now measures ~1.4-2.0s instead of ~0.65s, still at or below the old table probe (~1.6-2.4s) and fronted by the existing 60s cache. * fix(cli): probe Pi models from the home directory, not the runner cwd Under launchd/systemd the runner cwd is `/`, and starting Pi there is pathological: project discovery plus extensions that scan from the working directory walk the whole filesystem root. With extensions enabled (required so pi.registerProvider models still surface) the RPC probe took 16.8s at cwd=/, past the 15s timeout, so machine-level model discovery failed 100% of the time and the create-session form showed only 'Pi model discovery timed out'. Measured on macOS with 9 global extensions: cwd=/ extensions on 16.8s -> timeout cwd=/ extensions off 0.6s -> would lose extension providers cwd=$HOME extensions on 1.4s -> 29 models, complete The catalog is machine-scoped and does not depend on cwd, so probing from the home directory is both safe and representative. Verified from cwd=/ under the runner's exact launchd environment: 1.3-1.7s, 29 models. The old `pi --list-models` probe was immune because it never initialized a session; this regression arrived with the RPC probe and was missed because every earlier verification ran from a project directory. Tests assert the spawn cwd is homedir() and that --no-extensions stays absent; verified the cwd test fails when the option is removed. * fix(cli): keep the probe in the runner cwd, fall back to home only at a root Forcing every probe to the home directory fixed the launchd timeout but broke project-local discovery: a runner started inside a project stopped seeing that project's .pi/extensions providers, which the replaced --list-models probe did surface. Verified with a project-local provider: present from the project dir (30 models), absent from home (29). Use the runner cwd normally and fall back to home only when cwd is a filesystem root -- the launchd/systemd case where Pi startup walks the whole tree (16.8s, past PROBE_TIMEOUT_MS) and where there is no project to lose anyway. process.cwd() can also throw for a deleted directory, so that falls back to home too. Verified under the runner's launchd environment: from / -> home, 1.6s, 29 models; from a project dir -> that dir, 1.4s, 30 models including the project-local provider. Tests cover all three branches and were checked in both directions: pinning to home fails the project-cwd test, pinning to cwd fails the root test. |
||
|
|
f110845be6 |
merge: K6 fixtures batch 2 (events, permissions, sidechains, flavors, catalogs)
# Conflicts: # shared/fixtures/README.md # web/scripts/fixtures/generate.ts |
||
|
|
e1bfdd6265 |
feat(fixtures): SSE patch + pagination scenario fixtures (K7)
Two new golden-fixture suites, expectations machine-generated from the web
implementation (same source-of-truth principle as the chat suite):
- shared/fixtures/sse/ (12 cases): session-updated versioned-patch
application via applySessionDetailPatch — strict version gates for
metadata/agentState/todos/teamState, out-of-order arrival, teamState:null
clear, max-monotonic updatedAt, flat last-write-wins fields, sub-minute
activeAt keep-alive drop, scratchlistUpdatedAt trigger, and the pinned
actual behavior that activeTurnStartedAt is NOT applied by the patch path.
Inputs are stored schema-normalized and validated against SessionSchema /
strict SessionPatchSchema at generation time; per-patch applied/unchanged
verdicts are part of the contract.
- shared/fixtures/pagination/ (11 cases): op scripts driving the real
message-window store with a scripted ApiClient — latest page + SSE ingest,
before-cursor older pages, epoch-mismatch reset (window discard + recorded
internal latest request), reset:true replace preserving optimistic rows,
localId echo reconciliation, messages-consumed invokedAt stamping (no
cursor advance), message-cancelled removal, cancel-too-late invoked-row
ingest, hidden-row cursor advance, 400-row trim preserving queued rows,
and queued-state gap recovery. Documents pin the exact requests the store
issued, older-load outcomes, reconcile candidates, and a minimal window
projection (ids/order, queued/optimistic flags, hasMore, epoch, viewMode,
compound cursors).
Generator: web/scripts/fixtures/{sse,pagination}/ extend the K5 framework
(canonical serialization reused; suites pruned of stale files). Web
self-conformance tests replay every fixture against the real implementation
in bun run test:web. README documents schemas, projections, replay
contracts, and native consumption for both suites.
No time injection needed: the store's only Date.now() reads gate
notification throttling, which never reaches persisted or projected state.
|
||
|
|
c1ecb4798e |
feat(fixtures): batch 2 — events, permissions, sidechains, flavors, catalogs (K6)
34 new golden chat fixtures generated from the web pipeline, plus a machine-generated modes catalog: - Claude output: away_summary/microcompact_boundary/compact_boundary system subtypes, <task-notification> summary extraction, multi-block assistant message with an array-of-parts (text+image) tool_result - AgentEvent union completed: switch/message/error/title-changed(+dedupe)/ limit-warning(pipe text)/api-error folding/turn-duration/compact-summary/ abort-restore/unknown-type passthrough; thread-goal updated+cleared (silent) - Permission lifecycle via agentState.completedRequests: approved_for_session (mode+allowTools), denied with reason, canceled, AskUserQuestion flat answers, request_user_input nested answers (all with explicit createdAt) - Sidechain: Task tool_use with parentToolUseId-grouped subagent children - codex family: error, context_compacted, generated-image, review JSON, thread-goal messages, CodexBash call/result pair, exploration tool group - cli-output: cli-origin command block + stdout merge - agy flavor (output envelope): agy_message, agy_tool_action run_command mapping; cursor ACP plan snapshot (copilot skipped: identical codex shapes) - generator now also emits shared/fixtures/catalogs/modes.json from shared/src/modes.ts (per-flavor permission modes with labels/tones); self-conformance test covers it; README documents the catalogs contract |
||
|
|
9479f667b5 | feat(fixtures): golden chat fixtures generated from web pipeline (K4+K5) | ||
|
|
76e00f0091 |
feat(hub): add background-only ServerChan fallback (#1608)
* feat(hub): add background-only ServerChan fallback * fix(hub): validate ServerChan background setting type * fix(web): preserve pending visibility transitions * fix(web): guard visibility reports across subscriptions |
||
|
|
6c6f4b4929 |
feat(agy): replace fragile PTY/TUI wrapper with headless print-mode transport (#1591)
Replace the Antigravity (agy) integration — a PTY wrapping the TUI with
output-marker scraping ('? for shortcuts', 'Generating', trust dialogs,
/model picker navigation, quota-screen regex) — with a headless print-mode
transport: every user turn spawns `agy -p <msg> --conversation <uuid>
--output-format stream-json`, and NDJSON events (init/step_update/result)
map onto the existing transcript-entry channel (sendAgySessionMessage), so
hub/web rendering is unchanged. ~8.3k LOC (incl. tests) removed.
Fixes #1588. Design: docs/design/agy-headless-transport.md.
CLI:
- new cli/src/agy/headless/: agyNdjsonParser (pure functions, malformed-line
tolerance, step conversation-id adoption), AgyPlannerAccumulator (per-step
delta accumulation with settling retries), AgyHeadlessDriver (per-turn
spawn/kill loop, NDJSON chunk buffering, authoritative delivery ack via
user_input/result, interrupt + retry + shutdown lifecycle with consume/
restore, process-tree termination, SSH agent preserved, prompt log
redaction, per-turn model snapshot with conversation-DB fallback)
- runAgy/loop/session rewired; agy is remote-only (no PTY, no local mode,
no local-switch action); queued batches snapshot model/effort/mode
- deleted agyPty, agyPtyLauncher, agyHookCarrier(+scope cache), agyModelKeys,
agyQuestionKeys, agyAskQuestion, agySessionScanner, agyPermissionHandler
(+tests); buildAgyHooksJson removed; startHookServer agy-pre-invocation
route → 200 no-op
- runner: agy reopen/resume via generic --existing-session-id; commands/
agy.ts defaults remote; resume rejects ACTIVE agy sessions (remote-only,
in-flight turns cannot hand off)
- MCP stays user-managed (agy reads ~/.gemini/config/mcp_config.json and
workspace .agents/mcp_config.json natively in headless — verified)
Hub/web:
- machines.ts drops agy→pty forcing and rejects non-remote startingMode
- NewSession drops agy startingMode='pty'; terminal toggle disappears
automatically; RemoteModeDisplay hides the local-switch hint when absent
- docs/guide/agents.md updated: headless print mode, no PTY/hooks, MCP via
user's own mcp_config.json
Tests: 47 parser+driver tests (fake-binary e2e, chunk-split NDJSON, delivery
ack semantics, interrupt/retry/shutdown races, model attribution, EOF
framing, malformed envelopes); full suite green (cli ~2340, hub 1093,
web 2474, shared 262). Real-binary smoke on agy 1.1.13: single turn exit 0,
--conversation resume keeps the same conversation_id.
|
||
|
|
a6feb6e8ba |
feat: unified agent configuration descriptors (session config consolidation) (#1469)
* feat(config): add agent config descriptor protocol and advertise via runner capability Introduce shared agent configuration descriptors covering model, effort, permission, and secondary settings per agent flavor, plus the canonical HAPI YOLO -> native permission mode mapping. Runners advertise the builtin descriptors through the runner-state capability so hubs and web can render configuration without hardcoded flavor branches. Migrate the OpenCode create-session model picker from a bespoke radio list to the shared SelectControl combobox. * feat(web): render create-session permission from agent config descriptor Replace the flavor-branched Grok/Codex-family/YOLO permission block with a descriptor-driven PermissionField. Pi now reports permission as managed instead of silently ignoring the YOLO toggle, and YOLO-only flavors show the native permission mode the preference maps to. Removes the superseded GrokPermissionModeSelector and CodexFamilyPermissionModeSelector components. * ci: retry flaky claudeRemote 5s-timeout failure * fix(web): persist explicit OpenCode Default selection instead of restoring a concrete model The parent initialization effect treated every null selected model as 'uninitialized' and auto-picked a concrete advertised model, clobbering the user's explicit Default choice (and a restored Default preference). null now means explicit Default and is preserved; only undefined (no choice made yet) triggers probe-based initialization. Add parent-level regression tests for Default persistence and remembered-model restore. * fix(web): accept undefined selected model in OpencodeModelSelector props * feat(web+cli+hub): unify create-session model/effort fields and add Pi model/effort support Pi's agent config descriptor now advertises model (machine) and effort (static thinking levels) for create AND session availability: - cli: ListPiModelsForMachine RPC runs 'pi --list-models' (cached, inflight deduped) and parses the provider/model table; startup model match accepts provider-qualified ids - hub: GET /api/machines/:id/pi-models route + rpcGateway/syncEngine passthrough - web: NewSession renders Pi models grouped by provider through the generic ModelSelector and a new descriptor-driven EffortField (replaces the per-flavor LaunchEffortSelector/ReasoningEffortSelector pair); launch payload forwards Pi model + thinking-level effort (runner already supported --model/ --effort for pi) * fix(web): render Pi provider groups in ModelSelector and scope Grok availability warning - ModelSelector now renders grouped options as <optgroup> (Pi models are provider-grouped; identical modelIds from different providers stay distinct) - PermissionField only receives autoPermissionModeSupported for Grok — a cached Grok probe result no longer leaks the Grok warning onto other agents Addresses HAPI Bot Minor findings on #1469. * fix(web): drop Object.groupBy from ModelSelector; revalidate restored Pi models against the catalog - ModelSelector buckets options with a reduce instead of Object.groupBy (Safari < 17.4 has no polyfill — New Session would throw on those clients) - Pi restored model/effort are cleared when the value is absent from the live machine catalog, and Create waits for the catalog while a non-default Pi choice is being validated (mirrors Codex/Grok/Copilot handling) Addresses HAPI Bot findings on #1469. * fix(pi+web): serialize startup model before thinking level; hide Pi launch controls during history import - PiSession gains startupModelSettled; the startup set_thinking_level waits for the requested model's set_model attempt to settle first, so a level the default model rejects is not lost before the requested model is confirmed (set_model and set_thinking_level were already serialized by the runtime mutation lock; this pins the model-first ordering) - Create Session hides Pi model/effort controls while a Pi history import is selected — the import reopens the native session as-is and would silently ignore launch-only model/effort values Addresses HAPI Bot findings on #1469. * fix(pi): settle startup-model gate when model discovery fails or returns no models A failed or empty get_available_models response would leave the startupModelSettled gate unresolved, stranding a requested startup effort indefinitely. Resolve the gate on the error path and the empty-models path; adds regression tests for both. |
||
|
|
79ffe9aa0e |
feat(web): composer model/effort value buttons and first-class permission (part of #1438) (#1475)
* feat(web): composer model/effort value buttons and settings order Wide composers now show [model] and [effort] value buttons for non-Pi flavors (labels from the current session values), opening the settings sheet on click. Narrow viewports collapse to the settings button only via a new useNarrowViewport hook. The settings sheet reorders to Model -> Effort -> Permission -> other settings (Fast mode, collaboration, Copilot agent mode) so permission is first-class. Toolbar customization gains 'model'/'effort' items with settings labels. Pi keeps its dedicated model/thinking panels unchanged (unified descriptor-driven sheet is a follow-up). * fix(web): satisfy strict types in composer value-button test harness * fix(web): address review findings on composer model/effort value buttons - Normalize null/'auto'/'default' model wire values onto the value:null option so default-model sessions keep a localized label button (Major) - Exempt model/effort value buttons from the settings outside-click dismissal so a second click closes the sheet instead of reopening it - Read matchMedia synchronously in useNarrowViewport so narrow first paints never flash the wide toolbar - Add regression tests: model=null/'auto' labels, toggle-close behavior, and initial narrow-viewport render * feat(web): fold Pi into the generic composer model/effort value buttons Pi sessions previously exposed model/effort twice: dedicated 'Pi model' / 'Pi thinking level' toolbar buttons (PiModelPanel/PiThinkingLevelPanel) AND the settings sheet's generic Model/Effort sections. Consolidate so Pi looks exactly like every other flavor: - Pi now uses the generic model/effort value buttons; labels resolve from the provider-qualified piModels catalog (name -> modelId -> session id). - The settings sheet's Model section already renders provider-grouped Pi rows and Effort renders Pi thinking levels, so the dedicated panels and their toolbar slots are deleted. - Keep Pi's mid-turn control affordance (#1442): Pi turns hold thread.isDisabled for minutes, so Pi model/effort controls stay enabled while a turn is running (configurationControlsDisabled instead of controlsDisabled), including the sheet rows. - Drop 'piModel'/'piThinking' toolbar layout items; persisted layouts normalize them away automatically. Tests: pi value-button label/sheet tests, mid-turn model selection via the unified sheet, toolbar layout defaults. * fix(web): keep a session-settings trigger on narrow viewports and mid-turn Pi Address the HAPI review bot's two Minor findings on the Pi consolidation: - Narrow viewports collapse the model/effort value buttons into the settings sheet, so a persisted toolbar layout hiding the gear left no session-settings trigger at all. ComposerButtons now forces the gear back into the rendered layout on narrow viewports (wide layouts keep honoring the user's hidden choice). - The Pi mid-turn live-control rule only reached the value buttons and sheet rows; with those buttons gone on narrow, the gear was still disabled by controlsDisabled for the whole (minutes-long) Pi turn. HappyComposer now passes settingsDisabled={modelEffortControlsDisabled} so the gear stays clickable mid-turn for Pi exactly like the buttons. Tests: narrow + hidden-gear layout keeps Settings; narrow Pi mid-turn gear stays enabled and opens the provider-grouped sheet. * fix(web): address HAPI Bot Pi settings-sheet findings Three Minor findings from the review bot on the unified Pi sheet: - Provider-qualified selection: rows compared only modelId, so duplicate model IDs across providers all looked selected. Compare against piSelectedModel's provider+modelId when available. - Thinking-level reset: the removed Pi panel toggled the current level back to null; the unified effort rows only submitted concrete values. Re-clicking the selected effort row now clears it for Pi. - Memo staleness: the settings-sheet memo did not depend on modelEffortControlsDisabled, so a Pi disabled-state transition while the sheet was open left rows enabled from the prior render. Tests: colliding model IDs highlight only the matching provider row, re-clicking the selected effort row sends null, and a rerender with active=false disables the open sheet's rows. * fix(web): include piSelectedModel in settings-sheet memo deps The overlays memo reads piSelectedModel for provider-qualified row highlighting but only declared the derived selectedPiModel. When piSelectedModel hydrates from absent to a qualifier that resolves to the same catalog object, the memo is reused and duplicate model IDs stay highlighted across providers. Add the raw prop to the dep array. * fix(web): gate Pi model rows on catalog; reset drill-down via value button Address the HAPI Bot review on the unified composer settings sheet: - Pi no longer falls back to the generic synthesized modelOptions rows when its provider catalog is empty/loading. Selecting one of those would post a bare model id that runPi cannot resolve to a provider (first cached match or 409). The Model section now only renders for Pi when piModelGroups exists, and renders grouped rows exclusively. - Closing the sheet through the model/effort value button now goes through handleSettingsToggle, so a Cursor variant drill-down resets to the base model list on reopen (previously only the gear and outside-click paths cleared it). Tests: empty Pi catalog hides Model section; value-button close resets Cursor drill-down. * fix(web): hide Pi effort controls until the selected model resolves Address the HAPI Bot review: with the catalog still loading/failed there is no selectedPiModel to derive a capability map from, but the unified effort control stayed enabled (mid-turn Pi controls are intentionally live). Selecting a level would send set_thinking_level for a model that may not support reasoning, and the RPC can be rejected after the sheet closed. The old dedicated panel guarded this state. showEffortSettings now requires a resolved, reasoning-capable Pi model; the effort value button hides the same way. With an empty catalog Pi exposes no settings trigger at all, matching the old control states. Tests: unresolved Pi catalog mid-turn exposes no effort action. * fix(web): hide the Pi model trigger until the catalog resolves Address the HAPI Bot Minor finding: with an empty/loading Pi catalog the model value button fell back to the bare session model id and rendered an enabled trigger that opens no Model section. Show the button only once the provider-qualified catalog entry resolves. |