Commit Graph
591 Commits
Author SHA1 Message Date
SSU-WEI HUANGandGitHub d332cd5957 fix(acp): preserve permission state when tool input is missing (#1782) 2026-09-09 09:22:27 +08:00
NightWatcher314andGitHub 97b1dd4c34 fix(codex): preserve context config in remote sessions (#1633) 2026-09-06 14:19:44 +08:00
Junmo KimandGitHub b85c72f09f fix(claude): accept first fork child prompt on SessionStart:fork hook (#1671)
* test(claude): reproduce fork child deadlock when init never arrives

* fix(claude): accept first fork child prompt on SessionStart:fork hook

A forked child starts query() before any prompt exists, and the SDK
emits 'init' only after the first prompt is sent. Waiting for init
deadlocked fork children until forever. Materialization is already
signaled by the SessionStart:fork hook, so accept the child prompt on
that signal instead.

* chore: restore bun.lock
2026-09-06 14:19:32 +08:00
Junmo KimandGitHub 4603c883a0 fix(cli): identify the Windows runner process through CIM (#1755)
* fix(cli): identify the Windows runner process through CIM

isHapiRunnerProcess() checks that the PID persisted in runner state still
belongs to a HAPI runner. On Windows it read the command line through wmic
and returned true for any live PID when that spawn failed.

wmic is absent on current Windows 11 builds, so that branch is the only one
those hosts take and the check degrades to a liveness test. When a reboot
hands the stale runner PID to another process, runner start adopts it and
exits without starting a runner.

Query the command line through PowerShell Get-CimInstance Win32_Process
first and keep wmic as the fallback, matching getProcessStartMarker() in
this file. A probe that succeeds without reporting a command line falls
through to the next one, and an unreadable identity still preserves the
live state, so a real runner is never discarded on a probe that came back
empty.

* fix(cli): keep an unverifiable runner pid out of destructive paths

The identity check answered a boolean, so "alive but unidentifiable" had to
collapse into one of the two verdicts. It collapsed into "this is the
runner", and runner start acts on that by calling stopRunner(), which force
kills the persisted pid once the HTTP stop times out. A reused pid that no
probe can read was therefore still reachable by a kill.

Report the identity as runner, foreign, unknown or dead instead. Only a
confirmed runner is reported as running. An unknown pid is neither signalled
nor cleaned up, because clearing the state would also drop the lock that
keeps a second runner from starting, and the pid may still be a healthy
runner behind a probe that failed transiently.

The posix branch had the same fail-open through isProcessAlive and now takes
the shared unknown path.

* fix(cli): gate the runner force kill on a confirmed identity

Two paths reach stopRunner() with a boolean that cannot carry "unverified".
runner start stops an existing runner, and the spawned start-sync child reads
the same answer through isRunnerRunningCurrentlyInstalledHappyVersion(), where
a false result is read as a version mismatch and also calls stopRunner().
Propagating the unknown state to each caller would leave the next path open,
so check the identity where the process is actually signalled: only a
confirmed runner is force killed.

Also stop reading WMIC's column header as a command line. wmic get CommandLine
prints the header even when the property is empty or unreadable, so trimmed
stdout was never empty on success and an unreadable process was classified as
foreign, which cleared the state and the lock instead of preserving them.

* fix(cli): gate the whole runner stop on a confirmed identity

The identity check sat in front of the force kill, but stopRunner() first
sends an unauthenticated POST /stop to the port recorded in runner state.
An http request is a signal too: if the pid and its port were both reused,
an unrelated local service receives that request before the check runs.

Move the check to the start of stopRunner() so a pid that is not a confirmed
runner is never contacted at all. The later force-kill check is redundant
once the whole sequence is gated, so it goes away.
2026-09-06 14:19:14 +08:00
AnanovoandGitHub aa0c1dc808 fix(codex): fail closed on ambiguous Web Rewind boundaries (#1707)
* fix(codex): fail closed on ambiguous Web Rewind boundaries

* fix(codex): reject malformed native rewind history

* fix(codex): reject duplicate rewind identifiers

* feat(codex): offer Fork fallback for ambiguous rewind

* fix(codex): gate rewind Fork fallback on exact native boundary

* fix(codex): reject unresolved rewind boundaries safely

* chore: trigger PR checks

* fix(codex): gate safe rewind fallback on fork support

* fix(codex): require complete user ids for rewind fallback
2026-09-06 14:18:40 +08:00
f4553dd1ec fix(codex): ignore late app-server writes during disconnect (#1748)
* fix(codex): ignore late app-server writes during disconnect

* fix(codex): scope app-server responses to their process

---------

Co-authored-by: huxiang <huxiang@myai.tech>
2026-09-06 14:18:20 +08:00
Junmo KimandGitHub 980a921ba1 test(cli): provide a stub Claude CLI so runner spawn passes the agent availability preflight (#1696)
The runner integration suite spawns sessions whose new agent-availability
preflight requires an installable Claude CLI. CI runners have none, so every
spawn returned agent_unavailable and four suite tests failed.

The isolated test env now writes a minimal stub claude binary into the temp
home and points workers at it via HAPI_CLAUDE_PATH (the same override the
production launcher honors), keeping production behavior untouched.
2026-08-29 12:52:47 +01:00
Junmo KimandGitHub ec08959f07 fix(agy): show Gemini 3.7 Flash in the agy model list (#1585)
* refactor(agy): extract the agy models probe from the fetch flow

One invocation and the decision of what to do with its output were
tangled in a single promise. Split them so the probe can be given
different arguments, and run more than once, without duplicating the
stream and timeout handling.

No behaviour change: the same argument vector, timeout, stream
handling and fallbacks remain, and the existing tests are untouched.

* fix(agy): read the model list from agy's structured output

The picker recovered ids and labels from `agy models`, whose table is
meant for people to read: it has emitted display names only, then
bare ids, then tab-separated id/label pairs over the 1.1.x line, and
each shape change silently sent the probe back to the hardcoded
mirror. Since agy started offering Gemini 3.7 Flash, that mirror is
what users see, so the three new entries never appear.

agy publishes the same listing as a structured payload, so read that
instead: `agy --output-format=json models` returns
`command.data.models[] = {id, label}`. The flag is global, so it goes
before the subcommand and takes its `=` form.

Releases that predate the flag ignore it and print the table, so the
text parser stays as the fallback, and it now also understands the
tab-separated shape. Every release checked, 1.0.16 through 1.1.13,
ignores the flag rather than rejecting it; a build that rejected it
would emit no models at all, so ask once more without the flag when
the first invocation yields nothing either parser can read.

* fix(agy): follow the Gemini 3.7 Flash row in the model picker

agy now lists Gemini 3.7 Flash as the top row of the `/model` TUI
picker, but the hardcoded row table still starts at 3.6 Flash. A
session whose current model is that row can never change its model:
the picker cannot identify the current row and the change is
rejected.

Add the 3.7 row to MODEL_ROWS/TARGETS. Every existing row shifts
down by one, which leaves their relative deltas, and therefore the
navigation keys emitted between them, unchanged. Mirror the three
new ids into AGY_MODEL_LABELS so the session pickers offer them too.
2026-08-26 09:37:58 +08:00
weishu e5a8212f4a feat(session): validate agents and browse workspace directories 2026-08-25 16:12:29 +08:00
weishu a04275b51d chore: upgrade Bun to 1.4.0 2026-08-25 13:31:59 +08:00
SSU-WEI HUANGandGitHub be1ef2a2e4 feat(dsh): integrate DeepSeek Harness through ACP (#1632)
* feat(dsh): add DeepSeek Harness ACP flavor

* fix(dsh): update mobile flavor catalogs

* fix(dsh): keep mobile spawn policy managed

* fix(dsh): keep managed policy and prompt retry

* fix(dsh): suppress unsupported runner policy flags

* fix(dsh): align native managed-policy UX
2026-08-22 12:37:28 +08:00
Junmo KimandGitHub 661e9b4eb7 feat(web): show Claude round usage metadata (#1655) 2026-08-22 12:37:03 +08:00
AnanovoandGitHub 3d94e8eef3 fix(cli): preserve source extensions for generated media (#1650) 2026-08-20 19:36:36 +08:00
Junmo KimandGitHub 0aebf39c78 fix(opencode): keep one stored message per reasoning stream (#1643)
* fix(acp): carry the live reasoning marker on the wire payload

ACP agents stream thoughts a token at a time, so the handler coalesces
them into a buffer and re-sends the whole buffer under a stable stream
id every 250ms. The converter dropped the marker that says a payload is
one of those throttled snapshots, leaving the hub unable to tell a
replaceable snapshot from the settled message that closes the stream.

Mirrors how the text variant already forwards streamSnapshot.

* fix(hub): keep one stored message per reasoning stream

OpenCode reasoning arrives as a series of growing snapshots sharing one
stream id, and every snapshot was persisted as its own message. A 26h
session reached 48,844 rows and 63MB, and because the web budgets a
fixed number of messages, its 400-message window covered barely three
minutes of conversation — scrolling up walked through duplicate
snapshots instead of history.

Retire a stream's earlier live snapshots once their replacement is
stored. Sweeping only after the insert matters: the two statements are
separate transactions, so clearing first would leave a window where a
crash takes the whole stream. Only rows marked live are eligible and the
replacement is spared, so a stream always keeps at least one row and the
settled message that closes it is never removed.

Live rendering is unchanged: the web still receives every snapshot and
already folds them by stream id.

* fix(web): spend the message window on conversation, not repeated snapshots

The window budgets raw messages, but a reasoning stream renders as a
single folded block no matter how many snapshots it arrived in. On
sessions recorded before the hub started retiring them, those snapshots
fill the window on their own: in one 26h session the newest 400 messages
covered 202 seconds, so scrolling up paged through duplicates instead of
history.

Collapse each stream to its newest snapshot before trimming. Rendering
is unchanged — the timeline already folds them by stream id — and rows
without a stream id are never touched.

* fix(ios,android): port reasoning-snapshot compaction to the native windows

The window logic in HapiProtocol and :core:protocol is a one-to-one port
of the web store, so collapsing superseded reasoning snapshots only on
the web left the native windows budgeting raw snapshot rows. The hub
stores one row per stream now, but a client that already holds the older
snapshots still spends its window on them.

Add the same stream-id reader and compaction to both ports, in the shape
each already uses for agent-run rows, and pin the behaviour with a
pagination fixture. Both fixture suites enumerate shared/fixtures/pagination
from disk, so the ports cannot drift from the web again without CI saying
so.
2026-08-20 08:55:21 +08:00
SSU-WEI HUANGandGitHub 0f7a3da68b feat(cursor): mid-turn Steer via concurrent ACP session/prompt (#888) (#1609)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* feat(acp): split request dispatch from completion and add soft steer

- AcpStdioTransport.sendRequestWithDispatch separates stdin-accepted
  dispatch from the JSON-RPC response, keeping sendRequest behavior
  unchanged
- AcpSdkBackend tracks concurrent session/prompt requests with an
  activePromptRequests counter (main prompt + soft steers); response
  completion stays pending until every concurrent prompt settles
- beginSoftSteerPrompt kicks off a concurrent session/prompt (Cursor GUI
  Send semantics — no cancel, no handler swap) returning {dispatched,
  completed}; softSteerPrompt awaits the full response for direct callers

* feat(cursor): mid-turn soft steer via concurrent session/prompt (#888)

- CursorAcpRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, rejects control commands and mode mismatches,
  then soft-injects via beginSoftSteerPrompt without canceling the
  in-flight turn
- Acks the hub once stdin accepts the inject (not on turn completion) to
  stay inside the 30s RPC window; the launcher stays busy until the
  concurrent prompt settles so handlers are not swapped mid-inject
- steeringActive agent state mirrors the active-turn window; abort and
  cleanup reset it and invalidate pending steers
- Legacy stream-json Cursor sessions register a steer handler that
  reports unsupported

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* feat(shared): advertise cursor in the steer gate now that its handler lands

Cursor ACP sessions pass the web and hub steer gates; legacy stream-json
cursor sessions stay excluded.

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(cursor): consume the row at dispatch; keep waiters for prompt gating

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  the concurrent session/prompt; completion is background-only. A
  dispatched steer is never restored (no duplicate via the next prompt)
- softSteerWaiters are registered before awaiting dispatch so the main
  loop's finally cannot start the next prompt mid-inject; they still gate
  prompt handover on completion
- tests updated: post-dispatch ACP rejection keeps the row consumed

* fix(cursor,hub): never hang teardown on unresolved soft steer; align diagnostics

- Prompt-finally waits for soft-steer completion only when not exiting;
  the outer finally no longer waits at all — cleanup() disconnects the
  ACP transport, which rejects pending requests and settles the waiters
- syncEngine gate diagnostics and JSDoc name all supported flavors
  (Pi, Codex, Cursor ACP)
- regression test: Switch with an unresolved soft-steer completion still
  reaches teardown

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(cursor): abort drops soft-steer waiters so the next prompt never blocks

- Ordinary Abort (shouldExit false) now clears softSteerWaiters: the
  prompt finally cannot wait forever on a soft steer whose completion is
  unbounded; the ACP cancel rejects in-flight requests, and cleanup()
  settles leftovers on session end
- regression test: unresolved soft-steer completion after Abort no longer
  blocks the next prompt

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(cursor,acp): abort force-settles soft-steer bookkeeping

- AcpSdkBackend.abortSoftSteers() drops the concurrent-prompt counter and
  notifies response-complete so the next turn's waitForResponseComplete()
  cannot block on a soft steer that will never settle after abort
- handleAbort calls it before clearing the waiters; the main prompt's own
  finishPromptRequest stays guarded by Math.max(0, ...)
- unit tests cover counter release and no-op when idle

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* test(acp): match finishPromptRequest epoch signature in whitebox test

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(cursor): abort releases an in-progress soft-steer wait

- The prompt-finally wait races Promise.allSettled against the abort
  signal: an Abort that clears the waiters now also releases a wait that
  already started, so the launcher always reaches the next queued prompt

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(cursor,acp): commit on ACP acceptance; distinguish transport failures

- AcpStdioTransport marks transport-level failures (timeout, closed,
  stdin write) as indeterminate; explicit JSON-RPC error responses are not
- The steer handler commits + consumes on completion (ACP acceptance) and
  restores the row on an explicit rejection; an indeterminate transport
  failure keeps the row reserved so a delivered instruction is never
  re-sent, and the ACK reaches the hub even when abort reset the queue
- launcher/transport tests updated for the three outcomes

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* test(cursor): match tri-state cancel contract for dispatching steers

* fix(cursor): drop duplicate promptInFlight declaration after upstream merge

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(cursor,acp,web): indeterminate close marks, steer gating precision

- AcpStdioTransport.rejectAllPending marks close/protocol failures
  indeterminate, so an accepted-but-close-interrupted soft steer restores
  nothing (no duplicate delivery)
- the abort race in the soft-steer wait removes its listener in finally
  (no accumulation across repeated waits)
- SessionChat gates the Steer button on agentState.steeringActive for
  codex/cursor instead of the queued-grace thinking flag, so Steer is not
  exposed before the launcher can accept it
- codex pre-dispatch abort never enters reconciliation (merged from #1606)

* fix(steer): persist indeterminate outcomes without replay

* fix(cursor): hold ambiguous steers for explicit resolution

* fix(steer): make ambiguous delivery restart-safe

* fix(cursor): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(cursor): reject steers when prompt generation changes

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(cursor): preserve soft-steer reservations across abort

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(cursor): hold ambiguous dispatch failures

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(cursor): bound ACP dispatch acknowledgements

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(cursor): preserve state when dispatch ACK is uncertain

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(cursor): distinguish held cancel from removal

* fix(store): reserve schema v25 for steer delivery state

* fix(cursor): suppress late ACP updates after abort

* fix(steer): keep held cancel state and notify requeue

* fix(cursor): isolate late updates after abort

* fix(steer): release explicitly cancelled unknown reservations

* test(cursor): cover explicit held cancellation

* fix(codex): reject cancelled reservations before native steer

* fix(cursor): reject cancelled reservations before ACP steer

* fix(codex): make reservation restore atomic with state

* fix(cursor): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(cursor): hard-stop abandoned writes and update native queue state

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(cursor): isolate aborts and add native retry resolution

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(native): reconcile retry responses

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): resync busy cancel outcomes

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(cursor): hold restore failures for explicit resolution

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(cursor): drain foreground prompt after soft-steer abort

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones

* fix(cursor): drain soft steers before handler replacement

* fix(cursor): preserve buffered output on abort
2026-08-20 08:53:41 +08:00
weishu 1f5602ad6a Release version 0.29.0 2026-08-19 21:34:51 +08:00
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00
KorenKritaandGitHub 2fbd98dfe4 fix(pi): offer model-accurate thinking levels in the create-session form (#1626)
* fix(web): filter Pi effort options by model thinkingLevelMap in new session form

The create-session EffortField called getPiThinkingLevelOptions without the
selected model's thinkingLevelMap, so the Pi effort select always showed the
static off..high list: levels the model marks unsupported stayed visible and
xhigh/max never appeared even for models that opt in. Pass the map through
piSelectedModel (mirrors HappyComposer), and extend the stale-effort reset
effect so a level unsupported by the newly selected model falls back to auto.

* feat(cli): probe machine Pi models over RPC to carry thinkingLevelMap

The machine-level Pi model probe parsed the `pi --list-models` text table,
which only exposes provider/model/thinking-yes-no — thinkingLevelMap (and
name/contextWindow) never reached the create-session form, so model-accurate
thinking levels could not render there (xhigh/max are map-opt-in and were
permanently hidden; see the companion web commit).

Replace the table probe with a short-lived `pi --mode rpc` child
(--no-session --no-extensions --no-skills --no-prompt-templates --no-tools)
that issues get_available_models and reuses the session path's parsePiModels
schema, so machine and session catalogs share one wire contract. The
lightweight probe also measures faster than the table probe (~0.8s vs
~1.6-2.4s) and drops the whitespace-table parsing entirely. Cache/inflight
dedupe/timeout structure is unchanged.

* fix: address review findings on the Pi model probe and effort reset

Three review findings on the RPC probe / effort-map change:

- (high) Probe teardown killed only the direct child PID: with shell:true on
  Windows that is the shell, orphaning the interactive pi RPC process on
  every successful probe and on timeout; finish() also resolved before the
  process was confirmed gone. Use killProcessByChildProcess (taskkill /T on
  Windows, Unix tree kill, SIGKILL escalation) and settle only after the
  tree teardown completes.
- (medium) An explicit get_available_models success:false response was
  discarded, so Pi's own error text was lost and the interactive child hung
  until the generic 15s timeout. parsePiModelsProbeLine now returns a
  three-way result (unrelated/models/error), also validating the response
  id, and the probe rejects immediately with Pi's error.
- (medium) Switching an xhigh/max-capable model back to Default left the
  now-hidden effort in state and submitted it while the select visually fell
  back to auto. The reset effect now also covers model === 'auto' (undefined
  map) while still not resetting mid-resolve for a concrete model.

Tests: probe-line failure/foreign-id cases; NewSession restore-to-Default
reset and map-opt-in retention (the latter guards the former against a
false green from the restore path).

* fix(web): type the Pi model test mock as PiModelSummary

The inline mock element type omitted thinkingLevelMap, so the new
capability-driven tests failed typecheck (TS2353) and the required test job
stopped before the unit tests ran. Use the shared PiModelSummary type so the
mock cannot drift from the wire contract again.

* fix(web): reconcile hidden Pi effort when model discovery fails

A restored explicit Pi model never resolves when the machine catalog request
fails, so piSelectedModel stays null for good. The reset effect required a
resolved model, so it skipped reconciliation, while EffortField rendered with
an undefined map and hid xhigh/max. Creation is only gated on the loading
state, not on the error, so handleCreate could still forward the stale hidden
level for Pi to reject or clamp.

Treat a failed catalog as a settled selection (alongside Default and a
resolved model) and reconcile against the undefined map; keep skipping the
reset while a concrete model is still resolving without an error, so a
restored xhigh/max survives until the map can prove it valid.

Test asserts the spawn payload carries no effort after a failed catalog with
a restored xhigh; verified it fails when the error branch is reverted.

* fix(cli): honor a failed probe process-tree teardown

killProcessByChildProcess reports survivors by resolving false, but the probe
discarded that result and settled anyway. The Windows graceful path is
taskkill /T without /F and escalates nothing on its own, so a probe child that
refuses the signal would be reported as cleaned up and the catalog cached,
letting interactive Pi processes accumulate across refreshes.

Escalate to the forced teardown when the graceful one reports survivors, and
reject (caching nothing, so the next call re-probes) when even that fails.
Skip the check when the child has no pid: nothing can leak, and replacing a
spawn ENOENT with a teardown error would only obscure the real failure.

Adds probe lifecycle tests (spawn + teardown helper mocked) for the escalation
path, the reject-and-do-not-cache path, the no-pid spawn-failure path, and the
timeout path; verified they fail when the escalation is reverted.

* fix(cli): keep extensions enabled for the Pi model probe

--no-extensions silently dropped providers contributed through
pi.registerProvider, which the old `pi --list-models` probe did list. Users
with such an extension would have lost those models in the create-session
form only.

Verified with a project-local .pi/extensions provider in the same cwd: the
probe with --no-extensions returned 29 models, while both the default run and
the old table probe returned 30 including the extension's model (and its
thinkingLevelMap). Dropping the flag restores parity; the extension provider
now comes back with its map intact.

Discovery that cannot contribute models stays disabled (--no-session,
--no-skills, --no-prompt-templates, --no-tools). Cost: the probe now measures
~1.4-2.0s instead of ~0.65s, still at or below the old table probe (~1.6-2.4s)
and fronted by the existing 60s cache.

* fix(cli): probe Pi models from the home directory, not the runner cwd

Under launchd/systemd the runner cwd is `/`, and starting Pi there is
pathological: project discovery plus extensions that scan from the working
directory walk the whole filesystem root. With extensions enabled (required so
pi.registerProvider models still surface) the RPC probe took 16.8s at cwd=/,
past the 15s timeout, so machine-level model discovery failed 100% of the time
and the create-session form showed only 'Pi model discovery timed out'.

Measured on macOS with 9 global extensions:
  cwd=/      extensions on   16.8s  -> timeout
  cwd=/      extensions off   0.6s  -> would lose extension providers
  cwd=$HOME  extensions on    1.4s  -> 29 models, complete

The catalog is machine-scoped and does not depend on cwd, so probing from the
home directory is both safe and representative. Verified from cwd=/ under the
runner's exact launchd environment: 1.3-1.7s, 29 models.

The old `pi --list-models` probe was immune because it never initialized a
session; this regression arrived with the RPC probe and was missed because
every earlier verification ran from a project directory.

Tests assert the spawn cwd is homedir() and that --no-extensions stays absent;
verified the cwd test fails when the option is removed.

* fix(cli): keep the probe in the runner cwd, fall back to home only at a root

Forcing every probe to the home directory fixed the launchd timeout but broke
project-local discovery: a runner started inside a project stopped seeing that
project's .pi/extensions providers, which the replaced --list-models probe did
surface. Verified with a project-local provider: present from the project dir
(30 models), absent from home (29).

Use the runner cwd normally and fall back to home only when cwd is a
filesystem root -- the launchd/systemd case where Pi startup walks the whole
tree (16.8s, past PROBE_TIMEOUT_MS) and where there is no project to lose
anyway. process.cwd() can also throw for a deleted directory, so that falls
back to home too.

Verified under the runner's launchd environment: from / -> home, 1.6s, 29
models; from a project dir -> that dir, 1.4s, 30 models including the
project-local provider.

Tests cover all three branches and were checked in both directions: pinning to
home fails the project-cwd test, pinning to cwd fails the root test.
2026-08-19 09:20:40 +08:00
SSU-WEI HUANGandGitHub 3a436240c3 fix(pi): refresh context after compaction (#1634) 2026-08-19 09:19:26 +08:00
weishu 12847ddd0f Release version 0.28.0 2026-08-17 10:03:11 +08:00
AnanovoandGitHub d644d4f30f feat(cli): follow conversation language in status prompt (#1584) 2026-08-16 23:06:06 +08:00
SSU-WEI HUANGandGitHub 6c6f4b4929 feat(agy): replace fragile PTY/TUI wrapper with headless print-mode transport (#1591)
Replace the Antigravity (agy) integration — a PTY wrapping the TUI with
output-marker scraping ('? for shortcuts', 'Generating', trust dialogs,
/model picker navigation, quota-screen regex) — with a headless print-mode
transport: every user turn spawns `agy -p <msg> --conversation <uuid>
--output-format stream-json`, and NDJSON events (init/step_update/result)
map onto the existing transcript-entry channel (sendAgySessionMessage), so
hub/web rendering is unchanged. ~8.3k LOC (incl. tests) removed.

Fixes #1588. Design: docs/design/agy-headless-transport.md.

CLI:
- new cli/src/agy/headless/: agyNdjsonParser (pure functions, malformed-line
  tolerance, step conversation-id adoption), AgyPlannerAccumulator (per-step
  delta accumulation with settling retries), AgyHeadlessDriver (per-turn
  spawn/kill loop, NDJSON chunk buffering, authoritative delivery ack via
  user_input/result, interrupt + retry + shutdown lifecycle with consume/
  restore, process-tree termination, SSH agent preserved, prompt log
  redaction, per-turn model snapshot with conversation-DB fallback)
- runAgy/loop/session rewired; agy is remote-only (no PTY, no local mode,
  no local-switch action); queued batches snapshot model/effort/mode
- deleted agyPty, agyPtyLauncher, agyHookCarrier(+scope cache), agyModelKeys,
  agyQuestionKeys, agyAskQuestion, agySessionScanner, agyPermissionHandler
  (+tests); buildAgyHooksJson removed; startHookServer agy-pre-invocation
  route → 200 no-op
- runner: agy reopen/resume via generic --existing-session-id; commands/
  agy.ts defaults remote; resume rejects ACTIVE agy sessions (remote-only,
  in-flight turns cannot hand off)
- MCP stays user-managed (agy reads ~/.gemini/config/mcp_config.json and
  workspace .agents/mcp_config.json natively in headless — verified)

Hub/web:
- machines.ts drops agy→pty forcing and rejects non-remote startingMode
- NewSession drops agy startingMode='pty'; terminal toggle disappears
  automatically; RemoteModeDisplay hides the local-switch hint when absent
- docs/guide/agents.md updated: headless print mode, no PTY/hooks, MCP via
  user's own mcp_config.json

Tests: 47 parser+driver tests (fake-binary e2e, chunk-split NDJSON, delivery
ack semantics, interrupt/retry/shutdown races, model attribution, EOF
framing, malformed envelopes); full suite green (cli ~2340, hub 1093,
web 2474, shared 262). Real-binary smoke on agy 1.1.13: single turn exit 0,
--conversation resume keeps the same conversation_id.
2026-08-16 22:44:27 +08:00
SSU-WEI HUANGandGitHub a6feb6e8ba feat: unified agent configuration descriptors (session config consolidation) (#1469)
* feat(config): add agent config descriptor protocol and advertise via runner capability

Introduce shared agent configuration descriptors covering model, effort,
permission, and secondary settings per agent flavor, plus the canonical
HAPI YOLO -> native permission mode mapping. Runners advertise the
builtin descriptors through the runner-state capability so hubs and web
can render configuration without hardcoded flavor branches.

Migrate the OpenCode create-session model picker from a bespoke radio
list to the shared SelectControl combobox.

* feat(web): render create-session permission from agent config descriptor

Replace the flavor-branched Grok/Codex-family/YOLO permission block with
a descriptor-driven PermissionField. Pi now reports permission as managed
instead of silently ignoring the YOLO toggle, and YOLO-only flavors show
the native permission mode the preference maps to.

Removes the superseded GrokPermissionModeSelector and
CodexFamilyPermissionModeSelector components.

* ci: retry flaky claudeRemote 5s-timeout failure

* fix(web): persist explicit OpenCode Default selection instead of restoring a concrete model

The parent initialization effect treated every null selected model as
'uninitialized' and auto-picked a concrete advertised model, clobbering
the user's explicit Default choice (and a restored Default preference).
null now means explicit Default and is preserved; only undefined (no
choice made yet) triggers probe-based initialization. Add parent-level
regression tests for Default persistence and remembered-model restore.

* fix(web): accept undefined selected model in OpencodeModelSelector props

* feat(web+cli+hub): unify create-session model/effort fields and add Pi model/effort support

Pi's agent config descriptor now advertises model (machine) and effort
(static thinking levels) for create AND session availability:
- cli: ListPiModelsForMachine RPC runs 'pi --list-models' (cached, inflight
  deduped) and parses the provider/model table; startup model match accepts
  provider-qualified ids
- hub: GET /api/machines/:id/pi-models route + rpcGateway/syncEngine passthrough
- web: NewSession renders Pi models grouped by provider through the generic
  ModelSelector and a new descriptor-driven EffortField (replaces the
  per-flavor LaunchEffortSelector/ReasoningEffortSelector pair); launch payload
  forwards Pi model + thinking-level effort (runner already supported --model/
  --effort for pi)

* fix(web): render Pi provider groups in ModelSelector and scope Grok availability warning

- ModelSelector now renders grouped options as <optgroup> (Pi models are
  provider-grouped; identical modelIds from different providers stay distinct)
- PermissionField only receives autoPermissionModeSupported for Grok — a
  cached Grok probe result no longer leaks the Grok warning onto other agents

Addresses HAPI Bot Minor findings on #1469.

* fix(web): drop Object.groupBy from ModelSelector; revalidate restored Pi models against the catalog

- ModelSelector buckets options with a reduce instead of Object.groupBy
  (Safari < 17.4 has no polyfill — New Session would throw on those clients)
- Pi restored model/effort are cleared when the value is absent from the live
  machine catalog, and Create waits for the catalog while a non-default Pi
  choice is being validated (mirrors Codex/Grok/Copilot handling)

Addresses HAPI Bot findings on #1469.

* fix(pi+web): serialize startup model before thinking level; hide Pi launch controls during history import

- PiSession gains startupModelSettled; the startup set_thinking_level waits for
  the requested model's set_model attempt to settle first, so a level the
  default model rejects is not lost before the requested model is confirmed
  (set_model and set_thinking_level were already serialized by the runtime
  mutation lock; this pins the model-first ordering)
- Create Session hides Pi model/effort controls while a Pi history import is
  selected — the import reopens the native session as-is and would silently
  ignore launch-only model/effort values

Addresses HAPI Bot findings on #1469.

* fix(pi): settle startup-model gate when model discovery fails or returns no models

A failed or empty get_available_models response would leave the
startupModelSettled gate unresolved, stranding a requested startup effort
indefinitely. Resolve the gate on the error path and the empty-models path;
adds regression tests for both.
2026-08-16 22:43:39 +08:00
SSU-WEI HUANGandGitHub ca6a6f0939 fix(cursor): retry transient ACP connection errors (#1541)
* fix(cursor): retry transient ACP connection errors

* fix(cursor): keep retry classification conservative

* fix(cursor): avoid retrying completed tool effects

* fix(cursor): require terminal retry failures

* fix(cursor): track retry activity across extension events

* fix(cursor): honor permission abort before retry
2026-08-16 22:42:10 +08:00
AnanovoandGitHub 7909c46fff feat(search): support wildcard patterns across search fields (#1571)
* feat(search): add wildcard matching to search fields

* fix(search): harden wildcard matching and file globs

* fix(search): align file matching with shared wildcard semantics

* fix(search): bound file wildcard search in runner

* fix(search): normalize outline queries through shared matcher

* fix(web): remove duplicate markdown test context field
2026-08-16 22:41:24 +08:00
KorenKritaandGitHub bf6bae1924 fix(pi): settle autonomous agent lifecycles instead of swallowing them (#1563)
* fix(pi): settle autonomous agent lifecycles instead of swallowing them

When Pi starts an agent lifecycle on its own — a subagent completion
wake-up, scheduled work — no HAPI prompt is in flight, so the previous
prompt lifecycle has already delivered its settlement and
deliveredSettlement is still true. agent_start unconditionally set
thinking=true, but every settlement path (agent_settled delivery, the
legacy agent_end grace, the prompt-lifecycle fallback) was gated shut by
that stale flag, so the autonomous turn's completion was swallowed:
thinking stayed true forever, the FIFO pump stayed blocked
(piIsStreaming), new messages queued without ever being sent, and abort
waited on a settlement that could never arrive. The only escape was
killing the session.

Open a fresh settlement cycle when agent_start/turn_start arrives with
deliveredSettlement still true: reset deliveredSettlement,
agentEndObserved and activeAgentSettledSeen so the existing settlement
paths (direct agent_settled, legacy agent_end grace) apply to the
autonomous lifecycle. Prompt-driven lifecycles are unaffected because
beginPromptLifecycle has already reset the flag before their agent_start
arrives; mid-cycle retries are unaffected because their cycle has not
settled yet.

Evidence: hapi log 2026-08-13-18-39-10-pid-85758.log — 19:24:07
agent_start (no prompt accepted) → 19:25:40 agent_end + agent_settled
both swallowed → session stuck thinking=true for 27+ minutes until
killed.

* fix(pi): generation-scope the settled callback against autonomous lifecycle races

Review follow-up (HAPI Bot, Major): deliverSettlement() notifies
onAgentSettled only after an async conversationHistory.syncEntries().
An autonomous lifecycle can begin in that window; the stale finally
callback would then mark the new lifecycle's abort boundary as settled
before it emits agent_settled.

Capture lifecycleGeneration at settlement time and skip the
notification when it no longer matches, and advance the generation when
agent_start reopens a settlement cycle for an autonomous lifecycle so
the in-flight callback turns stale.

New regression test holds syncEntries open across the autonomous
agent_start and asserts the stale callback does not settle the new
boundary (fails without the fix).
2026-08-16 22:40:15 +08:00
SSU-WEI HUANGandGitHub 44aff5a924 fix(codex): cache codex model list to avoid spawning app-server per request (#1534)
* fix(codex): cache codex model list to avoid spawning app-server per request

listCodexModels() spawned a fresh `codex app-server` subprocess on every
request: exec a version probe, boot the app-server, validate the ChatGPT
session (token refresh over the network when needed), list models, then
kill the process. Fleet measurement showed 0.5-4.4s per call on healthy
machines and 33s (initialize timeout) on a machine with a slow OpenAI
network path, and the web refetches on every session open / dialog mount
(staleTime 30s).

Mirror the opencode model cache: cache successful non-empty lists for 5
minutes and coalesce concurrent requests into a single app-server spawn.
Failures and empty results are never cached, so a broken machine retries
on the next request.

Fixes #1533

* test(codex): cover TTL expiration of the model list cache
2026-08-16 22:39:57 +08:00
Shawn TianandGitHub 5c81ed69e2 fix(runner) preserve native instructions (#1557)
* fix(codex): preserve native instructions

Keep HAPI guidance in the developer layer so app-server retains Codex persistence rules. Align long-running goal states with the Codex 0.147 protocol.

* fix(codex): filter new goal statuses
2026-08-16 22:38:15 +08:00
SSU-WEI HUANGandGitHub 1f2fbcb142 fix(web): map codex-enveloped compact-summary to the chat block (#1582)
* fix(web): map codex-enveloped compact-summary to the chat block

The merged #1570 renders Pi compaction summaries as a dedicated chat
block via the event envelope. A compact-summary arriving in the codex
payload envelope (older import paths, future producers) would still be
silently dropped by the codex-content filter; map it to the same
agent-event the live pi wrapper emits, mirroring the existing
context_compacted handling. Adds a normalizeAgent regression test and
tightens the /compact thinking-state test to assert the last keepAlive
flips true (RPC outstanding) then false (settled).

Verified: bun typecheck clean; web normalizeAgent 12/12, cli runPi
57/57 in an isolated TMPDIR.

* fix(web): remove duplicate showSessionSummaryInChat in markdown-a fixture

Regression from #1530: the fixture object literal sets the key twice,
which breaks `bun typecheck` on upstream/main (Test workflow failing on
push). One-line cleanup; no behavior change.

Verified: bun typecheck exit 0 across cli/web/hub.
2026-08-16 22:35:31 +08:00
901f17d0ca feat(pi): support Pi slash commands from HAPI web (compact/session/model/help) (#1570)
* feat(pi): support Pi slash commands from HAPI web (compact/session/model/help)

Pi runs as 'pi --mode rpc' over piped stdio, so TUI slash commands typed in
web chat previously fell through to the LLM as plain text and silently did
nothing (notably /compact).

- shared: add Pi builtin slash command list (help/compact/session/model) so
  the web / menu exposes them; web test updated to match
- cli: intercept Pi builtin commands in runPi's user-message path
  * /compact [instructions] -> Pi compact RPC (120s timeout, works while
    streaming; summary + token delta reported back as chat messages)
  * /session -> get_session_stats formatted stats
  * /model [modelId] -> list/switch via set_model
  * /help -> supported-commands list
  * other Pi TUI builtins (/tree, /export, /reload, ...) -> explicit
    terminal-only notice instead of silent LLM pass-through
  * unknown slash text still passes through (extension commands, skills,
    templates keep working)
- gate the prompt pump with piCompactInFlight so queued prompts are not
  rejected by Pi mid-compaction; buffer commands until ready like prompts
- ListSlashCommands RPC merges HAPI builtins with Pi extension commands
- tests: parser unit tests + runPi integration tests (compact execution,
  streaming steer interception, failure reporting, FIFO blocking, model
  switch, unsupported commands, slash list merge)
- docs: document Pi slash command support in docs/guide/agents.md

* fix(pi): address review findings on slash command lifecycle

- compact timeout: fail the session (indeterminate outcome, runtime lease
  poisoned) instead of reopening the prompt FIFO into a possibly-compacting
  Pi; pump only when cleanup has not been initiated
- special commands: release the cancellation reservation before executing so
  a cancel landing mid-command is not acknowledged (hub would delete the
  queued row while the command still runs)
- tests: drop the duplicated slash-command describe block; add focused tests
  for compaction timeout with a queued prompt and cancellation during an
  in-flight special command

* fix(pi): route slash commands through the prompt FIFO and reject ambiguous models

- slash commands now share the prompt FIFO with ordinary messages: a
  /compact or /model typed after a queued prompt dispatches only after it
  (and after the active turn settles), instead of jumping the queue from
  the preparation chain
- the pump dispatches special entries out-of-band while piSpecialCommandInFlight
  keeps the FIFO blocked; steer promotion refuses slash commands
- /model <id> prefers an exact provider/modelId match and reports bare IDs
  shared by multiple providers as ambiguous instead of picking the first
- tests: FIFO ordering (queued prompt before /compact), steer-delivered
  /compact queued until settle, ambiguous/qualified model selection

* fix(pi): keep /compact interruptible, honor extension precedence, require token boundary

- head-of-line /compact dispatches even while Pi is streaming (Pi's
  compact() aborts the active generation itself); every other queued item
  still waits for the stream to settle, preserving FIFO order
- discovered extension commands / prompt templates override same-name
  builtins at message time, matching the slash-list merge precedence
- parsePiSpecialCommand requires a command-token boundary, so path-like
  text such as /compact.md or /model/config stays an ordinary prompt
- tests: interrupt rule, extension collision, reserved-name path prefixes,
  non-compact commands waiting for stream settle

* fix(pi): honor cancellation acknowledged during slash-command discovery

A cancel arriving while the chain awaits get_commands (cold cache) was
acknowledged via the preparing reservation but never re-checked, so a
canceled /compact could still execute. Re-check the cancellation marker
after discovery and drop the message before dispatch.

* fix(pi): qualify /model selectors and report failed slash RPCs once

- /model lists provider-qualified selectors (openai/gpt-5.2) so duplicate
  bare IDs remain usable and copy-pasteable; current model is qualified too
- compact/set_model failures are owned by the awaited slash/config handlers:
  the common response handler no longer emits the raw Pi error a second time
- tests: qualified listing with duplicate providers, single-message failure
  reporting for rejected /compact and /model

* fix(pi): consume slash-command queue row at dispatch

Special commands (/compact, /session, /model, /help) are executed by HAPI
itself and never delivered to Pi as prompts. Consumption was deferred until
the command finished, so a /compact run — an LLM summarization pass that can
take minutes — left the row stuck in the web queued bar for its whole
duration, then surfaced as a sent message. Consume the row the moment
dispatch starts; failures still surface via the explicit event message.

* fix(pi): guard special-command dispatch against unexpected rejections

* ci: retry Codex PR Review after infra failure (proxy 503)

* fix(pi): keep session queued-thinking grace during /compact dispatch

The queued-thinking grace is session-scoped, so clearing it while
acknowledging a dispatch-time /compact row also drops the grace for any
prompt queued behind it. /compact keeps running for minutes without
toggling Pi thinking state, which would leave the web session looking idle
while compaction and the following prompt are still pending. Only the
fast, synchronous commands (/session, /model, /help) clear the grace.

* fix(pi): render compaction summary as a dedicated chat block

The manual /compact RPC result was reported as two plain message
events ("📦 Compaction completed (tokens: …)" + "📦 Compaction
summary: …"), which the web chat renders as tiny centered status
lines — unusable for a real summary payload. Emit a structured
compact-summary event instead (summary + token delta) and render
it as an independent block: header with the delta and the summary
markdown in a scrollable panel.

Also emit the same structured event when importing Pi session
files (compaction entries), and queue the event lossless like
other user-visible messages so a disconnect cannot drop it.

Verified: bun typecheck clean; bun run test exit 0 (cli 2481
passed, web 2451 passed, hub/shared clean); runPi/loop/apiSession/
piSessions/presentation suites green.

* fix(pi): address HAPI Bot findings on compact dispatch and import

- Track compaction as thinking for its whole duration: /compact runs for
  minutes without a Pi streaming event, so the 15s queued-thinking grace
  alone left the web session looking idle while compaction and any queued
  prompts were still pending (updateThinkingState around the compact RPC).
- Imported Pi compaction summaries must use the event envelope
  (content.type: 'event') like the live wrapper's compact RPC result; the
  codex payload envelope is dropped by the web normalizer. Extend
  CodexImportedMessageSchema with the event variant.

* fix(pi): /model retries discovery when the model cache is empty

Startup model discovery can be late or fail once; using only the cached
catalog made /model report valid models as unknown. getPiModels() falls
back to the get_available_models RPC on an empty cache, used for both
listing and switching.

* fix(pi): interrupt in-flight /compact on Abort; surface startup model rejection

- The Abort action no longer waits on the runtime-mutation lease when a
  manual /compact is in flight (compaction can hold it for up to 120s,
  blowing the 25s abort deadline and failing closed). It sends the abort
  RPC directly so Pi cancels its compaction AbortController; the compact
  RPC's 'Compaction cancelled' error is not double-reported as a failure
  since Pi already emits the compaction_end(aborted) lifecycle event.
- A rejected detached startup set_model now emits a visible ⚠️ event into
  chat instead of only a debug log, restoring the pre-existing behavior.

* fix(pi): close the Abort race when /compact is queued on the mutation lock

Abort previously assumed an in-flight /compact always had its RPC issued;
the command is marked active at queue dispatch, but the compact RPC is sent
only after the runtime-mutation lock is acquired. An Abort landing in that
gap acknowledged success while the compact RPC still ran afterwards.
Track the compact's rpcStarted/cancelled state: Abort cancels a not-yet-
started compact in place (the queued callback skips it), and interrupts a
started one via the abort RPC as before.

* fix(pi): persist provider-qualified selection after /model switch

The success path updated currentModel/currentProvider and keepalive with a
bare model ID, leaving metadata.piSelectedModel on the previous provider.
The web picker prefers that metadata for selection, context-window
resolution, and effort options, so a switch like openai/gpt-5.2 ->
azure/gpt-5.2 was invisible. Persist piSelectedModel with the full
provider/modelId pair on every confirmed switch.

* fix(pi): retire pending extension UI requests when /compact interrupts a turn

The streaming-interrupt path sent the compact RPC without cancelling
pending extension UI requests first, unlike the Abort path. Editor
requests have no timeout, so the web could stay stuck on a stale
input/permission card and a later answer could be routed to the aborted
turn. Cancel all pending requests (with a response) before compacting.

* fix(pi): fail closed when the direct compact-abort RPC times out

The in-flight /compact abort branch awaited the abort RPC without the
ordinary Abort path's timeout handling: an unanswered abort left the
compaction outcome indeterminate (the compact RPC keeps the mutation
lease for up to 120s) while the wrapper still looked live. Fail the
session on PiRpcTimeoutError, mirroring the standard abort fail-closed
path.

---------

Co-authored-by: swear01 <swear01@users.noreply.github.com>
2026-08-15 11:21:45 +08:00
AnanovoandGitHub effc505a07 fix(mcp): clarify display_image user-output semantics (#1568) 2026-08-15 11:15:27 +08:00
df1a56e1db fix(cursor): exclusive agent spawn lease for list-models vs ACP (#1529)
* fix(cursor): exclusive agent spawn lease for list-models vs ACP (#1520)

Add a proper-lockfile spawn lease beside agent-acp-active so model probes
and ACP transport acquire mutual exclusion atomically before spawning
agent children, closing the post-#1518 check-then-act overlap window.

Fixes #1520

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(web): fix markdown-a test HappyChatContext mock for typecheck

Adds showSessionSummaryInChat to chatContext() so CI typecheck passes on
the PR branch (pre-existing main breakage unrelated to #1520).

Co-authored-by: Cursor <cursoragent@cursor.com>

* Revert "chore(web): fix markdown-a test HappyChatContext mock for typecheck"

This reverts commit 4973f321a31fda772e8990ea6cda20517ca906aa.

* fix(cursor): scope spawn lease to agent spawn window only (#1520)

Hold agent-cli.spawn only around spawn('agent') in AcpStdioTransport, not
for the full ACP session. Restores N concurrent cursor sessions per host;
list-models probe lease unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): tighten spawn lease lifecycle for babysit (#1520)

Acquire spawn lease before ACP marker publish; unregister on spawn failure.
Hold list-models probe lease until child exit on timeout. Add missing
showSessionSummaryInChat to markdown-a test mock (unblocks CI typecheck).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): re-check ACP marker after spawn lease acquire (#1520)

Close check-then-act window where ACP could publish its marker between
the inactive guard read and list-models spawn. Add regression test.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(cursor): fix ACP-after-acquire mock call order (#1520)

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): async spawn lease + force-kill probe timeout (#1520)

Add acquireAgentCliSpawnLease (setTimeout yields) for ACP create path;
AcpStdioTransport.create() async factory. Probe timeout uses
killProcessByChildProcess(force) while holding lease until child exit.

Addresses Bugbot Majors: session-lifetime mutex (fe07b708d), probe lease
release, post-acquire re-check (137baa779), sync loop starvation, timeout
escalation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(acp): coalesce concurrent initialize() transport spawns (#1520)

Await shared bootstrapTransport promise so overlapping initialize calls
do not spawn duplicate ACP children while create() is in flight.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 18:42:43 +01:00
weishu fea42212ea Release version 0.27.3 2026-08-12 10:26:06 +08:00
c69a88afae fix(cursor): close ACP list-models race and false exit 143 window (#1518)
* fix(cursor): close ACP list-models race and false exit 143 window

Register the agent-acp-active guard before spawn, hold it until stdio
close (not bare exit), record the ACP child PID, and align lock/cache
home with resolveHapiHomeDir so runner and session children agree.
Richer exit attribution distinguishes live-PID transport disruption
from confirmed child death. Fixes residual #1472 after #835.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(cursor): isolate ACP guard teardown from ~/.hapi

Reset afterEach under the temp HAPI_HOME only, and restore the
isolated home before teardown in the unset-HAPI_HOME case, so tests
cannot wipe a live agent-acp-active lock. Use distinct child PIDs in
registration tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): publish ACP lock pid before count

Fail-closed reservation order: write pids/<hostPid> before count so
concurrent reconcile cannot treat a mid-register lock as stale and
clear it for list-models. Keep a short mtime grace only when pids/ is
missing (mkdir gap). Regression covers mid-publish readers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): keep empty pids/ ACP reservation fail-closed

Between mkdir(pids) and the host pid writeFile, reconcile could see
liveCount=0 and clear the lock. Keep that window when count is still
absent and the lock is fresh; re-scan for pids published mid-reconcile.
Regression hooks the mkdir/write gap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): keep ACP lock across last-unregister publish race

Write a short-lived registering marker before pids/count so empty pids
with leftover count cannot erase a concurrent mid-addLockPid reservation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): use per-pid ACP registering markers

Crash/reboot must not pin list-models forever on a bare registering
file; prune dead owners and only keep live registrar PIDs.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 09:55:26 +08:00
7c4b3ab17b fix(cursor): stop session-list spinner flicker from ACP state_update (#1503)
* fix(cursor): stop session-list spinner flicker from ACP state_update

#1487 mapped Cursor ACP state_update running/idle onto hub thinking.
Cursor chatters those states while HAPI is queue-idle, so keepAlive flipped
thinking every ~1-2s and the session list spinner danced. Ignore
state_update for thinking; only bump on real activity, clear via prompt
finally/abort.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): clear thinking on ACP idle without bumping on running

Refine #1502 hotfix: ignore state_update running/requires_action (Cursor
chatter caused spinner flicker) but still clear on idle so mid-idle harness
wakes do not stick thinking=true forever.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): stop ACP background updates from flickering thinking

Ignore tool/content session updates for hub thinking (ACP allows them while
idle). Drive thinking from state_update only, debounce running 750ms, and skip
idle clears during an in-flight HAPI prompt so #1502 residual flicker dies.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 09:54:37 +08:00
SSU-WEI HUANGandGitHub 291e7bc40b fix(test): stop runner integration suite from leaking detached process trees (#1515) (#1521)
* fix(test): stop runner integration suite from leaking detached process trees (#1515)

The default CLI test run included runner.integration.test.ts, which spawns
real detached runner/session process trees. A failing, timed-out, or
interrupted test (or a plain runner stop) left those trees alive under
PID 1 — on the Mac this accumulated ~600 Node/Bun/agent processes and
several GiB of RSS over repeated runs.

Test harness changes only; production runner session-preservation
semantics are untouched:

- Exclude runner.integration.test.ts from the default parallel unit-test
  suite; move it into a dedicated serial integration project
  (vitest.integration.config.ts, 'bun run test:integration'). The
  20-session stress test is opt-in via HAPI_RUN_STRESS_TESTS=true.
- Add a test-owned process/session registry (processRegistry.ts): every
  runner, runner-spawned session, and terminal-style child is registered
  immediately after spawn; afterEach/afterAll run two-stage cleanup
  (logical stopRunnerSession first, then bounded process-tree kill),
  followed by a marker sweep for agent grandchildren reparented to PID 1.
- Add a per-run HAPI_TEST_MARKER env stamp + identity/secret env
  neutralization for test children (integrationEnv.ts) so outer HAPI/pi
  session variables never leak into test processes and the final audit
  can recognize test-owned processes by env alone.
- Final suite audit in globalSetup teardown: reap anything still
  carrying the run marker and fail with PID/command diagnostics if
  anything cannot be reaped, before removing the temp home.
- Regression coverage: a deliberately failing test registers a detached
  child and the follow-up audit must find zero test-owned processes.
- CI: replace the dead .env.integration-test step with a dedicated
  integration job running the serial project.

* refactor(test): drop unused killByChildProcess import and child field from registry

* chore(test): raise integration hookTimeout to 60s for slow teardown hosts

* fix(test): fail loudly when the process-table audit cannot scan; assert regression child death

Bot review #1521 findings:
- A failed `ps` scan (unsupported flags, buffer exhaustion, permissions)
  previously returned [] and silently disabled both teardown audit layers.
  It now throws; globalSetup teardown catches the scan error into the
  audit error (temp home is still removed) so the run fails visibly.
- The regression audit test cleaned the leak with the reaper before
  asserting, and force-killed the fresh marked runner. The failing
  test's direct child PID is now asserted dead in afterEach right after
  registry cleanup (before the marker sweep), and the audit test stops
  its own runner gracefully before reaping.

* fix(test): bound the logical cleanup phase so a hung runner cannot stall the hook

Bot review #1521: stopRunnerSession carries the worker's 60s HTTP timeout
(setup.ts raises HAPI_RUNNER_HTTP_TIMEOUT for the stress test), and the
integration hook timeout is also 60s — N sequential stops could exhaust
the hook budget before the process-tree fallback and marker sweep ran,
recreating the very leak this change prevents.

Logical shutdown is now parallel (Promise.allSettled over all tracked
sessions) and the whole phase (stops + PID resolution) races against a
15s budget, so stage-2 tree-kill and the marker sweep always get their
share of the hook window.

* fix(test): bound graceful runner stop in hooks; keep credentials out of audit diagnostics

Bot review #1521 (follow-up):
- stopRunner()'s HTTP stop can burn the worker-wide 60s timeout on a
  hung-but-live runner, starving the marker sweep within the hook budget.
  afterEach/afterAll now race the graceful stop against a 10s bound; a
  runner that does not stop in time is force-reaped by the sweep (it
  carries the run marker) and the next beforeEach's alive-PID guard
  ignores any stale state file.
- The env-bearing ps scan (ps eww) was also used for diagnostics, so the
  first 500 chars of a short-command process could print inherited
  credentials (CLI_API_TOKEN etc.) into teardown error logs. The scan now
  only identifies marked PIDs; command lines are fetched separately
  without 'e', falling back to '(command unavailable)' instead of the
  env dump.

* fix(test): reap runner model-probe orphans before the zero-survivor inspection

Bot review #1521 (Minor): inspect-before-reap. Applying it exposed a real
race: each test's runner legitimately spawns marker-carrying children at
startup (agent acp + agent --list-models model-catalog probes). Stopping
the runner orphans them (ppid 1) with the run marker, so the audit test's
OWN runner polluted the pure inspection with fresh probes spawned after
the failing test's sweep window.

- reapTestOwnedProcesses now re-kills every re-scan iteration instead of
  killing once and only re-scanning, so a process that survived its first
  SIGKILL (mid-exec) or spawned mid-kill is not given a free pass.
- The regression audit test stops its runner, reaps (clearing its own
  legitimate orphan probes), then inspects: anything still marked is a
  genuine survivor the bounded reaper could not remove and fails the
  suite. Killable leaks from the failing test are already asserted dead
  in afterEach before the sweep runs.

* fix(test): strictly bound the marker reaper; make per-test sweep unconditional and verified

Bot review #1521 (follow-up):
- The 10s reap deadline did not bound the awaited per-tree kills: each
  killProcessTreeByPid can wait up to 2s per PID, so several stuck
  processes could still exceed the 60s hook budget. Every process in a
  test-owned tree carries the marker (env is inherited), so tree-walking
  is unnecessary: the reaper now SIGKILLs every marked PID found by each
  scan, fire-and-forget, and re-scans every 250ms — the deadline strictly
  bounds the function.
- The per-test sweep was skipped when the direct-child assertion failed
  first, and its survivors were ignored. afterEach now snapshots the
  regression-child state BEFORE the unconditional sweep, then verifies
  both the registry result and the sweep leftovers.

* fix(test): replace it.fails regression with a direct assertion test

Bot review #1521 (Minor): Vitest applies the it.fails expected-failure
inversion after afterEach, so a broken registry assertion inside the hook
would be masked as an expected failure, and the marker sweep would erase
the evidence before the follow-up audit ran.

The regression is now a normal test that registers a detached child at
spawn time, deliberately performs NO per-test teardown, runs only the
spawn-time registered cleanup, and asserts the child PID is dead. The
afterEach no longer carries the registry-leak assertion (moved into the
test body where it cannot be inverted); the per-test sweep assertion and
the final audit test are unchanged.

* fix(test): bound registry stage-2 tree-kills; require live regression fixture

Bot review #1521 (follow-up):
- Stage-2 killProcessTreeByPid awaits per descendant serially and can
  consume the whole 60s hook for a large/stuck tree. Signals are all
  delivered synchronously (children first) before any waiting, so racing
  the awaits against a 5s budget bounds the phase without skipping any
  kill; waitForAllDead still verifies the outcome.
- The regression test could pass vacuously if its fixture exited during
  the startup delay (the registry exit listener would remove it before
  cleanup). It now asserts the child is alive before running cleanup.

* fix(test): kill registered roots with bare synchronous SIGKILL, no pgrep walk

Bot review #1521 (follow-up): racing the mapped killProcessTreeByPid
calls against a timer does not bound the phase — evaluating the map
invokes each call immediately, and each runs the recursive synchronous
pgrep walk before its first await, which can consume the hook before the
timer, runner stop, or marker sweep run.

Stage-2 now SIGKILLs registered roots directly (fire-and-forget, no
tree walk, no per-PID waits) and waits a bounded 5s for death.
Descendants are reaped by the unconditional marker sweep immediately
afterward — every descendant inherits the run marker, so tree-walking is
unnecessary.

* fix(test): drop duplicate process-death wait in registry cleanup

Bot review #1521 (Minor): the duplicated waitForAllDead delayed the
authoritative marker sweep by another 5s under the exact stuck-process
condition the harness must handle. Keep the single bounded wait; the
afterEach marker sweep remains the guarantee.
2026-08-12 09:29:59 +08:00
11964b4b0d fix(a2a): stamp causing inbound on work_ad notify ingest (#1510)
* fix(a2a): stamp causing inbound on work_ad notify ingest

Resolve the turn cause from the full session messages table at insert so
consumers do not guess from a truncated transcript. Stamp causeMessageId,
causeText, causeKind, and related_event_id, and link follows to the
previous work_ad.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): ignore scheduled and client-posted work_ads in cause chain

Future-scheduled inbounds are not this turn's cause. Only AGENT_NOTIFY_SUMMARY
rows chain related_event_id / sticky cause, so HTTP-posted work_ads cannot
steer hub-derived attribution.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): skip transcript echoes and reserve notify provenance

Claude jsonl echoes remote prompts as extra role=user CLI rows; mark them
and ignore them when choosing a work_ad cause. HTTP event writes cannot
claim AGENT_NOTIFY_SUMMARY provenance.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): stamp transcript echo only for hub-delivered prompts

Local Claude TTY prompts share isExternalUserMessage; matching pending
web/telegram text keeps those rows as work_ad causes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): match Claude queue text and bound cause message scan

Register pending transcript echoes at the formatted queue boundary
(and drop them on cancel). Later work_ad notifies scan after causeSeq
instead of decoding the full session transcript.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): note delivered Claude batch and skip uninvoked causes

Register transcript-echo text at SDK delivery after batch join and
skill expansion. Keep the marker when cancel misses the queue.
Cause candidates require invokedAt so queued localId rows wait.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): note parked Claude batches and chain work_ads by insert order

Stamp echo markers on the pending mode-switch delivery path. Advance
causeSeq past every invoked inbound in the same batch. List session
work_ads by rowid so backdated transcript timestamps cannot rewind the chain.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): pick latest invoked inbound after history hydration

Fork/merge copies leave old user rows without prior work_ads; choose
the newest invoked inbound before the assistant. Echo markers now keep
batch localIds so cancel can drop a restored-then-cancelled prompt.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): drop echo markers when a Claude batch is abandoned

The three-failure launcher cap discarded the in-flight prompt without
clearing pending transcript-echo text, so a later identical local
prompt could be misclassified.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): replace echo markers when a restored prompt is rebatched

Recoverable launch failure keeps the original marker; a later same-mode
prompt joins into new delivered text. Drop overlapping localId markers
so a later identical local prompt is not stamped isTranscriptEcho.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): keep notify cause chain across session merge

Re-key AGENT_NOTIFY_SUMMARY work_ads onto the surviving session and
bound later scans by causeCursorMessageId so merge seq-shift cannot
revive an already-consumed batch inbound as the sticky cause.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): keep notify rows on live source during history merge

mergeSessionHistory leaves the source socket alive. Only re-key
AGENT_NOTIFY_SUMMARY work_ads when mergeSessions deletes that id.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(a2a): drop id-less echo markers on rebatch and launch drop

API callers may omit localId. Replace the nameless in-flight marker
when a new id-less delivery is noted, and discard by delivered text
when the three-failure path abandons the batch. Do not mint hub
localIds (that would change invokedAt ack semantics).

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 18:08:02 +01:00
SSU-WEI HUANGandGitHub e6b9fd68e6 feat(pi): queue mid-turn messages by default; steer only via explicit per-message Steer button (#1480)
* feat(pi): queue mid-turn messages by default; steer only via explicit per-message Steer button

Pi (PyAgent) was the only flavor whose ordinary composer submission while
streaming bypassed the queue: the web resolved it to deliveryMode 'steer'
and the CLI dispatched a native steer into the running turn immediately,
with no waiting state. This makes Pi match Codex/Claude behavior (issue
#1466): mid-turn messages wait in the queue by default, and the operator
delivers one into the running turn with the new per-queued-message Steer
button.

- web: resolveMessageDeliveryMode now queues for every flavor; QueuedMessagesBar
  gains a Steer button (pi + thinking + remote-controlled + immediate rows)
  backed by a new useSteerQueuedMessage hook + api.steerMessage.
- hub: POST /sessions/:id/messages/:messageId/steer -> syncEngine.steerQueuedMessage
  (pi-only gate, remote-only, scheduled/absent/invoked rejection) -> RPC.
- cli: pi runner registers 'steer-queued-message'; a queued message is promoted
  into the active turn via the existing PiSteerDispatcher (target generation
  captured at promote time; turn-ended steers fall back to the prompt FIFO).
  Steers requested while the message is still preparing are deferred and
  promoted right after preparation completes.
- Removed the now-dead Alt+Enter / touch-hold queue gesture (its only purpose
  was opting out of the removed automatic steer).

Verified: bun typecheck; cli/hub/web/shared suites (env-dependent runner
integration + kimi wire-locator flakes reproduce on pristine upstream and
are unrelated to this diff).

* fix(pi): preserve queued messages on rejected steers and pin the steering generation

Addresses both Major findings from the HAPI Bot review of PR #1480.

- steerDispatcher: a deterministic native rejection (Pi responded error) now
  degrades the message to the ordinary prompt FIFO instead of emitting
  messages-consumed. A promoted queued message must not be lost just because
  the steer was rejected; the hub row stays queued until the FIFO delivers it.
  The indeterminate-timeout path keeps its fail-closed consume + escalate
  behavior (a duplicate delivery would be worse).
- runPi: the deferred-steer path now captures the streaming generation at RPC
  request time (Map<localId, generation>) instead of reading it after
  preparation completes, so a steer requested against turn G1 can never be
  injected into a turn G2 that started while the message was preparing — the
  dispatcher's generation-mismatch check degrades it to the FIFO.

Regression coverage: negative steer response preserves the entry via the FIFO
(no consume); generation rollover while preparing delivers as a normal prompt
at the next settle (no steer into the new turn).

Verified: bun typecheck; cli pi suites (48 tests), hub 1041, web 2301, shared
240 — all green; only the pre-existing environment-dependent runner
integration test fails locally (reproduces on pristine upstream).

* fix(pi): reject all scheduled steers and always clear deferred-steer bookkeeping

Addresses the two Minor findings from the HAPI Bot follow-up review.

- hub: steerQueuedMessage rejects every scheduled row — mature ones included —
  aligning the endpoint with the web UI (Steer is never offered on scheduled
  rows) and preserving scheduled-FIFO delivery semantics.
- cli: the deferred-steer bookkeeping map is now cleared in a finally on the
  preparation chain, covering the early exits (cancellation before/after
  attachment I/O, empty prepared message, preparation failure) that previously
  could leave a stale generation entry behind for the session lifetime.

Regression coverage: hub steer gate tests (mature scheduled row stays queued,
non-pi flavor rejected) and a runPi test proving cancellation wins over a
deferred steer (no steer/prompt/consume after preparation completes).

Verified: bun typecheck; hub 1043 pass, cli 2421 pass (only the pre-existing
environment-dependent runner integration suite fails locally), web 2301 and
shared 240 unchanged since their green runs.

* fix(web): reconcile stale queued rows when a steer returns invoked

Addresses the remaining Minor finding from the HAPI Bot follow-up review:
when the steer endpoint reports the message was already invoked and the
messages-consumed SSE was missed while the row was still queued, the hook
now marks the row consumed locally (mirroring useCancelQueuedMessage) so
the queued bar cannot keep a stale actionable row until the next sync.

Regression coverage: steer returning status 'invoked' reconciles the row
via markMessagesConsumed and shows no toast.

Verified: bun typecheck; web 2302 pass (hub/cli/shared unchanged since
their green runs).
2026-08-11 22:27:14 +08:00
1cd4d1137a feat(hub,cli,web): fleet runner version governance (skew, self-upgrade, soft-fail reopen) (#1108)
* fix(hub): govern runner capabilities so Cursor reopen soft-fails on skew

Hub↔runner protocol drift was reported as missing Cursor chat data when
cursor-chat-store-status was unregistered. Soft-fail reopen on probe errors,
advertise required machine capabilities, surface an unmissable upgrade banner,
and stop-runner when a newer CLI binary is already on disk.

Fixes #1084

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web,hub): make runner skew banner dismissible; gate auto-upgrade

Compact the out-of-date banner (minimize + 1h snooze + per-host Restart)
so it no longer blocks the session list. Auto stop-runner on skew stays
opt-in via HAPI_AUTO_UPGRADE_RUNNERS / autoUpgradeRunners (default off).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): tolerate full sessionStorage on skew banner minimize

QuotaExceededError from setItem aborted minimize before React state
updated, leaving the banner stuck over the session list. Persist to
memory when storage fails; only enable Restart when a newer CLI is
already on disk; clarify opt-in is stop-runner only, not package push.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): drop redundant autoUpgradeRunners; runners already self-restart

CLI version handoff already reloads the runner when the on-disk binary
mtime changes. Hub-driven stop-runner on skew duplicated that. Keep the
skew banner and manual Restart only as a stuck/disabled-handoff escape.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub,web): runner-only caps ads; gate Restart on supervisor

Address #1108 bot Majors on the thin tip: terminal/lazy bootstraps no
longer merge CURRENT_MACHINE_CAPABILITIES into the machine row (only
asRunner registration does). Banner Restart refuses unsupervised hosts
so stop-runner cannot leave a detached laptop offline; supervised
runners advertise supervisedRestart via HAPI_RUNNER_SUPERVISED=1.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,cli,web): clear sticky runner ads; docs SUPERVISED; i18n skew label

Omit-means-clear on runner registration so rollback cannot leave
supervisedRestart/capabilities sticky; always advertise boolean
supervisedRestart from asRunner. Document HAPI_RUNNER_SUPERVISED=1
and localize MachineSelector UPDATE REQUIRED.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Debian <heavygee@oos-linux.in.lockhouse>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 22:24:39 +08:00
24e0c76717 feat(cursor): bump hub thinking on ACP harness wake (#1487)
* feat(cursor): bump hub thinking on ACP harness wake

When Cursor resumes after idle (notify_on_output / mid-idle ACP activity
or a permission request), flip thinking via the existing session-alive
keepalive so the hub list matches reality. Fixes #1470.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): emit thinking true/false edges for ACP harness wake

Address Codex Major on #1487: activity listener now reports idle as
false, and the launcher only keepalives on actual thinking transitions
so streamed chunks do not spam session-alive.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): reattach activity thinking listener after session/new remap

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 10:03:42 +08:00
SSU-WEI HUANGandGitHub 2548eaf3ed fix(cli): raise flaky claudeRemote first-test timeout to 15s under CI load (#1493)
The first test in claudeRemote.test.ts imports the full remote-module
graph and intermittently exceeds vitest's default 5s timeout on loaded
CI runners, failing PRs that touch no cli/ files. Give that one test a
15s per-test timeout; verified: typecheck exit 0, claudeRemote suite
6/6 pass (~2s cold), full suite 223/224 files pass (runner.integration
fails identically on pristine base in this sandbox).

Fixes #1491
2026-08-11 10:03:27 +08:00
SSU-WEI HUANGandGitHub b572c35e34 fix(pi): surface mid-turn LLM errors in web chat (#1479)
Pi finalizes a failed assistant message with stopReason 'error' and
errorMessage on both message_end and turn_end, but PiMessageAccumulator
only handled text/thinking deltas, so the error was silently dropped and
the web showed nothing (spinner stop + partial text at most).

Emit an AgentMessage {type:'error'} once per failed message (message_end
wins, turn_end is the safety net for Pi builds that skip message_end);
the web already renders AGENT_MESSAGE_PAYLOAD_TYPE error payloads as
system messages. User-initiated aborts stay silent.

Closes #1478
2026-08-10 11:47:56 +08:00
Junmo KimandGitHub 495fa53465 fix(opencode): surface upstream errors and retries (#1433)
* fix(opencode): report why a prompt failed instead of pointing at logs

The provider's own explanation already reaches this process: the ACP
transport rejects session/prompt with the JSON-RPC error message
verbatim. The launcher caught it, logged it, and handed the user a fixed
"OpenCode prompt failed. Check logs for details." — a remote user is by
definition not at the machine holding those logs, so a rate-limited
session simply stopped with no stated reason.

Only the message is used; that channel carries no response headers,
cookies or body. It is unbounded though (a 20KB provider body produced a
20,210-character message), so it is capped at the same 200 characters
the compaction bridge already applies to provider text, and the JSON-RPC
"Internal error: " wrapper is stripped. A failure with nothing readable
to say still renders the sentence it always did.

* feat(opencode): surface upstream retries from the agent event stream

A provider rate limit leaves OpenCode retrying indefinitely, and it
announces that on exactly one channel: its own server event stream.
Measured against a provider stubbed to answer 429 — 40 minutes, 85
retries, zero ACP notifications, zero stderr bytes, session/prompt never
settling. HAPI showed a session that looked like it was thinking and
said nothing else.

Subscribes to that stream once the ACP session id is known and reports
retries as the same api_error system message Claude sessions already
use, so the web timeline folds a run of them into one block whose
attempt count climbs. The reason rides along in the payload, and the
presentation now appends it to "Retrying..." when an agent supplies one;
sessions that supply nothing render exactly as before.

Not every attempt is announced. The backoff tops out at 30 seconds and
OpenCode does not give up, so a session held against a daily quota would
otherwise persist two messages a minute for as long as it is left
running. The first few attempts are reported, then only attempt numbers
that are powers of two, which needs no clock to decide.

The subscription must be scoped with ?directory=: without it the
endpoint delivers heartbeats and no session events at all, with neither
an error nor a 404 to notice. Its session.error event is deliberately
not read — it carries the Authorization header, cookies and the full
response body verbatim, and the same failure already reaches the user
stripped to a message through the ACP prompt error.

Delegated turns are not covered: OpenCode's subagent tool runs them in a
child session with its own id, and this follows only the one it was
opened for.

No countdown is rendered, though the payload offers one: a timeline
block outlives the turn it describes. Turn state is left alone; a
retrying session really is busy.

* fix(opencode): only relay prompt failures OpenCode itself reported

AcpStdioTransport flattens a JSON-RPC error response to
new Error(response.error.message), so a rejected session/prompt is
indistinguishable by shape from an error the transport built locally.
Its process-close error appends up to 4KB of raw subprocess stderr to
the message, which the formatter would then have published to the hub
and every connected client.

Requiring the "Internal error: " wrapper OpenCode puts on every service
failure turns this into an allowlist: an unrecognised rejection renders
the sentence this path rendered before rather than whatever it happened
to contain. Retry text is collapsed to one line before it is judged,
since a whitespace-only message is truthy and embedded newlines break
the one-line contract the helper states.
2026-08-10 10:48:48 +08:00
AnanovoandGitHub 044d72fdc0 fix(web): show file metadata in preview header (#1450) 2026-08-10 10:47:58 +08:00
SSU-WEI HUANGandGitHub 427ac1ff95 fix(pi): sync native session name (#1454)
* fix(pi): sync native session name

* fix(pi): sync live native renames via session_info_changed

Pi emits session_info_changed on /name and set_session_name. Share the
native-title sync callback between get_state startup/resume and the live
rename event so HAPI metadata.summary.text mirrors Pi's authoritative
session name without a change_title flow. Explicit HAPI/web renames
(metadata.name) keep precedence per sessionTitle.ts.
2026-08-10 10:46:16 +08:00
SSU-WEI HUANGandGitHub 2a98b4425f fix(agy): sync native Anti-Gravity conversation titles (#1476)
* fix(agy): sync native Anti-Gravity conversation titles (#1439)

Read the Anti-Gravity CLI's conversation_summaries.db title for the
active brain UUID and mirror it into HAPI session metadata.summary,
reusing the existing native-title normalizer so placeholder/empty
titles are ignored and a user-defined metadata.name stays preferred
by the web title logic.

The existing AGY scanner polls every 5s while a brain is known, so the
title follows native generation/rename without PTY parsing. Lookup is
best-effort: missing/locked DB or schema drift degrades to no-op.

* test(agy): cover delayed/renamed native title polling and cleanup stop

Addresses HAPI Bot review suggestion on #1476: scanner re-reads the
native title on later scans (null -> generated -> renamed) and stops
reading after cleanup.
2026-08-10 10:46:04 +08:00
SSU-WEI HUANGandGitHub bb40a4e8d0 fix(agy): preserve PTY running status (#1456)
* fix(agy): preserve PTY running status

* fix(agent): let a trailing idle marker win over a same-chunk busy marker

runAgentPty treated any busy marker in a chunk as authoritative, so a
chunk carrying both 'Generating' and the idle footer discarded the idle
marker. With the silence watchdog disabled (agy), nothing re-armed
inputReady and the session stayed 'running' forever. Compare marker
positions: the last one in the chunk reflects the final repaint.

Adds lastMarkerIndex for string and non-global RegExp markers and a
null-watchdog regression with both markers in one ANSI-decorated chunk.

* fix(agent): scan the full chunk and keep busy-run completion bookkeeping

- Scan promptBuffer + data before truncating to PROMPT_BUFFER_SIZE, so a
  busy marker followed by more than 4096 bytes in one callback is still
  seen; otherwise the run never completes and a pending web delivery can
  block subsequent messages.
- When an idle marker wins over a same-chunk busy marker, keep the busy
  confirmation so completeAgentRun fires for the run that just ended.
- lastMarkerIndex now clones non-global RegExps with the g flag and scans
  the original string, preserving anchors and terminating on zero-width
  matches (previously /$/ looped forever).
2026-08-10 10:45:52 +08:00
SSU-WEI HUANGandGitHub e6e229c068 fix(codex): preserve plan approval in yolo (#1420) 2026-08-10 10:44:41 +08:00
SSU-WEI HUANGandGitHub b744414241 fix(cli): skip invalid set_mode for interactive Copilot sessions (#1463)
* fix(cli): skip invalid set_mode for interactive Copilot sessions

applyInitialAgentMode() unconditionally sent session/set_mode with
modeId 'interactive', which the Copilot ACP server rejects (Invalid
mode 'interactive' - supported: agent/plan/autopilot), killing every
default-mode Copilot session at startup.

- Short-circuit applyAgentMode('interactive') as a no-op success,
  mirroring buildCopilotAcpArgs which already omits --mode at spawn
- Classify 'Invalid mode' responses as unsupported runtime switching so
  server-side mode rejections degrade gracefully (restart to apply)

Verified: bun typecheck + bun run test (full suite green).

* fix(cli): map Copilot interactive mode to ACP agent mode

Address review findings: the previous no-op for interactive left the
backend in Plan/Autopilot after a runtime switch while HAPI reported
Interactive, and classifying 'Invalid mode' responses permanently
disabled switching for valid modes too.

- Map interactive -> the ACP 'agent' mode id (the spawn default) so
  session/set_mode always receives a valid mode and runtime switches
  back to Interactive stay in sync with the backend
- Keep 'Invalid mode' out of the capability-loss classifier: it is a
  per-mode rejection, not evidence that set_mode is unavailable
2026-08-10 10:44:24 +08:00
weishu c2cc73159c fix(codex): preserve transcript final answer ordering 2026-08-10 10:36:34 +08:00