Commit Graph
516 Commits
Author SHA1 Message Date
wushenghua 3eca03be18 feat(hub,cli,web): team config inheritance, reply visibility, task loop, scoped tokens
- spawned members inherit the caller's tool/model/thinking level/permission
  (spawn request schema + hub runtime + spawn_peer tool); explicit overrides win
- replies from turns triggered by teammates (not by the human) no longer vanish:
  system prompt and ACP briefs now require an explicit team_send(to=human) for
  human-facing conclusions produced on those turns
- task loop: team_task tool (list/update), dependsOn gating, deliverable evidence
  required when a member marks done, auto broadcast on change, web shows deps and
  deliverable on the task board
- team-scoped agent tokens (hapi_team_*): issued per team, injected as
  HAPI_TEAM_TOKEN at spawn, restricted by the auth middleware to this team's
  messages/tasks/status and audited via [TeamToken] hub log lines
2026-09-16 07:59:58 +08:00
wushenghuaandCodex 3f81120312 fix(cli): detect DSH through the Web runtime
Co-Authored-By: Codex <noreply@anthropic.com>
2026-09-16 01:58:51 +08:00
wushenghua 6c2301ae0f fix(cli): keep team prompt tests out of the workspace
teamPrompt tests ran with the real cwd and wrote .hapi/team memory into the
checkout (one charter.md was accidentally committed in ef33332d). Tests now
pass an explicit temp cwd; getTeamPromptBlock/withTeamInstruction accept a cwd.
Drops the stray artifact and ignores .hapi/ in this repo.
2026-09-16 01:28:47 +08:00
wushenghua 073e08ae47 feat(cli): per-team memory dir with legacy migration and gitignore
- team memory moves to <repo>/.hapi/teams/<slug>-<short-id>/ so several teams
  on one repository no longer share (and clobber) a single charter
- the pre-teams/ shared .hapi/team dir is moved into the first team dir that
  starts after the upgrade
- .hapi/ is appended to the repo .gitignore once (idempotent, best effort)
2026-09-16 01:27:43 +08:00
wushenghuaandCodex 388082ccd4 fix(deploy): gate unavailable agents and bound session scans
Co-Authored-By: Codex <noreply@anthropic.com>
2026-09-16 00:54:18 +08:00
wushenghua e4cc2f8eb1 fix(deploy): skip --hapi-starting-mode for codex runner spawns (lost in v0.30.3 rebase) 2026-09-15 14:13:58 +08:00
wushenghua e66dc1b0cc fix(deploy): keep sharedCodexRuntime across runner heartbeat (spread fileState) 2026-09-15 14:08:17 +08:00
wushenghua 41f0935bd3 fix(deploy): restore sharedCodexRuntime flag lost in v0.30.3 rebase; add compact to opencode slash builtins 2026-09-15 13:58:53 +08:00
wushenghua ef33332dc5 refactor(hub,cli,web): move team memory into the repo (runner side)
Team memory now lives with the code, like CLAUDE.md: <main repo>/.hapi/team/
on the machine where the agent runs. Worktrees resolve to the main repo so
every member shares one directory. The CLI creates charter.md + handoffs/ at
session start (team env only), points the agent at it in the system prompt,
and reports the path in session metadata; the web reads it through the
session file API, so remote runners work too.

Hub keeps only coordination data (teams.db): members/tasks/messages/status.
Its memory-dir management and the /memory API are removed.
2026-09-15 11:40:51 +08:00
wushenghua 94c4a2d223 fix(hub,cli,web): give web-created team sessions their team context at spawn
Root cause of 'lead never answers': sessions created through the web form
never received HAPI_TEAM_* env, so their system prompts had no team rules
(only spawn_peer-created members did). Now:
- /api/machines/:id/spawn accepts a team descriptor; the runner exports the
  env as before
- team mode creates the team first, spawns the lead with the team context,
  then attaches it as lead; member mode spawns with the team context before
  joining
- OpenCode system prompts (title + native tool instruction) carry the team
  block too
- human ping copy is action-first: reply directly (auto-synced to the
  group); do not run shell commands or explore the team
2026-09-15 11:09:32 +08:00
wushenghua b76ae314be feat(hub): bridge human<->member replies into the team group chat
Members replied in their own session and the group chat stayed empty,
which read as 'no response'. The hub now records a pending ping when a
human message is pushed to a member, and on the member's next
thinking->idle transition mirrors its last assistant text into the team
log (deduped: skipped if the member answered via team_send).
Prompts updated: normal replies sync to the group automatically; team_send
is for member collaboration and explicit reporting.
2026-09-15 10:34:04 +08:00
wushenghua b59c0608d4 feat(hub,web,cli): add Agent Team (sessions-as-teammates)
Hub
- independent teams.db (own user_version ladder); feature gate default OFF,
  so a disabled hub never creates the file and team routes simply do not exist
- /api/teams: list/detail/by-session/memory/messages/spawn/tasks,
  human message channel, team rename/archive/delete
- team runtime: spawn members via the existing runner RPC (worktree/yolo
  options), peer message routing (mention/dm push, broadcast pull-only),
  budgets (max members, message rate, reply chain depth), member-offline
  announcements that wake the lead or escalate to the human
- team-attention notifications: in-app toast when visible, web push otherwise
- team memory dir {HAPI_HOME}/teams/<id>/ with charter.md template; read-only
  browse API with path-traversal guards

CLI
- MCP tools: team_status / team_read / team_send / spawn_peer (claude, codex
  bridge and ACP flavors); spawn_peer stays behind manual approval
- HAPI_TEAM_* env for spawned members; system-prompt block for Claude/Codex;
  assignment brief carries team rules for flavors without prompt injection
- RUNNER_CAPABILITIES.agentTeam advertised to the hub

Web
- team group chat route with folded peer threads, key/task filters,
  human composer, member/decision/system message cards
- side panel: live member status, interactive task board (create/status/
  assignee), team memory browser, budget
- team create / add-member / settings dialogs; sidebar team section;
  hub settings toggle for the feature (restart to apply)
- SSE team-updated invalidation

Zero impact when disabled: no teams.db, no routes, no tool or prompt changes.
2026-09-15 07:16:57 +08:00
wushenghua d1ccd5f6db fix: record usage from OpenCode native responses 2026-09-13 16:43:15 +08:00
wushenghua db63079cd2 fix: preserve peer queue local ids 2026-09-13 16:36:14 +08:00
wushenghua 22e6fecdc0 fix: finish v0.30.3 local feature integration 2026-09-13 16:34:34 +08:00
wushenghua cb0dbd7d22 chore: rebase local deployment features onto hapi v0.30.3 2026-09-13 15:55:04 +08:00
weishu 917baf1201 fix(codex): sync shared plan proposals and web actions 2026-09-13 12:09:39 +08:00
weishu 1de9613df6 docs: align documentation with current implementation 2026-09-12 22:19:32 +08:00
weishu d1f4972079 fix(codex): restore shared sessions after heartbeat and archive 2026-09-12 18:29:57 +08:00
weishu f48b14a312 test(cli): render agent picker output in CI 2026-09-12 16:40:20 +08:00
weishu 18c454fb87 fix(cli): remove Codex startup warning and banner 2026-09-12 14:56:13 +08:00
weishu d18c01b4b6 revert(codex): remove Luna Reserve fallback (#1780)
Revert 8357da0a9d and its later shared-runtime integration.

Remove automatic model fallback, account usage polling, and the related UI, protocol fields, tests, and documentation without adding replacement quota handling.

Keep generic model pagination and method probing required by the current shared-session architecture. Cover idle sessions staying online without usage polling or agent-state churn.
2026-09-12 14:36:45 +08:00
weishu fc2fdfc25b feat(cli): add agent picker and remove Claude default
Separate top-level help and version flags from agent arguments. Require explicit agents in scripts and preserve command argument boundaries.

Validation: bun typecheck, bun run test, targeted runner integration tests, and PTY/source/compiled argv smoke checks.
2026-09-12 12:03:59 +08:00
weishu 0c4abcb3d1 feat(codex): share sessions across terminal and web
Use one native app-server for terminal, Web and phone clients while retaining the existing CLI and Runner lifecycle.

Synchronize native queues, permissions, question history and steering state; preserve explicit permission precedence and per-turn usage models. Resume inactive clear commands through Runner and reject independent child cold resumes.

Add shared-runtime regression tests, generated protocol fixtures and lifecycle documentation.
2026-09-12 10:59:17 +08:00
25af3f8e13 fix(cli): give Windows Codex MCP shim-spawn test a 20s budget (#1824)
Vitest's 5s default flakes under GHA Defender/cold-start on the only
unit test that real-spawns on Windows (#1823). Keep the global unit
default unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 15:56:56 +01:00
weishu a729456682 fix(claude): answer local permission prompts from web
Bridge main-session PermissionRequest hooks without suppressing the native
terminal dialog. Reconcile replies against native results and clean up on
timeout, cancellation, mode switches, and session changes.

Keep reply IDs distinct from native tool IDs across web and native clients;
add protocol fixtures and regression tests.

Refs #1796
2026-09-11 18:54:47 +08:00
Junmo KimandGitHub fd2822bab5 fix(agy): allow switching existing sessions to newly available models (#1814)
* fix(agy): say what the model picker is actually waiting on

The spinner in the New Session AGY picker read "Checking Antigravity
authentication…", but nothing at that point checks authentication — the
machine is running `agy models`, and the sign-in prompt is a separate
branch below it, shown only when agy reports the failure.

Name the wait after the work: "Fetching available models…", the same
words agy prints while it fetches.

* refactor(agy): describe a probe by its outcome, not by its response

The probe function returned a finished `AgyModelsResponse`, so "agy could
not be reached" and "agy listed no models" both arrived as a successful
response carrying the hardcoded mirror, and the caller could no longer
tell which had happened. Every policy decision about that answer has to
live inside the probe as a result.

Hand back what the probe observed — a live catalog, an auth failure, or
nothing usable — and let the caller turn it into a response. Same
behaviour: the mirror still stands in for both failure modes, and the
60s cache still holds whatever came out.

* fix(agy): serve the model catalog stale-while-revalidate

The `agy models` probe is a whole agy invocation — around 3s on a good
day, 15s when it times out — and the 60s window meant the New Session
picker paid that again a minute after the last look.

Keep the last listing agy actually returned and answer from it: fresh for
ten minutes, then still answered while a probe refreshes behind it, until
the entry is a day old and stops standing in for the machine at all. A
probe that times out or loses auth leaves that entry alone, so a blip no
longer empties a working picker, and the hardcoded mirror is no longer
recorded as if the machine had reported it.

Three things fall out of that and are handled here. A machine whose
sign-in has actually gone bad would otherwise look healthy for a day, so
an auth failure rides along with the catalog it can still serve — and,
because nothing else would re-probe a catalog that is still fresh, a
warning riding on the answer is itself a reason to look again. A failed
probe is not repeated on the very next request either, or a machine where
agy hangs would spawn it once per poll.

An explicit refresh always costs a probe, and never rides one that was
already running when it was asked for.

* fix(agy): let Retry force a fresh model catalog probe

With the catalog held for ten minutes, Retry would otherwise hand back the
answer it was pressed to replace, so the intent travels to the machine:
`?refresh=true` on the machine route, an optional RPC param, and a
one-shot flag on the query so ordinary mount and focus refetches stay
cheap. Every hop is optional, so a hub and a runner on different versions
still talk — the older side ignores it and answers from its cache.

Retry also has to be reachable, and honest, in the state that needs it.
The machine now answers with both a usable catalog and the sign-in failure
behind it, so the picker keeps the list and says why it may be out of
date, with the button right there rather than only once there is nothing
left to show. Pressing it runs agy, which can take tens of seconds, so the
button says so while it does.

The client contract covers both: `agy-models` is the one catalog route
that can carry an `error` on a successful response, and the one that
takes a refresh parameter.

* fix(agy): use the machine catalog in the in-session model picker

New Session already asks the machine what `agy models` lists, but a
session that is already open offered the built-in list in
`shared/src/models.ts`. That list is a hand-maintained mirror, so a model
agy started offering after the last release could be picked for a new
session and not for the one already running.

Point the composer at the same machine catalog. The mirror stays as the
fallback for the moment before the machine answers, and a model the
session is already on is kept selectable — and readable, when it is one
of the known presets — even after the catalog moves on without it.

* fix(agy): announce a model catalog re-check that changed the answer

Serving the last known catalog answers the picker instantly, but a picker
that was already open kept showing that answer until the user closed and
reopened it — the machine had no way to say it had found something newer.

Say it on the stream that already carries machine changes. The machine
daemon — the only process that answers `<machineId>:listAgyModels` —
emits it, and the hub forwards it as `machine-agy-models-updated` with
nothing but the machineId. Namespace resolution, per-machine delivery and
reconnect replay all come from the existing path.

What counts as a change is what the route would answer, not what sits in
the cache. That distinction carries the cases: a sign-in that lapsed or
came back changes no models yet changes what the user is told; a machine
whose agy was signed out has been answering from the hardcoded mirror, and
its first real listing is the largest change there is, for every client
except the one awaiting it.

* fix(agy): re-read the announced machine's model catalog

On `machine-agy-models-updated`, cancel and refetch that one machine's
catalog query — the app's global connection is always subscribed, so an
open picker redraws wherever it is.

Cancelling first is what makes it correct rather than merely likely.
query-core cancels an in-flight fetch only when the query already holds
data, so a picker opening for the first time would otherwise join the
request already on its way and settle on the listing the announcement
replaced. The refetch is answered from the machine's cache, so it starts
no probe and cannot bounce another announcement back.

A reconnect the hub could not replay takes the resync path, which clears
the agy catalogs the same way — that path has no announcement to fall back
on, so it is the one that can least afford to join a stale request.

* fix(agy): keep a model the user picked when the catalog moves under them

The catalog can now change while the New Session form is open, and the
form dropped any selection the machine no longer advertised — including
one the user had just made.

Keep that one, and list it as no longer listed so the form does not imply
agy is still offering it. A model restored from a draft or a saved
preference is still dropped: it may never have been runnable here.

* fix(agy): announce uncached authentication changes
2026-09-11 14:19:37 +08:00
e2518b6fba fix(codex): emit ready notifications after local turn completion (#1802)
Co-authored-by: Alireza Ghassemi <ravenblackdusk@gmail.com>
2026-09-09 21:14:08 +08:00
AnanovoandGitHub 2b402d24d9 fix(web): show Codex round usage metadata (#1685)
* fix(web): show Codex round usage metadata

* fix(web): skip imported Codex round duration
2026-09-09 09:32:43 +08:00
Junmo KimandGitHub 7031d60eb4 fix(opencode): expose model-specific reasoning effort options (#1716)
* fix(cli): discover opencode thought_level via set_config_option on model switch

* fix(cli): apply and refresh opencode model switches so thought_level stays discoverable

* fix(web): track opencode effort options across model switches

* feat(cli,hub,web): dynamic opencode effort options in new-session form

* test(cli): avoid platform-specific process event narrowing

* fix(opencode): address variant discovery review findings

* fix(opencode): synchronize effort options with model targets

* fix(opencode): roll back rejected model targets

* fix(cli): guard opencode variant probe workspace paths

* fix(web): clear stale opencode effort on model switch

* fix(cli): clear stale opencode effort metadata

* fix(web): reset stale effort options on model switch

* fix(web): reset opencode effort poll budget

* test(web): enforce opencode effort poll budget
2026-09-09 09:31:30 +08:00
905d89e03a feat(pi): auto-title Pi sessions via bundled hapi_change_title extension (#1719)
* feat(pi): auto-title Pi sessions via bundled hapi_change_title extension

Pi sessions never got automatic titles: HAPI set PI_RPC_EMIT_TITLE=1 but
Pi does not implement it, and unlike the Claude/Codex/OpenCode launchers
the Pi bridge neither registers a change_title tool nor injects a title
instruction. This materializes a bundled Pi extension at launch and
passes it via --extension, giving Pi sessions the same titling flow:

- hapi_change_title tool (namespaced to avoid collisions with user
  extensions, mirroring the OpenCode launcher naming)
- first-turn system-prompt instruction matching the Claude/Codex wording
- title lands via ctx.ui.setTitle(), which the existing extension UI
  bridge already syncs into session metadata

Removes the dead PI_RPC_EMIT_TITLE env var. Requires pi >= 0.35.0
(when --extension landed, 2026-01).

Closes #1669 by giving sessions titles from the session model itself,
with zero extra title-provider calls.

* fix(pi): publish title extension atomically and keep retitle rule

Review findings from the HAPI PR bot:

- [Major] Publish the generated extension atomically. All Pi sessions of a
  HAPI version share the same versioned path and Pi treats extension-load
  errors as fatal, so a plain writeFile could break a concurrent launch with
  a half-written file. The extension is now written to a unique temp file and
  moved into place with rename; if rename loses a race against another
  launcher that already published the file, the existing copy is kept and the
  temp file is cleaned up.
- [Minor] Inject the title instruction on every before_agent_start instead of
  stopping after the first title, preserving the objective-change retitle
  rule from the persistent Claude/Codex instruction.

Adds a concurrent-materialization test and a handler test that executes the
title tool and asserts a later turn still carries the instruction.

* test(pi): fix handler return type in title extension test

* fix(pi): keep bundled title extension compatible with older Pi releases

Follow-up review finding: the embedded extension assumed three behaviors
that only exist in recent Pi releases, while HAPI has no Pi minimum-version
floor, so accepted older installations either failed to load the extension
or reached the tool without setting a title:

- import TypeBox via '@sinclair/typebox', the specifier both the legacy
  extension loaders and the current one (which aliases it to the bundled
  'typebox') resolve;
- normalize the execute() context positionally: current Pi passes
  (toolCallId, params, signal, onUpdate, ctx), legacy releases passed
  (toolCallId, params, onUpdate, ctx, signal);
- call ctx.ui.setTitle() directly instead of gating on ctx.hasUI, which
  older RPC releases report as false even though setTitle emits the
  extension UI event (this extension only runs under HAPI's RPC bridge).

Tests now exercise both call orders, including hasUI: false.

---------

Co-authored-by: HongChenGG <HongChenGG@users.noreply.github.com>
2026-09-09 09:31:08 +08:00
SSU-WEI HUANGandGitHub 8357da0a9d feat(codex): support backend-authorized Luna Reserve fallback (#1780)
* feat(codex): reconcile Luna Reserve fallback and conditional usage

* fix(codex): preserve queued settings and reconcile Reserve sessions
2026-09-09 09:27:40 +08:00
AnanovoandGitHub 5c5c8b3a9f feat(codex): support user-configured MCP servers in remote sessions (#1789)
* feat(codex): load user MCP servers in HAPI sessions

* feat(codex): proxy Windows MCP command shims

* fix(codex): address MCP review feedback

* fix(codex): use HAPI-owned MCP proxy

* fix(codex): preserve Windows MCP argv boundaries

* fix(codex): keep remote MCP placement unchanged
2026-09-09 09:26:43 +08:00
SSU-WEI HUANGandGitHub d332cd5957 fix(acp): preserve permission state when tool input is missing (#1782) 2026-09-09 09:22:27 +08:00
NightWatcher314andGitHub 97b1dd4c34 fix(codex): preserve context config in remote sessions (#1633) 2026-09-06 14:19:44 +08:00
Junmo KimandGitHub b85c72f09f fix(claude): accept first fork child prompt on SessionStart:fork hook (#1671)
* test(claude): reproduce fork child deadlock when init never arrives

* fix(claude): accept first fork child prompt on SessionStart:fork hook

A forked child starts query() before any prompt exists, and the SDK
emits 'init' only after the first prompt is sent. Waiting for init
deadlocked fork children until forever. Materialization is already
signaled by the SessionStart:fork hook, so accept the child prompt on
that signal instead.

* chore: restore bun.lock
2026-09-06 14:19:32 +08:00
Junmo KimandGitHub 4603c883a0 fix(cli): identify the Windows runner process through CIM (#1755)
* fix(cli): identify the Windows runner process through CIM

isHapiRunnerProcess() checks that the PID persisted in runner state still
belongs to a HAPI runner. On Windows it read the command line through wmic
and returned true for any live PID when that spawn failed.

wmic is absent on current Windows 11 builds, so that branch is the only one
those hosts take and the check degrades to a liveness test. When a reboot
hands the stale runner PID to another process, runner start adopts it and
exits without starting a runner.

Query the command line through PowerShell Get-CimInstance Win32_Process
first and keep wmic as the fallback, matching getProcessStartMarker() in
this file. A probe that succeeds without reporting a command line falls
through to the next one, and an unreadable identity still preserves the
live state, so a real runner is never discarded on a probe that came back
empty.

* fix(cli): keep an unverifiable runner pid out of destructive paths

The identity check answered a boolean, so "alive but unidentifiable" had to
collapse into one of the two verdicts. It collapsed into "this is the
runner", and runner start acts on that by calling stopRunner(), which force
kills the persisted pid once the HTTP stop times out. A reused pid that no
probe can read was therefore still reachable by a kill.

Report the identity as runner, foreign, unknown or dead instead. Only a
confirmed runner is reported as running. An unknown pid is neither signalled
nor cleaned up, because clearing the state would also drop the lock that
keeps a second runner from starting, and the pid may still be a healthy
runner behind a probe that failed transiently.

The posix branch had the same fail-open through isProcessAlive and now takes
the shared unknown path.

* fix(cli): gate the runner force kill on a confirmed identity

Two paths reach stopRunner() with a boolean that cannot carry "unverified".
runner start stops an existing runner, and the spawned start-sync child reads
the same answer through isRunnerRunningCurrentlyInstalledHappyVersion(), where
a false result is read as a version mismatch and also calls stopRunner().
Propagating the unknown state to each caller would leave the next path open,
so check the identity where the process is actually signalled: only a
confirmed runner is force killed.

Also stop reading WMIC's column header as a command line. wmic get CommandLine
prints the header even when the property is empty or unreadable, so trimmed
stdout was never empty on success and an unreadable process was classified as
foreign, which cleared the state and the lock instead of preserving them.

* fix(cli): gate the whole runner stop on a confirmed identity

The identity check sat in front of the force kill, but stopRunner() first
sends an unauthenticated POST /stop to the port recorded in runner state.
An http request is a signal too: if the pid and its port were both reused,
an unrelated local service receives that request before the check runs.

Move the check to the start of stopRunner() so a pid that is not a confirmed
runner is never contacted at all. The later force-kill check is redundant
once the whole sequence is gated, so it goes away.
2026-09-06 14:19:14 +08:00
AnanovoandGitHub aa0c1dc808 fix(codex): fail closed on ambiguous Web Rewind boundaries (#1707)
* fix(codex): fail closed on ambiguous Web Rewind boundaries

* fix(codex): reject malformed native rewind history

* fix(codex): reject duplicate rewind identifiers

* feat(codex): offer Fork fallback for ambiguous rewind

* fix(codex): gate rewind Fork fallback on exact native boundary

* fix(codex): reject unresolved rewind boundaries safely

* chore: trigger PR checks

* fix(codex): gate safe rewind fallback on fork support

* fix(codex): require complete user ids for rewind fallback
2026-09-06 14:18:40 +08:00
f4553dd1ec fix(codex): ignore late app-server writes during disconnect (#1748)
* fix(codex): ignore late app-server writes during disconnect

* fix(codex): scope app-server responses to their process

---------

Co-authored-by: huxiang <huxiang@myai.tech>
2026-09-06 14:18:20 +08:00
Junmo KimandGitHub 980a921ba1 test(cli): provide a stub Claude CLI so runner spawn passes the agent availability preflight (#1696)
The runner integration suite spawns sessions whose new agent-availability
preflight requires an installable Claude CLI. CI runners have none, so every
spawn returned agent_unavailable and four suite tests failed.

The isolated test env now writes a minimal stub claude binary into the temp
home and points workers at it via HAPI_CLAUDE_PATH (the same override the
production launcher honors), keeping production behavior untouched.
2026-08-29 12:52:47 +01:00
Junmo KimandGitHub ec08959f07 fix(agy): show Gemini 3.7 Flash in the agy model list (#1585)
* refactor(agy): extract the agy models probe from the fetch flow

One invocation and the decision of what to do with its output were
tangled in a single promise. Split them so the probe can be given
different arguments, and run more than once, without duplicating the
stream and timeout handling.

No behaviour change: the same argument vector, timeout, stream
handling and fallbacks remain, and the existing tests are untouched.

* fix(agy): read the model list from agy's structured output

The picker recovered ids and labels from `agy models`, whose table is
meant for people to read: it has emitted display names only, then
bare ids, then tab-separated id/label pairs over the 1.1.x line, and
each shape change silently sent the probe back to the hardcoded
mirror. Since agy started offering Gemini 3.7 Flash, that mirror is
what users see, so the three new entries never appear.

agy publishes the same listing as a structured payload, so read that
instead: `agy --output-format=json models` returns
`command.data.models[] = {id, label}`. The flag is global, so it goes
before the subcommand and takes its `=` form.

Releases that predate the flag ignore it and print the table, so the
text parser stays as the fallback, and it now also understands the
tab-separated shape. Every release checked, 1.0.16 through 1.1.13,
ignores the flag rather than rejecting it; a build that rejected it
would emit no models at all, so ask once more without the flag when
the first invocation yields nothing either parser can read.

* fix(agy): follow the Gemini 3.7 Flash row in the model picker

agy now lists Gemini 3.7 Flash as the top row of the `/model` TUI
picker, but the hardcoded row table still starts at 3.6 Flash. A
session whose current model is that row can never change its model:
the picker cannot identify the current row and the change is
rejected.

Add the 3.7 row to MODEL_ROWS/TARGETS. Every existing row shifts
down by one, which leaves their relative deltas, and therefore the
navigation keys emitted between them, unchanged. Mirror the three
new ids into AGY_MODEL_LABELS so the session pickers offer them too.
2026-08-26 09:37:58 +08:00
weishu e5a8212f4a feat(session): validate agents and browse workspace directories 2026-08-25 16:12:29 +08:00
SSU-WEI HUANGandGitHub be1ef2a2e4 feat(dsh): integrate DeepSeek Harness through ACP (#1632)
* feat(dsh): add DeepSeek Harness ACP flavor

* fix(dsh): update mobile flavor catalogs

* fix(dsh): keep mobile spawn policy managed

* fix(dsh): keep managed policy and prompt retry

* fix(dsh): suppress unsupported runner policy flags

* fix(dsh): align native managed-policy UX
2026-08-22 12:37:28 +08:00
Junmo KimandGitHub 661e9b4eb7 feat(web): show Claude round usage metadata (#1655) 2026-08-22 12:37:03 +08:00
AnanovoandGitHub 3d94e8eef3 fix(cli): preserve source extensions for generated media (#1650) 2026-08-20 19:36:36 +08:00
Junmo KimandGitHub 0aebf39c78 fix(opencode): keep one stored message per reasoning stream (#1643)
* fix(acp): carry the live reasoning marker on the wire payload

ACP agents stream thoughts a token at a time, so the handler coalesces
them into a buffer and re-sends the whole buffer under a stable stream
id every 250ms. The converter dropped the marker that says a payload is
one of those throttled snapshots, leaving the hub unable to tell a
replaceable snapshot from the settled message that closes the stream.

Mirrors how the text variant already forwards streamSnapshot.

* fix(hub): keep one stored message per reasoning stream

OpenCode reasoning arrives as a series of growing snapshots sharing one
stream id, and every snapshot was persisted as its own message. A 26h
session reached 48,844 rows and 63MB, and because the web budgets a
fixed number of messages, its 400-message window covered barely three
minutes of conversation — scrolling up walked through duplicate
snapshots instead of history.

Retire a stream's earlier live snapshots once their replacement is
stored. Sweeping only after the insert matters: the two statements are
separate transactions, so clearing first would leave a window where a
crash takes the whole stream. Only rows marked live are eligible and the
replacement is spared, so a stream always keeps at least one row and the
settled message that closes it is never removed.

Live rendering is unchanged: the web still receives every snapshot and
already folds them by stream id.

* fix(web): spend the message window on conversation, not repeated snapshots

The window budgets raw messages, but a reasoning stream renders as a
single folded block no matter how many snapshots it arrived in. On
sessions recorded before the hub started retiring them, those snapshots
fill the window on their own: in one 26h session the newest 400 messages
covered 202 seconds, so scrolling up paged through duplicates instead of
history.

Collapse each stream to its newest snapshot before trimming. Rendering
is unchanged — the timeline already folds them by stream id — and rows
without a stream id are never touched.

* fix(ios,android): port reasoning-snapshot compaction to the native windows

The window logic in HapiProtocol and :core:protocol is a one-to-one port
of the web store, so collapsing superseded reasoning snapshots only on
the web left the native windows budgeting raw snapshot rows. The hub
stores one row per stream now, but a client that already holds the older
snapshots still spends its window on them.

Add the same stream-id reader and compaction to both ports, in the shape
each already uses for agent-run rows, and pin the behaviour with a
pagination fixture. Both fixture suites enumerate shared/fixtures/pagination
from disk, so the ports cannot drift from the web again without CI saying
so.
2026-08-20 08:55:21 +08:00
SSU-WEI HUANGandGitHub 0f7a3da68b feat(cursor): mid-turn Steer via concurrent ACP session/prompt (#888) (#1609)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* feat(acp): split request dispatch from completion and add soft steer

- AcpStdioTransport.sendRequestWithDispatch separates stdin-accepted
  dispatch from the JSON-RPC response, keeping sendRequest behavior
  unchanged
- AcpSdkBackend tracks concurrent session/prompt requests with an
  activePromptRequests counter (main prompt + soft steers); response
  completion stays pending until every concurrent prompt settles
- beginSoftSteerPrompt kicks off a concurrent session/prompt (Cursor GUI
  Send semantics — no cancel, no handler swap) returning {dispatched,
  completed}; softSteerPrompt awaits the full response for direct callers

* feat(cursor): mid-turn soft steer via concurrent session/prompt (#888)

- CursorAcpRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, rejects control commands and mode mismatches,
  then soft-injects via beginSoftSteerPrompt without canceling the
  in-flight turn
- Acks the hub once stdin accepts the inject (not on turn completion) to
  stay inside the 30s RPC window; the launcher stays busy until the
  concurrent prompt settles so handlers are not swapped mid-inject
- steeringActive agent state mirrors the active-turn window; abort and
  cleanup reset it and invalidate pending steers
- Legacy stream-json Cursor sessions register a steer handler that
  reports unsupported

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* feat(shared): advertise cursor in the steer gate now that its handler lands

Cursor ACP sessions pass the web and hub steer gates; legacy stream-json
cursor sessions stay excluded.

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(cursor): consume the row at dispatch; keep waiters for prompt gating

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  the concurrent session/prompt; completion is background-only. A
  dispatched steer is never restored (no duplicate via the next prompt)
- softSteerWaiters are registered before awaiting dispatch so the main
  loop's finally cannot start the next prompt mid-inject; they still gate
  prompt handover on completion
- tests updated: post-dispatch ACP rejection keeps the row consumed

* fix(cursor,hub): never hang teardown on unresolved soft steer; align diagnostics

- Prompt-finally waits for soft-steer completion only when not exiting;
  the outer finally no longer waits at all — cleanup() disconnects the
  ACP transport, which rejects pending requests and settles the waiters
- syncEngine gate diagnostics and JSDoc name all supported flavors
  (Pi, Codex, Cursor ACP)
- regression test: Switch with an unresolved soft-steer completion still
  reaches teardown

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(cursor): abort drops soft-steer waiters so the next prompt never blocks

- Ordinary Abort (shouldExit false) now clears softSteerWaiters: the
  prompt finally cannot wait forever on a soft steer whose completion is
  unbounded; the ACP cancel rejects in-flight requests, and cleanup()
  settles leftovers on session end
- regression test: unresolved soft-steer completion after Abort no longer
  blocks the next prompt

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(cursor,acp): abort force-settles soft-steer bookkeeping

- AcpSdkBackend.abortSoftSteers() drops the concurrent-prompt counter and
  notifies response-complete so the next turn's waitForResponseComplete()
  cannot block on a soft steer that will never settle after abort
- handleAbort calls it before clearing the waiters; the main prompt's own
  finishPromptRequest stays guarded by Math.max(0, ...)
- unit tests cover counter release and no-op when idle

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* test(acp): match finishPromptRequest epoch signature in whitebox test

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(cursor): abort releases an in-progress soft-steer wait

- The prompt-finally wait races Promise.allSettled against the abort
  signal: an Abort that clears the waiters now also releases a wait that
  already started, so the launcher always reaches the next queued prompt

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(cursor,acp): commit on ACP acceptance; distinguish transport failures

- AcpStdioTransport marks transport-level failures (timeout, closed,
  stdin write) as indeterminate; explicit JSON-RPC error responses are not
- The steer handler commits + consumes on completion (ACP acceptance) and
  restores the row on an explicit rejection; an indeterminate transport
  failure keeps the row reserved so a delivered instruction is never
  re-sent, and the ACK reaches the hub even when abort reset the queue
- launcher/transport tests updated for the three outcomes

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* test(cursor): match tri-state cancel contract for dispatching steers

* fix(cursor): drop duplicate promptInFlight declaration after upstream merge

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(cursor,acp,web): indeterminate close marks, steer gating precision

- AcpStdioTransport.rejectAllPending marks close/protocol failures
  indeterminate, so an accepted-but-close-interrupted soft steer restores
  nothing (no duplicate delivery)
- the abort race in the soft-steer wait removes its listener in finally
  (no accumulation across repeated waits)
- SessionChat gates the Steer button on agentState.steeringActive for
  codex/cursor instead of the queued-grace thinking flag, so Steer is not
  exposed before the launcher can accept it
- codex pre-dispatch abort never enters reconciliation (merged from #1606)

* fix(steer): persist indeterminate outcomes without replay

* fix(cursor): hold ambiguous steers for explicit resolution

* fix(steer): make ambiguous delivery restart-safe

* fix(cursor): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(cursor): reject steers when prompt generation changes

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(cursor): preserve soft-steer reservations across abort

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(cursor): hold ambiguous dispatch failures

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(cursor): bound ACP dispatch acknowledgements

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(cursor): preserve state when dispatch ACK is uncertain

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(cursor): distinguish held cancel from removal

* fix(store): reserve schema v25 for steer delivery state

* fix(cursor): suppress late ACP updates after abort

* fix(steer): keep held cancel state and notify requeue

* fix(cursor): isolate late updates after abort

* fix(steer): release explicitly cancelled unknown reservations

* test(cursor): cover explicit held cancellation

* fix(codex): reject cancelled reservations before native steer

* fix(cursor): reject cancelled reservations before ACP steer

* fix(codex): make reservation restore atomic with state

* fix(cursor): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(cursor): hard-stop abandoned writes and update native queue state

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(cursor): isolate aborts and add native retry resolution

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(native): reconcile retry responses

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): resync busy cancel outcomes

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(cursor): hold restore failures for explicit resolution

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(cursor): drain foreground prompt after soft-steer abort

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones

* fix(cursor): drain soft steers before handler replacement

* fix(cursor): preserve buffered output on abort
2026-08-20 08:53:41 +08:00
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00
KorenKritaandGitHub 2fbd98dfe4 fix(pi): offer model-accurate thinking levels in the create-session form (#1626)
* fix(web): filter Pi effort options by model thinkingLevelMap in new session form

The create-session EffortField called getPiThinkingLevelOptions without the
selected model's thinkingLevelMap, so the Pi effort select always showed the
static off..high list: levels the model marks unsupported stayed visible and
xhigh/max never appeared even for models that opt in. Pass the map through
piSelectedModel (mirrors HappyComposer), and extend the stale-effort reset
effect so a level unsupported by the newly selected model falls back to auto.

* feat(cli): probe machine Pi models over RPC to carry thinkingLevelMap

The machine-level Pi model probe parsed the `pi --list-models` text table,
which only exposes provider/model/thinking-yes-no — thinkingLevelMap (and
name/contextWindow) never reached the create-session form, so model-accurate
thinking levels could not render there (xhigh/max are map-opt-in and were
permanently hidden; see the companion web commit).

Replace the table probe with a short-lived `pi --mode rpc` child
(--no-session --no-extensions --no-skills --no-prompt-templates --no-tools)
that issues get_available_models and reuses the session path's parsePiModels
schema, so machine and session catalogs share one wire contract. The
lightweight probe also measures faster than the table probe (~0.8s vs
~1.6-2.4s) and drops the whitespace-table parsing entirely. Cache/inflight
dedupe/timeout structure is unchanged.

* fix: address review findings on the Pi model probe and effort reset

Three review findings on the RPC probe / effort-map change:

- (high) Probe teardown killed only the direct child PID: with shell:true on
  Windows that is the shell, orphaning the interactive pi RPC process on
  every successful probe and on timeout; finish() also resolved before the
  process was confirmed gone. Use killProcessByChildProcess (taskkill /T on
  Windows, Unix tree kill, SIGKILL escalation) and settle only after the
  tree teardown completes.
- (medium) An explicit get_available_models success:false response was
  discarded, so Pi's own error text was lost and the interactive child hung
  until the generic 15s timeout. parsePiModelsProbeLine now returns a
  three-way result (unrelated/models/error), also validating the response
  id, and the probe rejects immediately with Pi's error.
- (medium) Switching an xhigh/max-capable model back to Default left the
  now-hidden effort in state and submitted it while the select visually fell
  back to auto. The reset effect now also covers model === 'auto' (undefined
  map) while still not resetting mid-resolve for a concrete model.

Tests: probe-line failure/foreign-id cases; NewSession restore-to-Default
reset and map-opt-in retention (the latter guards the former against a
false green from the restore path).

* fix(web): type the Pi model test mock as PiModelSummary

The inline mock element type omitted thinkingLevelMap, so the new
capability-driven tests failed typecheck (TS2353) and the required test job
stopped before the unit tests ran. Use the shared PiModelSummary type so the
mock cannot drift from the wire contract again.

* fix(web): reconcile hidden Pi effort when model discovery fails

A restored explicit Pi model never resolves when the machine catalog request
fails, so piSelectedModel stays null for good. The reset effect required a
resolved model, so it skipped reconciliation, while EffortField rendered with
an undefined map and hid xhigh/max. Creation is only gated on the loading
state, not on the error, so handleCreate could still forward the stale hidden
level for Pi to reject or clamp.

Treat a failed catalog as a settled selection (alongside Default and a
resolved model) and reconcile against the undefined map; keep skipping the
reset while a concrete model is still resolving without an error, so a
restored xhigh/max survives until the map can prove it valid.

Test asserts the spawn payload carries no effort after a failed catalog with
a restored xhigh; verified it fails when the error branch is reverted.

* fix(cli): honor a failed probe process-tree teardown

killProcessByChildProcess reports survivors by resolving false, but the probe
discarded that result and settled anyway. The Windows graceful path is
taskkill /T without /F and escalates nothing on its own, so a probe child that
refuses the signal would be reported as cleaned up and the catalog cached,
letting interactive Pi processes accumulate across refreshes.

Escalate to the forced teardown when the graceful one reports survivors, and
reject (caching nothing, so the next call re-probes) when even that fails.
Skip the check when the child has no pid: nothing can leak, and replacing a
spawn ENOENT with a teardown error would only obscure the real failure.

Adds probe lifecycle tests (spawn + teardown helper mocked) for the escalation
path, the reject-and-do-not-cache path, the no-pid spawn-failure path, and the
timeout path; verified they fail when the escalation is reverted.

* fix(cli): keep extensions enabled for the Pi model probe

--no-extensions silently dropped providers contributed through
pi.registerProvider, which the old `pi --list-models` probe did list. Users
with such an extension would have lost those models in the create-session
form only.

Verified with a project-local .pi/extensions provider in the same cwd: the
probe with --no-extensions returned 29 models, while both the default run and
the old table probe returned 30 including the extension's model (and its
thinkingLevelMap). Dropping the flag restores parity; the extension provider
now comes back with its map intact.

Discovery that cannot contribute models stays disabled (--no-session,
--no-skills, --no-prompt-templates, --no-tools). Cost: the probe now measures
~1.4-2.0s instead of ~0.65s, still at or below the old table probe (~1.6-2.4s)
and fronted by the existing 60s cache.

* fix(cli): probe Pi models from the home directory, not the runner cwd

Under launchd/systemd the runner cwd is `/`, and starting Pi there is
pathological: project discovery plus extensions that scan from the working
directory walk the whole filesystem root. With extensions enabled (required so
pi.registerProvider models still surface) the RPC probe took 16.8s at cwd=/,
past the 15s timeout, so machine-level model discovery failed 100% of the time
and the create-session form showed only 'Pi model discovery timed out'.

Measured on macOS with 9 global extensions:
  cwd=/      extensions on   16.8s  -> timeout
  cwd=/      extensions off   0.6s  -> would lose extension providers
  cwd=$HOME  extensions on    1.4s  -> 29 models, complete

The catalog is machine-scoped and does not depend on cwd, so probing from the
home directory is both safe and representative. Verified from cwd=/ under the
runner's exact launchd environment: 1.3-1.7s, 29 models.

The old `pi --list-models` probe was immune because it never initialized a
session; this regression arrived with the RPC probe and was missed because
every earlier verification ran from a project directory.

Tests assert the spawn cwd is homedir() and that --no-extensions stays absent;
verified the cwd test fails when the option is removed.

* fix(cli): keep the probe in the runner cwd, fall back to home only at a root

Forcing every probe to the home directory fixed the launchd timeout but broke
project-local discovery: a runner started inside a project stopped seeing that
project's .pi/extensions providers, which the replaced --list-models probe did
surface. Verified with a project-local provider: present from the project dir
(30 models), absent from home (29).

Use the runner cwd normally and fall back to home only when cwd is a
filesystem root -- the launchd/systemd case where Pi startup walks the whole
tree (16.8s, past PROBE_TIMEOUT_MS) and where there is no project to lose
anyway. process.cwd() can also throw for a deleted directory, so that falls
back to home too.

Verified under the runner's launchd environment: from / -> home, 1.6s, 29
models; from a project dir -> that dir, 1.4s, 30 models including the
project-local provider.

Tests cover all three branches and were checked in both directions: pinning to
home fails the project-cwd test, pinning to cwd fails the root test.
2026-08-19 09:20:40 +08:00
SSU-WEI HUANGandGitHub 3a436240c3 fix(pi): refresh context after compaction (#1634) 2026-08-19 09:19:26 +08:00