Commit Graph
89 Commits
Author SHA1 Message Date
AnanovoandGitHub 17ee052d9a fix(web): preserve the visible chat window during rewind (#1766)
* fix(web): preserve chat window during rewind

* fix(web): scope rewind invalidation preservation

* fix(web): clear unknown rewind boundaries

* fix(web): deduplicate rewind invalidations

* fix(web): retain rewind dedupe history
2026-09-09 09:23:31 +08:00
matyasrathonyiandGitHub fa1caf7d71 docs: explain execution-host sleep in remote sessions (#1774) 2026-09-09 09:23:14 +08:00
3873e58496 fix(web): keep streamed reasoning/text block ids stable across snapshot rows (#1741)
* fix(web): keep streamed reasoning/text block ids stable across snapshot rows

Streaming snapshots of one stream (pi/codex reasoning and text) arrive as
separate message rows, and the window store retires older rows as newer
snapshots land. The timeline derived the block id from whichever row was
first seen, so the id (and the threadMessageId built from it) churned on
every snapshot, remounting the rendered reasoning panel mid-stream and
replaying its open animation — the panel visibly flashed/re-rendered on
every snapshot tick.

Derive the block id from the stream id when present (unique per stream,
stable across snapshot rows) so the block is updated in place and the
smooth streaming keeps appending to the previous text. Row-derived ids
remain the fallback for content without a stream id.

Also rerun gen:fixtures to refresh the two golden fixtures affected by
the new id shape.

* fix(ios,android): mirror stream-stable block ids in native chat ports

The native HapiKit (Swift) and protocol (Kotlin) chat pipelines are ports
of the web reducerTimeline and are pinned by the same golden fixtures in
shared/fixtures/chat. After the web-side change to derive streamed
reasoning/text block ids from the stream id, the ports still produced
row-derived ids, so the iOS/Android fixture conformance suites went red
on the two refreshed fixtures.

Apply the same streamId-first id derivation (row-derived fallback kept)
to both ports so all three pipelines project identical block ids.

* fix(web,ios,android): reject blank stream ids as block identity

Blank ('' or whitespace-only) stream ids are not streams per the wire
semantics in shared/src/messages.ts (readReasoningStreamId trims before
accepting). The previous nullish fallback let accepted payloads carrying
blank ids through, so every such row shared one empty block id: the
merge maps collided and assistant-ui occurrence suffixes churned with
list position, reintroducing remounts.

Normalize with a trim guard in all three pipelines (web, HapiKit,
protocol) and add a web regression test covering both empty and
whitespace-only ids.

* fix(ios): use normalized stream id for block construction identity

The blank-id guard was applied to lookup and map insertion but block
construction still read the raw optional, so accepted payloads carrying
blank/whitespace ids produced blocks sharing one blank SwiftUI identity
instead of falling back to row-derived ids (web/Android already used the
normalized local). Hoist the nonBlankStreamId result and reuse it for
lookup, block identity, and insertion in both the text and reasoning
branches.

Also add native coverage for stream identity: stream-id derivation for
text/reasoning plus blank ('' and whitespace-only) fallbacks, which the
golden fixtures do not exercise.

* fix(web): pin blank stream-id identity contract in golden fixtures

Update the two stale fixture descriptions (stream-keyed blocks are now
keyed by the stream id, not the first message) and add a generated
conformance fixture covering empty and whitespace-only codex data.id
values for both reasoning and text: blank ids are not stream identities,
so each payload keeps its own row-derived block id instead of collapsing
onto a shared blank identity. Web, iOS, and Android all run this same
golden fixture.

* feat(hub): make title provider max_tokens and timeout env-tunable

Reasoning models used as title providers (e.g. GLM thinking models) need
more than 64 completion tokens and more than the hardcoded 10s timeout to
emit a title, and the only workaround was patching the compiled binary
after every install.

Expose both knobs via HAPI_TITLE_PROVIDER_MAX_TOKENS and
HAPI_TITLE_PROVIDER_TIMEOUT_MS, following the existing
HAPI_TITLE_SUGGESTION_RATE_LIMIT pattern; defaults are unchanged.

* docs(hub): document title provider max_tokens/timeout env knobs

Add the two new HAPI_TITLE_PROVIDER_* variables to the title-provider
configuration table in the installation guide, and extend the provider
test to cover the timeout abort path (the signal fires and rejects the
in-flight request).

---------

Co-authored-by: HongChenGG <HongChenGG@users.noreply.github.com>
2026-09-06 14:20:19 +08:00
weishu 4948d23669 fix(android): enforce Google Play compliance 2026-08-25 16:44:03 +08:00
weishu e5a8212f4a feat(session): validate agents and browse workspace directories 2026-08-25 16:12:29 +08:00
weishu a04275b51d chore: upgrade Bun to 1.4.0 2026-08-25 13:31:59 +08:00
SSU-WEI HUANGandGitHub be1ef2a2e4 feat(dsh): integrate DeepSeek Harness through ACP (#1632)
* feat(dsh): add DeepSeek Harness ACP flavor

* fix(dsh): update mobile flavor catalogs

* fix(dsh): keep mobile spawn policy managed

* fix(dsh): keep managed policy and prompt retry

* fix(dsh): suppress unsupported runner policy flags

* fix(dsh): align native managed-policy UX
2026-08-22 12:37:28 +08:00
SSU-WEI HUANGandGitHub c9e243a466 docs(macOS): raise launchd runner file limit (#1646) 2026-08-20 19:44:01 +08:00
weishu c62cea9be8 docs: privacy policy as a VitePress page, replacing the static HTML
Markdown under docs/ renders through the existing docs pipeline
(deployed at /docs/privacy.html) and stays easy to maintain; the
hand-written website/public/privacy.html is gone. Footer gains a
Privacy Policy link. Content unchanged: self-hosted architecture, zero
developer-side collection, FCM transit on Android vs E2E envelopes on
iOS, camera/microphone usage, deletion, contact.
2026-08-19 20:50:10 +08:00
weishu 47c0768c27 feat(brand): refresh HAPI icons across platforms 2026-08-19 20:35:07 +08:00
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00
weishu c76af90d88 feat(hub): push config joins the settings.json system (env > file > default)
FCM and iOS/APNs push were the only hub knobs read straight from env,
bypassing the configuration rule every other field follows (env >
settings.json > default, env persisted on first sight). Fold them in:
serverSettings resolves fcmServiceAccountPath / iosPushMode /
iosPushRelayUrl / apnsKeyP8Path / apnsKeyId / apnsTeamId / apnsBundleId
/ apnsEnv under the shared rule, and the two resolvers now consume
configuration instead of process.env. FCM_PROJECT_ID is gone: operators
point at the service-account JSON and the project id comes from the
file — no copying values out of it. Paths accept ~.
2026-08-18 23:10:57 +08:00
weishu 1ad78f409a feat(hub): iOS push channel — direct APNs + relay with E2E envelope (P1) 2026-08-18 20:56:49 +08:00
weishu fada270424 docs(api): sse.md — activeTurnStartedAt is not patch-applied (matches reference impl + fixtures) 2026-08-17 12:12:32 +08:00
weishu 81146e68c5 docs(api): drop temporary dead-link ignores now that all contract pages exist 2026-08-17 11:28:14 +08:00
weishu 1a40db3b2d merge: K2 client contract docs (sse/pagination/messages) 2026-08-17 11:25:46 +08:00
weishu 1b9ca34fb5 docs(api): add native client contract (sse, pagination, messages) 2026-08-17 11:18:14 +08:00
weishu 16ebba12e8 docs(api): add native client contract (auth, rest, errors) 2026-08-17 11:15:00 +08:00
AnanovoandGitHub 76e00f0091 feat(hub): add background-only ServerChan fallback (#1608)
* feat(hub): add background-only ServerChan fallback

* fix(hub): validate ServerChan background setting type

* fix(web): preserve pending visibility transitions

* fix(web): guard visibility reports across subscriptions
2026-08-16 23:04:03 +08:00
SSU-WEI HUANGandGitHub 6c6f4b4929 feat(agy): replace fragile PTY/TUI wrapper with headless print-mode transport (#1591)
Replace the Antigravity (agy) integration — a PTY wrapping the TUI with
output-marker scraping ('? for shortcuts', 'Generating', trust dialogs,
/model picker navigation, quota-screen regex) — with a headless print-mode
transport: every user turn spawns `agy -p <msg> --conversation <uuid>
--output-format stream-json`, and NDJSON events (init/step_update/result)
map onto the existing transcript-entry channel (sendAgySessionMessage), so
hub/web rendering is unchanged. ~8.3k LOC (incl. tests) removed.

Fixes #1588. Design: docs/design/agy-headless-transport.md.

CLI:
- new cli/src/agy/headless/: agyNdjsonParser (pure functions, malformed-line
  tolerance, step conversation-id adoption), AgyPlannerAccumulator (per-step
  delta accumulation with settling retries), AgyHeadlessDriver (per-turn
  spawn/kill loop, NDJSON chunk buffering, authoritative delivery ack via
  user_input/result, interrupt + retry + shutdown lifecycle with consume/
  restore, process-tree termination, SSH agent preserved, prompt log
  redaction, per-turn model snapshot with conversation-DB fallback)
- runAgy/loop/session rewired; agy is remote-only (no PTY, no local mode,
  no local-switch action); queued batches snapshot model/effort/mode
- deleted agyPty, agyPtyLauncher, agyHookCarrier(+scope cache), agyModelKeys,
  agyQuestionKeys, agyAskQuestion, agySessionScanner, agyPermissionHandler
  (+tests); buildAgyHooksJson removed; startHookServer agy-pre-invocation
  route → 200 no-op
- runner: agy reopen/resume via generic --existing-session-id; commands/
  agy.ts defaults remote; resume rejects ACTIVE agy sessions (remote-only,
  in-flight turns cannot hand off)
- MCP stays user-managed (agy reads ~/.gemini/config/mcp_config.json and
  workspace .agents/mcp_config.json natively in headless — verified)

Hub/web:
- machines.ts drops agy→pty forcing and rejects non-remote startingMode
- NewSession drops agy startingMode='pty'; terminal toggle disappears
  automatically; RemoteModeDisplay hides the local-switch hint when absent
- docs/guide/agents.md updated: headless print mode, no PTY/hooks, MCP via
  user's own mcp_config.json

Tests: 47 parser+driver tests (fake-binary e2e, chunk-split NDJSON, delivery
ack semantics, interrupt/retry/shutdown races, model attribution, EOF
framing, malformed envelopes); full suite green (cli ~2340, hub 1093,
web 2474, shared 262). Real-binary smoke on agy 1.1.13: single turn exit 0,
--conversation resume keeps the same conversation_id.
2026-08-16 22:44:27 +08:00
901f17d0ca feat(pi): support Pi slash commands from HAPI web (compact/session/model/help) (#1570)
* feat(pi): support Pi slash commands from HAPI web (compact/session/model/help)

Pi runs as 'pi --mode rpc' over piped stdio, so TUI slash commands typed in
web chat previously fell through to the LLM as plain text and silently did
nothing (notably /compact).

- shared: add Pi builtin slash command list (help/compact/session/model) so
  the web / menu exposes them; web test updated to match
- cli: intercept Pi builtin commands in runPi's user-message path
  * /compact [instructions] -> Pi compact RPC (120s timeout, works while
    streaming; summary + token delta reported back as chat messages)
  * /session -> get_session_stats formatted stats
  * /model [modelId] -> list/switch via set_model
  * /help -> supported-commands list
  * other Pi TUI builtins (/tree, /export, /reload, ...) -> explicit
    terminal-only notice instead of silent LLM pass-through
  * unknown slash text still passes through (extension commands, skills,
    templates keep working)
- gate the prompt pump with piCompactInFlight so queued prompts are not
  rejected by Pi mid-compaction; buffer commands until ready like prompts
- ListSlashCommands RPC merges HAPI builtins with Pi extension commands
- tests: parser unit tests + runPi integration tests (compact execution,
  streaming steer interception, failure reporting, FIFO blocking, model
  switch, unsupported commands, slash list merge)
- docs: document Pi slash command support in docs/guide/agents.md

* fix(pi): address review findings on slash command lifecycle

- compact timeout: fail the session (indeterminate outcome, runtime lease
  poisoned) instead of reopening the prompt FIFO into a possibly-compacting
  Pi; pump only when cleanup has not been initiated
- special commands: release the cancellation reservation before executing so
  a cancel landing mid-command is not acknowledged (hub would delete the
  queued row while the command still runs)
- tests: drop the duplicated slash-command describe block; add focused tests
  for compaction timeout with a queued prompt and cancellation during an
  in-flight special command

* fix(pi): route slash commands through the prompt FIFO and reject ambiguous models

- slash commands now share the prompt FIFO with ordinary messages: a
  /compact or /model typed after a queued prompt dispatches only after it
  (and after the active turn settles), instead of jumping the queue from
  the preparation chain
- the pump dispatches special entries out-of-band while piSpecialCommandInFlight
  keeps the FIFO blocked; steer promotion refuses slash commands
- /model <id> prefers an exact provider/modelId match and reports bare IDs
  shared by multiple providers as ambiguous instead of picking the first
- tests: FIFO ordering (queued prompt before /compact), steer-delivered
  /compact queued until settle, ambiguous/qualified model selection

* fix(pi): keep /compact interruptible, honor extension precedence, require token boundary

- head-of-line /compact dispatches even while Pi is streaming (Pi's
  compact() aborts the active generation itself); every other queued item
  still waits for the stream to settle, preserving FIFO order
- discovered extension commands / prompt templates override same-name
  builtins at message time, matching the slash-list merge precedence
- parsePiSpecialCommand requires a command-token boundary, so path-like
  text such as /compact.md or /model/config stays an ordinary prompt
- tests: interrupt rule, extension collision, reserved-name path prefixes,
  non-compact commands waiting for stream settle

* fix(pi): honor cancellation acknowledged during slash-command discovery

A cancel arriving while the chain awaits get_commands (cold cache) was
acknowledged via the preparing reservation but never re-checked, so a
canceled /compact could still execute. Re-check the cancellation marker
after discovery and drop the message before dispatch.

* fix(pi): qualify /model selectors and report failed slash RPCs once

- /model lists provider-qualified selectors (openai/gpt-5.2) so duplicate
  bare IDs remain usable and copy-pasteable; current model is qualified too
- compact/set_model failures are owned by the awaited slash/config handlers:
  the common response handler no longer emits the raw Pi error a second time
- tests: qualified listing with duplicate providers, single-message failure
  reporting for rejected /compact and /model

* fix(pi): consume slash-command queue row at dispatch

Special commands (/compact, /session, /model, /help) are executed by HAPI
itself and never delivered to Pi as prompts. Consumption was deferred until
the command finished, so a /compact run — an LLM summarization pass that can
take minutes — left the row stuck in the web queued bar for its whole
duration, then surfaced as a sent message. Consume the row the moment
dispatch starts; failures still surface via the explicit event message.

* fix(pi): guard special-command dispatch against unexpected rejections

* ci: retry Codex PR Review after infra failure (proxy 503)

* fix(pi): keep session queued-thinking grace during /compact dispatch

The queued-thinking grace is session-scoped, so clearing it while
acknowledging a dispatch-time /compact row also drops the grace for any
prompt queued behind it. /compact keeps running for minutes without
toggling Pi thinking state, which would leave the web session looking idle
while compaction and the following prompt are still pending. Only the
fast, synchronous commands (/session, /model, /help) clear the grace.

* fix(pi): render compaction summary as a dedicated chat block

The manual /compact RPC result was reported as two plain message
events ("📦 Compaction completed (tokens: …)" + "📦 Compaction
summary: …"), which the web chat renders as tiny centered status
lines — unusable for a real summary payload. Emit a structured
compact-summary event instead (summary + token delta) and render
it as an independent block: header with the delta and the summary
markdown in a scrollable panel.

Also emit the same structured event when importing Pi session
files (compaction entries), and queue the event lossless like
other user-visible messages so a disconnect cannot drop it.

Verified: bun typecheck clean; bun run test exit 0 (cli 2481
passed, web 2451 passed, hub/shared clean); runPi/loop/apiSession/
piSessions/presentation suites green.

* fix(pi): address HAPI Bot findings on compact dispatch and import

- Track compaction as thinking for its whole duration: /compact runs for
  minutes without a Pi streaming event, so the 15s queued-thinking grace
  alone left the web session looking idle while compaction and any queued
  prompts were still pending (updateThinkingState around the compact RPC).
- Imported Pi compaction summaries must use the event envelope
  (content.type: 'event') like the live wrapper's compact RPC result; the
  codex payload envelope is dropped by the web normalizer. Extend
  CodexImportedMessageSchema with the event variant.

* fix(pi): /model retries discovery when the model cache is empty

Startup model discovery can be late or fail once; using only the cached
catalog made /model report valid models as unknown. getPiModels() falls
back to the get_available_models RPC on an empty cache, used for both
listing and switching.

* fix(pi): interrupt in-flight /compact on Abort; surface startup model rejection

- The Abort action no longer waits on the runtime-mutation lease when a
  manual /compact is in flight (compaction can hold it for up to 120s,
  blowing the 25s abort deadline and failing closed). It sends the abort
  RPC directly so Pi cancels its compaction AbortController; the compact
  RPC's 'Compaction cancelled' error is not double-reported as a failure
  since Pi already emits the compaction_end(aborted) lifecycle event.
- A rejected detached startup set_model now emits a visible ⚠️ event into
  chat instead of only a debug log, restoring the pre-existing behavior.

* fix(pi): close the Abort race when /compact is queued on the mutation lock

Abort previously assumed an in-flight /compact always had its RPC issued;
the command is marked active at queue dispatch, but the compact RPC is sent
only after the runtime-mutation lock is acquired. An Abort landing in that
gap acknowledged success while the compact RPC still ran afterwards.
Track the compact's rpcStarted/cancelled state: Abort cancels a not-yet-
started compact in place (the queued callback skips it), and interrupts a
started one via the abort RPC as before.

* fix(pi): persist provider-qualified selection after /model switch

The success path updated currentModel/currentProvider and keepalive with a
bare model ID, leaving metadata.piSelectedModel on the previous provider.
The web picker prefers that metadata for selection, context-window
resolution, and effort options, so a switch like openai/gpt-5.2 ->
azure/gpt-5.2 was invisible. Persist piSelectedModel with the full
provider/modelId pair on every confirmed switch.

* fix(pi): retire pending extension UI requests when /compact interrupts a turn

The streaming-interrupt path sent the compact RPC without cancelling
pending extension UI requests first, unlike the Abort path. Editor
requests have no timeout, so the web could stay stuck on a stale
input/permission card and a later answer could be routed to the aborted
turn. Cancel all pending requests (with a response) before compacting.

* fix(pi): fail closed when the direct compact-abort RPC times out

The in-flight /compact abort branch awaited the abort RPC without the
ordinary Abort path's timeout handling: an unanswered abort left the
compaction outcome indeterminate (the compact RPC keeps the mutation
lease for up to 120s) while the wrapper still looked live. Fail the
session on PiRpcTimeoutError, mirroring the standard abort fail-closed
path.

---------

Co-authored-by: swear01 <swear01@users.noreply.github.com>
2026-08-15 11:21:45 +08:00
AnanovoandGitHub ad72229923 feat(sessions): add on-demand AI title suggestions (#1577)
* feat(sessions): add on-demand AI title suggestions

* fix(sessions): address title suggestion review feedback

* fix(web): ignore stale title generation results
2026-08-15 11:13:49 +08:00
SSU-WEI HUANGandGitHub d396e9d6d4 feat(voice): curate dictation credential presets to ElevenLabs, OpenAI, Groq (#1474)
* feat(voice): curate dictation credential presets to ElevenLabs, OpenAI, Groq

Groq transcription was already wired end-to-end (GROQ_API_KEY,
whisper-large-v3, standard mode), but the credential onboarding panel
listed five providers with no hint that Groq is supported, so mobile
users could not discover it.

- Curate Settings > Voice > Dictation credential presets to ElevenLabs,
  OpenAI, and Groq (Deepgram / OpenAI-compatible remain fully supported
  via env and stay listed when configured)
- Name the three presets in the empty-state and manage hints (en + zh-CN)
- Lock the curated list in with a web preset test, a hub route test for
  the Groq whisper-large-v3 proxy, and shared provider-listing coverage
- Note the presets and no-restart save behavior in voice-assistant.md

Verified: bun typecheck (cli+web+hub) and targeted suites pass; full
test gate green except pre-existing load-sensitive runner stress tests.

* fix(voice): keep legacy dictation providers manageable when configured

HAPI Bot review finding (Major): curating the onboard panel to the three
presets made settings-managed Deepgram / OpenAI-compatible credentials
impossible to rotate or clear from the UI.

- Re-add deepgram / openai-compatible to the onboard provider list
  conditionally when credentials exist, restoring update/clear controls
- Fall back to the first preset if the selected provider leaves the list
- Cover the conditional list in the preset test

* fix(voice): surface partial OpenAI-compatible credentials in onboard panel

HAPI Bot follow-up finding (Major): hub marks openaiCompatible.configured
only when both base URL and model exist, so api-key-only or endpoint-only
stored settings lost the UI path to rotate or clear them.

- Gate the openai-compatible onboard entry on any stored field (base URL,
  model, or API key) via hasOpenAICompatibleCredentials()
- Cover api-key-only / base-url-only / model-only cases in tests
2026-08-11 22:27:44 +08:00
1cd4d1137a feat(hub,cli,web): fleet runner version governance (skew, self-upgrade, soft-fail reopen) (#1108)
* fix(hub): govern runner capabilities so Cursor reopen soft-fails on skew

Hub↔runner protocol drift was reported as missing Cursor chat data when
cursor-chat-store-status was unregistered. Soft-fail reopen on probe errors,
advertise required machine capabilities, surface an unmissable upgrade banner,
and stop-runner when a newer CLI binary is already on disk.

Fixes #1084

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web,hub): make runner skew banner dismissible; gate auto-upgrade

Compact the out-of-date banner (minimize + 1h snooze + per-host Restart)
so it no longer blocks the session list. Auto stop-runner on skew stays
opt-in via HAPI_AUTO_UPGRADE_RUNNERS / autoUpgradeRunners (default off).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): tolerate full sessionStorage on skew banner minimize

QuotaExceededError from setItem aborted minimize before React state
updated, leaving the banner stuck over the session list. Persist to
memory when storage fails; only enable Restart when a newer CLI is
already on disk; clarify opt-in is stop-runner only, not package push.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): drop redundant autoUpgradeRunners; runners already self-restart

CLI version handoff already reloads the runner when the on-disk binary
mtime changes. Hub-driven stop-runner on skew duplicated that. Keep the
skew banner and manual Restart only as a stuck/disabled-handoff escape.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub,web): runner-only caps ads; gate Restart on supervisor

Address #1108 bot Majors on the thin tip: terminal/lazy bootstraps no
longer merge CURRENT_MACHINE_CAPABILITIES into the machine row (only
asRunner registration does). Banner Restart refuses unsupervised hosts
so stop-runner cannot leave a detached laptop offline; supervised
runners advertise supervisedRestart via HAPI_RUNNER_SUPERVISED=1.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,cli,web): clear sticky runner ads; docs SUPERVISED; i18n skew label

Omit-means-clear on runner registration so rollback cannot leave
supervisedRestart/capabilities sticky; always advertise boolean
supervisedRestart from asRunner. Document HAPI_RUNNER_SUPERVISED=1
and localize MachineSelector UPDATE REQUIRED.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Debian <heavygee@oos-linux.in.lockhouse>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 22:24:39 +08:00
a0621194bb feat(web): native/deep-link ingest for /share (GET url, text, title) (#1413)
* feat(web): ingest GET /share?url=&text=&title= deep links

Native companions that cannot POST via Web Share Target can open the
same session picker by synthesizing the IndexedDB transfer client-side.
When id is present, the existing SW path still wins.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): preserve share deep-link whitespace; scrub content beside id

Keep GET content strings verbatim when non-empty (match POST form-data).
When id is present with leftover url/text/title, replace to ?id= only so
payload does not linger in the address bar. Query-param contract stays —
fragments would break shipped native companions.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): ingest /share deep links via URL fragment, not query

Shared url/text/title must not appear on the HTTP request line — hub
Hono logger (and any access log) records path+query. Native companions
open /share#url=&text=&title=; the client reads the fragment, scrubs it,
and continues with the existing ?id= picker path. Query validateSearch
keeps only id/error (Web Share Target redirect).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): keep /share hash ingest across StrictMode remount

Capture the fragment in useState and reuse a single putShareTransfer
promise so the first effect's scrub + cleanup cancel does not lose the
deep-link under React.StrictMode.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(web): fetch companion fileUrl into /share transfer files

Native shares cannot put binaries in the hash fragment. Companions host a
one-shot CORS URL and pass fileUrl/fileName/fileType; the share page fetches
bytes into the same IndexedDB files[] as Web Share Target POST.

* fix(web): cap share fileUrl fetch at the composer upload ceiling

Stream fileUrl downloads with Content-Length and body size checks matching
MAX_UPLOAD_BYTES so a crafted deep link cannot buffer unbounded bytes into
IndexedDB. Align native deep-link docs on the fileUrl hand-off vs POST.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): cast streamed fileUrl chunks to BlobPart for tsc

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 12:36:58 +01:00
24e0c76717 feat(cursor): bump hub thinking on ACP harness wake (#1487)
* feat(cursor): bump hub thinking on ACP harness wake

When Cursor resumes after idle (notify_on_output / mid-idle ACP activity
or a permission request), flip thinking via the existing session-alive
keepalive so the hub list matches reality. Fixes #1470.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): emit thinking true/false edges for ACP harness wake

Address Codex Major on #1487: activity listener now reports idle as
false, and the launcher only keepalives on actual thinking transitions
so streamed chunks do not spam session-alive.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor): reattach activity thinking listener after session/new remap

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 10:03:42 +08:00
28df974edd feat(settings): onboard hub provider credentials for dictation and voice (#1392)
* feat(settings): onboard hub transcription provider credentials in UI

Env-only keys made dictation invisible; Settings can now add/edit/clear
hub-side credentials (masked), with env still winning as override.
Refs tiann/hapi#1384.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(settings): onboard voice-assistant backends alongside dictation

Same Settings credential surface now covers ElevenLabs, Gemini Live, and
Qwen Realtime (alias env pairs), not only transcription providers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): address PR #1392 Major credential onboard findings

Alias env locks, non-destructive Save (omit empty fields), and
owner-only settings.json permissions for hub-stored provider secrets.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): harden credential onboard for second-pass Majors

Owner-namespace gate, stage-then-sync env after persist, and
per-field OpenAI-compatible editability under mixed env locks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): serialize settings RMW and clear partial compatible creds

Per-file settings lock for concurrent credential PUTs, and Clear shown
for partial OpenAI-compatible entries (key/url/model alone).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): serialize all settings writers via updateSettings

Route credentials, relay auth, generators, server settings, and CLI
token persistence through a locked RMW helper; reset Clear form state.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): share cross-process settings lock with CLI

Extract withSettingsFileLock for hub+CLI, keep owner-only 0o600
rewrites, and race hub credential updates against CLI-style writers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): keep UI secrets out of process.env; PID-own settings locks

Settings-backed provider credentials now live in an in-memory overlay
(getProviderEnvironment) so tunnel/ACP/Codex children do not inherit them.
Settings file locks record pid+token and only reclaim dead or legacy locks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): never reclaim ownerless settings lock sidecars

wx creates the lock path before the owner JSON is visible; unlinking
null owners let a waiter steal a live acquisition and collide on
settings.json.tmp (CI ENOENT). Only reclaim parsed owners with dead PIDs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): reclaim dead locks via rename; clean up failed publishes

Stale reclaim renames the sidecar to a unique break path and re-verifies
the expected dead owner before deleting it, so a loser cannot unlink a
successor's live lock. Failed owner writes unlink the wx sidecar.
Reclaim uses a sync owner read so contenders do not all observe one
dead owner across an await and race the exclusive create.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): reclaim dead locks under exclusive reaper sidecar

Stale reclaim now takes a fixed settings.json.lock.reap lock, re-validates
pid+token, then unlinks — so a delayed contender cannot move a successor's
live lock aside. Also document providerCredentials in settings.schema.json.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): fail closed on corrupt CLI settings; backoff busy reaper

CLI updateSettings now uses a strict read that rejects invalid JSON
instead of treating errors as {}, which could wipe providerCredentials.
Settings lock reclaim sleeps when another process holds .reap so retries
are not burned synchronously.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): publish locks via candidate+link; fix CLI vitest hoist

Acquire settings locks by writing a complete candidate then linkSync to
the fixed path so a crash cannot leave an empty live sidecar. Fix the
CLI persistence regression test to create its temp dir inside vi.hoisted.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): replace bespoke lock with proper-lockfile; hide tenant creds UI

Codex kept finding crash windows in hand-rolled lock sidecars. Switch the
shared settings lock to proper-lockfile's mkdir + mtime lease. Hide the
owner-only credentials editor from non-default namespaces on the voice page.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): adapt sessionSummaryContract to outcome updateSettings

Rebase onto main brought #1376 unique tmp + outcome-shaped writers;
wire sessionSummaryContract and the write-failure credential test to match.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: retrigger CI after rebase onto upstream/main

Empty commit — Meta reported no checks on da0c6c258 after tip-forward rebase.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-07 13:20:03 +08:00
c0b30bf916 feat(cli): MCP list_peers + runner hub auth for peer discovery (#1372)
* feat(cli): MCP list_peers + runner hub auth inheritance

Runner-spawned agents could not discover same-hub peers without
sitting on the hub host or pasting a session id. Add MCP list_peers
(in-process credentials), export HAPI_API_URL/CLI_API_TOKEN after
auth init for shell fallbacks, and clearer auth failure hints.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): do not export default hub URL into HAPI_API_URL

exportHapiHubAuthEnv was writing the implicit localhost default into
process.env, which made maybeAutoStartServer skip starting the bundled
hub. Only export HAPI_API_URL when the URL came from env or settings;
always still export CLI_API_TOKEN. Also fill missing deliveryMode on
abort restore so web typecheck matches RawSendError (main tip unblock).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): widen initializeApiUrl mock return type in test

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): never export CLI_API_TOKEN; exclude self from list_peers

Keep settings/prompt-backed hub secrets out of wrapped agent env so
shell JWT+curl cannot bypass peer-tool approval. Fresh hapi re-reads
settings; env-backed tokens already inherit. list_peers omits the
calling session from the shortlist.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): resolve peer labels via summary/path like web titles

list_peers was showing (unnamed) for ordinary sessions because titles
live in metadata.summary.text. Match web getSessionTitle and collapse
whitespace so each peer stays one agent-readable line.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub): emit full peer ids and honor GET /sessions?limit

Short 8-char prefixes collide across UUID namespaces; print full ids so
resolveSessionByPrefix stays unambiguous. Honor optional limit after sort
so listPeerSessions stops loading the whole namespace for scheduled counts.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): type sessions limit test mock as Map<string, number>

CI tsc rejected Map<string, null> for getNextScheduledAtBySessionIds.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub): unbounded ping resolve; peer list order=updatedAt

Keep GET /sessions?limit only for discovery callers. ping/inspect omit
limit so full UUIDs outside the first 500 stay resolvable. Peer lists
pass order=updatedAt so truncation matches newest-first. Basename
fallback splits Windows paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): auto-approve ACP title List Peer Sessions

Permission derivation prefers request.title; match the MCP tool title
form so default-mode ACP sessions do not prompt on discovery.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): pad list_peers fetch; split hub URL vs token hints

Fetch limit+2 when excluding the caller so overflow still surfaces at
limit=100. Clarify that auth login only saves the token, not HAPI_API_URL.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): use boolean overflow for ping-peer --list

Match MCP list_peers: fetch limit+1 and mark hasMore instead of claiming
an exact omitted count from a 200-row sample.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): tolerate mocked machineCache without expireInactive

CI flake: 5s inactivity tick hit test doubles that only stubbed
getOnlineMachinesByNamespace. Optional-call + stub the method.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-05 22:16:04 +08:00
weishu b7503d9309 docs: restructure docs site and fix drift against code
- split installation.md into installation/deployment/notifications
- merge cursor/grok guides into new agents.md with full support matrix
- sidebar: grouped sections; add namespace, deployment, notifications,
  native companion contract
- fix license footer (AGPL-3.0), settings schema fields and $id
- fix drift in pwa, faq, namespace, how-it-works, voice-assistant,
  quick-start, native-companion-contract
- move mermaid lightbox dogfood doc to localdocs (untracked)
- README: complete agent list, replace dead cursor/grok links
2026-08-05 07:51:16 +08:00
weishu de5fa4aecd docs: update voice assistant guide for multi-backend support 2026-08-05 06:46:26 +08:00
79f91e4b45 fix(acp/runner): Cursor worktree banner + skip nested --worktree hang (#1087)
Ignore Cursor's Using worktree stdout banner without masking other
non-JSON ACP frames (markClosed + kill). Skip --cursor-worktree when
spawn directory is already a linked git worktree so ACP can initialize.

Fixes #1085

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-04 10:58:24 +08:00
KorenKritaandGitHub c1b32b51fe fix(web): make browser-local speech probing Android-safe (#1349)
* fix(web): guard browser-local speech probes

* test(web): cover concurrent speech probes

* docs: clarify browser-local speech probing
2026-08-04 08:19:56 +08:00
weishu b67f4e56e5 feat(hub): per-hub relay auth keys with automatic recovery
The public relay used to accept a shared auth key compiled into every
hub, so its bandwidth was open to anyone. The relay now issues a
per-hub credential it can meter and revoke, and hubs obtain one on
their own.

- --relay resolves an auth key at startup: HAPI_RELAY_AUTH env, then a
  key persisted in settings.json, then a fresh key from the relay's
  /issue endpoint. There is no shared-key fallback; if no key can be
  obtained the tunnel does not start and the hub says why.
- A persisted key rejected by the relay (HTTP 403 after revocation or a
  secret rotation) is discarded and replaced once, then the tunnel is
  restarted, so a revoked hub recovers without manual edits. Keys given
  explicitly through the environment are never overwritten.
- Issuance is rate-limited per public IP; HTTP 429 is reported with the
  retry hint instead of being retried blindly, which matters for users
  sharing a CGNAT or corporate egress address.
- The tunnel URL now comes from upstream tunwg's slog JSON on stderr
  (msg="listener started"), replacing the fork's custom --json event,
  and --log_level=0 keeps per-request logs out of the hub console.

Requires a relay running tunwg with TUNWG_AUTH_SECRET configured.
2026-08-04 08:18:17 +08:00
SSU-WEI HUANGandGitHub c3a5522207 Add realtime dictation providers (#1329)
* feat: add realtime dictation providers

* fix: cancel realtime dictation startup

* fix: refresh local dictation availability

* fix: preserve dictation on disconnect

* fix: normalize OpenAI language hints
2026-08-03 10:03:06 +08:00
SSU-WEI HUANGandGitHub 9d07857570 Add provider-backed dictation mode (#1327) 2026-08-03 06:05:58 +08:00
quecai-niuandGitHub 0384a6e837 [codex] document Xiaomi microphone troubleshooting (#978)
* docs: add Xiaomi microphone troubleshooting

* docs: keep Xiaomi troubleshooting in English
2026-07-31 14:25:46 +01:00
226b2d066a feat(hub): native companion (FCM) push channel + device registry + pairing QR (#803)
* feat(hub): native companion (FCM) push channel + device registry

Adds opt-in FCM HTTP v1 notification delivery so a companion mobile/wearable
app can receive permission, ready, and task notifications end-to-end. The
channel is gated entirely on FCM_SERVICE_ACCOUNT_PATH + FCM_PROJECT_ID being
set; operators not running a companion see zero behavior change.

What lands:

- POST/DELETE /api/devices/register — JWT-authed FCM token registry,
  upsert on (namespace, deviceId, platform), platforms `phone` | `wear`.
- Sqlite v9 → v10 migration adds `fcm_devices` (idx on namespace + token).
- FcmService — minimal HTTP v1 client, RS256 service-account JWT via
  jose (dep already in tree), 5-minute access-token cache, 401 retry.
- FcmNotificationChannel — implements NotificationChannel, sends data-only
  FCM (so companion can route to phone+watch surfaces). Body composition
  parses an optional trailing `AGENT_NOTIFY_SUMMARY {json}` line for richer
  ready summaries; truncates plain assistant text to 280 chars otherwise.
  Tags each payload with `severity` (info/warning/success/error) so clients
  can color/categorise the notification.
- PushNotificationChannel gains a NativeFallbackProbe — when a namespace
  has at least one registered FCM device, web-push and SSE in-page toast
  are skipped so the operator does not double-notify on phone+browser.
  Probe is no-op when no FCM device is registered; PWA-only setups
  unchanged. Branch trace gated on HAPI_NOTIFY_DEBUG=1.
- shared/src/messages.ts — `extractAssistantPlainText` (codex + Claude SDK
  shapes) and `extractNotifySummary` (strict end-anchored line parser).
- hub/src/notifications/toolArgs.ts — tool-arg formatters lifted out of
  telegram/sessionView (kept duplicated there in this PR; refactor of
  Telegram is a follow-up).
- docs/api/native-companion-contract.md — payload + endpoints + env vars,
  versioned at contract v1.

Test coverage:

- 260 hub tests pass (incl. 23 new across FCM channel, push dedup,
  v10 migration, devices route).
- 60 shared tests pass (messages parsers).

Notes for reviewers:

- Reference companion implementation lives in a separate Android repo
  (Kotlin, phone APK + Wear OS APK) — this PR is hub-side only.
- No new runtime deps (`jose` and `zod` already declared in hub).

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): clarify scope - companion is remote-hub client, not hub-on-phone

Adds a Scope section to the native-companion contract so anyone
implementing it knows the audience: operators running the hub on a
server who want phone/watch as a notification surface, not users
expecting a Termux-bundled hub. Mirrors the framing now in
heavygee/hapi-companion README.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): correct Scope section - hub topology is unchanged

Removes the prior framing that referenced a non-existent 'Termux
hub-on-phone' alternative. This contract describes a native client to
the same hub the PWA talks to; it does not change where the hub runs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(web): companion app pairing QR in Settings

Companion section in Settings renders a QR code encoding the deeplink
hapicompanion://bind?hub=<base>&code=<token>. Scanning it from the HAPI
companion app (Android phone or Wear OS) auto-fills the bind form and
authenticates against this hub - no manual URL/token paste.

QR is gated behind a Show button so the access token doesn't sit visible
on screen by default; a Copy link affordance and the textual deeplink
are also exposed for manual onboarding.

Adds qrcode + @types/qrcode to web/ (already a hub dep, no new resolved
package - just a workspace declaration).

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(hub): terminal QR for companion app pairing alongside PWA QR

After the existing PWA access QR is rendered on tunnel start, also print
the hapicompanion://bind?hub=...&code=... deeplink and a matching QR.

Same tunnel + token, different scheme: phones with the companion app
installed pick up the deeplink via the manifest intent filter; phones
without it ignore it and fall back to the PWA QR above.

QR rendering failure is non-fatal in both cases - the textual deeplink
above the QR is sufficient for manual paste.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): address HAPI Bot review on PR #803

Two bugs surfaced by the upstream review bot:

1) Web Push silently dropped when FCM is not actually configured.
   The native-fallback probe only checked the device registry; it did
   not check whether resolveFcmConfig() actually succeeded. So an
   operator who previously enabled FCM, registered a phone, then later
   started the hub WITHOUT FCM_SERVICE_ACCOUNT_PATH would see the probe
   return true (devices still in DB) -> Web Push suppressed -> no FCM
   channel registered -> notifications go to /dev/null.

   Fix: extracted the probe construction into buildNativeFallbackProbe()
   which short-circuits to () => false when fcmConfig is missing. Probe
   never even consults the device store in the no-config branch, so
   stale rows can never matter.

2) Transient FCM failures permanently unregistered devices.
   sendToToken() returned a single boolean and sendToNamespace() removed
   any device whose send returned false. A 429 (rate limit), 503
   (server error), 401 (auth glitch), or even an ECONNREFUSED would
   delete the device row, after which the user would need to re-pair to
   get notifications again. The bot caught it; the fix is the obvious
   one.

   Fix: sendToToken() now returns 'sent' | 'invalid' | 'failed'.
   - 'invalid' is reserved for the responses that genuinely indicate a
     dead token: HTTP 404 with UNREGISTERED/NOT_FOUND, and HTTP 400
     with INVALID_ARGUMENT explicitly referencing the token field.
   - Everything else (429, 5xx, 401, 403, network errors) is 'failed'
     and counts toward the failed tally without removing the device.

   sendToNamespace() only calls removeDeviceByToken() on 'invalid'.

Tests: 11 new tests across two new files. fcmService.test.ts covers
all six branches (200, 404 unregistered, 429, 503, 401, network error)
plus a mixed-batch case that proves invalid tokens get removed in the
same call where transient-failure tokens survive. nativeFallbackProbe
.test.ts covers both no-config and configured branches plus the
explicit "no-config never touches the store" guarantee.

Hub test count: 273 -> 284 (all passing).

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): correct FCM visibility rule and remove unsupported event type

HAPI Bot review on PR #803 caught two contract-doc accuracy gaps:

1) Visibility rule was wrong. Doc said "FCM fires when Web Push would
   fire AND client not visible via SSE", but FcmNotificationChannel
   ALWAYS fires regardless of PWA visibility (deliberately - native
   companion is the canonical wrist-first surface, and there is a
   passing test asserting this). Companion app implementers reading
   the contract would have built foreground-suppression logic and
   then dropped notifications when the PWA tab was open.

2) Documented `session-completed` event doesn't exist. NotificationHub
   never calls into a 'session-completed' channel method on
   FcmNotificationChannel; the type would never reach a native client.
   Removed from the documented enum, leaving only the three actual
   events: ready, permission-request, task-notification.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): drop trailing whitespace, use blank line for paragraph break

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): persist CLI access token after Telegram bind so pairing QR works

The Settings -> Companion pairing QR reads the original CLI access token
from localStorage (hapi_access_token::<baseUrl>) so it can be encoded into
the hapicompanion://bind deeplink. For browser/CLI logins useAuthSource
already persists the token via setAccessToken, but the Telegram Mini App
bind path went through useAuth.bind() which exchanged the typed CLI token
for a JWT and never persisted it. Telegram users therefore always saw the
"signed in via Telegram..." fallback and got no usable QR.

After a successful client.bind() we now mirror useAuthSource's behavior
and write the same accessToken to the same localStorage key, restoring
parity between the two auth paths. No change for browser/CLI users.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): gate native-fallback probe on rolling FCM health

The native-fallback probe previously returned true whenever FCM was
configured AND devices were registered, which suppressed web-push for
the namespace. The HAPI Bot correctly pointed out the gap: if the FCM
pipeline silently breaks (expired service-account key, sustained 5xx,
OAuth token-fetch failure, network blackhole) the operator gets nothing
on either channel until they manually intervene.

Approach (deliberate, not the bot's exact suggested fix):

- FcmService now keeps a small rolling window (last 8 outcomes) of send
  attempts and exposes `isHealthy()`. The threshold is 5+/8 failures =
  unhealthy; the buffer starts empty so a freshly-booted hub is
  optimistic ("innocent until proven guilty") and does not double-fire
  on event #1.
- Token-fetch failure (`getFcmAccessToken` throws) now records exactly
  one health-failure (not one per device), short-circuits the send
  loop, and returns a result so `sendToNamespace` no longer leaks the
  exception.
- `invalid` token responses are explicitly excluded from the health
  buffer because they are per-device facts (rotated/uninstalled token),
  not pipeline failures - FCM was reachable, it just rejected one
  stale token.
- `buildNativeFallbackProbe` now optionally accepts the FcmService and
  short-circuits to "let web-push fire" when health is bad, before it
  even queries the device registry. The single-arg call shape is still
  supported for back-compat.

Why not the bot's exact suggestion ("invert: call FCM first, fall back
on result.sent === 0"):
- Couples PushNotificationChannel to FcmService and FcmSendPayload,
  reversing the clean parallel-channel architecture established earlier
  in this PR.
- Treats every transient single-event failure as fallback-worthy, which
  re-opens the duplicate-notification race that the suppression logic
  was added to close (FCM HTTP timeout that delivers later + the web
  push we sent in the meantime = two pings).
- A rolling health window only flips on sustained breakage, which is
  the actual operational scenario the bot is worried about.

The wrist-first design intent ("FCM fires unconditionally, web-push is
suppressed for the same namespace") documented in
docs/api/native-companion-contract.md is preserved on the happy path.
The probe only re-enables web-push when there is concrete evidence the
native pipeline is not delivering.

Tests:
- New FcmService.isHealthy suite covers empty-buffer, threshold flip,
  recovery as failures age out of the window, invalid-token exclusion,
  and network-error path.
- nativeFallbackProbe gains coverage for the unhealthy-but-registered,
  healthy-and-registered, and absent-fcmService (back-compat) cases.
- All 292 hub tests still pass; typecheck clean.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(telegram): drop duplicate tool-args formatter, use shared module

The Telegram session view had its own copy of formatToolArgumentsDetailed
identical to the one in hub/src/notifications/toolArgs.ts (already used by
the FCM channel). Replace the local copy with an import.

Removes ~70 lines of duplication, plus the now-unused MAX_TOOL_ARGS_LENGTH
constant and `truncate` import. The shared signature accepts an optional
opts arg whose default maxArgLength is 150 - matching the prior constant -
so the call site is unchanged.

Two benign upgrades come along for the ride from the shared module:
?? instead of || on field fallbacks (no real-world difference; permission
arguments never carry empty-string fields), and String(...) wrapping plus
a typeof object guard that makes non-string values render gracefully
instead of throwing into the catch block.

Hub tests: 311 pass / 0 fail. Telegram subset: 5 pass / 0 fail. typecheck
green.

Cold-reviewed by an out-of-context Claude Opus peer before push.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): require positive evidence in health window before suppressing web-push

Addresses HAPI Bot Major review on PR #803.

The previous health gate treated an empty outcome buffer as healthy
("innocent until proven guilty"). That created a silent-blackhole window
on cold start with broken FCM credentials: the push channel suppressed
SSE/Web Push for the first ~5 events while the FCM channel attempted
each delivery and recorded failures, until enough stacked to flip the
threshold. Every notification in that gap was silently lost.

New invariant: isHealthy() requires at least one successful FCM send in
the recent window (HEALTH_WINDOW=8) AND failures below threshold
(HEALTH_FAILURE_THRESHOLD=5). Both conditions are necessary; either
alone is insufficient evidence to safely suppress web-push fallback.

Trade-off: one duplicated notification per hub restart per namespace.
On the first event after restart, web-push fires alongside FCM (because
the gate has no positive evidence yet). Once FCM records that first
success, the gate engages and subsequent events are FCM-only. Worth it
for guaranteed delivery during cold-start outages.

Tests reworked to match new semantics:
- "starts UNHEALTHY with empty buffer" (was: healthy)
- "flips to healthy after first successful send" (new)
- "stays unhealthy across failures-only run" (new, exercises the exact
  blackhole scenario the bot flagged)
- "flips back to unhealthy after threshold breach with prior successes"
  (renamed, establishes successes first)
- "invalid tokens don't count against health" (reworked: send a mixed
  batch first to establish health, then verify invalids don't flip it)
- "network errors count as failures" (reworked: establish health first)

Hub tests: 313 pass / 0 fail. typecheck green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): bump FCM migration to V10→V11 after upstream service_tier V9→V10

Upstream/main landed sessions.service_tier at schema v10. The companion
FCM device registry now migrates at v11 so both changes compose cleanly
after the courtesy rebase onto current upstream/main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): per-dispatch native gate instead of stale FCM probe

FCM runs before web-push; PushNotificationChannel skips web/SSE only
when the same notify() dispatch already delivered via FCM. Removes the
isHealthy()+device-row probe that could suppress web-push after warm
FCM outages.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,web): cap notifySummary for FCM limits; fix PWA test cast

Rebase follow-up: truncate AGENT_NOTIFY_SUMMARY summary/action before
FCM data payload (bot Major). Fix usePwaUpdate.test.ts setTimeout mock
cast so bun typecheck passes on current main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): cap all FCM notifySummary fields and task bodies

Whitelist and truncate AGENT_NOTIFY_SUMMARY auxiliary fields before
JSON serialization; cap task-notification summaries to glance limit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): FCM fetch timeouts and cap Grep/Glob permission args

10s AbortSignal.timeout on OAuth + FCM send so sequential web-push
fallback is not blocked on hung Google endpoints; truncate Grep/Glob
pattern in permission detail formatter.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): bind FCM token to one namespace on re-pair

Delete stale fcm_devices rows sharing the same token when a native
install registers under a different namespace.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): localize Companion settings and pairing copy

Add en/zh-CN keys for the Companion section title and CompanionPairing
strings; matches locale-driven Settings pattern (bot Minor on #803).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): tighten FCM token-invalid detection and truncation edge cases

Parse FCM error JSON: only UNREGISTERED or token-field INVALID_ARGUMENT
unregister devices; generic NOT_FOUND stays transient. Guard limit<=3
in truncateReadyText so tiny action budgets cannot blow the glance cap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): parse FcmError details.errorCode for UNREGISTERED tokens

FCM v1 often returns HTTP 404 with root NOT_FOUND plus
details[].errorCode UNREGISTERED; prune those tokens while keeping
generic project/resource NOT_FOUND transient.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): mock AppContext for About Companion pairing in settings tests

Settings About now mounts CompanionPairing via useAppContext after the
#1027 hub redesign rebase; wrap the About route test with AppContext and
Companion mocks so the suite stays green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): point companion auth at POST /api/auth, not /api/bind

Pairing QR carries the CLI access token as `code`. /api/bind requires
Telegram initData; native companions must use /api/auth with accessToken.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): mount Companion pairing under Settings General

About is version/links only after the settings hub redesign; pairing is
setup, so keep Companion with language prefs and update the route tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Debian <heavygee@oos-linux.in.lockhouse>
2026-07-27 19:52:54 +08:00
SSU-WEI HUANGandGitHub 173f855b73 docs: remove sunset Gemini CLI launch references (#1132) 2026-07-23 08:42:11 +08:00
SSU-WEI HUANGandGitHub c87720ab4d fix(cli): load extra headers from settings (#1041)
* test: reproduce issue #786

* fix: load extra headers from settings (closes #786)

* test: cover extra header precedence and redaction

* fix: redact persisted extra headers in diagnostics

* test: cover runner extra header identity

* fix: restart runner when extra headers change
2026-07-16 12:27:50 +08:00
SSU-WEI HUANGandGitHub b9eed7c071 feat: add Grok Build support (#1030)
* test: define Grok Build integration behavior

* feat: add Grok Build agent integration

* test: cover Grok permissions and resume paths

* docs: add Grok Build setup guide

* fix: scope Grok ACP discovery to session cwd

* fix: align Grok permission UI semantics

* docs: clarify Grok runner setup

* test: require Grok create model and effort options

* feat: add Grok create model and effort pickers

* test: define Grok runtime parity behavior

* feat: add Grok runtime ACP controls and discovery

* fix: tighten Grok runtime controls

* fix: suppress nonfatal Grok title quota errors

* feat: support Grok Auto permission mode

* feat: forward ACP native session titles for Grok

* fix: guard Grok Windows shell arguments
2026-07-13 08:41:30 +08:00
73584e925a feat(cursor): multitask slash, autoReview mode, native worktree/add-dir (#1014)
* feat(cursor): multitask slash, autoReview mode, native worktree/add-dir

Close the highest-value Cursor Agent gaps for remote HAPI: expand ACP-safe
slash pass-through (/multitask, worktree, add-dir, …), add autoReview
permission mode (--auto-review spawn + mid-session slash), and route Cursor
New Session worktrees through agent --worktree instead of HAPI sibling trees.

Fixes #1013

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): accept --mode autoReview for hapi cursor

Align --mode parsing with CURSOR_PERMISSION_MODES so documented
`hapi cursor --mode autoReview` enables Smart Auto instead of silently
falling back to default.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-12 18:41:52 +08:00
65e1708c78 feat(web): mermaid diagram lightbox on click (#741)
* feat(web): mermaid diagram lightbox on click

Click rendered mermaid blocks in chat to open a zoomable full-screen viewer.
Re-renders from source in the modal with the current theme. Closes #737.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): fit mermaid lightbox to viewport on open

Auto-scale diagrams to fill the viewer instead of opening at intrinsic
mermaid size. Reset returns to fit; zoom label is relative to fit (100%).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): fit mermaid lightbox to device screen not inner panel

Use visualViewport for fit scale, full-screen pan layer, and a floating
toolbar so the diagram can use the whole display.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): show mermaid lightbox by reusing inline SVG

Second mermaid.render on open often left a 0×0 SVG while fit scale was
computed from the loading placeholder. Reuse the inline SVG in the modal
and measure viewBox with retried fit-to-screen.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): uniquify mermaid SVG ids in lightbox clone

Inlining the same mermaid markup twice duplicates element ids and breaks
url(#ref) resolution in the modal copy. Prefix ids and hrefs for lightbox only.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): give mermaid lightbox SVG explicit dimensions

Mermaid emits width="100%" with max-width in px; that collapses to 0×0
inside the centered lightbox layer. Derive width/height from viewBox for
the uniquified lightbox clone.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): render mermaid lightbox via isolated SVG data URL

String id rewrites broke mermaid's embedded CSS so only labels appeared
zoomed. Rasterize the inline SVG to a data-URL img instead of duplicating
markup in the DOM.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): lightbox re-renders SVG for sequence diagrams

Data-URL images drop or blank some mermaid diagram types (sequence).
Re-render with a modal-specific id into inline SVG on a code-bg panel,
and add sequence theme variables for dark/light.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): mermaid lightbox uses inline SVG in shadow DOM

Reuse the inline render in an isolated shadow root so sequence CSS stays
intact, and fit the viewport from viewBox dimensions instead of the loading
placeholder or width="100%" layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): Playwright lightbox coverage per mermaid diagram type

Add e2e harness and a script that opens the lightbox for each diagram
kind (flowchart through kanban). Fit uses inline getBBox() so compact
charts like gitGraph fill the viewport.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): bounded Playwright via webServer, fix gantt fit sizing

Playwright owns Vite lifecycle (no agent-spawned dev server). Fit uses
viewBox unless viewBox padding is excessive (gitGraph); wide charts use
width-based coverage in e2e.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(web): gitignore Playwright test-results

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): address PR 741 bot feedback (typecheck, fit floor, gitignore)

Guard lightbox open when svg is null; allow fit scale down to 0.01 while
keeping 0.25 minimum for manual zoom; ignore Playwright test-results/ correctly.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): Playwright asserts click expands diagram vs inline

Measure inline vs lightbox bounding box after click; require visible
growth (area ratio or max dimension) plus dialog + shadow SVG content.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(web): Playwright against live HAPI session for mermaid lightbox

Add seed script for a dedicated chat session, live hub Playwright suite
(HAPI_LIVE=1), and dogfood doc. Live tests fail until driver serves shadow-DOM
lightbox (catches gray-box regression on stale bundles).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): undo wrapper transform in lightbox fit; carry fit floor in zoom

Resolves PR #741 review threads (HAPI Bot Major):

1. measureSvgIntrinsicSize / measureContentSize prefer intrinsic dimensions
   (viewBox -> width/height attrs -> img.naturalSize) before getBoundingClientRect.
   When the rect is the only signal, divide by scaleRef.current so the 50/200ms
   refit retries stop compounding with the wrapper's scale(...) transform.
   Large diagrams no longer jump tiny or oversize after async render completes.

2. Interactive zoom (wheel/keys/buttons/pinch) now clamps with
   Math.min(MIN_SCALE, baseScaleRef.current). A diagram fitted below the
   normal 25% floor stays reachable instead of snapping back to 25% and
   clipping. Zoom-out button disabled threshold uses the same min.

3. Add Vitest coverage for both helpers (intrinsic precedence, scale-aware
   rect fallback, divide-by-zero guard) so regressions surface without
   needing the full Playwright stack.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(scripts): mermaid seed refuses to wipe non-fixture sessions

HAPI Bot Major (PR #741): SESSION_ID is documented as overridable,
and the script unconditionally deletes every message for the target
session before seeding fixtures. If pointed at a real session id,
that's silent data loss.

Refuse to proceed when an existing session id has a tag other than
'mermaid-lightbox-e2e'. New ids and the canonical fixture session
still seed normally; real sessions throw before any DELETE runs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): normalize mermaid svg for lightbox shadow root

Mermaid emits width="100%" on every diagram. Inside a shadow root whose
host has no explicit size, that collapses to zero in Chromium for most
diagram types - only ones that ship pixel attrs (e.g. journey) happen to
render. Operator confirmed on the live driver: every diagram except
journey opened to a grey rounded square.

MermaidLightboxSvg now runs normalizeMermaidSvgForStandaloneDisplay before
injecting (strips width/height="100%", bakes viewBox dims as pixels) and
sets :host{display:inline-block} so the host sizes to the SVG. Inline svg
in chat is unchanged - only the lightbox copy is normalized.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): keep mermaid lightbox content below the toolbar

Operator screenshot showed the diagram top (e.g. pie 'Pets' title)
clipped behind the toolbar bar. Two causes:

1. getScreenFitSize used the full viewport height, so the fit scale
   sized the diagram to fill an area the toolbar overlapped.
2. The viewport (drag/zoom area) was inset-0; content centered on the
   full viewport center, not the visible region's center, pushing the
   top behind the toolbar.

Measure the toolbar with a ResizeObserver, subtract its height from
the fit calculation (clamped at zero), and start the viewport region
below the toolbar (top: toolbarHeight). Fit scale recomputes whenever
toolbar height changes.

Adds Vitest coverage for getScreenFitSize reserved-top math.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): guard ResizeObserver before constructing it

HAPI Bot Major (PR #741): Vitest jsdom does not polyfill ResizeObserver,
so the toolbar measure effect throws ReferenceError when the existing
mermaid-diagram React tests open the lightbox. Same code path is also
brittle in any browser/webview without the API.

Fall back to plain window 'resize' listener when ResizeObserver is
absent. Toolbar height won't auto-update on element resize without it,
but the lightbox still renders and the resize listener catches the
common viewport-rotation case.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(scripts): live mermaid playwright wrapper runs from repo root

HAPI Bot Minor (PR #741): the wrapper sets cwd to scripts/, but the
test:mermaid-lightbox:live npm script lives in the repo-root
package.json, so spawning npm there exited before Playwright started.
Switch cwd to the repo root and drop the unused WEB_DIR constant.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): accept signed viewBox values in mermaid lightbox normalize

HAPI Bot Minor (PR #741): the viewBox regex only matched digits, dots,
and spaces, so a valid viewBox with negative origin (e.g. '-8 -8 640 480')
returned null. normalizeMermaidSvgForStandaloneDisplay then became a
no-op and left width='100%', re-introducing the zero-sized lightbox
render this PR is meant to fix for the affected diagrams.

Switch to the bot's suggested regex (signed numbers, single or double
quotes, comma or space separators) and reject NaN parts. Adds Vitest
coverage for signed origins, single quotes, comma separators, the
malformed/no-viewBox null paths, and an end-to-end normalize test that
fails against the old regex.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): align @playwright/test on 1.60.0 across workspaces

HAPI Bot Major (PR #741): web/package.json pinned @playwright/test at
1.49.1 while the root workspace and bun.lock were on 1.60.0. The
mismatch surfaced after rebasing onto upstream/main, where the root had
already moved to 1.60.0 while my web devDependency lagged from an older
commit. A frozen install would reject the lockfile and the new web e2e
script could resolve a different Playwright than root scripts.

Bump the web devDependency to 1.60.0 and regenerate bun.lock so all
workspaces share one Playwright version.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): move mermaid playwright fixtures out of public

HAPI Bot Minor (PR #741): the e2e and smoke fixtures lived under
web/public, so Vite copied them verbatim into web/dist and the hub
asset generator embedded them in production bundles. Both pages
import Vite dev-only paths (/@react-refresh and /src/dev/...), so
the production /mermaid-lightbox-{e2e,smoke}.html routes would 404
on those imports.

Move both fixtures to web/e2e-fixtures/ to match the existing
scratchlist-fixture pattern (relative ../src/dev import, served by
Vite at /e2e-fixtures/...) and update the Playwright spec to hit the
new path. Build now ships 112 PWA precache entries instead of 114
(both fixtures excluded from dist).

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-11 11:07:12 +08:00
5f27abddd4 feat(web): in-app PWA update prompt when new service worker is available (#946)
* feat(web): in-app PWA update prompt when new service worker is available (closes #938)

User-controlled reload with a persistent banner, visibility-triggered SW
checks, and an expandable rationale. Switches registerType to prompt.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): align vite.config with soup layers for clean driver merge

Keeps registerType prompt while matching garden IWER stubs and PWA
share_target shape expected by feat/pwa-share-target in the manifest.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Revert "fix(web): align vite.config with soup layers for clean driver merge"

This reverts commit 6f0915b0884d029a2413d8819a4dfe81d7c4e595.

* fix(web): make PWA reload apply waiting service worker updates

Handle SKIP_WAITING in injectManifest sw.ts and reload via controllerchange
with a timed fallback when vite-plugin-pwa prompt mode does not navigate.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): satisfy setTimeout mock typing in PWA reload tests

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): register PWA service worker before auth gates

Mount PwaUpdateProvider at app root and show the update banner on login
and error screens so registerSW runs for logged-out users too.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): offset PWA update banner below top status banners

Reserve top-12 when syncing or reconnecting so the reload prompt stays
visible above SyncingBanner and ReconnectingBanner.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): offset PWA update banner below voice error banner

Use PwaUpdateBannerWithStatusOffset inside VoiceProvider so voice errors
share the same top-12 reservation as sync and reconnect banners.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-18 10:15:10 +08:00
a2862a3300 docs(installation): add KillMode=process to runner systemd unit (closes #915) (#928)
The runner spawns child agent sessions with `detached: true`
(`cli/src/runner/run.ts:454`) so they survive runner restart, and
runner cleanup (`run.ts:1049`) does not iterate or kill tracked
children on shutdown. The runner is already designed as a long-lived
process whose exit leaves agent sessions intact.

But Node's `detached: true` calls `setsid()` (new process session),
which does NOT escape the parent's systemd cgroup. Without an
explicit `KillMode`, systemd defaults to `control-group`, which
SIGTERMs every PID in the runner's cgroup whenever the unit stops -
forcibly archiving every running session and discarding the detach
contract.

Adds `KillMode=process` to the reference runner unit and a note
explaining the contract. With this change, `systemctl restart
hapi-runner.service` (and any cascade-stop from `Requires=`) only
signals the main runner PID; the cleanup runs without killing
descendants; agent sessions stay alive; the new runner reconnects via
the existing socket.io reconnect path
(`cli/src/api/apiMachine.ts:385`) and re-establishes control via the
existing RPC layer.

This is the smallest fix for #915. The complementary safety net -
runner re-attaching to orphaned children on cold start when no
running runner exists - will be tracked in a separate issue and PR.

AI-disclosure (per CONTRIBUTING.md): drafted with claude-opus-4.7 as
peer agent during a fork-side post-mortem of a 7-hour outage that
this fix would have prevented.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-18 10:12:12 +08:00
HeavyGeeandGitHub 6d2d0d4707 fix(cursor): trim #784 safety patch to marker-only on legacy stream-json path (closes #822) (#828)
* fix(cursor): drop timing heuristic from #784 intercept; scan raw payload (#801 follow-up)

PR #801 shipped a two-strategy intercept for the synthetic AskQuestion
skip response in legacy stream-json mode. Real-traffic data from a
post-merge run shows the marker-match strategy never fires (the
converter's `extractToolResult` discards the marker for tool shapes it
does not recognize, returning `{}`) and the timing-signature
defense-in-depth strategy fires only on false positives - notably the
Anthropic Vertex Claude tool calls cursor-agent surfaces in legacy
sessions, which all land as `name=unknown` with the `{}` extracted
result and frequently complete under the 500 ms threshold.

Measured on a single legacy-resumed session (`7b769423`): 1,136
`name=unknown` tool calls, 16 rewritten as `no_input_surface`, zero
actual marker strings stored anywhere in the session. The 16 rewrites
were legitimate fast tool calls (Anthropic Vertex `toolu_vrtx_*` IDs)
mischaracterized as fabricated skip responses.

Changes:
- Remove the timing-signature heuristic and its supporting state
  (started-at map, elapsed-ms calculation, latency threshold, test-only
  state reset).
- Move the marker scan from the post-`extractToolResult` output to the
  raw `tool_call` payload, so it can see the marker on stream-json
  shapes the converter does not specifically recognize. Function-shaped
  tools exclude `function.arguments` from the scan to avoid matching
  agent-controlled input. Other shapes scan the full payload (no
  agent-input field exists at the top level).
- Refresh tests: drop timing-based positive cases, add a marker-in-raw-
  payload positive case for `name=unknown` shapes, and add a regression
  that legitimate fast `name=unknown` tool calls without the marker
  pass through with `status: completed`.
- Document scope: this intercept now lives only on the legacy stream-
  json path, which only resumed pre-ACP sessions hit. New cursor remote
  sessions go through `cursorAcpBackend` and the `cursor/ask_question`
  ACP extension method (#799) - immune to this bug. The intercept
  drains with the legacy session population.

Tracking: #784. Builds on #801, complements #799.

* fix(cursor): exclude agent input from marker scan; surface top-level Anthropic tool names (Codex P2)

Codex flagged a false-positive case on the fork-stage review of this
branch (heavygee/hapi#35, P2): an Anthropic tool_use shape with a
top-level `name` (e.g. `{id, name: 'TodoWrite', input: { ... }}`) gets
labelled `name=unknown` by the converter and passes the AskQuestion
gate. If the agent's `input` quotes the synthetic-skip marker - which
happens whenever an agent edits or documents this very bug - the
intercept would rewrite a perfectly fine TodoWrite as a fabricated
skip.

Two-part fix:

1. `extractToolName` now reads the top-level `name` field as a final
   fallback. A real `TodoWrite` / `Bash` / `str_replace_based_edit_tool`
   surfaces with its actual name and is rejected by the AskQuestion
   gate before the marker scan runs. The original AskQuestion
   fabrication case still surfaces as `unknown` (per #784 issue body
   the name is stripped in the fabricated payload) and remains
   detectable.

2. Defense in depth: introduce `AGENT_INPUT_KEYS = {input, args,
   arguments}` and exclude these from the non-function shape's marker
   scan. Even if a tool reaches this code path with `name=unknown` and
   the marker buried in its `input`, the intercept won't fire on agent-
   controlled text.

Two new regression tests:

- Anthropic tool_use shape `{id, name: 'TodoWrite', input: {todos: [
  marker]}}` → passes through with `status: 'completed'`.
- `name=unknown` shape with marker only inside `input` → passes through
  with `status: 'completed'`.

All 20/20 tests pass; typecheck clean (cli + web + hub).
2026-06-08 13:30:04 +08:00
3a8693f380 feat(cursor): migrate remote sessions to ACP with model/variant pickers (#799)
* feat(cli,web,hub): migrate Cursor remote sessions to ACP with model/effort pickers

Move stream-json remote launcher to legacy path and add ACP launcher with
set_config_option model/mode sync, optimistic keepalive on config changes, and
shared catalog caching. Web gets dual base/effort Cursor pickers for session and
new-session flows; hide composer status bar when Cursor sends no usage_update.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,web,shared): Cursor model picker — ACP wires + CLI sku variants

Enrich the web/mobile picker with agent --list-models SKUs grouped under
ACP wire bases, fix session-open base highlight, and keep catalog discovery
safe while the ACP transport holds the CLI lock.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cursor-acp): apply ACP default model when web resets to Default

Web sends model: null for Default; push session/set_config_option with the
ACP default[] wire so Cursor backend matches hub state. Regression tests
for setModel(null) and applyModelConfig(null).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(acp): clear stale agent-acp lock when owning process is gone

Check lock pid with signal 0; remove orphaned lock dirs after SIGKILL or
crash so listCursorModels can run cold probes again. Regression tests for
guard and catalog discovery.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(cursor): use live pid for ACP lock handler tests

Stale-lock cleanup clears dead pids; handler tests must simulate an
active lock with the current process pid to avoid cold probes/timeouts.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(acp): scope agent CLI lock guard to Cursor agent command only

Gemini/OpenCode/Kimi ACP sessions must not register agent-acp-active;
that blocked listCursorModels while unrelated backends were running.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,web): reject Cursor model changes for local sessions

Hub returns 409 when controlledByUser is set, matching Codex. Web hides
model and variant pickers for local Cursor sessions so users do not hit
a dead RPC path. Document pre-push-review in AGENTS.md.

Verified: bun typecheck; bun run test (919 cli + 243 hub + 768 web + 46 shared).
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): send stable ids for Cursor ask_question replies

Parse and submit question.id and option.id so ACP receives keys like
{ approach: ['a'] } instead of index/label. Verified: bun typecheck && bun run test.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 19:51:35 +08:00
HeavyGeeandGitHub dc0d21e05b fix(cursor): intercept fabricated Questions skipped AskQuestion result in headless mode (#784) (#801)
* fix(cursor): intercept fabricated 'Questions skipped' AskQuestion result in headless mode (#784)

When cursor-agent runs under `--print --output-format stream-json` (HAPI's
current Cursor remote launcher), the CLI returns a synthetic
`Questions skipped by the user, continue with the information you already have`
response for the `AskQuestion` tool in ~zero seconds with no error flag,
because there is no IDE surface to render the question. The underlying
model can interpret this as legitimate user consent and act on it.

This patch intercepts the synthetic result in
`cli/src/cursor/utils/cursorEventConverter.ts` and rewrites the
`tool_call`/completed event to a structured `no_input_surface` failure
(`status: 'failed'`, which downstream becomes `is_error: true`).

Detection has two strategies:

1. String match - any `tool_call`/completed payload whose serialized form
   contains the synthetic-skip marker is rewritten. This is robust to
   wherever cursor-agent stuffs the marker inside the `tool_call` object.
2. Timing + name heuristic (defense in depth) - any completion that arrives
   within 500 ms of its 'started' event with a trivial result, for a tool
   call named `AskQuestion`, `askQuestion`, `ask_question`, or the
   converter's `unknown` fallback, is also rewritten. This catches the case
   where cursor-agent changes the synthetic-string text in a future release.

The converter tracks per-call timestamps in a bounded `Map` (`<= 1024`
entries, oldest evicted on overflow) and clears entries when the
corresponding 'completed' event arrives. A small test-only reset hook
isolates state between Vitest cases.

This is a transitional safety patch. It auto-deletes when #781's ACP
launcher replaces the stream-json launcher and `cursor/ask_question`
becomes a proper bidirectional ACP method where fabrication is
structurally impossible.

Scope is intentionally tiny: only `cli/src/cursor/utils/cursorEventConverter.ts`,
its colocated Vitest file, and a section in `docs/guide/cursor.md`. No
changes to `cursorRemoteLauncher.ts`, ACP code, web normalizer, or
permission UI.

Refs: tiann/hapi#781 (long-term resolution via ACP migration)
Closes: tiann/hapi#784

* fix(cursor): gate AskQuestion intercept on tool name (#784 PR #801 review)

Address regression flagged by the HAPI auto-review bot on #801:

`containsSyntheticSkipMarker` previously stringified the entire `tool_call`
payload and matched the literal marker substring. Because this PR also adds
that exact marker to `docs/guide/cursor.md` (to document the intercept), a
Cursor `read_file` of that documentation page would surface the marker
inside `readToolCall.result.content` and be rewritten as a
`no_input_surface` failure, corrupting an unrelated, legitimate result.

The intercept is now gated on the tool name resolving to an
AskQuestion-shaped call (`AskQuestion`, `askQuestion`, `ask_question`, or
the converter's `unknown` fallback for unnamed function-shaped tools).
`read_file` / `write_file` tool calls - which have explicit `read_file`
and `write_file` names from `extractToolName` - no longer fall under the
intercept, regardless of what their payload contains.

The marker check itself now walks values recursively (string / array /
object), guarded by a `WeakSet` against cycles, instead of relying on
`JSON.stringify`. Slightly tidier; behaviour is otherwise unchanged for
the AskQuestion path.

Regression tests added:

- `read_file` result whose `content` contains the marker -> passes
  through with `status: 'completed'` and no `no_input_surface`.
- `write_file` whose serialized `args` contain the marker -> same.
- A non-AskQuestion function tool (`MyCustomTool`) whose result quotes
  the marker -> same.

All 846 cli tests pass (17 in this file). `bun run typecheck` exits 0.

* fix(cursor): scope synthetic-skip check to extracted result (#784 PR #801 review-2)

Address second Major finding from the HAPI auto-review bot on #801:

After the previous fix gated the intercept on the tool name, the marker
check still recursed into the entire `tool_call` object - which includes
`function.arguments`, the agent's own prompt text. A legitimate
AskQuestion whose prompt quotes the synthetic-skip marker (e.g. an agent
debugging this exact bug, or any prompt that pastes the marker verbatim)
would have been rewritten as `no_input_surface` even when the operator
actually answered.

Changes:

1. `extractToolResult` now extracts the cursor-side response from
   function-shaped tool calls. Previously it returned `{}` for anything
   that wasn't `readToolCall` or `writeToolCall`. It now returns
   `function.result` when present, otherwise every field of `function`
   except `name` and `arguments`. This excludes the agent's input from
   what downstream sees as the tool result, and as a side effect surfaces
   the actual cursor response for function-shaped tools (which was
   previously lost - see the #784 incident note about HAPI storing
   `output: {}` for AskQuestion in the message DB).

2. `shouldRewriteAsNoInputSurface` now searches only the extracted
   `result`, not the whole `tool_call`. The bot's exact recommendation.

3. Test added: an AskQuestion whose `arguments` quote the marker but
   whose `result` is a real user answer, with elapsed time past the
   500 ms threshold so the timing heuristic does not apply. Asserts the
   tool_result passes through with `status: 'completed'` and the
   operator's actual answer.

All 847 cli tests pass (18 in `cursorEventConverter.test.ts`).
`bun run typecheck` exits 0.

The widened `extractToolResult` scope is necessary for the marker check
to actually find the synthetic string (it lives inside `function.result`
or a sibling field), and is the bot's explicit recommendation. It also
removes the long-standing data-loss bug where AskQuestion responses were
surfaced to the message DB as opaque `{}` - regardless of fabrication.
2026-06-05 21:44:37 +08:00
lekoandGitHub db934a9d5b Update hapi runner command to use start-sync (#685)
* Update hapi runner command to use start-sync

* Update installation instructions for start command
2026-05-26 16:13:25 +08:00
weishu 35044cde50 remove unused docs 2026-05-22 09:01:14 +08:00
lekoandGitHub ce2e76a42e Add Windows remote terminal support (#642) 2026-05-19 07:54:12 +08:00