Commit Graph
11 Commits
Author SHA1 Message Date
AnanovoandGitHub 17ee052d9a fix(web): preserve the visible chat window during rewind (#1766)
* fix(web): preserve chat window during rewind

* fix(web): scope rewind invalidation preservation

* fix(web): clear unknown rewind boundaries

* fix(web): deduplicate rewind invalidations

* fix(web): retain rewind dedupe history
2026-09-09 09:23:31 +08:00
weishu e5a8212f4a feat(session): validate agents and browse workspace directories 2026-08-25 16:12:29 +08:00
SSU-WEI HUANGandGitHub f0e5ba9c0f feat(codex): mid-turn Steer via app-server turn/steer (#888) (#1606)
* feat(shared): steer capability gates and live steered signal schemas

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession gate which
  agents can deliver queued messages into the active turn (pi, codex,
  cursor ACP; legacy stream-json cursor excluded)
- AgentState.steeringActive, DecryptedMessage.steered and
  messages-consumed  live signal (never persisted by the hub)

* feat(cli): queue reservations and steered messages-consumed option

- MessageQueue2 gains takeByLocalId/restoreReservation/
  beginReservationDispatch/commitReservation so an async steer can reserve
  a queued row without racing the main loop's turn/start drain
- emitMessagesConsumed accepts steered: true to mark mid-turn delivery

* feat(codex): mid-turn steer via app-server turn/steer (#888)

- CodexAppServerClient.steerTurn + TurnSteerParams/Response types
- CodexRemoteLauncher registers the steer-queued-message RPC handler:
  reserves the queued row, validates it against the active turn (no
  control commands, matching mode hash), injects via turn/steer with an
  epoch guard that invalidates in-flight steers on abort/cleanup
- steeringActive agent state tracks the active-turn window
- hub syncEngine gate opens to codex; messages-consumed relays steered

* feat(web): Steered badge and steer gating for codex sessions

- HappyUserMessage shows a ↳ Steered badge fed by the live
  messages-consumed steered signal, preserved across server echoes and
  refetches (mergeMessages carries the optimistic marker)
- SessionChat gates canSteer via isSteeringSupportedForSession instead of
  the pi-only check
- clearStaleQueuedStatus normalizes a queued status on an invoked message
- fix(web): drop duplicate showSessionSummaryInChat in markdown test
  (upstream typecheck breakage)

* fix(codex,shared): address bot findings on steer gate and ambiguous turn/steer

- STEERING_SUPPORTED_FLAVORS / isSteeringSupportedForSession advertise
  codex and pi only; cursor joins when its soft-steer handler lands (#1609)
- turn/steer now splits dispatch (stdin accepted) from completion (turn
  finished): the hub RPC acks once dispatch succeeds — never on the
  concurrent turn's completion, which can exceed the 30s RPC window
- queue row commits only after the turn settles; a rejected/aborted steer
  restores the row so the message still delivers via turn/start, and a
  dispatched steer is never restored (no duplicate delivery)
- steer carries clientUserMessageId (echoed as userMessage.clientId) so
  ambiguous transport failures can reconcile the thread later
- client tests cover dispatch/complete split and stdin-write failure

* fix(codex): reconcile dispatched steers before restoring; align error copy

- A dispatched turn/steer whose completion fails (disconnect / protocol
  error) is now reconciled via thread/read by clientUserMessageId before
  the queued row is restored — the instruction is only re-delivered by
  turn/start when the thread never received it
- Reconcile targets the pinned steer thread, not whichever turn is
  current when completion fails
- syncEngine unsupported-flavor error now matches the capability gate
  (Pi and Codex only until the cursor handler lands)
- launcher tests cover steer success (ack on dispatch), reconcile-accepted
  and reconcile-rejected outcomes

* fix(codex): consume the row at dispatch; drop background reconcile

- The hub RPC acks and the queue row is consumed as soon as stdin accepts
  turn/steer; completion is background-only logging. A dispatched steer is
  never restored, so the same localId cannot be re-delivered via turn/start
  after the caller was told the steer succeeded
- Dispatch failure (stdin write error) still restores the row and reports
  failure
- steer.completed rejection is always handled (no unhandled rejection on
  the dispatch-failure path)
- tests updated: completion failure after dispatch keeps the row consumed;
  dispatch failure restores it

* fix(codex): distinguish definite rejection from indeterminate completion

- Transport-level failures (timeout, abort, disconnect, spawn, protocol)
  carry an indeterminate marker; explicit JSON-RPC error responses do not
- After a dispatched steer, turn completion resolves → commit + consumed;
  a definite app-server rejection restores the row (instruction was never
  accepted, so turn/start cannot duplicate it); an indeterminate outcome
  leaves the row reserved so it can never be delivered twice
- Completion handling registers before awaiting dispatch so the
  dispatch-failure path cannot leak an unhandled rejection
- client/launcher tests cover explicit rejection (restore), indeterminate
  outcome (row stays reserved) and dispatch failure

* fix(codex): reconcile indeterminate steers instead of a permanent reservation

- After an indeterminate completion (disconnect/protocol), reconcile the
  thread by clientUserMessageId immediately: accepted → commit + consumed,
  provably rejected → restore, still unreadable → keep the reservation and
  retry from the main-loop top on later passes (post-reconnect)
- A row never sits in dispatching forever: the hub cannot stamp it invoked
  while the instruction may never have been accepted
- tests: indeterminate keeps reserved while thread unreadable; accepted
  reconciliation consumes; rejected path restores

* fix(codex): accept all thread item shapes; retry reconcile; ack through abort

- Reconcile matcher accepts userMessage/user_message with clientId/
  client_id, matching the shapes the thread parser supports — an accepted
  steer can no longer be misclassified as rejected
- A pending reconciliation schedules a wakeLoop retry, so a temporary
  app-server outage cannot strand the reservation behind waitForTurnOrRecovery
- The success-path ACK no longer checks the steer epoch: the hub already
  reported steered on dispatch, so commit + messages-consumed must reach
  it even when an abort resets the queue in between

* fix(codex): reinit reconnected app-server; keep reconcile retries alive

- thread/read after a disconnect auto-connects a fresh app-server, which
  must be initialized before any request — reconcile now ensures
  connect + initialize (isConnected getter added to the client)
- every still-unknown loop-top reconciliation schedules the next retry,
  so recovery without external traffic is eventually observed
- launcher mock gains isConnected

* fix(codex): timer-driven reconciliation; init tracking; abort-safe ACK

- Reconciliation runs on a self-rescheduling 1s timer independent of the
  main loop (wakes it too), so idle loops and waitForTurnOrRecovery still
  observe app-server recovery; abort clears nothing implicitly — the ACK
  path commits and consumes even when the reservation was cancelled
- Absence of a durable client id is ambiguous: unmatched reads stay
  'unknown' and keep retrying instead of restoring the row
- CodexAppServerClient tracks initialized state (reset on disconnect/exit)
  so ensureAppServerInitialized re-initializes a fresh process before
  thread/read; initialize failures leave the flag false for the next retry
- tests: accepted reconciliation via scheduled timer, indeterminate
  keeps reserved, explicit rejection restores

* fix(codex): bind reconciliation to the launcher lifecycle

- runSteerReconciliation clears any armed retry timer on entry and never
  installs a second one, so loop-top and timer-driven passes cannot
  multiply
- shuttingDown is set when the main loop ends: timers are cleared and the
  pending map is dropped, so an unresolved steer can never respawn an
  app-server after cleanup (remote-to-local switch included)

* fix(codex): report steered only after app-server acceptance

- The handler now awaits steer.completed (the inject-acceptance response):
  an explicit JSON-RPC rejection surfaces as failed and restores the row
  for the normal turn/start path instead of a false steered
- Transport failure after dispatch reports 'Steer outcome is being
  reconciled' and keeps the row reserved while the timer-driven thread
  reconciliation runs
- dispatch-failure path also swallows the paired completion rejection

* fix(steer): tri-state cancel, clear-safe reservations, bounded acceptance wait

- MessageQueue2.cancelByLocalId returns 'in-flight' for a dispatching
  steer reservation: the hub neither deletes the row nor stamps invoked_at
  (new CancelMessageResponse 'busy' status; web restores the optimistic
  row); pushIsolateAndClear and reset/close share cancelReservations so
  /clear-style commands cannot have a rejected steer resurrect a discarded
  prompt
- turn/steer acceptance wait bounded at 25s (< hub 30s RPC timeout): a
  lost response is indeterminate and funnels into thread reconciliation
  instead of stranding the reservation
- tests updated for the tri-state cancel contract

* fix(codex,web): busy-aware edit flow; bound reconciliation reads

- QueuedMessagesBar edit flow treats a 'busy' cancel as unsuccessful: it
  never prefills the composer when the row is inside an async steer, so a
  second client cannot send a duplicate
- reconcileSteerByClientId bounds thread/read with a 5s timeout so a
  connected-but-silent app-server cannot hold the reservation in-flight
  indefinitely

* fix(steer): inFlight-dominated cancel acks; bounded reconciliation

- hub cancel-queued-message acks check inFlight before removed: a stale
  duplicate socket reporting removed can no longer delete the durable row
  while another socket is dispatching the steer
- reconciliation entries expire after 60s and mark delivered: after the
  rejection window, a dispatched steer that the app-server never proved
  (client ids dropped on restart) is committed instead of polling
  thread/read forever
- pre-dispatch failures (abort before write included) never enter
  reconciliation — they restore the row and report failure

* fix(steer): persist indeterminate outcomes without replay

* fix(steer): make ambiguous delivery restart-safe

* fix(steer): recover crash-held rows and preserve retry dedup

* fix(steer): ack retries and bound stdin dispatch

* fix(steer): reconcile indeterminate dispatches and serialize retries

* fix(codex): classify stdin callback failures as indeterminate

* fix(steer): recheck indeterminate cancels after ACK

* fix(steer): close retry and abort races

* fix(steer): serialize live retries and abort admission

* fix(steer): distinguish live dispatching from unknown

* fix(steer): keep ACK failures held and reconcile busy cancel

* fix(steer): distinguish held cancel from removal

* fix(store): combine schema v24 migrations

* fix(store): reserve schema v25 for steer delivery state

* fix(steer): keep held cancel state and notify requeue

* fix(steer): release explicitly cancelled unknown reservations

* fix(codex): reject cancelled reservations before native steer

* fix(codex): make reservation restore atomic with state

* fix(codex): terminate abandoned transport writes

* fix(steer): own abandoned app-server lifecycle and consume races

* fix(codex): confirm dispatch and recover abandoned turns

* test(codex): mock abandoned transport callback

* fix(codex): clear visible turn state on transport loss

* fix(steer): claim retries and cover native delivery state

* fix(native): preserve indeterminate state on Android hydration

* fix(steer): make retry claims single-winner

* fix(steer): serialize concurrent retry claims

* fix(socket): tolerate missing steer-state ACK callbacks

* fix(native): serialize retry operations

* docs(web): document unknown steer delivery and retry controls

* fix(steer): handle retry failures and abort-before-connect

* fix(steer): reinitialize after transport loss and finish iOS retry errors

* fix(steer): preserve indeterminate rows across reconnect gaps

* test(web): mock indeterminate queued recovery state

* fix(steer): recover consumed ACK tombstones

* fix(steer): expose consumed cancel tombstones
2026-08-19 20:07:39 +08:00
weishu c76af90d88 feat(hub): push config joins the settings.json system (env > file > default)
FCM and iOS/APNs push were the only hub knobs read straight from env,
bypassing the configuration rule every other field follows (env >
settings.json > default, env persisted on first sight). Fold them in:
serverSettings resolves fcmServiceAccountPath / iosPushMode /
iosPushRelayUrl / apnsKeyP8Path / apnsKeyId / apnsTeamId / apnsBundleId
/ apnsEnv under the shared rule, and the two resolvers now consume
configuration instead of process.env. FCM_PROJECT_ID is gone: operators
point at the service-account JSON and the project id comes from the
file — no copying values out of it. Paths accept ~.
2026-08-18 23:10:57 +08:00
weishu 1ad78f409a feat(hub): iOS push channel — direct APNs + relay with E2E envelope (P1) 2026-08-18 20:56:49 +08:00
weishu fada270424 docs(api): sse.md — activeTurnStartedAt is not patch-applied (matches reference impl + fixtures) 2026-08-17 12:12:32 +08:00
weishu 1a40db3b2d merge: K2 client contract docs (sse/pagination/messages) 2026-08-17 11:25:46 +08:00
weishu 1b9ca34fb5 docs(api): add native client contract (sse, pagination, messages) 2026-08-17 11:18:14 +08:00
weishu 16ebba12e8 docs(api): add native client contract (auth, rest, errors) 2026-08-17 11:15:00 +08:00
weishu b7503d9309 docs: restructure docs site and fix drift against code
- split installation.md into installation/deployment/notifications
- merge cursor/grok guides into new agents.md with full support matrix
- sidebar: grouped sections; add namespace, deployment, notifications,
  native companion contract
- fix license footer (AGPL-3.0), settings schema fields and $id
- fix drift in pwa, faq, namespace, how-it-works, voice-assistant,
  quick-start, native-companion-contract
- move mermaid lightbox dogfood doc to localdocs (untracked)
- README: complete agent list, replace dead cursor/grok links
2026-08-05 07:51:16 +08:00
226b2d066a feat(hub): native companion (FCM) push channel + device registry + pairing QR (#803)
* feat(hub): native companion (FCM) push channel + device registry

Adds opt-in FCM HTTP v1 notification delivery so a companion mobile/wearable
app can receive permission, ready, and task notifications end-to-end. The
channel is gated entirely on FCM_SERVICE_ACCOUNT_PATH + FCM_PROJECT_ID being
set; operators not running a companion see zero behavior change.

What lands:

- POST/DELETE /api/devices/register — JWT-authed FCM token registry,
  upsert on (namespace, deviceId, platform), platforms `phone` | `wear`.
- Sqlite v9 → v10 migration adds `fcm_devices` (idx on namespace + token).
- FcmService — minimal HTTP v1 client, RS256 service-account JWT via
  jose (dep already in tree), 5-minute access-token cache, 401 retry.
- FcmNotificationChannel — implements NotificationChannel, sends data-only
  FCM (so companion can route to phone+watch surfaces). Body composition
  parses an optional trailing `AGENT_NOTIFY_SUMMARY {json}` line for richer
  ready summaries; truncates plain assistant text to 280 chars otherwise.
  Tags each payload with `severity` (info/warning/success/error) so clients
  can color/categorise the notification.
- PushNotificationChannel gains a NativeFallbackProbe — when a namespace
  has at least one registered FCM device, web-push and SSE in-page toast
  are skipped so the operator does not double-notify on phone+browser.
  Probe is no-op when no FCM device is registered; PWA-only setups
  unchanged. Branch trace gated on HAPI_NOTIFY_DEBUG=1.
- shared/src/messages.ts — `extractAssistantPlainText` (codex + Claude SDK
  shapes) and `extractNotifySummary` (strict end-anchored line parser).
- hub/src/notifications/toolArgs.ts — tool-arg formatters lifted out of
  telegram/sessionView (kept duplicated there in this PR; refactor of
  Telegram is a follow-up).
- docs/api/native-companion-contract.md — payload + endpoints + env vars,
  versioned at contract v1.

Test coverage:

- 260 hub tests pass (incl. 23 new across FCM channel, push dedup,
  v10 migration, devices route).
- 60 shared tests pass (messages parsers).

Notes for reviewers:

- Reference companion implementation lives in a separate Android repo
  (Kotlin, phone APK + Wear OS APK) — this PR is hub-side only.
- No new runtime deps (`jose` and `zod` already declared in hub).

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): clarify scope - companion is remote-hub client, not hub-on-phone

Adds a Scope section to the native-companion contract so anyone
implementing it knows the audience: operators running the hub on a
server who want phone/watch as a notification surface, not users
expecting a Termux-bundled hub. Mirrors the framing now in
heavygee/hapi-companion README.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): correct Scope section - hub topology is unchanged

Removes the prior framing that referenced a non-existent 'Termux
hub-on-phone' alternative. This contract describes a native client to
the same hub the PWA talks to; it does not change where the hub runs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(web): companion app pairing QR in Settings

Companion section in Settings renders a QR code encoding the deeplink
hapicompanion://bind?hub=<base>&code=<token>. Scanning it from the HAPI
companion app (Android phone or Wear OS) auto-fills the bind form and
authenticates against this hub - no manual URL/token paste.

QR is gated behind a Show button so the access token doesn't sit visible
on screen by default; a Copy link affordance and the textual deeplink
are also exposed for manual onboarding.

Adds qrcode + @types/qrcode to web/ (already a hub dep, no new resolved
package - just a workspace declaration).

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(hub): terminal QR for companion app pairing alongside PWA QR

After the existing PWA access QR is rendered on tunnel start, also print
the hapicompanion://bind?hub=...&code=... deeplink and a matching QR.

Same tunnel + token, different scheme: phones with the companion app
installed pick up the deeplink via the manifest intent filter; phones
without it ignore it and fall back to the PWA QR above.

QR rendering failure is non-fatal in both cases - the textual deeplink
above the QR is sufficient for manual paste.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): address HAPI Bot review on PR #803

Two bugs surfaced by the upstream review bot:

1) Web Push silently dropped when FCM is not actually configured.
   The native-fallback probe only checked the device registry; it did
   not check whether resolveFcmConfig() actually succeeded. So an
   operator who previously enabled FCM, registered a phone, then later
   started the hub WITHOUT FCM_SERVICE_ACCOUNT_PATH would see the probe
   return true (devices still in DB) -> Web Push suppressed -> no FCM
   channel registered -> notifications go to /dev/null.

   Fix: extracted the probe construction into buildNativeFallbackProbe()
   which short-circuits to () => false when fcmConfig is missing. Probe
   never even consults the device store in the no-config branch, so
   stale rows can never matter.

2) Transient FCM failures permanently unregistered devices.
   sendToToken() returned a single boolean and sendToNamespace() removed
   any device whose send returned false. A 429 (rate limit), 503
   (server error), 401 (auth glitch), or even an ECONNREFUSED would
   delete the device row, after which the user would need to re-pair to
   get notifications again. The bot caught it; the fix is the obvious
   one.

   Fix: sendToToken() now returns 'sent' | 'invalid' | 'failed'.
   - 'invalid' is reserved for the responses that genuinely indicate a
     dead token: HTTP 404 with UNREGISTERED/NOT_FOUND, and HTTP 400
     with INVALID_ARGUMENT explicitly referencing the token field.
   - Everything else (429, 5xx, 401, 403, network errors) is 'failed'
     and counts toward the failed tally without removing the device.

   sendToNamespace() only calls removeDeviceByToken() on 'invalid'.

Tests: 11 new tests across two new files. fcmService.test.ts covers
all six branches (200, 404 unregistered, 429, 503, 401, network error)
plus a mixed-batch case that proves invalid tokens get removed in the
same call where transient-failure tokens survive. nativeFallbackProbe
.test.ts covers both no-config and configured branches plus the
explicit "no-config never touches the store" guarantee.

Hub test count: 273 -> 284 (all passing).

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): correct FCM visibility rule and remove unsupported event type

HAPI Bot review on PR #803 caught two contract-doc accuracy gaps:

1) Visibility rule was wrong. Doc said "FCM fires when Web Push would
   fire AND client not visible via SSE", but FcmNotificationChannel
   ALWAYS fires regardless of PWA visibility (deliberately - native
   companion is the canonical wrist-first surface, and there is a
   passing test asserting this). Companion app implementers reading
   the contract would have built foreground-suppression logic and
   then dropped notifications when the PWA tab was open.

2) Documented `session-completed` event doesn't exist. NotificationHub
   never calls into a 'session-completed' channel method on
   FcmNotificationChannel; the type would never reach a native client.
   Removed from the documented enum, leaving only the three actual
   events: ready, permission-request, task-notification.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): drop trailing whitespace, use blank line for paragraph break

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): persist CLI access token after Telegram bind so pairing QR works

The Settings -> Companion pairing QR reads the original CLI access token
from localStorage (hapi_access_token::<baseUrl>) so it can be encoded into
the hapicompanion://bind deeplink. For browser/CLI logins useAuthSource
already persists the token via setAccessToken, but the Telegram Mini App
bind path went through useAuth.bind() which exchanged the typed CLI token
for a JWT and never persisted it. Telegram users therefore always saw the
"signed in via Telegram..." fallback and got no usable QR.

After a successful client.bind() we now mirror useAuthSource's behavior
and write the same accessToken to the same localStorage key, restoring
parity between the two auth paths. No change for browser/CLI users.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): gate native-fallback probe on rolling FCM health

The native-fallback probe previously returned true whenever FCM was
configured AND devices were registered, which suppressed web-push for
the namespace. The HAPI Bot correctly pointed out the gap: if the FCM
pipeline silently breaks (expired service-account key, sustained 5xx,
OAuth token-fetch failure, network blackhole) the operator gets nothing
on either channel until they manually intervene.

Approach (deliberate, not the bot's exact suggested fix):

- FcmService now keeps a small rolling window (last 8 outcomes) of send
  attempts and exposes `isHealthy()`. The threshold is 5+/8 failures =
  unhealthy; the buffer starts empty so a freshly-booted hub is
  optimistic ("innocent until proven guilty") and does not double-fire
  on event #1.
- Token-fetch failure (`getFcmAccessToken` throws) now records exactly
  one health-failure (not one per device), short-circuits the send
  loop, and returns a result so `sendToNamespace` no longer leaks the
  exception.
- `invalid` token responses are explicitly excluded from the health
  buffer because they are per-device facts (rotated/uninstalled token),
  not pipeline failures - FCM was reachable, it just rejected one
  stale token.
- `buildNativeFallbackProbe` now optionally accepts the FcmService and
  short-circuits to "let web-push fire" when health is bad, before it
  even queries the device registry. The single-arg call shape is still
  supported for back-compat.

Why not the bot's exact suggestion ("invert: call FCM first, fall back
on result.sent === 0"):
- Couples PushNotificationChannel to FcmService and FcmSendPayload,
  reversing the clean parallel-channel architecture established earlier
  in this PR.
- Treats every transient single-event failure as fallback-worthy, which
  re-opens the duplicate-notification race that the suppression logic
  was added to close (FCM HTTP timeout that delivers later + the web
  push we sent in the meantime = two pings).
- A rolling health window only flips on sustained breakage, which is
  the actual operational scenario the bot is worried about.

The wrist-first design intent ("FCM fires unconditionally, web-push is
suppressed for the same namespace") documented in
docs/api/native-companion-contract.md is preserved on the happy path.
The probe only re-enables web-push when there is concrete evidence the
native pipeline is not delivering.

Tests:
- New FcmService.isHealthy suite covers empty-buffer, threshold flip,
  recovery as failures age out of the window, invalid-token exclusion,
  and network-error path.
- nativeFallbackProbe gains coverage for the unhealthy-but-registered,
  healthy-and-registered, and absent-fcmService (back-compat) cases.
- All 292 hub tests still pass; typecheck clean.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(telegram): drop duplicate tool-args formatter, use shared module

The Telegram session view had its own copy of formatToolArgumentsDetailed
identical to the one in hub/src/notifications/toolArgs.ts (already used by
the FCM channel). Replace the local copy with an import.

Removes ~70 lines of duplication, plus the now-unused MAX_TOOL_ARGS_LENGTH
constant and `truncate` import. The shared signature accepts an optional
opts arg whose default maxArgLength is 150 - matching the prior constant -
so the call site is unchanged.

Two benign upgrades come along for the ride from the shared module:
?? instead of || on field fallbacks (no real-world difference; permission
arguments never carry empty-string fields), and String(...) wrapping plus
a typeof object guard that makes non-string values render gracefully
instead of throwing into the catch block.

Hub tests: 311 pass / 0 fail. Telegram subset: 5 pass / 0 fail. typecheck
green.

Cold-reviewed by an out-of-context Claude Opus peer before push.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(fcm): require positive evidence in health window before suppressing web-push

Addresses HAPI Bot Major review on PR #803.

The previous health gate treated an empty outcome buffer as healthy
("innocent until proven guilty"). That created a silent-blackhole window
on cold start with broken FCM credentials: the push channel suppressed
SSE/Web Push for the first ~5 events while the FCM channel attempted
each delivery and recorded failures, until enough stacked to flip the
threshold. Every notification in that gap was silently lost.

New invariant: isHealthy() requires at least one successful FCM send in
the recent window (HEALTH_WINDOW=8) AND failures below threshold
(HEALTH_FAILURE_THRESHOLD=5). Both conditions are necessary; either
alone is insufficient evidence to safely suppress web-push fallback.

Trade-off: one duplicated notification per hub restart per namespace.
On the first event after restart, web-push fires alongside FCM (because
the gate has no positive evidence yet). Once FCM records that first
success, the gate engages and subsequent events are FCM-only. Worth it
for guaranteed delivery during cold-start outages.

Tests reworked to match new semantics:
- "starts UNHEALTHY with empty buffer" (was: healthy)
- "flips to healthy after first successful send" (new)
- "stays unhealthy across failures-only run" (new, exercises the exact
  blackhole scenario the bot flagged)
- "flips back to unhealthy after threshold breach with prior successes"
  (renamed, establishes successes first)
- "invalid tokens don't count against health" (reworked: send a mixed
  batch first to establish health, then verify invalids don't flip it)
- "network errors count as failures" (reworked: establish health first)

Hub tests: 313 pass / 0 fail. typecheck green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): bump FCM migration to V10→V11 after upstream service_tier V9→V10

Upstream/main landed sessions.service_tier at schema v10. The companion
FCM device registry now migrates at v11 so both changes compose cleanly
after the courtesy rebase onto current upstream/main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): per-dispatch native gate instead of stale FCM probe

FCM runs before web-push; PushNotificationChannel skips web/SSE only
when the same notify() dispatch already delivered via FCM. Removes the
isHealthy()+device-row probe that could suppress web-push after warm
FCM outages.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,web): cap notifySummary for FCM limits; fix PWA test cast

Rebase follow-up: truncate AGENT_NOTIFY_SUMMARY summary/action before
FCM data payload (bot Major). Fix usePwaUpdate.test.ts setTimeout mock
cast so bun typecheck passes on current main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): cap all FCM notifySummary fields and task bodies

Whitelist and truncate AGENT_NOTIFY_SUMMARY auxiliary fields before
JSON serialization; cap task-notification summaries to glance limit.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): FCM fetch timeouts and cap Grep/Glob permission args

10s AbortSignal.timeout on OAuth + FCM send so sequential web-push
fallback is not blocked on hung Google endpoints; truncate Grep/Glob
pattern in permission detail formatter.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): bind FCM token to one namespace on re-pair

Delete stale fcm_devices rows sharing the same token when a native
install registers under a different namespace.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): localize Companion settings and pairing copy

Add en/zh-CN keys for the Companion section title and CompanionPairing
strings; matches locale-driven Settings pattern (bot Minor on #803).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): tighten FCM token-invalid detection and truncation edge cases

Parse FCM error JSON: only UNREGISTERED or token-field INVALID_ARGUMENT
unregister devices; generic NOT_FOUND stays transient. Guard limit<=3
in truncateReadyText so tiny action budgets cannot blow the glance cap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): parse FcmError details.errorCode for UNREGISTERED tokens

FCM v1 often returns HTTP 404 with root NOT_FOUND plus
details[].errorCode UNREGISTERED; prune those tokens while keeping
generic project/resource NOT_FOUND transient.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): mock AppContext for About Companion pairing in settings tests

Settings About now mounts CompanionPairing via useAppContext after the
#1027 hub redesign rebase; wrap the About route test with AppContext and
Companion mocks so the suite stays green.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(contract): point companion auth at POST /api/auth, not /api/bind

Pairing QR carries the CLI access token as `code`. /api/bind requires
Telegram initData; native companions must use /api/auth with accessToken.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): mount Companion pairing under Settings General

About is version/links only after the settings hub redesign; pairing is
setup, so keep Companion with language prefs and update the route tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Debian <heavygee@oos-linux.in.lockhouse>
2026-07-27 19:52:54 +08:00