From 1de9613df60912efa9b5f623f322ecbddc298157 Mon Sep 17 00:00:00 2001 From: weishu Date: Sat, 12 Sep 2026 19:48:20 +0800 Subject: [PATCH] docs: align documentation with current implementation --- AGENTS.md | 40 +++++---- CONTRIBUTING.md | 10 ++- README.md | 4 +- android/README.md | 16 ++-- cli/README.md | 61 +++++++------ cli/src/runner/README.md | 114 ++++++++++++++----------- docs/.vitepress/config.ts | 3 +- docs/api/client-contract/auth.md | 8 +- docs/api/client-contract/errors.md | 10 ++- docs/api/client-contract/index.md | 13 ++- docs/api/client-contract/messages.md | 24 +++++- docs/api/client-contract/pagination.md | 18 ++-- docs/api/client-contract/rest.md | 14 +-- docs/api/client-contract/sse.md | 4 +- docs/api/native-companion-contract.md | 6 +- docs/guide/agents.md | 12 +-- docs/guide/deployment.md | 23 +++-- docs/guide/faq.md | 18 ++-- docs/guide/how-it-works.md | 29 +++---- docs/guide/installation.md | 7 ++ docs/guide/notifications.md | 19 ++++- docs/guide/quick-start.md | 4 +- docs/guide/why-hapi.md | 28 +++--- docs/privacy.md | 20 ++--- hub/README.md | 45 +++++++--- ios/README.md | 17 ++-- web/README.md | 20 +++-- 27 files changed, 360 insertions(+), 227 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 5fc80cf3..bfad0a6c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -6,13 +6,14 @@ Short guide for AI agents in this repo. Prefer progressive loading: start with t ## What is HAPI? -Local-first platform for running AI coding agents (Claude Code, Codex, Gemini) with remote control via web/phone. CLI wraps agents and connects to hub; hub serves web app and handles real-time sync. +Local-first platform for running AI coding agents with remote control via web/phone. CLI wraps agents and connects to hub; hub serves web app and handles real-time sync. See `docs/guide/agents.md` for launchable agents; Gemini remains a historical wire flavor, not a launchable integration. ## Repo layout ``` cli/ - CLI binary, agent wrappers, runner daemon hub/ - HTTP API + Socket.IO + SSE + Telegram bot +relay/ - Standalone encrypted native push relay (APNs + FCM) web/ - React PWA for remote control ios/ - Native SwiftUI app (in development) android/ - Native Kotlin Compose app (in development) @@ -22,7 +23,7 @@ docs/ - VitePress documentation site website/ - Marketing site ``` -Bun workspaces; `shared` consumed by cli, hub, web. `ios`/`android` outside workspaces (Xcode / Gradle toolchains). +Bun workspaces: cli, shared, hub, web, website, docs, relay. `shared` consumed by cli, hub, web as `@hapi/protocol`. `ios`/`android` outside workspaces (Xcode / Gradle toolchains). ## Architecture overview @@ -39,11 +40,13 @@ Bun workspaces; `shared` consumed by cli, hub, web. `ios`/`android` outside work ``` **Data flow:** -1. CLI spawns agent (claude/codex/gemini), connects to hub via Socket.IO +1. CLI starts the selected agent integration, connects to hub via Socket.IO 2. Agent events → CLI → hub (socket `message` event) → DB + SSE broadcast 3. Web subscribes to SSE `/api/events`, receives live updates 4. User actions → Web → hub REST API → RPC to CLI → agent +Web terminals use a separate JWT-authenticated Socket.IO `/terminal` namespace; `/cli` uses the CLI access token. + ## Reference docs - `README.md` - User overview, quick start @@ -66,8 +69,8 @@ Bun workspaces; `shared` consumed by cli, hub, web. `ios`/`android` outside work ## Common commands (repo root) ```bash -bun typecheck # All packages -bun run test # cli + hub + web + shared tests +bun typecheck # cli + hub + web + relay (shared checked through consumers) +bun run test # cli + hub + web + shared + relay tests bun run dev # hub + web concurrently bun run build:single-exe # All-in-one binary bun run gen:fixtures # Regenerate shared/fixtures/ from web pipeline @@ -82,7 +85,7 @@ iOS tests run in CI (`ios.yml`: macOS `swift test`); no local Xcode/Swift toolch - `api/` - Hub connection (Socket.IO client, auth) - `claude/` - Claude Code integration (wrapper, hooks) - `codex/` - Codex mode integration -- `agent/` - Multi-agent support (Gemini via ACP) +- `agent/` - Shared session/bootstrap support and ACP transport - `runner/` - Background daemon for remote spawn - `commands/` - CLI subcommands (auth, runner, doctor) - `modules/` - Tool implementations (ripgrep, difftastic, git) @@ -93,10 +96,12 @@ iOS tests run in CI (`ios.yml`: macOS `swift test`); no local Xcode/Swift toolch - `socket/` - Socket.IO setup - `socket/handlers/cli/` - CLI event handlers (session, terminal, machine, RPC) - `sync/` - Core logic (sessionCache, messageService, rpcGateway) -- `store/` - SQLite persistence (better-sqlite3) +- `store/` - SQLite persistence (bun:sqlite) - `sse/` - Server-Sent Events manager - `telegram/` - Bot commands, callbacks -- `notifications/` - Push (VAPID) and Telegram notifications +- `notifications/` - Notification dispatch and payload composition +- `push/`, `fcm/`, `push-ios/` - Web Push, Android, and iOS delivery +- `push-native/` - Shared encrypted envelope and push relay client - `config/` - Settings loading, token generation - `visibility/` - Client visibility tracking @@ -117,8 +122,8 @@ iOS tests run in CI (`ios.yml`: macOS `swift test`); no local Xcode/Swift toolch - `modes.ts` - Permission/model mode definitions ### iOS (`ios/`) -- `Packages/HapiKit/` - local SPM package: `HapiProtocol` (wire models + chat pipeline, fixtures-verified), `HapiClient` (API/auth/SSE/stores) -- `Hapi/` + `Hapi.xcodeproj` - thin SwiftUI app target +- `Packages/HapiKit/` - local SPM package: `HapiProtocol` (wire models + chat pipeline, fixtures-verified), `HapiClient` (API/auth/SSE/stores), `HapiUI` (rendering) +- `Hapi/` + `Hapi.xcodeproj` - SwiftUI app, native transcript, feature screens, push extension ### Android (`android/`) - `:core:protocol` - pure JVM wire types + chat pipeline (fixtures-verified) @@ -134,15 +139,15 @@ iOS tests run in CI (`ios.yml`: macOS `swift test`); no local Xcode/Swift toolch ## Pre-push self-review (agents) -Before commit/push/PR: use the **`pre-push-review`** skill (`~/.cursor/skills/pre-push-review/`). +Before commit/push/PR, run the repository checks directly: -1. **Mechanical:** `bun typecheck && bun run test` (matches `.github/workflows/test.yml`) +1. **Mechanical:** `bun typecheck && bun run test` (typecheck/unit-test portion of `.github/workflows/test.yml`; CI also runs selected Playwright and CLI integration tests) 2. **Logic:** skim `git diff origin/main...HEAD`; apply `.github/prompts/codex-pr-review.md` as a local Major checklist (no Codex required) 3. **Style:** optional ## Testing -- Test framework: Vitest (via `bun run test`) +- Test frameworks: Vitest for CLI/Web; Bun test for Hub/Shared/Relay (via `bun run test`) - Test files: `*.test.ts` next to source - Run: `bun run test` (from root) or `bun run test` (from package) - Hub tests: `hub/src/**/*.test.ts` @@ -153,8 +158,8 @@ Before commit/push/PR: use the **`pre-push-review`** skill (`~/.cursor/skills/pr | Task | Key files | |------|-----------| -| Add CLI command | `cli/src/commands/`, `cli/src/index.ts` | -| Add API endpoint | `hub/src/web/routes/`, register in `hub/src/web/index.ts` | +| Add CLI command | `cli/src/commands/`, register in `cli/src/commands/registry.ts`; public help in `help.ts` | +| Add API endpoint | `hub/src/web/routes/`, register in `hub/src/web/server.ts` | | Add Socket.IO event | `hub/src/socket/handlers/cli/`, `shared/src/socket.ts` | | Add web route | `web/src/routes/`, `web/src/router.tsx` | | Add web component | `web/src/components/` | @@ -167,8 +172,9 @@ Before commit/push/PR: use the **`pre-push-review`** skill (`~/.cursor/skills/pr - **RPC**: CLI registers handlers (`rpc-register`), hub routes requests via `rpcGateway.ts` - **Versioned updates**: CLI sends `update-metadata`/`update-state` with version; hub rejects stale -- **Session modes**: `local` (terminal) vs `remote` (web-controlled); switchable mid-session -- **Permission modes**: `default`, `acceptEdits`, `auto`, `bypassPermissions`, `plan` +- **Session modes**: `local` (terminal) vs `remote` (web-controlled) for handoff-capable integrations; Codex uses concurrent clients without ownership switching. See `docs/guide/codex-shared-sessions.md`. +- **Session identity**: Ordinary wrappers export `HAPI_SESSION_ID` after bootstrap. Shared Codex uses a per-root MCP bridge and `shell_environment_policy.set.HAPI_SESSION_ID`; never put one root's ID into the shared app-server environment (`cli/src/codex/shared/root.ts`, `runtime.ts`). +- **Permission modes**: Per-flavor catalogs in `shared/src/modes.ts`; session capabilities further constrain available controls - **Namespaces**: Multi-user isolation via `CLI_API_TOKEN:` suffix ## Adding new web features — consider an FUE diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index f71e31c6..1668f0c6 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -54,7 +54,7 @@ Have an idea to improve HAPI? Open an issue with: ## Getting Started 1. Fork and clone the repository -2. Install dependencies: +2. Install Bun 1.4.0, then install dependencies from the repo root: ```bash bun install ``` @@ -65,6 +65,14 @@ Have an idea to improve HAPI? Open an issue with: See the [README](README.md) for more build options. +From the repo root, `bun typecheck` checks CLI, Hub, Web, and Relay; +`bun run test` runs CLI/Web Vitest and Hub/Shared/Relay Bun tests. Native +toolchains and checks are documented in [iOS](ios/README.md) and +[Android](android/README.md). + +For documentation changes, run `bun run --cwd docs docs:build` and check +affected links and anchors in the generated site. + ## Questions? If you have questions, feel free to open an issue. We're here to help! diff --git a/README.md b/README.md index 92723506..42e61567 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ Run official Claude Code / Codex / Cursor Agent / Grok Build / OpenCode / Kimi / ## Features - **Seamless Handoff** - Work locally, switch to remote when needed, switch back anytime. No context loss, no session restart. -- **Shared Codex Sessions** - Terminal and Web/phone use the same Codex engine simultaneously; normal terminal/Runner lifecycle, with resume after exit. Requires Codex 0.154.0+. [Lifecycle and limits](docs/guide/codex-shared-sessions.md). +- **Shared Codex Sessions** - Use Codex from your terminal and phone at the same time. Requires Codex 0.154.0+. [Usage and limits](docs/guide/codex-shared-sessions.md). - **Native First** - HAPI wraps your AI agent instead of replacing it. Same terminal, same experience, same muscle memory. - **AFK Without Stopping** - Step away from your desk? Approve AI requests from your phone with one tap. - **Your AI, Your Choice** - Claude Code, Codex, Cursor Agent, Grok Build, OpenCode, Kimi, Copilot, Antigravity, Pi, DeepSeek Harness—different agents, one unified workflow. @@ -49,7 +49,7 @@ For self-hosted options (Cloudflare Tunnel, Tailscale), see [Installation](docs/ ## Native apps (iOS / Android) -Fully native SwiftUI and Kotlin Compose clients are in development under `ios/` and `android/`. They pair with your hub by scanning the same terminal QR code as the web app, and follow the same protocol — see the [client contract docs](docs/api/client-contract/index.md). +Fully native SwiftUI and Kotlin Compose clients are in development under `ios/` and `android/`, with chat, approvals, session creation, files, dictation, and push notifications. Pair using the companion QR code printed by the hub or shown in web Settings, or enter the hub URL and token manually. See the [client contract docs](docs/api/client-contract/index.md). ## Build from source diff --git a/android/README.md b/android/README.md index e6cf3a74..6eb03197 100644 --- a/android/README.md +++ b/android/README.md @@ -11,9 +11,9 @@ independent from the web app; shares only the protocol contract | Module | Type | Responsibility | |---|---|---| -| `:core:protocol` | **pure Kotlin/JVM** (no Android) | Hub wire types (kotlinx.serialization), chat pipeline port (normalize → reduce → tool groups), message-window/pagination logic, versioned patch application, modes catalog, git output parsers, `BindLink` pairing-link parsing. **M1a landed**: `wire/` (`HapiJson`, `Session`/`SessionPatch`/`SessionSummary`, `DecryptedMessage`, `AgentState`, `Machine`, 13-type `SyncEvent` union via `SyncEvents.parse`, `MessagesResponse`), `catalog/` (flavors + permission/collaboration modes), `patch/SessionPatching.kt` (exact port of `web/src/lib/sessionPatch.ts`), all fixture-verified. | -| `:core:data` | Android library | Transport + persistence. **M1b landed** — `auth/` (`JwtPeek`, `CredentialStore` interface + `EncryptedPrefsCredentialStore`/in-memory, `HubUrls` origin normalization, `HubRegistry` roster behind a storage seam, `AuthInterceptor` + single-flight `TokenAuthenticator` with `ensureFreshToken()` and terminal `AuthEvents`), `api/` (`HapiApi` — plain OkHttp + kotlinx.serialization, one suspend fun per v1 endpoint incl. generated-image bytes via a 256 MB OkHttp cache and the multipart transcription helper; `ApiError` with `(status, code)`), `HubSession` per-hub factory; MockWebServer-tested. **M1c landed** — `sse/`: `SseEngine` (per-key `global`/`session:` loops, `connection-changed` handshake gate with `ok`/`gap` resume verdict, per-key `Last-Event-ID` cursors advanced only after downstream hand-off (at-least-once), 10 s connect deadline, 90 s watchdog, 1 s→30 s→300 s backoff + jitter, background retry deferral + 45 s foreground stale check, one silent 401 re-auth per cycle), `OkHttpSseTransport` (dedicated client, `readTimeout=0`, incremental gzip decoding pinned by test, `acceptEncodingIdentity` fallback), `SyncEventRouter` → `SyncTargets` seam; virtual-time tested. Still to come: StateFlow stores + AtomicFile JSON snapshots (M2), FCM registration + WorkManager workers (M4). | -| `:app` | Android application | Compose UI, navigation, deep links (`hapicompanion://bind`), FCM service (M4), hand-rolled DI (`AppGraph`, no Hilt). **M1d landed** — `di/` (`AppGraph` process singletons: Preferences DataStore-backed `HubRegistryStorage`, `EncryptedPrefsCredentialStore`, `HubRegistry`, auth-terminal fan-out; `HubGraph` per active hub: `HubSession` + `SseEngine` wired to `ensureFreshToken`, recreated on hub switch; `LocalAppGraph` CompositionLocal + `viewModelFactory` helper), `feature/pairing/` (landing / zxing `ScanContract` QR scan / manual entry sharing one `PairingViewModel`: health + protocol check → `POST /api/auth` → persist + activate), `feature/home/` placeholder (hub switcher + sign-out), `Navigation.kt` (pairing ⇄ home, auth-terminal → pairing with banner), bind deep-link handling in `MainActivity`. | +| `:core:protocol` | **pure Kotlin/JVM** (no Android) | Hub wire types, chat pipeline, message pagination, versioned patches, agent/mode catalogs, git parsers, pairing links, and golden-fixture conformance tests. | +| `:core:data` | Android library | OkHttp API/SSE transport, per-hub authentication, secure credentials, StateFlow stores and disk snapshots, encrypted push registration/decoding, and background notification actions. | +| `:app` | Android application | Compose screens for pairing, sessions/chat, approvals, new sessions, files, Scratchlist, dictation, usage/storage, and settings; navigation, localization, FCM service, and WorkManager wiring. | Dependency direction: `:app` → `:core:data` → `:core:protocol`. @@ -30,9 +30,9 @@ tasks.test { } ``` -Fixture-driven tests (M2) read `System.getProperty("hapi.fixtures.dir")` — -no further build changes are needed when `shared/fixtures/**` lands. CI -re-runs this suite whenever `android/**` or `shared/fixtures/**` change. +Fixture-driven tests read `System.getProperty("hapi.fixtures.dir")` from the +checked-in golden fixture set. CI re-runs this suite whenever `android/**` +or `shared/fixtures/**` change. ## Building @@ -197,9 +197,9 @@ and restored hub state. The manifest also sets Sign-out (home → Sign out) deletes the stored credentials for that hub and drops it from the roster. -## Milestones (track B of the native-clients plan) +## Milestone history (track B of the native-clients plan) -- **M0** — this scaffold: modules, version catalog, CI, placeholder screen. +- **M0** — initial scaffold: modules, version catalog, CI, placeholder screen. - **M1** — foundations: wire types + modes catalog; auth + `HapiApi` (MockWebServer-tested); `SseEngine` reconnect state machine + versioned patches (gzip streaming verified); pairing UI + `hapicompanion://bind` deep link. - **M2** — read-only chat: chat pipeline port gated on fixtures all-green; session list; `MessageWindowStore` port; Markdown renderer; read-only chat screen (`LazyColumn` with chronological stable keys). - **M3** — interaction: composer (optimistic send/queue/steer/drafts), permission approvals UX, session controls (mode/model/abort/resume/rename/archive), new session, dictation. diff --git a/cli/README.md b/cli/README.md index ca520cee..0002ac71 100644 --- a/cli/README.md +++ b/cli/README.md @@ -28,10 +28,10 @@ Choose a supported coding agent from your terminal and control its sessions remo - `hapi` - Choose an agent interactively. Unavailable agents are shown with a reason and cannot be selected. - `hapi claude` - Start a Claude Code session (passes through Claude CLI flags). - `hapi codex` - Start Codex mode. See `src/codex/runCodex.ts`. -- `hapi codex resume ` - Resume existing Codex session. +- `hapi codex resume ` - Resume a Codex conversation by its native thread ID. For a HAPI session ID, use `hapi resume `. - `hapi cursor` - Start Cursor Agent mode. See `src/cursor/runCursor.ts`. Supports `hapi cursor resume `, `hapi cursor --continue`, `--mode plan|ask`, `--yolo`, `--model`. - Local and remote modes supported; remote uses `agent -p` with stream-json. + Local and remote modes supported; new remote sessions use `agent acp`. Pre-ACP sessions retain the legacy `agent -p` stream-json resume path. - `hapi grok` - Start Grok Build mode. See `src/grok/runGrok.ts`. - `hapi copilot` - Start GitHub Copilot mode. - `hapi kimi` - Start Kimi mode. @@ -69,14 +69,12 @@ hapi resume `hapi resume` lists resumable sessions for the current machine. `hapi resume ` hands off an active remote session and opens the same HAPI session in the local terminal. -**Codex exception:** Codex 0.154.0+ uses shared sessions, not handoff. `hapi -resume ` attaches another official TUI to its live execution; Web and other -terminals remain usable. The original terminal owns a terminal-created execution: -its exit stops that execution, leaving resumable history. Web-created/resumed -executions use the existing Runner; additional terminals only detach on exit. -**End session** archives the selected root. Native profile-v2 launch selection and in-place rewind are currently -unavailable. See [Codex shared sessions](../docs/guide/codex-shared-sessions.md) -for queue semantics, environment isolation, recovery, supported flags and tests. +For Codex, the terminal and Web stay usable at the same time. Closing the +original terminal stops a terminal-started run, but its history remains +resumable. Sessions started from the Web run under the Runner; closing a +terminal attached later does not stop them. **End session** archives the +selected conversation. See [Codex usage and limits](../docs/guide/codex-shared-sessions.md) +for details. ### Answer local Claude prompts from HAPI @@ -107,8 +105,9 @@ See `src/commands/auth.ts`. ### Runner management -- `hapi runner start` - Start runner as detached process. -- `hapi runner stop` - Stop runner gracefully. +- `hapi runner start` - Replace any existing runner and start a detached process with the supplied flags/environment. +- `hapi runner stop` - Stop runner gracefully; agent sessions stay alive. +- `hapi runner start-sync` - Run in the foreground (for a process supervisor). - `hapi runner status` - Show runner diagnostics. - `hapi runner list` - List active sessions managed by runner. - `hapi runner stop-session ` - Terminate specific session. @@ -144,9 +143,11 @@ See `src/ui/doctor.ts`. Codex sessions keep the MCP servers configured in the user's Codex `config.toml`. HAPI adds its own `hapi` bridge without replacing other user -servers. Runner-spawned Codex sessions copy only `config.toml` into their -temporary `CODEX_HOME`, so MCP settings are preserved while authentication -state remains isolated. The `hapi` server name is reserved by HAPI. +servers. When a runner spawn supplies a Codex auth token, it copies only +`config.toml` into a temporary `CODEX_HOME` and writes the supplied `auth.json`, +preserving MCP settings without copying unrelated authentication state. +Without a supplied token, Codex uses the runner's normal Codex home/auth. +The `hapi` server name is reserved by HAPI. On Windows, known package-manager shims (`uvx`, `npx`, `npm`, `pnpm`, `yarn`, `bunx`, and `.cmd`/`.bat` commands) use a short-lived HAPI stdio compatibility @@ -186,10 +187,10 @@ controls for DSH. ### Required - `CLI_API_TOKEN` - Shared secret; must match the hub. Can be set via env or `~/.hapi/settings.json` (env wins). -- `HAPI_API_URL` - Hub base URL (default: http://localhost:3006). ### Optional +- `HAPI_API_URL` - Hub base URL (default: http://localhost:3006; also configurable as `apiUrl` in settings). - `HAPI_HOME` - Config/data directory (default: ~/.hapi). - `HAPI_EXPERIMENTAL` - Enable experimental features (true/1/yes). - `HAPI_EXTRA_HEADERS_JSON` - JSON object of extra headers to send on CLI → hub requests, e.g. `{"Cookie":"CF_Authorization=..."}`. Can also be set as the `extraHeaders` object in `~/.hapi/settings.json` (environment variable wins). @@ -203,6 +204,9 @@ controls for DSH. - `HAPI_RUNNER_HEARTBEAT_INTERVAL` - Heartbeat interval in ms (default: 60000). - `HAPI_RUNNER_HTTP_TIMEOUT` - HTTP timeout for runner control in ms (default: 10000). +- `HAPI_RUNNER_WEBHOOK_TIMEOUT_MS` - Session-start webhook timeout in ms (default: 15000); raise for slow agent startup/resume. +- `HAPI_DISABLE_VERSION_HANDOFF` - Set to `1` to disable automatic runner replacement on CLI binary changes. +- `HAPI_RUNNER_SUPERVISED` - Set to `1` only when a supervisor restarts the runner after exit; enables the web Restart control's supervised path. ### Worktree (set by runner) @@ -214,17 +218,20 @@ controls for DSH. ### Set for the wrapped agent -- `HAPI_SESSION_ID` - The hub session id for the current run, exported into the wrapped agent/CLI child environment at spawn for every flavor (claude / codex / copilot / cursor / gemini / opencode / kimi / grok / pi), both runner-spawned and locally started sessions. Agents can read it to self-target "this chat" over the hub REST API or shell helpers without listing `/api/sessions`. Prefer the MCP `display_image` tool for inline media when it is available; use `HAPI_SESSION_ID` for hub REST / shell tooling where MCP is not. To **list** peers on the same hub/namespace, prefer MCP `list_peers` (works from runner-spawned sessions without sitting on the hub host; excludes the calling session). To **read** another session, prefer MCP `inspect_peer` or `hapi inspect-peer`. To **message** another session, prefer MCP `ping_peer` or `hapi ping-peer` — do not reinvent JWT+curl. User citations look like `[title](/sessions/)` or Copy-reference `See session "…" (/sessions/) for context`; pass that `` as `sessionIdPrefix`. Do not Grep/Glob `/sessions/` as a local filesystem path. On a remote runner, configure matching `HAPI_API_URL` + `CLI_API_TOKEN` (or `hapi auth login` / `~/.hapi/settings.json`) on the runner host so shell `hapi ping-peer --list` works; session CLI may export an explicit non-default hub URL into child env, but never mirrors `CLI_API_TOKEN` into wrapped agents. +- `HAPI_SESSION_ID` - The current HAPI session ID, available inside agent shells. Use it in scripts that target the current conversation without listing sessions. +- An explicitly configured `HAPI_API_URL` is also made available to agent shells. HAPI does not copy settings-backed `CLI_API_TOKEN` secrets into the agent environment; credentials already present in the parent environment may still be inherited. Web terminal PTYs strip hub secrets. - Lazy Codex (terminal) sessions export the id only after the hub row is materialized, which happens when the MCP bridge starts — before the agent process is spawned — so path-only self-targeting does not race a missing hub row. +For peer discovery and messaging, use the session's MCP `list_peers`, +`inspect_peer`, and `ping_peer` tools, or the corresponding CLI commands. +On a remote runner host, configure the matching hub URL and token so shell +commands reach the same hub (`hapi auth login` saves the token). - Example (shell fallback when MCP is unavailable) — path-only, self-targets the current session: +For example, this source-checkout helper displays an image in the current +session when MCP is unavailable: - ```bash - bun scripts/tooling/hapi-display-image.mjs /absolute/path/to/image.png "optional title" - ``` - - Explicit other session (prefix or full uuid) still works; that path may list sessions. +```bash +bun scripts/tooling/hapi-display-image.mjs /absolute/path/to/image.png "optional title" +``` ## Storage @@ -248,8 +255,8 @@ From the repo root: ```bash bun install -bun run build:cli -bun run build:cli:exe +bun run build:cli # Type-check the CLI; no executable output +bun run --cwd cli build:exe # Host-platform executable in cli/dist-exe// ``` For an all-in-one binary that also embeds the web app: @@ -260,7 +267,7 @@ bun run build:single-exe ## Source structure -- `src/api/` - Bot communication (Socket.IO + REST). +- `src/api/` - Hub communication (Socket.IO + REST). - `src/claude/` - Claude Code integration. - `src/codex/` - Codex mode integration. - `src/cursor/` - Cursor Agent integration. diff --git a/cli/src/runner/README.md b/cli/src/runner/README.md index 0d758434..5030c1ad 100644 --- a/cli/src/runner/README.md +++ b/cli/src/runner/README.md @@ -1,6 +1,10 @@ # HAPI CLI Runner: Control Flow and Lifecycle -The runner is a persistent background process that manages HAPI sessions, enables remote control from the mobile app, and handles auto-updates when the CLI version changes. +The runner is a persistent background process that starts and manages HAPI +sessions from web/phone. After you update the CLI, it can restart itself onto +the new binary; it does not download or install updates. + +Source paths below are relative to `cli/` unless prefixed with another package. ## 1. Runner Lifecycle @@ -9,8 +13,8 @@ The runner is a persistent background process that manages HAPI sessions, enable Command: `hapi runner start` Control Flow: -1. `src/index.ts` receives `runner start` command -2. Spawns detached process via `spawnHappyCLI(['runner', 'start-sync'], { detached: true })` +1. `src/commands/runner.ts` handles `runner start`, stopping any existing runner first so new flags/environment take effect +2. Spawns detached `runner start-sync`, forwarding configured workspace roots 3. New process calls `startRunner()` from `src/runner/run.ts` 4. `startRunner()` performs startup: - Sets up shutdown promise and handlers (SIGINT, SIGTERM, uncaughtException, unhandledRejection) @@ -19,8 +23,8 @@ Control Flow: - If same version running: exits with "Runner already running" - Lock acquisition: `acquireRunnerLock()` creates exclusive lock file to prevent multiple runners - Direct-connect setup: `authAndSetupMachineIfNeeded()` ensures `CLI_API_TOKEN` is set and `machineId` exists - - State persistence: writes PID, version, HTTP port, mtime to runner.state.json - HTTP server: starts Fastify on random port for local CLI control (list, stop, spawn) + - State persistence: writes PID, version, HTTP port, mtime to runner.state.json - WebSocket: establishes persistent connection to backend via `ApiMachineClient` - RPC registration: exposes `spawn-happy-session`, `stop-session`, `stop-runner` handlers - Heartbeat loop: every 60s (or `HAPI_RUNNER_HEARTBEAT_INTERVAL`) checks for version updates, prunes dead sessions, verifies PID ownership @@ -40,16 +44,17 @@ Control Flow: ### Version Detection & Auto-Update -The runner detects when CLI binary changes (e.g., after `npm upgrade hapi`): +The runner detects when the CLI binary changes (e.g., after `npm update -g @twsxtd/hapi`): 1. At startup, records `startedWithCliMtimeMs` (file modification time of CLI binary) 2. Heartbeat compares current CLI mtime with recorded mtime via `getInstalledCliMtimeMs()` -3. If mtime changed: - - Clears heartbeat interval - - Spawns new runner via `spawnHappyCLI(['runner', 'start'])` - - Waits 10 seconds to be killed by new runner -4. New runner starts, sees old runner running with different mtime -5. New runner calls `stopRunner()` which tries HTTP `/stop`, falls back to SIGKILL -6. New runner takes over +3. Replays the original runner arguments (including workspace roots), marking the replacement as an authorized handoff child +4. Releases the lock and waits up to 30 seconds for a different live runner PID in the state file +5. On confirmation, the old runner exits. On failure, it tries to reacquire the lock and stays online for a later retry; it exits if another process holds the lock + +`HAPI_DISABLE_VERSION_HANDOFF=1` disables this automatic replacement, not +the rest of the heartbeat. A foreground supervisor should run +`hapi runner start-sync`; advertise `HAPI_RUNNER_SUPERVISED=1` only when it +will restart the process after exit. ### Heartbeat System @@ -64,8 +69,11 @@ Every 60 seconds (configurable via `HAPI_RUNNER_HEARTBEAT_INTERVAL`): Command: `hapi runner stop` +This stops the runner, not its detached agent sessions. Use `stop-session` +to stop an individual session, or `doctor clean` for broader process cleanup. + Control Flow: -1. `stopRunner()` in `controlClient.ts` reads runner.state.json +1. `stopRunner()` in `controlClient.ts` reads runner.state.json and verifies the PID still belongs to a HAPI runner before contacting or signaling it 2. Attempts graceful shutdown via HTTP POST to `/stop` 3. Runner receives request, triggers shutdown with source `hapi-cli` 4. `cleanupAndShutdown()` executes: @@ -78,12 +86,16 @@ Control Flow: ## 2. Multi-Agent Support -The runner supports spawning sessions with different AI agents: +The runner supports the [current agent catalog](../../../docs/guide/agents.md). +If a spawn request omits `agent`, its fallback is still Claude; this is +separate from the interactive `hapi` picker, which waits for your choice +instead of launching Claude implicitly. +Examples of agent authentication: | Agent | Command | Token Environment | |-------|---------|-------------------| -| `claude` (default) | `hapi claude` | `CLAUDE_CODE_OAUTH_TOKEN` | -| `codex` | `hapi codex` | `CODEX_HOME` (temp directory with `auth.json` and a copy of user `config.toml`) | +| `claude` | `hapi claude` | Agent's existing login, or supplied `CLAUDE_CODE_OAUTH_TOKEN` | +| `codex` | `hapi codex` | Existing Codex home; a supplied token gets a temporary `CODEX_HOME` with `auth.json` and a copy of user `config.toml` | | `grok` | `hapi grok` | Grok CLI login or `XAI_API_KEY` | | `opencode` | `hapi opencode` | OpenCode config (no token injection) | @@ -103,11 +115,11 @@ Initiated by mobile app via backend RPC: 1. Backend forwards RPC `spawn-happy-session` to runner via WebSocket 2. `ApiMachineClient` invokes `spawnSession()` handler 3. `spawnSession()`: - - Validates/creates directory (with approval flow) + - Checks agent availability and workspace-root boundaries, then validates/creates the directory - Configures agent-specific token environment - Spawns detached HAPI process with `--hapi-starting-mode remote --started-by runner` - Adds to `pidToTrackedSession` map - - Sets up 15-second awaiter for session webhook + - Waits for the session-start webhook (15 seconds by default; `HAPI_RUNNER_WEBHOOK_TIMEOUT_MS` overrides) 4. New HAPI process: - Creates session with backend, receives `happySessionId` - Calls `notifyRunnerSessionStarted()` to POST to runner's `/session-started` @@ -116,9 +128,9 @@ Initiated by mobile app via backend RPC: ### Terminal-Spawned Sessions -User runs `hapi` directly: -1. CLI auto-starts runner if configured -2. HAPI process calls `notifyRunnerSessionStarted()` +User starts an agent from the terminal: +1. Session bootstrap registers with the hub; a runner is not required for terminal use +2. HAPI process calls `notifyRunnerSessionStarted()` if it can reach the local runner control server 3. Runner receives webhook, creates `TrackedSession` with `startedBy: 'hapi directly - likely by user from terminal'` 4. Session tracked for health monitoring @@ -137,9 +149,9 @@ When spawning a session, directory handling: ### Session Termination Via RPC `stop-session` or HTTP `/stop-session`: -1. `stopSession()` finds session by `happySessionId` or `PID-{pid}` format -2. Sends termination request via `killProcessByChildProcess()` or `killProcess()` (Windows uses `taskkill /T`) -3. `on('exit')` handler removes from tracking map +1. `stopSession()` locates the session, including persisted resume-process records +2. Stops its process tree and verifies exit; shared Codex uses a root-scoped stop so sibling conversations are not killed +3. Returns `stopped`, `already_gone`, or `still_alive`; uncertainty is not reported as successful termination ## 4. HTTP Control Server (Fastify) @@ -183,9 +195,11 @@ Terminates a specific session. ``` **Response (200):** ```json -{ "success": true } +{ "status": "stopped" } ``` +`status` is `stopped`, `already_gone`, or `still_alive`. + #### POST `/spawn-session` Creates a new session. @@ -259,21 +273,23 @@ Graceful runner shutdown. - `stop-session` - stop session by ID - `stop-runner` - request shutdown -All data is plain JSON over TLS; authentication is `CLI_API_TOKEN` (no end-to-end encryption). +Application payloads are plain JSON authenticated with `CLI_API_TOKEN`. +Transport protection depends on the hub URL: use HTTPS for remote access; +the built-in network relay protects traffic with WireGuard + TLS. ## 7. Process Discovery and Cleanup ### Doctor Command -`hapi doctor` uses `ps aux | grep` to find all HAPI processes: -- Production: matches `hapi` binary, `happy-coder` +`hapi doctor` uses `ps-list` to find HAPI processes: +- Production: matches `hapi` / `hapi.exe` - Development: matches `src/index.ts` (run via `bun`) - Categorizes by command args: runner, runner-spawned, user-session, doctor ### Clean Runaway Processes `hapi doctor clean`: -1. `findRunawayHappyProcesses()` filters for likely orphans +1. `findRunawayHappyProcesses()` selects runner and runner-spawned process categories (not only proven orphans); use with care 2. `killRunawayHappyProcesses()`: - Sends SIGTERM - Waits 1 second @@ -282,9 +298,10 @@ All data is plain JSON over TLS; authentication is `CLI_API_TOKEN` (no end-to-en ## 8. Integration Testing ### Test Environment -- Requires `.env.integration-test` -- Uses local hapi-hub (http://localhost:3006) -- Separate `~/.hapi-dev-test` home directory +- Run `bun run test:cli:integration` from the repo root (separate serial Vitest project) +- Global setup starts an isolated hub on a free loopback port, with a temporary home/database and generated token +- No `.env.integration-test` or running user hub is required +- Real detached process trees are owned and cleaned up by the test harness; stress coverage is opt-in via `HAPI_RUN_STRESS_TESTS=true` ### Key Test Scenarios - Session listing, spawning, stopping @@ -304,6 +321,8 @@ All data is plain JSON over TLS; authentication is `CLI_API_TOKEN` (no end-to-en ## Data Structure (Similar to Session's metadata + agentState) +Simplified excerpts; the complete wire schemas live in `shared/src/schemas.ts`. + ```typescript // Static machine information (rarely changes) interface MachineMetadata { @@ -328,10 +347,9 @@ interface RunnerState { ## 1. CLI Startup Phase -Checks if machine ID exists in settings: -- If not: creates ID locally only (so sessions can reference it) -- Does NOT create machine on hub - that's runner's job -- CLI doesn't manage machine details - all API & schema live in runner subpackage +Authentication/bootstrap ensures a machine ID exists in settings. Session +bootstrap also creates/loads that machine on the hub with its metadata; +the runner supplies live runner state, heartbeats, and machine-scoped RPCs. ## 2. Runner Startup - Initial Registration @@ -457,9 +475,14 @@ RPC method naming (machine-scoped) uses a `${machineId}:` prefix, for example: ## 6. Server Broadcasts to Clients +The Socket.IO examples below are for CLI machine subscribers. Web/native +clients instead receive `machine-updated` via SSE and refetch `/api/machines` +when the event has no machine data. Do not feed the CLI `update` envelope +directly into a native client's SSE decoder. + ### When runner state changes: ```json -// Server -> Mobile/Web clients +// Server -> CLI machine subscribers socket.emit('update', { "id": "update-id-xyz", "seq": 456, @@ -527,7 +550,7 @@ Authorization: Bearer - `runnerStateVersion`: For runner state updates - Allows concurrent updates without conflicts -3. **Security**: No end-to-end encryption (TLS only); CLI auth is a shared secret `CLI_API_TOKEN` +3. **Security**: Plain JSON at the application layer; remote transport uses HTTPS or the encrypted network relay. CLI auth is `CLI_API_TOKEN` 4. **Update Events**: Server broadcasts use same pattern as sessions: - `t: 'update-machine'` with optional metadata and/or runnerState fields @@ -537,15 +560,8 @@ Authorization: Bearer --- -# Improvements +# Operational notes -- runner.state.json file is getting hard removed when runner exits or is stopped. We should keep it around and have 'state' field and 'stateReason' field that will explain why the runner is in that state -- If the file is not found - we assume the runner was never started or was cleaned out by the user or doctor -- If the file is found and corrupted - we should try to upgrade it to the latest version? or simply remove it if we have write access - -- posts helpers for runner do not return typed results -- I don't like that runnerPost returns either response from runner or { error: ... }. We should have consistent envelope type - -- we loose track of children processes when runner exits / restarts - we should write them to the same state file? At least the pids should be there for doctor & cleanup - -- the runner control server binds to `127.0.0.1` on a random port; if we ever expose it beyond localhost, require an explicit auth token/header +- Normal shutdown removes `runner.state.json`; its absence does not prove the runner has never run. Use logs for shutdown history. +- Resume-spawn tracking persists separately in `runner.state.json.resume-processes.json`, with process-generation checks before recovery or termination. It is not a complete inventory of every terminal-started process. +- The local control server binds to `127.0.0.1` on a random port. It has no remote authentication layer; do not expose it through a public proxy. diff --git a/docs/.vitepress/config.ts b/docs/.vitepress/config.ts index c0afdb7f..d115facf 100644 --- a/docs/.vitepress/config.ts +++ b/docs/.vitepress/config.ts @@ -38,7 +38,8 @@ export default defineConfig({ { text: 'Agents', items: [ - { text: 'Agents', link: '/guide/agents' } + { text: 'Agents', link: '/guide/agents' }, + { text: 'Codex Usage & Limits', link: '/guide/codex-shared-sessions' } ] }, { diff --git a/docs/api/client-contract/auth.md b/docs/api/client-contract/auth.md index 4b1bb1d0..71562ffe 100644 --- a/docs/api/client-contract/auth.md +++ b/docs/api/client-contract/auth.md @@ -15,7 +15,7 @@ The hub's base token (`CLI_API_TOKEN`) is auto-generated on first run (32 random ## Pairing -Source of truth: `hub/src/startHub.ts` (lines ~317–366), `web/src/components/settings/CompanionPairing.tsx`. +Source of truth: `hub/src/startHub.ts`, `web/src/components/settings/CompanionPairing.tsx`. The hub terminal (when started with `--relay`) prints two QR codes; the web app's Settings → Companion pairing screen renders the second one as well: @@ -119,10 +119,10 @@ Clients should hide the usage/storage screens entirely when the paired namespace - Store the **access token** in platform-secure storage: iOS Keychain, Android `EncryptedSharedPreferences` (behind an interface so the mechanism can be swapped). Never plain files, never logs. - Key credentials **per hub base URL** (normalized), since a client can pair with several hubs. Web reference: localStorage key `hapi_access_token::` (`web/src/hooks/useAuth.ts`, `web/src/components/settings/CompanionPairing.tsx`). -- The JWT is a cache, not a secret worth keeping: it is fine to hold it in memory only and re-exchange on cold start. If persisted (to save one round-trip at launch), store it alongside the access token with the same protection. -- On unpair/sign-out: delete both credentials, and unregister FCM (`DELETE /api/devices/register`) first while you still hold a valid JWT. +- The JWT is a sensitive bearer credential, but need not be persisted: hold it in memory and re-exchange on cold start. If persisted (to save one round-trip at launch), store it alongside the access token with the same protection. +- On unpair/sign-out: delete both credentials, and unregister native push (`DELETE /api/devices/register`) first while you still hold a valid JWT. -## 401 error bodies +## 401 error bodies {#401-error-bodies} All are JSON with an `error` string; none carry a `code` field except Telegram's `not_bound` (which reuses `error` as the discriminator — natives never see it): diff --git a/docs/api/client-contract/errors.md b/docs/api/client-contract/errors.md index 6cfaa275..558cfb69 100644 --- a/docs/api/client-contract/errors.md +++ b/docs/api/client-contract/errors.md @@ -35,6 +35,7 @@ When no `code` is present, branch on status alone and treat the failure generica | 409 | `resume_unavailable` | `sessions.ts` resume/reopen result mapping | Session can't be resumed (e.g. unsupported state) | | 409 | `metadata_conflict` | `sessions.ts` reopen result mapping | Refetch session, retry once at most | | 409 | `runner_upgrade_required` | `machines.ts` Agent availability | Upgrade and restart the runner; disable session creation | +| 409 | `control_mode_not_applicable` | `sessions.ts` switch | Concurrent clients do not use ownership switching; hide takeover controls | | 409 | — (version conflict) | `sessions.ts` PATCH rename/summary, `machines.ts` PATCH rename — message mentions `version`/`concurrently`; **no code** | Concurrent edit — refetch and reapply | | 409 | — | `sessions.ts` delete-while-active, archive of plain inactive row, fork/rewind refusals, remote-only config on terminal-controlled sessions (`controlledByUser`) | Surface message; refresh session state | | 413 | — | `sessions.ts` upload (> 50 MB decoded), export too large (`{error, count, limit}`); `voice.ts` transcription (`Audio file too large`, 25 MB audio / ~26 MB body) | Reduce payload | @@ -56,7 +57,10 @@ Many endpoints do not answer from hub state — the hub relays the request over 2. **CLI offline / handler missing** → depends on the route: the model-catalog routes in `machines.ts` map `RpcTargetMissingError` to **503 `rpc_target_missing`**; `git.ts`-style routes fold it into the 200 `{success: false}` envelope; resume/reopen surface **503 `no_machine_online`**. 3. **Hub subsystems not up** → **503 `Not connected`** from `requireSyncEngine` (brief startup/shutdown window). -Practical rule: treat `success: false`, 503 `rpc_target_missing`, and 503 `no_machine_online` as the same user-facing condition — "the computer running this session is not reachable" — with the raw `error` string available in a details view. +Explicit `rpc_target_missing` / `no_machine_online` failures mean the execution +host or handler is unavailable. Do not infer that from `success: false` alone: +a reachable CLI can report a command, path, permission, or validation failure. +Preserve that error for display; HTTP success is not operation success. One route qualifies that rule. `GET /api/machines/:id/agy-models` keeps serving the last catalog the machine got out of `agy models` while the CLI re-checks in the background, so it can answer `success: true` **and** carry an `error`: the list is usable, and `error` says why it may be stale (typically the machine's agy sign-in has lapsed). Render it beside the catalog rather than instead of it, and offer `?refresh=true` as the way to ask again — a plain repeat is answered from the same cache. @@ -68,5 +72,5 @@ When that background re-check lands a different listing, the machine says so ove |-------|--------| | 400 / 403 / 404 / 409 / 413 / 422 | No (fix input, refresh state, or hide surface) | | 401 (middleware) | Once, after silent re-auth ([Auth](./auth.md#silent-re-auth-401-handling)) | -| 429 / 502 / 503 | Yes, with backoff | -| 200 `{success: false}` | Manual retry only (user-initiated) — the CLI answered and said no | +| 429 / 502 / 503 | Reads: retry with backoff. Mutations: follow the endpoint's recovery contract; do not replay a write whose outcome is unknown | +| 200 `{success: false}` | Inspect the error and endpoint contract; manual retry only when safe. This envelope can also represent a missing RPC target or timeout | diff --git a/docs/api/client-contract/index.md b/docs/api/client-contract/index.md index 952ac7c8..a0cd2617 100644 --- a/docs/api/client-contract/index.md +++ b/docs/api/client-contract/index.md @@ -2,7 +2,7 @@ **Audience:** Implementers of native HAPI clients — the iOS app (`ios/`), the Android app (`android/`), and any other non-web client that talks to a hub's client API. These pages are the primary spec for that work: every claim is grounded in hub/web source, and each section names its source file so implementers (human or AI agent) can verify against code. -**Scope:** The HTTP contract between a client and one hub — pairing and auth, REST endpoints, SSE streaming, message pagination, message decoding, and error semantics. A client using only this contract can replicate the web app's core feature set over **REST + SSE alone** (no Socket.IO — that transport is CLI↔hub internal). +**Scope:** The HTTP contract between a client and one hub — pairing and auth, REST endpoints, SSE streaming, message pagination, message decoding, and error semantics. Core session/chat features use **REST + SSE**. Web terminals additionally use the JWT-authenticated Socket.IO `/terminal` namespace, outside this contract; `/cli` is the internal CLI↔hub namespace and is not a native-client API. ## Pages @@ -39,8 +39,15 @@ Source of truth: `hub/src/web/server.ts` (`/health` route), `shared/src/version. The prose in [Messages](./messages.md) describes the decoding tree, but the *normative* artifact is `shared/fixtures/` — machine-generated golden files produced from the web implementation's chat pipeline (`web/src/chat/`). A native client's protocol module must reproduce those fixtures exactly; CI regenerates them whenever the web pipeline changes, so drift is caught automatically. -`shared/fixtures/` is a companion deliverable of this contract and may not exist yet when you first read this — the fixture generator and batches land in later work packages of the same track. Until then, `web/src/chat/` itself is the reference implementation. +The fixtures and native conformance suites are present in this repository. +Regenerate fixtures from the repo root with `bun run gen:fixtures`; never +hand-edit generated JSON. See the native package READMEs for conformance checks. ## Relationship to the companion push contract -[`docs/api/native-companion-contract.md`](../native-companion-contract.md) is the **FCM push contract**: device registration (`POST /api/devices/register`) and the outbound push payload the hub sends through Firebase. It predates this contract and is unchanged. A native client implements *both*: this contract for everything interactive, the companion contract for background push. Where the two overlap (auth, send-message, approve/deny), this contract is the more detailed spec. +The [native companion contract](../native-companion-contract.md) specifies +Android/iOS device registration (`POST /api/devices/register`), encrypted +push envelopes, direct FCM/APNs delivery, and the shared push relay. A native +client implements this contract for interactive features and that contract +for background push. Where they overlap (auth, send-message, approve/deny), +this contract is the more detailed spec. diff --git a/docs/api/client-contract/messages.md b/docs/api/client-contract/messages.md index 815d17f6..2fda7c09 100644 --- a/docs/api/client-contract/messages.md +++ b/docs/api/client-contract/messages.md @@ -21,7 +21,10 @@ type DecryptedMessage = { } ``` -`content` is deliberately `unknown` on the wire. **Decoding must be total**: malformed content degrades to a stringified fallback — a client must never drop or crash on a message it does not recognize (with the two precise exceptions listed in [Fallback rules](#fallback-rules)). +`content` is deliberately `unknown` on the wire. **Decoding must be total**: +never crash on unfamiliar content. Use stringified fallbacks for unknown +envelopes, and follow the family-specific skip/validation rules below for +known transport records (see [Fallback rules](#fallback-rules)). --- @@ -211,13 +214,22 @@ Several event rows are also synthesized by the other two families (system subtyp | `'output'` family, visible but unknown `data.type` | stringify as agent text | | `'event'` family, `data` lacks a string `type` | stringify as agent text | -"Stringify" = a stable JSON serialization (web: `safeStringify`) rendered as plain text. These are the only two legitimate drop paths; everything else must render something. +"Stringify" = a stable JSON serialization (web: `safeStringify`) rendered as +plain text. Known event types also have the validation/empty-content skip +rules listed above (for example, an image without an ID or unparseable usage). +The golden fixtures and normalizer are authoritative; do not turn those +transport-only records into fallback chat bubbles. --- ## Truncation marker -At ingest the hub head+tail-truncates any **string longer than 64 KiB found anywhere inside agent-role content** (`hub/src/store/contentCodec.ts`): the stored value becomes first 48 KiB + `\n…[hapi: truncated N chars]…\n` + last 12 KiB. User-role content is never truncated (it is delivered verbatim to the CLI). The operation is idempotent and applied deep (arrays/objects). +At ingest the hub head+tail-truncates strings longer than `64 * 1024` +**UTF-16 code units** (`String.length`, not bytes) inside agent-role content +(`hub/src/store/contentCodec.ts`): first `48 * 1024` units + +`\n…[hapi: truncated N chars]…\n` + last `12 * 1024` units. User-role content +is never truncated (it is delivered verbatim to the CLI). The operation is +idempotent and applied deep (arrays/objects). Clients must render truncated strings as-is (recognizing the `…[hapi: truncated N chars]…` marker is optional polish), must not assume tool results are complete, and must never choke on the marker. @@ -232,7 +244,7 @@ session.agentState = { requests?: Record completedRequests?: Record // flat (AskUserQuestion) @@ -243,6 +255,10 @@ session.agentState = { `agentState` updates arrive as a versioned SSE patch — apply it under the version gate described in [sse.md](./sse.md#versioned-patch-algorithm). Render pending `requests` as approval cards interleaved with the chat (the web reducer keys them to the matching `tool_use` when one exists); on resolution the entry moves to `completedRequests`, whose `status`/`answers` back-fill the tool card's permission state. Decide via `POST /api/sessions/:id/permissions/:requestId/approve` (`{mode?, allowTools?, decision?, answers?}`) or `…/deny` (`{decision?}`) — see [rest.md](./rest.md). Session-list badges come precomputed on `SessionSummary.pendingRequestsCount` / `pendingRequests` (≤ 5 entries). +`resolved` means native completion is known, but the winning answer/decision +is not. Render it neutrally; never infer approval or fill answers from an +unconfirmed local draft. + Correlate a permission with the transcript using **`entry.toolCallId ?? requestId`**, but always submit approval/denial using **`requestId`**. Claude local-mode requests use independent, one-shot reply IDs so stale responses cannot answer a later diff --git a/docs/api/client-contract/pagination.md b/docs/api/client-contract/pagination.md index 6dd5e4bc..03e0a721 100644 --- a/docs/api/client-contract/pagination.md +++ b/docs/api/client-contract/pagination.md @@ -140,7 +140,7 @@ Lifecycle: 2. On POST success: status → `queued` if the session is currently thinking, else `sent`. On failure: drop the row and restore the composer (or keep it as `failed` with a retry affordance when attachments are involved). 3. **Echo**: the hub emits `message-received` carrying the stored row (server `id`, real `seq`, same `localId`). Merging a stored row whose `localId` matches an optimistic row **replaces** the optimistic one, preserving the client-side `status` and any already-known `invokedAt` the server row lacks. Fallback when no `localId` echo matches: drop an optimistic `sent` row when a server user message lands within **10 s** of the same position. 4. **`messages-consumed {localIds, invokedAt}`** (SSE): stamp `invokedAt` and flip status to `sent` on matching rows (skip `failed` ones). This is what moves a message out of the queued bar and into the thread at its invocation position. -5. **`messages-indeterminate {localIds}`** (SSE): the steer outcome is unknown. Keep `invokedAt: null`, mark `deliveryState:'indeterminate'`, exclude the row from automatic replay, and show explicit Retry/Cancel actions. +5. **`messages-indeterminate {localIds}`** (SSE): a native dispatch or queue mutation has an unknown outcome. Keep `invokedAt: null`, mark `deliveryState:'indeterminate'`, and exclude the row from automatic replay. Retry/Cancel are explicit resolution actions and may remain unavailable until the native outcome can be reconciled. 6. **`messages-requeued {localIds}`** (SSE): an explicit Retry restored normal queue delivery; clear `deliveryState`. 7. **`message-cancelled {messageId, localId?}`** (SSE): remove the row (match either id). @@ -152,8 +152,8 @@ After a reconnect whose handshake said `resume: 'gap'` (an `ok` resume replayed 1. Finish a tail sync. 2. Collect candidate `localId`s: user rows with `invokedAt === null`, excluding optimistic rows still `sending`/`failed`. -3. `POST /api/sessions/:id/messages/queued-state` with `{"localIds": […]}` (max 1000 per call; batch above that) → `{queuedLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}`. -4. Apply `invokedLocalMessages` exactly like `messages-consumed`; drop candidates that are in **neither** list (deleted server-side). +3. `POST /api/sessions/:id/messages/queued-state` with `{"localIds": […]}` (max 1000 per call; batch above that) → `{queuedLocalIds: string[], indeterminateLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}`. +4. Apply `invokedLocalMessages` exactly like `messages-consumed`; mark `indeterminateLocalIds` as unresolved delivery. Retain both queued and indeterminate rows; drop only candidates absent from **all three** result groups. An in-flight native dispatch is reported as indeterminate, not as a deleted message. --- @@ -181,14 +181,20 @@ Response `{"ok": true}`. Sending to an inactive session returns `409 {"error":"S |---|---|---| | `{"status":"cancelled","localId":string\|null}` | Row deleted (or already gone). Bumps the epoch. | Remove the row. | | `{"status":"invoked","message":DecryptedMessage}` | Too late — the agent consumed it before the cancel landed. | **Ingest the returned message** as the authoritative row (correct `invokedAt`, status `sent`); do not resurrect the queued snapshot. | -| `{"status":"busy","localId":string}` | A live steer is still resolving. | Restore the row as indeterminate; reconcile queued state before allowing Retry/Cancel. | +| `{"status":"busy","localId":string}` | Native delivery/removal is unresolved; cancellation cannot be confirmed. | Restore the row as indeterminate; reconcile queued state before allowing Retry/Cancel. | Other subscribers learn the same outcome via `message-cancelled` / `messages-consumed` SSE events. -**Steer a queued message into the current turn**: `POST /api/sessions/:id/messages/:messageId/steer` (Pi sessions) → `SteerQueuedMessageResponseSchema`: +**Steer a queued message into the current turn**: `POST /api/sessions/:id/messages/:messageId/steer` → `SteerQueuedMessageResponseSchema`. Unlike the send-time `deliveryMode` option above, this endpoint supports Pi, Codex, and Cursor ACP sessions (`isSteeringSupportedForSession` in `shared/src/modes.ts`). It rejects all scheduled messages, and rejects terminal-controlled sessions unless they advertise `concurrentClients`. | Response | Client action | |---|---| | `{"status":"steered","localId"}` | Keep the row queued-side; it is being injected into the live turn. | | `{"status":"invoked","message"}` | Already consumed — ingest the message. | -| `{"status":"failed","error","localId":string\|null}` | Surface the error; the row remains queued. | +| `{"status":"failed","error","localId":string\|null}` | Surface the error. Do not infer delivery state from this alone; reconcile before retrying when the native outcome is unknown. | + +**Retry indeterminate delivery**: `POST /api/sessions/:id/messages/:messageId/retry` +is user-initiated only. `retried` or `already-queued` means normal queue delivery; +`invoked` carries the authoritative message; `not-found` means the row is gone. +`retry-unavailable` leaves the row unresolved: the hub could not prove that +retrying would avoid duplicate work. Reconcile instead of automatically retrying. diff --git a/docs/api/client-contract/rest.md b/docs/api/client-contract/rest.md index 786cd0ff..b2d85f1b 100644 --- a/docs/api/client-contract/rest.md +++ b/docs/api/client-contract/rest.md @@ -4,7 +4,7 @@ Endpoint tables for native clients, grouped by feature. Request/response shapes ## Conventions -- All paths below are relative to the hub base URL. Everything under `/api` requires `Authorization: Bearer ` ([Auth](./auth.md)). +- All paths below are relative to the hub base URL. `/api` routes require `Authorization: Bearer ` except the `/api/auth` exchange and Telegram `/api/bind` ([Auth](./auth.md)). - Path params (`:id`, `:messageId`, …) must be URL-encoded (the web client uses `encodeURIComponent` throughout). - Request bodies are JSON (`content-type: application/json`) with **one exception**: `POST /api/voice/transcription` is `multipart/form-data`. Responses are JSON unless noted (generated images and scratchlist attachments return raw bytes). - Bodies are validated with Zod; failures return `400` (see [Errors](./errors.md)). @@ -66,10 +66,10 @@ Source: `hub/src/web/routes/messages.ts`; schemas `MessagesQuerySchema`, `SendMe |---|---|---| | `GET /api/sessions/:id/messages` | Query: `limit?` (1–200, default 50), cursor pairs `beforeSeq+beforeAt` \| `afterSeq+afterAt` (+ optional `untilSeq+untilAt`, `epoch` with `after`) | `MessagesResponse` `{messages: DecryptedMessage[], page: {direction, limit, epoch, reset, nextBefore*/nextAfter*, snapshotHead*, hasMore}}` — full cursor semantics in [Pagination](./pagination.md) | | `POST /api/sessions/:id/messages` | `{text, localId?, attachments?, scheduledAt?, deliveryMode?: 'queue'\|'steer'}` — text or attachments required; `scheduledAt` requires `localId`, must be ≤ 7 days out, excludes attachments and steer | `{ok: true}` — the message itself arrives via SSE (`message-received`), reconciled by `localId` | -| `DELETE /api/sessions/:id/messages/:messageId` | — | `{status: 'cancelled', localId}` \| `{status: 'invoked', message}` \| `{status: 'busy', localId}` (cancel; `busy` = steer still resolving) | +| `DELETE /api/sessions/:id/messages/:messageId` | — | `{status: 'cancelled', localId}` \| `{status: 'invoked', message}` \| `{status: 'busy', localId}` (cancel; `busy` = native delivery/removal is unresolved) | | `POST /api/sessions/:id/messages/:messageId/steer` | — | `{status: 'steered', localId}` \| `{status: 'invoked', message}` \| `{status: 'failed', error, localId}` | | `POST /api/sessions/:id/messages/:messageId/retry` | — | `{status: 'retried', localId}` \| `{status: 'already-queued', localId}` \| `{status: 'retry-unavailable', localId}` \| `{status: 'invoked', message}` \| `{status: 'not-found'}` — explicit retry only; never automatic replay | -| `POST /api/sessions/:id/messages/queued-state` | `{localIds: string[]}` (≤ 1000, deduped) | `{queuedLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}` — resync optimistic sends after reconnect | +| `POST /api/sessions/:id/messages/queued-state` | `{localIds: string[]}` (≤ 1000, deduped) | `{queuedLocalIds: string[], indeterminateLocalIds: string[], invokedLocalMessages: [{localId, invokedAt}]}` — resync after reconnect; preserve indeterminate rows without auto-replaying them | The hub stamps `sentFrom: 'webapp'` on REST-sent messages server-side; the request body has no such field. @@ -97,8 +97,8 @@ Source: `hub/src/web/routes/sessions.ts`; flavor gates in `shared/src/modes.ts` | Method & path | Request | Applies to | |---|---|---| -| `POST /api/sessions/:id/permission-mode` | `{mode: PermissionMode}` | All flavors except `pi` (per-flavor allowed sets in `modes.ts`) | -| `POST /api/sessions/:id/model` | `{model: string \| {provider, modelId} \| null}` | All flavors (`supportsModelChange` is true for every current flavor); remote-only for codex/cursor/grok | +| `POST /api/sessions/:id/permission-mode` | `{mode: PermissionMode}` | All flavors except `pi` and `dsh` (per-flavor allowed sets in `modes.ts`) | +| `POST /api/sessions/:id/model` | `{model: string \| {provider, modelId} \| null}` | Flavors with `supportsModelChange` (not `dsh`); remote-only for codex/cursor/grok | | `POST /api/sessions/:id/effort` | `{effort: string \| null}` | claude, grok, pi (`supportsEffort`) | | `POST /api/sessions/:id/model-reasoning-effort` | `{modelReasoningEffort: string \| null}` | codex, opencode (remote-only) | | `POST /api/sessions/:id/service-tier` | `{serviceTier: 'fast' \| 'standard'}` | codex (remote-only) | @@ -110,6 +110,8 @@ Source: `hub/src/web/routes/sessions.ts`; flavor gates in `shared/src/modes.ts` ownership mode. `/switch` returns HTTP 409 with `code: 'control_mode_not_applicable'`; do not offer takeover. The remote-only Codex configuration restrictions above do not apply to these sessions. +Shared Codex rejects `safe-yolo` even though it remains in the historical +Codex permission-mode schema; do not offer it for concurrent sessions. Shared `/clear` has no global `supersededBySessionId` change; clients other than the caller stay on the original thread. Fork may return an already-bound shared @@ -262,7 +264,7 @@ These exist on the hub but v1 native clients must not implement or call them: | Codex Desktop import | `/api/codex/*` (`hub/src/web/routes/codexDesktop.ts`) | Desktop-import tooling | | Pi session import | `/api/pi/*`, `/api/sessions/:id/pi-*` (`hub/src/web/routes/piSessions.ts`, `sessions.ts`) | Import tooling (the `pi-models` catalog above is the one exception) | | Work graph | `/api/work-graph/*` (`hub/src/web/routes/workGraph.ts`) | Web-only feature | -| Web Push | `/api/push/*` (`hub/src/web/routes/push.ts`) | Browser Push API; natives use `/api/devices` (FCM) | +| Web Push | `/api/push/*` (`hub/src/web/routes/push.ts`) | Browser Push API; natives use `/api/devices/register` (Android/iOS push contract) | | Hub settings write | `PUT /api/hub-settings` | Owner-only hub administration | | Telegram | `POST /api/bind` | Telegram Mini App binding only | | Voice assistant | `/api/voice/token`, `/voices`, `/backend`, `/gemini-token`, `/qwen-token`, `/qwen-ws`, `/gemini-ws`, `/transcription/realtime-token`, `/telemetry`, credentials endpoints | Realtime assistant, not v1 dictation | diff --git a/docs/api/client-contract/sse.md b/docs/api/client-contract/sse.md index 3f6a7caa..3bb3a39a 100644 --- a/docs/api/client-contract/sse.md +++ b/docs/api/client-contract/sse.md @@ -192,9 +192,9 @@ Reference list sort (web): `globalPinned` > `pinned` > `active` > `pendingReques `POST /api/visibility` with body `{"subscriptionId": "", "visibility": "visible" | "hidden"}` → `{"ok": true}`. Errors: `400` invalid body, `404` unknown `subscriptionId` (or namespace mismatch), `503` hub not ready. Each new connection has a **new** `subscriptionId` — re-report after every reconnect (the web reference reports both of its connections on every foreground/background transition and retries a failed report after 2 s). -Semantics (`hub/src/visibility/visibilityTracker.ts`, `hub/src/push/pushNotificationChannel.ts`): when **any** connection in the namespace is visible, the hub delivers notification events (ready / permission request / task result) as in-app **`toast` SSE frames to the visible connections** and suppresses Web Push for the namespace; Web Push fires only when no visible connection exists (or toast delivery reached zero connections). Native FCM devices (`POST /api/devices/register`) are independent of visibility and fire unconditionally — see [native-companion-contract](../native-companion-contract.md). +Semantics (`hub/src/visibility/visibilityTracker.ts`, `hub/src/push/pushNotificationChannel.ts`): Android/iOS native delivery is independent of Web visibility. When a native provider accepts a notification for at least one device, the hub skips the Web Push/toast duplicate for that dispatch. Otherwise, if **any** connection in the namespace is visible, the hub first sends **`toast` SSE frames to visible connections**. Web Push is the fallback when no connection is visible or toast delivery reaches zero connections. Provider acceptance is not a handset delivery receipt — see the [native companion contract](../native-companion-contract.md). -Native rule: report `visible` on foreground and `hidden` on background, every time. A native client that stays `visible` while backgrounded suppresses its own (and every PWA's) hub-side push for the namespace, and receives its notifications only as toast frames nobody is looking at. +Native rule: report `visible` on foreground and `hidden` on background, every time. A stale `visible` report can divert the namespace's Web Push fallback into unseen toast frames; it does not disable native push. The native app separately suppresses its local notification when the corresponding chat is already open in the foreground. --- diff --git a/docs/api/native-companion-contract.md b/docs/api/native-companion-contract.md index 4fc59545..703c6d1b 100644 --- a/docs/api/native-companion-contract.md +++ b/docs/api/native-companion-contract.md @@ -9,7 +9,7 @@ binding (requires Telegram `initData`). ## Scope -A companion implementing this contract is a **native client to the same hub the PWA talks to**, surfacing notifications and reply / approve actions on a phone or wearable. Hub topology is unchanged - the hub still runs on the operator's dev machine. +A companion implementing this contract is a **native client to the same hub the PWA talks to**, surfacing notifications and reply / approve actions on a phone or wearable. The hub may run on the operator's development machine or a separate host; agents execute on their CLI/Runner machines. --- @@ -40,7 +40,7 @@ registry; no schema or database version change is required. **Response:** `{ "ok": true }` -Upsert on `(namespace, deviceId, platform)` - same device re-registering replaces the FCM token. +Upsert on `(namespace, deviceId, platform)` - same device re-registering replaces its push token. ### Unregister @@ -188,7 +188,7 @@ AAD = ASCII "hapi-push-v1" ``` Golden test vector (key `0x00..0x1f`, nonce `0x00..0x0b`): -[`shared/fixtures/push/envelope-v1.json`](../../shared/fixtures/push/envelope-v1.json) - +[`shared/fixtures/push/envelope-v1.json`](https://github.com/tiann/hapi/blob/main/shared/fixtures/push/envelope-v1.json) - the iOS implementation must reproduce it byte-for-byte. ### APNs request (what the device receives) diff --git a/docs/guide/agents.md b/docs/guide/agents.md index e314d0aa..79fccd3d 100644 --- a/docs/guide/agents.md +++ b/docs/guide/agents.md @@ -13,7 +13,7 @@ agent's integration; their supported syntax varies by agent. | Agent | Command | Integration | Local | Remote | Permission modes | Resume | |-------|---------|-------------|:-----:|:------:|------------------|:------:| | Claude Code | `hapi claude` | Terminal wrapper (local) + Claude Agent SDK (remote) | ✓ | ✓ | `default` `acceptEdits` `auto` `bypassPermissions` `plan` | ✓ | -| Codex | `hapi codex` | TUI wrapper (local) + `codex app-server` JSON-RPC (remote) | ✓ | ✓ | `default` `read-only` `safe-yolo` `yolo` (+ `plan` collaboration mode) | ✓ | +| Codex | `hapi codex` | Native terminal + `codex app-server` (Codex 0.154.0+) | ✓ | ✓ | `default` `read-only` `yolo` (+ `plan` collaboration mode) | ✓ | | Cursor Agent | `hapi cursor` | ACP (`agent acp`); legacy stream-json resume | ✓ | ✓ | `default` `plan` `ask` `debug` `autoReview` `yolo` | ✓ | | Grok Build | `hapi grok` | ACP (`grok agent stdio`) | ✓ | ✓ | `default` `auto` `plan` `bypassPermissions` | ✓ | | GitHub Copilot | `hapi copilot` | ACP (`copilot --acp --stdio`) | ✓ | ✓ | `default` `read-only` `safe-yolo` `yolo` | ✓ | @@ -38,10 +38,10 @@ Permission modes are per-agent — each flavor exposes its own set (see the matr ### Local and remote mode -Every session is either **local** (driven from the terminal) or **remote** (driven from web/phone). DSH is remote-only because its ACP server has no local terminal surface. Switching is seamless and keeps the same session state for flavors that support both: +Work **locally** in the terminal or **remotely** from web/phone, keeping the same conversation when you hand off. The support matrix shows which interfaces each agent offers; DSH, Pi, and Antigravity accept input only through HAPI's remote interface. -- **Remote → local:** press double-space in the terminal. -- **Local → remote:** send a message from the web UI or phone; the session switches automatically. +- **Remote → local:** continue in the terminal. If it shows the remote-control screen, press double-space to return to local input. +- **Local → remote:** send a message from the web UI or phone; HAPI handles the handoff. See [Seamless Handoff](./how-it-works.md#seamless-handoff) for details. @@ -52,7 +52,7 @@ hapi resume # Interactive picker of resumable sessions on this ma hapi resume # Resume a specific HAPI session ``` -`hapi resume` works for every resumable flavor except Gemini and fresh-session-only DSH. An active remote session is handed off to the local terminal first. Pi and Antigravity are the exceptions in the other direction: neither has a local input path, so their sessions always resume in remote mode. +`hapi resume` reopens the conversation on this machine, including active sessions you were using from your phone. Gemini and fresh-session-only DSH cannot be resumed. Pi and Antigravity resume with input still controlled from HAPI rather than the terminal. ## Cursor Agent @@ -271,7 +271,7 @@ overall permission policy. ## Other agents - **Claude Code** (`hapi claude`) — local sessions wrap the native TUI, remote sessions drive the Claude Agent SDK. [Claude Code docs](https://docs.anthropic.com/en/docs/claude-code) -- **Codex** (`hapi codex`) — OpenAI's Codex CLI; remote sessions talk to `codex app-server` over JSON-RPC, with a dedicated `plan` collaboration mode. [openai/codex](https://github.com/openai/codex) +- **Codex** (`hapi codex`) — OpenAI's Codex CLI, with terminal/Web control and a dedicated `plan` mode. See [Codex usage and limits](./codex-shared-sessions.md) for resume, terminal-exit behavior, and launch options. [openai/codex](https://github.com/openai/codex) - **GitHub Copilot** (`hapi copilot`) — Copilot CLI over ACP (`copilot --acp --stdio`). [GitHub Copilot](https://github.com/features/copilot) - **Kimi** (`hapi kimi`) — Moonshot AI's Kimi CLI over ACP (`kimi acp`). [MoonshotAI/kimi-cli](https://github.com/MoonshotAI/kimi-cli) - **OpenCode** (`hapi opencode`) — the open-source OpenCode agent over ACP (`opencode acp`). [opencode.ai](https://opencode.ai) diff --git a/docs/guide/deployment.md b/docs/guide/deployment.md index 36cfa77c..7d10a9db 100644 --- a/docs/guide/deployment.md +++ b/docs/guide/deployment.md @@ -50,22 +50,32 @@ https://tailscale.com/download ```bash sudo tailscale up -hapi hub ``` -Access via your Tailscale IP: +The default hub listens only on loopback. For browser access directly via +your Tailscale IP, start it with: + +```bash +HAPI_LISTEN_HOST=0.0.0.0 hapi hub +``` + +Restrict inbound access to the trusted network with your firewall, then open: ``` http://100.x.x.x:3006 ``` + +For native apps and HTTPS-only browser features, use an HTTPS endpoint +(for example, Tailscale Serve forwarding to the loopback hub) instead.
Public IP / Reverse Proxy -If the hub has a public IP, access directly via `http://your-hub-ip:3006`. - -Use HTTPS (via Nginx, Caddy, etc.) for production. +Keep the hub on its default `127.0.0.1:3006` and put an HTTPS reverse proxy +(Nginx, Caddy, etc.) in front of it. A public IP alone does not make the +loopback listener remotely accessible. Set `HAPI_LISTEN_HOST` only when you +need another bind address, and restrict direct access with your firewall. **Self-signed certificates (HTTPS)** @@ -104,6 +114,7 @@ Simple one-liner for quick background runs: ```bash # Hub +mkdir -p ~/.hapi/logs nohup hapi hub --relay > ~/.hapi/logs/hub.log 2>&1 & # Runner @@ -282,7 +293,7 @@ RestartSec=5 WantedBy=default.target ``` -> **Why `KillMode=process`?** The runner spawns each agent session as a detached child process (`detached: true` in `cli/src/runner/run.ts`) so that sessions stay alive when the runner exits. Without `KillMode=process`, systemd's default `KillMode=control-group` sends SIGTERM to every PID in the runner's cgroup when the unit stops, defeating the detach and forcibly archiving every running session. `KillMode=process` preserves the contract: stopping or restarting the runner only signals the runner itself; agent sessions stay alive, and a fresh runner re-establishes control via the existing socket.io reconnect path. This applies to runner upgrades, manual restarts, and any reboot in which the runner unit is stopped before agents have finished. +> **Why `KillMode=process`?** Agent sessions are detached from the runner. The default `KillMode=control-group` would terminate them when the runner service stops; `KillMode=process` preserves them during runner restarts and upgrades. It does not keep processes running through an operating-system shutdown or sleep. Enable and start: diff --git a/docs/guide/faq.md b/docs/guide/faq.md index f57ed5f7..5f16548f 100644 --- a/docs/guide/faq.md +++ b/docs/guide/faq.md @@ -72,11 +72,12 @@ Yes. Telegram is optional. You can use the web app directly in any browser or in ### How do I receive notifications? -HAPI supports three methods: +HAPI supports these notification channels: 1. **PWA Push Notifications** - Enable when prompted, works even when app is closed 2. **Telegram Bot** - See [Telegram Setup](./notifications.md#telegram-setup) -3. **FCM native push** - Used by the Android/Wear OS companion apps; notifications are delivered via Firebase Cloud Messaging +3. **Native app notifications** - Official Android and iOS apps use encrypted push delivery; pair your hub and allow notifications, with no push-provider setup required +4. **ServerChan (Server酱)** - Send notifications to WeChat and other channels; see [ServerChan Setup](./notifications.md#serverchan-server酱-setup) ### Can I start sessions remotely? @@ -99,7 +100,7 @@ Yes. Open any session and use the chat interface to send messages directly to th ### Why did my session look idle when the agent woke itself? -Some agents (especially Cursor) can resume after idle from harness signals such as background Shell `notify_on_output` or `/loop`, without you sending a new HAPI message. HAPI treats real ACP agent activity (and permission requests) as thinking again so the session list matches the agent - same keepalive path as a normal turn. This is different from session-attached jobs (`hapi job`), which show progress while the agent stays idle on purpose. +Some agents (especially Cursor) can resume after idle from harness signals such as background Shell `notify_on_output` or `/loop`, without you sending a new HAPI message. HAPI updates the session's thinking indicator when the agent resumes work or requests permission, so the list reflects that activity. ### Can I access a terminal remotely? @@ -115,10 +116,11 @@ The voice assistant supports three backends: ElevenLabs, Gemini Live, and Qwen R ### Is my data safe? -Yes. HAPI is local-first: -- All data stays on your machine -- Nothing is uploaded to external servers -- The database is stored locally in `~/.hapi/` +HAPI keeps session history on the hub you operate, in `~/.hapi/` by default, +rather than on a central HAPI account server. Your devices connect to that hub. +Coding agents still use their configured model providers; optional voice, +title generation, and notification features also contact external services. +See the [Privacy Policy](../privacy.md) for data handling and encrypted native push. ### How secure is the token authentication? @@ -261,7 +263,7 @@ hapi doctor clean | Design | Cloud-first | Local-first | | Users | Multi-user | Single user by default; lightweight multi-account isolation via [namespaces](./namespace.md) | | Deployment | Multiple services | Single binary | -| Data | Encrypted on server | Never leaves your machine | +| Session history | Encrypted on server | Stored on your own hub | See [Why HAPI](./why-hapi.md) for detailed comparison. diff --git a/docs/guide/how-it-works.md b/docs/guide/how-it-works.md index bbc0139d..fdbf92cd 100644 --- a/docs/guide/how-it-works.md +++ b/docs/guide/how-it-works.md @@ -1,9 +1,5 @@ # How it Works -**Codex 0.154.0+** uses a [shared app-server](./codex-shared-sessions.md): -terminal and Web/phone can act simultaneously, without switching ownership. -The local/remote handoff descriptions below apply to other agent integrations. - HAPI consists of three interconnected components that work together to provide remote AI agent control. ## Architecture Overview @@ -66,7 +62,7 @@ hapi runner start # Run background service for remote session spawning hapi ping-peer --list # Shell peer shortlist (prefer MCP list_peers in-session) ``` -MCP peer tools (same hub/namespace as the session): `list_peers` (discover), `inspect_peer` (read), `ping_peer` (message). These work from runner-spawned sessions even when the hub is on another host - see [Installation → Split hub + remote runner](./installation.md#split-hub--remote-runner-peer-discovery). +MCP peer tools (same hub/namespace as the session): `list_peers` (discover), `inspect_peer` (read), `ping_peer` (message). These work from runner-spawned sessions even when the hub is on another host - see [Installation → Split hub + remote runner](./installation.md#split-hub-remote-runner-peer-discovery). ### HAPI Hub @@ -86,9 +82,9 @@ A React-based PWA that provides the mobile interface: - **Chat Interface** - Send messages and view agent responses - **Permission Management** - Approve or deny tool access - **File Browser** - Browse project files and view git diffs -- **Terminal View** - Watch the full terminal output of a session +- **Terminal View** - Run commands on the working machine from your browser - **Voice Assistant** - Talk to your agent and approve permissions by voice (see [Voice input and assistant](./voice-assistant.md)) -- **Session Sharing** - Share a read-only view of a session via a link +- **Session References** - Copy a session reference or mention another conversation for context - **Remote Spawn** - Start new sessions on any connected machine ## Data Flow @@ -96,10 +92,10 @@ A React-based PWA that provides the mobile interface: ### Starting a Session ``` -1. User runs `hapi` in terminal +1. User runs `hapi` and chooses an agent │ ▼ -2. CLI starts Claude Code (or other agent) +2. CLI starts the selected agent │ ▼ 3. CLI connects to hub via Socket.IO @@ -188,7 +184,7 @@ When working in local mode, you have the full terminal experience — it is the - Direct keyboard input with instant response - Full terminal UI with syntax highlighting - Best for focused, uninterrupted coding sessions -- All AI processing happens locally on your machine +- Agent tools run on your machine; model requests use the provider configured in the agent ### Remote Mode @@ -213,14 +209,15 @@ Switch to remote mode when you need to step away: ``` **Local → Remote:** -- Receive a message from phone/web -- Session automatically switches to remote mode -- Terminal shows "Remote mode - waiting for input" +- Open the session on your phone/web and send a message +- HAPI keeps the conversation going on the same working machine **Remote → Local:** -- Press double-space in terminal -- Instantly regain local control -- Continue typing as if you never left +- Continue typing in the terminal +- If the terminal shows the remote-control screen, press double-space to return to local input + +Some agents keep both interfaces available at once, so no switch is needed. +For Codex terminal-exit and resume behavior, see [Usage and limits](./codex-shared-sessions.md). ### Use Cases diff --git a/docs/guide/installation.md b/docs/guide/installation.md index 606d29ec..96aa63c4 100644 --- a/docs/guide/installation.md +++ b/docs/guide/installation.md @@ -196,11 +196,18 @@ On first run, HAPI: | `HAPI_RELAY_FORCE_TCP` | `false` | - | Force TCP mode for relay | | `HAPI_OFFICIAL_WEB_URL` | `https://app.hapi.run` | - | Official web app origin, added to CORS when the relay is enabled | | `VAPID_SUBJECT` | `mailto:admin@hapi.run` | - | Web Push contact info | +| `HAPI_ANDROID_PUSH` | `auto` | `androidPushMode` | Android push: official relay by default, direct FCM when private credentials are configured; also accepts `relay`, `fcm`, `off` | +| `HAPI_IOS_PUSH` | `relay` | `iosPushMode` | iOS push: `relay`, direct `apns`, or `off` | +| `HAPI_PUSH_RELAY_URL` | `https://push.hapi.run` | `iosPushRelayUrl` | Shared Android/iOS push relay, independent of the network tunnel | +| `FCM_SERVICE_ACCOUNT_PATH` | - | `fcmServiceAccountPath` | Direct FCM credentials for private builds using the same Firebase project | | `HAPI_HOME` | `~/.hapi` | - | Config directory path | | `DB_PATH` | `~/.hapi/hapi.db` | - | Database file path | | `HAPI_EXPERIMENTAL` | - | - | CLI: enable experimental features (`true`/`1`/`yes`) | | `ELEVENLABS_API_KEY` | - | Settings / env | ElevenLabs API key for voice + dictation | | `ELEVENLABS_AGENT_ID` | Auto-created | - | Custom ElevenLabs agent ID | +| `GEMINI_API_KEY` / `GOOGLE_API_KEY` | - | Settings / env | Gemini Live voice assistant | +| `DASHSCOPE_API_KEY` / `QWEN_API_KEY` | - | Settings / env | Qwen Realtime voice assistant | +| `VOICE_BACKEND` | Auto-detected | - | Default assistant backend: `elevenlabs`, `gemini-live`, or `qwen-realtime` | | `OPENAI_API_KEY` | - | Settings / env | OpenAI API key for dictation (`gpt-transcribe` / `gpt-live-transcribe`) | | `DEEPGRAM_API_KEY` | - | Settings / env | Deepgram API key for dictation (`nova-3`) | | `GROQ_API_KEY` | - | Settings / env | Groq API key for dictation (`whisper-large-v3`) | diff --git a/docs/guide/notifications.md b/docs/guide/notifications.md index f8d3a131..b2cc09d0 100644 --- a/docs/guide/notifications.md +++ b/docs/guide/notifications.md @@ -1,8 +1,20 @@ # Notifications -Get notified when sessions need input, request permissions, fail, or complete — via Telegram, Server酱 (ServerChan), Web Push, or voice. +Get notified when sessions need input, request permissions, fail, or complete — via native app notifications, Telegram, Server酱 (ServerChan), Web Push, or voice. -Web Push works out of the box once you [install the PWA](./pwa.md); no configuration needed. The channels below are optional. +Web Push needs no provider configuration: [install the PWA](./pwa.md) and allow notifications. The channels below are optional. + +## Native app notifications + +For official Android and iOS apps, pair an updated hub and allow notifications. +No Firebase project or Apple developer account is needed. Notification content +is end-to-end encrypted through the official push relay, which also works when +you access the hub through Tailscale or your own HTTPS setup. Android requires +Google Play services and FCM connectivity. + +Private app builds can use their own matching Firebase/APNs credentials. +See the [native push contract](../api/native-companion-contract.md) for those +settings and delivery details. ## Telegram Setup @@ -19,7 +31,8 @@ export HAPI_PUBLIC_URL="https://your-public-url" hapi hub ``` -Then message your bot with `/start`, open the app, and enter your `CLI_API_TOKEN`. +Then message your bot with `/start`, open the app, and bind using +`CLI_API_TOKEN:` (for example, `your-token:default`). Related environment variables: diff --git a/docs/guide/quick-start.md b/docs/guide/quick-start.md index c5343d58..5708c02f 100644 --- a/docs/guide/quick-start.md +++ b/docs/guide/quick-start.md @@ -26,7 +26,9 @@ Details and local-only mode: [Hub setup](./installation.md#hub-setup) hapi ``` -This starts Claude Code wrapped with HAPI. The session appears in the web UI. +Choose an installed agent from the picker. Its session appears in the web UI. +To start one directly, use `hapi claude`, `hapi codex`, or another +[supported agent command](./agents.md). Scripts must specify the agent explicitly. ## Open the UI diff --git a/docs/guide/why-hapi.md b/docs/guide/why-hapi.md index 95863199..e785dac1 100644 --- a/docs/guide/why-hapi.md +++ b/docs/guide/why-hapi.md @@ -2,7 +2,7 @@ [Happy](https://github.com/slopus/happy) is an excellent project. So why build HAPI? -**The short answer**: Happy uses a centralized server that stores your encrypted data. HAPI is decentralized — each user runs their own hub, and the relay server only forwards encrypted traffic without storing anything. These different goals lead to fundamentally different architectures. +**The short answer**: Happy uses a centralized server that stores your encrypted data. HAPI is decentralized — each user runs their own hub, and the optional network relay forwards encrypted traffic rather than hosting your conversation history. These different goals lead to fundamentally different architectures. ## TL;DR @@ -10,8 +10,8 @@ |--------|-------|------| | **Architecture** | Centralized (cloud server stores encrypted data) | Decentralized (each user runs own hub) | | **Users** | Multi-user on shared server | Any number (each runs own hub) | -| **Data** | Encrypted on server (server cannot read) | Stays on your machine | -| **Encryption** | Application-layer E2EE (client encrypts before sending) | WireGuard + TLS via relay; or none needed if self-hosted | +| **Session history** | Encrypted on server (server cannot read) | Stored on your own hub | +| **Encryption** | Application-layer E2EE (client encrypts before sending) | WireGuard + TLS via relay; HTTPS for self-hosted remote access | | **Deployment** | Multiple services (PostgreSQL, Redis, app server) | Single binary | | **Complexity** | High (E2EE, key management, scaling) | Low (one command) | @@ -56,14 +56,14 @@ The server stores encrypted data — it never sees plaintext, but it does hold y Each user runs their own hub. HAPI offers two modes of remote access: -- **Self-hosted** (own server / Cloudflare Tunnel / Tailscale) — You control the full network path, no E2EE needed +- **Self-hosted** (own server / Cloudflare Tunnel / Tailscale) — You choose the host and HTTPS endpoint - **Public relay** (`hapi hub --relay`) — E2E encrypted via tunwg (WireGuard + TLS); the relay only forwards opaque packets - **Single embedded database** — SQLite, no external services - **One-command deployment** — Single binary, zero config #### Mode 1: Self-Hosted (own server or tunnel) -You control the entire path. No encryption beyond standard HTTPS is needed. +You operate the hub and choose how to expose it. Use HTTPS for remote access; a third-party proxy that terminates TLS is part of that trust boundary. ``` ┌────────────────────────────────────────────────────────────────────────┐ @@ -134,10 +134,10 @@ The relay server only forwards encrypted packets — it cannot read your data. | Aspect | Happy | HAPI | |--------|-------|------| -| **Where data lives** | Cloud server (encrypted blobs) | Your own machine | -| **Who stores it** | Central server holds encrypted data | Only your hub, locally | +| **Where session history lives** | Cloud server (encrypted blobs) | Your own hub | +| **Who stores it** | Central server holds encrypted data | Your hub; clients may cache data | | **Data at rest** | Encrypted (server cannot read) | Plaintext (protected by OS) | -| **Server's role** | Stores encrypted data + syncs devices | Relay only forwards (or no server at all if self-hosted) | +| **Server's role** | Stores encrypted data + syncs devices | Your hub stores history; optional relay forwards traffic | ### Deployment Model @@ -203,14 +203,14 @@ Goal: Multi-user cloud platform ``` Goal: Self-hosted tool — each user runs their own hub │ - ├──► Data never leaves your machine - │ └──► No application-layer E2EE needed + ├──► History stored on your own hub + │ └──► No central HAPI history store │ ├──► Each user has their own hub │ └──► No horizontal scaling needed; unlimited users in aggregate │ ├──► Self-hosted access (own server/tunnel) - │ └──► You control the full path — HTTPS sufficient + │ └──► Your hub behind your chosen HTTPS endpoint │ └──► Public relay access └──► WireGuard + TLS (tunwg) — relay forwards only @@ -223,7 +223,7 @@ Goal: Self-hosted tool — each user runs their own hub | Dimension | Happy | HAPI | |-----------|-------|------| | **Architecture** | Centralized cloud server | Decentralized (each user runs own hub) | -| **Server's role** | Stores encrypted data | Relay only forwards (or none if self-hosted) | +| **Server's role** | Stores encrypted data | Your hub stores history; optional relay forwards traffic | | **Data location** | Server (encrypted, zero-knowledge) | Local (plaintext, your machine) | | **Deployment** | Multiple services (PostgreSQL, Redis, Node.js) | Single binary (embedded SQLite) | | **Encryption** | Application-layer E2EE (client-side) | WireGuard + TLS (relay) or HTTPS (self-hosted) | @@ -236,6 +236,6 @@ The architectural differences stem from a centralized vs decentralized design: - **Happy**: Centralized cloud server that stores your encrypted data. The server never sees plaintext (zero-knowledge), but it does hold your data. This requires application-layer E2EE, key management, and distributed infrastructure (PostgreSQL, Redis, scaling). -- **HAPI**: Decentralized — each user runs their own hub. Your data stays on your machine. For remote access, you can self-host (own server or tunnel — no E2EE needed since you control the path) or use the public relay (WireGuard + TLS via tunwg — the relay only forwards encrypted packets it cannot read). This achieves one-command deployment with zero external dependencies. +- **HAPI**: Decentralized — you run the hub that stores your session history, on your workstation or another host you control. Remote access uses your own HTTPS endpoint or the built-in encrypted network relay. The CLI, hub, web app, and SQLite database ship in one binary. -The core tradeoff: Happy solves the "untrusted server" problem with sophisticated encryption. HAPI avoids the problem entirely by keeping your data on your own machine. +The core tradeoff: Happy encrypts data for storage on its server; HAPI puts the history store under your control. Coding agents and optional voice, title-generation, and notification features still use their configured providers. See the [Privacy Policy](../privacy.md) for those data flows and native push relay metadata. diff --git a/docs/privacy.md b/docs/privacy.md index 429b3bb2..c38bd768 100644 --- a/docs/privacy.md +++ b/docs/privacy.md @@ -5,7 +5,7 @@ aside: false # Privacy Policy -**Effective date: September 9, 2026** · Applies to the HAPI mobile companion apps (Android and iOS) and the self-hosted HAPI hub. +**Effective date: September 12, 2026** · Applies to the HAPI mobile companion apps (Android and iOS) and the self-hosted HAPI hub. ::: tip The short version HAPI is self-hosted software. Your app connects to a hub **you** operate; the HAPI project does not run a central account or application backend that receives your conversations or source code. Optional features can send data to services you or the app enable — notably the HAPI push relay, Firebase Cloud Messaging, Apple Push Notification service, voice-transcription providers, and the coding-agent/model providers configured on your machine. The HAPI push relay processes device and connection metadata to deliver encrypted notifications, prevent abuse, and diagnose delivery failures. HAPI contains no advertising or tracking SDKs and no product analytics. @@ -30,21 +30,21 @@ Your hub, coding agents, plugins, and command-line tools may send prompts, sourc ## Push notifications -**Android:** notifications use Google Firebase Cloud Messaging (FCM). When push is enabled, Firebase creates and manages a device token; the app registers that token and a random app device identifier with each paired hub. Your hub sends notification content and routing metadata — for example a session identifier, title, status, and action type — through FCM. HAPI does not end-to-end encrypt the Android notification payload, so Google processes this data under [Firebase's privacy terms](https://firebase.google.com/support/privacy). Builds without Firebase configuration do not register for or receive FCM push. +**Android:** notifications use Google Firebase Cloud Messaging (FCM). When push is enabled, the app registers its FCM token, random app device identifier, and device-generated encryption key with each paired hub. Official app delivery through the HAPI push relay is end-to-end encrypted: the relay and Google carry ciphertext and routing metadata, but cannot read notification content. Private builds using a hub configured for direct FCM retain an unencrypted notification payload, including session identifiers, title, status, and action type; Google processes that data under [Firebase's privacy terms](https://firebase.google.com/support/privacy). Builds without Firebase configuration do not register for or receive FCM push. **iOS:** notifications are end-to-end encrypted. Your hub encrypts the content with a key that exists only on your device and your hub; Apple's push service and the HAPI push relay carry ciphertext and routing metadata only, and cannot read the notification content. Self-hosters using a separately signed app build with matching APNs credentials can bypass the relay; a different developer account's credentials alone cannot send notifications to the official App Store build. -### HAPI iOS push relay +### HAPI native push relay -The default hub configuration uses the official relay at `https://push.hapi.run` for iOS push delivery. Encryption protects notification content, **not all metadata**: +The default hub configuration uses the official relay at `https://push.hapi.run` for iOS push and for Android push when no private Firebase credentials are configured. This service is separate from the optional network tunnel enabled by `hapi hub --relay`. Encryption protects notification content, **not all metadata**: -- **Data received:** an APNs device token, the encrypted notification envelope, and optional delivery settings such as priority and a notification-grouping identifier. The relay can observe request timing, envelope size, and the source IP of the connecting hub or proxy; that IP is not necessarily the phone's IP. -- **Uses:** forwarding notifications to Apple, limiting abusive requests, and diagnosing delivery failures. This information is not used for advertising, cross-app tracking, marketing, or product analytics. +- **Data received:** the target platform, an APNs or FCM device token, the encrypted notification envelope, and optional delivery settings such as priority and a notification-grouping identifier. The relay can observe request timing, envelope size, and the source IP of the connecting hub or proxy; that IP is not necessarily the phone's IP. +- **Uses:** forwarding notifications to Apple or Google, limiting abusive requests, and diagnosing delivery failures. This information is not used for advertising, cross-app tracking, marketing, or product analytics. - **Rate-limit state:** device tokens and source IPs are kept in bounded process-memory maps with counters and refill times. There is no fixed expiration timer: entries may remain until capacity-based eviction or process restart. These maps are not written to a database by the relay. -- **Operational logs:** the relay writes a stable, truncated SHA-256 hash of the device token and a delivery, error, or rate-limit outcome to its log output. Hashing avoids logging the usable token, but still allows events for a device to be correlated; it is not a claim of complete anonymity. The relay does not deliberately log notification envelopes, decrypted content, or raw device tokens. +- **Operational logs:** the relay writes a stable, platform-scoped, truncated SHA-256 hash of the device token and a delivery, error, or rate-limit outcome to its log output. Hashing avoids logging the usable token, but still allows events for a device to be correlated; it is not a claim of complete anonymity. The relay does not deliberately log notification envelopes, decrypted content, or raw device tokens. - **Hosting logs:** container logging and any reverse proxy may retain operational or connection logs according to their deployment configuration. The relay source code does not impose a retention period or automatically delete those external logs. Its lack of a database does not mean that the hosting environment retains no data. -Only the device and its paired hub hold the notification decryption key. Your hub stores its device registration, including the random app device identifier, APNs token, and encryption key; the random app device identifier and encryption key are not included in the relay push request. Apple receives the device token, encrypted envelope, and delivery settings needed to route the notification. +Only the device and its paired hubs hold the notification decryption key. Each hub stores its device registration, including the random app device identifier, APNs or FCM token, and encryption key; the random app device identifier and encryption key are not included in the relay push request. Apple or Google receives the device token, encrypted envelope, and delivery settings needed to route the notification. ## Camera @@ -56,7 +56,7 @@ Microphone access is used only when you start voice dictation. The app records a ## What the HAPI project collects -The HAPI project does not receive product-analytics events, advertising identifiers, conversations, source code, or hub access credentials through a central HAPI backend. The official iOS push relay does process the device and operational metadata described above; APNs device tokens are different from advertising identifiers. The apps contain no advertising, tracking, product-analytics, or third-party crash-reporting SDKs. If you contact us by email or GitHub, we receive the information you voluntarily include in that communication. +The HAPI project does not receive product-analytics events, advertising identifiers, conversations, source code, or hub access credentials through a central HAPI backend. The official native push relay does process the device and operational metadata described above; push device tokens are different from advertising identifiers. The apps contain no advertising, tracking, product-analytics, or third-party crash-reporting SDKs. If you contact us by email or GitHub, we receive the information you voluntarily include in that communication. Google Play, the Apple App Store, operating-system vendors, Firebase, and other services you enable may collect installation, device, diagnostic, notification, or service-usage data independently under their own policies. @@ -66,7 +66,7 @@ Unpairing a hub removes that hub's credentials from the app and attempts to unre Data stored on a hub remains under the hub operator's control and retention settings. Delete it from the hub, its underlying storage, and any configured provider as appropriate. HAPI has no central user account to delete. -For iOS relay data, stopping delivery does not erase earlier operational logs. Hub operators can disable iOS push with `HAPI_IOS_PUSH=off`; unpairing also attempts to remove that hub's device registration. Rate-limit entries are removed through the eviction/restart behavior described above, not by an app-side account-deletion action. External log deletion and rotation are controlled separately by the hosting operator. +For native relay data, stopping delivery does not erase earlier operational logs. Hub operators can disable push with `HAPI_IOS_PUSH=off` or `HAPI_ANDROID_PUSH=off` for the corresponding platform; unpairing also attempts to remove that hub's device registration. Rate-limit entries are removed through the eviction/restart behavior described above, not by an app-side account-deletion action. External log deletion and rotation are controlled separately by the hosting operator. To ask about the official relay's deployed log retention or request deletion of relay-related personal data, email [twsxtd@gmail.com](mailto:twsxtd@gmail.com). We will determine what records we can identify and the applicable deletion or retention requirements. We cannot identify a device from an email address alone and do not promise that all records can be located or that data on independently operated hubs or providers can be deleted by HAPI. Do not post device tokens, hub access credentials, or encryption keys in public issues or include them in an initial email; contact us privately to establish a safe way to handle your request. diff --git a/hub/README.md b/hub/README.md index aee3c8d7..ff694a4d 100644 --- a/hub/README.md +++ b/hub/README.md @@ -8,8 +8,8 @@ Telegram bot + HTTP API + realtime updates for hapi hub. - HTTP API for sessions, messages, permissions, machines, and files. - Server-Sent Events stream for live updates in the web app. - Socket.IO channel for CLI connections. -- Serves the web app from `web/dist` or embedded assets in the single binary. -- Persists state in SQLite. +- Serves the web app from `web/dist` or embedded assets in the single binary; network-relay mode uses the separately hosted official web app. +- Persists state in SQLite via `bun:sqlite`. ## Configuration @@ -32,6 +32,7 @@ Dictation and voice-assistant provider keys can also be added from **Settings - `ELEVENLABS_AGENT_ID` - Custom ElevenLabs agent ID (auto-created if not set). - `GEMINI_API_KEY` / `GOOGLE_API_KEY` - Gemini Live voice assistant. - `DASHSCOPE_API_KEY` / `QWEN_API_KEY` - Qwen Realtime voice assistant. +- `VOICE_BACKEND` - Default assistant backend (`elevenlabs`, `gemini-live`, or `qwen-realtime`). - `OPENAI_API_KEY` - OpenAI dictation (`gpt-transcribe` / `gpt-live-transcribe`). - `DEEPGRAM_API_KEY` - Deepgram dictation (`nova-3`, standard and realtime). - `GROQ_API_KEY` - Groq dictation (`whisper-large-v3`). @@ -86,9 +87,10 @@ bun run dev:hub ## HTTP API -See `src/web/routes/` for all endpoints. +The following is an overview. See the [client contract](../docs/api/client-contract/index.md) +for request/response shapes and error semantics, and `src/web/routes/` for all endpoints. -### Authentication (`src/web/routes/auth.ts`) +### Authentication (`src/web/routes/auth.ts`, `src/web/routes/bind.ts`) - `POST /api/auth` - Get JWT token (Telegram initData or `CLI_API_TOKEN[:namespace]`). - `POST /api/bind` - Bind a Telegram account using initData + `CLI_API_TOKEN:`. @@ -98,8 +100,10 @@ See `src/web/routes/` for all endpoints. - `GET /api/sessions` - List all sessions. Each summary includes `hasConversationContent`, derived from stored conversation messages (not titles or lifecycle events); full session SSE updates carry changes to this flag. - `GET /api/sessions/:id` - Get session details. - `POST /api/sessions/:id/abort` - Abort session. -- `POST /api/sessions/:id/switch` - Switch session to remote mode. +- `POST /api/sessions/:id/switch` - Hand off session control to the web. - `POST /api/sessions/:id/resume` - Resume inactive session. +- `POST /api/sessions/:id/reopen` - Reopen a session; follow the returned session ID. +- `POST /api/sessions/:id/clear` - Start a fresh conversation when supported. - `POST /api/sessions/:id/upload` - Upload file (base64, max 50MB). - `POST /api/sessions/:id/upload/delete` - Delete uploaded file. - `POST /api/sessions/:id/archive` - Archive active session. @@ -109,12 +113,17 @@ See `src/web/routes/` for all endpoints. - `GET /api/sessions/:id/skills` - List skills. - `POST /api/sessions/:id/permission-mode` - Set permission mode. - `POST /api/sessions/:id/model` - Set model preference. -- `POST /api/sessions/:id/effort` - Set Claude effort preference. +- `POST /api/sessions/:id/effort` - Set effort for Claude, Grok, or Pi; other config controls are flavor/capability-gated (see the client contract). +- `GET/POST /api/sessions/:id/scratchlist` - Read/create Hub-persisted scratchlist entries; update/delete and attachment routes share this prefix. ### Messages (`src/web/routes/messages.ts`) - `GET /api/sessions/:id/messages` - Get messages (paginated). - `POST /api/sessions/:id/messages` - Send message. +- `POST /api/sessions/:id/messages/queued-state` - Reconcile queued, indeterminate, and invoked local IDs. +- `DELETE /api/sessions/:id/messages/:messageId` - Cancel a queued message. +- `POST /api/sessions/:id/messages/:messageId/steer` - Steer a queued message when the session supports it. +- `POST /api/sessions/:id/messages/:messageId/retry` - Explicitly retry indeterminate delivery when safe; never auto-replay. ### Permissions (`src/web/routes/permissions.ts`) @@ -124,15 +133,21 @@ See `src/web/routes/` for all endpoints. ### Machines (`src/web/routes/machines.ts`) - `GET /api/machines` - List online machines. +- `PATCH /api/machines/:id` - Set/clear the machine display name. - `GET /api/machines/:id/agent-availability` - List installed/configured Agents. - `POST /api/machines/:id/spawn` - Spawn new session on machine. - `POST /api/machines/:id/list-directory` - Browse runner-scoped directories. - `POST /api/machines/:id/paths/exists` - Check if path exists. +- `POST /api/machines/:id/restart-runner` - Request runner restart (requires an available restart path). ### Usage (`src/web/routes/usage.ts`) - `GET /api/usage/summary` - Get cache-aware token usage for the owner namespace (`range=7d|30d|all`). +### Storage (`src/web/routes/storage.ts`) + +- `GET /api/storage/sqlite` - SQLite database/WAL/SHM sizes for the owner namespace. + ### Git/Files (`src/web/routes/git.ts`) - `GET /api/sessions/:id/git-status` - Git status. @@ -149,6 +164,9 @@ See `src/web/routes/` for all endpoints. ### Voice (`src/web/routes/voice.ts`) - `POST /api/voice/token` - Get ElevenLabs conversation token. +- `GET /api/voice/backend` - Discover configured assistant backends. +- `GET /api/voice/voices` - List voices for the selected backend. +- `GET/PUT /api/voice/transcription/credentials` - Read masked provider settings or update credentials (owner-only). - `GET /api/voice/transcription/providers` - List configured providers and supported modes. - `POST /api/voice/transcription` - Transcribe a bounded recording. - `POST /api/voice/transcription/realtime-token` - Mint a short-lived OpenAI, ElevenLabs, or Deepgram credential. @@ -159,6 +177,9 @@ See `src/web/routes/` for all endpoints. - `POST /api/push/subscribe` - Subscribe to push notifications. - `DELETE /api/push/subscribe` - Unsubscribe. +Native Android/iOS registration uses `POST`/`DELETE /api/devices/register` +(`src/web/routes/devices.ts`); see the [native push contract](../docs/api/native-companion-contract.md). + ### CLI (`src/web/routes/cli.ts`) - `POST /cli/sessions` - Create/load session. @@ -168,9 +189,10 @@ See `src/web/routes/` for all endpoints. ## Socket.IO -See `src/socket/handlers/cli.ts` for event handlers. +See `src/socket/handlers/cli/index.ts` and `src/socket/handlers/terminal.ts` for event handlers. -Namespace: `/cli` +Namespaces: `/cli` (raw CLI access token) and `/terminal` (client JWT). +Ordinary web/native session updates use SSE, not the CLI Socket.IO namespace. ### Client events (CLI to hub) @@ -184,14 +206,16 @@ Namespace: `/cli` - `rpc-register` - Register RPC handler. - `rpc-unregister` - Unregister RPC handler. -### Terminal events (web to hub) +### Terminal events (web to hub, `/terminal`) - `terminal:create` - Open terminal for session. - `terminal:write` - Send input. - `terminal:resize` - Resize dimensions. - `terminal:close` - Close terminal. +- `agent-terminal:subscribe` / `agent-terminal:unsubscribe` - Attach/detach the wrapped agent's terminal stream. +- `agent-terminal:input` / `agent-terminal:resize` - Interact with that terminal when supported. -### Hub events (hub to clients) +### Hub events (hub to CLI clients, `/cli`) - `update` - Broadcast session/message updates. - `rpc-request` - Incoming RPC call. @@ -236,6 +260,7 @@ See `src/store/index.ts` for SQLite persistence: - Machines with runner state. - Todo extraction from messages. - Users table for Telegram bindings (includes namespace). +- Scratchlist entries/attachments, usage, work graph, and push registrations. Message content is stored via `src/store/contentCodec.ts`: oversized strings inside agent messages (giant tool output) are head+tail truncated at ingest, diff --git a/ios/README.md b/ios/README.md index 6b5cc67c..eb80d664 100644 --- a/ios/README.md +++ b/ios/README.md @@ -27,12 +27,13 @@ the command line: xcodebuild build -project ios/Hapi.xcodeproj -scheme Hapi \ -destination 'generic/platform=iOS Simulator' CODE_SIGNING_ALLOWED=NO -# Package tests (also runs on macOS, the package is pure Foundation) +# Package tests (macOS; protocol/client plus UI package tests) swift test --package-path ios/Packages/HapiKit ``` -CI runs both on `macos-15` via `.github/workflows/ios.yml` (triggered by -changes under `ios/**` and `shared/fixtures/**`). +CI runs package tests, the simulator build, and app-hosted transcript tests +on `macos-15` via `.github/workflows/ios.yml` (triggered by changes under +`ios/**` and `shared/fixtures/**`). ### Localization catalog @@ -238,7 +239,7 @@ ios/ Hapi.xcodeproj/ Hand-rolled minimal project (objectVersion 77). Hapi/ App target sources. This is an Xcode 16 "synchronized folder": add files here and they join the target - without touching project.pbxproj. As of M2a: + without touching project.pbxproj. Main areas: Models/ AppModel (pairing state machine, hub switching, deep-link routing, scene phase) + HubSession (per-active-hub @@ -269,7 +270,7 @@ ios/ pinned section, applied-filter summary, pull-to-refresh, long-press pin/archive; row taps push the chat), - Chat/ (M2f read-only chat: ChatModel — + Chat/ (interactive chat: ChatModel — window state + session detail → ChatPipeline off-main, ~100 ms coalesced, last-seen stamping, header @@ -375,7 +376,7 @@ ios/ preferredColorScheme; system follows the OS, explicit modes override, OLED = dark on pure black); Language — - persist-only until the M5 i18n pass; + system/English/简体中文; applies on relaunch; owner-gated (JWT ns == "default", fail closed) Usage dashboard — range 7d/30d/all, stat tiles, Swift Charts @@ -862,9 +863,9 @@ for push entitlements: actions — can be exercised with `xcrun simctl push` using a payload whose `hapi.e` was produced with the device's registered key. -## Milestones (track A of the native-clients plan) +## Milestone history (track A of the native-clients plan) -- **M0** — this scaffold: project, HapiKit package, CI, one passing test. +- **M0** — initial scaffold: project, HapiKit package, CI, one passing test. - **M1** — foundations: HapiProtocol wire models + catalogs; APIClient + auth (Keychain, single-flight 401 refresh); SSEClient + reconnect state machine + versioned patch application (incl. gzip streaming check); pairing flow diff --git a/web/README.md b/web/README.md index ea8dc1b0..25444dd6 100644 --- a/web/README.md +++ b/web/README.md @@ -32,6 +32,7 @@ See `src/router.tsx` for route definitions. - `/sessions/$sessionId/files` - File browser with git status. - `/sessions/$sessionId/file` - File viewer with diff support. - `/sessions/$sessionId/terminal` - Terminal interface. +- `/browse` - Workspace browser, enabled by the runner's configured workspace roots. - `/share` - Share-target landing (Web Share Target POST → `?id=`, or native `/share#url=&text=&title=`). - `/settings` - Settings category hub (mobile) and responsive master-detail shell. - `/settings/general` - Language preferences. @@ -40,6 +41,8 @@ See `src/router.tsx` for route definitions. - `/settings/voice` - Everyday voice assistant preferences. - `/settings/voice/voices` - Full-page voice picker. - `/settings/voice/advanced` - Voice persona, tuning, and diagnostics. +- `/settings/machines` - Machine management and runner status. +- `/settings/storage` - SQLite storage sizes for the hub owner. - `/settings/usage` - Cache-aware token usage dashboard for the hub owner. - `/settings/about` - Application links and version information. @@ -51,21 +54,20 @@ See `src/router.tsx` for route definitions. - Session title from name, summary, or path. - Todo progress display. - Pending permission request count. -- Agent flavor label (claude/codex/gemini). -- Model mode display. +- Agent name and model display. ### Chat interface (`src/components/SessionChat.tsx`) - Message thread with infinite scroll. - Composer for sending messages. -- Permission mode toggle (default/acceptEdits/auto/bypassPermissions/plan). -- Model selection (default/sonnet/sonnet[1m]/opus/opus[1m]). -- Session abort and mode switch controls. +- Permission mode and model selection for supported agents. +- Session abort and handoff controls. - Context size display. - Per-session scratchlist (`src/components/AssistantChat/ScratchlistPanel.tsx`) - Workbench panel for held notes/drafts; **distinct from the queue**. - Add/delete/reorder entries; promote to composer (copy) or queue (send). - - Persists across reloads via `localStorage` keyed per session. + - Entries and attachments saved on the hub and synced across devices. + - Reordering affects only the current view and resets when entries refresh. - Keyboard shortcut: Ctrl/Cmd+Shift+S to focus the add-input. ### File browser (`src/routes/sessions/files.tsx`) @@ -82,12 +84,12 @@ See `src/router.tsx` for route definitions. ### Terminal (`src/routes/sessions/terminal.tsx`) - Remote terminal via xterm.js -- Real-time via Socket.IO +- Real-time via Socket.IO `/terminal` - Resize handling ### Voice assistant -- ElevenLabs integration (@elevenlabs/react) +- ElevenLabs (@elevenlabs/react), Gemini Live, and Qwen Realtime backends - Real-time voice control - Standard and realtime composer dictation with provider capability selection @@ -99,7 +101,7 @@ Modular session creation: - Directory input with recent paths - Agent type selector - Model selector -- Permission mode toggle (YOLO mode) +- Per-agent permission, effort, and collaboration controls when supported ## Authentication