Files
hapi/docs/guide/how-it-works.md
T
c0b30bf916 feat(cli): MCP list_peers + runner hub auth for peer discovery (#1372)
* feat(cli): MCP list_peers + runner hub auth inheritance

Runner-spawned agents could not discover same-hub peers without
sitting on the hub host or pasting a session id. Add MCP list_peers
(in-process credentials), export HAPI_API_URL/CLI_API_TOKEN after
auth init for shell fallbacks, and clearer auth failure hints.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): do not export default hub URL into HAPI_API_URL

exportHapiHubAuthEnv was writing the implicit localhost default into
process.env, which made maybeAutoStartServer skip starting the bundled
hub. Only export HAPI_API_URL when the URL came from env or settings;
always still export CLI_API_TOKEN. Also fill missing deliveryMode on
abort restore so web typecheck matches RawSendError (main tip unblock).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): widen initializeApiUrl mock return type in test

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): never export CLI_API_TOKEN; exclude self from list_peers

Keep settings/prompt-backed hub secrets out of wrapped agent env so
shell JWT+curl cannot bypass peer-tool approval. Fresh hapi re-reads
settings; env-backed tokens already inherit. list_peers omits the
calling session from the shortlist.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): resolve peer labels via summary/path like web titles

list_peers was showing (unnamed) for ordinary sessions because titles
live in metadata.summary.text. Match web getSessionTitle and collapse
whitespace so each peer stays one agent-readable line.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub): emit full peer ids and honor GET /sessions?limit

Short 8-char prefixes collide across UUID namespaces; print full ids so
resolveSessionByPrefix stays unambiguous. Honor optional limit after sort
so listPeerSessions stops loading the whole namespace for scheduled counts.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): type sessions limit test mock as Map<string, number>

CI tsc rejected Map<string, null> for getNextScheduledAtBySessionIds.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub): unbounded ping resolve; peer list order=updatedAt

Keep GET /sessions?limit only for discovery callers. ping/inspect omit
limit so full UUIDs outside the first 500 stay resolvable. Peer lists
pass order=updatedAt so truncation matches newest-first. Basename
fallback splits Windows paths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): auto-approve ACP title List Peer Sessions

Permission derivation prefers request.title; match the MCP tool title
form so default-mode ACP sessions do not prompt on discovery.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): pad list_peers fetch; split hub URL vs token hints

Fetch limit+2 when excluding the caller so overflow still surfaces at
limit=100. Clarify that auth login only saves the token, not HAPI_API_URL.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): use boolean overflow for ping-peer --list

Match MCP list_peers: fetch limit+1 and mark hasMore instead of claiming
an exact omitted count from a 200-row sample.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): tolerate mocked machineCache without expireInactive

CI flake: 5s inactivity tick hit test doubles that only stubbed
getOnlineMachinesByNamespace. Optional-call + stub the method.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-05 22:16:04 +08:00

228 lines
11 KiB
Markdown

# How it Works
HAPI consists of three interconnected components that work together to provide remote AI agent control.
## Architecture Overview
```
┌────────────────────────────────────────────────────────────────────────────┐
│ Your Machine (Local or Hub Host) │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ │ │ │ │ │ │
│ │ HAPI CLI │◄───────►│ HAPI Hub │◄───────►│ Web App │ │
│ │ │ Socket │ │ SSE │ (embedded) │ │
│ │ + AI Agent │ .IO │ + SQLite │ │ │ │
│ │ │ │ + REST API │ │ │ │
│ └──────────────┘ └──────┬───────┘ └──────────────┘ │
│ │ │
│ │ localhost:3006 │
└───────────────────────────────────┼────────────────────────────────────────┘
│
┌─────────▼─────────┐
│ Tunnel (Optional)│
│ Cloudflare/ngrok │
└─────────┬─────────┘
│
┌───────────────────────────────────┼────────────────────────────────────────┐
│ Public Internet │
│ │ │
│ ┌─────────────────────────┼─────────────────────────┐ │
│ │ ▼ │ │
│ │ ┌──────────────┐ ┌──────────────┐ │ │
│ │ │ │ │ │ │ │
│ │ │ Telegram │ │ PWA / │ │ │
│ │ │ Mini App │ │ Browser │ │ │
│ │ │ │ │ │ │ │
│ │ └──────────────┘ └──────────────┘ │ │
│ │ │ │
│ └───────────────────────────────────────────────────┘ │
│ Your Phone │
└────────────────────────────────────────────────────────────────────────────┘
```
> **Note:** The hub can run on your local desktop or a remote host (VPS, cloud, etc.). If deployed on a host with a public IP, tunneling is not required.
## Components
### HAPI CLI
The CLI is a wrapper around AI coding agents. It supports multiple agent flavors out of the box — see [Supported agents](./agents.md) for the full list. It:
- Starts and manages coding sessions
- Registers sessions with the HAPI hub
- Relays messages and permission requests
- Provides MCP (Model Context Protocol) tools
**Key Commands:**
```bash
hapi # Start a session (Claude Code by default)
hapi <agent> # Start a session with another agent flavor (see Supported agents)
hapi runner start # Run background service for remote session spawning
hapi ping-peer --list # Shell peer shortlist (prefer MCP list_peers in-session)
```
MCP peer tools (same hub/namespace as the session): `list_peers` (discover), `inspect_peer` (read), `ping_peer` (message). These work from runner-spawned sessions even when the hub is on another host - see [Installation → Split hub + remote runner](./installation.md#split-hub--remote-runner-peer-discovery).
### HAPI Hub
The hub is the central service that connects everything:
- **HTTP API** - RESTful endpoints for sessions, messages, permissions
- **Socket.IO** - Real-time bidirectional communication with CLI
- **SSE (Server-Sent Events)** - Live updates pushed to web clients
- **SQLite Database** - Persistent storage for sessions and messages
- **Telegram Bot** - Notifications and Mini App integration
### Web App
A React-based PWA that provides the mobile interface:
- **Session List** - View all active and past sessions
- **Chat Interface** - Send messages and view agent responses
- **Permission Management** - Approve or deny tool access
- **File Browser** - Browse project files and view git diffs
- **Terminal View** - Watch the full terminal output of a session
- **Voice Assistant** - Talk to your agent and approve permissions by voice (see [Voice input and assistant](./voice-assistant.md))
- **Session Sharing** - Share a read-only view of a session via a link
- **Remote Spawn** - Start new sessions on any connected machine
## Data Flow
### Starting a Session
```
1. User runs `hapi` in terminal
│
▼
2. CLI starts Claude Code (or other agent)
│
▼
3. CLI connects to hub via Socket.IO
│
▼
4. Hub creates session in database
│
▼
5. Web clients receive SSE update
│
▼
6. Session appears in mobile app
```
### Permission Request Flow
```
1. AI agent requests tool permission (e.g., file edit)
│
▼
2. CLI sends permission request to hub
│
▼
3. Hub stores request and notifies via SSE + Telegram
│
▼
4. User receives notification on phone
│
▼
5. User approves/denies in web app or Telegram
│
▼
6. Hub relays decision to CLI via Socket.IO
│
▼
7. CLI informs AI agent, execution continues
```
### Message Flow
```
User (Phone) Hub CLI
│ │ │
│──── Send message ──────►│ │
│ │─── Socket.IO emit ───►│
│ │ │
│ │ ├── AI processes
│ │ │
│ │◄── Stream response ───│
│◄─────── SSE ────────────│ │
│ │ │
```
## Communication Protocols
### CLI ↔ Hub: Socket.IO
Real-time bidirectional communication for:
- Session registration and heartbeat
- Message relay (user input → agent)
- Permission requests and responses
- Metadata and state updates
- RPC method invocation
### Hub ↔ Web: REST + SSE
- **REST API** for actions (send message, approve permission)
- **SSE stream** for real-time updates (new messages, status changes)
### External Access: Tunnel
For remote access outside your local network:
- **Built-in relay** (`hapi hub --relay`) - Managed tunwg tunnel (WireGuard + TLS), no third-party account required
- **Cloudflare Tunnel** (recommended) - Free, secure, reliable
- **Tailscale** - Mesh VPN for private networks
- **ngrok** - Quick setup for testing
## Seamless Handoff
HAPI's defining feature is the ability to seamlessly hand off control between local terminal and remote devices without losing session state.
### Local Mode
When working in local mode, you have the full terminal experience — it is the native agent CLI (Claude Code, Codex, OpenCode, and more):
- Direct keyboard input with instant response
- Full terminal UI with syntax highlighting
- Best for focused, uninterrupted coding sessions
- All AI processing happens locally on your machine
### Remote Mode
Switch to remote mode when you need to step away:
- Control via Web/PWA/Telegram from any device
- Approve permissions on the go
- Monitor progress while away from your desk
- Session continues running on your local machine
### How Switching Works
```
┌─────────────────┐ ┌─────────────────┐
│ Local Mode │◄──────────────────►│ Remote Mode │
│ (Terminal) │ │ (Phone/Web) │
└─────────────────┘ └─────────────────┘
│ │
│ ┌────────────────────────────┐ │
└─►│ Same Session, Same State │◄─────┘
└────────────────────────────┘
```
**Local → Remote:**
- Receive a message from phone/web
- Session automatically switches to remote mode
- Terminal shows "Remote mode - waiting for input"
**Remote → Local:**
- Press double-space in terminal
- Instantly regain local control
- Continue typing as if you never left
### Use Cases
1. **Remote Control While Away** - Start a session at your desk, continue from your phone during commute or coffee break
2. **Permission Approval** - AI requests file access, you get notified on phone, approve with one tap, session continues
3. **Multi-Device Collaboration** - View session progress on your phone while your desktop does the heavy lifting