* fix(test): stop runner integration suite from leaking detached process trees (#1515) The default CLI test run included runner.integration.test.ts, which spawns real detached runner/session process trees. A failing, timed-out, or interrupted test (or a plain runner stop) left those trees alive under PID 1 — on the Mac this accumulated ~600 Node/Bun/agent processes and several GiB of RSS over repeated runs. Test harness changes only; production runner session-preservation semantics are untouched: - Exclude runner.integration.test.ts from the default parallel unit-test suite; move it into a dedicated serial integration project (vitest.integration.config.ts, 'bun run test:integration'). The 20-session stress test is opt-in via HAPI_RUN_STRESS_TESTS=true. - Add a test-owned process/session registry (processRegistry.ts): every runner, runner-spawned session, and terminal-style child is registered immediately after spawn; afterEach/afterAll run two-stage cleanup (logical stopRunnerSession first, then bounded process-tree kill), followed by a marker sweep for agent grandchildren reparented to PID 1. - Add a per-run HAPI_TEST_MARKER env stamp + identity/secret env neutralization for test children (integrationEnv.ts) so outer HAPI/pi session variables never leak into test processes and the final audit can recognize test-owned processes by env alone. - Final suite audit in globalSetup teardown: reap anything still carrying the run marker and fail with PID/command diagnostics if anything cannot be reaped, before removing the temp home. - Regression coverage: a deliberately failing test registers a detached child and the follow-up audit must find zero test-owned processes. - CI: replace the dead .env.integration-test step with a dedicated integration job running the serial project. * refactor(test): drop unused killByChildProcess import and child field from registry * chore(test): raise integration hookTimeout to 60s for slow teardown hosts * fix(test): fail loudly when the process-table audit cannot scan; assert regression child death Bot review #1521 findings: - A failed `ps` scan (unsupported flags, buffer exhaustion, permissions) previously returned [] and silently disabled both teardown audit layers. It now throws; globalSetup teardown catches the scan error into the audit error (temp home is still removed) so the run fails visibly. - The regression audit test cleaned the leak with the reaper before asserting, and force-killed the fresh marked runner. The failing test's direct child PID is now asserted dead in afterEach right after registry cleanup (before the marker sweep), and the audit test stops its own runner gracefully before reaping. * fix(test): bound the logical cleanup phase so a hung runner cannot stall the hook Bot review #1521: stopRunnerSession carries the worker's 60s HTTP timeout (setup.ts raises HAPI_RUNNER_HTTP_TIMEOUT for the stress test), and the integration hook timeout is also 60s — N sequential stops could exhaust the hook budget before the process-tree fallback and marker sweep ran, recreating the very leak this change prevents. Logical shutdown is now parallel (Promise.allSettled over all tracked sessions) and the whole phase (stops + PID resolution) races against a 15s budget, so stage-2 tree-kill and the marker sweep always get their share of the hook window. * fix(test): bound graceful runner stop in hooks; keep credentials out of audit diagnostics Bot review #1521 (follow-up): - stopRunner()'s HTTP stop can burn the worker-wide 60s timeout on a hung-but-live runner, starving the marker sweep within the hook budget. afterEach/afterAll now race the graceful stop against a 10s bound; a runner that does not stop in time is force-reaped by the sweep (it carries the run marker) and the next beforeEach's alive-PID guard ignores any stale state file. - The env-bearing ps scan (ps eww) was also used for diagnostics, so the first 500 chars of a short-command process could print inherited credentials (CLI_API_TOKEN etc.) into teardown error logs. The scan now only identifies marked PIDs; command lines are fetched separately without 'e', falling back to '(command unavailable)' instead of the env dump. * fix(test): reap runner model-probe orphans before the zero-survivor inspection Bot review #1521 (Minor): inspect-before-reap. Applying it exposed a real race: each test's runner legitimately spawns marker-carrying children at startup (agent acp + agent --list-models model-catalog probes). Stopping the runner orphans them (ppid 1) with the run marker, so the audit test's OWN runner polluted the pure inspection with fresh probes spawned after the failing test's sweep window. - reapTestOwnedProcesses now re-kills every re-scan iteration instead of killing once and only re-scanning, so a process that survived its first SIGKILL (mid-exec) or spawned mid-kill is not given a free pass. - The regression audit test stops its runner, reaps (clearing its own legitimate orphan probes), then inspects: anything still marked is a genuine survivor the bounded reaper could not remove and fails the suite. Killable leaks from the failing test are already asserted dead in afterEach before the sweep runs. * fix(test): strictly bound the marker reaper; make per-test sweep unconditional and verified Bot review #1521 (follow-up): - The 10s reap deadline did not bound the awaited per-tree kills: each killProcessTreeByPid can wait up to 2s per PID, so several stuck processes could still exceed the 60s hook budget. Every process in a test-owned tree carries the marker (env is inherited), so tree-walking is unnecessary: the reaper now SIGKILLs every marked PID found by each scan, fire-and-forget, and re-scans every 250ms — the deadline strictly bounds the function. - The per-test sweep was skipped when the direct-child assertion failed first, and its survivors were ignored. afterEach now snapshots the regression-child state BEFORE the unconditional sweep, then verifies both the registry result and the sweep leftovers. * fix(test): replace it.fails regression with a direct assertion test Bot review #1521 (Minor): Vitest applies the it.fails expected-failure inversion after afterEach, so a broken registry assertion inside the hook would be masked as an expected failure, and the marker sweep would erase the evidence before the follow-up audit ran. The regression is now a normal test that registers a detached child at spawn time, deliberately performs NO per-test teardown, runs only the spawn-time registered cleanup, and asserts the child PID is dead. The afterEach no longer carries the registry-leak assertion (moved into the test body where it cannot be inverted); the per-test sweep assertion and the final audit test are unchanged. * fix(test): bound registry stage-2 tree-kills; require live regression fixture Bot review #1521 (follow-up): - Stage-2 killProcessTreeByPid awaits per descendant serially and can consume the whole 60s hook for a large/stuck tree. Signals are all delivered synchronously (children first) before any waiting, so racing the awaits against a 5s budget bounds the phase without skipping any kill; waitForAllDead still verifies the outcome. - The regression test could pass vacuously if its fixture exited during the startup delay (the registry exit listener would remove it before cleanup). It now asserts the child is alive before running cleanup. * fix(test): kill registered roots with bare synchronous SIGKILL, no pgrep walk Bot review #1521 (follow-up): racing the mapped killProcessTreeByPid calls against a timer does not bound the phase — evaluating the map invokes each call immediately, and each runs the recursive synchronous pgrep walk before its first await, which can consume the hook before the timer, runner stop, or marker sweep run. Stage-2 now SIGKILLs registered roots directly (fire-and-forget, no tree walk, no per-PID waits) and waits a bounded 5s for death. Descendants are reaped by the unconditional marker sweep immediately afterward — every descendant inherits the run marker, so tree-walking is unnecessary. * fix(test): drop duplicate process-death wait in registry cleanup Bot review #1521 (Minor): the duplicated waitForAllDead delayed the authoritative marker sweep by another 5s under the exact stuck-process condition the harness must handle. Keep the single bounded wait; the afterEach marker sweep remains the guarantee.
hapi CLI
Run Claude Code, Codex, Cursor Agent, Grok Build, or OpenCode sessions from your terminal and control them remotely through the hapi hub.
What it does
- Starts Claude Code sessions and registers them with hapi-hub.
- Starts Codex mode for OpenAI-based sessions.
- Starts Cursor Agent mode for Cursor CLI sessions.
- Starts Grok Build locally or via ACP for remote sessions.
- Starts OpenCode mode via ACP and its plugin hook system.
- Provides an MCP stdio bridge for external tools.
- Manages a background runner for long-running sessions.
- Includes diagnostics and auth helpers.
Typical flow
- Start the hub and set env vars (see ../hub/README.md).
- Set the same CLI_API_TOKEN on this machine or run
hapi auth login. - Run
hapito start a session. - Use the web app or Telegram Mini App to monitor and control.
Commands
Session commands
hapi- Start a Claude Code session (passes through Claude CLI flags). Seesrc/index.ts.hapi codex- Start Codex mode. Seesrc/codex/runCodex.ts.hapi codex resume <sessionId>- Resume existing Codex session.hapi cursor- Start Cursor Agent mode. Seesrc/cursor/runCursor.ts. Supportshapi cursor resume <chatId>,hapi cursor --continue,--mode plan|ask,--yolo,--model. Local and remote modes supported; remote usesagent -pwith stream-json.hapi grok- Start Grok Build mode. Seesrc/grok/runGrok.ts.hapi opencode- Start OpenCode mode via ACP. Seesrc/opencode/runOpencode.ts. Note: OpenCode supports local and remote modes; local mode streams via OpenCode plugins.hapi resume [sessionId]- List resumable sessions for this machine or resume one locally.hapi ping-peer <session-id-prefix> <message>- Resume (if needed) and message another session. Prefer this or MCPping_peer/list_peersover reinventing JWT+curl. Also--message-file/--list.hapi inspect-peer <session-id-or-prefix>- Read-only peer metadata + recent message text (no resume). Prefer this or MCPinspect_peerwhen a user cites[title](/sessions/<id>)or Copy-referenceSee session "…" (/sessions/<id>) for context./sessions/<id>is a hub path, not a local file. Optional--limit.
Resume a remote session locally
hapi resume
hapi resume <session-id>
hapi resume lists resumable sessions for the current machine. hapi resume <session-id> hands off an active remote session and opens the same HAPI session in the local terminal.
Authentication
hapi auth status- Show authentication configuration and token source.hapi auth login- Interactively enter and save CLI_API_TOKEN.hapi auth logout- Clear saved credentials.
See src/commands/auth.ts.
Runner management
hapi runner start- Start runner as detached process.hapi runner stop- Stop runner gracefully.hapi runner status- Show runner diagnostics.hapi runner list- List active sessions managed by runner.hapi runner stop-session <sessionId>- Terminate specific session.hapi runner logs- Print path to latest runner log file.
Both start and start-sync accept repeatable --workspace-root <path> (or --workspace-root=<path>). When set:
- The web
/browsepage surfaces scoped file trees rooted at those paths. - The runner refuses
list-directoryandspawn-sessionrequests for paths outside the configured roots. ~and~/fooare expanded.
Omitting the flag keeps the legacy behavior: no scoping, no /browse feature.
See src/runner/run.ts.
Diagnostics
hapi doctor- Show full diagnostics (version, runner status, logs, processes).hapi doctor clean- Kill runaway HAPI processes.
See src/ui/doctor.ts.
Other
hapi mcp- Start MCP stdio bridge. Seesrc/codex/happyMcpStdioBridge.ts.hapi hub- Start the bundled hub (single binary workflow).hapi server- Alias forhapi hub.
Configuration
See src/configuration.ts for all options.
Required
CLI_API_TOKEN- Shared secret; must match the hub. Can be set via env or~/.hapi/settings.json(env wins).HAPI_API_URL- Hub base URL (default: http://localhost:3006).
Optional
HAPI_HOME- Config/data directory (default: ~/.hapi).HAPI_EXPERIMENTAL- Enable experimental features (true/1/yes).HAPI_EXTRA_HEADERS_JSON- JSON object of extra headers to send on CLI → hub requests, e.g.{"Cookie":"CF_Authorization=..."}. Can also be set as theextraHeadersobject in~/.hapi/settings.json(environment variable wins).HAPI_CLAUDE_PATH- Path to a specificclaudeexecutable.HAPI_HTTP_MCP_URL- Default MCP target forhapi mcp.
Runner
HAPI_RUNNER_HEARTBEAT_INTERVAL- Heartbeat interval in ms (default: 60000).HAPI_RUNNER_HTTP_TIMEOUT- HTTP timeout for runner control in ms (default: 10000).
Worktree (set by runner)
HAPI_WORKTREE_BASE_PATH- Base repository path.HAPI_WORKTREE_BRANCH- Current branch name.HAPI_WORKTREE_NAME- Worktree name.HAPI_WORKTREE_PATH- Full worktree path.HAPI_WORKTREE_CREATED_AT- Creation timestamp (ms).
Set for the wrapped agent
-
HAPI_SESSION_ID- The hub session id for the current run, exported into the wrapped agent/CLI child environment at spawn for every flavor (claude / codex / copilot / cursor / gemini / opencode / kimi / grok / pi), both runner-spawned and locally started sessions. Agents can read it to self-target "this chat" over the hub REST API or shell helpers without listing/api/sessions. Prefer the MCPdisplay_imagetool for inline media when it is available; useHAPI_SESSION_IDfor hub REST / shell tooling where MCP is not. To list peers on the same hub/namespace, prefer MCPlist_peers(works from runner-spawned sessions without sitting on the hub host; excludes the calling session). To read another session, prefer MCPinspect_peerorhapi inspect-peer. To message another session, prefer MCPping_peerorhapi ping-peer— do not reinvent JWT+curl. User citations look like[title](/sessions/<id>)or Copy-referenceSee session "…" (/sessions/<id>) for context; pass that<id>assessionIdPrefix. Do not Grep/Glob/sessions/<id>as a local filesystem path. On a remote runner, configure matchingHAPI_API_URL+CLI_API_TOKEN(orhapi auth login/~/.hapi/settings.json) on the runner host so shellhapi ping-peer --listworks; session CLI may export an explicit non-default hub URL into child env, but never mirrorsCLI_API_TOKENinto wrapped agents.Lazy Codex (terminal) sessions export the id only after the hub row is materialized, which happens when the MCP bridge starts — before the agent process is spawned — so path-only self-targeting does not race a missing hub row.
Example (shell fallback when MCP is unavailable) — path-only, self-targets the current session:
bun scripts/tooling/hapi-display-image.mjs /absolute/path/to/image.png "optional title"Explicit other session (prefix or full uuid) still works; that path may list sessions.
Storage
Data is stored in ~/.hapi/ (or $HAPI_HOME):
settings.json- User settings (machineId, token, onboarding flag). Seesrc/persistence.ts.runner.state.json- Runner state (pid, port, version, heartbeat).logs/- Log files.
Requirements
- Claude CLI installed and logged in (
claudeon PATH). - Cursor Agent CLI installed (
agenton PATH) forhapi cursor. Install:curl https://cursor.com/install -fsS | bash(macOS/Linux),irm 'https://cursor.com/install?win32=true' | iex(Windows). - Grok Build CLI installed (
grokon PATH) forhapi grok. Authenticate withgrok login --device-authon headless runner machines, or setXAI_API_KEY. - OpenCode CLI installed (
opencodeon PATH). - Bun for building from source.
Build from source
From the repo root:
bun install
bun run build:cli
bun run build:cli:exe
For an all-in-one binary that also embeds the web app:
bun run build:single-exe
Source structure
src/api/- Bot communication (Socket.IO + REST).src/claude/- Claude Code integration.src/codex/- Codex mode integration.src/cursor/- Cursor Agent integration.src/grok/- Grok Build native TUI + ACP integration.src/agent/- Shared support for ACP-compatible agents.src/opencode/- OpenCode ACP + hook integration.src/runner/- Background service.src/commands/- CLI command handlers.src/ui/- User interface and diagnostics.src/modules/- Tool implementations (ripgrep, difftastic, git).
Related docs
../hub/README.md../web/README.md