Files
hapi/cli
SSU-WEI HUANGandGitHub 291e7bc40b fix(test): stop runner integration suite from leaking detached process trees (#1515) (#1521)
* fix(test): stop runner integration suite from leaking detached process trees (#1515)

The default CLI test run included runner.integration.test.ts, which spawns
real detached runner/session process trees. A failing, timed-out, or
interrupted test (or a plain runner stop) left those trees alive under
PID 1 — on the Mac this accumulated ~600 Node/Bun/agent processes and
several GiB of RSS over repeated runs.

Test harness changes only; production runner session-preservation
semantics are untouched:

- Exclude runner.integration.test.ts from the default parallel unit-test
  suite; move it into a dedicated serial integration project
  (vitest.integration.config.ts, 'bun run test:integration'). The
  20-session stress test is opt-in via HAPI_RUN_STRESS_TESTS=true.
- Add a test-owned process/session registry (processRegistry.ts): every
  runner, runner-spawned session, and terminal-style child is registered
  immediately after spawn; afterEach/afterAll run two-stage cleanup
  (logical stopRunnerSession first, then bounded process-tree kill),
  followed by a marker sweep for agent grandchildren reparented to PID 1.
- Add a per-run HAPI_TEST_MARKER env stamp + identity/secret env
  neutralization for test children (integrationEnv.ts) so outer HAPI/pi
  session variables never leak into test processes and the final audit
  can recognize test-owned processes by env alone.
- Final suite audit in globalSetup teardown: reap anything still
  carrying the run marker and fail with PID/command diagnostics if
  anything cannot be reaped, before removing the temp home.
- Regression coverage: a deliberately failing test registers a detached
  child and the follow-up audit must find zero test-owned processes.
- CI: replace the dead .env.integration-test step with a dedicated
  integration job running the serial project.

* refactor(test): drop unused killByChildProcess import and child field from registry

* chore(test): raise integration hookTimeout to 60s for slow teardown hosts

* fix(test): fail loudly when the process-table audit cannot scan; assert regression child death

Bot review #1521 findings:
- A failed `ps` scan (unsupported flags, buffer exhaustion, permissions)
  previously returned [] and silently disabled both teardown audit layers.
  It now throws; globalSetup teardown catches the scan error into the
  audit error (temp home is still removed) so the run fails visibly.
- The regression audit test cleaned the leak with the reaper before
  asserting, and force-killed the fresh marked runner. The failing
  test's direct child PID is now asserted dead in afterEach right after
  registry cleanup (before the marker sweep), and the audit test stops
  its own runner gracefully before reaping.

* fix(test): bound the logical cleanup phase so a hung runner cannot stall the hook

Bot review #1521: stopRunnerSession carries the worker's 60s HTTP timeout
(setup.ts raises HAPI_RUNNER_HTTP_TIMEOUT for the stress test), and the
integration hook timeout is also 60s — N sequential stops could exhaust
the hook budget before the process-tree fallback and marker sweep ran,
recreating the very leak this change prevents.

Logical shutdown is now parallel (Promise.allSettled over all tracked
sessions) and the whole phase (stops + PID resolution) races against a
15s budget, so stage-2 tree-kill and the marker sweep always get their
share of the hook window.

* fix(test): bound graceful runner stop in hooks; keep credentials out of audit diagnostics

Bot review #1521 (follow-up):
- stopRunner()'s HTTP stop can burn the worker-wide 60s timeout on a
  hung-but-live runner, starving the marker sweep within the hook budget.
  afterEach/afterAll now race the graceful stop against a 10s bound; a
  runner that does not stop in time is force-reaped by the sweep (it
  carries the run marker) and the next beforeEach's alive-PID guard
  ignores any stale state file.
- The env-bearing ps scan (ps eww) was also used for diagnostics, so the
  first 500 chars of a short-command process could print inherited
  credentials (CLI_API_TOKEN etc.) into teardown error logs. The scan now
  only identifies marked PIDs; command lines are fetched separately
  without 'e', falling back to '(command unavailable)' instead of the
  env dump.

* fix(test): reap runner model-probe orphans before the zero-survivor inspection

Bot review #1521 (Minor): inspect-before-reap. Applying it exposed a real
race: each test's runner legitimately spawns marker-carrying children at
startup (agent acp + agent --list-models model-catalog probes). Stopping
the runner orphans them (ppid 1) with the run marker, so the audit test's
OWN runner polluted the pure inspection with fresh probes spawned after
the failing test's sweep window.

- reapTestOwnedProcesses now re-kills every re-scan iteration instead of
  killing once and only re-scanning, so a process that survived its first
  SIGKILL (mid-exec) or spawned mid-kill is not given a free pass.
- The regression audit test stops its runner, reaps (clearing its own
  legitimate orphan probes), then inspects: anything still marked is a
  genuine survivor the bounded reaper could not remove and fails the
  suite. Killable leaks from the failing test are already asserted dead
  in afterEach before the sweep runs.

* fix(test): strictly bound the marker reaper; make per-test sweep unconditional and verified

Bot review #1521 (follow-up):
- The 10s reap deadline did not bound the awaited per-tree kills: each
  killProcessTreeByPid can wait up to 2s per PID, so several stuck
  processes could still exceed the 60s hook budget. Every process in a
  test-owned tree carries the marker (env is inherited), so tree-walking
  is unnecessary: the reaper now SIGKILLs every marked PID found by each
  scan, fire-and-forget, and re-scans every 250ms — the deadline strictly
  bounds the function.
- The per-test sweep was skipped when the direct-child assertion failed
  first, and its survivors were ignored. afterEach now snapshots the
  regression-child state BEFORE the unconditional sweep, then verifies
  both the registry result and the sweep leftovers.

* fix(test): replace it.fails regression with a direct assertion test

Bot review #1521 (Minor): Vitest applies the it.fails expected-failure
inversion after afterEach, so a broken registry assertion inside the hook
would be masked as an expected failure, and the marker sweep would erase
the evidence before the follow-up audit ran.

The regression is now a normal test that registers a detached child at
spawn time, deliberately performs NO per-test teardown, runs only the
spawn-time registered cleanup, and asserts the child PID is dead. The
afterEach no longer carries the registry-leak assertion (moved into the
test body where it cannot be inverted); the per-test sweep assertion and
the final audit test are unchanged.

* fix(test): bound registry stage-2 tree-kills; require live regression fixture

Bot review #1521 (follow-up):
- Stage-2 killProcessTreeByPid awaits per descendant serially and can
  consume the whole 60s hook for a large/stuck tree. Signals are all
  delivered synchronously (children first) before any waiting, so racing
  the awaits against a 5s budget bounds the phase without skipping any
  kill; waitForAllDead still verifies the outcome.
- The regression test could pass vacuously if its fixture exited during
  the startup delay (the registry exit listener would remove it before
  cleanup). It now asserts the child is alive before running cleanup.

* fix(test): kill registered roots with bare synchronous SIGKILL, no pgrep walk

Bot review #1521 (follow-up): racing the mapped killProcessTreeByPid
calls against a timer does not bound the phase — evaluating the map
invokes each call immediately, and each runs the recursive synchronous
pgrep walk before its first await, which can consume the hook before the
timer, runner stop, or marker sweep run.

Stage-2 now SIGKILLs registered roots directly (fire-and-forget, no
tree walk, no per-PID waits) and waits a bounded 5s for death.
Descendants are reaped by the unconditional marker sweep immediately
afterward — every descendant inherits the run marker, so tree-walking is
unnecessary.

* fix(test): drop duplicate process-death wait in registry cleanup

Bot review #1521 (Minor): the duplicated waitForAllDead delayed the
authoritative marker sweep by another 5s under the exact stuck-process
condition the harness must handle. Keep the single bounded wait; the
afterEach marker sweep remains the guarantee.
2026-08-12 09:29:59 +08:00
..
2025-12-16 15:03:50 +08:00
2025-12-16 15:03:50 +08:00
2025-12-16 15:03:50 +08:00
2026-01-03 22:22:45 +08:00
2025-12-16 15:03:50 +08:00
2026-01-04 20:45:15 +08:00

hapi CLI

Run Claude Code, Codex, Cursor Agent, Grok Build, or OpenCode sessions from your terminal and control them remotely through the hapi hub.

What it does

  • Starts Claude Code sessions and registers them with hapi-hub.
  • Starts Codex mode for OpenAI-based sessions.
  • Starts Cursor Agent mode for Cursor CLI sessions.
  • Starts Grok Build locally or via ACP for remote sessions.
  • Starts OpenCode mode via ACP and its plugin hook system.
  • Provides an MCP stdio bridge for external tools.
  • Manages a background runner for long-running sessions.
  • Includes diagnostics and auth helpers.

Typical flow

  1. Start the hub and set env vars (see ../hub/README.md).
  2. Set the same CLI_API_TOKEN on this machine or run hapi auth login.
  3. Run hapi to start a session.
  4. Use the web app or Telegram Mini App to monitor and control.

Commands

Session commands

  • hapi - Start a Claude Code session (passes through Claude CLI flags). See src/index.ts.
  • hapi codex - Start Codex mode. See src/codex/runCodex.ts.
  • hapi codex resume <sessionId> - Resume existing Codex session.
  • hapi cursor - Start Cursor Agent mode. See src/cursor/runCursor.ts. Supports hapi cursor resume <chatId>, hapi cursor --continue, --mode plan|ask, --yolo, --model. Local and remote modes supported; remote uses agent -p with stream-json.
  • hapi grok - Start Grok Build mode. See src/grok/runGrok.ts.
  • hapi opencode - Start OpenCode mode via ACP. See src/opencode/runOpencode.ts. Note: OpenCode supports local and remote modes; local mode streams via OpenCode plugins.
  • hapi resume [sessionId] - List resumable sessions for this machine or resume one locally.
  • hapi ping-peer <session-id-prefix> <message> - Resume (if needed) and message another session. Prefer this or MCP ping_peer / list_peers over reinventing JWT+curl. Also --message-file / --list.
  • hapi inspect-peer <session-id-or-prefix> - Read-only peer metadata + recent message text (no resume). Prefer this or MCP inspect_peer when a user cites [title](/sessions/<id>) or Copy-reference See session "…" (/sessions/<id>) for context. /sessions/<id> is a hub path, not a local file. Optional --limit.

Resume a remote session locally

hapi resume
hapi resume <session-id>

hapi resume lists resumable sessions for the current machine. hapi resume <session-id> hands off an active remote session and opens the same HAPI session in the local terminal.

Authentication

  • hapi auth status - Show authentication configuration and token source.
  • hapi auth login - Interactively enter and save CLI_API_TOKEN.
  • hapi auth logout - Clear saved credentials.

See src/commands/auth.ts.

Runner management

  • hapi runner start - Start runner as detached process.
  • hapi runner stop - Stop runner gracefully.
  • hapi runner status - Show runner diagnostics.
  • hapi runner list - List active sessions managed by runner.
  • hapi runner stop-session <sessionId> - Terminate specific session.
  • hapi runner logs - Print path to latest runner log file.

Both start and start-sync accept repeatable --workspace-root <path> (or --workspace-root=<path>). When set:

  • The web /browse page surfaces scoped file trees rooted at those paths.
  • The runner refuses list-directory and spawn-session requests for paths outside the configured roots.
  • ~ and ~/foo are expanded.

Omitting the flag keeps the legacy behavior: no scoping, no /browse feature.

See src/runner/run.ts.

Diagnostics

  • hapi doctor - Show full diagnostics (version, runner status, logs, processes).
  • hapi doctor clean - Kill runaway HAPI processes.

See src/ui/doctor.ts.

Other

  • hapi mcp - Start MCP stdio bridge. See src/codex/happyMcpStdioBridge.ts.
  • hapi hub - Start the bundled hub (single binary workflow).
  • hapi server - Alias for hapi hub.

Configuration

See src/configuration.ts for all options.

Required

  • CLI_API_TOKEN - Shared secret; must match the hub. Can be set via env or ~/.hapi/settings.json (env wins).
  • HAPI_API_URL - Hub base URL (default: http://localhost:3006).

Optional

  • HAPI_HOME - Config/data directory (default: ~/.hapi).
  • HAPI_EXPERIMENTAL - Enable experimental features (true/1/yes).
  • HAPI_EXTRA_HEADERS_JSON - JSON object of extra headers to send on CLI → hub requests, e.g. {"Cookie":"CF_Authorization=..."}. Can also be set as the extraHeaders object in ~/.hapi/settings.json (environment variable wins).
  • HAPI_CLAUDE_PATH - Path to a specific claude executable.
  • HAPI_HTTP_MCP_URL - Default MCP target for hapi mcp.

Runner

  • HAPI_RUNNER_HEARTBEAT_INTERVAL - Heartbeat interval in ms (default: 60000).
  • HAPI_RUNNER_HTTP_TIMEOUT - HTTP timeout for runner control in ms (default: 10000).

Worktree (set by runner)

  • HAPI_WORKTREE_BASE_PATH - Base repository path.
  • HAPI_WORKTREE_BRANCH - Current branch name.
  • HAPI_WORKTREE_NAME - Worktree name.
  • HAPI_WORKTREE_PATH - Full worktree path.
  • HAPI_WORKTREE_CREATED_AT - Creation timestamp (ms).

Set for the wrapped agent

  • HAPI_SESSION_ID - The hub session id for the current run, exported into the wrapped agent/CLI child environment at spawn for every flavor (claude / codex / copilot / cursor / gemini / opencode / kimi / grok / pi), both runner-spawned and locally started sessions. Agents can read it to self-target "this chat" over the hub REST API or shell helpers without listing /api/sessions. Prefer the MCP display_image tool for inline media when it is available; use HAPI_SESSION_ID for hub REST / shell tooling where MCP is not. To list peers on the same hub/namespace, prefer MCP list_peers (works from runner-spawned sessions without sitting on the hub host; excludes the calling session). To read another session, prefer MCP inspect_peer or hapi inspect-peer. To message another session, prefer MCP ping_peer or hapi ping-peer — do not reinvent JWT+curl. User citations look like [title](/sessions/<id>) or Copy-reference See session "…" (/sessions/<id>) for context; pass that <id> as sessionIdPrefix. Do not Grep/Glob /sessions/<id> as a local filesystem path. On a remote runner, configure matching HAPI_API_URL + CLI_API_TOKEN (or hapi auth login / ~/.hapi/settings.json) on the runner host so shell hapi ping-peer --list works; session CLI may export an explicit non-default hub URL into child env, but never mirrors CLI_API_TOKEN into wrapped agents.

    Lazy Codex (terminal) sessions export the id only after the hub row is materialized, which happens when the MCP bridge starts — before the agent process is spawned — so path-only self-targeting does not race a missing hub row.

    Example (shell fallback when MCP is unavailable) — path-only, self-targets the current session:

    bun scripts/tooling/hapi-display-image.mjs /absolute/path/to/image.png "optional title"
    

    Explicit other session (prefix or full uuid) still works; that path may list sessions.

Storage

Data is stored in ~/.hapi/ (or $HAPI_HOME):

  • settings.json - User settings (machineId, token, onboarding flag). See src/persistence.ts.
  • runner.state.json - Runner state (pid, port, version, heartbeat).
  • logs/ - Log files.

Requirements

  • Claude CLI installed and logged in (claude on PATH).
  • Cursor Agent CLI installed (agent on PATH) for hapi cursor. Install: curl https://cursor.com/install -fsS | bash (macOS/Linux), irm 'https://cursor.com/install?win32=true' | iex (Windows).
  • Grok Build CLI installed (grok on PATH) for hapi grok. Authenticate with grok login --device-auth on headless runner machines, or set XAI_API_KEY.
  • OpenCode CLI installed (opencode on PATH).
  • Bun for building from source.

Build from source

From the repo root:

bun install
bun run build:cli
bun run build:cli:exe

For an all-in-one binary that also embeds the web app:

bun run build:single-exe

Source structure

  • src/api/ - Bot communication (Socket.IO + REST).
  • src/claude/ - Claude Code integration.
  • src/codex/ - Codex mode integration.
  • src/cursor/ - Cursor Agent integration.
  • src/grok/ - Grok Build native TUI + ACP integration.
  • src/agent/ - Shared support for ACP-compatible agents.
  • src/opencode/ - OpenCode ACP + hook integration.
  • src/runner/ - Background service.
  • src/commands/ - CLI command handlers.
  • src/ui/ - User interface and diagnostics.
  • src/modules/ - Tool implementations (ripgrep, difftastic, git).
  • ../hub/README.md
  • ../web/README.md