mirror of
https://github.com/wu736139669/hapi.git
synced 2026-08-05 06:24:37 +00:00
* feat(runner): HAPI_DISABLE_VERSION_HANDOFF opt-out for mtime self-restart The heartbeat in cli/src/runner/run.ts triggers spawnHappyCLI(['runner','start']) + process.exit(0) when getInstalledCliMtimeMs() differs from startedWithCliMtimeMs. The same mtime guard fires in controlClient.isRunnerRunningCurrentlyInstalledHappyVersion when a fresh CLI invocation inspects the live runner. For operators who own process supervision (systemd, tmux, custom rebuild pipelines, etc.), source-file mtimes shift for reasons unrelated to npm upgrades. The clean exit defeats Restart=on-failure under systemd and leaves the machine offline. Setting HAPI_DISABLE_VERSION_HANDOFF=1 in the runner's environment now skips both checks while keeping the rest of the heartbeat (session pruning, state file persistence) intact. Default behavior is unchanged for npm consumers. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(runner): preserve original argv across self-restart and verify handoff The mtime-driven self-restart in cli/src/runner/run.ts spawned `hapi runner start` with no arguments, then process.exit(0)'d unconditionally after a 10s sleep. Two failure modes: 1. The forwarded `runner start-sync` lost the operator's --workspace-root flags (anything passed at the original invocation). Browse + spawn silently degraded to "no workspace roots". 2. If the replacement runner failed to come up at all (build was mid-flight, binary missing, etc.) the original runner still exited cleanly. Under systemd Restart=on-failure that means no runner is brought back, and the machine drops off the hub until manual intervention. Changes: - persistence.ts: add startedWithArgv?: string[] to RunnerLocallyPersistedState - run.ts: snapshot process.argv.slice(2) at startup, persist it on initial state write and on every heartbeat, replay it as the new runner's argv (default to ['runner','start-sync'] when nothing was captured) - controlClient.ts: new waitForRunnerHandoff(oldPid, {timeoutMs}) polls runner.state.json for a different live PID - run.ts: only clearInterval + process.exit(0) when handoff is confirmed. On spawn failure or 30s timeout, refresh the mtime baseline (so we don't respawn-loop on the same drift) and stay alive so the machine keeps serving. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(runner): address Codex review findings on #814 Two Major correctness fixes flagged by upstream Codex review on PR #814: (1) Stale-mtime poisoning on failed handoff (run.ts:854,867) The previous failure paths assigned startedWithCliMtimeMs = installedCliMtimeMs which the next heartbeat persisted to runner.state.json. Downstream isRunnerRunningCurrentlyInstalledHappyVersion() then reported the still-stale runner as current, masking the failure until the *next* genuine mtime change. Symptom: an mtime change that briefly failed to hand off would be silently forgotten. Fix: leave startedWithCliMtimeMs immutable. Gate handoff entry on a new nextHandoffAttemptAt timestamp; failure paths bump it by HANDOFF_RETRY_BACKOFF_MS (5 min) via deferHandoffRetry(). The heartbeat continues to write the honest "still on the old code" mtime, and the runner naturally re-attempts after the cooldown. (2) HAPI_DISABLE_VERSION_HANDOFF not honored by live runner (controlClient.ts:192, persistence.ts) The env var was only checked in the invoking CLI process. Under the documented systemd use case the env is set on the service unit but NOT on the operator's interactive shell - so a shell `hapi runner start` would still treat mtime drift as stale and kill the supervised runner during a rebuild. The exact regression this layer was built to prevent. Fix: capture HAPI_DISABLE_VERSION_HANDOFF at runner start time into state.startedWithVersionHandoffDisabled, persisted via the heartbeat. The controlClient mtime check now OR's the live env var with the persisted snapshot, so any caller honours the running runner's opt-out regardless of their own environment. Tests: cli typecheck clean; 14/14 runner unit tests pass. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(runner): address Codex #814 [Major] argv-capture + handoff race Two additional Major findings on the runner self-restart layer that were not addressed in a49fc57: 1. run.ts:672 - process.argv.slice(2) returns ['start-sync', ...] in compiled binary mode (raw argv is [hapi, runner, start-sync, ...]), so the handoff spawned `hapi start-sync ...` which resolveCommand treats as an unknown top-level and falls back to Claude. Replaced with getCliArgs() (the project's canonical argv normalizer) plus a defensive guard that falls back to ['runner', 'start-sync'] if the captured argv does not begin with 'runner'. 2. run.ts:892 - waitForRunnerHandoff did not actually keep the old runner alive. The child's startRunner() unconditionally called stopRunner() before acquiring the lock or writing its own state, so the parent's /stop handler resolved shutdown and exited BEFORE the child committed. If the child then failed (lock contention, auth error, anything between stopRunner and writeRunnerState), the machine went offline with no runner at all. New handoff protocol: - Parent sets HAPI_RUNNER_HANDOFF_FROM_PID=<pid> on the spawned child's env, then releases the lock BEFORE entering waitForRunnerHandoff (breaks the parent-holds-lock / child-needs-lock-to-write-state deadlock). - On wait-timeout the parent re-acquires the lock (long-retry, 30s) and defers retry; if re-acquire fails (third party took the lock) the parent exits cleanly so it does not stay alive without the lock invariant. - Child detects the env signal; if state.pid matches and that pid is alive, this is an authorized handoff: skip stopRunner(), skip the version-match early-exit, and acquire the lock with a longer retry window (60 attempts x 500ms) so it waits through the parent's asynchronous release. CLI typecheck clean. 14/14 runner unit tests still pass. The wider 46/664 failures in the CLI suite are pre-existing in this branch (unrelated: AppServerEventConverter, cursorEventConverter, hook server, Query) - baseline before this commit has 42+; my changes do not regress them. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
hapi CLI
Run Claude Code, Codex, Cursor Agent, Gemini, or OpenCode sessions from your terminal and control them remotely through the hapi hub.
What it does
- Starts Claude Code sessions and registers them with hapi-hub.
- Starts Codex mode for OpenAI-based sessions.
- Starts Cursor Agent mode for Cursor CLI sessions.
- Starts Gemini mode via ACP (Anthropic Code Plugins).
- Starts OpenCode mode via ACP and its plugin hook system.
- Provides an MCP stdio bridge for external tools.
- Manages a background runner for long-running sessions.
- Includes diagnostics and auth helpers.
Typical flow
- Start the hub and set env vars (see ../hub/README.md).
- Set the same CLI_API_TOKEN on this machine or run
hapi auth login. - Run
hapito start a session. - Use the web app or Telegram Mini App to monitor and control.
Commands
Session commands
hapi- Start a Claude Code session (passes through Claude CLI flags). Seesrc/index.ts.hapi codex- Start Codex mode. Seesrc/codex/runCodex.ts.hapi codex resume <sessionId>- Resume existing Codex session.hapi cursor- Start Cursor Agent mode. Seesrc/cursor/runCursor.ts. Supportshapi cursor resume <chatId>,hapi cursor --continue,--mode plan|ask,--yolo,--model. Local and remote modes supported; remote usesagent -pwith stream-json.hapi gemini- Start Gemini mode via ACP. Seesrc/agent/runners/runAgentSession.ts. Note: Gemini runs in remote mode only; it waits for messages from the hub UI/Telegram.hapi opencode- Start OpenCode mode via ACP. Seesrc/opencode/runOpencode.ts. Note: OpenCode supports local and remote modes; local mode streams via OpenCode plugins.hapi resume [sessionId]- List resumable sessions for this machine or resume one locally.
Resume a remote session locally
hapi resume
hapi resume <session-id>
hapi resume lists resumable sessions for the current machine. hapi resume <session-id> hands off an active remote session and opens the same HAPI session in the local terminal.
Authentication
hapi auth status- Show authentication configuration and token source.hapi auth login- Interactively enter and save CLI_API_TOKEN.hapi auth logout- Clear saved credentials.
See src/commands/auth.ts.
Runner management
hapi runner start- Start runner as detached process.hapi runner stop- Stop runner gracefully.hapi runner status- Show runner diagnostics.hapi runner list- List active sessions managed by runner.hapi runner stop-session <sessionId>- Terminate specific session.hapi runner logs- Print path to latest runner log file.
Both start and start-sync accept repeatable --workspace-root <path> (or --workspace-root=<path>). When set:
- The web
/browsepage surfaces scoped file trees rooted at those paths. - The runner refuses
list-directoryandspawn-sessionrequests for paths outside the configured roots. ~and~/fooare expanded.
Omitting the flag keeps the legacy behavior: no scoping, no /browse feature.
See src/runner/run.ts.
Diagnostics
hapi doctor- Show full diagnostics (version, runner status, logs, processes).hapi doctor clean- Kill runaway HAPI processes.
See src/ui/doctor.ts.
Other
hapi mcp- Start MCP stdio bridge. Seesrc/codex/happyMcpStdioBridge.ts.hapi hub- Start the bundled hub (single binary workflow).hapi server- Alias forhapi hub.
Configuration
See src/configuration.ts for all options.
Required
CLI_API_TOKEN- Shared secret; must match the hub. Can be set via env or~/.hapi/settings.json(env wins).HAPI_API_URL- Hub base URL (default: http://localhost:3006).
Optional
HAPI_HOME- Config/data directory (default: ~/.hapi).HAPI_EXPERIMENTAL- Enable experimental features (true/1/yes).HAPI_EXTRA_HEADERS_JSON- JSON object of extra headers to send on CLI → hub requests, e.g.{"Cookie":"CF_Authorization=..."}.HAPI_CLAUDE_PATH- Path to a specificclaudeexecutable.HAPI_HTTP_MCP_URL- Default MCP target forhapi mcp.
Runner
HAPI_RUNNER_HEARTBEAT_INTERVAL- Heartbeat interval in ms (default: 60000).HAPI_RUNNER_HTTP_TIMEOUT- HTTP timeout for runner control in ms (default: 10000).
Worktree (set by runner)
HAPI_WORKTREE_BASE_PATH- Base repository path.HAPI_WORKTREE_BRANCH- Current branch name.HAPI_WORKTREE_NAME- Worktree name.HAPI_WORKTREE_PATH- Full worktree path.HAPI_WORKTREE_CREATED_AT- Creation timestamp (ms).
Storage
Data is stored in ~/.hapi/ (or $HAPI_HOME):
settings.json- User settings (machineId, token, onboarding flag). Seesrc/persistence.ts.runner.state.json- Runner state (pid, port, version, heartbeat).logs/- Log files.
Requirements
- Claude CLI installed and logged in (
claudeon PATH). - Cursor Agent CLI installed (
agenton PATH) forhapi cursor. Install:curl https://cursor.com/install -fsS | bash(macOS/Linux),irm 'https://cursor.com/install?win32=true' | iex(Windows). - OpenCode CLI installed (
opencodeon PATH). - Bun for building from source.
Build from source
From the repo root:
bun install
bun run build:cli
bun run build:cli:exe
For an all-in-one binary that also embeds the web app:
bun run build:single-exe
Source structure
src/api/- Bot communication (Socket.IO + REST).src/claude/- Claude Code integration.src/codex/- Codex mode integration.src/cursor/- Cursor Agent integration.src/agent/- Multi-agent support (Gemini via ACP).src/opencode/- OpenCode ACP + hook integration.src/runner/- Background service.src/commands/- CLI command handlers.src/ui/- User interface and diagnostics.src/modules/- Tool implementations (ripgrep, difftastic, git).
Related docs
../hub/README.md../web/README.md