* fix(runner): prevent ghost sessions from orphaned spawn webhooks
spawnSession() registers a tracked session entry before the child's
"Session started" webhook arrives. When the 15s webhook timeout fires,
only pidToAwaiter / pidToErrorAwaiter are cleared; the
pidToTrackedSession entry is left in place and the detached child
keeps running. If the child eventually starts and reports its webhook,
onHappySessionWebhook() still finds the stale tracking entry and
promotes the orphan into a normal runner-managed session, surfacing
a message-less "ghost session" in the web UI.
This is easy to reproduce on opus[1m] --resume: observed real-world
case where four rapid "resume" clicks produced four orphan children
that all reported their webhooks ~60 minutes later, creating four
ghost sessions on the dashboard.
Three-part fix:
1. Make the webhook timeout configurable via
HAPI_RUNNER_WEBHOOK_TIMEOUT_MS (default unchanged at 15_000). Users
on slow models / large resumes can raise the ceiling so the
timeout never fires in the first place.
2. On timeout, also delete the pidToTrackedSession entry and SIGTERM
the child, so a late webhook cannot promote the orphan.
3. Defence in depth: if onHappySessionWebhook() receives a webhook
from a PID that is not tracked but whose payload claims
startedBy: 'runner', ignore it and SIGTERM the child. Genuine
terminal-launched children correctly report
startedBy: 'terminal', so this branch cannot false-positive on
them.
* fix(runner): use tree-kill on timeout and clean up worktree for orphans
Address review feedback on the kill path and worktree cleanup:
1. Timeout handler: replace bare `happyProcess.kill('SIGTERM')` with
`killProcessByChildProcess(happyProcess)` so the entire process tree
(wrapper + detached agent grandchildren) is reaped, matching the
existing `stopSession()` behaviour.
2. Orphan webhook handler: replace bare `process.kill(pid, 'SIGTERM')`
with `killProcess(pid)` for proper SIGTERM → SIGKILL escalation.
A ChildProcess reference is unavailable here (tracking entry already
removed), so tree-kill is not possible — but the timeout handler
should have already tree-killed the group; this is defence-in-depth.
3. Worktree leak: when a worktree session times out, register a
one-shot `exit` listener on the child process to run
`cleanupWorktree()` after the child actually exits. Previously
`maybeCleanupWorktree('spawn-error')` would skip cleanup because
the child was still alive at that point, and `onChildExited()` had
no worktree awareness after the tracking entry was deleted — so the
worktree leaked permanently.
---------
Co-authored-by: fengtian <fengtian@users.noreply.github.com>