Large media downloads (a 19MB video takes ~60s over the slow runner
tunnel) previously showed only a pulsing placeholder, so the load looked
dead and users gave up. Read the response stream against Content-Length
and render a percentage bar while the blob downloads; fall back to an
indeterminate "Loading …" label when the header is absent, and keep the
instant path for cache hits.
Loaded video/audio/file blobs stayed in memory for the whole session
(object URLs were only revoked on unmount), so opening several tens-of-MB
videos in one chat grew memory without bound. Explicitly loaded media now
goes through a shared cache that keeps at most three entries; the least
recently used object URL is revoked on overflow and its card falls back to
the Load button (a sticky eviction flag prevents a slow load from
re-publishing an already-revoked URL). Auto-loaded images keep their
existing per-card lifetime.
A 19MB video download traveled as one ~26MB base64 socket ack. On a remote
runner behind a slow tunnel (measured ~0.35MB/s k2lab->mini) that exceeds
the hub's RPC budget and Bun's HTTP idle cutoff, so the request died with
an empty reply at ~50s (and at ~240s after the first timeout bump).
The route now probes with a 1MB read that reports total size + mime, then
streams the rest chunk by chunk with Content-Length, honours Range/206 for
players and download tools (416 on unsatisfiable ranges), and falls back
to the legacy full payload when the session process predates this change
(it ignores offset/length and reports no size).
Also gives media/files/upload RPCs a 300s budget and raises the hub HTTP
idle timeout to 240s so a slow transfer is not killed while waiting on the
CLI round-trip.
Pasted images were stored inline in the message JSON as full-size base64
data URLs. Every conversation load (open, paginate, mobile reconnect) had
to download and parse those megabytes before anything rendered, with no
lazy loading and no per-image caching. Measured on this hub: 115 messages
holding 138MB of base64, the largest single message 18.3MB.
New messages now carry only a small thumbnail (browser-generated, 768px
webp/jpeg data URL); the original is uploaded as before and the hub keeps
a durable copy under ~/.hapi/attachments (table chat_attachments, schema
v27). The chat renders the thumbnail instantly and lazily swaps in the
original from GET /api/sessions/:id/attachments/:id when the viewer opens;
that endpoint answers sha256 ETag + immutable Cache-Control so repeated
opens are free. Sessions scope the URL, deletes remove the copy, and
hub/scripts/attachments.ts covers stats/verify/gc/export.
Old messages keep their inline previews (migration deliberately deferred).
Verified end-to-end against the local hub: upload returns the attachment
id/url, original 200 with ETag, revalidation 304, cross-session 404, delete
removes file+row (404 after). Typecheck clean; suite status matches the
pre-change baseline; fixtures unchanged.
Six text conflicts, all resolved:
- cli/apiMachine + runner/run: adopt upstream's MachinePathPolicy; keep the
local policy that browsing stays disabled until --workspace-root is set
(banner text + upstream's browse test adapted to assert the local gate).
- cli/codex root.test.ts: keep both the syncHistory split test and upstream's
change_title/name test; fixture return merges metadata + syncHistory.
- web/SessionList: combine the local personal/global pinned split with
upstream's search-relevance ranking (deps union).
- web/AskUserQuestionFooter: take upstream's draft-restore effect and thread
the local skippedByQuestion through the draft store; the local DSH skip test
moves to AskUserQuestionFooter.skip.test.tsx so upstream's vi.mock setup in
the original file cannot shadow it.
Also: bun install for the assistant-ui 0.15.21 bump, regenerated
shared/fixtures from the merged pipeline.
Test status vs the pre-merge baseline (d6cf54d7), identical pre-existing
failures only: cli 7, hub 2 (opencodeClear), web 35, shared 2 (DSH); relay
clean; typecheck clean across all packages.
The runner kills a runner-spawned child whose "session started" webhook
misses its bounded wait (15s by default), so a cold resume whose thread
has a large rollout always failed with "Session webhook timeout for PID
...": the codex shared runtime only reported after reconciling the full
native history (projection.history + refresh + refreshChildren), and
replaying tens of thousands of items took longer than that wait.
Measured on this machine before the fix: two ~800MB rollouts (36k/43k
lines) needed 17-20s to bind; one attempt was killed 2s before its late
webhook landed and was reaped as an orphan.
Split SharedCodexRoot.bind() into the local claim (queue, projection,
metadata, settings) and a separate syncHistory(); runSharedRuntime.bind()
now activates the session and reports it to the runner before the replay.
The session is controllable when the signal fires (controls registered),
and the replay continues in the background; re-sent history is idempotent
because projection uses stable message ids that the hub dedupes by
local_id.
Covered by root.test.ts: bind performs no native history reads, syncHistory
does the reconciliation.
Measured on this machine: from a Background session (a plain `ssh localhost` is
enough) `security show-keychain-info` already fails with "User interaction is
not allowed" (exit 36) and codesign returns errSecInternalComponent, while the
same command in Aqua reports the keychain unlocked and codesign succeeds. So the
key's ACL is not what blocks us - the session cannot reach the keychain at all -
and granting the key "allow all applications" would not help.
Correct the guidance in sign-build.sh and docs/local-deployment.md (both added
earlier today) accordingly: run the deploy from a Terminal window, or unlock the
keychain inside the Background session with the password
(`security unlock-keychain -p`, the CI-style recipe; the interactive form cannot
prompt there). Verified the failure path still prints the new text.
A fixed path plus a pinned identity does not by itself keep TCC grants valid:
each grant stores a code requirement, and a grant created while the binary was
ad-hoc signed stays pinned to that build's cdhash. Replacing the binary at the
fixed path then matches nothing, and since a record exists macOS never prompts
again - it denies silently. Processes already running keep the cached allow
until macOS re-evaluates them, then lose ~/Documents mid-life; new processes are
unaffected. That is exactly how long-running sessions stopped answering with
"OpenCode service failure" / "getcwd: Operation not permitted" while fresh
sessions on the same machine worked.
AGENTS.md gets the short version; docs/local-deployment.md gets the csreq check
(identity-based vs cdhash), the normal re-grant through System Settings, and the
fallback used on mini on 2026-09-17 (delete the stale row, restart tccd, insert a
record copied from a machine with the same signing certificate), including the
row shape and the other two services that rot with it.
The previous revision told people to run 'security unlock-keychain' first, which
cannot work from launchd's Background domain: the command cannot prompt for the
passphrase there, and an unlocked keychain alone does not let a background
process read the private key. It also guessed the session up front from
`launchctl managername`, which flagged legitimate deploys (an SSH session that
unlocked first signs fine).
Detect nothing in advance: run codesign, and when it fails with
errSecInternalComponent / "User interaction is not allowed", print the two
remedies that actually work - run the deploy from a Terminal window on this
machine, or allow the signing key for all applications in Keychain Access.
Verified both paths with the pinned identity (real sign succeeds; a shim that
returns errSecInternalComponent prints the guidance and exits 1).
Also ignore .opencode/ (OpenCode's per-project scratch dir, ~60MB here).
codesign reads the private key from the login keychain, and only the Aqua
session can reach it: agents/daemons run in launchd's Background domain,
where codesign fails with errSecInternalComponent even while Keychain
Access shows the keychain unlocked. sign-build.sh now warns with the fix
(run from a Terminal window, or `security unlock-keychain` first), and
docs/local-deployment.md records the rule.
The live pipeline records only the final model step of an OpenCode turn,
so the usage dashboard showed roughly 42% of OpenCode's own totals. Add
an offline reconciliation ledger instead of touching the streaming path
(a previous attempt that fed cumulative counters into the live pipeline
made the chat/status UI numbers jump):
- shared: read and aggregate `~/.local/share/opencode/opencode.db` per
(day, model), exposed as `@hapi/protocol/opencodeUsage`
- hub: additive `usage_reconciliation` table (no schema bump, safe to
roll back), local scan at startup and every 30 minutes, and the usage
summary prefers snapshots over live events for reconciled sessions
- cli: each runner scans its machine's store every 30 minutes and reports
snapshots over the new `opencode-usage-report` socket event, covering
remote machines and sessions whose process already exited
- the hub-local scan skips sessions whose store lives on another machine,
so it can never wipe remotely reported rows
Pressing Enter to confirm a Chinese IME candidate sent the message because the
team composer had no composition guard; it also ignored the shared
'hapi-composer-enter-behavior' setting. Now it never sends while composing and
follows the setting (send on Enter, or newline with Cmd/Ctrl+Enter).
Grouping every message into collapsed requirement cards made the team page read
as a wall of summaries. The default is now a chronological timeline (with the
AI chatter folding) and the requirement cards moved behind a 'By requirement'
filter; 'Key' and 'Task' views are unchanged.
Cursor renamed their CLI to cursor-agent; the old agent name is easily
shadowed by unrelated tools (on one machine the Grok TUI owns agent), which
made the cursor model probe run the wrong binary - four seconds of clap usage
errors and a broken ACP backend. Resolve cursor commands via PATH candidates
(cursor-agent first, then agent); HAPI_CURSOR_PATH overrides both.
A 2s fixed sleep after launchctl kickstart produced a false-negative health
failure on the mini (hub still booting) and triggered an unnecessary
rollback. Poll /health up to 20s before declaring failure.
Versioned filenames fixed the per-path code-signature cache kill (exit 137)
but made macOS TCC treat every build as a new app: each deploy re-prompted for
Documents/media-library access, and while the prompt was unanswered every
process touching those folders (model probes, session startup) blocked -
visible as slow/timeout requests in the app.
Go back to the fixed path ~/.hapi/bin/hapi (a real file) so TCC asks once and
remembers it across deploys. Guard the signature-cache hazard with explicit
verification: back up the installed binary to ~/.hapi/bin/backups/, install
with a fresh mtime, prove it execs repeatedly (--version x3) plus
codesign --verify, restart the hub / kickstart the runner, and restore the
backup on any failure.
- deploy-local.sh / deploy-remote.sh: fixed-path install, verify, rollback
- prune-backups.sh replaces prune-versions.sh (keeps newest N backups, drops
legacy versioned binaries)
- AGENTS.md, docs/local-deployment.md and the hapi-deploy skill updated
Every deploy writes a new ~110MB binary; without pruning the bin dir grows
unbounded (7.4GB across 48 versions on the mini). Keep the current symlink
target plus the N most recent previous versions (default 2, override with
HAPI_KEEP_VERSIONS).
- add scripts/prune-versions.sh: filename-sorted prune; safe while sessions
are live because unlinking does not affect running processes
- deploy-local.sh and deploy-remote.sh run it after a successful deploy
Cards kept showing 待开始 for requirements whose tasks were all finished,
because the status field is only written when the lead explicitly closes a
requirement. The card now falls back to the task states (blocked > done >
doing) and treats teammate activity without tasks as in progress; an explicit
status still wins.
The card showed the same sentence up to four times: as the title, in the
'your ask' block, in a separate 'original ask' block and again as the opening
human message in the timeline. Now the ask appears once at the top (clamped to
three lines while collapsed, full text when expanded) and the opening human
message is hidden inside the process view.
- cards show 'your ask' and a result block at a glance: the lead's conclusion
when present, otherwise a result derived from task states (n/m done plus the
latest deliverable); the original ask stays one click away
- the process opens to key milestones only (human messages, decisions, task
hand-offs, done/blocked tasks, system notices) with a 'show full process (+n)'
switch for the raw stream
- hub: one-time idempotent backfill files pre-requirement history under
requirements (fresh human asks become requirements; messages/tasks follow by
time window; short acknowledgements do not open requirements)
- task update broadcasts now carry the status in meta so the UI can pick
milestones without parsing text
- requirement cards keep the conclusion (result) and task progress visible while
collapsed, show the original ask when expanded, and offer a 补充说明 shortcut
that threads the next message under the requirement
- composer gains a requirement selector (auto-file / new / a specific one) and
the human message API accepts requirementId + newRequirement
- closing a requirement (or writing its conclusion) now notifies the human
out-of-band via a team-attention event
- AI messages inside a thread lose their heavy card styling (left rule only);
decisions keep the emphasised card
- requirements: a fresh human ask opens one automatically (replies inherit);
tasks and messages carry requirementId (spawn_peer / team_task / team_send
taskId); the lead closes one with team_requirement (status + conclusion)
- web: the team timeline groups messages into collapsible requirement cards
(title, status, conclusion, message count; a card with a pending decision
opens automatically); unassigned messages fall into an 'Other' card
- web: entering a team scrolls to the newest message, follows new messages only
while pinned to the bottom, and offers a 'jump to latest' button otherwise
- web: AI-to-AI chatter (status/chat/question/task-update, not addressed to the
human) folds into a one-line summary by default (runs of >= 2), with the
participants and latest line preview
- tests: requirement store/service coverage and grouping helper tests
to=human alone is an out-of-band notification, not a request for an answer;
the inbox (待你确认) now only contains kind=decision messages (legacy rows
with awaitingHuman still count). Prompts and tool descriptions now tell agents
to use kind=decision when the human must decide and plain broadcasts for
progress they merely need to see.
The stored team_members.status row is only written at join (working) and on
session down (offline), so it went stale and the team page chips showed
'working' for idle members while the app sidebar (live session state) showed
'idle'. Team detail / team list now derive status from the live session
(thinking -> working, inactive/missing -> offline, blocked preserved).
- team settings can edit the budget (max members / messages per minute / chain
depth); changes merge into the existing config and apply immediately
- default member cap 5 -> 8 (lead included)
- new DELETE /api/teams/:id/members/:sessionId (optional ?stopSession=1),
lead cannot be removed; removal is announced in the team log, the removed
session is notified, and its session is archived only when requested
- member list gets a human-confirmed remove action with a stop-session checkbox
- tests: service config merge + route budget/removal coverage
- extract the pinned-identity signing into scripts/sign-build.sh so local and
remote deploys ship the same stable signature (macOS TCC grants key on the
signing identity, not the path)
- add scripts/deploy-remote.sh: signs locally, copies the binary to a new
versioned file on the remote host, swaps the ~/.hapi/bin/hapi symlink,
kicks the launchd job (default com.hapi.runner, override via
HAPI_REMOTE_LAUNCHD_LABEL), and verifies the runner executes the new file;
running sessions survive
- document the remote flow in AGENTS.md and docs/local-deployment.md
The runner keeps its old machine RPC handlers and capability flags until
the process restarts: compiled binaries never self-update because the
heartbeat compares the mtime of the runner's own resolved exec path,
which is fixed for the life of the process.
- scripts/deploy-local.sh runs 'HAPI_CLI_EXECUTABLE=<stable symlink> hapi
runner start' when a runner is running; skips when none; warns and
exits 1 when the refresh fails (the hub deployment stays)
- AGENTS.md and docs/local-deployment.md document the runner refresh and
the manual sequence step
Ad-hoc signatures carry a cdhash-based TCC requirement that changes on
every build, so macOS treats each local binary as a new app and re-asks
protected permissions (Documents, Downloads, media library, ...) after
every deploy.
- add scripts/deploy-local.sh: pins one Apple Development identity
(SHA-1 stored in ~/.hapi/signing-identity, override via
HAPI_SIGN_IDENTITY), signs with identifier run.hapi.cli, copies to a
new versioned file, repoints ~/.hapi/bin/hapi, restarts the hub, and
rolls back when /health fails
- document the identity step and TCC rationale in AGENTS.md and
docs/local-deployment.md; agent sessions must prune TCC-protected
folders (~/Music, ~/Pictures, ~/Movies, ~/Library) from recursive
$HOME sweeps
- drop the one-off scripts/deploy-local-capacity-retry.sh (ad-hoc)
- spawned members inherit the caller's tool/model/thinking level/permission
(spawn request schema + hub runtime + spawn_peer tool); explicit overrides win
- replies from turns triggered by teammates (not by the human) no longer vanish:
system prompt and ACP briefs now require an explicit team_send(to=human) for
human-facing conclusions produced on those turns
- task loop: team_task tool (list/update), dependsOn gating, deliverable evidence
required when a member marks done, auto broadcast on change, web shows deps and
deliverable on the task board
- team-scoped agent tokens (hapi_team_*): issued per team, injected as
HAPI_TEAM_TOKEN at spawn, restricted by the auth middleware to this team's
messages/tasks/status and audited via [TeamToken] hub log lines
teamPrompt tests ran with the real cwd and wrote .hapi/team memory into the
checkout (one charter.md was accidentally committed in ef33332d). Tests now
pass an explicit temp cwd; getTeamPromptBlock/withTeamInstruction accept a cwd.
Drops the stray artifact and ignores .hapi/ in this repo.
teams.db gains team_pending_pings (additive table, no schema-version bump so
the previous binary can still open the file). The hub reloads unanswered human
pings on start, so a restart no longer drops the member-reply bridge.
- team memory moves to <repo>/.hapi/teams/<slug>-<short-id>/ so several teams
on one repository no longer share (and clobber) a single charter
- the pre-teams/ shared .hapi/team dir is moved into the first team dir that
starts after the upgrade
- .hapi/ is appended to the repo .gitignore once (idempotent, best effort)
- Team chat now talks to the lead by default; drop the 'everyone' recipient
(broadcast never woke anyone, which read as 'no response')
- Member decisions (kind=decision / to:human) are flagged awaitingHuman; the
human's reply carries inReplyTo and clears the decision, and a new dismiss
endpoint waves one off without replying
- Web: 'Needs your reply / Replied / Dismissed' chips on decisions, an inline
Reply button that binds the asking member, and a 'Needs you' inbox in the
side panel with a badge on the panel toggle
Old runners ignore the team payload, so the lead was spawned without
HAPI_TEAM_*: no team tools, no team prompt, no repo memory — it then
improvised (reverse-engineering the HAPI API, poking the project) instead
of replying. Gate team spawns in SyncEngine.spawnSession, share the
capability predicate via runnerCapabilities, and mark machines without it
disabled in team mode with an upgrade hint.
Team memory now lives with the code, like CLAUDE.md: <main repo>/.hapi/team/
on the machine where the agent runs. Worktrees resolve to the main repo so
every member shares one directory. The CLI creates charter.md + handoffs/ at
session start (team env only), points the agent at it in the system prompt,
and reports the path in session metadata; the web reads it through the
session file API, so remote runners work too.
Hub keeps only coordination data (teams.db): members/tasks/messages/status.
Its memory-dir management and the /memory API are removed.
Root cause of 'lead never answers': sessions created through the web form
never received HAPI_TEAM_* env, so their system prompts had no team rules
(only spawn_peer-created members did). Now:
- /api/machines/:id/spawn accepts a team descriptor; the runner exports the
env as before
- team mode creates the team first, spawns the lead with the team context,
then attaches it as lead; member mode spawns with the team context before
joining
- OpenCode system prompts (title + native tool instruction) carry the team
block too
- human ping copy is action-first: reply directly (auto-synced to the
group); do not run shell commands or explore the team
The add-member dialog only offered a hardcoded agent list and no model
picker. 'Add member' now opens the New Session form in member mode: a
banner with the team name plus role/task fields, then the full launch
options (every agent, every model, effort, permission mode, worktree).
On create, POST /api/teams/:id/members adopts the session, records the
task and delivers the assignment brief. The team session type is hidden
in this mode.
Member sessions (lead included) no longer scatter across the session
list: they live inside an expandable team row with role + live status,
and are filtered out of the standalone list. GET /api/teams now returns
members so the sidebar can render them without extra requests.
Members replied in their own session and the group chat stayed empty,
which read as 'no response'. The hub now records a pending ping when a
human message is pushed to a member, and on the member's next
thinking->idle transition mirrors its last assistant text into the team
log (deduped: skipped if the member answered via team_send).
Prompts updated: normal replies sync to the group automatically; team_send
is for member collaboration and explicit reporting.
Team creation now lives in the New Session form (session type: Team);
a second, less capable dialog in the sidebar only caused confusion.
The sidebar team section hides itself when there are no teams.
- New Session gains a third session type: Team. All launch options
(agent, every model, effort, permission mode, machine, directory) are
reused; only a team name is added.
- History-import affordances are hidden in team mode; the lead is spawned
as a plain session and the team is created with it as lead, then the app
navigates to the team chat.
- standalone TeamCreateDialog stays for quick creation from the sidebar.
Teams created without a lead previously dead-ended: members cannot be
spawned until a lead exists. PATCH /api/teams/:id now accepts
leadSessionId; the settings dialog exposes a lead picker and delivers the
lead brief on save.
- TeamCreateDialog: pick machine/directory/agent (optionally a worktree),
spawns the lead session, then creates the team - no more adopting an
existing session implicitly
- team chat header and side panel honor env(safe-area-inset-top) so the
mobile status bar / dynamic island no longer covers the title
- hub: deliver a lead brief to the lead session on team creation (a lead
spawned through the normal path has no HAPI_TEAM_* env)
- add-member dialog hints that members follow the lead's machine/directory
Hub
- independent teams.db (own user_version ladder); feature gate default OFF,
so a disabled hub never creates the file and team routes simply do not exist
- /api/teams: list/detail/by-session/memory/messages/spawn/tasks,
human message channel, team rename/archive/delete
- team runtime: spawn members via the existing runner RPC (worktree/yolo
options), peer message routing (mention/dm push, broadcast pull-only),
budgets (max members, message rate, reply chain depth), member-offline
announcements that wake the lead or escalate to the human
- team-attention notifications: in-app toast when visible, web push otherwise
- team memory dir {HAPI_HOME}/teams/<id>/ with charter.md template; read-only
browse API with path-traversal guards
CLI
- MCP tools: team_status / team_read / team_send / spawn_peer (claude, codex
bridge and ACP flavors); spawn_peer stays behind manual approval
- HAPI_TEAM_* env for spawned members; system-prompt block for Claude/Codex;
assignment brief carries team rules for flavors without prompt injection
- RUNNER_CAPABILITIES.agentTeam advertised to the hub
Web
- team group chat route with folded peer threads, key/task filters,
human composer, member/decision/system message cards
- side panel: live member status, interactive task board (create/status/
assignee), team memory browser, budget
- team create / add-member / settings dialogs; sidebar team section;
hub settings toggle for the feature (restart to apply)
- SSE team-updated invalidation
Zero impact when disabled: no teams.db, no routes, no tool or prompt changes.
* fix: preserve fallback models for legacy usage events
Use the session model when historical usage lacks event-level metadata, while retaining indexed attribution across model changes and epoch rebuilds. Explicit event models remain authoritative.\n\nvia [HAPI](https://hapi.run)\n\nCo-Authored-By: Codex <noreply@anthropic.com>
* fix: persist explicit usage models on replay
* fix(web): re-subscribe push when VAPID key changes (stale hub subscriptions)
* fix(web): prune stale push endpoint from the hub after VAPID re-subscribe
* fix(web): record VAPID key only after hub registration succeeds
* fix(web): preserve push registration on unsubscribe failure
via [HAPI](https://hapi.run)\n\nCo-Authored-By: HAPI <noreply@hapi.run>
* feat: add cache-aware token usage dashboard
Track normalized Claude, Codex, and ACP usage with incremental SQLite backfill. Exclude imported transcript history, rebuild usage after history rewrites, and expose an owner-only dashboard with cache-aware totals and breakdowns.
via [HAPI](https://hapi.run)
Co-Authored-By: HAPI <noreply@hapi.run>
* fix: preserve usage model and local dates
via [HAPI](https://hapi.run)
Co-Authored-By: HAPI <noreply@hapi.run>
* fix: normalize cached usage and timezone buckets
via [HAPI](https://hapi.run)
Co-Authored-By: HAPI <noreply@hapi.run>
---------
Co-authored-by: HAPI <noreply@hapi.run>
* feat(web): pin running sessions in an 'in progress' section with a live badge
* feat(web): show project name on pinned running session rows
* feat(web): make the pinned 'in progress' section collapsible
* fix(web): don't auto-expand directory groups when opening pinned running sessions
* fix(web): keep running section open while searching; clear auto-expand guard when selection leaves a group
* feat(web): show machine label on pinned running session rows
* fix(web): make running-section toggle keyboard-accessible with correct filtered state
* feat(web): split pinned running section into working/pending/idle groups with distinct badges