IDLE_TIMEOUT_MS dropped to 30s, a background ticker started, the tab closed. Officer logged the release
thirty seconds later and the job ticked straight through it — the exact point where deleteSession used
to call _claudeKill. Constant reverted.
Also worth writing down: restarting officer does not test this. The process dies outright and
releaseSession never runs, so that only exercises adoptOrphanedSession, which already worked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Officer's session records are in memory and die with `pm2 restart officer`, while the agent is a PM2
peer and keeps generating. `adoptOrphanedSession` rebuilds a binding — but only when a browser
reconnects to a session *by id*, which you can only do if you already knew the id. So a session that
survived a restart was invisible, and nothing could answer "what is running right now".
`claude:list` returns each live session with `isGenerating` and `pendingTasks` — the same two fields the
agent's own `armIdle` consults before collecting a session, so a caller can tell "busy" from "merely
open" the way it does. Surfaced as `GET /chat/live`, which sits beside `/chat/sessions`: those are
transcripts on disk, these are the ones with a process behind them.
`getActiveSessionKeys` is replaced rather than joined. It returned bare keys, could not distinguish a
session mid-turn from one merely open, and had never been called by anything.
`listLiveClaudeSessions` fails toward EMPTY, where `isClaudeGenerating` beside it fails toward alive.
The asymmetry is deliberate: not knowing there means leaving a spinner up, and not knowing here would
mean inventing sessions.
Step 2 of docs/chat-session-lifetime.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Officer's hour-long idle timer was doing two unrelated jobs: collecting its own in-memory binding, which
is its business, and terminating the agent, which is the sidecar's. It could not do the first without
the second, because `unsub` was a closure reachable only through `kill`.
So a browser that went away killed a live agent an hour later — including one the sidecar had
deliberately protected. The sidecar already refuses to collect a session that is mid-turn or holding
background tasks: `task:started` disarms its idle GC, and `armIdle` re-checks and re-arms rather than
firing once. Officer had no view of any of that. A laptop running out of battery overnight took a
`run_in_background` job with it for no reason.
`detach` now sits beside `kill` on both streaming handles, and `_sidecarUnsub` — declared and called for
a long time, never once assigned — is populated at all three sites. `releaseSession` unsubscribes and
forgets the record without killing; the idle timer points at it. `deleteSession` is unchanged, so an
explicit disconnect still ends the session.
The third assignment site was not in the plan: `adoptOrphanedSession` sets `_claudeKill` but nothing
else, so an adopted session that later idled out would have dropped its record while the listener stayed
subscribed — a leak of one per adopt-then-leave.
No double subscription: releasing unsubscribes first, so a returning browser either adopts with a fresh
listener or starts a first turn with none behind it.
Step 1 of docs/chat-session-lifetime.md. Step 2 (a list verb, so running sessions can be found after a
restart) is still open.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Investigation only, no code. The sidecar already protects sessions with background work in flight
(pendingTasks disarms its idle GC); officer's hour-long timer knows nothing about that and kills them
anyway, and does not survive its own restart. Plan is to have officer release its binding instead of
killing, and to add a list verb so running sessions can be found after a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The whole composer bar is the target, not just the textarea — a screenshot dragged out of the macOS
corner thumbnail is a small thing to aim with, and the bar is the biggest thing near where the cursor
already is. Paste already worked; this is the same attach path.
Three details, each of which breaks the drop silently if missed. `preventDefault` on dragover, or the
browser refuses the drop, never fires onDrop, and navigates to the file instead — taking whatever was
typed with it. A depth counter rather than a boolean, because dragenter/dragleave fire for every child
crossed and the highlight strobes as you move over the textarea. And only claiming drags that carry
files, so dragging selected text across the composer neither lights it up nor swallows the drop.
Non-image files in the same drag are ignored quietly: refusing the PDF among them with a toast would be
noise when the three screenshots you meant went in fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Queued prompts are almost always one thought arriving in pieces — a correction, then the thing you
forgot. Answering them one turn at a time made the agent reply to the first without knowing the second
existed, then re-answer once it did. Joined with a blank line between, in the order written, which is
how they read anyway.
The tray is unchanged: they stay separate rows, each removable right up until they go. What merges is
the delivery, not the queue.
Only a lone prompt can still be a slash command. Joined to anything else it is text that happens to
start with a slash, and running it as a command would silently drop everything queued behind it. The
drain now takes the whole queue at once, so it loops twice at most — again only if the batch was a
handled command and something arrived while it ran.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Send no longer refuses mid-turn. The prompt goes into a queue, and each turn's completion delivers the
next — one per turn, which is the whole drain loop. A slash command the client handles itself never
starts a turn, so delivery reports whether it did and the drain keeps going rather than waiting for a
completion that will not come.
Attachments are captured when the prompt is composed, not when it is delivered, so a queued message
keeps the files it was written with instead of picking up whatever is in the tray when its turn arrives.
The composer empties on queue as it does on send — a box that stayed full would read as "it didn't
take", and you would send it twice.
The send button turns amber with a different icon to say the press will not go anywhere yet, and sits
BESIDE stop rather than replacing it: typing a follow-up should not cost you the ability to interrupt.
A tray above the composer lists what is waiting, each item removable — without it a queued prompt is
invisible until its turn, which looks exactly like having lost it.
Stop clears the queue. Ending a turn is precisely the signal the drain waits for, so leaving it alone
fired the next prompt the instant you pressed the button meant to halt things. Nothing is lost: a queued
prompt was recorded in the prompt history when it was written, so Up brings it back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Shell-style. Up walks back, Down walks forward, and past the newest is the draft you were composing when
you left — stashed on the way in, because losing what you had typed to the key pressed to get back to it
would be the worst version of this.
Up only takes the key from the FIRST line and Down from the last. In a multi-line draft there is a line
to move to, and swallowing the arrow would strand the caret; on the edge there is nowhere to go, which is
exactly when history is what was meant. An empty list, or already at the oldest, leaves the key alone too.
Per tab and shared by every chat in it, in sessionStorage. The prompt most worth reaching for is often
one sent somewhere else — re-asking in a fresh chat, or in the other panel — and scoping it per session
would empty the history exactly when a new chat makes it most useful. Slash commands count; they are
prompts you sent.
The tests caught a real one: `record` wrote state while `step` read a ref that only refreshed on the next
render, so sending and immediately pressing Up walked the list as it was one prompt ago. The ref is
written first now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It was built on a belief carried over from Claude Code's terminal: that interrupting means the agent
never read the prompt, so handing it back lets you say it differently. That is not what happens here —
the prompt is delivered and read before escape can land, the transcript keeps it, and the agent answers
it on the next turn. So the composer refilled with something already sent, and sending it again sent it
twice.
Escape means "stop, I'll say it differently" or just "stop". Neither wants the old text back. Focus
still returns to the composer, which serves both.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two faults, and the visible symptom of both was the same: type a new title, press enter, watch it snap
back unchanged.
The rename 404'd. `useClaudeSessions()` was called with no cwd, so the PATCH went without `?cwd=` and
the server searched the default group — `findTranscript(email, cwd, id)` scopes by directory, so every
conversation living in a project was unfindable. Only chats in the default group could ever have been
renamed. The pane now passes the session's own cwd, as the list already did.
And the title it displayed could not have changed even on success. It came from `selected.title`, which
rides the `chat:selected-session` channel — published once when a row is clicked and never updated —
so the invalidation refreshed the row underneath while the header kept the old name. The title is now
resolved once in `ChatDetailPanel` from the sessions query and passed down, so the pane, the page title
and the row are one source. Renaming from the list's pencil retitles an open pane too, which it never
did before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things, one title.
The pane's header is now editable and calls the same `renameSession` the list's pencil does, so both
surfaces write the transcript's `summary` line and the invalidation that follows refreshes the row. The
id comes from the URL rather than `resumeSessionId`: they agree for an ordinary conversation and not for
a merged `/clear` chain, where the resume target is the tail while the list and the server address the
chain by its head — renaming the tail would have written a title nothing displays. `/chat/new` has no
transcript yet, so there the title is read-only.
And on `/chat/<id>` the conversation names the page, sitting between a typed tab name and the route
default: `label ?? override ?? titleForPath()`. Naming a window is deliberate and must still win. Not
gated on full screen, though that is where it earns its keep — the nav header is hidden there, so the
browser tab strip is the only thing telling two side-by-side windows apart. Tiled, the same value fills
the header's centre.
The edit interaction is now one `EditableTitle` shared with the nav header instead of a second copy of
it. `allowEmpty` is what keeps the header's "clear it to hand the tab back to the route name" working;
everywhere else empty means keep, since the rename endpoint 400s on it. `SessionList`'s row rename is
deliberately NOT folded in — it opens from a pencil and confirms with a check, so it is a different
interaction wearing the same styling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Maximize could only ever stop below the nav header — the content region is an `absolute z-2` stacking
context and the header a `fixed z-10` sibling, so that was the deepest a panel could get on its own.
Full screen reaches the rest by asking the shell to stand its header down, which leaves the shallow
version as a state nobody picks on purpose.
So it goes, and the two lights with it. `MaximizeMode` and the `{ id, mode }` session value collapse
back to a bare `fullscreenPanelId` — renamed because "maximized" would now be a lie about what it does
— under a new `FULLSCREEN_PANEL:` key, so a tab open across this reads nothing rather than an object
where a string belongs. `MaximizeButton` is gone; locked screens keep only the fullscreen toggle, which
writes no layout and so was always the one control the lock could permit.
Red stays. The amber used to REPLACE it while maximized so the way out could never be a way to delete;
with amber gone that guard would have cost the close button entirely, and red is already absent exactly
where it should be — locked screens render no traffic lights at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Maximize had one depth: fill the content region, leave the nav header visible. That was never a choice
about how much room to take — the region is an `absolute z-2` stacking context and the header is a
`fixed z-10` sibling, so no z-index a panel gives itself can paint over the nav. `top-[56px]` was the
workaround.
So full screen is cooperative rather than a bigger overlay. The panel asks, and the shell hides its own
header for it; `inset-0` is then genuinely the window. Still the same element and the same class swap —
no portal, no remount, so scroll position and playback survive the step between depths the way they
already survived maximize.
The mode rides beside the maximized panel id in sessionStorage as one value, so the two cannot drift;
a tab open across this change reads the old bare string, gets undefined for `.id`, and lands on
"nothing is maximized".
The toggle is offered from every state, so taking the window is one click from a tiled panel, and it
steps back to a maximized panel rather than all the way out. The amber light is present at both depths
and always goes all the way out, so neither is a trap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The client half exists in monorepo-mobile as of f29774a, with OffChat as the
proof of concept, so this document is no longer purely forward-looking. Adds a
section covering what shipped, the two places the advice here was wrong about
the client — one key per app is not possible on mobile (shared-session puts one
credential in front of all nineteen apps), and signout was never the problem,
distress signout was — and two decisions worth a second opinion: clearing the
credential on 401 only, and leaving the traded-in JWT to expire rather than
blacklisting it, because that handler clears vault tokens keyed on the user.
Platform chat ran at 200K while the same `claude` in a terminal got Opus 5's
full 1M. Nothing to do with the model, the account or compaction tuning — the
CLI gates 1M on `provider === 'firstParty' && Fp()`, and `Fp()` is satisfied
only when ANTHROPIC_BASE_URL is unset or its host is api.anthropic.com. We
point it at 127.0.0.1:5051 so the agent's traffic goes through the OAuth
proxy, which fails that host check and silently drops the window to 200K.
_CLAUDE_CODE_ASSUME_FIRST_PARTY_BASE_URL is the CLI's own escape hatch for a
first-party passthrough, and the assertion holds: the proxy forwards verbatim
to api.anthropic.com and already preserves anthropic-beta.
Read from the CLI binary (2.1.223), not inferred: claude-opus-5 carries
context:{window:1e6, native_1m:true}, so the model half of the gate always
passed. Only the hostname was wrong.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
both forwarders buffered the whole request body into an ArrayBuffer before
re-sending it to Immich. with maxRequestBodySize at 4GB that put a phone's
video upload in the sidecar's heap for a hop that never reads the bytes.
callUpstream now sets duplex: 'half' so a stream is a legal body, matching
what createSidecarProxy already does on the platform side. the four JSON
callers are unaffected.
the platform proxy forwards no content-length, so the body already reached
us chunked; this extends that one hop to Immich. verified against the live
instance (3.1.0): bulk-upload-check round-trips a streamed body and returns
the right verdict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Locked assets are gated on elevation, and elevation lives on a session — an API key
has no `auth.session`, so no key permission or allow-list entry can reach them. The
gate sits inside the generic owner-access check, so a locked asset's thumbnail and
original are covered too, not just its listings.
So `/_locked` mints a session at unlock, holds it in memory for the elevation window,
and closes it on lock, on idle, or when the active immich account changes. Nothing new
is written to photos_config: the pin and the password are never at rest, and a full
compromise of officer's database still does not open the folder. The cost is that
unlocking asks for the immich password as well as the pin.
`auth/*` stays refused wholesale in routes.ts. The four auth routes this needs are
reached through named endpoints that each do one thing, and the elevated forward
carries three resources rather than the main allow-list.
Not yet exercised at runtime — the sidecar has not run this code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
docs/mobile-api-keys.md — what the server accepts, what changes in
monorepo-mobile, and the 401-vs-403 distinction, which is the one that
bites: clearing a good key on a 403 turns a member's missing capability
into a logout loop they cannot escape.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
there were four doors, not two. the ws upgrade in server.tsx and the
vault notifications socket each verified the jwt themselves, so a key
that worked against /api would have 401'd on cliamp — signed in and can
play audio would have been two different questions for the music app.
both now call resolveAuthToken. verified: owner key upgrades cliamp
(101), bogus key 401, member key 403 on terminal exactly as their jwt
is.
reset-password and verify-token deliberately keep verify() — they read
a purpose-scoped reset token and a key must not be spendable as one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
settings > integrations > personal > api keys. mirrors the dav app
password panel, which is the same problem: a secret that exists for one
response, so the new key stays on screen until dismissed rather than in
a toast.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
a user can mint a long-lived key for an app or a device instead of
carrying a 30-day session, so multiple logins on the mobile apps are
per-device revocable rather than one shared token.
identity was being decided independently in userMiddleware and
originScopeMiddleware, each verifying the token itself. teaching only
one of them a new credential format is how those two stop agreeing, so
both now call resolveAuthToken and neither knows what a bearer string
is. verified: a member's key returns the same status as their jwt on
every route tried, 403s included.
a key carries its holder's full authority — not an escalation, it
equals what the password could already do. scoping wants a scopes
column, not a change here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Extracting resolveNotifyUser into its own module immediately caught a hole
in the fix from the previous commit: a header that was present but
unparseable fell through to the body, so a browser could send junk in the
header, name any user in the body and win.
PRESENCE of X-Officer-User is the signal, not its validity — a malformed
header means a proxied request went wrong, and falling through hands the
decision back to the caller we just declined to trust.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A turn could stop producing events and stay `isGenerating` indefinitely. Nothing
covered it: the idle timer answers the opposite question — how long a session with
NO turn in flight may sit before collection — and every client showed a spinner
with no timeout of its own, so a wedged turn presented as a chat that was still
thinking.
On 2026-08-08 one ran for seventeen minutes inside an auto-compaction, reached
over the socket to an iPad, and was indistinguishable there from a dead app. The
compaction is silent by design (the PreCompact hook is the only announcement, and
the code's own comment allows 2.5 minutes), so there was nothing to distinguish it
from.
A stall watchdog now rides every emitted event: any sign of life pushes the
deadline back, and expiry ends the turn the way a real failure would — isGenerating
off, idle re-armed, and an `error` the client can render. The agent process is
deliberately left alive, since it may still be working and the next turn resumes
it; what this guarantees is that the client is TOLD, which is the part that was
missing.
The budgets are generous rather than tight — ten minutes of silence normally,
twenty while compacting, re-armed from the PreCompact hook because that hook fires
as the long silence begins and the deadline the turn is holding was sized for
ordinary work. Killing a turn that was about to succeed is worse than the hang this
prevents.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On a Mac, Claude Code stores its credentials in the login Keychain and never
writes ~/.claude/.credentials.json — the only file this proxy knew how to read.
The workaround was to copy the Keychain blob into that file by hand, which is a
snapshot: a refresh ROTATES the refresh token and revokes the previous one, so
the two stores were not redundant copies but competitors, and whichever
refreshed second got `401 OAuth access token has been revoked`.
That is not hypothetical. On 2026-08-08 it took out every chat turn from the
iPad for six hours while the terminal CLI beside it worked fine — the harness
spawned, retried for three minutes and wrote the 401 into the transcript, which
from the app looks like an agent that simply never answers.
So on darwin the Keychain is the authority and the file is a mirror, holding the
same token rather than a different rotation of it. Everywhere else — every Linux
server — the file is still the authority and nothing changes. Detection is
process.platform, and a machine with no `security` binary or no such item falls
through to the file rather than failing.
Three recoveries, cheapest first:
- a watchdog checks every 30 minutes and refreshes when under an hour remains.
It checks rather than refreshing on a blind schedule because each refresh
rotates the token, so a needless one is another chance for the stores to
disagree.
- an upstream 401 now RE-READS before refreshing. When a token has genuinely
been revoked the machine usually already holds a good one, because Claude Code
refreshed it into the Keychain minutes ago; spending our own refresh token
there is what caused the divergence in the first place.
- only if nobody else has moved do we refresh ourselves.
The Keychain write goes through argv, which is the only non-interactive form
`security` offers, and matches on the service AND account pair — the account is
read off the existing item rather than assumed, or the update would silently
create a second entry instead of replacing the one Claude Code reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidecar fronts a REMOTE instance — its URL and token live in
service_connections, set from /gitea — so it needs nothing installed on the
laptop. That is what separates it from the sidecars left out of this profile,
which supervise a local daemon or container.
It also had to be classified either way: defineProfile throws at load on a name
that is in neither include nor exclude, so leaving it unlisted broke the profile
outright rather than merely omitting it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CLAUDE.md asserted "single-user is a hard invariant, not a stage" while
users held six rows and role_capabilities held grants. Every doc that
repeated it is corrected here, in prose and in the code comments that
carried the same claim.
The accurate statement is narrower: one owner who bypasses every check,
other accounts holding only what their role is granted, and a set of
capabilities — terminal, chat, files, tasks, items, desktop, browser — that
are structurally ungrantable because they execute as the owner's OS user.
TODO.md gains a Multi-user section for what the read turned up: no way to
create a second account, dashboards.id colliding across users, authorize.ts
untested, pty/vault/opencode taking no identity, Radicale still owner_only.
claude-sidecar-isolation.md's open question is answered rather than left
open — the per-email spawn model is dead weight, because chat is an
execution capability and no second account can ever reach it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
reset-password accepted any valid signed jwt as a reset token, including a
30-day session token — its sibling verify-token.ts already gated on
purpose === 'reset-password' and this handler did not. forgot-password mints
that claim, so the gate costs the legitimate flow nothing.
notify's DELETE /_officer/devices/:token deleted by token with no user
predicate: a token is the address of a device, not a secret, so any account
holding the notify capability could deregister another's device.
deletePushDevice now takes an optional userId — the route passes it, the
APNs/FCM dead-token paths deliberately do not.
POST /_officer/notify let a request body's userId override the
proxy-injected X-Officer-User. The header now wins where present, which is
what separates a signed-in browser from a loopback producer that has no
session to speak from.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Also found uncommitted in the shared tree; unrelated to agent panels, so it lands
on its own.
`reset` closed over `initialValue` from the first render, and its `useCallback` dep
list deliberately omitted it — with an eslint-disable to silence the warning that was
correctly pointing at the bug. Any caller whose default is computed (derived from
props, from a fetch, from another piece of state) got reset to whatever that default
happened to be on mount, which after the first render is the wrong value.
Reads through a ref instead, so reset always sees the current default. The
eslint-disable goes away because there is nothing left to suppress — the dep list is
honest now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Committing work that was left uncommitted in the shared tree. I did not write it;
I reviewed it in full, verified it against the running system, and am landing it at
the owner's explicit request because no one currently owns it.
This REPAIRS master. `useAgentPanel.ts` shipped in dbe585f and calls
`/chat/agent-panels`, but `registerAgentPanelRoutes` existed only in the working
tree — so on master as pushed, every one of those calls 404s. The feature has been
half-landed since that commit.
What it is. A panel on a dashboard can be given a name ("frontend", "code-reviewer").
Naming it mints two things: a `sessionKey`, which is the panel's permanent continuity
(it keys the sidecar's on-disk resume map and `chat_session_events`, so the same panel
reopens the same Claude session), and a `handoffToken`, a bearer credential scoped to
exactly one verb. The agent in that panel is then addressable by name, and can pass
work to a peer on the same dashboard over `/api/agent-handoff`.
Three doors, deliberately separate:
- `/chat/agent-panels` (browser, session-authed) — name / list / rename / forget.
Mounted on the chat router rather than given its own prefix: these routes create
and name Claude sessions, which is authority `chat` already grants. A second
top-level mount would have meant a second capability entry claiming the same
thing under a different name.
- `/api/agent-handoff` (agent, token-authed) — peers and send. Unprotected by the
session middleware and exempted in `capabilities/totality.ts` with its reasoning
written down, because the caller is a subprocess with a token, not a browser with
a cookie.
- The transcript stays where transcripts live. DELETE forgets the address and the
panel's claim on the session; it does not touch ~/.claude/projects.
Security, as verified rather than assumed:
- The sender is derived from the token, never from the request body — there is no
`from` field on the wire, so it cannot be forged.
- Every lookup is scoped to the token's `userId` AND `dashboardId`, so an agent can
only see and reach peers on its own dashboard.
- `toAgentPanelView` strips `handoffToken` and `userId`, and it is the only shape
the browser routes return. Confirmed by reading every return path.
- Live-tested: a real token on `GET /api/agent-handoff/peers` returns 200 with
correctly scoped peers; a bogus one returns 401.
Two judgement calls in the code worth knowing about, both already commented at their
site: the introduction turn inlines the handoff token into a runnable curl (a
single-owner MVP trade), and `agent_panels` carries no FK to `dashboards.id` because
that primary key is mid-rework to a composite.
Schema uses `uniqueIndex` throughout, never `unique().on(...)` — the rule that exists
because drizzle-kit mis-diffs named composite unique constraints and re-creates them,
which is what wiped seven tables on 2026-08-03.
NO `bun db:push` IS NEEDED. `agent_panels` is already live in Postgres with 6 rows;
the schema file is catching up to a database that already has it.
Verified: `bunx tsgo` clean, `bun test` 538 pass / 0 fail across 35 files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
a pin on each chip lifts it out of the strip and into a row above it, so the
one task you are actually waiting on stops sliding off the end as newer ones
arrive. more than one can be pinned; the pinned row scrolls like the other.
a pin outranks FINISHED_KEPT and the bulk clear both — it is an explicit
"keep this", and it would be useless if five newer tasks could still evict it.
pinning survives the task finishing, because the outcome is what you pinned it
for.
the tray only had a bulk clear, so getting rid of one finished chip meant
clearing all of them. each finished chip now carries its own close control.
the pill becomes a div wrapping two buttons — a button nested inside a button
is invalid and the browser eats one of the two clicks. running chips stay
undismissable: the tray is the only handle on work still going.
A recovered row was appended, so it landed at the bottom of the conversation instead of beside the
call that spawned it. It has no timestamp, but it does not need one: the harness stamps the task id
into the output of the tool call that started it, and live the task:started event arrives right
after that tool result — so anchoring there reproduces the position the row would have had.
First mention wins, and that is the correctness argument: the id is minted by the call that spawns
the task, so nothing earlier can contain it. Matching the most recent instead was wrong, and real
data caught it — a diagnostic that grepped the transcript printed both live ids and pulled the rows
down beside itself. That case is now a test.
Moved out of the hook into its own module since it is pure and has nothing to do with React.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The harness delivers a finished background task to the agent by writing it as the user's next
message, so Claude's file holds a raw <task-notification> envelope as a user turn. Live it never
shows, because the same event travels separately as task:notification — it appeared only when a
refresh rebuilt the conversation from the file, as a bubble on the owner's side he never typed.
Same defect as INTERRUPTION_MARKERS and the same fix. Anchored to the start of the message so
quoting one inside a real message stays yours. Also skipped when picking a session's title, where
it is no more a title than a slash command is.
Verified against a live transcript: 38 user bubbles before, 32 after, the 6 removed being exactly
the notifications.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
adding an email account failed with a bare "failed to add account" toast. three
defects stacked, each hiding the next.
the email sidecar's http.ts reconstructs what the platform's middleware used to
provide, but only did two of three — bodyParser was never remounted, so every
write route read ctx.get('body') as undefined and POST /accounts threw on
body.provider before ever reaching the credentials.
its onError then read `.status` off the thrown custom-error, which carries
`statusCode`. every deliberate 4xx fell through to the 500 branch and had its
message replaced with "internal error", so a rejected IMAP login and a genuine
crash looked identical. it also answered JSON where the rest of the api answers
errors as plain text. now mirrors hono.ts's handler rather than inventing a
second shape.
useClient threw a plain object, so the ~33 sites narrowing with
`err instanceof Error ? err.message : <fallback>` always took the fallback and
discarded the server's message. now throws an ApiError subclass keeping both
status and message, so those sites start surfacing real errors.
only email reads ctx.get('body'); every other sidecar is a pure proxy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A task row is officer's own invention, synthesised from the harness's system.task_started, and
nothing corresponding to it is ever written to Claude's transcript. So rebuildTranscript can only
produce user/tool/assistant rows, and sync:live deliberately carries no messages — which left the
background-task tray empty after a mid-task refresh even though the work was still running.
Fold the durable log on attach into started-minus-notified and hand that back on sync:live. The
same read now supplies the cursor, so this costs one query rather than two. Finished tasks are
excluded: replaying those would resurrect rows already seen to resolve.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Refreshing mid-turn appeared to kill the agent's output. It never did: the
session survives a dropped socket, the agent keeps generating into it and keeps
committing durable events, and `close` only detaches the socket and arms an
hour-long idle timer. What broke was purely delivery — and the reconnect path
that would have fixed it could not fire, because the browser came back having
forgotten officer's session key. It lived in page state. The only id left was
Claude's transcript uuid in the URL, and nothing accepted that.
So accept it. `attach` carries the uuid, and the agent's on-disk session map —
the single record relating the two — turns it back into the key everything else
is written in terms of. The uuid now also goes out at `system.init` rather than
only at `result`, which is what makes the first turn recoverable at all: until
now a chat had no address until it had finished, and a long first turn is
exactly the one worth reconnecting to.
`sync:live` deliberately carries no messages. The harness writes its transcript
as it goes, so the HTTP load on landing already supplies the past; sending the
server's record of the same messages on top of it would duplicate them, and
there is no shared id to reconcile the two by. Attach hands over the rest of the
turn, the half-written paragraph the transcript cannot hold, and the session's
cursor head — that last one so a *later* drop replays from the head instead of
re-delivering the whole conversation from zero.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tool rows used to close on a five-second timer, so the list shifted under you
while you were reading it and a call you had opened re-collapsed on its own.
open-ness is now derived — a row is open because it belongs to the live turn,
not because it rendered recently — and your own toggle lives above the
virtualiser, which was throwing it away every time a row scrolled out.
the previous turn now folds when you send the next message, and it folds
visibly: the new fold is born holding the rows it replaces, so the height is
unchanged across the swap, then closes over 260ms. reopening an old
conversation still renders collapsed; expanding work is only a service to
someone watching it happen.
autoscroll was failing for three compounding reasons. it was smooth, and a
smooth scroll is by definition away from the bottom for its whole duration,
so new content interrupting it left the view stranded. leaving the bottom was
read as intent regardless of cause, so the interrupted scroll — and any row
measuring past its 150px estimate — silently disarmed pinning until you
scrolled down by hand. and nothing watched for the ResizeObserver correction
that arrives after the estimate, which keeping tool rows open made much worse.
pinning is instant now, intent comes from real gestures, and the list re-pins
when the measured total changes.
25 tests over which controls exist in which mode and whether each calls the handler it is named after. interactive, locked, isMobile, isLastPanel, maximized and an app's own zoomable/transparent flags combine in six separate ternaries; a control present in a mode that should not have it is a way to edit a locked screen, and a control missing from one that should is what the close button was.
Drives PanelSlot directly rather than through WorkspaceView, so the workspace can be put into states a whole view cannot easily be pushed into — maximized, mid-swap, mobile. Nothing new found: the close button is the only defect the chrome had, and it was pinned last commit.
TrafficLights took onRemove and isLastPanel but called onClearApp either way — isLastPanel only chose the tooltip. Closing a panel from its own chrome was impossible: the panel stayed, emptied, and the context menu was the only working path. Worse than cosmetic now that identity lives in the layout — clearing the app drops the config that named the panel's agent while the panel survives to be renamed by whatever is put in it next.
Also wraps the ephemeral pane's sizing effect in a try/catch: react-resizable-panels asserts rather than no-ops when asked to size a group it has not laid out yet, and an assert thrown from an effect aborts the commit — taking the whole dashboard down for a file preview's geometry.
section 9 now carries a concrete instance of its own argument: a test found a bug that
two careful readings of the file had not, in code written three days earlier to prevent
exactly that failure. also records the two bun/testing-library harness facts that cost
more than the fix did, since both present as an unrelated file breaking for no reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the store every dashboard layout is written through had no tests. writing them found
a live defect on its rollback: the guard asked "does the cache still hold what i wrote?"
by reference, and setQueryData runs react query's structural sharing, which rebuilds an
object rather than storing the one it was handed. verified against 5.101.4 — an object
comes back !==, a string comes back ===. so the check was false for every container the
store exists to hold: every layout, every config.agentName. a refused write kept its
optimistic value while the toast said it had been rolled back, and the change vanished at
the next reload. only primitives ever reverted, which is why it went unnoticed.
replaced with a per-key write sequence, which asks the question the identity check meant
to ask — has anything written this key since — and does not depend on identity at all.
14 tests: readValue's kind guard, the optimistic write and its updater composition, and
five on revert including the object regression pin.
also moves testing-library's cleanup into test-setup. it auto-registers afterEach at
module import time, so bun attaches it to whichever file imports the library first and
every later file silently gets none. adding this test file was enough to break fourteen
assertions in DataTable.test.tsx, which does not import it. preload has no file scope, so
registering there removes the ordering from the question.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>