POST /api/users plus an Add-account form in Settings > User management. Until now
createUser had one call site — bootstrap, gated on an empty user table — so every
non-owner account anywhere had been inserted into Postgres by hand.
Created accounts are Active. The column defaults to Unverified and signin refuses
anything else with a bare UNAUTHORIZED, which is exactly what made the hand-INSERT
route look like a wrong password.
Also closes a hole found while reading the write path: a second Super Admin was
storable. The CHECK constraint pins user 1's role but cannot see other rows, and
getOwnerUser() was LIMIT 1 with no ORDER BY, so two holders would have made "who owns
this server" a question the query plan answered — and that answer feeds the agent
sidecar's identity, vault access and origin scoping. Both write paths now refuse the
role and getOwnerUser() orders by id.
USER_DIRS and provisionUserDirs move into data-path.ts so the create handler and
scripts/provision-user-dirs.ts cannot disagree about what an account's skeleton is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/app-store, built to the platform's own conventions: a locked WorkspaceView with two panels, the
selection in `?selected=` rather than a channel, and rows that are real links so cmd-click and a pasted
URL both work.
`?selected=` and not a /app-store/:id detail route, per docs/navigation-audit.md: this is a master list
with a live preview, and linking rows to a detail route would make the detail the whole page and destroy
the side-by-side. Both panels read the URL independently — the list and the detail cannot disagree if
neither is telling the other anything.
The install form is generated from the catalogue's fields rather than written per service, which is what
lets a sidecar shipping from its own repository present a form nobody here wrote. `existing` is first in
`modes` by catalogue rule, so the default selection is "I already have one" — the answer that avoids
starting a second copy of something already running.
States are distinguished rather than flattened. Blocked is amber and titled "Needs you", not an error:
everything worked and it is waiting for a token only a person can mint. Installed-and-enabled but with
a dead process shows a warning rather than a tick that lies. And the disable/uninstall copy says plainly
that data, configuration and tables are kept either way, because that is the question anyone hesitates
over before clicking.
The dock tile is CORE, not plugin-derived: the store is how every other feature arrives, so it must
never be one of the things that disappears.
Verified through the API the screen uses — 14 items, email reporting installed/enabled with its process
online, and /app-store present in the capability routes so the tile renders. NOT verified in a browser:
no page has been opened, so the rendering itself is reasoned rather than seen.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An instance may be on another host, behind a reverse proxy on 443 under a path prefix, on a tailnet
address, or on an arbitrary port because the usual one was taken. All ordinary self-hosted setups, and
each one a case where assuming otherwise produces a connection that fails later with no clue why.
Two places were sloppy about it. Jellyfin's placeholder read `http://localhost:8096` and Transmission's
`http://localhost:9091`, which quietly teach that a service must be local and on its project's default
port; both now show remote examples, and the field type says why. And nothing validated what was typed,
so a bare hostname or a URL with a token in the query string was stored as-is.
The rule: reject only what cannot work, normalise what is merely untidy, have no opinion about the rest.
No check that the host is local, that the port matches a default, or that the scheme is https — a
tailnet HTTP service is completely normal.
Trailing slashes go, because `${url}/api/x` otherwise doubles the separator: accepted by some servers
and 404 by others, which is the kind of difference that reproduces on one machine and not another.
Query strings go, because that is where a token hides, and it would sit in a column meant for a
location. Credentials in the URL are refused for the same reason — outside the encrypted secret, and in
every log line that ever prints it.
A missing scheme is named rather than called invalid: it is the commonest mistake, because it is what
people type into a browser.
Verified through the API: `memos.example.com` is blocked with the fix quoted back, and
`https://memos.example.com:8443/memos/` installs and stores normalised — remote host, non-standard port,
path prefix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by breaking it. Installing a sidecar that was already configured replaced its connection row with
the new install's, silently repointing a working service somewhere else — during testing that took the
live Transmission from :9091 to a scratch container on :18092, and the only symptom was that it stopped
working.
Install now blocks instead of overwriting, naming both URLs and offering the choice. Blocked rather
than failed because there is a sensible answer and the user is the only one who has it: keep what is
there, or reinstall with `replaceConnection` to change it deliberately. Harmless on a fresh machine;
this is entirely for the one with an existing setup.
Verified against the live row: an install pointed at a different URL is refused and the original
connection is still there afterwards.
Also makes "do you already have one?" structural rather than a UI convention. Three tests: anything that
can provision must also offer `existing`, `existing` must come first in `modes` since that is the order
the prompt uses, and it must ask for a URL. A new entry added later cannot quietly offer only "provision
one for me" — which is how someone with a working Immich ends up with a second one and finds out when
two libraries disagree.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tier two works end to end. Transmission installed through the API with a real container, verified on
this machine and then removed:
install preflight, provision, connect, schema, assets, process — container up, health-checked,
connection written from what the setup script printed
disable process stopped, then container Exited(0), data intact
enable container back up, then the process
uninstall container gone, install row gone, DATA UNTOUCHED — config, compose file, downloads and
watch directories all still present
Order matters in both directions and it is opposite each way. Enable brings the container up first: a
sidecar that starts before its upstream exists spends its first seconds failing health checks and
logging about a service that is merely not up yet. Disable stops the process first, for the same reason
in reverse.
`down`, never `down -v`, and no `--rmi`: the volumes are the user's data and the images are shared and
expensive to re-pull. Both are deliberate omissions, stated so nobody adds them later as a tidy-up.
Uninstall only brings down containers for `mode: 'provisioned'`. An `existing` install points at a
service the user runs themselves, and `down` there would stop a container Officer never started.
Every compose call tolerates a missing directory rather than failing. Three call sites can legitimately
arrive with nothing there — an `existing` install, a failed install that died before writing the file,
and a resumed uninstall re-running a completed step — and erroring would make a row impossible to
uninstall, which is the one state a user cannot escape.
Adds an `assets` step, before `process`: the dock reads manifests as soon as the install is recorded, so
an icon arriving a moment later shows as broken on the first render. And composeDir is recorded from the
install rather than derived later, because the directory is the user's and they may move it — uninstall
must not guess at a path it is about to run `down` in.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ALL_DOCK_ITEMS was a hardcoded list of everything, so a fresh machine offered Photos, Jellyfin,
Transmission and the rest — each leading to a screen reporting itself unavailable — and adding a sidecar
meant editing the shell. Neither survives sidecars shipping from their own repositories.
Split in two. CORE_DOCK_ITEMS is the baseline that exists on every install: chat, files, terminal, the
app's own screens, and Gitea, which is in the light profile because it fronts a remote instance.
Everything else is derived from installed sidecars' UI manifests, delivered with /capabilities.
Sent with the capability answer rather than fetched separately so the dock has ONE source. Two requests
means two moments, and a dock rendered between them shows a tile for something uninstalled or nothing
for something installed. Filtered by capability server-side too: a member is not handed the manifest of
a feature they cannot use, because "hidden in the client" is the kind of privacy that lasts until
someone opens the network tab.
Verified live. The owner — who bypasses every permission check — does not bypass this: /photos is absent
from routes and present in deniedRoutes because Photos is not installed. Flipping a row's `enabled`
makes its tile leave and return with no process touched.
Two things fell out. A manifest can declare extraTiles, because CalDAV is one sidecar presenting as
Calendar AND Contacts, and collapsing them to keep the model tidy would make the app worse. And
DEFAULT_DOCK_PATHS no longer pins /music: useDock drops a path with nothing behind it, so the default
dock came up a tile short on any machine where Music was never installed — a default that references an
optional feature is how an app looks subtly wrong on a fresh install for no stated reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Real PNG icons are coming, so this is the path they arrive by: a sidecar ships its assets beside its own
code, and install copies them to public/plugins/<id>/ where one static route serves them.
Copied rather than served in place because a sidecar shipping from its own repository has its assets
wherever that repository was unpacked, which is not a path the web server can be taught at build time.
One predictable destination means the serving rule never has to know how many plugins exist or where any
came from. It also makes assets a property of the INSTALL: uninstall removes them, and a plugin nobody
installed serves nothing.
Needed a new route, and the reason is a trap worth recording. `publicRoutes` in server.tsx is built by
globbing ./public at BOOT, so anything copied there afterwards is invisible to it — the first install of
a plugin would show a broken image until the server was restarted, and "install it, then restart to see
the icon" is not an install. `/plugins/*` resolves per request, like /novnc/* and /vendor/* already do.
Unlike those two it answers 404 rather than 500 for a missing file: an unpublished icon is an ordinary
state on a fresh machine, and a 500 would put a red line in the log for every dock render.
Proven end to end with a real asset: slskd's icon moved from public/slskd.png into the sidecar's own
assets/, published against an ALREADY RUNNING server, and fetched at 200 with the right bytes and
content-type — 404 before publishing, no restart between.
public/plugins/ is gitignored: it holds copies, and the originals live with each sidecar.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Real icons are coming and the manifest field already exists — slskd uses it. What is not decided is
where the bytes come from for a sidecar that ships from its own repository: /slskd.png works only
because it sits in the platform's public/, which a marketplace plugin cannot write to.
Records the three options and their trade — marketplace URL (loses icons offline), served by us from the
sidecar's directory (works offline, needs a route and caching), or a data URI (no fetch, but bloats
every manifest) — so the next person meets the question instead of assuming the current path generalises.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two gaps, both from the same root: the app knew what an account MAY use and not what this server
actually HAS.
Availability is now subtracted server-side in the /capabilities answer. "Installed" is orthogonal to
"permitted" and the owner is subject to it — the owner bypasses every permission check, but a capability
they hold unconditionally still means nothing if its sidecar was never installed. Without this the dock
on a fresh machine lists Photos, Jellyfin, Transmission and the rest, each leading to a screen that
reports itself unavailable.
Computed on the server rather than intersected in the client, so the rule lives in one place: the dock
already reads `/capabilities`, and making it read a second list and combine them is how a member's dock
and an owner's dock drift apart. `unavailable` is returned alongside `deniedRoutes` because the two mean
different things to a UI — "not yours" versus "not here yet, install it".
A disabled sidecar counts as unavailable: disable stops the process and its container, so the feature
genuinely does not work, and leaving its icon would make disable look broken rather than effective.
Reading install state failing subtracts NOTHING, matching useCapabilities' deliberate fail-open.
Each entry now also carries a UI manifest — name, icon, colour, rootRoute, routes — because a sidecar
shipping from its own repository has to be able to say what it looks like. The icon is a NAME rather
than an imported component: a manifest has to survive being JSON from marketplace.officer.dev, which a
lucide import cannot make. Tests pin the manifests against the capability registry, so a tile cannot
appear for a route the server guards differently, and against each other, so two sidecars cannot claim
one root route.
No backfill, by decision: this is proven on a blank machine first and applied to alpha from scratch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Email installs end to end now, which was the point of picking it as tier one: no container, no external
wiring, so the machinery is exercised without the provisioning half.
Verified against the running system, not asserted:
POST /api/app-store/email/install -> {"status":"installed","completed":["preflight","schema","process"]}
row -> email mode=config status=installed enabled=true
pm2 -> officer-email online
second install -> all three steps skipped, process not restarted
disable -> stopped
The server boots with the new router, which is the real test of the capability entry: totality.ts throws
before serve() if a mounted router has none, so booting IS the check passing.
pm2.ts shells out rather than importing pm2 as a library. PM2 is already the supervisor and the
ecosystem file is already the definition of how each process runs; a second thing in charge of that
means two supervisors disagreeing. It also means an owner can undo anything the app store did with a
command they already know. The one fact that matters: `pm2 start <name>` fails for a process PM2 has
never seen, so a first install starts from the ecosystem file with --only, and everything after goes by
name. Callers cannot know which case they are in, so startProcess decides.
Disable stops rather than deletes: a stopped process still shows in `pm2 list`, which is the honest
picture. Deleting would make a disabled sidecar indistinguishable from one never installed.
beginInstall returns the existing row instead of replacing it — that is what makes a retry a resume
rather than a re-provision — and clears lastError on the way in, so a UI never shows a stale failure
beside a working service.
The container half of enable/disable/uninstall is deliberately absent rather than stubbed silently: a
disable that leaves Immich running is a different thing from one that stops it, and the difference is
memory on the user's machine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Implements the half of the template contract that faces the platform: answers go in as environment and
never as prompts, and results come back as OFFICER_RESULT_<KEY>= lines on stdout.
A line protocol rather than JSON because the same stream is the user's live log — it goes to a terminal
panel while the install runs. A script that must emit clean JSON cannot also narrate, and one that emits
both needs a framing convention anyway. This mirrors the @@officer:progress@@ sentinel the job runner
already uses, with the same rule: marker lines are plucked out, everything else passes through.
parseResults is pure and tested against the realistic near-misses: a line that MENTIONS the prefix
without starting with it, an empty value (Transmission with no RPC auth returns exactly that, and blank
is a real answer), a value containing `=` (splitting on every one would truncate a credential), and a
prefix with no assignment (a script bug — skipped rather than stored as a blank key).
Verified end to end against a real script: environment reaches it, stderr is forwarded (docker compose
writes its progress there, so dropping it would hide most of what a user watches), OFFICER_NONINTERACTIVE
is set so a script that would block fails loudly instead of hanging behind a web form, and a non-zero
exit is reported with the tail.
Notes an artifact rather than hiding it: the two streams are pumped concurrently, so the error tail can
interleave differently from real time. The live log is correctly ordered; only the summary can read out
of order. Serialising the pumps would make a script that writes heavily to one stream block on the other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Install spans a container start, a health wait, an upstream API call and a process start. Any can fail,
and one of them — a token only a human can mint — is EXPECTED to stop the run. A straight-line function
has two bad options there: unwind everything, or leave a half-installed service that neither works nor
uninstalls, which is the state users cannot get out of.
So each step is named, completion is persisted, and running install again resumes. planSteps is a pure
function of (entry, mode) and the effects are injected, which makes ordering, resume, blocking and
failure testable with no Docker, Postgres, PM2 or Immich in sight. 15 tests cover exactly the behaviour
that only appears when something goes wrong.
Two rules are enforced by the plan rather than remembered at call sites: 'existing' never provisions, so
pointing at an instance the user already runs cannot start a container; and the members step is omitted
entirely for a service with no user concept, so a Transmission install does not report a step that did
nothing — which reads as a silent failure to anyone debugging a member's access.
`blocked` is a first-class outcome, not an error. For Immich the container is up and healthy and only
its own UI can mint a key; calling that a failure would make a normal install look broken and invite the
user to tear down a working container. The blocking step is deliberately NOT recorded as complete, so a
resume re-runs the step the human just answered.
Results feed forward — provision discovers the URL that connect writes down two steps later — over a
copy of the caller's values, so a failure halfway cannot rewrite what an earlier attempt achieved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There is now no uninstall option that deletes data, rather than a careful one that does. A user
uninstalling a sidecar is saying "stop running this", which is not the same sentence as "delete my photo
library", and for Immich or Jellyfin getting that wrong once is unrecoverable. No confirmation dialog
makes it a good default.
So: `docker compose down` without `-v`. Containers and networks go; the service directory and everything
under it stays exactly as it was.
The bind-mount convention already makes this hard to get wrong, which is worth noting because it means
the safety is structural rather than a rule someone has to keep following. Data lives on the host inside
the service directory, so `-v` — which only removes NAMED volumes — could not delete it even if a future
change added the flag back.
`mode: 'existing'` has no disposal question at all: we did not create that service, so uninstall removes
our sidecar and our rows and touches nothing else.
Reclaiming disk becomes its own feature later, with the sizes in front of the user — "Photos is using
340 GB, delete it?" — as a deliberate act rather than a checkbox inside an uninstall flow.
Removed two stale `down -v` references that survived the first pass, one in the schema comment and one
in the design doc's table. Leftovers like those are how a rule becomes permission again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The owner installs, but a server may already have members, and a member added next month needs the same
work. So the unit is (service × member) reachable from two triggers — install a service, provision
existing members; add a member, provision installed services — rather than a loop inside the installer.
Only handling the first works on day one and rots.
No new table. A member is provisioned exactly when they hold a service_connections row: their own
credential, url NULL, inheriting the instance from the owner's. That schema anticipated this before this
existed, and a second record of the same fact would only be able to disagree with the first.
Three outcomes, declared per catalogue entry so the installer never special-cases a service. `accounts`
is fully transparent. `none` is a single-tenant daemon with nothing to do — filtered before the
provisioning loop so callers can tell "nothing to do" from "did nothing", which look identical at a call
site and matter when someone is asking why a member cannot see a feature.
`invite` is not a weaker `accounts`, it is the correct outcome: Vaultwarden derives its encryption key
from the master password, so a credential we could mint would mean a vault we could read. Transparent
right up to where being transparent would be a defect.
The per-service work is an interface implemented beside each sidecar rather than a switch in core — a
central function growing a case per service is what would stop any of this shipping from its own
repository. Implementations must be idempotent, since both triggers can fire for the same pair and a
duplicate account upstream is not ours to undo. Deprovision is optional and defaults to leaving the
upstream account alone: deleting an Immich user deletes their photos.
Written assuming the vault's multi-user adaptation has landed. Today /api/vault is owner-only by an
explicit ownerGate, so a member is refused before Vaultwarden is reached — verified, and out of scope.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each provisionable service gets a directory holding a compose template and a setup.sh. Deliberately the
shape a sidecar needs once it lives in its own repository: metadata, compose, setup script, schema.
The contract (templates/README.md): answers come from the ENVIRONMENT, so the web form fills them in and
a person on a VPS is prompted only for what is missing, and only on a TTY — one script for both, not two
code paths. Idempotent, writes only inside its own directory, streams progress on stdout (the installer
pipes it to a terminal panel), and returns results as OFFICER_RESULT_<KEY>= lines so nothing has to
scrape a log.
House conventions throughout: relative bind mounts so data sits beside the compose file rather than
hiding behind `docker volume inspect`, containers running as the installing user so downloads are not
root-owned, loopback-only ports unless the service's whole job is inbound connections, and no external
networks — the owner's own composes attach to an `nginx` network that a fresh VPS does not have.
Transmission verified end to end on this machine, on non-conflicting ports, then torn down: renders,
starts, waits, reports. Its health check accepts 409 because Transmission rejects the first request by
design — only-200 would have waited out the full timeout against a working daemon. Re-run produced
exactly one container, and files landed owned by the user rather than root.
Vaultwarden covers the case where we GENERATE the credential rather than asking for one. An existing
token is reused, never rotated, because rotating during a resumed install would lock the owner out of
the admin page. The Argon2 hash has its `$` doubled or compose interpolation mangles it. The token is
not returned to the platform at all — the vault sidecar proxies the Bitwarden protocol and never needs
it, and a secret we do not hold is one we cannot leak.
Corrects the design doc, which assumed provisioning always knows the connection. Three shapes: we set
the credential, we generate it, or a human must mint it in the service's UI afterwards (Immich, Jellyfin,
Memos). The third makes "provisioned and running but not yet connected" a real state rather than a
failure.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An install that discovers a missing dependency halfway through has already made a directory, possibly
started a container and written a row, and then has to unwind — leaving the user with something that
neither works nor uninstalls. A 30ms check first is worth most of that.
Verified while writing this: nothing in scripts/ installs Docker, and nothing checks for it.
setup-dockers.sh invokes `docker compose` with no preflight, so a fresh host without Docker fails
partway through setup with a bare "command not found". Recorded in the design doc rather than fixed
here — the intended fix is a setup.sh per sidecar, which is also what a sidecar needs once it ships from
its own repository.
`docker compose version` is the probe, not `docker --version`: the latter passes with a dead daemon,
which is the failure people actually hit. "Not installed" and "daemon unreachable" are reported
separately because the remedies differ.
Checked per MODE, not per entry. A host without Docker can still install Photos by pointing at an Immich
somewhere else; refusing the whole entry is the over-strict check that makes people work around the
installer instead of using it.
Dropped `requires: 'docker'` from the catalogue type. Needing Docker is exactly "this entry can
provision", which `modes` already says, so declaring it twice invites the two to disagree. Derived by
needsDocker instead, and a test asserts the derivation matches every entry.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The install layout a machine should have, seasoned owner or not:
~/officerdev/
platform/ the app
data/ DATA_PATH
dockers/ services the app store provisioned
capabilities/ the file-based item store
One root, everything under it. OFFICER_ROOT derives from DATA_PATH rather than being a second variable
that has to agree with the first.
Deliberately not `~/dockers`, where a seasoned user already keeps their own estate — 47 services on this
machine. That separation buys two things. Containers the app store created are distinguishable from the
user's own structurally, rather than by a naming convention we would have to enforce and they could
break. And we never reason about someone else's compose files: the store does not scan, adopt or modify
anything outside its own directory.
That also simplifies "I already have one of these" — it is answered by the user giving a URL, never by
us finding a directory and guessing whose it is. An earlier draft had the installer adopting existing
directories, which meant reading, and potentially writing over, services Officer did not create.
This development machine predates the convention and derives an ugly-but-correct path, since the project
sits inside ~/dockers/officer.dev. Still isolated, still one root. New installs get the clean shape.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
First slice, on a worktree branch so none of it touches the tree the live server runs from.
`sidecar_installs` — server-level, no userId, because a sidecar is one process serving the machine.
That is the line that keeps the model coherent for several users: installed is server-level and
owner-only, configured is per user in service_connections. A member can use Gitea without being able to
install it or point it somewhere else.
`installed` and `enabled` are separate because they answer different questions, which is what gives the
reversible middle ground: disable stops the process and keeps container, config, schema and data.
`completedSteps` makes install resumable rather than merely retryable — the failure mode being designed
against is a half-installed service that neither works nor uninstalls.
The catalogue is data, not code: no functions, no compile-time coupling, because the same shape has to
arrive as JSON from marketplace.officer.dev later. Its test pins it to the real estate — it offers
exactly the processes the light profile excludes, names processes that exist, and claims capabilities
that exist. That last check earned itself immediately: it caught `vault` (no capability at all — it is
EXEMPT because Bitwarden clients carry a Vaultwarden bearer, not a platform JWT) and `notify` (which
does have one, where I had written null).
Docker templates follow the convention already in use across 47 services in ~/dockers: a directory per
service, compose inside, relative bind mounts so data sits beside it, USER_UID/USER_GID as the owner.
An existing directory is evidence of an existing install and must be adopted, never overwritten.
Records what Phase 0 must not foreclose: a remote marketplace, sidecars moving to their own
repositories, and third-party plugins — including the note that catalogue.test.ts pins Phase 0's
invariant rather than the design's, since that relationship inverts once sidecars leave this repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The previous commit migrated only when SQLite was completely empty, on the assumption that a non-empty
file is an authoritative one. A real account disproved that within minutes of it landing.
The older Gmail backfill had already written SOME keys into SQLite — last_sync_at and the uidvalidity
set — and never the imap_lastuid ones. So the file was non-empty and half-migrated at the same time,
all-or-nothing skipped the migration, and nine imap_lastuid keys stayed only in Postgres. A missing
lastuid makes the next sync refetch that folder from UID 1: on the mailbox this was found on, 18,755
messages and 6.9 GB.
Now merged per key with the file always winning a conflict. That keeps the property all-or-nothing was
protecting — a restored older emails.db still overrides a newer Postgres row for every key it has, so it
cannot be advanced past mail it does not contain — and adds the keys the file never had, which are
exactly the ones whose absence is expensive.
Verified against the live account: all 22 Postgres keys present afterwards, imap_lastuid:INBOX restored
to 208407, and last_sync_at left at the file's older value, so it re-checks a fortnight rather than
skipping it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The messages were in emails.db and the position — last_sync_at, and per-folder uidvalidity/lastuid —
was a jsonb column on email_accounts in Postgres. Two stores for one fact, with an edge that only shows
up when you try to move a mailbox to another machine.
The expensive part of an email account is the first sync: hours of IMAP for a large mailbox, which is
exactly why "copy emails.db to the new server" is the obvious way to bring one across. With the position
in Postgres that silently does not work — the new server's column is empty, !last_sync_at says first
sync, and the whole mailbox downloads again on top of the one just restored.
The other direction is quieter and worse. Restore an OLDER emails.db while Postgres holds a NEWER
position and the sidecar skips every message between the two, permanently, because nothing looks below
lastuid again. Re-syncing is slow; skipping mail is data loss nobody notices.
Not a new idea — the Gmail path already read SQLite and fell back to Postgres, backfilling so the
fallback was taken once. Only the IMAP path had not followed. This extracts that pattern so both use one
copy, and unifies the isFirstSync fork in accounts.ts, which is how the two drifted apart to begin with.
The file wins over Postgres, always, and only migrates when it holds nothing at all. Topping up a
partial position from Postgres would reintroduce precisely the divergence this removes.
email_accounts.sync_meta is kept and marked legacy rather than dropped: it is the one-time backfill
source for every account created before this, and dropping it would strand any that has not synced
since. Nothing writes to it now.
11 tests on the migration, aimed at both expensive failures — migrating when we should not, and failing
to migrate an account that predates the change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The serve has a second, newer API surface nobody here had looked at, and it publishes
exactly what the parity doc calls impossible under stdin ignore: delivery steer and queue
on POST /prompt, an interrupt that does not tear down, and a per-session event stream with
an after cursor — the durable-replay machinery officer hand-built for claude, as a
primitive. That would have made migrating obvious.
It does not execute. A prompt is accepted with an admittedSeq, stored, emits
prompt.admitted and prompted, and then never steps. Ruled out separately: the model, the
permissions (build is *:allow, no pending requests), the per-request location (the surface
is location-scoped via header or a deepObject query, supplied everywhere, no change), and a
config gate. The legacy POST /session/id/message?directory= generates fine in 17s, so the
serve itself works — only the new pipeline is inert. session.next.* is the tell.
And not a version problem, which is the part everything here had backwards: this Mac runs
1.18.11 and alpha runs 1.17.9, measured. The dead pipeline was tested on the NEWER binary.
The original "this server runs 1.17.9" meant alpha and was copied to a machine where it was
false; corrected in runner.ts and the test.
So building against it now would produce code that looks finished and does nothing, which
is the failure mode this project keeps rediscovering. One request reopens the question
after any upgrade, and the doc names it.
Also de-flakes the lifecycle tests: they spawn real processes, and a fixed sleep(750) went
red once on a machine busy running these probes. Presence assertions poll now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
They were permanently unnamed, and the two halves needed to name them already existed on
officer: the sidecar reports its own sessionKey because that is all it has, while the ses_ id
arrives separately over opencode:session and is recorded in opencode/state.ts. Nothing joined
them. /chat/live joins them now, so no protocol or sidecar change — widening
LiveOpenCodeSession would have meant sending the sidecar a fact it told officer in the first
place.
One list call names every row rather than one transcript load each, and it is skipped when
nothing is running or no id has been reported, so an idle Live panel never touches the serve.
Verified against a real turn, which also showed the design working as intended: the first
poll has no id yet and shows nothing, the next shows title and cwd. That window is real and
short, and showing nothing beats showing a key the user has never seen.
Worth knowing: opencode titles its own sessions "New session - <ISO timestamp>", so the row
is located but not meaningfully named. That is genuinely its title, not a bug here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B8, both halves. They share index.ts, so they share a commit.
In-flight turns: `opencode run` is spawned, not supervised, so pm2 restart officer-opencode
left every turn ALIVE — reparented, still spending tokens, still writing files as the agent,
with the only reader of its stdout gone. The transcript stopped mid-tool-call, which reads
as the agent hanging.
stopAllOpenCodeTurns kills them and settles each synchronously, because the caller is about
to process.exit and nothing waiting on proc.exited would ever run. Settling writes a reason,
so a reload after a restart explains itself instead of trailing off. Turns are stopped BEFORE
the connection is destroyed — that write travels over it — and the flush is bounded, since
losing the explanation is bad but hanging the restart is worse.
Stale serves: the sweep read /proc, so it was a no-op on macOS and orphaned serves piled up,
one per unclean exit, each holding a port. Added a pidfile sweep alongside it. A pid we wrote
ourselves needs no cwd guard to prove it is ours, which is the part ps cannot answer portably
(macOS would need lsof), and a serve started by hand is never in the file.
The guard checks command AND subcommand: matching the word serve anywhere in the line would
sweep a running turn whose prompt merely mentioned it. Fixtures are real ps output from both
machines, not invented. Split into serve-sweep.ts because index.ts spawns a serve at module
scope, so a test importing it would start one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B7. `msg.model || DEFAULT_MODEL` declared every session without an explicit model to be
claude-code, and the parity doc recorded only the visible half of what that cost.
The durable false cut-off is real: endTurnIfAgentIsGone asked the claude sidecar about a
key it had never held, was told false, and wrote "the agent went away" into a turn that
was running fine. It survives reload, because surviving reload is what that row is for.
The same default also handed the session to adoptOrphanedSession as a claude one, which
subscribes it to that sidecar bus and pins session.model — so an opencode turn output
never arrived, and stopping it called killClaude on a key that sidecar never had. A stop
button that silently does nothing.
decideResume makes both rules explicit: the server record beats the client claim, and an
unknown harness stays unknown — no adoption, no cut-off check, just the replay. Silence
is the safe failure when the wrong answer is written durably.
DEFAULT_MODEL stays in handleAttach and is now commented as to why: that path reached its
sessionId by asking the claude sidecar to resolve a claudeSessionId, so only claude could
have answered.
First test in api/chat, which had none. websocket.ts has no seam to drive the handler
through, so the decision is extracted and tested; the wiring around it is not covered.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two comments describing behaviour the code does not have.
LiveOpenCodeSession had been inserted between LiveClaudeSession and its docblock, so a
comment about isGenerating, pendingTasks and the idle GC read as documentation for the
OpenCode type — where it is contradicted by the correct comment directly beneath it.
Moved below, and it now states that it carries no ses_ id.
That absence is the point: /chat/live claimed title and cwd come from the session store
"so a turn whose id has not been reported yet shows unnamed". Nothing is looked up, and
there is no id here to look one up with. They are null permanently, not until-known.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A replaced turn is killed but dies asynchronously, so its proc.exited fired long after
the replacement was registered under the same sessionKey — and then ran the whole
completion path against it: emitted "OpenCode exited with code 143", which the sidecar
commits to chat_session_events so a false failure became permanent history, then deleted
the replacement from `running`. That blinded the new Live panel, made the stop button a
no-op and orphaned a process nothing could reach.
Mark the handle before killing it, retire it silently, and identity-check the delete —
a superseded turn does not own that key any more.
Two leaks in the same family, found while fixing it. An early return would not have been
enough: both watchdogs call finish, so the armed 10-minute hardTimer would have fired an
error at whichever turn held the key by then. And handleLine had no `done` guard, so
stdout still draining from the killed process was emitted under the replacement key.
Reproduced before fixing. The lifecycle tests need no real opencode — RunnerConfig.bin
takes a shell script that sleeps. The control test pins that an ordinary non-zero exit
still reports an error, so the guard cannot overreach.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Live panel asked `claude:list` and nothing else, so a running OpenCode turn was invisible — the panel
claimed to show what the agent is doing and silently omitted half of it.
Adds `opencode:list` / `opencode:sessions` and merges both harnesses in `/chat/live`, asked in parallel,
each failing toward empty so one sidecar being down contributes nothing rather than breaking the panel.
The OpenCode row is deliberately thinner than the Claude one rather than faked into parity:
isGenerating always true — a subprocess exists only while it generates, so there is no "merely open"
pendingTasks always 0 — `opencode run` has no background-task concept; reporting a number would
suggest a capability that does not exist
title / cwd null — the session store is keyed on the `ses_…` id the runner reports, not on
our sessionKey, so an unreported turn shows unnamed rather than guessed
This is the incremental option from docs/opencode-serve-path.md — enumeration without moving turns onto
the serve, so it buys the Live panel with no warm sessions, no SSE loop and no lifetime questions.
NOT verified end to end: no OpenCode turn was running to enumerate, so the verb is wired and typechecked
but has never returned a non-empty list. See COMMS/BLOCKERS.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Path names: the serve's cwd is DATA_PATH/opencode_server, not opencode-sidecar. Two comments said
otherwise and would send the next reader to a directory that does not exist.
Version pin: the comment claimed "verified live against 1.17.9" as though that were a property of the
code. It is a property of whichever binary is installed, and this project already runs two — 1.17.9 here,
1.18.11 on the other machine. Says so now, and points at the test as the thing that actually enforces it.
Tests, the first on the OpenCode path. `runner.ts`'s NDJSON → ChatEvent mapping was described as pure and
untested; it was untested but not pure — it lived inside `handleLine` as a closure over `emit`, the
accumulated cost and a reported-session flag, so it could not be called without spawning a binary.
Extracted as `mapRunLine`, genuinely pure: line in, {sessionId, events, costDelta} out. The two concerns
that span lines stay with the caller, because they are not properties of a line — emitting the session id
exactly once, and accumulating cost across steps. Behaviour is unchanged.
11 tests over what the mapping forwards, what it drops and what it must not turn into NaN. The last one
matters: a missing `cost` on a step_finish would otherwise propagate NaN into the turn total.
Phase 1 is complete: dead code deleted (previous commit), comments corrected, tests added.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 1 item 7, done in the order the parity doc asks: read the dead design into a note, then delete it.
`docs/opencode-serve-path.md` records what `event-mapper.ts` and the SSE half of `client.ts` did, and
what a rebuild would want back from them — the delta model and tool-state transitions, which are exactly
parity Phase 3's token streaming rather than new work.
Deleted: `event-mapper.ts` entirely, and `subscribe`/the shared `GET /event` SSE loop, `createSession`,
`postMessage`, `abort` from the client, plus `isServerHealthy` from server-manager. All had no callers.
`client.ts` goes 200-odd lines to 99. What stays is the REST reads the chat list and transcript use:
listSessions, getSession, getMessages, deleteSession, renameSession.
While in there, Phase 2's blocking question turned out to be cheap to settle, so it is answered rather
than left open. The review asked whether the serve can take a per-request directory, since without one a
serve-based turn path would reintroduce the single-directory coupling that shelved this work:
POST /session?directory=/tmp/oc-phase2-probe -> directory: "/tmp/oc-phase2-probe" honoured
POST /session with directory in the BODY -> directory: "<serve cwd>" ignored
It is a query parameter on every /session* route. So the coupling is gone on both architectures and the
blocker is cleared. The note does NOT start the migration: which of the three options to take is a
product call, and it lays them out rather than presuming one.
The first probe put `directory` in the body and appeared to prove the opposite. Recorded in the note,
because it is the obvious way to test this and it gives a confident wrong answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The sidecar seeded an AGENTS.md into the serve's project root telling the agent to read its working
directory from the "Working directory for this session" line in its system prompt. The run path sends no
system prompt, so there was no such line and the instruction had been inert since turns moved off the
serve to `opencode run --dir`.
What it was standing in for, `--dir` does properly — tested rather than assumed
(docs/opencode-phase0-review.md): `--dir` anchors the agent's own file operations, not just the process
cwd, and the anchor survives a multi-step turn with a write in the middle. Nothing replaces it.
The generated file is removed from disk too, not only from the code that wrote it; leaving it would have
kept feeding standing instructions to every session while looking, in the source, as though it were gone.
Also corrects the claim that all OpenCode sessions live in one server's project. True when turns
inherited the serve's directory, false now: one serve lists 7 sessions across several directories, which
is why the cwd filter has to read `directory` rather than assume a single one.
docs/opencode-phase0-review.md, item 2 — Phase 1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Regression from the fold change. Scroll-up paging is driven by a scroll listener, and a scroll listener
only fires on a list that actually scrolls. That was always true and never mattered, because the turn you
were looking at rendered in full and was tall enough on its own.
Folding on reload broke it. A window of twenty messages can be one turn with eighteen tool calls, which
collapses to three short rows: no overflow, no scroll event, and paging never starts. The whole
transcript above becomes unreachable — which reads as lost history rather than as a fetch that never
fired.
So don't wait for a scroll that cannot happen: after each render, if there is more to load and the
content does not overflow its viewport, load the next page. Terminates because each pass either fills the
viewport or exhausts the transcript.
This also fixes a latent case that predates folding — any first window short enough to fit on screen
could never be paged past.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reopened B4. Flipping `images: false` for OpenCode models made the metadata honest but changed nothing
on screen, because no code read the capability: the drop zone, the paste handler and the attach menu all
accepted images on every harness. The lie B4 described — drop a screenshot, watch it render in your own
bubble, have it discarded before the model sees it — was still there.
The flag is now load-bearing. Three entry points gated on `supportsImages`:
- the drop zone does not claim the drag at all (no highlight, no preventDefault), so the browser keeps
it rather than the composer swallowing a file it will drop on the floor
- an image paste falls through to the default
- the attach menu's Image entry is absent
`selectedModel || model` mirrors ModelSelector's `displayModel`, so the gate and the model name on screen
can never disagree. An unknown model allows images: a missing capability should not remove a working
control, and the flag is only false where we know it is false. Nothing to undo when images are plumbed
through OpenCodeRunParams later — the gate stops firing once the capability is true.
docs/opencode-phase0-review.md, item 1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three defects and one honest removal. Each was reproduced before being changed, as the doc asks.
B4 — images were offered and silently discarded. Every OpenCode model advertised `images: true`, the
composer gates on that flag, the bubble rendered the attachment, and `handleOpenCodeChat`'s message type
has no `images` field, so it never left officer. Flipped to false: 61 OpenCode models now decline, the
three Claude ones still accept. Plumbing them through OpenCodeRunParams stays Phase 4; advertising a
capability that does not exist is the part worth fixing today.
B5 — every OpenCode turn overwrote the previous turn's subscription handle without detaching it, so the
old session-scoped listener stayed attached and delivery doubled, tripled, and so on for any termination
that is not result/error/stopped. Deliberately NOT the Claude guard: Claude keeps one persistent session
and skips re-subscribing, while OpenCode spawns a fresh `opencode run` per turn, so a new subscription
each time is correct — detaching the old one is what was missing.
B6 — the sessionKey → `ses_…` map had no writer of deletions, so it grew for the process lifetime and a
reused key resumed a stale OpenCode session. Cleared in `deleteSession` only, never in `releaseSession`:
releasing means "let go, leave it running", and a returning browser must find the same `ses_…` again.
Phase 0 item 1 — the thinking toggle is removed rather than fixed. `thinking` is accepted on the wire
and forwarded by neither channel, so the control changed its own label and nothing else. Out of scope
for both harnesses by decision. The inert plumbing beneath it is left for a follow-up that touches the
socket contract; the props stay accepted-and-unread so no call site had to change.
Phase 0 is complete: B1, B2, B3 landed earlier; B4, B5, B6 and the selector here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resuming an OpenCode conversation dispatched it to the Claude CLI: a `ses_…` id handed to
`claude --resume`. Wrong harness, not degraded output.
The whole six-hop chain, confirmed rather than inferred, because the endpoints alone do not show which
hop drops the value:
/chat/:id fetches detail and sets `selected.model` = `opencode/big-pickle` ← the value exists
ChatDetailPanel renders <NewChat …> without a `model` prop ← dropped here
NewChat reads `initialModel={locationState?.model}` ← unrelated source
nothing in the tree ever writes `location.state.model` ← so always undefined
useChat therefore holds no model and the socket sends none
websocket.ts falls back to the user default, `isClaudeModel` is true
So the model was resolved correctly at the top and read from somewhere else at the bottom. `selected`
has carried `model` all along.
`locationState.model` stays as a fallback rather than being deleted: it is declared on
ChatLocationState and costs nothing to keep for a caller that navigates with one deliberately.
docs/opencode-parity.md B2, which flagged this as the one to verify hop by hop. Its account is accurate;
the added detail is that the value is produced and then dropped, not never produced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects, one cause: the session type declared two fields opencode 1.17.9 does not return.
`GET /session` returns `directory` at the top level. There is no `location` object and no `metadata`.
Re-verified by reading the live server rather than the type.
So `metadata.officer.cwd` was compared against `undefined` for every session and the list filter matched
nothing — and since `cwdOf` substitutes a default when no `?cwd=` is given, the `!cwd` escape never fired
either. There was no configuration in which an OpenCode session appeared in /chat. Confirmed against the
running server: 7 sessions present, 0 returned, and the `OpenCode` badge in SessionList was unreachable
code. Now 1 of 7 is listed under the default chat dir, the other 6 correctly filtered to their own
directories.
And `location?.directory ?? ''` was likewise always '', so resuming a session reported no cwd and
relocated the conversation to the default chat dir — which matters because OpenCode rebuilds its
working-directory system prompt every turn. Detail now reads the session's own record via a new
`getSession`, alongside the transcript.
`officerMeta` and the `metadata` tag are gone rather than fixed: the only writer of that tag
(`client.createSession`) has no callers, because the sidecar creates sessions with `opencode run --dir`.
Tagging would have been a second source of truth for something `directory` already answers.
docs/opencode-parity.md B1 and B3. Its suggested fix — derive from `location.directory` — was written
against a field that does not exist; the doc asked for the version to be checked first, and this is why.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`label ?? override ?? route` made a typed tab name permanent. That is right for navigation — you named
the window to find it again — and wrong the moment you rename the conversation itself: the tab kept the
old name, and kept it across reloads, because the stale one is in sessionStorage. The rename looked like
it had failed.
Both are deliberate acts, so the newer wins. The hard part is telling a rename from ordinary navigation:
from the outside, "same conversation, new title" and "different conversation, different title" are the
same event — a changed override. Clearing the tab name on any change would have wiped it every time you
clicked a chat.
So the override now carries the id of the thing it names. Same id with a new title is a rename and drops
the tab name; a new id is navigation and leaves it alone.
The alternative was to have the panel clear the label directly, which needs a QueryClient dragged across
the workspace boundary the bridge exists to avoid — the shell owns the tab name, so the shell decides
when to drop it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An upload over a few MB came back as a 400 with an empty body, no message anywhere, and nothing logged
by Immich, the sidecar or officer. It was a race, not a size limit. Immich judges an asset from its
first few KB and rejects immediately, then closes; both our hops were still writing the body; Node
treats the leftover bytes as a protocol violation and replaces the application's answer with a bodyless
`400 Bad Request` + `Connection: close`. The real message never reached the wire.
Measured before the change: streamed lost the message 1/4 at 8 MB and 4/4 at 32 MB — probability rising
with size, which is why small photos usually worked and a phone's video never did.
Both hops needed it. Fixing only the sidecar took 32 MB from 4/4 failing to 2/4, because the platform
proxy was losing it one hop up.
Bounded at 512 MB, above which the body streams exactly as before. That ceiling is not a refusal and is
deliberately not a 413: a file Immich ACCEPTS is read to the end and never races, so a 4 GB video is
unaffected. All that is given up above the cap is the error message on a file that was going to be
rejected anyway. A first attempt refused over-cap uploads outright and would have broken the working
4 GB case to improve diagnosis of the doomed one.
`bufferRequestBody` is opt-in and off by default: the vault and wallet proxies must keep streaming so a
passphrase or macaroon never lands in the platform's heap.
Also adds the proxy error logging that made this findable at all — status and two byte counts from
headers, never the bodies. `responseBytes: "unknown"` is what exposed the stripped response.
Verified live at 32/256 MB (buffered) and 640 MB (streamed, passes through).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sending a message used to collapse the turn above it, live, while you were still reading it. Watching
your own conversation fold up under you as you typed the next message is worse than the scrolling it
saved.
Folding now keys off `historicalCount` — how many messages at the front of the list came from the server
rather than from this sitting. Nothing collapses while you are watching, however many turns you send;
reload, and all of it has become history and folds at once, which is where the grouping actually earns
its place.
A count rather than a set of ids because everything historical is contiguous and at the front: the
preload seeds it, paging older messages prepends to it, live turns append past it, and a resume replaces
the list with a transcript that is history in its entirety.
Two consequences worth stating rather than discovering.
The last turn is no longer exempt. It used to be excluded from folding for being the live one by
definition; now it is an ordinary turn, so a reloaded conversation folds its final turn too — except
while it is still generating, since hiding work as it arrives is the exact thing being undone.
And folding is no longer a pure derivation over the message list, so a reload does NOT render identically
to a live session. That property was deliberate and is deliberately given up; it is the feature. The
state it costs is one number in useChat, never on the wire and never on disk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The agent's input is a streaming iterable, and a message pushed onto it mid-turn is picked up at the
next step boundary — the running turn reads it and carries on with the added context. Verified rather
than assumed: a probe pushed a sentinel four seconds into a turn busy with `sleep` calls, and that
turn's own final answer quoted the late instruction and obeyed it. One turn, one result, nothing
interrupted.
So the design this was heading for — stop the turn, then re-send the message wrapped in "please
continue, but…" — is not needed. Nothing is abandoned mid-flight, no tool call dies half-applied, and
the agent is never told to stop something it was part-way through.
Enter on an empty composer delivers the queue now instead of waiting for the turn to end. That keystroke
was free: handleSend has always returned immediately on empty input. The queue still fills and still
shows as it did, so the default behaviour is unchanged — this is the impatient path, not a replacement.
Officer needed nothing: handleChat already pushes onto the live session rather than opening a new one
whenever `_claudeKill` is set, which is exactly the injection. The only thing in the way was the
client's own refusal to send while generating.
The affordance is stated above the queue because the keystroke is otherwise undiscoverable — Enter on an
empty box has never done anything, so nobody would try it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The file-browser API does not filter them — readdir returns everything — so browsing for a working
directory opened onto .cache, .local, .npm and thirty more before anything worth picking.
Hidden by default, one toggle in the footer to reveal, and shown dimmed when revealed so they read as a
different class of thing. The count sits on the toggle and the empty state names it too: a folder
holding only dot-directories used to say 'No subfolders here', which is a lie with no way to notice it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every live row read 'Starting…' and linked nowhere, because the title lookup searched for a transcript
named after the session KEY. It isn't. The key is officer's handle for a conversation; the transcript is
named after Claude's own session id, and the mapping between them exists only inside the agent
(setClaudeSession/getClaudeSession). I assumed the two were the same and never checked — confirmed wrong
by looking for the ids from the officer log under ~/.claude/projects and finding nothing.
claude:list now reports claudeSessionId beside the key, the route resolves titles by that, and rows link
to it. Null means the first turn has not reported one yet, which is a genuinely unwritten conversation
and stays unlinked.
Needs the agent sidecar restarted to take effect — the new field comes from there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The titles were looked up client-side against the sessions query already in cache, which only covers the
group being browsed — so anything running in another directory rendered as a truncated uuid, which is
most of them.
Resolved server-side now. The agent reports keys and nothing else, so liveSessionTitle finds the
transcript by scanning the project slugs, takes the cwd off its own first entry, and hands that to
claudeSessionContext — the same path the list uses, so the two agree on naming, /clear chains merged
included, rather than offering a second opinion.
A null title means no transcript has been written yet. That row says 'Starting…' and is deliberately not
a link: pointing at a session you cannot open yet is worse than plainly not being a link.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The previous commit added it to defaultLayout, which anyone who has ever opened /chat never sees:
useDashboardState seeds its default only when the key is ABSENT, so a stored layout keeps the shape it
had when it was first written. appTypes/normalizeLayout does not cover this — it repairs which app a
panel runs, never the tree — so the change was visible only on a fresh account. It was shipped with a
note to reset the layout by hand, which is not a fix.
The screen now replaces a layout with no chat-live panel. Replacing outright is safe here specifically
because the screen is locked: the structure is dictated by code, and the only user contribution is
column sizes. Terminates because the replacement contains the panel it tests for.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The /chat sidebar is now a vertical split: live sessions on top, the transcript list below. They look
similar and answer completely different questions — the list reads conversations from disk, thousands of
them, while this reads the agent's in-memory map over `claude:list`. Only the second can tell you a
conversation is still working while nothing is on screen, which is exactly the state that has been
invisible: after a `pm2 restart officer`, or from a browser that has never seen the session, officer has
no record of a live turn and only the agent can say.
`pendingTasks` is surfaced per row because it is the load-bearing number. It is what keeps a session
alive with nothing on screen, and what makes restarting the agent sidecar unsafe at that moment.
Polled at 10s rather than pushed: liveness changes without officer being told — a turn ends, a
background task reports — so there is no single event to subscribe to. The request is one map read.
Titles come from the sessions query already in cache, so they cost nothing, but that query only covers
the group being browsed and a live session can be in any of them. Unmatched rows show a short key rather
than inventing a name, and an unsaved chat renders unlinked rather than pointing at a transcript that
does not exist yet.
Closes the UI half of step 2 in docs/chat-session-lifetime.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
resolveOwner already retried forever — but only when the query SUCCEEDED and returned nothing, which is
a fresh install waiting on bootstrap. A query that THREW escaped the function, rejected the top-level
await and exited the process, into exactly the PM2 restart loop its own comment says it exists to avoid.
So any Postgres restart (57P03 'the database system is starting up') or moment of unavailability killed
every live agent session on the machine and spun the sidecar until the database answered.
That is what took a session down on 2026-08-10, and why this process showed 468 restarts against 0 for
every peer that starts without needing the database.
The loop now catches as well as checks. Still retries forever, matching the case beside it: a database
coming back is a matter of time, and an agent that gave up would need a human to notice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Officer's session records are in memory and die with `pm2 restart officer`, while the agent is a PM2
peer and keeps generating. `adoptOrphanedSession` rebuilds a binding — but only when a browser
reconnects to a session *by id*, which you can only do if you already knew the id. So a session that
survived a restart was invisible, and nothing could answer "what is running right now".
`claude:list` returns each live session with `isGenerating` and `pendingTasks` — the same two fields the
agent's own `armIdle` consults before collecting a session, so a caller can tell "busy" from "merely
open" the way it does. Surfaced as `GET /chat/live`, which sits beside `/chat/sessions`: those are
transcripts on disk, these are the ones with a process behind them.
`getActiveSessionKeys` is replaced rather than joined. It returned bare keys, could not distinguish a
session mid-turn from one merely open, and had never been called by anything.
`listLiveClaudeSessions` fails toward EMPTY, where `isClaudeGenerating` beside it fails toward alive.
The asymmetry is deliberate: not knowing there means leaving a spinner up, and not knowing here would
mean inventing sessions.
Step 2 of docs/chat-session-lifetime.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Officer's hour-long idle timer was doing two unrelated jobs: collecting its own in-memory binding, which
is its business, and terminating the agent, which is the sidecar's. It could not do the first without
the second, because `unsub` was a closure reachable only through `kill`.
So a browser that went away killed a live agent an hour later — including one the sidecar had
deliberately protected. The sidecar already refuses to collect a session that is mid-turn or holding
background tasks: `task:started` disarms its idle GC, and `armIdle` re-checks and re-arms rather than
firing once. Officer had no view of any of that. A laptop running out of battery overnight took a
`run_in_background` job with it for no reason.
`detach` now sits beside `kill` on both streaming handles, and `_sidecarUnsub` — declared and called for
a long time, never once assigned — is populated at all three sites. `releaseSession` unsubscribes and
forgets the record without killing; the idle timer points at it. `deleteSession` is unchanged, so an
explicit disconnect still ends the session.
The third assignment site was not in the plan: `adoptOrphanedSession` sets `_claudeKill` but nothing
else, so an adopted session that later idled out would have dropped its record while the listener stayed
subscribed — a leak of one per adopt-then-leave.
No double subscription: releasing unsubscribes first, so a returning browser either adopts with a fresh
listener or starts a first turn with none behind it.
Step 1 of docs/chat-session-lifetime.md. Step 2 (a list verb, so running sessions can be found after a
restart) is still open.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The whole composer bar is the target, not just the textarea — a screenshot dragged out of the macOS
corner thumbnail is a small thing to aim with, and the bar is the biggest thing near where the cursor
already is. Paste already worked; this is the same attach path.
Three details, each of which breaks the drop silently if missed. `preventDefault` on dragover, or the
browser refuses the drop, never fires onDrop, and navigates to the file instead — taking whatever was
typed with it. A depth counter rather than a boolean, because dragenter/dragleave fire for every child
crossed and the highlight strobes as you move over the textarea. And only claiming drags that carry
files, so dragging selected text across the composer neither lights it up nor swallows the drop.
Non-image files in the same drag are ignored quietly: refusing the PDF among them with a toast would be
noise when the three screenshots you meant went in fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Queued prompts are almost always one thought arriving in pieces — a correction, then the thing you
forgot. Answering them one turn at a time made the agent reply to the first without knowing the second
existed, then re-answer once it did. Joined with a blank line between, in the order written, which is
how they read anyway.
The tray is unchanged: they stay separate rows, each removable right up until they go. What merges is
the delivery, not the queue.
Only a lone prompt can still be a slash command. Joined to anything else it is text that happens to
start with a slash, and running it as a command would silently drop everything queued behind it. The
drain now takes the whole queue at once, so it loops twice at most — again only if the batch was a
handled command and something arrived while it ran.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Send no longer refuses mid-turn. The prompt goes into a queue, and each turn's completion delivers the
next — one per turn, which is the whole drain loop. A slash command the client handles itself never
starts a turn, so delivery reports whether it did and the drain keeps going rather than waiting for a
completion that will not come.
Attachments are captured when the prompt is composed, not when it is delivered, so a queued message
keeps the files it was written with instead of picking up whatever is in the tray when its turn arrives.
The composer empties on queue as it does on send — a box that stayed full would read as "it didn't
take", and you would send it twice.
The send button turns amber with a different icon to say the press will not go anywhere yet, and sits
BESIDE stop rather than replacing it: typing a follow-up should not cost you the ability to interrupt.
A tray above the composer lists what is waiting, each item removable — without it a queued prompt is
invisible until its turn, which looks exactly like having lost it.
Stop clears the queue. Ending a turn is precisely the signal the drain waits for, so leaving it alone
fired the next prompt the instant you pressed the button meant to halt things. Nothing is lost: a queued
prompt was recorded in the prompt history when it was written, so Up brings it back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>