b4f88ec161a231b9af6f8c6cde56516c23f15587
978
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b4f88ec161 |
routes refuse at the route, and a new home is empty
Three things, from a member sitting on /music with no music capability on a server with no music sidecar: an empty library, and 403s in the console. PERMISSIONS AT THE ROUTE. `canVisit` filtered the dock and nothing else, so the tile was hidden and the route was wide open — typing the path, following an old link or restoring a tab rendered the screen anyway. RouteGate now wraps every screen in one place, inside the error boundary. It does not redirect. Sending someone to `/` erases what they asked for and reads as a bug: they clicked Music and landed on Home. It says why instead, and the URL stays put so a reload after installing the thing just works. And it says which of the two reasons applies, because they need different screens and send the reader to different places. `not-installed` is a fact about the SERVER — the owner gets a link to the app store. `not-granted` is a fact about the ACCOUNT, and only the owner can change it. Presenting either as the other sends you looking in the wrong place. ROUTES FOLLOW THE SIDECAR. Free, once the above exists: `deniedRoutes` already covers "held but its sidecar is not installed", so an uninstalled feature has no tile AND no screen. The dock, the Permissions list and the routes now agree because they read one answer. NO MORE SEEDING. Downloads/Documents/Music/Videos/Pictures are gone from both places that made them — the member's provisioning and, older and worse, `/ls`, which created folders in somebody's home as a side effect of LOOKING at it. A listing that invents its own contents is a listing you cannot trust, and the platform has no standing to choose a person's folder layout. A new home is empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4b058a6703 |
fix the sign-out reload loop I shipped an hour ago
The 401 handler ended with location.replace('/'), guarded by "unless the path starts
with /signin". There is no /signin route — the sign-in screen IS path="/". So every 401
on the signed-out landing page navigated to the page it was already on, fetched again,
401'd again. A hard refresh loop with no way out of the tab.
The reload was never what fixed anything: useAuth already renders the sign-in screen
when there is no token. It only existed to drop a stale query cache. So it is now the
last thing attempted and bounded three separate ways, any one of which breaks a loop
alone:
1. no token -> return. A 401 while already signed out is expected, not a revocation.
This one alone ends it, because a reloaded document has nothing left to clear.
2. once per document, module flag.
3. once per tab, sessionStorage marker — which also covers a host that re-injects the
token on every load, where clearing storage cannot help and guard 1 never fires.
Anyone stuck in the loop from the previous build: localStorage.clear() in the console.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d3bed0add9 |
the file browser can actually read a member's home, and plans is gone
"This folder is empty" was a lie. The five seeded directories were sitting there and the platform's readdir raised EACCES: a member's home is 700 and owned by them, which is correct for a shell and locks out the file browser, which runs inside the platform process. /ls caught the error and returned an empty listing, so a refusal looked exactly like data. Two doors, two boundaries, and that is the point rather than a compromise. The terminal and the agent RUN AS the member and the kernel is the boundary there. The file browser acts on the member's behalf from inside the platform, which already applies its own containment and is the owner's process on the owner's machine — it can read anything via sudo regardless. Giving it access describes who is doing the work. Done with named POSIX ACLs, because it has to hold in BOTH directions: a file the platform writes must be editable by the member and vice versa. Mode bits cannot say that — whichever party is neither owner nor group lands in "other", and widening "other" opens the home to every account on the box. A shared group fails the same way, since both parties would have to be in it and that puts every member in a group that can read every other member's home. Two named entries plus `d:` defaults grant exactly two users and are inherited by whatever either side creates, whatever their umask. Verified: platform lists the home, member edits a platform-written file, platform edits a member-written file, and a SECOND member is refused on both ls and cat. /ls now distinguishes EACCES from a missing directory. An empty result is data and must never be how a refusal looks. acl joins the core packages in setup.sh — the alternative is an account that provisions and then cannot list its own home. Also: the file browser's own useTasks/useAgents fired /tasks, /agents and both category endpoints on every render, which is where the last four 403s came from — they are the context menu's Run Task and agent submenus, execution-only. Gated. And plans is deleted: router, screen, routes, dock tile, hook, page title and its capability. It read markdown from <repo>/plans, which does not exist. Fresh-install Permissions is now Files alone, with Terminal to come. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6b4fed68fd |
gitea leaves the baseline and becomes an app-store install
It was in both light profiles on the reasoning that it fronts a REMOTE instance and so needs nothing installed locally. That is true and it was beside the point: a baseline process appears in the dock and in the Permissions screen whether or not anyone ever gave it a URL, so a fresh server offered to grant members access to a Gitea that did not exist. "Is Gitea here" had two answers that could disagree. Now it is `existing` mode with a URL and a token, like any other remote service, and the one place that says whether it is here is the install row. No compose template and no `provisioned` mode: Gitea is always something the owner already runs, and offering to spin one up would mean owning its migration, backup and upgrade story. members: 'none' — not because Gitea is single-tenant, it is the most per-user service in the catalogue, but because there is nothing for the INSTALLER to do. The owner's connection carries the instance; each member adds their own access token from /gitea and acts only as themselves upstream. A provisioner would need an admin token and would mint credentials on their behalf, which is more authority than this needs. The catalogue test already pinned "the store offers exactly what light leaves out", so removing it from the profile is what forced the entry to exist. Both light profiles changed together — the mac one carried the same comment and the same gap. Permissions on a fresh install is now Files and Plans. Plans stays because it reads the platform's own shipped markdown from <repo>/plans, not anyone's disk, so it needs nothing installed and exposes nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e393d0f5c2 |
a member's screens render, and the shell stops asking for things it cannot have
Three findings from granting Files to a role and signing in as the member. THE BLANK SCREEN. WorkspaceView returns null until workspace.isLoaded, and isLoaded was the success flag of GET /api/dashboards — which the `dashboards` capability gated. So a member with files granted got a completely blank Files screen and no request to /api/file-browser at all: the panel never mounted. Terminal, Chat and every other workspace screen were the same. /api/dashboards is not a feature. It is the per-user key-value store where every screen keeps its layout, entirely `personal`, every row keyed to the caller. Gating it does not restrict an account, it breaks it — which is the definition of `core` at the top of the registry. Moved there. And the failure mode was wrong independently: `isLoaded` now covers a failed fetch as well as a successful one, with `loadFailed` for the difference, so a screen that cannot remember its layout still renders with defaults instead of showing nothing and explaining nothing. THE STRAY REQUESTS. Six shell-level queries gated on isAuthenticated but not on capability, so a member's first paint fired 403s at /server-settings/settings, /jobs/counts (every three seconds, forever), /chat/models, /plans, /music/now-playing and the chat access policy. Each now checks the capability it needs. JobsIndicator and RescanButton also render nothing without `tasks` and `items` — the header was offering two links to a screen the member cannot open and a button that would 403. THE PERMISSIONS SCREEN. It listed all fourteen app capabilities on a server where none of their sidecars are installed. Offering to grant Photos on a machine with no Immich is not a permission decision. It now shows only what is installed, lists the rest as "nothing installed for these yet" so their absence reads as a fact rather than a bug, and marks confined rows as needing a Linux account. Fails open on a degraded read. Found while checking that: the headscale catalogue entry claimed only the `headscale` capability, but the same sidecar also serves `vpn` — a member enrolling their own device — so vpn was never subtracted. Hence `alsoServes`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2c9d4e55aa |
retry a linux account in place instead of deleting the person
POST /users/:id/provision-linux, and a terminal button on each user row. One
operation covering three needs that were all previously answered by "delete the
account and make it again":
backfill an account created before the feature existed, or while the host was not
set up for it
retry the first attempt failed for something since fixed — the traversable
ancestor chmod being the one everybody hits once
re-key replace authorized_keys with a new public key
Deleting to redo a retryable side effect throws away the password, the dashboards and
everything else keyed to the row.
The provisioning block moves out of create-user into provisionOsAccount, shared by
both entry points for the same reason app-store/members.ts is shaped that way: two
moments, one piece of work.
Found by testing the retry rather than the create: provisionUserDirs re-chmods every
directory including home, and home belongs to the MEMBER after the first successful
run — chmod requires ownership, so it threw EPERM and took every retry down before it
started. Those chmods are now a default for directories being created, not an
assertion about ones that already exist; os-user.ts sets the home's mode through sudo
and is the authority for it.
The route answers 200 with the error in the body, because the interesting cases are
partial: "the account exists and is confined but the keys failed" is not nothing
having happened, and the row shows both halves.
Verified end to end: blocked ancestor reports the chmod and leaves osUser null, the
retry after that chmod succeeds and records the row, and a re-key replaces
authorized_keys without rotating the outbound key.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ea0d2396f7 |
linux accounts use the chosen username, and refuse to take one over
Two changes, and the second is what makes the first safe. The officer_ prefix is gone: a member's account is the username the owner typed, so whoami says who they are and a commit from their checkout is attributed to something recognisable. Measured first — useradd on this host accepts everything validateUsername permits, including dots, hyphens, underscores and uppercase. The prefix was also load-bearing, though, and not for looks. ensureOsUser REUSES an existing account, which is what makes it re-runnable, and that was safe by construction while only we created officer_* names. Unprefixed, adoption becomes the dangerous path: a platform account named root would have found root in passwd, and every runAs for that member would have been a root shell. So adoption now requires the existing account's passwd home to be exactly the home we are about to confine — that is what makes it ours — and any uid below 1000 is refused outright. Verified: root and daemon refused as system accounts, and the owner's own username refused by name with its real home quoted back. Also, the ancestor trap from the first real install. A member's home is under DATA_PATH, which is under the OWNER'S home, and /home/<owner> is 750 on Debian and Ubuntu — so every mode bit on the account tree was right, the directory existed, and the member still could not reach it for want of x four levels up. It surfaced as "ssh-keygen: Could not stat …/.ssh: Permission denied", which points at the wrong thing entirely. firstUntraversableAncestor now walks the chain as the member before anything uses the home, and the error names the directory and the chmod. The dev machine was already 751, and the probe used /tmp, so it never crossed the ancestor that mattered. Worth remembering as a shape of mistake. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0281ca62d2 |
a deleted or blocked account loses its session on the next request
Reported from two browser windows: an account deleted from the dashboard survived a page refresh in the other one. Two independent halves. Server: userMiddleware looked the account up, then read the result as `dbUser?.passwordChangedAt` — so a DELETED account fell through the optional chain and the request proceeded on a token that is still cryptographically valid, for up to the full 30 days. `status` was the same hole from the other direction: signin refuses anything that is not Active, but nothing rechecked it afterwards, so marking someone Blocked did not end the session they already had, which is exactly when you would be doing it. Now the account must exist and be Active on every request. Client: nothing reacted to a 401 at all. onError fed the bug-report form and stopped there, so the window kept rendering off cached React Query data. A 401 now clears every storage key createClient reads and returns to the sign-in screen. /auth/ is exempt because a wrong password is also a 401 and reloading the form would look like a crash. window.officerBearerToken was declared non-optional, which made "there is no token" unspeakable. It has always been one of five sources, any of which may be absent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d513c0e13 |
files, for a member, in their own home
Introduces a fifth capability kind. `files` was `execution` — never grantable, because it meant the OWNER'S filesystem. It is now `confined`: execution-shaped, but the kernel enforces the boundary because the account has its own Linux user, its own home, and no permission above it. The rule that makes `confined` mean something lives in authorize.ts, once: a confined grant is DROPPED for an account with no osUser. So "granted but unconfined" resolves to no access rather than to the owner's home — which is what it would otherwise resolve to, since getOwnerHomeDir ignores the email it is handed whenever HOME_DIR is set. One rule covers the HTTP routes, the websocket doors and the dock, instead of each router remembering. resolveHomeDir(userId) is the new seam and it reads the row rather than the token, for the same reason authorize.ts re-reads role: provisioning a Linux account for an existing member has to take effect on the next request, not in thirty days. The file browser resolves it in middleware and puts it on ctx user, because getRootDir is called from fifteen places in that router. Making it async would have meant editing fifteen call sites, and the cost of missing one is serving the owner's home to a member. Now a handler cannot run without the answer. Two things a real run caught: - /ls seeds Downloads/Documents into the home as the service user, which is EPERM against a 700 home owned by the member — it took the whole listing down. Seeding is now best-effort there and happens at provision time instead, as the member. - .unique() on os_user made db:push ask whether to TRUNCATE users, which is unanswerable non-interactively. uniqueIndex instead, per databases/CLAUDE.md. Verified: a member without a Linux account is refused by name; with one, resolves to their own home and NOT to HOME_DIR; the owner still resolves to HOME_DIR; and every .. escape is refused while an absolute path is rebased under the root. Terminal is still execution — that is the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0fb9a29e64 |
ssh for a member's linux account, both directions
Inbound and outbound are two keys doing two jobs, and treating them as
alternatives breaks the goal:
inbound ~/.ssh/authorized_keys, from an optional public key the owner pastes
on the create form. Their private half stays on their laptop.
outbound ~/.ssh/id_ed25519, generated in their home, never leaves the machine.
"They pasted a key, so skip generating one" is the obvious simplification. Agent
forwarding covers a human in an interactive session, but a platform-spawned agent
has no agent socket to borrow — so an edge checkout it is asked to commit and push
needs a key that lives on the box. The inbound key is therefore optional and the
outbound one is not.
No linux password, ever: useradd sets none, which blocks password login and does
not block key auth. So "real user, reachable over SSH, no password anywhere" is
the resting state, and the platform password stays the platform's business.
Validation is about line count, not key shape. Every line of authorized_keys is a
credential, so a pasted value with a newline would install a SECOND key silently.
Multi-line refused, a private key refused by name, an options prefix refused.
Every write goes through sudo install: the home is 700 and the member's, so the
service user cannot even create .ssh. install sets content, owner and mode in one
step, and content travels as a temp path so nothing quotes a form value into a
shell. ssh-keygen runs AS the member so the private key is never briefly root's.
known_hosts is not seeded — StrictHostKeyChecking accept-new instead. The Gitea
SSH endpoint is not knowable at create time, and the default setting makes a first
connection prompt, which in a non-interactive agent turn is a hang rather than an
error. accept-new still refuses a changed host key.
The generated public key is stored on the row and shown twice: on the after-create
panel and behind a key button on the user's row. It has an errand attached that
nothing else will remind anyone about — it must be added to their Gitea account.
Verified with a real useradd: .ssh 700 and id_ed25519 600 both owned by the member
and usable by them, authorized_keys byte-identical to the paste, no key rotation on
a second run, and a multi-line paste refused with authorized_keys untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5c7ceb2283 |
per-user linux accounts, stage 1: the account and the privilege drop
A member gets a real Linux account whose home is the directory the platform already
provisions for them. Nothing uses it yet — this is the mechanism plus the account,
deliberately with no behaviour change, so the file browser and terminal can be moved
onto something already proven.
Bun.spawn silently ignores uid/gid. Verified on 1.3.10: from uid 1000,
Bun.spawn(['id','-u'], {uid: 65534}) exits 0 and prints 1000. No throw, no warning.
Bun's types don't declare the option so typed code can't reach it by accident, but the
runtime accepts it, and a silently absent isolation boundary is the worst outcome this
feature could have. So privilege drops go through sudo -n setpriv, and a test pins Bun's
behaviour — if it's ever implemented, that test tells us we may simplify.
sudo is required for the drop and not because of the uid: --init-groups fails with
"Operation not permitted" for an unprivileged caller even when reuid'ing to its own
account, because setgroups(2) is root-only. --reset-env is what stops the platform's
environment crossing; verified POSTGRES_URL is unset on the far side and HOME arrives
from the target's passwd entry.
Three bugs that only a real run with a real useradd could find:
- chmod after chown fails forever, because chmod needs ownership. Both orderings fail
unprivileged. Both operations now go through sudo, which is what makes it re-runnable.
- a member could read ANOTHER member's home: provisionUserDirs created at the default
umask (755) and only the account being created got confined. An unlistable parent is
no protection when the child is world-readable and emails are guessable. The skeleton
is now created closed, 711 on the account dir and 700 inside.
- platform/.env was 664 and a member's shell printed JWT_SECRET, which is enough to mint
an owner token and bypass every capability check. Now a boot check that refuses to
start with OFFICER_OS_USERS on while any .env in the project root is group- or
world-readable.
Design, the measured results and the staging plan: docs/per-user-linux-accounts.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
69a31051ac |
the owner can create accounts
POST /api/users plus an Add-account form in Settings > User management. Until now createUser had one call site — bootstrap, gated on an empty user table — so every non-owner account anywhere had been inserted into Postgres by hand. Created accounts are Active. The column defaults to Unverified and signin refuses anything else with a bare UNAUTHORIZED, which is exactly what made the hand-INSERT route look like a wrong password. Also closes a hole found while reading the write path: a second Super Admin was storable. The CHECK constraint pins user 1's role but cannot see other rows, and getOwnerUser() was LIMIT 1 with no ORDER BY, so two holders would have made "who owns this server" a question the query plan answered — and that answer feeds the agent sidecar's identity, vault access and origin scoping. Both write paths now refuse the role and getOwnerUser() orders by id. USER_DIRS and provisionUserDirs move into data-path.ts so the create handler and scripts/provision-user-dirs.ts cannot disagree about what an account's skeleton is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b7184283e0 |
app store: a screen, so this can be clicked instead of curled
/app-store, built to the platform's own conventions: a locked WorkspaceView with two panels, the selection in `?selected=` rather than a channel, and rows that are real links so cmd-click and a pasted URL both work. `?selected=` and not a /app-store/:id detail route, per docs/navigation-audit.md: this is a master list with a live preview, and linking rows to a detail route would make the detail the whole page and destroy the side-by-side. Both panels read the URL independently — the list and the detail cannot disagree if neither is telling the other anything. The install form is generated from the catalogue's fields rather than written per service, which is what lets a sidecar shipping from its own repository present a form nobody here wrote. `existing` is first in `modes` by catalogue rule, so the default selection is "I already have one" — the answer that avoids starting a second copy of something already running. States are distinguished rather than flattened. Blocked is amber and titled "Needs you", not an error: everything worked and it is waiting for a token only a person can mint. Installed-and-enabled but with a dead process shows a warning rather than a tick that lies. And the disable/uninstall copy says plainly that data, configuration and tables are kept either way, because that is the question anyone hesitates over before clicking. The dock tile is CORE, not plugin-derived: the store is how every other feature arrives, so it must never be one of the things that disappears. Verified through the API the screen uses — 14 items, email reporting installed/enabled with its process online, and /app-store present in the capability routes so the tile renders. NOT verified in a browser: no page has been opened, so the rendering itself is reasoned rather than seen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
dd2be8c8b3 |
app store: never assume a service is local, or on its usual port
An instance may be on another host, behind a reverse proxy on 443 under a path prefix, on a tailnet address, or on an arbitrary port because the usual one was taken. All ordinary self-hosted setups, and each one a case where assuming otherwise produces a connection that fails later with no clue why. Two places were sloppy about it. Jellyfin's placeholder read `http://localhost:8096` and Transmission's `http://localhost:9091`, which quietly teach that a service must be local and on its project's default port; both now show remote examples, and the field type says why. And nothing validated what was typed, so a bare hostname or a URL with a token in the query string was stored as-is. The rule: reject only what cannot work, normalise what is merely untidy, have no opinion about the rest. No check that the host is local, that the port matches a default, or that the scheme is https — a tailnet HTTP service is completely normal. Trailing slashes go, because `${url}/api/x` otherwise doubles the separator: accepted by some servers and 404 by others, which is the kind of difference that reproduces on one machine and not another. Query strings go, because that is where a token hides, and it would sit in a column meant for a location. Credentials in the URL are refused for the same reason — outside the encrypted secret, and in every log line that ever prints it. A missing scheme is named rather than called invalid: it is the commonest mistake, because it is what people type into a browser. Verified through the API: `memos.example.com` is blocked with the fix quoted back, and `https://memos.example.com:8443/memos/` installs and stores normalised — remote host, non-standard port, path prefix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ef4efd39b |
app store: ask before repointing a service that is already connected
Found by breaking it. Installing a sidecar that was already configured replaced its connection row with the new install's, silently repointing a working service somewhere else — during testing that took the live Transmission from :9091 to a scratch container on :18092, and the only symptom was that it stopped working. Install now blocks instead of overwriting, naming both URLs and offering the choice. Blocked rather than failed because there is a sensible answer and the user is the only one who has it: keep what is there, or reinstall with `replaceConnection` to change it deliberately. Harmless on a fresh machine; this is entirely for the one with an existing setup. Verified against the live row: an install pointed at a different URL is refused and the original connection is still there afterwards. Also makes "do you already have one?" structural rather than a UI convention. Three tests: anything that can provision must also offer `existing`, `existing` must come first in `modes` since that is the order the prompt uses, and it must ask for a URL. A new entry added later cannot quietly offer only "provision one for me" — which is how someone with a working Immich ends up with a second one and finds out when two libraries disagree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e83585dd42 |
app store: containers follow the sidecar through enable, disable and uninstall
Tier two works end to end. Transmission installed through the API with a real container, verified on
this machine and then removed:
install preflight, provision, connect, schema, assets, process — container up, health-checked,
connection written from what the setup script printed
disable process stopped, then container Exited(0), data intact
enable container back up, then the process
uninstall container gone, install row gone, DATA UNTOUCHED — config, compose file, downloads and
watch directories all still present
Order matters in both directions and it is opposite each way. Enable brings the container up first: a
sidecar that starts before its upstream exists spends its first seconds failing health checks and
logging about a service that is merely not up yet. Disable stops the process first, for the same reason
in reverse.
`down`, never `down -v`, and no `--rmi`: the volumes are the user's data and the images are shared and
expensive to re-pull. Both are deliberate omissions, stated so nobody adds them later as a tidy-up.
Uninstall only brings down containers for `mode: 'provisioned'`. An `existing` install points at a
service the user runs themselves, and `down` there would stop a container Officer never started.
Every compose call tolerates a missing directory rather than failing. Three call sites can legitimately
arrive with nothing there — an `existing` install, a failed install that died before writing the file,
and a resumed uninstall re-running a completed step — and erroring would make a row impossible to
uninstall, which is the one state a user cannot escape.
Adds an `assets` step, before `process`: the dock reads manifests as soon as the install is recorded, so
an icon arriving a moment later shows as broken on the first render. And composeDir is recorded from the
install rather than derived later, because the directory is the user's and they may move it — uninstall
must not guess at a path it is about to run `down` in.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
82e9fdacc2 |
dock: the shell keeps its own items, sidecars contribute theirs
ALL_DOCK_ITEMS was a hardcoded list of everything, so a fresh machine offered Photos, Jellyfin, Transmission and the rest — each leading to a screen reporting itself unavailable — and adding a sidecar meant editing the shell. Neither survives sidecars shipping from their own repositories. Split in two. CORE_DOCK_ITEMS is the baseline that exists on every install: chat, files, terminal, the app's own screens, and Gitea, which is in the light profile because it fronts a remote instance. Everything else is derived from installed sidecars' UI manifests, delivered with /capabilities. Sent with the capability answer rather than fetched separately so the dock has ONE source. Two requests means two moments, and a dock rendered between them shows a tile for something uninstalled or nothing for something installed. Filtered by capability server-side too: a member is not handed the manifest of a feature they cannot use, because "hidden in the client" is the kind of privacy that lasts until someone opens the network tab. Verified live. The owner — who bypasses every permission check — does not bypass this: /photos is absent from routes and present in deniedRoutes because Photos is not installed. Flipping a row's `enabled` makes its tile leave and return with no process touched. Two things fell out. A manifest can declare extraTiles, because CalDAV is one sidecar presenting as Calendar AND Contacts, and collapsing them to keep the model tidy would make the app worse. And DEFAULT_DOCK_PATHS no longer pins /music: useDock drops a path with nothing behind it, so the default dock came up a tile short on any machine where Music was never installed — a default that references an optional feature is how an app looks subtly wrong on a fresh install for no stated reason. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
209e916343 |
app store: publish a sidecar's assets to public/plugins/<id>/
Real PNG icons are coming, so this is the path they arrive by: a sidecar ships its assets beside its own code, and install copies them to public/plugins/<id>/ where one static route serves them. Copied rather than served in place because a sidecar shipping from its own repository has its assets wherever that repository was unpacked, which is not a path the web server can be taught at build time. One predictable destination means the serving rule never has to know how many plugins exist or where any came from. It also makes assets a property of the INSTALL: uninstall removes them, and a plugin nobody installed serves nothing. Needed a new route, and the reason is a trap worth recording. `publicRoutes` in server.tsx is built by globbing ./public at BOOT, so anything copied there afterwards is invisible to it — the first install of a plugin would show a broken image until the server was restarted, and "install it, then restart to see the icon" is not an install. `/plugins/*` resolves per request, like /novnc/* and /vendor/* already do. Unlike those two it answers 404 rather than 500 for a missing file: an unpublished icon is an ordinary state on a fresh machine, and a 500 would put a red line in the log for every dock render. Proven end to end with a real asset: slskd's icon moved from public/slskd.png into the sidecar's own assets/, published against an ALREADY RUNNING server, and fetched at 200 with the right bytes and content-type — 404 before publishing, no restart between. public/plugins/ is gitignored: it holds copies, and the originals live with each sidecar. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9084fabbf6 |
app store: note where a real PNG icon's bytes will have to live
Real icons are coming and the manifest field already exists — slskd uses it. What is not decided is where the bytes come from for a sidecar that ships from its own repository: /slskd.png works only because it sits in the platform's public/, which a marketplace plugin cannot write to. Records the three options and their trade — marketplace URL (loses icons offline), served by us from the sidecar's directory (works offline, needs a route and caching), or a data URI (no fetch, but bloats every manifest) — so the next person meets the question instead of assuming the current path generalises. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
27c1331b3c |
app store: a sidecar carries its own dock tile and routes, and an uninstalled one has neither
Two gaps, both from the same root: the app knew what an account MAY use and not what this server actually HAS. Availability is now subtracted server-side in the /capabilities answer. "Installed" is orthogonal to "permitted" and the owner is subject to it — the owner bypasses every permission check, but a capability they hold unconditionally still means nothing if its sidecar was never installed. Without this the dock on a fresh machine lists Photos, Jellyfin, Transmission and the rest, each leading to a screen that reports itself unavailable. Computed on the server rather than intersected in the client, so the rule lives in one place: the dock already reads `/capabilities`, and making it read a second list and combine them is how a member's dock and an owner's dock drift apart. `unavailable` is returned alongside `deniedRoutes` because the two mean different things to a UI — "not yours" versus "not here yet, install it". A disabled sidecar counts as unavailable: disable stops the process and its container, so the feature genuinely does not work, and leaving its icon would make disable look broken rather than effective. Reading install state failing subtracts NOTHING, matching useCapabilities' deliberate fail-open. Each entry now also carries a UI manifest — name, icon, colour, rootRoute, routes — because a sidecar shipping from its own repository has to be able to say what it looks like. The icon is a NAME rather than an imported component: a manifest has to survive being JSON from marketplace.officer.dev, which a lucide import cannot make. Tests pin the manifests against the capability registry, so a tile cannot appear for a route the server guards differently, and against each other, so two sidecars cannot claim one root route. No backfill, by decision: this is proven on a blank machine first and applied to alpha from scratch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ba71dc1957 |
app store: make it work — pm2, install state, and the routes
Email installs end to end now, which was the point of picking it as tier one: no container, no external
wiring, so the machinery is exercised without the provisioning half.
Verified against the running system, not asserted:
POST /api/app-store/email/install -> {"status":"installed","completed":["preflight","schema","process"]}
row -> email mode=config status=installed enabled=true
pm2 -> officer-email online
second install -> all three steps skipped, process not restarted
disable -> stopped
The server boots with the new router, which is the real test of the capability entry: totality.ts throws
before serve() if a mounted router has none, so booting IS the check passing.
pm2.ts shells out rather than importing pm2 as a library. PM2 is already the supervisor and the
ecosystem file is already the definition of how each process runs; a second thing in charge of that
means two supervisors disagreeing. It also means an owner can undo anything the app store did with a
command they already know. The one fact that matters: `pm2 start <name>` fails for a process PM2 has
never seen, so a first install starts from the ecosystem file with --only, and everything after goes by
name. Callers cannot know which case they are in, so startProcess decides.
Disable stops rather than deletes: a stopped process still shows in `pm2 list`, which is the honest
picture. Deleting would make a disabled sidecar indistinguishable from one never installed.
beginInstall returns the existing row instead of replacing it — that is what makes a retry a resume
rather than a re-provision — and clears lastError on the way in, so a UI never shows a stale failure
beside a working service.
The container half of enable/disable/uninstall is deliberately absent rather than stubbed silently: a
disable that leaves Immich running is a different thing from one that stops it, and the difference is
memory on the user's machine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4b9b98efda |
app store: run a sidecar's setup.sh, streaming it and collecting what it returns
Implements the half of the template contract that faces the platform: answers go in as environment and never as prompts, and results come back as OFFICER_RESULT_<KEY>= lines on stdout. A line protocol rather than JSON because the same stream is the user's live log — it goes to a terminal panel while the install runs. A script that must emit clean JSON cannot also narrate, and one that emits both needs a framing convention anyway. This mirrors the @@officer:progress@@ sentinel the job runner already uses, with the same rule: marker lines are plucked out, everything else passes through. parseResults is pure and tested against the realistic near-misses: a line that MENTIONS the prefix without starting with it, an empty value (Transmission with no RPC auth returns exactly that, and blank is a real answer), a value containing `=` (splitting on every one would truncate a credential), and a prefix with no assignment (a script bug — skipped rather than stored as a blank key). Verified end to end against a real script: environment reaches it, stderr is forwarded (docker compose writes its progress there, so dropping it would hide most of what a user watches), OFFICER_NONINTERACTIVE is set so a script that would block fails loudly instead of hanging behind a web form, and a non-zero exit is reported with the tail. Notes an artifact rather than hiding it: the two streams are pumped concurrently, so the error tail can interleave differently from real time. The live log is correctly ordered; only the summary can read out of order. Serialising the pumps would make a script that writes heavily to one stream block on the other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4a03b9f84e |
app store: the install step machine, resumable and testable without docker
Install spans a container start, a health wait, an upstream API call and a process start. Any can fail, and one of them — a token only a human can mint — is EXPECTED to stop the run. A straight-line function has two bad options there: unwind everything, or leave a half-installed service that neither works nor uninstalls, which is the state users cannot get out of. So each step is named, completion is persisted, and running install again resumes. planSteps is a pure function of (entry, mode) and the effects are injected, which makes ordering, resume, blocking and failure testable with no Docker, Postgres, PM2 or Immich in sight. 15 tests cover exactly the behaviour that only appears when something goes wrong. Two rules are enforced by the plan rather than remembered at call sites: 'existing' never provisions, so pointing at an instance the user already runs cannot start a container; and the members step is omitted entirely for a service with no user concept, so a Transmission install does not report a step that did nothing — which reads as a silent failure to anyone debugging a member's access. `blocked` is a first-class outcome, not an error. For Immich the container is up and healthy and only its own UI can mint a key; calling that a failure would make a normal install look broken and invite the user to tear down a working container. The blocking step is deliberately NOT recorded as complete, so a resume re-runs the step the human just answered. Results feed forward — provision discovers the URL that connect writes down two steps later — over a copy of the caller's values, so a failure halfway cannot rewrite what an earlier attempt achieved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b197c2aeba |
app store: settle the lifecycle — disable stops the container, uninstall keeps the schema
Disable now stops the container as well as the sidecar. There is no reason to leave Immich holding memory while Photos is switched off. For mode 'existing' there is no container of ours, so disable is only the sidecar. Uninstall stops both, removes the containers, and deletes the install row. It does NOT drop the sidecar's tables — pushing back on "maybe db schema too" for the same reason volumes are kept, because it is the same category. Music favourites, the Jellyfin server registry, photos configuration and saved connections are real data, and someone uninstalling Photos is saying "stop running this", not "forget which albums I favourited". Keeping them also makes reinstall a RESTORE: uninstall in June, reinstall in August, and the configuration is still there. Dropping the schema would hand back a blank service that looks subtly broken to someone who remembers setting it up. An unused table costs a row in information_schema and nothing else. Also removes a line left stale by the previous commit, which still said the user chooses disposal at uninstall time. There is no such choice any more, and a doc that describes an option the code does not have is how the option comes back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
936b96a36e |
app store: uninstall removes containers and never data
There is now no uninstall option that deletes data, rather than a careful one that does. A user uninstalling a sidecar is saying "stop running this", which is not the same sentence as "delete my photo library", and for Immich or Jellyfin getting that wrong once is unrecoverable. No confirmation dialog makes it a good default. So: `docker compose down` without `-v`. Containers and networks go; the service directory and everything under it stays exactly as it was. The bind-mount convention already makes this hard to get wrong, which is worth noting because it means the safety is structural rather than a rule someone has to keep following. Data lives on the host inside the service directory, so `-v` — which only removes NAMED volumes — could not delete it even if a future change added the flag back. `mode: 'existing'` has no disposal question at all: we did not create that service, so uninstall removes our sidecar and our rows and touches nothing else. Reclaiming disk becomes its own feature later, with the sizes in front of the user — "Photos is using 340 GB, delete it?" — as a deliberate act rather than a checkbox inside an uninstall flow. Removed two stale `down -v` references that survived the first pass, one in the schema comment and one in the design doc's table. Leftovers like those are how a rule becomes permission again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ee38257856 |
app store: make member provisioning a mechanism, not an install-time loop
The owner installs, but a server may already have members, and a member added next month needs the same work. So the unit is (service × member) reachable from two triggers — install a service, provision existing members; add a member, provision installed services — rather than a loop inside the installer. Only handling the first works on day one and rots. No new table. A member is provisioned exactly when they hold a service_connections row: their own credential, url NULL, inheriting the instance from the owner's. That schema anticipated this before this existed, and a second record of the same fact would only be able to disagree with the first. Three outcomes, declared per catalogue entry so the installer never special-cases a service. `accounts` is fully transparent. `none` is a single-tenant daemon with nothing to do — filtered before the provisioning loop so callers can tell "nothing to do" from "did nothing", which look identical at a call site and matter when someone is asking why a member cannot see a feature. `invite` is not a weaker `accounts`, it is the correct outcome: Vaultwarden derives its encryption key from the master password, so a credential we could mint would mean a vault we could read. Transparent right up to where being transparent would be a defect. The per-service work is an interface implemented beside each sidecar rather than a switch in core — a central function growing a case per service is what would stop any of this shipping from its own repository. Implementations must be idempotent, since both triggers can fire for the same pair and a duplicate account upstream is not ours to undo. Deprovision is optional and defaults to leaving the upstream account alone: deleting an Immich user deletes their photos. Written assuming the vault's multi-user adaptation has landed. Today /api/vault is owner-only by an explicit ownerGate, so a member is refused before Vaultwarden is reached — verified, and out of scope. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
890f57a7a6 |
app store: compose templates and their setup scripts, with two proven end to end
Each provisionable service gets a directory holding a compose template and a setup.sh. Deliberately the shape a sidecar needs once it lives in its own repository: metadata, compose, setup script, schema. The contract (templates/README.md): answers come from the ENVIRONMENT, so the web form fills them in and a person on a VPS is prompted only for what is missing, and only on a TTY — one script for both, not two code paths. Idempotent, writes only inside its own directory, streams progress on stdout (the installer pipes it to a terminal panel), and returns results as OFFICER_RESULT_<KEY>= lines so nothing has to scrape a log. House conventions throughout: relative bind mounts so data sits beside the compose file rather than hiding behind `docker volume inspect`, containers running as the installing user so downloads are not root-owned, loopback-only ports unless the service's whole job is inbound connections, and no external networks — the owner's own composes attach to an `nginx` network that a fresh VPS does not have. Transmission verified end to end on this machine, on non-conflicting ports, then torn down: renders, starts, waits, reports. Its health check accepts 409 because Transmission rejects the first request by design — only-200 would have waited out the full timeout against a working daemon. Re-run produced exactly one container, and files landed owned by the user rather than root. Vaultwarden covers the case where we GENERATE the credential rather than asking for one. An existing token is reused, never rotated, because rotating during a resumed install would lock the owner out of the admin page. The Argon2 hash has its `$` doubled or compose interpolation mangles it. The token is not returned to the platform at all — the vault sidecar proxies the Bitwarden protocol and never needs it, and a secret we do not hold is one we cannot leak. Corrects the design doc, which assumed provisioning always knows the connection. Three shapes: we set the credential, we generate it, or a human must mint it in the service's UI afterwards (Immich, Jellyfin, Memos). The third makes "provisioned and running but not yet connected" a real state rather than a failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7359867f7f |
app store: check the host before writing anything
An install that discovers a missing dependency halfway through has already made a directory, possibly started a container and written a row, and then has to unwind — leaving the user with something that neither works nor uninstalls. A 30ms check first is worth most of that. Verified while writing this: nothing in scripts/ installs Docker, and nothing checks for it. setup-dockers.sh invokes `docker compose` with no preflight, so a fresh host without Docker fails partway through setup with a bare "command not found". Recorded in the design doc rather than fixed here — the intended fix is a setup.sh per sidecar, which is also what a sidecar needs once it ships from its own repository. `docker compose version` is the probe, not `docker --version`: the latter passes with a dead daemon, which is the failure people actually hit. "Not installed" and "daemon unreachable" are reported separately because the remedies differ. Checked per MODE, not per entry. A host without Docker can still install Photos by pointing at an Immich somewhere else; refusing the whole entry is the over-strict check that makes people work around the installer instead of using it. Dropped `requires: 'docker'` from the catalogue type. Needing Docker is exactly "this entry can provision", which `modes` already says, so declaring it twice invites the two to disagree. Derived by needsDocker instead, and a test asserts the derivation matches every entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
654cb10711 |
app store: put provisioned containers under the officer root, not the user's
The install layout a machine should have, seasoned owner or not:
~/officerdev/
platform/ the app
data/ DATA_PATH
dockers/ services the app store provisioned
capabilities/ the file-based item store
One root, everything under it. OFFICER_ROOT derives from DATA_PATH rather than being a second variable
that has to agree with the first.
Deliberately not `~/dockers`, where a seasoned user already keeps their own estate — 47 services on this
machine. That separation buys two things. Containers the app store created are distinguishable from the
user's own structurally, rather than by a naming convention we would have to enforce and they could
break. And we never reason about someone else's compose files: the store does not scan, adopt or modify
anything outside its own directory.
That also simplifies "I already have one of these" — it is answered by the user giving a URL, never by
us finding a directory and guessing whose it is. An earlier draft had the installer adopting existing
directories, which meant reading, and potentially writing over, services Officer did not create.
This development machine predates the convention and derives an ugly-but-correct path, since the project
sits inside ~/dockers/officer.dev. Still isolated, still one root. New installs get the clean shape.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
fc48a572d2 |
app store: the catalogue, the install-state table, and what phase 0 must not foreclose
First slice, on a worktree branch so none of it touches the tree the live server runs from. `sidecar_installs` — server-level, no userId, because a sidecar is one process serving the machine. That is the line that keeps the model coherent for several users: installed is server-level and owner-only, configured is per user in service_connections. A member can use Gitea without being able to install it or point it somewhere else. `installed` and `enabled` are separate because they answer different questions, which is what gives the reversible middle ground: disable stops the process and keeps container, config, schema and data. `completedSteps` makes install resumable rather than merely retryable — the failure mode being designed against is a half-installed service that neither works nor uninstalls. The catalogue is data, not code: no functions, no compile-time coupling, because the same shape has to arrive as JSON from marketplace.officer.dev later. Its test pins it to the real estate — it offers exactly the processes the light profile excludes, names processes that exist, and claims capabilities that exist. That last check earned itself immediately: it caught `vault` (no capability at all — it is EXEMPT because Bitwarden clients carry a Vaultwarden bearer, not a platform JWT) and `notify` (which does have one, where I had written null). Docker templates follow the convention already in use across 47 services in ~/dockers: a directory per service, compose inside, relative bind mounts so data sits beside it, USER_UID/USER_GID as the owner. An existing directory is evidence of an existing install and must be adopted, never overwritten. Records what Phase 0 must not foreclose: a remote marketplace, sidecars moving to their own repositories, and third-party plugins — including the note that catalogue.test.ts pins Phase 0's invariant rather than the design's, since that relationship inverts once sidecars leave this repo. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
977d14922f |
email: migrate the sync position per key, not all-or-nothing
The previous commit migrated only when SQLite was completely empty, on the assumption that a non-empty file is an authoritative one. A real account disproved that within minutes of it landing. The older Gmail backfill had already written SOME keys into SQLite — last_sync_at and the uidvalidity set — and never the imap_lastuid ones. So the file was non-empty and half-migrated at the same time, all-or-nothing skipped the migration, and nine imap_lastuid keys stayed only in Postgres. A missing lastuid makes the next sync refetch that folder from UID 1: on the mailbox this was found on, 18,755 messages and 6.9 GB. Now merged per key with the file always winning a conflict. That keeps the property all-or-nothing was protecting — a restored older emails.db still overrides a newer Postgres row for every key it has, so it cannot be advanced past mail it does not contain — and adds the keys the file never had, which are exactly the ones whose absence is expensive. Verified against the live account: all 22 Postgres keys present afterwards, imap_lastuid:INBOX restored to 208407, and last_sync_at left at the file's older value, so it re-checks a fortnight rather than skipping it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
539aeca7ec |
reverse the fork decision: the pipeline works, my probe did not
I concluded a few commits ago that the serve new /api/session pipeline accepts prompts and never executes them, and kept turns on opencode run. Wrong. Every probe behind that passed an explicit model claude-sonnet-4-6, and THAT model silently does not run on the new surface — no error, no event, no assistant message. Drop the field and the same request completes. One broken variable in every experiment, read as a property of the system. Measured on 1.18.16, both machines upgraded today: delivery steer injects into a running turn (verified, output changed to order), delivery queue runs after it (verified, ONE then TWO, zero errors), and model selection works via POST /model — just not with sonnet. So the fork is reopened and worth taking, targeting the new surface rather than the legacy message path, which generates fine but has neither steer nor queue. Blocked only on why sonnet dies there while working under opencode run. Third time this project has hit the same trap: opencode accepts input it does not honour and says nothing — directory in the body, location.directory that never existed, now model. A probe that changes one thing and sees nothing has not learned the feature is missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0c90216c7a |
email: keep an account's sync position inside its own emails.db
The messages were in emails.db and the position — last_sync_at, and per-folder uidvalidity/lastuid — was a jsonb column on email_accounts in Postgres. Two stores for one fact, with an edge that only shows up when you try to move a mailbox to another machine. The expensive part of an email account is the first sync: hours of IMAP for a large mailbox, which is exactly why "copy emails.db to the new server" is the obvious way to bring one across. With the position in Postgres that silently does not work — the new server's column is empty, !last_sync_at says first sync, and the whole mailbox downloads again on top of the one just restored. The other direction is quieter and worse. Restore an OLDER emails.db while Postgres holds a NEWER position and the sidecar skips every message between the two, permanently, because nothing looks below lastuid again. Re-syncing is slow; skipping mail is data loss nobody notices. Not a new idea — the Gmail path already read SQLite and fell back to Postgres, backfilling so the fallback was taken once. Only the IMAP path had not followed. This extracts that pattern so both use one copy, and unifies the isFirstSync fork in accounts.ts, which is how the two drifted apart to begin with. The file wins over Postgres, always, and only migrates when it holds nothing at all. Topping up a partial position from Postgres would reintroduce precisely the divergence this removes. email_accounts.sync_meta is kept and marked legacy rather than dropped: it is the one-time backfill source for every account created before this, and dropping it would strand any that has not synced since. Nothing writes to it now. 11 tests on the migration, aimed at both expensive failures — migrating when we should not, and failing to migrate an account that predates the change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
41663bc207 |
decide the phase 2 fork: turns stay on opencode run
The serve has a second, newer API surface nobody here had looked at, and it publishes exactly what the parity doc calls impossible under stdin ignore: delivery steer and queue on POST /prompt, an interrupt that does not tear down, and a per-session event stream with an after cursor — the durable-replay machinery officer hand-built for claude, as a primitive. That would have made migrating obvious. It does not execute. A prompt is accepted with an admittedSeq, stored, emits prompt.admitted and prompted, and then never steps. Ruled out separately: the model, the permissions (build is *:allow, no pending requests), the per-request location (the surface is location-scoped via header or a deepObject query, supplied everywhere, no change), and a config gate. The legacy POST /session/id/message?directory= generates fine in 17s, so the serve itself works — only the new pipeline is inert. session.next.* is the tell. And not a version problem, which is the part everything here had backwards: this Mac runs 1.18.11 and alpha runs 1.17.9, measured. The dead pipeline was tested on the NEWER binary. The original "this server runs 1.17.9" meant alpha and was copied to a machine where it was false; corrected in runner.ts and the test. So building against it now would produce code that looks finished and does nothing, which is the failure mode this project keeps rediscovering. One request reopens the question after any upgrade, and the doc names it. Also de-flakes the lifecycle tests: they spawn real processes, and a fixed sleep(750) went red once on a machine busy running these probes. Presence assertions poll now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9118a9f76c |
name the live opencode rows
They were permanently unnamed, and the two halves needed to name them already existed on officer: the sidecar reports its own sessionKey because that is all it has, while the ses_ id arrives separately over opencode:session and is recorded in opencode/state.ts. Nothing joined them. /chat/live joins them now, so no protocol or sidecar change — widening LiveOpenCodeSession would have meant sending the sidecar a fact it told officer in the first place. One list call names every row rather than one transcript load each, and it is skipped when nothing is running or no id has been reported, so an idle Live panel never touches the serve. Verified against a real turn, which also showed the design working as intended: the first poll has no id yet and shows nothing, the next shows title and cwd. That window is real and short, and showing nothing beats showing a key the user has never seen. Worth knowing: opencode titles its own sessions "New session - <ISO timestamp>", so the row is located but not meaningfully named. That is genuinely its title, not a bug here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1892c23aef |
close bucket 0
B8 done, so every defect that made opencode behave wrongly is fixed. Notes what that does not mean: bucket 1 is capability gaps, and the visible ones are downstream of the phase 2 fork, which is still unstarted and still Andre to call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cdee320fed |
stop orphaning opencode turns and serves on restart
B8, both halves. They share index.ts, so they share a commit. In-flight turns: `opencode run` is spawned, not supervised, so pm2 restart officer-opencode left every turn ALIVE — reparented, still spending tokens, still writing files as the agent, with the only reader of its stdout gone. The transcript stopped mid-tool-call, which reads as the agent hanging. stopAllOpenCodeTurns kills them and settles each synchronously, because the caller is about to process.exit and nothing waiting on proc.exited would ever run. Settling writes a reason, so a reload after a restart explains itself instead of trailing off. Turns are stopped BEFORE the connection is destroyed — that write travels over it — and the flush is bounded, since losing the explanation is bad but hanging the restart is worse. Stale serves: the sweep read /proc, so it was a no-op on macOS and orphaned serves piled up, one per unclean exit, each holding a port. Added a pidfile sweep alongside it. A pid we wrote ourselves needs no cwd guard to prove it is ours, which is the part ps cannot answer portably (macOS would need lsof), and a serve started by hand is never in the file. The guard checks command AND subcommand: matching the word serve anywhere in the line would sweep a running turn whose prompt merely mentioned it. Fixtures are real ps output from both machines, not invented. Split into serve-sweep.ts because index.ts spawns a serve at module scope, so a test importing it would start one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
73ccf15a89 |
record that only B8 is left, and what B7 actually was
The bucket-0 table had no status anywhere; it lived in the report docs, which means the list itself still reads as eight open defects. Says B1-B7 are done and where. Also corrects B7 in place: the table describes the spurious cut-off only, and the same default was mis-adopting the session into the wrong harness entirely. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d9a3513cb5 |
stop resume-cursor guessing that a session is claude
B7. `msg.model || DEFAULT_MODEL` declared every session without an explicit model to be claude-code, and the parity doc recorded only the visible half of what that cost. The durable false cut-off is real: endTurnIfAgentIsGone asked the claude sidecar about a key it had never held, was told false, and wrote "the agent went away" into a turn that was running fine. It survives reload, because surviving reload is what that row is for. The same default also handed the session to adoptOrphanedSession as a claude one, which subscribes it to that sidecar bus and pins session.model — so an opencode turn output never arrived, and stopping it called killClaude on a key that sidecar never had. A stop button that silently does nothing. decideResume makes both rules explicit: the server record beats the client claim, and an unknown harness stays unknown — no adoption, no cut-off check, just the replay. Silence is the safe failure when the wrong answer is written durably. DEFAULT_MODEL stays in handleAttach and is now commented as to why: that path reached its sessionId by asking the claude sidecar to resolve a claudeSessionId, so only claude could have answered. First test in api/chat, which had none. websocket.ts has no seam to drive the handler through, so the decision is extracted and tested; the wiring around it is not covered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
be259f813a |
design: sidecars as installable apps
Agreed in conversation, nothing implemented. Light stops being a variant and becomes the baseline — chat, terminal, file browser — and the other fourteen sidecars arrive by the user asking for them from an app store, eventually including sidecars the user did not write. Mostly not a rewrite, for three reasons already true: every API route stays mounted regardless of which sidecars run, officer already spawns nothing, and service_connections already solves the multi-user case. What is new is provisioning, per-sidecar schema, and persisted install state. Docker: officer is the installer, never the owner. Real compose files in the user's own directory, started as him, found again by label. `docker compose down` works, and the containers outlive Officer. Per-sidecar schema is right here specifically because third-party plugins are a real goal, and the dependency graph makes it tractable: measured across 19 schema files, every sidecar depends on auth.ts and nothing else, with no sidecar-to-sidecar edges anywhere. So the plugin contract is "you may reference users.id" — which also makes full uninstall well-defined, since nothing else points at a plugin's tables. service_connections stays core and shared rather than per-service, because it already does the part nobody would get right alone: a NULL url means "inherit the instance", so the owner's row is the instance and members hold only their own credential, making "members never see the instance URL" a property of the schema instead of a filter someone has to remember. Records six open questions rather than settling them, including plugin migrations, ID namespacing for a marketplace, and where plugin-specific config lives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
72f4dcbdb3 |
mark the phase 1 review resolved, so it is not fixed twice
The review was written as a handover; it became a fixed tree instead. Records what landed, including the two leaks that only showed up while fixing it, and leaves the original reasoning untouched so it still reads as the argument it was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
42190f007f |
say what the live opencode rows actually do
Two comments describing behaviour the code does not have. LiveOpenCodeSession had been inserted between LiveClaudeSession and its docblock, so a comment about isGenerating, pendingTasks and the idle GC read as documentation for the OpenCode type — where it is contradicted by the correct comment directly beneath it. Moved below, and it now states that it carries no ses_ id. That absence is the point: /chat/live claimed title and cwd come from the session store "so a turn whose id has not been reported yet shows unnamed". Nothing is looked up, and there is no id here to look one up with. They are null permanently, not until-known. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c252f24f1d |
don't let a superseded turn finish somebody else's session
A replaced turn is killed but dies asynchronously, so its proc.exited fired long after the replacement was registered under the same sessionKey — and then ran the whole completion path against it: emitted "OpenCode exited with code 143", which the sidecar commits to chat_session_events so a false failure became permanent history, then deleted the replacement from `running`. That blinded the new Live panel, made the stop button a no-op and orphaned a process nothing could reach. Mark the handle before killing it, retire it silently, and identity-check the delete — a superseded turn does not own that key any more. Two leaks in the same family, found while fixing it. An early return would not have been enough: both watchdogs call finish, so the armed 10-minute hardTimer would have fired an error at whichever turn held the key by then. And handleLine had no `done` guard, so stdout still draining from the killed process was emitted under the replacement key. Reproduced before fixing. The lifecycle tests need no real opencode — RunnerConfig.bin takes a shell script that sleeps. The control test pins that an ordinary non-zero exit still reports an error, so the guard cannot overreach. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3216040d2f |
unbreak the linux light profile: classify officer-gitea
`ecosystem.light.config.cjs` did not mention officer-gitea, so `defineProfile` threw while the file was being loaded and pm2 could start NOTHING from it — not the app, not the agent, not the terminal. The profile has been dead on arrival since gitea was added to ecosystem.config.cjs and to the mac profile but not to this one. That is the drift check doing its job rather than a flaw in it: the alternative is a light install that silently starts less than it claims. The cost is that adding a sidecar breaks every profile until each one classifies it, which is the trade the file already documents. Included rather than excluded, matching the mac profile's reasoning: this sidecar fronts a REMOTE Gitea whose URL and token live in `service_connections`, so it installs nothing locally. That is the line between it and the excluded sidecars, which supervise a local daemon or container. Both light profiles now load and contain the same six processes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
96fd9e93a2 |
review phase 1, and report the bug the live panel made load-bearing
Phase 1 accepted and phase 2 answered well. One real defect: the supersede path in
runOpenCodeTurn kills a stale turn without marking it, so the dead process late-fires
finish() against the turn that replaced it — committing a false "OpenCode exited" to
chat_session_events, deleting the live handle from `running`, killing the stop button
and orphaning the process.
Predates this pass; reported now because
|
||
|
|
698784ae03 |
start a live document on sidecar bootstrapping
Investigation notes, written as they were gathered rather than after the fact. No code changes. Covers how a sidecar comes into existence: PM2 peer, dials in to /api/sidecar/register, reports an ephemeral port, officer proxies to it. Officer spawns nothing — the only startup problem left is ordering, handled by waiting on a capability rather than failing the first request. The finding worth having: `sidecar/claude/` is TWO processes, and they register as different sidecars. `claude/index.ts` is officer-anthropic-proxy and registers capability `proxy`; `claude/user-instance.ts` is officer-agent and registers capability `claude`. So `isConnected()` — defined as "a sidecar with capability proxy exists" — means the Anthropic proxy is up, not the agent, which is not what the name suggests. It has no callers today, so nothing is misreading it yet. Marks what is unverified and what is still open rather than presenting the lot as settled: the bind-read-release race in getFreePort, whether the `PORT ?? 5000` fallback is reachable, what a partial boot looks like, and whether the sidecar-side boilerplate is worth factoring the way create-proxy.ts factored officer's side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
140288031a |
report phases 0 and 1 back to the spec, and hand over phase 2's answer
Implementation report for whoever wrote opencode-parity.md and opencode-phase0-review.md: what landed, the three places the specs were wrong and how the reproduce-first rule caught each, the three places I deliberately did not follow them, and what is unverified. Phase 2's blocking question is answered in full — the serve takes a per-request `?directory=`, so the coupling is gone in both architectures. The migration is not started; that decision is framed in opencode-serve-path.md and left open. Flags `opencode:list` as the one new capability whose happy path is unproven, and B7 as newly more masked: the B2 fix makes the client send `model` more reliably, which hides the spurious cut-off rather than removing it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e8bd946272 |
show running opencode turns in the live panel
The Live panel asked `claude:list` and nothing else, so a running OpenCode turn was invisible — the panel
claimed to show what the agent is doing and silently omitted half of it.
Adds `opencode:list` / `opencode:sessions` and merges both harnesses in `/chat/live`, asked in parallel,
each failing toward empty so one sidecar being down contributes nothing rather than breaking the panel.
The OpenCode row is deliberately thinner than the Claude one rather than faked into parity:
isGenerating always true — a subprocess exists only while it generates, so there is no "merely open"
pendingTasks always 0 — `opencode run` has no background-task concept; reporting a number would
suggest a capability that does not exist
title / cwd null — the session store is keyed on the `ses_…` id the runner reports, not on
our sessionKey, so an unreported turn shows unnamed rather than guessed
This is the incremental option from docs/opencode-serve-path.md — enumeration without moving turns onto
the serve, so it buys the Live panel with no warm sessions, no SSE loop and no lifetime questions.
NOT verified end to end: no OpenCode turn was running to enumerate, so the verb is wired and typechecked
but has never returned a non-empty list. See COMMS/BLOCKERS.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8b409e8af8 |
finish phase 1: correct the stale comments, and put the ndjson mapping under test
Path names: the serve's cwd is DATA_PATH/opencode_server, not opencode-sidecar. Two comments said
otherwise and would send the next reader to a directory that does not exist.
Version pin: the comment claimed "verified live against 1.17.9" as though that were a property of the
code. It is a property of whichever binary is installed, and this project already runs two — 1.17.9 here,
1.18.11 on the other machine. Says so now, and points at the test as the thing that actually enforces it.
Tests, the first on the OpenCode path. `runner.ts`'s NDJSON → ChatEvent mapping was described as pure and
untested; it was untested but not pure — it lived inside `handleLine` as a closure over `emit`, the
accumulated cost and a reported-session flag, so it could not be called without spawning a binary.
Extracted as `mapRunLine`, genuinely pure: line in, {sessionId, events, costDelta} out. The two concerns
that span lines stay with the caller, because they are not properties of a line — emitting the session id
exactly once, and accumulating cost across steps. Behaviour is unchanged.
11 tests over what the mapping forwards, what it drops and what it must not turn into NaN. The last one
matters: a missing `cost` on a step_finish would otherwise propagate NaN into the turn total.
Phase 1 is complete: dead code deleted (previous commit), comments corrected, tests added.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d7b223127a |
delete the dead serve-turn client, and answer phase 2's blocking question first
Phase 1 item 7, done in the order the parity doc asks: read the dead design into a note, then delete it. `docs/opencode-serve-path.md` records what `event-mapper.ts` and the SSE half of `client.ts` did, and what a rebuild would want back from them — the delta model and tool-state transitions, which are exactly parity Phase 3's token streaming rather than new work. Deleted: `event-mapper.ts` entirely, and `subscribe`/the shared `GET /event` SSE loop, `createSession`, `postMessage`, `abort` from the client, plus `isServerHealthy` from server-manager. All had no callers. `client.ts` goes 200-odd lines to 99. What stays is the REST reads the chat list and transcript use: listSessions, getSession, getMessages, deleteSession, renameSession. While in there, Phase 2's blocking question turned out to be cheap to settle, so it is answered rather than left open. The review asked whether the serve can take a per-request directory, since without one a serve-based turn path would reintroduce the single-directory coupling that shelved this work: POST /session?directory=/tmp/oc-phase2-probe -> directory: "/tmp/oc-phase2-probe" honoured POST /session with directory in the BODY -> directory: "<serve cwd>" ignored It is a query parameter on every /session* route. So the coupling is gone on both architectures and the blocker is cleared. The note does NOT start the migration: which of the three options to take is a product call, and it lays them out rather than presuming one. The first probe put `directory` in the body and appeared to prove the opposite. Recorded in the note, because it is the obvious way to test this and it gives a confident wrong answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |