bba6d854bcd140d93d30479b950f427c518ec8d1
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bba6d854bc |
stop the run at the end of the rewritten sections
A WIP boundary after section 2 so the finished part can be run start to finish on its own, without the untouched sections below acting on the machine. It moves down as each section is worked through and goes away when the walk ends. Also ignores .setup-progress, which the script writes beside itself and is per-machine. The exit message names it, because with it in place a second run skips section 2 and the rewritten part cannot be re-felt from scratch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dbdef23d29 |
install what is missing and keep what is there, per package manager
lib/packages.sh, and section 2 wired to it.
The rule it exists to enforce: `apt-get install <present-package>` is not a
no-op, it upgrades the package if the repository has a newer one. On a machine
somebody already uses that silently moves a version they chose, and a setup
script is the last thing that should do that behind their back. pkg_install
queries the package database first and names only the genuinely absent packages
on the command line — a package already installed is never passed to apt at all.
It also says so out loud, every time, because a provisioning run should not be
opaque about what it is doing to the machine:
:: Core packages — installs what is missing, keeps what you already have
already here: curl ca-certificates gnupg git jq …
to install: btop tmux
Section 2's flat list of 19 is now pkgs_core(), split per package manager rather
than through a canonical-name table with overrides. The names genuinely disagree
(build-essential/base-devel, fd-find/fd) and three of them are not packages
elsewhere at all — apt-transport-https, lsb-release and software-properties-common
are apt concepts that exist to let later steps add the Docker repo and the
fastfetch PPA. A `case $PM` shows what each system actually gets, in one place.
Of those 19, six are load-bearing and the rest are the environment. Only
build-essential reaches beyond itself: it is a meta-package, so on a box with a
pinned gcc it pulls the distribution default alongside. Noted where it is
declared; it is the first thing to move out of core if that ever bites.
apt-get upgrade stays, but as its own announced step — it is the one place that
deliberately moves versions, rather than something that happens as a side effect
of asking for a tool.
DEBIAN_FRONTEND=noninteractive and NEEDRESTART_MODE=a now live inside the
helpers. needrestart has been on by default since Ubuntu 22.04 and stops to ask
which services to restart, which is how an unattended run ends up silently
waiting for a keypress.
dpkg-query on the status field rather than `dpkg -s`, which also succeeds for a
package removed but leaving its config behind — that state would read as present
and never be reinstalled.
Verified against this host's real dpkg database: all 19 report present, and a
mixed list correctly passes only the absent ones through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
41ff8030e9 |
ask what the machine is for, once, in pre-flight
MACHINE_ROLE is homelab, vps or dev, and several steps have a different right answer per role with no way to work it out themselves: whether the address is yours to pin (static IP), whether the box faces the open internet (fail2ban, SSH hardening, UFW), and whether it is allowed to sleep (suspend, logind). Asked in pre-flight rather than at each point of use. The steps that care run from swap through to the firewall, and being asked "is this a VPS?" for the fourth time halfway down a provisioning run is how people start answering without reading. The default offered is guessed from whether this machine's own address is in RFC1918 space, which beats asking whether it is virtualised — a homelab is very often a VM on Proxmox and would be misread as rented — and is the same fact most of the branches turn on anyway. A graphical session means dev; so does macOS. It is only ever a suggestion the user confirms. MACHINE_ROLE in the environment answers it ahead of time for an unattended run, which is why it is declared with :- rather than a plain assignment. The first version wiped the caller's value before ask_machine_role ever saw it; caught by running with MACHINE_ROLE=vps and watching the menu appear anyway. Verified: guesses vps on this host (public IPv4, no DISPLAY, no display manager), env override takes, and a bad value fails with the three valid ones named. Nothing consumes the role yet — the steps get wired as each is worked through. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
34bb8fc22a |
put the superseded setup scripts in setup-old, and repair what the move broke
scripts/setup/ is now what the new installer is being built in — machine-setup/ for the box, officer-setup.sh for the platform on top — and everything being replaced moved to scripts/setup-old/. It still works and is still what to run. Three things the move broke, and what each needed: starship.toml is not an old-setup artifact. os-user-shell.ts reads it at RUNTIME to seed a member's ~/.config/starship.toml when their Linux account is provisioned, and line 125 reads it inside a try whose catch returns "could not read the shell templates" — so account provisioning would have failed outright, not degraded. Moved back to scripts/setup/, which is where it belongs anyway (one file, both audiences) and which leaves the code correct with no edit. package.json's `setup` script pointed at a path that no longer exists. It now points at officer-setup.sh, where the installer is going, rather than at setup-old/ which is temporary. officer-setup.sh was created empty. An empty script exits 0, so `bun setup` would have reported success while doing nothing — worse than the broken path it replaced. It now explains that it is not written yet and exits 1, naming the setup-old script to run meanwhile. Also brought .tmux.conf and ufw-docker-rules.conf in beside machine-setup.sh, which reads both from SCRIPT_DIR and had been silently skipping them since the script was vendored. ssh-keys.zip deliberately stays out: it is key material, and *.zip is ignored. Comments in os-user-claude.ts, app-store/preflight.ts and two docs still name the old scripts/setup/setup.sh path. Left alone on purpose — repointing them at setup-old/ only to repoint them again when officer-setup.sh lands is churn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44141faf0a |
split machine-setup into an entry point and a base library
Structure before the work rather than during it: scripts/machine-setup.sh becomes
scripts/setup/machine-setup/, with the script itself as the entry point and
lib/base.sh holding what every part of it needs.
machine-setup.sh pre-flight and the numbered sections, for now
lib/base.sh shared state, output, the step/resume machine, prompts,
and OS detection
The rule for lib/ is definitions only — nothing there installs, writes or
restarts anything, so sourcing it is safe from anywhere. That is why the ERR
trap stayed in the entry point: a trap is a side effect on whoever sources it.
Behaviour is unchanged. Verified by diffing the moved region against the previous
commit: identical set of functions, and the only differences are added comments,
section banners, fail() reformatted onto three lines, and one new line — a guard
against double-sourcing, which matters because steps will source this directly
once they move out, and a second pass would reset SUMMARY.
The sections are still one 1111-line block below pre-flight; they move into
steps/ as each is worked through. The script also still reads ssh-keys.zip,
.tmux.conf and ufw-docker-rules.conf from SCRIPT_DIR, which is now this
directory, so those three steps warn and skip until the files follow it here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
dda214ffb0 |
detect the operating system before any step runs
The script assumed Ubuntu on x86_64 in every line of it. detect_os() now runs first and fills in OS, OS_NAME, OS_VERSION, PM, ARCH and IS_WSL, so the steps have something to branch on as support for other systems is added. Read from /etc/os-release rather than probing for a binary: a machine can have more than one package manager on PATH, and only os-release can say which distribution this actually is or give a version worth printing. Sourced in a subshell so its NAME, VERSION and ID do not leak in here. ID_LIKE is the fallback, so Pop!_OS, Mint and EndeavourOS resolve without being named. ARCH is normalised to amd64/arm64 in one place because upstream disagrees — Neovim ships aarch64, Go and Docker ship arm64, lazygit ships x86_64 — and several steps hardcode one spelling today. Windows exits with a message pointing at WSL2. WSL itself is detected and warned about rather than refused: it reports as Linux but has no real systemd session, so the suspend, logind and boot-hang steps do nothing there. Everything below pre-flight is still apt and systemd only, so a gate refuses pacman/dnf/brew by name rather than half-building a machine and stopping somewhere unhelpful. Relax that case one entry at a time as each grows a path. Verified on this host (Ubuntu 24.04.4, amd64, apt) and by stubbing uname and os_release for arch, manjaro/arm64, fedora, pop, macos, mingw and riscv64. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
483bb15d8a |
vendor the ubuntu machine provisioning script, verbatim
A byte-for-byte copy of /root/ubuntu-setup/setup-ubuntu.sh, the script that has provisioned every Ubuntu server here. Committed unchanged, before any edit, so that everything the setup-script rework does to it reads as a diff against what actually ran on real machines rather than against a tidied-up version of it. Nothing in the repo calls this yet. It also cannot find three files it reads from its own directory — ssh-keys.zip, .tmux.conf and ufw-docker-rules.conf all live beside the original in /root/ubuntu-setup, and SCRIPT_DIR is scripts/ here, so those steps warn and skip. The original stays where it is and stays authoritative until this one replaces it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
315073cf3b |
move the pm2 ecosystem files into ecosystem-files/ for reference
Temporary, and it breaks things — nothing has been repointed yet:
scripts/setup/setup.sh:59 joins a bare filename to $PROJECT_DIR
scripts/setup/setup_mac_light.sh:60 the same
src/servers/app-store/pm2.ts:23 starts sidecars from 'ecosystem.config.cjs'
src/servers/app-store/catalogue.test.ts:12-13 require('../../../ecosystem…')
ServersView.tsx:207 tells the owner to run pm2 start ecosystem.config.cjs
And one thing that changed silently rather than breaking: ecosystem.profile.cjs:53
pins cwd to __dirname, which was the repo root and is now ecosystem-files/, so the
.env that line exists to find is no longer beside it.
These are here to be read while the setup scripts are reworked, and get deleted
once that lands.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
79da78008a |
turn the field report into something an agent can follow
Expands the communications section from a list of what worked into the actual convention: the directory's lifetime and the rule that anything durable must move to docs/ before the merge; numbering, parity as attribution, non-consecutive numbers; slugs; reply-in-a-new-file and the one case where editing your own is right; referring to commits by sha because three remotes carried the same branch names. Records what a handoff must contain, with the verified/assumed split named as the rule that carried the most weight — a handoff confident about something untested is worse than none, because the reader builds on it. Adds a skeleton to copy. Documents termination as the four attempts it actually took, ending at the only checkable version: the exchange pauses when no open item is actionable by a participant. Adds the third state, deferred-with-a-reason, since a two-state protocol forces an agent to lie in one direction. Notes that a stall must be detectable because the human spotted both before either agent did. Adds a review-discipline section — check the enforcement rather than the description, run it against a real machine, a check never seen failing is not evidence, distrust vacuous passes, expect stacked bugs, distrust "inert today", and look at which way unknown resolves. Adds a failure-mode table to pattern-match against, and the git hygiene that bit us, including merge-verify-then-delete, which I got wrong. Closes with session economics, an ordered list of what to build, and the one thing not to automate: agents may coordinate on what is true and must not decide what is permitted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
015e280e5c |
document how to launch the watcher, since the mechanism is what did not transfer
The owner could not convey this to the second agent, who launched it differently and got something that looked identical and did not work. The script was never the hard part; the mechanism is. States the requirement so it survives a different harness — a detached shell process owned by the agent's harness, which exits when it has something to say, and whose exit re-invokes the agent — and notes that dropping any one of those three breaks it invisibly. Then the four wrong ways, each of which looks correct while running. Backgrounding with nohup or & produces a process that polls correctly, detects the push, exits, and never tells the agent, because the harness is not tracking it; I made that exact mistake and caught it only by re-reading my own command. A model-driven interval is functionally correct and pays a full context re-read per tick to learn nothing — the intuitive design, and the expensive one, which is why it is the first thing to warn a new agent about. A loop that does not exit on detection has no path to the agent at all. And per-tick logging is deferred cost that lands all at once on wake. Also records why 30s polling is free in a shell and ruinous in the model, including the five-minute prompt-cache TTL that makes any model-side wake beyond it pay for a full uncached read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e9d0261e87 |
a field report on two agents working one branch
docs/agent-coordination.md states the objective — several agents on one body of work, coordinating with each other rather than through the human — and was written in theory on 2026-08-07. On 2026-08-11/12 it ran for ten hours with two agents and the owner arbitrating. This is what happened, written as evidence rather than proposal. The load-bearing observation is narrower than "two reviewers are better than one": the person who writes the sentence explaining why something is safe is the worst-placed person to notice the code disagrees with it. One agent wrote "a wrong answer here must not happen by accident" and shipped that accident in the same commit; the other wrote a verification script that could not fail on the first one's machine. Neither was careless. Each was reading their own reasoning back and finding that it agreed with itself. Also records what only running found — an installer piped into the wrong shell, a parent directory created root:root, an ACL mask clamped so the file browser could not read a member's home, a chat cwd the member could not enter, ACL entries surviving a chown — all on first executions, all invisible to review. And what the communications channel got right and the five ways its termination rules broke, and why the repo watcher belongs in a shell loop rather than in the model. Names the identity gap as the first thing to build: both agents commit as the owner, so neither the log nor an agent can say who wrote a line. docs/agent-git-identity.md has called that an idea since 2026-08-10; it stopped being one tonight. Also corrects the deprovision spec's status, which still said "not yet run against a real account" after it had been run and verified clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2eedbcda54 |
bind nginx proxy manager to the tailscale address
It was the only service here publishing on 0.0.0.0, and a published docker port is not behind the firewall: docker writes its DNAT rules straight into the nat table, which ufw's INPUT chain never sees. `ufw default deny incoming` never covered 80/443/81 — ufw-docker-rules.conf on the host exists to patch exactly that, and patching a rule is weaker than never opening the socket. The address is read from `tailscale ip -4` at run time rather than passed in, because the host provisioning has already done `tailscale up` by the time this executes. It is validated against 100.64.0.0/10, the range tailscale and headscale both allocate from. SETUP_NPM_BIND overrides it. With neither, selecting NPM exits instead of falling back to 0.0.0.0 — a fallback would silently undo the point of the change. Two consequences worth knowing. tailscaled becomes a boot-order dependency, so the script warns when it is not enabled at boot; docker's restart policy covers the window but only if the tailnet comes up on its own. And HTTP-01 ACME challenges can no longer reach port 80, so any certificate NPM issues now needs DNS-01. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4e404f17c8 |
split the optional host dependencies out of setup.sh
8 Rust, 9 PulseAudio, 10 cliamp, 13 yt-dlp and 17 the remote desktop move to scripts/setup/setup-sidecars.sh, which nothing invokes — running it is a deliberate act. They are what the optional, sidecar-backed features need on the host, not what the app needs to serve itself. 11 Neovim, 12 the shell extras and 14 the npm globals are gone entirely. The host provisioning already installs node, npm, pm2, Claude Code, Neovim and the shell, and two installers racing for the same binaries is worse than one. That makes node, npm, pm2 and the agent CLIs prerequisites of this script rather than products of it, so the verification block still checks claude and pm2 — section 19 warns and skips rather than failing when pm2 is absent, which would otherwise finish "successfully" with nothing listening. eza is the one casualty: the provisioning installs lazygit, starship, oh-my-zsh and nvim, but not that. Section numbers keep their gaps so the two files read against each other. One line survives from the removed section 14 — the ~/.local/bin PATH export, which section 19's `has pm2` and the agent's claude lookup both still depend on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cec8fbe57e |
the acl check could not fail, because sudo drops DATA_PATH
Review of
|
||
|
|
76cd7c20bf |
sever the ACL as well as the ownership
severMemberTree reassigned the tree and left the access-control entries behind. confineUserTree
grants each member a named ACL on their whole tree — u:<uid>:rwx plus a default: copy — and
chown does not remove them: they are xattrs rather than ownership, and they record the uid
numerically.
Measured before this change: after chown -h -R to the service user, user:<uid>:rwx was still
present on the directory, on its children and in their defaults. The tree read as the
platform's while still granting the freed uid read and write on every byte, so the next account
allocated that number would inherit the previous member's home, keys, credential and container
storage — the hazard this file exists to prevent, reached through a door that find -uid cannot
see.
Now chown then setfacl -R -P -b. Proven on a scratch tree: owner 1001 with five entries naming
1001 becomes owner 1000 with none.
-b rather than removing the member's entries alone, because the service user owns all of it
afterwards and "no ACLs" is cheaper to verify than "no ACL naming one id". -P is already the
default for a recursive setfacl — verified, a symlink out of the tree was not followed — and is
stated for the same reason the chown above carries -h: a member chooses what their symlinks
point at, and this argv should not rest on a traversal default holding.
Found by running assert-uid-free.sh against a real tree; the spec and the checker had the same
blind spot and were corrected in
|
||
|
|
a2f63dc534 |
severing ownership does not sever the ACL, and neither the spec nor the checker said so
Reviewing
|
||
|
|
46799dada8 |
deprovision a member's linux account when the platform account goes
Implements docs/deprovision-os-account.md. Until now deleteUserHandler removed the row, cascaded the
database, and left the entire Linux side running — measured on production on 2026-08-12: working login
shell, healthy postgres container, 454M of data, uid queued for the next useradd to reissue along with
everything still owned by it.
The load-bearing rule from the spec: sever the data from the uid BEFORE releasing the uid, and if
severing fails, do not release. A failed deprovision is not a broken account, it is a trap for whoever
is created next.
Sequence: disable-linger, terminate-user, reap-and-prove, chown -R, userdel (never -r).
reap terminate-user is not a barrier. Production measured a three-hour-old `/bin/zsh -i` surviving
it AND the removal of /run/user/<uid>. So: pkill, bounded wait, pkill -9, bounded wait, and a
final count that must be zero or the account is not released.
chown fixes the uid and subuid halves in one pass — it rewrites every file it walks whatever owned
it. The range is still captured first, because userdel removes the /etc/subuid entry and after
that nothing on the machine remembers what it was. It is returned on every path including the
failures, and logged as the exact assert-uid-free.sh command line.
Two guards the spec did not ask for, both pure and unit-tested:
guardDeletable ensureOsUser's adoption rule backwards. Deletable only if the passwd home is the one
the platform would have confined, and uid >= 1000. Without it `userdel root` is one
bad users.osUser away and nothing else in the sequence would object.
guardMemberTree the tree must resolve to a direct child of DATA_PATH. The email reaches join() from a
database row and the result is the argument to a recursive chown.
chown runs with -h. Measured here that `chown -R` already declines to follow a symlink out of the tree and
re-owns the link itself, but the argv should say so rather than rest on traversal semantics — and
re-owning links is what makes `find -uid` (lstat) a meaningful check afterwards.
destroy exists, has no call site, and is chown-then-delete-as-the-service-user rather than sudo rm -rf, so
a recursive root delete built from a database column does not exist in this codebase.
deleteUserHandler now runs this FIRST and refuses to delete the row if it fails: the row is what remembers
there is anything to clean up, so deleting it first makes a failure unrecoverable through the UI.
NOT YET RUN AGAINST A REAL ACCOUNT. Only the pure guards have tests. The five-step validation is in the
doc; it needs the production host, a shell left open, and a container writing as a non-root user — the two
cases the quiet path passes vacuously.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
f34d7fef70 | merge: a checker for the deprovision spec, and the trap that makes it pass for free | ||
|
|
35715546e2 |
a checker for the deprovision spec, and the trap that makes it pass for free
scripts/assert-uid-free.sh is the verification half of docs/deprovision-os-account.md, written outside the implementation on purpose: a checker the function calls is a restatement of its own beliefs rather than an audit. Nine checks — passwd entry, uid reuse, both subid files, linger, runtime dir, live processes, files owned by the uid, and files owned anywhere in the freed subuid range. Two modes, because the range has to be captured BEFORE deletion. userdel removes the /etc/subuid entry along with the account, and after that there is no way to ask what range it held — so a checker that only runs afterwards silently drops the half most likely to be wrong. Exercised against green while fully provisioned: eight of nine checks fail, exit 1. A checker that has never been seen to fail is not evidence. And the trap worth knowing before anyone trusts a green result: the subuid check passes vacuously on most accounts. Files get a mapped owner only when a process inside a container runs as a NON-root user; an image whose files are root-owned maps to the member's own uid and leaves the range empty. Measured on green after a night of real use — claude installed, an image pulled, transcripts written — the range check found zero files and passed without testing anything. The spec now says how to build a specimen that actually exercises it, and to watch the check fail on that tree before trusting it to pass on a cleaned one. Docs and a script only; no behaviour change. On a branch, for whoever merges it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f5f509a99d |
scripts/setup is the initial install, nothing else
Two of the eight did not belong. cleanup-desktop.sh is the teardown — the inverse of an install, not part of one. provision-user-dirs.ts runs per account at invite time, on a machine that is already set up. Both are back at the top level, with their `../` derivations and usage strings put back. What is left is what a fresh machine runs once: the two installers (setup.sh, setup_mac_light.sh), the two things setup.sh calls (setup-dockers.sh, setup-desktop.sh), and the two files they deploy — starship.toml, copied to ~/.config, and officer-set-display.sh, which setup-desktop.sh installs to ~/.local/bin as a login-time mode setter. The last one is not a setup script and does not read like one; it is here because it is install payload, same as the toml, and setup-desktop.sh loads it by `$(dirname $0)`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9c353f5f0d |
move host setup into scripts/setup/
scripts/ was holding two unrelated kinds of thing: install-this-machine, and run-this-occasionally. The eight installers now live in scripts/setup/; what stays at the top level is the build steps (gen-index, prebuild, build/) and the two maintenance scripts (reindex-music, rebuild-soulseek-tree). The move is not just a rename. Three of these derive the repo root from their own location: setup.sh:51 PROJECT_DIR="$(dirname "$SCRIPT_DIR")" setup_mac_light.sh:51 same cleanup-desktop.sh:134 ENV_FILE="$(dirname "$0")/../.env" Left alone, all three would now resolve to scripts/ — and nothing downstream complains. PROJECT_DIR is where .env is written, where `bun install`, `gen:index` and `db:push` run, and what pm2 is pointed at, so a fresh install would have quietly provisioned scripts/ and reported success. cleanup-desktop.sh fails the other way: it would find no .env, print "No .env — skipping", and leave the real VNC_PASSWORD in the real file. All three are now `../..` with a comment saying why the level matters. provision-user-dirs.ts imports data-path.ts relatively; that one tsgo caught. Also disambiguated `setup.sh` where it had become two files. app-store/templates/<name>/setup.sh is a per-sidecar installer with its own contract, and preflight.ts + docs/sidecar-app-store.md discussed both in the same paragraph. The host one is now spelled with its full path at those sites. Verified: bash -n on all six shell scripts, tsgo clean, os-user tests pass, both derivations resolve to the repo root, starship.toml still resolves from os-user-shell.ts, and provision-user-dirs.ts runs under DRY_RUN. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
37adc65a12 |
delete the spent one-shot scripts
Eight scripts in scripts/ that nothing references and that mostly can no longer run. Kept in history;
none of them is recoverable knowledge that isn't already in the code they migrated to.
Three could not run at all against the current database:
migrate-items-to-files.ts SELECT * FROM tasks — that table was dropped when items became files
reset-user-data.ts deletes chat_sessions, chat_groups, projects; none exists. It has no
transaction, so it would wipe user_settings, user_state,
user_integrations and dock_configs and THEN throw. A half-wiped account
is worse than no script. It also misses chat_session_events, which is
where chat state actually lives now.
add-email-dock-user2.ts one-time, hardcoded to user 2, seeds a dock containing /projects
The rest are spent migrations whose destination is now the only implementation:
migrate-auth-to-pg.ts JSON -> Postgres, 2026-02
migrate-pg-to-files.ts Postgres -> JSON, the other leg of the same abandoned round trip
migrate-server-settings-to-pg.ts 2026-02
migrate-emails-to-sqlite.ts backfill into the email sidecar's store, 2026-07-31
seed-imap-uids.ts the sidecar writes imap_lastuid/imap_uidvalidity itself now
(sidecar/email/gmail-api.ts:533-535)
Kept, and why, since "unreferenced" was not the test: rebuild-soulseek-tree.ts is reusable by
construction — it runs the same buildTree the sidecar's ingest runs, so it answers any future change
in tree shape. reindex-music.ts is named in sidecar/music/index.ts:447. provision-user-dirs.ts shares
USER_DIRS with data-path.ts. cleanup-desktop.sh and officer-set-display.sh are called by
setup-desktop.sh.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a6acfea9a6 | carry the two threads todo.md was missing | ||
|
|
a730fc0fe0 | keep the three open threads the comms channel was holding | ||
|
|
dcaee2fc95 |
keep the three open threads the COMMS channel was holding
The channel was deleted when per-user Claude merged, which was right — it was conversation, not documentation. Three things in it were neither: found while proving the feature worked, understood, and unfinished. The web terminal renders a long URL unreadably. OSC 52 is fixed so "press c to copy" works, which is the path a user is meant to take; the rendering itself is not diagnosed. It matters because first-run login is every member's first five minutes, and the workaround was running claude under tmux on the server and reassembling the URL from a captured pane. Agent sessions do not survive a restart with their identity intact. That one property is behind three symptoms — the crash blast radius, the restart sweep having to skip sessions with no recorded userId, and the stuck "generating" spinner — and documenting them separately invites three separate fixes for one cause. And the ProcessTransport rejection is survivable but still unexplained. Recorded with the log markers that distinguish "the backstop is working" from "it stopped working", since the next occurrence is now evidence in a live process rather than a corpse. On a branch rather than straight onto master, docs-only, for whoever merges it next. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7040536f1f |
merge sidecar-app-store: per-user Claude
A member's agent turn now runs as their own Linux account, with their own claude install, their own ~/.claude credential, their own transcripts and sessions that record whose they are. Verified end to end on the production host: uid 1001, nine environment variables, zero ANTHROPIC_*, zero POSTGRES_URL, zero JWT_SECRET. The two owner-only refusals that held chat closed to members — the wholesale isSuperAdmin middleware in api/chat/chat.ts and the socket's 403 in server.tsx — are gone, removed together once the turn ran under runAs. Also carries a live credential fix that predates this work: mcp-host.json held the owner's 30-day JWT at 0644 inside a 755 directory on a host where every role has a shell. Now 0600 plus an explicit chmod, since writeFileSync's mode is ignored on an existing file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c73ffed806 |
close the comms channel, keep what was still open
The sidecar-app-store channel ran one night, from per-user Linux accounts to a member's first agent turn, and is deleted now the work has landed. A spent channel left in place gets read as current, which is worse than none. Three things lived only in those docs and move to TODO.md rather than disappearing: deprovisionOsAccount (observed on production — a deleted member kept a shell, a running container and 454M of data, with their uid free to reissue), the terminal replaying query sequences as keystrokes, and agent sessions not being durable, which is one missing property behind three symptoms. The deprovision spec itself already lives in docs/. CLAUDE.md's section is rewritten from "here is the current channel" to how to run one, since the answer to "which channels exist" is now none. What is worth keeping is the protocol that emerged: numbered alternating files, parity as the author, a reply even when there is nothing to say, and termination on a checkable condition rather than on someone deciding it feels finished. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ec1997fd0e |
a chat with no chosen directory runs in the caller's own home
The default was DATA_PATH/<email>/general_chat_sessions, a dedicated directory so /chat sessions formed their own Claude project group instead of cluttering the home. It is a sibling of the home, and confineUserTree makes every sibling the platform's at 0700 because the others are attachments and email_accounts. So it was unreachable for a member: the first live member turn started there and every Bash call failed on its own working directory before doing anything. A per-member copy inside each home fixed the symptom and left two rules to remember. The owner chose one rule instead — the account's own home, whoever they are — and accepted the trade knowingly: /chat sessions now share a project group with anything else run from that home, which was the reason the dedicated directory existed. Removed rather than left dangling: getGeneralChatSessionsCwd, ensureGeneralChatSessionsCwd, ensureMemberChatCwd, and general_chat_sessions from USER_DIRS so new accounts stop getting it. Existing directories are untouched and their transcripts stay where they are — Claude groups by cwd, so the owner's old /chat history remains under its own project slug rather than moving. The UI labels move with it: the default group now reads "home" rather than naming a directory that no longer has a role. ChatIdentity keeps carrying both email and home. The pairing was justified in the comment by general_chat_sessions being email-derived, which is now gone — but the distinction it encodes is real (the email says who, the home says where), so the comment explains that instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6c84c74c91 |
give a member a chat cwd they can actually enter
host captured the first live member turn: uid 1001, nine env vars, zero ANTHROPIC_*, zero POSTGRES_URL, zero JWT_SECRET. The privilege drop and the allowlist both held. One defect. The default chat cwd was DATA_PATH/<email>/general_chat_sessions — a sibling of the member's home, which confineUserTree deliberately makes the platform's at 0700 because the other siblings are attachments and email_accounts. So the turn ran in a directory the member cannot enter, and every Bash call failed on its own cwd. The agent reported its shell as broken, which was true. A member's default is now ~member/general_chat_sessions, created as them through runAs. mkdir -p, so it is idempotent per turn and needs no reprovision. The owner's path does not change, and the sibling stays 0700 — loosening it would trade a broken shell for an open directory holding attachments and mail. 29 said this path is email-derived and therefore stays email-derived. True, and it did not follow that it is usable: an email-derived path under DATA_PATH is precisely the set a member is locked out of. Splitting identity from filesystem path was right; assuming the identity side was inert was not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
661b3d761f |
42: the first member turn ran clean, and its cwd is unreachable by the member
The privilege drop and the allowlist both held in production. Captured from /proc during the first real member chat turn: uid 1001, parent sudo, HOME and CLAUDE_CONFIG_DIR both inside the member's home, and exactly nine environment variables — the allowlist plus what setpriv supplies. Zero ANTHROPIC_*, POSTGRES_URL or JWT_SECRET. The defect is the cwd. websocket.ts:134 defaults a chat turn to the email-derived general_chat_sessions, and confineUserTree makes every sibling of home the platform's at 0700. So the member's turn starts in a directory it cannot enter — verified, cd fails — and every Bash call in that turn dies instantly, which is what the owner saw as "the shell is unusable". 29 reasoned that those paths stay email-derived because they live under DATA_PATH rather than a home. That is true and it does not follow that they are usable: platform-owned by design means a member's turn can never run there. Fix is a member-owned default, and I would put general_chat_sessions inside their home rather than repointing at the home itself, since it keeps the existing shape for both parties and does not change the owner's path at all. The sibling's 0700 should not be loosened — it holds attachments and email_accounts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2a8f0049a3 |
handle OSC 52, so "press c to copy" reaches the user's clipboard
xterm.js does not handle OSC 52 unless something registers for it, and nothing did. A program offering "press c to copy" emitted the sequence and it vanished, so its confirmation was true about having sent it and false about anything arriving. Found while signing a member into Claude Code on the production host: its first-run login prints an OAuth URL too long to read off a wrapped pane and offers to copy it, "(Copied!)" appeared, and the clipboard was untouched. The URL had to be recovered by running claude under tmux on the server and reassembling it from the captured pane — which is not a thing a member can be asked to do, and first-run login is every new member's first five minutes. Writes only. A lone `?` in the data position is a read request — a program asking the terminal to hand over whatever the user has copied — and it is deliberately not answered: a shell should not be able to exfiltrate the clipboard of the person watching it. The clipboard API needs a secure context and generally a user gesture; the keypress that caused the sequence is that gesture. A refusal is swallowed rather than thrown, since a copy that does not land is the status quo rather than a reason to break the pane. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
532ad15ac1 |
drop the import the gate left behind
|
||
|
|
c59df4f866 |
let members use chat
The owner authorized this explicitly. Two refusals removed together, because
they were always one guard in two places: the wholesale isSuperAdmin middleware
in api/chat/chat.ts, and the chat socket's 403 in server.tsx.
They were right for the day they stood. A turn spawned claude as the OWNER and
every transcript path resolved through the owner's home, so a granted member
would have read the owner's sessions and run an agent as them.
What replaced them, rather than what deleted them:
the turn runs as the member spawnClaudeAsMember through sudo setpriv,
proven against a real account by reading file
ownership rather than trusting the process
the credential is theirs --reset-env plus an allowlist, so the owner's
proxy variables cannot cross
the transcripts are theirs ChatIdentity carries a home from resolveHomeDir
and claude-sessions cannot invent one
the sessions are theirs every session records its owner and all six
sidecar commands refuse a mismatch
Also adds the precondition host asked for in 10: a member whose claude is not
signed in gets the instruction rather than a turn that dies on an auth error and
reads as a broken agent. Not installed and not signed in are separate messages
because they need different actions.
registry.ts and registry.test.ts now describe chat as confined in fact rather
than ahead of its implementation. The comments at both former guards say what
had to exist first, and that a revert should go back to a refusal rather than to
a narrower one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ba64a412f2 |
40: nothing on the list is actionable by either agent
Read 39, nothing to fix. Refining the terminator now that tonight has tested it: "ends when the list is empty" is too strong, because the list will not be empty for days and yet neither agent has an item to act on. The condition that actually terminates is no item being actionable by a participant — everything left is the owner's or deliberately deferred with a stated reason. That state is reached, so this is where it stops, on a checkable condition rather than on either side judging itself done. Two things for tomorrow's protocol design: a stalled loop must be detectable, because open items plus no recent doc is watchable and silence is not; and "deferred with a reason" needs to be a first-class state distinct from open and done, since three times tonight the honest answer was "mine, and not now". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
53f36f07e4 |
adopt host's terminator: the list, not the guess
38 proved spawnClaudeAsMember against a real member account — the privilege drop lands and the SDK spawn survives it. Nothing to fix. Adopting host's correction to NO REPLY NEEDED: an exchange ends when the open list is empty, not when the sender thinks it is. Mine let either side close a thread with work still in it, which is the failure the alternation exists to prevent — silence and "I think we're done" read identically. deprovisionOsAccount stays mine, unblocked, and still not written: it is the function whose failure hands one member another's home, keys and container storage, and this is the end of the longest session either of us has had. Not blocked and not tonight are different statements and the list should carry both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0947d31330 |
38: the privilege drop is proven on a real member account
Ran the live test against green with both fixes in: 3 pass, 0 fail. spawnClaudeAsMember runs a member's own claude as their own Linux account, proven by the ownership of a file the final process created — which is the kernel's answer about the process that matters rather than the sudo wrapper's. That was the last thing that could have changed the design, and it did not. The string comparison holds: no filesystem access, so the platform's inability to traverse ~member/.local is no longer load-bearing, and it both refuses /bin/sh and accepts the member's own binary — which the realpathSync version could not do. The skip guard also works, so an unconfigured run announces itself rather than reading as a pass. Open items listed as state rather than as a judgement about whether a reply is needed, per the owner's point that a terminator based on the sender's guess can end an exchange while work remains. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c5522700ff |
compare the binary as a string, and prove the uid by file ownership
host ran the live test against green. Two results. SETPRIV WORKS. The privilege drop lands on the member and the SDK spawn survives it, so the design does not change shape and everything layered on the hook stays. That was the last question that could have moved the architecture. THE BINARY CHECK REFUSED A BYTE-IDENTICAL PATH. sameFile used realpathSync, which has to stat inside a 700 home the platform is `other` to, so it threw EACCES and the catch turned that into "not their binary" — every member turn refused forever, the moment the gates moved. Failing closed was the right direction and it made the feature impossible rather than unsafe. Now resolve(command) !== claudeBinIn(run.home). Both operands are computed by the platform from the same function, so string equality establishes exactly what the check is for and needs no access to their home. sameFile and its tests are deleted: a helper kept for a case that cannot arise is a trap for the next reader. The realpath version was defending against an upstream that normalises paths, and there is no such upstream — the platform controls both ends. The fact that decides this is that the check runs in the PLATFORM process, not the member's, and neither of us stated it until EACCES did. THE TEST WATCHED THE WRONG PROCESS. child.pid is sudo, whose real uid is legitimately the service user's until it execs down through setpriv, so asserting on it fails on a working drop. Split: --version through the hook proves their binary ran, and a second spawn creates a file in their home whose owner the test stats. Nothing self-reports, and a process cannot forge the uid that owns a file it created. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
85249a2756 |
36: the privilege drop works, and the binary check refuses every member turn
Ran the live test against green. Two results pointing opposite ways. setpriv survives. Verified independently by spawning the same argv and reading what the command printed: id -u = 1001. Nothing layered on the hook has to move. But the test asserts against /proc/<child.pid>/status where child.pid is sudo, whose real uid is legitimately 1000 until it execs down to setpriv, so it fails on a working drop. The intent — don't let the child self-report — is right; the fix is to observe the final process by having the child create a file and stat its owner, which a process cannot forge. The binary check refuses a byte-identical path. sameFile calls realpathSync, which throws EACCES for the service user because .local is 700 and the platform is "other", and the catch turns that into false. Every member turn would be refused the moment the gates move. That one is mine. In 16 I argued for leaving .local closed to the platform and reasoned about the file browser, without considering that spawn-as-member runs IN the platform process and must stat a path inside it. Recommended comparing resolve(command) to claudeBinIn(run.home) instead: both operands are platform-computed by the same function, so string equality establishes exactly what the check is for with no filesystem access. Granting traverse instead would need x on four directories, not one, and reverses 16 across the tree. All state restored and verified: .local back to 700, no residual ACL entries on any directory I touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9bab67aca5 |
test the privilege drop without lifting a gate
host caught a circularity I had written twice: do not lift the chat gates until a member turn has been watched running, but a member turn goes through chat and chat refuses non-owners. With the gates up there is nothing to watch; with them down the thing we wanted proven has already shipped. spawn-as-member.live.test.ts calls spawnClaudeAsMember directly against a real provisioned account — no gate, no chat, no SDK. The child's uid is read from /proc/<pid>/status, so it is the kernel's answer rather than anything the child chose to say, and it asserts >=1000 and not this process's uid: a failed privilege drop cannot pass by running as the service user. It also asserts the binary exited 0 having printed a version, which proves their install ran rather than merely being spawned, plus a negative that /bin/sh through the same hook throws. Opt-in via OFFICER_TEST_MEMBER and OFFICER_TEST_MEMBER_HOME, because it needs a provisioned member with claude installed — which exists on the production host and on no developer machine. A run without them skips loudly rather than reporting an empty file as a pass. Also adopted host's NO REPLY NEEDED terminator: "reply to everything" had no exit condition and cost the owner two agents being polite at each other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bae3ebf7ba |
34: the alternation needs a stop condition, and the unblock order is circular
33 is right that silence reads as a crashed agent, but "always reply" has no exit: each nothing-to-report obligates another, and every round costs the owner tokens for two agents to be polite at each other. Proposed an explicit NO REPLY NEEDED terminator, which cannot be confused with a crash and which either side can break by writing again. More importantly, 33's unblock order puts the gates coming off BEFORE the first member turn, while 19, 20, 22 and 27 all say the gates must not move until a member turn has been watched running. Both cannot hold: a member turn goes through chat, chat refuses non-owners, so with the gates up there is nothing to watch and with them down the thing we wanted proven first has already shipped. Two resolutions, and the better one is to exercise spawnClaudeAsMember directly against green's real account — asserting the process runs as uid 1001 with their HOME — which answers the only remaining question that can change the design, without a gate being involved. setpriv breaking the SDK transport should not first appear in a live chat. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e7346c8790 |
reply even when there is nothing to say
32 closed the last reviewable item, so I had nothing to report and reported nothing — which left host waiting on a reply that was never coming. The alternation is the protocol: a turn with no content is still a turn, and silence is indistinguishable from a crashed agent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
584074c845 |
32: all three callers fixed, nothing left passing an email
Verified by grepping every call site rather than only the three named: pipeline-executor.ts:499 and :591 and deliver.ts:37 all pass getOwnerHomeDir(email) now, and no caller anywhere passes an identity where a path is expected. Gates unchanged, 97 tests, 259 assertions. The "@param home — NOT an email" comment is the right residue: the compiler cannot distinguish the two strings and never will, so the warning has to live where a fourth caller would read it. Closes everything reviewable without a live member turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1575df3f78 |
stop building cwds out of an email address
host found a live regression from |
||
|
|
2bb8128619 |
30: resolveBaseCwd's outside callers still pass an email
The history layer itself checks out — claudeHome gone, ChatIdentity carries both halves, chatIdentity throws rather than falling back, identity resolved before cwd, and the opencode path keeping getOwnerHomeDir is correct and documented. But resolveBaseCwd's first parameter changed meaning from email to home, and three callers outside the commit still pass an email: pipeline-executor.ts:499 and :591, and agent-handoff/deliver.ts:36. Both parameters are string, so tsgo had nothing to say — exactly the wrong-but-well-typed case flagged as uncertainty (2). Before, the function resolved its own root via getOwnerHomeDir(email) and passing an email was correct. Now the argument IS the home, so any task step or handoff with a tilde, a relative cwd, or no cwd gets a relative path built from an email address, resolved against the platform process's working directory — the repo. Absolute paths still work, which will make it look intermittent. Live tonight on the owner's own features, not a member issue. Fix is to pass getOwnerHomeDir(email) at those three sites, the way agent-runner.ts now does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
95951fbe5a |
resolve transcripts and cwd against the caller's home, not the owner's
The history layer, and the last change that could be made without a live member.
claude-sessions.ts had `claudeHome = process.env.HOME_DIR ?? join(DATA_PATH,
email, 'home')`, which discards its argument whenever HOME_DIR is set — always,
on a real install. Every transcript read therefore resolved to the OWNER'S
~/.claude no matter who asked, and the comment above it asserted "single-user
platform" as though that were a property rather than an assumption. A member
reaching these functions would have been handed the owner's conversation list.
Now every read takes a ChatIdentity {email, home} with the home resolved from
resolveHomeDir(userId), and this file has no way to invent one. Both fields
travel together because they are genuinely different: general_chat_sessions
lives under DATA_PATH/<email>, not under a home. Collapsing them would be the
same class of mistake as undefined meaning "the owner".
websocket.ts's resolveCwd takes a home, so `~` expands against the caller's own.
Identity is resolved BEFORE the cwd — expanding `~` before knowing whose home it
is would be exactly the bug being removed — which also let a duplicate
resolveTurnIdentity call from
|
||
|
|
519109a342 |
28: agreed, with one correction — the backstop is not waiting on a reprovision
Doc only from 27, nothing to review. One correction to the handoff table: 21 is listed as waiting on the owner's reprovision and it is not. officer-agent was restarted at 00:30:29, the handler is loaded, it lives in the process ecosystem.config.cjs actually starts, and Bun 1.3.9 honours it. What it still needs is a rejection to fire, which is a different event. So the reprovision verifies 15 and 17 only. Marker for whoever looks: fatal "Bun v1.3" banners must stay at 4 and "UNHANDLED REJECTION" lines should start appearing instead. A fifth banner means the backstop did not take. Machine state for tomorrow: green provisioned on uid 1001 with claude 2.1.228 running as the member, rootless Docker up and isolated, file browser working, nobody signed in, both gates up, no stale accounts or orphaned uids, owner's containers untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a3431abeac |
stop before the history layer
Nothing to fix in 26 — the three-way identity is verified. Not starting the history layer: it is a ~10-signature refactor of how transcripts resolve, at the end of a long session, in the path whose failure mode is a member reading the owner's conversations. That is the shape host talked me out of earlier tonight, and the same argument applies whether or not I am the one making it. Tomorrow, after deprovisionOsAccount. Everything mechanical for a member turn is done and inert: provisioning, the login probe, agent-status, the privilege drop, the SDK wiring, session ownership, the six scoped commands, member populated, three-way turn identity, the rejection backstop. Both gates up, member unreachable in production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d96083c20 |
26: the three-way identity holds
Verified by reading the enforcement rather than the description. A failed resolveHomeDir and a null os_user both refuse now, isOwner is a positive branch, and the caller refuses before spawning anything and clears isGenerating. member is identity.kind === 'member' ? run : undefined, so undefined is reachable only from a positively established owner — which was the property worth having. Both refusal reasons are member-facing sentences that leak no paths. Gates unchanged, 84 tests pass here. Nothing further from me on this one. What remains needs the owner or a live member: the history layer, a member signing in, the first member turn, and the gates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6aeb304f56 |
never spell "I don't know whose turn this is" as "the owner"
host caught that resolveMemberRun failed open. Returning undefined means "run as the server owner" downstream — their binary, their ~/.claude credential, their HOME, their MCP config carrying OFFICER_AUTH_TOKEN — and three different inputs produced it: the caller being the owner, resolveHomeDir failing, and a member whose osUser is null. The last two mean "could not determine", and answering them with the owner's identity is the single thing this feature exists to prevent. 23's own comment said the caller must not fall back to the owner. The code did exactly that. The prose was right. Now a discriminated TurnIdentity: owner, member, or refuse-with-a-reason. The call site ends the turn on refuse instead of spawning. The owner's identity is reachable only by positively establishing isOwner, never by failing to establish anything else — resolveHomeDir already reported it as a positive fact and the funnel through undefined was the only thing discarding it. The null-osUser case is not hypothetical: provisionOsAccount is non-fatal at every stage and records the account either way, as its own source says. Tonight provisioning failed three separate ways on a real member and the account survived each time. No test yet, and the reason is in COMMS rather than hidden: it needs database fakes this repo has no pattern for, and inventing one at 01:00 to cover four branches is how the next defect gets written. The union is exhaustive, so tsgo catches a missing case — not the same thing, not nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
92e014c19f |
24: resolveMemberRun answers "I don't know" with the owner's identity
The plumbing is right where it matters: userId comes from ws.data, the authenticated socket, never the client message. Gates up, tests pass, path inert. The resolution is not. Three inputs collapse to undefined, and undefined means "run as the server owner": the caller IS the owner (correct), resolveHomeDir FAILED, and the account has no os_user. The last two are "I could not determine whose this is", and they are answered with the owner's binary, the owner's ~/.claude credential, the owner's HOME, and — since mcp-config branches on the same field — the owner's MCP config carrying OFFICER_AUTH_TOKEN. 23's own text says the caller must not fall back to running as the owner, and names a wrong answer here as the one thing that must not happen by accident. The code does exactly that. The no-os_user case is not hypothetical. provisionOsAccount is non-fatal at every stage and provision-os.ts records the account either way; provisioning failed three separate ways on a real member tonight while the row continued to exist. Such a member, once the gates lift, does not get an error — they get the owner's agent. Suggested a discriminated result — owner | member | refuse — so that the owner's identity can only be reached by positively establishing it, never by failing to establish anything else. resolveHomeDir already returns isOwner as a positive fact; only the funnel through undefined throws it away. Of everything tonight this is the one I would least want to discover after the gates moved, and I would fix it before the history layer: that one is a correctness bug when it lands wrong, this is a credential boundary that fails silently and looks like success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
311b2ea55c |
populate member from the authenticated socket
The last mechanical link: chat socket -> resolveMemberRun(userId) -> ClaudeSpawnStreamingParams.member -> claude-manager's branch -> spawnClaudeAsMember -> sudo setpriv. The path from a request to a privilege drop is now complete. Resolved from the authenticated socket, never from the client message — the same rule server.tsx applies to the pty sidecar, where it deletes any client-supplied osUser/home from the query string before setting its own. resolveMemberRun returns undefined rather than throwing when a home cannot be resolved, because undefined means "the owner" downstream: an account with no Linux user has nothing to confine a turn to, and falling back to the owner is the one wrong answer that must not happen by accident. A separate function with that reasoning attached rather than an inline ternary. Still inert. Both gates refuse non-owners before this line is reached, so the only path that reaches it today returns undefined via isOwner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b646140dbf |
22: the backstop is in the right process, and Bun honours it
Checked both things that would make a handler look like a fix while being one: it is in user-instance.ts, which ecosystem.config.cjs:23-25 confirms is what officer-agent runs — the proxy runs index.ts and would have been a perfect inert place to put it — and Bun 1.3.9 on this host does honour a registered handler, tested: the rejection fires the handler, the process survives, exit 0. Without one Bun terminates, which is the four crashes. Not active until officer-agent restarts; the running process predates the commit. Worth noting the restart is also the diagnostic. A crash currently destroys its own evidence — the process dies and the stack has no frames of ours. Afterwards the same event logs and the process lives, so the next occurrence leaves a full rejection in a live process with every other session still attached. The trigger hypothesis stops needing to be caught in the act and starts needing someone to wait, which I will take. On uncaughtException: agreed, and the asymmetry is not inconsistent. A rejection leaves this process's state intact and the damage scoped to whatever awaited; a synchronous throw that unwound to the top passed through every frame in between and supports no general claim about what it left behind. The two differ in what they imply about state, not in what they cost. And the durable-sessions instinct is the sharper half. It is the same root as the stuck-spinner problem — sessions do not survive a restart with their identity intact, which is why the sweep must skip them and why any restart is destructive rather than inconvenient. Three symptoms, one missing property, worth naming before they get fixed separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8c4f150c15 |
stop one conversation's transport hiccup from killing every session
host found this while we were elsewhere: four agent-sidecar crashes tonight, one
truncating the owner's turn mid-sentence.
error: ProcessTransport is not ready for writing
at write (…/claude-agent-sdk/sdk.mjs) <- no frames from our code
A floating rejection inside the SDK's own input pump, so no await of ours could
have caught it. With no handler anywhere in src/servers it reached the top
level, Bun exited, PM2 restarted, and every live session on the machine died —
not just the one whose transport failed.
That is
|
||
|
|
0bc7858302 |
20: scoping verified, stop before deprovisionOsAccount, and the sidecar is crashing
Both changes hold. The gid is threaded from account.gid with a comment that says why the field
exists beside uid. The six commands enforce for real — ownedSession compares session.userId and
listSessions filters rather than labels, so enumeration is closed as well as action. 97 tests,
gates unchanged.
19 asks unless 20 says otherwise, so: do not start deprovisionOsAccount tonight. Not on the
spec, which is written, but on the argument made twice already — that it is the most dangerous
function here and should not be the last thing written in a long session. 17 said it was the
last commit of the night and 19 followed it. Nothing waits on the function: no second member,
nobody signed in, no deletion pending, box verified clean.
Aside, outside this thread and at the owner's request. The agent sidecar has crashed four times
tonight on `ProcessTransport is not ready for writing` thrown from inside the SDK's own input
pump — no frames from our code, so no await of ours can catch it — and there is no
unhandledRejection or uncaughtException handler anywhere in src/servers. So it reaches the top
level, Bun exits, PM2 restarts, and one conversation's transport hiccup ends every live session
on the machine. That is
|
||
|
|
d59adbf1f2 |
scope the six sessionKey commands to their caller
The control surface half of
|
||
|
|
9833822625 |
pass the member's gid instead of reusing their uid
host caught that provisionRootlessDocker had no gid field, so the new install -d passed uid in the group position. Correct on this host only because useradd allocates a per-user group; wrong on any account whose gid is not its uid — one created by hand, one on a host whose login.defs uses a shared group, or one ensureOsUser adopted rather than created. Mode 700 means the group triad grants nothing, so nothing breaks today. That is what makes it worth fixing now rather than later: it would surface only after somebody widened the mode for an unrelated reason, and then not obviously. The call site already held account.gid from ensureOsUser — the same value the .local fix used correctly earlier the same night. Threaded through rather than derived, and the field carries a comment saying why it is separate from uid, since they are equal here and a reader would ask. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4ef99bf3e5 |
16: storage fix is right, but it group-owns docker storage by uid
Owning the ordering rather than testing for it is the right resolution to the strip race, and
the comment carries the reasoning. Unverified here — it needs a reprovision.
One finding. The `install -d` passes String(params.uid) in the `-g` position, and it is not a
typo: provisionRootlessDocker's params are { osUser, uid, home } with no gid, so the uid is
standing in for one. Correct on this host only because useradd allocated a matching group —
green is uid=1001 gid=1001. The .local fix in the same night used params.gid where it had it,
and provision-os.ts:90 already holds account.gid from ensureOsUser, so the fix is to thread it
through rather than derive it.
It matters because ensureOsUser ADOPTS an existing passwd entry when name and home match, and
an account made by hand, or a host whose login.defs uses a shared group, can have gid != uid.
Then a member's Docker storage is group-owned by a group that is not theirs. Mode 700 means
nothing breaks today, which is what makes it the kind of thing that surfaces after someone
widens the mode for an unrelated reason.
Also agreed to leave .local closed to the file browser, but on stronger grounds than symmetry:
the change would mean moving the ACL pass after directory creation — reordering the one
function that has produced three bugs tonight — to gain a directory holding an overlay2 tree
and a versions symlink that nobody wants to browse.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
893130940e |
create docker's storage ourselves, so the strip is not a race
host verified green's reprovision: claude 2.1.228 installs and runs as the member, the file browser reads their home, rootless Docker runs and sees 0 containers while the owner has 8. First end-to-end proof of any of this. One thing came out dirty. ~/.local/share/docker carried the home's inherited default ACLs after a "successful" strip, because the strip was guarded on existsSync and only the daemon creates that directory. On a first run the guard was false and the strip no-opped; the retry then started the daemon, which created the directory and inherited the defaults. The run meant to clean it up was the one that made it, and the guard could not tell "nothing to strip" from "nothing there yet". Now created by us before the daemon exists — member-owned, 700, nothing to inherit — and the strip is unconditional afterwards, repairing an account provisioned before this and no-opping on a clean one. A guard that depends on another process having got there first is a race however it is written; the fix is owning the order rather than testing for it. Third bug of this class tonight: an implicit parent directory, a strip guarded on another process's work, and an installer piped into the wrong shell. All three were invisible until a real member account existed, which is the argument for making the second one sooner than feels necessary. Left alone deliberately: .local being unreadable by the platform (a decision about intent, not a defect, and the owner's), and the -u 70 + bind mount observation, whose probe host already distrusts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e99949b1ac |
14: claude installs and runs for a member, first time anywhere
Verified against a real reprovision of green at 00:02. Both fixes in
|
||
|
|
ef000aaf51 |
pipe the installer into bash, and stop inventing a root-owned .local
Green's first provision failed three ways. host caught all three on the live
box; two are fixed here and the third is his to bisect.
THE INSTALLER IS BASH AND WE PIPED IT INTO SH. A script read on stdin never has
its shebang honoured — the interpreter you name is the one that runs it — and
install.sh declares #!/bin/bash and uses [[ =~ ]] on line 9. On Ubuntu /bin/sh
is dash, so it died with `Syntax error: "(" unexpected`, which reads like a
corrupt download rather than the wrong interpreter. scripts/setup.sh carried the
same line for the owner's own install and is fixed too.
INSTALL -D CREATED ~/.local AS ROOT. `install -d` makes missing parents but
applies -o/-g/-m only to the final component, so blessing ~/.local/dockers
invented a root:root .local inside the member's own home. Rootless Docker then
died on `mkdir …/.local/share: permission denied`, and the Claude installer
targets ~/.local/bin, so fixing the shell alone would have hit this next.
That is
|
||
|
|
52b021bbe2 |
12: first real provision failed three ways, and two are repeats
Answering 11's question: provisionClaudeCli ran for the first time anywhere and did not work.
Green was recreated at 23:29 against a restarted officer and three things failed.
The installer is bash and the pipe is dash. os-user-claude.ts:77 runs `curl … | sh`, which
ignores the script's #!/bin/bash and hands bash-only syntax to dash — /bin/sh is dash on
Ubuntu. Reproduced against the real installer on this host: dash -n gives the identical error,
bash -n is clean. scripts/setup.sh:858 carries the same line.
~/.local is created root:root. os-user.ts:398's `install -d -o -g -m 711 …/.local/dockers`
creates the missing parent but applies ownership only to the final component — the same defect
|
||
|
|
7cb402b25a |
chat sessions record whose they are, and refuse a mismatched caller
host found this reading 10: chat sessions carry no identity at all. state.ts
held a flat sessionKey -> transcript uuid map, the in-memory sessions Map was
keyed the same way, and websocket.ts takes sessionKey and resumeSessionId
straight off the client message.
|
||
|
|
4da82e7f91 |
spec deprovisionOsAccount, from a real teardown rather than from reading the code
Written in docs/ rather than COMMS because COMMS is deleted when per-user Claude lands and this describes a project that starts after it — a spec that gets deleted before it is implemented is not a spec. Everything measured on this host on 2026-08-11. The evidence for why it exists: after deleting a member through the UI, the row was gone and the Linux account, a working login shell, a healthy postgres container, 454MB of home and Docker storage, lingering, the runtime directory and the subuid ranges were all still there. Three things the spec carries that reading the code would not have produced. terminate-user is not a barrier. A member's /bin/zsh -i survived it by three hours, and userdel refuses while a process owned by the account is alive, so an implementation that trusts it works on a quiet account and fails on a member who left a shell open. The subuid half. Rootless Docker storage is owned by MAPPED ids, not the member's uid — postgres's data directory belonged to 231141, not 1002. userdel releases the range and a later account can be allocated it, so a check for "nothing owned by the freed uid" passes while hundreds of megabytes are still owned by the freed range. Verification has to scan the range. And a correction to the order I actually used: sever the data from the uid BEFORE releasing it. The teardown ran userdel first and removed data after, which leaves a window where the uid is free while files still carry it. The irreversible step goes last. Also specified: never userdel -r, preserve-by-chown as the default with destroy opt-in, refuse to release the uid if the sever failed, and do not run anything as the member after terminating — creating a session recreates the runtime directory the step just removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
73c359fd5f |
10 (amended again): chat sessions have no identity, and it should block the gates
Appending what is still missing to close per-user Claude, at the owner's request, into the
same unread file rather than opening 12.
The one worth reordering around: chat sessions never got the identity fix that
|
||
|
|
80e1a746c0 |
10 (amended): the teardown is done, and terminate-user does not reap everything
Amending 10 in place rather than adding 12: it is my own file, nobody has read it or acted on it, and it carried a PREDICTION about deleting green that is now a measurement. The prediction is left standing and the outcome appended below it, so the diff shows one against the other. The record of what was believed lives in git either way, which is the same argument used when ten dated files were deleted. Item 1 of 01 is now observed. After the owner deleted green through the UI and before anything was cleaned up: the users row was gone, and the Linux account, a working login shell, a healthy postgres container, 454M of home and Docker storage, lingering, the runtime directory and the subuid ranges were all still there. Nothing broke, which is what makes it dangerous. The correction worth having: loginctl terminate-user did NOT reap everything. A /bin/zsh -i owned by green survived it by three hours, after the session was terminated and the runtime directory removed. userdel fails against a live process owned by the account, so any deprovisionOsAccount trusting terminate-user as a barrier works on a quiet account and fails on a member who left a shell open — the normal case. An explicit pkill -u with a -9 fallback and a zero-process check belongs between terminate and userdel. Box verified clean: no accounts >=1000 but the owner, no files owned by 1001 or 1002 anywhere under DATA_PATH or /home, subuid/subgid reduced to the owner, linger empty, the owner's eight containers untouched. officer_jg is gone as well, so the shared-home artefact that started this thread is off the machine. Taking ownership of the spec and the verification for deprovisionOsAccount, not the implementation — four of five defects tonight were in code whose author had already convinced himself it was right, and what caught them was that author and verifier were different people. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d051aff4e0 |
10: operations done, chmodSync proven on a real install, green teardown warning
No 09 — the other side stopped, so the odd number goes unused; keeping parity.
|
||
|
|
07ab3f9e7a |
08: chmodSync verified, and the box now repairs itself at next bootstrap
Read
|
||
|
|
3e0daee611 |
chmod the mcp config, because writeFileSync's mode never fires on it
host measured what
|
||
|
|
fe30164452 |
06: writeFileSync's mode is ignored on an existing file, so the 0600 never fires
The mcp-config branch is right. The mode fix is correct in intent and inert everywhere it
matters: fs.writeFileSync passes mode to open(2), which honours it only when CREATING the
file. On an existing one the call truncates and the mode is ignored. Measured here — a 644
file stays 644 after writeFileSync with {mode:0o600}, while a fresh path comes out 600.
So user-instance.ts is fixed for new installs and a no-op for every deployed one, which is
the whole exposed population. 05 says the change takes effect at the next write; it will not.
That is the difference between "closed after a restart" and "never closed, and nobody is
watching any more". Verified after the commit: the file is still 0644 and green can still
read it.
Fix is an explicit chmodSync after the write, keeping the creation mode too — the first
closes the open-to-chmod window on a fresh write, the second repairs an already-leaking
install as a side effect of the next bootstrap, which is the only mechanism here that reaches
a deployed box.
Reordered the owner actions: the immediate chmod on the existing file stops the bleeding in a
second with no restart and no deploy, and it is what makes rotation final rather than a moving
target. I have not touched the file — it is the owner's and it is production.
Agreed on stopping. The next commit should be the rotation and the chmod, not feature code,
and no, do not move the history layer overnight on top of an open item.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
df4503180f |
stop handing a member's turn the owner's mcp config, and close the file
host found a live credential exposure while answering my question about what else the member branch missed. It was outside the diff, and predates all of it. MCP-CONFIG WAS NOT BRANCHED. mcpHostPath is module-level, written once at the owner's bootstrap, and was applied to every turn. Its env block carries OFFICER_AUTH_TOKEN, a 30-day JWT signing as the owner — so a member's turn would have spawned their MCP server holding it. Now inside the params.member ternary alongside the binary and the spawn, for the reason already written there: these values say whose turn this is and have to move together. A member gets none. What they should get instead is undecided, and undefined beats the owner's. THE FILE WAS 0644. Written with a bare writeFileSync into a 755 directory, on a host where `terminal` is granted to every role by default — so any member could cat it and hold owner-level API access on loopback. host verified that as a real member on the production host rather than reasoning about it. Now 0600. The mode is the only half of that which is code. The token has been world-readable and stays compromised until rotated, the directory chain above it is still 755, and neither is fixable from a commit. Both written up for the owner in COMMS 05, along with why I am stopping here rather than continuing: the next commit should be the rotation, not more feature work stacked on top of an open exposure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fe7bd49bc7 |
04: mcp-config is not branched, and it points at a world-readable owner token
Answering the question in 03 — whether the member branch misses another owner-derived value
the way cwd did. It does: extraArgs { 'mcp-config': mcpHostPath } at claude-manager.ts:373 is
outside the ternary and applies to every turn. But the larger finding is not in that diff.
LIVE ON THIS SERVER, and unrelated to per-user Claude: user-instance.ts:132 writes
mcp-host.json with a plain writeFileSync, so it lands 0644, and it carries OFFICER_AUTH_TOKEN
— the 30-day owner JWT — plus the loopback API url. Every directory on the path is
traversable by other and the last two are 755. Verified as green: the file reads. Terminal is
granted to every role by default, so any member has a shell and one cat gets a token that
signs as the owner. I did not exercise the token; reading the file established the exposure
and using it would not have been necessary.
Fix is the owner's: mode 0o600 on write, tighten DATA_PATH/<email> from 755, and rotate the
token, since mode bits do not retroactively unread it.
The two halves compound. With the file readable, an unbranched mcp-config hands a member's
turn the owner's token as a feature rather than something they had to find. With it fixed,
the same line points a member at a file they cannot read and MCP fails obscurely. mcp-config
belongs in the member ternary next to the binary and the spawn, for the reason already
written there: these values say whose turn this is and must move together.
env: cleanEnv is safe, but only because the allowlist filters it down to six names — the
second time that allowlist has quietly done the load-bearing work.
Rest of the wiring is correct. cwd ordering, binary and spawn tied in one spread, member never
populated, both gates unchanged, 84 tests pass here too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
fbabc22ed7 |
wire the member branch into the SDK spawn, still unreachable
ClaudeSpawnStreamingParams takes an optional member {osUser, home}; createSession
branches on it, using their binary and spawnClaudeAsMember together, or the
owner's CLAUDE_BIN as before.
THE BINARY AND THE PRIVILEGE DROP ARE ONE BRANCH ON PURPOSE. settingSources
makes ~/.claude authoritative for settings and ~ is whatever HOME the process
gets, so pointing the SDK at a member's binary while spawning as the service
user would read the OWNER'S settings and credential while running the member's
code — and it would look like it worked.
cwd defaults to member.home before HOST_HOME for the same reason: HOST_HOME is
this process's home, so a member would start in a directory they cannot read and
the failure would present as a broken agent rather than a wrong cwd.
Nothing populates `member`. Both gates refuse non-owners before any of this is
reached, so the delta is that spawnClaudeAsMember now has two importers instead
of one, and neither path a user can take changes. Verified rather than assumed,
since host made it a condition: both gates intact, 84 tests pass.
Not authorization: host gave an opinion on wire-first and deferred to the owner,
who has not ruled. Corrected in COMMS, where 01 had overstated it. The gates
come off on the owner's word alone; this reverts as one commit if the answer is
no.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
c15bd082b5 |
02: the probe fix verified live, and the owner has not ruled on wire-first
Read
|
||
|
|
baa2d29fa4 |
read the probe off stdout by position, and renumber comms
host's finding on
|
||
|
|
45df9aaa20 |
review 288679af: the sudo fix is right, its markers collide with the paths
The one-call change is correct and I confirmed the effect. One latent defect it introduced. claudeLoginState decides by substring on `probe.out`, and asMember returns stdout and stderr CONCATENATED — while `bin` is a substring of .local/bin/claude and `cred` of .credentials.json, both of which are passed as arguments. So anything writing either path to stderr flips the flag. Demonstrated here against green with neither file present: `sh -xc` traces the two paths and both booleans come back true, claiming a member is signed in when they have never logged in. Not live — the happy path measures empty stdout and stderr and the correct false/false — but it fails unsafe and is one debug flag away. Uppercase markers do not fix it: a trace echoes the script, so the literal lands on stderr too. The channel is the problem. Suggested stdout-only with a positional two-character answer, keeping stderr for diagnosis but out of the string being matched. Verified from the request list: the pertento host key matches the live server AND the known_hosts every push of mine has used for hours, so first-use acceptance was correct; both chat gates still up and spawnClaudeAsMember imported by zero files; 53 tests pass here, matching their count. Items 2 and 3 need `pm2 restart officer` and a provisioned member, which is outside what the owner scoped to me. Flagged as not-done rather than silent, and referred back to the owner along with the two questions that are theirs: who owns the parked items, and wire-first versus verify-first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
288679af09 |
one sudo call for both agent-status answers, not two
host's finding on
|
||
|
|
f4dc46d67a |
review fb2c5c28..6b7aad91: both guards fire now, one aggregate cost noted
Read |
||
|
|
6b7aad91db |
make both env guards able to fire, and follow the symlink
host was right twice, including about his own advice. Two dead guards had shipped here, both for the same reason: written inside the spawn closure, where the only way to reach them is to spawn — and the passing path spawns sudo. So nothing ever demonstrated either one firing. THE SUBSET CHECK WAS ALSO TAUTOLOGICAL. `permitted` came from the same constants memberEnv builds childEnv from, so it was empty under every edit where that holds — the exact criticism that retired NEVER_ENV. Worse, it lost the one live trigger the denylist had: a credential added to ALLOWED_ENV used to throw, and under the subset check widened the permitted set in the same motion and passed silently. That is the realistic future edit and it was the one left unguarded. Now both, and the denylist tests the LIST rather than the instance, so it fires on exactly that edit. Extracted as `assertEnvSafe` so a test can pass a poisoned allowlist — the guards being untestable in place is why they were decorative twice. THE BINARY CHECK WOULD HAVE THROWN ON EVERY TURN. Anthropic's installer puts a symlink at ~/.local/bin/claude into a versioned directory; resolve() does not follow symlinks, so the string compare matched only while `command` arrived as the symlink spelling, and would have failed the moment anything upstream normalised it — at exactly the point the hook gets wired. Compared through realpathSync on both sides now, per turn and never cached, since `claude update` moves the target. Nine tests pin all of it: a poisoned allowlist, a stray key, each NEVER_ENV name, and the symlink/target/missing-path cases. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
06bfcf9556 | Merge remote-tracking branch 'pertento/sidecar-app-store' into sidecar-app-store | ||
|
|
2bd96a9a98 |
tell a member why their agent is not working, instead of 403
|
||
|
|
fb2c5c28cb |
correct my own advice: the subset check cannot fire either
The inversion I suggested replaced one dead check with another. `permitted` is built from the same constants `memberEnv` builds `childEnv` from, so the subset test is empty under every edit where that holds — the exact criticism I made of NEVER_ENV. Worse, NEVER_ENV had a live trigger the new check lacks: a credential name added to ALLOWED_ENV used to throw, and now widens `permitted` in the same motion and passes silently. That is the realistic future edit, and it is the one now unguarded. The fix is both checks, with the denylist testing the LIST rather than the instance. Also verified here: Anthropic's installer puts a symlink at ~/.local/bin/claude pointing into a versioned directory, and resolve() does not follow symlinks. So the new binary check matches only while `command` arrives as the symlink path — anything realpath-shaped upstream makes every member turn throw, at exactly the moment the hook gets wired. Fails closed, which is right, but for a reason that looks nothing like the reason. Signing as `host` from here on, at the owner's request, to tell the two ends of this channel apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6a85ca5d1f |
the binary check that was a comment, and a guard that could fire
Four fixes from the live server's review. One was a real defect.
THE BINARY WAS NEVER CHECKED. spawn-as-member passed `command` from the SDK
through untouched while a comment claimed the member's own install was what ran.
Since claude-manager resolves the OWNER'S CLAUDE_BIN at module load, wiring the
hook would have exec'd the owner's binary as the member — the precise confusion
this file exists to prevent, asserted in prose and enforced nowhere. Now throws
unless the command resolves to claudeBinIn(run.home).
NEVER_ENV COULD NOT FIRE. It tested an environment that memberEnv builds from
ALLOWED_ENV, so a denied name was already impossible; it was also missing six
credential variables the installed SDK reads. Replaced with the subset check the
reviewer proposed: anything not in ALLOWED_ENV or {HOME, CLAUDE_CONFIG_DIR} is a
leak whatever it is called. Complete by construction, and it cannot rot as the
SDK grows variables — which the denylist provably had already.
Also: one derivation of the binary path instead of two (install resolved from
the email, exec from the home — fine until they disagree), and the constraint
that ALLOWED_ENV may never hold a secret written at the list itself, since
`env K=V` in the argv is visible in /proc/<pid>/cmdline to every account.
Not acted on, and said so in COMMS: their finding that the 711 in
|
||
|
|
6cd462caf6 |
close §4: officer_jg has no users row, and the adoption rule held
Queried the two rows the handoff asked for. There is no `users` row for `officer_jg`, and ids 2-4 are absent, which tells the whole story: an earlier row for jg@pertento.ai under the `officer_`-prefixed naming got a Linux account at uid 1001, the row was deleted without `userdel`, and the re-created account correctly refused to adopt it and took uid 1002. The home is derived from the email, which never changed — hence two accounts, one home. So this is the delete path, not a bypassed adoption rule, and it is observed rather than theorised. Inert today: the home belongs to green and its ACL names only pastilhas and green, so officer_jg cannot read it. The live hazard is uid 1001 going to the next member, which is what deprovisionOsAccount and its chown to the service user would close. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f09385c789 |
reply from the live server: the bind mount runs, and 711 is not what fixed it
Answers §1 of the per-user-accounts handoff and reviews the per-user-claude one.
A bind-mounted postgres:18-alpine starts, initialises and stays healthy under green's
rootless daemon — so 3bea46f's open question is closed. Two corrections though.
The mechanism in
|
||
|
|
ed52faae02 |
correct two claims about agents that the SDK disproves
docs/per-user-linux-accounts.md carried the reasoning that per-user agents were a large piece of work, and both halves of that reasoning were wrong. The SDK does have somewhere to put a uid — spawnClaudeCodeProcess, documented for running Claude Code in VMs and containers — so a member's turn does not have to become its own process. And the credential claim was backwards: the proxy holds the OWNER'S credential, reading the owner's own ~/.claude/.credentials.json, so pointing a member at it spends the owner's account on the member's turns. The previous handoff had already retracted that one; the doc had not caught up, which is how a retracted claim stays live. Corrected in place rather than deleted, with what was believed and why it was wrong, because the superseded version is the interesting part: the first claim is what made agents look like a later stage than they are. Adds the constraint that actually is out of scope, which the old text never stated: no platform process ever runs as a member, because the sidecar holds POSTGRES_URL and the JWT secret. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
701933d30d |
commit and push when it goes wrong too
The instruction was standing and nowhere in the repo, so every session started by waiting to be asked. Written down because the reason is not obvious: a branch held back because it is unfinished, untested or a dead end is exactly the branch whose history is worth having. A reverted commit and its message explain why an approach was abandoned; a quietly discarded attempt teaches the next person nothing, and they will try it again. The obligation that comes with it is saying what state the work is in — in the message, and in COMMS/ when another agent will pick it up — rather than letting a clean commit imply it is finished. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
62e98dff2e |
a member's claude is their own binary and their own login
First half of per-user Claude. Provisioning and the privilege drop, not yet
wired to a turn — the chat gates stay up and behaviour is unchanged for
everyone. Committed unfinished on purpose so the reasoning is on the record
before the server agent runs any of it; the state is written up in
COMMS/sidecar-app-store/2026-08-11-per-user-claude-handoff.md.
THE CLAIM THAT CHANGED. docs/per-user-linux-accounts.md:226-229 says the Agent
SDK "has nowhere to put a uid", so a member's turn has to become its own
process — a change of shape rather than a flag. It is a flag:
sdk.d.ts:951 exposes spawnClaudeCodeProcess, documented for exactly this ("run
Claude Code in VMs, containers, or remote environments"), and node's spawn
already satisfies the SpawnedProcess shape it wants. So no second sidecar, no
PM2 entry, no inverted transport, and none of the registry rework a second
instance would have forced (registration is name-keyed and evicts its
namesake; the nine claude verbs resolve by capability with no selector).
THE PLATFORM NEVER RUNS AS A MEMBER. The tempting reading of "each member runs
their own Claude" is a second officer-agent under their uid, and it is wrong:
that sidecar needs POSTGRES_URL and the JWT signing secret, so a member-uid
process holding them could read every account and sign a token as the owner —
strictly more than their shell can do, and already forbidden by the .env boot
check. The harness stays the service user's; the thing that runs the member's
code and holds the member's credential is theirs. That is the pty sidecar's
shape, not a new one.
PER-MEMBER BINARY, deliberately, over one shared /usr/local/bin/claude. The
private part is the credential, not the executable — but claude updates itself,
and a root-owned binary is one a member cannot update, which turns "my agent is
a version behind" into a request to the owner. Same installer the owner's own
install uses, run as them, in their home. Idempotent by skipping when present
rather than re-running: the retry button reprovisions on every press.
ALLOWLIST, NOT A FILTER, for the child's environment. At the moment of the call
the calling process holds POSTGRES_URL, the JWT secret and the owner's
ANTHROPIC_API_KEY; setpriv --reset-env means nothing crosses unless written
into the argv, so an allowlist is the complete answer to what a turn can see,
and a denylist would have to be right about every variable added later.
NEVER_ENV throws rather than leaks if someone widens it.
Login is the member's own act against their own account. The platform cannot do
it for them and must not try — the alternative is lending them the owner's
credential. claudeLoginState only reports whether the credential has appeared,
and reads it as the member, so a true answer means their process can reach it.
NOT VERIFIED: any of it at runtime. tsgo passes; nothing has been provisioned
and the spawn hook has never been called. If it turns out setpriv breaks how
the SDK reaches the process, this approach is wrong and the fallback is the
earlier plan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
48ed171b38 |
point CLAUDE.md at COMMS, so a new session finds it without being told
The channel is only useful if it is read, and relying on the owner to remember to say "check COMMS" in an opening prompt puts the mechanism back where it started. CLAUDE.md is loaded automatically, so the pointer belongs there: what the directory is for, that newest date wins, and which streams exist. Also states the split it is easy to get wrong — durable reasoning in docs/, coordination in COMMS — and that a spent handoff should be deleted rather than left to be mistaken for current. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
10fe3ffe65 |
COMMS/sidecar-app-store: a tracked channel between agents
Findings were being relayed through the owner by hand, from memory, at the end of long sessions. A file survives a context window and carries its reasoning; a message does not. The untracked COMMS/ at the workspace root stays what it is — state about one machine at one moment. This one is in the repo because any clone should carry it. First handoff covers what I would otherwise have asked the owner to pass on: the bind-mount container test I could not run here and how to retrofit green, the setup-dockers.sh PG18 layout left deliberately alone, the terminal replay bug and the deprovision/uid-reuse hole with a proposed fix, the shared-home question I cannot answer without the passwd and users rows, and the four things most likely to surprise a reader — bootstrap-only default grants, chat grantable but refused, Bun.spawn ignoring uid, and members never getting the owner's anthropic proxy. The README states the convention: dated files, verified separated from assumed, name lines, reply in a new file rather than editing someone else's, and delete a handoff when it is spent. Durable reasoning goes in docs/ or next to the code — this directory is for coordination, not for the record. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
401dcb710c |
a blessed directory for container bind mounts
From a live-server report: a bind-mounted postgres:18-alpine crash-looped with
`mkdir: can't create directory '…/18/docker'` on a directory that already existed.
|
||
|
|
f0af7237db |
terminal, chat and files are granted by default; permissions screen simplified
DEFAULTS. Every role now starts with the three confined capabilities at write, seeded in bootstrap. These are what the platform is FOR — an account that signs in and reaches none of them is not restricted, it is useless, and making the owner grant them by hand first is a step with no decision in it. Seeded as real rows rather than implied by absence, which keeps the table's one rule intact: a missing row means no access, always, with no exception to remember. Revoking one therefore works like revoking anything else — the row goes and nothing puts it back. Done in bootstrap because that happens exactly once per install, so seeding can never fight a later revocation. Non-fatal: an owner whose roles hold nothing is a one-click fix, while failing bootstrap over it leaves a platform with no account at all. `app` capabilities are deliberately not defaulted — they reach data the owner may not intend to share, and each needs a sidecar before it means anything. SCREEN. Role selection is tabs rather than a dropdown: three roles are the axis you move along, and a select hid two of them behind a click while giving no sense of which one you are editing. Row descriptions are gone — with three rows called Terminal, Chat and Files they explained nothing — and the "needs a Linux account" warning went with them, since every account now gets one at creation, so it was noise about a state that no longer occurs on its own. `needsOsAccount` is removed from the API too, not just hidden. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
aaeb3424ab |
the rootless docker fix is proven; correcting the record
|
||
|
|
3bea46f2d7 |
rootless docker per member — provisioning works, running a container does not yet
Not finished. Committed because the diagnosis is worth more than the code. WHY ROOTLESS AND NOT THE DOCKER GROUP. `usermod -aG docker <user>` is the one-line version and it is root: `docker run -v /:/host -it alpine chroot /host` is a root shell, which reads .env, every other member's home and the wallet seed. Every boundary from today, bypassed by one documented command. Rootless gives what was actually asked for — a daemon per account, containers in that account's user namespace, images in their own home. VERIFIED on this host: provisioning succeeds, the server reports 29.5.0, the daemon runs as the member, `docker pull` puts 403 MB under their own home, and `docker ps -a` shows nothing while the owner has four containers. That last line is the isolation, measured. NOT VERIFIED: actually running a container. It failed, and the cause is an interaction between two things built today: failed to copy xattrs: failed to set xattr "system.posix_acl_default" on …/volumes/…/_data Creating a volume copies xattrs, and the DEFAULT ACLs on a member's home — added so the file browser could read their files — are inherited by Docker's storage, where a mapped id inside a user namespace is not a valid id to set. Both features correct alone. The fix here strips default ACLs from ~/.local/share/docker only, leaving the access ACLs the file browser needs. That fix is UNPROVEN. The re-test failed for a different, environmental reason: probe users recycle uid 1001, and a stale lingering systemd user manager from a previous probe answered `systemctl --user`, so the unit appeared not to exist. Cleaned with `loginctl terminate-user`. Retest on a machine that has not had a uid-1001 user, or on a fresh uid. Also worth knowing before this ships: uid reuse after deleting a member is a real hazard, not just a test artefact — the next member gets the previous member's uid, and anything left lingering belongs to them. setup.sh gains uidmap and dbus-user-session as core packages; the shell template exports DOCKER_HOST from $XDG_RUNTIME_DIR when the socket exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
71589aee99 |
a member's terminal looks like the owner's
A new Linux account opens a shell with nothing: useradd copies /etc/skel, which on Ubuntu is a bash rc, and the account's shell is zsh — so it got no prompt, no history, no completion, no colour. "Their own account" should not mean a worse terminal than the owner's. src/servers/shell-skel/zshrc is the template, and scripts/starship.toml is reused rather than copied: setup.sh already deploys it for the owner, so one file serves both audiences and they cannot drift. Seeded by provisionOsAccount, which means the retry button applies it to accounts that already exist — no delete-and-recreate. The template depends on nothing but zsh. Starship, eza, nvim, bun, deno and cargo are each used only if present, and every path is $HOME-relative — the owner's own .zshrc has three absolute /home/pastilhas paths in it, which is exactly what a template must not inherit. Without starship it falls back to a zsh prompt showing the same information, because a shell that opens with a broken prompt reads as a broken machine. Never overwrites: written only when the file is ABSENT. ~/.zshrc.local is sourced last and never written, so there is somewhere to put your own config that no future template can reach. Three fixes found by running it: - install -D creates missing parents but applies -o/-g only to the FILE, so ~/.config came out root:root — readable but not writable by its owner, which would have surfaced weeks later as one tool mysteriously failing. The parent is now created explicitly. - useradd took its shell from process.env.SHELL, which under PM2 is whatever PM2 was launched from. A member's shell depended on how the server happened to be started. Now chosen from what is installed: zsh, else bash. - the pty sidecar spawned ITS $SHELL for a member, not theirs. It now execs their passwd shell via sh -c, so the login shell in /etc/passwd is the one they get. starship moves out of the light-profile skip. The light profile exists to serve a file browser, a terminal and chat — the terminal is one of its three reasons to be, and it is what every member gets. Leaving starship out meant the fallback prompt on exactly the installs most likely to have members. oh-my-zsh, eza and lazygit stay full-only. Verified in a real member shell: zsh from passwd, HISTFILE in their own home, eza-backed ll, starship active, EDITOR=nvim, and an edit to .zshrc surviving a reprovision. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e4acf19a35 | Merge remote-tracking branch 'gitea/master' into sidecar-app-store | ||
|
|
4d4a253f72 |
terminal runs as the member; chat is grantable and still refused
TERMINAL is confined now, and the shell is genuinely theirs. The pty sidecar spawns it through sudo setpriv as their own account, in their own home, with the platform's environment cleared. Verified end to end against the sidecar's own socket: id -u 1001, not 1000 file the shell wrote owned by ptyprobe ps -o user=,args= ptyprobe /bin/zsh -i env | grep -c POSTGRES 0 osUser and home are resolved in upgradeWs from the authenticated account, and whatever the browser sent under those names is DELETED first. The bridge forwards the query string to the sidecar untouched and the sidecar starts a shell from what it finds there, so trusting the client for either would let a member ask for the owner's uid in a query parameter. node-pty does support uid/gid, unlike Bun.spawn, and they are deliberately unused: they set the ids without applying the account's groups or resetting the environment, so the shell would keep the owner's groups and everything Bun loaded from .env. Also closes the pty identity blindness in TODO.md. Sessions record whose they are, list and kill scope to the caller, and re-attaching to a session belonging to another account is refused — otherwise a member resumes someone else's shell by guessing an id that travels in a query string. Measured: member killing the owner's session -> ok:false, owner killing it -> ok:true. CHAT is confined so the owner can grant it and the route resolves, and both execution doors refuse a non-owner: the router wholesale, and the socket in server.tsx. The agent has not moved — the SDK spawns claude itself with nowhere to put a uid, and every transcript path resolves through the owner's home, so a member would read the owner's session list and run an agent as the owner. Reads are refused too, because listClaudePwds returns the names of the owner's projects. A deliberate, temporary gap at the owner's request: permission and route now, function when a turn can be spawned under runAs with the member's own HOME. Both guards say so, and the registry test names them so a future edit cannot move one without the other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eda004a46d |
a naked platform does not describe what it does not have
Reversing my own call from an hour ago. I built the denied-route screen to EXPLAIN the absence — "Music is not installed", with a link to the app store — and argued a redirect erases what you asked for. The owner's correction is the better principle: a server should not know about a sidecar it does not have. Explaining Music is the app describing a feature that, as far as this install is concerned, does not exist, and it leaks the whole catalogue of what could be installed to any member who types a URL. So a denied path is now indistinguishable from an unknown one: redirect home, the same answer App.tsx's path="*" already gave. One behaviour for a member without a grant, an owner without the sidecar, and a typo. Nothing disclosed. The Permissions screen loses both explanatory blocks for the same reason. One listed every capability whose sidecar is absent — a catalogue of uninstallable features presented as a permissions decision. The other described chat, tasks, the desktop and the wallet as "not grantable" to an owner who may have none of them installed. `notInstalled` is gone from the API too, not just hidden in the UI. What is on that screen is what this server can actually do. Still short of what the owner described, and worth naming rather than implying otherwise: routes are DECLARED in App.tsx for every screen and this hides the ones that should not resolve. The end state is routes REGISTERED from the manifests of installed sidecars, so an uninstalled feature has no route to hide. The manifests already exist and the dock is already built from them; the router is not, yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b4f88ec161 |
routes refuse at the route, and a new home is empty
Three things, from a member sitting on /music with no music capability on a server with no music sidecar: an empty library, and 403s in the console. PERMISSIONS AT THE ROUTE. `canVisit` filtered the dock and nothing else, so the tile was hidden and the route was wide open — typing the path, following an old link or restoring a tab rendered the screen anyway. RouteGate now wraps every screen in one place, inside the error boundary. It does not redirect. Sending someone to `/` erases what they asked for and reads as a bug: they clicked Music and landed on Home. It says why instead, and the URL stays put so a reload after installing the thing just works. And it says which of the two reasons applies, because they need different screens and send the reader to different places. `not-installed` is a fact about the SERVER — the owner gets a link to the app store. `not-granted` is a fact about the ACCOUNT, and only the owner can change it. Presenting either as the other sends you looking in the wrong place. ROUTES FOLLOW THE SIDECAR. Free, once the above exists: `deniedRoutes` already covers "held but its sidecar is not installed", so an uninstalled feature has no tile AND no screen. The dock, the Permissions list and the routes now agree because they read one answer. NO MORE SEEDING. Downloads/Documents/Music/Videos/Pictures are gone from both places that made them — the member's provisioning and, older and worse, `/ls`, which created folders in somebody's home as a side effect of LOOKING at it. A listing that invents its own contents is a listing you cannot trust, and the platform has no standing to choose a person's folder layout. A new home is empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4b058a6703 |
fix the sign-out reload loop I shipped an hour ago
The 401 handler ended with location.replace('/'), guarded by "unless the path starts
with /signin". There is no /signin route — the sign-in screen IS path="/". So every 401
on the signed-out landing page navigated to the page it was already on, fetched again,
401'd again. A hard refresh loop with no way out of the tab.
The reload was never what fixed anything: useAuth already renders the sign-in screen
when there is no token. It only existed to drop a stale query cache. So it is now the
last thing attempted and bounded three separate ways, any one of which breaks a loop
alone:
1. no token -> return. A 401 while already signed out is expected, not a revocation.
This one alone ends it, because a reloaded document has nothing left to clear.
2. once per document, module flag.
3. once per tab, sessionStorage marker — which also covers a host that re-injects the
token on every load, where clearing storage cannot help and guard 1 never fires.
Anyone stuck in the loop from the previous build: localStorage.clear() in the console.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d3bed0add9 |
the file browser can actually read a member's home, and plans is gone
"This folder is empty" was a lie. The five seeded directories were sitting there and the platform's readdir raised EACCES: a member's home is 700 and owned by them, which is correct for a shell and locks out the file browser, which runs inside the platform process. /ls caught the error and returned an empty listing, so a refusal looked exactly like data. Two doors, two boundaries, and that is the point rather than a compromise. The terminal and the agent RUN AS the member and the kernel is the boundary there. The file browser acts on the member's behalf from inside the platform, which already applies its own containment and is the owner's process on the owner's machine — it can read anything via sudo regardless. Giving it access describes who is doing the work. Done with named POSIX ACLs, because it has to hold in BOTH directions: a file the platform writes must be editable by the member and vice versa. Mode bits cannot say that — whichever party is neither owner nor group lands in "other", and widening "other" opens the home to every account on the box. A shared group fails the same way, since both parties would have to be in it and that puts every member in a group that can read every other member's home. Two named entries plus `d:` defaults grant exactly two users and are inherited by whatever either side creates, whatever their umask. Verified: platform lists the home, member edits a platform-written file, platform edits a member-written file, and a SECOND member is refused on both ls and cat. /ls now distinguishes EACCES from a missing directory. An empty result is data and must never be how a refusal looks. acl joins the core packages in setup.sh — the alternative is an account that provisions and then cannot list its own home. Also: the file browser's own useTasks/useAgents fired /tasks, /agents and both category endpoints on every render, which is where the last four 403s came from — they are the context menu's Run Task and agent submenus, execution-only. Gated. And plans is deleted: router, screen, routes, dock tile, hook, page title and its capability. It read markdown from <repo>/plans, which does not exist. Fresh-install Permissions is now Files alone, with Terminal to come. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |