Four things the original got wrong, all of which only show up on a machine that
is not this one:
It detected swap with `swapon --show | grep -q '/'` — a test for a swap FILE.
A machine using zram or a swap partition reports no swap at all, and the step
would add a swapfile beside working swap. Reads SwapTotal from /proc/meminfo
now, which covers every kind.
It never looked at free disk. On a VPS with 4G free and 16G of RAM it would
fallocate 8G, fail, and take the run down under `set -e`. The recommendation is
now capped by what is actually there, keeping 5G back, and refuses rather than
shrinking to something useless.
fallocate was assumed to work. It produces a file that btrfs and zfs will not
swap on, so dd is the fallback — slow, but it always works.
swappiness was written by sed'ing /etc/sysctl.conf in place, tangling it with
whatever else lives there. It is a drop-in at /etc/sysctl.d now, so what this
script set is visible as its own file.
Role-dependent, which is the first use of MACHINE_ROLE: swappiness 10 on a
server, where swapping is the emergency valve and a page fault on a request path
is latency somebody is waiting for; the kernel default of 60 on dev, where
swapping out an application nobody has touched in an hour is exactly what you
want.
WSL is left alone entirely — WSL2 runs its own managed swap inside the VM, and a
swapfile written here is wasted disk the kernel will not use.
Sizes are reported rounded rather than floored. A 4 GiB swapfile is 4194300 kB,
which floors to 3 and reads as though a gigabyte went missing. Free disk stays
floored, deliberately: it decides how much to allocate, and rounding up invents
space.
Verified on this host — 4G RAM, 4G existing swap correctly detected and left
alone — and with the disk check stubbed at 6G free (caps to 1G) and 3G free
(refuses).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The refresh lived inside "System update", which `step` skips when its name is
already in the progress file. So a resumed run — the common case, since that is
what the progress file is for — installed core utils, added the fastfetch PPA and
set up the Docker repo against whatever the index happened to say hours or days
earlier. On a box left overnight that is a stale index and a "package not found"
somewhere unrelated.
It now runs in pre-flight, unconditionally, before any step exists to skip it.
`apt-get upgrade` stays where it was and stays confirmable: refreshing the index
changes nothing on the machine, upgrading is the one thing that does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The original took whatever was typed and handed it straight to timedatectl. An
unknown zone — a typo, a guess at the spelling — fails there, and under `set -e`
that takes the whole run down four steps in. Names are now checked against
/usr/share/zoneinfo before use, and a bad one just re-asks.
It also never showed what the machine was already set to, and defaulted to option
1 (UTC) on Enter, so pressing return on a correctly-configured box silently moved
it. Now the current zone is printed, Enter keeps it, and a zone equal to the
current one reports nothing to do rather than setting it again.
timezone_current reads three sources — timedatectl, /etc/timezone, then the
/etc/localtime symlink — because they differ in availability rather than in
answer: timedatectl needs systemd, /etc/timezone is Debian's, and the symlink is
the one that is always there. timezone_set writes through timedatectl where
there is a systemd to talk to and the files directly otherwise, which is what it
would have written anyway; that is also the WSL path, where timedatectl exists
but does nothing.
Europe/Berlin added to the shortlist; TIMEZONE in the environment answers the
prompt ahead of time and is validated the same way, failing early with the bad
value named.
Verified detection (UTC here), validation of four names, and the env-var
rejection path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Was three unconditional lines that ran on every pass and reported success either
way. Now it checks, says what it found, and asks.
The original tracked one fact where there are two:
what a new login shell is told to use LANG in /etc/default/locale
whether that locale actually exists whether it has been generated
Setting the first without the second is what produces "setlocale: LC_ALL: cannot
change locale" on every ssh login and every perl invocation. They fail
differently, so the step names whichever one is actually missing rather than
reporting a flat "locale not set".
Also fixes two things the original would have hit on a minimal image:
locale-gen comes from the `locales` package, which cloud base images do not
ship and which is not in core utils. It is installed on demand rather than
assumed, instead of failing with "locale-gen: command not found".
The locale is uncommented in /etc/locale.gen rather than only passed to
locale-gen as an argument. A locale generated by argument alone disappears the
next time anything regenerates from that file.
`locale -a` prints en_US.utf8 where the configuration spells it en_US.UTF-8, so
both sides are folded before comparing — a literal match reports a working locale
as missing.
LOCALE in the environment overrides the default. pacman, dnf and brew branches
are written but unreachable while the pre-flight gate is apt-only; macOS has no
system locale to set and says so.
Verified both paths on this host: en_US.UTF-8 reports already set and generated,
pt_PT.UTF-8 correctly reports both facts missing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"the only step that changes software already on this machine" is true, and
reads like a warning about something dangerous rather than a description of
apt upgrade. The reasoning stays in the section comment, where it explains why
this is its own step; the prompt just says what it does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The section asked "Proceed?" without saying what it was proposing to change —
the one question in the script where the answer matters most, since it is the
only step that moves versions of software already on the machine.
pkg_upgradable now names them, from `apt-get upgrade -s`: the same calculation
the real run does, as opposed to `apt list --upgradable`, which also lists
packages held back that would not actually move.
Nothing to upgrade means no prompt at all, and the summary says so rather than
claiming an upgrade happened. The list is capped at 25 with a count of the rest,
because a box untouched for months lists hundreds and a wall of names is no more
informative than the number.
Verified against this host (0 upgradable, so it reports current and does not
ask) and with a stubbed 40-package list for the cap.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three sections where there were two, and none of them touches the machine until
you say so:
2. System update upgrades what is already installed
3. Core utils what the distribution provides
4. Command-line tools lazydocker, lazygit, starship, fastfetch
The split matters because these are different kinds of change and deserve
separate answers. System update is the only step in the whole script that moves
versions of software already on the machine; core utils only ever adds what is
absent; and the four tools are upstream binaries the distribution does not ship
at all. Previously the update and the core packages were one step and the tools
were tacked onto the end of it, so agreeing to "essentials" meant agreeing to all
three at once.
Every section now prints what it will install and what it is leaving alone, then
asks. Enter means yes — unlike the machine-role question, which has no default,
because these are "do the thing you already asked for" and making twenty of them
require a deliberate keystroke would train people to hold the y key down.
ASSUME_YES=1 answers all of them for an unattended run, and EOF fails with that
named rather than spinning.
Refusing is recorded rather than glossed: LAST_SKIPPED feeds the summary, so a
declined section reads "Core utils: SKIPPED by request — cowsay neofetch" instead
of quietly reporting nothing installed.
Nothing to install means no prompt at all — there is nothing to agree to.
announce_plan takes the array NAMES rather than their contents, because once a
list has been through word splitting an empty one cannot be told from a missing
one.
Verified all three paths: accept, refuse, and nothing-to-do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The summary claimed credit for everything in a section's list, including the
packages it had just decided to leave alone — so a run that installed nothing
still ended with "Command-line tools: lazydocker lazygit starship fastfetch".
The announce above it said "nothing, all present" in the same breath.
pkg_install and tools_install now record LAST_INSTALLED and LAST_KEPT, and
summarise_last turns those into one honest line:
Core packages installed: btop tmux (17 already present)
Command-line tools: already present, nothing installed
Verified all three shapes — everything present, nothing present, and mixed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
No default, and it is the only question in the script like that. A guessed
default is right often enough to be trusted and wrong in exactly the case that
costs the most — pinning a static IP on a rented box, or leaving the firewall
open on one. Every branch downstream is about what this machine is exposed to,
so it is worth one deliberate keystroke rather than an Enter.
Empty and unrecognised answers re-ask rather than aborting; a failed read means
EOF rather than a wrong answer, and fails with the environment variable named,
because otherwise the loop spins forever the first time this runs unattended.
Drops guess_machine_role, which existed only to supply that default. default_iface
stays — the static IP section needs it when it is ported.
MACHINE_ROLE in the environment still answers it ahead of time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Your call, and the right one. Editing in place let mis-grouped code sit unnoticed
until it scrolled past in a live run — which is exactly how the four upstream
binaries buried in "System Update & Essentials" were found. Porting forces the
question of where each thing belongs before it runs, not after.
machine-setup.sh now contains only what has actually been worked through:
pre-flight, system update and core packages, command-line tools, and the summary.
1149 lines down to 172. The sections still to come are listed in a NOT PORTED YET
block, in order, and each arrives as its own commit.
The original is beside the other superseded scripts as
scripts/setup-old/setup-ubuntu.sh — verified byte-identical to the live
/root/ubuntu-setup copy — so porting reads from a file in the repo rather than
from root's home.
Two claims trimmed from the ported summary, because they were true of the old
script and not of this one yet: it reported the shell as "zsh (Oh My Zsh +
Starship)" unconditionally, and told you to reconnect as a user it had not
created. Replaced with what pre-flight actually knows — system, role, user.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
lazydocker, lazygit, starship and fastfetch were buried inside "System Update &
Essentials", after the package install and with no announcement — so a run
appeared to be installing system packages and then started pulling tarballs and
printing a five-shell starship tutorial. They are a different thing: upstream
binaries on their own release cadence, not anything the distribution ships. Now
their own step, announced in the same shape as the package section.
Each is checked before it is fetched. The original re-ran every installer on
every run, which is why a machine that already had starship got it reinstalled
along with its "add this to your ~/.zshrc" instructions — advice this script
does not want followed, since it writes the shell config itself. Its output is
now dropped; errors still surface.
Two real bugs fixed on the way:
lazygit's asset name was hardcoded to x86_64, so on arm64 the download 404s
and tar fails partway through the run. It now maps ARCH, and spells the
architectures the way lazygit does rather than the way we do.
The version was extracted with `tr -d 'v'`, which deletes every v in the
string rather than the leading one. `${version#v}` instead.
fastfetch stays a package but stops assuming the PPA is needed: Ubuntu picked
it up in 24.10, so the repository is now checked first and the PPA added only
where the archive has nothing. Verified on this host — noble genuinely has no
candidate, so the PPA is still the only source here.
Verified both branches of tools_install by stubbing the presence check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A WIP boundary after section 2 so the finished part can be run start to finish
on its own, without the untouched sections below acting on the machine. It moves
down as each section is worked through and goes away when the walk ends.
Also ignores .setup-progress, which the script writes beside itself and is
per-machine. The exit message names it, because with it in place a second run
skips section 2 and the rewritten part cannot be re-felt from scratch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
lib/packages.sh, and section 2 wired to it.
The rule it exists to enforce: `apt-get install <present-package>` is not a
no-op, it upgrades the package if the repository has a newer one. On a machine
somebody already uses that silently moves a version they chose, and a setup
script is the last thing that should do that behind their back. pkg_install
queries the package database first and names only the genuinely absent packages
on the command line — a package already installed is never passed to apt at all.
It also says so out loud, every time, because a provisioning run should not be
opaque about what it is doing to the machine:
:: Core packages — installs what is missing, keeps what you already have
already here: curl ca-certificates gnupg git jq …
to install: btop tmux
Section 2's flat list of 19 is now pkgs_core(), split per package manager rather
than through a canonical-name table with overrides. The names genuinely disagree
(build-essential/base-devel, fd-find/fd) and three of them are not packages
elsewhere at all — apt-transport-https, lsb-release and software-properties-common
are apt concepts that exist to let later steps add the Docker repo and the
fastfetch PPA. A `case $PM` shows what each system actually gets, in one place.
Of those 19, six are load-bearing and the rest are the environment. Only
build-essential reaches beyond itself: it is a meta-package, so on a box with a
pinned gcc it pulls the distribution default alongside. Noted where it is
declared; it is the first thing to move out of core if that ever bites.
apt-get upgrade stays, but as its own announced step — it is the one place that
deliberately moves versions, rather than something that happens as a side effect
of asking for a tool.
DEBIAN_FRONTEND=noninteractive and NEEDRESTART_MODE=a now live inside the
helpers. needrestart has been on by default since Ubuntu 22.04 and stops to ask
which services to restart, which is how an unattended run ends up silently
waiting for a keypress.
dpkg-query on the status field rather than `dpkg -s`, which also succeeds for a
package removed but leaving its config behind — that state would read as present
and never be reinstalled.
Verified against this host's real dpkg database: all 19 report present, and a
mixed list correctly passes only the absent ones through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MACHINE_ROLE is homelab, vps or dev, and several steps have a different right
answer per role with no way to work it out themselves: whether the address is
yours to pin (static IP), whether the box faces the open internet (fail2ban, SSH
hardening, UFW), and whether it is allowed to sleep (suspend, logind).
Asked in pre-flight rather than at each point of use. The steps that care run
from swap through to the firewall, and being asked "is this a VPS?" for the
fourth time halfway down a provisioning run is how people start answering
without reading.
The default offered is guessed from whether this machine's own address is in
RFC1918 space, which beats asking whether it is virtualised — a homelab is very
often a VM on Proxmox and would be misread as rented — and is the same fact most
of the branches turn on anyway. A graphical session means dev; so does macOS.
It is only ever a suggestion the user confirms.
MACHINE_ROLE in the environment answers it ahead of time for an unattended run,
which is why it is declared with :- rather than a plain assignment. The first
version wiped the caller's value before ask_machine_role ever saw it; caught by
running with MACHINE_ROLE=vps and watching the menu appear anyway.
Verified: guesses vps on this host (public IPv4, no DISPLAY, no display
manager), env override takes, and a bad value fails with the three valid ones
named. Nothing consumes the role yet — the steps get wired as each is worked
through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
scripts/setup/ is now what the new installer is being built in — machine-setup/
for the box, officer-setup.sh for the platform on top — and everything being
replaced moved to scripts/setup-old/. It still works and is still what to run.
Three things the move broke, and what each needed:
starship.toml is not an old-setup artifact. os-user-shell.ts reads it at
RUNTIME to seed a member's ~/.config/starship.toml when their Linux account is
provisioned, and line 125 reads it inside a try whose catch returns
"could not read the shell templates" — so account provisioning would have
failed outright, not degraded. Moved back to scripts/setup/, which is where it
belongs anyway (one file, both audiences) and which leaves the code correct
with no edit.
package.json's `setup` script pointed at a path that no longer exists. It now
points at officer-setup.sh, where the installer is going, rather than at
setup-old/ which is temporary.
officer-setup.sh was created empty. An empty script exits 0, so `bun setup`
would have reported success while doing nothing — worse than the broken path
it replaced. It now explains that it is not written yet and exits 1, naming
the setup-old script to run meanwhile.
Also brought .tmux.conf and ufw-docker-rules.conf in beside machine-setup.sh,
which reads both from SCRIPT_DIR and had been silently skipping them since the
script was vendored. ssh-keys.zip deliberately stays out: it is key material,
and *.zip is ignored.
Comments in os-user-claude.ts, app-store/preflight.ts and two docs still name the
old scripts/setup/setup.sh path. Left alone on purpose — repointing them at
setup-old/ only to repoint them again when officer-setup.sh lands is churn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Structure before the work rather than during it: scripts/machine-setup.sh becomes
scripts/setup/machine-setup/, with the script itself as the entry point and
lib/base.sh holding what every part of it needs.
machine-setup.sh pre-flight and the numbered sections, for now
lib/base.sh shared state, output, the step/resume machine, prompts,
and OS detection
The rule for lib/ is definitions only — nothing there installs, writes or
restarts anything, so sourcing it is safe from anywhere. That is why the ERR
trap stayed in the entry point: a trap is a side effect on whoever sources it.
Behaviour is unchanged. Verified by diffing the moved region against the previous
commit: identical set of functions, and the only differences are added comments,
section banners, fail() reformatted onto three lines, and one new line — a guard
against double-sourcing, which matters because steps will source this directly
once they move out, and a second pass would reset SUMMARY.
The sections are still one 1111-line block below pre-flight; they move into
steps/ as each is worked through. The script also still reads ssh-keys.zip,
.tmux.conf and ufw-docker-rules.conf from SCRIPT_DIR, which is now this
directory, so those three steps warn and skip until the files follow it here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The script assumed Ubuntu on x86_64 in every line of it. detect_os() now runs
first and fills in OS, OS_NAME, OS_VERSION, PM, ARCH and IS_WSL, so the steps
have something to branch on as support for other systems is added.
Read from /etc/os-release rather than probing for a binary: a machine can have
more than one package manager on PATH, and only os-release can say which
distribution this actually is or give a version worth printing. Sourced in a
subshell so its NAME, VERSION and ID do not leak in here. ID_LIKE is the
fallback, so Pop!_OS, Mint and EndeavourOS resolve without being named.
ARCH is normalised to amd64/arm64 in one place because upstream disagrees —
Neovim ships aarch64, Go and Docker ship arm64, lazygit ships x86_64 — and
several steps hardcode one spelling today.
Windows exits with a message pointing at WSL2. WSL itself is detected and
warned about rather than refused: it reports as Linux but has no real systemd
session, so the suspend, logind and boot-hang steps do nothing there.
Everything below pre-flight is still apt and systemd only, so a gate refuses
pacman/dnf/brew by name rather than half-building a machine and stopping
somewhere unhelpful. Relax that case one entry at a time as each grows a path.
Verified on this host (Ubuntu 24.04.4, amd64, apt) and by stubbing uname and
os_release for arch, manjaro/arm64, fedora, pop, macos, mingw and riscv64.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A byte-for-byte copy of /root/ubuntu-setup/setup-ubuntu.sh, the script that has
provisioned every Ubuntu server here. Committed unchanged, before any edit, so
that everything the setup-script rework does to it reads as a diff against what
actually ran on real machines rather than against a tidied-up version of it.
Nothing in the repo calls this yet. It also cannot find three files it reads from
its own directory — ssh-keys.zip, .tmux.conf and ufw-docker-rules.conf all live
beside the original in /root/ubuntu-setup, and SCRIPT_DIR is scripts/ here, so
those steps warn and skip.
The original stays where it is and stays authoritative until this one replaces it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It was the only service here publishing on 0.0.0.0, and a published docker port
is not behind the firewall: docker writes its DNAT rules straight into the nat
table, which ufw's INPUT chain never sees. `ufw default deny incoming` never
covered 80/443/81 — ufw-docker-rules.conf on the host exists to patch exactly
that, and patching a rule is weaker than never opening the socket.
The address is read from `tailscale ip -4` at run time rather than passed in,
because the host provisioning has already done `tailscale up` by the time this
executes. It is validated against 100.64.0.0/10, the range tailscale and
headscale both allocate from. SETUP_NPM_BIND overrides it.
With neither, selecting NPM exits instead of falling back to 0.0.0.0 — a
fallback would silently undo the point of the change.
Two consequences worth knowing. tailscaled becomes a boot-order dependency, so
the script warns when it is not enabled at boot; docker's restart policy covers
the window but only if the tailnet comes up on its own. And HTTP-01 ACME
challenges can no longer reach port 80, so any certificate NPM issues now needs
DNS-01.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
8 Rust, 9 PulseAudio, 10 cliamp, 13 yt-dlp and 17 the remote desktop move to
scripts/setup/setup-sidecars.sh, which nothing invokes — running it is a
deliberate act. They are what the optional, sidecar-backed features need on the
host, not what the app needs to serve itself.
11 Neovim, 12 the shell extras and 14 the npm globals are gone entirely. The
host provisioning already installs node, npm, pm2, Claude Code, Neovim and the
shell, and two installers racing for the same binaries is worse than one. That
makes node, npm, pm2 and the agent CLIs prerequisites of this script rather
than products of it, so the verification block still checks claude and pm2 —
section 19 warns and skips rather than failing when pm2 is absent, which would
otherwise finish "successfully" with nothing listening.
eza is the one casualty: the provisioning installs lazygit, starship, oh-my-zsh
and nvim, but not that.
Section numbers keep their gaps so the two files read against each other. One
line survives from the removed section 14 — the ~/.local/bin PATH export, which
section 19's `has pm2` and the agent's claude lookup both still depend on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review of 76cd7c2. The ACL finding is right and the fix is correct — verified here that `setfacl -R -P -b`
removes the default entries as well as the access ones, which the man page splits between -b and -k and
does not settle. `acl` is already a core package in setup.sh, so the new hard dependency is real.
But the checker it added cannot fail in the way it is documented to be run.
`assert-uid-free.sh` is invoked as `sudo ./assert-uid-free.sh --check ...`, and sudo's env_reset DROPS
DATA_PATH, so the script falls back to the hardcoded `/home/pastilhas/officerdev/data` — which is not this
machine's data directory and does not exist. Every check in the file is "look for X, report ok when nothing
is found", so a missing root reports clean without looking. Demonstrated: a tree carrying both
`user:65534:rwx` and `default:user:65534:rwx` was reported as `ok no ACL entries naming uid 65534`.
The ACL check is the one that fails silently and completely, because it is the only one scoped to DATA_PATH
alone — the uid and subuid scans still walk /home and would catch something. So the check just added to
catch the hazard ownership cannot see is the check a wrong DATA_PATH disables.
Fixed by refusing rather than passing:
require_roots every search root must exist, or exit 2 naming it and showing the sudo invocation
that preserves DATA_PATH
numeric guard uid/start/count must be numbers. deprovisionOsAccount logs '<no-subuid-range>' in
that position for an account with no /etc/subuid entry, and pasting that log line in
— which is exactly how it is meant to be used — made sub_end empty and turned the
range scan into a no-op.
The handler's audit line now prints DATA_PATH inside the command it tells the operator to copy, and says
so explicitly when there is no subuid range rather than emitting a command that cannot work.
Verified: bogus root exits 2, non-numeric range exits 2, and the ACL check FAILS on a specimen tree
carrying the entries — the "make it fail before trusting it to pass" step from the spec's own subuid
section, now done for the ACL half too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reviewing 46799dad against a real tree: severMemberTree reassigns ownership and leaves the
access-control entries behind. confineUserTree grants each member a named ACL on their whole
tree — u:<uid>:rwx plus a default: copy — and chown does not remove them. They are xattrs
rather than ownership, and they store the uid NUMERICALLY.
Measured: chown -h -R to the service user leaves user:<uid>:rwx intact on the directory, its
children and their defaults. So a preserved tree owned by the platform still grants the freed
uid read and write on every byte, and the next account allocated that number inherits the
previous member's home, SSH keys, credentials, transcripts and container storage. That is the
hazard the function exists to prevent, reached through ACLs instead of ownership.
This was my omission as much as the implementation's: the spec said "sever the data from the
uid" and specified only chown, and assert-uid-free.sh checked find -uid, which reads ownership
and cannot see an ACL. Both are fixed here — the spec now requires setfacl -R -b alongside the
chown, and the checker scans DATA_PATH with getfacl -R -n for entries naming the freed uid.
The checker was verified to catch it: run against green's live tree it now reports
"ACL entries still grant uid 1001 (user:1001:rwx)", which it did not before.
The code fix is one line in severMemberTree and is not mine to make.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
scripts/assert-uid-free.sh is the verification half of docs/deprovision-os-account.md, written
outside the implementation on purpose: a checker the function calls is a restatement of its own
beliefs rather than an audit. Nine checks — passwd entry, uid reuse, both subid files, linger,
runtime dir, live processes, files owned by the uid, and files owned anywhere in the freed
subuid range.
Two modes, because the range has to be captured BEFORE deletion. userdel removes the
/etc/subuid entry along with the account, and after that there is no way to ask what range it
held — so a checker that only runs afterwards silently drops the half most likely to be wrong.
Exercised against green while fully provisioned: eight of nine checks fail, exit 1. A checker
that has never been seen to fail is not evidence.
And the trap worth knowing before anyone trusts a green result: the subuid check passes
vacuously on most accounts. Files get a mapped owner only when a process inside a container runs
as a NON-root user; an image whose files are root-owned maps to the member's own uid and leaves
the range empty. Measured on green after a night of real use — claude installed, an image
pulled, transcripts written — the range check found zero files and passed without testing
anything. The spec now says how to build a specimen that actually exercises it, and to watch the
check fail on that tree before trusting it to pass on a cleaned one.
Docs and a script only; no behaviour change. On a branch, for whoever merges it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two of the eight did not belong. cleanup-desktop.sh is the teardown — the inverse of an install, not part
of one. provision-user-dirs.ts runs per account at invite time, on a machine that is already set up.
Both are back at the top level, with their `../` derivations and usage strings put back.
What is left is what a fresh machine runs once: the two installers (setup.sh, setup_mac_light.sh), the
two things setup.sh calls (setup-dockers.sh, setup-desktop.sh), and the two files they deploy —
starship.toml, copied to ~/.config, and officer-set-display.sh, which setup-desktop.sh installs to
~/.local/bin as a login-time mode setter. The last one is not a setup script and does not read like one;
it is here because it is install payload, same as the toml, and setup-desktop.sh loads it by
`$(dirname $0)`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
scripts/ was holding two unrelated kinds of thing: install-this-machine, and run-this-occasionally. The
eight installers now live in scripts/setup/; what stays at the top level is the build steps (gen-index,
prebuild, build/) and the two maintenance scripts (reindex-music, rebuild-soulseek-tree).
The move is not just a rename. Three of these derive the repo root from their own location:
setup.sh:51 PROJECT_DIR="$(dirname "$SCRIPT_DIR")"
setup_mac_light.sh:51 same
cleanup-desktop.sh:134 ENV_FILE="$(dirname "$0")/../.env"
Left alone, all three would now resolve to scripts/ — and nothing downstream complains. PROJECT_DIR is
where .env is written, where `bun install`, `gen:index` and `db:push` run, and what pm2 is pointed at, so
a fresh install would have quietly provisioned scripts/ and reported success. cleanup-desktop.sh fails
the other way: it would find no .env, print "No .env — skipping", and leave the real VNC_PASSWORD in the
real file. All three are now `../..` with a comment saying why the level matters.
provision-user-dirs.ts imports data-path.ts relatively; that one tsgo caught.
Also disambiguated `setup.sh` where it had become two files. app-store/templates/<name>/setup.sh is a
per-sidecar installer with its own contract, and preflight.ts + docs/sidecar-app-store.md discussed both
in the same paragraph. The host one is now spelled with its full path at those sites.
Verified: bash -n on all six shell scripts, tsgo clean, os-user tests pass, both derivations resolve to
the repo root, starship.toml still resolves from os-user-shell.ts, and provision-user-dirs.ts runs under
DRY_RUN.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eight scripts in scripts/ that nothing references and that mostly can no longer run. Kept in history;
none of them is recoverable knowledge that isn't already in the code they migrated to.
Three could not run at all against the current database:
migrate-items-to-files.ts SELECT * FROM tasks — that table was dropped when items became files
reset-user-data.ts deletes chat_sessions, chat_groups, projects; none exists. It has no
transaction, so it would wipe user_settings, user_state,
user_integrations and dock_configs and THEN throw. A half-wiped account
is worse than no script. It also misses chat_session_events, which is
where chat state actually lives now.
add-email-dock-user2.ts one-time, hardcoded to user 2, seeds a dock containing /projects
The rest are spent migrations whose destination is now the only implementation:
migrate-auth-to-pg.ts JSON -> Postgres, 2026-02
migrate-pg-to-files.ts Postgres -> JSON, the other leg of the same abandoned round trip
migrate-server-settings-to-pg.ts 2026-02
migrate-emails-to-sqlite.ts backfill into the email sidecar's store, 2026-07-31
seed-imap-uids.ts the sidecar writes imap_lastuid/imap_uidvalidity itself now
(sidecar/email/gmail-api.ts:533-535)
Kept, and why, since "unreferenced" was not the test: rebuild-soulseek-tree.ts is reusable by
construction — it runs the same buildTree the sidecar's ingest runs, so it answers any future change
in tree shape. reindex-music.ts is named in sidecar/music/index.ts:447. provision-user-dirs.ts shares
USER_DIRS with data-path.ts. cleanup-desktop.sh and officer-set-display.sh are called by
setup-desktop.sh.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Green's first provision failed three ways. host caught all three on the live
box; two are fixed here and the third is his to bisect.
THE INSTALLER IS BASH AND WE PIPED IT INTO SH. A script read on stdin never has
its shebang honoured — the interpreter you name is the one that runs it — and
install.sh declares #!/bin/bash and uses [[ =~ ]] on line 9. On Ubuntu /bin/sh
is dash, so it died with `Syntax error: "(" unexpected`, which reads like a
corrupt download rather than the wrong interpreter. scripts/setup.sh carried the
same line for the owner's own install and is fixed too.
INSTALL -D CREATED ~/.local AS ROOT. `install -d` makes missing parents but
applies -o/-g/-m only to the final component, so blessing ~/.local/dockers
invented a root:root .local inside the member's own home. Rootless Docker then
died on `mkdir …/.local/share: permission denied`, and the Claude installer
targets ~/.local/bin, so fixing the shell alone would have hit this next.
That is 71589ae for the second time — same function shape, same silent parent,
same class of consequence. Its own commit message said this surfaces "weeks
later as one tool mysteriously failing"; it took twenty minutes. Grepped the
other install -d/-D sites: os-user-shell already creates its parent explicitly,
os-user-ssh has no implicit parent.
NOT fixed: the file browser's ACL mask on a member home, where access mask is
--- while default:mask is rwx. That pattern means a chmod ran after the setfacl
and clamped only the access side, so the primitive is right and something later
is wrong. host has the live filesystem and has already half-excluded the
suspect; guessing from here would churn a working block. Noted that this commit
adds an install -d before the one he was about to bisect, so it wants a
reprovision first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Not finished. Committed because the diagnosis is worth more than the code.
WHY ROOTLESS AND NOT THE DOCKER GROUP. `usermod -aG docker <user>` is the one-line version
and it is root: `docker run -v /:/host -it alpine chroot /host` is a root shell, which reads
.env, every other member's home and the wallet seed. Every boundary from today, bypassed by
one documented command. Rootless gives what was actually asked for — a daemon per account,
containers in that account's user namespace, images in their own home.
VERIFIED on this host: provisioning succeeds, the server reports 29.5.0, the daemon runs as
the member, `docker pull` puts 403 MB under their own home, and `docker ps -a` shows nothing
while the owner has four containers. That last line is the isolation, measured.
NOT VERIFIED: actually running a container. It failed, and the cause is an interaction
between two things built today:
failed to copy xattrs: failed to set xattr "system.posix_acl_default" on …/volumes/…/_data
Creating a volume copies xattrs, and the DEFAULT ACLs on a member's home — added so the file
browser could read their files — are inherited by Docker's storage, where a mapped id inside
a user namespace is not a valid id to set. Both features correct alone. The fix here strips
default ACLs from ~/.local/share/docker only, leaving the access ACLs the file browser needs.
That fix is UNPROVEN. The re-test failed for a different, environmental reason: probe users
recycle uid 1001, and a stale lingering systemd user manager from a previous probe answered
`systemctl --user`, so the unit appeared not to exist. Cleaned with `loginctl terminate-user`.
Retest on a machine that has not had a uid-1001 user, or on a fresh uid.
Also worth knowing before this ships: uid reuse after deleting a member is a real hazard, not
just a test artefact — the next member gets the previous member's uid, and anything left
lingering belongs to them.
setup.sh gains uidmap and dbus-user-session as core packages; the shell template exports
DOCKER_HOST from $XDG_RUNTIME_DIR when the socket exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A new Linux account opens a shell with nothing: useradd copies /etc/skel, which on Ubuntu
is a bash rc, and the account's shell is zsh — so it got no prompt, no history, no
completion, no colour. "Their own account" should not mean a worse terminal than the
owner's.
src/servers/shell-skel/zshrc is the template, and scripts/starship.toml is reused rather
than copied: setup.sh already deploys it for the owner, so one file serves both audiences
and they cannot drift. Seeded by provisionOsAccount, which means the retry button applies
it to accounts that already exist — no delete-and-recreate.
The template depends on nothing but zsh. Starship, eza, nvim, bun, deno and cargo are each
used only if present, and every path is $HOME-relative — the owner's own .zshrc has three
absolute /home/pastilhas paths in it, which is exactly what a template must not inherit.
Without starship it falls back to a zsh prompt showing the same information, because a
shell that opens with a broken prompt reads as a broken machine.
Never overwrites: written only when the file is ABSENT. ~/.zshrc.local is sourced last and
never written, so there is somewhere to put your own config that no future template can
reach.
Three fixes found by running it:
- install -D creates missing parents but applies -o/-g only to the FILE, so ~/.config came
out root:root — readable but not writable by its owner, which would have surfaced weeks
later as one tool mysteriously failing. The parent is now created explicitly.
- useradd took its shell from process.env.SHELL, which under PM2 is whatever PM2 was
launched from. A member's shell depended on how the server happened to be started. Now
chosen from what is installed: zsh, else bash.
- the pty sidecar spawned ITS $SHELL for a member, not theirs. It now execs their passwd
shell via sh -c, so the login shell in /etc/passwd is the one they get.
starship moves out of the light-profile skip. The light profile exists to serve a file
browser, a terminal and chat — the terminal is one of its three reasons to be, and it is
what every member gets. Leaving starship out meant the fallback prompt on exactly the
installs most likely to have members. oh-my-zsh, eza and lazygit stay full-only.
Verified in a real member shell: zsh from passwd, HISTFILE in their own home, eza-backed
ll, starship active, EDITOR=nvim, and an edit to .zshrc surviving a reprovision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"This folder is empty" was a lie. The five seeded directories were sitting there and the
platform's readdir raised EACCES: a member's home is 700 and owned by them, which is
correct for a shell and locks out the file browser, which runs inside the platform
process. /ls caught the error and returned an empty listing, so a refusal looked exactly
like data.
Two doors, two boundaries, and that is the point rather than a compromise. The terminal
and the agent RUN AS the member and the kernel is the boundary there. The file browser
acts on the member's behalf from inside the platform, which already applies its own
containment and is the owner's process on the owner's machine — it can read anything via
sudo regardless. Giving it access describes who is doing the work.
Done with named POSIX ACLs, because it has to hold in BOTH directions: a file the
platform writes must be editable by the member and vice versa. Mode bits cannot say that
— whichever party is neither owner nor group lands in "other", and widening "other"
opens the home to every account on the box. A shared group fails the same way, since both
parties would have to be in it and that puts every member in a group that can read every
other member's home. Two named entries plus `d:` defaults grant exactly two users and are
inherited by whatever either side creates, whatever their umask.
Verified: platform lists the home, member edits a platform-written file, platform edits a
member-written file, and a SECOND member is refused on both ls and cat.
/ls now distinguishes EACCES from a missing directory. An empty result is data and must
never be how a refusal looks.
acl joins the core packages in setup.sh — the alternative is an account that provisions
and then cannot list its own home.
Also: the file browser's own useTasks/useAgents fired /tasks, /agents and both category
endpoints on every render, which is where the last four 403s came from — they are the
context menu's Run Task and agent submenus, execution-only. Gated.
And plans is deleted: router, screen, routes, dock tile, hook, page title and its
capability. It read markdown from <repo>/plans, which does not exist. Fresh-install
Permissions is now Files alone, with Terminal to come.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
POST /api/users plus an Add-account form in Settings > User management. Until now
createUser had one call site — bootstrap, gated on an empty user table — so every
non-owner account anywhere had been inserted into Postgres by hand.
Created accounts are Active. The column defaults to Unverified and signin refuses
anything else with a bare UNAUTHORIZED, which is exactly what made the hand-INSERT
route look like a wrong password.
Also closes a hole found while reading the write path: a second Super Admin was
storable. The CHECK constraint pins user 1's role but cannot see other rows, and
getOwnerUser() was LIMIT 1 with no ORDER BY, so two holders would have made "who owns
this server" a question the query plan answered — and that answer feeds the agent
sidecar's identity, vault access and origin scoping. Both write paths now refuse the
role and getOwnerUser() orders by id.
USER_DIRS and provisionUserDirs move into data-path.ts so the create handler and
scripts/provision-user-dirs.ts cannot disagree about what an account's skeleton is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The three non-owner accounts had no directory of their own — only the owner did,
grown organically as features created what they needed. Adds an idempotent script
that creates the skeleton, and runs it for those three.
DATA_PATH/<email>/{home,attachments,cache,dashboards,email_accounts,
general_chat_sessions,logs,sidecar}
Most of those are also created on demand by whichever feature owns them, so
pre-creating buys legibility rather than function: the tree now shows the shape a
user has without having to use it first. `home` is the exception and the reason
this exists — nothing creates it today, because getOwnerHomeDir returns
process.env.HOME_DIR whenever it is set, which it always is on a real install. It
is where a non-owner's sessions will run.
Takes emails as arguments rather than reading the user table. The layout does not
depend on the database, keeping the DB out means it runs with nothing else up,
and at invite time the caller already knows the email. It refuses arguments that
are not email addresses, since a stray one would create a junk directory sitting
next to real user roots and looking like one.
Deliberately NOT a revival of the provision.ts deleted two commits ago. That one
was ~90% per-container scaffolding — shell rc files, a generated CLAUDE.md and
settings.json describing the container — wrapped around the two useful lines this
keeps.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A light install is reached at localhost on the machine running it, so PUBLIC_URL
has exactly one right answer. setup.sh required it with a re-prompt loop; under
the light profile it now defaults to http://localhost:$PORT. The macOS installer
already did this — this is parity, and it removes the one prompt in a light run
whose answer a non-technical user could not be expected to produce.
setup_mac_light.sh also wrote PUBLIC_BUILD_ENV="development", justified in a
comment as "what makes plain http://localhost work". That is no longer true, and
the cost of it is not small. IS_DEV_BUILD gates exactly three things:
origin validation already off regardless — ALLOW_ANY_ORIGIN defaults to true
password rules validatePassword is skipped entirely on change-password
rate limiting the limiter returns next() before doing anything
So the only live effects were losing the last two, for a benefit that another
default already provided. It now writes "production", matching setup.sh. Nothing
about localhost needed relaxing: browsers treat http://localhost as a secure
context, so passkeys, getUserMedia and the clipboard all work over plain HTTP,
and passkeys in particular derive their RP ID from the request origin rather
than a configured domain.
That last point is the boundary worth knowing: http://192.168.x.x is NOT a
secure context, so reaching a light install from another device means putting
an HTTPS proxy in front of it. Recorded in the comments at both prompts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Renamed for parity now that Linux has a light profile too:
scripts/setup_mac.sh -> scripts/setup_mac_light.sh
ecosystem.mac.config.cjs -> ecosystem.mac.light.config.cjs
The macOS process list was still a hand-copied subset, which is the shape that
broke it: written 2026-07-28, within days it was running the Anthropic proxy
under the name officer-claude with nothing spawning `claude`, and pointing at a
pty entry point that had moved. Both silent. It now declares names and reasons
and reads script/args from ecosystem.config.cjs, so a launch change on the host
reaches it for free.
The include/exclude checks moved into ecosystem.profile.cjs rather than being
copied into the second profile — duplicating the guard rails would have repeated
the mistake they exist to catch. Both profiles were re-tested against a mutated
host ecosystem: renaming an included app and adding an unclassified sidecar each
throw in both, and an unmodified host loads five apps in both.
Kept as two files rather than collapsed into one, even though they currently
produce identical output. The exclusions do not mean the same thing: on macOS
officer-vnc CANNOT run, there being no Xorg; on a Linux light install it could
run fine and you have chosen not to. Merging them would lose that, and they
diverge the moment one profile gains something the other cannot have.
setup_mac_light.sh's verification loop now reads app names with node instead of
grepping for `name:` — the derived profile has no literal keys, so the grep
would have silently listed no services at all, which reads the same as a healthy
install with nothing configured.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
OFFICER_PROFILE=light installs the same thing the macOS build does, on Linux:
the file browser, the terminal, and Claude/opencode chat. It skips the archive
extras, the sudoers entry and auto-suspend disabling, Go, Rust, PulseAudio,
cliamp, Neovim, the shell tooling and yt-dlp, brings up Postgres alone of the
five Docker services, and starts ecosystem.light.config.cjs. Unset or `full`
behaves exactly as before.
The profile changes which processes start, not which code ships — every API
route stays mounted, so features whose sidecars are absent report themselves
unavailable rather than disappearing.
ecosystem.light.config.cjs DERIVES its apps from ecosystem.config.cjs rather
than copying them, because the hand-copied Mac list was broken within days of
being written by a sidecar split in two and a pty entry point that moved, and
both failures were silent. Here a script/args change on the host propagates for
free, and two consistency checks turn the silent cases loud:
- a name the profile needs that the host no longer defines throws at load
- an app added to the host that is in neither the include list nor the annotated
exclusion list throws, so a new sidecar cannot default to "not in the profile"
without someone deciding
Both were tested against a mutated copy of the host ecosystem: renaming
officer-agent and adding an unclassified sidecar each throw, and the unmodified
file loads five apps.
The verification block now reads app names with node instead of grepping for
`name:` — the derived file has no literal keys to match, so a grep would have
silently verified nothing — and skips the checks for tools the profile did not
install, so a clean light run does not report Go and cliamp as missing.
Full-profile behaviour is unchanged by construction: every guard wraps the
original code in an else branch. The preamble was tested across unset, full,
light and an invalid value.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ecosystem.mac.config.cjs was written on 2026-07-28 and broken within days by two
changes it never caught up with, leaving two of its four processes dead:
- it ran `officer-claude` against src/servers/sidecar/claude/index.ts. That file
is the Anthropic credential proxy; the process that actually spawns `claude`
is user-instance.ts, which was never started. Chat had a credential holder and
nothing driving it. Split into officer-anthropic-proxy + officer-agent to
match the host, and retired the `officer-claude` name that caused it.
- officer-pty pointed at src/servers/api/terminal/pty-sidecar.mjs, which no
longer exists; the sidecar moved to src/servers/sidecar/pty/index.mjs. The
terminal could not start at all.
No startup ordering is needed between the proxy and the agent: the agent reads
the proxy secret from disk and, when it is not there yet, warns and re-reads
before the next spawn.
Five sidecars have been added since the file was written (memos, caldav, photos,
notify, wallet). All are excluded, and every exclusion is now listed with its
reason so the next reader can tell "not applicable on macOS" from "forgotten" —
which is the failure this commit is fixing.
setup_mac.sh needed no path corrections; its verification loop derives the
process names from the ecosystem file, so it picks the new list up on its own.
Two gaps closed there:
- it called `pm2 save` without `pm2 startup`, writing a process list that
nothing reads — a reboot left the machine with nothing running. Added the
launchd equivalent behind SETUP_BOOT, optional because a laptop is not a
server, and non-fatal because pm2's launchd integration can want a prompt.
- the closing "not installed on macOS" note listed four omissions when there are
now fourteen.
Both files parse and every ecosystem entry point resolves to a file that exists.
The launchd path is the one thing not verified here — it cannot be exercised
from the Linux host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The symlink existed because the Claude sidecar hardcoded that path, and that
hardcoding came from the bwrap-sandboxed architecture: the jail ro-bound /usr
and could not see the installer's real target in ~/.local/bin. The sandbox is
gone, and claude-manager.ts now resolves the CLI itself — $CLAUDE_BIN, then
PATH, then ~/.local/bin/claude, /usr/local/bin/claude, /opt/homebrew/bin/claude.
Verified before removing rather than assumed:
- the only references left in the tree are the resolver's own fallback list and
this step; nothing in capabilities, no systemd unit, no crontab, no ecosystem
file and no shell rc mentions the path
- the agent sidecar's PATH under pm2 contains ~/.local/bin ahead of
/usr/local/bin, so Bun.which resolves to the installer's target and the
symlink is never consulted
- replaying the resolver in that exact environment with the symlink treated as
absent returns the same path, so it is not load-bearing
- resolveClaudeBin runs at claude-manager module scope, which ES import ordering
puts before user-instance.ts reassigns process.env.HOME — so the homedir()
candidate is evaluated against the real home, not the managed one
The install-and-verify step above is untouched, so a failed claude-code install
is still reported. Only the sudo-owned link into /usr/local/bin goes, a
directory macOS does not ship at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three files, all additive — nothing on master is modified by this beyond the
CLAUDE_BIN change below, and no file is deleted. The branch predates master by
about 180 commits, but it touches nothing master has touched since, so the
merge is clean.
The part that matters beyond macOS is claude-manager.ts. CLAUDE_BIN was pinned
to /usr/local/bin/claude, which dated from the bwrap-sandboxed architecture:
the jail ro-bound /usr and saw nothing else, so the installer's real target
(~/.local/bin/claude) had to be symlinked somewhere the sandbox could reach.
That sandbox is gone, and the hardcoded path left the sidecar unrunnable on any
host without it. It now resolves an explicit CLAUDE_BIN pin, then PATH, then the
locations Anthropic's installer actually writes to — mirroring how OPENCODE_BIN
is already resolved in the opencode sidecar.
ecosystem.mac.config.cjs is deliberately a trimmed set of processes rather than
a mac port of the full ecosystem. It is also stale in two specific ways, left
as-is here and worth fixing separately: it names officer-claude, which master
renamed to officer-anthropic-proxy, and its officer-pty runs
src/servers/api/terminal/pty-sidecar.mjs, which moved to
src/servers/sidecar/pty/index.mjs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The setup wrote /etc/sudoers.d/officer-service with tee and then chmod'd it.
Two problems, both with the same worst case: a malformed or wrongly-permissioned
file there breaks sudo completely, and you cannot sudo to repair it — on a
remote machine that means physical access or a rescue boot.
Generate into a temp file, gate on `visudo -c`, and only then install. Use
install(1) rather than tee+chmod so the content and the 0440 mode land in one
step; tee creates at the default umask first, and sudo refuses to read a sudoers
file with loose permissions, so the old ordering had a window where sudo could
reject its own configuration.
The re-run guard also grepped for the username anywhere in the file, so a
comment mentioning it counted as configured. Match the actual rule instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two guards that checked something other than the state they were protecting.
Section 17 skipped the entire remote desktop setup when `dpkg -s ubuntu-desktop`
succeeded, treating one package being present as proof that seven steps of
configuration had run. A host can have ubuntu-desktop and still be missing GDM
auto-login, the forced Xorg session, the captured EDID and its kernel command
line, and the login-time mode setter — which is exactly what this machine was
on 2026-08-02, while the guard cheerfully reported "skip". setup-desktop.sh is
idempotent throughout, so the guard bought nothing and cost a converged host.
The starship step had the opposite bug: it cp'd over ~/.config/starship.toml on
every run, so a customised config was silently destroyed. The nvim step two
sections down already guards on its config's existence; this now matches, and
distinguishes "absent" (deploy) from "identical" (skip) from "yours differs"
(keep, and say how to take ours).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
cleanup-desktop.sh still described the era when Officer installed a desktop of
its own — XFCE on a TigerVNC Xvnc — so tearing the remote desktop down meant
deleting a desktop nobody else used. That stopped being true when the setup
moved to mirroring the machine's existing session with x11vnc, and the script
was left actively dangerous: it purged dbus-x11 and Brave, deleted ~/.vnc, and
reinstalled gnome-keyring, all of which the current GNOME setup depends on or
deliberately removes. Running it today broke the desktop rather than cleaning
it up.
Rewritten as the actual inverse of setup-desktop.sh:
- removes what Officer added — x11vnc, ~/.vnc, the GDM auto-login and forced
Xorg keys, the forced EDID and its kernel command line, the login-time mode
setter, the legacy officer-vnc unit, the VNC entries in .env
- keeps ubuntu-desktop, gdm3 and dbus-x11, which are the machine's own desktop
and not Officer's to delete
- still purges XFCE and TigerVNC when present, since a host set up by the older
script carries them and they are precisely what this undoes
- Brave is opt-in behind --purge-brave: setup installs it, but by teardown time
it is usually just the user's browser with their profile in it
The GRUB and GDM edits were checked against copies of the real files: stripping
the EDID parameters leaves other kernel arguments intact wherever they sit in
the line, and the GDM revert does not disturb the commented examples that ship
in custom.conf. The package matcher names each xfce-family prefix rather than
globbing '^libxf', which would have taken libxfixes, libxft and libxfont with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Section 17 was still labelled "XFCE + VNC" while the step it runs installs
ubuntu-desktop and is guarded on it, so the heading described a setup the
script had already stopped producing.
Also spell out why setup-desktop.sh disables lightdm: it is not a display
manager this script ever installs, it is residue on hosts set up by an earlier
version that did install XFCE, and left enabled it beats GDM to the seat.
Comments and one echo string; no behaviour change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
scripts/build/runtime.ts had no caller after the build:editor scripts went. Its own usage
text gives away where it came from: --experiments, --tracking, --editor-setup, the same
phantom domain as the docs and examples cleaned up earlier. Nothing else references it —
the `./runtime` export in src/workspaces/types/package.json points at a different file,
which is untouched. helpers.ts stays; dashboard.ts and web.ts import it.
APP_CONVENTIONS.md carried its own copy of "React 19's compiler handles memoization. Never
use useCallback or useMemo." Correcting only CONVENTIONS.md would have left the two
contradicting each other, which is worse than either. Both now say the same thing and one
points at the other for the reasoning.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Email was the one sidecar built inside out. The platform held ~1,800 lines — the per-account
SQLite store, all 14 HTTP routes, account CRUD, resync, IMAP validation — while the 314-line
sidecar was a scheduler that reached BACK into the platform to do anything
(`import { performResync } from '../../api/email/resync'`).
The sidecar now serves its own HTTP listener and announces `email:server`, and
/api/email/* on the platform is createSidecarProxy like every other one: 1,801 lines down
to 22, with no mail knowledge left in it — not a message, not a folder, not a credential.
The routes moved verbatim, Hono and all. http.ts only reconstructs what the platform's
middleware used to provide: `user` on the context, from the X-Officer-User header the proxy
injects (trusted because this server binds loopback), and an error handler that turns
custom-errors into status codes.
The /email/events SSE stream went with them, which removes a whole round trip: the IDLE
watcher used to send `email:new` over the registration socket so the platform could push to
its SSE clients. Those clients are here now, so it calls broadcastEmailNew in-process and
`email:new` is gone from the wire protocol.
DELIBERATELY NOT DONE YET, and left backwards on purpose rather than half-moved:
- The two sync handlers (email-sync 381 lines, gmail-sync 712) still run in the platform's
queue and now import the store from its new home — a platform → sidecar import, which is
the wrong direction and is temporary. Moving them is option (A) from the plan: the sidecar
schedules its own syncs, independent of the platform Jobs list.
- accounts.ts still imports queue/init to enqueue a sync and to report sync status, and
index.ts still carries the queue-over-WS shim that inversion needs.
- The three channel handlers still open the mail store directly rather than asking over HTTP.
Two things worth knowing while testing: a from-scratch sync holds a proxied request open
well past the 60s idle default, hence timeoutSeconds on the proxy; and `gmail-sync` is
hardcoded in all three channel handlers even though the only account is provider=gmail with
auth_type=password, which routes to IMAP — so "sync emails" from a chat channel is
almost certainly already broken, and folds into the next stage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
slskd's browse response is flat: every folder is a full backslash-delimited
path. Rendered as-is, a filter for "pogues" gave 26 rows that all began with
the same 31 characters, and the real hierarchy — which is the only way to tell
an artist folder from an album folder — was invisible.
The shape is now derived once, at ingest, in the sidecar: buildTree() links
each path to its parent, synthesizes any ancestor slskd omitted (measured:
exactly one missing across ~30k folders on two real peers, but a single gap
would strand a whole subtree), and rolls subtree file counts and sizes up
bottom-up. A parent's own files are usually just cover art, so the number
worth showing on a collapsed row is the subtree's.
Storing the shape rather than recomputing it is what lets the UI open one
level at a time. Levels are still paged, because fan-out is brutal — the
widest folder measured has 1,181 children.
Filtering keeps the tree instead of falling back to a list: the search route
returns matches plus every ancestor, and the UI renders that skeleton
pre-expanded, so you see where a hit lives. Matching runs against the whole
path, so a matched folder implies its descendants match too and a matched
subtree arrives complete. The match cap is reported in the payload and shown
in the UI rather than passed off as the whole answer.
The two existing snapshots were backfilled by scripts/rebuild-soulseek-tree.ts,
which runs the same buildTree + finishSoulseekBrowse the ingest path runs — no
second implementation to drift, and no peer contact needed. Kept for the next
time the tree shape changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
setup.sh targets an ubuntu server and is left untouched. this is the
laptop equivalent: file browser, claude/opencode chat, terminal. no go,
rust, cliamp, pulseaudio, neovim, shell dotfiles, vnc desktop, sudoers
grant or power management — 510 lines against 1033.
every step is optional, prompted, and presettable non-interactively with
SETUP_* variables, so it also works as a repair tool for one piece.
nothing calls sudo. node@22 goes on PATH with brew link --force, pm2 into
~/.local, and the claude cli no longer needs a /usr/local/bin symlink now
that the sidecar resolves it.
postgres is detected before anything is installed — a server already
listening on 5432 (docker) is used as-is.
notes on the differences from setup.sh:
- set -e is on, but every optional step is guarded, so a failure warns and
the run continues to a summary instead of aborting mid-way.
- .env is written 0600 and the generated JWT_SECRET is length-checked
before use, since jwt.ts throws on anything under 32 chars.
- PUBLIC_BUILD_ENV=development, which is what lets plain http://localhost
work with no reverse proxy in front.
- xcode command line tools are not required to build: node-pty and argon2
both ship darwin prebuilds. they still matter for git on a fresh mac.
ecosystem.mac.config.cjs drops officer-vnc (x11vnc needs xorg),
officer-email and officer-music, and pins cwd on every app so bun picks
up .env wherever pm2 is started from.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
x11vnc mirrors :0, but with no monitor attached the connector has no EDID and no
CRTC, so GNOME renders nothing and the remote desktop is black. Reproduced the
hard way on this box: the desktop only worked because the session had started
while a screen was plugged in, and survived exactly until the next restart.
The step captures a connected monitor's EDID, installs it as
drm.edid_firmware with video=<connector>:1920x1080e so the connector reports
permanently attached, and adds a login-time hook to raise the resolution — the
replayed EDID's *preferred* mode is the captured panel's native one, which can be
tiny, and GNOME picks preferred. monitors.xml is the documented override but its
monitor matching did not take.
The EDID can only be captured from a screen that is plugged in, so a headless run
skips with instructions rather than pretending to succeed. Re-running is safe:
GRUB is left alone once the argument is present, and the mode setter no-ops when
the mode is already right.
officer-set-display.sh discovers the output and the largest mode within a cap
rather than hardcoding either, since setup runs before X exists and cannot know
them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
setup.sh installed pm2 but never ran anything with it, so a fresh install
finished with every dependency in place and nothing listening. That is not
cosmetic: /desktop returns 503 until officer-vnc is connected, and chat needs
officer-claude.
Adds a step that runs `pm2 startOrRestart ecosystem.config.cjs`, saves the
process list, and enables the boot unit when it is not already there. Using
startOrRestart rather than start means apps added to the ecosystem since the last
run get picked up — officer-music is in the ecosystem on this box but was never
running, for exactly that reason.
The verification block now reports which services are up, with the names read
from ecosystem.config.cjs so the list cannot drift as sidecars are added.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>