9f15de3448f38863798d26b33ec1b8545356398f
1243
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7a5cf89819 |
port the DNS section, as a choice rather than a decision
The original hardcoded Cloudflare plus Google with no way to say otherwise, and
rewrote /etc/systemd/resolved.conf wholesale — discarding DNSSEC, DNSOverTLS,
Domains and Cache if anything had set them, without mentioning it had. The
settings are a drop-in now, and the resolver is picked from a list with a
"keep what is there" that is the default.
The part worth having explicit is which layer is being changed. With
systemd-resolved there are two:
per-link what DHCP handed each interface, and what Tailscale installs on its
own. These answer for that link's domains — the provider's internal
names, the tailnet — and are printed by this step precisely to show
they are NOT being touched. Overriding them is how private
networking quietly stops resolving.
global the resolver used when no link claims the query. This is the one
the step sets.
On this host that distinction is live: eth0 has Hetzner's resolvers and
tailscale0 has 100.100.100.100, which is what answers ts.pastilhas.dev. Both are
left alone.
The drop-in is named 99- because systemd reads drop-ins in lexical order and the
LAST value wins. That is the opposite of sshd, whose drop-in three files away in
this same directory has to sort FIRST. Both are stated where they are written,
because getting it backwards fails silently in either direction.
resolv.conf is checked for actually pointing at resolved's stub before the
drop-in is trusted to do anything — a machine where something replaced the
symlink with a static file bypasses resolved entirely.
Resolution is tested afterwards rather than assumed. A resolver that does not
answer makes every later step fail for a reason that has nothing to do with it,
so that failure is reported and recorded rather than swallowed.
The choice names what each provider actually is, including that a resolver sees
every name the machine looks up.
Verified both paths against this host: keep reports unchanged, Quad9 renders the
right addresses, and the per-link display shows Hetzner and Tailscale correctly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
192293cdca |
offer to add an ssh key even when one already exists
The zip option was already gone — it never came across in the port, since the file is key material that cannot live in the repository and the script no longer sits next to it. Pasting a public key was already the first option. What was missing is the case where the account HAS a key: the section went straight to hardening, so there was no way to authorise a second machine, a rebuilt laptop or anyone else, and the original had no way to do it at all. The same menu is now offered either way. What differs is whether it can be declined without consequence: with no key, declining means the hardening below refuses too, and the run says so rather than quietly moving on. confirm() takes an optional default so this one can be [y/N]. Most questions in this script are "do the thing you already asked for" and Enter should mean yes; a genuine extra defaulting to yes is how people end up agreeing to things by reflex. A pasted key is trimmed before validation. Copying from a terminal or a password manager routinely brings leading or trailing whitespace, and ssh-keygen will not parse a key with it attached — which would have read as "that is not a valid key" for a key that is perfectly fine. Also corrected the reason unzip is in core utils, which still said it was there to open ssh-keys.zip. Verified: the add-another prompt appears and defaults to no, a whitespace-wrapped key is trimmed and accepted, and the already-hardened path is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9591f917f5 |
port ssh keys and hardening as one section, and make the hardening actually work
They were two sections, and being two is what let the second lock you out of a machine the first had failed to put a key on. Step 8 could warn-and-skip — no ssh-keys.zip, or an unrecognised menu choice, since its case had no default arm — and still mark itself done; step 9 then disabled password authentication and root login regardless. No key, no password, no root, on a box that may be in a datacentre. Nothing here turns off password authentication without first confirming a usable key is in place, and the refusal says why rather than skipping quietly. The hardening also did not do anything on a modern Ubuntu, and could not be seen not to: It sed'd /etc/ssh/sshd_config. Ubuntu includes /etc/ssh/sshd_config.d/*.conf from line 12 of that file, and sshd takes the FIRST value it obtains for a keyword rather than the last. Cloud images ship 50-cloud-init.conf containing `PasswordAuthentication yes`, read long before the line the sed edited. The run reported "SSH hardened" and password login stayed on. The settings now go in a drop-in named 01-machine-setup.conf, which is the only placement that wins under first-value-wins. It also sed'd ChallengeResponseAuthentication, renamed to KbdInteractiveAuthentication in OpenSSH 8.7. On 24.04 the old name is nowhere in the file, so that substitution matched nothing at all. State is read with `sshd -T`, which reports what sshd resolves across the main file and every drop-in — reading the config files tells you what is written, not what wins. Keys are counted by asking ssh-keygen to parse authorized_keys rather than by counting lines: comments, blanks and a half-finished paste all look like lines, and "there is a file" is not "there is a key that works". A pasted key is validated before it is stored, and matched on the key body rather than the whole line, so re-running does not authorise the same key four times over four runs. sshd -t validates the new config before anything is reloaded, and the drop-in is restored or removed if it does not parse — a config sshd refuses is a machine with no ssh after the next restart. Reload rather than restart, so the session this is running over is not the experiment, and the run says out loud to test a new connection before closing the current one. Generating a keypair now says the obvious thing the original did not: the private key is on the server, and a private key living on the machine it opens is a spare copy of the lock rather than a second factor. Verified against this host (1 key, already hardened, correctly does nothing) and with sshd_effective stubbed to a fresh-cloud-image state — the guard refuses and harden_sshd is never reached. Also verified key validation, dedup and 0700/0600 permissions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b1cc916258 |
detect root by uid, not by the name root
A provider whose image logs you in as "ubuntu" at uid 0 would have walked straight past the previous check, which compared the string. What makes an account root is uid 0; "root" is only the usual label for it. Two places now ask id -u rather than comparing names: the answer — an account at uid 0 is refused whatever it is called, and says which case it is rather than a bare "not root" the invoker — the warning about working as root fires when SUDO_USER is unset OR when SUDO_USER is itself uid 0. The second is the one that hides: sudo from a uid-0 account sets SUDO_USER to something that reads like an ordinary user and is not. The EUID check that requires the script to run as root was already uid-based and is unchanged. Verified by creating a real uid-0 account named ubuntu on this box: refused with the uid named, where the name check accepted it. That account has been removed — userdel refused it at first because it matches by uid and saw PID 1 running as uid 0, so -f was needed, and deliberately not -r, since its home was /root. Confirmed afterwards that root, /root, root's shadow entry and sudo are all intact and that root is once again the only uid-0 account. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d9d4085033 |
warn against working as root at the username prompt
root was already refused as an answer, but only the refusal said so — and only after somebody typed it. The advice now comes with the question, along with why: no safety net, a typo in a path that deletes instead of refusing, and nothing to distinguish you from a process that got out of hand. An extra warning when SUDO_USER is unset. That means the script was started as root rather than through sudo, which usually means root is how they log in — the exact situation the general advice is about, and the one where general advice is easiest to assume is aimed at somebody else. It says so plainly instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0eb3e3a85 |
create the user account first, and fix the sudoers filename bug that found
The user account section is now the first thing that acts, ahead of disk space.
The reason is a real defect, not tidiness. If the account does not exist yet,
USER_HOME is a path that is not there — and the ballast offers to put its file in
it, where ballast_create's `mkdir -p` runs as root and creates /home/<name> owned
by root:root. adduser afterwards finds the directory already present and does not
populate or chown it, so the account ends up with a home it cannot write to.
Making the account before any step can write into its home removes the ordering
entirely.
USER_HOME is re-read from getent after adduser runs. Until that point it is the
/home/<name> guess, because there is nothing to look up; adduser is free to have
used something else and every later step writes there.
Found while testing that, and worse than the thing it was testing:
local user="$1" dest="/etc/sudoers.d/99-${user}-nopasswd"
bash expands ${user} before the assignment to user has happened, so dest came out
as /etc/sudoers.d/99--nopasswd with the name missing. The rule inside was correct,
which is what made it invisible — visudo passes, sudo works, and the account
really does get passwordless sudo. What breaks is everything around it: every
account granted this way writes to that same file, so a second grant silently
overwrites the first and revokes it; and has_passwordless_sudo looks for
99-<user>-nopasswd, never finds it, and re-grants on every run forever.
Split into separate declarations, with the reason recorded where it happened, and
an empty username is now refused outright. Scanned the other lib files for the
same shape — the remaining multi-assignment locals only read positional
parameters, which is safe.
The stray /etc/sudoers.d/99--nopasswd this created on the dev box during testing
has been removed and visudo -c re-verified.
Verified with a real throwaway account: correctly reports not-granted before,
writes 99-msdemo-nopasswd as root:root 0440 with the right rule, reports granted
after, and leaves sudoers valid.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fbb90c917d |
default the username to whoever ran sudo, and look up their real home
Three changes to the question at the top of the run. USER_HOME is looked up rather than assumed. The original built "/home/$USERNAME", which is only the usual answer — an account created with a different home, or one whose home was moved, had every later step writing to a directory that was not theirs. getent passwd knows; the /home guess remains only as the fallback for an account that does not exist yet, where there is nothing to look up. The default is now whoever invoked sudo. On a re-run, or on a machine that is already somebody's, that is the answer every time, and retyping it is a chance to typo it into creating a second account. root invoking the script directly offers no default, since root is never the account being set up — and is refused if typed. The name is validated against the portable shape of a Linux account name before anything else happens. Letting adduser refuse it later means several questions have already been answered against a name that was never going to work. It also no longer goes through prompt_value, which obeys any environment variable matching the name it is filling in. USERNAME is set by some login environments, and a variable this script silently takes as an answer should not be one that might already be set for unrelated reasons. SETUP_USERNAME is the explicit override. The tmux config write moved out of the user section entirely. It was there only because the original copied it right after adduser, where $USER_HOME first exists. It is a dotfile and belongs with .zshrc and the starship config in the shell section. Verified: defaults to the sudo invoker, resolves daemon's home to /usr/sbin rather than /home/daemon, falls back to /home for an account that does not exist, and rejects a name with a space, a leading digit, one over 32 characters, and root. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3fb0e5c887 |
port the user account section, and stop clobbering files in the home
Two real defects fixed on the way across.
The sudoers write was in the wrong order. The original echoed the rule straight
into /etc/sudoers.d, validated it afterwards, and chmod'd it later still. A
malformed file there breaks sudo COMPLETELY — and you cannot sudo to repair it,
so on a remote machine that is a rescue console — and so does one with loose
permissions, because sudo refuses to read its own configuration. Both of those
windows were live in the original ordering. grant_passwordless_sudo now writes a
temp file, runs visudo -c against it, and only then places it with install(1),
which applies the content and the 0440 mode in one step. Nothing reaches
/etc/sudoers.d that has not already been validated.
The .tmux.conf copy overwrote whatever was in the home on every run. lib/files.sh
adds the two shapes that stop this whole class of thing:
install_config installs when absent, does nothing when identical, and keeps
what the user wrote when it differs — printing the cp to take
ours, so the choice stays theirs
append_once wraps a block in named markers so a second run recognises its
own work; also lets a human see which lines came from this
script and remove them as a unit
append_once is what the five unguarded `cat >>` into .zshrc need when those
sections are ported — a second pass currently duplicates the starship init, the
nvim PATH, bun, deno and the aliases.
Passwordless sudo is asked separately from creating the account, because it is a
security posture rather than part of making a user, and the cost is stated: a key
that can log into this account is root without a further step. Officer's actual
requirement is stated too — os-user-shell.ts runs `sudo -n`, and a prompt it
cannot answer surfaces as a permissions error rather than a question — and
refusing records that consequence in the summary instead of a bare "skipped".
Verified: all three install_config outcomes, append_once writing exactly once
across two runs, visudo rejecting junk before anything is installed, and the
section reporting correctly against this host's existing account.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
458510a0a8 |
scope sleep to homelab, and port the boot hang fix
Sleep and suspend is homelab only now. The two exclusions are for different reasons and both are stated in the run rather than left implicit: dev — a laptop should sleep; disabling it is a hot bag and a flat battery. vps — not merely unnecessary, harmful. A virtual machine has no lid and no power button, but the provider's Shut down control works by sending an ACPI power button event. HandlePowerKey=ignore makes the VM ignore it, so graceful shutdown requests silently do nothing and the instance is hard-killed instead. systemd defaults that key to poweroff for exactly this reason. The boot hang fix is everything except vps, where systemd-networkd genuinely manages the network and the unit is load-bearing. Rather than asking whether boot "feels slow" — a question people answer from memory of the worst time it happened — the step prints what the unit actually cost on this boot, from systemd's own accounting. On this host that is 14ms, which ends the discussion. On a NetworkManager desktop it is two minutes, which also ends it. On dev the wording says outright that a small number here means there is nothing to do. The original's live guard is kept and is what actually decides: NetworkManager active and networkd not. Anything else, including "cannot tell", is left alone, and the reason is printed. The warning against disabling systemd-networkd outright is carried across into the library, where the alternative would be attempted. Verified all three roles on this host, which is a networkd machine: homelab and dev both correctly refuse and explain, vps skips as not applicable, and the timing helper reads 14ms out of systemd-analyze. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c87127ff93 |
port the sleep and suspend section
The original ran unconditionally, so a laptop that went through it stopped suspending — a hot bag and a flat battery. It is a server concern: dev is skipped with the reason printed, like the ballast. Four things wrong with it beyond the role: HandleLidSwitchDocked was never set. A laptop used as a homelab server, docked and closed, still suspends — which is the exact machine this setting exists for. Added. RuntimeDirectorySize=10% was set alongside the sleep handlers. It is the size of /run, has nothing to do with sleeping, and 10% is systemd's own default, so the line never did anything. Dropped. systemd-logind was restarted on every pass whether or not anything changed, disturbing live sessions for nothing. The step now checks first and does not reach the restart when the machine is already configured. (The platform's own scripts/setup-old/setup.sh already had this guard; the machine script did not.) The settings were sed'd into logind.conf in place. They are a drop-in at /etc/systemd/logind.conf.d/99-machine-setup.conf now, so what this script set is one file that can be read or removed on its own. Current state is printed before anything is asked — whether the targets are masked, and what the lid, idle and power-key handlers actually do. logind_effective reads the main file and every drop-in and takes the last match, since a drop-in overrides logind.conf; reading only the main file reports a configured machine as unconfigured. The power-button consequence is stated rather than left to be discovered: after this, pressing power physically does nothing and a clean shutdown is `sudo poweroff`. WSL has no logind and cannot suspend, and says so. Verified both roles on this host, which the old script had already configured — correctly reports the targets masked and the handlers set, and correctly reports itself not fully configured because HandleLidSwitchDocked is missing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6ffd3534bd |
do not offer the ballast on a dev machine, and route its alerts through one place
Servers only now. On a machine you sit at, a filling disk announces itself — the editor refuses to save, the browser complains — and you are there to deal with it. The reserve is for the box nobody is watching, where the first sign is a service that stopped working hours ago. Skipped rather than asked, but said out loud with the reason and recorded in the summary. A section that silently produces no output is indistinguishable from one that failed. The cron this section installs was already there and is unchanged: /etc/cron.d runs the checker as root every ten minutes, and it deletes the ballast when free space falls under the threshold. What changed is where its message goes. Both alerts now run through one notify() inside the generated checker rather than calling logger directly, so there is a single place to add a second channel. Today it is still syslog only — the message lands in the journal and nowhere else, so nobody learns about it until they go looking, which is precisely the wrong moment. Push, mail or Officer's own notify sidecar hook in there. It also echoes to stderr now, so running the checker by hand shows the message instead of appearing to do nothing. Verified: dev reports not-applicable and asks nothing, vps still asks and records a refusal, and the regenerated checker parses and reports status. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d19dc5a92a |
ask where the ballast goes and how big it is
Three questions instead of one, because the two the original never asked are the
two that decide whether the thing is useful.
1. Do you want one, with the explanation first.
2. Where. Home (easiest to find again months from now), beside Officer, or a
path typed in. This is not tidiness: the checker measures its own directory,
so a ballast only protects the filesystem it sits on. Choosing where it goes
is choosing which mount is covered.
3. How much, as 5/10/20% — with the actual numbers, and with what would be LEFT
rather than only what is taken:
[1] 5% — reserves 2.9GB leaving 54.3GB free
[2] 10% — reserves 5.8GB leaving 51.4GB free
[3] 20% — reserves 11.5GB leaving 45.7GB free
A percentage on its own is unanswerable. The number that decides it is the
one on the right: the reserve has to be big enough to matter and small
enough not to be the thing that filled the disk.
The size is computed against the filesystem the chosen path lands on, after the
location is known, so the percentages are of the right disk. ballast_free_kb
walks up to a directory that exists, since nothing has created the target yet.
An existing ballast in either default location is found and left alone rather
than a second one being made beside it.
Verified end to end in a temp home: created at the chosen path, 2.9G for 5% of
57.1GB, checker installed and reporting it present.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d4b9b7b334 |
ask where Officer should be installed, in pre-flight
OFFICER_ROOT, defaulting to <user home>/officerdev. One directory holding the four things Officer is made of, per docs/sidecar-app-store.md — the app, its data, the item store, and any containers the app store provisions — so the whole installation can be moved, backed up or deleted as a unit. Asked at the start with the other questions rather than at the point it is first needed. It decides the shape of several later steps: where the repository is cloned, where DATA_PATH sits beside it, and which filesystem the app store's bind mounts come out of. Asking once up front also means the run can be described before it starts rather than discovered as it goes. A leading ~ is expanded explicitly. It arrives as a literal from a read or an environment variable — nothing expands it there — and would otherwise create a directory actually named "~" in whatever the working directory happened to be. Relative paths are refused with the value named, and a trailing slash is trimmed so the path composes cleanly with what gets appended to it. Nothing creates the directory yet; that belongs to officer-setup. This records the answer and reports it, including whether it already exists. Verified: Enter takes the default, ~ expands, trailing slash trims, OFFICER_ROOT in the environment skips the prompt, and a relative path fails with the value named. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6947284f1b |
check the root filesystem is actually using the whole drive
New first section, before anything else that changes the machine, because the
swapfile and the ballast both size themselves from free disk.
Ubuntu Server's installer on its defaults gives the root logical volume a fixed
size and leaves the rest of the drive as unallocated extents in the volume group.
On a 2TB disk that is a ~100G root with nothing to indicate a problem: lsblk
shows the whole drive, df shows 100G, and the two are never seen side by side
until the day it fills. Growing a virtual disk at a provider leaves the same
shape one layer down, and so does resizing a partition without telling the
filesystem inside it.
Three layers, any of which can be the short one, so all three are measured and
printed together:
drive: 76.3GB /dev/sda
volume: 76.1GB /dev/sda1
filesystem: 76.1GB ext4, mounted at /
Seeing them in one place is most of the value. The fix is then whichever layer is
short: lvextend for free extents, growpart for a partition that stops early
(followed by pvresize and lvextend when LVM is in the way), or resize2fs alone
when only the filesystem is behind.
Only ever grows. Nothing here shrinks, creates or deletes a partition, and ext4,
xfs and btrfs all grow while mounted — so no unmount, no reboot, and a failure
part-way leaves a smaller filesystem on a larger container, which is the state it
started in.
growpart is the authority on whether a partition can move — it exits 1 with
NOCHANGE when the partition already reaches the end — but it comes from
cloud-guest-utils, which is not on every image. Installing a package purely to
ask a question is too eager, so plain arithmetic on the device sizes decides
whether it is even worth looking, and only then is growpart fetched.
A gigabyte of slack before anything is reported: a filesystem is always slightly
smaller than its container, and reporting journal and reserved-block overhead as
reclaimable space would make this section cry wolf on every machine.
Verified on this host: plain ext4 partition filling its disk, correctly reports
nothing to reclaim, and every helper returns the right device and size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9b0e05bfd8 |
add three resource-pressure sections, each asked rather than assumed
Swap covers memory pressure. These are its neighbours: 8. Emergency disk ballast the same valve, for disk 9. earlyoom what happens when swap runs out too 10. inotify watch limit the silent one All three follow the rule this script now works to: the role sets which way the recommendation points, never whether the question is asked. A dev machine is still offered the ballast, with the recommendation pointing the other way; a server is still offered the inotify raise, because anything running `bun --watch` or serving a file browser is a watcher too. The ballast is section 22 of the original, moved up beside swap where it belongs and moved out of the user's home. The original wrote the checker into $USER_HOME/.local/bin and ran it from a root cron — a root cron executing a script in a directory its owner can write is a privilege escalation waiting to be noticed. Moot on a box where that user already has passwordless sudo, but wrong. Both the checker and the file are in root-owned system paths now. Two bugs found by running the generated checker rather than reading it: It df'd the ballast's own directory, which does not exist before the ballast is created — and with `set -euo pipefail` that meant cron mailing an error every ten minutes. It now walks up to a directory that exists, and the installer creates the directory itself rather than depending on the create step. The inotify text claimed a default of 8192. This host is at 29461: Ubuntu raised it, and stating a number the reader can see is wrong on their own screen undermines the rest of the explanation. It now describes the failure instead and prints the machine's actual value. earlyoom is a distro package and a systemd unit, so it is checked with `systemctl is-active` and reports honestly when it installs but fails to start. Verified: checker --status and its no-op path both exit 0 with no directory present, and the helpers report correctly against this host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
716a6e2750 |
port the swap section, and size it against the disk
Four things the original got wrong, all of which only show up on a machine that is not this one: It detected swap with `swapon --show | grep -q '/'` — a test for a swap FILE. A machine using zram or a swap partition reports no swap at all, and the step would add a swapfile beside working swap. Reads SwapTotal from /proc/meminfo now, which covers every kind. It never looked at free disk. On a VPS with 4G free and 16G of RAM it would fallocate 8G, fail, and take the run down under `set -e`. The recommendation is now capped by what is actually there, keeping 5G back, and refuses rather than shrinking to something useless. fallocate was assumed to work. It produces a file that btrfs and zfs will not swap on, so dd is the fallback — slow, but it always works. swappiness was written by sed'ing /etc/sysctl.conf in place, tangling it with whatever else lives there. It is a drop-in at /etc/sysctl.d now, so what this script set is visible as its own file. Role-dependent, which is the first use of MACHINE_ROLE: swappiness 10 on a server, where swapping is the emergency valve and a page fault on a request path is latency somebody is waiting for; the kernel default of 60 on dev, where swapping out an application nobody has touched in an hour is exactly what you want. WSL is left alone entirely — WSL2 runs its own managed swap inside the VM, and a swapfile written here is wasted disk the kernel will not use. Sizes are reported rounded rather than floored. A 4 GiB swapfile is 4194300 kB, which floors to 3 and reads as though a gigabyte went missing. Free disk stays floored, deliberately: it decides how much to allocate, and rounding up invents space. Verified on this host — 4G RAM, 4G existing swap correctly detected and left alone — and with the disk check stubbed at 6G free (caps to 1G) and 3G free (refuses). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5948f9673 |
refresh the package index in pre-flight, not inside a skippable step
The refresh lived inside "System update", which `step` skips when its name is already in the progress file. So a resumed run — the common case, since that is what the progress file is for — installed core utils, added the fastfetch PPA and set up the Docker repo against whatever the index happened to say hours or days earlier. On a box left overnight that is a stale index and a "package not found" somewhere unrelated. It now runs in pre-flight, unconditionally, before any step exists to skip it. `apt-get upgrade` stays where it was and stays confirmable: refreshing the index changes nothing on the machine, upgrading is the one thing that does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6480788979 |
port the timezone section
The original took whatever was typed and handed it straight to timedatectl. An unknown zone — a typo, a guess at the spelling — fails there, and under `set -e` that takes the whole run down four steps in. Names are now checked against /usr/share/zoneinfo before use, and a bad one just re-asks. It also never showed what the machine was already set to, and defaulted to option 1 (UTC) on Enter, so pressing return on a correctly-configured box silently moved it. Now the current zone is printed, Enter keeps it, and a zone equal to the current one reports nothing to do rather than setting it again. timezone_current reads three sources — timedatectl, /etc/timezone, then the /etc/localtime symlink — because they differ in availability rather than in answer: timedatectl needs systemd, /etc/timezone is Debian's, and the symlink is the one that is always there. timezone_set writes through timedatectl where there is a systemd to talk to and the files directly otherwise, which is what it would have written anyway; that is also the WSL path, where timedatectl exists but does nothing. Europe/Berlin added to the shortlist; TIMEZONE in the environment answers the prompt ahead of time and is validated the same way, failing early with the bad value named. Verified detection (UTC here), validation of four names, and the env-var rejection path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
163f8d5899 |
port the locale section
Was three unconditional lines that ran on every pass and reported success either way. Now it checks, says what it found, and asks. The original tracked one fact where there are two: what a new login shell is told to use LANG in /etc/default/locale whether that locale actually exists whether it has been generated Setting the first without the second is what produces "setlocale: LC_ALL: cannot change locale" on every ssh login and every perl invocation. They fail differently, so the step names whichever one is actually missing rather than reporting a flat "locale not set". Also fixes two things the original would have hit on a minimal image: locale-gen comes from the `locales` package, which cloud base images do not ship and which is not in core utils. It is installed on demand rather than assumed, instead of failing with "locale-gen: command not found". The locale is uncommented in /etc/locale.gen rather than only passed to locale-gen as an argument. A locale generated by argument alone disappears the next time anything regenerates from that file. `locale -a` prints en_US.utf8 where the configuration spells it en_US.UTF-8, so both sides are folded before comparing — a literal match reports a working locale as missing. LOCALE in the environment overrides the default. pacman, dnf and brew branches are written but unreachable while the pre-flight gate is apt-only; macOS has no system locale to set and says so. Verified both paths on this host: en_US.UTF-8 reports already set and generated, pt_PT.UTF-8 correctly reports both facts missing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1beb357f2e |
drop the ominous wording from the system update prompt
"the only step that changes software already on this machine" is true, and reads like a warning about something dangerous rather than a description of apt upgrade. The reasoning stays in the section comment, where it explains why this is its own step; the prompt just says what it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
454faf5406 |
list what the system update would actually upgrade
The section asked "Proceed?" without saying what it was proposing to change — the one question in the script where the answer matters most, since it is the only step that moves versions of software already on the machine. pkg_upgradable now names them, from `apt-get upgrade -s`: the same calculation the real run does, as opposed to `apt list --upgradable`, which also lists packages held back that would not actually move. Nothing to upgrade means no prompt at all, and the summary says so rather than claiming an upgrade happened. The list is capped at 25 with a count of the rest, because a box untouched for months lists hundreds and a wall of names is no more informative than the number. Verified against this host (0 upgradable, so it reports current and does not ask) and with a stubbed 40-package list for the cap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f896d4882f |
one section per concern, each announced and confirmed before it acts
Three sections where there were two, and none of them touches the machine until you say so: 2. System update upgrades what is already installed 3. Core utils what the distribution provides 4. Command-line tools lazydocker, lazygit, starship, fastfetch The split matters because these are different kinds of change and deserve separate answers. System update is the only step in the whole script that moves versions of software already on the machine; core utils only ever adds what is absent; and the four tools are upstream binaries the distribution does not ship at all. Previously the update and the core packages were one step and the tools were tacked onto the end of it, so agreeing to "essentials" meant agreeing to all three at once. Every section now prints what it will install and what it is leaving alone, then asks. Enter means yes — unlike the machine-role question, which has no default, because these are "do the thing you already asked for" and making twenty of them require a deliberate keystroke would train people to hold the y key down. ASSUME_YES=1 answers all of them for an unattended run, and EOF fails with that named rather than spinning. Refusing is recorded rather than glossed: LAST_SKIPPED feeds the summary, so a declined section reads "Core utils: SKIPPED by request — cowsay neofetch" instead of quietly reporting nothing installed. Nothing to install means no prompt at all — there is nothing to agree to. announce_plan takes the array NAMES rather than their contents, because once a list has been through word splitting an empty one cannot be told from a missing one. Verified all three paths: accept, refuse, and nothing-to-do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5ef9f7662 |
report what a section actually installed, not what it was asked for
The summary claimed credit for everything in a section's list, including the packages it had just decided to leave alone — so a run that installed nothing still ended with "Command-line tools: lazydocker lazygit starship fastfetch". The announce above it said "nothing, all present" in the same breath. pkg_install and tools_install now record LAST_INSTALLED and LAST_KEPT, and summarise_last turns those into one honest line: Core packages installed: btop tmux (17 already present) Command-line tools: already present, nothing installed Verified all three shapes — everything present, nothing present, and mixed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dd655577e3 |
make the machine-role question require an answer
No default, and it is the only question in the script like that. A guessed default is right often enough to be trusted and wrong in exactly the case that costs the most — pinning a static IP on a rented box, or leaving the firewall open on one. Every branch downstream is about what this machine is exposed to, so it is worth one deliberate keystroke rather than an Enter. Empty and unrecognised answers re-ask rather than aborting; a failed read means EOF rather than a wrong answer, and fails with the environment variable named, because otherwise the loop spins forever the first time this runs unattended. Drops guess_machine_role, which existed only to supply that default. default_iface stays — the static IP section needs it when it is ported. MACHINE_ROLE in the environment still answers it ahead of time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cb243eed9 |
port into a clean script instead of editing the original in place
Your call, and the right one. Editing in place let mis-grouped code sit unnoticed until it scrolled past in a live run — which is exactly how the four upstream binaries buried in "System Update & Essentials" were found. Porting forces the question of where each thing belongs before it runs, not after. machine-setup.sh now contains only what has actually been worked through: pre-flight, system update and core packages, command-line tools, and the summary. 1149 lines down to 172. The sections still to come are listed in a NOT PORTED YET block, in order, and each arrives as its own commit. The original is beside the other superseded scripts as scripts/setup-old/setup-ubuntu.sh — verified byte-identical to the live /root/ubuntu-setup copy — so porting reads from a file in the repo rather than from root's home. Two claims trimmed from the ported summary, because they were true of the old script and not of this one yet: it reported the shell as "zsh (Oh My Zsh + Starship)" unconditionally, and told you to reconnect as a user it had not created. Replaced with what pre-flight actually knows — system, role, user. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7622239949 |
split the upstream binaries out of the package section
lazydocker, lazygit, starship and fastfetch were buried inside "System Update &
Essentials", after the package install and with no announcement — so a run
appeared to be installing system packages and then started pulling tarballs and
printing a five-shell starship tutorial. They are a different thing: upstream
binaries on their own release cadence, not anything the distribution ships. Now
their own step, announced in the same shape as the package section.
Each is checked before it is fetched. The original re-ran every installer on
every run, which is why a machine that already had starship got it reinstalled
along with its "add this to your ~/.zshrc" instructions — advice this script
does not want followed, since it writes the shell config itself. Its output is
now dropped; errors still surface.
Two real bugs fixed on the way:
lazygit's asset name was hardcoded to x86_64, so on arm64 the download 404s
and tar fails partway through the run. It now maps ARCH, and spells the
architectures the way lazygit does rather than the way we do.
The version was extracted with `tr -d 'v'`, which deletes every v in the
string rather than the leading one. `${version#v}` instead.
fastfetch stays a package but stops assuming the PPA is needed: Ubuntu picked
it up in 24.10, so the repository is now checked first and the PPA added only
where the archive has nothing. Verified on this host — noble genuinely has no
candidate, so the PPA is still the only source here.
Verified both branches of tools_install by stubbing the presence check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bba6d854bc |
stop the run at the end of the rewritten sections
A WIP boundary after section 2 so the finished part can be run start to finish on its own, without the untouched sections below acting on the machine. It moves down as each section is worked through and goes away when the walk ends. Also ignores .setup-progress, which the script writes beside itself and is per-machine. The exit message names it, because with it in place a second run skips section 2 and the rewritten part cannot be re-felt from scratch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dbdef23d29 |
install what is missing and keep what is there, per package manager
lib/packages.sh, and section 2 wired to it.
The rule it exists to enforce: `apt-get install <present-package>` is not a
no-op, it upgrades the package if the repository has a newer one. On a machine
somebody already uses that silently moves a version they chose, and a setup
script is the last thing that should do that behind their back. pkg_install
queries the package database first and names only the genuinely absent packages
on the command line — a package already installed is never passed to apt at all.
It also says so out loud, every time, because a provisioning run should not be
opaque about what it is doing to the machine:
:: Core packages — installs what is missing, keeps what you already have
already here: curl ca-certificates gnupg git jq …
to install: btop tmux
Section 2's flat list of 19 is now pkgs_core(), split per package manager rather
than through a canonical-name table with overrides. The names genuinely disagree
(build-essential/base-devel, fd-find/fd) and three of them are not packages
elsewhere at all — apt-transport-https, lsb-release and software-properties-common
are apt concepts that exist to let later steps add the Docker repo and the
fastfetch PPA. A `case $PM` shows what each system actually gets, in one place.
Of those 19, six are load-bearing and the rest are the environment. Only
build-essential reaches beyond itself: it is a meta-package, so on a box with a
pinned gcc it pulls the distribution default alongside. Noted where it is
declared; it is the first thing to move out of core if that ever bites.
apt-get upgrade stays, but as its own announced step — it is the one place that
deliberately moves versions, rather than something that happens as a side effect
of asking for a tool.
DEBIAN_FRONTEND=noninteractive and NEEDRESTART_MODE=a now live inside the
helpers. needrestart has been on by default since Ubuntu 22.04 and stops to ask
which services to restart, which is how an unattended run ends up silently
waiting for a keypress.
dpkg-query on the status field rather than `dpkg -s`, which also succeeds for a
package removed but leaving its config behind — that state would read as present
and never be reinstalled.
Verified against this host's real dpkg database: all 19 report present, and a
mixed list correctly passes only the absent ones through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
41ff8030e9 |
ask what the machine is for, once, in pre-flight
MACHINE_ROLE is homelab, vps or dev, and several steps have a different right answer per role with no way to work it out themselves: whether the address is yours to pin (static IP), whether the box faces the open internet (fail2ban, SSH hardening, UFW), and whether it is allowed to sleep (suspend, logind). Asked in pre-flight rather than at each point of use. The steps that care run from swap through to the firewall, and being asked "is this a VPS?" for the fourth time halfway down a provisioning run is how people start answering without reading. The default offered is guessed from whether this machine's own address is in RFC1918 space, which beats asking whether it is virtualised — a homelab is very often a VM on Proxmox and would be misread as rented — and is the same fact most of the branches turn on anyway. A graphical session means dev; so does macOS. It is only ever a suggestion the user confirms. MACHINE_ROLE in the environment answers it ahead of time for an unattended run, which is why it is declared with :- rather than a plain assignment. The first version wiped the caller's value before ask_machine_role ever saw it; caught by running with MACHINE_ROLE=vps and watching the menu appear anyway. Verified: guesses vps on this host (public IPv4, no DISPLAY, no display manager), env override takes, and a bad value fails with the three valid ones named. Nothing consumes the role yet — the steps get wired as each is worked through. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
34bb8fc22a |
put the superseded setup scripts in setup-old, and repair what the move broke
scripts/setup/ is now what the new installer is being built in — machine-setup/ for the box, officer-setup.sh for the platform on top — and everything being replaced moved to scripts/setup-old/. It still works and is still what to run. Three things the move broke, and what each needed: starship.toml is not an old-setup artifact. os-user-shell.ts reads it at RUNTIME to seed a member's ~/.config/starship.toml when their Linux account is provisioned, and line 125 reads it inside a try whose catch returns "could not read the shell templates" — so account provisioning would have failed outright, not degraded. Moved back to scripts/setup/, which is where it belongs anyway (one file, both audiences) and which leaves the code correct with no edit. package.json's `setup` script pointed at a path that no longer exists. It now points at officer-setup.sh, where the installer is going, rather than at setup-old/ which is temporary. officer-setup.sh was created empty. An empty script exits 0, so `bun setup` would have reported success while doing nothing — worse than the broken path it replaced. It now explains that it is not written yet and exits 1, naming the setup-old script to run meanwhile. Also brought .tmux.conf and ufw-docker-rules.conf in beside machine-setup.sh, which reads both from SCRIPT_DIR and had been silently skipping them since the script was vendored. ssh-keys.zip deliberately stays out: it is key material, and *.zip is ignored. Comments in os-user-claude.ts, app-store/preflight.ts and two docs still name the old scripts/setup/setup.sh path. Left alone on purpose — repointing them at setup-old/ only to repoint them again when officer-setup.sh lands is churn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44141faf0a |
split machine-setup into an entry point and a base library
Structure before the work rather than during it: scripts/machine-setup.sh becomes
scripts/setup/machine-setup/, with the script itself as the entry point and
lib/base.sh holding what every part of it needs.
machine-setup.sh pre-flight and the numbered sections, for now
lib/base.sh shared state, output, the step/resume machine, prompts,
and OS detection
The rule for lib/ is definitions only — nothing there installs, writes or
restarts anything, so sourcing it is safe from anywhere. That is why the ERR
trap stayed in the entry point: a trap is a side effect on whoever sources it.
Behaviour is unchanged. Verified by diffing the moved region against the previous
commit: identical set of functions, and the only differences are added comments,
section banners, fail() reformatted onto three lines, and one new line — a guard
against double-sourcing, which matters because steps will source this directly
once they move out, and a second pass would reset SUMMARY.
The sections are still one 1111-line block below pre-flight; they move into
steps/ as each is worked through. The script also still reads ssh-keys.zip,
.tmux.conf and ufw-docker-rules.conf from SCRIPT_DIR, which is now this
directory, so those three steps warn and skip until the files follow it here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
dda214ffb0 |
detect the operating system before any step runs
The script assumed Ubuntu on x86_64 in every line of it. detect_os() now runs first and fills in OS, OS_NAME, OS_VERSION, PM, ARCH and IS_WSL, so the steps have something to branch on as support for other systems is added. Read from /etc/os-release rather than probing for a binary: a machine can have more than one package manager on PATH, and only os-release can say which distribution this actually is or give a version worth printing. Sourced in a subshell so its NAME, VERSION and ID do not leak in here. ID_LIKE is the fallback, so Pop!_OS, Mint and EndeavourOS resolve without being named. ARCH is normalised to amd64/arm64 in one place because upstream disagrees — Neovim ships aarch64, Go and Docker ship arm64, lazygit ships x86_64 — and several steps hardcode one spelling today. Windows exits with a message pointing at WSL2. WSL itself is detected and warned about rather than refused: it reports as Linux but has no real systemd session, so the suspend, logind and boot-hang steps do nothing there. Everything below pre-flight is still apt and systemd only, so a gate refuses pacman/dnf/brew by name rather than half-building a machine and stopping somewhere unhelpful. Relax that case one entry at a time as each grows a path. Verified on this host (Ubuntu 24.04.4, amd64, apt) and by stubbing uname and os_release for arch, manjaro/arm64, fedora, pop, macos, mingw and riscv64. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
483bb15d8a |
vendor the ubuntu machine provisioning script, verbatim
A byte-for-byte copy of /root/ubuntu-setup/setup-ubuntu.sh, the script that has provisioned every Ubuntu server here. Committed unchanged, before any edit, so that everything the setup-script rework does to it reads as a diff against what actually ran on real machines rather than against a tidied-up version of it. Nothing in the repo calls this yet. It also cannot find three files it reads from its own directory — ssh-keys.zip, .tmux.conf and ufw-docker-rules.conf all live beside the original in /root/ubuntu-setup, and SCRIPT_DIR is scripts/ here, so those steps warn and skip. The original stays where it is and stays authoritative until this one replaces it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
315073cf3b |
move the pm2 ecosystem files into ecosystem-files/ for reference
Temporary, and it breaks things — nothing has been repointed yet:
scripts/setup/setup.sh:59 joins a bare filename to $PROJECT_DIR
scripts/setup/setup_mac_light.sh:60 the same
src/servers/app-store/pm2.ts:23 starts sidecars from 'ecosystem.config.cjs'
src/servers/app-store/catalogue.test.ts:12-13 require('../../../ecosystem…')
ServersView.tsx:207 tells the owner to run pm2 start ecosystem.config.cjs
And one thing that changed silently rather than breaking: ecosystem.profile.cjs:53
pins cwd to __dirname, which was the repo root and is now ecosystem-files/, so the
.env that line exists to find is no longer beside it.
These are here to be read while the setup scripts are reworked, and get deleted
once that lands.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
79da78008a |
turn the field report into something an agent can follow
Expands the communications section from a list of what worked into the actual convention: the directory's lifetime and the rule that anything durable must move to docs/ before the merge; numbering, parity as attribution, non-consecutive numbers; slugs; reply-in-a-new-file and the one case where editing your own is right; referring to commits by sha because three remotes carried the same branch names. Records what a handoff must contain, with the verified/assumed split named as the rule that carried the most weight — a handoff confident about something untested is worse than none, because the reader builds on it. Adds a skeleton to copy. Documents termination as the four attempts it actually took, ending at the only checkable version: the exchange pauses when no open item is actionable by a participant. Adds the third state, deferred-with-a-reason, since a two-state protocol forces an agent to lie in one direction. Notes that a stall must be detectable because the human spotted both before either agent did. Adds a review-discipline section — check the enforcement rather than the description, run it against a real machine, a check never seen failing is not evidence, distrust vacuous passes, expect stacked bugs, distrust "inert today", and look at which way unknown resolves. Adds a failure-mode table to pattern-match against, and the git hygiene that bit us, including merge-verify-then-delete, which I got wrong. Closes with session economics, an ordered list of what to build, and the one thing not to automate: agents may coordinate on what is true and must not decide what is permitted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
015e280e5c |
document how to launch the watcher, since the mechanism is what did not transfer
The owner could not convey this to the second agent, who launched it differently and got something that looked identical and did not work. The script was never the hard part; the mechanism is. States the requirement so it survives a different harness — a detached shell process owned by the agent's harness, which exits when it has something to say, and whose exit re-invokes the agent — and notes that dropping any one of those three breaks it invisibly. Then the four wrong ways, each of which looks correct while running. Backgrounding with nohup or & produces a process that polls correctly, detects the push, exits, and never tells the agent, because the harness is not tracking it; I made that exact mistake and caught it only by re-reading my own command. A model-driven interval is functionally correct and pays a full context re-read per tick to learn nothing — the intuitive design, and the expensive one, which is why it is the first thing to warn a new agent about. A loop that does not exit on detection has no path to the agent at all. And per-tick logging is deferred cost that lands all at once on wake. Also records why 30s polling is free in a shell and ruinous in the model, including the five-minute prompt-cache TTL that makes any model-side wake beyond it pay for a full uncached read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e9d0261e87 |
a field report on two agents working one branch
docs/agent-coordination.md states the objective — several agents on one body of work, coordinating with each other rather than through the human — and was written in theory on 2026-08-07. On 2026-08-11/12 it ran for ten hours with two agents and the owner arbitrating. This is what happened, written as evidence rather than proposal. The load-bearing observation is narrower than "two reviewers are better than one": the person who writes the sentence explaining why something is safe is the worst-placed person to notice the code disagrees with it. One agent wrote "a wrong answer here must not happen by accident" and shipped that accident in the same commit; the other wrote a verification script that could not fail on the first one's machine. Neither was careless. Each was reading their own reasoning back and finding that it agreed with itself. Also records what only running found — an installer piped into the wrong shell, a parent directory created root:root, an ACL mask clamped so the file browser could not read a member's home, a chat cwd the member could not enter, ACL entries surviving a chown — all on first executions, all invisible to review. And what the communications channel got right and the five ways its termination rules broke, and why the repo watcher belongs in a shell loop rather than in the model. Names the identity gap as the first thing to build: both agents commit as the owner, so neither the log nor an agent can say who wrote a line. docs/agent-git-identity.md has called that an idea since 2026-08-10; it stopped being one tonight. Also corrects the deprovision spec's status, which still said "not yet run against a real account" after it had been run and verified clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2eedbcda54 |
bind nginx proxy manager to the tailscale address
It was the only service here publishing on 0.0.0.0, and a published docker port is not behind the firewall: docker writes its DNAT rules straight into the nat table, which ufw's INPUT chain never sees. `ufw default deny incoming` never covered 80/443/81 — ufw-docker-rules.conf on the host exists to patch exactly that, and patching a rule is weaker than never opening the socket. The address is read from `tailscale ip -4` at run time rather than passed in, because the host provisioning has already done `tailscale up` by the time this executes. It is validated against 100.64.0.0/10, the range tailscale and headscale both allocate from. SETUP_NPM_BIND overrides it. With neither, selecting NPM exits instead of falling back to 0.0.0.0 — a fallback would silently undo the point of the change. Two consequences worth knowing. tailscaled becomes a boot-order dependency, so the script warns when it is not enabled at boot; docker's restart policy covers the window but only if the tailnet comes up on its own. And HTTP-01 ACME challenges can no longer reach port 80, so any certificate NPM issues now needs DNS-01. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4e404f17c8 |
split the optional host dependencies out of setup.sh
8 Rust, 9 PulseAudio, 10 cliamp, 13 yt-dlp and 17 the remote desktop move to scripts/setup/setup-sidecars.sh, which nothing invokes — running it is a deliberate act. They are what the optional, sidecar-backed features need on the host, not what the app needs to serve itself. 11 Neovim, 12 the shell extras and 14 the npm globals are gone entirely. The host provisioning already installs node, npm, pm2, Claude Code, Neovim and the shell, and two installers racing for the same binaries is worse than one. That makes node, npm, pm2 and the agent CLIs prerequisites of this script rather than products of it, so the verification block still checks claude and pm2 — section 19 warns and skips rather than failing when pm2 is absent, which would otherwise finish "successfully" with nothing listening. eza is the one casualty: the provisioning installs lazygit, starship, oh-my-zsh and nvim, but not that. Section numbers keep their gaps so the two files read against each other. One line survives from the removed section 14 — the ~/.local/bin PATH export, which section 19's `has pm2` and the agent's claude lookup both still depend on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cec8fbe57e |
the acl check could not fail, because sudo drops DATA_PATH
Review of
|
||
|
|
76cd7c20bf |
sever the ACL as well as the ownership
severMemberTree reassigned the tree and left the access-control entries behind. confineUserTree
grants each member a named ACL on their whole tree — u:<uid>:rwx plus a default: copy — and
chown does not remove them: they are xattrs rather than ownership, and they record the uid
numerically.
Measured before this change: after chown -h -R to the service user, user:<uid>:rwx was still
present on the directory, on its children and in their defaults. The tree read as the
platform's while still granting the freed uid read and write on every byte, so the next account
allocated that number would inherit the previous member's home, keys, credential and container
storage — the hazard this file exists to prevent, reached through a door that find -uid cannot
see.
Now chown then setfacl -R -P -b. Proven on a scratch tree: owner 1001 with five entries naming
1001 becomes owner 1000 with none.
-b rather than removing the member's entries alone, because the service user owns all of it
afterwards and "no ACLs" is cheaper to verify than "no ACL naming one id". -P is already the
default for a recursive setfacl — verified, a symlink out of the tree was not followed — and is
stated for the same reason the chown above carries -h: a member chooses what their symlinks
point at, and this argv should not rest on a traversal default holding.
Found by running assert-uid-free.sh against a real tree; the spec and the checker had the same
blind spot and were corrected in
|
||
|
|
a2f63dc534 |
severing ownership does not sever the ACL, and neither the spec nor the checker said so
Reviewing
|
||
|
|
46799dada8 |
deprovision a member's linux account when the platform account goes
Implements docs/deprovision-os-account.md. Until now deleteUserHandler removed the row, cascaded the
database, and left the entire Linux side running — measured on production on 2026-08-12: working login
shell, healthy postgres container, 454M of data, uid queued for the next useradd to reissue along with
everything still owned by it.
The load-bearing rule from the spec: sever the data from the uid BEFORE releasing the uid, and if
severing fails, do not release. A failed deprovision is not a broken account, it is a trap for whoever
is created next.
Sequence: disable-linger, terminate-user, reap-and-prove, chown -R, userdel (never -r).
reap terminate-user is not a barrier. Production measured a three-hour-old `/bin/zsh -i` surviving
it AND the removal of /run/user/<uid>. So: pkill, bounded wait, pkill -9, bounded wait, and a
final count that must be zero or the account is not released.
chown fixes the uid and subuid halves in one pass — it rewrites every file it walks whatever owned
it. The range is still captured first, because userdel removes the /etc/subuid entry and after
that nothing on the machine remembers what it was. It is returned on every path including the
failures, and logged as the exact assert-uid-free.sh command line.
Two guards the spec did not ask for, both pure and unit-tested:
guardDeletable ensureOsUser's adoption rule backwards. Deletable only if the passwd home is the one
the platform would have confined, and uid >= 1000. Without it `userdel root` is one
bad users.osUser away and nothing else in the sequence would object.
guardMemberTree the tree must resolve to a direct child of DATA_PATH. The email reaches join() from a
database row and the result is the argument to a recursive chown.
chown runs with -h. Measured here that `chown -R` already declines to follow a symlink out of the tree and
re-owns the link itself, but the argv should say so rather than rest on traversal semantics — and
re-owning links is what makes `find -uid` (lstat) a meaningful check afterwards.
destroy exists, has no call site, and is chown-then-delete-as-the-service-user rather than sudo rm -rf, so
a recursive root delete built from a database column does not exist in this codebase.
deleteUserHandler now runs this FIRST and refuses to delete the row if it fails: the row is what remembers
there is anything to clean up, so deleting it first makes a failure unrecoverable through the UI.
NOT YET RUN AGAINST A REAL ACCOUNT. Only the pure guards have tests. The five-step validation is in the
doc; it needs the production host, a shell left open, and a container writing as a non-root user — the two
cases the quiet path passes vacuously.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
f34d7fef70 | merge: a checker for the deprovision spec, and the trap that makes it pass for free | ||
|
|
35715546e2 |
a checker for the deprovision spec, and the trap that makes it pass for free
scripts/assert-uid-free.sh is the verification half of docs/deprovision-os-account.md, written outside the implementation on purpose: a checker the function calls is a restatement of its own beliefs rather than an audit. Nine checks — passwd entry, uid reuse, both subid files, linger, runtime dir, live processes, files owned by the uid, and files owned anywhere in the freed subuid range. Two modes, because the range has to be captured BEFORE deletion. userdel removes the /etc/subuid entry along with the account, and after that there is no way to ask what range it held — so a checker that only runs afterwards silently drops the half most likely to be wrong. Exercised against green while fully provisioned: eight of nine checks fail, exit 1. A checker that has never been seen to fail is not evidence. And the trap worth knowing before anyone trusts a green result: the subuid check passes vacuously on most accounts. Files get a mapped owner only when a process inside a container runs as a NON-root user; an image whose files are root-owned maps to the member's own uid and leaves the range empty. Measured on green after a night of real use — claude installed, an image pulled, transcripts written — the range check found zero files and passed without testing anything. The spec now says how to build a specimen that actually exercises it, and to watch the check fail on that tree before trusting it to pass on a cleaned one. Docs and a script only; no behaviour change. On a branch, for whoever merges it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f5f509a99d |
scripts/setup is the initial install, nothing else
Two of the eight did not belong. cleanup-desktop.sh is the teardown — the inverse of an install, not part of one. provision-user-dirs.ts runs per account at invite time, on a machine that is already set up. Both are back at the top level, with their `../` derivations and usage strings put back. What is left is what a fresh machine runs once: the two installers (setup.sh, setup_mac_light.sh), the two things setup.sh calls (setup-dockers.sh, setup-desktop.sh), and the two files they deploy — starship.toml, copied to ~/.config, and officer-set-display.sh, which setup-desktop.sh installs to ~/.local/bin as a login-time mode setter. The last one is not a setup script and does not read like one; it is here because it is install payload, same as the toml, and setup-desktop.sh loads it by `$(dirname $0)`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9c353f5f0d |
move host setup into scripts/setup/
scripts/ was holding two unrelated kinds of thing: install-this-machine, and run-this-occasionally. The eight installers now live in scripts/setup/; what stays at the top level is the build steps (gen-index, prebuild, build/) and the two maintenance scripts (reindex-music, rebuild-soulseek-tree). The move is not just a rename. Three of these derive the repo root from their own location: setup.sh:51 PROJECT_DIR="$(dirname "$SCRIPT_DIR")" setup_mac_light.sh:51 same cleanup-desktop.sh:134 ENV_FILE="$(dirname "$0")/../.env" Left alone, all three would now resolve to scripts/ — and nothing downstream complains. PROJECT_DIR is where .env is written, where `bun install`, `gen:index` and `db:push` run, and what pm2 is pointed at, so a fresh install would have quietly provisioned scripts/ and reported success. cleanup-desktop.sh fails the other way: it would find no .env, print "No .env — skipping", and leave the real VNC_PASSWORD in the real file. All three are now `../..` with a comment saying why the level matters. provision-user-dirs.ts imports data-path.ts relatively; that one tsgo caught. Also disambiguated `setup.sh` where it had become two files. app-store/templates/<name>/setup.sh is a per-sidecar installer with its own contract, and preflight.ts + docs/sidecar-app-store.md discussed both in the same paragraph. The host one is now spelled with its full path at those sites. Verified: bash -n on all six shell scripts, tsgo clean, os-user tests pass, both derivations resolve to the repo root, starship.toml still resolves from os-user-shell.ts, and provision-user-dirs.ts runs under DRY_RUN. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
37adc65a12 |
delete the spent one-shot scripts
Eight scripts in scripts/ that nothing references and that mostly can no longer run. Kept in history;
none of them is recoverable knowledge that isn't already in the code they migrated to.
Three could not run at all against the current database:
migrate-items-to-files.ts SELECT * FROM tasks — that table was dropped when items became files
reset-user-data.ts deletes chat_sessions, chat_groups, projects; none exists. It has no
transaction, so it would wipe user_settings, user_state,
user_integrations and dock_configs and THEN throw. A half-wiped account
is worse than no script. It also misses chat_session_events, which is
where chat state actually lives now.
add-email-dock-user2.ts one-time, hardcoded to user 2, seeds a dock containing /projects
The rest are spent migrations whose destination is now the only implementation:
migrate-auth-to-pg.ts JSON -> Postgres, 2026-02
migrate-pg-to-files.ts Postgres -> JSON, the other leg of the same abandoned round trip
migrate-server-settings-to-pg.ts 2026-02
migrate-emails-to-sqlite.ts backfill into the email sidecar's store, 2026-07-31
seed-imap-uids.ts the sidecar writes imap_lastuid/imap_uidvalidity itself now
(sidecar/email/gmail-api.ts:533-535)
Kept, and why, since "unreferenced" was not the test: rebuild-soulseek-tree.ts is reusable by
construction — it runs the same buildTree the sidecar's ingest runs, so it answers any future change
in tree shape. reindex-music.ts is named in sidecar/music/index.ts:447. provision-user-dirs.ts shares
USER_DIRS with data-path.ts. cleanup-desktop.sh and officer-set-display.sh are called by
setup-desktop.sh.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a6acfea9a6 | carry the two threads todo.md was missing | ||
|
|
a730fc0fe0 | keep the three open threads the comms channel was holding |