09303299cc0234e4367a673162b3540dd2e4be53
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
09303299cc |
port Tailscale as its own section, early, and stop it hanging
Moved to position 7 — after core utils, which give it curl, and well before SSH
hardening, which is the step that can lock you out. The argument is that
Tailscale is a second way into the machine, so it wants to exist before anything
that can go wrong does.
── Why the original hung, and what stops it now ──
Its prompt accepted an empty auth key and passed it anyway. `tailscale up
--authkey ""` falls back to the interactive flow: it prints a URL and blocks,
with no timeout, forever. From the outside that is a script that has frozen.
Nothing here passes an empty key — the flag is omitted entirely, and the run says
in advance that a URL is coming and that it will wait. Every call carries
--timeout=60s, and a timeout is reported with the command to run by hand rather
than left as silence. State is read with `tailscale status --json` before
anything is run, so a node that is already up is offered a reconfigure instead of
having `up` fired at it blindly.
Diagnosed on this host rather than guessed at, and honestly the diagnosis is
partial: the exit-node branch left no trace at all — no /etc/sysctl.d file,
networkd-dispatcher present but with zero mentions in apt history, so it came
with the image. ip_forward=1 came from the unconditional part of the section, not
the branch. That points at `tailscale up` as where it stopped, and the empty-key
path is the candidate that fits, but I could not reproduce it to be certain.
── What the section now covers ──
control plane Tailscale's own service by default; a self-hosted headscale as
an explicit choice with NO suggested URL. The original defaulted
to headscale.pastilhas.eu, so a stranger running it pointed
their machine at somebody else's control plane.
auth key, or the browser flow, stated as an equal option
Tailscale SSH ssh over the tailnet with no keys, governed by tailnet ACLs —
and pointed out as a way back in if the sshd hardening later in
the run goes wrong
subnet router homelab only, defaulting to this machine's actual LAN CIDR
exit node with what it means for whose traffic goes where
forwarding sysctls and Tailscale's recommended NIC offload settings, and
only when an exit node or a route actually needs them
Approval is mentioned: an advertised route or exit node does nothing until it is
approved in the admin console, which is otherwise a silent non-event.
Also added --only <step> and --list, because this section in particular needs to
be run on its own while it is being worked on. A step run that way ignores the
progress file and does not record itself — asking for one step is not progress
through the script.
One bash quirk fixed on the way: "${VAR:-Tailscale's own service}" does not
parse. An apostrophe inside a ${:-} default opens a quoted section that swallows
the closing brace, and the error surfaces as "unexpected EOF" 400 lines away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
192293cdca |
offer to add an ssh key even when one already exists
The zip option was already gone — it never came across in the port, since the file is key material that cannot live in the repository and the script no longer sits next to it. Pasting a public key was already the first option. What was missing is the case where the account HAS a key: the section went straight to hardening, so there was no way to authorise a second machine, a rebuilt laptop or anyone else, and the original had no way to do it at all. The same menu is now offered either way. What differs is whether it can be declined without consequence: with no key, declining means the hardening below refuses too, and the run says so rather than quietly moving on. confirm() takes an optional default so this one can be [y/N]. Most questions in this script are "do the thing you already asked for" and Enter should mean yes; a genuine extra defaulting to yes is how people end up agreeing to things by reflex. A pasted key is trimmed before validation. Copying from a terminal or a password manager routinely brings leading or trailing whitespace, and ssh-keygen will not parse a key with it attached — which would have read as "that is not a valid key" for a key that is perfectly fine. Also corrected the reason unzip is in core utils, which still said it was there to open ssh-keys.zip. Verified: the add-another prompt appears and defaults to no, a whitespace-wrapped key is trimmed and accepted, and the already-hardened path is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b1cc916258 |
detect root by uid, not by the name root
A provider whose image logs you in as "ubuntu" at uid 0 would have walked straight past the previous check, which compared the string. What makes an account root is uid 0; "root" is only the usual label for it. Two places now ask id -u rather than comparing names: the answer — an account at uid 0 is refused whatever it is called, and says which case it is rather than a bare "not root" the invoker — the warning about working as root fires when SUDO_USER is unset OR when SUDO_USER is itself uid 0. The second is the one that hides: sudo from a uid-0 account sets SUDO_USER to something that reads like an ordinary user and is not. The EUID check that requires the script to run as root was already uid-based and is unchanged. Verified by creating a real uid-0 account named ubuntu on this box: refused with the uid named, where the name check accepted it. That account has been removed — userdel refused it at first because it matches by uid and saw PID 1 running as uid 0, so -f was needed, and deliberately not -r, since its home was /root. Confirmed afterwards that root, /root, root's shadow entry and sudo are all intact and that root is once again the only uid-0 account. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d9d4085033 |
warn against working as root at the username prompt
root was already refused as an answer, but only the refusal said so — and only after somebody typed it. The advice now comes with the question, along with why: no safety net, a typo in a path that deletes instead of refusing, and nothing to distinguish you from a process that got out of hand. An extra warning when SUDO_USER is unset. That means the script was started as root rather than through sudo, which usually means root is how they log in — the exact situation the general advice is about, and the one where general advice is easiest to assume is aimed at somebody else. It says so plainly instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0eb3e3a85 |
create the user account first, and fix the sudoers filename bug that found
The user account section is now the first thing that acts, ahead of disk space.
The reason is a real defect, not tidiness. If the account does not exist yet,
USER_HOME is a path that is not there — and the ballast offers to put its file in
it, where ballast_create's `mkdir -p` runs as root and creates /home/<name> owned
by root:root. adduser afterwards finds the directory already present and does not
populate or chown it, so the account ends up with a home it cannot write to.
Making the account before any step can write into its home removes the ordering
entirely.
USER_HOME is re-read from getent after adduser runs. Until that point it is the
/home/<name> guess, because there is nothing to look up; adduser is free to have
used something else and every later step writes there.
Found while testing that, and worse than the thing it was testing:
local user="$1" dest="/etc/sudoers.d/99-${user}-nopasswd"
bash expands ${user} before the assignment to user has happened, so dest came out
as /etc/sudoers.d/99--nopasswd with the name missing. The rule inside was correct,
which is what made it invisible — visudo passes, sudo works, and the account
really does get passwordless sudo. What breaks is everything around it: every
account granted this way writes to that same file, so a second grant silently
overwrites the first and revokes it; and has_passwordless_sudo looks for
99-<user>-nopasswd, never finds it, and re-grants on every run forever.
Split into separate declarations, with the reason recorded where it happened, and
an empty username is now refused outright. Scanned the other lib files for the
same shape — the remaining multi-assignment locals only read positional
parameters, which is safe.
The stray /etc/sudoers.d/99--nopasswd this created on the dev box during testing
has been removed and visudo -c re-verified.
Verified with a real throwaway account: correctly reports not-granted before,
writes 99-msdemo-nopasswd as root:root 0440 with the right rule, reports granted
after, and leaves sudoers valid.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fbb90c917d |
default the username to whoever ran sudo, and look up their real home
Three changes to the question at the top of the run. USER_HOME is looked up rather than assumed. The original built "/home/$USERNAME", which is only the usual answer — an account created with a different home, or one whose home was moved, had every later step writing to a directory that was not theirs. getent passwd knows; the /home guess remains only as the fallback for an account that does not exist yet, where there is nothing to look up. The default is now whoever invoked sudo. On a re-run, or on a machine that is already somebody's, that is the answer every time, and retyping it is a chance to typo it into creating a second account. root invoking the script directly offers no default, since root is never the account being set up — and is refused if typed. The name is validated against the portable shape of a Linux account name before anything else happens. Letting adduser refuse it later means several questions have already been answered against a name that was never going to work. It also no longer goes through prompt_value, which obeys any environment variable matching the name it is filling in. USERNAME is set by some login environments, and a variable this script silently takes as an answer should not be one that might already be set for unrelated reasons. SETUP_USERNAME is the explicit override. The tmux config write moved out of the user section entirely. It was there only because the original copied it right after adduser, where $USER_HOME first exists. It is a dotfile and belongs with .zshrc and the starship config in the shell section. Verified: defaults to the sudo invoker, resolves daemon's home to /usr/sbin rather than /home/daemon, falls back to /home for an account that does not exist, and rejects a name with a space, a leading digit, one over 32 characters, and root. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3fb0e5c887 |
port the user account section, and stop clobbering files in the home
Two real defects fixed on the way across.
The sudoers write was in the wrong order. The original echoed the rule straight
into /etc/sudoers.d, validated it afterwards, and chmod'd it later still. A
malformed file there breaks sudo COMPLETELY — and you cannot sudo to repair it,
so on a remote machine that is a rescue console — and so does one with loose
permissions, because sudo refuses to read its own configuration. Both of those
windows were live in the original ordering. grant_passwordless_sudo now writes a
temp file, runs visudo -c against it, and only then places it with install(1),
which applies the content and the 0440 mode in one step. Nothing reaches
/etc/sudoers.d that has not already been validated.
The .tmux.conf copy overwrote whatever was in the home on every run. lib/files.sh
adds the two shapes that stop this whole class of thing:
install_config installs when absent, does nothing when identical, and keeps
what the user wrote when it differs — printing the cp to take
ours, so the choice stays theirs
append_once wraps a block in named markers so a second run recognises its
own work; also lets a human see which lines came from this
script and remove them as a unit
append_once is what the five unguarded `cat >>` into .zshrc need when those
sections are ported — a second pass currently duplicates the starship init, the
nvim PATH, bun, deno and the aliases.
Passwordless sudo is asked separately from creating the account, because it is a
security posture rather than part of making a user, and the cost is stated: a key
that can log into this account is root without a further step. Officer's actual
requirement is stated too — os-user-shell.ts runs `sudo -n`, and a prompt it
cannot answer surfaces as a permissions error rather than a question — and
refusing records that consequence in the summary instead of a bare "skipped".
Verified: all three install_config outcomes, append_once writing exactly once
across two runs, visudo rejecting junk before anything is installed, and the
section reporting correctly against this host's existing account.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d4b9b7b334 |
ask where Officer should be installed, in pre-flight
OFFICER_ROOT, defaulting to <user home>/officerdev. One directory holding the four things Officer is made of, per docs/sidecar-app-store.md — the app, its data, the item store, and any containers the app store provisions — so the whole installation can be moved, backed up or deleted as a unit. Asked at the start with the other questions rather than at the point it is first needed. It decides the shape of several later steps: where the repository is cloned, where DATA_PATH sits beside it, and which filesystem the app store's bind mounts come out of. Asking once up front also means the run can be described before it starts rather than discovered as it goes. A leading ~ is expanded explicitly. It arrives as a literal from a read or an environment variable — nothing expands it there — and would otherwise create a directory actually named "~" in whatever the working directory happened to be. Relative paths are refused with the value named, and a trailing slash is trimmed so the path composes cleanly with what gets appended to it. Nothing creates the directory yet; that belongs to officer-setup. This records the answer and reports it, including whether it already exists. Verified: Enter takes the default, ~ expands, trailing slash trims, OFFICER_ROOT in the environment skips the prompt, and a relative path fails with the value named. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f896d4882f |
one section per concern, each announced and confirmed before it acts
Three sections where there were two, and none of them touches the machine until you say so: 2. System update upgrades what is already installed 3. Core utils what the distribution provides 4. Command-line tools lazydocker, lazygit, starship, fastfetch The split matters because these are different kinds of change and deserve separate answers. System update is the only step in the whole script that moves versions of software already on the machine; core utils only ever adds what is absent; and the four tools are upstream binaries the distribution does not ship at all. Previously the update and the core packages were one step and the tools were tacked onto the end of it, so agreeing to "essentials" meant agreeing to all three at once. Every section now prints what it will install and what it is leaving alone, then asks. Enter means yes — unlike the machine-role question, which has no default, because these are "do the thing you already asked for" and making twenty of them require a deliberate keystroke would train people to hold the y key down. ASSUME_YES=1 answers all of them for an unattended run, and EOF fails with that named rather than spinning. Refusing is recorded rather than glossed: LAST_SKIPPED feeds the summary, so a declined section reads "Core utils: SKIPPED by request — cowsay neofetch" instead of quietly reporting nothing installed. Nothing to install means no prompt at all — there is nothing to agree to. announce_plan takes the array NAMES rather than their contents, because once a list has been through word splitting an empty one cannot be told from a missing one. Verified all three paths: accept, refuse, and nothing-to-do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dd655577e3 |
make the machine-role question require an answer
No default, and it is the only question in the script like that. A guessed default is right often enough to be trusted and wrong in exactly the case that costs the most — pinning a static IP on a rented box, or leaving the firewall open on one. Every branch downstream is about what this machine is exposed to, so it is worth one deliberate keystroke rather than an Enter. Empty and unrecognised answers re-ask rather than aborting; a failed read means EOF rather than a wrong answer, and fails with the environment variable named, because otherwise the loop spins forever the first time this runs unattended. Drops guess_machine_role, which existed only to supply that default. default_iface stays — the static IP section needs it when it is ported. MACHINE_ROLE in the environment still answers it ahead of time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
41ff8030e9 |
ask what the machine is for, once, in pre-flight
MACHINE_ROLE is homelab, vps or dev, and several steps have a different right answer per role with no way to work it out themselves: whether the address is yours to pin (static IP), whether the box faces the open internet (fail2ban, SSH hardening, UFW), and whether it is allowed to sleep (suspend, logind). Asked in pre-flight rather than at each point of use. The steps that care run from swap through to the firewall, and being asked "is this a VPS?" for the fourth time halfway down a provisioning run is how people start answering without reading. The default offered is guessed from whether this machine's own address is in RFC1918 space, which beats asking whether it is virtualised — a homelab is very often a VM on Proxmox and would be misread as rented — and is the same fact most of the branches turn on anyway. A graphical session means dev; so does macOS. It is only ever a suggestion the user confirms. MACHINE_ROLE in the environment answers it ahead of time for an unattended run, which is why it is declared with :- rather than a plain assignment. The first version wiped the caller's value before ask_machine_role ever saw it; caught by running with MACHINE_ROLE=vps and watching the menu appear anyway. Verified: guesses vps on this host (public IPv4, no DISPLAY, no display manager), env override takes, and a bad value fails with the three valid ones named. Nothing consumes the role yet — the steps get wired as each is worked through. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44141faf0a |
split machine-setup into an entry point and a base library
Structure before the work rather than during it: scripts/machine-setup.sh becomes
scripts/setup/machine-setup/, with the script itself as the entry point and
lib/base.sh holding what every part of it needs.
machine-setup.sh pre-flight and the numbered sections, for now
lib/base.sh shared state, output, the step/resume machine, prompts,
and OS detection
The rule for lib/ is definitions only — nothing there installs, writes or
restarts anything, so sourcing it is safe from anywhere. That is why the ERR
trap stayed in the entry point: a trap is a side effect on whoever sources it.
Behaviour is unchanged. Verified by diffing the moved region against the previous
commit: identical set of functions, and the only differences are added comments,
section banners, fail() reformatted onto three lines, and one new line — a guard
against double-sourcing, which matters because steps will source this directly
once they move out, and a second pass would reset SUMMARY.
The sections are still one 1111-line block below pre-flight; they move into
steps/ as each is worked through. The script also still reads ssh-keys.zip,
.tmux.conf and ufw-docker-rules.conf from SCRIPT_DIR, which is now this
directory, so those three steps warn and skip until the files follow it here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|