Servers only now. On a machine you sit at, a filling disk announces itself — the
editor refuses to save, the browser complains — and you are there to deal with
it. The reserve is for the box nobody is watching, where the first sign is a
service that stopped working hours ago.
Skipped rather than asked, but said out loud with the reason and recorded in the
summary. A section that silently produces no output is indistinguishable from
one that failed.
The cron this section installs was already there and is unchanged: /etc/cron.d
runs the checker as root every ten minutes, and it deletes the ballast when free
space falls under the threshold.
What changed is where its message goes. Both alerts now run through one notify()
inside the generated checker rather than calling logger directly, so there is a
single place to add a second channel. Today it is still syslog only — the
message lands in the journal and nowhere else, so nobody learns about it until
they go looking, which is precisely the wrong moment. Push, mail or Officer's own
notify sidecar hook in there. It also echoes to stderr now, so running the
checker by hand shows the message instead of appearing to do nothing.
Verified: dev reports not-applicable and asks nothing, vps still asks and records
a refusal, and the regenerated checker parses and reports status.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three questions instead of one, because the two the original never asked are the
two that decide whether the thing is useful.
1. Do you want one, with the explanation first.
2. Where. Home (easiest to find again months from now), beside Officer, or a
path typed in. This is not tidiness: the checker measures its own directory,
so a ballast only protects the filesystem it sits on. Choosing where it goes
is choosing which mount is covered.
3. How much, as 5/10/20% — with the actual numbers, and with what would be LEFT
rather than only what is taken:
[1] 5% — reserves 2.9GB leaving 54.3GB free
[2] 10% — reserves 5.8GB leaving 51.4GB free
[3] 20% — reserves 11.5GB leaving 45.7GB free
A percentage on its own is unanswerable. The number that decides it is the
one on the right: the reserve has to be big enough to matter and small
enough not to be the thing that filled the disk.
The size is computed against the filesystem the chosen path lands on, after the
location is known, so the percentages are of the right disk. ballast_free_kb
walks up to a directory that exists, since nothing has created the target yet.
An existing ballast in either default location is found and left alone rather
than a second one being made beside it.
Verified end to end in a temp home: created at the chosen path, 2.9G for 5% of
57.1GB, checker installed and reporting it present.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Swap covers memory pressure. These are its neighbours:
8. Emergency disk ballast the same valve, for disk
9. earlyoom what happens when swap runs out too
10. inotify watch limit the silent one
All three follow the rule this script now works to: the role sets which way the
recommendation points, never whether the question is asked. A dev machine is
still offered the ballast, with the recommendation pointing the other way; a
server is still offered the inotify raise, because anything running `bun --watch`
or serving a file browser is a watcher too.
The ballast is section 22 of the original, moved up beside swap where it belongs
and moved out of the user's home. The original wrote the checker into
$USER_HOME/.local/bin and ran it from a root cron — a root cron executing a
script in a directory its owner can write is a privilege escalation waiting to be
noticed. Moot on a box where that user already has passwordless sudo, but wrong.
Both the checker and the file are in root-owned system paths now.
Two bugs found by running the generated checker rather than reading it:
It df'd the ballast's own directory, which does not exist before the ballast is
created — and with `set -euo pipefail` that meant cron mailing an error every
ten minutes. It now walks up to a directory that exists, and the installer
creates the directory itself rather than depending on the create step.
The inotify text claimed a default of 8192. This host is at 29461: Ubuntu
raised it, and stating a number the reader can see is wrong on their own screen
undermines the rest of the explanation. It now describes the failure instead
and prints the machine's actual value.
earlyoom is a distro package and a systemd unit, so it is checked with
`systemctl is-active` and reports honestly when it installs but fails to start.
Verified: checker --status and its no-op path both exit 0 with no directory
present, and the helpers report correctly against this host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four things the original got wrong, all of which only show up on a machine that
is not this one:
It detected swap with `swapon --show | grep -q '/'` — a test for a swap FILE.
A machine using zram or a swap partition reports no swap at all, and the step
would add a swapfile beside working swap. Reads SwapTotal from /proc/meminfo
now, which covers every kind.
It never looked at free disk. On a VPS with 4G free and 16G of RAM it would
fallocate 8G, fail, and take the run down under `set -e`. The recommendation is
now capped by what is actually there, keeping 5G back, and refuses rather than
shrinking to something useless.
fallocate was assumed to work. It produces a file that btrfs and zfs will not
swap on, so dd is the fallback — slow, but it always works.
swappiness was written by sed'ing /etc/sysctl.conf in place, tangling it with
whatever else lives there. It is a drop-in at /etc/sysctl.d now, so what this
script set is visible as its own file.
Role-dependent, which is the first use of MACHINE_ROLE: swappiness 10 on a
server, where swapping is the emergency valve and a page fault on a request path
is latency somebody is waiting for; the kernel default of 60 on dev, where
swapping out an application nobody has touched in an hour is exactly what you
want.
WSL is left alone entirely — WSL2 runs its own managed swap inside the VM, and a
swapfile written here is wasted disk the kernel will not use.
Sizes are reported rounded rather than floored. A 4 GiB swapfile is 4194300 kB,
which floors to 3 and reads as though a gigabyte went missing. Free disk stays
floored, deliberately: it decides how much to allocate, and rounding up invents
space.
Verified on this host — 4G RAM, 4G existing swap correctly detected and left
alone — and with the disk check stubbed at 6G free (caps to 1G) and 3G free
(refuses).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The original took whatever was typed and handed it straight to timedatectl. An
unknown zone — a typo, a guess at the spelling — fails there, and under `set -e`
that takes the whole run down four steps in. Names are now checked against
/usr/share/zoneinfo before use, and a bad one just re-asks.
It also never showed what the machine was already set to, and defaulted to option
1 (UTC) on Enter, so pressing return on a correctly-configured box silently moved
it. Now the current zone is printed, Enter keeps it, and a zone equal to the
current one reports nothing to do rather than setting it again.
timezone_current reads three sources — timedatectl, /etc/timezone, then the
/etc/localtime symlink — because they differ in availability rather than in
answer: timedatectl needs systemd, /etc/timezone is Debian's, and the symlink is
the one that is always there. timezone_set writes through timedatectl where
there is a systemd to talk to and the files directly otherwise, which is what it
would have written anyway; that is also the WSL path, where timedatectl exists
but does nothing.
Europe/Berlin added to the shortlist; TIMEZONE in the environment answers the
prompt ahead of time and is validated the same way, failing early with the bad
value named.
Verified detection (UTC here), validation of four names, and the env-var
rejection path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Was three unconditional lines that ran on every pass and reported success either
way. Now it checks, says what it found, and asks.
The original tracked one fact where there are two:
what a new login shell is told to use LANG in /etc/default/locale
whether that locale actually exists whether it has been generated
Setting the first without the second is what produces "setlocale: LC_ALL: cannot
change locale" on every ssh login and every perl invocation. They fail
differently, so the step names whichever one is actually missing rather than
reporting a flat "locale not set".
Also fixes two things the original would have hit on a minimal image:
locale-gen comes from the `locales` package, which cloud base images do not
ship and which is not in core utils. It is installed on demand rather than
assumed, instead of failing with "locale-gen: command not found".
The locale is uncommented in /etc/locale.gen rather than only passed to
locale-gen as an argument. A locale generated by argument alone disappears the
next time anything regenerates from that file.
`locale -a` prints en_US.utf8 where the configuration spells it en_US.UTF-8, so
both sides are folded before comparing — a literal match reports a working locale
as missing.
LOCALE in the environment overrides the default. pacman, dnf and brew branches
are written but unreachable while the pre-flight gate is apt-only; macOS has no
system locale to set and says so.
Verified both paths on this host: en_US.UTF-8 reports already set and generated,
pt_PT.UTF-8 correctly reports both facts missing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>