1eecc400a79af1fe88ee8405fcdbb0e629d2ef2c
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1eecc400a7 |
switch off Vaultwarden's routers, pending extraction into a plugin
Same treatment as the browser relay: mounts commented, code left on disk. Its
tables were already commented out of the schema earlier tonight, which is what
made this necessary — /api/vault was mounted against tables db:push no longer
creates, so a fresh core install shipped an endpoint that could only fail with a
Postgres "relation does not exist".
Vaultwarden is not one mount. Eight places had to go, and grepping for `vault`
found them only because several are not named after a router:
hono.ts /api/vault the authenticated reverse-proxy
/vaultwarden the unauthenticated one for the browser extension
VAULT_ONLY_PREFIXES loop /identity, /notifications, /icons, /events
the isBitwardenClient diverter an /api/* middleware that hands Bitwarden
clients to the vault router before anything else sees them
./api/vault/sidecar-server a SIDE-EFFECT import capturing the sidecar's port
UNPROTECTED_API_PREFIXES the '/vault' entry
server.tsx the 'vault' ws provider, its handler, and the notifications upgrade route
The side-effect import is the one worth naming: it registers a sidecar listener
and appears in no route table, so nothing about unmounting the routers would have
stopped it running.
No capability registry change, unlike browser and task-logs. Vaultwarden is
exempt from totality on both halves — EXEMPT_API_PREFIXES has '/vault'
("Bitwarden protocol clients authenticate to Vaultwarden, not to Officer") and
EXEMPT_WS_PROVIDERS has 'vault'. So nothing claims it and nothing breaks by
unmounting it. I said the opposite before checking; the check is what settled it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
595dd082a7 |
delete Task Logs
A complete read path over a table nothing could write to. task-logger.ts exported createTaskLog, appendToLog and finalizeLog, and none of the three was called anywhere in the tree — so `task_logs` could never gain a row, and the screen was permanently empty for everyone. The read half was fully wired: mounted router, capability claim, dock icon, two routes and a page-title rule. That is why it looked alive. Gone, in the order it was reached: Screens/Dashboard/TaskLogs/ the screen App.tsx /task-logs and /task-logs/:id Dashboard/index.tsx the export Layout/Dock.tsx the 'Logs' icon, and ScrollText with it state/usePageTitle.ts the title rule api/task-logs/task-logs.ts the router, and its mount in hono.ts api/task-logger.ts 101 lines of orphaned writer officer_db/src/operations/ the directory officer_db/src/schema.ts the export line officer_db/src/types.ts TaskLogSelect / TaskLogInsert capabilities/registry.ts loses '/task-logs' from the `tasks` capability's `api` AND `routes`. The api half is not optional: assertCapabilityTotality check 2 refuses to boot on a capability claiming a prefix nothing mounts, so unmounting the router while leaving the claim would have stopped the server starting. db:push now creates 21 tables, down from 43 at the start of the evening. officer_db/src/operations was one of the two lopsided directories the schema/queries merge exposed. integrations/ is the remaining one, and it is legitimate — it spans server and user-data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f9cd0a798a |
drop two dead tables: queue_jobs and terminal_containers
Neither had a reader or a writer anywhere in the tree. db:push created them on every fresh install and nothing ever touched them. queue_jobs is from when the background job engine kept its state in Postgres; it works on files now (src/servers/queue/storage.ts). terminal_containers held a docker id and port per user, from the architecture where every account ran in its own container — gone, as data-path.ts already records. Their types went with them: QueueJobSelect/Insert and TerminalContainerSelect/ Insert were unreferenced too. db:push now creates 22 tables. operations/schema.ts keeps task_logs and a note about what left and why, along with the thing still wrong there: it has no queries.ts, because task_logs is reached as `schema.taskLogs` from src/servers/api/task-logger.ts, past this package's boundary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
288b4bf683 |
each feature declares its own exports; the barrel is one line per feature
src/index.ts is 47 lines and reads as a list: `export * from './agent-panels';` and 22 more, under `db` and `schema`. It was 297 lines of hand-written named exports. What a feature exports now lives in <feature>/index.ts, beside the schema and queries it describes. Adding a query function is one file in one directory rather than that file plus a list three levels up that nothing enforces — a function missing from that list was invisible to all 107 consumers while existing and compiling perfectly. The old file had drifted in the ways a hand-maintained list does: twelve features were listed twice because values and types were separate statements repeating the path, soulseek three times, two features used an inline `type` specifier instead, and `db` and `schema` — the package's most fundamental exports — sat at line 270 with notify, types and app-store appended after them. The surface is byte-for-byte the same set. Checked rather than asserted: 247 exported names before, 247 after, no missing and no extra. The per-feature index files carry the same named lists the barrel did, so `export *` widens nothing. `operations` is deliberately not in the list, and the root file says why: it has a schema and no queries, its task_logs is reached as `schema.taskLogs` from src/servers past this package's boundary, and its other two tables are read by nothing at all. Verified: every file in the package parses, every relative import resolves to a real file, and all 107 consumers still parse. Not typechecked — empty node_modules, frozen installs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
68f2c55ecf |
one directory per feature: schema.ts and queries.ts together
src/databases/officer_db/src/<feature>/{schema.ts,queries.ts}, replacing the
parallel schema/ and queries/ trees. 24 feature directories, 46 files moved with
git mv so history follows.
The parallel trees had drifted, which is what the restructure is really fixing:
four features were named differently on each side — app-store/sidecar-installs,
email/email-accounts, server/server-config
operations had a schema and NO query file: its task_logs is reached directly
from src/servers/api/task-logger.ts, bypassing this package's own boundary
integrations had queries and NO schema, because it spans two features'
tables — server_integrations and user_integrations
Both lopsided cases survive as directories holding one file, which states the
problem instead of hiding it across two trees.
Nothing outside the package changed how it imports. `officerdb`, `officerdb/types`
and `officerdb/db` resolve exactly as before; index.ts absorbed the path changes.
Added `"./*": "./src/*"` so the new layout is reachable — `officerdb/soulseek/schema`
— which one script needed, because soulseek is a plugin and therefore commented
out of the aggregator.
schema/index.ts became src/schema.ts, keeping the core/plugin split from earlier
tonight. drizzle.config.ts and the package's "./schema" export follow it.
Verified rather than assumed: all 52 files in the package parse, every relative
import resolves against the new layout (checked by walking each specifier to a
real file, since parsing does not check paths), and everything in the tree
importing officerdb still parses. Not typechecked — empty node_modules, frozen
installs.
One rewrite bug worth recording: the rule mapping a query module's sibling import
also matched the './schema' this pass had just written, turning it into
'../schema/queries' in 22 files. Caught by the resolver check, not by parsing —
both spellings parse fine.
Also corrects every path reference the move invalidated: src/databases/CLAUDE.md's
layout diagram, the root CLAUDE.md data section, three docs, and seven sidecar
comments naming queries/<x>.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4cebae1c85 |
schema barrel: core tables only, plugin tables commented
db:push now creates 24 tables instead of 43. The 19 belonging to sidecars a light install does not run are commented out in schema/index.ts, kept as the record of what each table is called and which file defines it. Core is ecosystem.light plus officer-headscale, plus the app store and service_connections — app-store/effects.ts reads the latter, so it is core however few plugins exist. Commented: email, music, notify, dav, photos, jellyfin, invoiceshelf, soulseek, vault, wallet. The reason this is only a barrel edit is worth writing down. drizzle.config.ts points `schema` at schema/index.ts, so that file is drizzle-kit's view of the schema — but it is NOT the runtime's. Every query imports its table object from a schema file and calls db.select().from(table), and nothing anywhere uses drizzle's relational API (db.query.X), which is the only thing the `schema` passed to drizzle() in db.ts is for. So a commented line removes a table from the database without removing a line of code. That only held after moving nine query modules off the barrel and onto their own schema file — email-accounts, music, notify, photos, jellyfin, invoiceshelf, soulseek, vault and wallet all imported their tables from '../schema', so commenting the barrel would have broken the query module, then its export, then its consumers. dav already did it the right way. With that done the cascade is gone: everything still compiles, and the tables simply are not created. Two corrections to what was asked, both checked rather than assumed. The query barrel (src/index.ts) has no bearing on db:push — drizzle never reads it — so commenting it would not have kept a single table out of the database. And vault is not plugin-only from the platform's side: /api/vault is mounted top-level in hono.ts for the Bitwarden client, and api/vault/router.ts imports getVaultTokens. Its TABLES are out, but that route still exists and will fail against them. Not typechecked — empty node_modules, frozen installs — and db:push was not run, since there is no database here. Every file in the package parses, as does everything in the tree importing officerdb, and nothing outside the package reached a plugin table through the exported schema namespace. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9d6c196848 |
PUBLIC_URL comes back, and gen:index takes it as an argument
Removing PUBLIC_URL earlier today was wrong, and the Build section is where it would have surfaced: gen-index.ts exits 1 without it, so `bun gen:index` fails, index.gen.html is never written, and `bun start` has no page to serve. It came out because origin validation was being discontinued — but that was one of four consumers and the only one that is gone. gen-index needs it for OpenGraph tags, which crawlers fetch standalone and cannot resolve relative; task-api-env builds OFFICER_API_HOST from it; and dav/router hard-requires it, https only, to build an iOS profile. It is also the one value this machine genuinely cannot derive, which is what separates it from DATA_PATH and the rest that left today. gen:index now takes a URL as its first argument, ahead of the environment and .env: `bun gen:index https://officer.example.com`. Changing the public address is one command rather than an edit plus a regenerate, and a second address can be generated for without touching the install's .env. It also validates now. A relative or scheme-less value substituted silently and produced OpenGraph tags nothing can resolve — invisible until someone shares a link and the preview comes back blank. .env is PORT, PUBLIC_URL, POSTGRES_URL. Verified by running the section; all three paths through gen:index exercised (absent, valid argument, invalid). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8207824a81 |
build the secret store: one key per purpose, none in .env
.env now holds PORT and POSTGRES_URL. Every encryption and signing key lives in $OFFICER_ROOT/secrets/officer-keys.db — 0600, 0700 directory, owned by the service user, created on first use. The design doc planned to move ONE at-rest key into the store. What shipped splits it: headscale, wallet, photos, jellyfin, invoiceshelf, vault and service-connections each get their own, plus jwt. VAULT_STORE_KEY encrypted all seven, so one leak opened all of them — and it was named after whichever plugin needed it first, which is why it read as safe to change if you did not run a vault. A core install bootstraps two, jwt and headscale; the rest appear when their plugin first asks. The file IS the secret. No second key unlocks it, because a key beside the store it opens buys nothing. The gain was never secrecy, it is blast radius: bun auto-loads .env into all twenty pm2 processes, so a key there is readable from /proc/<pid>/environ of twenty processes — officer-music held the key that decrypts wallet seed envelopes. Two defects found by testing the store rather than reading it, both of which would have shipped: The WAL was 0644. Enabling WAL creates -wal and -shm at 0644 rather than inheriting the database's mode, and a freshly written key lives in the WAL before checkpoint — so the 0600 on the database was decorative. The 0700 directory covered it, but only until someone loosened the directory. PRAGMA journal_mode = WAL takes an exclusive lock, and busy_timeout was set AFTER it. With twelve concurrent openers, six died on that line with SQLITE_BUSY. Every sidecar opens this store at boot, so they open it simultaneously by definition: most of them would have failed to start on a cold boot and none on a warm one. Fixed by ordering the pragmas; re-tested with twelve racing processes, one key, one row. crypto.ts takes a purpose as its first argument now, which the design doc had explicitly promised would not happen — 32 call sites across seven query modules. That promise is corrected in the doc rather than quietly dropped. Also live, not just comments: wallet/upstream.ts gated wallet storage on process.env.VAULT_STORE_KEY and would have reported "unconfigured" forever. It asks the store now, and the question it answers changed — not "did somebody set a variable" but "can this process open the store", since the key is created on demand. assertSecretsClosed covers the store, its directory and its WAL. The jwt key mints owner tokens, so a member's shell reading it is strictly worse than the .env leak that check was written for. Not typechecked: node_modules is empty and installs are frozen, so the officerdb/secret-store subpath could not be resolved at runtime here — verified that officerdb/types fails identically, so it is the empty tree and not the new export. The store module itself was tested directly: creation, idempotence across processes, hasKey not creating, permissions, and the twelve-way race. Every changed file parses; the setup section runs and degrades correctly when the import is unavailable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1ccb21e67f |
CLAUDE.md: five capability kinds, not four
The file described a stricter model than the code enforces. It said terminal, chat, files, tasks, desktop and browser are all `kind: 'execution'` and therefore never shareable — but terminal, chat and files moved to `confined` on 2026-08-11 with per-user Linux accounts, and are grantable. The distinction matters and is now written down: a confined grant means nothing without a Linux user. authorize.ts:97 drops it for an account whose `osUser` is null, so "granted but unconfined" resolves to no access rather than to the owner's home — which is what it would otherwise resolve to, since getOwnerHomeDir ignores the email it is passed. Verified in the code, not inferred from the comment. Also brings the Data section in line with today: DATA_PATH, OFFICER_ITEMS_DIR and HOME_DIR are no longer environment variables, the install root is derived from cwd, and assertInstallLayout is why a wrong cwd fails instead of relocating the install. And `bun setup` runs officer-setup.sh, which is sections 1-6 of 10 — worth saying, since the entry read as though it were finished. Registry needs real work for where this is going. This is only the docs catching up to what is there now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9864fb1a49 |
switch off the browser relay, pending extraction into a plugin
BROWSER_RELAY_PORT is gone and the second listener no longer starts. The Chrome
extension, api/browser/ and the /browser screen all stay on disk — this is going
to be extracted, and deleting it means writing it again.
Three things had to move together, and the middle one would have failed the boot
on its own:
server.tsx the listener, commented out with the variable name recorded
hono.ts the /api/browser mount, closed
registry.ts the 'browser' capability's claim on /browser, dropped
assertCapabilityTotality checks both directions: check 2 refuses to start on a
capability claiming a prefix nothing serves. Unmounting the router alone would
have left the registry describing it, and the server would not have come up.
The capability itself survives because it also claims /scrape, which shares
nothing with the relay — it launches its own headless chromium through playwright
and never speaks to the extension.
The comment in server.tsx carries the two facts that are not recoverable by
reading the remaining code. First, the port is an INPUT TO A CREDENTIAL:
relay-auth.ts derives each extension's token as HMAC(JWT_SECRET,
'officer-browser-relay-v1:${port}:${userId}:${salt}'), so bringing the relay back
on a different number silently invalidates every paired browser — reported by the
extension as "Relay not reachable", which SETUP.md blames on a wrong address,
port or token. Second, it cannot come back as a kernel-assigned port:0 like the
other sidecars: the extension is configured by hand and stores the value, so a
port that moves each restart breaks the pairing each restart.
Left alone deliberately: the /browser route in App.tsx, its Dock entry, and the
Settings → Browser Relay panel. They will not work against a closed endpoint.
Removing them is frontend work for the extraction, not part of switching the
listener off.
.env is down to PORT and POSTGRES_URL.
Not typechecked (empty node_modules, frozen installs). Every changed file parses;
the setup section was run and writes two variables.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
571d0a62ff |
delete the bug-report discord webhook, and the last dead HOME_DIR reads
DISCORD_BUG_REPORT_WEBHOOK is gone, with sendToDiscord and its helpers. Reports still land in DATA_PATH/bug-reports — the disk write always happened first and the webhook was only a ping about it, so nothing about the report is lost. It was a personal notification channel living in deployment config, on a platform whose owner is the only person who files reports. It was also never in .env.example: the setup script wrote a variable nothing documented, which is the same drift as PORT, in the other direction. Note DISCORD_WEBHOOK_URL is a DIFFERENT variable — the notify sidecar's own channel — and is untouched. Then a parity sweep of setup / .env.example / what the code reads, which turned up two leftovers from earlier today: HOME_DIR was still read in six files, each with its own `?? homedir()` fallback. Dead since nothing sets it, but a dead read is worse than none — it reads as a supported override. They take homedir() directly now. user-instance.ts gets a comment on why its line stays where it is: it sits above `process.env.HOME = homeDir`, and homedir() reads $HOME, so a read moved below that assignment would return whichever member was last spawned into. Two of the six had fallback chains ending in process.cwd() and '' — the second would have silently disabled whatever consumed it rather than failing. VAULTWARDEN_URL was uncommented in .env.example among the variables setup writes, though it is a plugin variable setup has never written. Commented out with the other plugin entries. The three files now agree: setup writes PORT, BROWSER_RELAY_PORT and POSTGRES_URL; .env.example lists those plus JWT_SECRET and VAULT_STORE_KEY, which are required by code and deliberately unwritten until the secret store lands. Not typechecked (empty node_modules, frozen installs). Every changed file parses; the setup section was run and writes three variables. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f063fc0c08 |
remove origin validation
ALLOW_ANY_ORIGIN, ALLOW_ANY_ORIGIN_MUSIC, and everything they gated. The flag defaulted to ON, so none of it ran on a real install — what comes out is documented defence in depth that was already switched off. The file said so itself: "Both flags and their call sites come out once the tailnet is the perimeter." Origin was never authentication here in any case. An app's `officer://<hex>` origin is chosen by the client, forgeable outside a browser, and extractable from a shipped binary. Gone: the two flags, isOriginAllowed, isOriginCheckDisabled, isMusicOriginExempt, originValidationMiddleware, ORIGIN_RULES and the whole OFFICER_<APP>_ORIGIN scheme, PUBLIC_URL's origin/host derivation, and origin-validation.test.ts, which existed only to pin them. CORS now echoes whatever Origin it is given, which is what every install already did. What SURVIVES is the reason this needed care. origin-validation.ts held two unrelated things, and the second was the global authorization gate — a valid non-owner token reaches only what its role grants, deliberately NOT under the flag because it is account-based rather than origin-based. Its own comment called it "the airtight half". Deleting the file wholesale would have deleted authorization. So it moves to _middlewares/capability-gate.ts as capabilityGateMiddleware, with the name matching what it does: nothing in it reads an Origin header any more. hono.ts mounts it in the same position, ahead of every router. origin-middleware.ts stays and is untouched — it extracts the Origin for six auth handlers that log it, and for passkeys. Extraction, not validation. Also updates every claim that rested on the old model: CLAUDE.md's security section and repo map, docs/secret-store.md, docs/mobile-api-keys.md, and five messages in machine-setup's Tailscale section which told the owner to set ALLOW_ANY_ORIGIN=false when declining a tailnet. That advice is now impossible to follow, and the honest version is different: with no tailnet the token is the whole lock, so put a proxy in front and restrict who can reach it. Not typechecked (empty node_modules, frozen installs). Every changed file parses; the setup section was run and writes four variables now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5adb4aa08 |
the anthropic proxy binds PORT + 1
ANTHROPIC_PROXY_PORT is gone. It was 5051 hardcoded in four files: proxy.ts, which binds it, and three others that guessed the same constant to find it. It is the only sidecar that binds a fixed port, and that part is a real constraint rather than an oversight. Every other one binds `port: 0`, lets the kernel choose and reports back over the registration socket — which works because their consumer is the platform. The proxy's consumer is `claude`, spawned by a different pm2 process that needs ANTHROPIC_BASE_URL at spawn time and has no channel to ask what port the proxy landed on. Two processes with nothing between them have to agree in advance. So the number must be predictable, but it need not be 5051 — a value chosen against nothing, in the registered range, free to collide with anything the owner installs later. The symptom of that collision would have been chat failing while the rest of the platform looked healthy. PORT + 1 keeps the predictability and drops both the constant and the variable. Nothing to set, no second number to keep in agreement with the first, and the pair moves together when the install moves. Also corrects .env.example, which said the proxy "holds the API credential, which lives in the host env". It does not. The upstream credential is the OAuth token claude writes to ~/.claude/.credentials.json, and the ANTHROPIC_API_KEY the agent presents is the proxy's own generated secret. Verified the derivation at PORT=9000 and PORT=10000; all four consumers now import it; every edited file parses. Still not typechecked — empty node_modules, frozen installs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3c7f52ab77 |
PORT is read in one place, and it has no default
officer-url.mjs is now the only file in the tree that touches process.env.PORT.
Twenty-two others read it and supplied their own default; a value with
twenty-two sources is not configuration, it is twenty-two things to keep in sync,
and they had already drifted three ways.
It throws when PORT is unset rather than guessing. A default only covers the case
where .env was never loaded — which is not a machine anyone wants running,
because POSTGRES_URL is missing in the same breath. What the default bought was a
process that starts, binds somewhere unexpected, and fails later for a reason
that does not name the cause. Same posture as jwt.ts with JWT_SECRET.
It is .mjs, not .ts, and that is the whole reason this could be one file. pm2
launches officer-pty with node (ecosystem.config.cjs) and everything else with
bun; node cannot import TypeScript, so a .ts module would have left the pty
sidecar holding the only surviving copy of the default — precisely the thing
being removed. allowJs is already on, so the TS callers still get types. Verified
both runtimes import it, and that PUBLIC_URL-style overrides still work.
It also exports API_URL and OFFICER_API_URL, because nineteen sidecars were
independently building `ws://127.0.0.1:${PORT}` and two more were building the
http form. Those are one listener described in two protocols — no sidecar binds
anything — so they belong beside the port rather than being rediscovered per
file.
server.tsx now takes PORT as a number, so Number(PORT) at the serve site is gone.
Not typechecked (empty node_modules, frozen installs). Every edited file parses
under `bun build --no-bundle`; node and bun both load the new module; the unset
and non-numeric paths were exercised; the pm2 profile still loads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e72cae4830 |
one default port, and it is 9000
Every PORT fallback in the tree now says 9000. There were three answers to one question, each defensible where it was written and none of them visible from the others: server.tsx and 19 sidecars 5000 a default from before there was an installer user-instance.ts 9010 what scripts/setup-old/setup.sh really wrote .env.example 9000 what we told people to write 5000 goes first because macOS binds it — AirPlay Receiver has owned it since Monterey, so a dev server there fails to bind or gets shadowed by something that answers. All 22 sites moved together, which is the point. Changing the app alone would have turned a consistent-but-wrong default into a split one: the app on 9000 while nineteen sidecars still dialled 5000. 9010 was the interesting one. It was the only value that ever matched a real machine, because it is what the old installer wrote — and it was in the single file whose disagreement would have broken chat alone, with nothing else looking wrong. Its own comment records the same bug being fixed once already, within the file, by a change that left it disagreeing with everything outside it. Note what these defaults actually are: the sidecars bind nothing. user-instance.ts has no listener at all — it builds ws:// and http:// URLs that both address the app's single listener. So every one of these numbers is a guess at where the app is, for a value that .env always supplies. Worth removing rather than aligning, which is a separate change. Not typechecked (empty node_modules, frozen installs). Every edited file parses under `bun build --no-bundle`; the pm2 profile loads. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3bc06bba3c |
spell the root derivation as resolve rather than dirname
resolve(process.cwd(), '..') instead of dirname(process.cwd()). Identical on every input — checked including trailing slash and filesystem root — and it reads as the path arithmetic it is. resolve was already imported here for SEED_PATH. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f071c0b24 |
the install root is derived, not configured
Seven variables out of .env. DATA_PATH, OFFICER_ITEMS_DIR and HOME_DIR are gone from the code entirely; PUBLIC_URL, PUBLIC_BUILD_ENV, JWT_SECRET and VAULT_STORE_KEY are no longer written by the setup script. data-path.ts now derives OFFICER_ROOT as dirname(process.cwd()), with data/, capabilities/ and dockers/ as fixed names under it. The direction used to run the other way — DATA_PATH from env, then OFFICER_ROOT = dirname(DATA_PATH) in app-store/paths.ts — which meant three environment variables that had to agree with each other and with the tree on disk. Eight files re-read process.env.DATA_PATH independently, each with its own `?? cwd()/data` fallback. They import the one value now, which is what made removing it safe: otherwise each would have derived its own and drifted. Three things this turned up. The cwd pin in ecosystem.profile.cjs was broken. It set `cwd: __dirname` under a comment asserting "__dirname is the repo root — this file sits beside ecosystem.config.cjs", which stopped being true when these files moved into ecosystem-files/. It walks up to the platform's package.json now, which holds wherever the file lives. That was a live bug before this change and a load-bearing one after it, since cwd now decides where the install is. assertInstallLayout joins the other two boot assertions. A wrong cwd does not error — it computes a plausible root somewhere else and writes managed homes and agent runs into it, so the install looks empty and the data looks lost with nothing naming the cause. It throws before serve(), first of the three, because a wrong answer there makes the other two check the wrong files. getOwnerHomeDir captures homedir() once at module load rather than per call. Measured on bun 1.3.10: both os.homedir() and os.userInfo().homedir return $HOME when set rather than reading passwd, and user-instance.ts assigns process.env.HOME on its way to spawning an agent. A lazy read would have returned the owner's home on the first call and a member's afterwards. data-path.ts imports only node builtins, so it is evaluated before any of that runs. JWT_SECRET and VAULT_STORE_KEY leaving .env means an install made by this script does not boot — jwt.ts throws at module load without one. That is the agreed sequencing: they move to the SQLite store (docs/secret-store.md), and writing them here meanwhile would create a second origin for a secret the store then has to be reconciled with. Said plainly in .env.example and in lib/env.sh rather than left to be discovered. Not typechecked: node_modules is empty here and installs are frozen. Every edited file parses under `bun build --no-bundle`; the profile loads and pins the right cwd; assertInstallLayout was exercised from both the repo and /tmp; the setup section was run and writes five variables. Prettier was NOT run — 3.9.6 via bunx is not the pinned resolution and reformatted unrelated unions and line wraps in six files, so those were reverted and the edits re-applied by hand. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
86079adb9a |
mail transport is configured in the app, not in .env
MAIL_TRANSPORT was a fallback left from the old registration flow that sent
confirmation mail. That flow is gone; the variable outlived it.
It was never the primary source anyway. getTransport reads server_config
('server-settings' → smtp) first, which already backs a full UI at Settings →
Server → SMTP and its API in api/server-settings/smtp.ts, supporting resend,
smtp and mailhog. The env var only answered when that was absent — which is a
second source of truth for something the owner can already set, with the failure
mode that a stale URL in .env silently answers for a server whose settings row
is simply empty.
Removed from transport.ts, .env.example and the setup script's Environment
section, which no longer asks for it. setup-old/ still mentions it; that is the
archive and is left alone.
Also split the try. It wrapped the read AND the transport construction and
swallowed both, so three different problems produced one message. Unreachable
database, nothing configured, and stored settings that do not build a transport
now say different things, because the fix for each is different and this message
is all the caller ever sees.
The two consumers — queue/engine.ts and auth/forgot-password.ts — now raise
until SMTP is set in the UI, which is the honest answer rather than a regression.
Not typechecked: node_modules is empty here and installs are frozen. transport.ts
parses under `bun build --no-bundle`; the setup script was run and no longer
prompts for or writes the variable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
040ea41dbc |
per-user linux accounts are not optional any more
OFFICER_OS_USERS is gone. The platform behaves as it always would have with the flag on, and there is nothing to enable. Six conditionals, five of which were dead weight — provisionOsAccount, deprovisionOsAccount and the create/delete paths each opened with an early "not enabled on this server" return, and the API told the frontend whether to render the Linux controls at all. Those go, along with the 'disabled' DeprovisionResult stage, which nothing can produce now. The sixth is the one with teeth. assertSecretsClosed opened with `if (!OS_USERS_ENABLED) return`, described in its own comment as "a no-op when the feature is off, so an existing install is unaffected until the owner opts in". It is now unconditional: the server refuses to boot while any .env in the project root is group- or world-readable. A member's shell reading .env and printing JWT_SECRET was confirmed exploitable when this check was written, and a prerequisite that only holds when somebody remembers to set a variable is not a prerequisite. Nothing to remove on the environment side — the flag was never in .env.example or in the setup script. Not typechecked: node_modules is empty in this tree and installs are frozen, so tsgo could not run. All six files parse under `bun build --no-bundle`, and the changes are deletions of dead branches plus one removed early return. Formatted with prettier 3.9.6 via bunx rather than the pinned resolution, for the same reason; its one unrelated reformat was reverted by hand. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
32af97e260 |
secret-store: the anthropic proxy secret moves in too
Third value for the store, agreed in conversation. Neither an encryption nor a signing key — a bearer credential the agent presents to the proxy on localhost — but it qualifies on the same properties: generated once, shared between two core processes, fatal to regenerate silently. It makes the case better than the other two, because it is not in .env. It is in $DATA_PATH/sidecar/claude-state.json, which is the exact location decision 3 rules out by name: DATA_PATH is what gets backed up. Also records the rename. ANTHROPIC_API_KEY is wrong in both halves — not Anthropic's, not an API key; Anthropic's real credential is the OAuth token in ~/.claude/.credentials.json that the proxy swaps this one for. It is anthropic-proxy-secret everywhere we control, and keeps the CLI's name only on the assignment `claude` itself reads. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cb67b22f80 |
officer-setup: the environment section
Writes .env, and the whole point of the section is the two values it must not write twice. JWT_SECRET and VAULT_STORE_KEY are read back from any existing .env and kept. The original script reminted JWT_SECRET on every run that agreed to regenerate .env, which logs every device out with no stated reason, and never wrote VAULT_STORE_KEY at all — so a scripted install had no at-rest key and the vault and wallet refused to store anything. VAULT_STORE_KEY is the more dangerous of the two now that it is being written. It is not Vaultwarden's despite the name: it encrypts every secret column in Postgres, and the wallet seed envelope on top of the owner passphrase. Changing it is unrecoverable for the seed, because the passphrase opens the inner envelope and that is the outer one. Said in the section, in the file it writes, and in .env.example, which described it as Vaultwarden's and understated it. DATA_PATH and OFFICER_ITEMS_DIR are derived from $OFFICER_ROOT rather than asked — two questions that had to agree with each other and with the app store. ALLOW_ANY_ORIGIN is written explicitly from whether tailscale0 exists, rather than left to the platform default. The default is ON, which CLAUDE.md says is only defensible because the tailnet is the perimeter; with no tailnet there is no perimeter, so it goes out as false. Added to .env.example, which omitted it. PORT defaults to 9000, matching .env.example. The old script used 9010; nothing depends on either, and it is a prompt. write_env restores the prior umask. It was set to 077 so the secrets are never briefly world-readable, but umask is not scoped to a function and would have made every file the later sections create owner-only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9eabc3ee4c |
design: the secret store
Written from the conversation of 2026-08-12. Nothing implemented; every claim about the current tree was checked on that date. The short version: secrets stay in Postgres, the keys move out of .env into a small SQLite store, and rotation becomes an operation instead of data loss. What is actually wrong today is narrower than "secrets in .env" and worth stating precisely, because the separation that exists is correct and must survive a refactor: secrets live in Postgres and the key that opens them does not. The problem is blast radius across processes — Bun auto-loads .env, so VAULT_STORE_KEY sits in the environment of all twenty pm2 processes, and officer-music holds the key that decrypts wallet seed envelopes for no reason. The document records the decisions and, more usefully, what was ruled out: Keys cannot go in Postgres. A dump would carry the ciphertext and the thing that opens it. Encrypting the key with a second key only moves the question — one secret has to be readable without any other, and the only decision is where it lives. The store is not encrypted at rest, and this was tested rather than assumed: stock SQLite silently ignores unknown pragmas, so `PRAGMA key` succeeds, encrypts nothing, and the value is readable with `strings`. bun:sqlite ships stock SQLite 3.53.0. Whole-file encryption needs SQLCipher, which is a second native dependency, and this project already knows what one of those costs. SQLite rather than a flat file for rotation, not secrecy: rotation needs key VERSIONS, since an interrupted rotation needs the old key and the new one to both exist. The file must not live in $OFFICER_ROOT/data/ — that is what people back up, and a key store in the same tarball as a database dump rebuilds the problem. It also records the core/plugin split the design assumes: light plus officer-headscale is the core, because CLAUDE.md rests the security model on the tailnet and a model that rests on the tailnet cannot treat administering it as optional. Vaultwarden and the wallet become plugins. Moving headscale into light removes it from the app store automatically, since catalogue.test.ts asserts the catalogue equals full minus light. Five open questions are left open rather than guessed at. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0bceace6f1 |
an officerdev docker network, and the reasoning next to the port binding
One network for everything Officer provisions, created before anything joins it and declared external in the compose file. Postgres needs nothing from it today — the platform is a host process reaching it over loopback — but a reverse proxy in front of the web UI does, and so does any app-store service that talks to another. Creating it now means the later ones do not have to be migrated onto it. Two things written next to the line they explain, rather than assumed: Why loopback. Publishing a port makes Docker write its own DNAT and ACCEPT rules into iptables, and those are evaluated BEFORE ufw sees the packet — so `ports: "5432:5432"` is reachable from the internet while `ufw status` reports everything denied. That is the same mechanism the machine-setup firewall section hooks DOCKER-USER to close. Binding to 127.0.0.1 sidesteps it: the DNAT rule only matches traffic arriving on loopback. Why the password is not decoration. Loopback means nothing off this machine, but every account ON it can open 127.0.0.1:5432 — including the per-user Linux accounts Officer gives its members. What stops them is that they cannot authenticate. The password is the boundary between the platform and anyone with a login here, which is why it stays random and why both files holding it are 0600. A unix socket would remove even that, and was ruled out for a specific reason: postgres.js only treats a host as a socket path when the host FIELD contains a slash (src/index.js:468), and officer_db/src/db.ts passes a bare URL string. It would take a change to db.ts, which is not a setup-script change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a64d5610e6 |
officer-setup section 5: Postgres, and only Postgres
The original offered five containers. Of those:
Postgres the only database Officer has — account, passkeys, settings,
dashboards, email accounts, the queue. Required.
Redis not referenced anywhere in the platform. No import of the client,
no environment variable, no mention; the only "redis" string in
src/ is the word "rediscover" in a comment. Dropped. (It is still
in package.json and comes out in the dependency pass.)
SearXNG zero references anywhere. Dropped. If it ever arrives it brings
its own compose file and its own Redis with it.
Mailhog a development convenience, offered separately rather than here.
Nginx PM a deployment choice — Caddy, Traefik, nginx or the tailnet — and
not something a setup script should pick.
Provisioned into $OFFICER_ROOT/dockers/postgres/, the same convention the app
store uses: one directory per service, the compose file in it, relative bind
mounts so the data sits beside the compose file.
Bound to 127.0.0.1, deliberately and with the reason in the compose file itself.
Docker publishes ports by writing iptables rules underneath ufw, so "5432:5432"
is reachable from the internet whatever the firewall reports — the same mechanism
the machine-setup firewall section exists to close. The platform runs on this
machine, so loopback is all it needs.
The password lives in a 0600 .env beside the compose file rather than inside it,
so the compose file can be read or copied without carrying a credential. A second
run reuses it rather than minting a new one, which would leave the container and
the URL disagreeing.
Readiness is waited for rather than assumed: Postgres initialises its data
directory on first start, and db:push against a database that is still starting
fails in a way that reads as a schema problem.
Choosing an existing database checks the URL but does not insist on it — the URL
may be right and the database not yet started, and refusing to continue over that
would be worse than saying so.
Also carries the whitespace fix for the comment removed in the previous commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
911b79b6a2 |
drop a comment describing a machine that no longer exists
paths.ts carried a parenthetical explaining that the development machine had the layout inverted — the project inside ~/dockers/officer.dev/, so the root derived to officer.dev and the app store's directory came out as a dockers inside a dockers. That machine is gone. The project sits at ~/officerdev/platform, which is the clean shape the comment said new installs would get. Anyone reading it now goes looking for a directory that is not there and comes away unsure whether the derivation can be trusted. The rule above it is unchanged and is the whole contract: data/ is a direct child of the root, and OFFICER_ROOT is dirname(DATA_PATH). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0b083db8f |
officer-setup section 2: one root, and nothing configurable underneath it
$OFFICER_ROOT/
platform/ the app
data/ managed homes, attachments, job logs
dockers/ anything the app store provisions
capabilities/ skills, tools, tasks, processes
The original asked separately for DATA_PATH and OFFICER_ITEMS_DIR and left the
app store's directory implicit — three answers that had to agree with each other,
given by somebody with no reason to know they had to. One question now, at the
top of the run, and the rest follows from it.
This is also what the code already assumes rather than a new convention:
app-store/paths.ts derives OFFICER_ROOT as dirname(DATA_PATH) and DOCKERS_DIR as
OFFICER_ROOT/dockers, so writing DATA_PATH=<root>/data is the whole of what makes
the layout correct. No code changes.
Anybody who wants data/ on a bigger volume can symlink it. That is a decision
about storage, not about how Officer is laid out, and it does not need a prompt.
The one check worth having: a directory that exists but belongs to somebody else.
That happens when an earlier run created it as root, and everything written into
it afterwards fails in a way that reads as a permissions bug in the platform
rather than as a bad directory. Reported with what writes there and offered as a
chown.
Placed before the repository, because the checkout lands inside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8a0946ea39 |
officer-setup section 3: dependencies
bun install, as the account, in the checkout. Two things stated because a failure here is otherwise opaque. The lockfile is frozen — bunfig.toml sets [install] frozenLockfile = true — so bun resolves from bun.lock and nothing else. A package.json that disagrees with it is a hard failure rather than a quiet resolution, which is deliberate: the friction exists so an unexplained lockfile change shows up in a diff. If the install fails complaining about the lockfile, the section says that is the frozen lockfile working and that it wants a human to read the diff, rather than reporting a generic failure. node-pty has no Linux prebuild, so this compiles it from source on every machine. That is what build-essential and python3 are in machine-setup's core utils for, and the section says so — the failure would otherwise surface much later as a terminal that never starts. Success is checked by the artefact rather than by the exit status: bun can complete while the native module is not built, because it skips a dependency's lifecycle scripts unless it trusts the package. So the section looks for node_modules/node-pty/build/Release/*.node and, when it is missing, names the consequence and the command that fixes it instead of reporting success. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ee21fa16f0 |
officer-setup: the repository URL is https
https://gitea.officer.dev/officerdev/platform.git, not the ssh form. The reachability check is now one test for either scheme: `git ls-remote` with both prompts disabled. That is the real question — not whether the host answers but whether this account can read the repository — and neither prompt fails cleanly on its own. Over https git asks for a username nobody is there to type; over ssh it asks for a password or stops on host-key verification. With GIT_TERMINAL_PROMPT=0 and BatchMode both off, an unreadable repository is an immediate non-zero rather than a hang. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5a8cd4e9c |
officer-setup section 2: the repository
Clones from ssh://git@gitea.officer.dev:2222/officerdev/platform.git, or uses the checkout already at $OFFICER_ROOT/platform. Cloned as the account, never as root. A repository owned by root is one the owner cannot pull, cannot commit in, and whose node_modules they cannot write — and every later section in this script writes into that directory as them. Three things it refuses to do quietly: It does not repoint an existing remote. This checkout points at gitea.pastilhas.dev rather than the new gitea.officer.dev; that is reported with the command to change it, because where somebody's work pushes to is their decision. It does not pull over uncommitted changes. A dirty tree means the pull is skipped and said so, rather than failing halfway or burying the work. It pulls with --ff-only, so a failure means the branch has diverged rather than that the network was down, and the message says which. SSH reachability is checked before the clone, not after. An ssh URL with no usable key does not fail cleanly: git prompts for a password nobody is there to type, or stops on host-key verification. BatchMode turns both into an immediate answer, and the check reads the server's response rather than the exit code — Gitea greets a successful authentication and then exits 1, so exit status alone reports success as failure. When the key is missing it offers the https form of the same URL, which works without a key if the repository is readable anonymously, and otherwise stops and says to add the key. Verified against the new host: ssh authentication from this account already works. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b724e3ffbe |
start officer-setup: pre-flight, inheriting what machine-setup already asked
The second half of the install, and a much smaller script than the original: of the old setup.sh's thirteen sections, seven are machine-setup's job now and two more were already removed. What is left is the repository, dependencies, the database, .env, the schema, the build and pm2. Pre-flight asks nothing on a normal run. machine-setup saves the account, the Officer path and the role beside itself, and this reads the same file — so machine-setup then officer-setup is two scripts and one set of answers. It prompts only where that file is absent, which is a supported case rather than an error: somebody may have provisioned the box their own way. It then checks the machine is actually ready — git, node, bun and pm2 required, docker optional — and reports all of them together with what each is for. Finding out about a missing bun three sections in, after a repository has been cloned and a database started, is a worse way to learn it. A missing required tool stops the run and names machine-setup. Found by running it: a remembered answer can go stale. My own earlier testing had left SETUP_USERNAME=gitfresh in that file, for a throwaway account I then deleted, and the run dead-ended on it. A remembered account that no longer exists is a reason to ask again, not a reason to stop — so it is checked before it is trusted, reported, and replaced. Docker being absent is a warning rather than a failure: Postgres can be one you already run, and the app store simply cannot provision until Docker is there. Sections 2 to 9 are listed and not built. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e120dfa36e |
full read: fix the set -e footguns a full run would have hit
Read the whole thing — 2392 lines of entry point and 2700 of libraries — looking for what shellcheck cannot see. shellcheck itself is clean at error level; its warnings are cross-file false positives and one deliberate tilde in a display string. Everything below is a real defect. ── The Git section aborted on any machine where git was not already configured ── `git config --global --get <key>` exits NON-ZERO when the key is simply unset, and `VAR="$(git_get …)"` propagates that under `set -e`. So on a fresh machine — the case this script exists for — the section died at its first assignment, before printing anything, and took the remaining nine sections with it. It passed every earlier test because those harnesses sourced the section under a `bash -c` with no `set -e`. Verified now against a genuinely fresh account with the real script: the section completes and writes a correct .gitconfig. ── An optional step failing aborted the whole run ── Twelve functions ended on a command that can fail — `systemctl enable --now earlyoom`, `systemctl restart systemd-logind`, `chsh`, `sysctl -w`, `chown -R`, the oh-my-zsh installer, and others. Called as plain commands under `set -e`, any one of them failing ends the script, so a masked unit or a container without systemd would abort a 28-section run over an optional improvement. They now return 0 explicitly and the callers verify the outcome instead — which also fixed a lie: the sleep section printed "sleep disabled, logind reloaded" whether or not the restart had worked. It now checks the targets and the logind values and reports honestly. ── chown user:user assumed the primary group is named after the user ── True on Debian and Ubuntu, which create a group per user. Not true for an account from LDAP, or made with `useradd -g users`, or on an image with a shared group — there `install -g <user>` fails with "invalid group" and the step aborts. Proved it against an account whose primary group is `oddgroup`: the old form fails, the new one gets ownership right. Eight call sites now ask `id -gn`. ── Also hardened ── agent_path and current_editor gained `|| true` for the same reason git_get needed it: "nothing is set" is an answer, not a failure. Verified afterwards: shellcheck clean at error level, every section runs standalone without aborting, and the two apparent failures in that sweep are correct behaviour — Timezone and Git refusing an empty answer from /dev/null. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
30052e3295 |
port the firewall, last, and bind the Docker rules to the real interface
Last in the run for the reason the original gave: enabling a firewall is the one step that can cut the connection it is running over. The security bug is in the shipped rules. ufw-docker-rules.conf hardcodes eth0 in all three of its rules. Docker publishes container ports by writing its own iptables rules underneath ufw — DOCKER-USER is the hook that lets ufw have a say at all — so on a machine with predictable interface names (ens18, enp1s0, most VPS images) none of those rules match, the final DROP never fires, and every published port is open to the internet while `ufw status` reports active. A firewall that says it is working and is not is worse than no firewall. The rules are now substituted with the interface the machine actually uses, verified by applying them against a stubbed ens18. Order inside the section is the other thing that matters: OpenSSH is allowed BEFORE anything is enabled, unconditionally, because a firewall enabled without an ssh rule on a machine reached over ssh needs a console to fix. The prompt says so, and says to open a second session before closing the current one. tailscale0 is checked and offered, because the default is deny inbound and the tailnet is an inbound interface like any other — without that rule Officer is unreachable over the tailnet while Tailscale reports itself connected. A correction to something I said while writing this: I reported that this host was missing its tailscale0 rule. It is not. I had run `ufw status verbose | head -8`, which cut the output above the rule list. The full status shows it allowed, and nothing was wrong. That leaves the NOT PORTED list empty. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d6a78d7b5a |
put the shell configuration in one place, after everything it configures
The Shell section moves to 26, after Neovim, the runtimes and the agent CLIs. Everything that writes to .zshrc now happens there and only there: the starship init, the PATH for the agent CLIs (moved out of that section), the aliases, and the editor. That ordering is what the editor choice needs — it offers whichever of nvim, vim and nano are actually present, so it has to run after Neovim is installed rather than naming an editor that is not there. Which was the original's mistake in the other direction: it set core.editor to nvim four sections before installing it. The default editor is the setting git's core.editor was deliberately left out in favour of. EDITOR, VISUAL and SUDO_EDITOR go in the account's shell, and the Debian `editor` alternative is set too — an account's shell config cannot reach root or sudoedit, and those are exactly the cases where the wrong editor is most annoying. Recorded a limitation of append_once while cleaning up after it: renaming a marker orphans the block that used the old name, and changing a block's content does nothing because the marker is still found. Both need the old block removed by hand. This run left exactly that — a `local-bin` block superseded by `agent-clis` — in the dev box's .zshrc, now removed. UFW is deliberately still unported and will be last, for the reason the original gave: it is the one step that can cut the connection the run is happening over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
607a962115 |
port the agent CLIs, using Anthropic's installer rather than npm
Claude Code goes in through https://claude.ai/install.sh, matching what the platform already does for members in os-user-claude.ts and chosen there for the auto-update npm does not give. The comment at os-user-claude.ts:28 claiming setup.sh already did this was simply wrong — setup-ubuntu.sh used `npm install -g @anthropic-ai/claude-code`. Two things that installer insists on, both of which a naive port gets wrong and both of which os-user-claude.ts had already found: It REFUSES to run under sudo from a regular user's shell — it checks for uid 0 with SUDO_USER set, because everything it writes goes under $HOME and under sudo that is root's. This script runs as root, so the install has to be done AS the account. It declares #!/bin/bash and uses [[ … =~ … ]], so it must be piped to bash. On Ubuntu /bin/sh is dash and `| sh` fails. Two bugs found by running it rather than reading it: opencode does not install to ~/.local/bin. It goes to ~/.opencode/bin, which is what sidecar/opencode/index.ts:22 hardcodes. The first version looked in the wrong place, reported a working install as missing, and installed it again — the run said "did not complete" while the installer had plainly succeeded. claude on this machine came from npm, so `command -v claude` found it and the section would have left a copy that never updates. It now detects an npm install by resolving the binary into node_modules, says so, and offers to reinstall through the official installer — naming the npm copy and how to remove it rather than deleting something it did not put there. ~/.local/bin and ~/.opencode/bin are both added to the account's PATH. The sidecars do not need it — they check the exact paths — but a user who cannot run `claude` in their own terminal reasonably concludes it was never installed. PI stays optional and says outright that nothing in the platform spawns it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cb3e7062b9 |
ensure the bun symlink on every run, not only after installing it
The link was made as part of install_bun, so a machine that already had bun never got one. That is this machine: bun 1.3.14 in ~/.bun/bin, no /usr/local/bin/bun, and `bun` resolving to nothing at all for root. Nothing has broken yet only because `pm2 startup` has never been run here — the moment boot persistence is enabled, all twenty ecosystem apps that say `script: 'bun'` fail at boot and work perfectly when started by hand. ensure_bun_symlink now runs whether or not this script did the install, and says which of the three things happened: made it, found it already correct, or could not find bun to link. The last records an error, since a missing link is a reboot-shaped failure rather than a cosmetic one. Safe across upgrades, which was the question: a symlink resolves by path, not by inode, and `bun upgrade` replaces the file at $BUN_INSTALL/bin/bun rather than moving it. Demonstrated by replacing a target with a new file — new inode, link still resolves. It breaks only if the home directory goes, which breaks bun anyway. Also fixed the status line, which reported "not installed" on a machine with bun in the user's home: it asked root's PATH, which is exactly what has no bun before the link exists. bun_version now asks whichever copy is there. This run created the link on this machine — /usr/local/bin/bun -> the account's copy, and root can now run bun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00e58931ff |
port the JS runtimes: node, bun and pm2 installed rather than offered
Not a choice. Officer does not run without them, so asking would be asking
whether to install Officer — which was settled by running the script. They are
installed and reported, with the reason each is load-bearing stated once:
node pm2 is a Node application, and officer-pty compiles node-pty against
whatever Node is installed. There is no Linux prebuild, so this is not
an ABI question — it is a build dependency on every machine.
bun the platform itself and nineteen of the twenty pm2 apps.
pm2 supervises all of them, and the ecosystem files are written for it.
Node now tracks the current LTS, asked of nodejs.org, rather than the pinned
setup_22.x the original used — which ages into "the version we happened to pick"
the moment a new LTS lands. Resolves to v24.19.0 (Krypton) today, and NodeSource
publishes setup_24.x, checked with a HEAD request before anything is piped into a
shell.
Deno is the one genuine choice and stays optional, defaulting to no. Nothing in
Officer imports it — verified across the whole tree, the only references left are
in the old setup script — so the prompt says that outright and offers it for the
user's own work rather than pretending it is part of the platform.
bun is installed as the account and then symlinked into /usr/local/bin. pm2
started at boot by systemd has no login shell and therefore no ~/.bun/bin on
PATH; without the symlink every bun-based sidecar fails on reboot and works when
started by hand, which is a miserable thing to debug.
Every install is verified after it runs rather than trusting an exit status. A
NodeSource run can succeed while apt holds an older nodejs back, and reporting
the version asked for instead of the one present is how a machine ends up
disagreeing with its own setup log. Tested with an installer stubbed to succeed
and change nothing: both report failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
999362948f |
allow Node 22 or newer, and record why it was pinned to exactly 22
The preinstall check demanded exactly 22 — `v < 22 || v > 22` — which refuses Node 24, the current LTS. Relaxed to `>= 22`. Recording the reason it was exact, because it was deliberate and the details are gone: some months before now there was a real node-pty build failure that pinning to 22 solved. Nobody remembers what it was. That is exactly the kind of decision that gets undone twice, so it is written down here, in CLAUDE.md, and in the project memory rather than living in one person's recollection. What the evidence says now: node-pty 1.1.0 ships prebuilt binaries for darwin-arm64, darwin-x64, win32-arm64 and win32-x64 — and nothing for Linux. So its install script always falls through to `node-gyp rebuild` and compiles against whatever Node is installed. There is no prebuilt binary, so there is no ABI to mismatch, and node-pty declares no engines field. That reasoning is sound and completely untested: nothing here has built node-pty against 24, and node_modules has never existed on this machine. If `bun install` fails building it, or officer-pty cannot load its native module, restore the exact pin — CLAUDE.md says so, with the line to put back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
068a310bd3 |
add python3 to core utils — node-gyp needs it, and node-gyp is not optional
node-pty ships prebuilt binaries for darwin-arm64, darwin-x64, win32-arm64 and win32-x64. That is the complete list — there are no Linux prebuilds. So on Linux its install script always falls through to `node-gyp rebuild` and compiles from source, every time, on every machine. node-gyp needs Python 3. python3 was in the old setup.sh and I dropped it when rewriting the package list as "what the script itself would break without" — which missed that the thing it breaks is not this script but `bun install`, later, with an error about a Python that was never mentioned. The terminal sidecar then does not come up, and the reason is three steps removed from the symptom. build-essential was already there and is the other half of the same requirement; they are now noted together where they are declared. Found while answering whether node-pty constrains the Node version. It does not — with no prebuilt binary there is no ABI to mismatch, so it builds against whatever Node is installed, including 24. The exact-22 pin in package.json:11 is the platform's own choice, not node-pty's requirement. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0286cc6db6 |
port the Neovim section
Kept as it worked — upstream tarball, symlink, a config repo cloned into the account's ~/.config/nvim — with the defects fixed rather than the design changed. The one that mattered: the asset name. Neovim publishes nvim-linux-x86_64.tar.gz and nvim-linux-arm64.tar.gz. The original mapped aarch64 to "aarch64", which is not a name Neovim has ever published, so on an arm machine it downloaded a 404 and handed the HTML error page to tar. Verified against the release API — same class of bug as lazygit's hardcoded x86_64, and the second one this port has found in an arch mapping. The tarball is now checked with `tar -tzf` before anything is removed, so a bad download says what is wrong instead of failing inside tar. The rest: The tarball went to the working directory, via `curl -LO`, and stayed there if tar failed. It goes to /tmp and is cleaned up. The old /opt install was removed before the new one was known to be good. The download and its sanity check now come first, so a failed fetch leaves the working copy alone. The custom-repo option defaulted to git@gogs:andrepadez/nvim-config.git — a private repository nobody else can clone, and the same mistake as defaulting the login server to a personal headscale. No default now. git clone runs from /, for the reason git config does: the script's working directory is usually under the invoking user's home at 0750, which the target account cannot stat. The ~/.config/nvim/.git removal is now conditional on it being the starter. That is a template and dropping its history is right; a config of the user's own is something they will want to keep pulling. An existing config is left alone and said so, rather than moved to a .bak that silently overwrote the previous .bak. The PATH line the original appended to .zshrc is gone. /usr/local/bin/nvim is symlinked and already on PATH, so it was doing nothing except growing the file on every run. Verified on this host (already current, existing config left alone) and against a fresh account (LazyVim starter cloned, owned correctly, .git dropped). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5a33a517e7 |
move Tailscale ahead of the command-line tools
Now section 6, with the tools at 7. No dependency in either direction: Tailscale needs curl, which core utils installs at 5, and nothing in it touches lazydocker, lazygit, starship or fastfetch. Same reasoning as putting it early in the first place — it is a second way into the machine, so it should exist before anything that can go wrong does, and fetching four upstream binaries is a longer gap than it needs to sit behind. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2e70a6dd21 |
ask before replacing a config the user already has, and show the difference
Keeping theirs silently was safe but unhelpful: they never learn a newer version exists, and the only hint was a cp command printed in a warning. Now it asks. [1] keep yours — nothing changes [2] use ours — yours is kept as <file>.before-machine-setup [3] show me the difference first The diff is labelled "yours" and "ours" rather than by path, so - is what you would lose and + is what you would gain, and it goes through the pager because a config diff is routinely longer than a screen. Choosing to replace always keeps the old file beside the new one; nothing is destroyed. An unattended run — ASSUME_YES, or no terminal on stdin — keeps theirs and says so. "Yes to everything" cannot sensibly mean "overwrite configuration nobody was present to defend", so this is the one prompt ASSUME_YES answers conservatively rather than affirmatively. Applies to every file that goes through install_config, which is .tmux.conf and starship.toml today and is where any other dotfile should go. On the starship question: there is one file now, scripts/setup/starship.toml, and the inline copy is gone. Owner and members get the same prompt, which is what os-user-shell.ts always claimed. Verified through a pty, since the -t 0 guard correctly makes the interactive path untestable over a pipe: the diff renders, replacing writes the backup, and the live file ends up byte-identical to ours. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b9e16c11e |
port the shell section, and make one starship config serve both audiences
os-user-shell.ts:33 calls scripts/setup/starship.toml "the prompt config the owner's own install uses — one file, both audiences". It was not: the original machine script wrote a DIFFERENT config inline, so the owner got a prompt that only disabled language modules while every member got the repo file with its custom format. Two prompts, one comment claiming otherwise. This deploys the same file the platform does, which makes the comment true. Verified with cmp against a fresh account: byte-identical to what a member gets. Nothing overwrites any more: .config/starship.toml and .tmux.conf go through install_config, so they are written when absent, skipped when identical, and KEPT when they differ — with the cp printed, so taking ours stays the reader's decision. On this host that is what happens: the existing config differs and is left alone. The starship line in .zshrc is marker-wrapped by append_once. Verified over three consecutive runs: one block, not three. The original appended it unguarded every time. The login shell is now its own question. Having zsh on the machine and being handed it at every login are different decisions, and `chsh` made the second one silently. It also adds the shell to /etc/shells first, which chsh requires. .tmux.conf lives here now, with the rest of the dotfiles, rather than in user creation where the original put it only because that is where $USER_HOME first exists. One bug found by running it: install_config returns 2 for "kept yours", which is an outcome rather than a failure — but still non-zero, so calling it as a plain command under `set -e` ended the run before `case $?` could read it. Captured with && / || at both call sites, and the contract is documented where the function is defined so the next caller does not repeat it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
03cc01a135 |
say that stopping and coming back costs nothing
Choosing option 1 means leaving to set up a server elsewhere, which could take an afternoon. The run now says outright that Ctrl-C is fine and that returning picks up here — completed steps skipped, answers kept — so nobody feels they have to finish in one sitting or start over. The same promise once at the top, where the resume notice already was. That notice is now phrased as what it means rather than as file paths: "2 step(s) already done, and they will be skipped", with the command to start over instead of a bare mention of the file. On the question of why pre-flight kept re-asking: it does not, and I caused what you saw. Nearly every test command I have run today ended with `sudo rm -f .setup-answers`, so the file was deleted between your runs. Proved the round trip — a run with the variables set writes all three, and a second run with no environment at all asks nothing and prints what it remembered. The file is gitignored, so there was never a reason to be deleting it. Stopped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4fcc34de18 |
frame the public server as common to all three, not a cost of offscale
"offscale runs on a publicly reachable server" read as a demand offscale makes and the easy route does not. It is not. Tailscale's coordination server is publicly reachable too — they run it for you, and that is the entire difference between option 3 and hosting it yourself. Said that way round, the requirement stops being a reason not to self-host and becomes what self-hosting means. Same sentence in all three places: the menu entry, the branch taken when 1 is chosen, and the long answer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a989f8fbfa |
say where offscale runs, which matters more than how it installs
"One command to install" was the wrong emphasis and read as though it happens here. It does not: a coordination server has to be reachable by every device that joins, including phones on mobile data and laptops in other buildings, so it needs an address that resolves from anywhere. It goes on a small public VPS of its own — not this machine, and not behind a home router. That is the thing people get wrong, and getting it wrong produces a private network unreachable from exactly the devices it exists to reach. It is a property of being the thing everyone checks in with, so it is true of headscale too, and the long answer now says so. Choosing option 1 leads with it, links the install anchor rather than the page, and says plainly that nothing below will work until that server is up and answering — so somebody who has not done it stops here instead of typing an address that does not exist yet. "Installs in one command" survives in the goodies list, where it belongs: it is one command ON THAT SERVER, and the sentence now says so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c629a849d9 |
make all four network options work, including having no network at all
Option 1 no longer refuses. It assumes offscale is already running — installing
it is one command, documented on the site — and asks for its address and a key,
which is mechanically what option 2 does. The two share a branch because the
difference between them is what to say, not what to do: one is "you already have
a server", the other is "set one up first, here is where".
Option 4 is new: no private network. Presented as a real choice rather than a
failure to choose, with what it costs stated before it is taken and paged so it
is read rather than scrolled past:
· anything reachable remotely has to be published deliberately and kept closed
otherwise
· TLS certificates are yours to obtain and renew
· every exposed service needs its own authentication, since there is no longer
a boundary in front of it
· the machine will be found — anything on a public address is scanned within
minutes
And the one that is specific to this platform rather than general advice:
ALLOW_ANY_ORIGIN defaults ON, which is deliberate and only defensible because
the tailnet is the perimeter. With no tailnet it must be set to false with an
HTTPS proxy in front, or Officer runs with a check disabled on an assumption that
is no longer true. The summary line says so, so it survives the run.
Declining option 4 redraws the menu rather than dropping to a bare prompt.
Tailscale is still installed when 4 is chosen, and the run says how to connect it
later.
Verified all four end to end with tailscale stubbed, plus the decline-and-choose-
again path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
06492c3297 |
name the control server explicitly, and log out before moving between them
Two gaps found by tracing the assembled command rather than assuming it, and the first meant option 3 did not work at all on a machine like this one. `tailscale up` with no --login-server keeps whatever ControlURL is already stored. So on a node already pointed at a self-hosted server — which this host is — choosing "the easy route" left it exactly where it was. No error, no message, and a summary line claiming it had connected. The URL is now passed explicitly in both cases, TS_DEFAULT_CONTROL_URL for Tailscale's own service. And a node logged in to one coordination server cannot simply be pointed at another; it has to be logged out first. That is now detected by comparing the stored URL with the target, and offered rather than done quietly — the tailnet drops while it happens, and the run says so, because on a machine reached over the tailnet that is the session you are reading this in. Declining leaves the node where it is and records that. Verified all three paths with tailscale stubbed: switching logs out then connects to controlplane.tailscale.com, declining leaves it on offscale, and reconnecting to the SAME server offers no logout at all. Also noted while tracing: lan_cidr correctly finds nothing on this host, since a /32 with host routes has no subnet to advertise. That means the homelab subnet-router prompt is the one path here that has not been exercised on real hardware. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
22a661131d |
page the ? output
The long answer is now well past a screen, so it scrolls the question off the top and the reader lands at a prompt having lost what they were choosing between. Piped through a pager, so it is read a screen at a time and the menu is redrawn underneath it afterwards. `more` rather than `less`: it exits at the end of the file instead of sitting there waiting to be quit, which is right for something asked for once. Only when stdout is a terminal. Redirected or piped — a transcript, a log, the test harness — it comes through whole, since a pager there either blocks or mangles the output. Applied to confirm()'s help hook as well, so every ? in the script pages, not just this one. Verified both ways: driven through a pty it shows --More--, and with output piped all four help sections come through in full. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
85d8fa62f1 |
mention that Officer administers the network it just told you to set up
The section explains three ways to get a coordination server and then says nothing about running one afterwards, which is the part that decides whether self-hosting is a good idea. Officer's Headscale app is the answer to it, and it works against headscale and offscale alike. Written from what the app actually does rather than from the pitch: several servers registered and switched between, each PROBED rather than remembered — the comment in ServersView.tsx is explicit that a "not checked" dot is the one thing that list must never show — and, on the active one, nodes, users, pre-auth keys, invites and the ACL policy with an assistant, plus a console and diagnostics. Placed in FOR OFFICER rather than under offscale, because it is true of either self-hosted option and is the reason picking one is not a commitment to administering it over ssh. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
766c8ee655 |
say what the mobile app actually is, rather than overclaiming it
"our implementation of the same protocol rather than a wrapper around theirs" was too strong. It is the Tailscale client with our branding, and one real difference: it takes an invite from the server directly. The corrected version is not a weaker claim, it is a more specific one. "Our own implementation" invites the question of whether it is trustworthy and whether it keeps up; "the Tailscale client, our branding, and it takes an invite directly" answers both — it is their client, so it is as good as their client, and the thing it adds is the thing that was hard. The reason it is hard stays in, since it is what makes the difference worth naming: the official app has to be talked into using a server that is not Tailscale's, and that is where people abandon self-hosted headscale. And it still says there are no desktop apps of our own, now with what to do instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d544db07e |
write the offscale copy from officer.dev, and link it in both places
Read https://officer.dev/infrastructure/offscale.html and used its own words for the protocol claim — "the protocol on the wire is Tailscale's, the encryption is WireGuard's" — which is stronger than my paraphrase and is the sentence that stops "our own distribution" reading as a fork. The goodies are named rather than gestured at. No placeholders left: · installs in one command, with the certificates handled · health, logs, restarts and access policies from the app, instead of a config file and a CLI · enrolling a device is a link and a tap — the key is minted and handed over for you · several networks at once, and services reachable across them The URL appears twice, as asked: in the menu entry, where somebody deciding between three options can reach it without typing ?, and at the end of the long answer for somebody who read the whole thing and wants more. The mobile-app paragraph stays and is not from the page — the page does not name its client platforms. It is your account of them, kept because it answers the obvious objection to self-hosting, and it still says plainly that there are no desktop apps of our own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
46ce58dd12 |
correct the client claim: the mobile apps are ours
"talking to stock Tailscale clients" was wrong, and wrong in a way that gave away the strongest thing offscale has. On computers it is the stock client. On iPhone, iPad and Android it is our own app — our implementation of the same protocol, not a wrapper around theirs. No desktop app of our own yet, and the copy says so. Stated as a differentiator rather than a footnote, because it is the specific that answers the obvious objection to self-hosting. Getting the official mobile app to talk to a self-hosted server is the part of running headscale people give up at; having an app that simply does is worth more than any sentence about extras. The protocol claim is unchanged and still leads, because it is what makes the mobile app reassuring rather than alarming: our own client is our implementation of Tailscale's protocol, not a private one. "Our own distribution" plus "our own app" reads as a fork unless the first thing said is that there is nothing to fork. Marker narrowed from "goodies" to "more goodies" — one is now named. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2803bc34b7 |
draft the offscale copy: same protocol, different amount of work
Two things said plainly, because they are the two a reader needs before choosing: Nothing about the protocol changes. offscale is headscale's open-source code speaking to stock Tailscale clients, so a machine on an offscale network behaves exactly as it would on either of the others. Worth stating outright — "our own distribution" reads as a fork, and a fork of a network protocol is something to be wary of. There is no offscale protocol to be locked into, because there is no offscale protocol. What changes is the work. Running headscale yourself is a project: install it, put TLS in front of it, keep it upgraded, administer it through a config file and a CLI. offscale makes that a step in a setup script. "our own sugar on top" is gone. One marker left, in both the menu and the long answer: the goodies are unnamed. Two or three specifics would be worth more than the sentence they replace — every product claims extras, and the claim is only interesting when it says which. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
538a2d3b6d |
give every network option its own exposition, not just offscale
Each of the three now carries a couple of lines under it, separated by a blank line, so the choice can be made from the menu itself rather than by typing ? and reading a page. 2 and 3 are written; offscale's is a marked placeholder with your one-liner standing in until the longer copy arrives. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ff7035a47a |
correct the mechanism recorded for the set -e failure
The previous commit blamed the loop body. That is wrong, and I only found out by trying to reproduce it: the same `[[ … ]] && assign` inside a case inside a while loop survives `set -e` perfectly well at top level. What actually happened is one level further out. The failing assignment was the last thing the case ran, the case was the last thing the loop body ran, and the loop was the last thing THE FUNCTION ran — so load_answers returned non-zero, and calling a function that returns non-zero is a plain command failure, which does end the script. Worth getting right because the general rule is different from the one I wrote: it is not "avoid && in loops", it is "a function whose last statement can return non-zero fails when it is called, however innocuous the statement looks". Scanned the libraries for that shape. The only hit is lan_cidr, which ends in an awk pipeline and returns 0. Predicate functions ending in a bare test — ballast_exists, has_authorized_key and the rest — are meant to return non-zero and are only ever called in conditions, which set -e exempts. The fix itself was already correct and is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4d478cc0f1 |
add the offscale line, and fix a set -e bug the test exposed
The menu is one function now rather than being written out twice — it is shown
again after ? prints the long answer — and option 1 carries your line: offscale is
just tailscale and headscale, with our own sugar on top.
The bug it surfaced is the more useful half. load_answers used
[[ -z "${MACHINE_ROLE:-}" ]] && MACHINE_ROLE="$value"
as the last statement in a while-read loop body. When the variable is already set
the test is false, the compound returns non-zero, and as the final statement in a
loop body under `set -e` that ends the script. The failure is silent about its
cause: the trap prints "Step: unknown" and a line number inside the library,
before pre-flight has run.
It needed both conditions to appear — an answers file on disk AND the variables
already set in the environment — which is why every earlier test missed it and
running with env overrides hit it immediately. Written as if/then now, and the
other lib files scanned for the same shape.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6a29c39b74 |
restructure the Tailscale network choice, with offscale left to be written
Four options, in the order you gave: 1 set up your own network (offscale) 2 use a network you already run (headscale, offscale) 3 the easy route (tailscale.com) ? what are tailscale, headscale and offscale? ? prints the long answer and then shows the options again, rather than dropping the reader back at a bare prompt having forgotten what they were choosing between. The explanation frames all three as one question — who keeps the list of your machines and hands out the keys — and says plainly that the coordination server never carries traffic, since that is the thing people assume it does. Tailscale and headscale are written. OFFSCALE is a marked placeholder, and so is what option 1 actually does; both are yours to fill in and the run says so rather than pretending. Option 3 is the plain flow: no --login-server at all, and the auth-key prompt says what that means — leave it blank and Tailscale prints a link that either creates the account or adds this machine to an existing one. Option 2 keeps the "no suggested URL" rule, because a coordination server URL is somebody's private infrastructure. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
059f0f6df2 |
lead the Tailscale section with the question, not the explanation
Ten lines of prose before the first prompt assumed the reader had never heard of Tailscale. Anyone already running it does not need to be told what it is, and having to scroll past it every run is the cost of writing for the other reader. The prompt comes first now, and `?` is an answer. Typing it prints the full description and asks again; not typing it costs nothing. confirm() takes an optional help function as its third argument. Where one is given the prompt becomes [Y/n/?], so the explanation announces that it is available without taking up room. The same hook is there for any other section that wants it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a63a327065 |
remember the pre-flight answers between runs
The --only flag worked, in that it reached the section — but reaching it meant answering four pre-flight questions first, every time, which is not usable for working on one section. The same problem was already there without --only: a resumed run re-asked the role, the account and the Officer path that it had been told on the previous pass. Answers are saved beside the progress file and loaded before anything is asked. The environment still wins over what was saved, so SETUP_USERNAME=x on the command line overrides it, and --reask throws the file away and asks again. Read as assignments rather than sourced. The file sits next to the script and is read by a run that is already root; sourcing it would make it executable content in a place nothing guards. Second run now goes straight through pre-flight, printing what it remembered, to the one step asked for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
09303299cc |
port Tailscale as its own section, early, and stop it hanging
Moved to position 7 — after core utils, which give it curl, and well before SSH
hardening, which is the step that can lock you out. The argument is that
Tailscale is a second way into the machine, so it wants to exist before anything
that can go wrong does.
── Why the original hung, and what stops it now ──
Its prompt accepted an empty auth key and passed it anyway. `tailscale up
--authkey ""` falls back to the interactive flow: it prints a URL and blocks,
with no timeout, forever. From the outside that is a script that has frozen.
Nothing here passes an empty key — the flag is omitted entirely, and the run says
in advance that a URL is coming and that it will wait. Every call carries
--timeout=60s, and a timeout is reported with the command to run by hand rather
than left as silence. State is read with `tailscale status --json` before
anything is run, so a node that is already up is offered a reconfigure instead of
having `up` fired at it blindly.
Diagnosed on this host rather than guessed at, and honestly the diagnosis is
partial: the exit-node branch left no trace at all — no /etc/sysctl.d file,
networkd-dispatcher present but with zero mentions in apt history, so it came
with the image. ip_forward=1 came from the unconditional part of the section, not
the branch. That points at `tailscale up` as where it stopped, and the empty-key
path is the candidate that fits, but I could not reproduce it to be certain.
── What the section now covers ──
control plane Tailscale's own service by default; a self-hosted headscale as
an explicit choice with NO suggested URL. The original defaulted
to headscale.pastilhas.eu, so a stranger running it pointed
their machine at somebody else's control plane.
auth key, or the browser flow, stated as an equal option
Tailscale SSH ssh over the tailnet with no keys, governed by tailnet ACLs —
and pointed out as a way back in if the sshd hardening later in
the run goes wrong
subnet router homelab only, defaulting to this machine's actual LAN CIDR
exit node with what it means for whose traffic goes where
forwarding sysctls and Tailscale's recommended NIC offload settings, and
only when an exit node or a route actually needs them
Approval is mentioned: an advertised route or exit node does nothing until it is
approved in the admin console, which is otherwise a silent non-event.
Also added --only <step> and --list, because this section in particular needs to
be run on its own while it is being worked on. A step run that way ignores the
progress file and does not record itself — asking for one step is not progress
through the script.
One bash quirk fixed on the way: "${VAR:-Tailscale's own service}" does not
parse. An apostrophe inside a ${:-} default opens a quoted section that swallows
the closing brace, and the error surfaces as "unexpected EOF" 400 lines away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9d53ff506e |
port Docker, with the group-versus-rootless choice spelled out
Three options, each explained rather than named, because the difference between
them is a security posture and the default is the one that sounds harmless.
1 docker group, the default. The text says what the group actually is: anyone
in it can run `docker run -v /:/host -it alpine chroot /host` and have a root
shell. It is not "access to Docker", it is root by a longer route — the same
framing os-user-docker.ts already uses for why members never get it.
Whether that matters is conditional, and the run works it out rather than
asserting either way: on an account that already has sudo it is a shorter
path to something they can reach anyway, and it says so; on an account that
does not, it is a real escalation, and it says that instead. Caught in
testing, where the reassuring sentence was being printed for a throwaway
account with no sudo at all — the exact case where it is untrue.
2 rootless, with the thing nobody would find out stated at the prompt:
Officer's app store cannot provision containers with it. compose.ts,
preflight.ts and system-monitor all spawn `docker` with no environment of
their own, so they reach /var/run/docker.sock; DOCKER_HOST is set only for
member commands, in os-user-docker.ts. pm2 started at boot by systemd has no
session either, so exporting it in a shell rc does not reach the process
that matters. The consequence is recorded in the summary, not just spoken.
3 neither, and what that costs.
Also fixed in the port: the repository codename came from `lsb_release -cs`, which
is wrong on every derivative — Mint reports "vanessa", Pop reports its own, and
Docker publishes neither, so `apt update` fails against a repository that does not
exist. os-release carries UBUNTU_CODENAME on exactly those systems for exactly
this reason; it is preferred now, with VERSION_CODENAME as the fallback, and the
ubuntu/debian half of the URL comes from ID_LIKE rather than being hardcoded.
The shared `services` network is created only when missing, checked with
`docker network inspect` rather than by running create and discarding the error.
Verified on this host, and against a throwaway account both with and without
sudo.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f9c6b7925a |
record the default-editor section, to come last
Asks for nano, vim or nvim and sets EDITOR/VISUAL in the shell config, plus the Debian `editor` alternative so root and anything reading the system default agree with it. Has to come after the Neovim section, or nvim cannot honestly be offered as one of the choices — which is the same ordering mistake the original made by setting core.editor to nvim four sections before installing it. This is the setting core.editor was left out in favour of: one preference that git, crontab -e, visudo and systemctl edit all follow. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4339ab829d |
set pull.rebase, and say why core.editor is not set
pull.rebase true, which is the setting whose absence stopped this very repository mid-session today: git refuses to pull when branches have diverged and asks which of three things you meant, every time, until told once. Rebasing replays local commits on top of what was fetched rather than adding a merge commit that records nothing but the fact that you had not pulled yet. core.editor stays out, deliberately, and the run says so rather than leaving its absence to look like an oversight. Git's fallback chain is GIT_EDITOR → core.editor → $VISUAL → $EDITOR → system default, so core.editor is a git-only override sitting above $EDITOR. The original set both it and `export EDITOR` in the shell section — two settings for one preference, which drift apart the moment either is changed and leave git using an editor nothing else does. Setting only $EDITOR means git, crontab -e, visudo and systemctl edit all follow one answer. Verified on a throwaway account: .gitconfig comes out with user, init.defaultBranch master and pull.rebase true. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
28673869c3 |
port the git section, and stop it overwriting an identity that already exists
Detects first, then asks. If there is a name and email configured it shows them and offers to change them, defaulting to no — the common reason to re-run this script is everything except this. If there is nothing configured it just asks. Default branch is master rather than main, unless the answer says otherwise. Three problems in the original, beyond running unconditionally: prompt_value accepts an empty answer, so pressing Enter wrote `user.name = ""`. An empty name is worse than none: unset makes git refuse to commit and say why, empty makes it commit with a blank author and never mention it. ask_required re-asks instead. core.editor was set to nvim four sections before Neovim is installed, so anything invoking the editor in between failed. Dropped for now rather than moved — it is a preference, and worth deciding separately. Nothing checked whether the writes worked. That last one was not theoretical. Testing against a throwaway account, all three writes failed and the section still printed "OK: written". `git config --global` needs no repository, but git stats the working directory on the way, looking for one — and the script runs from under the invoking user's home, which is 0750, so the target account cannot stat it: fatal: failed to stat '<cwd>': Permission denied The wrappers now run in a subshell from /, which every account can stat, and the caller checks the exit status and reads the value back before claiming success. Also recorded where it is written: docs/agent-git-identity.md says every agent Officer runs commits as the owner, because it runs as the owner. This is not only the human's identity, it is what git log attributes agent commits to — which is worth knowing while choosing it. Verified both paths: this host's existing identity is shown and left alone by default, and a fresh account gets a correct .gitconfig owned by that account with defaultBranch master. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
08e7de1b6d |
move unattended-upgrades into core utils, and make sure it is actually on
apt only. It is a Debian and Ubuntu package — dnf's equivalent is dnf-automatic
and pacman has no equivalent at all — so it is not a name to translate across the
other lists.
Installing the package is not by itself enough to switch it on. The apt-daily
timers read /etc/apt/apt.conf.d/20auto-upgrades, and on this host no package owns
that file: `dpkg -S` says it came from nothing, which means the original script
wrote it. So the section checks for it and offers to write it, rather than
assuming the install did.
Beyond that it only reports, because the interesting facts about unattended
upgrades are not whether it installed:
It never reboots on its own, deliberately. A kernel or libc update is installed
and then not used, and the machine keeps running the old one until it restarts.
Nothing announces that except /var/run/reboot-required, which nobody reads. The
section prints it, names the packages waiting, and puts it in the summary — it
is the failure people do not notice for months.
Ubuntu's Allowed-Origins includes plain ${distro_codename} as well as
-security, so this takes ordinary updates too, not only security ones.
Verified both paths: this host reports enabled with no reboot pending, and
pointing AUTO_UPGRADES at a temp file exercises the enable path and writes the
three periodic settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
0da184d082 |
verify the fail2ban defaults instead of recalling them
The previous commit hedged on the ban policy because I thought fail2ban was not installed here. It is — the earlier ubuntu-setup run installed it — so the claim could be checked rather than remembered. Checked, and it was right: /etc/fail2ban/jail.d/defaults-debian.conf ships `[sshd] enabled = true`, and the running jail reports maxretry 5, findtime 600, bantime 600. This host has banned 6 addresses already. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f176b92378 |
move fail2ban into core utils, and reduce its section to a status report
Your call, and the reasoning holds: it configures nothing of its own, an existing install with its own jails is untouched because pkg_install never names a package that is already present, and it is worth having by default. One thing recorded where it is declared, because it makes fail2ban unlike every other entry in that list: it is a daemon, not a binary. Installing it starts it, and Debian and Ubuntu ship an enabled sshd jail — so from that moment an address that fails to log in five times in ten minutes is blocked for ten. That is the point of it, and it includes you, from wherever you are connecting. (Recalled rather than verified: fail2ban is not installed on this host and the sandbox would not let me unpack the .deb to check the shipped jail.d file.) The section no longer installs anything. It reports whether fail2ban is running, which jails are active, and how to unban an address — because a daemon quietly blocking connections is worth knowing about before it blocks yours, and a run that installs it as one name in a list of twenty gives no hint that anything started. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bdf13331ae |
port the static IP section, and offer the actual fix for the reboot-changes-IP problem
Homelab only, as it should always have been. On a vps the provider's DHCP is
authoritative and already stable, and pinning an address there is how an instance
is stranded; on dev the machine moves between networks and a fixed address is the
opposite of what is wanted. Both say so rather than skipping quietly.
The more useful change is that a static address is no longer the only answer
offered, because it is not the right one for the problem it was added to solve.
A fresh Ubuntu box taking a new IP on every reboot is not the router
misbehaving. systemd-networkd's ClientIdentifier defaults to `duid` — man
systemd.network is explicit — so the machine introduces itself to DHCP with an
RFC 4361 client ID built from an IAID and a DUID. This host shows it:
DHCP4 Client ID: IAID:0x56504d98/DUID
Consumer routers key leases and reservations on the MAC. The two never match, so
the router does not recognise the machine as one it has seen and hands out the
next free address — and a reservation pinned to the MAC is never honoured, which
is the part that makes the router look broken.
`dhcp-identifier: mac` in netplan sets ClientIdentifier=mac and the router sees
what it expects. DHCP keeps working, reservations start being honoured, and
nothing is pinned on the machine. That is now the first option, with the static
address second and still carrying the original's warnings.
It is written as its own 99- netplan file and merged with whatever the installer
or cloud-init already wrote, rather than this script parsing and rewriting their
YAML. Deliberately NOT applied: it takes effect at the next reboot, which is the
moment the problem shows up anyway, so there is nothing to gain by dropping the
network now. `netplan generate` validates before either file is kept, and the
file is removed again if it does not.
Found and fixed while testing: dhcp_client_identifier parsed the value with
awk -F': *' and took field 2 — but the value is itself "IAID:0x…/DUID", so the
split yielded "IAID", the DUID test failed, and the helper reported "mac" on a
machine that was plainly sending a DUID. It would have told the user the opposite
of the truth about their own problem. Reads everything after the first colon now.
Verified on this host: correctly reports duid, explains why, and renders both the
homelab and vps paths.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7a5cf89819 |
port the DNS section, as a choice rather than a decision
The original hardcoded Cloudflare plus Google with no way to say otherwise, and
rewrote /etc/systemd/resolved.conf wholesale — discarding DNSSEC, DNSOverTLS,
Domains and Cache if anything had set them, without mentioning it had. The
settings are a drop-in now, and the resolver is picked from a list with a
"keep what is there" that is the default.
The part worth having explicit is which layer is being changed. With
systemd-resolved there are two:
per-link what DHCP handed each interface, and what Tailscale installs on its
own. These answer for that link's domains — the provider's internal
names, the tailnet — and are printed by this step precisely to show
they are NOT being touched. Overriding them is how private
networking quietly stops resolving.
global the resolver used when no link claims the query. This is the one
the step sets.
On this host that distinction is live: eth0 has Hetzner's resolvers and
tailscale0 has 100.100.100.100, which is what answers ts.pastilhas.dev. Both are
left alone.
The drop-in is named 99- because systemd reads drop-ins in lexical order and the
LAST value wins. That is the opposite of sshd, whose drop-in three files away in
this same directory has to sort FIRST. Both are stated where they are written,
because getting it backwards fails silently in either direction.
resolv.conf is checked for actually pointing at resolved's stub before the
drop-in is trusted to do anything — a machine where something replaced the
symlink with a static file bypasses resolved entirely.
Resolution is tested afterwards rather than assumed. A resolver that does not
answer makes every later step fail for a reason that has nothing to do with it,
so that failure is reported and recorded rather than swallowed.
The choice names what each provider actually is, including that a resolver sees
every name the machine looks up.
Verified both paths against this host: keep reports unchanged, Quad9 renders the
right addresses, and the per-link display shows Hetzner and Tailscale correctly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
192293cdca |
offer to add an ssh key even when one already exists
The zip option was already gone — it never came across in the port, since the file is key material that cannot live in the repository and the script no longer sits next to it. Pasting a public key was already the first option. What was missing is the case where the account HAS a key: the section went straight to hardening, so there was no way to authorise a second machine, a rebuilt laptop or anyone else, and the original had no way to do it at all. The same menu is now offered either way. What differs is whether it can be declined without consequence: with no key, declining means the hardening below refuses too, and the run says so rather than quietly moving on. confirm() takes an optional default so this one can be [y/N]. Most questions in this script are "do the thing you already asked for" and Enter should mean yes; a genuine extra defaulting to yes is how people end up agreeing to things by reflex. A pasted key is trimmed before validation. Copying from a terminal or a password manager routinely brings leading or trailing whitespace, and ssh-keygen will not parse a key with it attached — which would have read as "that is not a valid key" for a key that is perfectly fine. Also corrected the reason unzip is in core utils, which still said it was there to open ssh-keys.zip. Verified: the add-another prompt appears and defaults to no, a whitespace-wrapped key is trimmed and accepted, and the already-hardened path is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9591f917f5 |
port ssh keys and hardening as one section, and make the hardening actually work
They were two sections, and being two is what let the second lock you out of a machine the first had failed to put a key on. Step 8 could warn-and-skip — no ssh-keys.zip, or an unrecognised menu choice, since its case had no default arm — and still mark itself done; step 9 then disabled password authentication and root login regardless. No key, no password, no root, on a box that may be in a datacentre. Nothing here turns off password authentication without first confirming a usable key is in place, and the refusal says why rather than skipping quietly. The hardening also did not do anything on a modern Ubuntu, and could not be seen not to: It sed'd /etc/ssh/sshd_config. Ubuntu includes /etc/ssh/sshd_config.d/*.conf from line 12 of that file, and sshd takes the FIRST value it obtains for a keyword rather than the last. Cloud images ship 50-cloud-init.conf containing `PasswordAuthentication yes`, read long before the line the sed edited. The run reported "SSH hardened" and password login stayed on. The settings now go in a drop-in named 01-machine-setup.conf, which is the only placement that wins under first-value-wins. It also sed'd ChallengeResponseAuthentication, renamed to KbdInteractiveAuthentication in OpenSSH 8.7. On 24.04 the old name is nowhere in the file, so that substitution matched nothing at all. State is read with `sshd -T`, which reports what sshd resolves across the main file and every drop-in — reading the config files tells you what is written, not what wins. Keys are counted by asking ssh-keygen to parse authorized_keys rather than by counting lines: comments, blanks and a half-finished paste all look like lines, and "there is a file" is not "there is a key that works". A pasted key is validated before it is stored, and matched on the key body rather than the whole line, so re-running does not authorise the same key four times over four runs. sshd -t validates the new config before anything is reloaded, and the drop-in is restored or removed if it does not parse — a config sshd refuses is a machine with no ssh after the next restart. Reload rather than restart, so the session this is running over is not the experiment, and the run says out loud to test a new connection before closing the current one. Generating a keypair now says the obvious thing the original did not: the private key is on the server, and a private key living on the machine it opens is a spare copy of the lock rather than a second factor. Verified against this host (1 key, already hardened, correctly does nothing) and with sshd_effective stubbed to a fresh-cloud-image state — the guard refuses and harden_sshd is never reached. Also verified key validation, dedup and 0700/0600 permissions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b1cc916258 |
detect root by uid, not by the name root
A provider whose image logs you in as "ubuntu" at uid 0 would have walked straight past the previous check, which compared the string. What makes an account root is uid 0; "root" is only the usual label for it. Two places now ask id -u rather than comparing names: the answer — an account at uid 0 is refused whatever it is called, and says which case it is rather than a bare "not root" the invoker — the warning about working as root fires when SUDO_USER is unset OR when SUDO_USER is itself uid 0. The second is the one that hides: sudo from a uid-0 account sets SUDO_USER to something that reads like an ordinary user and is not. The EUID check that requires the script to run as root was already uid-based and is unchanged. Verified by creating a real uid-0 account named ubuntu on this box: refused with the uid named, where the name check accepted it. That account has been removed — userdel refused it at first because it matches by uid and saw PID 1 running as uid 0, so -f was needed, and deliberately not -r, since its home was /root. Confirmed afterwards that root, /root, root's shadow entry and sudo are all intact and that root is once again the only uid-0 account. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d9d4085033 |
warn against working as root at the username prompt
root was already refused as an answer, but only the refusal said so — and only after somebody typed it. The advice now comes with the question, along with why: no safety net, a typo in a path that deletes instead of refusing, and nothing to distinguish you from a process that got out of hand. An extra warning when SUDO_USER is unset. That means the script was started as root rather than through sudo, which usually means root is how they log in — the exact situation the general advice is about, and the one where general advice is easiest to assume is aimed at somebody else. It says so plainly instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0eb3e3a85 |
create the user account first, and fix the sudoers filename bug that found
The user account section is now the first thing that acts, ahead of disk space.
The reason is a real defect, not tidiness. If the account does not exist yet,
USER_HOME is a path that is not there — and the ballast offers to put its file in
it, where ballast_create's `mkdir -p` runs as root and creates /home/<name> owned
by root:root. adduser afterwards finds the directory already present and does not
populate or chown it, so the account ends up with a home it cannot write to.
Making the account before any step can write into its home removes the ordering
entirely.
USER_HOME is re-read from getent after adduser runs. Until that point it is the
/home/<name> guess, because there is nothing to look up; adduser is free to have
used something else and every later step writes there.
Found while testing that, and worse than the thing it was testing:
local user="$1" dest="/etc/sudoers.d/99-${user}-nopasswd"
bash expands ${user} before the assignment to user has happened, so dest came out
as /etc/sudoers.d/99--nopasswd with the name missing. The rule inside was correct,
which is what made it invisible — visudo passes, sudo works, and the account
really does get passwordless sudo. What breaks is everything around it: every
account granted this way writes to that same file, so a second grant silently
overwrites the first and revokes it; and has_passwordless_sudo looks for
99-<user>-nopasswd, never finds it, and re-grants on every run forever.
Split into separate declarations, with the reason recorded where it happened, and
an empty username is now refused outright. Scanned the other lib files for the
same shape — the remaining multi-assignment locals only read positional
parameters, which is safe.
The stray /etc/sudoers.d/99--nopasswd this created on the dev box during testing
has been removed and visudo -c re-verified.
Verified with a real throwaway account: correctly reports not-granted before,
writes 99-msdemo-nopasswd as root:root 0440 with the right rule, reports granted
after, and leaves sudoers valid.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fbb90c917d |
default the username to whoever ran sudo, and look up their real home
Three changes to the question at the top of the run. USER_HOME is looked up rather than assumed. The original built "/home/$USERNAME", which is only the usual answer — an account created with a different home, or one whose home was moved, had every later step writing to a directory that was not theirs. getent passwd knows; the /home guess remains only as the fallback for an account that does not exist yet, where there is nothing to look up. The default is now whoever invoked sudo. On a re-run, or on a machine that is already somebody's, that is the answer every time, and retyping it is a chance to typo it into creating a second account. root invoking the script directly offers no default, since root is never the account being set up — and is refused if typed. The name is validated against the portable shape of a Linux account name before anything else happens. Letting adduser refuse it later means several questions have already been answered against a name that was never going to work. It also no longer goes through prompt_value, which obeys any environment variable matching the name it is filling in. USERNAME is set by some login environments, and a variable this script silently takes as an answer should not be one that might already be set for unrelated reasons. SETUP_USERNAME is the explicit override. The tmux config write moved out of the user section entirely. It was there only because the original copied it right after adduser, where $USER_HOME first exists. It is a dotfile and belongs with .zshrc and the starship config in the shell section. Verified: defaults to the sudo invoker, resolves daemon's home to /usr/sbin rather than /home/daemon, falls back to /home for an account that does not exist, and rejects a name with a space, a leading digit, one over 32 characters, and root. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3fb0e5c887 |
port the user account section, and stop clobbering files in the home
Two real defects fixed on the way across.
The sudoers write was in the wrong order. The original echoed the rule straight
into /etc/sudoers.d, validated it afterwards, and chmod'd it later still. A
malformed file there breaks sudo COMPLETELY — and you cannot sudo to repair it,
so on a remote machine that is a rescue console — and so does one with loose
permissions, because sudo refuses to read its own configuration. Both of those
windows were live in the original ordering. grant_passwordless_sudo now writes a
temp file, runs visudo -c against it, and only then places it with install(1),
which applies the content and the 0440 mode in one step. Nothing reaches
/etc/sudoers.d that has not already been validated.
The .tmux.conf copy overwrote whatever was in the home on every run. lib/files.sh
adds the two shapes that stop this whole class of thing:
install_config installs when absent, does nothing when identical, and keeps
what the user wrote when it differs — printing the cp to take
ours, so the choice stays theirs
append_once wraps a block in named markers so a second run recognises its
own work; also lets a human see which lines came from this
script and remove them as a unit
append_once is what the five unguarded `cat >>` into .zshrc need when those
sections are ported — a second pass currently duplicates the starship init, the
nvim PATH, bun, deno and the aliases.
Passwordless sudo is asked separately from creating the account, because it is a
security posture rather than part of making a user, and the cost is stated: a key
that can log into this account is root without a further step. Officer's actual
requirement is stated too — os-user-shell.ts runs `sudo -n`, and a prompt it
cannot answer surfaces as a permissions error rather than a question — and
refusing records that consequence in the summary instead of a bare "skipped".
Verified: all three install_config outcomes, append_once writing exactly once
across two runs, visudo rejecting junk before anything is installed, and the
section reporting correctly against this host's existing account.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
458510a0a8 |
scope sleep to homelab, and port the boot hang fix
Sleep and suspend is homelab only now. The two exclusions are for different reasons and both are stated in the run rather than left implicit: dev — a laptop should sleep; disabling it is a hot bag and a flat battery. vps — not merely unnecessary, harmful. A virtual machine has no lid and no power button, but the provider's Shut down control works by sending an ACPI power button event. HandlePowerKey=ignore makes the VM ignore it, so graceful shutdown requests silently do nothing and the instance is hard-killed instead. systemd defaults that key to poweroff for exactly this reason. The boot hang fix is everything except vps, where systemd-networkd genuinely manages the network and the unit is load-bearing. Rather than asking whether boot "feels slow" — a question people answer from memory of the worst time it happened — the step prints what the unit actually cost on this boot, from systemd's own accounting. On this host that is 14ms, which ends the discussion. On a NetworkManager desktop it is two minutes, which also ends it. On dev the wording says outright that a small number here means there is nothing to do. The original's live guard is kept and is what actually decides: NetworkManager active and networkd not. Anything else, including "cannot tell", is left alone, and the reason is printed. The warning against disabling systemd-networkd outright is carried across into the library, where the alternative would be attempted. Verified all three roles on this host, which is a networkd machine: homelab and dev both correctly refuse and explain, vps skips as not applicable, and the timing helper reads 14ms out of systemd-analyze. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c87127ff93 |
port the sleep and suspend section
The original ran unconditionally, so a laptop that went through it stopped suspending — a hot bag and a flat battery. It is a server concern: dev is skipped with the reason printed, like the ballast. Four things wrong with it beyond the role: HandleLidSwitchDocked was never set. A laptop used as a homelab server, docked and closed, still suspends — which is the exact machine this setting exists for. Added. RuntimeDirectorySize=10% was set alongside the sleep handlers. It is the size of /run, has nothing to do with sleeping, and 10% is systemd's own default, so the line never did anything. Dropped. systemd-logind was restarted on every pass whether or not anything changed, disturbing live sessions for nothing. The step now checks first and does not reach the restart when the machine is already configured. (The platform's own scripts/setup-old/setup.sh already had this guard; the machine script did not.) The settings were sed'd into logind.conf in place. They are a drop-in at /etc/systemd/logind.conf.d/99-machine-setup.conf now, so what this script set is one file that can be read or removed on its own. Current state is printed before anything is asked — whether the targets are masked, and what the lid, idle and power-key handlers actually do. logind_effective reads the main file and every drop-in and takes the last match, since a drop-in overrides logind.conf; reading only the main file reports a configured machine as unconfigured. The power-button consequence is stated rather than left to be discovered: after this, pressing power physically does nothing and a clean shutdown is `sudo poweroff`. WSL has no logind and cannot suspend, and says so. Verified both roles on this host, which the old script had already configured — correctly reports the targets masked and the handlers set, and correctly reports itself not fully configured because HandleLidSwitchDocked is missing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6ffd3534bd |
do not offer the ballast on a dev machine, and route its alerts through one place
Servers only now. On a machine you sit at, a filling disk announces itself — the editor refuses to save, the browser complains — and you are there to deal with it. The reserve is for the box nobody is watching, where the first sign is a service that stopped working hours ago. Skipped rather than asked, but said out loud with the reason and recorded in the summary. A section that silently produces no output is indistinguishable from one that failed. The cron this section installs was already there and is unchanged: /etc/cron.d runs the checker as root every ten minutes, and it deletes the ballast when free space falls under the threshold. What changed is where its message goes. Both alerts now run through one notify() inside the generated checker rather than calling logger directly, so there is a single place to add a second channel. Today it is still syslog only — the message lands in the journal and nowhere else, so nobody learns about it until they go looking, which is precisely the wrong moment. Push, mail or Officer's own notify sidecar hook in there. It also echoes to stderr now, so running the checker by hand shows the message instead of appearing to do nothing. Verified: dev reports not-applicable and asks nothing, vps still asks and records a refusal, and the regenerated checker parses and reports status. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d19dc5a92a |
ask where the ballast goes and how big it is
Three questions instead of one, because the two the original never asked are the
two that decide whether the thing is useful.
1. Do you want one, with the explanation first.
2. Where. Home (easiest to find again months from now), beside Officer, or a
path typed in. This is not tidiness: the checker measures its own directory,
so a ballast only protects the filesystem it sits on. Choosing where it goes
is choosing which mount is covered.
3. How much, as 5/10/20% — with the actual numbers, and with what would be LEFT
rather than only what is taken:
[1] 5% — reserves 2.9GB leaving 54.3GB free
[2] 10% — reserves 5.8GB leaving 51.4GB free
[3] 20% — reserves 11.5GB leaving 45.7GB free
A percentage on its own is unanswerable. The number that decides it is the
one on the right: the reserve has to be big enough to matter and small
enough not to be the thing that filled the disk.
The size is computed against the filesystem the chosen path lands on, after the
location is known, so the percentages are of the right disk. ballast_free_kb
walks up to a directory that exists, since nothing has created the target yet.
An existing ballast in either default location is found and left alone rather
than a second one being made beside it.
Verified end to end in a temp home: created at the chosen path, 2.9G for 5% of
57.1GB, checker installed and reporting it present.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d4b9b7b334 |
ask where Officer should be installed, in pre-flight
OFFICER_ROOT, defaulting to <user home>/officerdev. One directory holding the four things Officer is made of, per docs/sidecar-app-store.md — the app, its data, the item store, and any containers the app store provisions — so the whole installation can be moved, backed up or deleted as a unit. Asked at the start with the other questions rather than at the point it is first needed. It decides the shape of several later steps: where the repository is cloned, where DATA_PATH sits beside it, and which filesystem the app store's bind mounts come out of. Asking once up front also means the run can be described before it starts rather than discovered as it goes. A leading ~ is expanded explicitly. It arrives as a literal from a read or an environment variable — nothing expands it there — and would otherwise create a directory actually named "~" in whatever the working directory happened to be. Relative paths are refused with the value named, and a trailing slash is trimmed so the path composes cleanly with what gets appended to it. Nothing creates the directory yet; that belongs to officer-setup. This records the answer and reports it, including whether it already exists. Verified: Enter takes the default, ~ expands, trailing slash trims, OFFICER_ROOT in the environment skips the prompt, and a relative path fails with the value named. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6947284f1b |
check the root filesystem is actually using the whole drive
New first section, before anything else that changes the machine, because the
swapfile and the ballast both size themselves from free disk.
Ubuntu Server's installer on its defaults gives the root logical volume a fixed
size and leaves the rest of the drive as unallocated extents in the volume group.
On a 2TB disk that is a ~100G root with nothing to indicate a problem: lsblk
shows the whole drive, df shows 100G, and the two are never seen side by side
until the day it fills. Growing a virtual disk at a provider leaves the same
shape one layer down, and so does resizing a partition without telling the
filesystem inside it.
Three layers, any of which can be the short one, so all three are measured and
printed together:
drive: 76.3GB /dev/sda
volume: 76.1GB /dev/sda1
filesystem: 76.1GB ext4, mounted at /
Seeing them in one place is most of the value. The fix is then whichever layer is
short: lvextend for free extents, growpart for a partition that stops early
(followed by pvresize and lvextend when LVM is in the way), or resize2fs alone
when only the filesystem is behind.
Only ever grows. Nothing here shrinks, creates or deletes a partition, and ext4,
xfs and btrfs all grow while mounted — so no unmount, no reboot, and a failure
part-way leaves a smaller filesystem on a larger container, which is the state it
started in.
growpart is the authority on whether a partition can move — it exits 1 with
NOCHANGE when the partition already reaches the end — but it comes from
cloud-guest-utils, which is not on every image. Installing a package purely to
ask a question is too eager, so plain arithmetic on the device sizes decides
whether it is even worth looking, and only then is growpart fetched.
A gigabyte of slack before anything is reported: a filesystem is always slightly
smaller than its container, and reporting journal and reserved-block overhead as
reclaimable space would make this section cry wolf on every machine.
Verified on this host: plain ext4 partition filling its disk, correctly reports
nothing to reclaim, and every helper returns the right device and size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9b0e05bfd8 |
add three resource-pressure sections, each asked rather than assumed
Swap covers memory pressure. These are its neighbours: 8. Emergency disk ballast the same valve, for disk 9. earlyoom what happens when swap runs out too 10. inotify watch limit the silent one All three follow the rule this script now works to: the role sets which way the recommendation points, never whether the question is asked. A dev machine is still offered the ballast, with the recommendation pointing the other way; a server is still offered the inotify raise, because anything running `bun --watch` or serving a file browser is a watcher too. The ballast is section 22 of the original, moved up beside swap where it belongs and moved out of the user's home. The original wrote the checker into $USER_HOME/.local/bin and ran it from a root cron — a root cron executing a script in a directory its owner can write is a privilege escalation waiting to be noticed. Moot on a box where that user already has passwordless sudo, but wrong. Both the checker and the file are in root-owned system paths now. Two bugs found by running the generated checker rather than reading it: It df'd the ballast's own directory, which does not exist before the ballast is created — and with `set -euo pipefail` that meant cron mailing an error every ten minutes. It now walks up to a directory that exists, and the installer creates the directory itself rather than depending on the create step. The inotify text claimed a default of 8192. This host is at 29461: Ubuntu raised it, and stating a number the reader can see is wrong on their own screen undermines the rest of the explanation. It now describes the failure instead and prints the machine's actual value. earlyoom is a distro package and a systemd unit, so it is checked with `systemctl is-active` and reports honestly when it installs but fails to start. Verified: checker --status and its no-op path both exit 0 with no directory present, and the helpers report correctly against this host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
716a6e2750 |
port the swap section, and size it against the disk
Four things the original got wrong, all of which only show up on a machine that is not this one: It detected swap with `swapon --show | grep -q '/'` — a test for a swap FILE. A machine using zram or a swap partition reports no swap at all, and the step would add a swapfile beside working swap. Reads SwapTotal from /proc/meminfo now, which covers every kind. It never looked at free disk. On a VPS with 4G free and 16G of RAM it would fallocate 8G, fail, and take the run down under `set -e`. The recommendation is now capped by what is actually there, keeping 5G back, and refuses rather than shrinking to something useless. fallocate was assumed to work. It produces a file that btrfs and zfs will not swap on, so dd is the fallback — slow, but it always works. swappiness was written by sed'ing /etc/sysctl.conf in place, tangling it with whatever else lives there. It is a drop-in at /etc/sysctl.d now, so what this script set is visible as its own file. Role-dependent, which is the first use of MACHINE_ROLE: swappiness 10 on a server, where swapping is the emergency valve and a page fault on a request path is latency somebody is waiting for; the kernel default of 60 on dev, where swapping out an application nobody has touched in an hour is exactly what you want. WSL is left alone entirely — WSL2 runs its own managed swap inside the VM, and a swapfile written here is wasted disk the kernel will not use. Sizes are reported rounded rather than floored. A 4 GiB swapfile is 4194300 kB, which floors to 3 and reads as though a gigabyte went missing. Free disk stays floored, deliberately: it decides how much to allocate, and rounding up invents space. Verified on this host — 4G RAM, 4G existing swap correctly detected and left alone — and with the disk check stubbed at 6G free (caps to 1G) and 3G free (refuses). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5948f9673 |
refresh the package index in pre-flight, not inside a skippable step
The refresh lived inside "System update", which `step` skips when its name is already in the progress file. So a resumed run — the common case, since that is what the progress file is for — installed core utils, added the fastfetch PPA and set up the Docker repo against whatever the index happened to say hours or days earlier. On a box left overnight that is a stale index and a "package not found" somewhere unrelated. It now runs in pre-flight, unconditionally, before any step exists to skip it. `apt-get upgrade` stays where it was and stays confirmable: refreshing the index changes nothing on the machine, upgrading is the one thing that does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6480788979 |
port the timezone section
The original took whatever was typed and handed it straight to timedatectl. An unknown zone — a typo, a guess at the spelling — fails there, and under `set -e` that takes the whole run down four steps in. Names are now checked against /usr/share/zoneinfo before use, and a bad one just re-asks. It also never showed what the machine was already set to, and defaulted to option 1 (UTC) on Enter, so pressing return on a correctly-configured box silently moved it. Now the current zone is printed, Enter keeps it, and a zone equal to the current one reports nothing to do rather than setting it again. timezone_current reads three sources — timedatectl, /etc/timezone, then the /etc/localtime symlink — because they differ in availability rather than in answer: timedatectl needs systemd, /etc/timezone is Debian's, and the symlink is the one that is always there. timezone_set writes through timedatectl where there is a systemd to talk to and the files directly otherwise, which is what it would have written anyway; that is also the WSL path, where timedatectl exists but does nothing. Europe/Berlin added to the shortlist; TIMEZONE in the environment answers the prompt ahead of time and is validated the same way, failing early with the bad value named. Verified detection (UTC here), validation of four names, and the env-var rejection path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
163f8d5899 |
port the locale section
Was three unconditional lines that ran on every pass and reported success either way. Now it checks, says what it found, and asks. The original tracked one fact where there are two: what a new login shell is told to use LANG in /etc/default/locale whether that locale actually exists whether it has been generated Setting the first without the second is what produces "setlocale: LC_ALL: cannot change locale" on every ssh login and every perl invocation. They fail differently, so the step names whichever one is actually missing rather than reporting a flat "locale not set". Also fixes two things the original would have hit on a minimal image: locale-gen comes from the `locales` package, which cloud base images do not ship and which is not in core utils. It is installed on demand rather than assumed, instead of failing with "locale-gen: command not found". The locale is uncommented in /etc/locale.gen rather than only passed to locale-gen as an argument. A locale generated by argument alone disappears the next time anything regenerates from that file. `locale -a` prints en_US.utf8 where the configuration spells it en_US.UTF-8, so both sides are folded before comparing — a literal match reports a working locale as missing. LOCALE in the environment overrides the default. pacman, dnf and brew branches are written but unreachable while the pre-flight gate is apt-only; macOS has no system locale to set and says so. Verified both paths on this host: en_US.UTF-8 reports already set and generated, pt_PT.UTF-8 correctly reports both facts missing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1beb357f2e |
drop the ominous wording from the system update prompt
"the only step that changes software already on this machine" is true, and reads like a warning about something dangerous rather than a description of apt upgrade. The reasoning stays in the section comment, where it explains why this is its own step; the prompt just says what it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
454faf5406 |
list what the system update would actually upgrade
The section asked "Proceed?" without saying what it was proposing to change — the one question in the script where the answer matters most, since it is the only step that moves versions of software already on the machine. pkg_upgradable now names them, from `apt-get upgrade -s`: the same calculation the real run does, as opposed to `apt list --upgradable`, which also lists packages held back that would not actually move. Nothing to upgrade means no prompt at all, and the summary says so rather than claiming an upgrade happened. The list is capped at 25 with a count of the rest, because a box untouched for months lists hundreds and a wall of names is no more informative than the number. Verified against this host (0 upgradable, so it reports current and does not ask) and with a stubbed 40-package list for the cap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f896d4882f |
one section per concern, each announced and confirmed before it acts
Three sections where there were two, and none of them touches the machine until you say so: 2. System update upgrades what is already installed 3. Core utils what the distribution provides 4. Command-line tools lazydocker, lazygit, starship, fastfetch The split matters because these are different kinds of change and deserve separate answers. System update is the only step in the whole script that moves versions of software already on the machine; core utils only ever adds what is absent; and the four tools are upstream binaries the distribution does not ship at all. Previously the update and the core packages were one step and the tools were tacked onto the end of it, so agreeing to "essentials" meant agreeing to all three at once. Every section now prints what it will install and what it is leaving alone, then asks. Enter means yes — unlike the machine-role question, which has no default, because these are "do the thing you already asked for" and making twenty of them require a deliberate keystroke would train people to hold the y key down. ASSUME_YES=1 answers all of them for an unattended run, and EOF fails with that named rather than spinning. Refusing is recorded rather than glossed: LAST_SKIPPED feeds the summary, so a declined section reads "Core utils: SKIPPED by request — cowsay neofetch" instead of quietly reporting nothing installed. Nothing to install means no prompt at all — there is nothing to agree to. announce_plan takes the array NAMES rather than their contents, because once a list has been through word splitting an empty one cannot be told from a missing one. Verified all three paths: accept, refuse, and nothing-to-do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5ef9f7662 |
report what a section actually installed, not what it was asked for
The summary claimed credit for everything in a section's list, including the packages it had just decided to leave alone — so a run that installed nothing still ended with "Command-line tools: lazydocker lazygit starship fastfetch". The announce above it said "nothing, all present" in the same breath. pkg_install and tools_install now record LAST_INSTALLED and LAST_KEPT, and summarise_last turns those into one honest line: Core packages installed: btop tmux (17 already present) Command-line tools: already present, nothing installed Verified all three shapes — everything present, nothing present, and mixed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dd655577e3 |
make the machine-role question require an answer
No default, and it is the only question in the script like that. A guessed default is right often enough to be trusted and wrong in exactly the case that costs the most — pinning a static IP on a rented box, or leaving the firewall open on one. Every branch downstream is about what this machine is exposed to, so it is worth one deliberate keystroke rather than an Enter. Empty and unrecognised answers re-ask rather than aborting; a failed read means EOF rather than a wrong answer, and fails with the environment variable named, because otherwise the loop spins forever the first time this runs unattended. Drops guess_machine_role, which existed only to supply that default. default_iface stays — the static IP section needs it when it is ported. MACHINE_ROLE in the environment still answers it ahead of time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cb243eed9 |
port into a clean script instead of editing the original in place
Your call, and the right one. Editing in place let mis-grouped code sit unnoticed until it scrolled past in a live run — which is exactly how the four upstream binaries buried in "System Update & Essentials" were found. Porting forces the question of where each thing belongs before it runs, not after. machine-setup.sh now contains only what has actually been worked through: pre-flight, system update and core packages, command-line tools, and the summary. 1149 lines down to 172. The sections still to come are listed in a NOT PORTED YET block, in order, and each arrives as its own commit. The original is beside the other superseded scripts as scripts/setup-old/setup-ubuntu.sh — verified byte-identical to the live /root/ubuntu-setup copy — so porting reads from a file in the repo rather than from root's home. Two claims trimmed from the ported summary, because they were true of the old script and not of this one yet: it reported the shell as "zsh (Oh My Zsh + Starship)" unconditionally, and told you to reconnect as a user it had not created. Replaced with what pre-flight actually knows — system, role, user. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7622239949 |
split the upstream binaries out of the package section
lazydocker, lazygit, starship and fastfetch were buried inside "System Update &
Essentials", after the package install and with no announcement — so a run
appeared to be installing system packages and then started pulling tarballs and
printing a five-shell starship tutorial. They are a different thing: upstream
binaries on their own release cadence, not anything the distribution ships. Now
their own step, announced in the same shape as the package section.
Each is checked before it is fetched. The original re-ran every installer on
every run, which is why a machine that already had starship got it reinstalled
along with its "add this to your ~/.zshrc" instructions — advice this script
does not want followed, since it writes the shell config itself. Its output is
now dropped; errors still surface.
Two real bugs fixed on the way:
lazygit's asset name was hardcoded to x86_64, so on arm64 the download 404s
and tar fails partway through the run. It now maps ARCH, and spells the
architectures the way lazygit does rather than the way we do.
The version was extracted with `tr -d 'v'`, which deletes every v in the
string rather than the leading one. `${version#v}` instead.
fastfetch stays a package but stops assuming the PPA is needed: Ubuntu picked
it up in 24.10, so the repository is now checked first and the PPA added only
where the archive has nothing. Verified on this host — noble genuinely has no
candidate, so the PPA is still the only source here.
Verified both branches of tools_install by stubbing the presence check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bba6d854bc |
stop the run at the end of the rewritten sections
A WIP boundary after section 2 so the finished part can be run start to finish on its own, without the untouched sections below acting on the machine. It moves down as each section is worked through and goes away when the walk ends. Also ignores .setup-progress, which the script writes beside itself and is per-machine. The exit message names it, because with it in place a second run skips section 2 and the rewritten part cannot be re-felt from scratch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dbdef23d29 |
install what is missing and keep what is there, per package manager
lib/packages.sh, and section 2 wired to it.
The rule it exists to enforce: `apt-get install <present-package>` is not a
no-op, it upgrades the package if the repository has a newer one. On a machine
somebody already uses that silently moves a version they chose, and a setup
script is the last thing that should do that behind their back. pkg_install
queries the package database first and names only the genuinely absent packages
on the command line — a package already installed is never passed to apt at all.
It also says so out loud, every time, because a provisioning run should not be
opaque about what it is doing to the machine:
:: Core packages — installs what is missing, keeps what you already have
already here: curl ca-certificates gnupg git jq …
to install: btop tmux
Section 2's flat list of 19 is now pkgs_core(), split per package manager rather
than through a canonical-name table with overrides. The names genuinely disagree
(build-essential/base-devel, fd-find/fd) and three of them are not packages
elsewhere at all — apt-transport-https, lsb-release and software-properties-common
are apt concepts that exist to let later steps add the Docker repo and the
fastfetch PPA. A `case $PM` shows what each system actually gets, in one place.
Of those 19, six are load-bearing and the rest are the environment. Only
build-essential reaches beyond itself: it is a meta-package, so on a box with a
pinned gcc it pulls the distribution default alongside. Noted where it is
declared; it is the first thing to move out of core if that ever bites.
apt-get upgrade stays, but as its own announced step — it is the one place that
deliberately moves versions, rather than something that happens as a side effect
of asking for a tool.
DEBIAN_FRONTEND=noninteractive and NEEDRESTART_MODE=a now live inside the
helpers. needrestart has been on by default since Ubuntu 22.04 and stops to ask
which services to restart, which is how an unattended run ends up silently
waiting for a keypress.
dpkg-query on the status field rather than `dpkg -s`, which also succeeds for a
package removed but leaving its config behind — that state would read as present
and never be reinstalled.
Verified against this host's real dpkg database: all 19 report present, and a
mixed list correctly passes only the absent ones through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
41ff8030e9 |
ask what the machine is for, once, in pre-flight
MACHINE_ROLE is homelab, vps or dev, and several steps have a different right answer per role with no way to work it out themselves: whether the address is yours to pin (static IP), whether the box faces the open internet (fail2ban, SSH hardening, UFW), and whether it is allowed to sleep (suspend, logind). Asked in pre-flight rather than at each point of use. The steps that care run from swap through to the firewall, and being asked "is this a VPS?" for the fourth time halfway down a provisioning run is how people start answering without reading. The default offered is guessed from whether this machine's own address is in RFC1918 space, which beats asking whether it is virtualised — a homelab is very often a VM on Proxmox and would be misread as rented — and is the same fact most of the branches turn on anyway. A graphical session means dev; so does macOS. It is only ever a suggestion the user confirms. MACHINE_ROLE in the environment answers it ahead of time for an unattended run, which is why it is declared with :- rather than a plain assignment. The first version wiped the caller's value before ask_machine_role ever saw it; caught by running with MACHINE_ROLE=vps and watching the menu appear anyway. Verified: guesses vps on this host (public IPv4, no DISPLAY, no display manager), env override takes, and a bad value fails with the three valid ones named. Nothing consumes the role yet — the steps get wired as each is worked through. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
34bb8fc22a |
put the superseded setup scripts in setup-old, and repair what the move broke
scripts/setup/ is now what the new installer is being built in — machine-setup/ for the box, officer-setup.sh for the platform on top — and everything being replaced moved to scripts/setup-old/. It still works and is still what to run. Three things the move broke, and what each needed: starship.toml is not an old-setup artifact. os-user-shell.ts reads it at RUNTIME to seed a member's ~/.config/starship.toml when their Linux account is provisioned, and line 125 reads it inside a try whose catch returns "could not read the shell templates" — so account provisioning would have failed outright, not degraded. Moved back to scripts/setup/, which is where it belongs anyway (one file, both audiences) and which leaves the code correct with no edit. package.json's `setup` script pointed at a path that no longer exists. It now points at officer-setup.sh, where the installer is going, rather than at setup-old/ which is temporary. officer-setup.sh was created empty. An empty script exits 0, so `bun setup` would have reported success while doing nothing — worse than the broken path it replaced. It now explains that it is not written yet and exits 1, naming the setup-old script to run meanwhile. Also brought .tmux.conf and ufw-docker-rules.conf in beside machine-setup.sh, which reads both from SCRIPT_DIR and had been silently skipping them since the script was vendored. ssh-keys.zip deliberately stays out: it is key material, and *.zip is ignored. Comments in os-user-claude.ts, app-store/preflight.ts and two docs still name the old scripts/setup/setup.sh path. Left alone on purpose — repointing them at setup-old/ only to repoint them again when officer-setup.sh lands is churn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44141faf0a |
split machine-setup into an entry point and a base library
Structure before the work rather than during it: scripts/machine-setup.sh becomes
scripts/setup/machine-setup/, with the script itself as the entry point and
lib/base.sh holding what every part of it needs.
machine-setup.sh pre-flight and the numbered sections, for now
lib/base.sh shared state, output, the step/resume machine, prompts,
and OS detection
The rule for lib/ is definitions only — nothing there installs, writes or
restarts anything, so sourcing it is safe from anywhere. That is why the ERR
trap stayed in the entry point: a trap is a side effect on whoever sources it.
Behaviour is unchanged. Verified by diffing the moved region against the previous
commit: identical set of functions, and the only differences are added comments,
section banners, fail() reformatted onto three lines, and one new line — a guard
against double-sourcing, which matters because steps will source this directly
once they move out, and a second pass would reset SUMMARY.
The sections are still one 1111-line block below pre-flight; they move into
steps/ as each is worked through. The script also still reads ssh-keys.zip,
.tmux.conf and ufw-docker-rules.conf from SCRIPT_DIR, which is now this
directory, so those three steps warn and skip until the files follow it here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
dda214ffb0 |
detect the operating system before any step runs
The script assumed Ubuntu on x86_64 in every line of it. detect_os() now runs first and fills in OS, OS_NAME, OS_VERSION, PM, ARCH and IS_WSL, so the steps have something to branch on as support for other systems is added. Read from /etc/os-release rather than probing for a binary: a machine can have more than one package manager on PATH, and only os-release can say which distribution this actually is or give a version worth printing. Sourced in a subshell so its NAME, VERSION and ID do not leak in here. ID_LIKE is the fallback, so Pop!_OS, Mint and EndeavourOS resolve without being named. ARCH is normalised to amd64/arm64 in one place because upstream disagrees — Neovim ships aarch64, Go and Docker ship arm64, lazygit ships x86_64 — and several steps hardcode one spelling today. Windows exits with a message pointing at WSL2. WSL itself is detected and warned about rather than refused: it reports as Linux but has no real systemd session, so the suspend, logind and boot-hang steps do nothing there. Everything below pre-flight is still apt and systemd only, so a gate refuses pacman/dnf/brew by name rather than half-building a machine and stopping somewhere unhelpful. Relax that case one entry at a time as each grows a path. Verified on this host (Ubuntu 24.04.4, amd64, apt) and by stubbing uname and os_release for arch, manjaro/arm64, fedora, pop, macos, mingw and riscv64. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |