d85f089817f3e427b8df9f698e69144ad8ff795e
198
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
11710f283a |
wire the reverse proxy in as officer-setup 12
~/npm-setup-draft/setup-npm.sh, adapted to the script's own helpers and placed last
— it is the only step that needs Officer already running.
It ignores --unattended, as asked. Every other question in this script has a
defensible default; a domain name, a DNS provider and that provider's API
credentials do not, and the step is opt-in besides. Its prompts read stdin directly
instead of going through confirm()/ask_required(), and they are NAMED APART
(proxy_confirm, proxy_ask) so nobody later consolidates them into the shared helpers
and quietly makes --unattended agree to publishing a public hostname.
The valve is a TTY check rather than the flag: with no terminal there is nobody to
ask, so it skips and prints the manual instructions. A cron-driven install still
works.
Five fixes to the draft:
- `${OFFICER_REPO}/scripts/store-npm-credential.ts` — OFFICER_REPO is a git URL,
not a directory, so that path was https://…/platform.git/scripts/… and the -f
test could never pass. The whole persist-to-platform branch was dead code
falling through to the print. Dropped it: the comment beside it already argued
that not storing this password is a legitimate outcome, since only a human
logging into the admin UI needs it.
- NOT re-runnable, despite saying so. claim_admin returned early on an already
claimed instance without setting NPM_EMAIL/NPM_PASSWORD, and get_token
dereferenced both under set -u. Second run died on an unbound variable. It now
asks for the existing credentials.
- $HOME/dockers → $OFFICER_ROOT/dockers, matching data-path.ts. And the network
is SETUP_DOCKER_NETWORK (`services`), not a second bridge called `officerdev`.
- dig → getent hosts. dnsutils is not installed by this platform, so the check was
command-not-found on a fresh VPS — and an empty answer is indistinguishable from
"not resolving yet", so it waited the full 30 minutes before failing.
- python3 → jq for host-side JSON. jq is already in the core package list; the one
remaining python3 runs INSIDE the NPM container to read its own credential
template, which is the point of reading it from there.
Failure is contained: every function warns and returns non-zero rather than exiting,
so a proxy that does not come up leaves a finished Officer install behind. Retry
with `--only Proxy`.
Verified: bash -n, shellcheck -S warning clean, --list shows Proxy, all seven
external commands present, and the jq filters checked against sample payloads
including the multi-line DNS credential.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
1d95ad3d1b |
install the rootless docker prerequisites with docker itself
Reported from a member's daemon failing: "rootless Docker needs these packages on the host: uidmap". They were being installed — but only inside branch [2] "rootless Docker for <owner>" in section 22. The owner's choice is not the only one that matters: every Developer account the platform provisions gets its own rootless daemon whatever the owner picked for themselves. So on a machine where the owner chose the docker group, the host never got them and every member's daemon failed. Moved into install_docker_engine, so they arrive with Docker rather than with one particular answer to a question about the owner. Three packages, not the one in the error. checkDockerPrerequisites in os-user-docker.ts is the authority and wants uidmap (newuidmap, newgidmap) AND docker-ce-rootless-extras (dockerd-rootless-setuptool.sh); dbus-user-session is what keeps a member's systemd --user alive without a login session. rootless-extras is only RECOMMENDED by docker-ce — installed by default, so usually there by luck, and absent on any host configured with --no-install-recommends. Named explicitly. Reproduced on this machine while checking: rootless-extras present via Recommends, uidmap absent, newuidmap and newgidmap missing. Exactly the reported failure, on a box that chose the docker group. The rootless branch still installs uidmap and dbus-user-session behind its pkg_is_installed guard. Redundant now, kept deliberately: it is the only thing that fixes a machine whose Docker was installed by an older run of this script. Verified: bash -n on both files, and all three packages present in the noble archive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3862558b92 |
real aliases for the owner, and eza to go with them
The owner's `aliases` block was one line — `alias sz`. Members got a full set from shell-skel/zshrc and the owner got that. Replaced with the eza ls family, the oh-my-zsh standards, and n/vim/sz/ld/httpserver. eza added to all four core package lists. It is in the noble archive at 0.18.2-1, so this is a package rather than a binary fetch, and Core utils is section 5 — well before Shell at 26, so `command -v eza` is already true when the block is written. The eza aliases are GUARDED behind `command -v eza` and the rest are not, and the asymmetry is deliberate: these replace `ls`. Unguarded, a machine where eza failed to install has no working `ls` in any new shell, which reads as a broken machine rather than a missing package. `alias ld=lazydocker` without lazydocker is one command-not-found when you type it — that can degrade honestly. Same principle shell-skel/zshrc already holds to. python3, not python, for httpserver: Ubuntu ships no `python` binary at all, so as given it would have been a command-not-found on every machine this targets. Checked the editor block first — it only exports EDITOR/VISUAL/SUDO_EDITOR, so n and vim do not collide with anything already appended. KNOWN: append_once returns 1 when its marker is already present, so a machine that has already run this keeps the old one-line block and gets none of the above. That is the function working as designed — it exists so a second run does not duplicate its work, and it cannot tell a stale block from one the owner edited. Fix by hand: delete the `# >>> machine-setup: aliases >>>` block from ~/.zshrc and re-run `machine-setup.sh --only Shell`. Verified: bash -n, zsh -n on the block, the eza guard leaving ls unset when eza is absent, and vim resolving through n to nvim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
977e30e782 |
point the setup default at public https
was ssh://git@gitea.pastilhas.dev:2222/officerdev/platform.git now https://gitea.officer.dev/officerdev/platform.git Bigger than a URL swap. The SSH default could not clone on a genuinely fresh machine: the key machine-setup generates there is brand new and Gitea has never seen it, so `--repo` was effectively mandatory on a first install — which is the problem that flag was added for two hours ago. HTTPS needs no key and no agent, so the default now works on a blank box. The old comment explained SSH-because-private and set the condition for changing it: "back to HTTPS when the repository is public". It now is — verified with an anonymous `git ls-remote`, which lists refs with no credentials. Rewrote the comment to record why it moved and what to do if it ever goes private again, since that reasoning is the part worth keeping. clone_repo already runs GIT_TERMINAL_PROMPT=0, so a private repo would fail fast rather than hang on a username prompt. No change needed there. repo.sh is still the only place that sets this, and --repo / OFFICER_REPO still override it. Verified both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
edbe446b34 |
revert "allow port 22 through the docker-user allowlist"
this reverts
|
||
|
|
e36c6bb431 |
allow port 22 through the docker-user allowlist
the DOCKER-USER chain is the only thing gating docker-published ports from the internet — docker writes its own DNAT/FORWARD rules and bypasses ufw, so `ufw allow <port>` has no effect on a published container port. the allowlist permitted only 80 and 443, so a machine provisioned from this template dropped gitea ssh silently. the failure is hard to spot: the port looks open locally and docker ps shows it published, but external clients hang at TCP connect with no refusal. local tests pass because they arrive via lo and match the loopback RETURN before reaching the DROP. comments added so the next person recognises it faster. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d14e11f6c |
--unattended: every question that has a default answers itself
51 yes/no prompts and ~20 free-text ones, of which about six actually need a human.
The line drawn is "a question with a default answers itself; a question with no
possible default still asks", so it stays attended without being a conversation.
Half of it already existed: ASSUME_YES=1 was implemented and honoured by confirm()
in both scripts, returning each question's OWN default — so a "do the thing you
asked for" question goes yes and a genuine extra goes no. --unattended sets it.
The new part is menu_answer(), for the eight numbered menus. It sets the variable
EMPTY rather than passing a default in, because every menu already consumes its
choice as `${CHOICE:-<n>}` — the default lives next to the options it selects
between, which is the right place, and a second copy in the helper could drift from
the one the prompt advertises. Verified all eight consume that way before touching
them. `read <<<''` rather than eval or `declare -g`, which is bash 4.2+ and rules
out the bash 3.2 macOS still ships.
officer-setup's ask_required takes its default too, except where there is none — the
owning account on a machine machine-setup never ran on, where a guess would install
as the wrong user.
STILL ASKS, deliberately: the username; the Tailscale control plane, login server
and auth key; the git identity; and an SSH public key when the account has none.
That last one is a trap I nearly walked into — on a fresh VPS KEY_COUNT==0 forces
ADD_KEY=true with no confirm, and the menu's default is "[1] paste a public key",
which then prompts with no default at all. Auto-answering that menu would hang or
fail, so it is excluded by name. adduser also still asks for a password; that is
the tool, not us.
Two pre-existing bugs fixed on the way: machine-setup's sudo re-exec passed "$@"
after `shift` had emptied it, so --only and --reask stopped existing the moment it
escalated — same bug as officer-setup had. And UNATTENDED/ASSUME_YES are named in
all three sudo lists, because env_reset would otherwise drop the flag at
escalation, which is now the fourth variable lost that way.
Verified: bash -n on five files, --help on all three, and menu_answer + confirm
under the flag showing a menu resolving to its default and a no-default confirm
correctly answering no.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4c33ef7206 |
fix the silent death after installing zsh
Reported from a fresh Hetzner VPS: the run stopped dead right after apt finished installing zsh, printing nothing at all — just install.sh's "machine setup did not finish". install_oh_my_zsh carried a comment saying it "Returns 0 whatever happens". It did not. Under `set -e` a failing command inside a function aborts the SHELL at that line when the function is called plainly; `return 0` underneath is never reached. The command is also `>/dev/null 2>&1`, so the cause was invisible — which is why the transcript just ends. `|| true` is what actually makes it non-fatal. The file already uses that idiom correctly in four other places, so this was a slip rather than a misunderstanding. set_login_shell had the identical bug on `chsh`, which the same run would have hit on the very next question. Fixed differently and deliberately: `|| true` there would let the caller announce a login shell that was never set, so it returns chsh's real status and the CALLER guards the call — which is also what keeps set -e out of it. A refusal now reports, names the manual chsh command, and carries on, because a machine with zsh installed and bash at login still works. Does not explain WHY oh-my-zsh failed on that host — the output was discarded. It will now say "oh-my-zsh did not install" and continue, which is enough to see it. Verified: bash -n on both files, and a reduced case proving broken() exits 1 while fixed() survives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6c13e0d8f6 |
tell the operator they are still root, once
Neither script ever becomes the user it sets the machine up for — a process cannot change its own uid, so both run as root and drop privileges per command instead. Everything Officer owns ends up belonging to that user and every pm2 process runs as them, but the session you are left holding is root's. Two things that fixes are invisible until they bite: group membership is fixed at LOGIN, so the `docker` group just granted is not in the current session, and the shell configuration was written into their home and is not loaded in root's. Both present as "the machine is broken" rather than "log in again". Printed by whichever half runs LAST. The first attempt put it at the end of both, which says it twice on a full install — and the first time it is wrong, because officer-setup is about to run and still needs the root session it tells you to leave. install.sh is the only thing that knows whether anything follows, so it sets OFFICER_SETUP_FOLLOWS and machine-setup stays quiet. Also drops "Pre-flight complete. The remaining sections are not built yet." from the end of officer-setup. All 11 sections exist; that line last made sense when 6 did. Verified: bash -n on all three, the set -e behaviour of `$RUN_OFFICER && export` under --machine-only, and the suppression across all five ways in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cbfe376a42 |
accept the repo URL as an argument
bun setup -- --repo https://github.com/you/platform.git The default is a private Gitea over SSH, which only authenticates on a machine whose key it already knows — so a genuinely fresh server could not clone at all without editing lib/repo.sh or knowing OFFICER_REPO existed. Added to both entry points. install.sh exports it rather than forwarding an argument it does not own; officer-setup.sh sets it before lib/repo.sh is sourced, which reads `${OFFICER_REPO:-<default>}`, so an absent flag still defaults. Two bugs found doing it, both pre-existing: - officer-setup.sh ALREADY had an arg parser, at the top, before the sources. My first attempt added a second one further down that was unreachable — every argument had already been consumed and `*)` would have exited 2 on --repo. Caught because `--help` printed the wrong usage. - both scripts re-execute through sudo passing `"$@"`, which the parse loop had already emptied with `shift`. So `officer-setup.sh --only build` run as a normal user silently became a FULL run the moment it escalated, and `install.sh --officer-only` re-ran the machine half. Nothing said so; the flag just stopped existing. ORIGINAL_ARGS is captured before the loop now. `${ORIGINAL_ARGS[@]+"${ORIGINAL_ARGS[@]}"}` is the set -u safe form — expanding an empty array is an error on bash before 4.4, and this runs on whatever the machine came with. Verified: bash -n on both, --help/--list/--repo/--repo=/unknown-option on both, the set -e behaviour of the guarded export, and that args survive the shift loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
081920c61f |
drop fastfetch from machine-setup
It stopped a real install. The guard covered the wrong half: a PPA that fails to ADD is caught and skipped, but one that adds cleanly while carrying no package for the running codename gets past that and dies on `pkg_install_now fastfetch`. It was also the only tool in the set with no source but a third-party PPA on Ubuntu 24.04 and older. A neofetch clone is not worth a branch in a script whose whole job is to survive machines nobody has seen. Removed from tools_default, the tool_command mapping and its installer. No shell config invoked it, so nothing is left calling a missing binary. software-properties-common stays in the core apt list for now, with a note: it provides add-apt-repository, the fastfetch PPA was its only caller, and Docker writes its own sources.list.d entry by hand — so it is now dead weight. Left as a separate decision rather than folded into this one. Verified: bash -n on all four scripts, and no live reference remains. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
77f1284925 |
install report: first cut, generated by the helpers
Every run writes a timestamped install-report.md recording what was installed, changed, kept, skipped, started and run as root. Written for an adversarial read: the person who just ran a setup script off the internet hands it to an agent of their choosing and asks whether it did anything it should not have. Recorded by the HELPERS rather than by the sections. pkg_install and install_config report themselves, so anything installed or written through them appears whether or not a section author remembered — a section that has to remember is a section that will forget, and an incomplete report is worse than none because it reads as a full account. "Kept" is recorded as carefully as "changed". Leaving somebody's .zshrc alone is the claim a reviewer most wants substantiated, and it is invisible unless stated. Secrets are redacted at the moment of recording rather than filtered at render, so a credential never sits in memory formatted for printing. Verified against a POSTGRES_URL and an api_key/password pair. REPORT_FILE is passed through the sudo re-exec. It was not, first time, and the report silently vanished — the third variable this evening lost to env_reset. Unfinished on purpose, paused mid-task at the owner's request: machine-setup's 26 sections still only report through the two shared helpers, so the sections that change system state directly — systemd units, netplan, ufw, sshd drop-ins — are not yet recorded. That is the half a reviewer would care most about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1cfc0042f |
rename officer-agent to officer-claude-code
The old name said nothing about what the process runs, and it sits directly beside officer-anthropic-proxy — a different process doing a different job — so "the agent" was ambiguous exactly where it mattered. CLAUDE.md already had to spend a paragraph insisting the two are not the same thing. It spawns `claude`; the name says so now. Only two references were functional: the generator's CORE_PROCESSES and the CORE list in catalogue.test.ts. Everything else was prose or comments. Left alone deliberately: `x-officer-agent-token`. It looks like the same string and is not — it is the agent-handoff HTTP header, naming a per-panel bearer token, unrelated to any pm2 process. Renaming it would have changed a wire protocol to tidy a label. Historical docs keep the old name. claude-sidecar-isolation.md and open-threads-after-per-user-claude.md are dated investigations that record the PREVIOUS rename, from officer-claude to officer-agent, and rewriting them would make that history unreadable. CLAUDE.md notes the change instead, where somebody reading those will be looking. Also worth recording, from the owner: merging this with officer-anthropic-proxy into one sidecar was investigated tonight and rejected. They stay separate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
88c9e96895 |
recover earlier answers on resume
A skipped section leaves its variables unset and later sections read them, so on a resume — which skips every section before the one that stopped — Build announced "PUBLIC_URL <not set — run the Environment section first>" on a machine whose .env had been written twenty minutes earlier. Three variables cross a section boundary: ENV_PORT and ENV_PUBLIC_URL from Environment, POSTGRES_URL from Database. They are read back once near the top, from the file that already holds the answers, rather than per-section — the next variable to cross would otherwise have to remember to do it again. Only fills what is empty, so a value passed on the command line still wins and a section that actually runs still overwrites it. Build had its own late read-back that made the generation work while the screen said it would not. Removed, now that the value is there before anything prints. Verified against the real install at /home/pastilhas/officerdev-test: --only Build now reports the URL that run chose. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e5fe966308 |
count the schema tables instead of printing "?"
The Schema section read ${SCHEMA_TABLES:-?} and nothing ever assigned it, so it
announced "? tables" — which reads as "the count could not be determined" rather
than "nobody set this". schema_table_count existed in lib/build.sh and was never
called.
Counted from schema.ts rather than hardcoded, so the number stays true when a
plugin line is uncommented. Reports 21 against the current tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
60ab531294 |
use the short tailnet name, not the FQDN
http://officer-dev:9000 rather than http://officer-dev.ts.pastilhas.dev:9000. All three forms resolve inside the tailnet — short name, FQDN, raw 100.x — and the short one is what anybody actually types. PUBLIC_URL is read by people too: gen:index bakes it into the page's OpenGraph tags. It depends on the tailnet's search domain, which every Tailscale client sets when MagicDNS is on. A device that has lost it resolves the FQDN instead, and the answer there is to type the longer one rather than to default everybody to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f33e7474b0 |
default PUBLIC_URL to the tailnet address, not localhost
localhost is wrong on a machine with a tailnet, and quietly so: it works from the machine itself and nowhere else, so the mistake surfaces on the first phone rather than during setup. And PUBLIC_URL is not decoration — gen:index bakes it into the page's OpenGraph tags, the task API hands it to scripts as OFFICER_API_HOST, and the CalDAV profile builder refuses without it. The tailnet is where Officer is actually reached, and it is the perimeter the whole security model rests on now that origin checking is gone. Its address is the honest default. Prefers the MagicDNS name over the raw 100.x address — both work, but the name survives a node being re-registered and is something a person can type. Falls back to localhost with no tailnet, which is right rather than merely tolerable: a machine with no private network has no better address to guess. Suggests http://officer-dev.ts.pastilhas.dev:9000 on this machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bedc420d4d |
clone over SSH while the repository is private
An HTTPS clone of a private repo prompts for a username, and under sudo with no interactive terminal that hangs or dies with "could not read Username" — which is what gitea.pastilhas.dev does right now. ssh://git@gitea.pastilhas.dev:2222/officerdev/platform.git instead, temporarily. Back to HTTPS when it is public; nothing else in the script cares which. Tested the path the script actually takes, not just the URL: the clone runs as the OWNER rather than root, and sudo drops SSH_AUTH_SOCK, so there is no agent to answer a passphrase. `sudo -u pastilhas env -u SSH_AUTH_SOCK git ls-remote` returns HEAD, so the key works unaided on this machine. A passphrase-protected key that relies on an agent would not. Still overridable with OFFICER_REPO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9b56e04b5e |
point the clone at gitea.pastilhas.dev
gitea.officer.dev is not serving yet. Overridable with OFFICER_REPO, as it always was, so a fork or a mirror needs no edit here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ea87ffed04 |
ask for privileges, do not demand them
Run any of the three as yourself. On Linux they now ask through sudo and re-execute, rather than refusing until you type it. Typing `sudo` still works and changes nothing — it just stops being the price of starting. This also fixes a trap that had nothing to do with taste. `sudo` strips the environment by default (`env_reset`), so `OFFICER_ROOT=/somewhere sudo ./install.sh` silently loses the variable and installs to the default path instead. The re-exec passes OFFICER_ROOT, SETUP_USERNAME and MACHINE_ROLE to sudo BY NAME rather than relying on -E, which env_reset ignores. This project has already lost a variable to that once — see the DATA_PATH commit. The invoking account is recovered the way it always was: sudo sets SUDO_USER, which lib/base.sh already reads, including the check for a SUDO_USER that is itself uid 0 on providers whose default account is root under an ordinary name. macOS never escalates, in any of the three. Homebrew refuses to run as root, the account running the script IS the owner, and the sections that needed root are the ones the macOS path skips. --help and argument errors still work with no privileges at all, since arguments are parsed before any of this. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3cca07187e |
scripts/install.sh — one command for both halves
`bun setup` runs it. Machine setup first, then officer setup, stopping if the first does not finish rather than running the second against a machine that is not ready. They stay two scripts because they answer two different questions and are worth running apart — a machine you already trust needs only the second, one you are rebuilding needs only the first. --machine-only and --officer-only say so directly, and both halves remain runnable by path. Privileges are checked here, before anything is done, because the two systems want opposite things: Linux needs root for apt, systemd, useradd, netplan and ufw and for creating directories owned by the service account; macOS must NOT be root, since Homebrew refuses to run as one. Each script already enforces its own rule, so this is only about failing early instead of halfway. officer-setup.sh gained the same OS-aware check. It required root unconditionally, which on macOS would have failed immediately after machine-setup — which must run as the user — and for no reason: there the account running it IS the owner, so there is nothing to chown and nothing to drop privileges to. Arguments are parsed before privileges, so --help works without sudo and an unknown option is rejected before anybody is asked for a password. It did not, first time round. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5aa2e0387 |
keep the shell templates together in scripts/setup
starship.toml, tmux.conf and zshrc side by side. tmux.conf and zshrc moved up out of machine-setup/. starship.toml could not have moved down to join them: the PLATFORM reads it, at src/servers/os-user-shell.ts:34, to deploy to every member's Linux account. That is a runtime path rather than an import, so moving it would have broken member provisioning silently — no build error, members simply get no starship config. So the templates collect where the shared one already had to be. Left as a marker for tomorrow rather than resolved: there are now TWO zshrc templates, this one for the owner and shell-skel/zshrc for members, while starship.toml is deliberately one file for both. Either the owner needs different shell config from a member or they should be the same file. tmux.conf has the same question waiting, since it is going into provisioning too. Noted in the zshrc header where whoever picks it up will be looking. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0da15d78a1 |
add an empty zshrc template for the owner
Somewhere to put what the owner actually wants, to be filled in and wired up tomorrow. Not referenced by the Shell section yet. The header records the decision that has to be made when it is: the section does not install a .zshrc today, it appends four marker-wrapped blocks — starship, agent, aliases, editor — through append_once. A template that is installed AND appended to ends up with the same lines twice, so those blocks either move into this file or stay out of it, not both. zshrc, no leading dot: a template in a repository, not a dotfile in a home directory, matching tmux.conf beside it and src/servers/shell-skel/zshrc which has been spelled that way all along. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e2faad40a3 |
rename the tmux template to tmux.conf, no leading dot
It is a template in the repository, not a dotfile in a home directory — and the destination it is increasingly installed to, ~/.config/tmux/tmux.conf, has no dot either. Naming the source after a path it may not be written to is how the wrong file gets read. The two remaining dotted references are correct and stay: they name the DESTINATION ~/.tmux.conf, which does have a dot when that is where tmux looks. .setup-answers and .setup-progress keep theirs too. They are runtime state in a directory, not templates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
61a3ae720f |
install the tmux config where tmux actually reads it
tmux 3.1 added an XDG location and it takes PRECEDENCE over ~/.tmux.conf. Verified on 3.4 here by writing a different marker into each and asking tmux which one it ended up with: both present -> ~/.config/tmux/tmux.conf only ~/.tmux.conf -> ~/.tmux.conf only the XDG one -> the XDG one So the Shell section writing ~/.tmux.conf on a machine that already has the XDG file produced a file tmux will never read, and reported "tmux config installed" having changed nothing anybody could observe. That is the worst shape a config step can have: it looks done. tmux_config_target now picks the path tmux will actually load — the existing XDG file if there is one, otherwise ~/.tmux.conf, which is still what every guide names and what a machine with neither should get. When both exist the section says so out loud before targeting the winner, because "your other file wins" is not something anyone infers from a success message. Also removed an untracked duplicate at scripts/setup/.tmux.conf. The one the script installs is scripts/setup/machine-setup/.tmux.conf — SCRIPT_DIR is the machine-setup directory — and two identical copies with only one of them read is the drift this whole evening has been about. Nothing to change about the config itself: the tracked copy is already byte-for- byte the owner's own ~/.tmux.conf. Not yet wired into per-user provisioning. src/servers/shell-skel/ seeds a zshrc for a member and has no tmux config beside it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
085afe7604 |
wait for Postgres instead of failing startup work once
The failure here was not a crash — it was the opposite, and that is why it would never have been noticed. server.tsx fired initQueue() and cleanupOnStartup() as bare promises with a .catch() that logged. postgres-js connects lazily, so nothing fails at import; the first query does. If Postgres is a few seconds behind — exactly what a reboot looks like, with pm2's resurrect racing Docker starting the container — both log one line during boot and do nothing else. cleanupOnStartup is the one that matters. It marks jobs interrupted by the previous shutdown and promotes the queued backlog, so failing it once leaves those jobs marked running forever: nothing retries, nothing complains again, and the only thing that would have corrected them has already run. officerdb now exports waitForDatabase(timeoutMs = 60s): polls `select 1`, logs once while waiting, resolves true or false rather than throwing. Bounded on purpose — an unbounded wait holds a process open with no way to tell starting from hung, and the caller decides what giving up means. Deliberately NOT awaited before serve(). The listener is already up by that point and holding it closed would turn a database thirty seconds late into a reverse proxy answering connection-refused instead of a page. Requests needing the database fail honestly in the meantime. The rest of the estate was already fine, which is worth recording so nobody "fixes" it again: postgres() opens no socket at construction, officer-agent's resolveOwner is an unbounded 5s retry loop written after this exact failure cost a session, and opencode, pty, headscale and the anthropic proxy touch no database at boot at all. officer-setup's Services section also waits for pg_isready before starting pm2. Not because starting early breaks anything, but because Verify would then report a failure that is really a race — and a red line that is usually noise is a red line people stop reading. Verified waitForDatabase against a dead port: logged once, returned false after the timeout, did not throw. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
62cbf510cb |
no ecosystem files in git; officer-setup generates one
Deleted all four — ecosystem.config.cjs, .light., .mac.light. and the .profile. they derived from. The repository now contains no ecosystem file at all, and .gitignore keeps it that way. officer-setup writes one at the end, describing exactly the six processes a core install runs: officer, officer-anthropic-proxy, officer-agent, officer-opencode, officer-pty and officer-headscale. No profiles, no derivation, no plugins. The four existed because a profile has to subtract from something, so the full list had to name every plugin's process whether or not anybody installed it — and a test then had to assert the two files still agreed. Generating one file removes the subtraction, the second list and the test that policed them. It is .cjs, not the .js PM2's docs use, and that is not a preference: package.json declares "type": "module", so a .js file here is ESM and `module.exports` throws. PM2 require()s the config. Sections 10 (Services) and 11 (Verify) are built on top of it — write, startOrRestart, save, optional boot hook, then check every process is online with a sane restart count AND that the API actually answers on PORT. A process can be `online` and serving nothing, so the port is asked directly rather than inferred. That completes all eleven sections. catalogue.test.ts required both deleted files at import, so it could not even load. Its central assertion — "the store offers exactly what light leaves out" — has no meaning without a full list to subtract from, which is the point of the change. Replaced by two weaker but real checks: the store must not offer a core process, and every process it names must have a sidecar directory to run. The second catches the same typo the old one did without needing a manifest of everything; verified it holds for all 15 catalogue entries. app-store/pm2.ts starts a sidecar with `--only` against this file, which now holds core alone — so it can stop a plugin but cannot start one that was never written in. Marked `[open]` there rather than left to be discovered: appending a plugin's entry is the plugin system's job. Generated one in a scratch directory and required it with node: six apps, correct cwd on each, valid CommonJS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b075f1f882 |
macOS is a dev machine, and machine-setup now treats it as one
Officer on a Mac is a dev helper on a laptop somebody sits at. It is never the homelab or VPS case, so the role is not asked for there — it is `dev`, and every section that exists to make a machine a good server is skipped. Seventeen of twenty-six sections skip, listed once in MACOS_SKIP in lib/base.sh with a reason each, rather than an `if macos` threaded through each section. Most would simply fail — no systemd, no ufw, no netplan, no useradd, no /etc/ssh/sshd_config.d — but a few would SUCCEED and be wrong, which is worse: stopping a laptop from sleeping, or freezing the address of a machine that moves between networks daily. Nine run: System update, Core utils, Tailscale, Command-line tools, Git, Docker, Neovim, JavaScript runtimes, Agent CLIs. The blocker was root. Linux needs it for nearly everything; Homebrew REFUSES to run as root and says so, so the whole script under sudo would have failed at the first brew install having already taken a password. It is now required on Linux and refused on macOS, which works precisely because the macOS path skips everything that needed it. Docker is checked, not installed. Docker Desktop is a GUI app that wants opening, permissions and a running window — not a shell script's business — and colima and lima both cost an evening the first time something does not resolve. So the step reports whether the daemon answers and points at the download otherwise. The group-vs-rootless choice below it is Linux only: Desktop runs containers in a VM owned by whoever is logged in, so there is no group to join. Added the Xcode command line tools as a macOS-only step, before anything that builds. node-pty ships no prebuilt binary on any platform and always falls through to node-gyp, so `bun install` cannot finish without a compiler — and it fails deep in a dependency tree naming neither Xcode nor node-pty. `xcode-select --install` opens a dialogue and returns immediately, so the step says to come back rather than pretending to have waited. Tailscale takes the cask, not install.sh — that script is a Linux package-manager wrapper. The cask ships a usable CLI; the Mac App Store build is sandboxed and does not. Not run on a Mac. There isn't one here, so this is read from the code and from what each tool documents, not observed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
68f2c55ecf |
one directory per feature: schema.ts and queries.ts together
src/databases/officer_db/src/<feature>/{schema.ts,queries.ts}, replacing the
parallel schema/ and queries/ trees. 24 feature directories, 46 files moved with
git mv so history follows.
The parallel trees had drifted, which is what the restructure is really fixing:
four features were named differently on each side — app-store/sidecar-installs,
email/email-accounts, server/server-config
operations had a schema and NO query file: its task_logs is reached directly
from src/servers/api/task-logger.ts, bypassing this package's own boundary
integrations had queries and NO schema, because it spans two features'
tables — server_integrations and user_integrations
Both lopsided cases survive as directories holding one file, which states the
problem instead of hiding it across two trees.
Nothing outside the package changed how it imports. `officerdb`, `officerdb/types`
and `officerdb/db` resolve exactly as before; index.ts absorbed the path changes.
Added `"./*": "./src/*"` so the new layout is reachable — `officerdb/soulseek/schema`
— which one script needed, because soulseek is a plugin and therefore commented
out of the aggregator.
schema/index.ts became src/schema.ts, keeping the core/plugin split from earlier
tonight. drizzle.config.ts and the package's "./schema" export follow it.
Verified rather than assumed: all 52 files in the package parse, every relative
import resolves against the new layout (checked by walking each specifier to a
real file, since parsing does not check paths), and everything in the tree
importing officerdb still parses. Not typechecked — empty node_modules, frozen
installs.
One rewrite bug worth recording: the rule mapping a query module's sibling import
also matched the './schema' this pass had just written, turning it into
'../schema/queries' in 22 files. Caught by the resolver check, not by parsing —
both spellings parse fine.
Also corrects every path reference the move invalidated: src/databases/CLAUDE.md's
layout diagram, the root CLAUDE.md data section, three docs, and seven sidecar
comments naming queries/<x>.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9d6c196848 |
PUBLIC_URL comes back, and gen:index takes it as an argument
Removing PUBLIC_URL earlier today was wrong, and the Build section is where it would have surfaced: gen-index.ts exits 1 without it, so `bun gen:index` fails, index.gen.html is never written, and `bun start` has no page to serve. It came out because origin validation was being discontinued — but that was one of four consumers and the only one that is gone. gen-index needs it for OpenGraph tags, which crawlers fetch standalone and cannot resolve relative; task-api-env builds OFFICER_API_HOST from it; and dav/router hard-requires it, https only, to build an iOS profile. It is also the one value this machine genuinely cannot derive, which is what separates it from DATA_PATH and the rest that left today. gen:index now takes a URL as its first argument, ahead of the environment and .env: `bun gen:index https://officer.example.com`. Changing the public address is one command rather than an edit plus a regenerate, and a second address can be generated for without touching the install's .env. It also validates now. A relative or scheme-less value substituted silently and produced OpenGraph tags nothing can resolve — invisible until someone shares a link and the preview comes back blank. .env is PORT, PUBLIC_URL, POSTGRES_URL. Verified by running the section; all three paths through gen:index exercised (absent, valid argument, invalid). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8207824a81 |
build the secret store: one key per purpose, none in .env
.env now holds PORT and POSTGRES_URL. Every encryption and signing key lives in $OFFICER_ROOT/secrets/officer-keys.db — 0600, 0700 directory, owned by the service user, created on first use. The design doc planned to move ONE at-rest key into the store. What shipped splits it: headscale, wallet, photos, jellyfin, invoiceshelf, vault and service-connections each get their own, plus jwt. VAULT_STORE_KEY encrypted all seven, so one leak opened all of them — and it was named after whichever plugin needed it first, which is why it read as safe to change if you did not run a vault. A core install bootstraps two, jwt and headscale; the rest appear when their plugin first asks. The file IS the secret. No second key unlocks it, because a key beside the store it opens buys nothing. The gain was never secrecy, it is blast radius: bun auto-loads .env into all twenty pm2 processes, so a key there is readable from /proc/<pid>/environ of twenty processes — officer-music held the key that decrypts wallet seed envelopes. Two defects found by testing the store rather than reading it, both of which would have shipped: The WAL was 0644. Enabling WAL creates -wal and -shm at 0644 rather than inheriting the database's mode, and a freshly written key lives in the WAL before checkpoint — so the 0600 on the database was decorative. The 0700 directory covered it, but only until someone loosened the directory. PRAGMA journal_mode = WAL takes an exclusive lock, and busy_timeout was set AFTER it. With twelve concurrent openers, six died on that line with SQLITE_BUSY. Every sidecar opens this store at boot, so they open it simultaneously by definition: most of them would have failed to start on a cold boot and none on a warm one. Fixed by ordering the pragmas; re-tested with twelve racing processes, one key, one row. crypto.ts takes a purpose as its first argument now, which the design doc had explicitly promised would not happen — 32 call sites across seven query modules. That promise is corrected in the doc rather than quietly dropped. Also live, not just comments: wallet/upstream.ts gated wallet storage on process.env.VAULT_STORE_KEY and would have reported "unconfigured" forever. It asks the store now, and the question it answers changed — not "did somebody set a variable" but "can this process open the store", since the key is created on demand. assertSecretsClosed covers the store, its directory and its WAL. The jwt key mints owner tokens, so a member's shell reading it is strictly worse than the .env leak that check was written for. Not typechecked: node_modules is empty and installs are frozen, so the officerdb/secret-store subpath could not be resolved at runtime here — verified that officerdb/types fails identically, so it is the empty tree and not the new export. The store module itself was tested directly: creation, idempotence across processes, hasKey not creating, permissions, and the twelve-way race. Every changed file parses; the setup section runs and degrades correctly when the import is unavailable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9864fb1a49 |
switch off the browser relay, pending extraction into a plugin
BROWSER_RELAY_PORT is gone and the second listener no longer starts. The Chrome
extension, api/browser/ and the /browser screen all stay on disk — this is going
to be extracted, and deleting it means writing it again.
Three things had to move together, and the middle one would have failed the boot
on its own:
server.tsx the listener, commented out with the variable name recorded
hono.ts the /api/browser mount, closed
registry.ts the 'browser' capability's claim on /browser, dropped
assertCapabilityTotality checks both directions: check 2 refuses to start on a
capability claiming a prefix nothing serves. Unmounting the router alone would
have left the registry describing it, and the server would not have come up.
The capability itself survives because it also claims /scrape, which shares
nothing with the relay — it launches its own headless chromium through playwright
and never speaks to the extension.
The comment in server.tsx carries the two facts that are not recoverable by
reading the remaining code. First, the port is an INPUT TO A CREDENTIAL:
relay-auth.ts derives each extension's token as HMAC(JWT_SECRET,
'officer-browser-relay-v1:${port}:${userId}:${salt}'), so bringing the relay back
on a different number silently invalidates every paired browser — reported by the
extension as "Relay not reachable", which SETUP.md blames on a wrong address,
port or token. Second, it cannot come back as a kernel-assigned port:0 like the
other sidecars: the extension is configured by hand and stores the value, so a
port that moves each restart breaks the pairing each restart.
Left alone deliberately: the /browser route in App.tsx, its Dock entry, and the
Settings → Browser Relay panel. They will not work against a closed endpoint.
Removing them is frontend work for the extraction, not part of switching the
listener off.
.env is down to PORT and POSTGRES_URL.
Not typechecked (empty node_modules, frozen installs). Every changed file parses;
the setup section was run and writes two variables.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
571d0a62ff |
delete the bug-report discord webhook, and the last dead HOME_DIR reads
DISCORD_BUG_REPORT_WEBHOOK is gone, with sendToDiscord and its helpers. Reports still land in DATA_PATH/bug-reports — the disk write always happened first and the webhook was only a ping about it, so nothing about the report is lost. It was a personal notification channel living in deployment config, on a platform whose owner is the only person who files reports. It was also never in .env.example: the setup script wrote a variable nothing documented, which is the same drift as PORT, in the other direction. Note DISCORD_WEBHOOK_URL is a DIFFERENT variable — the notify sidecar's own channel — and is untouched. Then a parity sweep of setup / .env.example / what the code reads, which turned up two leftovers from earlier today: HOME_DIR was still read in six files, each with its own `?? homedir()` fallback. Dead since nothing sets it, but a dead read is worse than none — it reads as a supported override. They take homedir() directly now. user-instance.ts gets a comment on why its line stays where it is: it sits above `process.env.HOME = homeDir`, and homedir() reads $HOME, so a read moved below that assignment would return whichever member was last spawned into. Two of the six had fallback chains ending in process.cwd() and '' — the second would have silently disabled whatever consumed it rather than failing. VAULTWARDEN_URL was uncommented in .env.example among the variables setup writes, though it is a plugin variable setup has never written. Commented out with the other plugin entries. The three files now agree: setup writes PORT, BROWSER_RELAY_PORT and POSTGRES_URL; .env.example lists those plus JWT_SECRET and VAULT_STORE_KEY, which are required by code and deliberately unwritten until the secret store lands. Not typechecked (empty node_modules, frozen installs). Every changed file parses; the setup section was run and writes three variables. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f063fc0c08 |
remove origin validation
ALLOW_ANY_ORIGIN, ALLOW_ANY_ORIGIN_MUSIC, and everything they gated. The flag defaulted to ON, so none of it ran on a real install — what comes out is documented defence in depth that was already switched off. The file said so itself: "Both flags and their call sites come out once the tailnet is the perimeter." Origin was never authentication here in any case. An app's `officer://<hex>` origin is chosen by the client, forgeable outside a browser, and extractable from a shipped binary. Gone: the two flags, isOriginAllowed, isOriginCheckDisabled, isMusicOriginExempt, originValidationMiddleware, ORIGIN_RULES and the whole OFFICER_<APP>_ORIGIN scheme, PUBLIC_URL's origin/host derivation, and origin-validation.test.ts, which existed only to pin them. CORS now echoes whatever Origin it is given, which is what every install already did. What SURVIVES is the reason this needed care. origin-validation.ts held two unrelated things, and the second was the global authorization gate — a valid non-owner token reaches only what its role grants, deliberately NOT under the flag because it is account-based rather than origin-based. Its own comment called it "the airtight half". Deleting the file wholesale would have deleted authorization. So it moves to _middlewares/capability-gate.ts as capabilityGateMiddleware, with the name matching what it does: nothing in it reads an Origin header any more. hono.ts mounts it in the same position, ahead of every router. origin-middleware.ts stays and is untouched — it extracts the Origin for six auth handlers that log it, and for passkeys. Extraction, not validation. Also updates every claim that rested on the old model: CLAUDE.md's security section and repo map, docs/secret-store.md, docs/mobile-api-keys.md, and five messages in machine-setup's Tailscale section which told the owner to set ALLOW_ANY_ORIGIN=false when declining a tailnet. That advice is now impossible to follow, and the honest version is different: with no tailnet the token is the whole lock, so put a proxy in front and restrict who can reach it. Not typechecked (empty node_modules, frozen installs). Every changed file parses; the setup section was run and writes four variables now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f071c0b24 |
the install root is derived, not configured
Seven variables out of .env. DATA_PATH, OFFICER_ITEMS_DIR and HOME_DIR are gone from the code entirely; PUBLIC_URL, PUBLIC_BUILD_ENV, JWT_SECRET and VAULT_STORE_KEY are no longer written by the setup script. data-path.ts now derives OFFICER_ROOT as dirname(process.cwd()), with data/, capabilities/ and dockers/ as fixed names under it. The direction used to run the other way — DATA_PATH from env, then OFFICER_ROOT = dirname(DATA_PATH) in app-store/paths.ts — which meant three environment variables that had to agree with each other and with the tree on disk. Eight files re-read process.env.DATA_PATH independently, each with its own `?? cwd()/data` fallback. They import the one value now, which is what made removing it safe: otherwise each would have derived its own and drifted. Three things this turned up. The cwd pin in ecosystem.profile.cjs was broken. It set `cwd: __dirname` under a comment asserting "__dirname is the repo root — this file sits beside ecosystem.config.cjs", which stopped being true when these files moved into ecosystem-files/. It walks up to the platform's package.json now, which holds wherever the file lives. That was a live bug before this change and a load-bearing one after it, since cwd now decides where the install is. assertInstallLayout joins the other two boot assertions. A wrong cwd does not error — it computes a plausible root somewhere else and writes managed homes and agent runs into it, so the install looks empty and the data looks lost with nothing naming the cause. It throws before serve(), first of the three, because a wrong answer there makes the other two check the wrong files. getOwnerHomeDir captures homedir() once at module load rather than per call. Measured on bun 1.3.10: both os.homedir() and os.userInfo().homedir return $HOME when set rather than reading passwd, and user-instance.ts assigns process.env.HOME on its way to spawning an agent. A lazy read would have returned the owner's home on the first call and a member's afterwards. data-path.ts imports only node builtins, so it is evaluated before any of that runs. JWT_SECRET and VAULT_STORE_KEY leaving .env means an install made by this script does not boot — jwt.ts throws at module load without one. That is the agreed sequencing: they move to the SQLite store (docs/secret-store.md), and writing them here meanwhile would create a second origin for a secret the store then has to be reconciled with. Said plainly in .env.example and in lib/env.sh rather than left to be discovered. Not typechecked: node_modules is empty here and installs are frozen. Every edited file parses under `bun build --no-bundle`; the profile loads and pins the right cwd; assertInstallLayout was exercised from both the repo and /tmp; the setup section was run and writes five variables. Prettier was NOT run — 3.9.6 via bunx is not the pinned resolution and reformatted unrelated unions and line wraps in six files, so those were reverted and the edits re-applied by hand. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
86079adb9a |
mail transport is configured in the app, not in .env
MAIL_TRANSPORT was a fallback left from the old registration flow that sent
confirmation mail. That flow is gone; the variable outlived it.
It was never the primary source anyway. getTransport reads server_config
('server-settings' → smtp) first, which already backs a full UI at Settings →
Server → SMTP and its API in api/server-settings/smtp.ts, supporting resend,
smtp and mailhog. The env var only answered when that was absent — which is a
second source of truth for something the owner can already set, with the failure
mode that a stale URL in .env silently answers for a server whose settings row
is simply empty.
Removed from transport.ts, .env.example and the setup script's Environment
section, which no longer asks for it. setup-old/ still mentions it; that is the
archive and is left alone.
Also split the try. It wrapped the read AND the transport construction and
swallowed both, so three different problems produced one message. Unreachable
database, nothing configured, and stored settings that do not build a transport
now say different things, because the fix for each is different and this message
is all the caller ever sees.
The two consumers — queue/engine.ts and auth/forgot-password.ts — now raise
until SMTP is set in the UI, which is the honest answer rather than a regression.
Not typechecked: node_modules is empty here and installs are frozen. transport.ts
parses under `bun build --no-bundle`; the setup script was run and no longer
prompts for or writes the variable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
cb67b22f80 |
officer-setup: the environment section
Writes .env, and the whole point of the section is the two values it must not write twice. JWT_SECRET and VAULT_STORE_KEY are read back from any existing .env and kept. The original script reminted JWT_SECRET on every run that agreed to regenerate .env, which logs every device out with no stated reason, and never wrote VAULT_STORE_KEY at all — so a scripted install had no at-rest key and the vault and wallet refused to store anything. VAULT_STORE_KEY is the more dangerous of the two now that it is being written. It is not Vaultwarden's despite the name: it encrypts every secret column in Postgres, and the wallet seed envelope on top of the owner passphrase. Changing it is unrecoverable for the seed, because the passphrase opens the inner envelope and that is the outer one. Said in the section, in the file it writes, and in .env.example, which described it as Vaultwarden's and understated it. DATA_PATH and OFFICER_ITEMS_DIR are derived from $OFFICER_ROOT rather than asked — two questions that had to agree with each other and with the app store. ALLOW_ANY_ORIGIN is written explicitly from whether tailscale0 exists, rather than left to the platform default. The default is ON, which CLAUDE.md says is only defensible because the tailnet is the perimeter; with no tailnet there is no perimeter, so it goes out as false. Added to .env.example, which omitted it. PORT defaults to 9000, matching .env.example. The old script used 9010; nothing depends on either, and it is a prompt. write_env restores the prior umask. It was set to 077 so the secrets are never briefly world-readable, but umask is not scoped to a function and would have made every file the later sections create owner-only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0bceace6f1 |
an officerdev docker network, and the reasoning next to the port binding
One network for everything Officer provisions, created before anything joins it and declared external in the compose file. Postgres needs nothing from it today — the platform is a host process reaching it over loopback — but a reverse proxy in front of the web UI does, and so does any app-store service that talks to another. Creating it now means the later ones do not have to be migrated onto it. Two things written next to the line they explain, rather than assumed: Why loopback. Publishing a port makes Docker write its own DNAT and ACCEPT rules into iptables, and those are evaluated BEFORE ufw sees the packet — so `ports: "5432:5432"` is reachable from the internet while `ufw status` reports everything denied. That is the same mechanism the machine-setup firewall section hooks DOCKER-USER to close. Binding to 127.0.0.1 sidesteps it: the DNAT rule only matches traffic arriving on loopback. Why the password is not decoration. Loopback means nothing off this machine, but every account ON it can open 127.0.0.1:5432 — including the per-user Linux accounts Officer gives its members. What stops them is that they cannot authenticate. The password is the boundary between the platform and anyone with a login here, which is why it stays random and why both files holding it are 0600. A unix socket would remove even that, and was ruled out for a specific reason: postgres.js only treats a host as a socket path when the host FIELD contains a slash (src/index.js:468), and officer_db/src/db.ts passes a bare URL string. It would take a change to db.ts, which is not a setup-script change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a64d5610e6 |
officer-setup section 5: Postgres, and only Postgres
The original offered five containers. Of those:
Postgres the only database Officer has — account, passkeys, settings,
dashboards, email accounts, the queue. Required.
Redis not referenced anywhere in the platform. No import of the client,
no environment variable, no mention; the only "redis" string in
src/ is the word "rediscover" in a comment. Dropped. (It is still
in package.json and comes out in the dependency pass.)
SearXNG zero references anywhere. Dropped. If it ever arrives it brings
its own compose file and its own Redis with it.
Mailhog a development convenience, offered separately rather than here.
Nginx PM a deployment choice — Caddy, Traefik, nginx or the tailnet — and
not something a setup script should pick.
Provisioned into $OFFICER_ROOT/dockers/postgres/, the same convention the app
store uses: one directory per service, the compose file in it, relative bind
mounts so the data sits beside the compose file.
Bound to 127.0.0.1, deliberately and with the reason in the compose file itself.
Docker publishes ports by writing iptables rules underneath ufw, so "5432:5432"
is reachable from the internet whatever the firewall reports — the same mechanism
the machine-setup firewall section exists to close. The platform runs on this
machine, so loopback is all it needs.
The password lives in a 0600 .env beside the compose file rather than inside it,
so the compose file can be read or copied without carrying a credential. A second
run reuses it rather than minting a new one, which would leave the container and
the URL disagreeing.
Readiness is waited for rather than assumed: Postgres initialises its data
directory on first start, and db:push against a database that is still starting
fails in a way that reads as a schema problem.
Choosing an existing database checks the URL but does not insist on it — the URL
may be right and the database not yet started, and refusing to continue over that
would be worse than saying so.
Also carries the whitespace fix for the comment removed in the previous commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a0b083db8f |
officer-setup section 2: one root, and nothing configurable underneath it
$OFFICER_ROOT/
platform/ the app
data/ managed homes, attachments, job logs
dockers/ anything the app store provisions
capabilities/ skills, tools, tasks, processes
The original asked separately for DATA_PATH and OFFICER_ITEMS_DIR and left the
app store's directory implicit — three answers that had to agree with each other,
given by somebody with no reason to know they had to. One question now, at the
top of the run, and the rest follows from it.
This is also what the code already assumes rather than a new convention:
app-store/paths.ts derives OFFICER_ROOT as dirname(DATA_PATH) and DOCKERS_DIR as
OFFICER_ROOT/dockers, so writing DATA_PATH=<root>/data is the whole of what makes
the layout correct. No code changes.
Anybody who wants data/ on a bigger volume can symlink it. That is a decision
about storage, not about how Officer is laid out, and it does not need a prompt.
The one check worth having: a directory that exists but belongs to somebody else.
That happens when an earlier run created it as root, and everything written into
it afterwards fails in a way that reads as a permissions bug in the platform
rather than as a bad directory. Reported with what writes there and offered as a
chown.
Placed before the repository, because the checkout lands inside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8a0946ea39 |
officer-setup section 3: dependencies
bun install, as the account, in the checkout. Two things stated because a failure here is otherwise opaque. The lockfile is frozen — bunfig.toml sets [install] frozenLockfile = true — so bun resolves from bun.lock and nothing else. A package.json that disagrees with it is a hard failure rather than a quiet resolution, which is deliberate: the friction exists so an unexplained lockfile change shows up in a diff. If the install fails complaining about the lockfile, the section says that is the frozen lockfile working and that it wants a human to read the diff, rather than reporting a generic failure. node-pty has no Linux prebuild, so this compiles it from source on every machine. That is what build-essential and python3 are in machine-setup's core utils for, and the section says so — the failure would otherwise surface much later as a terminal that never starts. Success is checked by the artefact rather than by the exit status: bun can complete while the native module is not built, because it skips a dependency's lifecycle scripts unless it trusts the package. So the section looks for node_modules/node-pty/build/Release/*.node and, when it is missing, names the consequence and the command that fixes it instead of reporting success. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ee21fa16f0 |
officer-setup: the repository URL is https
https://gitea.officer.dev/officerdev/platform.git, not the ssh form. The reachability check is now one test for either scheme: `git ls-remote` with both prompts disabled. That is the real question — not whether the host answers but whether this account can read the repository — and neither prompt fails cleanly on its own. Over https git asks for a username nobody is there to type; over ssh it asks for a password or stops on host-key verification. With GIT_TERMINAL_PROMPT=0 and BatchMode both off, an unreadable repository is an immediate non-zero rather than a hang. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5a8cd4e9c |
officer-setup section 2: the repository
Clones from ssh://git@gitea.officer.dev:2222/officerdev/platform.git, or uses the checkout already at $OFFICER_ROOT/platform. Cloned as the account, never as root. A repository owned by root is one the owner cannot pull, cannot commit in, and whose node_modules they cannot write — and every later section in this script writes into that directory as them. Three things it refuses to do quietly: It does not repoint an existing remote. This checkout points at gitea.pastilhas.dev rather than the new gitea.officer.dev; that is reported with the command to change it, because where somebody's work pushes to is their decision. It does not pull over uncommitted changes. A dirty tree means the pull is skipped and said so, rather than failing halfway or burying the work. It pulls with --ff-only, so a failure means the branch has diverged rather than that the network was down, and the message says which. SSH reachability is checked before the clone, not after. An ssh URL with no usable key does not fail cleanly: git prompts for a password nobody is there to type, or stops on host-key verification. BatchMode turns both into an immediate answer, and the check reads the server's response rather than the exit code — Gitea greets a successful authentication and then exits 1, so exit status alone reports success as failure. When the key is missing it offers the https form of the same URL, which works without a key if the repository is readable anonymously, and otherwise stops and says to add the key. Verified against the new host: ssh authentication from this account already works. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b724e3ffbe |
start officer-setup: pre-flight, inheriting what machine-setup already asked
The second half of the install, and a much smaller script than the original: of the old setup.sh's thirteen sections, seven are machine-setup's job now and two more were already removed. What is left is the repository, dependencies, the database, .env, the schema, the build and pm2. Pre-flight asks nothing on a normal run. machine-setup saves the account, the Officer path and the role beside itself, and this reads the same file — so machine-setup then officer-setup is two scripts and one set of answers. It prompts only where that file is absent, which is a supported case rather than an error: somebody may have provisioned the box their own way. It then checks the machine is actually ready — git, node, bun and pm2 required, docker optional — and reports all of them together with what each is for. Finding out about a missing bun three sections in, after a repository has been cloned and a database started, is a worse way to learn it. A missing required tool stops the run and names machine-setup. Found by running it: a remembered answer can go stale. My own earlier testing had left SETUP_USERNAME=gitfresh in that file, for a throwaway account I then deleted, and the run dead-ended on it. A remembered account that no longer exists is a reason to ask again, not a reason to stop — so it is checked before it is trusted, reported, and replaced. Docker being absent is a warning rather than a failure: Postgres can be one you already run, and the app store simply cannot provision until Docker is there. Sections 2 to 9 are listed and not built. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e120dfa36e |
full read: fix the set -e footguns a full run would have hit
Read the whole thing — 2392 lines of entry point and 2700 of libraries — looking for what shellcheck cannot see. shellcheck itself is clean at error level; its warnings are cross-file false positives and one deliberate tilde in a display string. Everything below is a real defect. ── The Git section aborted on any machine where git was not already configured ── `git config --global --get <key>` exits NON-ZERO when the key is simply unset, and `VAR="$(git_get …)"` propagates that under `set -e`. So on a fresh machine — the case this script exists for — the section died at its first assignment, before printing anything, and took the remaining nine sections with it. It passed every earlier test because those harnesses sourced the section under a `bash -c` with no `set -e`. Verified now against a genuinely fresh account with the real script: the section completes and writes a correct .gitconfig. ── An optional step failing aborted the whole run ── Twelve functions ended on a command that can fail — `systemctl enable --now earlyoom`, `systemctl restart systemd-logind`, `chsh`, `sysctl -w`, `chown -R`, the oh-my-zsh installer, and others. Called as plain commands under `set -e`, any one of them failing ends the script, so a masked unit or a container without systemd would abort a 28-section run over an optional improvement. They now return 0 explicitly and the callers verify the outcome instead — which also fixed a lie: the sleep section printed "sleep disabled, logind reloaded" whether or not the restart had worked. It now checks the targets and the logind values and reports honestly. ── chown user:user assumed the primary group is named after the user ── True on Debian and Ubuntu, which create a group per user. Not true for an account from LDAP, or made with `useradd -g users`, or on an image with a shared group — there `install -g <user>` fails with "invalid group" and the step aborts. Proved it against an account whose primary group is `oddgroup`: the old form fails, the new one gets ownership right. Eight call sites now ask `id -gn`. ── Also hardened ── agent_path and current_editor gained `|| true` for the same reason git_get needed it: "nothing is set" is an answer, not a failure. Verified afterwards: shellcheck clean at error level, every section runs standalone without aborting, and the two apparent failures in that sweep are correct behaviour — Timezone and Git refusing an empty answer from /dev/null. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
30052e3295 |
port the firewall, last, and bind the Docker rules to the real interface
Last in the run for the reason the original gave: enabling a firewall is the one step that can cut the connection it is running over. The security bug is in the shipped rules. ufw-docker-rules.conf hardcodes eth0 in all three of its rules. Docker publishes container ports by writing its own iptables rules underneath ufw — DOCKER-USER is the hook that lets ufw have a say at all — so on a machine with predictable interface names (ens18, enp1s0, most VPS images) none of those rules match, the final DROP never fires, and every published port is open to the internet while `ufw status` reports active. A firewall that says it is working and is not is worse than no firewall. The rules are now substituted with the interface the machine actually uses, verified by applying them against a stubbed ens18. Order inside the section is the other thing that matters: OpenSSH is allowed BEFORE anything is enabled, unconditionally, because a firewall enabled without an ssh rule on a machine reached over ssh needs a console to fix. The prompt says so, and says to open a second session before closing the current one. tailscale0 is checked and offered, because the default is deny inbound and the tailnet is an inbound interface like any other — without that rule Officer is unreachable over the tailnet while Tailscale reports itself connected. A correction to something I said while writing this: I reported that this host was missing its tailscale0 rule. It is not. I had run `ufw status verbose | head -8`, which cut the output above the rule list. The full status shows it allowed, and nothing was wrong. That leaves the NOT PORTED list empty. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d6a78d7b5a |
put the shell configuration in one place, after everything it configures
The Shell section moves to 26, after Neovim, the runtimes and the agent CLIs. Everything that writes to .zshrc now happens there and only there: the starship init, the PATH for the agent CLIs (moved out of that section), the aliases, and the editor. That ordering is what the editor choice needs — it offers whichever of nvim, vim and nano are actually present, so it has to run after Neovim is installed rather than naming an editor that is not there. Which was the original's mistake in the other direction: it set core.editor to nvim four sections before installing it. The default editor is the setting git's core.editor was deliberately left out in favour of. EDITOR, VISUAL and SUDO_EDITOR go in the account's shell, and the Debian `editor` alternative is set too — an account's shell config cannot reach root or sudoedit, and those are exactly the cases where the wrong editor is most annoying. Recorded a limitation of append_once while cleaning up after it: renaming a marker orphans the block that used the old name, and changing a block's content does nothing because the marker is still found. Both need the old block removed by hand. This run left exactly that — a `local-bin` block superseded by `agent-clis` — in the dev box's .zshrc, now removed. UFW is deliberately still unported and will be last, for the reason the original gave: it is the one step that can cut the connection the run is happening over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
607a962115 |
port the agent CLIs, using Anthropic's installer rather than npm
Claude Code goes in through https://claude.ai/install.sh, matching what the platform already does for members in os-user-claude.ts and chosen there for the auto-update npm does not give. The comment at os-user-claude.ts:28 claiming setup.sh already did this was simply wrong — setup-ubuntu.sh used `npm install -g @anthropic-ai/claude-code`. Two things that installer insists on, both of which a naive port gets wrong and both of which os-user-claude.ts had already found: It REFUSES to run under sudo from a regular user's shell — it checks for uid 0 with SUDO_USER set, because everything it writes goes under $HOME and under sudo that is root's. This script runs as root, so the install has to be done AS the account. It declares #!/bin/bash and uses [[ … =~ … ]], so it must be piped to bash. On Ubuntu /bin/sh is dash and `| sh` fails. Two bugs found by running it rather than reading it: opencode does not install to ~/.local/bin. It goes to ~/.opencode/bin, which is what sidecar/opencode/index.ts:22 hardcodes. The first version looked in the wrong place, reported a working install as missing, and installed it again — the run said "did not complete" while the installer had plainly succeeded. claude on this machine came from npm, so `command -v claude` found it and the section would have left a copy that never updates. It now detects an npm install by resolving the binary into node_modules, says so, and offers to reinstall through the official installer — naming the npm copy and how to remove it rather than deleting something it did not put there. ~/.local/bin and ~/.opencode/bin are both added to the account's PATH. The sidecars do not need it — they check the exact paths — but a user who cannot run `claude` in their own terminal reasonably concludes it was never installed. PI stays optional and says outright that nothing in the platform spawns it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cb3e7062b9 |
ensure the bun symlink on every run, not only after installing it
The link was made as part of install_bun, so a machine that already had bun never got one. That is this machine: bun 1.3.14 in ~/.bun/bin, no /usr/local/bin/bun, and `bun` resolving to nothing at all for root. Nothing has broken yet only because `pm2 startup` has never been run here — the moment boot persistence is enabled, all twenty ecosystem apps that say `script: 'bun'` fail at boot and work perfectly when started by hand. ensure_bun_symlink now runs whether or not this script did the install, and says which of the three things happened: made it, found it already correct, or could not find bun to link. The last records an error, since a missing link is a reboot-shaped failure rather than a cosmetic one. Safe across upgrades, which was the question: a symlink resolves by path, not by inode, and `bun upgrade` replaces the file at $BUN_INSTALL/bin/bun rather than moving it. Demonstrated by replacing a target with a new file — new inode, link still resolves. It breaks only if the home directory goes, which breaks bun anyway. Also fixed the status line, which reported "not installed" on a machine with bun in the user's home: it asked root's PATH, which is exactly what has no bun before the link exists. bun_version now asks whichever copy is there. This run created the link on this machine — /usr/local/bin/bun -> the account's copy, and root can now run bun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00e58931ff |
port the JS runtimes: node, bun and pm2 installed rather than offered
Not a choice. Officer does not run without them, so asking would be asking
whether to install Officer — which was settled by running the script. They are
installed and reported, with the reason each is load-bearing stated once:
node pm2 is a Node application, and officer-pty compiles node-pty against
whatever Node is installed. There is no Linux prebuild, so this is not
an ABI question — it is a build dependency on every machine.
bun the platform itself and nineteen of the twenty pm2 apps.
pm2 supervises all of them, and the ecosystem files are written for it.
Node now tracks the current LTS, asked of nodejs.org, rather than the pinned
setup_22.x the original used — which ages into "the version we happened to pick"
the moment a new LTS lands. Resolves to v24.19.0 (Krypton) today, and NodeSource
publishes setup_24.x, checked with a HEAD request before anything is piped into a
shell.
Deno is the one genuine choice and stays optional, defaulting to no. Nothing in
Officer imports it — verified across the whole tree, the only references left are
in the old setup script — so the prompt says that outright and offers it for the
user's own work rather than pretending it is part of the platform.
bun is installed as the account and then symlinked into /usr/local/bin. pm2
started at boot by systemd has no login shell and therefore no ~/.bun/bin on
PATH; without the symlink every bun-based sidecar fails on reboot and works when
started by hand, which is a miserable thing to debug.
Every install is verified after it runs rather than trusting an exit status. A
NodeSource run can succeed while apt holds an older nodejs back, and reporting
the version asked for instead of the one present is how a machine ends up
disagreeing with its own setup log. Tested with an installer stubbed to succeed
and change nothing: both report failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|