1292a5c5abbfe2c8aae602aa4cc32cc6be01e335
1275
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1292a5c5ab |
opencode is owner-only until it carries an identity
a turn on the opencode harness ran as the owner, in the owner's home, whoever asked. handleOpenCodeChat resolves its cwd against getOwnerHomeDir(email), which discards the email it is given, and the sidecar runs one shared `opencode serve` as the service user — sendOpenCodeStreaming accepts userId/email/username and forwards none of them. it carried a comment calling itself owner-only; nothing enforced it. reachable by any account with the `chat` grant, which every role holds by default (DEFAULT_ROLE_CAPABILITIES), and isClaudeModel is a startsWith, so a typo'd model string landed there too. the model is client-supplied and never checked against the catalogue. the same gap on the read side: opencode's session store has no per-user scoping at all, so loadOpenCodeSession/delete/rename take an id and no identity, and the list and live routes returned other people's conversations. so: ChatIdentity carries isOwner as its own fact (not inferred from osUser === null, which holds only while resolveHomeDir refuses a member without one), and every opencode door in chat.ts checks it — list, load, live, delete, rename — plus a refusal on the execution path in handleChat. /chat/models hides opencode from non-owners as a courtesy; the socket refuses regardless. a stopgap, not a design. the fix is to thread identity through the opencode sidecar the way spawnClaudeAsMember does, and TODO.md has been saying so. not fixed here, and worth knowing: a member's session list is still empty and /chat/pwds still 500s, because readdirSync on their ~/.claude/projects is EACCES — claude creates it at mode 700, which zeroes the ACL mask. visible in officer-error.log right now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4dc7cd90c2 |
a plugin declares permissions, not capabilities, and has no kind
'capabilities' already means three things in this codebase — the permission
registry, the officer-items store, and the sidecar's routing keys. a fourth
would be one too many, and the field is really just permissions. the name is
free: the old permissions table went in
|
||
|
|
b18601530f |
a manifest for offscale, and the rule it immediately broke
written against the real plugin rather than invented as a field list, on the theory that an abstract one includes what nothing needs and misses what is awkward. that paid off on the first field that mattered. the rule here said a plugin may declare `app` and nothing else. offscale's capability is `admin` — owner only — and should stay that way, so the rule was wrong. the distinction is direction, not privilege: `core` means every account and not deniable, so claiming it grants yourself to everyone; `admin` means owner only, which is a plugin restricting itself. corrected table in the doc. core, execution and confined stay the platform's to assign. `publisher` is the only input to the mount prefix, through one function, so first-party and third-party cannot drift into two code paths. sidecar.runtime is a field because officer-pty needs node for node-pty's abi while everything else is bun — one plugin already needs it, so not speculative. dependsOn is informational and unenforced. code dependencies need no declaration now that a plugin builds inside the workspace, and service dependencies already degrade; this exists so the store can say the console section wants the terminal plugin, rather than the section silently doing nothing. health is marked deferred rather than open, with the reasoning, so it does not get re-raised. migrations likewise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6b2905cc7 |
how the frontend ships, and what plugins may depend on
everything moves to the plugin, frontend included, so federation stopped being
a later problem and had to be answered. it is answered by not needing it: bun
builds the spa into build/ at start and rebuilds it on install, serving from
that directory instead of compiling through the html import. Bun.build is a
runtime call, so an install needs no restart — just a refresh. same origin
throughout, which is why there is no cors work and no rewrite of useClient.
App.tsx keeps core routes and gains one map over `plugins`, each mounted at a
wildcard delegating to the plugin's own router. that list comes from a generated
Plugins.tsx, because a bundler cannot follow import(runtimeString) — the
specifier has to be concrete before the build. the six places the shell
currently hardcodes headscale collapse into that one file, dock included; the
runtime dockItemsFromPlugins path follows rather than competing with it.
presentation moves to build time, permission stays runtime.
dependencies turned out to be two different problems wearing one word. a service
dependency (assist → anthropic-proxy) is a wire call and already degrades. a
code dependency (ConsoleView → TerminalView) is in the bundle and cannot. rule:
may depend, must degrade. service calls go through the api carrying the user's
token, with the user's own permissions, which also deletes the state-file read
claude-proxy uses today to lift the proxy's secret.
no per-plugin permission list: a plugin is part of the app and bounded by the
account calling it. that makes marketplace review a security boundary rather
than a naming one, which is worth knowing rather than discovering.
and the developer environment is a platform checkout — clone it, run dev, build
the plugin inside. the 13 workspace packages resolve by name because bun links
them, so `import { useClient } from 'hooks/useClient'` just works with no
registry and no versioning. dev-time and build-time become the same mechanism.
also writes down the headscale inventory now that it has been read end to end,
including that assist.ts travels unwired as a marker and must not be tidied away
as dead code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7f26f0b4b8 |
offscale is headscale plus the companion, not a rename
the name looks like branding on someone else's project, which is exactly how it gets 'corrected' back later. it is not: offscale is the stock headscale server plus the companion that ships beside it, and the invite flow is the first thing that only exists there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
01a20fff4e |
the invite flow replaced device enrolment; it is not a gap
closes the one open item left by deleting /api/vpn. removing the vpn capability leaves no member-grantable headscale surface and that is correct: the owner mints an invite from the headscale app, the companion turns it into the redirect the phone claims, and the device joins. no per-member permission on officer is involved at any step. recorded as decided rather than open so nobody reintroduces a member-facing enrolment route believing something was lost. nothing was — /api/vpn/enroll was the design the invite flow replaced, and it never had a UI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88a44ec4a7 |
delete /api/vpn
it had no caller. verified three ways before removing: nothing in the mobile monorepo reaches it (enrollVpn's only call site is behind `if (embedded)`, and the one app rendering VpnScreen never passes embedded), nothing in the officer web app references it, and the live database holds no vpn grants. the companion was checked separately by its own author — zero references there either. and it will not come back. offscale is permanently standalone: the thing that gets you to the platform cannot itself need the platform, or a broken tailnet locks you out of both. gone: api/vpn/router.ts, its mount, and the `vpn` capability. the registry keeps a comment where the capability was, because its removal has a cost worth recording — headscale is admin-only, so no member-grantable headscale surface remains, and reintroducing one is a deliberate act rather than an oversight. kept: the sidecar's enroll.ts. its bare POST /_officer/enroll handler is now unreachable, but the file is also the dispatcher for /enroll/invites, which is live and fundamental. the header comment now says so, so nobody deletes it looking for dead code. also records the third component in the doc. two of the three have an "enroll" surface and only one is ours: /api/v1/enroll/* belongs to the companion, is where the phone actually goes, and must not be collapsed into /api/offscale/*. capabilities tests: 17 pass / 8 fail both before and after, stash-verified — the 8 are pre-existing, in totality and path-to-capability, which is precisely the machinery dynamic mounting will rework. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bbc60b34ac |
headscale leaves the baseline, and offscale gets a design doc
first step of extracting headscale into a plugin. CORE_PROCESSES is five now, and the catalogue.test CORE[] mirror follows it — not optional, since that list asserts "the catalogue must not offer a core process" and would have blocked adding offscale to the catalogue later. the local generated ecosystem file lost its entry too, and officer-headscale was stopped and deleted from pm2 by hand. the platform still mounts /api/headscale and still declares the headscale and vpn capabilities, so the feature is present-but-unavailable rather than gone. docs/offscale-plugin.md is a live document for the rest of it. what it records that nothing else does: core is now `officer` alone and everything else is a plugin; routes are /api/<app-name> for ours and /api/p/<creator>/<app-name> for third parties, derived by one function so the two can never become two systems; tables stay in public with an app-name prefix; mounting becomes genuinely dynamic, which retires the "every route stays mounted" premise and relocates assertCapabilityTotality from a boot check to a per-mount transaction. it also records a rejected experiment with evidence — a postgres schema per plugin works completely, including cross-schema FK, idempotent push and DROP SCHEMA CASCADE as uninstall — and the reason not to: drizzle-kit 0.31.8 needs schemaFilter naming every schema, contradicting its own docs, and without it push reports "No changes detected" and creates nothing. a plugin install that reports success and makes no tables is the exact failure shape we have hit three times this week. and /api/vpn is dead: no caller in the mobile monorepo, none in the web app, no grants in the database. offscale is permanently standalone, so it never comes back. the invite flow is unaffected — the phone claims from the Companion, not from officer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d000cedf2f |
let anyone toggle hidden files again
the show/hide dotfiles button was dead for members. they open the browser at their own home, that home is `/`, and the toggle is disabled at `/`. it was never meant to apply to them. |
||
|
|
336e718463 |
read a member's transcripts as the member
a provisioned member could chat normally and had no conversation list. every
refresh came back empty, so nothing could be resumed, and a new chat never
became a saved one.
nothing was wrong with the logic. the turn runs as them, writes its transcript
into their home, and the platform looks in exactly the right place — it just
cannot read what it finds.
confineUserTree grants the service user a named acl entry on every member home,
with d: defaults so anything created later inherits it. that entry is real and
getfacl shows it. it does not survive a file created at mode 600, because posix
derives the acl mask from the group bits of the creation mode:
user:officer:rwx #effective:---
mask::---
claude writes every transcript at exactly that mode — .claude and projects/ are
775, every *.jsonl is 600. so readdir and stat worked, every read raised eacces,
and summarizeTranscript catches eacces and returns null. the sessions did not
fail, they vanished.
no acl can fix this. the creation mode ands the mask down, so d: defaults cannot
raise it, and the only way up is through `other`, which is every account on the
box. a 600 file has two readers: its owner, and root.
so read as the owner of the file, through the same runAsArgv the terminal and
the agent already use. spawnSync keeps it synchronous, which is what lets it
drop into a 914-line synchronous parser reached from five modules instead of
rippling await through all of it.
the privileged surface turned out to be seven call sites, not the file: stat
needs traverse and readdir needs read, and the 775 directories give both. only
content needed identity.
also fixes a 500. parseClaudeTranscript read the file uncaught after an
existsSync that passes, so deep-linking /chat/<id> as a member threw rather than
404ing. it returns null now, like the list path always did.
verified against a throwaway linux account provisioned the same way a member is
— 700 home, named acl, transcript written as them at 600. before: 0 sessions and
loadClaudeSession null. after: the session, its title, its messages, and a
rename that leaves the file owned by the member at 600.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
fe0012635a |
stop leaking the parent claude session into the one we spawn
pm2 inherits the environment of whoever ran pm2 start, so restarting this sidecar from inside a claude code terminal — which is how it is restarted most of the time — bakes that terminal's session into the daemon. right now this process is carrying CLAUDE_CODE_MESSAGING_SOCKET for an unrelated pid that has been alive for an hour and a half. three of these were already stripped; the rest arrived with 2.x and were never added. this is hygiene, not the fix for today's hang — a spawn was verified to succeed with the whole set present — but a child attaching to a stranger's ipc socket is not a failure anyone would recognise from the symptom. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
547662842b |
a dead claude process no longer hangs the chat forever
the sdk runs two independent tasks per session: the consumer loop (`for await (const msg of q)`) and an input pump that writes the queue to the child's stdin. the consumer loop's `finally` is what removes a session from the map — but when the CHILD dies it is the input pump that fails, with `ProcessTransport is not ready for writing`, and that rejection neither ends the consumer loop nor is caught anywhere. so the loop stayed parked on a stream with no writer, `finally` never ran, the session stayed in the map, and spawnClaudeStreaming handed every later turn to the same corpse. each one pushed a message onto a queue nobody drained: no error, no result, no timeout. the client spun forever and the only trace was one unhandledRejection line in the sidecar log. observed on the host today; the only cure was pm2 restart officer-claude-code. a member's turn already supplied its own spawn function because it has to go through setpriv. the owner had none, and therefore no place to observe the child — which is exactly why its death was invisible. so give the owner one too, and wrap both in watchChild: on exit or error, drop the session from the map and, if a turn was in flight, tell the client. emitting only while generating is deliberate. a child that exits between turns is invisible to the user, and an error bubble arriving in a chat nobody is looking at would be noise — dropping the map entry is the whole repair there, because the next turn builds a fresh session and resumes the transcript by id. the stall timer now tears the session down as well. it used to keep it — "it may still be working, and the next turn resumes it" — which is right for a slow agent and wrong for a wedged one: the session stayed broken, so every later turn hung the same way and "send again to continue" was a lie. sessions with background tasks outstanding are still left alone, since a job can be silent far longer than ten minutes and still land its notification. verified live against the real manager: killed the child mid-turn, saw the error surface and the next turn rebuild the session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eb1fd8c31a |
headscale is core, so stop offering to install it
Reported as: the routes are unreachable and there is no dock tile, on a server where officer-headscale is up and healthy. Both symptoms, one cause. Availability is derived ONLY from the sidecar_installs table — `usable` is the rows with status='installed' AND enabled, and every capability mapped to a sidecar outside that set is added to `unavailable`. A CORE sidecar never gets a row there, because core processes are started by pm2 from the generated ecosystem file and never go through the app store. So `headscale` and `vpn` were permanently unavailable, which withheld the dock manifest AND put /headscale into deniedRoutes for the route guard. The design already knew. catalogue.test.ts has a test called "does not offer to install the baseline", and it has been FAILING since headscale was promoted: Expected to not contain: "officer-headscale" docs/secret-store.md predicted it in as many words — "moving headscale into the light profile also removes it from the app store automatically: catalogue.test.ts asserts the catalogue equals full − light, so the test fails until the entry is deleted". The entry was never deleted, and the failing test was never read. So: entry removed, and the tile moved to CORE_DOCK_ITEMS, where the other things that are always present live. DashboardLayout filters every tile through canVisit(), so a member still never sees it — the capability is kind: 'admin'. The entry's existingFields (URL + API key) are not lost. Servers are added from the Servers view inside the app — ServersView.tsx, ServerForm.tsx, useHeadscaleServers.ts — which is where they were really configured; the app-store form was a second place to type the same two values. Verified: catalogue.test.ts 19 pass/1 fail → 20 pass/0 fail, tsgo clean. PRE-EXISTING, not touched: 10 other tests fail on master, 8 of them in src/servers/capabilities. Confirmed identical before and after this change by stashing it and re-running. Worth a look but not this change's business. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d85f089817 |
tsgo is clean
All five errors gone. Both were real resolution bugs rather than dead code that happened to be noisy — the unreachable parts were unreachable for the wrong reason. officerdb's export map gave the wildcard no extension: "./types": "./src/types.ts" explicit entries carry it "./*": "./src/*" the wildcard did not so `officerdb/soulseek/schema` resolved to `src/soulseek/schema`, which is not a file, while `src/soulseek/schema.ts` sat right there. Now "./src/*.ts". Verified every subpath still resolves at RUNTIME with Bun.resolveSync — an exports map is exactly the thing where a typecheck fix can break the running app, and three of the four paths are load-bearing. types.ts inferred EmailAccount* and PushDevice* from `./schema`, the aggregator that drizzle-kit reads — where both tables are commented out because they belong to plugins. But inferring a TYPE has nothing to do with whether the table exists in the live database: these describe rows the plugin's own code passes around, and that code compiles whether or not the plugin is installed. Reading them off the aggregator coupled the two, so commenting a plugin out of schema.ts broke the build of code that was already unreachable. They now come from ./email/schema and ./notify/schema directly — the same move the query modules made when the split landed, and the thing that lets a table leave the aggregator without breaking anything. src/databases/CLAUDE.md already describes this as the rule; types.ts was the one file that had not followed it. Verified: bunx tsgo --noEmit, zero output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
11710f283a |
wire the reverse proxy in as officer-setup 12
~/npm-setup-draft/setup-npm.sh, adapted to the script's own helpers and placed last
— it is the only step that needs Officer already running.
It ignores --unattended, as asked. Every other question in this script has a
defensible default; a domain name, a DNS provider and that provider's API
credentials do not, and the step is opt-in besides. Its prompts read stdin directly
instead of going through confirm()/ask_required(), and they are NAMED APART
(proxy_confirm, proxy_ask) so nobody later consolidates them into the shared helpers
and quietly makes --unattended agree to publishing a public hostname.
The valve is a TTY check rather than the flag: with no terminal there is nobody to
ask, so it skips and prints the manual instructions. A cron-driven install still
works.
Five fixes to the draft:
- `${OFFICER_REPO}/scripts/store-npm-credential.ts` — OFFICER_REPO is a git URL,
not a directory, so that path was https://…/platform.git/scripts/… and the -f
test could never pass. The whole persist-to-platform branch was dead code
falling through to the print. Dropped it: the comment beside it already argued
that not storing this password is a legitimate outcome, since only a human
logging into the admin UI needs it.
- NOT re-runnable, despite saying so. claim_admin returned early on an already
claimed instance without setting NPM_EMAIL/NPM_PASSWORD, and get_token
dereferenced both under set -u. Second run died on an unbound variable. It now
asks for the existing credentials.
- $HOME/dockers → $OFFICER_ROOT/dockers, matching data-path.ts. And the network
is SETUP_DOCKER_NETWORK (`services`), not a second bridge called `officerdev`.
- dig → getent hosts. dnsutils is not installed by this platform, so the check was
command-not-found on a fresh VPS — and an empty answer is indistinguishable from
"not resolving yet", so it waited the full 30 minutes before failing.
- python3 → jq for host-side JSON. jq is already in the core package list; the one
remaining python3 runs INSIDE the NPM container to read its own credential
template, which is the point of reading it from there.
Failure is contained: every function warns and returns non-zero rather than exiting,
so a proxy that does not come up leaves a finished Officer install behind. Retry
with `--only Proxy`.
Verified: bash -n, shellcheck -S warning clean, --list shows Proxy, all seven
external commands present, and the jq filters checked against sample payloads
including the multi-line DNS credential.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b6feca8350 |
owner can reset a member's platform password
The gap at the other end of create-user.ts: the owner could set a password once, at creation, and never again. Losing it meant a hand-written UPDATE with an argon2 hash — the same "edit Postgres by hand" hole that creating accounts used to have. POST /api/users/:id/password, owner-gated, with a button on the row. GENERATED, not typed. The failure this exists for is "I created the account and forgot to copy the password down", and an owner typing a replacement can lose it the same way on the second go. Shown once in a dialog built to be copied — a dialog and not a toast, because a toast that times out while somebody finds a pen loses the one thing they came for. The generator satisfies validatePassword BY CONSTRUCTION rather than by luck: one character drawn from each of the four required classes, the rest from the union, then Fisher-Yates shuffled so the first four positions are not always lower/upper/digit/special. Rejection sampling throughout — `% n` on a byte biases the early characters. Then it runs validatePassword on its own output, so if the rules ever gain a requirement the alphabets do not cover it throws at the one call site instead of minting passwords the login form rejects. Measured: 20,000 generations, all four classes present every time. l, I, 1, O and 0 are absent from the alphabets. This gets read off a screen and typed somewhere else. Signs them out everywhere, as asked: passwordChangedAt = now, and userMiddleware already refuses any token whose iat predates it. That overwrites the null create-user leaves to mean "the owner chose this, not them" — checked, nothing reads that column except the token check. The Linux account is deliberately untouched, and the dialog says so. Members have no Linux password and never had one: ensureOsUser runs useradd with no -p, so it is created locked. Their terminal goes through setpriv, which does not authenticate; their SSH is the key the owner pasted; `su - <member>` as root does not ask. And machine-setup sets PasswordAuthentication no — verified on this host — so one could not be used to log in even if it existed. Setting one would be a new way in, not a repair. The owner is excluded: they have change-password, which asks for the current one, and resetting themselves here would end the session doing it. Verified: transpiles, all lucide icons exist, 20k generator runs. tsgo next. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d95ad3d1b |
install the rootless docker prerequisites with docker itself
Reported from a member's daemon failing: "rootless Docker needs these packages on the host: uidmap". They were being installed — but only inside branch [2] "rootless Docker for <owner>" in section 22. The owner's choice is not the only one that matters: every Developer account the platform provisions gets its own rootless daemon whatever the owner picked for themselves. So on a machine where the owner chose the docker group, the host never got them and every member's daemon failed. Moved into install_docker_engine, so they arrive with Docker rather than with one particular answer to a question about the owner. Three packages, not the one in the error. checkDockerPrerequisites in os-user-docker.ts is the authority and wants uidmap (newuidmap, newgidmap) AND docker-ce-rootless-extras (dockerd-rootless-setuptool.sh); dbus-user-session is what keeps a member's systemd --user alive without a login session. rootless-extras is only RECOMMENDED by docker-ce — installed by default, so usually there by luck, and absent on any host configured with --no-install-recommends. Named explicitly. Reproduced on this machine while checking: rootless-extras present via Recommends, uidmap absent, newuidmap and newgidmap missing. Exactly the reported failure, on a box that chose the docker group. The rootless branch still installs uidmap and dbus-user-session behind its pkg_is_installed guard. Redundant now, kept deliberately: it is the only thing that fixes a machine whose Docker was installed by an older run of this script. Verified: bash -n on both files, and all three packages present in the noble archive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e881015df5 |
members get the same aliases as the owner
The set that just went into the owner's zshrc, mirrored into shell-skel/zshrc, so a
shell on this machine and a shell in a member's account behave alike rather than
diverging by who you happen to be.
Not a copy-paste. Three differences, each because this file has rules the owner's
appended block does not:
- `n` and `vim` are NOT repeated. The Editor block above already sets them, and
only when nvim is actually installed — better than the owner's unguarded pair.
- the eza family keeps its `else` branch rather than only being guarded. A member
with no eza still gets a coloured, grouped listing instead of bare `ls`, and
every alias in the family has a real fallback: lll, lh, ltr and l were added to
that branch too rather than silently existing only when eza does.
- lazydocker is guarded like its neighbours duf and lazygit, per this file's
stated rule that nothing is required beyond zsh itself.
eza needs no separate install for members: they share the host, and machine-setup
puts it in the core package list.
KNOWN, same shape as append_once: seedShellConfig only rewrites .zshrc while it is
still byte-for-byte the template, so a member provisioned before this keeps the old
one. No members exist right now, so nothing to migrate.
Verified: zsh -n, and the fallback branch resolving all ten aliases with eza absent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3862558b92 |
real aliases for the owner, and eza to go with them
The owner's `aliases` block was one line — `alias sz`. Members got a full set from shell-skel/zshrc and the owner got that. Replaced with the eza ls family, the oh-my-zsh standards, and n/vim/sz/ld/httpserver. eza added to all four core package lists. It is in the noble archive at 0.18.2-1, so this is a package rather than a binary fetch, and Core utils is section 5 — well before Shell at 26, so `command -v eza` is already true when the block is written. The eza aliases are GUARDED behind `command -v eza` and the rest are not, and the asymmetry is deliberate: these replace `ls`. Unguarded, a machine where eza failed to install has no working `ls` in any new shell, which reads as a broken machine rather than a missing package. `alias ld=lazydocker` without lazydocker is one command-not-found when you type it — that can degrade honestly. Same principle shell-skel/zshrc already holds to. python3, not python, for httpserver: Ubuntu ships no `python` binary at all, so as given it would have been a command-not-found on every machine this targets. Checked the editor block first — it only exports EDITOR/VISUAL/SUDO_EDITOR, so n and vim do not collide with anything already appended. KNOWN: append_once returns 1 when its marker is already present, so a machine that has already run this keeps the old one-line block and gets none of the above. That is the function working as designed — it exists so a second run does not duplicate its work, and it cannot tell a stale block from one the owner edited. Fix by hand: delete the `# >>> machine-setup: aliases >>>` block from ~/.zshrc and re-run `machine-setup.sh --only Shell`. Verified: bash -n, zsh -n on the block, the eza guard leaving ls unset when eza is absent, and vim resolving through n to nvim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
977e30e782 |
point the setup default at public https
was ssh://git@gitea.pastilhas.dev:2222/officerdev/platform.git now https://gitea.officer.dev/officerdev/platform.git Bigger than a URL swap. The SSH default could not clone on a genuinely fresh machine: the key machine-setup generates there is brand new and Gitea has never seen it, so `--repo` was effectively mandatory on a first install — which is the problem that flag was added for two hours ago. HTTPS needs no key and no agent, so the default now works on a blank box. The old comment explained SSH-because-private and set the condition for changing it: "back to HTTPS when the repository is public". It now is — verified with an anonymous `git ls-remote`, which lists refs with no credentials. Rewrote the comment to record why it moved and what to do if it ever goes private again, since that reasoning is the part worth keeping. clone_repo already runs GIT_TERMINAL_PROMPT=0, so a private repo would fail fast rather than hang on a username prompt. No change needed there. repo.sh is still the only place that sets this, and --repo / OFFICER_REPO still override it. Verified both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
edbe446b34 |
revert "allow port 22 through the docker-user allowlist"
this reverts
|
||
|
|
e36c6bb431 |
allow port 22 through the docker-user allowlist
the DOCKER-USER chain is the only thing gating docker-published ports from the internet — docker writes its own DNAT/FORWARD rules and bypasses ufw, so `ufw allow <port>` has no effect on a published container port. the allowlist permitted only 80 and 443, so a machine provisioned from this template dropped gitea ssh silently. the failure is hard to spot: the port looks open locally and docker ps shows it published, but external clients hang at TCP connect with no refusal. local tests pass because they arrive via lo and match the loopback RETURN before reaching the DROP. comments added so the next person recognises it faster. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d14e11f6c |
--unattended: every question that has a default answers itself
51 yes/no prompts and ~20 free-text ones, of which about six actually need a human.
The line drawn is "a question with a default answers itself; a question with no
possible default still asks", so it stays attended without being a conversation.
Half of it already existed: ASSUME_YES=1 was implemented and honoured by confirm()
in both scripts, returning each question's OWN default — so a "do the thing you
asked for" question goes yes and a genuine extra goes no. --unattended sets it.
The new part is menu_answer(), for the eight numbered menus. It sets the variable
EMPTY rather than passing a default in, because every menu already consumes its
choice as `${CHOICE:-<n>}` — the default lives next to the options it selects
between, which is the right place, and a second copy in the helper could drift from
the one the prompt advertises. Verified all eight consume that way before touching
them. `read <<<''` rather than eval or `declare -g`, which is bash 4.2+ and rules
out the bash 3.2 macOS still ships.
officer-setup's ask_required takes its default too, except where there is none — the
owning account on a machine machine-setup never ran on, where a guess would install
as the wrong user.
STILL ASKS, deliberately: the username; the Tailscale control plane, login server
and auth key; the git identity; and an SSH public key when the account has none.
That last one is a trap I nearly walked into — on a fresh VPS KEY_COUNT==0 forces
ADD_KEY=true with no confirm, and the menu's default is "[1] paste a public key",
which then prompts with no default at all. Auto-answering that menu would hang or
fail, so it is excluded by name. adduser also still asks for a password; that is
the tool, not us.
Two pre-existing bugs fixed on the way: machine-setup's sudo re-exec passed "$@"
after `shift` had emptied it, so --only and --reask stopped existing the moment it
escalated — same bug as officer-setup had. And UNATTENDED/ASSUME_YES are named in
all three sudo lists, because env_reset would otherwise drop the flag at
escalation, which is now the fourth variable lost that way.
Verified: bash -n on five files, --help on all three, and menu_answer + confirm
under the flag showing a menu resolving to its default and a no-default confirm
correctly answering no.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4c33ef7206 |
fix the silent death after installing zsh
Reported from a fresh Hetzner VPS: the run stopped dead right after apt finished installing zsh, printing nothing at all — just install.sh's "machine setup did not finish". install_oh_my_zsh carried a comment saying it "Returns 0 whatever happens". It did not. Under `set -e` a failing command inside a function aborts the SHELL at that line when the function is called plainly; `return 0` underneath is never reached. The command is also `>/dev/null 2>&1`, so the cause was invisible — which is why the transcript just ends. `|| true` is what actually makes it non-fatal. The file already uses that idiom correctly in four other places, so this was a slip rather than a misunderstanding. set_login_shell had the identical bug on `chsh`, which the same run would have hit on the very next question. Fixed differently and deliberately: `|| true` there would let the caller announce a login shell that was never set, so it returns chsh's real status and the CALLER guards the call — which is also what keeps set -e out of it. A refusal now reports, names the manual chsh command, and carries on, because a machine with zsh installed and bash at login still works. Does not explain WHY oh-my-zsh failed on that host — the output was discarded. It will now say "oh-my-zsh did not install" and continue, which is enough to see it. Verified: bash -n on both files, and a reduced case proving broken() exits 1 while fixed() survives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6c13e0d8f6 |
tell the operator they are still root, once
Neither script ever becomes the user it sets the machine up for — a process cannot change its own uid, so both run as root and drop privileges per command instead. Everything Officer owns ends up belonging to that user and every pm2 process runs as them, but the session you are left holding is root's. Two things that fixes are invisible until they bite: group membership is fixed at LOGIN, so the `docker` group just granted is not in the current session, and the shell configuration was written into their home and is not loaded in root's. Both present as "the machine is broken" rather than "log in again". Printed by whichever half runs LAST. The first attempt put it at the end of both, which says it twice on a full install — and the first time it is wrong, because officer-setup is about to run and still needs the root session it tells you to leave. install.sh is the only thing that knows whether anything follows, so it sets OFFICER_SETUP_FOLLOWS and machine-setup stays quiet. Also drops "Pre-flight complete. The remaining sections are not built yet." from the end of officer-setup. All 11 sections exist; that line last made sense when 6 did. Verified: bash -n on all three, the set -e behaviour of `$RUN_OFFICER && export` under --machine-only, and the suppression across all five ways in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cbfe376a42 |
accept the repo URL as an argument
bun setup -- --repo https://github.com/you/platform.git The default is a private Gitea over SSH, which only authenticates on a machine whose key it already knows — so a genuinely fresh server could not clone at all without editing lib/repo.sh or knowing OFFICER_REPO existed. Added to both entry points. install.sh exports it rather than forwarding an argument it does not own; officer-setup.sh sets it before lib/repo.sh is sourced, which reads `${OFFICER_REPO:-<default>}`, so an absent flag still defaults. Two bugs found doing it, both pre-existing: - officer-setup.sh ALREADY had an arg parser, at the top, before the sources. My first attempt added a second one further down that was unreachable — every argument had already been consumed and `*)` would have exited 2 on --repo. Caught because `--help` printed the wrong usage. - both scripts re-execute through sudo passing `"$@"`, which the parse loop had already emptied with `shift`. So `officer-setup.sh --only build` run as a normal user silently became a FULL run the moment it escalated, and `install.sh --officer-only` re-ran the machine half. Nothing said so; the flag just stopped existing. ORIGINAL_ARGS is captured before the loop now. `${ORIGINAL_ARGS[@]+"${ORIGINAL_ARGS[@]}"}` is the set -u safe form — expanding an empty array is an error on bash before 4.4, and this runs on whatever the machine came with. Verified: bash -n on both, --help/--list/--repo/--repo=/unknown-option on both, the set -e behaviour of the guarded export, and that args survive the shift loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
64f3de59fb |
CLAUDE.md caught up on how setup is run
It said `bun setup` runs officer-setup.sh and was "IN PROGRESS, sections 1-6 of 10". It runs scripts/install.sh, and both halves are finished — machine-setup has 28 sections, officer-setup 11. That line is probably why the orchestrator got doubted: the one document you would check to find out how to install says the wrong entry point. Added what install.sh actually is — an orchestrator that runs the two halves and nothing else, either half runnable alone, both re-runnable, run it as yourself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
081920c61f |
drop fastfetch from machine-setup
It stopped a real install. The guard covered the wrong half: a PPA that fails to ADD is caught and skipped, but one that adds cleanly while carrying no package for the running codename gets past that and dies on `pkg_install_now fastfetch`. It was also the only tool in the set with no source but a third-party PPA on Ubuntu 24.04 and older. A neofetch clone is not worth a branch in a script whose whole job is to survive machines nobody has seen. Removed from tools_default, the tool_command mapping and its installer. No shell config invoked it, so nothing is left calling a missing binary. software-properties-common stays in the core apt list for now, with a note: it provides add-apt-repository, the fastfetch PPA was its only caller, and Docker writes its own sources.list.d entry by hand — so it is now dead weight. Left as a separate decision rather than folded into this one. Verified: bash -n on all four scripts, and no live reference remains. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb0286e1a9 |
put rootless docker back, behind the developer gate
Uncommented at all three sites: the call and import in provisionOsAccount, the ~/.local/dockers bind-mount directory in confineUserTree, and the DOCKER_HOST block in the member zshrc. Not restored unconditionally, which is how it was before. It now sits behind the same Developer check as the Postgres role — the gate that prompted disabling it in the first place. So rolePermitsDatabase is renamed rolePermitsDevTools: it gates two things now and a name saying "database" while deciding whether you get containers is the kind of comment that goes stale silently. ~/.local/dockers is created for EVERY account rather than only Developers. It is two install calls, and confineUserTree is the function that places the layout, not the one that knows who is a Developer — so a member promoted later finds it already correct. Postgres role work is untouched and still in place. Verified: transpiles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d6f01862fe |
fix a stack overflow in the wallet copy button
Mine, from the clipboard sweep. format.ts exported copyToClipboard wrapping navigator.clipboard; the sweep replaced the body call with copyToClipboard(value), so the function called itself. CopyField.tsx is the caller, so every copy button in the Wallet was an infinite recursion. Removed the wrapper rather than repointing it — helpers/clipboard already does more (execCommand fallback on an insecure origin) and CopyField imports it directly now. Found by finally running tsgo, in the officerdev-test tree, which has node_modules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bf7e919593 |
only a Developer gets a postgres role
The gate rootless Docker was under, which I did not know about when the Postgres role replaced it. It inherits the rule along with the purpose: a member not trusted to run containers is not thereby trusted to run databases. rolePermitsDatabase() is the single place that rule is written. Admin is deliberately NOT included — administering the platform is not developing on it, and they are separate roles precisely so they can be held separately. Say so if that is wrong. provisionOsAccount now takes the role. Two call sites: create-user passes what the owner picked, provision-linux-route reads it from the row — which makes that route the way a member promoted to Developer gets the database role they did not qualify for when their account was made. The half that makes the gate real is in updateUserRoleHandler. Without it the rule would decide what a Developer gets at creation and never look again, so demoting one would leave their role, their databases and a working password in their ~/.zshenv — a permission surviving its own revocation, with the UI then saying something untrue. Revoke before recording, grant after: the drop runs BEFORE updateUser so a failure aborts with the role unchanged and the whole thing retryable. Promotion runs after and is non-fatal, like every other provisioning step. Demotion keeps their data, same as deletion — databases are reassigned to the platform role, not dropped — and logs where it went, so it does not look deleted. KNOWN, not handled: a demoted member keeps a stale ~/.pgpass and ~/.zshenv naming a role that no longer exists. Harmless (the connection just fails) but untidy, and it means `cat ~/.zshenv` shows a password that no longer works. Verified: transpiles, both call sites updated. Still no tsgo. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ac58e2007 |
drop the postgres role when a member is decommissioned
It was written and never called — dropPostgresRole had zero call sites, so deleting a
member left their role and databases on the cluster.
That is the uid trap in a different id space, and worse. provisionPostgresRole ADOPTS
an existing role, so a role left behind is inherited whole, with its databases, by the
next member who gets the same username. useradd hands out the lowest free uid by
accident; the owner hands out usernames on purpose, so reuse is likelier here, not
less.
Rewritten to PRESERVE rather than destroy. The first version dropped the databases,
which is inconsistent with severMemberTree three files away — that chowns a member's
files to the service user rather than deleting them, and a database is the same kind
of thing. The owner removing an account has not necessarily asked to destroy the work
in it, and dropping is the one choice that cannot be walked back.
Needs the full idiom, per database, and both halves matter:
ALTER DATABASE .. OWNER TO REASSIGN OWNED does not move database ownership
REASSIGN OWNED BY .. TO .. moves tables, schemas, functions
DROP OWNED BY .. removes what is left, which after a reassign is
only the GRANTS — without it DROP ROLE still
refuses, an ACL entry is a dependency too
Both statements act only on the database they are connected to, so it is a connection
per database rather than a loop over `db`.
Ordered after the Linux teardown (which can fail and abort, and must not do so after
something irreversible) and before deleteUser (the row is what remembers there is
anything to clean up).
Verified live: a member with a database, a table, a row and a schema. Naive DROP ROLE
refused with "2 objects in database carol_app". After the sequence: role gone, row
intact, table owned by postgres. Then recreated the same username and confirmed she is
REFUSED from the old database — the login trigger holds because ownership moved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
9f15de3448 |
put the postgres credentials somewhere findable
The password now lands in ~/.zshenv as PGHOST/PGPORT/PGUSER/PGPASSWORD, as well as in
~/.pgpass. Asked for on the grounds that it is an easier place to remember, which is a
real requirement — a credential you cannot find is one you will ask about every time.
.zshenv rather than the .zshrc that was asked for, for two reasons, neither about
secrecy:
- zsh sources .zshrc for INTERACTIVE shells only. Verified: `zsh -c` prints an empty
PGUSER when it is set there, and the right one from .zshenv. A script, a cron entry
or an agent turn running psql would silently get nothing.
- .zshrc is a shared template and seedShellConfig only updates it while it still
matches byte-for-byte, so members keep their edits. A per-member password in it
would strand every member on the template they were created with — a silent
maintenance break rather than a tradeoff.
Both files are still written because they are not redundant: .pgpass is what libpq
reads with no shell involved, so it is the one that works for psycopg, a systemd unit
or a compiled binary. The rotate condition now covers both — either missing means we
cannot reconstruct it from the other, so we regenerate.
Also drops the `head -1 ~/.pgpass` parsing I had put in the zshrc template. It was a
hack, and the values are in the environment before that file is read now anyway.
Verified: zsh sourcing order and both shell types, transpiles. Still no tsgo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3ca7f5f331 |
close the cross-member database leak with a login event trigger
The residue the last commit documented is gone. A database one member creates is now refused to every other member at connection time, so the catalogue metadata never becomes readable in the first place. Both obvious routes are dead ends, measured rather than assumed: datacl is not inherited from the template, and CREATE DATABASE fires no event trigger because it is a global object. What works is a `login` event trigger (PG17+) installed in template1 — event triggers live in a per-database catalogue and CREATE DATABASE copies the template's catalogues, so every member-created database carries it automatically. No naming convention, no sweep, no window. The first version put the function in `public` and a member defeated it in one statement: DROP FUNCTION public.officer_owner_only() CASCADE; -- takes the trigger with it They could not drop or disable the trigger itself, but in PG15+ `public` is owned by pg_database_owner — which resolves to THEM in their own database — and a schema owner may drop objects in it they do not own. Both objects were owned by postgres and it made no difference. Caught because I tried it rather than reasoned about it. Moved into a platform-owned schema with PUBLIC revoked. Every route then refused: DROP EVENT TRIGGER, ALTER .. DISABLE, DROP FUNCTION, DROP SCHEMA, ALTER SCHEMA .. OWNER TO, CREATE OR REPLACE over the top, and PGOPTIONS=-c event_triggers=off (that GUC is superuser-only). A superuser can still set it, which is the recovery path. ensureTemplateIsolation opens its own short-lived connection because a connection cannot change database and CREATE DATABASE requires no other session on the template — a pooled connection to template1 would make every member's `createdb` fail. Verified live, end to end: alice in her own, bob refused, alice refused from bob's, postgres in, owner reads and writes normally, and both members refused CONNECT on `officer`. Probe roles and databases dropped; template1 keeps the trigger, which is the intended state. Still not typechecked — node_modules is empty in this tree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2273941e71 |
give each member a postgres role instead of a docker daemon
A Postgres login role named the same as their Linux account, with CREATEDB, plus a
~/.pgpass so psql never prompts. This is what replaces rootless Docker: the case it
was really there for was "let me run a database to develop against", and a container
per member answered it with a daemon, an image cache and a subuid range each.
Measured on postgres:18-alpine before writing any of it, because three things I
asserted turned out to be wrong:
- a fresh LOGIN role CAN connect to `officer` (datacl NULL = PUBLIC has CONNECT),
but CANNOT read any application table — privileges are owner-only, so
has_table_privilege('users','UPDATE') is false. The capability model was never
reachable from here.
- the `trust` line in pg_hba does not cover host connections: Docker's NAT rewrites
the source, so they fall through to scram-sha-256. Verified with a wrong password.
- revoking from the ROLE does nothing. Privileges are additive and there is no DENY;
only revoking from PUBLIC is a lock.
So ensureAppDatabaseClosed revokes CONNECT+TEMPORARY on the platform's own database
from PUBLIC, and it runs inside provisionPostgresRole rather than in the setup script
— an install set up before today, or restored from a dump, then still cannot end up
with a member who can connect to `officer`.
Password is generated per member, 40 chars, rejection-sampled over an alphanumeric
alphabet: CREATE ROLE is a utility statement and cannot take a bind parameter, so the
safety comes from the alphabet rather than from escaping. Not stored anywhere — it
lives in their 600 ~/.pgpass, the same posture as their SSH key, where we keep only
the public half. Only (re)set when .pgpass is missing, so a reprovision does not
rotate a credential they may have pasted into an app config.
KNOWN RESIDUE, not handled: a database one member creates is metadata-readable by
another. datacl is not inherited from the template (measured: closing template1 and
creating from it still produced NULL), and CREATE DATABASE fires no event trigger, so
nothing can close it at creation. A second member can read table and column NAMES from
the catalogue. They cannot read a row and cannot create anything. Closing it needs a
sweep or a pg_hba rule per member; both are decisions, not details.
Verified live against the running cluster: role creation, refusal on `officer`,
creating and using two databases, and the DROP ... WITH (FORCE) teardown. All probe
roles and databases dropped afterwards.
NOT verified: bunx tsgo, still — node_modules is empty in this tree. The drizzle
return shape was checked by reading PostgresJsQueryResultHKT (RowList<T[]>, extends
Array) rather than by running it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
41cdd1e63a |
stop creating the container bind-mount directory too
~/.local/dockers existed only so a rootless container's inner uid could traverse to a bind source. No daemon, no need. Commented out with its reasoning intact in confineUserTree. .local itself stays — it is not Docker's. The claude installer targets ~/.local/bin, and the comment above it records that directory being created root-owned and blocking the install. Nothing to change in deprovisionOsAccount: it never called os-user-docker.ts. Its only Docker-shaped part is capturing /etc/subuid before userdel, which already treats a missing entry as normal and is worth keeping for any rootless tooling. Verified: transpiles. No test referenced composeDir. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7b32f5bc2f |
don't provision rootless docker for new members
Commented out at the call site in provisionOsAccount, with the import. The code in os-user-docker.ts stays and is now unreferenced — turning it back on is uncommenting two blocks. Only affects NEW accounts. Members provisioned before this keep their daemon, and deprovisionOsAccount still tears one down, which is what those accounts need. Verified: transpiles. tsgo has still not run in this tree (node_modules is empty). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cc1eab7794 |
docs: a triage map of the documentation
42 documents, 13,000 lines, and no way to tell from a filename which describe the system as it is and which record an afternoon in July. This sorts them: living, stale, historical, and two clusters that want consolidating. Says plainly how much was verified — mostly filenames, status lines and greps for what changed today — so it reads as a starting point rather than a verdict. Names the two obvious consolidations without performing them. Nine opencode documents for one migration that has landed (verified: `opencode serve` is in the sidecar, so the plan's "nothing here is implemented" is false), and three mobile-dav documents that are one correspondence. Both need all of them read first, which is not a 4am job. Marks the historical ones as not-to-be-rewritten. claude-sidecar-isolation.md records the officer-claude to officer-agent rename that preceded tonight's rename to officer-claude-code; editing it to match today's code would destroy the reasoning it exists to hold. And notes what most of them share: they were written when the estate was twenty processes and everything was simply present. A core install is six. The fix is usually one line — say whether the thing is core or a plugin — not a rewrite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8826c2847e |
docs: working-on-officer catches up to the code
It described four capability kinds and said terminal, chat and files can never be granted. There are five, and those three moved to `confined` on 2026-08-11 — the kernel enforces the boundary because the account has its own Linux user, and a grant means nothing without one. The layout diagram was missing dockers/ and secrets/, and implied the paths are configured. They are derived from the working directory, which is why the pm2 cwd pin and assertInstallLayout exist. Adds what is switched off as of tonight: six core processes, every plugin router commented out beside its capability claim, the ecosystem files now generated, and .env down to three values with the keys in the secret store. First of a documentation sweep. 42 docs; this one first because it is the operational guide somebody actually reaches for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
236a3a5481 |
docs: first container test pass, and what it found
Ubuntu 24.04, Debian 12, Arch and Fedora 41. OS and package-manager detection
correct on all four; --help works unprivileged; and the install report is written
end to end in a container that had never seen this code — task 1's mechanism
confirmed off the machine it was written on.
Three real findings.
--only does not isolate a step. Running --only "Core utils" still created a user
account, because ask_username and the account creation sit in the preamble above
the step framework, so everything before the first `step` runs every time. It is
defensible and it is not what the flag appears to promise.
.setup-answers travels with a copy of the tree. Correctly gitignored and 0600,
but it lives inside the repository directory, so `cp -r` carries it — a container
that had never run setup came up already knowing the username and created that
account. Nothing secret in it; it is a surprise, which in an installer is the
expensive kind.
adduser leaks its own interactive prompt ("Try again? [y/N]") on the
account-creation path. Harmless here because the run had already stopped, but a
hang on a real unattended install.
Also records what containers cannot reach: no init means systemd, netplan, ufw
and the sshd drop-ins are only verifiable as "wrote the right file"; Docker and
Postgres are untested; macOS is unreachable entirely and everything about it is
reasoned rather than executed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d7d64cd6d6 |
docs: the install-variant tree, for tomorrow's decision
Enumerates the forty-seven prompts the two scripts actually ask and sorts them into branches, consents and values — the distinction that decides what a generated leaf can remove. Four real branches (OS, role, tailnet state, which half), fourteen consents that let a leaf omit a section entirely, and a set of values that must stay prompts because baking them in would mean publishing somebody's hostname. Names the two things that need deciding rather than deciding them: Whether a leaf strips dead code or sets constants and calls the base. They are different artifacts and the plan rests on which one is meant — the first is what makes it auditable by being short, the second is what keeps it maintainable. And the combinatorics: 4 OS x 3 roles x 3 tailnet states is 36 leaves before consents, so the tree cannot be the full product. Publishing a few opinionated leaves keeps the static-file-anyone-can-diff property; generating on demand does not, which is the property per-leaf scripts existed for. Also notes that --unattended and a generated leaf are the same mechanism seen twice, and that install_config's existing behaviour — keep the user's file when there is no tty — is the conservatism every unattended answer needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cd209483e3 |
fix the clipboard over http, and audit the rest
navigator.clipboard is secure-context only, like crypto.randomUUID before it —
over plain http on a tailnet address the object does not exist. Twenty call
sites across eighteen files, in three states that all looked fine in review:
bare calls that threw and killed the handler, optional-chained calls that
silently did nothing, and one carrying the comment "Officer is always behind
HTTPS", which it is not.
The optional-chained ones are the worst of the three: a copy button that reports
success and copies nothing is indistinguishable from a working one until someone
pastes.
helpers/clipboard.ts falls back to document.execCommand('copy') over an
off-screen textarea — deprecated, and it works on any origin because it predates
the secure-context rule. Off-screen rather than hidden, because display:none and
visibility:hidden elements cannot be selected and the copy fails silently.
Reading the clipboard has no equivalent: execCommand('paste') was never permitted
from script. The file browser's paste-a-file path now checks canReadClipboard()
and explains itself instead of throwing.
docs/http-secure-context-audit.md is the full sweep the owner asked for: what was
fixed, what cannot be, and what was checked and found clear. crypto.subtle is
used nowhere in the frontend, which was the one worth confirming since it has no
cheap fallback. Notification's six matches are type names, not the API.
geolocation and navigator.share are already guarded. getUserMedia is in four
files and is being removed — but QrTransfer uses it for the CAMERA, not a
microphone, so "remove audio" does not cover it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
77f1284925 |
install report: first cut, generated by the helpers
Every run writes a timestamped install-report.md recording what was installed, changed, kept, skipped, started and run as root. Written for an adversarial read: the person who just ran a setup script off the internet hands it to an agent of their choosing and asks whether it did anything it should not have. Recorded by the HELPERS rather than by the sections. pkg_install and install_config report themselves, so anything installed or written through them appears whether or not a section author remembered — a section that has to remember is a section that will forget, and an incomplete report is worse than none because it reads as a full account. "Kept" is recorded as carefully as "changed". Leaving somebody's .zshrc alone is the claim a reviewer most wants substantiated, and it is invisible unless stated. Secrets are redacted at the moment of recording rather than filtered at render, so a credential never sits in memory formatted for printing. Verified against a POSTGRES_URL and an api_key/password pair. REPORT_FILE is passed through the sudo re-exec. It was not, first time, and the report silently vanished — the third variable this evening lost to env_reset. Unfinished on purpose, paused mid-task at the owner's request: machine-setup's 26 sections still only report through the two shared helpers, so the sections that change system state directly — systemd units, netplan, ufw, sshd drop-ins — are not yet recorded. That is the half a reviewer would care most about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74b4c7908a |
fix the crash over http: crypto.randomUUID is secure-context only
Chat took the whole page down at the end of every turn with `TypeError: crypto.randomUUID is not a function`. `crypto.randomUUID()` is SECURE-CONTEXT ONLY — over plain http on anything that is not localhost it is not defined at all. Officer is reached at http://officer-dev:9000, which is neither, so all eighteen call sites in the frontend were throwing. The stack shows why it was fatal rather than merely broken: it was called inside a `useState` initialiser, so the throw happened during render and unmounted the tree. The assistant message that has no id yet is created at the end of a turn, which is exactly when it fired. No TLS needed. `crypto.getRandomValues()` carries no such restriction — it is on `Crypto`, not `SubtleCrypto`, and works in an insecure context. helpers/random-id uses randomUUID when it exists and otherwise assembles a v4 from the same CSPRNG: same 122 bits, same version and variant bits. Verified both paths produce a UUID matching the v4 pattern, including with randomUUID deleted. `crypto.subtle` is not used anywhere in the frontend, so randomUUID was the whole of the problem. Audio recording is a different matter — getUserMedia genuinely requires a secure context and cannot be polyfilled. Nine files, eighteen call sites. The vendored hls.mjs is left alone. Two of my own mistakes on the way, both caught by parsing rather than by reading: the rewrite added an import of the helper TO the helper, and inserted another one inside a multi-line import block — the same trap as the officerdb move earlier tonight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1cfc0042f |
rename officer-agent to officer-claude-code
The old name said nothing about what the process runs, and it sits directly beside officer-anthropic-proxy — a different process doing a different job — so "the agent" was ambiguous exactly where it mattered. CLAUDE.md already had to spend a paragraph insisting the two are not the same thing. It spawns `claude`; the name says so now. Only two references were functional: the generator's CORE_PROCESSES and the CORE list in catalogue.test.ts. Everything else was prose or comments. Left alone deliberately: `x-officer-agent-token`. It looks like the same string and is not — it is the agent-handoff HTTP header, naming a per-panel bearer token, unrelated to any pm2 process. Renaming it would have changed a wire protocol to tidy a label. Historical docs keep the old name. claude-sidecar-isolation.md and open-threads-after-per-user-claude.md are dated investigations that record the PREVIOUS rename, from officer-claude to officer-agent, and rewriting them would make that history unreadable. CLAUDE.md notes the change instead, where somebody reading those will be looking. Also worth recording, from the owner: merging this with officer-anthropic-proxy into one sidecar was investigated tonight and rejected. They stay separate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
88c9e96895 |
recover earlier answers on resume
A skipped section leaves its variables unset and later sections read them, so on a resume — which skips every section before the one that stopped — Build announced "PUBLIC_URL <not set — run the Environment section first>" on a machine whose .env had been written twenty minutes earlier. Three variables cross a section boundary: ENV_PORT and ENV_PUBLIC_URL from Environment, POSTGRES_URL from Database. They are read back once near the top, from the file that already holds the answers, rather than per-section — the next variable to cross would otherwise have to remember to do it again. Only fills what is empty, so a value passed on the command line still wins and a section that actually runs still overwrites it. Build had its own late read-back that made the generation work while the screen said it would not. Removed, now that the value is there before anything prints. Verified against the real install at /home/pastilhas/officerdev-test: --only Build now reports the URL that run chose. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e5fe966308 |
count the schema tables instead of printing "?"
The Schema section read ${SCHEMA_TABLES:-?} and nothing ever assigned it, so it
announced "? tables" — which reads as "the count could not be determined" rather
than "nobody set this". schema_table_count existed in lib/build.sh and was never
called.
Counted from schema.ts rather than hardcoded, so the number stays true when a
plugin line is uncommented. Reports 21 against the current tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
60ab531294 |
use the short tailnet name, not the FQDN
http://officer-dev:9000 rather than http://officer-dev.ts.pastilhas.dev:9000. All three forms resolve inside the tailnet — short name, FQDN, raw 100.x — and the short one is what anybody actually types. PUBLIC_URL is read by people too: gen:index bakes it into the page's OpenGraph tags. It depends on the tailnet's search domain, which every Tailscale client sets when MagicDNS is on. A device that has lost it resolves the FQDN instead, and the answer there is to type the longer one rather than to default everybody to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f33e7474b0 |
default PUBLIC_URL to the tailnet address, not localhost
localhost is wrong on a machine with a tailnet, and quietly so: it works from the machine itself and nowhere else, so the mistake surfaces on the first phone rather than during setup. And PUBLIC_URL is not decoration — gen:index bakes it into the page's OpenGraph tags, the task API hands it to scripts as OFFICER_API_HOST, and the CalDAV profile builder refuses without it. The tailnet is where Officer is actually reached, and it is the perimeter the whole security model rests on now that origin checking is gone. Its address is the honest default. Prefers the MagicDNS name over the raw 100.x address — both work, but the name survives a node being re-registered and is something a person can type. Falls back to localhost with no tailnet, which is right rather than merely tolerable: a machine with no private network has no better address to guess. Suggests http://officer-dev.ts.pastilhas.dev:9000 on this machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bedc420d4d |
clone over SSH while the repository is private
An HTTPS clone of a private repo prompts for a username, and under sudo with no interactive terminal that hangs or dies with "could not read Username" — which is what gitea.pastilhas.dev does right now. ssh://git@gitea.pastilhas.dev:2222/officerdev/platform.git instead, temporarily. Back to HTTPS when it is public; nothing else in the script cares which. Tested the path the script actually takes, not just the URL: the clone runs as the OWNER rather than root, and sudo drops SSH_AUTH_SOCK, so there is no agent to answer a passphrase. `sudo -u pastilhas env -u SSH_AUTH_SOCK git ls-remote` returns HEAD, so the key works unaided on this machine. A passphrase-protected key that relies on an agent would not. Still overridable with OFFICER_REPO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |