7b4137ccca4ac528346f6b3cfe68e28ae8de0b84
1298
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7b4137ccca |
a plugin's frontend, generated and rebuilt without a restart
the last piece. installing a plugin now brings its UI with it.
a bundler cannot follow import(runtimeString), so which plugins have a frontend
cannot be answered from the database at render time — it has to be written into
source first. Plugins.gen.tsx is that file: concrete imports, generated from
what is installed, gitignored because it describes THIS machine.
App.tsx keeps its core routes and gains one map. the wildcard hands the whole
subtree to the plugin's own router, which react-router nests natively.
serving moved to build/ in production. the html import is bundled once when the
module graph loads and can never change after, which is precisely why a plugin's
frontend needed a restart; Bun.build measures ~900ms for a 25MB bundle, so an
install can just rebuild. development keeps the html import, because that is
what gives HMR and bun --watch restarts on every source change anyway.
verified end to end against a running server, no restart at any point: install
regenerated the module, rebuilt the bundle (chunk hash changed), and the
plugin's own markup was in it; /example and /example/deeper both served; disable
took it back out of both the module and the bundle and 404'd the api; enable put
it back.
three things worth recording because they were found rather than reasoned:
the shell output is named after the ENTRYPOINT — index.gen.html, not index.html
— and naming: { entry: '[name].[ext]' } does not change it because [name] is
'index.gen'. found as a 503 on the first boot after the switch.
App.tsx already destructured a `plugins`, from useServerSettings — the DEAD
plugin system that scans a directory which does not exist and always returns [].
it silently shadowed the import. the new one is `installedPlugins` and says why.
seedAppRegistry takes plugin panels as an argument rather than importing them:
officerdev is a dependency of the shell, so importing upward would invert that.
756 pass, same 10 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a00116b2c0 |
a plugin's permissions become real capabilities, and survive a restart
two gaps between "a plugin can mount routes" and "a plugin is part of the platform". both closed. FIRST: nothing registered a plugin's declared permissions, so the capability gate could not resolve a plugin path at all. it resolved to null, and null is denied — the owner never noticed because isSuperAdmin short-circuits every check, which is exactly the shape of bug that reaches a member first. the registry is now rebuildable the same way the hono app is: CORE_REGISTRY holds the platform's own, CAPABILITIES is core plus whatever the installed plugins declare, and setPluginCapabilities replaces the plugin half wholesale rather than diffing it. two invariants hold by construction — DEFAULT_ROLE_- CAPABILITIES and CORE_CAPABILITIES derive from CORE_REGISTRY, so a plugin can never put itself in the fresh-install baseline and can never become `core` (every account, undeniable). a key colliding with a core one is refused and logged, because a plugin able to redefine `chat` could widen it. ownerOnly maps to admin, everything else to app. those are the only kinds a manifest can express, and it has no field for a kind at all. capabilities are registered BEFORE routes are mounted: the gate runs ahead of every router, so mounting a route whose permission is not yet registered would 403 the freshly installed plugin until something else happened to refresh. SECOND: nothing mounted plugins at boot. honoServer is built with none at import, because discovery reads disk and database and neither can be awaited at module scope, and every install verb rebuilt — so it tested perfectly and would have silently unmounted everything on the first restart. server.tsx now refreshes before serve(), so there is no window where an installed plugin 404s, and a plugin that will not load is logged rather than fatal. verified: after a restart, [plugins] mounted /example, the row survived, officer-example came back online from its ecosystem entry, capabilityForApiPath resolves /api/example/ping to the example capability at kind=app, and it appears in the owner's grantable list. 756 pass, same 10 pre-existing failures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
62ee0d1e60 |
the beat between steps is feedback, not decoration
the comment called it cosmetic and worth being honest about, which reads like an apology and invites the next reader to delete it as a pointless sleep. the real reason is better. some of this work is genuinely slow — pm2 start measures ~770ms — and some is effectively instant. without a pause the fast steps land in one frame, the log jumps from empty to finished, and you cannot tell 'it worked' from 'nothing happened'. the interval is what makes a step something you saw happen rather than something you found already done. only applied when something is listening, so the json path still runs flat out. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d4606d4a2 |
close the owner's dotfiles once somebody else has a shell
ubuntu's default umask is 002 with user-private groups, so everything the owner
creates lands 775/664. alone on a machine that is harmless. it stops being
harmless the moment a member has a login — and member homes are NESTED inside
the owner's, so the owner's home must stay traversable AND readable (the
ancestor-read requirement bun exposed today) and every dotfile in it is legible
by default.
measured as green before writing this: ~/.pm2/logs (all 12 files, every log the
platform has written), ~/.pm2/dump.pm2, ~/.claude/projects (names every
directory the owner works in), ~/.config, ~/.local, ~/.cache, ~/.npm, ~/.bun,
~/.opencode — all listable. assertSecretsClosed was already holding the line
that matters: .env, .ssh, .zsh_history, .claude.json and the credentials are
denied, and dump.pm2 turned out to hold no secret values because bun loads .env
at runtime rather than through pm2.
so this is the tier below fatal: not tokens, but logs and the shape of the
owner's work.
it runs from PROVISIONING, not from setup, and that is the point. ~/.claude does
not exist until the agent has run once; a chmod at install time finds half the
list missing and silently does nothing — the same failure mode as the ACL mask
earlier today. every member's arrival re-closes whatever appeared since.
two directories are left open on purpose, and both are the same latent bug:
/usr/local/bin/bun -> /home/pastilhas/.bun/bin/bun
/usr/local/bin/gh -> /home/pastilhas/.local/bin/gh
system-wide tools installed into one user's home, so every member resolves them
through it. i found this by closing them and breaking bun and gh for green.
~/.local/share and ~/.local/state ARE closed; only the bin directory is
reachable. the honest fix is installing them outside the owner's home.
verified both directions on this host: green is denied .claude, .pm2, .config,
.local/share, .local/state, .cache — and still has working bun, gh, psql, their
own claude, and their own project tree.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a220342b22 |
stream the install, so it reads like a log instead of a spinner
each verb now reports its steps as they complete, over server-sent events, and the detail panel renders them arriving. POST rather than GET, so EventSource is unavailable — it sends no Authorization header and these routes are owner-only. The client reads the body and parses frames by hand, which is what useCompanionLogStream already does for the headscale container logs; the parser only has to understand what our own endpoint emits. the runner does not know whether anyone is listening. it takes an optional onStep and calls it, so the non-streaming path is the same code with no callback rather than a second implementation of the same four verbs. there is a 220ms beat between steps and it is cosmetic — worth saying out loud. pm2 start genuinely takes ~770ms, measured, but writing a row and rebuilding the router do not, and four lines landing in one frame look like a stall followed by a jump. small enough not to matter to a script, long enough to follow. writing to a closed stream is caught rather than fatal: navigating away mid-install must not abort the install, because by then it is the server's work and half an install is the one outcome the ordering was designed to avoid. verified over the wire with timestamps — frames arrive incrementally, the sidecar step showing its real duration rather than the beat. afterwards pm2 holds the five core apps, plugin_installs is zero, and ecosystem.config.cjs is byte-identical. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
02e049cae8 |
terminal: stop replaying questions, stop opening two sockets, bind the word keys
three separate faults behind "reconnecting gets weird and the keyboard is not natural". ── the replay typed into the shell ── the pty buffer was stored raw and replayed verbatim on every re-attach. anything in it that ASKS the terminal a question — DSR, DA, DECRQM, XTVERSION, XTGETTCAP, the OSC colour queries — got asked again, and xterm answered correctly by writing the reply to its input. the pty receives that as a keystroke nobody typed. stripped on the way IN, since the buffer is the thing that gets replayed and a live client already answered them once when they were legitimately asked. only questions are removed; everything that draws is untouched. where a control shares its final byte with one that draws, the parameter is enumerated rather than wildcarded — CSI 18 t asks the window size, CSI 22 t pushes the title, and stripping the second would change what a replay renders. 36 tests, both directions, because both fail silently. ── two sockets on one session ── handleClose armed a reconnect timer; handleVisibilityChange fired on tab focus whenever readyState was CLOSED — which is exactly what a pending timer leaves. both ran. every keystroke went twice, two replay frames fought over the screen, and only one socket was ever cleaned up because __terminalCleanup is overwritten by whichever connect ran last. connect() is now the single guard, and a stale socket's close no longer speaks for the session. ── the keyboard ── alt-arrow was dead for everyone: xterm.js 5 rewrote it into the ctrl-arrow sequence, xterm.js 6 removed that rewrite and emits the honest ^[[1;3C/D (verified — the string 1;3D does not appear anywhere in the 6.0 bundle). nothing bound it. so it broke on a dependency bump, with no shell config changed. bound in zsh rather than translated in the browser, deliberately: tmux.conf claims M-Left/M-Right for pane switching, and a client-side rewrite would send ^[b to tmux and break it. the real sequence lets tmux handle it inside a session and zsh outside. ctrl-arrow was worse and more embarrassing: it worked for MEMBERS and not for the OWNER. shell-skel/zshrc has had the bindings all along; the owner's .zshrc is assembled in machine-setup and never got them. the owner had a strictly worse shell than the accounts they provision. confirmed with `zsh -i -c bindkey` before and after. also: escape-time 10 in tmux.conf. the 500ms default delays every Alt chord and every Escape, which is most of what "not natural" felt like. applied to this host by hand — setup only runs at install. cmd+arrow is left alone: xterm emits nothing for it, so there is no sequence to bind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2634df7a04 |
the install runner: ecosystem entry, sidecar, row, mounts
closes the hole app-store/pm2.ts has carried since 2026-08-13 — "installing a
plugin has to append its entry here before starting it, that is the plugin
system's job and it is not built". this is that job, and it is why nothing in
the app-store catalogue installs end to end either.
verified against a running server with a sidecar in the tree:
install ecosystem added · sidecar started · recorded · mounted /example
route 200, pm2 online
disable sidecar stopped · unmounted
route 404, pm2 stopped
enable sidecar started · mounted
route 200, pm2 online
uninstall record removed · unmounted · sidecar stopped, deleted, entry gone
route 404, not in pm2, tables untouched
afterwards ecosystem.config.cjs is byte-identical to before, pm2 holds the same
five core apps, and plugin_installs is back to zero rows.
the ecosystem file is edited rather than regenerated: the core entries come from
officer-setup's shell array, so the platform does not know that list and a copy
here would be a second thing to drift. the header above module.exports is
preserved verbatim too — officer-setup's explains that bun auto-loads .env from
the working directory and that data-path derives the install root from its
PARENT, so a wrong cwd relocates the whole install rather than failing. losing
that to a plugin install would be a poor trade.
order is the design. bringing up goes outside-in, taking down goes inside-out,
so the worst intermediate state is "recorded but not running" — visible, and
fixed by a retry — never "running but forgotten", which nothing can see.
each verb returns what it actually did, in order, and the detail panel shows it.
an install that mounted routes but could not start a sidecar is a different
outcome from one that worked, and a spinner that stops cannot say which.
the schema push is still deliberately not wired, and the reason is now in the
code: db:push DROPS tables absent from the schema it is given, so an uninstall
that regenerated the barrel would delete a plugin's data as a side effect of
stopping it. offscale does not need it — headscale_servers already ships in the
platform schema.
720 pass, same 10 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ed195e0904 |
record what got built tonight
the plugin system works end to end for a plugin with an api/router.ts, at runtime, with no restart. what is wired, what is not (schema push, the sidecar's pm2 entry, websocket providers, totality across plugin routes), and what was deliberately left: offscale is not extracted, because moving it deletes working code across ~50 files and that wants someone watching. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
98c400bf33 |
a /plugins screen to install, enable, disable and uninstall
the management surface for what the last commit made possible. two panels either side of a selection that lives in ?selected= and is read by both independently, so neither can be telling the other something stale — rows are real Links, not buttons holding the name in a closure. the detail panel shows what the tree declared (api, schema, sidecar, web), because "installed and nothing happened" is otherwise a mystery, and it names what uninstall does NOT do: neither disable nor uninstall deletes anything the plugin stored, and the screen says so rather than leaving someone to guess whether a button destroys their data. a directory whose manifest will not parse is listed with its error rather than skipped. a malformed plugin that simply does not appear is indistinguishable from one nobody wrote. `outdated` is surfaced as an Update button: the version on disk moving after an install is the normal state on a developer's machine, and it should be visible rather than inferred. the four mutations are written out rather than generated in a loop — useMutation is a hook, and a hook called from inside a helper is a rules-of-hooks violation even when the call order happens to be stable. caught before it shipped. verified against a running server: the spa builds (19.8 MB bundle containing the new screen), / serves 200, /api/plugins answers authenticated and 401s without a token. full suite 719 pass, same 10 pre-existing failures. live server and plugin_installs left untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
282a64a637 |
plugins install, enable, disable and uninstall at runtime
the rest of the mechanism, and it works end to end. against a real server, with
no restart at any point:
/api/example/ping BEFORE install 404
AFTER install 200 {"plugin":"example","ok":true}
AFTER disable 404
AFTER enable 200
AFTER uninstall 404
core route throughout 200
plugin_installs is a new table rather than a reuse of sidecar_installs. that one
belongs to the app store's model, where installing means provisioning a
container or pointing at a remote instance, and it carries mode, compose_dir and
completed_steps to say so. a plugin install has none of those, and reusing it
would have meant a `mode` that lies about every plugin. the two models coexist
until the app store is rebuilt on this one.
the row is needed because presence is not installation: plugins live in the
repository, so a developer writing one has the directory there and has installed
nothing. the tree says what could run, the table says what does.
mount.ts joins the two and rebuilds. an install row whose directory has gone is
dropped from the snapshot rather than reported — but the row is left in the
database, because deleting it there would turn "somebody moved the checkout"
into silent data loss. a plugin whose router will not load stays unmounted and
says why, rather than taking the other nine down with it.
/api/plugins is owner-only in its own right, like /api/app-store, and its
capability guards the MANAGEMENT surface only — a plugin's own permissions come
from its manifest, so a member can hold one at read without being able to
install anything.
plugins/example is the reference implementation and is meant to be read: the
smallest thing that is still a real plugin, with the directory layout as its own
documentation.
not wired yet, and marked [open] in the router: the schema push and the
sidecar's pm2 entry. a plugin with db/schema.ts or sidecar/ needs both before it
works end to end.
full suite: 719 pass, same 10 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0701aba902 |
brotli in the core package set
installed on this host already; verified round-tripping from a member shell. same package name on apt, pacman and dnf. on brew it is there because macOS ships the library but not the CLI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2e6c263751 |
the hono app is built, not assembled once
first piece of the plugin system: the platform can now be rebuilt with a
different set of plugins mounted, at runtime, without restarting.
hono cannot do this the obvious way. its default SmartRouter throws "Can not add
a route since the matcher is already built" the moment a route is added after
serving begins, RegExpRouter does the same, and hono has no api to REMOVE a
route at all — so uninstall was impossible even with TrieRouter, which does
allow adding. tested all four.
so nothing is added to a live app. buildHonoApp(plugins) constructs a fresh one
and honoServer is reassigned, which keeps the default fast router and makes
uninstall expressible. server.tsx now serves it through a closure rather than
the bound honoServer.fetch — that one line is the whole mechanism, since the
bound method would capture whichever app existed at serve() and every rebuild
would silently do nothing.
buildHonoApp is pure: everything it needs arrives as an argument, so an app for
a hypothetical plugin set can be built without a database, a filesystem or a
running server.
alongside it, discovery. plugins live at platform/plugins/<app-name>/ — inside
the repo, because bun links the workspace packages into the root node_modules
and that is what lets a plugin author write `import { useClient } from
'hooks/useClient'` with no publishing and no version negotiation. verified with
Bun.resolveSync from a directory there.
discovery is by convention and presence is the declaration: api/router.ts,
db/schema.ts, sidecar/index.ts, web/Router.tsx. the app name comes from the
directory, so it cannot disagree with where the code sits, and the sidecar
runtime comes from the extension — .mjs is node, .ts is bun — which is already
the rule here and cannot contradict the file it describes.
a broken plugin is collected, never thrown: one unreadable manifest must not
stop the boot or hide the nine beside it that are fine.
verified by booting the refactored server on a spare port — /api answers 200,
protected routes still 401. full suite: 719 pass, and the same 10 failures as
before this change (8 in capabilities, plus cliamp and pty), stash-verified
earlier as pre-existing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b2349b5480 |
install the psql client, matching the server it talks to
there was no psql on this machine. the server runs in a container, so nothing ever put a client on the host, and `docker exec officer-postgres psql` is the owner's tool — a member has their own Postgres role and no access to the owner's Docker socket. the version is derived from PG_IMAGE rather than typed again, because the pairing is load-bearing: pg_dump refuses a server newer than itself, and Ubuntu 24.04 ships client 16 against this 18 server. so the archive package is not merely old, it is unusable for dumps. that is also why this sits beside the server definition instead of in machine-setup's package list — one constant, one place to bump. PGDG added the same way docker.sh adds Docker's: key in its own file, one sources.list.d entry, no add-apt-repository. non-fatal, and the exit status is not the gate — apt can succeed while holding an older client back, so the check is that psql is present AND is the major we asked for. installed by hand on this host already: psql/pg_dump 18.6, verified as green connecting with their own role. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0ae0a5dc58 |
music is where the richer permission model gets designed
offscale is deliberately the simple case — one shared resource, read or write. music is the next extraction and the right place to build the in-plugin visibility system, because it has real per-user data (favourites, playlists, now-playing) on top of a real shared one (a single global library index). so 'whose is this row' has a non-uniform answer there, where offscale's is just 'the owner's'. not designed yet and deliberately not designed here. recorded so the intent survives the gap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
acd51c969c |
the platform grants read or write; richer rules belong to the plugin
the platform's contract is what it already has: a role holds read or write on a capability, stored in role_capabilities and enforced by the gate. anything beyond — who sees whose rows, per-user isolation, record ownership, visibility of any kind — is the plugin author's job, inside the plugin. the platform should not grow machinery for it. a plugin knows what its data means; the platform only knows whether this account got through the door. offscale v1 uses that exactly. one shared resource: read sees what the owner sees, write can change it including deleting a server the owner registered. that is dangerous on purpose — the stored credential is a headscale admin key with no read-only equivalent, so write is close to full control of the tailnet, and that is the owner's call. expected use is read for most roles. two consequences, both inside the plugin. the queries stop scoping by the caller and resolve to the owner's id, leaving the per-user shape in the table unused as the seam if isolation is ever wanted. and two POSTs are really reads — /ssh-test probes and /policy/assist explicitly never saves — so they need readOnlyWrites, or a read-level account finds a broken feature where a withheld permission should be. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4c3682dae6 |
let a member read the two directories above their own home
`bun run` from anywhere inside a member's home died with
error loading current directory
error: An internal error occurred (CouldntReadCurrentDirectory)
before it looked at package.json, bun.lock or .git — all of which were present.
it is not walking up looking for a workspace root. it primes its resolver cache
by walking DOWN from / and opening every component of the cwd for READING:
openat("/home/pastilhas/officerdev/") = 6
openat(".../officerdev/data/") = -1 EACCES
openat(".../officerdev/data/<email>/") = -1 EACCES
those two are 711 — traversable, not listable — which is enough to cd into a
home and not enough for a program that reads its ancestors. `getcwd` succeeds;
the ancestor read is what fails. `O_PATH` would need only `x`, so this is
arguably bun's bug, but it presents as a member's project being mysteriously
unbuildable and nothing here can fix it from the other side.
so DATA_PATH and the account dir now carry a named ACL entry per member. that
gives up the property the old comment named — a member can now `ls` DATA_PATH
and learn the other accounts' email addresses — and keeps everything that
matters: every home is still 700 and owned by its member, every platform
sibling still 700 and owned by the service user. verified as green: the account
list is visible, and email_accounts, another home, the repo .env, the owner's
ssh key, .pgpass and ~/.claude/.credentials.json are all still denied.
the mask is set explicitly to rx alongside the entry. chmod recomputes the mask
from the group bits, which for 711 is --x, so without that the next member's
provisioning would silently clamp every earlier member back to traverse-only.
applied by hand to the one existing account; provisioning covers new ones.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
56bb383c6d |
a plugin installs with no questions unless it says otherwise
offscale needs none of the install fields the catalogue carries — no modes, no existingFields, no configFields, no composeTemplate, no members. nothing to provision, nothing to point at. install is put the code there, push the schema, start the sidecar, swap the routes, and it is available. configuration happens afterwards inside the app, which is already how headscale works: a server is registered at runtime and lands in offscale_servers. so no-questions is the default rather than offscale's special case, and the prompting machinery gets designed against the first extracted plugin that actually needs docker or a remote instance. part of why this was the right pilot — it exercises mounting, schema and sidecar without install being a variable at the same time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
327783532e |
reach a member's transcripts by listing them, not just by reading them
yesterday's fix routed transcript CONTENT through the member's identity and
stopped there, on the strength of a comment in ChatIdentity saying enumeration
never needed it — "their directories are 775 and the platform holds an ACL
entry, so readdirSync and statSync have always worked".
there are no 775 directories on this path. claude creates ~/.claude/projects/
and every project group at mode 700, and a 700 directory clamps the ACL mask to
--- exactly as a 600 file does:
user:officer:rwx #effective:---
mask::---
measured against a real member home:
existsSync(projects) -> true (stat only needs traverse on .claude)
readdirSync(projects) -> EACCES
existsSync(projects/<slug>) -> false
statSync(<transcript>) -> EACCES
existsSync answering false rather than throwing is why this was invisible: every
caller read it as "no such session". one root cause, three reported symptoms —
an empty conversation list, no title on a new chat, and a /chat/<id> deep link
that never restored the conversation. a fourth nobody had reported yet: delete
removed nothing and still answered ok, because unlink needs w+x on the group
directory too.
so enumeration goes through the same door as content, as ONE call rather than a
spawn per entry: listTranscriptsAs runs a single `find` as the member and
returns every transcript with its mtime, which readdir+stat could not do without
dozens of setpriv forks per request and a matching pile of auth.log lines. the
owner keeps a fork-free path — that process already IS the owner. removeAs does
the same for unlink, and readTailAs no longer stats a file it cannot stat.
summarizeTranscript now takes the mtime it is given instead of stat'ing again,
which is both the fix and one less syscall per file.
verified against jg@pertento.ai on this machine: 6 conversations listed with
titles from their first prompts, and a deep link by id alone loads 71 messages.
owner path re-checked and unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
9f903479ce |
websocket routes reload too, so nothing needs a restart
the last gap. six ws providers live in bun's route table rather than hono's, so
the app swap does not reach them — but server.reload({routes}) does, and in both
directions: refused before, connected after install, refused again after
uninstall, with core routes untouched throughout.
so a plugin can own a socket from the start, and no part of an install needs the
process restarted.
still untested: whether connections already open across a reload survive it.
that matters before an install is allowed to interrupt somebody's terminal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
6ab838c77f |
mounting at runtime after all: rebuild the app and swap it
this went round twice — runtime dynamic, then generated-plus-restart on the
belief that hono could not mount after serving, then back once that was actually
tested. the doc keeps the route rather than just the destination.
tested: SmartRouter (hono's default) and RegExpRouter both throw 'Can not add a
route since the matcher is already built'. TrieRouter and PatternRouter accept
it. so runtime adding is possible but costs the fast matcher, and hono has no
remove-route api at all, which uninstall needs.
what solves both is not adding routes but rebuilding: construct a fresh app from
the current plugin set and reassign the variable. the fetch closure reads it per
request, so the reassignment is the swap — atomic, no dropped connections, no
server.reload, and the default SmartRouter is kept. verified 404 before install,
200 after, 404 again after uninstall, with core routes unaffected throughout.
the mechanical cost is one line: server.tsx:322 is '/api/*': honoServer.fetch, a
bound method evaluated once at serve(), and has to become a closure or the swap
does nothing.
websockets stay open: six providers live in bun's route table rather than
hono's, so a plugin owning a socket needs server.reload({routes}), untested.
offscale has none.
and totality stops being a boot check — buildApp() is now the single place
routes are mounted, so it is where the assertion belongs, refusing the swap
rather than refusing the boot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
13437e0e48 |
mounting: generated then restart, reversing the call for runtime dynamic
this reverses the earlier decision for C (mount and unmount at runtime) and says so rather than quietly overwriting it. the requirement behind C was that the platform must not need to know a plugin in advance. that is met either way: what it reads is a generated file listing the installed routers, analogous to Plugins.tsx on the frontend — nothing hardcoded, nothing read from a table at boot, the imports made concrete at install. C would have bought only the absence of a restart. and a restart is close to free here, because sidecars are pm2 peers rather than children — a property that was fought for, since officer used to spawn the agent and pm2's tree-kill took the owner's chat down on every restart. what a restart costs is websockets, which reconnect, and in-memory session records, which claude:list already recovers. the happy consequence is that assertCapabilityTotality stays a boot check instead of becoming a per-mount transaction. it does need to be fed the route table rather than Object.keys(handlers) first — generated mounts widen that gap rather than closing it, so that is a prerequisite and not a tidy-up beside it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7befaf032a |
the manifest holds only what the tree cannot say
lean it to identity facts and human choices: publisher, version, platform range, the four presentation fields, and permissions. everything structural becomes convention, where presence is the declaration — sidecar/, api/router.ts, db/schema.ts, web/Router.tsx, web/panels.ts. appName comes from the directory name, so the id cannot disagree with where the code sits. the dock tile and page title needed no fields at all: the tile is label + icon + color + mountPrefix, and the title is label. writing them again was duplication that could only drift. runtime is the file extension. index.mjs is node, index.ts is bun — implicit, but already the rule here, since officer-pty runs under node for node-pty's abi and everything else is bun. better than a field that can contradict the file. dependsOn is gone; nothing read it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7ebc4d0ccd |
note the totality/route-table drift for later
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1292a5c5ab |
opencode is owner-only until it carries an identity
a turn on the opencode harness ran as the owner, in the owner's home, whoever asked. handleOpenCodeChat resolves its cwd against getOwnerHomeDir(email), which discards the email it is given, and the sidecar runs one shared `opencode serve` as the service user — sendOpenCodeStreaming accepts userId/email/username and forwards none of them. it carried a comment calling itself owner-only; nothing enforced it. reachable by any account with the `chat` grant, which every role holds by default (DEFAULT_ROLE_CAPABILITIES), and isClaudeModel is a startsWith, so a typo'd model string landed there too. the model is client-supplied and never checked against the catalogue. the same gap on the read side: opencode's session store has no per-user scoping at all, so loadOpenCodeSession/delete/rename take an id and no identity, and the list and live routes returned other people's conversations. so: ChatIdentity carries isOwner as its own fact (not inferred from osUser === null, which holds only while resolveHomeDir refuses a member without one), and every opencode door in chat.ts checks it — list, load, live, delete, rename — plus a refusal on the execution path in handleChat. /chat/models hides opencode from non-owners as a courtesy; the socket refuses regardless. a stopgap, not a design. the fix is to thread identity through the opencode sidecar the way spawnClaudeAsMember does, and TODO.md has been saying so. not fixed here, and worth knowing: a member's session list is still empty and /chat/pwds still 500s, because readdirSync on their ~/.claude/projects is EACCES — claude creates it at mode 700, which zeroes the ACL mask. visible in officer-error.log right now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4dc7cd90c2 |
a plugin declares permissions, not capabilities, and has no kind
'capabilities' already means three things in this codebase — the permission
registry, the officer-items store, and the sidecar's routing keys. a fourth
would be one too many, and the field is really just permissions. the name is
free: the old permissions table went in
|
||
|
|
b18601530f |
a manifest for offscale, and the rule it immediately broke
written against the real plugin rather than invented as a field list, on the theory that an abstract one includes what nothing needs and misses what is awkward. that paid off on the first field that mattered. the rule here said a plugin may declare `app` and nothing else. offscale's capability is `admin` — owner only — and should stay that way, so the rule was wrong. the distinction is direction, not privilege: `core` means every account and not deniable, so claiming it grants yourself to everyone; `admin` means owner only, which is a plugin restricting itself. corrected table in the doc. core, execution and confined stay the platform's to assign. `publisher` is the only input to the mount prefix, through one function, so first-party and third-party cannot drift into two code paths. sidecar.runtime is a field because officer-pty needs node for node-pty's abi while everything else is bun — one plugin already needs it, so not speculative. dependsOn is informational and unenforced. code dependencies need no declaration now that a plugin builds inside the workspace, and service dependencies already degrade; this exists so the store can say the console section wants the terminal plugin, rather than the section silently doing nothing. health is marked deferred rather than open, with the reasoning, so it does not get re-raised. migrations likewise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6b2905cc7 |
how the frontend ships, and what plugins may depend on
everything moves to the plugin, frontend included, so federation stopped being
a later problem and had to be answered. it is answered by not needing it: bun
builds the spa into build/ at start and rebuilds it on install, serving from
that directory instead of compiling through the html import. Bun.build is a
runtime call, so an install needs no restart — just a refresh. same origin
throughout, which is why there is no cors work and no rewrite of useClient.
App.tsx keeps core routes and gains one map over `plugins`, each mounted at a
wildcard delegating to the plugin's own router. that list comes from a generated
Plugins.tsx, because a bundler cannot follow import(runtimeString) — the
specifier has to be concrete before the build. the six places the shell
currently hardcodes headscale collapse into that one file, dock included; the
runtime dockItemsFromPlugins path follows rather than competing with it.
presentation moves to build time, permission stays runtime.
dependencies turned out to be two different problems wearing one word. a service
dependency (assist → anthropic-proxy) is a wire call and already degrades. a
code dependency (ConsoleView → TerminalView) is in the bundle and cannot. rule:
may depend, must degrade. service calls go through the api carrying the user's
token, with the user's own permissions, which also deletes the state-file read
claude-proxy uses today to lift the proxy's secret.
no per-plugin permission list: a plugin is part of the app and bounded by the
account calling it. that makes marketplace review a security boundary rather
than a naming one, which is worth knowing rather than discovering.
and the developer environment is a platform checkout — clone it, run dev, build
the plugin inside. the 13 workspace packages resolve by name because bun links
them, so `import { useClient } from 'hooks/useClient'` just works with no
registry and no versioning. dev-time and build-time become the same mechanism.
also writes down the headscale inventory now that it has been read end to end,
including that assist.ts travels unwired as a marker and must not be tidied away
as dead code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7f26f0b4b8 |
offscale is headscale plus the companion, not a rename
the name looks like branding on someone else's project, which is exactly how it gets 'corrected' back later. it is not: offscale is the stock headscale server plus the companion that ships beside it, and the invite flow is the first thing that only exists there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
01a20fff4e |
the invite flow replaced device enrolment; it is not a gap
closes the one open item left by deleting /api/vpn. removing the vpn capability leaves no member-grantable headscale surface and that is correct: the owner mints an invite from the headscale app, the companion turns it into the redirect the phone claims, and the device joins. no per-member permission on officer is involved at any step. recorded as decided rather than open so nobody reintroduces a member-facing enrolment route believing something was lost. nothing was — /api/vpn/enroll was the design the invite flow replaced, and it never had a UI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88a44ec4a7 |
delete /api/vpn
it had no caller. verified three ways before removing: nothing in the mobile monorepo reaches it (enrollVpn's only call site is behind `if (embedded)`, and the one app rendering VpnScreen never passes embedded), nothing in the officer web app references it, and the live database holds no vpn grants. the companion was checked separately by its own author — zero references there either. and it will not come back. offscale is permanently standalone: the thing that gets you to the platform cannot itself need the platform, or a broken tailnet locks you out of both. gone: api/vpn/router.ts, its mount, and the `vpn` capability. the registry keeps a comment where the capability was, because its removal has a cost worth recording — headscale is admin-only, so no member-grantable headscale surface remains, and reintroducing one is a deliberate act rather than an oversight. kept: the sidecar's enroll.ts. its bare POST /_officer/enroll handler is now unreachable, but the file is also the dispatcher for /enroll/invites, which is live and fundamental. the header comment now says so, so nobody deletes it looking for dead code. also records the third component in the doc. two of the three have an "enroll" surface and only one is ours: /api/v1/enroll/* belongs to the companion, is where the phone actually goes, and must not be collapsed into /api/offscale/*. capabilities tests: 17 pass / 8 fail both before and after, stash-verified — the 8 are pre-existing, in totality and path-to-capability, which is precisely the machinery dynamic mounting will rework. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bbc60b34ac |
headscale leaves the baseline, and offscale gets a design doc
first step of extracting headscale into a plugin. CORE_PROCESSES is five now, and the catalogue.test CORE[] mirror follows it — not optional, since that list asserts "the catalogue must not offer a core process" and would have blocked adding offscale to the catalogue later. the local generated ecosystem file lost its entry too, and officer-headscale was stopped and deleted from pm2 by hand. the platform still mounts /api/headscale and still declares the headscale and vpn capabilities, so the feature is present-but-unavailable rather than gone. docs/offscale-plugin.md is a live document for the rest of it. what it records that nothing else does: core is now `officer` alone and everything else is a plugin; routes are /api/<app-name> for ours and /api/p/<creator>/<app-name> for third parties, derived by one function so the two can never become two systems; tables stay in public with an app-name prefix; mounting becomes genuinely dynamic, which retires the "every route stays mounted" premise and relocates assertCapabilityTotality from a boot check to a per-mount transaction. it also records a rejected experiment with evidence — a postgres schema per plugin works completely, including cross-schema FK, idempotent push and DROP SCHEMA CASCADE as uninstall — and the reason not to: drizzle-kit 0.31.8 needs schemaFilter naming every schema, contradicting its own docs, and without it push reports "No changes detected" and creates nothing. a plugin install that reports success and makes no tables is the exact failure shape we have hit three times this week. and /api/vpn is dead: no caller in the mobile monorepo, none in the web app, no grants in the database. offscale is permanently standalone, so it never comes back. the invite flow is unaffected — the phone claims from the Companion, not from officer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d000cedf2f |
let anyone toggle hidden files again
the show/hide dotfiles button was dead for members. they open the browser at their own home, that home is `/`, and the toggle is disabled at `/`. it was never meant to apply to them. |
||
|
|
336e718463 |
read a member's transcripts as the member
a provisioned member could chat normally and had no conversation list. every
refresh came back empty, so nothing could be resumed, and a new chat never
became a saved one.
nothing was wrong with the logic. the turn runs as them, writes its transcript
into their home, and the platform looks in exactly the right place — it just
cannot read what it finds.
confineUserTree grants the service user a named acl entry on every member home,
with d: defaults so anything created later inherits it. that entry is real and
getfacl shows it. it does not survive a file created at mode 600, because posix
derives the acl mask from the group bits of the creation mode:
user:officer:rwx #effective:---
mask::---
claude writes every transcript at exactly that mode — .claude and projects/ are
775, every *.jsonl is 600. so readdir and stat worked, every read raised eacces,
and summarizeTranscript catches eacces and returns null. the sessions did not
fail, they vanished.
no acl can fix this. the creation mode ands the mask down, so d: defaults cannot
raise it, and the only way up is through `other`, which is every account on the
box. a 600 file has two readers: its owner, and root.
so read as the owner of the file, through the same runAsArgv the terminal and
the agent already use. spawnSync keeps it synchronous, which is what lets it
drop into a 914-line synchronous parser reached from five modules instead of
rippling await through all of it.
the privileged surface turned out to be seven call sites, not the file: stat
needs traverse and readdir needs read, and the 775 directories give both. only
content needed identity.
also fixes a 500. parseClaudeTranscript read the file uncaught after an
existsSync that passes, so deep-linking /chat/<id> as a member threw rather than
404ing. it returns null now, like the list path always did.
verified against a throwaway linux account provisioned the same way a member is
— 700 home, named acl, transcript written as them at 600. before: 0 sessions and
loadClaudeSession null. after: the session, its title, its messages, and a
rename that leaves the file owned by the member at 600.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
fe0012635a |
stop leaking the parent claude session into the one we spawn
pm2 inherits the environment of whoever ran pm2 start, so restarting this sidecar from inside a claude code terminal — which is how it is restarted most of the time — bakes that terminal's session into the daemon. right now this process is carrying CLAUDE_CODE_MESSAGING_SOCKET for an unrelated pid that has been alive for an hour and a half. three of these were already stripped; the rest arrived with 2.x and were never added. this is hygiene, not the fix for today's hang — a spawn was verified to succeed with the whole set present — but a child attaching to a stranger's ipc socket is not a failure anyone would recognise from the symptom. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
547662842b |
a dead claude process no longer hangs the chat forever
the sdk runs two independent tasks per session: the consumer loop (`for await (const msg of q)`) and an input pump that writes the queue to the child's stdin. the consumer loop's `finally` is what removes a session from the map — but when the CHILD dies it is the input pump that fails, with `ProcessTransport is not ready for writing`, and that rejection neither ends the consumer loop nor is caught anywhere. so the loop stayed parked on a stream with no writer, `finally` never ran, the session stayed in the map, and spawnClaudeStreaming handed every later turn to the same corpse. each one pushed a message onto a queue nobody drained: no error, no result, no timeout. the client spun forever and the only trace was one unhandledRejection line in the sidecar log. observed on the host today; the only cure was pm2 restart officer-claude-code. a member's turn already supplied its own spawn function because it has to go through setpriv. the owner had none, and therefore no place to observe the child — which is exactly why its death was invisible. so give the owner one too, and wrap both in watchChild: on exit or error, drop the session from the map and, if a turn was in flight, tell the client. emitting only while generating is deliberate. a child that exits between turns is invisible to the user, and an error bubble arriving in a chat nobody is looking at would be noise — dropping the map entry is the whole repair there, because the next turn builds a fresh session and resumes the transcript by id. the stall timer now tears the session down as well. it used to keep it — "it may still be working, and the next turn resumes it" — which is right for a slow agent and wrong for a wedged one: the session stayed broken, so every later turn hung the same way and "send again to continue" was a lie. sessions with background tasks outstanding are still left alone, since a job can be silent far longer than ten minutes and still land its notification. verified live against the real manager: killed the child mid-turn, saw the error surface and the next turn rebuild the session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eb1fd8c31a |
headscale is core, so stop offering to install it
Reported as: the routes are unreachable and there is no dock tile, on a server where officer-headscale is up and healthy. Both symptoms, one cause. Availability is derived ONLY from the sidecar_installs table — `usable` is the rows with status='installed' AND enabled, and every capability mapped to a sidecar outside that set is added to `unavailable`. A CORE sidecar never gets a row there, because core processes are started by pm2 from the generated ecosystem file and never go through the app store. So `headscale` and `vpn` were permanently unavailable, which withheld the dock manifest AND put /headscale into deniedRoutes for the route guard. The design already knew. catalogue.test.ts has a test called "does not offer to install the baseline", and it has been FAILING since headscale was promoted: Expected to not contain: "officer-headscale" docs/secret-store.md predicted it in as many words — "moving headscale into the light profile also removes it from the app store automatically: catalogue.test.ts asserts the catalogue equals full − light, so the test fails until the entry is deleted". The entry was never deleted, and the failing test was never read. So: entry removed, and the tile moved to CORE_DOCK_ITEMS, where the other things that are always present live. DashboardLayout filters every tile through canVisit(), so a member still never sees it — the capability is kind: 'admin'. The entry's existingFields (URL + API key) are not lost. Servers are added from the Servers view inside the app — ServersView.tsx, ServerForm.tsx, useHeadscaleServers.ts — which is where they were really configured; the app-store form was a second place to type the same two values. Verified: catalogue.test.ts 19 pass/1 fail → 20 pass/0 fail, tsgo clean. PRE-EXISTING, not touched: 10 other tests fail on master, 8 of them in src/servers/capabilities. Confirmed identical before and after this change by stashing it and re-running. Worth a look but not this change's business. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d85f089817 |
tsgo is clean
All five errors gone. Both were real resolution bugs rather than dead code that happened to be noisy — the unreachable parts were unreachable for the wrong reason. officerdb's export map gave the wildcard no extension: "./types": "./src/types.ts" explicit entries carry it "./*": "./src/*" the wildcard did not so `officerdb/soulseek/schema` resolved to `src/soulseek/schema`, which is not a file, while `src/soulseek/schema.ts` sat right there. Now "./src/*.ts". Verified every subpath still resolves at RUNTIME with Bun.resolveSync — an exports map is exactly the thing where a typecheck fix can break the running app, and three of the four paths are load-bearing. types.ts inferred EmailAccount* and PushDevice* from `./schema`, the aggregator that drizzle-kit reads — where both tables are commented out because they belong to plugins. But inferring a TYPE has nothing to do with whether the table exists in the live database: these describe rows the plugin's own code passes around, and that code compiles whether or not the plugin is installed. Reading them off the aggregator coupled the two, so commenting a plugin out of schema.ts broke the build of code that was already unreachable. They now come from ./email/schema and ./notify/schema directly — the same move the query modules made when the split landed, and the thing that lets a table leave the aggregator without breaking anything. src/databases/CLAUDE.md already describes this as the rule; types.ts was the one file that had not followed it. Verified: bunx tsgo --noEmit, zero output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
11710f283a |
wire the reverse proxy in as officer-setup 12
~/npm-setup-draft/setup-npm.sh, adapted to the script's own helpers and placed last
— it is the only step that needs Officer already running.
It ignores --unattended, as asked. Every other question in this script has a
defensible default; a domain name, a DNS provider and that provider's API
credentials do not, and the step is opt-in besides. Its prompts read stdin directly
instead of going through confirm()/ask_required(), and they are NAMED APART
(proxy_confirm, proxy_ask) so nobody later consolidates them into the shared helpers
and quietly makes --unattended agree to publishing a public hostname.
The valve is a TTY check rather than the flag: with no terminal there is nobody to
ask, so it skips and prints the manual instructions. A cron-driven install still
works.
Five fixes to the draft:
- `${OFFICER_REPO}/scripts/store-npm-credential.ts` — OFFICER_REPO is a git URL,
not a directory, so that path was https://…/platform.git/scripts/… and the -f
test could never pass. The whole persist-to-platform branch was dead code
falling through to the print. Dropped it: the comment beside it already argued
that not storing this password is a legitimate outcome, since only a human
logging into the admin UI needs it.
- NOT re-runnable, despite saying so. claim_admin returned early on an already
claimed instance without setting NPM_EMAIL/NPM_PASSWORD, and get_token
dereferenced both under set -u. Second run died on an unbound variable. It now
asks for the existing credentials.
- $HOME/dockers → $OFFICER_ROOT/dockers, matching data-path.ts. And the network
is SETUP_DOCKER_NETWORK (`services`), not a second bridge called `officerdev`.
- dig → getent hosts. dnsutils is not installed by this platform, so the check was
command-not-found on a fresh VPS — and an empty answer is indistinguishable from
"not resolving yet", so it waited the full 30 minutes before failing.
- python3 → jq for host-side JSON. jq is already in the core package list; the one
remaining python3 runs INSIDE the NPM container to read its own credential
template, which is the point of reading it from there.
Failure is contained: every function warns and returns non-zero rather than exiting,
so a proxy that does not come up leaves a finished Officer install behind. Retry
with `--only Proxy`.
Verified: bash -n, shellcheck -S warning clean, --list shows Proxy, all seven
external commands present, and the jq filters checked against sample payloads
including the multi-line DNS credential.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b6feca8350 |
owner can reset a member's platform password
The gap at the other end of create-user.ts: the owner could set a password once, at creation, and never again. Losing it meant a hand-written UPDATE with an argon2 hash — the same "edit Postgres by hand" hole that creating accounts used to have. POST /api/users/:id/password, owner-gated, with a button on the row. GENERATED, not typed. The failure this exists for is "I created the account and forgot to copy the password down", and an owner typing a replacement can lose it the same way on the second go. Shown once in a dialog built to be copied — a dialog and not a toast, because a toast that times out while somebody finds a pen loses the one thing they came for. The generator satisfies validatePassword BY CONSTRUCTION rather than by luck: one character drawn from each of the four required classes, the rest from the union, then Fisher-Yates shuffled so the first four positions are not always lower/upper/digit/special. Rejection sampling throughout — `% n` on a byte biases the early characters. Then it runs validatePassword on its own output, so if the rules ever gain a requirement the alphabets do not cover it throws at the one call site instead of minting passwords the login form rejects. Measured: 20,000 generations, all four classes present every time. l, I, 1, O and 0 are absent from the alphabets. This gets read off a screen and typed somewhere else. Signs them out everywhere, as asked: passwordChangedAt = now, and userMiddleware already refuses any token whose iat predates it. That overwrites the null create-user leaves to mean "the owner chose this, not them" — checked, nothing reads that column except the token check. The Linux account is deliberately untouched, and the dialog says so. Members have no Linux password and never had one: ensureOsUser runs useradd with no -p, so it is created locked. Their terminal goes through setpriv, which does not authenticate; their SSH is the key the owner pasted; `su - <member>` as root does not ask. And machine-setup sets PasswordAuthentication no — verified on this host — so one could not be used to log in even if it existed. Setting one would be a new way in, not a repair. The owner is excluded: they have change-password, which asks for the current one, and resetting themselves here would end the session doing it. Verified: transpiles, all lucide icons exist, 20k generator runs. tsgo next. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d95ad3d1b |
install the rootless docker prerequisites with docker itself
Reported from a member's daemon failing: "rootless Docker needs these packages on the host: uidmap". They were being installed — but only inside branch [2] "rootless Docker for <owner>" in section 22. The owner's choice is not the only one that matters: every Developer account the platform provisions gets its own rootless daemon whatever the owner picked for themselves. So on a machine where the owner chose the docker group, the host never got them and every member's daemon failed. Moved into install_docker_engine, so they arrive with Docker rather than with one particular answer to a question about the owner. Three packages, not the one in the error. checkDockerPrerequisites in os-user-docker.ts is the authority and wants uidmap (newuidmap, newgidmap) AND docker-ce-rootless-extras (dockerd-rootless-setuptool.sh); dbus-user-session is what keeps a member's systemd --user alive without a login session. rootless-extras is only RECOMMENDED by docker-ce — installed by default, so usually there by luck, and absent on any host configured with --no-install-recommends. Named explicitly. Reproduced on this machine while checking: rootless-extras present via Recommends, uidmap absent, newuidmap and newgidmap missing. Exactly the reported failure, on a box that chose the docker group. The rootless branch still installs uidmap and dbus-user-session behind its pkg_is_installed guard. Redundant now, kept deliberately: it is the only thing that fixes a machine whose Docker was installed by an older run of this script. Verified: bash -n on both files, and all three packages present in the noble archive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e881015df5 |
members get the same aliases as the owner
The set that just went into the owner's zshrc, mirrored into shell-skel/zshrc, so a
shell on this machine and a shell in a member's account behave alike rather than
diverging by who you happen to be.
Not a copy-paste. Three differences, each because this file has rules the owner's
appended block does not:
- `n` and `vim` are NOT repeated. The Editor block above already sets them, and
only when nvim is actually installed — better than the owner's unguarded pair.
- the eza family keeps its `else` branch rather than only being guarded. A member
with no eza still gets a coloured, grouped listing instead of bare `ls`, and
every alias in the family has a real fallback: lll, lh, ltr and l were added to
that branch too rather than silently existing only when eza does.
- lazydocker is guarded like its neighbours duf and lazygit, per this file's
stated rule that nothing is required beyond zsh itself.
eza needs no separate install for members: they share the host, and machine-setup
puts it in the core package list.
KNOWN, same shape as append_once: seedShellConfig only rewrites .zshrc while it is
still byte-for-byte the template, so a member provisioned before this keeps the old
one. No members exist right now, so nothing to migrate.
Verified: zsh -n, and the fallback branch resolving all ten aliases with eza absent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3862558b92 |
real aliases for the owner, and eza to go with them
The owner's `aliases` block was one line — `alias sz`. Members got a full set from shell-skel/zshrc and the owner got that. Replaced with the eza ls family, the oh-my-zsh standards, and n/vim/sz/ld/httpserver. eza added to all four core package lists. It is in the noble archive at 0.18.2-1, so this is a package rather than a binary fetch, and Core utils is section 5 — well before Shell at 26, so `command -v eza` is already true when the block is written. The eza aliases are GUARDED behind `command -v eza` and the rest are not, and the asymmetry is deliberate: these replace `ls`. Unguarded, a machine where eza failed to install has no working `ls` in any new shell, which reads as a broken machine rather than a missing package. `alias ld=lazydocker` without lazydocker is one command-not-found when you type it — that can degrade honestly. Same principle shell-skel/zshrc already holds to. python3, not python, for httpserver: Ubuntu ships no `python` binary at all, so as given it would have been a command-not-found on every machine this targets. Checked the editor block first — it only exports EDITOR/VISUAL/SUDO_EDITOR, so n and vim do not collide with anything already appended. KNOWN: append_once returns 1 when its marker is already present, so a machine that has already run this keeps the old one-line block and gets none of the above. That is the function working as designed — it exists so a second run does not duplicate its work, and it cannot tell a stale block from one the owner edited. Fix by hand: delete the `# >>> machine-setup: aliases >>>` block from ~/.zshrc and re-run `machine-setup.sh --only Shell`. Verified: bash -n, zsh -n on the block, the eza guard leaving ls unset when eza is absent, and vim resolving through n to nvim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
977e30e782 |
point the setup default at public https
was ssh://git@gitea.pastilhas.dev:2222/officerdev/platform.git now https://gitea.officer.dev/officerdev/platform.git Bigger than a URL swap. The SSH default could not clone on a genuinely fresh machine: the key machine-setup generates there is brand new and Gitea has never seen it, so `--repo` was effectively mandatory on a first install — which is the problem that flag was added for two hours ago. HTTPS needs no key and no agent, so the default now works on a blank box. The old comment explained SSH-because-private and set the condition for changing it: "back to HTTPS when the repository is public". It now is — verified with an anonymous `git ls-remote`, which lists refs with no credentials. Rewrote the comment to record why it moved and what to do if it ever goes private again, since that reasoning is the part worth keeping. clone_repo already runs GIT_TERMINAL_PROMPT=0, so a private repo would fail fast rather than hang on a username prompt. No change needed there. repo.sh is still the only place that sets this, and --repo / OFFICER_REPO still override it. Verified both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
edbe446b34 |
revert "allow port 22 through the docker-user allowlist"
this reverts
|
||
|
|
e36c6bb431 |
allow port 22 through the docker-user allowlist
the DOCKER-USER chain is the only thing gating docker-published ports from the internet — docker writes its own DNAT/FORWARD rules and bypasses ufw, so `ufw allow <port>` has no effect on a published container port. the allowlist permitted only 80 and 443, so a machine provisioned from this template dropped gitea ssh silently. the failure is hard to spot: the port looks open locally and docker ps shows it published, but external clients hang at TCP connect with no refusal. local tests pass because they arrive via lo and match the loopback RETURN before reaching the DROP. comments added so the next person recognises it faster. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d14e11f6c |
--unattended: every question that has a default answers itself
51 yes/no prompts and ~20 free-text ones, of which about six actually need a human.
The line drawn is "a question with a default answers itself; a question with no
possible default still asks", so it stays attended without being a conversation.
Half of it already existed: ASSUME_YES=1 was implemented and honoured by confirm()
in both scripts, returning each question's OWN default — so a "do the thing you
asked for" question goes yes and a genuine extra goes no. --unattended sets it.
The new part is menu_answer(), for the eight numbered menus. It sets the variable
EMPTY rather than passing a default in, because every menu already consumes its
choice as `${CHOICE:-<n>}` — the default lives next to the options it selects
between, which is the right place, and a second copy in the helper could drift from
the one the prompt advertises. Verified all eight consume that way before touching
them. `read <<<''` rather than eval or `declare -g`, which is bash 4.2+ and rules
out the bash 3.2 macOS still ships.
officer-setup's ask_required takes its default too, except where there is none — the
owning account on a machine machine-setup never ran on, where a guess would install
as the wrong user.
STILL ASKS, deliberately: the username; the Tailscale control plane, login server
and auth key; the git identity; and an SSH public key when the account has none.
That last one is a trap I nearly walked into — on a fresh VPS KEY_COUNT==0 forces
ADD_KEY=true with no confirm, and the menu's default is "[1] paste a public key",
which then prompts with no default at all. Auto-answering that menu would hang or
fail, so it is excluded by name. adduser also still asks for a password; that is
the tool, not us.
Two pre-existing bugs fixed on the way: machine-setup's sudo re-exec passed "$@"
after `shift` had emptied it, so --only and --reask stopped existing the moment it
escalated — same bug as officer-setup had. And UNATTENDED/ASSUME_YES are named in
all three sudo lists, because env_reset would otherwise drop the flag at
escalation, which is now the fourth variable lost that way.
Verified: bash -n on five files, --help on all three, and menu_answer + confirm
under the flag showing a menu resolving to its default and a no-default confirm
correctly answering no.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4c33ef7206 |
fix the silent death after installing zsh
Reported from a fresh Hetzner VPS: the run stopped dead right after apt finished installing zsh, printing nothing at all — just install.sh's "machine setup did not finish". install_oh_my_zsh carried a comment saying it "Returns 0 whatever happens". It did not. Under `set -e` a failing command inside a function aborts the SHELL at that line when the function is called plainly; `return 0` underneath is never reached. The command is also `>/dev/null 2>&1`, so the cause was invisible — which is why the transcript just ends. `|| true` is what actually makes it non-fatal. The file already uses that idiom correctly in four other places, so this was a slip rather than a misunderstanding. set_login_shell had the identical bug on `chsh`, which the same run would have hit on the very next question. Fixed differently and deliberately: `|| true` there would let the caller announce a login shell that was never set, so it returns chsh's real status and the CALLER guards the call — which is also what keeps set -e out of it. A refusal now reports, names the manual chsh command, and carries on, because a machine with zsh installed and bash at login still works. Does not explain WHY oh-my-zsh failed on that host — the output was discarded. It will now say "oh-my-zsh did not install" and continue, which is enough to see it. Verified: bash -n on both files, and a reduced case proving broken() exits 1 while fixed() survives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6c13e0d8f6 |
tell the operator they are still root, once
Neither script ever becomes the user it sets the machine up for — a process cannot change its own uid, so both run as root and drop privileges per command instead. Everything Officer owns ends up belonging to that user and every pm2 process runs as them, but the session you are left holding is root's. Two things that fixes are invisible until they bite: group membership is fixed at LOGIN, so the `docker` group just granted is not in the current session, and the shell configuration was written into their home and is not loaded in root's. Both present as "the machine is broken" rather than "log in again". Printed by whichever half runs LAST. The first attempt put it at the end of both, which says it twice on a full install — and the first time it is wrong, because officer-setup is about to run and still needs the root session it tells you to leave. install.sh is the only thing that knows whether anything follows, so it sets OFFICER_SETUP_FOLLOWS and machine-setup stays quiet. Also drops "Pre-flight complete. The remaining sections are not built yet." from the end of officer-setup. All 11 sections exist; that line last made sense when 6 did. Verified: bash -n on all three, the set -e behaviour of `$RUN_OFFICER && export` under --machine-only, and the suppression across all five ways in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cbfe376a42 |
accept the repo URL as an argument
bun setup -- --repo https://github.com/you/platform.git The default is a private Gitea over SSH, which only authenticates on a machine whose key it already knows — so a genuinely fresh server could not clone at all without editing lib/repo.sh or knowing OFFICER_REPO existed. Added to both entry points. install.sh exports it rather than forwarding an argument it does not own; officer-setup.sh sets it before lib/repo.sh is sourced, which reads `${OFFICER_REPO:-<default>}`, so an absent flag still defaults. Two bugs found doing it, both pre-existing: - officer-setup.sh ALREADY had an arg parser, at the top, before the sources. My first attempt added a second one further down that was unreachable — every argument had already been consumed and `*)` would have exited 2 on --repo. Caught because `--help` printed the wrong usage. - both scripts re-execute through sudo passing `"$@"`, which the parse loop had already emptied with `shift`. So `officer-setup.sh --only build` run as a normal user silently became a FULL run the moment it escalated, and `install.sh --officer-only` re-ran the machine half. Nothing said so; the flag just stopped existing. ORIGINAL_ARGS is captured before the loop now. `${ORIGINAL_ARGS[@]+"${ORIGINAL_ARGS[@]}"}` is the set -u safe form — expanding an empty array is an error on bash before 4.4, and this runs on whatever the machine came with. Verified: bash -n on both, --help/--list/--repo/--repo=/unknown-option on both, the set -e behaviour of the guarded export, and that args survive the shift loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
64f3de59fb |
CLAUDE.md caught up on how setup is run
It said `bun setup` runs officer-setup.sh and was "IN PROGRESS, sections 1-6 of 10". It runs scripts/install.sh, and both halves are finished — machine-setup has 28 sections, officer-setup 11. That line is probably why the orchestrator got doubted: the one document you would check to find out how to install says the wrong entry point. Added what install.sh actually is — an orchestrator that runs the two halves and nothing else, either half runnable alone, both re-runnable, run it as yourself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |