diff --git a/TODO.md b/TODO.md index 4720814c..7931a76a 100644 --- a/TODO.md +++ b/TODO.md @@ -14,13 +14,14 @@ the owner's OS user and can never be granted. Indirection there really is accide ## Multi-user -- [ ] **`deprovisionOsAccount` does not exist, and deleting a member leaves their whole Linux side.** - `deleteUserHandler` removes the row and cascades the database; `userdel` never runs. Observed on the - production host on 2026-08-12: a member deleted through the UI kept a working login shell, a running - Postgres container and 454M of data, and their uid was free for the next `useradd` to reissue. Spec in - `docs/deprovision-os-account.md`. The ordering that matters: reap processes explicitly (`terminate-user` - does **not** reap a stale shell, and `userdel` fails while one lives), then `chown -R` to the service - user, then `userdel` — sever before release, and abort if the `chown` fails. +- [ ] **`deprovisionOsAccount` is written but has never run against a real account.** Landed 2026-08-12 in + `os-user-deprovision.ts` and wired into `deleteUserHandler`, which now refuses to delete the row when + the Linux teardown fails — so a failure is retryable instead of forgotten. Only the pure guards + (`guardDeletable`, `guardMemberTree`, `parseSubUidEntry`) have tests; the reap loop, the `chown -R` + sever and `userdel` have been exercised by nobody. `docs/deprovision-os-account.md` → "What is still + unproven" has the five-step validation, and it has to happen on the production host with a throwaway + account that has **a shell left open** and **a container writing as a non-root user** — those are the + two cases the quiet path passes vacuously. - [ ] **The terminal replays terminal QUERIES, which get typed into the shell.** `sidecar/pty/sessions.mjs` replays the whole scrollback on attach; query sequences in the buffer get re-asked, xterm.js answers, diff --git a/docs/deprovision-os-account.md b/docs/deprovision-os-account.md index b9aa0896..4eb31b63 100644 --- a/docs/deprovision-os-account.md +++ b/docs/deprovision-os-account.md @@ -1,11 +1,12 @@ # Deprovisioning a member's Linux account -**Status:** specification. Not implemented. Written from a manual teardown performed on the production host on -2026-08-11, so the ordering constraints below are measured rather than reasoned. +**Status:** implemented 2026-08-12 in `src/servers/os-user-deprovision.ts`, called by `deleteUserHandler`. +**Not yet run against a real account** — see "What is still unproven" at the bottom before trusting it. +Written from a manual teardown performed on the production host on 2026-08-11, so the ordering constraints +below are measured rather than reasoned. -**Trigger for implementing:** before the first account that does not belong to the server owner. Not "after -per-user Claude" — the risk opens when a real person has an account that might later be deleted, which may or -may not be the same moment. +The spec is kept as written rather than rewritten in the past tense: it is the reasoning the implementation +has to keep satisfying, and the failure modes it names are still the ones a change would reintroduce. --- @@ -228,3 +229,49 @@ subuid ranges, and the final verified-clean state (no accounts ≥ 1000 but the The one thing not measured is the preserve path. The teardown used `rm -rf`, because the data was a disposable test database. `chown -R` as a severing mechanism is reasoned, not observed. + +--- + +## What the implementation decided, where the spec left it open + +- **Q1, where severed data goes:** in place. A `deleted/` location is a second thing that can fail between + severing and releasing, and the ordering rule already says nothing may come between them. +- **Q2, destroy in the UI:** no. `policy: 'destroy'` exists and has no call site; `deleteUserHandler` always + preserves. It is also implemented as *chown, then delete as the service user* rather than `sudo rm -rf`, so + a recursive delete as root built from a database column does not exist in the codebase at all. +- **Q3, monotonic uid allocation:** not done. Severing addresses the same hazard and also fixes files orphaned + by any other route; the two are not exclusive and this one is still available later. +- **Q4, a preserved member's Docker images:** unchanged — they stay on disk, owned by the service user, + readable by nobody who wants them. Correct and wasteful, as the spec predicted. + +Two guards were added that the spec did not ask for, both exported and unit-tested: + +- `guardDeletable` — the adoption rule from `ensureOsUser` read backwards. An account is only deletable if its + passwd home is the home the platform would have confined, and its uid is ≥ 1000. Without it, `userdel root` + is one bad `users.osUser` value away, and nothing else in the sequence would object. +- `guardMemberTree` — the member tree must resolve to a direct child of `DATA_PATH`. The email reaches + `join()` from a database row and the result is the argument to a recursive `chown`. + +`chown` runs with `-h`. Measured on this host that `chown -R` already declines to follow a symlink out of the +tree and re-owns the link itself, but the flag states it in the argv rather than resting on traversal +semantics — and re-owning links is what makes `find -uid` (which uses `lstat`) a meaningful check. + +## What is still unproven + +The whole thing has been exercised only by unit tests over the pure guards. Nothing has run against a live +account, and this dev machine is deliberately not the place to try it. + +To validate on the production host, against a throwaway account: + +1. Create a member, open a terminal as them, and **leave a shell running** — that is the case + `terminate-user` does not handle, and the reap loop is the part most likely to be wrong. +2. Give them a container that writes as a non-root user, so the subuid range is genuinely populated: + `postgres:18-alpine` drops to uid 70 and its data directory lands on `subuid_start + 70`. +3. `./scripts/assert-uid-free.sh --capture ` — **before** deleting, or the range is gone. +4. Confirm the range check *fails* on that tree while the account still exists. A checker that has never + failed has not been tested. +5. Delete through the UI, then + `sudo ./scripts/assert-uid-free.sh --check `. + +The delete handler logs that exact command line with the captured values after a successful deprovision, +because after `userdel` nothing else on the machine remembers the range. diff --git a/src/servers/api/users/manage-users.ts b/src/servers/api/users/manage-users.ts index 7d65b998..d3eda6c2 100644 --- a/src/servers/api/users/manage-users.ts +++ b/src/servers/api/users/manage-users.ts @@ -3,6 +3,7 @@ import { getUsers, getUserById, updateUser, deleteUser, USER_ROLES, OWNER_USER_I import type { UserRole } from 'officerdb'; import * as errors from '@@/custom-errors'; import { OS_USERS_ENABLED } from '@@/os-user'; +import { deprovisionOsAccount } from '@@/os-user-deprovision'; // Owner-only management of the other accounts. Everything here is gated by ownerGate in // users-router.ts; these handlers assume the caller is the Super Admin. @@ -104,6 +105,48 @@ export const deleteUserHandler: Handler = async function (ctx) { const existing = await getUserById(id); if (!existing) throw errors.NOT_FOUND('User not found'); + // ── The Linux side goes FIRST, and its failure stops the delete ── + // + // This used to be the whole handler: remove the row, cascade the database, done. Measured on the + // production host on 2026-08-12, immediately after deleting a member through this route: their Linux + // account was still alive with a working login shell, their rootless Docker daemon was still running a + // healthy postgres container, and 454 MB of their data was intact — while the platform had forgotten + // they existed. `useradd` hands out the lowest free uid, so that number was queued up to be reissued to + // the next member along with everything still owned by it. + // + // Ordered this way round because the row is what remembers there is anything to clean up. Delete it + // first and a failed deprovision is unrecoverable through the UI: no row, no osUser, nothing to retry + // against. Keeping the account on failure is also the safer half of the trade — an account that still + // exists is inert, whereas a freed uid whose files still carry it is the hazard itself. + if (OS_USERS_ENABLED && existing.osUser) { + const deprovisioned = await deprovisionOsAccount({ email: existing.email, osUser: existing.osUser }); + if (!deprovisioned.ok) { + // Loud on purpose. A missing Docker install warns into a log; this one names the account, the stage + // and the freed range — which after a failed release is the only surviving record of it. + console.error( + `[users] DEPROVISION FAILED for ${existing.email} at stage '${deprovisioned.stage}': ${deprovisioned.error}` + + (deprovisioned.freed + ? ` — uid ${deprovisioned.freed.uid}, subuid ${deprovisioned.freed.subUid?.start ?? 'none'}` + + ` ${deprovisioned.freed.subUid?.count ?? ''}` + : ''), + ); + throw errors.INTERNAL_SERVER_ERROR( + `Could not remove ${existing.email}'s Linux account: ${deprovisioned.error} ` + + `The platform account was NOT deleted, so this can be retried.`, + ); + } + for (const warning of deprovisioned.warnings) console.warn(`[users] ${existing.email}: ${warning}`); + if (deprovisioned.freed) { + // The audit line. `scripts/assert-uid-free.sh --check` takes exactly these arguments, and after + // `userdel` this log is the only place the freed subuid range still exists. + const { osUser, uid, subUid } = deprovisioned.freed; + console.info( + `[users] deprovisioned ${osUser} — verify with: sudo ./scripts/assert-uid-free.sh --check ` + + `${osUser} ${uid} ${subUid?.start ?? ''} ${subUid?.count ?? ''}`, + ); + } + } + // Deleting a user cascades: passkeys, dashboards, screens, email accounts, playlists, everything keyed // to them. There is no undo, which is why the UI asks first. await deleteUser(id); diff --git a/src/servers/os-user-deprovision.test.ts b/src/servers/os-user-deprovision.test.ts new file mode 100644 index 00000000..abcc07e1 --- /dev/null +++ b/src/servers/os-user-deprovision.test.ts @@ -0,0 +1,102 @@ +import { describe, expect, test } from 'bun:test'; +import { guardDeletable, guardMemberTree, parseSubUidEntry } from './os-user-deprovision'; + +// The guards that stand between a database column and `sudo userdel` / `sudo chown -R`. +// +// Everything else in os-user-deprovision.ts needs a machine with accounts to delete on. These three are pure +// precisely so the dangerous decisions can be tested without one — they are the part where being wrong is not +// recoverable, and the manual teardown that produced the spec is not a thing anyone should repeat to check a +// refactor. + +describe('guardDeletable', () => { + const home = '/data/member@example.com/home'; + + test('accepts an account the platform made', () => { + expect(guardDeletable({ osUser: 'green', uid: 1002, passwdHome: home, expectedHome: home })).toEqual({ ok: true }); + }); + + test('refuses root, and every other system account', () => { + // The scenario: a `users` row whose osUser column says 'root'. Nothing else in the sequence would stop it + // — `loginctl`, `pkill` and `userdel` would all simply do as they were told. + const result = guardDeletable({ osUser: 'root', uid: 0, passwdHome: '/root', expectedHome: home }); + expect(result.ok).toBe(false); + if (!result.ok) expect(result.error).toContain('system account'); + + expect(guardDeletable({ osUser: 'daemon', uid: 1, passwdHome: home, expectedHome: home }).ok).toBe(false); + expect(guardDeletable({ osUser: 'nobody', uid: 999, passwdHome: home, expectedHome: home }).ok).toBe(false); + }); + + test('refuses a same-named account that is not ours', () => { + // A human account that happens to share a member's username. Its home is its own, so it fails the only + // test that proves ownership. This is `ensureOsUser`'s adoption rule read backwards. + const result = guardDeletable({ osUser: 'green', uid: 1002, passwdHome: '/home/green', expectedHome: home }); + expect(result.ok).toBe(false); + if (!result.ok) expect(result.error).toContain('not an account the platform created'); + }); + + test('refuses an account with no passwd home at all', () => { + expect(guardDeletable({ osUser: 'green', uid: 1002, passwdHome: null, expectedHome: home }).ok).toBe(false); + }); + + test('the uid floor is not sufficient on its own', () => { + // Both halves are required: a high uid with a foreign home is still somebody else's account. + expect(guardDeletable({ osUser: 'green', uid: 4000, passwdHome: '/srv/green', expectedHome: home }).ok).toBe(false); + }); +}); + +describe('guardMemberTree', () => { + test('resolves a normal account directory', () => { + const result = guardMemberTree('/data', 'member@example.com'); + expect(result).toEqual({ ok: true, tree: '/data/member@example.com' }); + }); + + test('refuses traversal out of DATA_PATH', () => { + // The path this produces is the argument to `chown -R` and, under destroy, to a recursive delete. The + // email reaches join() from a database row. + expect(guardMemberTree('/data', '../../etc').ok).toBe(false); + expect(guardMemberTree('/data', '..').ok).toBe(false); + expect(guardMemberTree('/data', 'a/../../b').ok).toBe(false); + }); + + test('refuses a nested path even without traversal', () => { + // A member tree is one level down. Anything deeper is not an account directory, whatever it is. + expect(guardMemberTree('/data', 'a/b').ok).toBe(false); + }); + + test('refuses DATA_PATH itself', () => { + expect(guardMemberTree('/data', '').ok).toBe(false); + expect(guardMemberTree('/data', '.').ok).toBe(false); + }); + + test('refuses an empty DATA_PATH', () => { + // Unset DATA_PATH would otherwise make every member tree a path under the process's cwd. + expect(guardMemberTree('', 'member@example.com').ok).toBe(false); + }); +}); + +describe('parseSubUidEntry', () => { + const file = 'pastilhas:100000:65536\ngreen:231072:65536\nofficer_jg:165536:65536\n'; + + test('reads the range for the named account only', () => { + expect(parseSubUidEntry(file, 'green')).toEqual({ start: 231072, count: 65536 }); + expect(parseSubUidEntry(file, 'officer_jg')).toEqual({ start: 165536, count: 65536 }); + }); + + test('a missing entry is null, not a guess', () => { + // Normal: rootless Docker is tolerated when it fails, and accounts predating it have no entry. A guessed + // range would make the audit check a different range than the one that was actually freed. + expect(parseSubUidEntry(file, 'nobody')).toBeNull(); + expect(parseSubUidEntry('', 'green')).toBeNull(); + }); + + test('does not match on a prefix', () => { + // 'green' must not match the 'green2' line — the ranges are different and the wrong one verifies nothing. + expect(parseSubUidEntry('green2:300000:65536\n', 'green')).toBeNull(); + }); + + test('rejects a malformed line rather than producing NaN', () => { + // NaN would flow into `assert-uid-free.sh --check` as an argument and quietly scan nothing. + expect(parseSubUidEntry('green:notanumber:65536\n', 'green')).toBeNull(); + expect(parseSubUidEntry('green:231072:0\n', 'green')).toBeNull(); + }); +}); diff --git a/src/servers/os-user-deprovision.ts b/src/servers/os-user-deprovision.ts new file mode 100644 index 00000000..b4f61ca5 --- /dev/null +++ b/src/servers/os-user-deprovision.ts @@ -0,0 +1,359 @@ +import { rm } from 'node:fs/promises'; +import { existsSync } from 'node:fs'; +import { userInfo } from 'node:os'; +import { join, resolve } from 'node:path'; +import { DATA_PATH } from './data-path'; +import { OS_USERS_ENABLED, lookupOsUser, osUserHome } from './os-user'; + +// Taking a member's Linux account away again. +// +// The specification, and every measured claim in it, is `docs/deprovision-os-account.md` — written from a +// manual teardown on the production host on 2026-08-11. Read it before changing the ORDER of anything here. +// +// ── Why this is not the mirror image of provisioning ── +// +// `provisionOsAccount` may fail freely: an account that did not get its Linux side is merely unusable, and +// the owner presses retry. This one cannot. `useradd` allocates the lowest free uid, so the moment `userdel` +// returns, that number is available to the next member — while the previous member's home, keys and +// container storage may still be owned by it. A half-finished deprovision does not leave a broken account; +// it leaves a **trap**, and the next person to create an account walks into it. +// +// Hence the one rule the sequence exists to enforce: +// +// Sever the data from the uid BEFORE releasing the uid. If severing fails, DO NOT release. +// +// The manual teardown got exactly this backwards — it ran `userdel` first and cleaned up afterwards — and +// the window it opened is the reason this file's steps are ordered rather than grouped. +// +// ── The subuid half ── +// +// A member's rootless Docker files are not owned by their uid. Container processes map through +// `/etc/subuid`, so on the production host `green`'s postgres data directory was owned by 231141 +// (their range start + 70, postgres's uid inside the Alpine image). `userdel` frees the whole range along +// with the uid. So "nothing is owned by uid 1002" can be true while hundreds of megabytes are still owned by +// ids a future member's containers will map onto. +// +// `chown -R` fixes both at once — it rewrites every file it walks regardless of who owned it — which is why +// severing is a single operation rather than one pass per id space. But the range still has to be CAPTURED +// before `userdel`, because that is the last moment anyone can ask what it was. It is returned to the caller +// for exactly that reason; `scripts/assert-uid-free.sh --check` takes it as an argument. + +/** What to do with the member's files. `preserve` is the default and the only one with a call site. */ +export type DeprovisionPolicy = 'preserve' | 'destroy'; + +/** + * The identity of the account as it was before deletion. + * + * Returned even on failure when it was captured, because it is unrecoverable afterwards: `userdel` removes + * the `/etc/subuid` entry, and nothing on the machine then remembers which range the account held. + */ +export type FreedIdentity = { + osUser: string; + uid: number; + /** Null when the account had no `/etc/subuid` entry — rootless Docker was never provisioned for it. */ + subUid: { start: number; count: number } | null; +}; + +export type DeprovisionResult = + | { + ok: true; + /** False when there was nothing to do — no Linux account by that name. Idempotent re-runs land here. */ + removed: boolean; + freed: FreedIdentity | null; + /** Non-fatal residue. The uid is safe to reissue; something cosmetic outlived the teardown. */ + warnings: string[]; + } + | { + ok: false; + error: string; + /** Which step refused. `sever` in particular means the account is intact and MUST stay that way. */ + stage: 'disabled' | 'guard' | 'reap' | 'sever' | 'release'; + freed: FreedIdentity | null; + }; + +async function sudo(args: string[]): Promise<{ ok: boolean; out: string }> { + const proc = Bun.spawn(['sudo', '-n', ...args], { stdout: 'pipe', stderr: 'pipe' }); + const [out, err] = await Promise.all([new Response(proc.stdout).text(), new Response(proc.stderr).text()]); + return { ok: (await proc.exited) === 0, out: `${out}${err}`.trim() }; +} + +async function plain(args: string[]): Promise<{ ok: boolean; out: string }> { + const proc = Bun.spawn(args, { stdout: 'pipe', stderr: 'pipe' }); + const [out, err] = await Promise.all([new Response(proc.stdout).text(), new Response(proc.stderr).text()]); + return { ok: (await proc.exited) === 0, out: `${out}${err}`.trim() }; +} + +/** The account this process runs as — who severed data is reassigned to. */ +const serviceUser = (): string => userInfo().username; + +/** + * The account's `/etc/subuid` range, read before anything is removed. + * + * Absent is normal, not an error: rootless Docker is provisioned per member but tolerated when it fails, and + * an account made before that existed has no entry at all. + */ +export function parseSubUidEntry(fileContent: string, osUser: string): { start: number; count: number } | null { + for (const line of fileContent.split('\n')) { + const [name, start, count] = line.split(':'); + if (name !== osUser) continue; + const parsedStart = Number(start); + const parsedCount = Number(count); + if (!Number.isInteger(parsedStart) || !Number.isInteger(parsedCount) || parsedCount <= 0) return null; + return { start: parsedStart, count: parsedCount }; + } + return null; +} + +/** + * How many processes the uid still owns. + * + * `pgrep` exits 1 on no match, which is the expected answer here — so the exit code says nothing and the + * line count is the measurement. + */ +async function processCount(uid: number): Promise { + const found = await plain(['pgrep', '-u', String(uid)]); + if (!found.out) return 0; + return found.out.split('\n').filter((line) => line.trim()).length; +} + +/** + * Refuse to touch an account that is not ours to touch. + * + * This is `ensureOsUser`'s adoption rule pointed the other way, and it is the most important function in the + * file. Adoption asks "may I take this account over"; if it is wrong, a member gets somebody else's home. + * This asks "may I DELETE this account" — and if it is wrong, `userdel root` is a single sudo away, driven by + * a string in a database row. + * + * The passwd home is the proof of ownership. `ensureOsUser` refuses to adopt an account whose home is not the + * one it was going to confine, so any account the platform created or adopted has `home == osUserHome(email)`. + * Anything else on this machine, by any name, does not — including every system account. + * + * Exported so it can be tested against the cases that matter without a machine to delete users on. + */ +export function guardDeletable(params: { + osUser: string; + uid: number; + passwdHome: string | null; + expectedHome: string; +}): { ok: true } | { ok: false; error: string } { + if (params.uid < 1000) { + return { ok: false, error: `refusing to deprovision '${params.osUser}': uid ${params.uid} is a system account` }; + } + if (params.passwdHome !== params.expectedHome) { + return { + ok: false, + error: + `refusing to deprovision '${params.osUser}': its home is ${params.passwdHome ?? 'unknown'}, not ` + + `${params.expectedHome}. This is not an account the platform created.`, + }; + } + return { ok: true }; +} + +/** + * The member's tree must be a direct child of `DATA_PATH`, spelled with no traversal. + * + * The email arrives from a database row and reaches `join()`, so `../..` in it would walk out of `DATA_PATH` + * — and this path is the argument to a recursive `chown` and, under `destroy`, to a recursive delete. Cheap + * to assert, and the failure it prevents has no upper bound. + */ +export function guardMemberTree( + dataPath: string, + email: string, +): { ok: true; tree: string } | { ok: false; error: string } { + if (!dataPath) return { ok: false, error: 'DATA_PATH is empty; refusing to resolve a member tree' }; + const tree = resolve(join(dataPath, email)); + const parent = resolve(dataPath); + if (tree === parent || !tree.startsWith(`${parent}/`) || tree.slice(parent.length + 1).includes('/')) { + return { ok: false, error: `refusing to operate on ${tree}: not a direct child of ${parent}` }; + } + return { ok: true, tree }; +} + +/** + * Remove a member's Linux account, severing their data from the uid first. + * + * Idempotent. An account that is already gone returns `{ ok: true, removed: false }`, and every step + * tolerates having been done before — so a partial failure can be retried by calling this again rather than + * by finishing it by hand. + * + * Never throws; returns a result. But unlike the provisioning functions, `ok: false` here is not + * cosmetic — see the note at the top of the file, and `stage` on the failure result. + */ +export async function deprovisionOsAccount(params: { + email: string; + osUser: string; + policy?: DeprovisionPolicy; +}): Promise { + const policy: DeprovisionPolicy = params.policy ?? 'preserve'; + const warnings: string[] = []; + + if (!OS_USERS_ENABLED) { + return { + ok: false, + stage: 'disabled', + freed: null, + error: 'per-user Linux accounts are not enabled on this server', + }; + } + + // Absent is success. This is what a re-run after a completed deprovision looks like, and what an account + // that never had a Linux side looks like — neither is a problem to report. + const ids = await lookupOsUser(params.osUser); + if (!ids) return { ok: true, removed: false, freed: null, warnings }; + + const passwd = await plain(['getent', 'passwd', params.osUser]); + const passwdHome = passwd.ok ? (passwd.out.split(':')[5] ?? null) : null; + const expectedHome = osUserHome(params.email); + + const allowed = guardDeletable({ osUser: params.osUser, uid: ids.uid, passwdHome, expectedHome }); + if (!allowed.ok) return { ok: false, stage: 'guard', freed: null, error: allowed.error }; + + const treeGuard = guardMemberTree(DATA_PATH, params.email); + if (!treeGuard.ok) return { ok: false, stage: 'guard', freed: null, error: treeGuard.error }; + + // ── Capture, before anything can destroy the evidence ── + // + // Read now because `userdel` removes the /etc/subuid entry with the account, and after that there is no way + // to ask what range it held. Carried on every return path below, including the failures, so an operator + // auditing a partial teardown still has the numbers `assert-uid-free.sh --check` needs. + const subuidFile = await sudo(['cat', '/etc/subuid']); + const freed: FreedIdentity = { + osUser: params.osUser, + uid: ids.uid, + subUid: subuidFile.ok ? parseSubUidEntry(subuidFile.out, params.osUser) : null, + }; + if (!subuidFile.ok) warnings.push('could not read /etc/subuid; the freed range is unknown and unverifiable'); + + // ── 1. Linger off, before terminating anything ── + // + // Lingering keeps a systemd user manager alive with no login session. Terminate first and linger can bring + // it straight back; disable first and nothing can respawn in the gap. A non-lingering account makes this a + // no-op, which is why its failure is a warning rather than a stop. + const linger = await sudo(['loginctl', 'disable-linger', params.osUser]); + if (!linger.ok) warnings.push(`disable-linger reported: ${linger.out}`); + + // ── 2. Terminate, then PROVE it ── + // + // Measured on the production host: `terminate-user` is not a barrier. A three-hour-old `/bin/zsh -i` owned + // by the member survived it, and survived /run/user/ being removed. `userdel` refuses while any + // process owned by the account lives, so trusting this call works on a quiet account and fails on a member + // who left a shell open — which is the normal case, not the edge one. + await sudo(['loginctl', 'terminate-user', params.osUser]); + + const reaped = await reapProcesses(ids.uid); + if (!reaped.ok) return { ok: false, stage: 'reap', freed, error: reaped.error }; + + // ── 3. Sever the data from the uid, BEFORE releasing it ── + const severed = await severMemberTree({ tree: treeGuard.tree, policy }); + if (!severed.ok) { + // Deliberately returns here rather than continuing. An account left intact is inert; a freed uid whose + // files are still owned by it is the hazard this whole file exists to prevent. + return { ok: false, stage: 'sever', freed, error: severed.error }; + } + + // ── 4. Release ── + // + // Never `-r`. It would delete the home, which contradicts `preserve` and would make `destroy` a consequence + // of a flag rather than of an explicit decision. Plain `userdel` was measured to remove the passwd, shadow + // and group entries and both the /etc/subuid and /etc/subgid ranges. + const del = await sudo(['userdel', params.osUser]); + if (!del.ok) { + return { + ok: false, + stage: 'release', + freed, + error: + `userdel failed for ${params.osUser}: ${del.out}. Their data has already been reassigned to ` + + `${serviceUser()}, so the account is inert — but it still exists. Retry.`, + }; + } + + // Residue: reported, not fatal. /run/user/ is a tmpfs systemd normally reaps with the session; if it + // outlives one, it is empty, cleared at reboot, and no reason to call a correct teardown a failure. + if (existsSync(`/run/user/${ids.uid}`)) { + warnings.push(`/run/user/${ids.uid} still exists; it is tmpfs and clears on reboot`); + } + if (existsSync(`/var/lib/systemd/linger/${params.osUser}`)) { + warnings.push(`linger marker for ${params.osUser} survived disable-linger`); + } + + return { ok: true, removed: true, freed, warnings }; +} + +/** + * Kill everything the uid owns, escalating, and refuse to return success while any of it lives. + * + * The bounded waits are the point. `pkill` returns as soon as the signal is delivered, not when the process + * has gone, so a check that follows it immediately reads the state before the kill took effect. + */ +async function reapProcesses(uid: number): Promise<{ ok: true } | { ok: false; error: string }> { + // Re-asserted here rather than trusted from the caller. This is the one function that signals by uid alone + // — the name is not in the argv — so if the number were ever wrong, it would be wrong about somebody else's + // processes with nothing else to catch it. + if (uid < 1000) return { ok: false, error: `refusing to signal uid ${uid}: not a member account` }; + + if ((await processCount(uid)) === 0) return { ok: true }; + + await sudo(['pkill', '-u', String(uid)]); + for (let attempt = 0; attempt < 10 && (await processCount(uid)) > 0; attempt++) { + await Bun.sleep(200); + } + + if ((await processCount(uid)) > 0) { + await sudo(['pkill', '-9', '-u', String(uid)]); + for (let attempt = 0; attempt < 10 && (await processCount(uid)) > 0; attempt++) { + await Bun.sleep(200); + } + } + + const remaining = await processCount(uid); + if (remaining > 0) { + return { + ok: false, + error: `${remaining} process(es) owned by uid ${uid} survived SIGKILL; not releasing the account`, + }; + } + return { ok: true }; +} + +/** + * Break the link between the member's files and their uid. + * + * `preserve` reassigns; `destroy` reassigns and then deletes. Destroying via reassignment rather than via + * `sudo rm -rf` is deliberate: after the chown the service user owns every byte and can delete the tree + * itself, so a recursive delete as root — with a path built from a database column — never has to exist in + * this codebase. + * + * `-h` because a symlink must be re-owned rather than followed. Measured on this host: `chown -R` already + * behaves this way (a symlink pointing outside the tree was left untouched, and the link's own ownership was + * rewritten), but the flag states it in the argv instead of resting on traversal semantics nobody documented + * — and if it ever did follow, the target would be whatever a member chose to point at. + * + * Re-owning the symlinks is also what makes the audit meaningful: `find -uid` uses `lstat`, so a link left + * owned by the freed uid is a failing check. + */ +async function severMemberTree(params: { + tree: string; + policy: DeprovisionPolicy; +}): Promise<{ ok: true } | { ok: false; error: string }> { + // Nothing to sever is success — an account whose directories were already removed by hand still needs its + // passwd entry released, and refusing here would leave that undone forever. + if (!existsSync(params.tree)) return { ok: true }; + + const who = serviceUser(); + const chowned = await sudo(['chown', '-h', '-R', `${who}:${who}`, params.tree]); + if (!chowned.ok) return { ok: false, error: `could not reassign ${params.tree} to ${who}: ${chowned.out}` }; + + if (params.policy === 'preserve') return { ok: true }; + + try { + await rm(params.tree, { recursive: true, force: true }); + } catch (ex) { + return { + ok: false, + error: `reassigned ${params.tree} but could not delete it: ${ex instanceof Error ? ex.message : String(ex)}`, + }; + } + return { ok: true }; +}