Written in docs/ rather than COMMS because COMMS is deleted when per-user Claude lands and this describes a project that starts after it — a spec that gets deleted before it is implemented is not a spec. Everything measured on this host on 2026-08-11. The evidence for why it exists: after deleting a member through the UI, the row was gone and the Linux account, a working login shell, a healthy postgres container, 454MB of home and Docker storage, lingering, the runtime directory and the subuid ranges were all still there. Three things the spec carries that reading the code would not have produced. terminate-user is not a barrier. A member's /bin/zsh -i survived it by three hours, and userdel refuses while a process owned by the account is alive, so an implementation that trusts it works on a quiet account and fails on a member who left a shell open. The subuid half. Rootless Docker storage is owned by MAPPED ids, not the member's uid — postgres's data directory belonged to 231141, not 1002. userdel releases the range and a later account can be allocated it, so a check for "nothing owned by the freed uid" passes while hundreds of megabytes are still owned by the freed range. Verification has to scan the range. And a correction to the order I actually used: sever the data from the uid BEFORE releasing it. The teardown ran userdel first and removed data after, which leaves a window where the uid is free while files still carry it. The irreversible step goes last. Also specified: never userdel -r, preserve-by-chown as the default with destroy opt-in, refuse to release the uid if the sever failed, and do not run anything as the member after terminating — creating a session recreates the runtime directory the step just removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
211 lines
9.4 KiB
Markdown
211 lines
9.4 KiB
Markdown
# Deprovisioning a member's Linux account
|
|
|
|
**Status:** specification. Not implemented. Written from a manual teardown performed on the production host on
|
|
2026-08-11, so the ordering constraints below are measured rather than reasoned.
|
|
|
|
**Trigger for implementing:** before the first account that does not belong to the server owner. Not "after
|
|
per-user Claude" — the risk opens when a real person has an account that might later be deleted, which may or
|
|
may not be the same moment.
|
|
|
|
---
|
|
|
|
## What happens today
|
|
|
|
`deleteUserHandler` removes the `users` row and cascades the database. `userdel` never runs. Measured on a
|
|
real member (`green`, uid 1002) immediately after deleting them through the UI, before any cleanup:
|
|
|
|
| | after `deleteUserHandler` |
|
|
|---|---|
|
|
| `users` row | gone |
|
|
| Linux account | alive, uid 1002 |
|
|
| Login shell | `id -u` → 1002 — the deleted account still had a working login |
|
|
| Rootless Docker | daemon running, `postgres` container `Up 2 hours (healthy)` |
|
|
| Home + Docker storage | 454 MB intact |
|
|
| linger, `/run/user/1002`, `/etc/subuid`, `/etc/subgid` | all present |
|
|
|
|
Nothing breaks, which is what makes it dangerous. The account keeps working; only the platform forgets it
|
|
exists.
|
|
|
|
## Why it matters: uid reuse
|
|
|
|
`useradd` allocates the lowest free uid. Delete a member and the uid is free while their files still carry it,
|
|
so the next member created inherits the previous member's home, keys, Docker storage and anything else owned
|
|
by that number. By uid, not by any decision anyone made.
|
|
|
|
This is not hypothetical. `officer_jg` (uid 1001) and `green` (uid 1002) both had login shells pointing at one
|
|
home on this host, and the `users` table had no row for `officer_jg` at all — an earlier account for the same
|
|
email, deleted from the platform, whose Linux side survived. The adoption rule in `ensureOsUser` was never
|
|
bypassed; the account simply outlived the row.
|
|
|
|
**The invariant this function exists to guarantee:**
|
|
|
|
> After deprovisioning, no file anywhere is owned by the freed uid **or by any id in its freed subuid range**,
|
|
> and no passwd entry, linger marker, runtime directory or process refers to it.
|
|
|
|
## The subuid half, which is easy to miss
|
|
|
|
A member's rootless Docker storage is **not** owned by their uid. Container processes map through
|
|
`/etc/subuid`, so the files are owned by ids in that range — on this host, `green` had `231072:65536`, and
|
|
postgres's data directory was owned by `231141` (231072 + 70, postgres's inner uid in the Alpine image).
|
|
|
|
`userdel` releases the subuid range along with the uid, and a later account can be allocated the same range.
|
|
So a check for "nothing is owned by the freed uid" **passes while hundreds of megabytes are still owned by the
|
|
freed subuid range**, and a future member's containers would map onto another member's leftover files.
|
|
|
|
Any verification has to cover the range, not just the uid.
|
|
|
|
---
|
|
|
|
## The sequence
|
|
|
|
Ordering is load-bearing. Each step explains what breaks if it moves.
|
|
|
|
### 1. Disable linger, before stopping anything
|
|
|
|
```
|
|
loginctl disable-linger <user>
|
|
```
|
|
|
|
Lingering keeps a systemd user manager alive with no login session. Terminate first and linger can bring it
|
|
back; disable first and nothing can re-spawn between the two steps.
|
|
|
|
*(The manual teardown ran these in the opposite order and worked. This order is specified because it removes a
|
|
race rather than because the other one failed.)*
|
|
|
|
### 2. Terminate the session, then **verify it actually died**
|
|
|
|
```
|
|
loginctl terminate-user <user>
|
|
```
|
|
|
|
**`terminate-user` is not a barrier.** Measured: a `/bin/zsh -i` owned by the member survived it — three hours
|
|
old, still running after the session was terminated and `/run/user/<uid>` was removed. `userdel` refuses while
|
|
a process owned by the account is alive, so an implementation that trusts `terminate-user` works on a quiet
|
|
account and fails on a member who left a shell open, which is the normal case.
|
|
|
|
Required after terminating:
|
|
|
|
```
|
|
pkill -u <user> # wait, then re-check
|
|
pkill -9 -u <user> # only if the count is still non-zero
|
|
```
|
|
|
|
with a bounded wait between and a final assertion that the process count is zero. **Do not proceed while it is
|
|
not.**
|
|
|
|
### 3. Sever the data from the uid — *before* releasing it
|
|
|
|
Two policies. The platform's default is **preserve**:
|
|
|
|
```
|
|
chown -R <service-user>:<service-group> <member-tree>
|
|
```
|
|
|
|
Destroying a member's data because their account was deleted is a separate decision from removing their
|
|
access, and the platform has no standing to make it silently. Reassigning ownership severs the uid link while
|
|
keeping every byte.
|
|
|
|
**Destroy** is opt-in, for a deliberate rebuild:
|
|
|
|
```
|
|
rm -rf <member-tree>
|
|
```
|
|
|
|
**This step must complete before step 4.** That is the one ordering choice the manual teardown got wrong: it
|
|
released the uid first and removed the data afterwards, which leaves a window where the uid is free while
|
|
files still carry it. If the process dies in that window, the next `useradd` inherits them. Sever first, then
|
|
release — the irreversible step goes last, and only once nothing points at it.
|
|
|
|
### 4. Release the account
|
|
|
|
```
|
|
userdel <user> # NEVER -r
|
|
```
|
|
|
|
`-r` deletes the home, which contradicts the preserve policy and would make the destroy policy depend on a
|
|
flag rather than on an explicit decision. Measured: plain `userdel` removes the passwd, shadow and group
|
|
entries **and** the `/etc/subuid` and `/etc/subgid` ranges.
|
|
|
|
### 5. Verify, and refuse to call it done otherwise
|
|
|
|
See the checklist below. A deprovision that half-succeeded is worse than one that failed cleanly, because the
|
|
uid is free and something still owns files.
|
|
|
|
---
|
|
|
|
## Verification: what "clean" means
|
|
|
|
All of these must hold for the freed uid *and* its freed subuid range:
|
|
|
|
- `getent passwd <user>` → nothing
|
|
- no entry in `/etc/subuid` or `/etc/subgid`
|
|
- `find <DATA_PATH> /home -uid <uid>` → nothing
|
|
- `find <DATA_PATH> /home -uid <subuid-start> -o ... ` over the freed range → nothing
|
|
*(a range scan, not a single id — the mapped ids are spread across it)*
|
|
- `/var/lib/systemd/linger/<user>` absent
|
|
- `/run/user/<uid>` absent
|
|
- no processes owned by the uid
|
|
|
|
Worth extracting as `assertUidFree(uid, subuidRange)` and reusing it as the post-condition of the function and
|
|
as a test.
|
|
|
|
**Trap for the verifier:** do not use `sudo -u <user> …` to check anything after step 2. Creating a session
|
|
starts a user manager and recreates `/run/user/<uid>`, so the check would undo the step it is verifying.
|
|
|
|
---
|
|
|
|
## Behaviour requirements
|
|
|
|
**Idempotent.** Every step tolerates already-done. Re-running on a clean box is a no-op, and re-running after a
|
|
partial failure completes it. The delete handler should be able to call it, fail, and have an operator press
|
|
retry.
|
|
|
|
**Never throws; returns a result.** Same posture as `provisionOsAccount`, `provisionSshAccess` and
|
|
`provisionRootlessDocker`.
|
|
|
|
**But a failed deprovision is not the same as a failed provision.** An account that fails to provision is
|
|
merely unusable. An account that fails to *de*provision may have a freed uid with files still owned by it,
|
|
which is the hazard itself. So:
|
|
|
|
- if step 3 (sever) fails, **do not proceed to step 4**. Leaving the account intact is strictly safer than
|
|
freeing a uid that still owns data.
|
|
- a partial failure must be surfaced loudly, not warned into a log the way a missing Docker install is.
|
|
- the `users` row should not be considered fully deleted while the OS side is in a partial state, or the
|
|
platform forgets about a mess it created.
|
|
|
|
**Must not:** run `userdel -r`; delete data under the preserve policy; touch any account other than the one
|
|
named; run anything as the member after step 2.
|
|
|
|
---
|
|
|
|
## Call sites
|
|
|
|
- `deleteUserHandler` — the reason this exists.
|
|
- An admin-triggered retry, for an account left in a partial state.
|
|
- Worth considering: a startup reconciliation that reports Linux accounts with `os_user` set and no
|
|
corresponding `users` row. That is exactly how `officer_jg` would have been noticed months earlier, and it
|
|
is a report rather than an action — nothing should be deleted automatically at boot.
|
|
|
|
## Open questions for whoever implements it
|
|
|
|
1. **Where does severed data go?** Reassigned in place under the member's old path, or moved somewhere that
|
|
reads as archival? In place is simpler; a `deleted/` location makes it obvious the data is orphaned.
|
|
2. **Is destroy ever exposed in the UI**, or is it always a deliberate operator action outside the platform?
|
|
3. **Should uid allocation avoid reuse entirely** as defence in depth — a monotonic counter rather than
|
|
`useradd`'s lowest-free? The previous discussion concluded severing is better, and it is, because it also
|
|
fixes orphaned files. The two are not exclusive.
|
|
4. **What happens to a member's rootless Docker images and volumes** under preserve? They become unreadable to
|
|
any live account once chowned, which is correct but means the disk stays occupied by data nobody can open.
|
|
|
|
---
|
|
|
|
## Provenance
|
|
|
|
Every measured claim here comes from a real teardown on the production host on 2026-08-11: the surviving
|
|
account and container after a UI delete, the shell that outlived `terminate-user`, `userdel` releasing the
|
|
subuid ranges, and the final verified-clean state (no accounts ≥ 1000 but the owner, no files owned by 1001 or
|
|
1002 anywhere under `DATA_PATH` or `/home`, linger empty, the owner's eight containers untouched).
|
|
|
|
The one thing not measured is the preserve path. The teardown used `rm -rf`, because the data was a disposable
|
|
test database. `chown -R` as a severing mechanism is reasoned, not observed.
|