# Waits: how an agent waits for something without burning context **Status:** draft, 2026-08-13. One mechanism proven (git remote polling, run twice); everything else here is specification and reasoning. Claims are marked **measured** or **reasoned** — do not let that slip. `docs/two-agent-field-report-2026-08-12.md` describes this for one purpose: one agent waiting on another's push. That was where it was discovered, not where it belongs. This file is about the primitive itself, because the same shape answers "wait for CI", "wait for the job to finish", "wait for the container to go healthy", "wait for a reply", and a dozen other things Officer already needs. --- ## The primitive > A **wait** is a harness-owned process that blocks until a condition holds, then exits — and whose exit > re-invokes the agent. Three properties. Drop any one and it breaks in a way that is not visible from watching it run: 1. **The waiting happens below the model.** No inference per tick. The agent is suspended. 2. **The harness owns the process**, so its exit is an event the harness delivers. A process the harness is not tracking can finish perfectly and tell nobody. 3. **It exits when it has something to say.** The exit *is* the notification. A wait that detects and keeps running has informed no one. Everything below follows from those three. --- ## The cost model, which decides everything else This is the part that is easy to get half-right, and half-right is what leads people to build the expensive version. | | cost | |---|---| | a tick while waiting | **nothing** — no model runs | | a thousand ticks | **nothing** | | **each wake** | a full context read, uncached | **Measured** (field report, 2026-08-12): an idle watcher produced 85 bytes over seven minutes with zero inference. **Measured** tonight: two fires, each costing exactly one wake. **Reasoned, and the part usually missed:** a wake re-reads the entire conversation, and conversations only grow. So the cost of a wait is not `duration` — it is `fires × context-at-the-time`. Idle is free forever; the tenth notification in a long session costs several times the first. Worse, waits are the exact workload the prompt cache cannot help. The TTL is about five minutes; anything worth waiting for takes longer than that. **Every wake is an uncached read, by construction.** Two consequences that should drive design: - **Say less on wake.** The output that survives to the wake enters the context permanently. One line per tick over 24h is 2,880 lines that land at once and then stay. - **Prefer many short sessions to one long one.** A wait in a fresh session costs a constant amount per event. The same wait in an immortal session costs monotonically more. This is the single strongest argument for event-driven agents over resident ones. --- ## Prefer blocking over polling. Prefer events over both. The git watcher polls because a git remote can only be *asked*. Most things Officer waits on are not like that, and a poll is the worst of the three options that usually exist. **Tier 1 — block on the kernel.** Zero syscalls while waiting, and detection is immediate rather than average-half-an-interval late. | waiting for | how to block | |---|---| | a file appearing or changing | `inotifywait -q -e close_write,create,moved_to ` | | a process to exit | `tail --pid= -f /dev/null` | | a lock to release | `flock true` | | a line on a pipe or log | `read -r line < ` | | an inbound HTTP callback | a listener that blocks on `accept()` | | whichever of several finishes first | `wait -n` over background pids | **Tier 2 — block on the service.** Some services will hold a connection open and tell you. | waiting for | how | |---|---| | a row to change | Postgres `LISTEN` / `NOTIFY` — the connection blocks, the database pushes | | new mail | IMAP `IDLE` | | a container to change state | `docker events --filter …` (streams, blocks) | | a systemd unit | `systemctl --wait` / journal follow | Officer keeps almost everything in one Postgres. `LISTEN`/`NOTIFY` is therefore the highest-leverage unbuilt piece here: job completion, a new chat message, a status flip, all become blocking waits with no polling anywhere. **Tier 3 — poll, because the source can only be asked.** A git remote, a third-party HTTP API, a health endpoint. Then the rules are: read-only calls (`git ls-remote`, never `git fetch` — a fetch mutates refs under a working tree that may be mid-edit), a `timeout` on every call so a hung network call cannot leave the wait alive and blind, and an interval matched to how fast the thing actually changes. **Never poll in the model.** A scheduled wake-up, a `/loop 30s`, a "check every minute" — these are the same shape wearing the same clothes and they pay a full uncached context read *per tick* to learn nothing. This is the intuitive design and its expense is invisible, which is why it needs saying first. --- ## Make firing mean something The rest of this file is about how to wait cheaply. This section is about the other half, and it is the one that decides whether a fleet of these is affordable. **Most waits find nothing, almost always.** A daily release check answers "no" 360 days a year. A branch watcher wakes on every push, including everyone else's. So the number that matters is not the cost of a useful wake — it is the cost of a useless one, multiplied by how many there will be. The fix is not a cheaper wake. It is to **push the relevance test into the wait condition**, so that firing already implies relevance: - **Do not** wait on "a push", then wake and check whether it carries a `COMMS//NN-*.md`. Wait on a push *that contains one* — a filename test the shell can do with no model at all. - **Do not** wait on "the releases page changed", then wake and read it. Wait on "the version string differs from my cursor" — a string compare. Three tiers, and almost everything should die at the first: | tier | cost | for | |---|---|---| | **shell condition** | zero | anything expressible as a filename, a diff, a version, a status | | **fresh minimal agent** | one small cold read | relevance genuinely needs judgement, but not history | | **escalate with real context** | a full read of a long session | the event has to be interpreted against what came before | A session fork that inherits context but returns nothing to it (Claude Code's `/btw`) is tier two done well. It is still a context read, so it is the fallback when a shell test cannot express relevance — not the default. **Corollary for the platform:** a wait's condition should be part of its declaration, not something the agent evaluates after waking. `wait for: push to touching COMMS/**` is a cheaper and more honest thing to build than `wait for: push` plus an agent that decides. ## The contract a wait must honour Specification. None of this is built yet. **Exit codes are the vocabulary.** ``` 0 fired — the condition holds; payload on stdout 1 timed out — the bounded lifetime elapsed, nothing happened 2 broke — the wait itself failed and is no longer trustworthy ``` `1` and `2` must be distinguishable. "Nothing happened" and "I stopped being able to tell" are opposite facts and a wait that conflates them is worse than no wait, because absence reads as reassurance. **Output is a payload, not a log.** One line on arm so there is a record of what was watched; silence while waiting; a minimal structured payload on fire. Everything printed is permanent context. **A cursor, persisted.** The wait is armed at a position — a SHA, a byte offset, a row id, a timestamp — and that position belongs on disk, not only in the process. Then a re-arm after a restart neither misses events nor re-reports old ones. The git watcher currently holds its base only in memory, which is why a session restart loses the thread. **Bounded lifetime, and the bound is not "forever".** `seq 1 2880` is a runaway backstop, not a policy. A wait that times out should re-arm from its cursor rather than die silently. **Liveness must be externally checkable.** A dead wait and a quiet one are indistinguishable, and that ambiguity has already cost two missed pushes. Cheapest fix: touch a heartbeat file each tick, so `mtime` answers "is it alive" without asking the process. In a UI that shows running processes — as Officer's chat does — the chip itself is the signal, which is a real advantage and should be kept. **Idempotent re-arm, and self-trip protection.** An agent that acts and then wakes on its own action is a loop. Re-arm from the position *after* your own change, and never run two waits on the same condition. --- ## Where this applies in Officer The reason to generalise. Each of these is a place something currently either blocks a turn, gets polled by a human, or is discovered late. | wait | tier | notes | |---|---|---| | a pipeline/script job finishes | 1 or 2 | `data/jobs/.log` is a file — inotify. Or `NOTIFY` on the row | | a download completes | 1 | same, and the progress sentinel already exists | | a container becomes healthy | 2 | `docker events` | | a member logs into `claude` for the first time | 1 | `~/.claude/.credentials.json` appearing — currently polled by `/agent-status` | | new mail arrives | 2 | IMAP IDLE, in the email sidecar | | CI, a deploy, a remote build | 3 | poll, with a timeout | | a push to any repo | 3 today, **event tomorrow** | Gitea is ours: a webhook removes the wait entirely | | a long `db:push` or migration finishes | 1 | process wait | | disk crosses a threshold | 3 | slow-moving; poll infrequently | | **a human replies** | 1 | an approval gate: the agent arms a wait and stops costing anything until answered | That last row is the one worth dwelling on. An agent that needs a decision currently either blocks a session or asks and forgets. A wait makes "stopped, pending your answer" cost nothing while it lasts. --- ## Choosing a lifetime | shape | when | cost | |---|---|---| | **wait inside a live session** | the agent holds context the event needs interpreting against | free while idle, growing per fire | | **wait, then hand off** | context matters up to the fire, not after | one growing session, then reset | | **no wait — event spawns a fresh agent** | the event carries everything needed (a SHA, a job id) | constant per event, forever | The third is the destination for anything recurring. The first is right for tonight's watcher, where the value is that I already know what the commits mean. The rule: **if the payload plus the repo is enough to act on, do not keep a session alive to receive it.** --- ## Failure modes | pattern | what it looks like | |---|---| | **launched outside the harness** | `nohup … &` — runs, detects, exits, and no one is told. Looks perfect | | **model-driven poll** | correct behaviour, full context read per tick | | **detects but does not exit** | prints "found it" into a file nobody reads | | **chatty** | per-tick output, deferred, all landing at once on wake | | **silent death** | session restarts, wait dies, quiet branch and dead watcher look identical | | **self-trip** | agent's own push wakes it, usually because an old wait was never stopped | | **timeout mistaken for quiet** | exit 1 treated as "nothing happened" when it means "I stopped looking" | | **mutating poll** | `git fetch` in a loop, moving refs under a working tree | --- ## Open questions 1. **Is a wait a platform feature or an agent habit?** Officer has a job runner, a Gitea instance and a sidecar pattern. `POST /waits {condition, payload}` returning when it fires is a plausible platform primitive — and would make waits available to capabilities, not only to agents. 2. **What arms a wait for an agent that is not running?** The webhook shape needs the platform to spawn the agent, which is `send-claude-code` plus a trigger. Most of that exists. 3. **Should waits be declarative?** `wait for: file:` / `pg:notify:` / `git:/` — a small vocabulary compiled to the right tier, so nobody hand-writes a poll for something inotify could have blocked on. 4. **How does a wait survive a session restart** without either missing its event or re-firing on an old one? The cursor answers half of it; the other half is who re-arms. 5. **What is the right granularity of notification?** One wake per push, or one wake per batch after a quiet period? Batching trades latency for context, and context is the scarce thing. --- ## Provenance The mechanism, the three properties and the four wrong ways to launch it come from `docs/two-agent-field-report-2026-08-12.md`, which recorded them after they were learned the hard way. What this file adds is the cost model stated as a formula rather than an anecdote, the block-over-poll hierarchy, the exit-code contract, and the argument that the destination is event-spawned short-lived agents rather than resident ones. Nothing in "the contract" or "where this applies" has been implemented. The only thing running today is a tier-3 git poll, which is the good version of the wrong shape.