mints a dav app password, renders a configuration profile carrying both the
caldav and carddav payloads, and parks it behind a single-use five-minute token
that safari can fetch without a session.
one profile with both payloads is not a convenience: ios keys accounts by
server+username, so adding carddav separately gets folded into the existing
caldav account and contacts silently never appear.
the profile holds the password in plaintext, so it is held in memory only —
persisting it would falsify createDavAppPassword's "not stored" guarantee.
signing is opt-in via DAV_PROFILE_SIGN_CERT/_KEY/_CHAIN and off by default;
this box has no tls certificate, tls terminates upstream. signed at mint time
reading the cert from disk, so a renewal needs no restart and no hook.
the download route is registered before the /dav mount because hono matches in
registration order and the sync door's /* would otherwise demand http basic.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first iteration gave every user their own Docker container: the user's whole
world lived inside it, and only the super admin could see the real filesystem.
That model is gone, but its scaffolding was still in the tree, and it had already
cost time today — the /usr/local/bin/claude symlink removed a few commits ago
existed only because the bwrap jail ro-bound /usr and could not see the
installer's target.
Deleted:
generate-container-context.ts built the CLAUDE.md and settings.json that told
an agent what its container looked like. Its
only importer was the provisioning removed in
the previous commit, so it had zero consumers.
getUserPiConfigDir pointed into the managed container home. No
consumers anywhere in the tree.
Renamed:
DATA_PATH/<email>/.container-context -> agent-config. It holds one file, the
MCP server config handed to the CLI, and has nothing to do with containers. The
path is written and consumed through a return value, so nothing else reads it;
an old directory left on disk is inert.
Documented rather than removed, because both still have live callers and pulling
them out is a refactor rather than a cleanup:
getHomeDir the container's home. Nothing executes there now — terminals,
chats and task runs all use getOwnerHomeDir — but it survives
as that function's fallback and in pipeline-executor.
toShellUsername named for deriving a Linux username inside the container,
32-char limit and all. Nothing creates a Linux user now; the
value ends up only as a claim in the signed task token, so it
is a sanitiser wearing an old name. Unpicking it means
changing that token and WSData.
Nothing to clean on disk: DATA_PATH/<email> has no home/ tree and no
.container-context/. The docs that still mention any of this are the two marked
"Historical" at the top, which are records of what was true then and should keep
saying so.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
provisionUserEnvironment seeded a managed home under DATA_PATH/<email> — shell
configs from a template dir, plus a generated CLAUDE.md and settings.json — and
bootstrap called it fire-and-forget when the owner account was created. It is
outdated: the managed home is not where anything runs. Task runs, terminals and
chats all execute in getOwnerHomeDir (the real login home, HOME_DIR), not the
getHomeDir tree this was populating.
Removed the function, its four template files, and the call in bootstrap. No
setup script referenced it — checked all of scripts/*.sh — and bootstrap was its
only caller anywhere in the tree.
getHomeDir and toShellUsername stay in data-path.ts: pipeline-executor and
pipeline-job-manager still use them.
This leaves src/servers/generate-container-context.ts with zero consumers, since
provision was the only thing importing it. Left in place rather than deleted in
the same commit — it is 173 lines that build a CLAUDE.md and a settings.json for
an agent, which is plausibly wanted somewhere else, and that is a call for the
owner rather than a side effect of this cleanup.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three files, all additive — nothing on master is modified by this beyond the
CLAUDE_BIN change below, and no file is deleted. The branch predates master by
about 180 commits, but it touches nothing master has touched since, so the
merge is clean.
The part that matters beyond macOS is claude-manager.ts. CLAUDE_BIN was pinned
to /usr/local/bin/claude, which dated from the bwrap-sandboxed architecture:
the jail ro-bound /usr and saw nothing else, so the installer's real target
(~/.local/bin/claude) had to be symlinked somewhere the sandbox could reach.
That sandbox is gone, and the hardcoded path left the sidecar unrunnable on any
host without it. It now resolves an explicit CLAUDE_BIN pin, then PATH, then the
locations Anthropic's installer actually writes to — mirroring how OPENCODE_BIN
is already resolved in the opencode sidecar.
ecosystem.mac.config.cjs is deliberately a trimmed set of processes rather than
a mac port of the full ecosystem. It is also stale in two specific ways, left
as-is here and worth fixing separately: it names officer-claude, which master
renamed to officer-anthropic-proxy, and its officer-pty runs
src/servers/api/terminal/pty-sidecar.mjs, which moved to
src/servers/sidecar/pty/index.mjs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
a DAV client tells you almost nothing when it fails — iOS reports every setup
failure as "Cannot connect using SSL" regardless of cause — so the only way to
know whether a phone ever asked for an address book is to record that it did.
one line per proxied request with the client's own user-agent, and a warning
when a credential is present but wrong. a missing credential is the normal
opening move and stays unlogged; a wrong one is indistinguishable from it at
the client end, where both just say "password incorrect".
no body is ever logged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
two bugs, both of which iOS reports as "Cannot connect using SSL" — a message
about TLS for a problem that has nothing to do with TLS. the certificate was
never involved.
the well-known routes were registered with .get. RFC 6764 §6 has the client
probe the well-known URI with the method it actually intends to use, and iOS
sends PROPFIND — which fell through to the SPA catch-all and 404'd. they
answered correctly in a browser, which is why they looked healthy.
hono's cors() answers every OPTIONS itself as a preflight and never calls the
route beneath it, so OPTIONS /dav/ returned a bare 204 with no DAV header.
OPTIONS is not a preflight to a DAV client — it is how the client asks what the
server can do, and iOS refuses an account whose server does not advertise
calendar-access. cors now skips the DAV paths entirely; a CalDAV client is not
a browser and has no origin to check.
verified from the public internet: PROPFIND on both well-known paths 301s to
/dav/, and an authenticated OPTIONS now returns
`DAV: 1, 2, 3, calendar-access, addressbook, extended-mkcol`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The nightly full rebuilds every album by definition, so "it completed" says
nothing about whether it is still needed. Diff the from-scratch build against
the index that was already live, just before the slot swaps in, and log the
delta: albums added, removed, and entries whose contents changed.
A run reporting NONE did no useful work. A string of those is the evidence for
retiring the nightly; a delta that keeps coming back names the albums to go and
look at instead of guessing.
Entries are compared field by field rather than by JSON.stringify: the optional
fields are spread conditionally, so two entries built by the same code can
serialise with different key order, and stringify would report every album as
changed every night. A cache-format upgrade is labelled, since that rebuilds
everything legitimately and would otherwise read as total rot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
restarting officer under a live turn left the browser connected but permanently
silent. the sidecars are pm2 peers, so the agent kept generating and kept
committing to chat_session_events — what died was officer's binding to it. on
`resume-cursor` the server only re-attached the socket when an in-memory session
still existed, so after a restart there was no session and, critically, no
session-scoped subscription relaying sidecar events to the client. the client got
its durable replay and then nothing, which reads exactly like the agent stopping.
adopt the session instead: recreate the record and re-open the subscription
without spawning anything. `_claudeKill` has to be set as part of that — handleChat
treats its absence as "first turn" and would open a second subscription, doubling
every message.
the client now echoes the model and cwd from its session:init back in the
handshake, since after a restart it is the only party that still remembers them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
step 3 of docs/nextcloud-replacement.md. collections, events and contacts as
plain json, so the browser never has to parse multistatus xml to draw a list.
this talks dav to radicale over loopback rather than reading its on-disk format.
the storage layout is radicale's private business and changes between versions;
propfind is its supported interface and costs one in-process hop. parsing the
storage directly would be faster and would break silently on upgrade, which is a
bad trade for a calendar.
three bugs found and fixed while verifying, all of which fail quietly rather
than loudly:
calendar-data and address-data are NOT webdav live properties. rfc 4791 and
6352 define them as report-only, and radicale correctly returns an empty prop
for them under propfind — so the first version returned a 207 full of nothing,
which reads exactly like "your calendar is empty".
the principal resource matched the calendar test, because `<C:calendar-home-set/>`
satisfies /calendar\b/. the principal showed up in the list as a calendar
called "1". the tag has to be required to end.
the internal fetcher omitted X-Script-Name, so the ui was handed /1/work/ for
the same collection a phone sees as /dav/1/work/, and nothing downstream could
have matched them up.
ical.ts is deliberately small: it unfolds lines and pulls out the fields a list
shows. it does NOT expand rrule or resolve vtimezone — radicale owns correctness
there, and rrule is passed through raw so the ui can say "repeats" without
either of us pretending to know when.
verified against the running stack with a real vevent and a real vcard, then the
fixtures were removed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
wraps the self-hosted memos instance, same shape as transmission and slskd. no
schema change was needed: service_connections already says `service` is text
because "adding a service should not be a schema change", and memos is the
one-instance-per-owner case that table was built for.
the sidecar holds the url and the personal access token; the platform side is
16 lines of createSidecarProxy and holds neither.
/_api/* is a pass-through onto the instance's own /api/v1 rather than a
hand-written wrapper per endpoint — memos generates its rest api from protobufs
and it moves between minor versions, so re-describing it here would be a second
thing to keep in sync. the allow-list is the one piece of policy, and it keeps
this from being a general ssrf hop. auth routes are excluded: signin/signout
would mint sessions on the instance, and this authenticates with a stored token.
probing is two calls on purpose. /healthz answers unauthenticated, so a bad url
is distinguishable from a bad token — memos returns 200 and an empty list for
unauthenticated reads rather than 401, so "the list came back" proves nothing.
verified against the live container: unconfigured reports not-connected, a bad
token is rejected WITH the reason and nothing is stored, and the platform mount
401s without a session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
from a full audit of all 179 route definitions under src/servers/api, tracing
consumers through useClient, raw fetch, EventSource, the capabilities repo and
the mobile monorepo. only routes with zero consumers anywhere are removed.
server-settings/applications.ts whole file — an app install/update registry
with no settings section to drive it
server-settings/claude-code.ts whole file — the ai settings screen talks
to chat-providers/* exclusively
GET browser/extension-download superseded by a static asset; BrowserRelay
links at /browser-relay-extension.zip
GET integrations/ a stub returning []
GET chat-providers/auth /api-keys says the same thing with more detail
GET desktop/vnc-status and with it the vnc:status command and reply,
which existed only to serve this route.
docs/sidecar-audit-2026-07.md called this
one dead months ago
deliberately KEPT, because "no caller" turned out not to mean "dead":
POST activity/announce not orphaned — it is the missing PRODUCER for the
detached[] list GET activity/tasks already returns
and ActivityScreen already renders. an unbuilt
feature, not dead code, and finishing or dropping
it is a product decision.
GET agents/runs three days old. part of agent grounds, still being
built. "not yet consumed" is not "dead".
PUT/GET vault/unlock-key six days old, storage half of a feature whose
client half is unwritten. the vault is off limits.
DELETE integrations/google/connection caller exists but is deliberately
commented out of the tree. dormant on purpose.
vnc-manager's getSession is now orphaned too, but it is sidecar-internal and
was not in scope; noted rather than chased.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
first two steps of docs/nextcloud-replacement.md — the half that has to work on
a phone, because that is the half that cannot be faked.
radicale is supervised by the sidecar rather than reimplemented. nextcloud does
not implement caldav either; it vendors sabre/dav. icalendar and vcard are a
weekend, but sync-collection, rrule expansion, vtimezone and ctag/etag are not,
and when they are subtly wrong a phone does not error — it silently stops
syncing, or silently duplicates every event.
two doors, because a browser should not speak dav:
/dav/* top-level, http basic against a scoped app password, every
verb and every dav header forwarded verbatim. this is what
davx5 and ios talk to. same reasoning as /api/vault being
mounted outside protectedRouter.
/api/caldav/* the ordinary sidecar proxy, for officer's own ui. json.
the shared proxy factory could not carry the dav door: it forwards three
headers and dav dies without Depth, and it derives the user from a jwt a phone
cannot hold. so it is a separate file, per that factory's own instruction never
to grow per-app logic.
new `dav_app_passwords` — a phone cannot do jwt, and the alternative is the
account password living in a phone's account manager. argon2, shown once,
revocable per device, and accepted ONLY by /dav.
.well-known/caldav and carddav redirect to the dav root. they are most of what
makes adding an account feel transparent, and they need naming explicitly in
server.tsx or the SPA `/*` fallback answers the phone with html.
verified end to end against the running stack: 401 + WWW-Authenticate
unauthenticated; 207 with calendar-access and addressbook advertised; MKCALENDAR,
PUT and GET of a real VEVENT; calendar-query and sync-collection REPORTs; MKCOL,
PUT and GET of a real vCard. X-Script-Name is set because radicale otherwise
generates hrefs at / and the client follows them into the SPA.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the harness stamps every message a subagent produces with parent_tool_use_id.
the sidecar wrote it outgoing and nothing ever read it coming back, so a
subagent's prose and tool calls were spliced into the main transcript as if the
agent you are talking to had produced them — and worse, its deltas were appended
to the same text buffer, so two voices were concatenated inside one bubble.
both buffering layers (stream-parser's textBuffer and turn-stream's buffer) are
now maps keyed by parent, and parentToolUseId rides on ChatEvent, ServerMessage
and Message. useChat nests parented output under the Task row that spawned it;
ToolActivity draws the trace inside the expanded panel.
background tasks get the same treatment from the other end: task:started and
task:notification were two unrelated fake assistant bubbles minutes apart, and
are now one role:'task' row correlated by taskId that appears pending and
resolves in place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chain was: a Stop hook in Claude's settings curls POST /api/hooks/claude-done,
the platform POSTs /_officer/panel-refresh to the pty sidecar, the sidecar sends a
`panel-refresh` frame to every attached terminal, and the Claude Code panel bumps
`preview:refresh` and `files:refresh-signal`.
It has never fired. generateClaudeSettings writes settings.json into the MANAGED
home under DATA_PATH, but HOME_DIR points terminals at the owner's real login home
— which is where Claude reads its settings from. Verified on this machine: no
claude-done hook exists in ~/.claude/settings.json, and DATA_PATH/*/home/.claude
does not exist at all.
Deleting rather than repairing it, because the Chat panel already does exactly this
job from onTurnComplete — in-process, conditioned on the turn having made tool
calls, with no hook, no HTTP round trip, and no endpoint. The chat UI is where agent
work happens; the terminal TUI is not the destination.
Also removes /api/hooks/claude-done, which was mounted above protectedRouter and so
was the one unauthenticated write-ish endpoint on the API surface.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
unregisterSidecar rejected every entry in the pending map, not just the ones
belonging to the sidecar that went away. Restarting any single sidecar failed
in-flight work on every other: `pm2 restart officer-music` could kill a running
agent turn with `Sidecar "music" disconnected` — a message pointing at a process
that had nothing to do with it. Pending entries now carry their owning sidecar
id and the rejection loop skips the rest.
Also deletes src/servers/api/anthropic-proxy.ts. It had no importers; PM2's
officer-anthropic-proxy runs src/servers/sidecar/claude/index.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/api/vpn/enroll minted pre-auth keys itself, from HEADSCALE_URL, HEADSCALE_API_KEY
and HEADSCALE_USER in the host env. Three globals describe one server; Officer keeps
a registry of many in headscale_servers with one active, so the env could contradict
the server the owner had selected — and HEADSCALE_USER filed every joining device
under the same name on all of them.
The two credential vars had already been removed from the environment and nothing
noticed: the route checks `if (!base || !apiKey)` first, so it had been answering
503 to every enrollment attempt, silently. HEADSCALE_USER was read but never reached.
Enrollment moves into the sidecar that owns the registry and acts on the active
server. The owning user is resolved rather than hardcoded: an explicit userId wins,
one user on the server needs no choice, several is a 409 listing them instead of a
silent guess. The platform route keeps its path and response shape — both are a
contract with enrollVpn() in the mobile core — and is now a bare forward holding no
Headscale URL, key or user name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the key in chain-source-esplora.test.ts was labelled a published test vector and
was not one — it was the owner's live bip84 account xpub, pulled from the
database during an earlier verification and pasted in.
it cannot spend, but it discloses every address that wallet will ever use and
its whole history, permanently. replaced with the bip84 spec's own vector,
derived in the file from the published mnemonic so its provenance can be checked
rather than taken on trust, and pinned by an assertion against the spec's first
address so a future substitution fails loudly.
this does not remove the key from history. that needs a rewrite of 5c38236.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
esplora and nbxplorer are now both selectable from wallet settings. the two are
stored as separate service_connections rows but are mutually exclusive: saving
either retires the other, so "which endpoint is in use" is never decided by a
precedence rule.
the nbxplorer probe cross-checks the chain it reports indexing against the
configured network, so pointing a mainnet wallet at a testnet node is refused at
the form rather than discovered later as an unexplained zero balance. esplora
cannot report this, so there is nothing to check there.
/_health now runs the same probe the form does, instead of its own hardcoded
esplora path — the two can no longer disagree about what a working endpoint is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OnchainBackend depended on the concrete EsploraChain class, so the only wallet
it could ever have was an Esplora-backed one. The seam is now WalletChainSource,
and it is drawn at the scan rather than at the HTTP client: Esplora is
address-level and has to walk the gap limit, NBXplorer is wallet-level and has
no per-address endpoint at all, so there is nothing to share one level down.
The backend keeps the keys and the money — derivation, snapshot cache, coin
selection, PSBT construction, signing — and owns no HTTP. Which indexer answers
is a constructor argument.
Also adds the NBXplorer implementation of the seam, verified end to end against
the owner's own pruned node, and the first tests over any of this: a stub
Esplora drives a real backend through the gap-limit walk, balance summation,
UTXO mapping, transaction scoring and address issuance. Nothing covered the
scan before it was moved, which is the wrong time to have no tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The second chain source. Esplora is address-level, so finding a wallet's coins means
walking the gap limit ourselves — 43-160 requests per refresh. NBXplorer is scheme-level:
register the account xpub once and every question after that is a single call. So this is
deliberately NOT an implementation of EsploraChain's interface; there is no per-address
query worth emulating, and emulating one would throw the advantage away.
Verified end to end against the owner's live 2.6.9 instance: status, track, balance,
utxos, transactions, unused address and fee estimates all round-trip, and NBXplorer
derives bc1qwnvlm8... for BIP84 0/0 — byte-identical to what keys.ts derives.
Three things the upstream docs get wrong or leave out, all confirmed by hand:
- single-sig taproot is `-[taproot]`, absent from NBXplorer's own scheme table
- querying an UNTRACKED scheme returns 200 with every figure zeroed, which reads as a
real empty wallet; track() is therefore on every read path, not just at setup
- a rejected broadcast returns 200 with success:false, never an HTTP error
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WALLET_ESPLORA_URL was the one wallet setting an owner actually has to change — off a
public explorer that rate-limits and sees every address, onto their own indexer — and it
was the one they could only change with a shell and a restart. It now lives in
service_connections under 'esplora' and is edited at Wallet -> Settings -> Chain source,
probed against /blocks/tip/height before it is stored.
The URL joins the backend fingerprint, so re-pointing rebuilds every on-chain backend and
drops the gap-limit scan taken through the old endpoint. /_health probes what the wallets
actually use rather than the built-in default, and /_officer/config no longer reports a URL
it cannot know.
Also two receive-screen defects the Blockstream 429 exposed: a query error rendered as
"No address available", and the "new address" button called refetch() on the ?peek=true
query, so it re-fetched the same address instead of advancing the index.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the PATCH route already accepted a name; nothing in the UI ever sent one. adds an
inline editor on the settings header (pencil → input, enter saves, escape cancels)
and a rename mutation. renaming touches only the label, so it needs neither the
passphrase nor an unlocked wallet.
the route took the name unvalidated — it now trims and refuses a blank one, with a
64-char cap matched on create so a name you can create is one you can type back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
both sidecars read their upstream from a new service_connections table instead of
process.env: one row per (user, service), the secret encrypted at rest, upserted
through a /_config route the app drives. transmission gains a Connection section,
soulseek gains one too, and both take over the whole app while nothing is stored.
TRANSMISSION_URL/USER/PASS/RPC_PATH and SLSKD_URL/API_KEY can come out of .env.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The same registry photos got: any number of labelled instances stored encrypted in
invoiceshelf_accounts, one selected, switchable from the nav. The token is write-only across
the sidecar boundary — the list has no field that could carry it back — and nothing reads
INVOICESHELF_URL/TOKEN/COMPANY_ID any more, so officer's own process.env no longer holds a
credential only the sidecar can use.
The company is pinned on the account row rather than resolved per request. InvoiceShelf's
`company` header does not error on a wrong or missing value; it silently returns another
company's books. So the choice is made once, at add time, and a token that can act for several
answers 409 with the list instead of guessing.
Both apps also take an email and password now, because neither service makes a key easy to get:
InvoiceShelf 2.4.2 ships no screen that issues tokens at all (POST /auth/login is the only way),
and Immich's is buried in account settings. The sidecar does the exchange — InvoiceShelf mints a
Sanctum token, Immich logs in, creates an all-permissions API key and closes the session again —
and stores only what comes back. The password is never persisted. Pasting a key still works.
Verified against the live instances: InvoiceShelf 2.4.2 and Immich 3.1.0, routes and DTOs read
from the running containers. The two sign-in paths are untested end to end — no second login to
try them with.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
IMMICH_URL/IMMICH_API_KEY lived in the platform-wide .env, which was wrong
twice over: bun auto-loads .env into every process started in this directory,
so `officer` itself held an immich credential it has no code to use — and
connecting a library was a shell task on the server rather than something the
owner could do from the app.
it is a registry, not a single connection: any number of labelled accounts with
one selected, the same shape headscale_servers uses. two keys against the same
instance (one per immich user) is the ordinary case, so the label is what has to
be unique, not the url. one active account per owner is enforced by a partial
unique index rather than by convention.
keys are encrypted at rest and write-only across the sidecar boundary — no route
returns one, masked or otherwise. every save is validated against the live
instance first, so a wrong or under-scoped key is a 400 with the reason instead
of a stored row that makes every later screen fail mysteriously.
the drizzle snapshot under migrations/ is regenerated; nothing applies it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Right-click an entry and agents whose triggers match appear under their own Run Agent
submenu — deliberately not folded in with tasks, because an agent run is a chat session and
not a job, and one menu promising both would lie about what a click does. The modal shows the
absolute target path, autofills entry_path with it (tilde expansion stays the server's job),
and on Run links to /chat?cwd=<runs dir> instead of a queue entry: there is no job row to view.
Trigger matching and category grouping are now shared with tasks rather than duplicated, and
the task input form is reused as-is.
Rescan counts agents and invalidates their caches, so a new AGENT.md shows up on the button
rather than after the 60s staleTime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The sidecar knew which transcript file a turn had landed in and kept it to itself —
setClaudeSession fed --resume and nothing else. A session officer started was therefore
unaddressable from the platform side. Put it on the result event so it crosses the wire.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
officer-photos owns the whole Immich contract: the instance URL and the API
key live there and nowhere else, and the platform side is an auth-gated
forwarder holding no credentials. The route surface is an allow-list keyed on
the first path segment, so admin, auth, api-keys, sessions, jobs, system-config
and libraries are unreachable by construction rather than by enumeration.
The UI mirrors Immich's own sidebar — timeline, explore, map, search, albums,
people, favorites, sharing, archive, trash — because the point of a sidecar
screen is to reproduce what the upstream already ships, then extend it. The
timeline reads Immich's columnar time-bucket format directly; selection lives
in the URL per docs/navigation-audit.md.
Two things worth knowing for anyone touching this later:
- `duration` is an integer count of milliseconds in Immich 3.0. It was an
HH:MM:SS.mmm string before, and every stale example still shows that form.
- the map container is sized with h-full/w-full, never `absolute inset-0`.
maplibre's stylesheet sets `position: relative; overflow: hidden` on the
element it is given, and an unlayered vendor rule beats Tailwind 4's layered
`.absolute` regardless of source order — so the div collapses to height 0 and
clips its own canvas away. Nothing errors: the GL context is healthy, tiles
download and pixels are drawn into a buffer nobody ever composites.
maplibre-gl is pinned to 5.x deliberately; 6.0 resolves a separate worker file
from import.meta.url, which Officer's index.html fallback answers with HTML.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bun's recursive fs.watch takes one inotify watch per ENTRY, files included —
~92k for this library against a 65536 ceiling — so the watch could never be
established. The ENOSPC came back asynchronously as an FSWatcher 'error' event
with no listener, which rethrew and killed the sidecar 17k times, draining the
per-UID watch pool for every other process on the machine along the way.
Reindexing is triggered instead (the browser button, the phone's pull-to-refresh,
the nightly full); an incremental over 6273 folders measures 1.8s.
Three index defects the nightly full had been papering over:
- outputsExist verified meta.json/cover.jpg/discography.json but neither lyrics/
nor posters/, so a lost lyrics file kept a matching v and a passing check and
the album was skipped on every incremental forever — only a full restored it.
Record both counts in the manifest and compare them (CACHE_VERSION 2 -> 3).
- walk() read a failed readdir as "the folder is gone", and runBuild prunes
whatever is missing from next — so one transient EIO on the library disk
deleted that folder and its whole subtree from the index. Carry the previous
entries forward for every error but ENOENT/ENOTDIR.
- a from-scratch build has no previous entries to carry, so it now refuses to
publish a slot when any folder was unreadable, leaving the live index alone.
A disk that hiccups during the nightly costs a skipped night, not a hole.
reindexNow builds in place, so a cache-format upgrade is handed to the staged
path rather than rewriting 6k albums underneath live readers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Re-applies to the mirror what f22c667 did for Xvnc, which I dropped by restoring vnc-manager from an
older commit. The bug is in the lifecycle, not the server, so it came straight back: the running mirror
lives in module state, a sidecar restart forgets it while the process keeps running, and waitForPort
accepted ANY listener on 5900 as proof of a healthy start.
It bit immediately. An Xvnc left over from the virtual-desktop experiment kept hold of 5900, so the
restarted sidecar reported a healthy mirror while the browser was being served the stale XFCE session
underneath it — which read as "the revert did not work".
reclaimPort frees the port before spawning (TERM, then KILL after two seconds) and waitForPort now also
fails when the process we spawned has exited.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reverts the Xvnc virtual-desktop work (83bf746, f22c667). The separate desktop was the right answer to
"4K on the TV and 1080p remote", but it brought a chain of its own problems — a dock left stranded
below the bottom of the screen after every resize, because xfce4-panel does not follow a RandR change
reliably — and the owner would rather have one session that works than two that need supervision.
So: one GNOME session, mirrored, with the TV set to 1920x1080. That is under SCALE_ABOVE_WIDTH, so
x11vnc serves it 1:1 with no scaling, and both ends see the same 1920x1080 desktop.
resizeSession is now off on the client — a mirror reflects a physical screen and cannot be resized.
scaleViewport stays ON and is load-bearing: noVNC maps a click as
(clientX - canvasRect.left) / display._scale, and autoscale() is the only code that sets the canvas's
displayed size and _scale in the same call. Disabling it pins _scale at 1 while the canvas is displayed
at some other size, which doubled every pointer coordinate. That was my change and my bug.
Kept from the Xvnc detour, because all of it applies to the mirror too: the display is discovered by
socket ownership rather than assuming :0 (GDM gives :0 to its greeter), -noxdamage is gone, the clip to
the primary output stays, noVNC loads as one bundle, and showDotCursor covers the invisible pointer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the TODO item about orphaned/duplicated VNC servers across sidecar restarts. The diagnosis
there was correct and outlived the move off x11vnc, because the shape of the bug is in the lifecycle,
not the server: the running desktop lives in module state, a sidecar restart forgets it while the
process keeps running, and waitForPort accepted ANY listener on 5900 as proof of a healthy start.
The next start would then spawn a server that could not bind the port, see the ORPHAN listening, and
report success — leaving the platform convinced it had started a desktop the browser was not
looking at.
Two changes. reclaimPort frees the port before spawning: TERM whatever holds it, KILL after two
seconds. waitForPort now also fails when the process we spawned has exited, so a foreign listener
cannot be mistaken for our own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The TV runs at 4K so it can play 4K video; a usable remote desktop wants about 1080p. One framebuffer
cannot be both, so mirroring meant every remote session was a scaled-down 4K desktop — dense to read
and expensive to encode. This gives remote its own display at 1920x1080 and leaves the TV alone.
Xvnc (TigerVNC) rather than x11vnc: it is the X server AND the VNC server in one process, so nothing
polls or scales — the server knows which rectangles changed and encodes them directly, where x11vnc
had to diff a framebuffer it did not own. It also implements RandR SetDesktopSize, so the client's
existing resizeSession makes the desktop resize itself to the browser panel. No scaling on either
side at any panel size, which removes the density problem rather than trading it for blur.
XFCE rather than GNOME, and NOT because it is lighter. Ubuntu's GNOME is managed by per-USER systemd
units — org.gnome.Shell@x11.service, gnome-session-manager@ubuntu.service and the whole
org.gnome.SettingsDaemon.* set all sit under user@<uid>.service, and gnome-session@.target is marked
RefuseManualStart. A second GNOME session for the SAME user collides with every one of them. That is
almost certainly what 106e5bd recorded as "a ghost logind session that broke lightdm login" — a
property of the session model, not something care avoids. XFCE has no such per-user units and
coexists with the TV session: same user, same home, same files, different shell.
The cookie goes in ~/.vnc/Xauthority, not ~/.Xauthority, which belongs to the TV session and is not
ours to write. Display is chosen by scanning for a free number from :2 up, checking the lock file as
well as the socket because a stale lock alone stops an X server starting. Teardown kills the session
before the server, since killing Xvnc first leaves XFCE's children reparented and running.
Verified end to end on a scratch display before committing: Xvnc listened, XFCE came up with xfwm4,
xfdesktop, xfce4-panel and xfsettingsd, RandR reported a 1920x1080 VNC-0 output, and the GNOME
session on the TV was untouched throughout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A screen that is plugged in but switched off still reports "connected" to xrandr while contributing
nothing to the framebuffer — which is exactly the state this machine is now in, with the unused KVM
output disabled. Counting those made the mirror take the multi-output path and clip to the primary
when there was only one live screen; harmless here because the clip equalled the whole framebuffer,
but wrong, and it would have masked a real single-screen case.
Only an enabled output carries a WxH+X+Y geometry, so requiring that in the pattern is what tells
the two apart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
-noxdamage was set in 106e5bd, the commit that introduced the mirror. That message explains every
other flag it chose — -scale, -shared, -forever, -localhost — and says nothing about this one, and
no doc or TODO mentions it either. It looks defensive rather than diagnosed.
It is not free. Without the DAMAGE extension x11vnc is never told which rectangles changed, so it
polls the entire framebuffer over and over to discover it. Cost then scales with screen AREA,
continuously, instead of with what actually moved. On a 4K mirror that dominates: it is why the
remote desktop here felt slow next to a machine with half the hardware and a sixth of the pixels.
xdpyinfo on this server reports DAMAGE, MIT-SHM and XFIXES, so the optimisation is available and
was simply switched off.
Kept as a comment rather than deleted silently: if stale patches ever appear (some drivers do
under-report damage) putting it back is the fix, and the next person should know that is the trade
rather than rediscovering it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
X composes every attached output into one framebuffer, so with a 4K monitor at +0+0 and a 1080p TV
at +3840+0 the framebuffer is 5760x2160 and mirroring it whole sent BOTH screens side by side,
then halved them for being over the scale threshold. The remote desktop showed a squashed double-width
image with the second monitor hanging off the right — correct, and useless.
Clip to the primary output instead: 3840x2160+0+0 here, which then scales to a clean 1920x1080.
Only clips when a primary is actually marked AND more than one output is connected. With a single
output the framebuffer already IS that screen, so clipping would add a failure mode for no gain.
The scale decision now keys off what is really being served — the clip when there is one, the whole
framebuffer otherwise — instead of a framebuffer width that may span screens.
Never showed up under LightDM because only one output was ever live there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mirror hardcoded :0. That held under LightDM, which gave the user session :0. GDM does not —
it keeps :0 for its own greeter and starts the user's Xorg with -displayfd, letting the number be
picked at runtime; on this machine the session lands on :1. So after migrating to Ubuntu Desktop
the mirror failed on every attempt with "Can't open display :0", with a healthy session sitting
one number over.
Resolve by socket ownership: /tmp/.X11-unix/X<n> is owned by whoever runs that X server, so the
socket owned by us is the owner's session and anything else is the greeter's. Falls back to the
lowest socket (a root-run Xorg, as LightDM had) and finally to 0, so it is never worse than the
constant it replaces.
The display is now threaded through rather than read from a module constant, so getFramebufferWidth
measures the display actually being mirrored and both log lines name it.
Found while verifying the XFCE-to-GNOME migration on this machine — the session was up and x11 with
the cookie in the GDM path, and only the display number was wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second half of c69cda4. The agents router imports getAgentRunsDir from data-path, which was still
an uncommitted edit, so a clone got past the missing-module error only to fail on a missing export:
SyntaxError: Export named 'getAgentRunsDir' not found in module '.../src/servers/data-path.ts'
My check before c69cda4 verified that every module the agents files import is TRACKED, but not that
the SYMBOLS they import actually EXIST in the committed version of those modules. Module resolution
and named-export resolution are separate failures and the first check only covered the former.
The whole diff to this file is one feature — 'agents' joins ItemType/ITEM_TYPES, and
getAgentRunsDir is added — so it goes in whole rather than split.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3f22a80 committed hono.ts with `import { agentsRouter } from './api/agents/agents'` while
src/servers/api/agents/ was still untracked, so master has been unbootable for any clone:
error: Cannot find module './api/agents/agents' from '.../src/servers/hono.ts'
That import line was work in progress from a parallel session that happened to be sitting in
hono.ts; staging the file to mount the notify router swept it in. The machine it was committed
from kept working because the files were there on disk, which is exactly why it went unnoticed.
Committing the three files completes what that commit already assumed. Verified first that every
module they import is tracked, and that the one cross-boundary import (`TurnMessage` from
../chat/types) is `import type`, so it is stripped at runtime and does not depend on the still
uncommitted edit to that file.
The frontend half of the same feature (AgentRunnerDialog, AgentRunnerModal, useAgents) is still
untracked and deliberately left that way — no committed file references it, so it cannot break a
clone, and it is not mine to commit.
A repo-wide scan of all 1278 tracked TS files in HEAD for imports resolving to untracked or missing
modules now comes back clean apart from index.gen.html, which is generated at boot by design.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first producer, so the pipe can be exercised from the UI before any real event is wired
to it. Marked TEMPORARY in the source and meant to be deleted when a real producer replaces
it.
POST /_officer/notify now takes the user from the X-Officer-User header the proxy injects
when the body omits it. A producer inside the tailnet says who to notify; a browser reaching
this through /api/notify cannot know its own id, and the platform has already authenticated
whoever sent it. Body still wins where present.
The button sends { type: 'test' }, which fans out to every configured channel — Discord
today, APNs and FCM the moment their credentials exist — and toasts what each one reported,
so "no channels configured" is distinguishable from "sent".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 4, and the last transport. Google directly via FCM HTTP v1 — no Expo, no firebase-admin.
Unlike APNs the signed JWT is not the credential: it is an assertion exchanged at
oauth2.googleapis.com for a 1-hour access token, so there are two things to cache and a
network round trip on the cold path. Concurrent pushes share one in-flight exchange rather
than each starting their own, and the token refreshes five minutes early so a request cannot
race the expiry.
The detail that breaks most first FCM integrations: `message.data` values must all be
STRINGS. A number or boolean is rejected with a bare INVALID_ARGUMENT that does not say which
field. Everything is stringified on the way out — id, ok and count included — and a test
walks every value asserting its type rather than trusting the code that wrote it.
Also v1-specific: there is no multicast (the /batch endpoint is deprecated), so N devices is
N requests, which happens to match the APNs shape anyway.
Errors are read for `error.status` only. FCM's `message` field can echo the device token, so
logging the whole body would put device addresses in the logs. UNREGISTERED / INVALID_ARGUMENT
/ NOT_FOUND delete the row; anything else counts a strike.
21 tests across both channels, and verified against the real endpoint: a throwaway key gets
400 invalid_grant "account not found" from Google, meaning the endpoint, form encoding,
grant_type, RS256 signature and claim structure were all accepted and only the account is
missing. A malformed assertion would have failed earlier, with a different error.
The doorbell rule is asserted on this channel too: a producer passing subject/from cannot get
either into the serialized message.
Still needed for a real send: FCM_SERVICE_ACCOUNT, a Firebase project, and google-services.json
in the Android build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 3. Apple directly over HTTP/2 — no Expo, no library, just node:http2 and node:crypto.
Three things here are the difference between working and a silent failure:
- The provider JWT must be signed with dsaEncoding 'ieee-p1363'. Node's default is DER, which
is a perfectly valid ECDSA signature that Apple rejects, and the rejection says nothing
about why. A test asserts the signature is 64 bytes rather than trusting the flag.
- Apple rejects a token minted more than once per 20 minutes and any token older than 60, so
it is cached and refreshed at 40 — between the two walls, not per request.
- APNs expects ONE long-lived HTTP/2 session carrying many requests. A session per push gets
throttled, so sessions are kept per host and re-created only when they die, with an error
handler so a transport failure cannot take the sidecar down as an unhandled rejection.
Sandbox and production are separate hosts and separate token namespaces, so devices are sent
per their stored `environment` — a debug-build token against production fails with
BadDeviceToken and no other symptom.
Dead tokens (BadDeviceToken, Unregistered, DeviceTokenNotForTopic) delete their row
immediately; everything else counts a strike. Going direct means Apple answers inline, so
none of Expo's deferred receipt-polling is needed.
The doorbell rule is now enforced by a test, not just by convention: a producer that passes
subject/from/body — which the type forbids but JavaScript permits — cannot get any of it into
the serialized payload. `aps.alert` carries a generic title composed from the category, and
the custom `officer` key carries ids the app fetches by.
12 tests, 100% of the JWT and payload paths. Verified against the real endpoint too: a
throwaway key gets 403 InvalidProviderToken from api.sandbox.push.apple.com, which means the
connection, path, apns-topic, push-type and JWT structure are all accepted and only the
credential is missing.
Still needed for a real send: APNS_KEY_P8, APNS_KEY_ID, APNS_TEAM_ID, and a device token
from the app.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 2 of docs/push-notifications.md. One real channel working end to end before any Apple
or Google credential exists, so the pipe is proven before the hard part.
officer-notify is a PM2 peer with its own loopback listener, announced as notify:server and
proxied at /api/notify. It is a sidecar rather than platform code because the producers are
spread across sidecars — the queue, email, the agent — and a platform-owned notifier would
force every one of them to call back into the platform. That is the inversion just removed
from email; this avoids recreating it.
Channels sit behind one interface (types.ts) so APNs and FCM slot in beside Discord rather
than replacing anything. Each is awaited with its own error boundary and the dispatcher
always resolves: a job that finished has finished whether or not a banner appeared, so a
channel must never be able to break its producer.
text.ts is where the doorbell rule is actually enforced. APNs and FCM both need a title to
render a banner, so "send nothing" was never available — what we control is that the string
is composed HERE from the category alone. A producer sends { type: 'mail', count: 3 } and
the wire carries "3 new emails". It cannot carry a subject line because there is nowhere to
put one.
Device registration lives behind X-Officer-User, trusted because the listener binds loopback.
Platform and environment are validated rather than defaulted: an iOS token from a debug build
fails against production APNs with a silent BadDeviceToken, so a wrong value is a device that
never receives anything and never says why. GET /_officer/devices returns only the last 8
characters of a token — enough to identify a row, not enough to push to it.
Verified end to end against a fake webhook: /_health reports configured channels, a test
notification arrives as {"content":"Officer"}, { type: 'mail', count: 3 } arrives as
{"content":"3 new emails"}, and every validation path returns its own error.
Deletes src/servers/notify/discord.ts, which this supersedes and which had no other callers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Opening a wallet meant waiting for a full gap-limit scan before any number appeared, and a
restart threw that work away. Worse, an unreachable Esplora rendered identically to an empty
wallet — as a zero balance — which is alarming for the one case where it is not true.
The backend now keeps a snapshot and serves it stale-while-revalidate: a snapshot inside the
TTL is served as-is, an older one is served immediately with a refresh started behind it, and
only a wallet that has genuinely never been read blocks on the network. wallet_chain_cache
holds one row per wallet so a refresh is a single atomic upsert.
The snapshot, not the endpoint, is the unit of caching. Balances, UTXOs and history were
three fetches over a shared scan, so the three queries a wallet screen fires on mount could
each observe a different moment; building them together costs the same requests and fixes
that incidentally.
Only the chain's own facts are stored. Addresses, scripts and pubkeys are re-derived from the
account xpub on load — cheaper than persisting them, and it means a restored snapshot cannot
disagree with the wallet's actual keys. Stored coordinates are validated rather than trusted,
and a snapshot at an unknown version is discarded, not migrated.
Two reads deliberately opt out. sendCoins takes a fresh snapshot because selecting coins from
a cached UTXO set builds a transaction spending outputs that may already be gone, and that
failure arrives as a broadcast rejection after signing. nextUnused does too, because handing
out an address whose stale record says "unused" is silent address reuse — a privacy leak the
owner cannot see or undo. Receive-address generation is therefore the one read that stops
working while the upstream is down, on purpose.
Failures are recorded alongside the last good snapshot rather than replacing it; wiping data
on failure would reproduce the exact bug this exists to fix. Every cache operation is
best-effort, so a database problem degrades to a slow load and can never fail a wallet
request. A sync block on balances, transactions and utxos carries the age to the UI, which
now distinguishes "empty" from "never read".
Verified against three live mainnet wallets: snapshots persisted and reloaded, and two
wallets kept their data and age through a real Esplora rate-limit failure while recording the
error separately.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A wallet whose Esplora endpoint was slow or down rendered as a wallet with no data at all,
rather than as an error. The cause was a timeout inversion: Bun.serve's default idleTimeout
is 10s, which is shorter than the Esplora client's own 20s per-request timeout. Bun killed
the response before the scan could either finish or report why it had not, so the failure
never reached the handler that would have surfaced it.
The ceiling has to sit above the whole gap-limit walk, not one request — a scan is many
sequential rounds of address queries. Raised to 255s, Bun's maximum, which is what every
other sidecar already uses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An LND seed is not a BIP39 mnemonic. It shares the 24-word shape and the English wordlist,
which is exactly why it fails confusingly: the words validate as plausible input, the
checksum does not, and the user is told their own seed is invalid.
aezeed is a different construction — a 19-byte payload (version, birthday, 16-byte entropy)
sealed with AEZ under scrypt(passphrase, salt), with the salt carried in the mnemonic
itself. So it needs its own decipher, not a flag on the BIP39 path. The vendored aez/aezeed
implementation under sidecar/wallet/aezeed does that, and the recovered entropy becomes the
BIP32 root the same way BIP39 output does.
The seed envelope gains a kind ('bip39' | 'aezeed') so an unlock knows which derivation to
run rather than guessing from word count, which cannot distinguish them.
Also note the aezeed passphrase is not a BIP39 passphrase: it decrypts the seed rather than
salting the derivation, so a wrong one fails the checksum outright instead of silently
producing a different wallet. The UI can therefore tell the user they typed it wrong, which
is not possible for BIP39.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signing in from a server-to-server client returned 500. originMiddleware leaves `origin`
undefined when a request has neither Origin nor Referer — exactly what such a client sends —
and signin.ts took it as a string, passed it into getPasskeysByUserIdAndOrigin, and Postgres
rejected the undefined parameter. It would have thrown on origin.startsWith() a few lines
later too.
It used to be unreachable: originValidationMiddleware rejected origin-less requests before
the handler ran, so the value was always a string by the time anything touched it. Turning
the origin checks off removed that gate without the code behind it ever having needed to
cope. This is the second thing that flag has surfaced rather than caused.
The two call sites want different answers, so they get different ones:
- signin normalises to ''. No passkey is registered against the empty origin, so an
origin-less caller gets an empty list and falls through to password auth, which is what it
is asking for.
- the four passkey routes now refuse with 400 "Passkey operations require an Origin header".
WebAuthn is defined in terms of an origin — a passkey is registered against one and is only
verifiable against the same one — so substituting '' there would be quietly wrong.
Also flips ALLOW_ANY_ORIGIN to default ON: the checks are off unless it is explicitly
'false'. A deliberate inversion of fail-closed, safe because of where this runs — the
perimeter is the tailnet, devices are admitted by hand, and every protected route still
requires a valid token. Verified both directions: unset lets a foreign origin through,
ALLOW_ANY_ORIGIN=false rejects it. The origin test suite pins the flag off, so it still
asserts the checks reject things rather than passing vacuously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All three existed to drive the platform from a chat app. The phone app does that now, so
they are dead weight — three bot gateways, three command parsers, account pairing, admin
config screens and bot tokens sitting in the database.
Gone entirely: telegram/ and whatsapp/, discord/'s bot and command handler, the shared
channel plumbing they were the only users of (pairing.ts, send-and-await.ts, types.ts,
routes.ts and its 20 config/pairing/status endpoints), the six settings components, and
their sections in the integrations screen.
What Discord keeps is the one piece worth keeping — pushing a message out — as
notify/discord.ts, configured by DISCORD_WEBHOOK_URL in the env. No UI, no pairing, no
stored credential, and it never throws: a notification that fails to send is logged and
dropped. Unset means notifications are silently skipped, which is the default state.
Kept deliberately: send-claude-code.ts and send-opencode.ts. They live under channels/ but
have nothing to do with chat apps — they are how /chat and the pipeline executor drive an
agent turn.
Also drops four dependencies with no remaining importer (discord.js,
node-telegram-bot-api, whatsapp-web.js, qrcode), and the /sync-now route added to the email
sidecar an hour ago, whose only consumer was the channel handlers.
The "Channel Models" settings section stays, with its description corrected — it is keyed
off the general access policy rather than anything channel-specific, so it governs non-owner
accounts, not chat apps. Whether that whole class still earns its place is the open
question already noted against origin-validation.
Not touched: the telegram, whatsapp and discord rows in server_integrations, which still
hold their bot tokens. Deleting rows is a different kind of decision and the SQL is in the
handover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stage 2, and the end of the inversion. The two sync handlers (1,093 lines) ran in the
platform's queue, which meant the sidecar reached back over its registration socket to ask
the platform to enqueue work, and the credentials travelled through Postgres job metadata to
get there. Option (A) from the plan: they run here now, and the Jobs screen is left to the
things it actually describes.
The handlers moved almost unedited. Their bodies were already a list of steps taking a
context, so sync-runner.ts synthesizes that context and runs them; what went away is the
JobHandler wrapper and the registration. `job.userId` is the OWNER'S EMAIL rather than a
numeric id — the queue's naming — and it resolves the mail store path, so it is called out
in the type. That is the same field whose absence made the mailbox read as empty two commits
ago; it is set from user.email and checked this time.
Deliberately not a queue: one run per account, no persistence, no retry. A failure is picked
up by the ten-minute cron like any other, and a sync interrupted by a restart resumes from
the stored cursor rather than the beginning. PermanentError survives as a local class — it
signalled "do not retry" to the queue and now just carries its message to the sync state.
accounts.ts asks the runner whether an account is syncing instead of scanning job rows, and
the queue-over-WS shim in index.ts is gone: enqueueViaWs, listJobsViaWs, the pending-response
map and the queue branch in the command handler. Nothing but a port crosses that socket now.
The three chat channels stop opening the mail store directly. They each carried their own
copy of count-rows / enqueue / poll / count-again, coupling three chat bridges to the mail
schema — and they enqueued `gmail-sync` unconditionally, the OAuth path, for an
app-password account that syncs over IMAP, so the command was already broken. One shared
helper calls a new POST /sync-now on the sidecar, which syncs and reports what arrived.
queue/handlers/ is now empty; both handlers there were email. The queue is untouched and
still serves the Jobs screen.
Not moved, and fine where they are: scripts/migrate-emails-to-sqlite.ts and
scripts/seed-imap-uids.ts are one-off maintenance scripts that open the store directly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mailbox read as empty after the migration — "No emails synced yet", and Sync Now
reporting no new mail against 18,430 messages that were sitting on disk untouched.
The store is keyed by the owner's email: DATA_PATH/<owner>/email_accounts/<account>/emails.db.
The platform's userMiddleware put the entire user row on the context, so `user.email`
resolved. The proxy injects only an id, and the middleware I wrote to replace it set
`{ id }` and nothing else — so `user.email` was undefined, the path never resolved to the
real mailbox, reads found nothing, and the sync compared against an empty set and concluded
there was nothing new. Nothing was written to the wrong place and no data was touched.
The sidecar loads the full row via getUserById now, matching what the middleware provided.
Cached: it runs on every request and single-user is a hard invariant.
Worth keeping in mind for the sidecars still to be extracted — moving routes across a
process boundary silently changes what is on the context, and it type-checks either way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>