photo sync turned out to be mostly built already. officer proxies immich through
officer-photos, and that sidecar's allow-list already permits the `assets`
resource for GET/POST/PUT/DELETE — so immich's own upload and bulk-upload-check
endpoints are already reachable with officer's auth in front and the immich key
never leaving the sidecar. the deliverable is therefore the contract, not a new
ingest service.
three gaps are written down rather than papered over:
- immich is not currently connected in officer (_health says configured:false),
so none of it could be verified live. the api key moved out of .env and was
never re-entered in the ui. owner action.
- upload bodies are buffered twice, once in createSidecarProxy and once in the
photos sidecar. nothing fails at phone-photo sizes; a video library would be
unpleasant. fixing it touches the shared factory, so it is a decision.
- no resumable upload. immich's own app has the same limitation.
file sync is design-only, as agreed. the recommendation is syncthing supervised
as a sidecar rather than reimplementing nextcloud's sync protocol — a mount is
not a sync, and the local-first replica is the whole feature.
the iOS answer is stated plainly because it changes the design: continuous
background sync is not possible there, which is why nextcloud's own iOS app is
manual too. if iOS matters, webdav becomes first-class rather than optional, and
that is the owner's call to make before any code is written.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
wraps the self-hosted memos instance, same shape as transmission and slskd. no
schema change was needed: service_connections already says `service` is text
because "adding a service should not be a schema change", and memos is the
one-instance-per-owner case that table was built for.
the sidecar holds the url and the personal access token; the platform side is
16 lines of createSidecarProxy and holds neither.
/_api/* is a pass-through onto the instance's own /api/v1 rather than a
hand-written wrapper per endpoint — memos generates its rest api from protobufs
and it moves between minor versions, so re-describing it here would be a second
thing to keep in sync. the allow-list is the one piece of policy, and it keeps
this from being a general ssrf hop. auth routes are excluded: signin/signout
would mint sessions on the instance, and this authenticates with a stored token.
probing is two calls on purpose. /healthz answers unauthenticated, so a bad url
is distinguishable from a bad token — memos returns 200 and an empty list for
unauthenticated reads rather than 401, so "the list came back" proves nothing.
verified against the live container: unconfigured reports not-connected, a bad
token is rejected WITH the reason and nothing is stored, and the platform mount
401s without a session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
from a full audit of all 179 route definitions under src/servers/api, tracing
consumers through useClient, raw fetch, EventSource, the capabilities repo and
the mobile monorepo. only routes with zero consumers anywhere are removed.
server-settings/applications.ts whole file — an app install/update registry
with no settings section to drive it
server-settings/claude-code.ts whole file — the ai settings screen talks
to chat-providers/* exclusively
GET browser/extension-download superseded by a static asset; BrowserRelay
links at /browser-relay-extension.zip
GET integrations/ a stub returning []
GET chat-providers/auth /api-keys says the same thing with more detail
GET desktop/vnc-status and with it the vnc:status command and reply,
which existed only to serve this route.
docs/sidecar-audit-2026-07.md called this
one dead months ago
deliberately KEPT, because "no caller" turned out not to mean "dead":
POST activity/announce not orphaned — it is the missing PRODUCER for the
detached[] list GET activity/tasks already returns
and ActivityScreen already renders. an unbuilt
feature, not dead code, and finishing or dropping
it is a product decision.
GET agents/runs three days old. part of agent grounds, still being
built. "not yet consumed" is not "dead".
PUT/GET vault/unlock-key six days old, storage half of a feature whose
client half is unwritten. the vault is off limits.
DELETE integrations/google/connection caller exists but is deliberately
commented out of the tree. dormant on purpose.
vnc-manager's getSession is now orphaned too, but it is sidecar-internal and
was not in scope; noted rather than chased.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
first two steps of docs/nextcloud-replacement.md — the half that has to work on
a phone, because that is the half that cannot be faked.
radicale is supervised by the sidecar rather than reimplemented. nextcloud does
not implement caldav either; it vendors sabre/dav. icalendar and vcard are a
weekend, but sync-collection, rrule expansion, vtimezone and ctag/etag are not,
and when they are subtly wrong a phone does not error — it silently stops
syncing, or silently duplicates every event.
two doors, because a browser should not speak dav:
/dav/* top-level, http basic against a scoped app password, every
verb and every dav header forwarded verbatim. this is what
davx5 and ios talk to. same reasoning as /api/vault being
mounted outside protectedRouter.
/api/caldav/* the ordinary sidecar proxy, for officer's own ui. json.
the shared proxy factory could not carry the dav door: it forwards three
headers and dav dies without Depth, and it derives the user from a jwt a phone
cannot hold. so it is a separate file, per that factory's own instruction never
to grow per-app logic.
new `dav_app_passwords` — a phone cannot do jwt, and the alternative is the
account password living in a phone's account manager. argon2, shown once,
revocable per device, and accepted ONLY by /dav.
.well-known/caldav and carddav redirect to the dav root. they are most of what
makes adding an account feel transparent, and they need naming explicitly in
server.tsx or the SPA `/*` fallback answers the phone with html.
verified end to end against the running stack: 401 + WWW-Authenticate
unauthenticated; 207 with calendar-access and addressbook advertised; MKCALENDAR,
PUT and GET of a real VEVENT; calendar-query and sync-collection REPORTs; MKCOL,
PUT and GET of a real vCard. X-Script-Name is set because radicale otherwise
generates hrefs at / and the client follows them into the SPA.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tailwind's preflight resets `list-style: none` on every ul/ol, and none of the
three prose scopes put it back. the indent was there, so a bulleted list just
looked tight — but an ORDERED list rendered with no numbers at all, which reads
as the model having emitted broken markdown. pasting the same text into an
editor showed it numbered correctly, which is the tell.
the `li::marker` rules were colouring a marker that was never drawn.
fixed in .chat-md, .file-viewer-md and .skill-md, plus the inline
.markdown-preview block in MarkdownEditor, which had the same hole. nested
levels follow the usual convention (disc/circle/square, decimal/alpha/roman) and
task-list items drop the bullet, since the checkbox is already the marker.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
reloading /chat/<id>?cwd=<dir> landed on an empty general_chat_sessions instead
of the conversation. one line did both halves of it:
if (replaceUrl) window.history.replaceState(null, '', `/chat/${msg.sessionId}`)
that ran on session:init, and session:init's sessionId is officer's own
per-connection key — websocket.ts mints it as `msg.sessionId || randomUUID()`.
/chat/sessions/:id resolves a CLAUDE TRANSCRIPT uuid, so the address bar ended
up naming something no lookup could find; the detail fetch 404'd and the catch
dropped you into a blank chat. the template also had no location.search, so
?cwd= — added later, for agent grounds — was thrown away every time the socket
connected, which is why the pwd picker fell back to the default group.
the transcript uuid is only known once the turn reports it, and it only started
crossing the wire in 7b6ca5f, so move the rewrite to `result`, use
claudeSessionId, and carry the query string through untouched. new chats gain a
working permalink too — /chat/new used to become an unresolvable id the same way.
mobile back had the same query-string hole; it now keeps the search.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the harness stamps every message a subagent produces with parent_tool_use_id.
the sidecar wrote it outgoing and nothing ever read it coming back, so a
subagent's prose and tool calls were spliced into the main transcript as if the
agent you are talking to had produced them — and worse, its deltas were appended
to the same text buffer, so two voices were concatenated inside one bubble.
both buffering layers (stream-parser's textBuffer and turn-stream's buffer) are
now maps keyed by parent, and parentToolUseId rides on ChatEvent, ServerMessage
and Message. useChat nests parented output under the Task row that spawned it;
ToolActivity draws the trace inside the expanded panel.
background tasks get the same treatment from the other end: task:started and
task:notification were two unrelated fake assistant bubbles minutes apart, and
are now one role:'task' row correlated by taskId that appears pending and
resolves in place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chain was: a Stop hook in Claude's settings curls POST /api/hooks/claude-done,
the platform POSTs /_officer/panel-refresh to the pty sidecar, the sidecar sends a
`panel-refresh` frame to every attached terminal, and the Claude Code panel bumps
`preview:refresh` and `files:refresh-signal`.
It has never fired. generateClaudeSettings writes settings.json into the MANAGED
home under DATA_PATH, but HOME_DIR points terminals at the owner's real login home
— which is where Claude reads its settings from. Verified on this machine: no
claude-done hook exists in ~/.claude/settings.json, and DATA_PATH/*/home/.claude
does not exist at all.
Deleting rather than repairing it, because the Chat panel already does exactly this
job from onTurnComplete — in-process, conditioned on the turn having made tool
calls, with no hook, no HTTP round trip, and no endpoint. The chat UI is where agent
work happens; the terminal TUI is not the destination.
Also removes /api/hooks/claude-done, which was mounted above protectedRouter and so
was the one unauthenticated write-ish endpoint on the API surface.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
unregisterSidecar rejected every entry in the pending map, not just the ones
belonging to the sidecar that went away. Restarting any single sidecar failed
in-flight work on every other: `pm2 restart officer-music` could kill a running
agent turn with `Sidecar "music" disconnected` — a message pointing at a process
that had nothing to do with it. Pending entries now carry their owning sidecar
id and the rejection loop skips the rest.
Also deletes src/servers/api/anthropic-proxy.ts. It had no importers; PM2's
officer-anthropic-proxy runs src/servers/sidecar/claude/index.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/api/vpn/enroll minted pre-auth keys itself, from HEADSCALE_URL, HEADSCALE_API_KEY
and HEADSCALE_USER in the host env. Three globals describe one server; Officer keeps
a registry of many in headscale_servers with one active, so the env could contradict
the server the owner had selected — and HEADSCALE_USER filed every joining device
under the same name on all of them.
The two credential vars had already been removed from the environment and nothing
noticed: the route checks `if (!base || !apiKey)` first, so it had been answering
503 to every enrollment attempt, silently. HEADSCALE_USER was read but never reached.
Enrollment moves into the sidecar that owns the registry and acts on the active
server. The owning user is resolved rather than hardcoded: an explicit userId wins,
one user on the server needs no choice, several is a 409 listing them instead of a
silent guess. The platform route keeps its path and response shape — both are a
contract with enrollVpn() in the mobile core — and is now a bare forward holding no
Headscale URL, key or user name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
git add -A swept up a throwaway verification script. it held no key material —
it read addresses from a file outside the repo — but it had no business being
committed. ignoring *.tmp.ts so the next one cannot repeat it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the key in chain-source-esplora.test.ts was labelled a published test vector and
was not one — it was the owner's live bip84 account xpub, pulled from the
database during an earlier verification and pasted in.
it cannot spend, but it discloses every address that wallet will ever use and
its whole history, permanently. replaced with the bip84 spec's own vector,
derived in the file from the published mnemonic so its provenance can be checked
rather than taken on trust, and pinned by an assertion against the spec's first
address so a future substitution fails loudly.
this does not remove the key from history. that needs a rewrite of 5c38236.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
esplora and nbxplorer are now both selectable from wallet settings. the two are
stored as separate service_connections rows but are mutually exclusive: saving
either retires the other, so "which endpoint is in use" is never decided by a
precedence rule.
the nbxplorer probe cross-checks the chain it reports indexing against the
configured network, so pointing a mainnet wallet at a testnet node is refused at
the form rather than discovered later as an unexplained zero balance. esplora
cannot report this, so there is nothing to check there.
/_health now runs the same probe the form does, instead of its own hardcoded
esplora path — the two can no longer disagree about what a working endpoint is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OnchainBackend depended on the concrete EsploraChain class, so the only wallet
it could ever have was an Esplora-backed one. The seam is now WalletChainSource,
and it is drawn at the scan rather than at the HTTP client: Esplora is
address-level and has to walk the gap limit, NBXplorer is wallet-level and has
no per-address endpoint at all, so there is nothing to share one level down.
The backend keeps the keys and the money — derivation, snapshot cache, coin
selection, PSBT construction, signing — and owns no HTTP. Which indexer answers
is a constructor argument.
Also adds the NBXplorer implementation of the seam, verified end to end against
the owner's own pruned node, and the first tests over any of this: a stub
Esplora drives a real backend through the gap-limit walk, balance summation,
UTXO mapping, transaction scoring and address issuance. Nothing covered the
scan before it was moved, which is the wrong time to have no tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The second chain source. Esplora is address-level, so finding a wallet's coins means
walking the gap limit ourselves — 43-160 requests per refresh. NBXplorer is scheme-level:
register the account xpub once and every question after that is a single call. So this is
deliberately NOT an implementation of EsploraChain's interface; there is no per-address
query worth emulating, and emulating one would throw the advantage away.
Verified end to end against the owner's live 2.6.9 instance: status, track, balance,
utxos, transactions, unused address and fee estimates all round-trip, and NBXplorer
derives bc1qwnvlm8... for BIP84 0/0 — byte-identical to what keys.ts derives.
Three things the upstream docs get wrong or leave out, all confirmed by hand:
- single-sig taproot is `-[taproot]`, absent from NBXplorer's own scheme table
- querying an UNTRACKED scheme returns 200 with every figure zeroed, which reads as a
real empty wallet; track() is therefore on every read path, not just at setup
- a rejected broadcast returns 200 with success:false, never an HTTP error
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WALLET_ESPLORA_URL was the one wallet setting an owner actually has to change — off a
public explorer that rate-limits and sees every address, onto their own indexer — and it
was the one they could only change with a shell and a restart. It now lives in
service_connections under 'esplora' and is edited at Wallet -> Settings -> Chain source,
probed against /blocks/tip/height before it is stored.
The URL joins the backend fingerprint, so re-pointing rebuilds every on-chain backend and
drops the gap-limit scan taken through the old endpoint. /_health probes what the wallets
actually use rather than the built-in default, and /_officer/config no longer reports a URL
it cannot know.
Also two receive-screen defects the Blockstream 429 exposed: a query error rendered as
"No address available", and the "new address" button called refetch() on the ?peek=true
query, so it re-fetched the same address instead of advancing the index.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
the PATCH route already accepted a name; nothing in the UI ever sent one. adds an
inline editor on the settings header (pencil → input, enter saves, escape cancels)
and a rename mutation. renaming touches only the label, so it needs neither the
passphrase nor an unlocked wallet.
the route took the name unvalidated — it now trims and refuses a blank one, with a
64-char cap matched on create so a name you can create is one you can type back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
both sidecars read their upstream from a new service_connections table instead of
process.env: one row per (user, service), the secret encrypted at rest, upserted
through a /_config route the app drives. transmission gains a Connection section,
soulseek gains one too, and both take over the whole app while nothing is stored.
TRANSMISSION_URL/USER/PASS/RPC_PATH and SLSKD_URL/API_KEY can come out of .env.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The same registry photos got: any number of labelled instances stored encrypted in
invoiceshelf_accounts, one selected, switchable from the nav. The token is write-only across
the sidecar boundary — the list has no field that could carry it back — and nothing reads
INVOICESHELF_URL/TOKEN/COMPANY_ID any more, so officer's own process.env no longer holds a
credential only the sidecar can use.
The company is pinned on the account row rather than resolved per request. InvoiceShelf's
`company` header does not error on a wrong or missing value; it silently returns another
company's books. So the choice is made once, at add time, and a token that can act for several
answers 409 with the list instead of guessing.
Both apps also take an email and password now, because neither service makes a key easy to get:
InvoiceShelf 2.4.2 ships no screen that issues tokens at all (POST /auth/login is the only way),
and Immich's is buried in account settings. The sidecar does the exchange — InvoiceShelf mints a
Sanctum token, Immich logs in, creates an all-permissions API key and closes the session again —
and stores only what comes back. The password is never persisted. Pasting a key still works.
Verified against the live instances: InvoiceShelf 2.4.2 and Immich 3.1.0, routes and DTOs read
from the running containers. The two sign-in paths are untested end to end — no second login to
try them with.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
IMMICH_URL/IMMICH_API_KEY lived in the platform-wide .env, which was wrong
twice over: bun auto-loads .env into every process started in this directory,
so `officer` itself held an immich credential it has no code to use — and
connecting a library was a shell task on the server rather than something the
owner could do from the app.
it is a registry, not a single connection: any number of labelled accounts with
one selected, the same shape headscale_servers uses. two keys against the same
instance (one per immich user) is the ordinary case, so the label is what has to
be unique, not the url. one active account per owner is enforced by a partial
unique index rather than by convention.
keys are encrypted at rest and write-only across the sidecar boundary — no route
returns one, masked or otherwise. every save is validated against the live
instance first, so a wrong or under-scoped key is a 400 with the reason instead
of a stored row that makes every later screen fail mysteriously.
the drizzle snapshot under migrations/ is regenerated; nothing applies it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
clicking a cluster only zoomed, so a circle marked 700 was unopenable at any
zoom level that still grouped them. it now opens a lazy list of exactly the
assets that cluster covers, independent of zoom.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Section 17 was still labelled "XFCE + VNC" while the step it runs installs
ubuntu-desktop and is guarded on it, so the heading described a setup the
script had already stopped producing.
Also spell out why setup-desktop.sh disables lightdm: it is not a display
manager this script ever installs, it is residue on hosts set up by an earlier
version that did install XFCE, and left enabled it beats GDM to the seat.
Comments and one echo string; no behaviour change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Right-click an entry and agents whose triggers match appear under their own Run Agent
submenu — deliberately not folded in with tasks, because an agent run is a chat session and
not a job, and one menu promising both would lie about what a click does. The modal shows the
absolute target path, autofills entry_path with it (tilde expansion stays the server's job),
and on Run links to /chat?cwd=<runs dir> instead of a queue entry: there is no job row to view.
Trigger matching and category grouping are now shared with tasks rather than duplicated, and
the task input form is reused as-is.
Rescan counts agents and invalidates their caches, so a new AGENT.md shows up on the button
rather than after the 60s staleTime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Which project group you are looking at is addressable state, so /chat?cwd=<dir> has to be a
link anyone can hand out — it survives a refresh and an agent run can point straight at its
own runs directory. Was usePanelChannel('chat:active-cwd'), which per the navigation audit is
for signals and refresh buses only. Row links and New Chat now carry the query string along.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The sidecar knew which transcript file a turn had landed in and kept it to itself —
setClaudeSession fed --resume and nothing else. A session officer started was therefore
unaddressable from the platform side. Put it on the result event so it crosses the wire.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
officer-photos owns the whole Immich contract: the instance URL and the API
key live there and nowhere else, and the platform side is an auth-gated
forwarder holding no credentials. The route surface is an allow-list keyed on
the first path segment, so admin, auth, api-keys, sessions, jobs, system-config
and libraries are unreachable by construction rather than by enumeration.
The UI mirrors Immich's own sidebar — timeline, explore, map, search, albums,
people, favorites, sharing, archive, trash — because the point of a sidecar
screen is to reproduce what the upstream already ships, then extend it. The
timeline reads Immich's columnar time-bucket format directly; selection lives
in the URL per docs/navigation-audit.md.
Two things worth knowing for anyone touching this later:
- `duration` is an integer count of milliseconds in Immich 3.0. It was an
HH:MM:SS.mmm string before, and every stale example still shows that form.
- the map container is sized with h-full/w-full, never `absolute inset-0`.
maplibre's stylesheet sets `position: relative; overflow: hidden` on the
element it is given, and an unlayered vendor rule beats Tailwind 4's layered
`.absolute` regardless of source order — so the div collapses to height 0 and
clips its own canvas away. Nothing errors: the GL context is healthy, tiles
download and pixels are drawn into a buffer nobody ever composites.
maplibre-gl is pinned to 5.x deliberately; 6.0 resolves a separate worker file
from import.meta.url, which Officer's index.html fallback answers with HTML.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bun's recursive fs.watch takes one inotify watch per ENTRY, files included —
~92k for this library against a 65536 ceiling — so the watch could never be
established. The ENOSPC came back asynchronously as an FSWatcher 'error' event
with no listener, which rethrew and killed the sidecar 17k times, draining the
per-UID watch pool for every other process on the machine along the way.
Reindexing is triggered instead (the browser button, the phone's pull-to-refresh,
the nightly full); an incremental over 6273 folders measures 1.8s.
Three index defects the nightly full had been papering over:
- outputsExist verified meta.json/cover.jpg/discography.json but neither lyrics/
nor posters/, so a lost lyrics file kept a matching v and a passing check and
the album was skipped on every incremental forever — only a full restored it.
Record both counts in the manifest and compare them (CACHE_VERSION 2 -> 3).
- walk() read a failed readdir as "the folder is gone", and runBuild prunes
whatever is missing from next — so one transient EIO on the library disk
deleted that folder and its whole subtree from the index. Carry the previous
entries forward for every error but ENOENT/ENOTDIR.
- a from-scratch build has no previous entries to carry, so it now refuses to
publish a slot when any folder was unreadable, leaving the live index alone.
A disk that hiccups during the nightly costs a skipped night, not a hole.
reindexNow builds in place, so a cache-format upgrade is handed to the staged
path rather than rewriting 6k albums underneath live readers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Re-applies to the mirror what f22c667 did for Xvnc, which I dropped by restoring vnc-manager from an
older commit. The bug is in the lifecycle, not the server, so it came straight back: the running mirror
lives in module state, a sidecar restart forgets it while the process keeps running, and waitForPort
accepted ANY listener on 5900 as proof of a healthy start.
It bit immediately. An Xvnc left over from the virtual-desktop experiment kept hold of 5900, so the
restarted sidecar reported a healthy mirror while the browser was being served the stale XFCE session
underneath it — which read as "the revert did not work".
reclaimPort frees the port before spawning (TERM, then KILL after two seconds) and waitForPort now also
fails when the process we spawned has exited.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reverts the Xvnc virtual-desktop work (83bf746, f22c667). The separate desktop was the right answer to
"4K on the TV and 1080p remote", but it brought a chain of its own problems — a dock left stranded
below the bottom of the screen after every resize, because xfce4-panel does not follow a RandR change
reliably — and the owner would rather have one session that works than two that need supervision.
So: one GNOME session, mirrored, with the TV set to 1920x1080. That is under SCALE_ABOVE_WIDTH, so
x11vnc serves it 1:1 with no scaling, and both ends see the same 1920x1080 desktop.
resizeSession is now off on the client — a mirror reflects a physical screen and cannot be resized.
scaleViewport stays ON and is load-bearing: noVNC maps a click as
(clientX - canvasRect.left) / display._scale, and autoscale() is the only code that sets the canvas's
displayed size and _scale in the same call. Disabling it pins _scale at 1 while the canvas is displayed
at some other size, which doubled every pointer coordinate. That was my change and my bug.
Kept from the Xvnc detour, because all of it applies to the mirror too: the display is discovered by
socket ownership rather than assuming :0 (GDM gives :0 to its greeter), -noxdamage is gone, the clip to
the primary output stays, noVNC loads as one bundle, and showDotCursor covers the invisible pointer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reverts the scaleViewport=false half of 69c750d. That change was wrong and it caused a worse bug than
the one it was aimed at.
noVNC maps a click as (clientX - canvasRect.left) / display._scale. autoscale(), which only runs when
scaleViewport is enabled, sets the canvas's displayed size and _scale in the same call, so the two are
consistent by construction. With scaleViewport off, _scale is pinned at 1 while the canvas continues to
be displayed at some other size and nothing reconciles them.
Measured on the owner's session against a 1109x715 desktop: pointing a quarter of the way across
registered at x=575 (51%), and pointing at the middle saturated at x=1108, the right-hand edge. A clean
factor of two in both axes. Clicks near the right and bottom edges still appeared to work, because
doubling an already-large coordinate clamps back onto the edge it was aimed at — which is why the
panel clock and the dock stayed clickable while the middle of the screen did not.
My reasoning for the original change — that resizeSession and scaleViewport are alternatives — was
simply wrong. They are complementary: resize matches the session to the container, scaling covers the
interval before the server honours it, and scaling is also what keeps the coordinate maths honest.
showDotCursor stays: the invisible pointer was a real and separate defect.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The pointer was invisible, not misplaced. noVNC hides the browser's own cursor over the canvas and
draws the remote cursor in its place, but initialises that image to RFB.cursors.none — so until the
server sends a shape, nothing is drawn AND the real cursor stays hidden. The pointer simply vanishes.
Measured before changing anything: against a 1109x715 session the pointer covered x 5..1108 and
y 11..714, a clean 1:1 map with no clamping or scaling. Clicks were landing exactly where they were
aimed the whole time; the cursor just could not be seen, which is indistinguishable from a desync
when you are the one trying to click something.
showDotCursor renders a dot in exactly that case. It does not mask a real problem: when the server
does supply a shape, the shape still wins.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both were enabled. They are alternative strategies, not complementary ones: resizeSession asks the
server to become the container's size, scaleViewport scales whatever the server sends to fit. Running
both means noVNC resizes AND then applies a scale factor, and any gap between the size requested and
the size actually granted leaves a fractional scale that every pointer coordinate is mapped through.
The visible symptom was a cursor offset from the real pointer, with clicks landing somewhere else.
It only started mattering with the move to Xvnc. x11vnc could not resize, so resizeSession was inert
and scaleViewport did all the work; Xvnc implements RandR SetDesktopSize, so now both are live and
they interfere.
Scaling is redundant now the session genuinely becomes the container size: pixels are 1:1 and there
is no coordinate arithmetic left to get wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the TODO item about orphaned/duplicated VNC servers across sidecar restarts. The diagnosis
there was correct and outlived the move off x11vnc, because the shape of the bug is in the lifecycle,
not the server: the running desktop lives in module state, a sidecar restart forgets it while the
process keeps running, and waitForPort accepted ANY listener on 5900 as proof of a healthy start.
The next start would then spawn a server that could not bind the port, see the ORPHAN listening, and
report success — leaving the platform convinced it had started a desktop the browser was not
looking at.
Two changes. reclaimPort frees the port before spawning: TERM whatever holds it, KILL after two
seconds. waitForPort now also fails when the process we spawned has exited, so a foreign listener
cannot be mistaken for our own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The TV runs at 4K so it can play 4K video; a usable remote desktop wants about 1080p. One framebuffer
cannot be both, so mirroring meant every remote session was a scaled-down 4K desktop — dense to read
and expensive to encode. This gives remote its own display at 1920x1080 and leaves the TV alone.
Xvnc (TigerVNC) rather than x11vnc: it is the X server AND the VNC server in one process, so nothing
polls or scales — the server knows which rectangles changed and encodes them directly, where x11vnc
had to diff a framebuffer it did not own. It also implements RandR SetDesktopSize, so the client's
existing resizeSession makes the desktop resize itself to the browser panel. No scaling on either
side at any panel size, which removes the density problem rather than trading it for blur.
XFCE rather than GNOME, and NOT because it is lighter. Ubuntu's GNOME is managed by per-USER systemd
units — org.gnome.Shell@x11.service, gnome-session-manager@ubuntu.service and the whole
org.gnome.SettingsDaemon.* set all sit under user@<uid>.service, and gnome-session@.target is marked
RefuseManualStart. A second GNOME session for the SAME user collides with every one of them. That is
almost certainly what 106e5bd recorded as "a ghost logind session that broke lightdm login" — a
property of the session model, not something care avoids. XFCE has no such per-user units and
coexists with the TV session: same user, same home, same files, different shell.
The cookie goes in ~/.vnc/Xauthority, not ~/.Xauthority, which belongs to the TV session and is not
ours to write. Display is chosen by scanning for a free number from :2 up, checking the lock file as
well as the socket because a stale lock alone stops an X server starting. Teardown kills the session
before the server, since killing Xvnc first leaves XFCE's children reparented and running.
Verified end to end on a scratch display before committing: Xvnc listened, XFCE came up with xfwm4,
xfdesktop, xfce4-panel and xfsettingsd, RandR reported a 1920x1080 VNC-0 output, and the GNOME
session on the TV was untouched throughout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
public/novnc is unbundled noVNC source. The dynamic import pointed at rfb.js, so the browser walked
the module graph natively: fetch a file, parse it, discover its imports, fetch those, repeat. The
graph is 42 modules and six levels deep, so opening the desktop cost six SEQUENTIAL round trips and
42 requests before the VNC handshake could even start — and a hard refresh pays it in full every
time. Over the tailnet that is the "takes a long time to load", not the pixels.
Bundled with bun: 52 modules to one 190 KB file, 56 KB gzipped, one request. The regeneration
command is in a comment next to the constant so a future noVNC bump does not silently keep serving
a stale bundle.
The source tree stays: it is what gets bundled, and keeping it makes the diff of a version bump
readable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A screen that is plugged in but switched off still reports "connected" to xrandr while contributing
nothing to the framebuffer — which is exactly the state this machine is now in, with the unused KVM
output disabled. Counting those made the mirror take the multi-output path and clip to the primary
when there was only one live screen; harmless here because the clip equalled the whole framebuffer,
but wrong, and it would have masked a real single-screen case.
Only an enabled output carries a WxH+X+Y geometry, so requiring that in the pattern is what tells
the two apart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
-noxdamage was set in 106e5bd, the commit that introduced the mirror. That message explains every
other flag it chose — -scale, -shared, -forever, -localhost — and says nothing about this one, and
no doc or TODO mentions it either. It looks defensive rather than diagnosed.
It is not free. Without the DAMAGE extension x11vnc is never told which rectangles changed, so it
polls the entire framebuffer over and over to discover it. Cost then scales with screen AREA,
continuously, instead of with what actually moved. On a 4K mirror that dominates: it is why the
remote desktop here felt slow next to a machine with half the hardware and a sixth of the pixels.
xdpyinfo on this server reports DAMAGE, MIT-SHM and XFIXES, so the optimisation is available and
was simply switched off.
Kept as a comment rather than deleted silently: if stale patches ever appear (some drivers do
under-report damage) putting it back is the fix, and the next person should know that is the trade
rather than rediscovering it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
X composes every attached output into one framebuffer, so with a 4K monitor at +0+0 and a 1080p TV
at +3840+0 the framebuffer is 5760x2160 and mirroring it whole sent BOTH screens side by side,
then halved them for being over the scale threshold. The remote desktop showed a squashed double-width
image with the second monitor hanging off the right — correct, and useless.
Clip to the primary output instead: 3840x2160+0+0 here, which then scales to a clean 1920x1080.
Only clips when a primary is actually marked AND more than one output is connected. With a single
output the framebuffer already IS that screen, so clipping would add a failure mode for no gain.
The scale decision now keys off what is really being served — the clip when there is one, the whole
framebuffer otherwise — instead of a framebuffer width that may span screens.
Never showed up under LightDM because only one output was ever live there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mirror hardcoded :0. That held under LightDM, which gave the user session :0. GDM does not —
it keeps :0 for its own greeter and starts the user's Xorg with -displayfd, letting the number be
picked at runtime; on this machine the session lands on :1. So after migrating to Ubuntu Desktop
the mirror failed on every attempt with "Can't open display :0", with a healthy session sitting
one number over.
Resolve by socket ownership: /tmp/.X11-unix/X<n> is owned by whoever runs that X server, so the
socket owned by us is the owner's session and anything else is the greeter's. Falls back to the
lowest socket (a root-run Xorg, as LightDM had) and finally to 0, so it is never worse than the
constant it replaces.
The display is now threaded through rather than read from a module constant, so getFramebufferWidth
measures the display actually being mirrored and both log lines name it.
Found while verifying the XFCE-to-GNOME migration on this machine — the session was up and x11 with
the cookie in the GDM path, and only the display number was wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second half of c69cda4. The agents router imports getAgentRunsDir from data-path, which was still
an uncommitted edit, so a clone got past the missing-module error only to fail on a missing export:
SyntaxError: Export named 'getAgentRunsDir' not found in module '.../src/servers/data-path.ts'
My check before c69cda4 verified that every module the agents files import is TRACKED, but not that
the SYMBOLS they import actually EXIST in the committed version of those modules. Module resolution
and named-export resolution are separate failures and the first check only covered the former.
The whole diff to this file is one feature — 'agents' joins ItemType/ITEM_TYPES, and
getAgentRunsDir is added — so it goes in whole rather than split.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3f22a80 committed hono.ts with `import { agentsRouter } from './api/agents/agents'` while
src/servers/api/agents/ was still untracked, so master has been unbootable for any clone:
error: Cannot find module './api/agents/agents' from '.../src/servers/hono.ts'
That import line was work in progress from a parallel session that happened to be sitting in
hono.ts; staging the file to mount the notify router swept it in. The machine it was committed
from kept working because the files were there on disk, which is exactly why it went unnoticed.
Committing the three files completes what that commit already assumed. Verified first that every
module they import is tracked, and that the one cross-boundary import (`TurnMessage` from
../chat/types) is `import type`, so it is stripped at runtime and does not depend on the still
uncommitted edit to that file.
The frontend half of the same feature (AgentRunnerDialog, AgentRunnerModal, useAgents) is still
untracked and deliberately left that way — no committed file references it, so it cannot break a
clone, and it is not mine to commit.
A repo-wide scan of all 1278 tracked TS files in HEAD for imports resolving to untracked or missing
modules now comes back clean apart from index.gen.html, which is generated at boot by design.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Filtering a big search is how you find an album, not just how you hide
files: you type one track you know, and the folder that comes back is the
one you want. But the filter had also stripped that folder down to the one
track, and "download folder" then queued only that track — so finding the
album and losing it were the same act.
A folder now keeps its whole self alongside its matching files, and says
"3 of 24". Expanding shows the album, keeps the matched track highlighted,
and widens every download button in it to the full folder. A peer whose
other folders matched nothing can be expanded the same way, since the album
you found usually sits next to the rest of the artist.
Nothing queues what isn't on screen: each button downloads exactly what is
shown under it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A history row is only as fresh as the last list load, and the list was only
read on mount — so a search you had just started sat at 0 responses looking
stalled until you navigated away and back. slskd's own web UI goes straight
to the search when you submit, and that view already polls while the search
runs, so follow it.
Which search is open now lives in ?search=<id> rather than local state: rows
are real links, Back returns to the history, and a reload keeps your place.
The list also keeps re-reading while any search is still running.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TEMPORARY / EXPERIMENTAL, at the owner's request, after deedy/qr-data-transfer (QRFerry).
Two panels: one loops a file as QR frames, the other scans them through the camera and
rebuilds it. Entirely client-side — nothing about a transfer reaches the server, which is
the point of the technique.
NOT RaptorQ, and that is the one real design decision here. QRFerry carries RFC 6330
fountain-coded symbols so a receiver can rebuild from ANY sufficient set of frames. This
uses a plain indexed carousel instead, for two reasons — the second being the deciding one:
1. RFC 6330 is days of work and unpleasant to debug.
2. The sender and receiver are being reimplemented on iOS and Android. A format one person
can re-derive from protocol.ts in an afternoon is worth more here than optical
efficiency. Every frame is independent, self-describing, and parses with a string split.
The cost is honest and written down: without fountain coding you must eventually capture each
specific frame, so a miss waits for the next pass rather than being covered by surplus. Fine
for a few hundred KB on a steady camera; it degrades where RaptorQ would start to pay for
itself.
Details that matter for the phone implementations:
- base64url, no padding — ':' and '/' would collide with the field separator.
- CRC-32 of the whole file in the meta frame, checked after reassembly. The test pins the
reference value for "123456789" (cbf43926) so any stock implementation will agree.
- The meta frame repeats every 12 frames, so a receiver joining late learns the filename and
total without waiting a full cycle.
- Error correction level L: frames are short-lived and repeated forever, so QR capacity is
better spent staying sparse enough to scan than on recovery.
16 tests over the protocol, which is the spec the other implementations should match.
The receiver needs a secure origin for camera access; over the tailnet with HTTPS that holds,
and it says so plainly rather than failing silently on plain http.
Adds qrcode and jsqr. @types/qrcode was already present and orphaned — its runtime package had
been removed with the chat-channel cleanup.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first producer, so the pipe can be exercised from the UI before any real event is wired
to it. Marked TEMPORARY in the source and meant to be deleted when a real producer replaces
it.
POST /_officer/notify now takes the user from the X-Officer-User header the proxy injects
when the body omits it. A producer inside the tailnet says who to notify; a browser reaching
this through /api/notify cannot know its own id, and the platform has already authenticated
whoever sent it. Body still wins where present.
The button sends { type: 'test' }, which fans out to every configured channel — Discord
today, APNs and FCM the moment their credentials exist — and toasts what each one reported,
so "no channels configured" is distinguishable from "sent".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 4, and the last transport. Google directly via FCM HTTP v1 — no Expo, no firebase-admin.
Unlike APNs the signed JWT is not the credential: it is an assertion exchanged at
oauth2.googleapis.com for a 1-hour access token, so there are two things to cache and a
network round trip on the cold path. Concurrent pushes share one in-flight exchange rather
than each starting their own, and the token refreshes five minutes early so a request cannot
race the expiry.
The detail that breaks most first FCM integrations: `message.data` values must all be
STRINGS. A number or boolean is rejected with a bare INVALID_ARGUMENT that does not say which
field. Everything is stringified on the way out — id, ok and count included — and a test
walks every value asserting its type rather than trusting the code that wrote it.
Also v1-specific: there is no multicast (the /batch endpoint is deprecated), so N devices is
N requests, which happens to match the APNs shape anyway.
Errors are read for `error.status` only. FCM's `message` field can echo the device token, so
logging the whole body would put device addresses in the logs. UNREGISTERED / INVALID_ARGUMENT
/ NOT_FOUND delete the row; anything else counts a strike.
21 tests across both channels, and verified against the real endpoint: a throwaway key gets
400 invalid_grant "account not found" from Google, meaning the endpoint, form encoding,
grant_type, RS256 signature and claim structure were all accepted and only the account is
missing. A malformed assertion would have failed earlier, with a different error.
The doorbell rule is asserted on this channel too: a producer passing subject/from cannot get
either into the serialized message.
Still needed for a real send: FCM_SERVICE_ACCOUNT, a Firebase project, and google-services.json
in the Android build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 3. Apple directly over HTTP/2 — no Expo, no library, just node:http2 and node:crypto.
Three things here are the difference between working and a silent failure:
- The provider JWT must be signed with dsaEncoding 'ieee-p1363'. Node's default is DER, which
is a perfectly valid ECDSA signature that Apple rejects, and the rejection says nothing
about why. A test asserts the signature is 64 bytes rather than trusting the flag.
- Apple rejects a token minted more than once per 20 minutes and any token older than 60, so
it is cached and refreshed at 40 — between the two walls, not per request.
- APNs expects ONE long-lived HTTP/2 session carrying many requests. A session per push gets
throttled, so sessions are kept per host and re-created only when they die, with an error
handler so a transport failure cannot take the sidecar down as an unhandled rejection.
Sandbox and production are separate hosts and separate token namespaces, so devices are sent
per their stored `environment` — a debug-build token against production fails with
BadDeviceToken and no other symptom.
Dead tokens (BadDeviceToken, Unregistered, DeviceTokenNotForTopic) delete their row
immediately; everything else counts a strike. Going direct means Apple answers inline, so
none of Expo's deferred receipt-polling is needed.
The doorbell rule is now enforced by a test, not just by convention: a producer that passes
subject/from/body — which the type forbids but JavaScript permits — cannot get any of it into
the serialized payload. `aps.alert` carries a generic title composed from the category, and
the custom `officer` key carries ids the app fetches by.
12 tests, 100% of the JWT and payload paths. Verified against the real endpoint too: a
throwaway key gets 403 InvalidProviderToken from api.sandbox.push.apple.com, which means the
connection, path, apns-topic, push-type and JWT structure are all accepted and only the
credential is missing.
Still needed for a real send: APNS_KEY_P8, APNS_KEY_ID, APNS_TEAM_ID, and a device token
from the app.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 2 of docs/push-notifications.md. One real channel working end to end before any Apple
or Google credential exists, so the pipe is proven before the hard part.
officer-notify is a PM2 peer with its own loopback listener, announced as notify:server and
proxied at /api/notify. It is a sidecar rather than platform code because the producers are
spread across sidecars — the queue, email, the agent — and a platform-owned notifier would
force every one of them to call back into the platform. That is the inversion just removed
from email; this avoids recreating it.
Channels sit behind one interface (types.ts) so APNs and FCM slot in beside Discord rather
than replacing anything. Each is awaited with its own error boundary and the dispatcher
always resolves: a job that finished has finished whether or not a banner appeared, so a
channel must never be able to break its producer.
text.ts is where the doorbell rule is actually enforced. APNs and FCM both need a title to
render a banner, so "send nothing" was never available — what we control is that the string
is composed HERE from the category alone. A producer sends { type: 'mail', count: 3 } and
the wire carries "3 new emails". It cannot carry a subject line because there is nowhere to
put one.
Device registration lives behind X-Officer-User, trusted because the listener binds loopback.
Platform and environment are validated rather than defaulted: an iOS token from a debug build
fails against production APNs with a silent BadDeviceToken, so a wrong value is a device that
never receives anything and never says why. GET /_officer/devices returns only the last 8
characters of a token — enough to identify a row, not enough to push to it.
Verified end to end against a fake webhook: /_health reports configured channels, a test
notification arrives as {"content":"Officer"}, { type: 'mail', count: 3 } arrives as
{"content":"3 new emails"}, and every validation path returns its own error.
Deletes src/servers/notify/discord.ts, which this supersedes and which had no other callers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Design in docs/push-notifications.md. The short version: Apple and Google are unavoidable —
iOS suspends apps so only APNs can wake one, and Android only accepts pushes from FCM — but
Expo is not. Its push service is a relay in front of both and does not remove either
credential, so we talk to Apple and Google ourselves.
Both protocols were verified in Bun before designing around them: node:http2 works as a
client (APNs is HTTP/2-only), and ES256 signing produces the raw 64-byte r||s form Apple
requires rather than Node's default DER, which is silently rejected. No push library is
needed for either channel.
The load-bearing decision is the payload: a push is a doorbell, not a message. Apple and
Google see metadata regardless, so they must not also see content — a notification carries a
category and an id, never a subject, sender or error, and the app composes the visible text
locally and fetches the real thing over the tailnet on tap.
This commit is the registry: push_devices, holding native APNs/FCM tokens. environment is a
column because APNs sandbox and production are different hosts AND different token
namespaces — a debug-build token fails against production with a silent BadDeviceToken, so
guessing is not an option. Registration upserts on (token, bundle_id) because tokens rotate
and the app re-registers every launch. Failure counting prunes dead tokens; a hard rejection
deletes at once.
Nothing sensitive lands here: a token is useless without the APNs key or FCM service
account, both of which stay in the sidecar's env.
NOTE: `bun db:push` will fail until the telegram/whatsapp/discord rows are deleted from
server_integrations — the CHECK constraint added earlier refuses while they exist. That is
the enforcement working, not a problem to route around.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
nav-test-checklist.md and email-migration-checklist.md moved to the workspace root, out of
version control.
They are working notes for a migration in progress — which checks have been clicked through
on this machine and which have not. That is state about one host at one moment, not
something a clone of the repo should carry, and it goes stale the instant the migration
finishes. The reference docs they were sitting next to are the opposite: they describe the
system and should travel with it.
The root CLAUDE.md, which pointed at both by their tracked paths, now points at the new
location and records the convention so the next one is put in the right place.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>