Where things stand
An agent server today is configured with exactly one hypermedia endpoint (SEED_AGENTS_HM_SERVER_URL) plus an IPFS gateway URL for raw reads. It talks to that endpoint using the typed DAG-CBOR /api/* protocol — the same one the CLI and the @seed-hypermedia/client library use — which is served by the web app or the desktop's API bridge, which in turn call the daemon over gRPC. The daemon's own gRPC port is not interchangeable with that endpoint.
Three facts about this arrangement shape everything that follows.
The daemon already has a real permission system. Every blob carries a visibility set — blob_visibility (id, space) — and a blob is public only if it has a row for space 0. Private documents seed rows for their owning space, and visibility propagates down the link graph so that every Change, Ref, DagPB file and Raw blob reachable from a private document inherits it. Access to a private blob is decided inside the blockstore: the caller must be the space owner or hold a Capability from the owner with role WRITER or AGENT. The AGENT role is already transitive for one hop — owner grants AGENT to a key, that key can grant onward, and the daemon's writer check walks the chain. Callers prove who they are with Daemon.Authenticate (a signed timestamp that yields a short-lived bearer token) over HTTP, or with a signed ephemeral capability over libp2p, and the RBSR sync layer filters both items and range fingerprints by the peer's authorized spaces. This is described from the storage side in Portable Agent Sessions.
The agent server does not use it. Its Seed client never authenticates. The /api/* bridge faithfully forwards an Authorization: Bearer header to the daemon, but the agent server never sends one, so every read is anonymous and every private document is invisible — including documents the agent's owner could read, and documents the agent itself just wrote privately. Writes work only because the agent server holds full account keys (the "signing identities" stored as secrets) and signs Changes and Refs with them directly, resolving a WRITER/AGENT capability to embed in the Ref.
The agent server has a second, disjoint permission system. Its signed-action envelopes are authorized against a local account_authorizations table with roles OWNER and AGENT, populated by RegisterSigner with a capability blob. That table answers "may this key act on this account on this agent server"; the daemon's capabilities answer "may this key write to this space on the network". Nothing connects them, and neither one says anything about reading.
So the work is not to invent permissions; it is to extend the daemon's model with what agents need and then make the agent server a proper client of it.
What agent servers need
Read what the owner would let them read. An agent working on a private project must see the private documents in that project. Today it cannot.
Least privilege, not account keys. A server holding a user's account key can do anything that user can, forever, everywhere. An agent should act under a grant that is narrower than the owner, scoped to paths, and revocable without rotating the owner's key.
Attributable identity. Readers of a document should be able to tell that a change was made by an agent acting for someone, not by that person. That means a distinct key per agent whose authority traces back to the owner through signed capabilities.
Server identity. A hosting server has its own key so that session logs, tool-call results and audit trails can be signed by the process that produced them, separate from the agent key that signs published content.
Revocation and expiry. Owners must be able to cut off a server or an agent, and grants should be able to expire on their own.
Multi-space, multi-source. An agent may legitimately work across several spaces that live on different daemons. The model has to answer "which data, from where, under which grant" without the agent server becoming a router.
One source of truth. The agent server's authorization decisions and the network's should derive from the same signed facts.
The proposed model: a three-key capability chain
Keys
Owner account key. Stays on the user's devices. Signs the root capability and nothing else in this flow.
Agent server key. One per deployment, generated on first boot, stored like any other server secret. It authenticates the server to daemons and peers, signs session-log entries, and issues second-hop capabilities. It is not a publishing identity.
Per-agent key. One per agent, held by the server. It signs the Changes, Comments and Refs the agent publishes and is what appears as the author. Because it is a distinct key, a reader can resolve its capability chain and render "Research assistant, acting for Eric" rather than just "Eric".
Capabilities
Hop 1 — owner → server. The owner signs a Capability with delegate = server key, role = AGENT, a path limiting the subtree, and an expires timestamp. This is the only thing the owner has to do, and it is done from a device that holds the owner key — the desktop app, or the web app's delegation flow. The existing desktop "add agent server" step is exactly where this belongs.
Hop 2 — server → agent. When an agent is created, the server signs a Capability with delegate = agent key and a role and path no broader than hop 1. The daemon's writer check already accepts this shape (owner → AGENT holder → delegate). The one rule to enforce on validation is that a hop-2 path must be equal to or under the hop-1 path.
Changes to the daemon's capability blob
Four additions, all backward compatible because existing capabilities simply lack the fields:
A READER role. Grants access to private blobs under the path without write rights. dbBlobCanCallerAccess and GetSpacesByAccount learn to accept it; the writer check ignores it. This is the piece that fixes agent reads.
expires. Capability validation treats an expired grant as absent. Bearer tokens issued to a delegate expire no later than the governing capability.
Path narrowing across hops. Hop-2 capabilities must be scoped within hop 1. The transitive query gains a prefix test.
Revocation by tombstone. Owners publish a Ref tombstone for a capability CID, the same mechanism used to delete documents. Every capability beneath a tombstoned one becomes invalid transitively. Because tombstones are blobs, revocation propagates through normal sync.
Optionally, a COMMENTER role (may publish Comments but not Changes) covers the very common "agent replies in discussions but cannot edit" case without reaching for a general scope language.
What the agent server changes
Generates a server key and exposes its principal in /api/health and the settings UI.
Replaces account_authorizations as a source of truth: the envelope authorizer resolves the signer's capability chain on the network (cached, with the tombstone check) instead of a local table. RegisterSigner becomes "publish this capability and cache it."
Authenticates its Seed client. createSeedClient gains an auth option that runs Daemon.Authenticate with a key and attaches the bearer token; the agent server authenticates as the per-agent key for reads and writes on that agent's behalf.
Stops needing account keys at all for agents that have a hop-2 capability. Existing "signing identity" secrets remain for users who explicitly want an agent to publish as them, but that becomes the exception rather than the only way.
How a private read works end to end
The agent server signs an Authenticate request with the agent key. The daemon resolves the chain — owner → server (AGENT, /projects/alpha) → agent (READER, /projects/alpha/notes) — and returns a bearer token. Subsequent requests carry it; the blockstore sees a private blob, finds an authenticated caller, and dbBlobCanCallerAccess returns true because READER on an ancestor path covers it. The response carries Cache-Control: private. Nothing about the data path is new; only the caller's identity is.
The same chain is what GetAuthorizedSpacesForPeer consults when a daemon syncs with another daemon, so a daemon operated by the agent server receives exactly the private blobs those grants cover and no others.
Topology: where should the daemon be?
A — one remote HM API (today). Fine for the hosted gateway and for the desktop (where the "remote" is the app's own local daemon). With bearer auth added it can serve private reads. Its limits are structural: the agent sees only what that one daemon has, every read is a round-trip, and the agent host must trust the API host with its bearer tokens.
C — many remote HM APIs. Make the endpoint per-account or per-agent and route requests by space. This reaches data no single daemon holds, but it means N trust relationships, N bearer tokens, a routing table, and reimplementing discovery and sync in the agent server. It is topology B rebuilt over HTTP with fewer guarantees.
B — a co-located daemon (recommended). Each agent server runs its own Seed daemon as a sidecar and talks to it over loopback. The daemon authenticates on the p2p network with the server key and the agent keys, subscribes to the spaces their capabilities reach, and serves them locally. This is precisely what a desktop already is: a daemon that holds keys, syncs what those keys may see, and answers local queries. The hypermedia network's discovery, authorized RBSR sync and visibility filtering do the multi-source work that option C would reimplement.
B also closes the loop with session logs: the agent server writes session blobs straight into its daemon's blockstore, with visibility rows for the owner's space, and they replicate to the owner's devices and to any other authorized server through ordinary sync. Moving an agent stops being a special workflow; it is authorizing a second server and letting sync run.
The cost is one more process per agent host. The hosted deployment already runs Docker Compose with several containers, and the desktop already spawns both a daemon and an agent server, so the operational shape is familiar. For a thin self-hosted deployment that genuinely cannot run a daemon, option A with bearer auth remains supported.
Interface: gRPC or the typed /api/* protocol?
The agent server currently speaks the /api/* DAG-CBOR protocol. The desktop and web apps speak gRPC (connect) to the daemon, and the web app implements /api/* on top of it. With a co-located daemon the agent server could use either.
gRPC / connect to the daemon | Typed | |
|---|---|---|
Served by | the daemon itself | web app / desktop bridge (not the daemon) |
Auth | bearer via | bearer forwarded to the daemon, already works |
Surface | full: documents, comments, capabilities, sync, subscriptions, daemon admin | read queries, resource fetch, |
Streaming | yes (server streams, websocket subscriptions) | polling |
Who else uses it | desktop, web server, CLI's gRPC skill | CLI, client library, third-party integrations |
Transport to remote hosts | gRPC-web only where exposed; not a public API | designed as the public, stable API |
Recommendation: the agent server talks connect/gRPC to its co-located daemon, exactly as the desktop does, and keeps the /api/* client for the remote-endpoint fallback (topology A) and for reaching public gateways. This gives agents the full daemon surface — subscriptions for triggers instead of polling, capability issuance, sync control, peer authentication — without the agent server depending on a web server being present.
The longer-term simplification is for the daemon to serve /api/* natively. It is a small, typed, CBOR-over-HTTP protocol; the web app's implementation is a thin mapping onto gRPC calls. Moving it into Go means every consumer — agents, CLI, client library, web — hits one process with one authorization path, and "HM server URL" always means "a daemon." That is a separate project, but the permission model here does not depend on it and would not need to change when it lands.
Rollout
Daemon: READER role, expiry, path narrowing, tombstone revocation. Extend the Capability blob and the three SQL checks (SQLCanWriteRootByOwnerID, isValidWriter, dbBlobCanCallerAccess). Regression tests for: READER can read private, cannot write; expired denies; hop-2 wider than hop-1 denies; tombstoned hop-1 invalidates hop-2.
Client library: authenticated client. createSeedClient(url, {signer}) runs Authenticate, caches the bearer, refreshes before expiry, and sends it on every request. Ship in @seed-hypermedia/client so the CLI gets it too.
Agent server: server key + per-agent keys. Generate on boot / on agent create. Issue hop-2 capabilities. Authenticate reads as the agent key. Keep account-key signing identities as an opt-in.
Desktop: grant flow. "Connect agent server" prompts for the server principal, chooses a path and expiry, and publishes the hop-1 capability — reusing the existing capability UI.
Unify authorization. The agent server's envelope check resolves chains from the daemon instead of account_authorizations; RegisterSigner publishes instead of storing.
Co-located daemon. Add the daemon to the hosted and self-hosted Compose files and to the desktop's embedded agents mode; switch the agent server's default client to connect/gRPC on loopback, with /api/* remote as a fallback when SEED_AGENTS_HM_SERVER_URL is set explicitly.
Sessions as blobs. With a local daemon, write session logs into the blockstore with owner-space visibility, per Portable Agent Sessions.
Steps 1–3 deliver private reads and least privilege on the current topology. Steps 4–5 make the grant user-facing and remove the second permission system. Steps 6–7 change the topology and are where the multi-source and session-portability benefits arrive.
Open questions
Key custody on the server. Per-agent keys live on the agent server, so a compromised server exposes every agent key it holds — but only the rights those keys were granted, for as long as the grants last, and all of it revocable at hop 1. That is a strictly better position than holding account keys. Hardware-backed or enclave custody can come later without changing the model.
Capability discovery. A daemon verifying a chain needs the capability blobs. Capabilities are already indexed and synced as structural blobs; the new part is that a daemon must fetch a hop-1 capability it has never seen when an unfamiliar agent key authenticates. Discovery by the delegate's principal is the right primitive.
Label and consent UX. Users should see, in plain language, what they are granting: "Agents on acme.example may read and write under /projects/alpha until January 2027." The label field already exists; a standard rendering of role + path + expiry should accompany it everywhere capabilities are listed.
Role granularity. READER, COMMENTER, WRITER, AGENT cover today's needs. A general scope language (per-operation grants, rate limits) is tempting and premature; roles plus paths plus expiry are enough to start, and they are what the SQL checks can evaluate cheaply.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime