The Seed Agents security model rests on five things: signed account-scoped actions, server-side owner and collaborator checks, encrypted and redacted secrets, signed WebSocket subscriptions, and a model-facing tool surface whose authority is fixed at a handful of verbs. For the Hypermedia side of trust, see integrity and permissions.
Trust boundaries
Desktop app ↔ agents server.
Agents server ↔ model providers.
Agents server ↔ Seed Hypermedia and content servers.
Agents server ↔ SQLite storage.
Desktop renderer ↔ desktop daemon signing API.
Agent runtime ↔ sandbox microVM (the only place model-written code runs).
HTTP action authentication
Every HTTP action is a signed SignedActionEnvelope, described in the signed API:
{
type: 'AgentsAction', signer, sig, account, action
}A signer is authorized when:
signer === account; or
account_authorizations has role OWNER or AGENT for (account, signer).
WebSocket authentication
The socket receives no private data until it sends a signed Subscribe action. The server checks that the subscription key belongs to the signed account. See WebSocket subscriptions.
A socket cannot switch accounts after a successful subscription.
Server-to-client events are not individually signed.
Agent collaboration authorization
An agent stays owned by agents.account_id. agent_collaborators adds explicit per-agent access for another signed Seed account, first as a pending invitation and then as an accepted reader or writer row:
pending invitees see only invitation metadata and cannot access agent contents;
readers can inspect all agent-scoped state, including memory, tools, prompts, sessions, transcripts, attachments, and runs;
writers can also change agent-scoped state and interact with sessions;
only the owner can invite or revoke members, or delete the agent;
two owner-set agent flags open the agent beyond membership: public_read makes every signed account a reader by agent id. public_chat (settable only while public_read is on, and cleared with it) promotes those public readers to chatter. Chatters may create, message, attach to, and stop sessions, but every writer-level action is still refused: agent, memory, tool, and trigger edits, session rename and delete, InvokeSessionTool, run cancel and signal. Use public chat to expose an agent to everyone. Use writer for people you trust to change it;
provider, secret, OAuth, and signing-identity changes stay scoped to the collaborator's own account. Optional agentId on provider and identity list actions exposes only the owner's redacted records needed to render and edit that shared agent.
Agent, session, and run WebSocket subscriptions use the same access check. Service events fan out to every accepted collaborator's account subscription. Revocation takes effect before the next request or subscription.
Account isolation
Account-owned tables include account_id:
model providers;
secrets;
agents;
agent collaborators (the invited account is also stored explicitly);
sessions;
runs;
tool documents;
action idempotency;
authorizations.
Session events do not store the account ID directly. The server verifies ownership through the parent session.
Agent isolation
Agents of the same account do not read each other's state. The thread: address (reads and listings) and continuation session_events sources reach only the calling agent's own threads. Memory and tools are per-agent by construction (state_dir). Until a deliberate inter-agent interaction model exists, agents communicate over public interfaces: documents and comments on the hypermedia network. They never inspect one another's transcripts.
Secrets
SetSecret accepts key bytes, encrypts them, and returns only redacted metadata. CreateSigningIdentity generates a new Ed25519 account key on the server for the agent, publishes its profile, and stores the raw seed through the same encrypted secret path. The owner's own key stays on their device and grants the agent a capability. ImportSigningIdentity accepts the seed of an existing key, decrypted on the client. Agent keys live on the server, so treat them as less secure than a personal identity. Use a key made for the agent, and self-host the agents server if you don't want a hosted server to hold it. ListSigningIdentities returns only redacted metadata for account-scoped secrets tagged with kind: 'hm-account-key'. Plaintext key material is never returned, and cross-account keys are not visible.
The shared client refuses to send secrets to non-local plain HTTP servers (isSafeAgentServerSecretTarget(), frontend/packages/ui/src/agents/client.ts). Remote servers must use HTTPS.
Do not log:
plaintext secrets;
decrypted API keys;
provider secret config;
signed request bodies;
full model prompts or responses;
full session content;
large or sensitive tool outputs.
Secret encryption limitations
Current key storage:
AES-GCM key lives in server_config in the same SQLite DB.
This prevents accidental disclosure through the API or logs. It does not protect against full DB compromise.
Future production work should add:
OS keychain or KMS-backed key storage;
key rotation;
secret versioning;
backup and restore guidance.
Provider endpoint safety
Each provider type has a code-owned spec in PROVIDER_SPECS (agents/src/api-service.ts) with a fixed default base URL and an allowCustomBaseUrl flag. The base URL is resolved by resolveProviderBaseUrl():
Pinned providers (openai, anthropic, google, openrouter, deepseek, groq, xai): the spec default base URL always wins, and any stored baseUrl override is ignored. This keeps a stored API key from being redirected to an arbitrary host.
Self-hosted and custom providers (ollama, custom): the user-supplied baseUrl is honored, because pointing at a local or private endpoint is the whole purpose. These accept requests without an API key. The base-URL value is set by the authenticated account owner through signed SetModelProvider actions, so the endpoint and any key it carries share a single trust owner. The desktop save flow still refuses to send an API key to a remote plain-HTTP agent server. Outbound SSRF to whatever a custom baseUrl names is accepted by design. Tightening this (for example, blocking link-local ranges for non-local custom endpoints) is future hardening.
Subscription ("Sign in with ChatGPT") providers hold OAuth credentials in place of an API key. Pi re-resolves and refreshes them per request through the persisted auth backend. An expired or revoked sign-in fails the run with an explicit re-auth message, never a cryptic provider 401 (api-service.ts:4262). The flow is off unless the operator opts in with SEED_AGENTS_SUBSCRIPTION_AUTH, because it needs a client that can catch the provider's localhost redirect (agents/src/config.ts:20).
Adding a pinned provider type needs only a PROVIDER_SPECS entry. It inherits the pinned-URL policy automatically. See model providers.
The tool surface is the authority boundary
The model sees the five verbs (read, write, call, delegate, plan) plus the two session verbs, status and continue_session, which touch only the session's own title and lineage. Only promotion widens that set, and promotion has two bounds (api-service.ts:4322):
promotion is derived only from durable tool_call events in this session's own transcript, so it survives restarts and live state cannot inject it;
the promoted list is intersected with enabledCallableTools() before it reaches Pi. This filter is a security control. A hallucinated or injected call {tool: 'bash'} durably stores that name, and an unfiltered allowlist would hand bash to Pi and activate Pi's own host bash/edit builtins outside the sandbox.
noTools: 'builtin' on the Pi session means Pi's own tool suite never loads. The only executable surface is the Seed-owned custom tools.
Events written by a user's own verb calls carry actor: 'user' and are explicitly skipped when computing promotion (api-service.ts:7237), so a user's palette activity never reshapes the agent's active toolset.
Grants
Two things are granted per agent, both stored in definition.tools. The verbs themselves are never grants.
The callable set: which of search, query, attributes, web_search, execute the agent may dispatch. execute also drops out when the host cannot run sandboxes.
Publish: the pseudo-tool publish (legacy write-group names still count). Without it, write to hm:// or ipfs:// returns 403 (api-service.ts:7454, api-service.ts:7499). Memory writes are never gated. That is the intended line: private files are the agent's workspace, and signed public content is a disclosure.
A delegate child's tools narrowing intersects against the parent's full callable set, so delegation can only reduce authority (api-service.ts:2623).
MCP servers: definition.mcpServers names the account MCP servers an agent may call. Enabling a server is a grant on par with execute: its tools run with whatever the remote server can do. The projected mcp tool documents only cache this grant. executeMcpTool re-checks definition.mcpServers before any call, so a stale document cannot reach a server the owner turned off. See mcp.md.
The promotion filter admits the enabled callable set and the agent's own enabled non-builtin documents (lambdas and MCP projections), re-derived from the definition at run start. A promoted tool document executes through the same call dispatch and the same checks as an explicit call.
MCP server safety
A connected MCP server is reached with account-configured URLs and headers from the agents host (agents/src/mcp.ts). Only http(s) remote transports exist. The service never spawns stdio processes.
Risks:
SSRF and private-network access: the same unmitigated posture as web reading. An owner can point a server record at any reachable address. Loopback is allowed on purpose, because local MCP proxies are common in development.
Tool authority: the remote server decides what its tools do. Model-driven calls carry model-authored arguments. A server that acts on the world (files, issues, payments) should be enabled only for agents whose prompts warrant it.
Prompt injection: tool results are untrusted text, exactly like fetched web pages. See the prompt injection map.
Mitigations present:
header values are encrypted account secrets, redacted from every response. The desktop refuses to send one to a non-HTTPS remote agent server;
server names are slugs and tool names are sanitized to [A-Za-z0-9_-] and capped at 64 characters, so a remote name can never collide with a verb, shadow a builtin, or break a provider's tool-name rules. An authored lambda keeps its name against a remote tool of the same name;
input is validated against the projected contract before a call leaves the host, and a miss returns the contract;
results are bounded (256 KiB text, 4 MiB per inline image) and server errors become tool_result.error;
connections are per run and closed with it. A call has a 120s timeout and a connect has 20s;
deleting a server scrubs it from every agent and deletes the header secrets it owns.
Agent-managed triggers
Agents manage their own triggers directly. write ~/triggers/<name> creates, edits, enables, disables, or deletes a trigger, and enabled is honored exactly as written (defaulting to true). This is a deliberate product decision by the owner (2026-08-19): "do this every morning" said in chat should just work, with no separate approval step in the desktop.
The threat model consequence: a trigger is standing authority to act with nobody present. An agent can be steered by a prompt injection in content it reads, and it can now grant that authority to itself. An earlier event-bus design (now only in git history) proposed gating activation on a user gesture. That gate was built and then removed on the owner's direction. The remaining mitigations give visibility. None of them prevent it:
every trigger write is a durable, actor-stamped tool_call and tool_result pair on the session Log;
trigger writes emit trigger-updated account events, so the desktop Triggers tab reflects changes live;
a trigger fires only what the agent could already do. Its callable set and publish grant still bound the blast radius, and delegation still only narrows authority;
the firing-chain loop guard (TRIGGER_CHAIN_MAX_HOPS) still stops runaway trigger chains.
If consent is ever wanted back, the enforcement point is writeTriggerAddress and the history is in git.
Framing injection (<user_action>, <plan_state>)
Two model-facing frames carry text the model or a fetched page authored, handed back inside tags whose syntax the model knows. Both rewrite every < as its unicode escape through escapeActionFraming() (api-service.ts:9179). The escape stays valid inside JSON and still renders as < to a human reader, but can never form a tag:
<user_action> and <user_action_result>: the user's own verb calls and their results, including fetched web content. Without escaping, a page containing </user_action_result> could close the frame and forge trusted user actions for everything after it.
<plan_state>: the live checklist injected fresh each turn. Step ids and labels are whatever the model last wrote, so a label carrying </plan_state> would otherwise turn the rest of the block into instructions nothing vouched for (api-service.ts:409).
A related guard: an actor-less tool_result answering a user-actor tool_call (a synthetic written before the synthesizer knew about actors) is dropped, and never replayed as an orphan provider tool result (api-service.ts:4868).
InvokeSessionTool itself is bounded: only read, write, and call are accepted, it is rejected with 409 while the session has a live run, and execution failures append to the log where anyone can see them (api-service.ts:2345).
read safety
The read verb reaches memory, tool contracts, hypermedia, IPFS, the public web, the activity feed, attachments, other threads, and run records.
thread: reads and the bare thread: listing are scoped to the calling agent: an agent reads only its own conversations. run: reads are scoped by account_id alone, so an agent can read the run records of every other agent on the same account, including their tool inputs and results. read ~/self exposes only what the account owner already configured: definition, grants, signing-key names. It never exposes secret material or provider keys. Decide whether run records should also be scoped to the agent before treating agents on one account as isolated from each other.
Risks:
SSRF and private-network access if the server runs in a sensitive network;
model-driven reads of arbitrary web resources;
large or sensitive tool outputs.
Mitigations present:
address parsing rejects anything outside the supported forms;
output size limit (256 KiB) on every tool result;
durable, actor-stamped tool events for everything the agent and user do;
hypermedia-first resolution for https:// addresses. It falls through to the web reader only on an explicit not-hypermedia marker or a 404, so scraped HTML never silently replaces document content;
no shelling out to the CLI.
Future mitigations:
outbound allow and deny policy;
private-network blocklist;
audit log.
Web reading safety
Reading https:// addresses performs a server-side fetch (the static tier) and, when configured, has Crawl4AI fetch the page in a real browser (agents/src/web-tools.ts). web_search is backed by self-hosted SearXNG. Neither carries a third-party API key.
Risks:
SSRF and private-network access. On a host with access to a private network or cloud metadata endpoint, a model could request internal addresses. This is unmitigated beyond http(s) scheme enforcement. Before exposing this on a sensitive network, add a private-network and metadata blocklist and an outbound allow and deny policy.
Model-driven retrieval of arbitrary web content into the conversation is the main prompt-injection surface. Treat fetched page text as untrusted input, never as instructions.
Crawl4AI executes a real browser. Keep it on the internal network only, and never publish it to the host or internet.
Mitigations present:
http and https scheme enforcement and URL validation;
bounded markdown output (200 KiB, truncated on a byte boundary);
raw mode refuses non-text content types;
the Crawl4AI shared token (SEED_AGENTS_CRAWLER_TOKEN) gates the crawler so only the agents service can use it;
web_search is a granted callable (it is not an always-on verb) and throws a clean error when no SearXNG backend is configured;
failures degrade to tool_result.error (or a partial flag for incomplete search coverage). Nothing is silently made up.
Agent memory safety
Agent memory (agents/src/agent-memory.ts) exposes a real filesystem directory to model-controlled input, so path handling is strict:
every path is validated before use: string-only, no null bytes, no .. segments, backslashes normalized, depth (16) and length (512 bytes) bounded, and the resolved absolute path re-verified to sit inside <stateDir>/memory;
symlinks are refused as read and write targets and skipped in listings, so memory operations cannot follow a planted link out of the sandbox;
ownership is checked through the agent row before any filesystem operation, so accounts cannot touch other accounts' memory;
binary memory files are never sent to the model: a memory read returns only metadata for binary content. Their raw bytes are returned to the owning user over the signed API for preview and download;
memory content is visible to the model and the user by design. Do not store secrets in agent memory;
write ipfs:// and UploadAgentMemoryFileToIpfs chunk files as UnixFS and publish the blocks through the typed HM API. The daemon stores the blocks, but its gateway refuses them (blob … is not public) until public Hypermedia content links to the CID. DagPB visibility propagates from a referencing Change, Comment, or Profile (backend/storage/schema.sql visibility rules). So a standalone write ipfs:// is not retrievable yet. It becomes retrievable the moment any public blob references it. SEED_AGENTS_IPFS_SERVER_URL only selects the gateway for later reads. Treat publishing as irreversible disclosure anyway. Showing a memory file to the owner in chat does not go through IPFS: the transcript renders  over the signed API;
session attachments (files dropped into the chat composer) are session-private: stored under <stateDir>/session-attachments/<sessionId>/, keyed by content SHA-256, capped at 100 MiB each, never auto-copied into cross-session memory or published, deleted with the session, and exposed to the model as metadata until it reads one by address.
Memory has no size quota, and never has. Earlier revisions of this document described a 1 MiB text write cap, a 100 MiB per-file cap, a 1 GiB per-agent total, and a 2000-entry limit. None of those exist in the code, at HEAD or in the commit that introduced the memory filesystem (8326a22e7). agent-memory.ts bounds only path length and depth. downloadToMemory allows downloads of any size and aborts only on a 60-second stall (agent-memory.ts:290). The sandbox mount is created with mount.bind(memoryRoot) and no .quota() call (code-exec.ts:402), even though the builder supports one. Disk exhaustion by a runaway model write, download, or sandbox program is unbounded today. The only enforced size limits nearby are per-attachment (100 MiB, session-attachments.ts:26) and per-chunked-upload (2 GiB, api-service.ts:158).
Code execution safety
execute and every authored lambda run model-written code, so isolation comes from hardware virtualization. Process sandboxing is not used (agents/src/code-exec.ts):
each execution runs in a fresh ephemeral microVM (embedded microsandbox runtime: libkrun on macOS/Linux, WHP on Windows) with the restricted in-guest security profile. The VM boundary is the isolation line. Seccomp and containers are not;
the only host filesystem exposure is the agent's own memory directory, bind-mounted at /workspace. Code cannot see other agents' memory, the SQLite DB, or secrets;
guest-created symlinks inside the memory directory cannot trick host-side reads. Memory reads refuse symlinks and listings skip them;
nothing is interpreted by a shell unless the runtime is the shell. The sandbox takes an argv array, so model code containing quotes, newlines, or $ needs no escaping (code-exec.ts:470);
sandbox networking is on by default but constrained to a non-local egress policy (NetworkPolicy.fromProfiles(['public']), with nonLocal() as the older-SDK dialect, code-exec.ts:117). Code reaches the public internet but not the host's private network or cloud-metadata endpoints. DNS uses an explicit resolver set (SEED_AGENTS_EXEC_DNS) and never the host's. SEED_AGENTS_EXEC_ALLOW_NETWORK=false removes the NIC entirely. On-by-default egress widens the exfiltration surface compared to an offline sandbox. The isolation is the non-local policy plus the memory-only mount. There is no air gap;
CPU count, guest memory, per-exec timeout (clamped to ≤ 300s) and total sandbox lifetime (timeout + 30s) are capped server-side. stdout and stderr are bounded to 64 KiB each before reaching the model;
a lambda's input is baked into its program as a double-JSON.stringify literal, so no call input can escape into code (code-exec.ts:505);
an authored lambda rides on the same grant the execute tool needs (api-service.ts:7696). Without that check, writing a tool document would be a way around an owner who turned code execution off;
resource note: each concurrent execution boots a microVM with its configured guest memory. There is no per-account concurrency limit yet.
Script (workflow) safety
Script children are untrusted, model-authored JavaScript. The posture is defense in depth (agents/src/workflow-host.ts):
Zero-ambient-authority realm: each run gets a fresh QuickJS-WASM context with no Date, Math.random, timers, fetch, imports, or process access. A submission-time lint rejects those tokens up front, and the realm removes them at runtime. The only way to affect the world is the journaled ctx bridge.
Every effect is validated, bounded, and journaled: ctx.call is checked against the read and write verbs plus the agent's enabled callables (api-service.ts:3801) and the tool's input schema. Results are size-bounded by the tool caps. The journal is a flight recorder: you can list every external effect after the fact via GetRunJournal.
No new authority: a script can do exactly what its agent could do call-by-call in chat, under the agent's own signing identities and configured HM server. The new factor is scale, bounded by spawn depth (3), fan-out (10 children per run), the separate workflow concurrency pool, compute fuel between awaits, VM memory, and journal caps.
Child outputs re-enter parents as data (schema-validated when a typed result was declared), inside tool results. They are never trusted instructions. A prompt-injected child can corrupt only its own return value.
Kill switch: CancelRun on any root cascades to every descendant. Queued runs never start, waiting runs never wake, live agent runs abort through Pi, live script VMs are interrupted. StopSession on the launching chat does the same for its whole tree.
Accepted gaps: there are no cost (dollar or token) budgets yet. Wall-time, depth, fan-out, and concurrency caps are the blast-radius controls (live usage is persisted per run and visible to clients). A ctx.call interrupted between execution and its journaled result re-executes on resume (at-least-once). That is fine for idempotent tools, but a write that crashed at exactly that point could publish twice. Idempotency keys are the roadmap fix.
Honest-record guarantees
Several behaviors keep the log from quietly disagreeing with reality. That is a security property and a product property:
the runtime settles a plan step only on evidence (every attached child succeeded) and never derives anything from failure (api-service.ts:2723);
resolvedBy: 'runtime' cannot be forged from model input and is carried across rewrites only while the step stays done (api-service.ts:2174);
a run that exhausts its continuations leaves an actor-system notice naming exactly what was left open, and nothing is ticked off on the agent's behalf (api-service.ts:2659);
a typed child that never delivered fails, because its parent is blocked on a result that is never coming.
Replay protection status
Implemented:
idempotency for create and message actions with client IDs;
every signed action carries a signed action.ts Unix epoch millisecond timestamp;
HTTP and WebSocket envelopes are rejected when action.ts is missing, invalid, or more than five minutes from server local time (MAX_ACTION_CLOCK_SKEW_MS in agents/src/auth.ts). The window bounds both clock skew and time spent queued behind a busy server.
Not implemented:
nonce caching, so a captured request can still be replayed within the five-minute timestamp window.
Nonce caching remains a high-priority hardening project.
Logging security
Recent diagnostic logs are designed to include:
account, agent, session, and run IDs;
partial IDs;
event counts;
byte lengths;
status codes;
content types;
durations;
active tool names and model and provider identity.
They should not include secret values or full message content. Keep future logs at this level unless doing explicit local-only debugging.
Security checklist for new work
For every new action:
Verify signature.
Verify signer authorization.
Normalize inputs at the boundary.
Scope DB queries by account ownership.
Redact sensitive data.
Add unauthorized and cross-account tests, including reader-vs-writer and pending-vs-accepted behavior for agent-scoped actions.
Decide idempotency and replay semantics.
Decide WebSocket fanout policy.
Update docs.
For every new tool or address form:
Add or update the canonical registry entry in agents/protocol/src/tool-registry.ts so prompt metadata, input schema, and rendering metadata are reviewed together.
Decide the grant: callable set, publish, or ungated.
Validate inputs at the runtime boundary and return the contract on a miss.
Bound output size.
If it can be promoted, confirm the promotion filter still holds.
If model-authored text re-enters the prompt inside a frame, escape it.
Avoid sensitive logs.
Add tests for missing credentials and provider or tool errors.
Update security.md, model-providers.md, or tools.md.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime