This plan grows the agents service past one box in three phases. Each phase can stop and hold. It is written against the 2026-08-29 production baseline: one 4-vCPU host running agents-stable/-staging/-dev, Caddy, SearXNG, and Crawl4AI, saturated by a single heavy dev agent. The finished perf-squeeze plan, now in git history, measured this.
Why this shape
All durable state is per-account. It is SQLite rows and a state directory, both keyed by account. Nothing global needs a shared database, so scaling means sharding by account plus stateless helpers. There is no distributed-consensus problem. The expensive, bursty work (sandbox microVMs, headless Chromium) also separates cleanly from the latency-sensitive work: the signed API, WebSocket streaming, and trigger polling.
Phase 1: one box, contained (now)
Set Compose CPU limits per container. Move Crawl4AI to its own small node when memory gets tight. The architecture does not change. This holds while total sandbox demand fits in about 2 to 3 dedicated vCPUs.
Phase 2: split execution from the control plane
Add an exec host: a cheap dedicated-vCPU KVM machine running a thin daemon. The daemon exposes the existing CodeExecutor contract (execute(request) → result plus availability) over HTTP, with a bearer token on the internal network. SEED_AGENTS_EXEC_BACKEND=remote plus a URL selects it. The microsandbox embedding, the watchdog, and the planned long-lived VM pool all sit behind the same interface, so agents code does not change. See the code execution backend in operations.
The control plane is small and steady again: API, WS streaming, polling, SQLite.
Exec capacity grows by adding exec hosts. The control plane round-robins agents across them, with sticky assignment per agent so the long-lived VM pool stays warm.
The agent's memory workspace must reach the exec host. Sync directory snapshots on demand (rsync-style, content-addressed) and skip a network filesystem. Executions are bursty and local, and the workspace is already the durable copy.
This phase lets ion "ramp up" without touching anyone's chat latency. One exec host is about one more ion running flat out.
Phase 3: shard the control plane by account
When one control plane saturates (thousands of accounts, or heavy WS fan-out), run N agents servers. Each owns a subset of accounts with its own SQLite and data dir. A routing layer (a Caddy map on account id, or a 20-line router service) sends signed envelopes and WS subscriptions to the owning shard. To move an account, copy its rows and state dir, then flip the route. There is no shared state and no cross-shard traffic. Triggers and runs are already scoped to accounts.
Scaling and cost model (Hetzner-style pricing, 2026)
Role | Machine | ~€/mo | Capacity |
|---|---|---|---|
Control plane | 4 shared vCPU / 8 GB (CPX31) | ~16 | thousands of accounts; light, steady CPU |
Exec host | 4 dedi vCPU / 16 GB (CCX23) | ~25 | 3 to 4 concurrent heavy sandboxes (1 vCPU each) |
Crawl node | 2 vCPU / 8 GB | ~8 | SearXNG + Crawl4AI |
Rules of thumb:
Cost scales with concurrent sandbox vCPUs. The number of accounts barely matters. Budget roughly €6 to €8 per month per always-on concurrent sandbox vCPU. An idle agent costs pennies (rows in SQLite). A compiling agent costs a vCPU.
Provider tokens cost far more than infrastructure in every phase. One ion-class agent at about 220 requests per hour spends more on tokens per day than an exec host costs per month. So infra should never be the reason to throttle an agent that the model budget wants running. See model providers.
Phase 2 with one exec host costs about €50/mo total and supports today's load with an ion running continuously. Each further ion-equivalent adds about €25/mo. Phase 3 adds about €16/mo per control-plane shard, needed only at large account counts.
What we avoid on purpose
Shared or networked SQLite, or a move to Postgres "for scale". Per-account sharding makes it unnecessary far beyond current plans.
Kubernetes. Three roles with Compose and Terraform stay easy to read. Revisit only if exec hosts are autoscaled.
Scheduling sandboxes on the control plane "because there is spare CPU". That spare CPU is the latency budget for everyone's triggers and streams. Phase 2 exists to protect it.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime