The problem
The agents server runs everything on one JavaScript event loop: the HTTP API, the WebSocket broadcast fan-out, the background poll loops, and the execution of every agent run and workflow. That loop is a single thread, so any CPU-bound stretch of run execution blocks the API and WebSocket for its whole duration.
Production profiling (bun's JSC sampler, symbolized through the build's source map) showed that under real load the loop is CPU-saturated: about 95% userspace JS, with one core pinned. /api/health, a handler that only replies, stalled for 6 to 23 seconds while 2 agent runs and 1 workflow were active. The costs are spread across the run path: CBOR encode and decode of session events, #piMessages re-decoding the whole session each turn, per-session serialization, and the QuickJS workflow VM.
Earlier PRs removed large fixed costs: a missing index (ListSessions 80× cheaper), native Ed25519 verify (idle CPU about 8× lower), and batched WebSocket broadcasts. Those made the server fast at rest. They cannot fix saturation under concurrent runs, because yielding does not help when a single thread pins its core. The structural fix is real parallelism: move run execution off the request-serving loop.
Why this takes a staged project
#executeAgentRun and #runPiAgent are about 1,150 lines. They reference 40+ distinct Service members (#db, #emit, #runQueue, session and plan management, sub-session spawning, trigger firing, title generation, and more) inside a 14,700-line Service class. Moving agent-run execution to another thread means either moving most of that class or building a message proxy for those 40+ methods. It also means coordinating SQLite writes and the microsandbox native addon across threads. That needs a staged migration.
Two facts make it doable:
Code execution already runs out of process. execute runs in microsandbox microVMs (separate msb processes), and the JS loop only orchestrates them with async I/O. So the work to isolate is the JS orchestration and CBOR/serialization. The sandboxes stay where they are.
Workflows already have a narrow effect boundary. The QuickJS workflow VM is a pure, synchronous compute context. Its only host interaction is a small, serializable WorkflowAdapters contract. That makes it the best first candidate. See script.
Target architecture
┌────────────────────────── main thread ──────────────────────────┐
HTTP / WS ─────► │ API handlers · WebSocket broadcast · RunQueue · SQLite writes │
│ ▲ │ │
│ effect results effect requests │
└─────────────────────────┼────────────┼────────────────────────────┘
│ ▼
┌──────────────── run worker(s) ────────┼────────────────────────────┐
│ CPU-bound execution: pi agent loop / QuickJS VM, CBOR, serialize │
│ no DB writes, no WS — every side effect crosses back as a message │
└────────────────────────────────────────────────────────────────────┘Main thread keeps everything shared and I/O-bound: the RunQueue and its leases, all SQLite writes, WebSocket broadcast, and trigger firing. It stays responsive because it does no CPU-bound run work.
Run workers do the CPU-bound execution and hold no shared state. Every side effect (append an event, spawn a child, execute a tool, update a plan) is a message to main, which performs it and replies. The workflow engine already uses this shape internally: compute in the worker, effects on the owner.
Effect bridge
The worker never touches the DB or the socket. It sends typed effect requests and awaits typed results:
worker → main { id, op: 'callTool', args } main → worker { id, ok, result }
worker → main { id, op: 'spawnAgent', args } main → worker { id, ok, result }
worker → main { id, op: 'appendEvent', args } main → worker { id, ok }
...Ordering is FIFO per run. Effects that must be durable before the run counts as advanced (event appends, journal entries) are acked. The worker flushes all pending acks before it reports a terminal outcome.
Cancellation
Run cancellation and the QuickJS fuel and interrupt check are synchronous and hot, so they cannot wait for a round trip. A one-byte SharedArrayBuffer per run carries the cancel flag. Main sets it, and the worker's interrupt handler reads it inline. There is no message latency and no polling.
SQLite
SQLite stays single-writer on main. Workers write nothing directly. Their effects become main-thread writes. A worker that needs a read-only query (rare) can use a separate read-only WAL connection, but by default all DB access is an effect. This avoids multi-writer coordination. See persistence.
microsandbox
Unchanged. execute stays a callTool effect, and main drives the microVM as it does today. The addon never loads in a worker.
Phased migration
Phase | Scope | Risk | Ships behind |
|---|---|---|---|
1 (this PR's POC) | Workflow QuickJS VM in a worker via proxy-adapters | Low: clean existing boundary, default-off flag, workflows only |
|
2 | Harden phase 1: worker pool/reuse, backpressure, crash recovery, metrics; enable by default | Low to med | flag flips default-on |
3 | Extract the agent-run effect surface: enumerate the 40+ | Med: API-shape-preserving refactor, no behavior change | internal |
4 | Run agent runs in workers over that contract, one worker per run, capped by the existing concurrency limits | High: the payoff and the hard part |
|
5 | Worker pool sizing, lifecycle, observability; enable by default; delete the in-process path | Med | default-on |
Each phase ships on its own behind a flag and is validated in production before the next. No phase changes the wire API. Message shapes and durable formats stay the same throughout, so every client keeps working, and so does a rollback to the in-process path.
Phase 1 proof of concept (in this PR)
The workflow VM goes first because runWorkflowVM(adapters) already takes its entire host interaction as a parameter. The POC runs that exact function, unchanged, inside a worker. It supplies a proxy WorkflowAdapters whose every method posts a message to main and, for the async ones, awaits the reply. Main supplies the real adapters it already builds in #executeWorkflowRun. The journal stays on main. The worker gets the loaded entries at start and streams appends back. Cancellation uses a SharedArrayBuffer.
Guarantees:
Default off. Without SEED_AGENTS_WORKFLOW_WORKER=1, execution is byte-for-byte the current in-process path. The worker path is opt-in and reversible per deploy.
Identical semantics. Same runWorkflowVM, same journal, same determinism and replay, same outcomes. The worker path only changes the transport.
Measurable win. With the flag on, a workflow's QuickJS compute no longer blocks the main loop. /api/health stays responsive while a workflow burns CPU in its worker.
This proves the effect bridge and SharedArrayBuffer cancel pattern end to end on the safest surface, so phases 3 and 4 (agent runs) build on a tested base.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime