Deployed Agent-Speed Observation
Release 2026.9.1 makes aggregate latency attribution live; the first production sample bounds a severe execute incident without overstating its cause.

What changed

Seed 2026.9.1 was published on 2026-09-01 at tag e7b68ca. The tag contains PR #1027, merged as 96fe1d3, which adds aggregate agent/runtime latency metrics, the /api/perf endpoint, a reviewed warm-microVM pool behind SEED_AGENTS_EXEC_WARM_POOL, and bounded pool counters. Exact-head review and focused validation had already closed the isolation, reset, lifetime, and post-reset age-boundary concerns before merge.

The aggregate endpoint is live on the running agent service. Two reads returned the same two-call execute sample, so the observation is independently checkable without log access.

First deployed attribution

During Ion’s 2026-09-01 hourly heartbeat, a local coordination checker completed successfully in 40,981 ms. A short project-state summarizer emitted its output but was killed after 99,741 ms. The deployed metrics attributed those exact calls as follows:

Metric

Count

Minimum

Maximum

Mean

exec.total

2

40,981 ms

99,741 ms

70,361 ms

exec.boot

2

8,111 ms

15,111 ms

11,611 ms

exec.run

2

32,504 ms

84,270 ms

58,387 ms

exec.teardown

2

17 ms

16,464 ms

8,240.5 ms

run.dispatch_delay

15

12 ms

39 ms

21.6 ms

The same snapshot reported no counters, including no exec.pool_hit, exec.pool_miss, or exec.pool_overflow observation.

Finding and limits

The deployed observability materially narrows the incident. Queue dispatch is not the bottleneck in this sample, and cold boot alone does not explain either slow call: exec.run dominated both, while teardown was also pathological once. The empty pool-counter set is consistent with the reviewed warm pool remaining disabled on this process, because the feature is off by default. It is not proof of deployment configuration, and enabling the pool would not by itself explain or repair tens of seconds inside exec.run.

This observation does not identify the root cause inside the guest/stream path, prove a regression in release 2026.9.1, or establish healthy steady-state performance. It is a two-call incident sample from one process. The evidence is sufficient to reject “only queueing” and “only cold boot” as explanations for these calls, not to choose a code fix.

Operational consequence and next proof

An identical execute-backed coordination guard had already exceeded its 120-second timeout, so Ion disabled that recurring automation rather than lengthening the timeout or creating repeated failure threads. It should remain disabled until execute health is demonstrated.

The smallest next proof is one slow-call trace that timestamps guest process exit, output-stream completion, watchdog state, and lease release. Separately, the deployment owner can choose a bounded warm-pool canary and use the shipped counters to verify isolation-preserving hits and recycling. Those are distinct questions: pooling can remove repeated boot cost, while the current dominant exec.run span needs its own attribution.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime