Home / Benchmarks
Benchmarks
All numbers on this page were measured on 2026-09-26 on a DigitalOcean
m-16vcpu-128gb node (16 vCPU, nested KVM, nyc3), driving the API
with the e2b Python SDK 2.40.0. Two self-hosted Firecracker substrates were
exercised: AgentENV 0.2.2 and CubeSandbox v0.7.2. Methodology and analysis:
docs/SUBSTRATE-BAKEOFF.md and
docs/FINDINGS.md
(branch main). Raw logs:
results/ in the same branch.
Sequential lifecycle (n=20)
| Operation | p50 | p99 |
|---|---|---|
| create | 136 ms | 160 ms |
| exec | 98 ms | 150 ms |
| pause | 186 ms | 274 ms |
| resume | 161 ms | 184 ms |
| exec after resume | 78 ms | 92 ms |
| kill | 129 ms | 150 ms |
CubeSandbox v0.7.2, base template, run 164241. A pause keeps memory: a Python variable set before pause survived a later connect.
| Operation | p50 | p99 |
|---|---|---|
| create | 97 ms | 133 ms |
| first exec | 70 ms | 84 ms |
| pause | 56 ms | 64 ms |
| resume | 93 ms | 102 ms |
| kill | 31 ms | 36 ms |
AgentENV 0.2.2 (Firecracker), ubuntu:22.04 template, results/20260926_040606-*.
Idle density on one 16-vCPU node
| Alive (idle) | Host CPU busy | Host RAM / sandbox | exec p50 | exec p99 |
|---|---|---|---|---|
| 250 | 25.8% | 19.0 MB | 190 ms | 220 ms |
| 500 | 43.0% | 19.6 MB | 237 ms | 382 ms |
| 1,000 | 81.6% | 20.8 MB | 323 ms | 2,248 ms |
| 1,500 | 100% | 19.9 MB | 48.4 s | 52.1 s |
- CubeSandbox,
basetemplate, runs 164241 + 165002. ~20 MB of host RAM per idle sandbox; the usable single-node ceiling is ~1,000 idle — exec p99 exceeds 2 s at 1,000 and the host saturates at 1,500. - AgentENV reference (same host shape, guest DAMON off): 1,000 idle at 34% CPU, exec p50 49 ms, 34–41 MB host RAM per sandbox.
- Pausing everything returns memory: 1,093 paused sandboxes in 75 s, host usage 42.5 GB → 9.9 GB.
- Derived idle cost: ~$0.001 per idle sandbox-hour on a $1/h node.
100 concurrent realistic users, 10 minutes
| Action | n | p50 | p99 | errors |
|---|---|---|---|---|
| create | 100 | 786 ms | 916 ms | 0 |
| shell | 312 | 43 ms | 198 ms | 65 — git missing in this template (template gap, not platform) |
| file r/w (64 KB) | 317 | 46 ms | 124 ms | 0 |
| python script | 312 | 223 ms | 486 ms | 0 |
| internet check | 100 | 1,005 ms | 6,292 ms | 0 |
| pause | 582 | 17.7 s | 39.0 s | 0 |
| resume | 582 | 882 ms | 3,252 ms | 1 (10 s timeout) |
| resume check (file intact) | 581 | 120 ms | 384 ms | 0 |
| kill | 99 | 392 ms | 4,939 ms | 1 (30 s timeout) |
- CubeSandbox,
codetemplate, run 172511. Settings: 60 s ramp, 600 s per user, 5–60 s think time; 60% of users take a long idle pause/resume cycle; 10% run heavy bursts; every user runs an internet check. - 0 create errors; 2 platform errors in ~4,800 actions (~0.04%). Host peak: 79% CPU, 11 GB RAM.
- Honest caveat: concurrent pause is slow (p50 17.7 s vs 186 ms sequential) — memory snapshots queue on disk. The gateway will make pause asynchronous and rate-limit idle auto-pause.
300 concurrent real coding agents, 7 harnesses
| Concurrency | Runs | Passed | Infra errors | Agent time p50 / p90 |
|---|---|---|---|---|
| 50 | 56 | 55 | 0 | 77 s / 100 s |
| 100 | 105 | 104 | 0 | 130 s / 175 s |
| 200 | 203 | 201 | 0 | 186 s / 299 s |
| 300 | 301 | 298 | 0 | 302 s / 489 s |
- AgentENV, template
agents-2g(2 vCPU / 2 GiB). Task: fix real bugs, pytest oracle. 7 harnesses interleaved: Claude Code, Codex CLI, opencode, aider, cline, grok-cli, mini-swe-agent — bring your own model key; runs used a self-hosted Qwen3.6-35B on one H100. - 665 runs, 658 passed (98.9%), zero infrastructure errors. All 7 failures were the agent ending with tests still failing — model quality, not the sandbox.
- Replication on CubeSandbox at c300: 301 runs, 299 passed; the 2 failures were host-CPU-saturation timeouts in
cline auth, not sandbox faults. - Bottleneck was the 16-vCPU sandbox host, not the GPU; throughput plateaued at ~1,900 tasks/h.
Reproduce: benchmark code (bench/), raw results/ logs and both analysis docs are in the
agent-infra repo on branch
main.