Compare top sandbox providers on the same target hardware shape for real developer and CI/CD workloads.
Leaderboard · Methodology · Dataset · Architecture · Contributing
4 vCPU · 8 GiB RAM · 40 GB disk
We measure the end-to-end time that developers and agents actually experience when using a sandbox to complete software engineering tasks. Going from a ticket to a PR is a multi-phase workflow - clone a repo, install dependencies, lint, build, test, etc.
A sandbox provider can top a creation time or CPU performance chart and still lose badly on:
- dependency installation is thousands of small, random file writes, and a network-attached or bandwidth-capped disk turns that into the longest step of your run.
- cloning a repo has the opposite profile: mostly sequential writes, bounded by network.
- single-threaded developer tools are limited by single-thread CPU not threads
- isolation technology and in-sandbox Docker capability
Our synthetic benchmarks use versioned Phoronix Test Suite profiles published through OpenBenchmarking.org, the long-standing standard for Linux hardware and software benchmarking.
Each workload has an inspectable definition for installation, arguments, repetition, parsing, units, and result direction, and can be reproduced independently outside this repository. The profiles are vendored, their metric definitions are generated rather than transcribed, and CI rejects drift. Custom instrumentation is limited to measurements PTS cannot represent: provider lifecycle latency, sandbox capabilities, pricing, and complete developer workflows. Those workflows are authored as repo-local PTS profiles, so first-party workloads inherit the same execution and parsing model as upstream benchmarks.
Each result is produced from fresh, independently created sandboxes. Workloads, toolchains, arguments, and target resources are pinned; raw samples and observed machine properties are retained; normalized runs are schema-validated before publication.
Within-sandbox passes and between-sandbox replicates are tracked separately. Failed, missing, unsupported, and resource-mismatched results are disclosed rather than treated as zero or silently excluded.
Read the full methodology.
mise install # pinned non-Bun tools (typos, shellcheck, hadolint)
bun install --frozen-lockfile
bun run typecheck
bun run test
bun run lintThe full command contract, workspace layout, and enforced dependency DAG are in Architecture.
Provider benchmarks, dataset publication, and toolchain releases require protected credentials and run only from maintainer-controlled workflows; pull requests never receive provider secrets. See CI & secrets.
Contributions must preserve three invariants:
- Every provider performs equivalent work.
- Every number is traceable to raw samples and exact workload provenance.
- Missing or non-comparable results remain visible.
See CONTRIBUTING.md and SECURITY.md.