Coding benchmarks have a dirty secret: to score a model you have to run what it wrote. A pass@k run over a few hundred problems with ten samples each is thousands of freshly generated programs — and you'd never curl | bash a single one of them from a stranger.
Most people start on their laptop with a subprocess.run and a timeout. That works until a sample writes a 40 GB file, forks until the machine stalls, or quietly reads your home directory. The fix isn't a cleverer timeout. It's running the samples somewhere you don't care about.
The setup that works
A sandbox is a Firecracker microVM with Python 3.12, Node.js 22 and bash, started in about a second and deleted when you're done. Outbound internet works, inbound doesn't, and nothing of yours lives inside.
The trap is treating a sandbox like a function call. Each one is billed for at least 60 seconds and takes a second to start, so one sandbox per sample is slow and wasteful. Put the harness inside instead:
from eqvps import Sandbox
with Sandbox.create(tariff="standard", ttl=4 * 3600) as sb:
sb.upload("/root/harness.py", open("harness.py").read())
sb.upload("/root/samples.jsonl", open("slice_03.jsonl").read())
task = sb.exec("python3 /root/harness.py /root/samples.jsonl > /root/results.jsonl",
background=True)
task.wait(poll_interval=10)
results = sb.download_text("/root/results.jsonl")
Inside, harness.py runs each sample in a subprocess with its own short timeout and records pass, fail or timeout. One misbehaving sample kills its own subprocess, not the run. If a sample manages to wreck the sandbox itself, you lose one slice, recreate the sandbox and continue.
Files go up to 5 MB per transfer, so split big datasets into slices — which you want anyway for parallelism.
Throughput, honestly
An account can hold up to 20 sandboxes, but 2 commands execute at the same time. A background task counts while it runs. So the shape of a fast eval is two sandboxes, each grinding through its slice in the background, not twenty sandboxes fighting over two slots.
For typical code benchmarks that's plenty: most samples finish in under a second, and one sandbox works through thousands in an hour. If you need to evaluate many models at once against a huge suite, you'll hit the ceiling. At that point a VPS running your own isolation is the better tool.
The CPU rule you should know about
Sandboxes are built for bursty work. One held at more than 90% CPU for over 15 minutes is treated as abuse — an ephemeral sandbox is stopped. Normal eval runs, which alternate between quick executions and bookkeeping, don't come near it. A benchmark that burns full CPU for hours (compiling a large project per sample, numeric stress tests) will. For that, take a VPS by the month: the AI agent plan gives 4 vCPU and 4 GB for $10.
What a run costs
You pay per second for the tariff's vCPU and RAM, 60 seconds minimum, from a prepaid balance.
| Workload | Tariff | Time | Approx. cost |
|---|---|---|---|
| 1,000 short Python samples | small (0.5 vCPU, 1 GB) | ~20 min | $0.011 |
| 5,000 samples, numpy allowed | standard (1 vCPU, 2 GB) | ~2 h | $0.13 |
| Builds + tests per sample | plus (2 vCPU, 4 GB) | ~3 h | $0.40 |
The model calls that generate those samples will cost you more than the execution — usually by an order of magnitude.
Keys and reproducibility
If the harness calls a model API from inside the sandbox — say, for LLM-as-judge scoring — pass the key as a sandbox environment variable. Values are stored encrypted, never logged, and the API only returns variable names.
For reproducibility, pin package versions in your harness and record the sandbox tariff with the results. Every sandbox starts from the same clean image, which removes the "worked on my laptop" variable from your numbers. Honestly, that alone is worth the switch.
Start from the sandbox page and the sandbox docs. New accounts get $1 of sandbox time — enough for a first eval run of a few thousand samples on the small tariff. The SDK reference covers background tasks and file transfers, and sandbox limits and billing has the full quota table.
Related: sandbox for AI agents and your first agent task in a sandbox.
Comments
No comments yet. Be the first.