−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS
Get started

Sandbox for LLM evaluation: run 1,000 model-written samples for about a cent

Scoring a model on coding tasks means executing thousands of programs nobody wrote by hand. Run them in disposable microVM sandboxes instead of on your laptop — per-second billing, with an honest note on throughput and CPU limits.

Coding benchmarks have a dirty secret: to score a model you have to run what it wrote. A pass@k run over a few hundred problems with ten samples each is thousands of freshly generated programs — and you'd never curl | bash a single one of them from a stranger.

Most people start on their laptop with a subprocess.run and a timeout. That works until a sample writes a 40 GB file, forks until the machine stalls, or quietly reads your home directory. The fix isn't a cleverer timeout. It's running the samples somewhere you don't care about.

The setup that works

A sandbox is a Firecracker microVM with Python 3.12, Node.js 22 and bash, started in about a second and deleted when you're done. Outbound internet works, inbound doesn't, and nothing of yours lives inside.

The trap is treating a sandbox like a function call. Each one is billed for at least 60 seconds and takes a second to start, so one sandbox per sample is slow and wasteful. Put the harness inside instead:

from eqvps import Sandbox

with Sandbox.create(tariff="standard", ttl=4 * 3600) as sb:
    sb.upload("/root/harness.py", open("harness.py").read())
    sb.upload("/root/samples.jsonl", open("slice_03.jsonl").read())
    task = sb.exec("python3 /root/harness.py /root/samples.jsonl > /root/results.jsonl",
                   background=True)
    task.wait(poll_interval=10)
    results = sb.download_text("/root/results.jsonl")

Inside, harness.py runs each sample in a subprocess with its own short timeout and records pass, fail or timeout. One misbehaving sample kills its own subprocess, not the run. If a sample manages to wreck the sandbox itself, you lose one slice, recreate the sandbox and continue.

Files go up to 5 MB per transfer, so split big datasets into slices — which you want anyway for parallelism.

Throughput, honestly

An account can hold up to 20 sandboxes, but 2 commands execute at the same time. A background task counts while it runs. So the shape of a fast eval is two sandboxes, each grinding through its slice in the background, not twenty sandboxes fighting over two slots.

For typical code benchmarks that's plenty: most samples finish in under a second, and one sandbox works through thousands in an hour. If you need to evaluate many models at once against a huge suite, you'll hit the ceiling. At that point a VPS running your own isolation is the better tool.

The CPU rule you should know about

Sandboxes are built for bursty work. One held at more than 90% CPU for over 15 minutes is treated as abuse — an ephemeral sandbox is stopped. Normal eval runs, which alternate between quick executions and bookkeeping, don't come near it. A benchmark that burns full CPU for hours (compiling a large project per sample, numeric stress tests) will. For that, take a VPS by the month: the AI agent plan gives 4 vCPU and 4 GB for $10.

What a run costs

You pay per second for the tariff's vCPU and RAM, 60 seconds minimum, from a prepaid balance.

WorkloadTariffTimeApprox. cost
1,000 short Python samplessmall (0.5 vCPU, 1 GB)~20 min$0.011
5,000 samples, numpy allowedstandard (1 vCPU, 2 GB)~2 h$0.13
Builds + tests per sampleplus (2 vCPU, 4 GB)~3 h$0.40

The model calls that generate those samples will cost you more than the execution — usually by an order of magnitude.

Keys and reproducibility

If the harness calls a model API from inside the sandbox — say, for LLM-as-judge scoring — pass the key as a sandbox environment variable. Values are stored encrypted, never logged, and the API only returns variable names.

For reproducibility, pin package versions in your harness and record the sandbox tariff with the results. Every sandbox starts from the same clean image, which removes the "worked on my laptop" variable from your numbers. Honestly, that alone is worth the switch.

Start from the sandbox page and the sandbox docs. New accounts get $1 of sandbox time — enough for a first eval run of a few thousand samples on the small tariff. The SDK reference covers background tasks and file transfers, and sandbox limits and billing has the full quota table.

Related: sandbox for AI agents and your first agent task in a sandbox.

Ready to deploy? Pay with crypto, no KYC — live in about a minute.

Deploy now →

FAQ

Why does an eval need a sandbox at all?

Because you're executing thousands of programs written by a model. Most are harmless, some loop forever, a few delete files or open network connections. On your workstation that's a risk to your data; in a sandbox it's a failed sample.

Should I create one sandbox per sample?

Usually not. Creating a sandbox takes about a second and is billed for at least 60 seconds, so one sandbox per sample wastes both. Run a harness inside one sandbox that executes many samples, each with its own timeout, and recreate the sandbox between models or when something breaks it.

How fast can I go?

An account runs at most 2 commands at the same time across up to 20 sandboxes. The practical pattern is a few sandboxes, each running your harness as a background task that works through a slice of the dataset.

Can I run a long benchmark for hours?

As a background task, yes, up to the sandbox lifetime (24 hours for an ephemeral one). But a sandbox held at full CPU for more than 15 minutes is treated as abuse. Bursty eval runs are fine; steady hours-long number crunching belongs on a VPS.

Can the sandbox call my model's API?

Outbound internet is open, so yes. Pass the API key as a sandbox environment variable — it's stored encrypted and never returned by the API — rather than pasting it into generated code.

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.