A diff shows what a patch claims to do. It doesn't show whether the bug is actually gone, whether the new branch of an if ever runs, or whether the dependency bump breaks import on a clean install. Reviewers know this, which is why so many reviews end with "LGTM, assuming tests pass".
The gap gets wider with AI-written patches. A model produces plausible code fast, and plausible is exactly the kind that slips past a tired reviewer. The cheap fix is to run the change before anyone approves it — somewhere it can't hurt anything.
The check: base, head, tests
The most convincing evidence a patch can offer is a reproduction that fails on the old code and passes on the new one. A sandbox makes that a 20-line script. It's a Firecracker microVM with Python 3.12, Node.js 22, git and curl, started in about a second:
import os
from eqvps import Sandbox
REPO, BASE, HEAD = os.environ["REPO_URL"], os.environ["BASE_SHA"], os.environ["HEAD_SHA"]
def repro(sb, ref):
sb.exec(f"cd /root/app && git checkout -q {ref}", timeout=55)
return sb.exec("cd /root/app && python3 /root/repro.py", timeout=55).exit_code
with Sandbox.create(tariff="standard", ttl=1800) as sb:
sb.exec(f"git clone -q {REPO} /root/app", timeout=55)
sb.exec("cd /root/app && pip install -q -r requirements.txt", timeout=55)
sb.upload("/root/repro.py", open("repro.py").read())
before, after = repro(sb, BASE), repro(sb, HEAD)
tests = sb.exec("cd /root/app && python3 -m pytest -q", background=True).wait()
print(f"repro on base: {before}, on head: {after}; tests: {tests.state}, exit {tests.exit_code}")
repro.py is the snippet from the bug report. If it exits non-zero on base and zero on head, the patch fixes what it says it fixes. Post that line, plus the test summary, as a comment on the pull request, and the reviewer starts from facts.
The test suite runs as a background task because a single synchronous command stops at 55 seconds. The TTL of 30 minutes caps how long a stuck run can bill you.
Review agents that run code
The same idea works for an AI reviewer. Connected through the MCP server, an agent has sandbox tools: create a sandbox, run a command, upload and download files. Instead of writing "this might break Python 3.8 compatibility", it checks out the branch, runs the code and quotes the traceback — or reports that it's fine.
Give such an agent a spending cap and a read-only token, and decide which tools it may call without asking. MCP guardrails lists every tool by risk.
Which tariff
| Repository | Tariff | Typical pass | Approx. cost |
|---|---|---|---|
| Small library, pure Python or JS | small (0.5 vCPU, 1 GB) | 2 min | $0.0011 |
| Web app with a test suite | standard (1 vCPU, 2 GB) | 5 min | $0.0055 |
| Native extensions, compiled deps | plus (2 vCPU, 4 GB) | 15 min | $0.033 |
Ephemeral sandboxes bill per second with a 60-second minimum. If a reviewer wants to come back to the same environment over a couple of days, create it as persistent instead: it keeps its disk for up to 30 days and is billed per started hour, so a standard sandbox kept for a two-day review costs about $3.17. Delete it when the pull request is merged.
Limits worth knowing
- Two commands at a time per account. Plenty for reviews, which run one after another. If you check dozens of pull requests at once, they queue.
- No inbound ports. You can't open the app in a browser. Start it inside and test it with
curl localhostfrom a second command. - No Docker inside. Tests that spin up containers belong on a VPS runner.
- Outbound internet is open. That's what makes
pip installwork. Don't put anything into the sandbox you wouldn't hand to the author of the patch.
Getting started
New accounts get $1 of sandbox time, which covers more than a hundred review passes on the standard tariff. The sandbox page has the tariffs, the SDK reference every method used above, and Python SDK in 5 minutes the setup.
Related: isolated CI test runs for the full pipeline, and sandbox for AI agents for agents that write the code in the first place.
Comments
No comments yet. Be the first.