−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS
Get started

Sandbox for code review and PR checks: prove a patch works before you approve it

Reading a diff tells you what a patch claims. Running it tells you whether it's true. Check out a pull request in a disposable microVM, reproduce the bug before and after, run the tests, and review with evidence — for about half a cent per pass.

A diff shows what a patch claims to do. It doesn't show whether the bug is actually gone, whether the new branch of an if ever runs, or whether the dependency bump breaks import on a clean install. Reviewers know this, which is why so many reviews end with "LGTM, assuming tests pass".

The gap gets wider with AI-written patches. A model produces plausible code fast, and plausible is exactly the kind that slips past a tired reviewer. The cheap fix is to run the change before anyone approves it — somewhere it can't hurt anything.

The check: base, head, tests

The most convincing evidence a patch can offer is a reproduction that fails on the old code and passes on the new one. A sandbox makes that a 20-line script. It's a Firecracker microVM with Python 3.12, Node.js 22, git and curl, started in about a second:

import os
from eqvps import Sandbox

REPO, BASE, HEAD = os.environ["REPO_URL"], os.environ["BASE_SHA"], os.environ["HEAD_SHA"]

def repro(sb, ref):
    sb.exec(f"cd /root/app && git checkout -q {ref}", timeout=55)
    return sb.exec("cd /root/app && python3 /root/repro.py", timeout=55).exit_code

with Sandbox.create(tariff="standard", ttl=1800) as sb:
    sb.exec(f"git clone -q {REPO} /root/app", timeout=55)
    sb.exec("cd /root/app && pip install -q -r requirements.txt", timeout=55)
    sb.upload("/root/repro.py", open("repro.py").read())
    before, after = repro(sb, BASE), repro(sb, HEAD)
    tests = sb.exec("cd /root/app && python3 -m pytest -q", background=True).wait()
    print(f"repro on base: {before}, on head: {after}; tests: {tests.state}, exit {tests.exit_code}")

repro.py is the snippet from the bug report. If it exits non-zero on base and zero on head, the patch fixes what it says it fixes. Post that line, plus the test summary, as a comment on the pull request, and the reviewer starts from facts.

The test suite runs as a background task because a single synchronous command stops at 55 seconds. The TTL of 30 minutes caps how long a stuck run can bill you.

Review agents that run code

The same idea works for an AI reviewer. Connected through the MCP server, an agent has sandbox tools: create a sandbox, run a command, upload and download files. Instead of writing "this might break Python 3.8 compatibility", it checks out the branch, runs the code and quotes the traceback — or reports that it's fine.

Give such an agent a spending cap and a read-only token, and decide which tools it may call without asking. MCP guardrails lists every tool by risk.

Which tariff

RepositoryTariffTypical passApprox. cost
Small library, pure Python or JSsmall (0.5 vCPU, 1 GB)2 min$0.0011
Web app with a test suitestandard (1 vCPU, 2 GB)5 min$0.0055
Native extensions, compiled depsplus (2 vCPU, 4 GB)15 min$0.033

Ephemeral sandboxes bill per second with a 60-second minimum. If a reviewer wants to come back to the same environment over a couple of days, create it as persistent instead: it keeps its disk for up to 30 days and is billed per started hour, so a standard sandbox kept for a two-day review costs about $3.17. Delete it when the pull request is merged.

Limits worth knowing

  • Two commands at a time per account. Plenty for reviews, which run one after another. If you check dozens of pull requests at once, they queue.
  • No inbound ports. You can't open the app in a browser. Start it inside and test it with curl localhost from a second command.
  • No Docker inside. Tests that spin up containers belong on a VPS runner.
  • Outbound internet is open. That's what makes pip install work. Don't put anything into the sandbox you wouldn't hand to the author of the patch.

Getting started

New accounts get $1 of sandbox time, which covers more than a hundred review passes on the standard tariff. The sandbox page has the tariffs, the SDK reference every method used above, and Python SDK in 5 minutes the setup.

Related: isolated CI test runs for the full pipeline, and sandbox for AI agents for agents that write the code in the first place.

Ready to deploy? Pay with crypto, no KYC — live in about a minute.

Deploy now →

FAQ

How is this different from running CI on the pull request?

CI answers whether the existing tests pass. A review check answers whether the patch does what it says: it runs the reproduction from the bug report on the old code and on the new code, tries the edge case the reviewer is worried about, and keeps a sandbox around for the reviewer to poke at. They work well together.

Can an AI review agent use it?

Yes, and that's where it pays off most. Over MCP the agent gets tools to create a sandbox, run commands and read files, so instead of guessing whether a change breaks something, it runs the code and quotes the output in its review.

Is it safe to check out a pull request from a stranger?

That's the point of the sandbox. The code runs in a separate microVM with its own kernel and nothing of yours inside except what you put there. Don't pass a token that can write to your repository; a read-only token for private repos is enough.

Can I open the app from the pull request in my browser?

No. Sandboxes have no inbound ports, so nothing outside can connect to them. You can start the app inside and test it with curl from another command, or deploy previews to a VPS if reviewers need to click through a UI.

What does a review pass cost?

A typical pass — clone, install, reproduce twice, run tests — takes a few minutes on the standard tariff and costs about $0.005. You pay per second with a 60-second minimum, from a prepaid balance.

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.