Self-hosted CI runners have one awkward property: they remember. Caches, tokens, a deploy key someone added "temporarily" last spring. That's fine for your own branches. It's a problem the moment a pull request arrives from someone you've never heard of — or from your own AI agent, which wrote a patch nobody has read yet.
The clean answer is a machine per run. Clone, install, test, throw the machine away.
How the run looks
A sandbox is a Firecracker microVM that starts in about a second with Python 3.12, Node.js 22, git and curl. Outbound internet works, so git clone and pip install -r requirements.txt behave as usual. There are no inbound ports — nothing can call into the test environment while it runs.
From a CI job, the whole thing is a short script. Here it is with the Python SDK, called from any CI job:
import os, sys
from eqvps import Sandbox
repo, ref = os.environ["REPO_URL"], os.environ["PR_SHA"]
with Sandbox.create(tariff="standard", ttl=1800) as sb:
sb.exec(f"git clone {repo} /root/app && cd /root/app && git checkout {ref}", timeout=55)
sb.exec("cd /root/app && pip install -r requirements.txt", timeout=55)
task = sb.exec("cd /root/app && python3 -m pytest -q", background=True)
result = task.wait(on_output=lambda out, err: print(out, end=""))
sys.exit(result.exit_code or 0)
The token lives in your CI secrets as EQVPS_API_KEY. Nothing else from your CI — no deploy keys, no cloud credentials — goes into the sandbox. When the with block ends, the sandbox is deleted, also if the job crashed halfway.
The test run itself is a background task, because a single synchronous command stops at 55 seconds. A background task streams output while it runs and can go on until the sandbox's TTL — set here to 30 minutes so a hung test can't run up a bill.
The limits you'll actually hit
Being honest about these saves you an afternoon.
Concurrency. An account can hold 20 sandboxes, but only 2 commands execute at the same time. For CI that's two jobs truly running in parallel. Ten PRs arriving at once will queue. If your pipeline shards tests across 16 workers, this is the wrong tool for the main pipeline — keep that on your own runners on a VPS and use sandboxes for the untrusted lane.
No containers inside. Docker doesn't ship in the sandbox. Unit and integration tests that need a Postgres container won't work as-is; tests against SQLite or an in-memory fake will.
Files. Uploads and downloads go up to 5 MB per file through the API. Pull the code with git inside the sandbox rather than uploading a tarball.
What it costs
You pay for the tariff's vCPU and RAM per second, 60 seconds minimum, from a prepaid balance — no subscription.
| Run | Tariff | Approx. cost |
|---|---|---|
| Lint + unit tests, 1 min | small (0.5 vCPU, 1 GB) | $0.0006 |
| Full suite, 3 min | standard (1 vCPU, 2 GB) | $0.0033 |
| Build + tests, 10 min | plus (2 vCPU, 4 GB) | $0.022 |
At these prices the interesting question isn't cost, it's whether your tests finish within the time you set. Give the TTL some headroom over your slowest green run.
A sensible split
Our recommendation: keep trusted branches on your fast, cached runner. Send everything you didn't write — forks, external contributors, agent-generated patches — through a sandbox first. If the sandbox run is green and a human has looked at the diff, promote it to the trusted pipeline.
That gives you one lane where speed and caching matter and one where a clean machine matters, without forcing one tool to be both.
Start with the sandbox overview and the connection guide — new accounts get $1 of sandbox time, which covers a few hundred short test runs. Every method used above is in the SDK reference, and the exact billing rules are in sandbox limits and billing.
Related: sandbox for code review and PR checks and sandbox for AI agents.
Comments
No comments yet. Be the first.