Every agent that writes code eventually has to run it. That's the whole point of a code-writing agent — and also the uncomfortable part. The model just produced a script it has never executed, which installs packages nobody reviewed and touches files nobody listed. Running that next to your SSH keys and your production .env is a bet you make many times a day.
A sandbox makes the bet cheap. The code gets its own Linux machine for a few seconds or minutes, then the machine is gone.
What the agent gets
Each sandbox is a Firecracker microVM: its own kernel, its own filesystem, its own network namespace. It is not a container sharing your kernel, and it is not a folder on your server. One starts in about a second, which matters when an agent spins one up per task instead of per day.
Inside there's Python 3.12 with pip, Node.js 22 with npm, bash, git and curl. Outbound internet works — the agent can pip install pandas or clone a repo. Inbound doesn't: no open ports, no SSH. Commands, files and output travel through our API, and that's the only door.
Three ways in:
- MCP — if your agent lives in Claude, Cursor, Cline or anything MCP-aware, it gets sandbox tools next to the VPS tools. No code on your side.
- Python or TypeScript SDK —
pip install eqvpsornpm i @eqvps/sdk, then create, run, delete. - Plain REST — documented in the sandbox docs and the OpenAPI spec.
A minimal loop in Python looks like this:
from eqvps import Sandbox
with Sandbox.create(tariff="small") as sb:
sb.upload("/root/task.py", agent_generated_code)
result = sb.exec("python3 /root/task.py", timeout=55)
print(result.exit_code, result.stdout[-2000:])
# the sandbox is deleted here, also if anything above raised
A non-zero exit code comes back as a normal result, not an exception — which is what you want when the agent needs to read the traceback and try again.
The dangerous cases, and what happens to them
A hostile or just broken package. It runs inside a VM that holds nothing of yours. Delete the sandbox and it's gone.
Secrets. Don't put your main API keys into the code the agent writes. If a task genuinely needs one, pass it as a sandbox environment variable: values are stored encrypted, never logged, and the API only ever returns their names.
Runaway loops. Each command stops at 55 seconds unless you started it as a background task. An idle ephemeral sandbox deletes itself after its idle timeout (5 minutes by default). A daily or monthly spending limit stops new sandboxes when the agent has spent enough for one day.
Exfiltration. Here we have to be straight: outbound internet is open, because without it pip install doesn't work. A sandbox protects your machine; it doesn't stop code from sending whatever it can read inside the sandbox to the internet. So the rule is simple — only put into the sandbox what you're fine losing.
Picking a tariff
| Tariff | vCPU / RAM | Per hour | Good for |
|---|---|---|---|
| micro | 0.25 / 512 MB | $0.0165 | one-off scripts, data checks |
| small | 0.5 / 1 GB | $0.033 | most agent code, pip installs |
| standard | 1 / 2 GB | $0.066 | pandas, test suites |
| plus | 2 / 4 GB | $0.132 | heavy installs (torch), builds |
Billing is per second with a 60-second minimum, from the same prepaid balance as our VPS. A typical agent task — create, install two packages, run, delete — lands under a minute on small, so it's billed the 60-second minimum: about $0.0006. Honestly, for an agent that runs code a few hundred times a day, the sandbox bill is the smallest line in the budget; the LLM tokens cost far more.
When a sandbox is the wrong tool
If the agent needs to serve something — a web app, a webhook receiver, a bot that listens for messages — a sandbox won't do it, because nothing can reach it from outside. Same for a long-lived service that should run for months. That's a VPS job: see VPS for AI agents. A common split is an agent living on a small VPS and creating sandboxes for every piece of code it hasn't seen before.
Sustained heavy compute is the other limit. A sandbox pinned at full CPU for more than 15 minutes is treated as abuse, so long number-crunching jobs belong on a server you rent by the month.
Getting started
Create an account, take a token and run the first sandbox — the connection guide walks through it, and new accounts get $1 of sandbox time to try. Then point the agent at the sandbox page or the MCP tools, and let it break things where breaking things is free.
Step by step: your first agent task in a sandbox. Related: isolated CI test runs, LLM evaluation and code review and PR checks.
Comments
No comments yet. Be the first.