−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS
Get started

Sandbox for AI agents: run untrusted code in a microVM that starts in a second

Your agent writes code it has never run before. A Firecracker microVM sandbox gives that code a disposable Linux box with no access to your machine, your keys or your network — started in about a second, billed per second.

Every agent that writes code eventually has to run it. That's the whole point of a code-writing agent — and also the uncomfortable part. The model just produced a script it has never executed, which installs packages nobody reviewed and touches files nobody listed. Running that next to your SSH keys and your production .env is a bet you make many times a day.

A sandbox makes the bet cheap. The code gets its own Linux machine for a few seconds or minutes, then the machine is gone.

What the agent gets

Each sandbox is a Firecracker microVM: its own kernel, its own filesystem, its own network namespace. It is not a container sharing your kernel, and it is not a folder on your server. One starts in about a second, which matters when an agent spins one up per task instead of per day.

Inside there's Python 3.12 with pip, Node.js 22 with npm, bash, git and curl. Outbound internet works — the agent can pip install pandas or clone a repo. Inbound doesn't: no open ports, no SSH. Commands, files and output travel through our API, and that's the only door.

Three ways in:

  • MCP — if your agent lives in Claude, Cursor, Cline or anything MCP-aware, it gets sandbox tools next to the VPS tools. No code on your side.
  • Python or TypeScript SDK — pip install eqvps or npm i @eqvps/sdk, then create, run, delete.
  • Plain REST — documented in the sandbox docs and the OpenAPI spec.

A minimal loop in Python looks like this:

from eqvps import Sandbox

with Sandbox.create(tariff="small") as sb:
    sb.upload("/root/task.py", agent_generated_code)
    result = sb.exec("python3 /root/task.py", timeout=55)
    print(result.exit_code, result.stdout[-2000:])
# the sandbox is deleted here, also if anything above raised

A non-zero exit code comes back as a normal result, not an exception — which is what you want when the agent needs to read the traceback and try again.

The dangerous cases, and what happens to them

A hostile or just broken package. It runs inside a VM that holds nothing of yours. Delete the sandbox and it's gone.

Secrets. Don't put your main API keys into the code the agent writes. If a task genuinely needs one, pass it as a sandbox environment variable: values are stored encrypted, never logged, and the API only ever returns their names.

Runaway loops. Each command stops at 55 seconds unless you started it as a background task. An idle ephemeral sandbox deletes itself after its idle timeout (5 minutes by default). A daily or monthly spending limit stops new sandboxes when the agent has spent enough for one day.

Exfiltration. Here we have to be straight: outbound internet is open, because without it pip install doesn't work. A sandbox protects your machine; it doesn't stop code from sending whatever it can read inside the sandbox to the internet. So the rule is simple — only put into the sandbox what you're fine losing.

Picking a tariff

TariffvCPU / RAMPer hourGood for
micro0.25 / 512 MB$0.0165one-off scripts, data checks
small0.5 / 1 GB$0.033most agent code, pip installs
standard1 / 2 GB$0.066pandas, test suites
plus2 / 4 GB$0.132heavy installs (torch), builds

Billing is per second with a 60-second minimum, from the same prepaid balance as our VPS. A typical agent task — create, install two packages, run, delete — lands under a minute on small, so it's billed the 60-second minimum: about $0.0006. Honestly, for an agent that runs code a few hundred times a day, the sandbox bill is the smallest line in the budget; the LLM tokens cost far more.

When a sandbox is the wrong tool

If the agent needs to serve something — a web app, a webhook receiver, a bot that listens for messages — a sandbox won't do it, because nothing can reach it from outside. Same for a long-lived service that should run for months. That's a VPS job: see VPS for AI agents. A common split is an agent living on a small VPS and creating sandboxes for every piece of code it hasn't seen before.

Sustained heavy compute is the other limit. A sandbox pinned at full CPU for more than 15 minutes is treated as abuse, so long number-crunching jobs belong on a server you rent by the month.

Getting started

Create an account, take a token and run the first sandbox — the connection guide walks through it, and new accounts get $1 of sandbox time to try. Then point the agent at the sandbox page or the MCP tools, and let it break things where breaking things is free.

Step by step: your first agent task in a sandbox. Related: isolated CI test runs, LLM evaluation and code review and PR checks.

Ready to deploy? Pay with crypto, no KYC — live in about a minute.

Deploy now →

FAQ

Why not just run the agent's code on my own server?

Because the agent doesn't know what the code does either. A pip package with a hostile install script, a loop that fills the disk, a command that reads ~/.ssh — on your own machine each of those touches your keys and your data. In a sandbox the worst case is a broken sandbox you delete.

What runs inside a sandbox?

Linux with Python 3.12 and pip, Node.js 22 and npm, bash, git and curl. Outbound internet works, so pip install, npm install and git clone are fine. There are no inbound ports and no SSH — commands and files go through the API, the SDK or MCP.

How long can one command run?

Up to 55 seconds per call. Anything longer — a training script, a long test suite — starts as a background task and you poll its output; a task can run for as long as the sandbox lives (up to 24 hours for an ephemeral one).

What does it cost?

You pay for the vCPU and RAM of the tariff, per second, minimum 60 seconds. The smallest tariff (micro) is $0.0165 an hour; small, which fits most agent code, is $0.033. It comes from the same prepaid balance as VPS, and new accounts get $1 of sandbox time to try it.

Can the agent create sandboxes on its own?

Yes. Over MCP the agent gets sandbox tools directly; from code it's three lines with the Python or TypeScript SDK. You can cap spending per day or month, so an agent stuck in a loop can't burn through the balance.

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.