−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS
Get started

Your first agent task in a sandbox: an LLM writes code, a microVM runs it

Build the smallest useful AI code execution loop: a model writes a Python script, a Firecracker sandbox runs it, errors go back to the model until the script works. About 40 lines, plus the same task through MCP with no code at all.

An agent that writes code is only useful if something runs that code safely and reports back. This guide builds that loop end to end: a model writes a Python script, a sandbox runs it, and if it fails the error goes back to the model. Then the same task without writing any code, through MCP.

You need Python 3.8+, an EQVPS account token (see connecting your account) and an API key for a model.

1. Set a spending cap first

An agent creates sandboxes on its own, so give it a ceiling before it starts. Two dollars a day and twenty a month is plenty for experiments:

curl -X PUT https://api.eqvps.com/api/v1/eqvps/sandboxes/budget \
  -H "Authorization: Bearer $EQVPS_API_KEY" -H "Content-Type: application/json" \
  -d '{"daily_usd": 2, "monthly_usd": 20}'

When a cap is hit, new sandboxes get 429 budget_exceeded and running ephemeral ones are deleted. Details are in sandbox limits and billing.

2. Install the two libraries

pip install eqvps anthropic
export EQVPS_API_KEY="your-eqvps-token"
export ANTHROPIC_API_KEY="your-model-key"

The example uses one model provider, but nothing depends on it. Only the ask() function talks to the model.

3. The write, run, fix loop

Save this as agent.py:

import re
import anthropic
from eqvps import Sandbox

TASK = "Count the prime numbers below 1,000,000 and print only the number."
client = anthropic.Anthropic()

def ask(messages):
    msg = client.messages.create(model="claude-sonnet-5", max_tokens=2000, messages=messages)
    return msg.content[0].text

def extract_code(text):
    m = re.search(r"`{3}(?:python)?\n(.*?)`{3}", text, re.S)
    return m.group(1) if m else text

messages = [{"role": "user", "content": "Write a Python 3 script for this task. "
             "Reply with a single code block and nothing else.\n\nTask: " + TASK}]

with Sandbox.create(tariff="small", idle_timeout=120) as sb:
    for attempt in range(1, 4):
        reply = ask(messages)
        r = sb.run(extract_code(reply), timeout=55)
        print(f"attempt {attempt}: exit code {r.exit_code}")
        if r.ok:
            print(r.stdout.strip())
            break
        messages += [
            {"role": "assistant", "content": reply},
            {"role": "user", "content": f"The script failed with exit code {r.exit_code}.\n"
             f"stderr:\n{r.stderr[-3000:]}\nFix it. Reply with a single code block."},
        ]

Run it:

python3 agent.py

A typical run prints attempt 1: exit code 0 and then 78498. If the first script fails, you'll see the second attempt with the error fixed.

What each part does:

  • Sandbox.create(..., idle_timeout=120) starts a microVM in about a second. If your script dies halfway, the sandbox deletes itself after 2 idle minutes anyway.
  • sb.run(code, timeout=55) runs the code and returns exit code, stdout and stderr. A crash in the code is a normal result, not an exception, so the loop can show it to the model.
  • r.stderr[-3000:] sends only the end of the error. A full traceback from a large script can waste a lot of the model's context.
  • Three attempts is a deliberate limit. A model that hasn't fixed a script in three tries usually needs a better task description, not a fourth try.

4. When the task needs more than 55 seconds

A synchronous call can run for 55 seconds at most. For longer work, start it in the background and wait:

task = sb.run(code, background=True)
result = task.wait(timeout=1800)

The task keeps running even if your wait times out, and you can pick it up again with sb.task(task_id).

5. The same task through MCP, no code

If you use an MCP client such as Claude Desktop or Cursor with the EQVPS MCP server, the agent already has the sandbox tools: create_sandbox, run_code, exec_command, upload_file, download_file and kill_sandbox. A prompt is enough:

Create a small ephemeral sandbox. Write a Python script that counts the primes below 1,000,000, run it there, fix it if it fails, tell me the result and delete the sandbox.

The agent makes the same calls as the script above. Which tools are safe to allow without confirmation is covered in MCP guardrails.

What it cost

The sandbox lived for well under a minute and was billed the 60-second minimum: $0.00055 on small. The model call costs more than the sandbox. A new account's $1 trial credit covers about 1,800 runs like this one.

Where to go next

FAQ

Why not just run the model's code on my own machine?

Because you don't know what it will do. Generated code can delete files, read your keys or loop forever. In a sandbox it runs in a separate microVM with its own kernel, and the sandbox is deleted afterwards.

Does this work with models other than the one in the example?

Yes. The loop only needs a function that takes messages and returns text. Swap the ask() function for any chat API, including a local model.

What stops an agent from creating sandboxes forever?

Set a daily and monthly spending cap before you start. When the cap is reached, new sandboxes are refused with 429 budget_exceeded. Two commands per account can also run at the same time, so a runaway loop can't fan out.

Can the generated code reach the internet?

Yes, sandboxes have outbound internet access so code can install packages and fetch data. They have no inbound ports. Don't put secrets into the prompt or the code; pass the ones a script really needs as environment variables.

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.