Back to blog

Give Your AI Agent a Playground, Not a Sandbox

Aviram Hassan · September 30, 2026 · 7 min read

If you’ve spent even 15 minutes scrolling tech Twitter recently you might’ve noticed that everyone is building (or looking for) sandboxes for their AI agents. Sandboxes give an agent somewhere to experiment, try things, and build without the fear of breaking anything real. But the problem with a sandbox is its starting state, which is either synthetic or empty.

Take two agents, or even two humans, and give them the same task: fix an issue in a repository that already has a live, working environment behind it. Give one of them access to that environment. Give the other a sandbox and a page of instructions on how to build an environment to test against, from scratch. The second one will spend a part of its time/tokens reconstructing Redis, a billing API, and every other dependency before it can even run a test. The bigger issue is that what it builds is still an approximation of the real environment, and every gap in that approximation is a bug that passes in the sandbox and fails in production. The first one skips all of that and tests against the real thing from the start, so if its code works there, it will work in production as well.

The instinct of most engineering teams today is to fix an underperforming agent by giving it more: more context, more docs, more architecture diagrams. That’s not wrong, exactly, it’s just not sufficient on its own. You can hand an agent your business logic, your architecture, your entire set of docs, and it will still need to test its output against something real, because there’s always going to be a gap between what the docs say and what the system actually does.

Sandboxes are great for greenfield projects, things being built from scratch with no history and no legacy to trip over, but that’s not what most of us are shipping into. What an AI agent needs to work inside your enterprise stack isn’t a sandbox, it’s a playground. A sandbox gives the agent an isolated space and leaves it to fill that space with whatever approximation of your system it can build. A playground gives it the real thing: the services your app depends on, already up and running, with real data in the databases, real configuration, and real third-party integrations. And it’s still safe to experiment in, because the agent’s changes stay isolated from everyone else using it.

Turning staging into a playground

You probably already have an environment built that can function as the agent’s playground, you just don’t think of it that way: your staging cluster. It has the real services, integrations, data, and all other dependencies that your application uses in production. mirrord lets your AI agent connect the code it’s iterating on to that cluster, so instead of testing against mocks, it’s testing against the real thing, with none of the risk of breaking the environment for other developers or agents who might be using it.

Let me walk through a real example from my own dogfooding I did recently.

How I gave my agent a playground to test its changes in

One of our customers recently pointed out that they couldn’t see their invoices on our site. So I threw an AI agent at it and had a PR in under 10 minutes. But let’s break down what that PR actually needed to do:

  1. Call our billing API (Paddle, in our case)
  2. Modify the frontend to display the data
  3. Add backend endpoints for the frontend to consume
  4. Handle real-world scenarios: batch download, selecting individual invoices, pagination

After generating the code, the agent had no way to actually verify any of it worked. It doesn’t have our Paddle credentials wired up locally, it doesn’t know the shape our invoice data has actually drifted into in production, and nothing in the agent’s sandbox resembles the production Redis instance the billing service reads from. So it handed me code that looked plausible, and I had to test it myself, which made testing the bottleneck. My options were:

  1. Pull the PR to my local environment and test it there
  2. Build an image for the changed service and deploy it to staging

Both of those take longer than it took to generate the code itself.

Let the agent test its work

How do we enable the agent to test and iterate on its own, instead of that falling to me? We give it access to our staging environment, the same one I’d use to test it myself, so it can see how the code it writes actually behaves under production-like conditions instead of guessing at it.

What that translates into for the agent is running the changed service locally through mirrord so it’s wired into the real cluster, then testing it through a browser (using something like Playwright) the same way a user would and checking whether the result is actually correct. Here’s roughly what that setup looked like:

1. A mirrord config scoped to the service being changed

// .mirrord/invoices.json
{
  "target": {
    "path": "deployment/billing-api",
    "namespace": "staging"
  },
  "feature": {
    "network": {
      "incoming": {
        "mode": "steal",
        "http_filter": {
          "header_filter": "^baggage: .*mirrord-session={{ key }}.*"
        }
      }
    }
  }
}

That header filter is what keeps this agent’s traffic separate from everyone else using the staging cluster, human or other agents. The {{ key }} template gets replaced by mirrord at runtime with whatever key you pass via --key, so this same config file works for every session instead of a hardcoded value routing everyone to the same one. Only requests carrying baggage: mirrord-session=<key> get routed to the locally running process and everything else keeps hitting the deployed service.

2. The mirrord Operator, installed in the cluster, because if you want to use your staging environment then you need to let multiple agents (and developers) use it concurrently, which is what the operator enables

3. A rules file that tells the agent how to do all this. I use AGENTS.md for this, and this is what it looked like:

# Testing billing-api changes

You MUST validate any change to billing-api against staging using mirrord
before you consider the task done. Do NOT ask me to deploy or review a diff
without proof it actually ran.

1. Run the service locally through mirrord, passing this session's key:
   mirrord exec -f .mirrord/invoices.json --key invoices-fix -- go run billing-api/main.go

2. Run the Playwright suite against the staging URL, tagged with our
   baggage header so only your traffic hits your local process:
   npx playwright test invoices.spec.ts --headed

3. Attach the resulting screenshots and the recorded video to your summary.

And a Playwright config that carries the same header, so the browser’s requests get routed to the agent’s local process instead of the deployed one:

export default defineConfig({
  use: {
    baseURL: "https://staging.example.com",
    extraHTTPHeaders: { baggage: "mirrord-session=invoices-fix" },
    video: "on",
    screenshot: "on",
  },
});

With that in place, the agent isn’t just telling me the tests pass, it’s handing me screenshots and a recording of someone clicking through the actual invoices page, backed by our actual Paddle data, running in our actual staging cluster. That is a way better artifact to review than a green checkmark based on some unit tests in CI.

Taking it further with Preview Environments

Once I’d seen the screenshots, I wanted to play with the change myself. But I still didn’t want to spin up my own setup for it, and I definitely didn’t want to deploy to staging and step on my teammates who are also working against staging. That’s where mirrord Preview Environments come in handy: I wired a step into CI/CD that builds an image and uses mirrord to spin up an extra deployment of only the changed service in the cluster. This deployment interacts with all the existing services it needs in the cluster but only gets the traffic tagged for it so everyone else using staging isn’t impacted.

Roughly, the CI step looks like this:

on:
  pull_request:
    types: [labeled]

jobs:
  preview:
    if: contains(github.event.pull_request.labels.*.name, 'preview')
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      # ... configure kubeconfig for the cluster ...
      - uses: metalbear-co/mirrord-preview@main
        with:
          action: start
          target: deployment/billing-api
          namespace: staging
          image: myrepo/billing-api:${{ github.sha }}
          filter: "baggage: mirrord-session={{ key }}"
          key: pr-${{ github.event.pull_request.number }}

I label my PR “preview” and get back a link I can open directly because mirrord’s share-ingress feature serves the preview off its own URL. I click through the flow end to end: view the invoices, download a batch, and paginate. Cost stays minimal too, since idle Preview Environments scale their pods down to zero while nobody’s sending them traffic and scale back up the moment someone does. We’re only paying for what’s actually running, which for a low-traffic feature like this is close to nothing.

If you’re interested in setting up something similar for your team register for our self-serve trial or schedule a demo.

What is mirrord?

mirrord is a Kubernetes development platform that lets developers and AI coding agents test code in a production-like environment before deploying it. Your service runs wherever you're working, locally, in CI, or in an agent's sandbox, while mirrord proxies its traffic, environment variables, and files to and from a shared staging cluster, so it behaves as if it were deployed without actually being deployed.

Engineering teams at companies like monday.com, National Australia Bank, and SurveyMonkey use mirrord to iterate and ship faster, while spending less on dev environment infrastructure.

Want to dig deeper?

With mirrord, cloud developers can run local code in the context of their Kubernetes cluster — streamlining coding, debugging, testing, and troubleshooting.