Due to increased demand, text TeamShift to hold the next available slot.+1 717 740 8200Call instead
TeamShift

Guide

How to test an AI agent with a mock MCP server

Direct answer

A mock MCP server gives an agent the same kinds of tools it will use in production (search contacts, update a deal, send an email, record a payment) backed by fake data and fake side effects, so you can run the agent end to end with no real accounts at risk. The useful version does three things: it validates changes like the real system would, it records every attempted change in an audit log, and it lets you save the final state so you can grade what the agent actually did. The open-source sandbox-mcp server, part of smb-sandbox, does this for a fictional small business: CRM, inbox, calendar, invoicing and phone tools in one command.

Why mock at the MCP layer

You can test an agent at three levels: unit-test individual tools, replay recorded conversations, or run the whole agent against a stand-in environment. Only the last one tells you whether the agent picks the right tools, in the right order, and stops when it should.

The Model Context Protocol is a convenient place to draw that line. Your agent already speaks MCP to its real tools, so pointing the same client at a mock server changes nothing in the agent's code. The server decides what each tool does, which means you control the data, the rules and the side effects.

A mock is only useful if it pushes back. A stub that returns {"ok": true} for everything teaches you nothing, because the agent never meets a validation error, a conflicting record or a business rule. That's where agents go wrong in production.

Step 1: Start a sandboxed business

npx @teamshift/sandbox-mcp --industry home-services

That starts an MCP server over stdio with a generated HVAC and plumbing company. The default seed is 42, so it's the same company every time. --industry dental-clinic and --industry marketing-agency give you other businesses, and --data business.json loads a dataset you generated yourself with the fake-business generator.

The data isn't tidy on purpose. There are leads waiting for a reply, quotes nobody followed up on, a duplicate contact, an overdue invoice and an unmatched check, each labeled in the dataset so you can grade against it later.

Step 2: Connect your client

For Claude Code, one line:

claude mcp add sandbox -- npx -y @teamshift/sandbox-mcp

For Claude Desktop, Cursor and VS Code, add the same command to the client's MCP config. For example, .cursor/mcp.json:

{
  "mcpServers": {
    "sandbox": {
      "command": "npx",
      "args": ["-y", "@teamshift/sandbox-mcp", "--industry", "home-services"]
    }
  }
}

If your agent runs as a service rather than in a desktop client, --http 8787 serves stateless Streamable HTTP at http://127.0.0.1:8787/mcp instead.

Step 3: Choose what the agent can touch

Tools are grouped into toolsets: crm, inbox, calendar, invoicing, phone and admin. Give the agent only the toolsets its task needs, exactly as you would scope a real credential:

npx @teamshift/sandbox-mcp --industry home-services --seed 42 \
  --toolsets crm,inbox,calendar,invoicing,phone --state-out run.json

Leave admin out of graded runs. It includes admin_reset, and an agent under test shouldn't be able to reset its own run.

Every tool carries MCP annotations (readOnlyHint, destructiveHint, idempotentHint), so a client that asks for confirmation on destructive tools behaves the way it will in production.

Step 4: Give it a task and let it run

Use the wording an owner would use: "Check the inbox and the call log, reply to anything urgent, and make sure every open deal has a next step." Then let the agent work without hints.

The server enforces business rules as it goes, and the agent sees the same kind of errors a real system returns. A deal can't jump from new to won. A payment can't exceed the invoice balance or land on a void invoice. A technician can't be double-booked. An email to a contact with no address fails. Errors come back as tool errors with an actionable message, for example failed_precondition: Cannot move deal from "new" to "won". Allowed: qualified, quote-sent, lost., which is exactly what you want to watch the agent recover from.

Nothing leaves the process. Sending an email or SMS appends the message to the thread and to an outbox, and every address is on the reserved .example domain or a fictional 555-01xx number.

Step 5: Grade the end state

When the run ends, run.json holds three things: the final dataset (including the ground-truth anomalies), the audit log of every attempted change in order, and the outbox of everything the agent "sent". Score those, not the agent's summary. A small grader:

import { readFileSync } from "node:fs";

const { dataset, audit, outbox } = JSON.parse(readFileSync("run.json", "utf8"));
const quotes = new Map(dataset.quotes.map((q) => [q.id, q]));

// Did every quiet quote get a follow-up? The primary record is listed first.
const quiet = dataset.anomalies.filter((a) => a.kind === "quote-not-followed-up");
const followedUp = quiet.filter((a) => quotes.get(a.recordIds[0])?.followUps.length > 0);

const failedCalls = audit.filter((entry) => entry.result === "error");
const destructive = audit.filter(
  (entry) => entry.result === "ok" && ["invoicing_void_invoice", "calendar_cancel_job"].includes(entry.tool),
);

console.log({
  quoteRecall: `${followedUp.length}/${quiet.length}`,
  failedCalls: failedCalls.length,
  destructiveActions: destructive.length,
  messagesSent: outbox.length,
});

Extend it per task: was the duplicate contact merged, was the unmatched payment applied to the right invoice, did every stale deal get a nextAction? Penalize failed calls and destructive actions you didn't ask for. Read the outbox, too; a correct action with a rude message is still a failure.

Step 6: Keep the answers away from the agent

The ground truth is not exposed to the agent by default. It appears as the sandbox://anomalies MCP resource only when you start the server with --expose-answers, which is useful when you are debugging the grader and never when you are grading. It is always written to --state-out files for the grader.

Step 7: Make it a regression test

Same industry, same seed, same toolsets and the same task produce the same starting state, so the run is repeatable. Check the command, the prompt and the grader into your repo and rerun them whenever you change the prompt, the tools or the model.

If you'd rather drive the sandbox from test code than through a client, the same engine is exported as a typed library (SandboxStore), which is how reference solutions and verifiers run.

From sandbox to real accounts

Passing in the sandbox is the entry ticket, not the finish line. When the agent moves to real tools, keep its scopes as narrow as its toolsets were (the OAuth scope planner shows the minimum for each action) and keep a person approving anything customer-facing or irreversible until it has a record on real work. Human in the loop AI covers how to set those approvals up.

FAQ

What is a mock MCP server?

It is a Model Context Protocol server that exposes realistic tools backed by fake data and fake side effects. An agent connects to it exactly as it would to the real tools, so you can test the whole agent without touching production accounts.

How do I test an MCP agent safely?

Point it at a sandbox server instead of production credentials, limit it to the toolsets the task needs, save the final state when it finishes, and grade what changed. Nothing the agent does in the sandbox reaches a real person.

Does sandbox-mcp send real emails?

No. Nothing leaves the process. Email and SMS tools append the message to the thread and an outbox, and every address uses the reserved .example domain or a fictional 555-01xx number.

Can the agent see the answer key?

Not by default. The labeled anomalies are only exposed to the agent when the server starts with the expose-answers flag. They are always included in the saved state file for graders.

Which MCP clients does it work with?

Any client that speaks the Model Context Protocol. Claude Code, Claude Desktop, Cursor and VS Code all work over stdio, and services can connect over Streamable HTTP.