Guide
How to test an AI agent without real customer data
Direct answer
To test an AI agent without real customer data, give it a fictional business instead of your production accounts: generate a realistic, seeded dataset (customers, deals, quotes, invoices, messages, calls), load it into the system the agent will use, give the agent a real task, and grade what it changed against a list of known problems planted in the data. Because the dataset is seeded, you can rerun the exact same scenario after every prompt, tool or model change and compare scores. The open-source fake-business generator does the data half in one command, and its labeled anomalies are the answer key.
Why production data is the wrong place to start
There are three separate problems with testing an agent on real accounts, and they compound.
- Privacy. Customer names, emails, phone numbers and invoice amounts end up in prompts, logs, traces and whatever eval tooling you use. Each of those is another place personal data now lives.
- Side effects you can't take back. An agent that is still being tuned will, sooner or later, email a customer something wrong, merge the wrong two contacts, or record a payment twice. On real accounts, those mistakes reach real people.
- No answer key. Even if nothing goes wrong, you can't score the run. Did the agent find every stale quote? You don't know how many there were, so you can't tell a careful agent from a lucky one.
Spreadsheets of random names don't fix the third problem. An agent that follows up on quotes needs quotes linked to deals, deals linked to contacts, emails that mention the quote number, and a timeline where the follow-up is visibly overdue. Random rows don't have any of that, so the test only proves the agent can read a table.
Step 1: Generate a business with a known set of problems
The fake-business package generates one coherent fictional company. Records reference each other, timestamps are causally ordered, and eleven kinds of everyday mess are planted on purpose and listed in an anomalies array: duplicate contacts, quotes nobody followed up, missed calls never returned, stale deals, overdue invoices, an invoice attached to the wrong job, unmatched payments, and more.
# A small HVAC and plumbing company, the same one every time
npx @teamshift/fake-business --industry home-services --seed 42 --size small --summary
# The full dataset as JSON
npx @teamshift/fake-business --industry home-services --seed 42 > business.json
You can try it without installing anything on the in-browser generator. Three industries ship today: home-services, dental-clinic and marketing-agency. All emails use the reserved .example domain and phone numbers sit in the fictional 555-0100 to 555-0199 range, so the dataset is safe to commit to a repo or paste into a bug report.
Step 2: Load it where the agent works
Put the data wherever the agent will read and write:
- A database.
--format sql-sqlite --out business.sql, thensqlite3 business.db < business.sql. Usesql-postgresfor PostgreSQL. - A CRM or accounting sandbox.
--format hubspotand--format quickbookswrite import CSVs shaped for each product's import wizard. They are not official templates, so check the column mapping and use a sandbox or test account. - Mock tools over MCP. The sandbox-mcp server exposes the same business as CRM, inbox, calendar, invoicing and phone tools with no network access at all. The mock MCP server guide walks through that setup.
Step 3: Give the agent a real task
Write the task the way an owner would say it, not the way a test author would. "Follow up on every estimate that has gone quiet" is a better test than "call crm_follow_up_quote for each quote with status sent", because the first one checks whether the agent can find the work on its own.
Keep a short list of tasks that map to anomaly kinds you care about:
| Task | Anomaly kinds it should resolve |
|---|---|
| Follow up on quiet estimates | quote-not-followed-up |
| Return every missed call from a new caller | missed-call-no-callback, lead-never-contacted |
| Chase overdue invoices politely | overdue-invoice |
| Clean up the CRM | duplicate-contact, missing-contact-info, conflicting-status |
| Reconcile payments | unmatched-payment, invoice-wrong-job |
Step 4: Grade the end state, not the transcript
An agent's summary of what it did is not evidence. Read back the data afterwards and compare it with the answer key. Each anomaly lists its records with the primary record first, so precision and recall are a few lines of code:
import { generate } from "@teamshift/fake-business";
const before = generate({ industry: "home-services", seed: 42, size: "small" });
const truth = new Set(
before.anomalies.filter((a) => a.kind === "quote-not-followed-up").map((a) => a.recordIds[0]),
);
// `flagged` = quote ids your agent followed up, read back from wherever it wrote them
function score(flagged: string[]) {
const hits = flagged.filter((id) => truth.has(id)).length;
return { precision: hits / Math.max(flagged.length, 1), recall: hits / Math.max(truth.size, 1) };
}
Then check the things a score misses: did the agent touch records it had no reason to touch, did it send anything a person should have approved, and did its actions happen in a sensible order? The dataset's events timeline helps with the last one.
Step 5: Test the clean case too
Generate the same business with --messiness 0. It has no planted problems, so anything the agent "fixes" is a false positive. An agent that follows up on quotes that were already followed up, or merges contacts that were never duplicates, is the one that will annoy your customers later.
Step 6: Rerun on every change
The output is byte-identical for the same options, so the scenario is a regression test. Keep the seed, the task wording and the grader in your repo, and rerun after you change the prompt, the tools or the model. When a score drops, you are looking at the same business and the same problems, so the cause is your change.
Where human review still belongs
A good sandbox score earns an agent a place in your real systems; it doesn't earn it unsupervised access. Keep a person approving anything that emails customers, moves money or deletes records until the agent has a track record on real work. Human in the loop AI covers how to set those gates without turning every action into a ticket, and the OAuth scope planner shows the minimum permissions and approval tier for each action an agent takes.
FAQ
Can I test an AI agent on production data if I anonymize it?
You can, but anonymization tends to break the links that make the data useful (names in email bodies, phone numbers in call logs), and it still gives you no answer key. Generated data with labeled problems is usually faster to set up and easier to score.
Is synthetic business data realistic enough for agent testing?
It is when the records are connected. A generator that links leads to deals, quotes, jobs, invoices and messages, and plants specific known problems, tests the reasoning an agent needs. Random rows only test whether it can read a table.
How do I measure an agent's accuracy on a business task?
Compare the records it changed with a known list of records that needed changing. Precision is how many of its changes were needed; recall is how many needed changes it made. Run it on a clean dataset as well to count false positives.
Do I need to install anything to try this?
No. The fake-business generator runs in the browser at teamshift.io/tools/fake-business, and the command-line version runs with npx on Node.js 20 or later.