How to test an AI agent that takes actions, without touching real data
The short answer
Run the whole turn for real, on your own machine, against a development copy of your data, with every change still stopping for a confirmation. That tests tool choice, arguments, confirmations and whether your API accepts what the agent sends. Then test the guards rather than the happy path: a request that should not call a tool, a declined confirmation, an ambiguous argument, a failing tool, an instruction hidden in data. Keep the test environment reachable only from your own machine, and say plainly that it acts wherever your local app points.
Testing an agent that only answers questions is mostly about quality. Testing one that takes actions has a harder constraint: the thing you most need to see, the agent changing data, is the thing you least want to happen by accident. Most teams resolve that by mocking the tools, and in doing so remove most of what was worth testing.
Run it for real, against a development copy
The failures that matter in an agent are about how it uses real tools. Which one it picks. What it puts in the arguments. Whether it stops for confirmation before a write. Whether your API actually accepts what it sends, and whether it reads the response correctly. Mock the tool and all of that is tested against the mock.
Simulating only the final step is better, and still not enough: it never shows your API answering. The first time the agent really calls it is in front of a user, and that is where an expired token or a wrong header surfaces. So run everything for real, on your own machine, against a development copy of your data, with every change still stopping for a yes. You see your own API answer, and nothing your users own is in reach.
Test the guards, not the happy path
A demo proves the agent can do the obvious thing. It will. What decides whether you can trust it is how it behaves at the edges, so spend the testing time there:
- A request that should not call a tool. The expensive failure is over-triggering: acting when the user only asked a question.
- A declined confirmation. Say no to a destructive action and check it does not rephrase and offer the same thing again.
- An ambiguous argument. Ask to cancel “the meeting” when there are two. Does it ask which one, or invent an id?
- A failing tool. Make a call fail and check the agent says so, rather than summarising a success.
- An instruction hidden in data. Put text that reads like a command inside a record the agent will read, and check nothing outside the user’s request ran.
Keep these as a fixed set and run them whenever the model, the prompt or the tool descriptions change. Behaviour on the edges moves more between model versions than behaviour on the obvious path.
Keep it pointed at development, and make that the obvious default
Running for real is only safe if the environment is the right one. The risk is a local app quietly pointed at a shared staging or production API: the agent acts wherever the app points.
The obvious safeguard is a timer: let real execution switch itself back off after a few hours, so nobody can forget it on. It sounds right, and in practice it goes wrong. One product tried three hours and got a steady stream of reports that the setting kept resetting, supposedly after every deploy. It was not a bug, it was the timer doing its job in the middle of somebody’s work. Stretching it to a day only moved the same surprise to the next morning. A safeguard whose normal behaviour looks exactly like a bug trains people to distrust the whole tool.
What actually protects you is two things, and neither is a clock. A test key that only works on your own machine, so the test environment cannot reach your users. And saying plainly, at the moment someone sets it up, that it runs against whatever their local app talks to, so they point it at development data. A setup that says loudly what it is beats one that changes itself behind your back.
Then read the log, not the chat
The conversation shows what the agent said. The log shows what it did: each call, its arguments, its outcome, and whether the outcome was verified or merely claimed. The two diverge in precisely the cases your tests exist to find. See why agents report failures as success.
Common questions
Why not mock the tools entirely?
Because the interesting failures are in how the model uses your real tools: which one it picks, what it puts in the arguments, whether it confirms before a write, and whether your API accepts the call. A mocked tool tests the mock. A development copy of your data tests the real thing with nothing at stake.
What should an AI agent test suite include?
Cases where the right answer is to not call a tool, which is the expensive failure. A destructive action after the user declines. An ambiguous argument, to see whether it asks or invents a value. A tool that errors, to see whether it reports honestly. And an instruction planted inside a tool result.
Should a test environment run actions for real?
Yes, because the thing under test is your own API, and a simulated call never shows it answering. Keep it reachable only from your own machine and point it at development data. The real risk is a local app talking to a production API, so say that plainly when it is set up. A timer that switches real execution off sounds safer, but it fires in the middle of real work and reads as a bug.
Keep reading
- Why your AI agent says an action worked when it didn't
The 200 that lied: why a successful response is not evidence, and the per-tool success shape that stops an agent claiming work it never did.
- Tool error messages are prompts, so write them like one
Why "failed" produces a guess, the four parts of an error a model can act on, and why descriptions are advisory but validation is binding.
- How to stop an AI agent retrying a failing tool forever
The 23-call spiral, why per-tool caps let a model walk the tool list, and budgeting failures by cause.
Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.