Verb

Writing

How to test an AI agent that takes actions, without touching real data

The short answer

Run the whole turn for real, on your own machine, against a development copy of your data, with every change still stopping for a confirmation. That tests tool choice, arguments, confirmations and whether your API accepts what the agent sends. Then test the guards rather than the happy path: a request that should not call a tool, a declined confirmation, an ambiguous argument, a failing tool, an instruction hidden in data. Keep the test environment reachable only from your own machine, and say plainly that it acts wherever your local app points.

Testing an agent that only answers questions is mostly about quality. Testing one that takes actions has a harder constraint: the thing you most need to see, the agent changing data, is the thing you least want to happen by accident. Most teams resolve that by mocking the tools, and in doing so remove most of what was worth testing.

Run it for real, against a development copy

The failures that matter in an agent are about how it uses real tools. Which one it picks. What it puts in the arguments. Whether it stops for confirmation before a write. Whether your API actually accepts what it sends, and whether it reads the response correctly. Mock the tool and all of that is tested against the mock.

Simulating only the final step is better, and still not enough: it never shows your API answering. The first time the agent really calls it is in front of a user, and that is where an expired token or a wrong header surfaces. So run everything for real, on your own machine, against a development copy of your data, with every change still stopping for a yes. You see your own API answer, and nothing your users own is in reach.

Test the guards, not the happy path

A demo proves the agent can do the obvious thing. It will. What decides whether you can trust it is how it behaves at the edges, so spend the testing time there:

Keep these as a fixed set and run them whenever the model, the prompt or the tool descriptions change. Behaviour on the edges moves more between model versions than behaviour on the obvious path.

Keep it pointed at development, and make that the obvious default

Running for real is only safe if the environment is the right one. The risk is a local app quietly pointed at a shared staging or production API: the agent acts wherever the app points.

The obvious safeguard is a timer: let real execution switch itself back off after a few hours, so nobody can forget it on. It sounds right, and in practice it goes wrong. One product tried three hours and got a steady stream of reports that the setting kept resetting, supposedly after every deploy. It was not a bug, it was the timer doing its job in the middle of somebody’s work. Stretching it to a day only moved the same surprise to the next morning. A safeguard whose normal behaviour looks exactly like a bug trains people to distrust the whole tool.

What actually protects you is two things, and neither is a clock. A test key that only works on your own machine, so the test environment cannot reach your users. And saying plainly, at the moment someone sets it up, that it runs against whatever their local app talks to, so they point it at development data. A setup that says loudly what it is beats one that changes itself behind your back.

Then read the log, not the chat

The conversation shows what the agent said. The log shows what it did: each call, its arguments, its outcome, and whether the outcome was verified or merely claimed. The two diverge in precisely the cases your tests exist to find. See why agents report failures as success.

Common questions

Why not mock the tools entirely?

Because the interesting failures are in how the model uses your real tools: which one it picks, what it puts in the arguments, whether it confirms before a write, and whether your API accepts the call. A mocked tool tests the mock. A development copy of your data tests the real thing with nothing at stake.

What should an AI agent test suite include?

Cases where the right answer is to not call a tool, which is the expensive failure. A destructive action after the user declines. An ambiguous argument, to see whether it asks or invents a value. A tool that errors, to see whether it reports honestly. And an instruction planted inside a tool result.

Should a test environment run actions for real?

Yes, because the thing under test is your own API, and a simulated call never shows it answering. Keep it reachable only from your own machine and point it at development data. The real risk is a local app talking to a production API, so say that plainly when it is set up. A timer that switches real execution off sounds safer, but it fires in the middle of real work and reads as a bug.

Keep reading

Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.