Verb

Writing

Why your AI agent says an action worked when it didn't

The short answer

Because the tool told it so. Plenty of calls succeed at the transport level while doing nothing: an SDK that resolves with an error object instead of throwing, a GraphQL response that is 200 with an errors array, an update that matched zero rows. The model reads success and reports it. The fix is to define what success looks like for each tool as a shape, and to report error whenever a call cannot prove it worked.

An agent that fails loudly is annoying. An agent that fails and reports success is dangerous, because the user stops looking for the problem. The second kind is far more common than it should be, and it is almost never the model’s fault. The model is repeating what the tool told it.

Seven successes and an empty editor

One production coding agent wrote code into a user’s editor through a browser bridge. The bridge resolved its promise every time, and inside the payload it said, quite clearly, that nothing had been written: a field set to false. HTTP was fine. The promise was fine. The job was not done.

The agent read “ok” and told the user it had worked. It did that seven times in one session, against an editor that was empty, and the session ended with a customer’s 598 line script gone. Nothing in that sequence threw an exception. Every individual step looked like success.

Where the 200 that lies comes from

That was a bridge, but the pattern is everywhere an agent takes actions in somebody else’s system:

In every one of these the model does its job correctly. It was told the call succeeded, and it summarised a success.

Define success as a shape, per tool

The fix is not a smarter prompt. It is deciding, for each tool, what evidence of success looks like, and refusing to report ok without it. For a create, an id came back. For an update, rows affected is at least one, or the changed field reads back with the new value. For a client bridge, the flag that says the work happened is true.

Then run that check on every result the transport calls successful, before the model sees anything:

For adapter code you hand to customers, fix the most-copied example first. The snippet that teaches the pattern should unwrap the SDK and throw: const { data, error } = await query; if (error) throw new Error(error.message);

Never let the model report a write it did not verify

This is the rule the rest serves. Record the outcome from the response your code inspected, not from the model’s account afterwards, because those are different facts and they diverge precisely in the cases that matter. Where a write cannot be confirmed, the honest message is that it could not be confirmed, with what the user can check themselves.

A refusal costs the user a minute. A false success costs them the work, and the trust, and they find out later, somewhere else, from someone else. For the rest of the write path, see safe database writes for AI agents.

Common questions

Isn't checking the HTTP status enough?

No. The status says the request was accepted, not that the thing you wanted happened. A write routed to a read replica, an update whose filter matched nothing, and a queue that accepted a message nobody consumes all return success.

Why doesn't Supabase throw when a query fails?

By design, supabase-js resolves to an object carrying data and error, and a missing row or a denied policy comes back with error populated and data null. Code that awaits the call and returns the result hands a failure to the model as though it were an answer. Destructure it and throw on error.

What should the agent say when it cannot verify a write?

That it could not confirm the change, plainly, with what the user can check themselves. A false success is the most expensive thing an agent can say, because the user stops looking for the problem.

Keep reading

Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.