Verb

Writing

How to stop an AI agent retrying a failing tool forever

The short answer

Give each turn a small failure budget, around three, and count failures by their cause rather than by tool name. A per-tool cap lets the model walk the list: when four interchangeable tools all fail for the same reason, three tries each is twelve wasted calls before anything stops it. Group tools that share a dependency into one budget, treat a missing precondition as waiting rather than failing, and never retry an expensive call just because it timed out.

When a tool fails, a model’s instinct is to adjust the arguments and try again. That is often right once. It is rarely right five times, and left alone it does not stop. One production agent logged a single turn that called the same tool 23 times.

Every one of those calls cost money, and the user spent the whole time looking at a spinner. The fix sounds obvious, a retry cap, and the obvious version does not work.

Start with a failure budget per turn

Give each turn a small number of tool failures it may spend, and three is enough. Past that the model is no longer converging on the right call, it is sampling guesses. When the budget runs out, stop calling tools and tell the user what failed, in words they can act on.

One rule underneath it: a failure is a result, not an exception. A tool that throws should become an error the model reads and routes around, never something that ends the turn. The user asked for something, and one broken tool should not throw away the rest of the answer.

Why a cap per tool name is not enough

The natural implementation counts failures per tool. The model walks straight around it, because when something underneath is broken, several tools fail for the same reason. Three real cases:

The rule that follows: budget by failure cause, not by tool name. Tools that fail for the same underlying reason draw from one budget. Otherwise the model tries each in turn and you pay for every step.

Treat preconditions separately

A missing precondition should not spend the failure budget at all, and it should not read like a failure either. Its message should state the step that satisfies it. That turns a retry loop into one correct call followed by the one that was waiting.

Cap the expensive things by count, not by failure

Some calls should never be repeated inside a turn, successful or not: a long generation, a sub-agent, anything slow and metered. Give those a hard count, often one, and no recursion.

Timeouts deserve special suspicion. A timeout on an expensive call looks exactly like a transient failure worth retrying, but the first attempt may still be running and may still complete. Retrying pays twice, and on a write, it can do the thing twice.

Then make the error text do the work

Most retry spirals start with an error that said nothing useful, so the model tried the only move available, which was changing the arguments. An error that says what to do instead prevents the loop before any budget is needed. That is covered in how to write tool error messages for an AI agent.

Common questions

How many tool failures should an agent be allowed per turn?

Three is enough in practice. Past that the model is not converging, it is guessing, and every further attempt costs latency and money while the user watches a spinner. Stop and say what went wrong.

Why isn't a limit per tool enough?

Because the model routes around it. If the underlying problem is shared, say the service all your write tools depend on is down, a per-tool cap lets it try each tool in turn. Tools that fail for the same reason must share one budget.

Should a timeout be retried?

Not automatically, and never on an expensive call. A timeout looks like a transient failure worth retrying, but on a slow or costly operation the first attempt may still complete, and retrying pays twice or does the thing twice.

Keep reading

Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.