Verb

Writing

Your agent works. That was never the hard part.

The short answer

The demo and the deployment are different projects, and only the second one has a deadline nobody can estimate. Most in-app agents stall because they run as a single service account, so the audit log records that the agent did something rather than who asked for it, and no reviewer can sign off on a system where 'who authorised this refund' has no answer. Fix that first: run every action with the asking user's own access, make the confirmation something your interface renders from the real pending call, and keep a log that names a person. Do that and the rest of the review is ordinary.

Two years ago, getting a model to call a function reliably was the whole project. It is not any more. With a coding agent open beside you, an assistant that reads a record from your product and changes it is an evening, and it will be a good evening: the thing will work, your team will gather round, and somebody will say it feels like the future.

Then it does not ship. Not loudly, and usually not because anyone decided against it. It goes into a review, comes back with questions, goes into a backlog, and nine months later the branch does not merge any more. Gartner, polling more than 3,400 organisations, expects over 40% of agentic AI projects to be cancelled by the end of 2027, and lists inadequate risk controls alongside cost and unclear value. That framing is generous. In our experience the risk control that kills it is one specific thing, and it is decided by an architectural choice made in the first hour.

The mistake is made before the demo

When you first wire an agent into your product, it needs credentials. The fast answer, and the one every tutorial reaches for, is to give it its own: a service account, an admin key, a token in an environment variable. It works immediately. Every tool can do everything, and you get to spend your attention on the interesting part.

You have just made your agent one user. Everything it does for anybody, it does as that user. Two consequences follow, and neither is visible in a demo.

The first is that you now own a second permission system. Your API already knows that Priya can refund an order and Jamie cannot. The agent does not, because it is not Priya or Jamie, it is the agent, and the agent can do everything. So you start writing checks inside the agent to re-decide what your API had already decided. That copy will drift from the original, and the drift will not be caught by tests, because both are correct in isolation.

The second is worse, and it is the one that stops the project. Your audit log now says the agent did it. Not who asked.

“The AI agent did this” is not an auditable event

Put yourself in the reviewer’s chair, because someone will sit in it. A customer disputes a refund. You open the log. It says that at 14:02 the assistant issued a refund against order 4521. The obvious next question is who asked for that, and you cannot answer it from the log, because the log recorded the actor and the actor is a robot that acts for everyone.

You can usually reconstruct it. Somebody joins the action to a conversation, and the conversation to a session, and the session to a person, and forty minutes later you have an answer you would rather not have to defend. That is not an audit trail. An audit trail is something you can hand to someone who does not trust you.

This is the whole blocker, and it is why so many of these die quietly rather than being rejected. Nobody says no. They say “can we see the access model”, and the access model has one account in it.

Run every action as the person who asked

The alternative is less work, not more, which surprises people. Instead of giving the agent an identity, give it nothing, and have each action execute with the credential the asking user already holds. The same session, the same token, the same rights they had a second earlier when they were clicking.

Three problems disappear at once. Your API makes exactly the same authorisation decision it makes for any other request from that person, so there is no second permission system to maintain and nothing to drift. The blast radius of a prompt injection shrinks to whatever that one user could already have done by hand, which is a bad afternoon rather than a breach. And the log names a person without you doing anything, because the request genuinely came from them.

It also changes what you can say in the review. “It cannot do anything the person asking could not already do themselves” is a sentence a reviewer can verify against your existing API, which they already trust. Every other answer asks them to trust something new.

The confirmation has to be built by your interface

The second thing a reviewer asks is what stops the agent doing something nobody asked for. The tempting answer is to instruct the model to check first. That produces a sentence that reads like a confirmation, with nothing behind it: the model says “shall I cancel that?”, the user says yes, and the only thing that ever gated the action was the model’s own politeness.

A real gate is structural. The call is pending, your interface renders a card from the actual pending call showing the actual arguments, and nothing runs until a human clicks. The distinction matters more than it looks: one of them holds when the model is confused or adversarially prompted, and the other holds exactly as long as the model is well behaved. There is a longer version of this in confirming destructive actions, and the containment argument is in prompt injection for agents that take actions.

Show the real values, too. A card that says “Confirm this action?” teaches people to click through it, and a gate everybody clicks through is decoration.

You will find out late where you have been testing

The last question is where you tested it, and this is the one we got wrong ourselves, so it is worth telling properly.

Our deploy pipeline runs the full test suite before anything goes live, which is the right instinct. On 2026-09-18 a deploy failed on a retention test: it expected one record and found five. The test was correct. The reason it failed is that the suite had been connecting to the live database, and it had been doing so for as long as the pipeline had existed. Fixing it meant discovering that an earlier run had deleted four real conversations.

Nothing about that was an AI problem. It is an ordinary configuration mistake, of the kind every team has made. The reason it belongs in this post is the shape of it: the dangerous version of this mistake is invisible until the day something writes. A suite that only reads can point at production for a year and never tell you. Agents write. Anything you build to exercise an agent will eventually delete something, and the only safe assumption is that your test environment is production until you have proved otherwise.

So build the sandbox before you need it, and make it obvious from the interface which one you are in. Not a configuration flag someone has to remember to check. Something visible on screen, so that being in the wrong place is something you see rather than something you work out afterwards.

Budget the second half

If you are estimating this work, the useful thing is to notice which half you are estimating. The demo is a weekend and everybody can picture it. The list that follows is the one that takes the quarter:

None of that is interesting, and all of it is the difference between a thing you showed people and a thing your users have. Whether you buy it or build it matters much less than whether you costed it. Teams that plan for this ship. Teams that plan for the demo have a very good demo.

Common questions

Why do AI agent projects get cancelled?

Gartner polled more than 3,400 organisations and expects over 40% of agentic AI projects to be cancelled by the end of 2027, naming escalating cost, unclear business value and inadequate risk controls. In practice the third one arrives first: the agent works, then it meets a reviewer who asks who is accountable for what it did, and the honest answer is a service account nobody can attribute to a person.

Should an AI agent have its own service account?

Only if you are prepared to rebuild your entire permission model inside the agent, and to accept an audit trail that cannot tell you who asked. It is much cheaper to run each action with the credential the asking user already holds, because then your API makes the same authorisation decision it makes for every other request, and the log names a person by construction.

How long does it take to build an in-app AI agent?

A working demo is a weekend, and with a coding agent beside you it may be an evening. That is genuinely not the estimate that matters. The work between that demo and something a security reviewer will approve is identity, a structural confirmation gate, an audit trail, per-tool limits and somewhere safe to test, and teams routinely find a quarter in there that nobody planned for.

What does a security reviewer ask about an AI agent?

Four questions, reliably. Whose permissions does it act with. What stops it doing something the user did not ask for. What does the log say afterwards, and does it name a person. And where did you test this, because if the answer is production the conversation is over.

Keep reading

Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.