Prompt injection in an agent that takes actions: contain it, don't hope
The short answer
Assume an injected instruction will eventually be followed, and design so that following it is bounded. Enforce which tools may run in the executor, not only in the list you offer the model. Dispatch by a literal allowlist rather than by name lookup. Run every call as the signed-in user so an injection can do nothing that person could not. Take the tenant from the session, never from arguments. And put every write behind a confirmation rendered from the real call.
An agent that only answers questions and gets prompt-injected says something wrong. An agent that takes actions and gets prompt-injected does something wrong, in your product, as one of your users. The second is the problem worth engineering for, and the useful framing is not prevention. It is containment.
Start from the assumption that it will obey
An agent reads text from places you do not control: a customer’s record, a support ticket, a document, the result of any tool that returns user-written content. Any of it can contain something shaped like an instruction. Filters and careful system prompts reduce how often the model follows such text. None of them get it to zero, and a defence that works most of the time is a defence an attacker only has to beat once.
So ask a different question. Not “how do I stop the model obeying?” but “if it obeys, what is the most it can do?” Every layer below makes that answer smaller, and none of them depends on the model behaving.
Offering a tool is not enforcing it
Most agents limit the tools the model is given for the current context. That is a strong signal, since a well-behaved model does not call what it was not offered. It is not enforcement.
In one production agent, tools were partitioned by mode, and the partition controlled only which tools each mode offered. A turn in one mode, confirmed by the server log, successfully ran two tools that belonged exclusively to the other, because the executor looked names up in the full registry with no mode check. Nothing stopped a call to any registered tool from running.
The gate belongs in the executor. Before any call runs, check the name against what this context permits, and reject everything else, whatever the model emitted and however it came to emit it.
The switch is the allowlist
When tool calls are dispatched in client code, write the dispatch as a literal switch on known names. A dynamic lookup, handlers[toolName](args), turns a string the model produced into a function call in your codebase. That is the same class of bug as SQL injection, with the model standing where the untrusted query used to be. With a switch, a name that has no case cannot run, whatever an injected instruction asks for.
Run as the signed-in user, and nothing more
An agent with a service account has the union of every user’s access, so an injection can reach all of it. An agent that calls your API with the signed-in user’s own credential can do exactly what that person could do by clicking, and no more. An instruction to export every customer fails for the same reason it would fail for that user, at the same place, with no second permission system for anyone to argue past.
The same goes for tenancy. The tenant comes from the signed session, resolved on the server, never from an argument the model fills in. A tool that accepts an organisation id is a tool that can be talked into naming a different one. Below that, row-level security is the floor: see Postgres RLS for a multi-tenant agent.
A human sees every write
The last layer catches what gets through the rest. Anything that changes data stops at a confirmation rendered by your code from the real pending call, showing the real arguments. An injected “delete these records” surfaces as a card saying exactly that, in front of a person who did not ask for it. That only works if the card is structural and not the model asking in prose, covered in confirming destructive actions.
Test it like any other guard
Put an injection case in your evaluation set: a tool result containing an instruction, and a check that nothing outside the user’s intent ran. Run it on every model change, because susceptibility differs between models and a swap that improves everything else can quietly make this one worse.
Common questions
Can prompt injection be prevented?
Not reliably at the model. Anything the model reads, including a tool result, a customer record or a support ticket, can contain text that reads like an instruction. Prevention is a filter that will miss something; containment is a set of limits that hold whether or not it misses.
Isn't limiting the tools the model is offered enough?
No. A model can emit the name of a tool it was never given, and in one production agent tools restricted to another mode ran anyway, because the executor looked the name up in the full registry. The check has to live where the call is executed.
What does running as the signed-in user protect?
It caps the blast radius at that user's own access. An injected instruction to export every customer fails at your API for the same reason it would fail for that person clicking around, with no separate permission system for the agent to talk its way past.
Keep reading
- Letting an AI agent write to your database, safely
Classification, confirmation with real values, and the difference between an action verified and an action merely claimed.
- How to confirm a destructive action, so the confirmation is real
Why a model asking 'shall I?' is not a confirmation, what the card must show, and what to do when nobody answers it.
- Postgres row-level security for a multi-tenant AI agent
The isolation boundary, the tables that deliberately sit outside it, and why the loud failure is the one worth engineering for.
Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.