Letting an AI agent write to your database, safely
The short answer
Four rules. Route writes through your own API rather than the database, so your existing validation applies. Classify each one as write or destructive and require an explicit human confirmation that shows the real arguments, not a generic 'are you sure?'. Verify the write from the response rather than from the model's account of it. And when it fails, say so plainly, because a false success costs more than a refusal ever does.
Reads are the easy half. An agent that fetches the wrong record wastes a call. An agent that writes the wrong record has changed something somebody now has to find and undo, and the finding is usually the expensive part.
What follows is the write path as it actually has to be built, including the two failures that are easy to design past and expensive to hit.
1. Writes go through your API, not your database
Your API already validates input, enforces authorisation, fires the side effects and keeps the audit trail. A tool that writes SQL directly bypasses every one of those, which means you now maintain two definitions of what a valid write is. They agree on the day you write them and not much longer.
Route through the endpoint your own front end calls. The agent becomes another client of rules you already trust.
2. Classify, and treat destructive as its own class
Read, write, destructive. Three classes, not two. Creating a project and deleting a customer are different kinds of event and deserve different friction, different wording, and in most products different default states: destructive tools should not be in the list at all until somebody has asked for one twice.
The classification also gives you something to put in the log later, which is what makes “what has this thing been doing” answerable in one query.
3. The confirmation must show real arguments
A gate that says “Confirm this action?” teaches people to click through it. A gate that says Cancel order 1041 makes them read. The difference is not politeness, it is whether the confirmation does any work at all: the entire value of the step is that a human looked at the specific thing about to happen.
This means the card has to be rendered by your interface from the pending call, with the arguments the call will actually run with. Not a summary. Not the model’s paraphrase. The values.
4. Verify the write. Do not take the model’s word for it
This is the failure worth designing the whole system around, and it is not exotic. The call fails, is denied, or times out, and the model reports the turn as complete because a helpful summary is what it was trained to produce. Nothing errored. Nobody was attacked. The user simply believes something happened that did not.
So keep two separate facts: what the API returned, and what the model said. Log the first. Show the user the first. A system that cannot tell “verified” from “claimed” is a system whose logs cannot answer the only question anybody will ever ask them.
5. Make retries safe
Agents retry. Users press the button twice. Connections drop between the write landing and the response arriving, which is the case that produces duplicates while looking exactly like a failure. An idempotency key on write endpoints turns that ambiguous case into a safe one, and it is much cheaper to add before you need it.
The failure this all exists to prevent
Not a breach. A quiet, confident report that something is done when it is not. Breaches are rare and loud; false completions are common and silent, and they are what make people stop trusting an assistant after using it twice.
Every rule above is downstream of that: route through the API so failures are real failures, confirm with real values so somebody saw it, verify from the response so the record is true, and make retries safe so the ambiguous case has an answer.
Common questions
Is a confirmation prompt enough to make AI writes safe?
Only if the confirmation is generated by your code from the actual tool call, and shows the real arguments. A model asked to confirm in its own words will produce a sentence that looks like a confirmation and has no button under it, so nothing gates the action. The gate has to be structural, not conversational.
How do I know the action actually happened?
Record the outcome from the API response, not from what the model says afterwards. Those are different facts and they come apart exactly when it matters: a call that timed out, was denied, or never ran can still be summarised as done by a model trying to be helpful.
What about deletes?
Treat destructive separately from write. A delete deserves its own classification, its own confirmation wording naming what is about to be lost, and in most products it should not be in the agent's tool list at all until someone has asked for it twice.
Should AI writes be idempotent?
Where you can manage it, yes. Agents retry, users press buttons twice, and connections drop between the write landing and the response arriving. An idempotency key turns the ambiguous case into a safe one.
Keep reading
- How to confirm a destructive action, so the confirmation is real
Why a model asking 'shall I?' is not a confirmation, what the card must show, and what to do when nobody answers it.
- Postgres row-level security for a multi-tenant AI agent
The isolation boundary, the tables that deliberately sit outside it, and why the loud failure is the one worth engineering for.
- Prompt injection in an agent that takes actions: contain it, don't hope
Why offering a tool list is not enforcement, the switch that is the allowlist, and bounding what an obeyed injection can reach.
Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.