# Where agents should — and shouldn’t — act on their own.

> In this field note, veridive sets out a practical test for AI agent autonomy: let an agent act alone only when the action is reversible, inexpensive, well logged and within a narrow scope. It describes four levels of autonomy, the guardrails each needs, the risk of prompt injection, and how to extend autonomy with evidence.

Agents are powerful when actions are reversible, inexpensive and well logged. For everything else, they should prepare and a person should decide.

## Key takeaways

- Let an agent act alone only when the action is reversible, inexpensive, logged and narrow in scope.
- Design each task at one of four levels: suggest, draft, act with approval, or act and report.
- Put the guardrails outside the model: permissions, allowed actions, budgets, a stop switch and an audit trail.
- Treat every document, email and web page an agent reads as untrusted input.
- Extend autonomy one narrow task at a time, when the record shows the agent gets that task right.

The word “agent” covers a lot of ground. In this note we mean software that uses a model to decide its next step and then takes it: it reads a request, looks things up, calls tools and changes something in another system. That last part is what makes agents useful, and what makes them risky.

Capable models can plan multi-step work surprisingly well. So the practical question is not whether an agent can do a task ([or whether you need one at all](https://veridive.com/insights/ai-agents-vs-workflows/)). It is where an agent should act without asking, and where it should prepare the work and let a person decide. We use a simple test.

## What makes an action safe to automate?

We look for four properties. An action that has all four is a candidate for autonomy. An action that is missing even one should wait for a person.

- **Reversible.** Can it be undone quickly and completely? Tagging a ticket, filling a draft field or moving a case between queues can be undone. Paying a supplier, deleting a record or emailing a customer cannot.
- **Inexpensive.** If it goes wrong, is the damage small? A misrouted internal request costs minutes. A wrong refund, a wrong price or a missed legal deadline costs far more.
- **Logged.** Is every step recorded, with the inputs the agent saw and the reason it gave? If you can’t reconstruct what happened, you can’t learn from a mistake or show what went right.
- **Narrow.** Is the scope tightly defined, with a known set of tools and a known kind of case? “Route delivery questions to the right queue” is narrow. “Handle customer service” is not.

The test is deliberately conservative. It is easier to give an agent more room later than to repair trust after an expensive mistake.

## What are the levels of autonomy?

Autonomy is not a switch. We design each task at one of four levels, and different tasks in the same workflow often sit at different levels.

1. **Suggest.** The agent recommends and a person acts. For example, it proposes a category for a request and a person chooses.
2. **Draft.** The agent prepares the full piece of work and a person reviews it. For example, a reply with the order data and the policy paragraph attached, waiting for a support agent to send.
3. **Act with approval.** The agent carries out the action once a person confirms it, often in one click. For example, an ERP entry that posts after an accountant approves the draft.
4. **Act and report.** The agent acts on its own and records what it did, while people review samples and exceptions. For example, tagging and routing incoming requests.

Most first projects live at levels two and three. Level four is earned, one narrow task at a time.

## What does this look like in one workflow?

Take an illustrative accounts-payable workflow in which an agent handles incoming supplier invoices. The tasks inside it sit at different levels:

- **Reading and classifying the invoice:** act and report. It is reversible and logged, and a person sees the result later anyway.
- **Matching it to the purchase order and delivery note:** draft. The agent proposes the match and explains any mismatch; a clerk confirms.
- **Preparing the ERP entry:** act with approval. The entry posts only after an accountant approves it.
- **Changing supplier bank details:** never automated. Any such request goes to a person, flagged as a known fraud pattern.
- **Payment:** outside the agent’s permissions entirely.

One workflow, five tasks, four different answers. The mistake is to ask whether “the invoice agent” should be autonomous. The right question is which task, under which limits.

> An agent should earn autonomy one task at a time, with the record to prove it.

## Where should a person always decide?

Some decisions stay with people however good the agent becomes:

- anything that moves money beyond an agreed threshold;
- commitments to customers, suppliers or other companies, such as prices, delivery dates or contract terms;
- decisions about people, including hiring, performance and access rights;
- anything that deletes data or can’t be reversed;
- legal, regulatory or safety-related judgments;
- any case where the agent signals low confidence, or where the input looks unusual.

This is not a comment on the quality of the models. It is about accountability. When something goes wrong in these areas, the organization needs a named person who made the call with the evidence in front of them.

## What should a handover to a person contain?

When an agent stops and hands a case to a person, the handover decides whether the agent saved time or created work. A good handover says what the request was, what the agent checked, what it found, which sources it used, why it stopped and what it would suggest next. The person should be able to decide from one screen, without redoing the search. If handovers routinely need rework, that is a design problem to fix before the agent’s role grows.

## Which guardrails don’t depend on the model?

A model can be persuaded, confused or simply wrong. So the most important guardrails live outside it, [in ordinary software](https://veridive.com/insights/what-are-ai-guardrails/):

- **Least privilege.** The agent gets only the permissions its task needs: read access where reading is enough, write access to specific fields, never an administrator account.
- **Allowed actions.** A fixed list of [tools](https://veridive.com/insights/ai-agent-tool-design/) and actions. Anything else is refused by the system, not by the prompt.
- **Budgets and rate limits.** A ceiling on spending, on the number of actions per hour and on the size of any single action.
- **A stop switch.** A way for the owner to pause the agent immediately, with its cases falling back to people.
- **Its own identity.** The agent acts under its own service account, so every change in every system can be traced back to it.
- **Repeat-safe actions.** Actions designed so that running one twice by accident does no harm, such as updating a draft rather than sending a payment.
- **An audit trail.** Every input, step, tool call and output recorded and searchable.
- **A sandbox first.** New tasks run against test systems before they touch production.

With these in place, the worst an agent can do is bounded by design, whatever the model decides.

## Why is everything an agent reads untrusted?

Agents read text written by other people: emails, documents, web pages, form fields. Any of it can contain instructions. An invoice might carry hidden text asking the agent to change the supplier’s bank details. A customer email might ask it to forward the order history of another account. This is called [prompt injection](https://veridive.com/insights/prompt-injection-prevention/), and it matters more for agents than for chat assistants, because agents can act. Agents that browse the web deserve extra care: a page can be written specifically to steer them.

The defenses are practical. Keep instructions and data separate, so the agent treats a document as something to analyze, never as something to obey. Limit tools and permissions, so an injected instruction has nowhere to go. Require approval for sensitive actions, so an injected request meets a person. And test for it: adversarial examples belong in the evaluation set and run again after every change.

## How do you extend autonomy with evidence?

Before an agent touches live work, run it in shadow mode: it processes real cases alongside the team, but its actions go nowhere. Comparing its proposals with what people actually did is the cheapest test there is.

Then start every new task at suggest or draft. Record what the agent proposed and what the person decided. After enough cases, you have a measured agreement rate for that specific task, broken down by case type.

When agreement is consistently high on a narrow slice of work, such as routing delivery questions in one channel, that slice can move up a level. The rest stays where it is. Keep sampling after the change, and move the task back down if quality slips or the inputs change. Autonomy granted this way is specific, reversible and documented, which is exactly what you would ask of the agent itself.

## The short version

Let agents act alone where actions are reversible, inexpensive, logged and narrow. Everywhere else, let them do the gathering, checking and drafting, which is often most of the work, and let a person decide. That division is not a limitation. It is how you get the speed of automation without giving up the judgment your customers and regulators expect. And if you are unsure where a task belongs, put it one level lower than you think, measure it for a month, and let the record decide.

## Frequently asked questions

### When is it safe to let an AI agent act without approval?

When the action is reversible, inexpensive, fully logged and limited to a narrow scope, and when the agent has a measured track record on that exact task. Tagging, routing or filling a draft field that a person will see later are typical examples. Anything that moves money, commits the company or reaches a customer should wait for a person.

### What is prompt injection, and why does it matter for agents?

Prompt injection is text inside something an AI reads, such as an email, a document or a web page, that tries to change what the AI does. It matters more for agents because they can take actions: an injected instruction could trigger a payment or leak data. Defenses include limited permissions, approval steps and testing with adversarial inputs.

## Sources

1. LLM06:2025 Excessive Agency. OWASP Gen AI Security Project. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
2. LLM01:2025 Prompt Injection. OWASP Gen AI Security Project. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
