# How to design the tools an AI agent is allowed to use.

> In this field note, veridive explains how to design the tools an AI agent may use. An agent is only as safe as its tools, so each should be narrow, clearly described, scoped to minimum permissions, safe to call twice and informative when it fails. It compares good and bad definitions and covers testing with the model in the loop.

An agent is only as safe and reliable as its tools. Make each tool narrow, clearly described, scoped to minimum permissions, safe to call twice and informative when it fails.

## Key takeaways

- From the model’s side a tool is a name, a description and parameters; the real limits live in the code behind it.
- Prefer narrow tools with fixed choices, each running with minimum permissions under the agent’s own service account.
- Make actions safe to repeat with idempotency keys, set-not-increment operations and draft-first writes a person approves.
- Return errors that tell the model what to do next, log every call, and test tools with the model in the loop.

From the model’s side, a tool is three things: a name, a description and a list of parameters. The model reads them, decides whether a call would help and writes the arguments. Everything else, including what the tool can reach, whether it checks its inputs and what happens if it runs twice, is decided by the code behind it.

So an agent is only as safe and reliable as its tools. Make each tool narrow, clearly described, scoped to minimum permissions, safe to call twice and informative when it fails.

## What does a tool look like from the model’s side?

A tool definition has a name, a plain-language description and a parameter schema with types, required fields and allowed values. The model has no other knowledge of what the tool does: it chooses tools by reading their descriptions, the way a new colleague reads an internal wiki. Your code runs the call and returns the result as text, which the model reads before deciding its next step.

Three consequences follow. The description is a prompt and deserves the same care. The schema is a contract, enforced in code on every call. And the result is input to the next decision, so a result that echoes customer text or a web page is untrusted, like anything else the agent reads. Whether a task needs an agent at all is a prior question, covered in [do you need an AI agent, or a workflow](https://veridive.com/insights/ai-agents-vs-workflows/).

## Why should tools be narrow?

Take two illustrative tool definitions for a support agent:

| Illustrative tool | `update_ticket_status` | `run_database_query` |
|---|---|---|
| Parameters | Ticket ID; status from a fixed list: open, waiting on customer, resolved | Any query text |
| What the model must get right | Which ticket, and one of three statuses | Tables, joins, filters and syntax |
| Worst case if misused | One ticket in the wrong status, easily reverted | Any record read, changed or deleted |
| Testing | A handful of cases per status | No finite set of tests covers it |

The narrow tool leaves the model fewer ways to be wrong, and it makes real limits possible. Every call runs under the agent’s own service account, with permissions scoped per tool: the status tool can write one field, a lookup tool can only read, and nothing runs with administrator rights. Where the agent reads on behalf of a user, it sees only what that user may see. An injected instruction that reaches a narrow tool can do only what the tool allows.

Narrow doesn’t mean endless, though. Dozens of overlapping tools make selection harder. Give each agent only the tools its job needs, and merge tools that differ by a single parameter.

## How should a tool be described?

Like instructions for a careful new colleague:

- **What it does and when to use it.** “Changes a ticket’s status once the customer’s issue is resolved.”
- **When not to use it.** “Don’t use it to close duplicates; use `merge_tickets`.”
- **Each parameter, with its format.** Identifier patterns, dates as year-month-day, units, and allowed values as a fixed list.
- **Side effects, stated plainly.** “Creates a draft only” or “sends an email to the customer” changes how carefully the model should use it.

Keep names distinct and descriptions non-overlapping. Two tools that sound alike will be confused, and the logs will show it.

## How do you make actions safe to repeat?

Agents repeat themselves. A call times out and is retried, a network error hides a success, or the model calls again to make sure. For a status update that is harmless; for a refund it is not. Three habits make repeats safe:

- **Idempotency keys.** Every write carries a key derived from the case and the action. A second call with the same key returns the first result instead of acting again.
- **Set, don’t increment.** “Set status to resolved” can run twice; “add one to the quantity” can’t.
- **Draft first.** Where possible, a tool creates a draft reply, refund or ERP entry that a person approves before anything final happens. Drafts can be overwritten safely; sent emails can’t.

> An agent is only as safe as the tools you give it.

## What should a tool return when something goes wrong?

An error message is an instruction to the model, so write it as one. “Error 500” leaves the model to guess or retry blindly. “Ticket not found; search by customer email with `find_ticket`, or ask the user for the ticket number” gives it a next step.

- **Say whether to retry.** A timeout can be retried once; a permission error should never be retried, and the case should go to a person.
- **Name the field.** A validation error says which parameter failed and what is allowed.
- **Leak nothing.** No stack traces, credentials or other customers’ data in error text.

Log every call: the tool, its inputs and outputs, errors, duration, the case and run it belongs to, and the prompt and model versions. That log is how you debug a strange run, answer an auditor and find the next evaluation cases.

## How do you test tools with the model in the loop?

First without the model. Unit-test each tool’s validation, permissions, idempotency (call it twice) and error messages, like any API.

Then with the model, in a sandbox against test systems. Run a set of realistic tasks several times, because paths vary between runs, and check:

- **Selection.** Did the agent pick the right tool, and leave the wrong ones alone?
- **Arguments.** Were they valid and correct?
- **Recovery.** Inject timeouts, missing records and permission errors. Does the agent recover or hand over cleanly?
- **Abuse.** Put instructions in ticket text that ask for a tool the task doesn’t need, as covered in [which prompt injection defenses hold](https://veridive.com/insights/prompt-injection-prevention/).

Track wrong-tool and invalid-argument rates, steps per task and handovers, and keep the failing runs as evaluation cases.

## Audit your tools

List every tool your agent can call. For each, write down the worst case if it were called with wrong arguments, called twice, or called because of an injected instruction. Narrow any tool whose worst case you wouldn’t accept. Tool design is where the [guardrail questions](https://veridive.com/approach/#guardrails) become code, and it is at the heart of the bounded agents we build as [custom AI software](https://veridive.com/services/custom-ai-software/).

## Frequently asked questions

### What makes a good tool for an AI agent?

A good tool does one narrow job, has a clear description of when to use it and when not to, takes parameters with fixed choices where possible, and runs with minimum permissions under the agent’s own account. It is safe to call twice, and when it fails it returns a message that tells the model what to do next.

### What is function calling in LLMs?

Function calling, also called tool calling, lets a language model request an action from your software. The model reads each tool’s name, description and parameter schema, decides whether a call would help and writes the arguments; your code runs the tool and returns the result. The model never executes anything itself, so the limits belong in the code behind each tool.

### How do you stop an AI agent from repeating an action?

Design every write action to be safe to repeat. Give each call an idempotency key derived from the case and the action, so a second call with the same key returns the first result instead of acting again. Prefer operations that set a value over ones that increment it, and create drafts that a person approves rather than final actions.

## Sources

1. LLM06:2025 Excessive Agency. OWASP Gen AI Security Project. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
2. LLM01:2025 Prompt Injection. OWASP Gen AI Security Project. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
