veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEngineering

What are AI guardrails? A map of where each one lives.

“Guardrails” covers very different controls: checks before the model sees a request, limits on what the system can do, validation of what it produces, and people at approval points. The most dependable are ordinary software, not instructions in a prompt.

veridive5 min read

Ask a vendor what guardrails its assistant has and you may hear about a content filter. Ask the security team and you hear about permissions. The operations lead talks about approvals, and whoever wrote the prompt points to a paragraph that begins “You must never…”.

All four are called guardrails. They are different controls, living in different places, and they fail in different ways. They sit in four layers (input, capability, output and people), and the dependable ones are ordinary software, not instructions to a model.

What do people mean by guardrails?

A guardrail is any control that limits what an AI system receives, does or produces, or that puts a person at a decision point. Mapped by where they live:

LayerWhere it livesTwo examples
InputBefore the model sees a requestMasking personal data; rejecting out-of-scope requests
CapabilityAround the tools and data the system can reachRead-only access; a fixed list of allowed actions
OutputBetween the model and whoever uses its answerSchema and business-rule checks; citation checks
PeopleAt the decision pointsApproval above a threshold; escalation when unsure

Notice what is missing: instructions in the prompt. A line such as “never promise a refund” shapes behavior most of the time, and it is worth writing. But a model can misread it, lose it in a long conversation or be talked out of it by text hidden in a document. It is a preference, not a control. If breaking a rule would cost money or trust, the rule needs a check in code or a person behind it.

An instruction in a prompt is a preference. A permission is a control.

What happens before the model sees a request?

Input filtering. Checks on a request before it reaches the model: topics out of scope, abusive language, known injection phrasings, unexpected file types. A useful first screen, and easy to evade when it is the only one. Ask: what happens to the request the filter misses?

Data masking. Replacing personal data such as names, identity numbers or account numbers with placeholders before text leaves your systems. It reduces what a provider sees but doesn’t make the data anonymous. Ask: which fields are masked, and has detection been tested on our own formats?

Filters are particularly weak against hidden instructions, for reasons covered in which prompt injection defenses hold.

What limits what the system can do?

Least privilege. The system acts under its own account with only the access its task needs: read where reading is enough, write to specific fields, never an administrator role. Ask: what is the most damaging thing this account could do?

Allowed actions and limits. A fixed list of actions the system may take, with parameters checked in code and ceilings on amounts, volumes and spend. Anything else is refused by the system, whatever the model asks for. Ask: who can change the list, and is every change logged?

What checks an answer before anyone sees it?

Output validation. Checking output against a schema and business rules before it is used: required fields present, amounts within range, identifiers that exist in master data. Ask: where does an output that fails go, to a retry loop or to a person?

Grounding checks. Confirming that each claim is supported by a source the user is allowed to see, and that the citation says what the answer claims. Ask: what does the system say when no source supports an answer?

Where do people fit in?

Approval points. A person confirms any action that moves money, reaches a customer, commits the company or can’t be undone. Ask: does the reviewer see the evidence, or only the recommendation?

Escalation. When the system is unsure or a check fails, the case goes to a person with the work gathered so far. Ask: what exactly does that person receive? How much an agent may do before a person steps in is covered in where agents should and shouldn’t act on their own.

Take an illustrative refund assistant for an online retailer, with one guardrail per layer:

  • Input: card numbers in customer messages are masked before any model call.
  • Capability: it can read orders and the refund policy and create refund drafts, but it has no permission to issue a refund.
  • Output: a draft whose amount exceeds the order total, or that cites no policy paragraph, is blocked.
  • People: refunds above a set amount, and every case the system marks as uncertain, wait for a team member.

How do you know the guardrails work?

Test them like any other feature. Every guardrail exists to prevent a specific failure, so write that failure down as a test case: a message containing a card number, a document asking for the refund to go to a different account, a draft above the order total, a case that should escalate. Put these adversarial cases in the evaluation set and run them on every change. Test permissions directly too, by attempting the forbidden action with the system’s own credentials.

Then watch them in production. Count how often each guardrail fires, and review a sample of what it caught and what it let through. A guardrail that never fires is either working or broken; only the sample tells you which.

Map one system

Take one AI system you run and fill in the four-layer table for it. Empty cells, and cells that hold only prompt instructions, are where to start. The four guardrail questions on our approach page are a shorter version of the same map. They shape every custom AI software project we build and every system we look after through AI reliability work.

Sources

  1. OWASP Top 10 for Large Language Model Applications OWASP Gen AI Security Project genai.owasp.org/llm-top-10
  2. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

Ask an assistant about this note

EngineeringGuardrailsGovernance

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What are AI guardrails?

AI guardrails are controls that limit what an AI system can receive, do or produce, plus the people who approve its consequential actions. They sit in four layers: checks before the model sees a request, limits on the tools and data it can reach, validation of its output, and human approval points. The most dependable ones are ordinary software rather than instructions in a prompt.

Is a system prompt a guardrail?

Only a weak one. Instructions such as “never promise a refund” shape a model’s behavior most of the time and are worth writing, but the model can misread them, lose them in a long conversation or be talked out of them. Treat a prompt instruction as a preference, and back every rule that matters with a permission, a check in code or a person.

How do you test AI guardrails?

Write the failure each guardrail should prevent as a test case, such as a request containing a card number or a draft with an amount above the order total, and add these adversarial cases to the evaluation set so they run on every change. Test permissions directly with the system’s own credentials, and monitor how often each guardrail fires in production.