Red-teaming an AI system: a practical test plan before launch.
Red-teaming means structured attempts to make the system misbehave: leak data, skip an approval, give off-policy answers or run up costs. Do it with people who know the business, log every finding, and turn each one into a test.
veridive6 min read
Most testing before launch asks whether the system does what it should. Red-teaming asks what it does when someone tries to make it misbehave: a customer who pastes instructions into a complaint, a colleague who asks the assistant for someone else’s salary, a supplier invoice with hidden text.
AI red teaming is a structured attempt to make the system leak data, skip an approval, give off-policy answers or run up costs, before real users and attackers try. Do it with people who know the business as well as security, log every finding, and turn each one into a test that runs on every change.
What is red-teaming for an AI system?
A small team, a time box, a written plan with categories, and a log. It differs from a classic penetration test, which looks for weaknesses in infrastructure and code, such as exposed services or broken authentication. Red-teaming an AI system tests behavior: what it can be talked into, what data it reveals and which controls it can be persuaded to skip. You need both.
It also differs from the evaluation set, which checks quality on normal work; red-team findings add the adversarial cases to it. Passing prompt-injection and data-leak tests is one of the conditions we check before go-live.
Who should be on the red team?
Four to six people, for half a day or a day:
- Business insiders, such as a senior customer-service agent or a finance reviewer. They know which wrong answers would hurt, like a refund above policy or a promised delivery date, and how real customers phrase awkward requests.
- Someone from security, who knows attack techniques and how similar systems get abused.
- An engineer who knows the system: its tools, permissions and data sources.
- A note-taker, so the testers can keep testing.
Don’t staff it only with the builders: without meaning to, they know where not to push.
- Done when: each person has categories assigned and access to a test environment.
- What goes wrong: a security-only team finds clever jailbreaks that don’t matter to the business, and misses the off-policy answers that would.
Which attack categories should the plan cover?
Seven categories make a solid starting plan:
| Category | What testers try | What holding looks like |
|---|---|---|
| Direct prompt injection | Telling the assistant to ignore its rules, reveal its instructions or play another role | It refuses and reveals nothing internal |
| Indirect prompt injection | Hiding instructions in an email, document or web page the system reads | The text is treated as data; nothing is triggered |
| Data leakage between users | Asking for another customer’s order, guessing IDs, asking about “the last conversation” | Only the user’s own data appears |
| Approval bypass | Talking it into acting without approval, or splitting one large refund into small ones | The approval point can’t be skipped |
| Off-policy advice | Pushing for compensation, promises, or legal or medical advice | It stays on policy or hands over to a person |
| Abusive inputs | Insults, distress, malformed or unusual text | Calm handling, and escalation where needed |
| Cost exhaustion | Very long inputs, requests that trigger loops of tool calls or retries | Limits stop it and an alert fires |
Test in every language the system supports: an attack that fails in English can work in Turkish. Our notes on prompt injection defenses and where guardrails live cover the controls under test, and our guardrails set out the ones we design in.
How do you run and log a session?
- Prepare a test environment that mirrors production, with the same prompts, tools and permissions, test data instead of real customer records, and accounts for several users so leakage can be tested.
- Brief the team: scope, categories, what is off-limits, and the rule that every success or half-success is logged.
- Attack in pairs, starting from the plan and then improvising on whatever looks weak.
- Log verbatim: the exact input and output, the category, a severity, the trace ID and who found it.
- Close with fifteen minutes to agree severities and owners while everyone is still in the room.
The findings log is a simple table: ID, category, input and output, severity, proposed fix, owner, status, and whether the evaluation case has been added.
- Done when: every finding has a severity and an owner before the session ends.
- What goes wrong: paraphrased logs nobody can reproduce, or testing a demo configuration that differs from production.
What happens to the findings?
Triage by severity. Critical findings, such as another user’s data appearing, block launch until fixed. High ones are fixed before launch, or the launch goes ahead with limits. Medium and low ones are scheduled. Fix each at the layer where it belongs: permissions, tool limits, validation and approval design first, prompt wording last, because a reworded prompt can be undone by a reworded attack. Then turn every finding into an evaluation case with the expected safe behavior, and re-run the set after the fix.
Take an illustrative half-day session on a customer-service assistant for an online retailer, with five people. Three findings go into the log:
- Indirect injection, high. A customer email with a fake “note from the manager” made the draft reply offer a full refund. Fix: the review screen shows the refund against the policy limit, and instructions inside customer text are treated as data. Test added.
- Data leakage, critical. Asking “what did I order last time?” from a guest session with a similar email address returned another customer’s order. Fix: order lookups use the verified customer ID from the session, never an address typed into the chat. Test added, plus a permission test in the regular suite.
- Cost exhaustion, medium. A very long pasted document sent the assistant into repeated summarizing and retries. Fix: an input length limit and a cap on tool calls per conversation. Test added.
A finding that doesn’t become a test will come back.
How often should you repeat it?
Before launch, as part of go-live. After major changes: a new model, new tools or permissions, a new channel, language or data source. On a regular cycle, such as quarterly, with some new faces so the testing doesn’t go stale. And after any incident, since an incident is a red-team finding you didn’t plan. Between sessions, the evaluation cases from earlier findings run on every change, which keeps old holes closed, and it is the kind of routine AI reliability takes on.
- Done when: the next session’s trigger is written into the runbook.
- What goes wrong: one pre-launch session treated as a permanent pass.
Sources
- OWASP Top 10 for Large Language Model Applications OWASP Gen AI Security Project genai.owasp.org/llm-top-10
- MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) MITRE atlas.mitre.org
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
Ask an assistant about this note