What an AI audit trail should record, and who should read it.
When a decision is questioned months later, you need to rebuild what the system saw, which sources it used, which model and prompt answered, and who approved it. Log that by design, and treat the logs as sensitive data.
veridive5 min read
Take an illustrative case. A customer disputes a refund decision months after it was made: the assistant recommended a partial refund for a returned sofa, an agent approved it, and the customer says the policy promised a full one. Someone now has to answer precisely.
With a good audit trail, that takes minutes. The record shows the customer’s message and photos as received, the order, the returns policy in the version in force that day, the passage the system cited, the model and prompt version behind the recommendation, and the agent who approved it. It may show the system was right. It may show that it cited an outdated version, which is just as important to know.
Without one, the team reconstructs from memory. An AI audit trail should be designed so every decision can be rebuilt, and the logs themselves treated as sensitive data.
What questions should an audit trail answer?
Five, for any case, without asking the people involved:
- What did the system see?
- Which sources did it use, in which versions?
- Which model, prompt and settings produced the output?
- What did it recommend, and what did it flag as uncertain?
- Who decided, what did they do, and when?
If the log can’t answer all five, it is a debugging log, not an audit trail.
What should be recorded for each case?
A field checklist, with the reason for each:
- Case ID and timestamps for arrival, processing, review and action, so events can be ordered and matched with other systems.
- Input as received, or a reference to it plus a fingerprint (a hash), so you can show what the system saw even if the source changes later.
- Retrieved sources with versions: document IDs, versions or effective dates and the passages used, because policies change and “which version?” is often the whole dispute.
- Records looked up: the order, amounts and customer status the system read, since these are the facts behind the output.
- Model and prompt version: the model identifier the provider returned, the prompt version from your repository and key settings, because behavior changes when any of them does.
- Output as shown to the reviewer, citations included, since that is what the person approved.
- Flags and checks: validation results, uncertainty flags, rules triggered and why a case was escalated.
- Reviewer and action taken: who reviewed, what they changed, approved or rejected, and their reason for any override.
- Final outcome in the system of record: what was actually sent, paid or posted.
Recording what the person saw matters as much as recording what the model said. An approval only means something if you can show what was in front of the approver.
How long should logs be kept?
Long enough to answer disputes and audits, and no longer than your data protection rules allow. The two pull in opposite directions, and the answer is not yours alone: logs contain personal data, so agree retention and access with your DPO, and ask counsel about sector or contractual rules on record keeping.
Splitting the record often helps. Keep the decision record, meaning IDs, versions, flags, reviewer and action, for the longer period, and full text for a shorter one, or store references to the source system rather than copies. Decide in advance how a deletion request interacts with the trail.
Who should have access to them?
Fewer people than have access to the application. A workable split:
- The workflow owner and auditors read individual cases when there is a reason to.
- The reliability team sees aggregated, masked data for monitoring.
- Engineers get time-limited access for debugging, masked where possible.
Log access to the logs, too. A trail is often less protected than the systems it describes, which makes it an easy place to leak from. It is the same question as “who can access what?” in our guardrails.
How do logs support improvement as well as audit?
The same records feed everyday quality work. Override rates, flags, cost per case and response times come from them, and they are what monitoring after launch runs on. Every override, with the reviewer’s correction, is a candidate case for the evaluation set. When something goes wrong, the trail is where an incident review starts. It also answers what a board or risk committee should ask: which systems influence which decisions, how often people override them, and what the last incident taught; see what boards should ask about AI. A trail designed only for auditors tends to be read only by auditors; one designed for improvement gets looked at every week, which is also how its gaps get found.
How do you test that the trail is complete?
Run a reconstruction drill. Pick one past case at random and one awkward one, and ask someone who worked on neither to rebuild each decision from the logs alone: what the system saw, which sources and versions it used, which model and prompt answered, what it recommended and who approved it. Time it, and note every question the logs couldn’t answer.
If you can’t rebuild a decision from the log, you can’t defend it or learn from it.
Repeat the drill before go-live and after every major change, such as a new model or data source, because version fields are the first to break silently.
Rebuild a past decision
Pick one decision your system made a while ago and try to rebuild it from the logs before anyone asks you to. Each question you can’t answer is a field to add. Building audit trails in from the start is part of how we approach custom AI software.
This note is general information, not legal advice.
Sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- ISO/IEC 42001:2023 Information technology — Artificial intelligence — Management system International Organization for Standardization (ISO) www.iso.org/standard/42001
- OECD AI Principles OECD oecd.ai/en/ai-principles
Ask an assistant about this note