Every override is a lesson: building a feedback loop that improves AI.
Thumbs-up buttons collect noise. Structured feedback from the review step (what was changed and why), triaged weekly by the owner, turns daily corrections into better sources, prompts and evaluation cases.
veridive5 min read
Most AI tools ship with a thumbs-up and a thumbs-down. A few months in, someone exports the clicks: a scattering of thumbs-down with no comment, some enthusiastic thumbs-up from the same few people, and nothing anyone can act on.
Meanwhile the useful feedback is thrown away every day. Each time a reviewer corrects a draft or overrides a recommendation, they know exactly what was wrong. Capture it at the review step (what was changed and why), have the owner triage it every week, and daily corrections turn into better sources, prompts and evaluation cases.
Why do feedback buttons rarely improve anything?
- They are optional. Only the annoyed and the delighted click, so the sample says little about ordinary cases.
- They are vague. A thumbs-down on a refund draft could mean the wrong policy, the wrong amount or the wrong tone.
- Nobody owns them. No routine reads the clicks, so nothing changes, and people notice.
- They sit beside the work, not in it. The edit a person actually made is the real signal, and buttons don’t record it.
The correction a person makes is worth more than any rating they give.
What feedback should the review step capture?
Three things, recorded automatically or with one click:
- The edit. The difference between the draft and what was finally used.
- The reason. For every override, one choice from a short fixed list: wrong source, missing data, policy changed, tone or other, with optional free text.
- The context. Case type, sources cited, and the prompt and model versions.
Keep the list short enough to choose from in a second, and add a reason only when “other” keeps collecting the same explanation. How to fit this into the screen without slowing people down is covered in designing a review screen.
Who triages it, and how often?
The workflow owner, weekly, with one experienced reviewer and the engineer who maintains the system. An hour is usually enough. The routine:
- Group the week’s overrides by reason and case type.
- Read a sample from each group; counts show where to look, examples show what happened.
- Decide for each pattern: fix it, watch it, or accept it as a legitimate judgment call.
- Record every confirmed error as an evaluation case with the corrected answer, before anyone touches a fix.
- Assign each fix an owner and the evaluation cases it must pass.
Step four is the one teams skip. It proves the fix works, and it stops a later change from quietly bringing the error back. Accepting is a legitimate outcome too: some overrides are judgment calls the system should never make, and the right change is to send those cases straight to a person.
What kinds of fixes does feedback lead to?
Most fixes aren’t to the model. They fall into five kinds:
- Source content. An outdated or conflicting document is retired or corrected by its owner.
- Retrieval. The right document exists but isn’t found, so indexing, metadata or search settings change.
- Prompt. Instructions, examples or output format change, as a new tested version.
- Rules. A threshold or eligibility check moves out of the prompt and into code.
- Training for people. Reviewers or users need to learn a new policy or how to use the escalation path.
Prefer the fix closest to the cause. A prompt patch for an outdated document hides the problem until the next document changes.
Here is an illustrative example, with invented numbers: one week of overrides from a returns team, sorted into fixes.
| Reason | Cases | What the sample showed | Fix |
|---|---|---|---|
| Wrong source | 14 | Last season’s outlet policy still being cited | Source: retire the old page |
| Policy changed | 9 | A new rule for opened electronics, not yet published | Source and rule: publish it, add a check in code |
| Tone | 6 | Replies to repeat customers sounded curt | Prompt: two real examples of the right tone |
| Missing data | 5 | Photos from one marketplace channel never arrived | Retrieval: fetch that channel’s attachments |
| Other | 3 | Goodwill refunds, decided case by case | None: a person’s call |
Every row with a fix also produced new evaluation cases before the fix shipped.
How do you show people their feedback mattered?
Tell them, briefly and regularly. After each triage, post a short “what changed” note where the team works: what reviewers flagged, what was fixed, what was deliberately left alone and why. Name the catches that led to fixes. When a fix goes live, say so, so people can check it themselves. In the returns example, the note might read: “You flagged answers citing the old outlet policy. The page is retired and those cases are now in our tests. New tone examples for repeat customers are live.”
People who see their corrections lead somewhere keep giving reasons; people who don’t, stop. Making this a routine rather than a courtesy is part of turning a tool into everyday practice, the work behind AI enablement.
How do you know the loop is working?
- The override rate falls for the case types you fixed, and stays down.
- The same reason doesn’t come back from the same cause.
- The evaluation set grows with real cases every month.
- The time from spotting a pattern to shipping its fix gets shorter.
- Most overrides still carry a reason, and “other” stays small.
One caution: a falling override rate can also mean reviewers stopped looking. Check with known-answer cases, as described in why reviewers stop checking AI output.
The first weekly triage
Add the five reasons to your review step, book an hour a week with the workflow owner, and send the first “what changed” note after the first triage. The same rhythm, reviewing what changed with the owners and choosing the next improvement on evidence, runs through AI reliability.
Ask an assistant about this note