Before the AI assistant: a clean-up checklist for your documents.
An assistant can’t be more current or consistent than its library. Before building, name an owner for each document set, retire duplicates and old versions, fix what scans and tables hide, and mark what is draft, final or expired.
veridive5 min read
The first wrong answer from a new policy assistant is rarely invented. More often it is a faithful quote from the wrong document: a travel policy replaced two revisions ago, still sitting in a shared folder, now retrieved and cited with complete confidence.
An assistant can’t be more current or consistent than its library. So the most useful work before building is unglamorous: name an owner for each document set, retire duplicates and old versions, fix what scans and tables hide, and mark what is draft, final or expired. The fifteen checks below come in five groups of three: scope, versions, format, ownership and permissions.
Why does document quality decide answer quality?
Because a document assistant answers from what its search finds, and search doesn’t know which documents are right. It finds old versions as readily as new ones, can’t read a scan without a text layer, and treats a draft in a shared folder as policy. RAG, explained for business teams covers the mechanics.
Contradictions do quieter damage. When two current documents disagree, the assistant quotes whichever passage ranks higher for that phrasing, so two people asking the same question in different words can get different answers, each with a valid-looking citation.
Every problem in the library becomes a problem in the answers, delivered fluently and with a citation.
Which documents should be in scope first?
- One document set with frequent questions and a named owner, such as HR policies or standard contracts. A bounded set can be cleaned and tested; “the whole intranet” can’t.
- A written list of what is in and what stays out. Drafts, personal folders, email archives and meeting notes stay out unless someone decides otherwise, because the assistant treats everything in scope as approved.
- The questions people actually ask, collected from the inbox or help desk. They show which documents matter, and they become the test set.
How do you deal with versions and duplicates?
- One current version per document. Superseded versions move to an archive outside the assistant’s scope, because search can’t tell old from new.
- Duplicates removed or replaced with links. Copies saved in several folders drift apart and compete in every result list.
- Metadata that helps retrieval: owner, effective date, status (draft, final or expired), audience and language. Search can filter on it, and citations can show which version an answer came from.
Here is an illustrative case. An HR team preparing its policy library finds three travel policies: the current one on the intranet, the previous version in a shared folder, and a copy with handwritten notes that finance scanned. The per-diem annex exists only as a scanned PDF with a rate table. Before go-live, the owner confirms the current version, the previous one moves to an archive outside the assistant’s scope, and the scanned copy is deleted. The annex is rebuilt as a real table and checked line by line against the scan, and effective dates and status go into the metadata. The first test run still turns up a stray: an FAQ page summarizing the old per-diem rates, which the owner retires.
What about scans, tables and images?
- Scans turned into checked text. OCR (optical character recognition) makes them searchable. Check a sample by eye, especially Turkish characters such as ş, ğ and ı, and stamps that land on top of text.
- Tables kept as tables. A rate table or approval matrix flattened into a line of numbers loses which figure belongs to which row. Check tables after conversion.
- Text for what images carry. A process drawn as a flowchart or a form shown as a screenshot may be invisible to search; add the steps or fields as text.
Who owns each document set, and who may open it?
Ownership
- A named owner per set: a person, not a department, who says which version is current and approves changes.
- Contradictions settled by the owner. When two documents disagree, that is a policy question, not a technical one; the assistant can’t be more consistent than the rules it reads.
- A route for changes. New versions are published in the source system, and the old one is marked expired the same day.
Permissions
- Access reviewed before indexing. The assistant inherits the source system’s permissions, mistakes included.
- Over-shared folders fixed. A file that “everyone” could technically open but nobody could find becomes findable in seconds.
- Roles tested. Ask as an employee, a manager and an HR specialist, and check what each can retrieve; access control for assistants shows how.
How do you keep it clean after launch?
A clean library decays unless someone owns the routine:
- Unanswered questions become a content backlog. Every “no source found”, thumbs-down and escalation goes to the owner as a list: a missing document, an unclear paragraph or a question outside scope.
- Review dates trigger reviews. A document past its review date is flagged to its owner instead of lingering.
- Owners see which documents are cited most, so review effort goes where the answers come from.
- New sets pass the same fifteen checks before they are added.
- Failures become test cases. A wrong answer traced to a document problem joins the evaluation set, so the fix stays fixed; why document assistants give wrong answers shows how to trace one.
Clean the busiest set first
Pick the document set behind the questions people ask most, and run the fifteen checks on it before anyone configures an assistant. That same choice, one set with a clear owner, is the first step of a knowledge and document intelligence project, and keeping sources connected, cleaned and current is part of data and AI foundations.
Ask an assistant about this note