Field notesStrategy & leadership
An AI glossary for business teams, in plain words.
The terms you hear in every vendor meeting, each defined in two plain sentences with the one question it should make you ask. Grouped by the decision each term affects, not by the alphabet.
veridive8 min read
The vendor meeting moves fast. The answers are “grounded”, the agent is “fully autonomous” and the model can be “trained on your data”. Everyone nods, because asking what a word means feels like admitting you are behind.
It isn’t. Each of these words hides a question that decides whether the system will work for you. Here each term gets two plain sentences and that question, grouped by the decision it affects, not by the alphabet.
How should you use this glossary?
Read the group that matches the decision on the table; terms with their own field note link to it. Product names, versions and prices are left out, so the definitions stay true as tools change.
Take an illustrative example: a vendor calls its assistant “grounded in your data”. Three entries below turn that into questions: which sources does it search, does every answer cite one, and does it respect who may see what?
The useful part of every AI term is the question it should make you ask.
Which terms describe the models themselves?
These decide which model you use and where it runs.
- Large language model (LLM). A model trained on vast amounts of text to predict what comes next, so it can draft, summarize, classify and answer. It sounds just as sure when it is wrong. Ask: who checks its output?
- Token. The unit a model reads and writes, often a word or part of one. Usage is priced and limited in tokens, and Turkish often needs more of them than English. Ask: what does one finished task cost?
- Context window. The most text a model can consider in one request, instructions and documents included. A bigger window costs more per call and doesn’t mean the model uses it all well. Ask: what gets left out?
- Prompt. The instructions, examples and material sent with each request. In a business system it is part of the software, so version and test it like code. Ask: who can change it?
- Open-weight model. A model whose trained parameters are published, so you can run it on your own servers or cloud. You control where data is processed, and take on hosting, security and updates. Ask: who will run it?
- Fine-tuning. Extra training of a model on your examples, to change its style, format or behavior on a narrow task. It teaches patterns, not changing facts, so retrieval usually fits documents better. Ask: why not retrieval?
Which terms describe how systems find information?
These decide whether answers come from your approved sources, the core of data and AI foundations.
- Embedding. A list of numbers that captures what a text means, so similar meanings sit close together. It lets a system match a question to a passage that uses different words. Ask: was it tested on our language and jargon?
- Semantic search. Search by meaning rather than exact words. It finds paraphrases but can miss codes, names and numbers, which is why keyword search still matters. Ask: does it find a part number?
- RAG (retrieval-augmented generation). The system first searches your sources, then the model answers from what it found. It is the usual way to answer from your documents without retraining anything. Ask: which sources, and does every answer cite one?
- Grounding. Tying an answer to material retrieved for that question, not to what the model absorbed in training. A grounded answer can still be wrong if the source is out of date. Ask: what happens when no source fits?
- Citation. A pointer from a claim to the exact passage, page or moment behind it. It lets a reviewer check a claim in seconds. Ask: does the passage support the claim?
- Permission-aware retrieval. Retrieval that returns only what the person asking may already open. Without it, an assistant can quote a file to someone who could never open it. Ask: are permissions checked at question time?
Which terms describe systems that act?
These decide what a system may do without a person.
- Agent. Software that uses a model to choose its next step and take it: looking things up, calling tools, changing records. It suits paths that can’t be known in advance, and is harder to test. Ask: what can it do alone?
- Workflow. A fixed sequence of steps with model calls at chosen points, such as classify, extract, then draft. Many business processes are better served by one than by an agent. Ask: which step needs a model?
- Tool calling. How a model asks the surrounding software to run a defined action, such as looking up an order. The software, not the model, decides whether the call is allowed. Ask: which tools, with which permissions?
- Structured output. Output in a fixed format, such as named fields, that software can validate before use. It is how an invoice becomes a draft entry. Ask: what happens when a field fails a check?
- Human in the loop. A person reviews, approves or corrects the work at defined points. It works only if the reviewer has the evidence, the time and the authority to say no. Ask: where does a person decide?
- Autonomy level. How far a system may go alone: suggest, draft, act with approval, or act and report. Tasks in one workflow often sit at different levels. Ask: what evidence moves a task up?
Which terms describe quality and risk?
These decide how you know it is good enough, and safe.
- Evaluation set. Real examples with the answers an expert would accept, plus scoring rules. It is the acceptance test before launch and the regression test after every change. Ask: who wrote the answers?
- Acceptance criteria. Written thresholds for quality, cost and speed that a system must meet to go live, including errors that are never acceptable. They turn “it looks good” into a decision. Ask: were they agreed before the demo?
- Hallucination. Output that is fluent and confident but not supported by the sources or the facts. It can’t be switched off, only reduced and caught by design. Ask: how often, on our cases?
- Guardrails. The checks and limits around a model: permissions, input and output checks, approvals, spending limits. The ones that matter live in ordinary software, not in the prompt. Ask: which still hold when the model is wrong?
- Prompt injection. Instructions hidden in text the system reads, such as an email, that try to change what it does. It matters most when the system can act or see sensitive data. Ask: how was it tested?
- LLM as a judge. One model scoring another’s answers against a rubric, so evaluations can run often. The grader needs checking against people too. Ask: does it agree with our reviewers?
Which terms describe running a system day to day?
These decide what it costs and whether it stays good after launch.
- Latency. The time from a request to a usable answer. A live chat needs seconds and an overnight batch can wait, which changes design and cost. Ask: what does the workflow need at peak?
- Cost per task. The full cost of one finished piece of work: model usage, infrastructure, review time and upkeep. Compare this, mistakes included, rather than token prices. Ask: what does review add?
- Routing. Sending each case to the cheapest option that handles it reliably: a small model, a larger one or a person. The thresholds come from the evaluation set, not instinct. Ask: which cases go where?
- Drift. A slow change in quality with nobody touching the system, because documents, the case mix or the provider’s model changed. It often shows first as more overrides. Ask: would we notice within days?
- Model version. The specific release a system uses; providers update and retire them. Pinning one keeps behavior stable until you choose to move. Ask: how much notice do we get?
- Audit trail. A record of what the system saw, which sources it used, what it produced and who approved it. It lets you rebuild a decision long afterwards. Ask: could we reconstruct a disputed case?
Which words should make you ask a follow-up question?
Some words sound like answers but hide the question. Ask the one beside it.
| When you hear | Ask instead |
|---|---|
| “Fully autonomous” | Which actions happen without a person, and can they be undone? |
| “AI-powered” | Which step uses a model, and what happens when that step is wrong? |
| “No hallucinations” | Measured how, on whose cases? What does it do when it finds no source? |
| “Trained on your data” | Fine-tuned, retrieved or stored? Where is it processed, and does it train anything else? |
| “Enterprise-grade security” | Where is data processed and logged, and who can see the logs? |
| “It learns as you use it” | What changes, who approves it, and can it be rolled back? |
Before the next vendor meeting
Send the relevant questions ahead of a vendor meeting; good vendors answer in writing. To apply them to one of your own workflows, what we do starts there, and how we work shows where evaluation sets and approval points fit.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) arXiv arxiv.org/abs/2005.11401
- LLM01:2025 Prompt Injection OWASP Gen AI Security Project genai.owasp.org/llmrisk/llm01-prompt-injection
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
Ask an assistant about this note