veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesData & models

Why keyword search still matters in AI retrieval.

Vector search finds meaning; keyword search finds exact things such as part numbers, invoice numbers, legal terms and names. Business questions need both, plus reranking, and Turkish word forms make the keyword side harder and more important.

veridive6 min read

A maintenance engineer types “HX-4471 torque setting” into the new document assistant. Retrieval returns three passages about torque settings for similar parts and none about the HX-4471. The answer that follows is fluent, cited and about the wrong part.

Nothing is wrong with the model. Retrieval, the search step of RAG, was built on vector search alone, which is good at meaning and poor at exact strings, and business questions are full of exact strings: part numbers, invoice numbers, clause numbers, names. Retrieval for real work needs both kinds of search, a reranker to order the shortlist, and filters for things that aren’t about relevance at all, such as dates and permissions. In Turkish, the keyword side needs extra care.

What does vector search do well, and what does it miss?

Vector search turns each passage and each question into an embedding, a list of numbers that places texts with similar meaning close together, and returns the passages nearest the question. It is good at paraphrase: “can my spouse come on a business trip?” finds the passage headed “accompanying family members”. Depending on the embedding model, it can also match a Turkish question to an English passage.

It struggles with strings that carry little meaning of their own. Illustrative queries where it tends to fail:

QueryWhat vector search tends to returnWhy
“HX-4471 torque setting”Torque values for similar partsCodes that differ by one character look almost identical as vectors
“clause 9 termination”Termination clauses from any contractNumbers carry little meaning in an embedding
“Selin Aksoy approval limit”Approval limits for managers in generalA name is weak evidence next to the topic words

It also blurs small differences in wording: “with approval” and “without approval” sit close together.

What does keyword search still do better?

Keyword search, also called lexical search, scores passages by the words they share with the query. The common scoring method, BM25, rewards terms that are rare across the collection and frequent in the passage, which is exactly right for identifiers: “HX-4471” appears in a handful of passages, so those passages score high. Keyword search also wins on exact phrases and terms of art (“force majeure”, “mücbir sebep”), acronyms, invoice and order numbers, and names.

Its weaknesses mirror vector search: it misses synonyms and paraphrases, and it depends on text being split into words the way users type them. The tokenizer decides whether “HX-4471” stays one token or splits into two, so test that codes stay intact and still match “HX4471” and “hx 4471”.

Vectors find what a passage means. Keywords find what it says, character for character.

How does hybrid search combine the two?

Run both searches on the same question, then merge the result lists. Two merge methods are common:

  • Score combination: normalize each list’s scores to the same range and take a weighted sum. Simple, but the weights need tuning, and score scales drift as the collection changes.
  • Rank fusion: score each passage by its position in each list; reciprocal rank fusion, for example, adds up one divided by a constant plus the rank. It ignores raw scores, so it is harder to break.

Two refinements matter in practice. Query routing: a query that looks like an identifier, such as a part-number or invoice-number pattern, can go to an exact lookup first or get more keyword weight. Metadata filters: document type, owner, effective date, language and, above all, permissions are applied as filters inside both searches, before any ranking. Permissions are never a ranking signal: a passage the user may not open must not appear in either list.

Where does reranking fit?

After the merge, take a shortlist of a few dozen passages and rescore each with a reranker: a model that reads the question and the passage together and judges how well one answers the other. It is slower than either search, which is why it sees only the shortlist, and usually more accurate, because it compares the actual words rather than precomputed vectors. The top few passages go to the language model.

Rerankers have blind spots too. Check that yours respects exact matches (it should not push the HX-4471 passage below a fluent one about a similar part) and that it handles Turkish as well as English.

What changes for Turkish and other richly inflected languages?

Turkish builds meaning with suffixes, so one word appears in many forms: “sözleşme” (contract) also appears as “sözleşmenin”, “sözleşmedeki” and “sözleşmelerimizde”. An index of raw tokens misses most of them. What makes Turkish hard for AI covers the language itself; for retrieval, four points matter.

  • Analyze documents and queries the same way. Stemming (cutting suffixes by rule) or lemmatization (mapping words to their dictionary form) must run on both sides, or query and index never meet.
  • Watch over-stemming. Aggressive rules merge unrelated words and flood results with noise, so compare analyzers on real queries.
  • Split apostrophe suffixes. Turkish attaches suffixes to names and codes after an apostrophe, as in “HX-4471’in tork değeri” or “Aksoy’un onay limiti”. If the tokenizer keeps them attached, the exact match is lost.
  • Keep an ASCII-folded field. Users type “sozlesme” for “sözleşme”, so index a folded copy for recall and rank exact matches higher.

Embedding models and rerankers vary in Turkish quality as well, so test the vector side on the same real queries.

How do you measure retrieval quality?

Separately from answer quality. If the right passage never reaches the model, no prompt can fix the answer.

  1. Mark the gold passages: for each question in the evaluation set, the passage or passages that contain the answer.
  2. Measure recall at k: was a gold passage in the top k results, where k is the number of passages you send to the model?
  3. Measure rank: how high the first gold passage appears; mean reciprocal rank summarizes it across questions.
  4. Break results down by query type: identifiers, names, paraphrases, Turkish, English and mixed.
  5. Compare configurations on the same set: vector only, keyword only, hybrid, and hybrid with reranking.

Re-run the comparison whenever the analyzer, the embedding model, the chunking or the reranker changes.

Test the search itself

Pull fifty real queries from your search logs, mark those with codes, clause numbers or names, and check whether the right passage lands in the top five under each configuration. We build retrieval, and the evaluation that measures it, in data and AI foundations projects; knowledge and document intelligence shows where it ends up. For the other places answers go wrong, see why document assistants give wrong answers.

Sources

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) arXiv arxiv.org/abs/2005.11401

Ask an assistant about this note

Data & modelsSearchRAG

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is hybrid search in RAG?

Hybrid search runs keyword search and vector search on the same question and merges their results before the language model sees any passage. Keyword search finds exact strings such as part numbers, clause numbers and names; vector search finds passages that say the same thing in other words. Merged and then reranked, the two cover each other’s blind spots, which you can confirm on your own evaluation set.

Why does vector search miss product codes?

Because embeddings represent meaning, and codes carry little of it. Two part numbers that differ by one character look almost identical as vectors, and a rare code or a person’s name can be represented poorly. Keyword search matches the exact string and rewards rare terms, so it finds the passage that mentions that code. Business retrieval needs both kinds of search.