veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesTurkish & multilingual

What makes Turkish hard for AI, and how to design around it.

Turkish builds meaning with suffixes, has a dotted and dotless i that breaks naive software, and often shares a sentence with English. None of that stops a good system; it means testing on real Turkish text and designing search, casing and tone on purpose.

veridive6 min read

A document assistant is demonstrated in English, the answers are good, and the team switches it to Turkish for the pilot. Within a day the complaints are specific: searches miss passages that are plainly there, names in capitals come out with the wrong letter, and a reply calls the customer “siz” in one sentence and “sen” in the next.

None of this means Turkish is out of reach; the system was configured with English defaults. Turkish builds meaning with suffixes, has two letters i with their own capitals, marks formality in grammar and, at work, often shares a sentence with English. Each has a known design answer, and each is easy to miss until real Turkish text arrives.

So design search, casing and tone for Turkish on purpose, and test on your own Turkish text with native reviewers instead of assuming a model handles it.

What is different about Turkish for language technology?

Five features cause most of the trouble.

  • Suffixes carry the grammar. Turkish is agglutinative: plural, possession, case and tense are suffixes chained onto a root. “Fatura” (invoice) also appears as “faturalar”, “faturayı”, “faturanın” and “faturalarımızdan”, each a different string to a computer. Negation is a suffix too: “ödenecek” means “will be paid”, “ödenmeyecek” “will not be paid”.
  • Suffixes change shape. Vowel harmony gives one suffix several spellings (-ler or -lar; -de, -da, -te or -ta), and roots change at the edges: “kitap” becomes “kitabı”, and “izin” becomes “izni”.
  • Four letters i. Dotted i with capital İ, and dotless ı with capital I; default software knows only the English pair.
  • Two ways to say you. Formal “siz” and informal “sen” are marked in verbs and possessives, so register is grammar, not just word choice.
  • English inside Turkish. “PO’yu onayladım”, “deadline’ı kaçırdık”, “case’i escalate ettim”: English words take Turkish suffixes, usually after an apostrophe.

How do suffixes affect search and retrieval?

Keyword search matches exact tokens, so “fatura” misses “faturanın”. Wildcards catch regular forms but fail where the root changes (“kitap*” never finds “kitabı”) and overmatch on short roots: “mal” (goods) also finds “maliyet” (cost) and “malzeme” (material).

Two techniques reduce a word to its base. Stemming strips suffixes by rule: fast, and fine for regular forms. Lemmatization uses a morphological analyzer and a dictionary, so it also handles roots that change, such as “iznini” back to “izin”. Search engines usually include a Turkish analyzer; the common mistake is applying it to documents but not to queries.

Semantic search compares meaning rather than strings and copes better with word forms, but it is weak on exact items such as codes, names and clause numbers, and on one-word queries. Business search needs both, which is what hybrid search combines.

Many users type without Turkish characters: “siparis”, “sikayet”, “odeme”. Index an ASCII-folded copy for recall and rank exact matches higher, because folding merges real words such as “sık” (frequent) and “şık” (elegant).

Consider an illustrative case. An employee types “izin” into an HR policy assistant to learn whether unused leave carries over, and gets an answer about system access, because “izin” also means permission. The passage she needed, headed “Kullanılmayan izinlerin devri”, never reached the model: the index treated “izinlerin” as a different word, and semantic search matched the one-word query to everything about permissions. Two changes fixed it: a Turkish analyzer on documents and queries, so “izinlerin” and “iznini” match “izin”, and hybrid search, so the keyword match and the passage’s meaning count together. The passage now ranks first, and those queries are permanent test cases.

What is the dotted and dotless i problem?

In Turkish, i capitalizes to İ and ı to I; default casing, built for English, maps i to I. Some results are visible, some quiet:

  • “iade” uppercased by default becomes “IADE”, which a Turkish reader sees as a misspelling of “İADE”.
  • “KAPI” lowercased by default becomes “kapi”, not “kapı”.
  • “İADE” lowercased by many default implementations becomes an i followed by a separate combining dot, so a case-insensitive search for “iade” never matches the heading.

The fix is Turkish-aware casing for Turkish text and locale-independent casing for code, identifiers and system keys; apply Turkish rules to code, and “id” becomes “İD” and lookups fail.

Older files hide the same family of problems. A legacy Turkish encoding read as Western European turns “İş Sözleşmesi” into “Ýþ Sözleþmesi”; ö, ü and ç survive, so the damage is easy to miss. Documents also mix languages, with English headings over Turkish paragraphs: detect language per passage, not per file, and normalize encodings before indexing.

How do models handle quality and tone in Turkish?

It depends on the model, the task and the domain, and English results predict little. Test, and look for specific failures:

  • Grammar at the edges. Wrong case endings or clumsy suffix chains make text read as machine-made even when the facts are right.
  • Mixed address. “Siparişiniz kargoya verildi, takip numaranı aşağıda bulabilirsin” starts formal and ends informal.
  • Guessed courtesy titles. Turkish attaches Bey or Hanım to the first name, and a model guessing from a unisex name such as Deniz will sometimes be wrong.
  • Terminology. English terms your teams use daily, translated into Turkish nobody uses, or the reverse.
  • Number formats. “1.250,50 TL” uses a decimal comma; read with English conventions, “1.250” is one and a quarter.

Tokenizers built mostly on English often split Turkish into more pieces than equivalent English, so measure cost on Turkish text. Spotting these failures is a skill worth practicing in role-based training in Turkish.

How do you evaluate a system in Turkish?

Build the evaluation set from real Turkish inputs, including the awkward forms: helpdesk questions, search logs, emails and documents. Tag each example by phenomenon (inflected query, ASCII-typed, uppercase source, mixed language, formal or informal), so results show which one fails.

Native Turkish speakers who know the work should write the reference answers and score the output against a rubric covering grammar, form of address and terminology as well as facts. Run every candidate model on the same set, report Turkish results separately, and re-run the set whenever a model, prompt or analyzer changes.

“Supports Turkish” is a claim. Finding “izinlerin” when someone types “izin” is evidence.

What should you ask a vendor about Turkish?

  1. Which Turkish analyzer do search and retrieval use, on documents and on queries? Show a search for one word in three forms.
  2. How is casing handled? Show “iade” finding “İADE”, and “ışık” finding “IŞIK”.
  3. What happens to Turkish typed without Turkish characters?
  4. How are legacy encodings detected and repaired?
  5. Which models were tested in Turkish, on whose examples, and who scored them?
  6. How are form of address and terminology controlled and tested?

A good answer is a demonstration on your own files, not “our model supports Turkish”.

Real queries, real misses

Pull a few dozen real Turkish queries from search logs or the helpdesk, misspellings included, and note which passage each should find. Run them against your current search; the misses show whether the analyzer, casing or encoding needs attention first. The same checks apply to any document assistant answering in Turkish; the analyzers, casing and evaluation underneath are what data and AI foundations covers.

Ask an assistant about this note

Turkish & multilingualSearchEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

Why is Turkish difficult for AI and search systems?

Mainly because Turkish builds meaning by adding suffixes to a root, so one word appears in many forms and exact-match search misses most of them. Casing rules for the dotted and dotless i differ from English defaults, older files use legacy encodings, and English terms often appear inside Turkish sentences. Each has a known fix, but only testing on real Turkish text shows whether it was applied.

What is the Turkish i problem in software?

The Turkish i problem is a casing bug. Turkish has a dotted i with the capital İ and a dotless ı with the capital I, while default software casing maps i to I. Casing Turkish text with default rules produces misspellings and missed search matches, and applying Turkish rules to code breaks identifiers. Use locale-aware casing for text and locale-independent casing for code.

How do you test an AI system’s quality in Turkish?

Build an evaluation set from real Turkish questions and documents, including inflected words, text typed without Turkish characters, mixed Turkish and English, and formal and informal address. Have native Turkish speakers who know the work write and score the reference answers, and run every candidate model on the same set. Report Turkish results separately, so an average doesn’t hide weak spots.