Field notesTurkish & multilingual
Reading Turkish documents with AI: OCR, tables, stamps and signatures.
Most document errors start before the model: a scanner that reads ş as s, a table flattened into a paragraph, a stamp over an amount. Route each document type deliberately, measure errors on Turkish characters and key fields, and send unreadable pages to a person.
veridive6 min read
When an AI system misreads a document, the model usually gets the blame. Look at what it was given: text in which every “ş” became “s”, a table of line items flattened into one paragraph, an amount hidden under a company stamp. The model did its best with a damaged copy.
Most document errors start before the model, and Turkish documents add their own traps: letters that English-trained OCR drops, number formats that move the decimal point, and stamps and wet signatures on delivery notes and contracts. The fix is a pipeline that picks its route per document type, measures errors on Turkish characters and key fields, and sends unreadable pages to a person instead of guessing.
Where do document pipelines lose information?
At five points, and each looks like a model error by the time it reaches the output:
- Intake. Photos at an angle, low-resolution or re-scanned copies, several documents merged into one PDF.
- Text. Misread characters, and the wrong reading order: a bilingual contract with Turkish on the left and English on the right gets read across both columns, line by line.
- Structure. Tables flattened into text, rows split across pages, headers repeated into the body.
- Overlays. Stamps, signatures and handwritten corrections on top of printed text.
- Normalization. Dates, amounts and encodings converted with the wrong conventions.
Done when: every extracted value traces back to its page region and to the text the model saw.
What is different about Turkish characters?
The marks that set ç, ğ, ı, İ, ö, ş and ü apart are small: a cedilla, a breve, a dot or its absence, two dots. Low-resolution scans, faxes and blurry photos lose them first. OCR configured only for English can’t output them at all and substitutes the nearest Latin letter, so check the language setting first. The damage spreads downstream:
- Names and addresses. “Gökçe Şahin” becomes “Gokce Sahin”, and a search for the correct spelling misses the record.
- Matching. The supplier in the ERP is spelled with Turkish characters and the extracted name isn’t. Compare ASCII-folded forms to find candidates, but let a rule or a person confirm the match.
- Valid wrong words. Some misreads produce another real word, such as “cam” (glass) for “çam” (pine), so a dictionary check won’t catch every error.
- Casing. OCR that drops the dot in “İRSALİYE” produces “IRSALIYE”, and correct Turkish lowercasing then turns it into “ırsalıye”. Casing rules can’t repair an OCR error.
Numbers need the same care. “4.500,00” means four thousand five hundred; a parser expecting English formats reads “4.500” as four and a half. IBANs and Turkish identity numbers carry check digits, so validate them instead of trusting the text.
Done when: errors on Turkish letters are measured separately, and names and addresses are checked against master data.
How do you choose between digital text, OCR and vision models?
Detect what you have first, then route it:
| Input | How to recognize it | Route |
|---|---|---|
| Digital PDF | A clean text layer with valid characters | Read the text layer directly, keeping positions |
| Scanned PDF | No text layer, or a garbled one from an old OCR run | OCR with Turkish enabled, plus layout analysis |
| Phone photo | An image with perspective, shadows or glare | Straighten and crop, then OCR or a vision model |
| Structured e-document | Arrives as data | Parse the data; never OCR a printout of it |
A text layer is not proof of a digital PDF: someone may have run a poor OCR over it long ago, so test its quality. Structured electronic invoices already arrive as data, as invoice automation in Türkiye explains; reading their rendered copy only adds errors.
Vision models, which read the page image directly, handle messy layouts well but fill gaps plausibly: allow “unreadable” as an answer and cross-check their numbers.
Done when: each document type has a named route chosen on a sample of real files, and anything that fits no route goes to a person.
How do you keep tables intact?
Line items, quantities and prices live in tables. Flattened into text, they leave the model guessing which number belongs to which row. Extract tables as tables, with rows, columns and cell positions, and handle merged cells, multi-line descriptions and tables that continue on the next page.
Then validate. Quantity times unit price should equal the line total, the lines should add up to the subtotal, and the subtotal, KDV and any withholding should reconcile to the total. When a check fails, flag the table with the failing cells highlighted; never let the pipeline adjust a number to make the sum work. The same discipline applies to any structured output from a model.
Done when: every extracted table passes its arithmetic checks or is flagged.
What about stamps, signatures and handwriting?
Turkish business documents carry a kaşe, the company stamp, and an imza, often a wet signature, and both regularly land on printed text. Treat them as objects to detect, not to interpret: present or absent, where, and whether they cover a field that must be read. Whether a signature is genuine, or a stamp belongs to the right company, is a check for people. Handwritten corrections, such as a crossed-out quantity with a new one written beside it, are exceptions for a person.
Picture an illustrative scanned delivery note from a regional supplier. OCR reads the delivery address as “Sisli” instead of “Şişli”, and a blue kaşe covers the quantity on the second line. Without checks, “Sisli” goes into the address field and the model guesses the quantity from context. With them, three things happen. The address check finds “Şişli” in the customer master data and proposes it as a marked correction. The stamp detector reports the overlap, so the quantity is set to “unreadable” instead of filled in. And the line total divided by the unit price suggests a value, shown as a suggestion only. The note goes to a person with the page image and the stamp region highlighted, who confirms the quantity against the original.
A field marked “unreadable” is a correct answer. A guessed one is not.
How do you measure extraction quality?
Field by field, on an evaluation set of real documents covering your suppliers, scan qualities and input types, including phone photos and stamped, handwritten pages. The people who process the documents label the correct values. Then measure:
- Field accuracy for each key field, after normalizing formats.
- Errors on Turkish letters, reported separately.
- Table accuracy: cells correct and rows aligned.
- Stamp and signature detection: present or absent, correctly.
- Silent errors: wrong values nobody flagged, the expensive ones, tracked apart from flagged cases.
Report by document type and route, and agree a threshold per field with the workflow owner.
One document type first
Take one document type, such as delivery notes, and collect a sample that includes the worst copies you receive. Label the key fields, run your current pipeline, and count the Turkish character errors and silent errors; the results show which route each input needs. Drafting ERP entries from these documents is the core of ERP and enterprise workflows, the pipelines underneath are data and AI foundations work, and the same parsing serves any document assistant.
Ask an assistant about this note