# Turkish speech-to-text for calls and meetings: what affects accuracy.

> In this field note, veridive explains what affects Turkish speech recognition in calls and meetings: phone audio, accents, overlapping speech, numbers, names and English terms. It covers measuring word error rate and key-term accuracy on your own recordings, keeping speaker labels and timestamps, processing questions for the DPO, and designing review around the words that matter.

Transcription quality decides everything downstream in call and meeting intelligence. Errors cluster in phone audio, accents, crosstalk, numbers, names and English terms inside Turkish speech; measure on your own recordings and design review around the words that matter.

## Key takeaways

- Every summary, score and search result inherits the transcript’s errors, so transcription quality comes first.
- Errors cluster in phone audio, accents, overlapping speech, numbers, addresses, product names and English terms inside Turkish speech.
- Measure word error rate on a sample of your own calls, and exact accuracy on the key terms that drive decisions.
- Keep speaker labels and timestamps, validate numbers against systems, and send uncertain key terms to a person with the audio.

The call summary says the customer’s payment went through. The recording says “ödeme yapılmadı”, the payment was not made. The transcript dropped one short syllable, and with it the negation. Everything built on that transcript — the summary, the quality score, the follow-up task — is now confidently wrong.

Speech-to-text, or automatic speech recognition, turns audio into text with a timestamp for each word. In call and meeting intelligence it is the first step, and every later step reads the transcript, not the audio. So its quality decides everything downstream. In Turkish calls, errors cluster in predictable places: phone audio, accents, crosstalk, numbers, names and English terms. Measure on your own recordings and design the review around the words that matter.

## Why does transcription quality decide everything downstream?

A language model summarizing a transcript can’t hear the call. It works with the words it is given and writes fluently either way, so a transcription error becomes a well-phrased wrong finding. Quality scores, compliance checks, meeting minutes and search all inherit the same errors.

The errors are not spread evenly. Greetings and common phrases are easy; the hard words are the ones that carry decisions: amounts, account numbers, names, product names and negations. An overall accuracy figure can look good while exactly those words are wrong.

## What makes Turkish calls hard to transcribe?

Six things, most of them common to every contact center and a few specific to Turkish:

- **Phone-quality audio.** Phone lines carry a narrow band of frequencies and compress the voice, removing some of the detail that separates sounds such as “s” and “ş”.
- **Accents.** Regional accents from across Türkiye, and callers whose first language is not Turkish.
- **Overlapping speech.** Interruptions, both sides talking at once, and constant backchannel sounds such as “hı hı” and “evet evet”.
- **Numbers.** Turkish numbers are spoken as compound words, and grouping changes the digits: “on iki” is twelve, but with a pause after “on” it can be written as “10 2”, three digits where two belong.
- **Names and addresses.** Surnames, street and neighborhood names are rare words. When callers spell them, they use city names for letters (“Adana’nın A’sı”), and a careless transcript fills up with cities.
- **Product names and English terms.** Plan names, brand names and English words inside Turkish sentences, often with Turkish suffixes attached.

Turkish grammar adds one more wrinkle: meaning sits in suffixes, so a single missed syllable can flip a statement, as “yapıldı” and “yapılmadı” show.

## How do you measure accuracy on your own recordings?

Start with word error rate (WER): the substitutions, deletions and insertions in a system’s transcript, divided by the number of words in a careful human transcript of the same audio. Take a sample of your own calls across queues, line quality, times of day and agents, and have people transcribe it to a written style guide that covers how to write numbers, filler words and English terms.

Three details decide whether the number means anything:

1. **Normalize both transcripts the same way** before scoring, including Turkish-aware casing, punctuation and numbers. Otherwise you measure formatting, not recognition.
2. **Read WER with character error rate.** Long suffixed words turn one wrong letter into a whole wrong word, so WER looks harsh in Turkish. Character error rate shows near-misses, but don’t let it flatter the system: “yapıldı” and “yapılmadı” are close in characters and opposite in meaning.
3. **Measure key terms separately.** List the words that drive decisions, such as order numbers, amounts, IBANs, product names, negations of key verbs and required compliance phrases, and count how often each is exactly right.

Compare engines on the same sample, and report by queue and line quality rather than as one average.

## What about speakers, timestamps and numbers?

**Speakers.** Diarization decides who spoke when. If your platform can record agent and customer on separate channels, do it, and speaker labels become exact. In mono recordings and meetings, diarization errors put words in the wrong mouth, so a refund promise can be credited to the customer. Measure speaker attribution on the sample too.

**Timestamps.** Keep them for every word or segment, so every finding links to the moment it was said and a reviewer can hear it in seconds. That is the idea behind [answers with receipts](https://veridive.com/insights/answers-with-receipts/), applied to audio.

**Numbers.** Convert spoken numbers to digits consistently, then validate them: order numbers against the order system, IBANs and identity numbers against their check digits.

Take an illustrative example: a customer asks for a refund to a different account, reads out an IBAN in groups (“TR, on iki, otuz dört…”) and spells the street name letter by letter. The review should check four things. The IBAN’s check digits show whether any digit was misheard; if the check fails, the IBAN is marked unverified and a person replays that moment. The spelled letters appear as letters, not as a list of cities. The IBAN is attributed to the customer, not to the agent reading it back. And the timestamp takes the reviewer straight to the second it was said. Once verified, the IBAN is masked in the stored transcript.

> A transcript can be mostly right and wrong about the one number that matters.

## Where should recordings be processed?

Recordings carry voices, names, account details and sometimes health information, so settle these questions with your data protection officer before anything is processed:

- Where are audio and transcripts processed and stored, and does using a provider abroad count as a transfer under KVKK or GDPR?
- Does the recording notice callers hear cover transcription and automated analysis?
- How long are audio, transcripts and derived data such as scores kept, and who can access each?
- Could voice data count as biometric data in your use, and how are health details mentioned in calls handled?
- Agents are recorded too: which questions does this raise for HR and counsel?

Deployment changes the answers: processing in a specific region, in your own cloud, or on-premises with an open-weight speech model.

## How do you design around the remaining errors?

No engine is perfect, so design for the errors that remain:

- **Review key terms, not whole transcripts.** Highlight amounts, account numbers and negations with their confidence, and send uncertain ones to a person with the audio snippet.
- **Validate against systems.** Check order numbers, product names and amounts against the records they refer to.
- **Teach the vocabulary.** Where the engine allows it, add your product names, plan names and common street names as custom vocabulary.
- **Don’t score what the transcript can’t support.** Tone of voice and sarcasm need a person; a [call quality scorecard](https://veridive.com/insights/call-quality-scorecard/) should say which criteria stay human.
- **Re-measure after changes.** New products, campaigns and engines change the error profile.

## Measure on your own calls

Pick a few hours of calls from one queue, have them transcribed carefully by people, and list the key terms that matter for that queue. Run your current engine, or two candidates, and measure WER and key-term accuracy side by side. For contact centers this is the first step of [customer operations](https://veridive.com/solutions/customer-operations/) work on calls, and it underpins every finding in [voice and meeting intelligence](https://veridive.com/solutions/voice-meeting-intelligence/).

This note is general information, not legal advice.

## Frequently asked questions

### How accurate is Turkish speech recognition on phone calls?

It depends on your audio, your callers and the engine, so the only reliable answer comes from your own recordings. Transcribe a sample of real calls by hand, compare each engine’s output with it and calculate the word error rate. Then check the terms that drive decisions, such as numbers, names and amounts, separately, because a good average can hide errors exactly where they matter.

### What is word error rate in speech recognition?

Word error rate measures how many words a speech recognition system gets wrong compared with a careful human transcript: substitutions, deletions and insertions, divided by the number of words in the reference. Lower is better. In Turkish, long suffixed words make one small error count as a whole wrong word, so read it alongside character error rate and key-term accuracy.

### What is speaker diarization?

Speaker diarization is the step that decides who spoke when, so each part of a transcript is labeled with its speaker. In contact centers, recording the agent and the customer on separate channels makes the labels exact. In mono recordings and meetings, diarization errors can put a promise or a complaint in the wrong person’s mouth, so measure them too.
