veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesTurkish & multilingual

Arabic and English in one AI system: what to plan for in the Gulf.

In the Gulf, work moves between Arabic and English, and Arabic itself spans formal written Arabic, dialects and Latin-script chat. Plan for script and direction, name spellings across scripts and native reviewers, and test quality per language instead of assuming it.

veridive5 min read

A customer record in the CRM holds a name in Arabic script. The same customer emails in English and signs with a Latin spelling of that name, then follows up on WhatsApp in Gulf dialect, partly in Latin letters. The policy that answers the question is in English, and the reply should go out in Arabic. One conversation, two languages, three ways of writing one of them.

That is ordinary bilingual work in the Gulf, and an AI system built for English handles it badly in predictable places. It doesn’t take a special model. It takes planning for script and direction, for names across scripts and for native review, and testing quality in each language instead of assuming it.

What does bilingual Arabic and English work look like?

In many organizations, systems, contracts and internal documents run in English, while customers and some colleagues write in Arabic. The patterns resemble Turkish and English work: questions in one language, sources in the other, and messages that mix both. Two things are different. Arabic uses its own script, written right to left, and “Arabic” covers several varieties that a system must treat as distinct inputs.

What is different about Arabic for AI systems?

  • The script. Letters join and change shape by position, and text runs right to left.
  • Short vowels are usually not written. One written word can have several readings, resolved by context.
  • Spelling variants are common. Different forms of alef, final letters such as tāʾ marbūṭa and hāʾ, and decorative stretching characters are typed interchangeably. Search needs normalization for them, applied carefully, because merging letters also merges some distinct words.
  • Words carry attached parts. The article, some prepositions and conjunctions attach to the word, so one Arabic word can equal an English phrase. As with Turkish suffixes, search needs an Arabic-aware analyzer, on documents and on queries alike.
  • Address is gendered. “You” and the verb forms that go with it differ for men and women, so a reply must know how to address the customer or be phrased to avoid it. Defaulting to the masculine form is an easy mistake for generated text.

None of this is exotic for well-built software, but systems configured with English defaults miss all of it without raising an error.

How do dialects and Latin-script Arabic affect results?

Modern Standard Arabic is the formal written language of contracts, official letters and most documentation. Everyday speech, voice notes and informal chat use dialects, which differ from it and from each other in vocabulary and grammar. A customer in the Gulf may write formally in an email and in dialect on WhatsApp.

Chat adds Arabic written in Latin letters, often called Arabizi, with digits standing in for sounds the Latin alphabet lacks: a 3 for the letter ain, a 7 for haa. It is spelled inconsistently, mixes freely with English, and is easy for a system to misread as broken English. Don’t assume a model that handles formal Arabic handles dialect or Arabizi; collect real examples of each and test.

Arabic is not one input. Formal text, dialect and Arabizi each need their own tests.

What changes in the interface?

  • Layout. Right-to-left pages mirror alignment and navigation, but not everything mirrors: logos, media controls and some charts keep their direction. Decide element by element.
  • Mixed-direction text. An Arabic sentence containing an English product name, an order number or an email address can display with punctuation or number groups in the wrong place. Isolate each inserted value’s direction, and test with real strings, not placeholder text.
  • Numerals. Arabic text may use Western digits or Arabic-Indic digits. Parsers and search must accept both, and one number should never mix them.
  • Every surface. Check the chat widget, email, PDF exports and messaging apps separately, because each renders direction in its own way.

How do you match names across scripts?

One Arabic name has many Latin spellings, and a CRM may store a short form while an email signature shows a longer one with a father’s name. Match in layers: normalize both sides (hyphens, spacing, “Al” and “El” prefixes, doubled letters), generate transliteration candidates from the Arabic form, and score the similarity. Then confirm with something other than the name, such as a phone number, email address or customer number, because common names are shared by many people. Uncertain matches go to a person, and no record is merged on a name alone.

Here is an illustrative case. The CRM holds محمد المنصوري in Arabic script. An email arrives signed “Mohamed Almansouri”. The matcher normalizes the Latin name to “mohamed mansouri”, generates candidates from the Arabic form (Mohammed, Muhammad and Mohamed; Al Mansoori, Almansouri and Al-Mansuri), and finds a strong match. The sender’s email address is not on file, so instead of linking automatically, the system proposes the match to an agent with the evidence shown, and the agent confirms it using the order number in the email.

How do you evaluate quality in both languages?

Split the evaluation set by language and variety: English, formal Arabic, dialect, Arabizi and mixed messages, with sources in both languages. Native Arabic reviewers, familiar with the dialects your customers actually use, score the Arabic answers for meaning, tone and form of address; English reviewers score the English ones. Add interface checks to the same set, including screenshots of mixed-direction replies. Report every slice separately, and run each candidate model on the same set, because quality in one language says little about another.

Test before choosing a model

Collect a few dozen real messages from each variety your customers use, and a list of names as they appear in your CRM and in email. Test retrieval, answers, name matching and rendering on them before choosing a model. The layers underneath, such as analyzers, name matching and evaluation, are data and AI foundations work, and for customer messages the design runs inside customer operations.

Ask an assistant about this note

Turkish & multilingualArabicEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

Can one AI assistant work in both Arabic and English?

Yes, if it is designed for both from the start. That means answering in the user’s language, handling formal Arabic, dialects and Arabic typed in Latin letters, rendering right-to-left and mixed-direction text correctly, and matching names across scripts. Quality in each language has to be tested on your own examples with native reviewers, not assumed from English results.

What is Arabizi and why does it matter for AI?

Arabizi is Arabic written in Latin letters, often with digits standing in for sounds the Latin alphabet lacks, such as 3 for the letter ain and 7 for haa. It appears in informal chat and is spelled inconsistently. A system that expects only Arabic script or English can misread these messages, so include real examples in the evaluation set.