A caller in Barka switches from Arabic to English mid sentence, drops a word only people from the interior use, and talks twice as fast as a news anchor. Most Arabic speech AI was trained on none of that. Every time it guesses wrong, the call still lands back on your desk, and the minutes it costs you add up on every shift.
Two Arabics, one phone line
Modern Standard Arabic, the formal Arabic used in news bulletins, textbooks and official letters, is not how anyone actually talks on the phone. What your customers speak is Omani Arabic: a spoken dialect with its own words, its own rhythm and its own shortcuts, and it shifts a little from Muscat to the interior to Dhofar. Add code switching, jumping between Arabic and English inside the same sentence, and you have the everyday sound of a real customer call.
This gap is not a small accent difference. A system trained only on formal Arabic can miss ordinary dialect words entirely, misread a fast sentence as noise, or freeze the moment a caller drops in an English word for a product name or a time. During khareef, when Dhofar's phone lines fill with visitors mixing Gulf Arabic, English and everything between, that gap gets tested on nearly every call.
Why the models stumble
Most Arabic language AI, including the voice systems already sitting behind many phone menus, learned its Arabic from whatever text and audio was easiest to gather and label at scale: news broadcasts, religious texts, government documents, formal writing. That pile barely contains spoken dialect, and it almost never contains a caller switching languages mid sentence.
The Arabic your customers speak on the phone was almost never in the data any language model learned from.
The result shows up as small, expensive failures. A system asks the caller to repeat something they already said clearly. It hears an English word inside an Arabic sentence and stalls. It cannot keep pace with a fast, casual sentence the way it follows a slow, formal one. Every one of those moments either forces a transfer to a human agent or leaves the caller repeating themselves, and both cost time your desk is already short on. For the deeper mechanics of how speech turns into text and back into a voice, that groundwork is covered elsewhere.
What building for the dialect actually takes
Getting this right is not a setting you switch on. It takes real recorded calls, not scripted studio readings, labelled by people who actually speak the dialect, including its regional variants and its code switch points. It takes a feedback loop: logging every call the system misunderstood and correcting it against exactly that failure, again and again, because slang and word choice shift by generation and by region.
Oman is making its own attempt at this from the government side. Maeen, announced at COMEX 2025, is described as the country's first advanced national language model, built on local data with the explicit aim of reflecting Omani language and culture rather than generic Arabic. It launched for government use first, and the project is still early, but the announcement itself is a useful signal: officials are saying openly that generic Arabic AI does not fully serve how people here actually talk.
A handful of voice systems built specifically for phone lines take the same approach for customer calls: trained on dialect recordings, tested against real code switching, and measured on how often a caller has to repeat themselves. CustomerCare.OM's Omani Arabic customer care service is one example built this way, for businesses that already staff a phone line and want fewer of those transfers, not a shop deciding whether to answer the phone at all.
| Item | Example number | Your business |
|---|---|---|
| Calls answered per day | 300 | ___ |
| Share that trip up a generic system (dialect words, code switching, fast speech) | 1 in 5 (60 calls) | ___ |
| Extra minutes per bounced call (repeat, transfer, re-explain) | 3 minutes | ___ |
| Loaded cost of a minute of agent time | OMR 0.100 | ___ |
| Extra cost per day | OMR 18.000 | ___ |
| Extra cost per month (26 working days) | OMR 468.000 | ___ |
Half our transfers are just the system asking someone to repeat what they already said clearly.
What this means for you
You do not need to understand language models to check this. Pull five real calls from yesterday and count how many times a caller had to repeat something, switched to English out of frustration, or got bounced to a human because the system stalled. That number is your dialect gap, and it is usually bigger than owners expect.
If you are evaluating any AI system for your phone line, ask for a live demo using your own recorded calls, in the dialect and mix of languages your customers actually use, not a polished script in formal Arabic. A vendor that cannot handle that test on the spot will not handle it on a real Tuesday afternoon either.
Getting Omani Arabic right is not a nice extra for a phone line, it is the difference between a call that ends and one that escalates.
Is Modern Standard Arabic useless for customer service then?
No. It still works well for written replies, emails and formal documents. The gap is specifically in spoken, real time calls, where customers do not talk that way.
What is Maeen exactly?
Oman's own national Arabic language model project, announced in September 2025 and aimed first at government services, built to reflect local language and culture rather than generic Arabic.
Does a system need to catch every local word to be useful?
No. It needs to handle the common patterns of dialect and code switching confidently enough that most calls finish without needing a human rescue.
The bottom line
Omani Arabic on the phone is faster, more mixed and more local than the Arabic most AI ever studied, and that gap shows up as real minutes lost on your desk. It is fixable: test any system on your own real calls, ask hard questions of any vendor, and watch a national project like Maeen mature. None of that requires you to understand a single line of code.
Sources checked for this article
Practical information, not legal advice. Rules and dates were checked on 25 August 2026; verify current official positions before acting.
