Information Age Consulting Request a briefing

Home / Services / Practice 01

Arabic language processing

Systems that read Modern Standard Arabic accurately and Gulf dialect natively — deployable inside your own network, with no text leaving the country.

Typical clientMinistries, regulators, broadcasters, banks
Engagement6–16 weeks to production
DeploymentOn-premise or Kuwait-region cloud
LanguagesMSA, Kuwaiti, Gulf dialects

The problem

Your text is not the Arabic these tools were trained on

Commercial NLP is trained overwhelmingly on newswire Modern Standard Arabic. Real institutional text is not that. Citizen feedback, social posts, support tickets and internal correspondence are written in dialect, with inconsistent spelling, mixed script, and grammatical particles that do not exist in MSA at all.

The failures are not subtle. Below is a single Kuwaiti sentence run through a standard MSA pipeline and through ours.

TokenStandard MSA pipelineOur analyser
الخدماتnoun, plural ✓noun, plural ✓
صارتverb, past ✓verb, past ✓
وايدunknown token — droppedintensifier adverb, Kuwaiti — carries the sentiment
هالسنةunknown token — droppeddemonstrative + noun, Kuwaiti — resolves the time reference

Two dropped tokens out of six. Both of them the ones that told you how the citizen actually felt and when they meant. Multiply that across a hundred thousand comments and the report you brief your minister on is measuring something other than public opinion.

Capabilities

What we build

Component

Spelling & grammar APIمدقق إملائي

A REST endpoint that checks Arabic orthography and morphology, with a browser plug-in for staff and a bulk mode for correcting archives. Already deployed in publishing and academic environments.

Component

Morphological analysisالتحليل الصرفي

Root, pattern, part of speech and diacritisation for every token, including dialectal forms. This is the layer everything else sits on, and the reason our downstream accuracy holds up on real text.

Component

Sentiment & dialect identificationتحليل المشاعر واللهجة

Polarity and intensity scoring calibrated on Kuwaiti and wider Gulf usage, plus automatic dialect labelling so you can segment an audience by how they write, not only by what they say.

Component

Entity & topic extractionاستخراج الكيانات والموضوعات

Names of people, entities, laws and places pulled out of unstructured Arabic and normalised against your own reference lists, so the same ministry is not counted under four spellings.

Component

Retrieval for Arabic archivesالبحث الدلالي

Semantic search and question answering over your own documents, wired to an LLM of your choosing — including models that run entirely inside your data centre.

How we work

Three stages, and you can stop after the first

Stage 01

Sample assessment

You send a representative extract of your text. We run it through our pipeline and return a written assessment: what is extractable, what accuracy to expect, what would need building, and whether an off-the-shelf tool would in fact serve you better.

2 weeks · fixed fee
Stage 02

Pilot on live data

A working system on a bounded slice of your operation — one department, one campaign, one archive — with measured accuracy against a human-annotated benchmark you can audit.

4–6 weeks
Stage 03

Deployment & handover

Production install inside your environment, integration with your existing dashboards, documentation in Arabic and English, and training for the team that will own it after we leave.

6–10 weeks

Start here

Tell us what your text is doing wrong.

Send us a sample — a set of citizen comments, a document archive, a support inbox — and we will come back with a written read on what is achievable, what it would take, and whether you need us at all.