Programme B
Arabic Language and AI
Arabic is the language of the sources and of a living, diverse community of speakers. It deserves AI systems that handle it with care and quote it faithfully. This programme develops open ways to check whether they do.

In progress The Citation Fidelity Standard: work has begun, and the citation check on a defined pilot corpus is being documented.
Planned Open test data for classical Arabic, Modern Standard Arabic and dialects, developed with linguists and philologists.
Aim
Language models are increasingly used in Arabic: to answer questions, to summarise and translate, and to explain classical texts. How well they handle the language matters, for the precision of classical Arabic, for Modern Standard Arabic and for the dialects people speak. When they quote, they sometimes attribute to a book or an author words that do not appear there. Research on Arabic language models names this misattribution as a central problem, especially where religious sources are quoted.
The aim of this programme is to make the quality of Arabic in AI systems checkable, with open methods and test data that any researcher or institution can use. Its first building block is an open standard for citation fidelity, because trust in an answer begins with a faithful quotation.
Fields of work
Citation fidelity
An open test set and method to check that an AI system attributes to a source only what the source actually says.
Classical and contemporary Arabic
Test data covering classical Arabic, Modern Standard Arabic and major dialects, developed together with linguists and philologists.
Arabic first
Methods and documentation written in Arabic first, so that the community of the language can examine them without translation.
The Citation Fidelity Standard
The Standard is built from three components.
Open test set and method
Any operator of a language model can measure its system against the same test set. The result is a test report for that system, published with its operator's consent. No ranking of countries is produced.
Reference implementation for sovereign operation
The checking routine is intended to run on an institution's own infrastructure, with the model that the institution chooses, so that its data remain under its control.
Provenance model with link back
For each accepted quotation, the source, edition, location and method of verification are recorded. The quotation leads back to the source and to the institution that holds it, which remains the owner and is named.
How a quotation is checked
The rule is to reject rather than guess. A quotation is shown only if it can be found word for word in a named edition. If it cannot, the answer is withheld and the reason is recorded.
A negative result has a narrow meaning. It says only that the passage was “not found in the stated edition at the stated location”. It passes no judgement on whether a text is authentic, and it never leads to a text being described as “fabricated” or “inauthentic”. A person checks the wording before it is published.
Qur’anic text is shown only word for word from a named standard edition of the Qur’an. Any deviation from that edition leads to rejection.
Building on existing work
The Standard is meant to complement existing work, not to replace it, including the open corpora published by OpenITI.
Others have taken first steps in the same direction, among them the IslamicEval 2025 shared task and Fanar-Sadiq, which validates quotations and shows the steps of its verification. Benchmarks for Arabic text recognition, such as KITAB-Bench, address the reading of documents rather than citation fidelity in AI answers.
For whom
- Teams developing Arabic language models who want to test how faithfully their systems quote.
- Researchers in Arabic natural language processing and the digital humanities.
- Operators of platforms that answer questions about classical texts.
- Linguists and philologists working on Arabic and its dialects.
First verifiable output
A documented citation check on a defined pilot corpus: examples of accepted and rejected answers with the reason for each, and negative tests showing that the check fails when it should.
Publication date: expected in the first quarter of 2027.
How progress is measured
- Rejection rate: the share of answers withheld because a quotation could not be matched.
- Verbatim citation fidelity, measured on a published sample.
- The number of external systems measured with the consent of their operators.
Values will be published only together with a method report.
What this programme does not do
- We pass no verdicts on hadith. We show where a text is found, not what standing it has.
- We issue no fatwas and no religious rulings, make no theological assessments and do not rank schools of law.
- When a system that also answers questions of fiqh is measured, we check citation fidelity only, never the correctness of a ruling.
- We do not measure any system without the consent of its operator.
- We publish no rankings of countries or of dialects.
- We do not replace philological scholarship. People with that expertise review every publication before it appears.