Brief
Iraqi Arabic benchmark finds MSA tests saturate while dialect track separates models
A benchmark called Mizan evaluates large language models on Iraqi Arabic and the Iraqi civic context. In a pilot of 27 systems, the Modern Standard Arabic track saturated while the Iraqi track produced a consistent 14-18-point gap between models.
The benchmark pairs a Modern Standard Arabic (MSA) baseline track with an Iraqi track across six axes: dialect comprehension, dialect generation, bidirectional MSA-Iraqi translation, Iraq-specific knowledge, official-document field extraction, and safety. It was built from 340 originally authored, dually reviewed items, with statistically audited answer positions and Wilson intervals on every published score. A pilot evaluation of 27 systems found the MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders. Official-document extraction confined every system to 32-56. Two dedicated Arabic models scored below a size-matched generalist on the Iraqi track. The safety-hardened tier of the newest model family deterministically refused innocuous dialect-comprehension items as policy violations.
Our reading
Our reading is that the benchmark's value lies in exposing dialect-specific gaps and over-refusal that MSA leaderboards may miss, though the pilot's 27-system scope means the results are preliminary.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- It has an MSA baseline track paired with an Iraqi track across six axes.
- It was built from 340 originally authored, dually reviewed items with statistically audited answer positions and Wilson intervals on every published score.
- A pilot evaluation of 27 systems found the MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders.
- Official-document extraction confines every system to 32-56.
- Two dedicated Arabic models score below a size-matched generalist on the Iraqi track.
- The safety-hardened tier of the newest model family deterministically refuses innocuous dialect-comprehension items as policy violations.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 6 of 7 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.