BriefPulse Science · Research reporting with methods and limits kept visible. RSS · BriefPulse network
BriefPulse Science

Findings, methods and limits explained with the evidence in view.

16 September 2026

Brief

Iraqi Arabic benchmark finds MSA tests saturate while dialect track separates models

A benchmark called Mizan evaluates large language models on Iraqi Arabic and the Iraqi civic context. In a pilot of 27 systems, the Modern Standard Arabic track saturated while the Iraqi track produced a consistent 14-18-point gap between models.

The benchmark pairs a Modern Standard Arabic (MSA) baseline track with an Iraqi track across six axes: dialect comprehension, dialect generation, bidirectional MSA-Iraqi translation, Iraq-specific knowledge, official-document field extraction, and safety. It was built from 340 originally authored, dually reviewed items, with statistically audited answer positions and Wilson intervals on every published score. A pilot evaluation of 27 systems found the MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders. Official-document extraction confined every system to 32-56. Two dedicated Arabic models scored below a size-matched generalist on the Iraqi track. The safety-hardened tier of the newest model family deterministically refused innocuous dialect-comprehension items as policy violations.

Our reading

Our reading is that the benchmark's value lies in exposing dialect-specific gaps and over-refusal that MSA leaderboards may miss, though the pilot's 27-system scope means the results are preliminary.

Source details and supporting facts

Each line is stated by the page named above it.

Stated by arXiv

  • It has an MSA baseline track paired with an Iraqi track across six axes.
  • It was built from 340 originally authored, dually reviewed items with statistically audited answer positions and Wilson intervals on every published score.
  • A pilot evaluation of 27 systems found the MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders.
  • Official-document extraction confines every system to 32-56.
  • Two dedicated Arabic models score below a size-matched generalist on the Iraqi track.
  • The safety-hardened tier of the newest model family deterministically refuses innocuous dialect-comprehension items as policy violations.

Sources

  1. arXivText stored 16 September 2026

How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.

What that means
  • 6 of 7 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
  • Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
  • The check reads stored text only: no claim rests on a fresh look that did not happen.
  • Where the reporting was silent, the text says so instead of filling the gap.

More from Science