Brief
AI judges improve AI-written patent drafts, but match a human attorney only partly
A new arXiv preprint describes a patent-drafting testbed in which a separate AI judge scores drafts and feeds back revisions. The reported gains are measured largely by that same judge, and its agreement with a professional patent attorney was meaningful but depended heavily on which metric was used.
The setup: an AI agent drafts a patent, a separately invoked AI judge scores the draft and returns structured feedback, and the agent revises. Across several inventions and drafting-agent configurations, this loop consistently improved the judge's own quality scores, while revision without judge feedback tended to plateau. The authors also report that iterative judge feedback let a lower-reasoning agent approach the performance of a substantially more expensive high-reasoning one.
The check: the judge was compared with an independent evaluation by a professional patent attorney. Agreement was meaningful but strongly metric-dependent, with systematic calibration differences. That is the boundary to hold on to — the improvement is largely measured by the judge being tested, not by a legal standard.
Our reading
Our reading: the gains are real inside the paper's own scoring loop, but the attorney comparison suggests the judge is not yet a substitute for professional judgment.
What to do or watch
The precise unresolved question is whether judge-guided drafting gains persist when drafts are scored by a professional patent attorney rather than by the same judge providing the feedback, since the attorney comparison showed meaningful but strongly metric-dependent agreement and systematic calibration differences. Watch for follow-up work that reports attorney-scored outcomes of the judge-revision loop, rather than the judge's own quality scores.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision.
- Judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate.
- Iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent.
- The judge was validated against independent evaluation by a professional patent attorney, finding meaningful but strongly metric-dependent agreement and systematic calibration differences.
Sources
- arXivText stored 15 September 2026
How this story was checked. Written from the 1 page listed above, stored 15 September 2026; claims checked against that stored text on 15 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.