Brief
Vendor-Native Coding Harnesses Show No Clear Average Advantage in Paired Test
A new arXiv paper compares agentic coding harnesses paired with the same models on a private, contamination-controlled suite. It reports no resolved average advantage for either harness, with opposite results across task types and unresolved cost ordering.
The study ran 80 tasks under claude-agent-sdk and deepagents on claude-opus-4-8, and under openai-codex SDK and deepagents on gpt-5.5. 792 of 800 planned runs were graded by an isolated oracle. For Opus 4.8, the average difference was -1.25 percentage points (48.8% vs 50.0%, 95% CI [-10.0, +7.5]); for GPT-5.5, +1.25 pp (55.6% vs 54.4%, CI [-4.4, +6.9]). The Opus average hid opposite strata: native trailed by 9.0 pp on 61 repository tasks and led by 23.7 pp on 19 contest tasks, a partition chosen after seeing the data. Cost per solved task favored the neutral harness in observed usage, but missing usage records leave the billed ordering unresolved.
Our reading
Our reading is that the result undercuts a simple vendor-native advantage story, but the task-type split and cost uncertainty mean it is not a settled verdict.
What to do or watch
Watch for a designed replication that tests the repository-versus-contest task split before treating the harness comparison as settled, since that partition was chosen after seeing the data. The precise unresolved question is the billed cost ordering per solved task, which hinges on how the 58 runs on the Anthropic account that left no usage record are allocated.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- The same 80 tasks ran under claude-agent-sdk and under deepagents on claude-opus-4-8, and under the openai-codex SDK and deepagents on gpt-5.5.
- 792 of 800 planned runs were graded by an isolated oracle.
- Neither contrast resolves an average advantage for either harness: -1.25 pp for Opus 4.8 (48.8% vs 50.0%, task-bootstrap 95% CI [-10.0, +7.5]) and +1.25 pp for GPT-5.5 (55.6% vs 54.4%, CI [-4.4, +6.9]).
- The Opus average combines opposite strata: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks (label-permutation p = 0.003).
- The partition was chosen after seeing the data and needs a designed replication.
- Re-priced from raw per-turn usage at frozen list prices, the neutral harness cost 1.3 to 1.6 times as much per solved task on Opus 4.8 and 1.2 times on GPT-5.5.
- On the Anthropic account 58 runs left no usage record, and allocating that spend to either cell would move the Opus ratio between 0.7 and 2.3, so the billed ordering is unresolved.
- This revision corrects an August 2026 manuscript whose cost figures rested on a usage-semantics defect in our own telemetry (Section 5.1).
Sources
- arXivText stored 14 September 2026
How this story was checked. Written from the 1 page listed above, stored 14 September 2026; claims checked against that stored text on 14 September 2026.
What that means
- 8 of 8 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.