Brief
A 35B 'co-work' model lands at the cheap end of a cost-performance curve, its authors report
A preprint describes Occamy-1.0, a 35-billion-parameter agent model built by further training an existing checkpoint, and reports it is competitive with larger systems on the authors' own benchmarks. The claims rest on the team's stated evaluation and pricing protocol, not independent testing.
Occamy-1.0 is aimed at 'co-work' agents — systems that gather information, call tools, write code and move files across many model invocations, where cost and latency accumulate over a whole episode. The authors built execution-grounded data and environments, captured replayable long-horizon trajectories, and used staged post-training on the post-trained Qwen3.6-35B-A3B checkpoint. Across a suite of co-work benchmarks, they report it is consistently among the strongest comparably sized models and competitive with substantially larger frontier systems on several tasks. Supporting evaluations in tool calling, coding and instruction following are said to preserve broad agentic capability. The model weights and a subset of the training data are released.
Our reading
Our reading: the headline numbers come from the authors' own protocol, so the durable contribution here is the released weights and training-data subset rather than the ranking.
What to do or watch
The unresolved question is whether Occamy-1.0's placement at the low-cost knee of the cost-performance curve holds outside the authors' stated evaluation and pricing protocol, since no independent testing is reported. Readers who want to judge the claims can look to the released model weights and training-data subset, which the authors describe as the durable artifact.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- Occamy-1.0 is a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint.
- The authors release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.
- Under the authors' stated evaluation and pricing protocol, aggregate performance across four representative benchmarks places the model at the low-cost knee of the observed cost-performance Pareto frontier.
- The authors report that supporting evaluations in tool calling, coding, and instruction following show the specialization preserves broad agentic capability.
Sources
- arXivText stored 14 September 2026
How this story was checked. Written from the 1 page listed above, stored 14 September 2026; claims checked against that stored text on 14 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.