Brief
GraphEcho benchmark: agents that stop repeating paths can end up reaching fewer sources
A preprint describes a benchmark that varies how many paths lead to the same evidence while holding the evidence itself fixed, and reports that less repetition does not always mean better evidence use. Only the abstract is available to us, so no effect sizes, model names or agent counts can be checked.
GraphEcho holds evidence content fixed while varying how many paths lead to it and where that evidence came from, then scores both the agent's judgments and its exploration. Across the frozen agents evaluated, adding redundant supporting paths raised the share of repeated walks.
Provenance-aware post-training reduced revisits and improved accuracy on the synthetic tasks, but reached fewer distinct sources. On scientific claims, repetition fell again while accuracy declined. The authors read this as a gap between efficient exploration and effective evidence use. The results come from controlled synthetic experiments, and the reported test on scientific claims shows accuracy falling, so the useful detail is a trade-off, not a solved problem.
Our reading
This desk covers preprints alongside peer-reviewed results, but only when the measured finding can be separated from the authors' reading of it — here it can, narrowly. The measured part is the trade-off between fewer revisits and fewer distinct sources; the claim that this exposes a gap between efficient exploration and effective evidence use is the authors' own. Anyone building or buying graph-…
What to do or watch
Watch whether this preprint clears peer review and whether the same trade-off shows up outside synthetic benchmarks. The unresolved question is whether an agent that looks efficient in exploration is actually reaching distinct evidence in real research use.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- A large language model agent can follow more graph paths without acquiring more independent evidence.
- GraphEcho tests whether agents mistake repeated encounters with the same paths for additional corroboration.
- The benchmark varies path counts and evidential origins while holding evidence content fixed, and evaluates both judgments and active exploration.
- Redundant supporting paths increase the share of repeated walks across all evaluated frozen agents.
- Provenance-aware post-training reduces revisits and improves synthetic accuracy, yet covers fewer distinct sources.
- On scientific claims, provenance-aware post-training continues to reduce repetition while accuracy declines.
Sources
- arXivText stored 17 September 2026
How this story was checked. Written from the 1 page listed above, stored 17 September 2026; claims checked against that stored text on 17 September 2026.
What that means
- 6 of 6 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.