英文标题:Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
作者:Jiaqiang Li, Yajie Yang, Zhiheng Xi, Jiadong Chen, Enyu Zhou, Senjie Jin, Yang Nan, Jiazheng Zhang, Han Wang, Yanxin Li, Dingwei Zhu, Bicheng Deng, Yuhui Wang, Xiang Zheng, Qi Zhang, Lei Bai, Xingjun Ma, Tao Gui
arXiv ID:2609.11243 | 分类:cs.AI | 发表:2026-09-10
许可:CC-BY
摘要 自主研究智能体日益被期望能够检索文献、分析实验证据并生成科学假设。这些能力要求进行多步证据支撑的推理,即在得出结论之前逐步获取、整合并验证证据。然而,现有的多模态基准主要评估最终答案的准确性,未能回答预测是否真正由可追溯的科学证据所支撑。我们提出了 Sci-MMR,一个基于结构化论证图构建的多步证据支撑科学推理基准,该论证图将科学主张、基于引用的知识、视觉证据以及支撑区域联系起来。Sci-MMR 包含 235 个多跳推理任务,涵盖四个科学学科,每个任务平均包含九个图面板。通过评估八个前沿多模态模型,我们发现答案准确率始终比完整证据恢复率高出 20% 以上,揭示了一个仅凭答案评估在结构上无
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。