英文标题:New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models
作者:Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
arXiv ID:2609.11022 | 分类:cs.CV | 发表:2026-09-10
许可:CC0
摘要 考虑一个简单的力学任务,涉及滑动的物块、弹跳的物体或连接在弹簧上的质量。{{NL}}模型首先看到来自一次测量实验的图像;例如,物块滑行了多远,并且必须回答关于新试验的问题,例如物块在固定推力后是否会越过目标。{{NL}}第一次实验可能提供足够的信息来回答,或者模型可能需要另一次测量,例如物体的质量、摩擦、恢复系数或弹簧刚度。{{NL}}我们研究视觉语言模型能否决定何时立即回答,以及当需要更多证据时,应进行哪项实验。{{NL}}当前的物理推理基准通常只对最终答案评分,因此它们并不直接评估这一决策。{{NL}}我们引入一种受控评估,其中每个问题呈现一张测量图像,以及由两个可能的质量值和另一个
A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measurement image and four possible physical worlds created by combining two possible masses and two possible values of another relevant property. The model must either stop and answer or select the cheapest additional experiment that can resolve the question. We construct matched problem pairs where changing either the observed measurement or the question changes the optimal action. Since all possible worlds and experiment costs are known, we can explicitly determine the optimal choice. Across six open models and 144 physical parameter sets, direct responses repeat the same action for 95.1% to 100% of image pairs even when the correct action changes. Brief reasoning improves action switching, but the best model makes both decisions correctly for only 5.9% of image pairs. Additional analysis reveals failures in measurement interpretation, physical reasoning, and response formatting. By evaluating evidence selection separately from final answers, our benchmark reveals limitations in physical reasoning that conventional answer accuracy can overlook.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。