英文标题:What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
作者:Marek Hradil, Danae Sánchez Villegas
arXiv ID:2608.23474 | 分类:cs.CL | 发表:2026-09-12
许可:CC-BY
摘要 视觉语言模型(VLM)在视频和图像序列基准上取得了强劲的性能,但它们是否捕捉到了时间结构仍不清楚。为研究这一问题,我们将时间定位形式化为一个异常检测问题,提供了一种简单且可控的评估方法,直接测试对时间一致性的敏感性。我们提出了 TimeCatch,其中时间异常通过交换相邻帧来构造,帧级异常则通过用高斯噪声替换某一帧来构造。模型在四个合成和真实世界数据集上接受异常检测与定位任务的评估,并辅以一项人类研究。{{NL}}我们的评估揭示了帧级异常检测与时间异常检测之间存在显著差距。虽然 VLM 能够一致地检测帧级异常并常常准确定位它们,但在我们的主要评估设置下,它们在时间异常检测上通常表现接近随
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, under our main evaluation setting they generally perform near chance on temporal anomaly detection and show limited localization performance. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity show that performance can improve under some conditions, while substantial gaps in temporal anomaly detection and localization remain. Together, these findings reveal a gap between frame-level and temporal anomaly detection. TimeCatch provides a controlled benchmark for evaluating temporal consistency in vision-language models.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。