英文标题:BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
作者:Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He, Xiao Cheng, Jicheng Wang, Wenbo Chen, Alex Yates, Yinzhe Zhao, Bingran You, Yuan Gao, Ayush Munot, Shubham Gaur, Zhe Ye, Hao Wang, Xiangyi Li, Dawn Song, Christophe Hauser
arXiv ID:2609.11028 | 分类:cs.CR | 发表:2026-09-10
许可:CC-BY
摘要。 LLM智能体基准日益作为交互式评估{{NL}}基础设施发挥作用。智能体观察状态、调用工具、修改工作区、提交{{NL}}产物,并从结果流程中获得奖励。这种交互性使{{NL}}评估易受奖励黑客攻击:智能体通过利用与奖励相关的轨迹而非解决预期{{NL}}任务来提高其测量得分。现有防御主要依赖任务特定的补丁、提示{{NL}}指令或事后检测器。它们无法提供可复用的证据,证明{{NL}}一次具体运行始终处于其预期的评估边界之内。本文{{NL}}提出BenchShield,一种面向LLM智能体评估中奖励完整性的模型支撑插桩层。BenchShield将检测建立在评估奖励相关事件的有限生命周期{{NL}
LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not provide reusable evidence that a concrete run remained within its intended evaluation boundary. This paper presents BenchShield, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation. BenchShield grounds detection in a finite lifecycle model of an evaluation's reward-relevant events. Within the benchmark infrastructure, two complementary analyses operate over this model. A static, phase-aware taint analysis exposes reward-hacking paths before a run. Its runtime counterpart uses infrastructure-side evidence to attribute concrete agent use and emit evidence-backed claims. We construct BenchShield Trajectories, a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks. Compared with an agentic hackability scanner baseline on the same tasks and model, BenchShield improves full-chain recall from 23-94% to 77-100%, same-vector coverage from 16-56% to 43-78%, and reduces per-task cost by up to 65%. Its runtime analysis achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。