英文标题:CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
作者:Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
arXiv ID:2609.07944 | 分类:cs.AI | 发表:2026-09-07
许可:CC-BY
摘要 现有的面向LLM的因果推断基准通常评估方法的文本描述,或评估生成的代码能否运行。它们很少检验所执行的工作流是否恢复了目标因果估计。CausalVerify通过将现实解读与可验证计算相分离,研究结构化计量经济学因果估计工作流中的这一验证问题。该基准基于259篇已发表的经济学论文,每篇论文由重构的研究问题、数据描述和制度背景表示,并配有100个固定种子的合成场景,这些场景为双重差分(DID)、事件研究(ES)、工具变量(IV)和断点回归设计(RDD)生成了CSV数据集。实验A(真实论文文本一致性)针对四LLM共识标签,对方法族和方向一致性进行评分。实验B(合成执行)运行模型编写的R代码,并检
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall $Ï=0.81$ and Spearman $Ï=0.93$, versus Kendall $Ï$ between $-0.20$ and $0.10$ for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。