英文标题:What Fixed-Rollout pass@k Evaluations Can Identify
作者:Pranav Singh, Prashant Singh
arXiv ID:2609.09245 | 分类:stat.ML | 发表:2026-09-08
许可:CC-BY
摘要 重复采样评估越来越多地将 pass@{{PT_MATH_1}} 外推到远超每个问题所采集样本数 {{PT_MATH_2}} 的范围。我们证明,在合并/随机任务的条件二项模型中,固定 {{PT_MATH_3}} 的成功计数只能识别潜在逐任务成功分布的 {{PT_MATH_4}} 个自由矩。因此,直接 pass@{{PT_MATH_5}} 在 {{PT_MATH_6}} 时可被识别,但一般的 extrapolated pass@{{PT_MATH_7}}、尾部指数和尾部常数在 {{PT_MATH_8}} 时不可识别——即使在同一 rollout 预算下拥有任意多个可交换任务也是如此。这比通常
Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k <= n, but generic extrapolated pass@k, tail exponents, and tail constants are not identified for k > n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the fixed-depth count-law experiment. We give exact count-law-preserving constructions with incompatible extrapolations, state the exceptional unique-extension case, and compute sharp population identified intervals through Hausdorff principal representations. On the public 10,000-rollout-per-problem release of Brown et al., counterfactual n = 16 evaluations leave failure at k = 1000 ambiguous by factors from 1.5 to over 2,600 across four MATH/GSM8K/CodeContests configurations. The calibration shows that intermediate-scale failure share alone does not determine width. Our result does not reject parametric inference-time scaling laws; it supplies the nonparametric baseline against which their assumptions can be evaluated. We give an exact, conservative one-coordinate finite-task confidence certificate and a reporting standard separating direct estimates, identified sets, and model-conditioned forecasts.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。