英文标题:Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
作者:Jack Hopkins, Dipika Khullar, Fabien Roger
arXiv ID:2607.08173 | 分类:cs.AI | 发表:2026-09-07
许可:CC-BY
摘要 语言模型的黑盒审计是部署前的重要工具,但它可能遗漏微妙的失调形式和隐藏信息。为了在审计过程中更好地引出隐藏信息,我们引入了过度思考:利用推理任务向量放大推理模型“边思考边输出”倾向的过程。给定非推理指令模型 {{PT_MATH_1}} 和推理蒸馏模型 {{PT_MATH_2}} 的参数,我们将过度思考模型定义为 {{PT_MATH_3}},其中 {{PT_MATH_4}} 将推理放大到纯推理模型 {{PT_MATH_5}} 之上。{{NL}}此外,我们引入了新的逐层衰减策略,该策略选择性地放大推理,同时不损失模型输出的质量和连贯性。我们证明,在四种实验设置中,跨 2B-32B 模型,过度
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of a non-reasoning instruct model $M$ and reasoning-distilled model $R$, we define the \emph{overthinking model} as $\boldsymbolθ_{\mathcal{O}_α} = \boldsymbolθ_{\mathcal{M}} + α(\boldsymbolθ_{\mathcal{R}} - \boldsymbolθ_{\mathcal{M}})$, where $α> 1$ amplifies reasoning beyond the pure reasoning model $R$. Additionally, we introduce new layer-wise attenuation strategies that selectively amplify reasoning without losing quality and coherence of model outputs. We demonstrate that overthinking models are more likely to reveal hidden information across four experimental settings, across 2B-32B models. Our findings suggest that reasoning amplification may surface secrets or unintended behaviors acquired during training up to $10\times$ more frequently than the original reasoning model. How secrets surface depends on the secret type: some require perturbation along the reasoning direction, while others yield to any sufficiently large weight perturbation.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。