英文标题:Grounded in Sound: Reinforcement Learning with a Frozen Acoustic Judge to Curb ASR Insertion Hallucinations
作者:Tingzhen Xiong, Rilin Chen, Weiwei Li, Wentao Zhang, Qicong Xie
arXiv ID:2609.14455 | 分类:cs.SD | 发表:2026-09-13
许可:CC-BY
摘要 当强化学习(RL)被用于自动语音识别(ASR)的后训练时,奖励几乎总是存在于文本空间:它将假设与参考进行比较,而从不检查假设是否得到音频的支持。在高度规律的语音上,这为一种捷径提供了许可——依靠强大的语言先验进行猜测,而非真正聆听。一旦声学条件退化,这种捷径便不受约束地运行,输出流畅但无根基的词,即插入错误。我们提出一种声学保真度奖励:一种GRPO奖励,辅以一个单独预训练、永久冻结、非自回归的字符级wav2vec2-CTC声学评判器,该评判器严格仅在训练时使用,推理时不存在(单一模型以贪心方式解码)。在LibriSpeech上训练,并在包含真实AMI会议语音(33,282个话语-条件实例
When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never checks whether the hypothesis is supported by the audio. On highly regular speech this licenses a shortcut - guessing from a strong language prior rather than listening. Once the acoustics degrade, the shortcut runs unchecked and emits fluent but ungrounded words, i.e., insertion errors. We propose an acoustic-fidelity reward: a GRPO reward augmented with a separately pretrained, permanently frozen, non-autoregressive character-level wav2vec2-CTC acoustic judge, used strictly at training and absent at inference, where a single model decodes greedily. Trained on LibriSpeech and evaluated across a six-tier difficulty gradient including real AMI meeting speech (33,282 utterance-condition instances), the method reduces insertion errors by 28.3% on close-talking AMI-IHM and 22.3% on far-field AMI-SDM, while lowering WER on AMI-SDM from 35.89% to 34.71% and showing no detectable WER difference on the other five tiers, against a schedule-matched WER-GRPO baseline. The insertion reduction holds under a meeting-level clustered bootstrap. Four prespecified analyses support content-conditioned insertion calibration: output collapses 85-90% on unintelligible audio that preserves energy and voice activity; the gain is not recovered by the evaluated 32-best CTC rescoring configuration, yet RL internalizes it into a single greedy decoding run; and policy-only confidence yields lower insertion-AURC in all four evaluated settings. We frame this as a mechanism paper, demonstrated in one instantiation: a 7B speech LLM with a 0.3B CTC judge.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。