英文标题:Domain-Incremental Learning for Multi-Channel Replay Speech Detection
作者:Michael Neri
arXiv ID:2609.11194 | 分类:eess.AS | 发表:2026-09-10
许可:CC-BY
摘要 重放攻击是对语音控制系统最易实施的威胁,而暴露这些攻击的声学线索受到攻击所处环境的强烈调制。因此,部署在实际场景中的检测器必须随时间吸收新的声学条件,理想情况下无需重新访问过去的录音,因为无限期保留语音既昂贵又受到法律限制。我们将此问题建模为跨声学环境的{{PT_A_1}}({{PT_A_2}}),并提出了首个面向多通道重放语音检测的持续学习基准,在ReMASC语料库的所有{{PT_MATH_1}}种环境排序上以五个随机种子评估了一种基于波束成形器的先进检测器。顺序微调会严重遗忘,使先前学习环境的错误率上升{{PT_MATH_2}}个百分点。弹性权重巩固(EWC)将遗忘减半但丧失了可塑性
Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordings, since retaining speech indefinitely is both expensive and legally constrained. We frame this as Domain-Incremental Learning (DIL) over acoustic environments and present the first continual learning benchmark for multi-channel replay speech detection, evaluating a state-of-the-art beamformer-based detector over all 24 environment orderings of the ReMASC corpus with five seeds. Sequential fine-tuning forgets severely, raising the error rate on previously learned environments by 18.8 points. Elastic weight consolidation (EWC) halves forgetting but loses plasticity, gradient projection memory (GPM) is statistically indistinguishable from naive fine-tuning, and the proposed task-specific beamformer (TSB) that keeps one spatial front-end per environment significantly improves final and incremental accuracy. We further show that the last environment of the sequence dominates final performance. Code, results, and analysis are available at https://github.com/michaelneri/replay-speech-continual.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。