英文标题:Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
作者:Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto, Taichi Nishimura, Huang Xie, Tuomas Virtanen
arXiv ID:2609.12484 | 分类:eess.AS | 发表:2026-09-11
许可:CC-BY
摘要 本文介绍了声学场景与事件检测与分类(DCASE)2026 挑战赛任务 6——长音频中的音频时刻检索(AMR)的概述。{{NL}}给定一段长达数分钟的音频录音和一个自由形式的文本查询,AMR 旨在检索录音中与查询匹配的时间时刻,其中每个时刻由一对起始和结束时间戳表示。{{NL}}该任务需要有效的跨模态对齐和长程时间建模。{{NL}}我们描述了任务定义、评估指标、开发集与评估集,以及一个将预训练的 MS-CLAP 特征提取器与基于检测 Transformer(DETR)的时刻检测网络相结合的基线系统。{{NL}}在开发数据上,基于人工标注数据集和合成数据集训练的基线取得了 13.56% 的
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 6, Audio Moment Retrieval (AMR) from Long Audio. Given a several-minute-long audio recording and a free-form text query, AMR aims to retrieve temporal moments in the recording that match the query, where each moment is represented by a pair of start and end timestamps. This task requires effective cross-modal alignment and long-range temporal modeling. We describe the task definition, the evaluation metrics, the development and evaluation datasets, and a baseline system that combines a pre-trained MS-CLAP feature extractor with a Detection Transformer (DETR)-based moment-detection network. On the development data, the baseline trained on a manually annotated dataset and a synthetic dataset achieved Recall1@0.7 of 13.56%, indicating that AMR in long audio remains a challenging problem. The challenge attracted 21 teams, which submitted 59 systems in total. The three best systems achieved Recall1@0.7 of 48.59%, roughly 3.5 times the baseline score. The results show that strengthening the audio-text feature extractor and the moment-detection network led to substantial performance improvements. Furthermore, the top three teams boosted performance by applying confidence score calibration or ensembling across different temporal resolutions of features.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。