英文标题:SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
作者:Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
arXiv ID:2609.09672 | 分类:cs.CL | 发表:2026-09-09
许可:CC-BY
摘要 音频与多模态大语言模型的快速发展解锁了具有变革性的语音理解能力,然而评估框架仍以英语为中心,导致东南亚(SEA)语言严重缺乏代表性。我们提出 SEA-SpeechBench,据我们所知,这是首个大规模多任务基准,通过 99 个评估集、97,194 个样本以及 597 小时精选音频数据,评估 11 种东南亚语言的语音理解能力。我们的基准包含 3 大类共 9 项多样化任务:语音处理(自动语音识别、语音翻译、口语问答)、副语言分析(情感、性别、年龄、说话人识别),以及时间理解——一个新颖维度,涵盖带时间戳的内容查询与长达 3 分钟的扩展音频序列内的时间定位。我们实现了以东南亚母语和英语进行的多
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We introduce SEA-SpeechBench, to the best of our knowledge, the first large-scale multitask benchmark that evaluates speech understanding in 11 SEA languages through 97,194 samples across 99 evaluation sets and 597 hours of curated audio data. Our benchmark comprises 9 diverse tasks across 3 categories: speech processing (automatic speech recognition, speech translation, spoken question answering), paralinguistic analysis (emotion, gender, age, speaker recognition), and temporal understanding, a novel dimension featuring timestamped content queries and temporal localization within extended audio sequences up to 3 minutes. We implement multilingual prompting in both native SEA languages and English to reflect user interactions with audio-language models. Evaluation of leading open-source and proprietary systems reveals marked performance gaps. Across all models, performance remains underwhelming on temporal understanding, emotion recognition, and speech translation. Prompting in low-resource languages such as Burmese and Tamil lags behind English by up to 41 percentage points. Our findings expose critical model limitations and underscore the need for inclusive model development. The SEA-SpeechBench benchmark is available at https://zwenyu.github.io/SEA-SpeechBench/.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。