英文标题:EMBLEM: Enhancing Multi-script Table Detection through Masking
作者:Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan
arXiv ID:2609.08330 | 分类:cs.LG | 发表:2026-09-08
许可:CC-BY
摘要 表格检测是文档分析中的一项核心任务,支撑着信息检索、文档重建和视觉问答等下游应用。现有深度学习模型在英文和中文文档上表现出色,但由于文字多样性和标注数据有限,在多语言、多文字文档上表现不佳。为解决这一问题,我们引入了 MANDALA(用于表格检测的多文字标注文档),这是一个经人工整理的数据集,包含 {{PT_MATH_1}} 个含表格页面,涵盖 18 种语言和 15 种文字,覆盖多个领域。我们还提出了 EMBLEM,一种用于多文字表格检测(MTD)的基于掩码的范式。EMBLEM 生成掩码图像,隐藏文字和字体特有的细节,使在大量英文文档上预训练的模型能够专注于与文字无关的页面布局。在三种表
Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and visual question answering. While existing deep learning models perform well on English and Chinese documents, they struggle with multilingual, multi-script documents due to script diversity and the limited availability of labeled data. To address this challenge, we introduce MANDALA (Multi-script Annotated Documents for Table Detection), a manually curated dataset of 2,323 table-containing pages spanning 18 languages and 15 scripts across diverse domains. We also propose EMBLEM, a masking-based paradigm for Multi-script Table Detection (MTD). EMBLEM generates masked images that conceal script- and font-specific details, enabling models pre-trained on abundant English documents to focus on script-agnostic page layout. Experiments across three table detection architectures show that EMBLEM consistently outperforms strong baselines on MANDALA while remaining competitive on five standard English-dominant benchmarks. Using only English masked images for fine-tuning, with no multi-script training data, EMBLEM achieves an absolute F1-score gain of 20.8% on MANDALA. We release MANDALA along with the accompanying code and models at https://github.com/IITB-LEAP-OCR/EMBLEM.git.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。