英文标题:ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making
作者:Jun Xiang, Zhijie Bao, Rong Hu, Kaizhou Qin, Wei Chen, Zhongyu Wei
arXiv ID:2609.07601 | 分类:cs.CL | 发表:2026-09-07
许可:CC-BY
摘要 大语言模型(LLM)在个性化医疗助手中的应用日益受到关注。然而,现有的医学基准主要依赖带有预选证据的静态问答,这使得LLM能否从真实的纵向电子健康记录(EHR)中做出可靠的临床决策仍不明确。为弥补这一空白,我们提出了ObGynLongBench,一个基于规则的、面向妇产科决策的长上下文EHR基准,包含来自976份真实妊娠EHR病史的1,500个临床决策点案例以及可追溯的规则。每个案例都锚定于一位患者、一个妊娠时间线节点以及一个决策前信息边界,从而支持仅证据(Evidence-only)、就诊级EHR(Visit-level EHR)和病史级EHR(History-level EHR)三种
The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic health records (EHRs). To bridge this gap, we introduce ObGynLongBench, a rule-grounded long-context EHR benchmark for obstetric and gynecologic decision-making, comprising 1,500 clinical decision-point cases from 976 real pregnancy EHR histories and traceable rules. Each case is anchored to a patient, a pregnancy-timeline point, and a pre-decision information boundary, enabling Evidence-only, Visit-level EHR, and History-level EHR evaluation. Evaluating 17 LLMs reveals a substantial Evidence-to-EHR Gap: models perform well when evidence is directly provided, but accuracy drops when evidence must be extracted from same-day records or full pre-decision EHR histories. Further analyses identify evidence utilization as a key bottleneck: performance decreases with longer EHR contexts and more complex evidence requirements, and earlier failures often predict later failures within the same patient history. Finally, active-search agents perform best among EHR access strategies, highlighting patient-specific evidence utilization as a central challenge for reliable personalized medical assistants. Resources are available at https://github.com/xiangjun2003/ObgynLongbench.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。