英文标题:ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation
作者:Zesheng Wei, Mengfan Li, Wenhao Liu, Yixin Zhang, Zilei Wang, Yang Deng
arXiv ID:2609.11101 | 分类:cs.CL | 发表:2026-09-10
许可:CC-BY
摘要 纠纷调解对于维护社会和谐与韧性至关重要,然而培养技艺娴熟的调解员既昂贵又耗时。现有的基于大语言模型的调解研究仍受限于不切实际的任务设定、低保真数据集以及粗粒度的评估指标,这些指标掩盖了逐轮对话的动态变化。为弥补这些不足,我们提出了 ProMediConv,一个新颖的基准测试框架,将调解建模为一个主动式、多阶段且感知当事人的对话过程,融合了 11 种调解策略和四种当事人行为模式(BP)状态。基于 972 个完整的真实案例,我们构建了一个高保真调解数据集,包含话语级别的策略与 BP 状态标注。此外,为更好地评估智能体的影响,我们提出了 MAD(平均属性差异),一种能够捕捉对话全程 BP 变化
Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelity datasets, and coarse evaluation metrics that obscure turn-by-turn dynamics. To address these gaps, we introduce ProMediConv, a novel benchmarking framework that models mediation as a proactive, multi-stage, and party-aware dialogue process incorporating 11 mediation strategies and four party behavior pattern (BP) states. Using 972 complete real-world cases, we construct a high-fidelity mediation dataset with utterance-level annotations of strategies and BP states. Furthermore, to better assess agent impact, we propose MAD (Mean Attribute Difference), a fine-grained metric that captures BP shifts throughout the dialogue. Leveraging this framework, we establish a comprehensive benchmark by evaluating diverse models alongside our tailored baseline ProMediAgent. Extensive empirical analyses reveal critical behavioral phenomena and underscore the persistent challenges current models face in dynamic, multi-party mediation. Ultimately, ProMediConv provides a rigorous foundation and a vital quantitative standard for advancing AI-assisted conflict resolution. Our dataset and codebase are accessible at https://github.com/ZsWei66/ProMediConv_repo.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。