英文标题:The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
作者:Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari
arXiv ID:2609.11489 | 分类:cs.AI | 发表:2026-09-11
许可:CC-BY
摘要 合作型AI智能体是针对其他AI进行评估的,然而人类合作依赖于隐式约定——即超越字面信息去解读意义的共享协议——而AI-AI基准测试可能无法捕捉到这一点。我们提出约定差距(convention gap),即由沟通的字面内容所预测的失败概率与观测到的失败率之间的差异,作为隐式沟通的一种度量。在纸牌游戏花火(Hanabi)中,有限的牌堆和确定性的提示约束使该后验概率可以被精确计算。我们重放了来自三个公开数据集——人类-人类(hanab.live)、AI-AI(HOAD)和人类-AI(HanabiData)对局——的约101,000个出牌动作。该差距在人类配对中为+26.2个百分点(pp),在A
Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。