英文标题:MIRAGE: Measuring Interpolation and Redundancy in Affinity GEneralization
作者:Mehdi Yazdani-Jahromi, Sanjay Padhi, Ivan Garibay
arXiv ID:2609.14491 | 分类:q-bio.QM | 发表:2026-09-13
许可:CC-BY
摘要 深度学习模型如今在基于结构的药物设计中扮演着核心角色,从预测{{NL}}蛋白质–配体复合物与结合亲和力,到配体排序与构象生成。近期{{NL}}共折叠模型据报道在计算成本大幅降低的情况下,达到了接近自由能微扰的精度。然而,基于单一留出相关性或汇总{{NL}}构象成功率的常规评估,无法区分可迁移的结合原理与对公共结构数据库中{{NL}}相关蛋白家族的反复接触。这一区分至关重要,因为实际{{NL}}成功取决于在真正新颖靶点上的表现。 我们提出 MIRAGE,即测量亲和力泛化中的插值与冗余(Measuring Interpolation and Redundancy in{{NL}}Affini
Deep learning now underpins structure-based drug design, from complex and affinity prediction to ligand ranking and pose generation. Recent co-folding models reportedly approach free-energy-perturbation accuracy at far lower cost. Yet standard evaluation, a single held-out correlation or pooled pose-success rate, cannot separate transferable binding principles from repeated exposure to related protein families in public databases, and practical success depends on genuinely novel targets. We introduce MIRAGE (Measuring Interpolation and Redundancy in Affinity GEneralization), a plug-in benchmark treating historical public family support (through 2019) as an explicit variable, applying a family-support axis to affinity and pose prediction via matched strata, family-disjoint controls, ligand-only baselines, and temporal evaluation. Co-folder affinity accuracy rises sharply with family support, while shallow controls that cannot exploit the test family stay flat, large for co-folders and near zero for every family-disjoint or trivial control. For Nesso-1 it survives covariate, conditioning, balancing, and clustering checks; Boltz-2's endpoint is limited by coverage. It localizes to family support rather than ligand chemistry, approaching a level from family identity alone. Rankings reverse on novel families, where a family-disjoint random forest leads both co-folders, significantly vs Nesso-1. On one external low-support target, neither co-folder beats molecular weight, corroborative rather than population-level evidence. gnina shows significant support dependence in rescoring whereas smina does not; MSA-free pose engines show larger gaps than smina redocking. This redundancy-driven inflation differs from conventional leakage. We propose reporting performance across family support plus excess over a support-insensitive baseline, and release MIRAGE as an installable benchmark and dataset.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。