英文标题:CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
作者:Kechen Liu, Ola Shorinwa
arXiv ID:2608.27406 | 分类:cs.RO | 发表:2026-09-11
许可:CC-BY
摘要 当前最先进的以动作为条件的视频模型通常局限于单一机器人具身,使其无法利用包含丰富信号、可用于学习可泛化物理规律的庞杂异构视频数据语料。为弥合这一差距,我们提出了 CLAP,一个跨具身的以动作为条件的视频生成框架,能够在涵盖人类与机器人智能体的多样化、互联网规模视频上进行训练。CLAP 基于这样一个洞见:无论行为主体为何,普适的物理定律支配着时空动力学。然而,跨具身学习并非易事,因为动作表征在不同机器人平台之间差异极大,且在人类视频中通常缺失。CLAP 通过三项核心贡献应对这一根本性挑战。{{NL}}第一,CLAP 利用末端执行器位姿、自然语言指令以及学习到的潜在动作表征,调和了人类与机器
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。