日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
操作学習/クロスエンボディ転移arXiv:2609.05892

A4A: 人間のデモからの行動指向4Dアフォーダンスのクロスエンボディ転移

A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations

シェア:XThreadsFacebookLINEはてブBluesky

人間のデモから行動指向の4Dアフォーダンス(相互作用点の将来軌跡)を学習し、ロボットポリシーの事前学習に利用することで、多様なVLAポリシーの操作性能を向上させる手法を提案した。

著者: Yifan Han, Litao Liu, Yuqi Gu, Ye Lu, Hanqing Wang, Sidney Wai, Ishaan Myrie, Qi Zhang, Jingjin Yu, Gen Li

分類: cs.RO

原文アブストラクト

Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be transferred effectively to robot control. Existing affordance representations are typically formulated as 2D masks, 3D regions, contact points, or actionability scores, and therefore primarily identify where interaction may occur. However, effective manipulation also requires modeling how interaction-relevant geometry evolves during task execution. To bridge this gap, we introduce action-oriented 4D affordances, which represent the language-conditioned future trajectories of interaction-relevant 3D points. These trajectories capture task-conditioned geometric evolution rather than embodiment-specific actions, enabling transferable interaction priors across humans and robots. Based on this representation, we construct a large-scale action-oriented 4D affordance dataset from existing human--object interaction video data and complementary RGB-D demonstrations, and introduce A4A, an affordance-to-action framework that uses 4D affordance trajectory prediction to pretrain robot policies before manipulation fine-tuning. Experiments in both simulation and the real world validate the effectiveness of A4A, showing that pretraining with action-oriented 4D affordance data consistently improves the manipulation performance of diverse VLA policies. These results establish action-oriented 4D affordances as an effective cross-embodiment representation for transferring manipulation knowledge from human demonstrations to robot control.