日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08119

AutodidactWAM: 生成動画からロボット動作へのクロスモーダル自己蒸留

AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions

シェア:XThreadsFacebookLINEはてブBluesky

世界行動モデルが生成した動画から手姿勢推定と逆運動学で動作を復元し、自己蒸留でロボットの動作精度を大幅に改善した研究。

詳しい要約

1. どんなもの?

World-action models (WAMs) の一種である Cosmos 3 を、未見のロボットである Unitree G1 humanoid (BrainCo 製五指ハンド) に LoRA fine-tune で適応させる研究。video-action asymmetry (生成 video は妥当だが co-generated action が systematic に mis-targeted) を明らかにし、teleoperation なしで hand-pose estimator + inverse kinematics により生成 video から action を復元する AutodidactWAM を提案。復元 action を preferred target として action 関連層のみを fine-tune する self-distillation 手法。

2. 先行研究と比べてどこがすごい?

従来の WAM 適応は teleoperation による task-specific データを要するが、本手法は one-time embodiment adaptation 後は追加の teleoperation 不要。native action の成功率 (pre-grasp 17%, grasp 10%, pick-and-place 7%) に対し、復元 action gate は約 75%, 47%, 42% と大幅改善。held-out object でも native より良い。

3. 技術・手法の肝は?

生成 video 上で動作する hand-pose estimator (teleoperation なしで学習) と inverse kinematics により action estimate を復元。native prediction と pair にして preferred target とし、action 関連層のみを fine-tune (video は teacher-forced)。supervised relabeling と Diffusion-DPO の rectified-flow 適応を比較。最終的に preference supervision, supervised target fitting, Cartesian trajectory anchoring を組み合わせた DPO+SFT+DTW が最良。

4. どうやって有効だと検証した?

closed-loop real-robot trials を pre-grasp, grasp, pick-and-place の3段階で評価。native action は約 17%, 10%, 7% 成功。復元 action gate は約 75%, 47%, 42%。DPO+SFT+DTW は training object (Oreo) で pre-grasp 90%, full-task 20%、held-out object で 80%, 30%。Plain Flow-DPO は validation preference accuracy 1.000 ながら成功率 0%。

5. 議論はある?

Plain Flow-DPO が preference accuracy 1.000 でも成功率 0% であることから、contrastive objective 単独ではなく training objectives の組み合わせが性能向上を駆動すると考察。video-action asymmetry の存在と、生成 video を action 復元に活用する有効性を示唆。

6. 次に読むべき論文は?

Cosmos 3 (WAM の基盤モデル)、Diffusion-DPO、rectified-flow、LoRA fine-tune、inverse kinematics 関連の研究。特に Diffusion-DPO の rectified-flow 適応や、WAM の video-action asymmetry を扱う研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sergei Kurchev, Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov, Dzmitry Tsetserukou

分類: cs.RO

原文アブストラクト

World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction. Adapting one such model with a lightweight LoRA fine-tune to a previously unseen robot, a Unitree G1 humanoid with five-fingered BrainCo hands, exposes a video-action asymmetry: the video renders plausible task executions, while the co-generated action is systematically mis-targeted. We evaluate closed-loop real-robot trials at three cumulative stages: pre-grasp, grasp, and pick-and-place. The native action succeeds only approximately 17%, 10%, and 7% of the time, respectively, and performs worse on held-out objects. We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates. Paired with the native prediction, these recovered actions provide preferred targets for fine-tuning only the action-related layers, while the generated video is teacher-forced. We compare supervised relabeling with a rectified-flow adaptation of Diffusion-DPO. After one-time embodiment adaptation, self-distillation requires no additional task-specific teleoperation. The recovered-action gate reaches approximately 75%, 47%, and 42% pre-grasp, grasp, and pick-and-place success, compared with 17%, 10%, and 7% for the native action. A hybrid objective combining preference supervision, supervised target fitting, and Cartesian trajectory anchoring (DPO+SFT+DTW) performs best: on Oreo, the training object, it reaches 90% pre-grasp and 20% full-task success; on a held-out object, it reaches 80% and 30%. Plain Flow-DPO reaches 0% success despite 1.000 validation preference accuracy, indicating that the combination of training objectives, rather than the contrastive objective alone, drives the observed gains.

関連論文

PR本紙発行元 EmplifAI