日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2610.10846

潜在行動によるクロスエンボディメント・ロボット基盤ワールドモデル

Cross-Embodiment Robot Foundation World Models with Latent Actions

シェア:XThreadsFacebookLINEはてブBluesky

多様なロボットの身体と行動空間を統一した潜在行動空間で扱うワールドモデルLAC-WMを提案し、未知の身体への適応性能が最大46.7%向上することを示した。

詳しい要約

1. どんなもの?

- 多様なrobot embodimentsとaction spacesを横断して汎化するrobot world modelの構築は困難。 - 本研究はLatent Action-Conditioned Robot World Model (LAC-WM)を提案。 - 多様なembodiments間で共有される学習済みunified latent action space内で動作する。 - このunified action spaceにより、未見のrobot embodimentsへの適応時のworld model性能が向上する。 - 明示的なmotion labelsに条件づけるExplicit Action-Conditioned World Model (EAC-WM)と比較。

2. 先行研究と比べてどこがすごい?

- 従来のexplicit action conditioning (EAC-WM)では、embodiments間でaction representationsがdisjointになり、新規robotへの適応時の下流性能が制限される。 - LAC-WMはunified latent action spaceを学習することでこの問題を回避。 - 下流性能はdexterous manipulationで最大46.7%、LIBEROで11.7%向上。 - さらに、pretraining時のembodiment数増加に対し、LAC-WMの下流性能は正にスケールする。 - 一方EAC-WMはdisjoint action spaceのため、pretraining embodiments数の増加に伴い性能が低下する。

3. 技術・手法の肝は?

- 多様なembodiments間で共有されるunified latent action spaceを学習。 - そのlatent action space内で動作するLatent Action-Conditioned Robot World Model (LAC-WM)を構築。 - 比較対象のEAC-WMは明示的なmotion labelsに条件づける。 - 両モデルをdexterous manipulation tasksとmodified LIBERO benchmarkで評価。 - 具体的なネットワーク構造や学習手順の詳細は要旨からは不明。

4. どうやって有効だと検証した?

- dexterous manipulation tasksとmodified LIBERO benchmarkで両モデルを評価。 - LAC-WMはEAC-WMに対し、dexterous manipulationで最大46.7%、LIBEROで11.7%下流性能を改善。 - pretraining embodiments数の増加に対するスケーリングを検証。 - LAC-WMはembodiment数増加で下流性能が正にスケール。 - EAC-WMはembodiment数増加で性能が低下することを確認。

5. 議論はある?

- unified action spaceが効率的なcross-embodiment learningに重要であることを示唆。 - これはroboticsにおける主要な課題に対処するもの。 - explicit action conditioningはdisjoint action representationsを生み、適応を制限する。 - 具体的な限界や失敗事例、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: Explicit Action-Conditioned World Model (EAC-WM)、LIBERO benchmark。 - 関連手法としてrobot world models、cross-embodiment learning、latent action spaceに関する研究が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Tsung-Yen Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier

分類: cs.RO

原文アブストラクト

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/

関連論文

PR本紙発行元 EmplifAI