日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.21461

AtomEgo: 一人称視点ロボット統合による身体性基盤モデル事前学習の探求

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

シェア:XThreadsFacebookLINEはてブBluesky

約2,659時間の一人称視点人間データとロボットデータを組み合わせ、身体性基盤モデルの事前学習における効果的な統合手法を体系的に検討した研究。

詳しい要約

1. どんなもの?

- 本研究は、embodied foundation modelの事前学習において、egocentric human interaction dataをrobot dataと統合する方法を体系的に検討したものである。 - 約2,659時間のキュレーションされたコーパスとスケーラブルなデータ処理パイプラインを構築し、ego-robot co-trainingの可能性を探る。 - vision-language-actionモデルとworld-actionモデルのアーキテクチャを対象に、3つの代表的なパラダイム(ドメイン固有のaction headを持つjoint co-training、embodiment alignmentによるprogressive ego-to-robot transfer、joint video-action modeling)を調査。 - マルチタスク実ロボット実験と言語条件付きcross-embodiment表現分析を通じて評価。 - 結果として「Data Scale * Alignment Quality --> Capability Gain」という原則…

2. 先行研究と比べてどこがすごい?

- 従来のembodied foundation modelはロボットデモンストレーションの規模と多様性の制約を受けていた。 - 本研究は、大規模egocentric human interaction dataを活用する方向性を示し、その統合方法を体系的に比較検討した点が新しい。 - 人間とロボットのembodimentおよびaction-spaceのギャップが大きいため、効果的な組み込み方が不明確だったが、本研究はその課題に対する実践的ガイドラインを提供する。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- 約2,659時間のキュレーションされたコーパスとスケーラブルなデータ処理パイプラインを構築。 - 3つのパラダイムを検討: - ドメイン固有のaction headを持つjoint co-training - embodiment alignmentによるprogressive ego-to-robot transfer - joint video-action modeling - vision-language-actionモデルとworld-actionモデルのアーキテクチャに適用。 - マルチタスク実ロボット実験と言語条件付きcross-embodiment表現分析で評価。

4. どうやって有効だと検証した?

- マルチタスク実ロボット実験を実施。 - 言語条件付きcross-embodiment表現分析を実施。 - これらの評価を通じて、egocentric dataの汎化への影響とアラインメントの重要性を検証。 - 具体的なタスクや指標の詳細は要旨からは不明。

5. 議論はある?

- 結果から「Data Scale * Alignment Quality --> Capability Gain」という原則を提示。 - egocentric dataは汎化を改善し得るが、その価値はアラインメントと活用の効果に依存する。 - この原則はスケーラブルなego-robot pre-trainingの実践的ガイドラインとなり得る。 - 限界や今後の課題についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、vision-language-actionモデル、world-actionモデル、embodiment alignment、cross-embodiment representation learningなどが挙げられる。 - 同分野の定番として、RT-1、RT-2、PaLM-E、OpenVLAなどのembodied foundation modelに関する論文が次に読むべき候補となる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie, Xiaoquan Sun, Junyang Zheng, Zhuoyang Song, Jiaxing Zhang, Jiayu Chen

分類: cs.RO, cs.AI

原文アブストラクト

Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision--language--action and world--action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video--action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality --> Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego--robot pre-training.

関連論文

PR本紙発行元 EmplifAI