日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.25627

MachEmbodied-U0:身体知能のための統合理解・生成モデル

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

シェア:XThreadsFacebookLINEはてブBluesky

理解と生成の専門家をMixture-of-Transformersで統合し、視覚力学と行動生成をフローマッチングで結びつけた身体知能基盤モデルを提案。LIBEROで99.0%の成功率を達成。

詳しい要約

1. どんなもの?

- 汎用ロボット制御のための統合 embodied foundation model - 名称: MachEmbodied-U0 (ME-U0) - 理解と生成の expert を Mixture-of-Transformers で接続 - 対象タスク - 下流のロボット操作制御 - subtask prediction, affordance grounding, visual dynamics の zero-shot 実行 - 入力/出力 - 視覚・言語から行動を生成 - 未来の RGB, depth, surface normals, optical flow を予測

2. 先行研究と比べてどこがすごい?

- Vision-language-action models - 強い意味 prior を持つが scene dynamics を明示的にモデル化しない - World-action models - 視覚予測と制御を結合するが、細粒度操作に必要な意味・空間構造を必ずしも露出しない - ME-U0 の利点 - 理解と生成を統合し、subtask prediction と affordance grounding を生成のガイドに使用 - 複数の visual dynamics を補完監督として活用 - 下流 supervision なしで zero-shot 能力を発揮

3. 技術・手法の肝は?

- Mixture-of-Transformers アーキテクチャ - understanding expert と generation expert を接続 - 生成のガイド - subtask prediction と affordance grounding - flow matching による joint visual-dynamics and action generation - Visual dynamics - future RGB, depth, surface normals, optical flow を予測 - appearance, geometry, motion の補完監督 - Multi-rate Rotary Position Encoding (MRPE) - visual dynamics と fine-grained control を整列 - 事前学習 - 約 4,200 時間の curated demonstrations - robotic datasets と egocentric datasets を…

4. どうやって有効だと検証した?

- シミュレーションベンチマーク - RoboDojo: 平均スコア 17.66 - LIBERO: 平均成功率 99.0% - LIBERO-Plus: 平均成功率 82.5% - 実世界ロボット操作タスク - シミュレーション外でも有効性を確認 - Zero-shot 検証 - 対応する下流 supervision なしで subtask prediction, affordance grounding, visual dynamics を simulated/real-world observations で確認

5. 議論はある?

- 主張 - 競争力のある下流制御性能と、転移可能な task-grounding / visual-dynamics 能力を両立 - simulation と real world の両方で有効 - 制限・議論点 - 要旨からは不明 - 具体的な失敗例、計算コスト、データ依存性、一般化限界には触れていない

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究 - Vision-language-action models - World-action models - 関連手法 - flow matching - Mixture-of-Transformers - Multi-rate Rotary Position Encoding (MRPE) - ベンチマーク - RoboDojo - LIBERO - LIBERO-Plus

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie

分類: cs.RO, cs.CV

原文アブストラクト

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.

関連論文

PR本紙発行元 EmplifAI