MobiAgent: 長期的モバイルマニピュレーションのための二重ループ再帰的政策自己改善
MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation
VLMを基盤に高レベル推論と低レベル制御を分離し、展開中の自己改善ループと自動スキル獲得ループを組み合わせて長期的モバイルマニピュレーションを実現するエージェントフレームワーク。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi
分類: cs.RO, cs.AI
原文アブストラクト
Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $π_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.
関連論文
- MM-ABC: 見て、協調し、想像する汎用モバイルマニピュレーションモバイルマニピュレーション
- SAKI: 人間の動画からのスキル組み立てと運動学的模倣による長期的モバイルマニピュレーションモバイルマニピュレーション
- OpenArmベースの実験室モバイルマニピュレーションにおける表現ハンドオフモバイルマニピュレーション
- FloAff-Kitchen:正準的かつ段階的な床面アフォーダンス学習によるナビゲーションと操作の橋渡しモバイルマニピュレーション
- DynaMOMA: 動的物体のモバイル操作のための把持姿勢の瞬時予測モバイルマニピュレーション
- 開世界モバイルマニピュレーションのための関節3Dシーングラフモバイルマニピュレーション