日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モバイルマニピュレーションarXiv:2610.03476

MobiAgent: 長期的モバイルマニピュレーションのための二重ループ再帰的政策自己改善

MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

VLMを基盤に高レベル推論と低レベル制御を分離し、展開中の自己改善ループと自動スキル獲得ループを組み合わせて長期的モバイルマニピュレーションを実現するエージェントフレームワーク。

詳しい要約

1. どんなもの?

- 長期的な mobile manipulation を対象とした dual-loop agentic framework『MobiAgent』 - Inner Loop は VLM による receding-horizon planning と visual reflection で atomic skills を動的に構成 - Outer Loop は deployment rollouts を自動で分割・検証・クラスタリングし、skill library を継続的に fine-tuning - 評価は RoboCasa、BEHAVIOR-1K、実世界タスクで実施

2. 先行研究と比べてどこがすごい?

- 従来の Vision-Language-Action models は短期的タスクに強く、多段階目的の階層推論が不足 - 既存 hierarchical agents は sub-task mapping が固定的で replanning が柔軟でなく、継続学習も欠如 - MobiAgent は BEHAVIOR-1K で $π_{0.5}$-TA を 22.5 percentage points 上回る - 実行失敗からの robust recovery を実現し、自律データ再利用で成功率を改善

3. 技術・手法の肝は?

- Inner Loop: VLM による receding-horizon planning と visual reflection で atomic skills を動的構成 - 高レベル推論と低レベル制御を分離し、composable atomic skills を活用 - skills は unified VLM backbone を共有する specialized flow-matching experts が実行 - Outer Loop: deployment rollouts を自動分割・検証・クラスタリングして atomic skills を発見 - 人間のアノテーションなしで skill library を継続的に fine-tuning

4. どうやって有効だと検証した?

- RoboCasa、BEHAVIOR-1K、実世界タスクで評価 - BEHAVIOR-1K で $π_{0.5}$-TA を 22.5 percentage points 上回る - 実行失敗からの robust recovery を確認 - 自律データ再利用により RoboCasa で 7.50% から 27.50%、Astribot S1 で 32.5% から 57.5% へ成功率改善

5. 議論はある?

- 長期的 mobile manipulation における compounding execution errors と capacity interference が課題 - locomotion と arm control 間の干渉を unified VLM backbone 共有で緩和 - 既存手法の rigid sub-task mapping、inflexible replanning、継続学習欠如を克服 - 具体的な限界や失敗事例、計算コスト、スケーラビリティの議論は要旨からは不明

6. 次に読むべき論文は?

- $π_{0.5}$-TA - Vision-Language-Action models - hierarchical agents - flow-matching experts - RoboCasa - BEHAVIOR-1K - Astribot S1

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi

分類: cs.RO, cs.AI

原文アブストラクト

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $π_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.

関連論文

PR本紙発行元 EmplifAI