MotorMind: 汎用視覚言語モデルをゼロショットロボットマニピュレーションに活用するフレームワーク
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
汎用VLMが直接観察から行動を生成しフィードバックで適応するだけでロボット操作を実現するハーネスを提案し、専用学習や外部ツールなしでLIBERO-PROや実機xArm6で高い成功率を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang
分類: cs.RO
原文アブストラクト
Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.
関連論文
- ENCORE: 少数の実演から操作戦略を発見するエージェントマニピュレーション
- 汎用ロボットポリシーのためのシンプルなエージェント記憶マニピュレーション
- 文脈内ロボット学習の簡素化:マニピュレーションタスクのための民主化レシピマニピュレーション
- FORM: 直接的な材料則同定によるロボットマニピュレーションマニピュレーション
- MVG-WAM: ロボットマニピュレーションのための多視点幾何認識型ワールドアクションモデルマニピュレーション
- RoboHarn-Evo: 階層的物理知識を進化させ自己改善するロボットマニピュレーションマニピュレーション