日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.38078

MotorMind: 汎用視覚言語モデルをゼロショットロボットマニピュレーションに活用するフレームワーク

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

汎用VLMが直接観察から行動を生成しフィードバックで適応するだけでロボット操作を実現するハーネスを提案し、専用学習や外部ツールなしでLIBERO-PROや実機xArm6で高い成功率を達成した。

詳しい要約

1. どんなもの?

一般の Vision Language Model (VLM) をそのままロボット操作に用いるためのハーネス MotorMind を提案。VLM が観測から直接 mid-level action を提案し、決定論的制御とフィードバックに接続、非同期監視と背景メモリ更新を行う。タスク特化の policy 学習や coding agent、SAM3 などの grounding tool を必要とせず zero-shot で操作を実現する。

2. 先行研究と比べてどこがすごい?

従来の VLA モデルは zero-shot 汎化が限られ、専用学習のため汎用 VLM の進歩を直接活かせない。また VLM を高レベル推論や coding agent に使う agentic システムは外部モデル・ツール依存で複雑・高コスト。MotorMind は外部モデルや SAM3 等の grounding tool なしで、LIBERO-PRO で 66.7%(摂動下 53.8%)を達成し、比較した prior zero-shot 手法の最大 13.3%/19.2% を大きく上回る。

3. 技術・手法の肝は?

VLM が提案する mid-level action を決定論的ロボット制御とフィードバックに接続するハーネス。非同期監視と背景メモリ更新を備え、実行フィードバックに継続的に適応する。タスク特化 policy 学習、coding agent、SAM3 等の追加 grounding tool を使わない。

4. どうやって有効だと検証した?

LIBERO-PRO スイートで base 66.7%、摂動下 53.8% の成功率を達成し、prior zero-shot 手法(最大 13.3%、19.2%)と比較。実機 xArm6 で直接操作と人間による摂動設定を含め平均 95% 成功。より強い VLM に置換すると性能が向上し、残存失敗(視覚 grounding、embodied reasoning、action knowledge 由来)が VLM 能力向上で減少することを確認。

5. 議論はある?

残存する失敗は主に visual grounding、embodied reasoning、action knowledge に起因し、VLM 能力向上に伴い減少する。より強い VLM への置換で性能が向上することから、適切な mid-level action 表現と非同期実行ハーネスを備えれば汎用 VLM が zero-shot ロボット操作を実行可能と主張。

6. 次に読むべき論文は?

要旨で参照・比較されている prior zero-shot 手法、VLA モデル、agentic robotic systems、coding agent、grounding tool の SAM3 など。具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

関連論文

PR本紙発行元 EmplifAI