日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
移動マニピュレーションarXiv:2610.05414

RMMBench:ロボット移動マニピュレーションのための包括的ベンチマーク

RMMBench: A Comprehensive Benchmark for Robotic Mobile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

移動とマニピュレーションを組み合わせた70種類のタスクでVLMの身体性能力を評価するベンチマークを提案し、最先端VLMでも空間定位に課題があることを示した。

詳しい要約

1. どんなもの?

- RMMBenchは、ロボットのモバイルマニピュレーションを評価するための包括的なベンチマークである。 - 言語指示を理解し、連続空間で長期的なタスクを実行する能力を評価する。 - 高レベルと低レベルの身体化タスクを統合フレームワークにシームレスに統合する。 - 「ナビゲーション-マニピュレーション」タスクスイートを構築し、局所的なマニピュレーションから長期的な複合ナビゲーションまで70の標準的なタスクシナリオを含む。

2. 先行研究と比べてどこがすごい?

- 既存のベンチマークは、多様なロボットタスクを評価する包括的な方法を欠いている。 - 評価指標が比較的制限されており、VLMの身体化能力を徹底的かつ細粒度で評価することが難しい。 - RMMBenchは、これらの問題に対処し、高レベルと低レベルのタスクを統合し、多様なシナリオを提供する。

3. 技術・手法の肝は?

- 言語指示を理解し、連続空間で長期的なタスクを実行するロボットを必要とする評価ベンチマークを提案。 - 高レベルと低レベルの身体化タスクを統一フレームワークに統合。 - 70の標準的なタスクシナリオからなる「ナビゲーション-マニピュレーション」タスクスイートを構築。 - 局所的なマニピュレーションから長期的な複合ナビゲーションまでをカバー。

4. どうやって有効だと検証した?

- 最先端のVLMを評価し、モバイルマニピュレーションタスクにおいて空間定位に大きな課題があることを明らかにした。 - 長期的なインタラクション中にロボットの空間知覚能力を強化する必要性を強調した。 - ベンチマークはhttps://mxxq-stack.github.io/rmmbench-project/でアクセス可能。

5. 議論はある?

- 最先端のVLMはモバイルマニピュレーションタスクにおいて空間定位に依然として大きな課題がある。 - 長期的なインタラクション中にロボットの空間知覚能力を強化する必要性が示唆された。 - その他の議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、Vision-Language Models (VLMs) や Robotic Mobile Manipulation に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huapeng Li, Fuxiang Feng, Jinqiu Fan, Shuo Yang, Fengjiao Chen, Xuezhi Cao, Ran Song, Wei Zhang

分類: cs.RO

原文アブストラクト

Although the advancement of vision-language models (VLMs) has endowed robots with enhanced environmental understanding and task reasoning, a comprehensive evaluation methodology is important to advance the integration of VLMs in robotic navigation and manipulation. However, current benchmarks lack a comprehensive method to evaluate diverse robotic tasks, and evaluation metrics remain relatively constrained, making it difficult to assess the embodied capabilities of VLMs in a thorough and fine-grained manner. To address this issue, we propose RMMBench, an evaluation benchmark that requires robots to understand language instructions and perform long-horizon tasks in continuous spaces. RMMBench seamlessly integrates high- and low-level embodied tasks into a unified framework, constructing a "navigation-manipulation" task suite comprising 70 canonical task scenarios that range from localized manipulation to long-horizon composite navigation. The results reveal that leading VLMs still face major challenges in spatial localization when performing mobile manipulation tasks, and also highlight the necessity of enhancing the spatial perception capability of robots during long-horizon interactions. RMMBench can be accessed at https://mxxq-stack.github.io/rmmbench-project/

関連論文

PR本紙発行元 EmplifAI