日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチロボット協調arXiv:2610.02161

DuoMind: 意味的通信による分散マルチロボット協調

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

シェア:XThreadsFacebookLINEはてブBluesky

VLMによる高レベル調整とVLAによる低レベル実行を組み合わせた分散階層型フレームワークで、ロボット間の意味的メッセージ交換により長期的な協調作業を実現した。

詳しい要約

1. どんなもの?

- 複数ロボットの協調を目的とした分散階層フレームワーク DuoMind を提案。 - 各ロボットは VLA ベースの action model で低レベル実行、VLM ベースの orchestrator で高レベル推論とエージェント間調整を担う。 - 各 planning step で orchestrator はタスク指示、局所観測、他ロボットからのメッセージを推論し、action model への低レベル指示と peer ロボットへの semantic message を生成。 - 分散制御下での長期的 manipulation タスクを評価する benchmark RoboPoly も開発。

2. 先行研究と比べてどこがすごい?

- 従来の VLM/VLA の進展は主に single-robot 設定に集中していた。 - 複数ロボット系への拡張は、長期的行動の調整と信頼性の高い細粒度実行の両立が難しかった。 - DuoMind は分散階層構成と semantic communication により、この multi-robot coordination の課題に対処。 - さらに multi-robot coordination 用 benchmark が不足していたため RoboPoly を構築。

3. 技術・手法の肝は?

- 各ロボットに VLA-based action model と VLM-based orchestrator を配置する分散階層アーキテクチャ。 - orchestrator はタスク指示、局所観測、受信メッセージを入力に推論。 - 出力として action model 用の低レベル指示と、他ロボット向けの semantic message を生成。 - VLM の semantic reasoning と VLA の precise action-generation を組み合わせ、pretrained model の補完的強みを活用。

4. どうやって有効だと検証した?

- RoboPoly と RoboTwin 上で実験を実施。 - DuoMind が multi-robot task performance を改善することを示した。 - ablation studies により hierarchical orchestration と semantic communication の寄与を確認。 - 詳細は project page で公開。

5. 議論はある?

- 要旨からは不明。 - ただし multi-robot coordination 用 benchmark の不足を指摘し、RoboPoly を開発した点は議論の出発点として示唆される。

6. 次に読むべき論文は?

- RoboPoly(本論文で提案された benchmark) - RoboTwin(実験で使用された benchmark) - VLM および VLA の基盤研究(一般名として vision-language model, vision-language-action model)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang

分類: cs.RO, cs.AI

原文アブストラクト

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.

関連論文

PR本紙発行元 EmplifAI