日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.36588

強化学習ファインチューニングによる協調マルチエージェント視覚言語行動モデル

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みVLAを3段階の強化学習ファインチューニングでマルチロボット協調に適応させ、初期化対応データ収集・オフライン信用フィルタリング・潜在空間オンラインRLにより成功率を大幅に向上させた。

詳しい要約

1. どんなもの?

- 協調型マルチエージェント VLA モデルのための強化学習 (RL) 手法の研究。 - VLA は大規模単一エージェントデータで事前学習されるため、ロボット間協調に必要な細粒度の調整スキルを欠く。 - マルチロボット実演による SFT では実演データに性能が制限され、自己経験から改善できない。 - 3 段階の reinforced fine-tuning (RFT) パイプラインを提案。 - π0 と π0.5 をバックボーンに、RoboTwin、RoboFactory、実世界の Franka 2 台での 11 タスクで評価。

2. 先行研究と比べてどこがすごい?

- 従来の SFT はマルチロボット実演に依存し、性能が実演データに縛られ自己経験から改善できない。 - 既存の VLA 向けオンライン RL は難しいマルチエージェントタスクで効果が低く、その原因を noisy co-exploration と不安定な更新に帰属。 - 提案手法は初期化を考慮したデータ収集で人間コストを削減しつつ初期化シフトへの頑健性を獲得。 - offline credit-filtered tuning によりエージェント単位で credit を割り当て、正の advantage を持つ軌道のみで微調整。 - オンライン latent-space fine tuning により VLA を凍結し latent noise space で RL を実行。

3. 技術・手法の肝は?

- 第 1 段階: initialization-aware data collection。初期配置を掃引し、事前学習 VLA が繰り返し失敗した時のみ人間実演を呼び出す。 - 第 2 段階: offline credit-filtered tuning。個々のエージェントに credit を割り当て、joint rollout 全体ではなく正の advantage を持つ per-agent trajectory で微調整。 - 第 3 段階: online latent-space fine tuning。VLA を凍結し、その latent noise space で RL を実行。 - これにより noisy co-exploration と不安定な更新を回避。 - バックボーンは π0 と π0.5。

4. どうやって有効だと検証した?

- RoboTwin、RoboFactory、実世界の Franka 2 台によるマニピュレーションの計 11 タスクで評価。 - 平均成功率が RoboTwin で +23.1%、RoboFactory で +16.4%、実世界タスクで +44% 向上。 - コードは https://anonymous.4open.science/r/mavla_rft-2BC0/ で公開。

5. 議論はある?

- 既存の VLA 向けオンライン RL が難しいマルチエージェントタスクで効果が低い理由を noisy co-exploration と不安定な更新に帰属。 - その代替として online latent-space fine tuning を採用。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- π0 - π0.5 - RoboTwin - RoboFactory - SFT (supervised fine-tuning) - online RL for VLAs

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li

分類: cs.RO, cs.AI, cs.MA

原文アブストラクト

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.

関連論文

PR本紙発行元 EmplifAI