日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA/連合学習arXiv:2609.34968

RoboFL: 世界行動モデルのための連合エキスパートアセンブリ

RoboFL: Federated Expert Assembly for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

各クライアントが訓練したLoRAアダプタをサーバー側MoEのエキスパートとして組み込み、ルーティングとエキスパートを協調学習する連合学習フレームワークを提案。実機Frankaで集中学習を12.23%上回り、通信量を最大86.81%削減した。

詳しい要約

1. どんなもの?

- Vision-language-action (VLA) と world-action モデルを対象に、物理インタラクションデータの不足・機関分断・タスク不均一を解決する federated learning 手法。 - 各 client が共有基盤モデルを parameter-efficient fine-tuning (PEFT) で適応し、full-model 更新を交換しない。 - 提案 RoboFL は MoSAIC (Mixture of Slotted Adapters) を federated world-action learning に実装。 - ローカル学習済み LoRA adapter を server MoE の expert branch として直接設置。 - server-side router が token assignment を学習し、routing と expert parameter を同時精緻化。

2. 先行研究と比べてどこがすごい?

- 従来の federated PEFT は naive aggregation で incompatible な更新が混ざる問題があった。 - MoE-style routing を federated aggregation に組み込むと specialization が薄まり expert selection が不安定になる課題があった。 - RoboFL は structured expert assembly により、centralized PEFT の InternVLA-A1 を Franka arm で 12.23% 上回る。 - MoE-based federated VLA baseline と比べ、1 round あたりの client communication を最大 86.81% 削減。

3. 技術・手法の肝は?

- MoSAIC: ローカル LoRA adapter を server MoE の expert branch として直接インストール。 - server-side router が prior-informed branch 上で token assignment を学習。 - Foresight-to-Action Routing Distillation (FARD): モデルの3つの path 間で routing を整合。 - Path-Consensus Expert Aggregation (PCEA): 完全な expert 更新を compact な global adapter に変換し、personalized redistribution を実現。

4. どうやって有効だと検証した?

- RoboTwin 2.0、RLBench、実機 Franka robot arm で実験。 - Franka arm で centralized PEFT の InternVLA-A1 を 12.23% 上回る。 - MoE-based federated VLA baseline 比で per-round client communication を最大 86.81% 削減。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- InternVLA-A1 (centralized PEFT baseline) - MoE-based federated VLA baselines - LoRA - Mixture of Experts (MoE) - RoboTwin 2.0 - RLBench

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Rongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang, Shenli Zheng, Chenrui Wu, Yili Jin, Li Du, Dan Wang, Yuan Du, Shanghang Zhang

分類: cs.RO, cs.AI

原文アブストラクト

Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.

PR本紙発行元 EmplifAI