日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00913

eRLT: 行動に関連するトークンルーティングによる効率的なVLA強化学習

eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing

シェア:XThreadsFacebookLINEはてブBluesky

凍結したVLAモデルからタスク固有の行動関連特徴をトークンと層をまたいで抽出し、強化学習のサンプル効率を向上させる手法を提案。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルを下流タスクへ効率的に適応させる強化学習 (RL) 手法。 - 凍結した VLA の内部表現から、タスク固有の action-relevant な特徴を抽出し、actor と critic の state representation を構築する。 - トークンと層をまたぐ routing により、サンプル効率を改善する。

2. 先行研究と比べてどこがすごい?

- 既存手法は VLA 非依存の visual encoder を使うか、VLA 内部表現を固定圧縮するかのいずれかで、タスク固有の action-relevant 特徴を明示的に抽出していなかった。 - eRLT は routing tokens と layer router により、複数深さの視覚言語特徴を動的に集約し、action refinement と action-value estimation に有用な特徴を抽出する。 - その結果、LIBERO と RoboTwin の 7 タスクで平均正規化学習曲線 AUC を最大 23.7% 改善。

3. 技術・手法の肝は?

- 学習された routing tokens が凍結 VLA の複数層の視覚言語特徴を動的に集約する。 - 軽量な layer router がこれらの要約を固定次元の RL token に統合する。 - routing module は expert demonstrations で初期化され、expert action を予測する特徴を捉え、その後 online interaction からの critic feedback で action-value estimation 向けに洗練される。

4. どうやって有効だと検証した?

- LIBERO と RoboTwin の 7 タスクで評価し、代表的なベースラインと比較して平均正規化学習曲線 AUC を最大 23.7% 改善。 - 実機実験として USB connector insertion と motherboard ribbon-cable insertion を行い、最強ベースラインに対して AUC がそれぞれ 108.9% と 46.7% 改善。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、VLA モデル (例: RT-2, OpenVLA)、online RL による VLA 適応 (例: 凍結 VLA の RL fine-tuning)、LIBERO ベンチマーク、RoboTwin ベンチマークなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dehao Huang, Jianbang Liu, Jianpan Gao, Chao Tang, Zilang Cen, Zedong Dan, Jiaheng Wang, Tingguang Li, Yue Wang, Hong Zhang

分類: cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.

関連論文

PR本紙発行元 EmplifAI