VLA Grounder: ブラックボックスVLAモデルのための言語条件付け空間最適化
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
凍結されたVLAモデルの動作を改善するため、人間の指示をVLAに適したコマンドへ変換する言語条件付け空間ポリシーを強化学習で最適化する手法を提案。
著者: Damir Shodiev, Aleksei Staroverov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov
分類: cs.AI
原文アブストラクト
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input. We study whether frozen VLA policies can be improved by optimizing language space rather than updating action weights. Our method introduces a language-conditioning space policy that translates a human instruction into a short VLA-grounded command using object appearance, spatial relations, and target-grounding cues. The language-conditioning space policy is initialized with a failure-derived command-space prior and optimized with reinforcement learning from sparse task-completion rewards, while the downstream VLA remains fully frozen. This yields language-conditioning space optimization: RL discovers which VLA-grounded commands best elicit successful behavior from the frozen action policy. Experiments on RL4VLA and VL-Think show that language-conditioning space optimization improves success on instruction-sensitive, symbolic, and multi-object manipulation tasks, demonstrating that language can serve as an optimizable variable for a robot foundation models. Website: https://tttonyalpha.github.io/vla_grounder
関連論文
- CounterAlign: 視覚言語行動モデルのための反事実的監督VLA/強化学習
- EXIMO: VLMによるVLAポリシー探索のガイドVLA/強化学習
- StructRL: フローベースVLAのための構造化アクション空間探索VLA/強化学習
- 新規ロボット形態に対するOpenVLA-OFTのRLブートストラッピングVLA/強化学習
- WCM: 視覚・言語・行動強化学習のためのワールドクリティックモデルVLA/強化学習
- 少ないデータから多くを学ぶ:後知恵からの強化学習VLA/強化学習