PredVLA: サブミリオンパラメータの予測符号化ポリシーによるロボット操作
PredVLA: A Sub-Million-Parameter Predictive-Coding Policy for Robot Manipulation
この論文は、わずか68万パラメータで言語条件付き操作を実現する予測符号化ポリシーPredVLAを提案し、LIBEROベンチマークで高い成功率を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Hiroki Sawada, Shunichi Kasahara
分類: cs.RO
原文アブストラクト
Large pretrained vision-language-action models dominate modern robot-manipulation benchmarks, but it remains unclear how much model scale is necessary for strong language-conditioned control, or whether fundamentally different control architectures can remain competitive at much smaller parameter budgets. We present PredVLA, a language-conditioned predictive-coding policy with only 0.68 million trainable network parameters and no robot-data pretraining, whose hierarchical generative recurrent dynamics predict visual features and proprioception while observations influence latent state only through online inference from the resulting sensory prediction errors. On LIBERO, PredVLA achieves an 86.9% mean success rate across the three short-horizon suites and 75.4% when the long-horizon suite is included. Under a controlled comparison using the same frozen front end, demonstrations, action decoder, and evaluation protocol, PredVLA achieves 3.7x and 7.4x mean success rates of parameter-matched Transformer and LSTM policies, respectively. The predictive-coding formulation also makes the contribution of observation-driven correction directly measurable: because observations influence the recurrent state only through prediction-error-based latent inference, disabling this inference yields an exact open-loop control condition. Together, these results show that a sub-million-parameter recurrent generative policy can achieve strong performance on modern language-conditioned manipulation benchmarks while providing an explicit mechanism for prediction-error-driven online state correction.
関連論文
- FLARE: 視覚言語ロボット操作における自動修正と回復のための障害認識フレームワークマニピュレーション
- リー群制約付きMeanFlowによる高速生成的把持マニピュレーション
- VISTA: 視覚から推定する空間接触注意による高密度接触操作マニピュレーション
- PRISM: 投影統合型サンプリングベースMPCとベイズコスト調整による双腕マニピュレーションマニピュレーション
- ロボット操作のための軌道レベル連続行動表現マニピュレーション
- ブラウザ制御と混合ステッパードライバ構成による再現可能な視覚誘導6自由度ロボットアームマニピュレーション