できないことを見つけよう:自己改善型VLAモデルのためのエージェント的実世界強化学習
Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models
視覚言語行動モデルが実世界で自律的に練習し、失敗から学んで自己改善する強化学習フレームワークFINDを提案。人間によるリセットや報酬ラベルなしで成功率を55%から71.9%に向上させた。
著者: Yuan Fang, Zechu Li, Haolei Tong, Puze Liu, Georgia Chalvatzaki
分類: cs.RO
原文アブストラクト
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $π_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: FIND.github.io.