ESP: ワンステップ多峰性行動生成のためのエネルギー・スコア方策
ESP: Energy-Score Policy for One-Step Multimodal Action Generation
拡散・フローマッチング方策の反復サンプリングを不要にし、エネルギー・スコアで行動ヘッドを訓練して1回のネットワーク評価で行動チャンクを生成する手法を提案。シミュレーションと実機のマニピュレーションで、遅延を大幅に削減しつつ競争力のある成功率を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Lilika Makabe, Heecheol Kim, Yasuyuki Matsushita
分類: cs.RO
原文アブストラクト
Generative action models based on diffusion and flow matching have been increasingly adopted in vision-language-action (VLA) policies for their ability to capture diverse behaviors, including multiple valid action sequences under the same observation and instruction. Their iterative sampling procedures, however, require repeated network evaluations to generate each action chunk, increasing inference latency in closed-loop control. We propose ESP (Energy-Score Policy), a teacher-free approach that maps policy context and noise directly to an action chunk in a single network evaluation. ESP trains the action head with the energy score rather than mean squared error. Whereas squared-error regression targets the conditional mean, the energy score is strictly proper: its expected value is uniquely minimized by the target distribution. This provides a principled objective for learning multimodal action distributions without iterative sampling, with exact recovery at the population optimum when the model can represent the target distribution. Experiments on both simulation and real-world manipulation tasks demonstrate competitive task success with substantially lower action-generation latency than the flow matching baseline. These results support direct distributional learning as an efficient alternative to iterative generative robot policies.