学習不要の行動クローニング
Training-free Behavior Cloning
行動予測制御(BPC)を提案し、行動認識型検索指標とハンケル行列に基づく行動継続事前分布、閉形式の1ステップ残差補正を組み合わせることで、方策の学習なしに実演データから直接制御を合成する。シミュレーションと実機で学習方策に匹敵する性能を示し、方策適合を数時間から数秒に短縮、Jetson Orin Nano上で75Hz以上の閉ループ制御を実現した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
分類: cs.RO
原文アブストラクト
Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $π_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.