RE-0: 局所オンポリシー蒸留による身体性コード・アズ・ポリシーエージェントの検証済み再帰的改善
RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation
教師の失敗を考慮し、学生の失敗履歴に対する教師の局所修正を環境で検証してからオンポリシー蒸留の監督に使うことで、コード生成型身体性エージェントを再帰的に改善する手法を提案。
著者: Jiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou, Chengye Song, Hongzhou Wang
分類: cs.RO, cs.AI, cs.LG
原文アブストラクト
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.