Copper-Policy: ロバストなロボットマニピュレーションのための表現学習
Copper-Policy: Focus on the Representation for Robust Robot Manipulation
方針と共にコンパクトな世界表現を学習し、ピクセル再構成なしで未来の観測埋め込みを予測することで、効率的な学習と高い制御性能を両立したロボットマニピュレーション手法。
著者: Zexin Feng, Yixu Feng, Lingyu Xiao, Shang Su, Kexin Zheng, Chang Xu, Mengkai Shi, Shuo Feng, Xintao Yan
分類: cs.RO, cs.LG
原文アブストラクト
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $π_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.