SUN: 言語に基づく制御から学習、実機への永続的プログラム
SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies
モデルベース制御と学習ポリシーを橋渡しするため、幾何・接触関係を一度定義してMPCコストや報酬、遷移ガードなどにコンパイルする型付き実行可能プログラム「SUN」を導入し、大規模言語ビジョンモデルで自動合成するシステムKuafuを提案。デモなしで高成功率を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
分類: cs.RO, cs.AI
原文アブストラクト
Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what control verified.We introduce Semantically UNified (SUN) Programs, typed executables where geometric and contact relations are defined once and compiled into aligned Model Predictive Control (MPC) costs, satisfaction predicates, RL rewards, transition guards, and diagnostics. Our system, Kuafu, driven by large vision language systems, automatically synthesizes SUN Programs from language and scene semantics, screens feasibility via MPC, and retains semantics while training stage-conditioned policies. Across nine tasks, Kuafu achieves 82.03% macro-success, outperforming sparse-reward (35.67%) and Stage-BC (24.75%) baselines. At 8192-way scale, it generates 10.57x the successful trajectory time per hour of human teleoperation. With 500 trajectories per task, Kuafu data trains DP3 policies to 46.0% simulation success (vs. 22.4% for alternatives) and 34.7% on physical Franka and Kinova robots. These results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.
関連論文
- Peg-in-Bench: 高精度ロボット挿入のためのモジュール式ベンチマークマニピュレーション
- 否定制約付き器用把持のためのポテンシャル誘導粒子ステアリングマニピュレーション
- Facet-0: 接触を伴う精密操作のためのロボット基盤モデルマニピュレーション
- Motus2: 巧みな操作のための自己進化型汎用世界モデルマニピュレーション
- Zeva: 文脈内因果学習による汎用身体操作の実現マニピュレーション
- 制約付きLLMによる意味と物理の橋渡し:安全で信頼できるロボットマニピュレーションマニピュレーション