日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2608.31167v1

SUN: 言語に基づく制御から学習、実機への永続的プログラム

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

シェア:XThreadsFacebookLINEはてブBluesky

モデルベース制御と学習ポリシーを橋渡しするため、幾何・接触関係を一度定義してMPCコストや報酬、遷移ガードなどにコンパイルする型付き実行可能プログラム「SUN」を導入し、大規模言語ビジョンモデルで自動合成するシステムKuafuを提案。デモなしで高成功率を達成した。

詳しい要約

1. どんなもの?

本論文は、長期にわたる操作タスクにおいて、モデルベース制御と学習ポリシーを統合するための新しいフレームワークであるSemantically UNified (SUN) Programsを提案している。SUN Programsは、幾何学的・接触関係を一度定義し、Model Predictive Control (MPC)コスト、充足述語、RL報酬、遷移ガード、診断情報にコンパイルされる型付き実行可能ファイルである。システムKuafuは、大規模視覚言語システムを用いて、言語とシーンセマンティクスからSUN Programsを自動合成し、MPCで実現可能性をスクリーニングし、セマンティクスを保持したままステージ条件付きポリシーを訓練する。

2. 先行研究と比べてどこがすごい?

従来の制御と学習の橋渡し手法では、タスクセマンティクスが破棄され、報酬が手動設計され、制御が検証した動作から学習ポリシーが逸脱する問題があった。SUN Programsは、セマンティクスを明示的に保持し、制御と学習の間で整合性を取る点が新しい。また、デモンストレーションや手動の密な報酬を必要とせず、シミュレーションでスクリーニングしたタスクセマンティクスを学習に活用することで、ロバストなポリシーを実現している。

3. 技術・手法の肝は?

手法の核は、SUN Programsの設計とKuafuシステムのパイプラインである。SUN Programsは、幾何学的・接触関係を型付き実行可能ファイルとして定義し、MPCコスト、RL報酬、遷移ガードなどにコンパイルする。Kuafuは、大規模視覚言語システムを用いて言語とシーンセマンティクスからSUN Programsを自動合成し、MPCで実現可能性をスクリーニングし、セマンティクスを保持したままステージ条件付きポリシーを訓練する。

4. どうやって有効だと検証した?

9つのタスクで評価し、Kuafuは82.03%のマクロ成功率を達成し、スパース報酬(35.67%)とStage-BC(24.75%)のベースラインを上回った。8192ウェイスケールでは、人間のテレオペレーション1時間あたりの成功軌道時間が10.57倍に向上した。タスクあたり500トラジェクトリで、Kuafuデータで訓練したDP3ポリシーはシミュレーション成功率46.0%(代替手法は22.4%)、物理FrankaおよびKinovaロボットで34.7%を達成した。

5. 議論はある?

要旨からは、議論の余地や限界についての詳細は不明である。ただし、シミュレーションでスクリーニングしたセマンティクスが実ロボットに転移する際のギャップや、大規模視覚言語システムの合成精度への依存などが潜在的な課題として考えられるが、要旨には明記されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、Model Predictive Control (MPC)、RL報酬設計、Behavior Cloning (BC)、DP3 (Diffusion Policy) などが挙げられる。次に読むべき論文としては、これらの手法の詳細や、言語条件付きロボット操作に関する最近の研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong

分類: cs.RO, cs.AI

原文アブストラクト

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what control verified.We introduce Semantically UNified (SUN) Programs, typed executables where geometric and contact relations are defined once and compiled into aligned Model Predictive Control (MPC) costs, satisfaction predicates, RL rewards, transition guards, and diagnostics. Our system, Kuafu, driven by large vision language systems, automatically synthesizes SUN Programs from language and scene semantics, screens feasibility via MPC, and retains semantics while training stage-conditioned policies. Across nine tasks, Kuafu achieves 82.03% macro-success, outperforming sparse-reward (35.67%) and Stage-BC (24.75%) baselines. At 8192-way scale, it generates 10.57x the successful trajectory time per hour of human teleoperation. With 500 trajectories per task, Kuafu data trains DP3 policies to 46.0% simulation success (vs. 22.4% for alternatives) and 34.7% on physical Franka and Kinova robots. These results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.

関連論文