日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.20575

サンプリングベースMPCによる視覚ポリシー学習の加速

Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control

シェア:XThreadsFacebookLINEはてブBluesky

サンプリングベースのモデル予測制御と一次勾配ポリシー最適化を組み合わせ、深度画像から歩行・操作ポリシーを効率的に学習し、実機ロボットへゼロショット転移する手法を提案。

詳しい要約

1. どんなもの?

- 視覚に基づくロコモーションとマニピュレーションのポリシー学習を効率化する手法。 - Sampling-Guided Policy Search (SGPS) を提案。 - サンプリングベースの Model Predictive Control (MPC) と一階勾配ポリシー最適化を組み合わせる。 - 行動クローニングで初期化し、訓練ではサンプリングによる改良と短ホライズンの FoPG 更新を交互に行う。 - 視覚ポリシー訓練にはレンダリングを計算グラフから除外した decoupled FoPG を使用。 - 単一 GPU で Unitree Go2 と G1 のシミュレーションにおいてロコモーション、障害物走破、クレート押し、両腕運搬のポリシーを学習。 - 実機 Go2 にゼロショット転移し、オンボード深度で自律的にトロット、クロール、ハードル越え、行動切替を実現。

2. 先行研究と比べてどこがすごい?

- 従来の一階勾配ポリシー勾配 (FoPG) は微分可能シミュレーションで訓練コストを削減するが、局所最適化により意図しない接触パターンに収束する問題があった。 - SGPS はサンプリングベース MPC による反復的な行動目標改良を組み合わせることで、この短所を克服。 - 実験では、改良が初期化や追従のみを上回るポリシー学習を実現することを示した。 - また、decoupled FoPG によりレンダリングを計算グラフから除外し、状態ポリシーティーチャーなしで深度観測から直接学習可能にした点が新しい。

3. 技術・手法の肝は?

- サンプリングベースの Model Predictive Control (MPC) と一階勾配ポリシー最適化を組み合わせた Sampling-Guided Policy Search (SGPS) を提案。 - 行動クローニングでポリシーをサンプリング行動から初期化。 - 訓練では、サンプリングベースの改良と短ホライズンの FoPG 更新を、摂動初期状態とランダム化ダイナミクスの下で交互に実行。 - 視覚ポリシー訓練には decoupled FoPG 定式化を採用し、レンダリングを計算グラフから除外。 - これにより深度観測から直接学習し、状態ポリシーティーチャーを不要に。

4. どうやって有効だと検証した?

- 単一 GPU 上で Unitree Go2 と G1 ロボットのシミュレーションにおいて、ロコモーション、障害物走破、クレート押し、両腕運搬のポリシーを学習。 - 実験により、改良が初期化や追従のみを上回るポリシー学習を実現することを示した。 - ハードウェア展開では、蒸留ポリシーが実機 Go2 にゼロショット転移し、オンボード深度を用いて自律的にトロット、クロール、ハードル越え、行動切替を実行。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、First-order Policy Gradients (FoPG)、Sampling-based Model Predictive Control (MPC)、Behavior Cloning、Decoupled FoPG が挙げられる。 - 同分野の定番として、Differentiable Simulation、Model-Based Reinforcement Learning、Visual Policy Learning などが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yilang Liu, Haoxiang You, Qian Wang, Daniel Rakita, Ian Abraham

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact patterns. To address this shortfall, we propose Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization. Behavior cloning initializes the policy from sampled actions; training then alternates sampling-based refinement with short-horizon FoPG updates under perturbed initial states and randomized dynamics. For visual policy training, we use a decoupled FoPG formulation that excludes rendering from the computation graph, enabling direct learning from depth observations without a state-policy teacher. On a single GPU, SGPS learns policies for locomotion, obstacle traversal, crate pushing, and bimanual carrying on simulated Unitree Go2 and G1 robots. Our experiments further show that refinement improves policy learning beyond initialization and tracking alone. For hardware deployment, the distilled policy transfers zero-shot to a real Go2 and uses onboard depth to autonomously trot, crawl, clear hurdles, and switch between these behaviors.

関連論文

PR本紙発行元 EmplifAI