日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.25631

DynaForge: 動的マニピュレーションのデモ生成のための計画誘導型残差学習

DynaForge: Planning-Guided Residual Learning for Dynamic Manipulation Demonstration Generation

シェア:XThreadsFacebookLINEはてブBluesky

低周波の大域計画と高周波の物体中心逆運動学を組み合わせ、残差方策で接触時の動的相互作用を補正することで、動的物体操作の高品質なデモンストレーションを生成するフレームワーク。

詳しい要約

1. どんなもの?

- 動的な物体操作のための高品質なデモンストレーション生成手法 - 計画誘導型フレームワークで、残差補正を学習 - 低頻度のグローバル計画と高頻度の物体中心逆運動学を組み合わせ - 動的相互作用中の行動を残差ポリシーで修正 - 暗黙的カリキュラムでロールアウトをグループ化し、混合成功グループを選択 - 残差強化学習を進化する能力フロンティアに集中 - シミュレーション9タスクで平均成功率を41.30%から78.37%に向上 - 実世界3タスクで30-60%の成功率を達成

2. 先行研究と比べてどこがすごい?

- 静的タスク向け手法は動的設定に容易に転用できない - 計画ベース手法は接触付近で失敗する可能性 - DOMINOスタイルのリプレイは動的相互作用を単純化し、ポリシー学習の経験を制限 - DynaForgeは計画と残差学習を組み合わせ、接触付近の失敗を克服 - 同じ名目環境ステップ予算で、vanilla GRPOの0.73倍の最適化ステップでより高い最終成功率 - シミュレーションで計画先行の41.30%から78.37%に改善 - DP3ポリシーでDOMINOデータの7.07%に対し49.11%の平均成功率 - 実世界でDOMINOの0-10%に対し30-60%の成功率

3. 技術・手法の肝は?

- 計画誘導型フレームワーク - 低頻度グローバル計画と高頻度物体中心逆運動学をタスクフェーズ間で組み合わせ - 残差ポリシーを適用し、動的相互作用中の行動を修正 - 暗黙的カリキュラム:マッチした条件下でロールアウトをグループ化 - 混合成功グループを選択し、残差強化学習を能力フロンティアに集中 - 残差学習により計画の限界を補正

4. どうやって有効だと検証した?

- シミュレーション9タスクで評価 - 平均デモ生成成功率が41.30%から78.37%に向上 - CanとBottleで、同じ名目環境ステップ予算でvanilla GRPOの0.73倍の最適化ステップでより高い最終成功率 - タスクあたり800デモでDP3ポリシーを訓練 - DynaForgeデータで49.11%の平均成功率、DOMINOデータで7.07% - 実世界3動的タスクで30-60%の成功率、DOMINO訓練ポリシーは0-10% - シミュレーションから実世界への転移能力を実証

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- DOMINO - DP3 - GRPO

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yiyang Jin, Yu Zheng, Xiao He, Hesheng Wang

分類: cs.RO

原文アブストラクト

Dynamic object manipulation is essential for robots operating in real-world environments, yet methods for generating high-quality demonstrations remain limited. Methods designed for static tasks do not readily transfer to dynamic settings. Among dynamic demonstration generators, planning-based methods can fail near contact, while DOMINO-style replay simplifies dynamic interactions and may limit the experience available for policy learning. We present DynaForge, a planning-guided framework that learns residual corrections for dynamic manipulation demonstration generation. DynaForge combines low-frequency global planning with high-frequency object-centric inverse kinematics across task phases, and applies a residual policy to correct actions during dynamic interaction. An implicit curriculum groups rollouts under matched conditions and selects mixed-success groups, focusing residual reinforcement learning on the evolving competence frontier. On Can and Bottle, it uses 0.73x as many optimizer steps as vanilla GRPO at the same nominal environment-step budget, with higher observed final success rates. Across nine simulation tasks, DynaForge increases mean demonstration-generation success from 41.30% of the planning prior to 78.37%. With 800 demonstrations per task, DP3 policies trained on DynaForge data achieve 49.11% mean success, compared with 7.07% for DOMINO data. On three real-world dynamic tasks, DynaForge-trained policies achieve 30-60% success, compared with 0-10% for DOMINO-trained policies, showing the ability of DynaForge for sim-to-real transfer.

関連論文

PR本紙発行元 EmplifAI