日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35715

X-Reset: クロスエンボディメント・リセットによる物体中心強化学習のスケーリング

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

シェア:XThreadsFacebookLINEはてブBluesky

人間の手と物体のデモンストレーションをロボットのリセット状態に変換し、汎用物体中心報酬での強化学習の探索問題を解決するフレームワークを提案。

詳しい要約

1. どんなもの?

- X-Resetは、シミュレーションにおける強化学習(RL)で、器用な操作ポリシーを訓練するフレームワーク。 - タスクに依存しない報酬で汎用ポリシーを訓練する際の探索問題を、人間の手と物体のデモンストレーションを用いて解決する。 - 20物体、3つのembodiment(22-DoFハンドを2つの異なるアームに搭載、およびparallel-jaw gripper)で汎用ポリシーを訓練。 - ポリシーは物体状態とゴールのみに依存し、デモンストレーションはリセット分布を通じて訓練に組み込まれる。

2. 先行研究と比べてどこがすごい?

- 先行研究では、高品質なロボットデモンストレーション、タスクごとの報酬整形、または狭い行動モードへの制限によって探索を容易にしていた。 - X-Resetは、人間の手と物体のデモンストレーションを模倣やリターゲットされた人間動作の追従ではなく、リセットとして利用する点が新しい。 - タスクに依存しない汎用報酬で、多様な物体の操作をゼロから学習する探索問題を解決。 - 訓練物体数に対してスケールし、未見物体に汎化、不完全な手姿勢推定からも学習可能、sim-to-realのゼロショット転移を実現。

3. 技術・手法の肝は?

- 人間の手と物体の状態を運動学的にリターゲットしてノイズの多いロボット状態を生成。 - シミュレーションで不安定な状態をフィルタリングし、残りをRL訓練中のリセットとしてサンプリング。 - 汎用の物体中心報酬(object-centric rewards)を用いて訓練。 - ポリシーは物体状態とゴールのみを入力とし、デモンストレーションはリセット分布を通じて間接的に利用。

4. どうやって有効だと検証した?

- 20物体、3つのembodiment(22-DoFハンドを2つの異なるアームに搭載、およびparallel-jaw gripper)で汎用ポリシーを訓練。 - 訓練物体数に対するスケーリング、未見物体への汎化、不完全な手姿勢推定からの学習、sim-to-realのゼロショット転移を検証。 - ゼロからのRLの探索課題を解決することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法として、人間の手と物体のデモンストレーションを用いた模倣学習、リターゲティング、タスクに依存しない報酬設計、物体中心RL、sim-to-real転移などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Prithwish Dan, Chenyang Ma, Wei Zhan

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.

関連論文

PR本紙発行元 EmplifAI