日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35318

DexAgent: 自己進化するツールライブラリを備えたエージェント型Human2Sim2Robotフレームワークによる巧みな操作

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点の人間動画とタスク指示から、物理的に妥当なロボット軌道を生成して方策学習を可能にするエージェント型フレームワークを提案。ツールライブラリを自己進化させ、多様な物体や長期的タスクに対応する。

詳しい要約

1. どんなもの?

単一の egocentric human video と task prompt から、物理的に妥当な robot trajectories を生成し policy training を可能にする agentic Human2Sim2Robot framework「DexAgent」を提案。 - 4段階で動作: 人間動画の semantic understanding、property-based simulation reconstruction、robot trajectory optimization、robot data generation。 - 各段階で tool library から skill を選択、必要なら新規 skill を開発。 - property-specific verifiers が各段階の成果を物理的妥当性と task 要件で評価し、refinement の feedback を提供。 - 最終段階で object と robot の状態を変化させ、単一動画から多様な robot trajectories を生成し、rendered observati…

2. 先行研究と比べてどこがすごい?

既存の Human2Sim2Robot pipelines は predefined procedures に依存し、多様な object properties や interactions、特に articulated や deformable objects への対応が困難。 - DexAgent は agentic に各段階で skill を選択・開発し、verification-guided に適応するため、多様な物体と long-horizon tasks を処理可能。 - 単一の human video から多様な robot trajectories を生成し、sim-to-real transfer のための retexturing を行う点が新しい。 - 11 real-world tasks で、DexAgent 生成データで訓練した policy は competing baselines より 3.5x 高い success rate を達成。

3. 技術・手法の肝は?

4段階の agentic pipeline が肝。 - semantic understanding of human videos: 人間動画からタスク意図を理解。 - property-based simulation reconstruction: 物体特性に基づく simulation 再構築。 - robot trajectory optimization: 物理的に妥当な軌道最適化。 - robot data generation: 物体・ロボット状態を変化させ多様な軌道を生成し、rendered observations を retexture。 - tool library から skill を選択、必要に応じて新規 skill を開発。 - property-specific verifiers が各段階の成果を評価し、feedback で refinement。 - 新規 skill と verifier を tool library に保持し self-evolving。

4. どうやって有効だと検証した?

11 real-world tasks で検証。 - DexAgent 生成データで訓練した policy は competing baselines より 3.5x 高い success rate を達成。 - 詳細な評価指標やベースラインの内訳は要旨からは不明。 - project website で追加情報を提供。

5. 議論はある?

要旨からは不明。 - 想定される議論点: tool library の self-evolving による処理時間短縮の程度、verifier の汎化性能、sim-to-real gap の残存、articulated/deformable objects への対応限界など。 - 要旨ではこれらに関する具体的な議論や限界は述べられていない。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法として Human2Sim2Robot pipelines、sim-to-real transfer、dexterous manipulation、agentic frameworks が挙げられる。 - 同分野の定番として、domain randomization、imitation learning、reinforcement learning、digital twin などが次に読むべき候補。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Youhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu, Huang Huang

分類: cs.RO

原文アブストラクト

Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. This adaptive, verification-guided process allows DexAgent to process diverse objects and long-horizon tasks. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities. This reduces processing time as DexAgent encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5x higher success rate than competing baselines. Project website: https://dexagent123.github.io/.

関連論文

PR本紙発行元 EmplifAI