日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
データ生成arXiv:2609.36171

SkillWeaver: 神経インタラクションスキルを活用したエージェント的探索によるスケーラブルなロボットデータ生成

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

シェア:XThreadsFacebookLINEはてブBluesky

VLMエージェントが学習済みの閉ループ操作スキルを再利用・パラメータ化しながら探索し、長期的な操作データを自動生成するフレームワークを提案。

詳しい要約

1. どんなもの?

- ロボット学習のためのデータ生成フレームワーク「SkillWeaver」を提案。 - エージェントがNeural Interaction Skills (NIS)を探索し、自律的にロボット経験を生成。 - NISは再利用可能でパラメータ化された閉ループポリシー。 - VLMエージェントがタスクと環境を推論し、NISを呼び出して相互作用。 - 検証、反省、記憶を生成し、探索をガイド。 - 39.1Kのデモンストレーションを14.1Kのシーンで生成し、視覚運動ポリシーに蒸留。

2. 先行研究と比べてどこがすごい?

- 従来のデータ生成はオープンループ制御、スクリプト化されたスキルシーケンス、タスク固有プログラムに依存。 - SkillWeaverはエージェントベースで自律的に探索し、閉ループのNISを活用。 - 事前に定められた実行パイプラインを必要とせず、長期的な行動を発見可能。 - シミュレーションと実世界での汎化性能を大幅に向上。

3. 技術・手法の肝は?

- Neural Interaction Skills (NIS)を強化学習ポリシーとして実装。 - 閉ループで接触を伴うマニピュレーションを実現。 - 探索を検証器ガイド付き木探索として組織。 - VLMエージェントがNISを呼び出し、パラメータ化し、結果を観察。 - 検証、反省、記憶を生成して次の探索に活用。

4. どうやって有効だと検証した?

- シミュレーションベンチマークと実世界マニピュレーションで評価。 - SkillWeaver生成経験で訓練した視覚運動ポリシーが、新規物体、空間配置、タスク、環境への汎化を大幅改善。 - ゼロショットおよび少数ショットのsim-to-simとsim-to-real転移を実現。

5. 議論はある?

- エージェントによる神経相互作用スキルの探索が、スケーラブルなロボットデータ生成の代替手段となる可能性を示唆。 - 限界や課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、強化学習ポリシー、VLMエージェント、検証器ガイド付き木探索、視覚運動ポリシー蒸留などが挙げられる。 - 同分野の定番として、テレオペレーションによるデータ収集、シミュレーションでのデータ生成パイプライン、オープンループ制御、スクリプト化スキルシーケンスなどが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki

分類: cs.RO

原文アブストラクト

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes NIS to interact with the environment, observes their outcomes, and generates verification, reflection, and memory to guide subsequent exploration. We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation and organize exploration as verifier-guided tree search, enabling the agent to discover successful long-horizon behaviors without relying on predetermined execution pipelines. SkillWeaver scales autonomously to 39.1K demonstrations across 14.1K scenes, which we distill into visuomotor policies. Across simulation benchmarks and real-world manipulation, training on SkillWeaver-generated experience substantially improves generalization to novel objects, spatial configurations, tasks, and environments, and enables zero- and few-shot sim-to-sim and sim-to-real transfer. Our results suggest agentic exploration over neural interaction skills as a scalable alternative for robot data generation.

関連論文

PR本紙発行元 EmplifAI