日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.28258

世界モデルによる汎化可能なロボット挿入

Generalizable Robotic Insertion with World Models

シェア:XThreadsFacebookLINEはてブBluesky

手首カメラの視覚と固有感覚を統合した単一の世界モデルを最大90の挿入タスクで訓練し、未知形状の物体へのゼロショット挿入を実現した研究。

詳しい要約

1. どんなもの?

- ロボットによる挿入タスクを汎用化するフレームワーク。 - world models を用い、robot proprioceptive information と wrist-mounted camera の raw visual observations を組み合わせる。 - 単一の world model を最大90の挿入タスクで訓練し、未知形状の物体に対する zero-shot 成功率56%を達成。 - model-free baseline の7%を大幅に上回る。 - 訓練データセットに物体を追加するほど性能が向上し、スケーラビリティを示す。 - held-out objects での finetuning によりデータ効率が向上し、場合によっては漸近性能も改善。 - 完全に data-driven な方法で未知物体の組立を可能にした初のシステムと主張。

2. 先行研究と比べてどこがすごい?

- 従来の挿入タスクは各タスクに特化した policy に依存し、新規問題への展開が面倒で時間がかかる。 - 提案手法は単一の world model で多様な挿入タスクを扱い、未知物体への zero-shot 成功率56%を達成。 - model-free baseline は7%にとどまり、大幅な性能向上を示す。 - 訓練データセットの物体数を増やすと性能が向上するスケーラビリティを持つ。 - finetuning によりデータ効率が向上し、場合によっては漸近性能も向上する点が先行研究と異なる。 - 完全に data-driven に未知物体を組立可能な初のシステムであると主張。

3. 技術・手法の肝は?

- world models を採用し、robot proprioceptive information と wrist-mounted camera からの raw visual observations を統合。 - 単一の world model を最大90の挿入タスクで訓練。 - モデルベースアプローチにより、未知物体への zero-shot 汎化を実現。 - 訓練データセットに多様な物体を含めることで性能が向上するスケーラブルな学習。 - held-out objects での finetuning によりデータ効率を改善。 - 具体的なネットワーク構造や学習アルゴリズムの詳細は要旨からは不明。

4. どうやって有効だと検証した?

- 最大90の挿入タスクで単一の world model を訓練し、未知形状の物体に対する zero-shot 成功率を評価。 - 提案手法が56%の zero-shot 成功率を達成する一方、model-free baseline は7%であることを示す。 - 訓練データセットに物体を追加するにつれて性能が向上することを確認し、スケーラビリティを検証。 - held-out objects で finetuning を行い、データ効率と漸近性能を評価。 - 実験設定や評価指標の詳細は要旨からは不明。

5. 議論はある?

- 提案手法は未知物体への zero-shot 成功率56%を達成し、model-free baseline の7%を大きく上回る。 - 訓練データセットの物体数を増やすと性能が向上し、スケーラビリティが示唆される。 - finetuning によりデータ効率が向上し、場合によっては漸近性能も改善する。 - 完全に data-driven な未知物体組立の初のシステムと位置づけ、スケーラブルで汎用的なロボット組立への重要な一歩と主張。 - 限界や課題、今後の展望についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は model-free baseline のみ。 - 関連手法として world models、model-based reinforcement learning、robotic insertion、zero-shot generalization に関する論文が挙げられる。 - 具体的な論文名は要旨に記載がないため、同分野の定番として world models (Ha & Schmidhuber, 2018) や model-based RL のサーベイを読むことを推奨。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Nicklas Hansen, Iretiayo Akinola, Yijie Guo, Jie Xu, Bingjie Tang, Hao Su, Xiaolong Wang, Abhishek Gupta, Dieter Fox, Yashraj Narang

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.

関連論文

PR本紙発行元 EmplifAI