日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.32762

視点一般化可能な視覚模倣学習ポリシーに何が重要か:実証研究

An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning

シェア:XThreadsFacebookLINEはてブBluesky

視覚模倣学習において、密な視覚トークンの保持と幾何学的推論を行う行動ヘッドが視点一般化に有効であることを実験的に示し、シミュレーションから実世界へのゼロショット転移も実現した。

著者: Mino Nakura, Sriram Krishna, Yufei Wang, Shubham Tulsiani, Zackory Erickson, David Held

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.

関連論文

PR本紙発行元 EmplifAI