日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.28818

KeyGen: カテゴリーレベルの方策汎化のための教師なしキーポイントに基づく物体中心表現

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

シェア:XThreadsFacebookLINEはてブBluesky

点群から物体の意味的3Dキーポイントを教師なしで学習し、それを物体中心表現として拡散方策に組み込むことで、未知の物体形状や姿勢に対しても汎化するマニピュレーション手法を提案。

著者: Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg

分類: cs.RO, cs.AI

原文アブストラクト

Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.

関連論文

PR本紙発行元 EmplifAI