日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.30404

POIL: 安定力学系を用いた点ベースのワンショット模倣学習

POIL: Point-based One-Shot Imitation Learning with Stable Dynamical Systems

シェア:XThreadsFacebookLINEはてブBluesky

物体の機能部位上の3D点群を共有表現として使い、1回の実演から軌道を転移しつつ、閉ループ制御で外乱や把持姿勢の変化に頑健に実行するワンショット模倣学習手法を提案。

詳しい要約

1. どんなもの?

POILは、stable dynamical systemsを備えたpoint-based one-shot imitation learningフレームワークである。 - one-shot imitationは多数のdemonstration収集を避けられるが、新規物体へのtrajectory転移と、scene条件・grasp configuration・外乱下でのrobustな実行の両方が必要。 - POILはこの両問題を、物体のfunctional part上の3D point集合という共有表現で扱う。 - この点集合をtrajectory transferとclosed-loop executionの両方に用いる。

2. 先行研究と比べてどこがすごい?

stable dynamical modelsをSE(3) poseからpoint setへ拡張した点が新しい。 - 既存のstable dynamical systemsはSE(3) poseを対象としていたが、POILは既知の3D modelやpose estimatorを必要とせずpoint set上でclosed-loop駆動する。 - 単一demonstrationをobject category・grasp pose・goal geometryをまたいで転移しつつ、実行中の外乱から回復できる点を示す。 - 要旨からは、比較対象となる具体的な先行研究名は明示されていない。

3. 技術・手法の肝は?

肝は共有表現としてのfunctional part上の3D point集合である。 - point correspondencesによりdemonstrated trajectoryのone-shot transferを実現。 - multi-modal large language modelで共有functional partをgroundingし、viewpoint・pose・object category変化をまたいでtrajectoryを転移。 - 実行時はmulti-view trackingで同じ点をonline観測。 - Point-set BCSDMが、tracked pointsのみから計算した単一rigid-body twistへper-point velocitiesを射影し、closed loopで駆動。 - goalではcontrollerが古典的SO(3) potential上のgradient flowとなり、rigid-object仮定下でそのpotentialのalmost-global convergenceを終端相が継承する。

4. どうやって有効だと検証した?

simulationとreal-robot experimentsの両方で検証。 - 単一demonstrationをobject category・grasp pose・goal geometryをまたいで転移できることを示す。 - 実行中のexternal disturbancesからの回復を示す。 - 具体的なタスク・指標・ベースラインは要旨からは不明。

5. 議論はある?

要旨からは不明。 - 限界や失敗ケース、計算コスト、multi-modal large language modelのgrounding誤りへの感度などは明示されていない。 - 理論面では、goalでcontrollerが古典的SO(3) potential上のgradient flowとなり、rigid-object仮定下でalmost-global convergenceを継承することが述べられる。

6. 次に読むべき論文は?

要旨で参照・比較されている具体的研究は明示されていない。 - 関連手法として、stable dynamical systems、BCSDM、SE(3)上のstable dynamical models、one-shot imitation learning、multi-modal large language modelによるgroundingが挙げられる。 - 同分野の定番として、Behavior Cloning、Imitation Learning、Dynamical Movement Primitives、SE(3) pose estimation、point cloud correspondenceなどが次の読むべき候補。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sang Min Kim, Jinwoo Seo, Hyeongjun Heo, Junho Lee, Yonghyeon Lee, Young Min Kim

分類: cs.RO

原文アブストラクト

We present POIL, a point-based one-shot imitation learning framework with stable dynamical systems. While one-shot imitation avoids collecting extensive demonstrations, successful one-shot manipulation requires not only transferring a demonstrated trajectory to a novel object but also executing it robustly under changing scene conditions, grasp configurations, and external disturbances. POIL addresses both problems through a shared representation: a set of 3D points on the object's functional part, used jointly for trajectory transfer and closed-loop execution. The one-shot transfer from the demonstrated trajectory is enabled with point correspondences. POIL grounds the shared functional part with a multi-modal large language model, and transfers the trajectory across viewpoint, pose, and object category changes. During execution, multi-view tracking observes the same points online, and Point-set BCSDM drives them in closed loop by projecting per-point velocities onto a single rigid-body twist computed from the tracked points alone. This extends stable dynamical models from an SE(3) pose to a point set without requiring a known 3D model or pose estimator. We show that at the goal the controller becomes a gradient flow on the classical SO(3) potential, so its terminal phase inherits the almost-global convergence of that potential under a rigid-object assumption. Across simulation and real-robot experiments, POIL transfers a single demonstration across object category, grasp pose, and goal geometry, while recovering from external disturbances during execution. Project page: https://sangminkim-99.github.io/poil

関連論文

PR本紙発行元 EmplifAI