日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.19796

LIFD: ロボットマニピュレーションのための3D認識シーン記憶を実現するアンカード拡散モデル

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

単一RGB視点とリカレントメモリから3D認識のシーントークンを拡散モデルで補完し、マニピュレーションポリシーに接続するシーン記憶フレームワークを提案。LIBEROやMetaWorld、実機UR5eで高い成功率を達成した。

詳しい要約

1. どんなもの?

LIFD(Look, Imagine, Focus, and Do)は、部分的観測下のロボット操作のための持続的で3D-awareなシーン記憶フレームワーク。 - 単一RGB視野とrecurrent memoryからscene-token表現を学習・補完する。 - 多視点合意からscene-tokenを学習し、rectified-flow modelで生成する。 - Anchor-Guided Cross-Attentionで現在の幾何特徴に条件付けして補完する。 - コンパクトなslot featuresを操作policyに接続する。 - 展開時は1台のRGBカメラ、proprioception、タスク指示のみ必要。

2. 先行研究と比べてどこがすごい?

先行研究と比べて、部分的観測下で観測履歴を保持しつつ欠損内容を可視証拠と結びつけて推論する点が特徴。 - 幾何対応RGB特徴は可視構造を記述するが、観測済み領域が移動で消失する問題に対処。 - 多視点・幾何supervisionで表現学習し、展開時は単一RGB視野で済む。 - LIBERO平均成功率をJoint training比で3.1ポイント改善(91.6%)。 - MetaWorldで79.8%、UR5e 4タスクで56.0%(OpenVLA-7Bの40.5%を上回る)。

3. 技術・手法の肝は?

技術の肝は、scene-token表現の学習と生成、およびpolicy接続。 - 多視点合意からscene-token表現を学習。 - 単一RGB視野とrecurrent memoryからrectified-flow modelでtokenを生成。 - Anchor-Guided Cross-Attentionで現在の幾何特徴に条件付けして補完。 - コンパクトなslot featuresを操作policyに接続。 - 表現学習に多視点・幾何supervisionを使用。

4. どうやって有効だと検証した?

LIBERO、MetaWorld、UR5eタスクで検証。 - LIBERO平均成功率91.6%(Staged)、Joint training比+3.1ポイント。 - MetaWorldで79.8%。 - UR5e 4タスクファミリ、各10デモで平均成功率56.0%。 - OpenVLA-7Bの40.5%と比較。

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、一般化性に関する議論は記述されていない。

6. 次に読むべき論文は?

要旨で参照・比較されている研究としてOpenVLA-7B、Joint training、LIBERO、MetaWorld、UR5eタスクが挙げられる。 - 関連手法としてrectified-flow model、Anchor-Guided Cross-Attention、scene-token表現、slot featuresが挙げられる。 - 同分野の定番として3D-aware scene memoryや部分的観測下のロボット操作に関する研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu

分類: cs.RO

原文アブストラクト

Robotic manipulation under partial observability requires spatial information that extends beyond the current view. Geometry-aware RGB features describe visible structure, but previously observed regions may disappear as the robot or scene moves. Maintaining a useful scene representation therefore requires retaining observation history while inferring missing content without losing its connection to visible evidence. We introduce LIFD (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns a scene-token representation from multi-view agreement and completes it from a single RGB view and recurrent memory. A rectified-flow model generates the tokens while Anchor-Guided Cross-Attention conditions completion on current geometric features. Compact slot features connect this representation to a manipulation policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6% average success on LIBERO and 79.8% on MetaWorld, improving LIBERO average success by 3.1 percentage points over Joint training. On four UR5e task families with ten demonstrations per family, it achieves 56.0% mean success, compared with 40.5% for OpenVLA-7B.

関連論文

PR本紙発行元 EmplifAI