日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.31337

表現ガイドによるロボットマニピュレーション用実行可能プログラムの生成と統合

Representation-Guided Generation and Integration of Executable Programs for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

VLMを用いて、物体中心表現を共有しながら知覚・計画・制御プログラムを生成・統合し、ロボットマニピュレーションを実現するフレームワークRIVETを提案。

詳しい要約

1. どんなもの?

- VLM code generation を用いてロボットマニピュレーションシステムを自動構築するフレームワーク RIVET を提案。 - 共有の object-centric representation を中心に、知覚・計画・制御を統合する。 - 対象タスクは cube stacking、tangram rearrangement、3D assembly。 - シミュレーションと実機で評価し、実世界で 83% の成功率を達成。

2. 先行研究と比べてどこがすごい?

- 従来の VLM code generation では、独立に生成されたコンポーネントが互換性のない幾何・タスク情報で動作する問題があった。 - RIVET は共有表現を介して生成プログラムを統合し、この不整合を解消する。 - 一度生成したプログラムを未知の開始・目標配置に再利用でき、コード再生成が不要。 - 異なる幾何・関係・順序要件のタスクに共通フレームワークを適応可能。

3. 技術・手法の肝は?

- 共有表現は per-object 6D poses と relation graph を組み合わせる。 - 6D poses は action grounding に必要な metric 情報を保持。 - relation graph は planning に必要なタスクレベル構造を提供。 - この表現に導かれ、VLM が知覚・レンダリング・関係推論・計画の協調プログラムを生成。 - 各プログラムはタスク固有計算と利用可能なパッケージを組み合わせる。

4. どうやって有効だと検証した?

- シミュレーションと実機で cube stacking、tangram rearrangement、3D assembly を評価。 - オフライン生成したシステムを再利用し、実世界で 83% の全体成功率を達成。 - 結果から、表現誘導型プログラム生成が異なる要件のタスクに適応できることを示す。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、一般化範囲についての議論は要旨に記載なし。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法として VLM code generation、object-centric representation、6D pose estimation、relation graph、task and motion planning が挙げられる。 - 同分野の定番として、Code as Policies、SayCan、Inner Monologue などが次に読む候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruixiao Yang, Mingxin Yu, Chuchu Fan

分類: cs.RO

原文アブストラクト

Building a robotic manipulation system requires connecting perception, planning, and control through carefully designed representations and interfaces. VLM code generation offers a way to automate this construction, but independently generated components may operate on incompatible geometric and task-level information. We present Representation-guided Integration of VLM-generated Executable Task programs (RIVET), a framework for generating complete manipulation systems around a shared object-centric representation. The representation combines per-object 6D poses, which preserve the metric information required for action grounding, with a relation graph that exposes the task-level structure required for planning. Guided by this representation, a VLM generates cooperating perception, rendering, relation-inference, and planning programs, each combining task-specific computation with available packages where useful. The resulting programs are authored once for a manipulation domain and reused on unseen start and goal configurations without code regeneration. We evaluate RIVET on cube stacking, tangram rearrangement, and three-dimensional assembly in simulation and on a physical robot, where we achieve 83% overall success rate in the real world by reusing offline-generated systems. Our results demonstrate that representation-guided program generation can adapt a common manipulation framework to tasks with different geometric, relational, and sequential requirements.

関連論文

PR本紙発行元 EmplifAI