日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18358

GraphPoint: 意味的実体グラフと点軌道による組合せ的ロボットマニピュレーション

GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

言語指示に応じて物体と動作を組み合わせて汎化するため、意味的実体グラフと把持点の未来軌道予測を組み合わせた操作方策GraphPointを提案し、新ベンチマークCoManiで検証した。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションのための新しいフレームワーク GraphPoint を提案 - 意味的エンティティグラフと幾何学的制御を接続 - 将来のグリッパ点軌道を予測し、ロボット幾何学を用いて行動に変換 - サブタスク内およびサブタスク間の構成的再利用を調査 - ベンチマーク CoMani を導入し、両能力を評価

2. 先行研究と比べてどこがすごい?

- 従来のポリシーはデモンストレーションを超えた一般化が困難 - 言語とシーンが強く相関すると、固定された視覚-行動マッピングを学習 - GraphPoint は言語指示に依存した一般化を実現 - サブタスク内とサブタスク間の両レベルで構成的再利用を評価 - CoMani ベンチマークで制御された分割を提供

3. 技術・手法の肝は?

- 意味的エンティティグラフを構築し、グリッパとオブジェクトを意味的役割で整理 - 行動タイプと修飾子に基づいて相互作用を条件付け - 将来のグリッパ点軌道を予測し、ロボット幾何学で行動に変換 - 予測された進捗が実行中の遷移をガイド - 言語指示に応じた行動生成を可能にする

4. どうやって有効だと検証した?

- CoMani ベンチマーク上で実験とアブレーションを実施 - 制御された分割により、指示依存の一般化を評価 - サブタスク内およびサブタスク間の両レベルで有効性を検証 - 初期シーンを一致させ、単一の意味的要因を変化させることで視覚的ショートカットを排除 - 結果、GraphPoint の有効性が確認された

5. 議論はある?

- 言語とシーンの相関が強い場合のポリシーの問題を指摘 - 構成的再利用の重要性を議論 - ベンチマーク CoMani の設計意図を説明 - アブレーションにより各コンポーネントの寄与を分析 - 詳細な議論は要旨からは不明

6. 次に読むべき論文は?

- CoMani ベンチマークに関連する研究 - 構成的ロボットマニピュレーションの先行研究 - 言語条件付きロボットポリシーの一般化に関する論文 - 意味的グラフと軌道予測を組み合わせた手法 - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kang Luo, Hesheng Wang

分類: cs.RO

原文アブストラクト

Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We investigate compositional reuse at two levels: within a subtask, combining familiar entities, action types, and action modifiers; and across subtasks, reusing learned subtasks in unseen long-horizon tasks. We introduce CoMani, a benchmark with controlled splits for evaluating both capabilities. Matched initial scenes and controlled changes to a single semantic factor encourage reliance on language rather than visual shortcuts. We further propose GraphPoint, which connects semantic entity graphs to geometric control by predicting future gripper point trajectories and converting them into actions using robot geometry. The framework organizes the gripper and objects by semantic roles and conditions their interactions on action types and modifiers, while predicted progress guides transitions during execution. Experiments and ablations on CoMani validate the effectiveness of our method for instruction-dependent generalization at both levels. Code will be released at GraphPoint.

関連論文

PR本紙発行元 EmplifAI