日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24896

モジュール設計と三点インターフェースによる操舵可能で反応的な把持

Steerable and Reactive Grasping Through Modular Design with a Three-Point Interface

シェア:XThreadsFacebookLINEはてブBluesky

物体形状から把持候補点を選び、モデルベース制御で接近、最後は強化学習方策が指関節情報のみで把持を安定化するモジュール型把持フレームワークを提案。

詳しい要約

1. どんなもの?

- 物体把持のためのモジュール型フレームワーク - 3点インターフェースで把持位置決定、到達、接触維持を統合 - 物体形状と言語コマンドから把持候補を生成 - モデルベース反応制御とRLポリシーを組み合わせ - シミュレーションと実機で検証

2. 先行研究と比べてどこがすごい?

- 従来のsqueezeベースラインやend-to-endベースラインと比較 - モジュール設計により幾何推論と局所制御を分離 - 単一のRLポリシーを多物体・多把持構成に共有 - 視覚や物体形状に依存せず固有感覚のみで把持安定化

3. 技術・手法の肝は?

- 事前計算されたgrasp-affordance heatmapから接触3点組をサンプリング - モデルベース反応コントローラが物体追跡、衝突回避、手先誘導 - 最終数センチでRLポリシーが固有感覚フィードバックで把持を洗練・安定化 - RLポリシーは指関節状態と最近の行動のみ観測

4. どうやって有効だと検証した?

- シミュレーションでgrasp-and-lift成功率をsqueezeおよびend-to-endベースラインと比較 - 到達収束特性を評価 - 把持ステアリングを実証 - 実機で2つの訓練物体と1つの未見物体に対するパイプライン全体をデモ

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- squeezeベースライン、end-to-endベースライン、grasp-affordance heatmap、Reinforcement Learning (RL) ポリシー

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Andrew Nguyen, Yonghyeon Lee, Sangbae Kim

分類: cs.RO

原文アブストラクト

Dexterous grasping requires deciding where to grasp, reaching the target, and maintaining stable contact. We connect these stages through a compact three-point interface that separates global geometric reasoning from local contact control. Given object geometry and optional language commands, our framework samples contact triples from a precomputed grasp-affordance heatmap. A model-based reactive controller tracks the object, avoids collisions, and guides the hand toward the selected contacts. In the final centimeters, a Reinforcement Learning (RL) policy uses proprioceptive feedback to refine and stabilize the grasp despite reaching and perception errors. It observes only finger joint states and its recent actions, with no target points, visual observations, or object geometry, so a single policy is shared across objects and grasp configurations. In simulation, we compare grasp-and-lift success against squeeze and end-to-end baselines, characterize reaching convergence, and demonstrate grasp steering; hardware demonstrations on two training objects and one unseen object illustrate the full pipeline. Our modular framework uses geometry to guide the reach and local feedback to secure the grasp.

関連論文

PR本紙発行元 EmplifAI