日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24068

触覚はいつ重要か? 混雑環境での巧みな把持における視覚とインタラクションのギャップを描く

When Does Touch Matter? Charting the Vision-Interaction Gap in Cluttered Dexterous Grasping

シェア:XThreadsFacebookLINEはてブBluesky

混雑した卓上での巧みな把持において、視覚のみ・力覚・触覚・それらの統合を実機で比較し、インタラクション感覚が接触の早期棄却や再把持を可能にし、把持成功率を大きく向上させることを示した研究。

詳しい要約

1. どんなもの?

- 雑然とした(cluttered)環境でのdexterous graspingを対象とした実世界研究。 - vision、per-finger/wristのwrench推定、分散配置されたfingertip taxelsを組み合わせたdexterous systemを構築。 - 5つのtabletop scene条件で、vision-only、wrench、taxel、combinedの各policyとrepresentation/fusion baselineを比較。 - 接触やocclusionがgrasp品質を隠す状況で、interaction signalがいつ有効かを制御的に評価する。

2. 先行研究と比べてどこがすごい?

- 著者らの知る限り、clutter下のtarget-oriented dexterous graspingにおいて、これらのinteraction modalityを組み合わせかつ個別に評価した初の実世界研究。 - 従来はvision-onlyが主流だったが、本研究はwrenchとtaxelの寄与を分離して示す。 - combined policyは24/25成功(vision-onlyは14/25)、3つのconfined条件では15/15(vision-onlyは6/15)。 - vision-interaction gapが条件依存的に広がることを示し、cluttered graspingをbenchmarkとして位置づける。

3. 技術・手法の肝は?

- vision、per-fingerとwristのwrench推定、distributed fingertip taxelsを統合したdexterous system。 - demonstrations、visual observations、action space、compliant controlを固定し、sensing modalityのみを比較する制御実験設計。 - vision-only、wrench、taxel、combinedのpolicyに加え、representationとfusionのbaselineを比較。 - ablationによりwrenchとtaxel feedbackの相補性を検証。

4. どうやって有効だと検証した?

- 5つのtabletop scene条件での実世界trial比較。 - combined policyは24/25成功、vision-onlyは14/25。 - 3つのconfined条件ではcombinedが15/15、vision-onlyが6/15。 - ablationでwrenchとtaxelの相補性を確認。 - 行動比較により、interaction feedbackが不適切なcontactの早期棄却、lift前のregrasping、より安定したgraspを可能にすることを示す。

5. 議論はある?

- wrenchとtaxel feedbackが相補的であることをablationで示す。 - interaction feedbackによりcontact棄却、regrasping、grasp安定性が改善。 - vision-interaction gapが広がる条件をchartingし、cluttered dexterous graspingをinteraction sensingの必要性を決めるbenchmarkとする。 - 限界や失敗要因の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法としてvision-based dexterous grasping、tactile sensing、wrench estimation、multimodal fusion、compliant controlの代表的文献を次に読むべき。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hao Jiang, Luis Dominguez, Daniel Seita

分類: cs.RO

原文アブストラクト

Dexterous grasping in clutter poses a basic sensing question: when do tactile measurements and external wrench estimates improve on visual geometry? Occlusion and contact can obscure grasp quality, motivating a controlled evaluation of these interaction signals. We present a controlled real-world study over five tabletop scene conditions on a dexterous system that combines vision, per-finger and wrist wrench estimates, and distributed fingertip taxels. With demonstrations, visual observations, action space, and compliant control fixed, we compare vision-only, wrench, taxel, and combined policies plus representation and fusion baselines. The combined policy succeeds in 24/25 trials versus 14/25 for vision only, and 15/15 versus 6/15 across the three confined conditions. Ablations show that wrench and taxel feedback are complementary. Behavioral comparisons show that interaction feedback enables earlier rejection of inadequate contacts, regrasping before lift, and more stable grasps. To our knowledge, this is the first real-world study to combine and separately evaluate these interaction modalities for target-oriented dexterous grasping in clutter. These results chart a widening vision-interaction gap and position cluttered dexterous grasping as a benchmark for determining when the learned policy needs interaction sensing. Project website: https://interaction-dex-grasp.github.io/

関連論文

PR本紙発行元 EmplifAI