日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24180

GraspTune: 触覚駆動による把持実行の微調整で堅牢な把持を実現

GraspTune: Tactile-Driven Execution Refinement for Robust Grasping

シェア:XThreadsFacebookLINEはてブBluesky

視覚的な把持提案を触覚フィードバックで実行段階で修正し、安定した把持成功率を大幅に向上させるフレームワークを提案。

詳しい要約

1. どんなもの?

本論文はGraspTuneという触覚駆動の実行段階リファインメントフレームワークを提案する。視覚的なgrasp proposalを安定した物理的graspへ変換することを目的とし、approach、contact formation、final grasp executionの各段階でbounded residual TCP motionsを適用する。局所depth、tactile signals、state、historyから制御向けcontact semanticsを学習し、contact change、contact risk、post-close readinessをmulti-task supervisionで予測する。diffusion-pretrained residual policyを条件付けし、PPOでclosed-loop executionに整合させる。

2. 先行研究と比べてどこがすごい?

視覚的grasp proposal生成は急速に進歩したが、選択されたproposalを安定した物理的graspへ変換する実行段階は中心的課題として残っていた。GraspTuneは実行層のリファインメントとして、GraspNet、Contact-GraspNet、AnyGrasp、VGNという4つのproposal generatorに対して安定grasp成功率をそれぞれ+19.22、+9.55、+12.45、+20.70ポイント向上させた。また、4-fold held-out category studyで未見物体の実行を54.58%から70.33%へ改善し、カテゴリ不連続な汎化を示した。

3. 技術・手法の肝は?

GraspTuneはnominal proposalから開始し、approach、contact formation、final grasp execution中にbounded residual TCP motionsを適用する。state-conditioned expert contact queriesとmulti-task supervisionを用いて、局所depth、tactile signals、state、historyから制御向けcontact semanticsを学習する。この表現がdiffusion-pretrained residual policyを条件付け、PPOと整合させてclosed-loop executionを実現する。

4. どうやって有効だと検証した?

20物体カテゴリにわたる60,000回以上のシミュレーション実行で、4つのproposal generatorに対する実行層の利点を検証した。4-fold held-out category studyで未見物体の実行成功率を54.58%から70.33%へ向上させた。Xense fingertip sensorsを備えたUR5eセットアップでの1,000回以上の実ロボット試行で、GraspNet実行を71.0%から84.3%へ改善し、実世界ポリシーのfine-tuningなしでの直接転移を検証した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

GraspNet、Contact-GraspNet、AnyGrasp、VGN。これらは本論文で比較・基盤として用いられたproposal generatorである。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Juntao Li, Xingke Xia, Sichao Liu, Daqiang Guo

分類: cs.RO

原文アブストラクト

Visual grasp proposal generation has advanced rapidly, yet converting a selected proposal into a stable physical grasp remains a central execution-stage challenge. This paper introduces GraspTune, a tactile-driven execution-stage refinement framework that starts from a nominal proposal and applies bounded residual TCP motions during approach, contact formation, and final grasp execution. GraspTune learns control-facing contact semantics from local depth, tactile signals, state, and history using state-conditioned expert contact queries and multi-task supervision for contact change, contact risk, and post-close readiness. The representation conditions a diffusion-pretrained residual policy and is aligned with PPO for closed-loop execution. Across more than 60,000 simulated executions over 20 object categories, GraspTune establishes an execution-layer benefit across four proposal generators, raising stable grasp success by +19.22, +9.55, +12.45, and +20.70 percentage points for GraspNet, Contact-GraspNet, AnyGrasp, and VGN. A four-fold held-out category study raises unseen-object execution from 54.58% to 70.33%, showing category-disjoint generalization of contact correction. Across more than 1,000 real-robot trials on a UR5e setup with Xense fingertip sensors, GraspTune raises GraspNet execution from 71.0% to 84.3%, validating direct transfer without realworld policy fine-tuning. Together, these results turn visually plausible proposals into stable physical grasps for downstream contact-rich manipulation. A supplementary video is available at https://youtu.be/kcq7fSLNtzU.

関連論文

PR本紙発行元 EmplifAI