日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2609.21726

ZeroTouch: 触覚教師あり視覚接触推定による接触リッチマニピュレーション

ZeroTouch: Tactile-Supervised Visual Contact Estimation for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

訓練時のみ触覚を教師信号として使い、展開時は手首カメラ映像から接触変形・6軸力・把持圧縮目標を推定する枠組みを提案し、物体把持タスクで高い成功率を達成した。

詳しい要約

1. どんなもの?

- ロボットの把持において、接触時の物理的相互作用を推定し、把持に依存した圧縮目標を選ぶための枠組み。 - ZeroTouch は tactile-supervised framework で、wrist RGB 観測、gripper state、局所重力方向から、密な接触変形、瞬間的な six-axis wrench、把持依存の desired compression target を予測する。 - tactile 測定は訓練時の privileged supervision としてのみ使用し、展開時には不要。

2. 先行研究と比べてどこがすごい?

- 従来は tactile sensors が直接的な相互作用測定を提供するが、展開時に専用ハードウェアを必要とした。 - ZeroTouch は tactile を訓練時のみの監督に使い、展開時は wrist RGB などから接触推定を行う点が新しい。 - 完全なアーキテクチャは、state-only baseline の normal-force MAE 2.017 N を 0.531 N に低減。 - 物理評価では、未見物体 95%、見た物体/未見把持 80%、内容/負荷シフト 90% の成功率。 - 同一評価プロトコルで OpenVLA は 25%、40%、55%、SmolVLA は 10%、25%、35% であり、比較優位を示す。

3. 技術・手法の肝は?

- wrist RGB 観測、gripper state、局所重力方向を入力とする。 - 出力は密な接触変形、瞬間的な six-axis wrench、把持依存の desired compression target。 - tactile 測定を訓練時の privileged supervision として利用し、展開時には tactile を必要としない。 - 完全なアーキテクチャの詳細な構成は要旨からは不明。

4. どうやって有効だと検証した?

- 完全な validation set で normal-force MAE を評価し、state-only baseline の 2.017 N から 0.531 N への低減を確認。 - 物理評価を 20 trials per condition で実施。 - 条件は unseen object、seen-object/unseen-grasp、content/load shift。 - 成功率はそれぞれ 95%、80%、90%。 - 同一プロトコルで OpenVLA と SmolVLA と比較。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、tactile supervision の一般化可能性、計算コストなどについての議論は要旨に記載がない。

6. 次に読むべき論文は?

- OpenVLA - SmolVLA - 関連手法として tactile-supervised learning、visual contact estimation、contact-rich manipulation の定番研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dmitriy Kosenkov, Daniia Zinniatullina, Miguel Altamirano Cabrera, Iana Zhura, Mikhail Derevianchenko, Dzmitry Tsetserukou

分類: cs.RO

原文アブストラクト

Reliable robotic grasping benefits from estimating the evolving physical interaction and selecting a grasp-dependent compression target. Tactile sensors provide direct interaction measurements but require dedicated hardware at deployment. We introduce ZeroTouch, a tactile-supervised framework that predicts dense contact deformation, the instantaneous six-axis wrench, and a grasp-dependent desired compression target from wrist RGB observations, gripper state, and local gravity direction. Tactile measurements are used only as privileged supervision during training and are not required at deployment. On the full validation set, the complete architecture reduces normal-force MAE from 2.017 N for a state-only baseline to 0.531 N. In physical evaluation with 20 trials per condition, ZeroTouch achieves 95% success on an unseen object, 80% in a seen-object/unseen-grasp condition, and 90% under a content/load shift. Under the same evaluation protocol, OpenVLA achieves 25%, 40%, and 55%, while SmolVLA achieves 10%, 25%, and 35%, respectively.

関連論文

PR本紙発行元 EmplifAI