日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.30959

VisTacAlign: 触覚付き人間・ロボット実演による巧みな操作方策の共訓練

VisTacAlign: Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations

シェア:XThreadsFacebookLINEはてブBluesky

人間の手の動きと触覚をロボットハンドにリターゲット・信号空間で整合させ、視覚・触覚・固有感覚を入力とする拡散トランスフォーマ方策を人間とロボットの実演で共訓練する枠組みを提案。

詳しい要約

1. どんなもの?

- 人間とロボットの実演データを共同訓練する枠組み VisTacAlign を提案。 - 3D視覚触覚の巧みな操作ポリシーを学習。 - 17-DoF 触覚ロボットハンドを使用。 - 3つの実世界タスク(Lego組立、イチゴ摘み、電動ドリル操作)で評価。

2. 先行研究と比べてどこがすごい?

- 人間とロボットのギャップを視覚・触覚の両モダリティで埋める。 - 人間の手を消去しロボットハンドメッシュに置換、ステレオ基盤モデルを再実行し点群の誤差と可視性を一致させる。 - 触覚グローブをロボットの触覚センサに信号空間で整列。 - ロボットのみのポリシーより性能向上。

3. 技術・手法の肝は?

- グローブ追跡した人間の手の動きを17-DoF触覚ロボットハンドにリターゲット(一度の指先補正)。 - ステレオビューから人間の手を消去し、ロボット記録のピクセルで彩色したロボットハンドメッシュを配置。 - 合成画像にリアルタイムステレオ基盤モデルを再実行し、人間の点群にロボットと同じステレオ誤差と可視性を持たせる。 - 容量性触覚グローブをロボットの指先センサに信号空間で整列し、解釈可能な指ごとの力表現を得る。 - Diffusion Transformer が点群、固有受容、指ごとの触覚トークンを消費。

4. どうやって有効だと検証した?

- 3つの実世界タスク(Lego組立、サイズの異なるイチゴ摘み、電動ドリルの起動と持ち上げ)で検証。 - 既存のロボットデータに整列した人間の実演を追加すると、ロボットのみのポリシーより改善。 - アブレーションにより触覚入力と視覚整列の両方が必要であることを示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、ステレオ基盤モデル、Diffusion Transformer、触覚グローブ、リターゲティングなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Julien Poffet, Matthew Strong, Ankush Dhawan, Baiyu Shi, Shalika Neelaveni, Yujia Yuan, Zhenan Bao, Monroe Kennedy

分類: cs.RO

原文アブストラクト

Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human--robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot's fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force -- Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill -- adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary. Project page: https://vis-tac-align.github.io

関連論文

PR本紙発行元 EmplifAI