TACIT: 巧みな操作における空間的注意のための触覚接触スーパービジョン
TACIT: Tactile Contact Supervision for Spatial Attention in Dexterous Manipulation
テレオペレーションの触覚接触データを用いて空間的注意を教師なしで学習し、視触覚拡散ポリシーに組み込むことで、物体配置やペグ挿入タスクの成功率を大幅に向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yanhou Lai, Fucai Zhu, Ruiqiang Wang, Koichi Hashimoto
分類: cs.RO
原文アブストラクト
Visuomotor policies trained from a few demonstrations may reproduce demonstrated trajectories without reliably following changes in object position. Existing approaches with explicit attention typically obtain spatial priors from human annotation or visual models. We introduce TACIT (tactile contact informs attention), which uses measured tactile contacts from teleoperated demonstrations to supervise spatial attention without additional point annotation. Gaussian targets over preceding camera point clouds supervise an attention head whose pooled output conditions a visuotactile diffusion policy. Targets are used only during training; tactile observations remain inputs at inference. In the primary real-robot benchmark, with ten demonstrations per task and five demonstrated placement regions, TACIT achieves 66.7% success on ball placement and 73.3% on peg insertion, compared with 10.0% and 20.0% for input-matched 3D visuotactile fusion and 20.0% and 43.3% for vision-only DP3. TACIT enters the 150 mm palm-to-object approach region within 12 seconds in all 30 trials per task; all remaining failures occur after arrival. Across three training seeds on real ball and simulated peg, TACIT outperforms input-matched fusion and an architecture-matched control without explicit attention supervision, supporting the contribution of supervision beyond branch capacity. Pre-contact and contact-time supervision show no consistent ordering. These results demonstrate that measured tactile contact provides effective spatial supervision for approach behavior from few demonstrations within the evaluated workspace.
関連論文
- DexTacWAM: 巧みな操作のための視触覚ワールドアクションモデルマニピュレーション
- 平面果樹園における視覚運動ロボット剪定のためのハイブリッド強化学習マニピュレーション
- CAST: 衝突を考慮した建設ロボットによる同時軌道推定と計画マニピュレーション
- ロボット構成空間における異種制約のための実行可能性距離場マニピュレーション
- InsertAnything: シミュレーションから現実への汎化可能な接触リッチ精密挿入マニピュレーション
- フレーズ単位のロボット古琴演奏:両腕動作計画と音触覚インタラクション監視マニピュレーション