差動増幅器に着想を得たAmpAttentionによる多視点ロボット操作
Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation
多視点ロボット操作における注意機構のノイズを抑制するため、アナログ回路の差動増幅器に着想を得た新しい注意機構AmpAttentionを提案し、これを統合したRVAFモデルがRLBenchタスクで最高成功率を達成しつつ訓練時間を33.3%削減することを示した。
著者: Jin Yang, Ping Wei, Nanning Zheng
分類: cs.RO, cs.AI, cs.CV
原文アブストラクト
Multi-view robotic manipulation methods with the attention mechanism have recently achieved significant progress in both training efficiency and task performance. However, the inherent redundancy, occlusion, and viewpoint dependency in robotic view images often lead to severe attention drift. To address this challenge, we propose AmpAttention, a novel attention mechanism inspired by differential amplifiers in analog circuits. It aims to suppress attention noise and capture high signal-to-noise ratio signals for more reliable perception. Based on this, we introduce the RVAF model, which integrates task-guided intra-view and inter-view AmpAttention. Compared to previous state-of-the-art methods, RVAF achieves the optimal average success rate across 18 RLBench tasks (249 variations) while reducing training time by 33.3\%. RVAF also demonstrates strong potential in real-world high-precision tasks, exemplified by its ability to pick up a dart and accurately insert it into the red bullseye. Furthermore, we extend RVAF to RVAF++ by incorporating the SAM2 image encoder. RVAF++ achieves substantial gains on high-precision tasks, achieving a 91\% success rate on the `insert peg' task. More qualitative results are provided at the anonymous project website https://anonymous.4open.science/w/RVAF-Anonymization.
関連論文
- Peg-in-Bench: 高精度ロボット挿入のためのモジュール式ベンチマークマニピュレーション
- 否定制約付き器用把持のためのポテンシャル誘導粒子ステアリングマニピュレーション
- Facet-0: 接触を伴う精密操作のためのロボット基盤モデルマニピュレーション
- Motus2: 巧みな操作のための自己進化型汎用世界モデルマニピュレーション
- Zeva: 文脈内因果学習による汎用身体操作の実現マニピュレーション
- SUN: 言語に基づく制御から学習、実機への永続的プログラムマニピュレーション