レジスタ経由遅延融合:視覚運動模倣におけるショートカットしやすい観測融合の再配線
Register-Routed Delayed Fusion: Rewiring Shortcut-Prone Observation Fusion in Visuomotor Imitation
視覚と言語・固有感覚などのコンパクト信号の融合経路を制御するRRDFを提案し、レジスタを介した遅延融合で視覚応答性と模倣性能を改善した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jieting Long, Weidong Cai, Weiming Zhi
分類: cs.RO
原文アブストラクト
Visuomotor imitation policies combine high-dimensional visual observations with compact signals such as proprioception, and their fusion topology determines when and through which tokens these streams interact. In dense token fusion, visual tokens may attend directly to compact tokens from the first encoder layer, allowing action-predictive compact cues to influence spatial visual representations early in their formation. We ask whether controlling this route improves visual responsiveness and policy behavior. We introduce Register-Routed Delayed Fusion (RRDF), which masks direct compact-visual attention and stages cross-modal interaction through a learned register workspace. Its isolate-collect-route schedule protects an early stream-separated prefix and later permits only register-mediated exchange, while compact conditioning remains available to the native action generator. Across five simulation tasks and three real-robot tasks, RRDF matches or improves dense ACT under nominal conditions. Appearance-shift evaluations on four simulation tasks and held-out-position evaluations on three real-robot tasks also favor RRDF. Phase-matched input probes show lower measured state-to-image sensitivity, while ablations indicate that adding registers alone does not reproduce the full performance gain. These results support controlling cross-modal propagation while retaining compact action conditioning.