日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02813

レジスタ経由遅延融合:視覚運動模倣におけるショートカットしやすい観測融合の再配線

Register-Routed Delayed Fusion: Rewiring Shortcut-Prone Observation Fusion in Visuomotor Imitation

シェア:XThreadsFacebookLINEはてブBluesky

視覚と言語・固有感覚などのコンパクト信号の融合経路を制御するRRDFを提案し、レジスタを介した遅延融合で視覚応答性と模倣性能を改善した。

詳しい要約

1. どんなもの?

視覚運動模倣ポリシーにおける視覚と固有感覚などのコンパクト信号の融合トポロジーを制御する手法。Register-Routed Delayed Fusion (RRDF) を提案し、視覚トークンが初期層からコンパクトトークンに直接注意するのを防ぎ、学習された register workspace を介した遅延融合を行う。

2. 先行研究と比べてどこがすごい?

dense token fusion を用いる ACT と比較し、名目条件下で同等以上、appearance-shift や held-out-position 評価でも優位。直接的な compact-visual attention をマスクし、register を介した経路に置き換える点が新しい。

3. 技術・手法の肝は?

isolate-collect-route スケジュールを採用。初期はストリーム分離を保ち、後段で register を介した交換のみ許可。コンパクト条件付けはネイティブの action generator に利用可能なまま。

4. どうやって有効だと検証した?

5つのシミュレーションタスクと3つの実ロボットタスクで評価。4つのシミュレーションタスクで appearance-shift、3つの実ロボットタスクで held-out-position 評価を実施。phase-matched input probes で state-to-image 感度の低下を確認。ablation で register 追加だけでは性能向上が再現しないことを示す。

5. 議論はある?

cross-modal propagation を制御しつつコンパクトな action conditioning を保持することの有効性を示唆。ただし、register 単独では不十分であり、スケジュール全体の寄与が重要。

6. 次に読むべき論文は?

ACT (dense token fusion) や register を用いた手法、cross-modal attention 制御に関する研究。要旨で参照されている具体的な論文名は不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jieting Long, Weidong Cai, Weiming Zhi

分類: cs.RO

原文アブストラクト

Visuomotor imitation policies combine high-dimensional visual observations with compact signals such as proprioception, and their fusion topology determines when and through which tokens these streams interact. In dense token fusion, visual tokens may attend directly to compact tokens from the first encoder layer, allowing action-predictive compact cues to influence spatial visual representations early in their formation. We ask whether controlling this route improves visual responsiveness and policy behavior. We introduce Register-Routed Delayed Fusion (RRDF), which masks direct compact-visual attention and stages cross-modal interaction through a learned register workspace. Its isolate-collect-route schedule protects an early stream-separated prefix and later permits only register-mediated exchange, while compact conditioning remains available to the native action generator. Across five simulation tasks and three real-robot tasks, RRDF matches or improves dense ACT under nominal conditions. Appearance-shift evaluations on four simulation tasks and held-out-position evaluations on three real-robot tasks also favor RRDF. Phase-matched input probes show lower measured state-to-image sensitivity, while ablations indicate that adding registers alone does not reproduce the full performance gain. These results support controlling cross-modal propagation while retaining compact action conditioning.

関連論文

PR本紙発行元 EmplifAI