日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.33707

敵対的学習はマルチビューVLAの汎化を改善するか?ビュー崩壊の解明と緩和

Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse

シェア:XThreadsFacebookLINEはてブBluesky

マルチビューVLAに敵対的学習を適用すると、特定の視点変化には強くなる一方で「ビュー崩壊」と呼ばれる現象が起こり、手首視点に過度に依存することを明らかにし、ビュースワップ介入で緩和を試みた研究。

著者: Futa Waseda, Shuhei Kurita, Isao Echizen

分類: cs.AI, cs.RO

原文アブストラクト

Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.

関連論文

PR本紙発行元 EmplifAI