日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37165

ドメイン不変潜在先読みによるVLAモデルの偽相関の解消

Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語行動モデルが視覚分布シフトで脆くなる原因である偽相関を、ドメイン変換軌道から学んだドメイン不変の未来潜在表現で方策を監督することで軽減するDILLを提案し、LIBERO-Plusで平均成功率69.1%を達成した。

詳しい要約

1. どんなもの?

Vision-Language-Action (VLA) モデルの視覚的分布シフトに対する脆弱性を、spurious correlation(ショートカット学習)の観点から緩和する表現学習フレームワーク Domain-Invariant Latent Lookahead (DILL) を提案する研究。 - 目的: VLA ポリシーがドメイン固有の視覚的要因に依存せず、タスク関連構造に基づくようにする。 - 構成: Task-Domain Encoder と lookahead prediction、domain disentanglement を組み合わせる。 - 評価: 反事実的タスク視点評価、LIBERO-Plus、実世界マニピュレーション。

2. 先行研究と比べてどこがすごい?

従来の VLA モデルは視覚分布シフト下で brittleness を示し、タスク関連構造ではなくドメイン固有要因に結びついた spurious correlation に依存しがちだった。 - DILL はドメイン変換済み軌道データから domain-invariant な future latent を学習し、ポリシーを監督する点が新しい。 - 結果として LIBERO-Plus で平均成功率 69.1%、最強ベースラインを 11.4 ポイント上回る。 - 反事実的タスク視点評価でショートカット依存の低減を示す。

3. 技術・手法の肝は?

中核は domain-invariant future latent を予測・活用する表現学習。 - Task-Domain Encoder を contrastive objective と Gaussian disentanglement regularization で訓練し、タスク関連構造とドメイン固有視覚変動を分離。 - 学習済みエンコーダが future latent を提供し、lookahead prediction と domain disentanglement を通じて VLA ポリシー学習を監督。 - これによりポリシーが偶発的な視覚要因ではなくタスク関連構造に注目するよう促す。

4. どうやって有効だと検証した?

複数の評価で有効性を検証。 - Counterfactual task-view evaluations: ショートカット依存の低減を確認。 - LIBERO-Plus: 視覚的ロバスト性が向上し、平均成功率 69.1%(最強ベースライン比 +11.4 ポイント)。 - 実世界マニピュレーション実験: 制御されたシミュレーション外でも適用可能性を支持。 - 潜在空間診断: タスク整合構造を保持しつつドメイン固有変動を抑制する表現が、行動改善と相伴うことを示す。

5. 議論はある?

要旨からは、限界や失敗事例、計算コスト、ハイパーパラメータ感度などの詳細な議論は不明。 - 主張されているのは、ショートカット依存低減、視覚ロバスト性向上、実世界適用可能性、潜在空間診断との整合。 - これらを踏まえた一般的な議論(例: ドメイン変換の設計依存性、他タスクへの一般化)は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている具体的な先行研究名は明記されていない。 - 関連手法として Vision-Language-Action (VLA) モデル、contrastive learning、Gaussian disentanglement regularization、lookahead prediction が挙げられる。 - ベンチマークとして LIBERO-Plus が用いられている。 - 同分野の定番として、VLA ポリシー学習や視覚的ロバスト性に関する研究を次に読む候補として挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang

分類: cs.RO, cs.LG

原文アブストラクト

Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.

関連論文

PR本紙発行元 EmplifAI