日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35003

視覚中断下での行動を視覚言語行動モデルで学習する

Learning to Act under Visual Interruptions with Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

カメラ映像が途切れてもロボットが作業を続けられるよう、視覚言語行動モデルを訓練し、推論時にオプティカルフローや世界モデルで欠損視覚を補完する手法MINTを提案し、ベンチマークMAIL-Benchで評価した。

詳しい要約

1. どんなもの?

- VLAモデルはロボット操作で高い能力を示すが、通常はタスク実行中に全カメラストリームが利用可能な前提で開発・評価される。 - カメラがフレーム供給を停止した場合、ポリシーは欠損ビューの以降の観測なしで行動を続ける必要がある。 - 本研究はこの視覚中断がclosed-loop manipulationに与える影響を調べる。 - MAIL-Benchというベンチマークを導入し、VLAモデルでの視覚中断を評価する。 - さらにMINTを提案し、視覚入力欠損下でも機能するようVLAポリシーを訓練する。

2. 先行研究と比べてどこがすごい?

- 従来のVLAモデルは全カメラストリームが利用可能な前提で開発・評価されており、カメラ停止時の影響は十分に理解されていない。 - 本研究は視覚中断下でのclosed-loop manipulationを体系的に評価するMAIL-Benchを導入した点が新しい。 - また、欠損視覚入力下でも機能するよう訓練し、推論時に欠損観測を補完するMINTを提案。 - π_{0.5}やGR00T N1.5でカメラ損失下のタスク成功率を大幅に改善。 - AgiBot G2での実機展開も示す。

3. 技術・手法の肝は?

- MAIL-Bench:各ポリシーの成功参照軌道の複数段階で異なるカメラを中断し、視覚入力が利用不能になった際の能力保持を測定。 - MINT:まず視覚入力欠損下でも機能するようVLAポリシーを訓練。 - 推論時、optical-flow extrapolationまたはaction-conditioned world modelを用いて欠損観測を選択的に補完。 - 予測ビューが信頼できない場合は撤回する。

4. どうやって有効だと検証した?

- π_{0.5}とGR00T N1.5での実験により、MINTが元モデルよりカメラ損失下のタスク成功率を大幅に改善することを示す。 - AgiBot G2での実験により、カメラ損失下の実機展開も実証。 - ベンチマークは https://minglejiang.github.io/Mail-Bench/ で公開。

5. 議論はある?

- 視覚中断がclosed-loop manipulationに与える影響は実用的に重要だが、十分に理解されていない。 - カメラ停止時のポリシー挙動や、欠損観測補完の信頼性判断に関する議論が含まれる可能性があるが、要旨からは詳細不明。 - 具体的な限界や失敗事例、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- π_{0.5} - GR00T N1.5 - AgiBot G2 - optical-flow extrapolation - action-conditioned world model - VLAモデル全般(関連手法として)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingle Jiang, Rui Xu, Yunke Wang, Chang Xu

分類: cs.RO, cs.AI

原文アブストラクト

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $π_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/

関連論文

PR本紙発行元 EmplifAI