日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.07946

視覚外乱に適応するVLAモデルの自己教師ありテスト時適応

Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution

シェア:XThreadsFacebookLINEはてブBluesky

ロボット実行中に未知の視覚外乱が起きても、直前の未実行行動チャンクを自己教師信号としてVLA方策をテスト時に適応させ、成功率を向上させる手法SALTを提案。

詳しい要約

1. どんなもの?

- ロボット実行中に発生する未知の視覚的disruptionに対応するtest-time adaptation手法SALTを提案。 - vision-language-action (VLA) policyがdisruptionの種類やタイミングを知らずに応答する問題を扱う。 - leftover trajectory(前回のaction chunkの未実行部分)を自己教師信号として利用。 - 連続するchunkが時間的に重なる性質を利用し、同一未来制御区間の時間的に整合したtargetを得る。 - disruption注釈・expert action・target-domain demonstrationを必要としない。

2. 先行研究と比べてどこがすごい?

- 従来のtest-time adaptationはdisruption注釈やexpert action、target-domain demonstrationを要することが多いが、SALTはpolicy自身の予測のみからsupervisionを得る。 - 未知のdisruption種別・タイミングに対し、注釈なしで適応できる点が新しい。 - Transition Anchoringによりshift前の計画を保持し、適応をshift全体に固定できる。 - Sequential Correction Propagationで補正を実行trajectoryに沿って伝播させる。 - 具体的な先行研究名は要旨からは不明。

3. 技術・手法の肝は?

- leftover trajectoryを自己教師信号としてtest-time adaptationに用いる。 - 連続chunkの時間的重なりを利用し、現在予測と同じ未来制御区間のtargetをleftoverから得る。 - Transition Anchoring: visual shift開始時にleftoverがcorruption前の計画を保持し、それへpolicyを更新して適応を固定。 - Sequential Correction Propagation: 適応後policyを保持し現在chunkを再生成、そのleftoverが次回replanのtargetとなり補正を伝播。 - lightweight adaptation gateをnominal trajectoryのみでcalibrateし、更新開始を判断。

4. どうやって有効だと検証した?

- LIBERO-10で5種類のpersistent visual corruptionに対し評価。 - SmolVLAで平均successが43.9%から53.2%へ向上。 - GR00T N1.7で58.7%から66.0%へ向上。 - nominal performanceをほぼ維持。 - 実機でdigitalおよびphysical disruption平均のtask progressが0.49から0.61へ向上。

5. 議論はある?

- 要旨からは不明。 - 想定される論点として、adaptation gateのcalibrationやnominal性能保持、実機での汎化性が考えられるが、要旨に明記なし。

6. 次に読むべき論文は?

- SmolVLA - GR00T N1.7 - LIBERO-10 - test-time adaptation - vision-language-action (VLA) policy

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ahin Lee, Jinwoo Seo, Youngsoo Jang, Taesik Gong

分類: cs.RO, cs.AI

原文アブストラクト

Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.

関連論文

PR本紙発行元 EmplifAI