日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09462

TMT: 未見タスクにおける視覚言語行動ポリシーの実行時バックドア検出

TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks

シェア:XThreadsFacebookLINEはてブBluesky

良性ロールアウトで訓練したトークン多様体と潜在遷移モデルにより、VLAポリシーのバックドア起動を実行時に検出し、自己蒸留による浄化も探索する手法を提案。

詳しい要約

1. どんなもの?

VLA policy に対する runtime backdoor detector「TMT」を提案。benign rollout のみで学習し、unseen task でも backdoor を検出する。 - Token Manifold と latent Transition modeling の2分岐で構成。 - 入力トークン構造と隣接層 latent dynamics の予測誤差を評価。 - 疑わしい rollout を token manifold で特定し、latent 偏差で確認後、transition selection を更新。 - さらに self-distillation による policy purification も探索。

2. 先行研究と比べてどこがすごい?

従来の backdoor detector を VLA 向けに適応し、anomaly/failure detection 手法も VLA backdoor detector として再目的化して比較。 - 10 baselines との post-hoc 比較で、3つの VLA backdoor attack に対し unseen task で state-of-the-art の検出性能を達成。 - 個々には plausible な action からなる malicious behavior や、unfamiliar task による正当な変化を区別できる点が優位。

3. 技術・手法の肝は?

benign rollout のみで学習する2分岐 detector。 - Token Manifold 分岐:入力トークン構造を評価し suspicious rollout を特定。 - Latent Transition 分岐:隣接層 latent dynamics の prediction error を評価し、token 分岐の疑いを確認。 - 確認後、transition selection を更新して後続 monitoring を導く。 - Purification:frozen backdoored policy が benign-input action を提供し、benign/triggered 観測ペアで student を self-distillation 監督。clean reference policy 不要。

4. どうやって有効だと検証した?

3つの VLA backdoor attack に対し unseen task で評価。 - 従来 backdoor detector を適応し、anomaly/failure detection 手法を VLA backdoor detector として再目的化した10 baselines と post-hoc 比較。 - TMT が state-of-the-art の backdoor detection 性能を達成。 - 詳細な実験設定・指標は要旨からは不明。

5. 議論はある?

要旨からは不明。 - 想定される議論:unseen task での汎化、benign rollout のみでの学習限界、self-distillation による purification の有効性と副作用、計算コスト、他 attack への頑健性など。 - これらは要旨に明記されていない。

6. 次に読むべき論文は?

要旨で参照/比較されている研究:traditional backdoor detectors、anomaly detection methods、failure detection methods、VLA backdoor attacks。 - 関連手法:Token Manifold、latent Transition modeling、self-distillation。 - 同分野の定番:VLA policy、backdoor attack/defense、runtime detection。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zirun Zhou, Jingfeng Zhang, HaoChuan Xu, Xizhe Zhang, Elliott Wen, Jing Sun, Hong Jia

分類: cs.RO

原文アブストラクト

Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears. Detecting such activation is difficult because malicious behavior can comprise individually plausible actions, while unfamiliar tasks introduce legitimate changes in observations and behavior. We introduce TMT, a runtime backdoor detector based on Token Manifold and latent Transition modeling. Trained on benign rollouts, its two branches assess input-token structure and prediction errors in adjacent-layer latent dynamics. A suspicious rollout identified by the token manifold branch, once confirmed through latent deviations, guides transition selection for subsequent monitoring. We further explore policy purification through self-distillation: a frozen copy of the backdoored policy provides benign-input actions to supervise a student on paired benign and triggered observations, without requiring a separate clean reference policy. For evaluation, we adapt traditional backdoor detectors and repurpose anomaly and failure detection methods as VLA backdoor detectors. In a post-hoc comparison with ten baselines, TMT achieves state-of-the-art backdoor detection performance on unseen tasks across three VLA backdoor attacks. Our project page is available at https://zzr42.github.io/tmt/.

関連論文

PR本紙発行元 EmplifAI