FailPatch: 視覚言語行動モデルのための失敗残差パッチング
FailPatch: Failure Residual Patching for Vision-Language-Action Models
VLA方策の失敗しやすい状態を検出し、凍結した方策の隠れ表現に残差エキスパートを注入して成功率を改善する手法を提案。
著者: Peng Yu, Jiacheng Wang, Ziheng Zhang, Xuchong Zhang, Baoting Li, Zhuoyuan Yu, Yuxiang Chen, Tiancai Wang, Hongbin Sun
分類: cs.RO
原文アブストラクト
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.