日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.23580

TaskAnchor: 長期的マニピュレーションのための反応型VLAにおけるタスク状態の接地

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

視覚的に似た観測でも実行段階によって行動が異なる「タスク状態の曖昧さ」を解消するため、実行履歴に基づく軽量アダプタTaskAnchorを提案し、VLAの成功率を大幅に向上させた。

著者: Hengyan Liu, Wenlve Zhou, Bo Yue, Yongyi Su, Ruixiang Wang, Zhanqi Zhang, Dekun Lu, Wei Gao, Xiaofen Xing, Kui Jia

分類: cs.RO

原文アブストラクト

Reactive vision--language--action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9--5.5$\times$ the average success rates of the published $π_{0.5}$ and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08\,ms per action chunk for $π_{0.5}$.

関連論文

PR本紙発行元 EmplifAI