日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.25636

RoboFollow:身体性エージェントにおける指示追従の幻想を暴く

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

シェア:XThreadsFacebookLINEはてブBluesky

視覚シーンが一つのタスクしか許さない「低シーンエントロピー」問題を指摘し、言語理解を厳密に診断する4段階ベンチマークRoboFollowを提案。9つのVLA/WAMポリシーを評価し、指示追従能力が実は脆弱であることを示した。

詳しい要約

1. どんなもの?

- 現代の embodied agents は高い成功率を示すが、実際の instruction following 能力はその数値ほど強くないという錯覚を暴く研究。 - 錯覚の原因を low scene entropy という構造的性質に帰着。 - 視覚 scene が一つの有効タスクしか許さない場合、language が冗長になり、policy が language をほとんど使わず高スコアを出せる。 - これを診断する benchmark「RoboFollow」を提案。 - 3原則: High Scene Entropy、Hierarchical Diagnostic Protocol (L0–L3)、Confound-Controlled Diagnosis。 - 9つの VLA および WAM policies を評価し、instruction following が見過ごされたボトルネックであることを示す。

2. 先行研究と比べてどこがすごい?

- 従来の embodied agents 評価は成功率に依存し、language の寄与を分離できていなかった。 - RoboFollow は low scene entropy という視点を導入し、vision だけでは解けない高エントロピー scene を設計。 - これにより language への依存を強制し、instruction following を直接診断。 - 4段階の L0–L3 protocol で視覚 layout と semantics を段階的に摂動し、等価な指示の一貫性と異なる指示の識別性を空間関係・属性・軌道制約・論理にわたり検証。 - 相互作用対象を単純化し、行動を訓練済みレパートリーに制限、Intent と Execution を段階別に報告することで、理解と運動実行を分離。 - 既存の緩和策 (stronger VLM backbones, QA co-training, LangForce, Classifier-Free Guidance) がこのギャップを埋められないことを示した点が新しい。

3. 技術・手法の肝は?

- High Scene Entropy: 各訓練 scene が複数の運動学的に異なるタスク分岐を支持するよう設計。 - vision だけでは不十分にし、language への依存を強制。 - Hierarchical Diagnostic Protocol: 4段階 (L0–L3) で視覚 layout と semantics を漸進的に摂動。 - 空間関係、属性、軌道制約、論理に関する指示の一貫性と識別性を調べる。 - Confound-Controlled Diagnosis: 相互作用対象を単純化し、行動を訓練済みレパートリーに制限。 - 段階別に Intent と Execution スコアを報告し、comprehension と motor execution を分離。 - 評価対象は 9つの VLA および WAM policies。

4. どうやって有効だと検証した?

- RoboFollow benchmark 上で 9つの VLA および WAM policies を評価。 - 強い L0 性能が、fine-tuning 設定下で L1–L3 に確実に転移しないことを示した。 - 代表的な緩和策 (stronger VLM backbones, QA co-training, LangForce, Classifier-Free Guidance) を試したが、いずれもこのギャップを埋められなかった。 - これにより instruction following が critical かつ overlooked なボトルネックであることを実証。

5. 議論はある?

- 現代の embodied agents の高い成功率は instruction following 能力を過大評価させる錯覚であると議論。 - low scene entropy がこの錯覚の構造的原因であり、language が冗長になることで policy が language を無視しても高スコアを出せる。 - 既存の緩和策では L0 から L1–L3 への性能ギャップを解消できないことを指摘。 - instruction following が見過ごされた重要ボトルネックであると結論。 - 具体的な限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: VLA policies, WAM policies, stronger VLM backbones, QA co-training, LangForce, Classifier-Free Guidance。 - 関連手法として、embodied instruction following や vision-language-action モデル、benchmark 設計に関する研究が次に読むべき候補。 - 具体的な論文名は要旨に明記されていないため、同分野の定番として VLA (Vision-Language-Action) モデルや instruction following benchmark に関する研究を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chang Guo, Yukun Xie, Bohan Tan, Zheng Chang, Zhaokai Yin, Qianli Ma, Yingqiao Wang, Chao Liang, Zhipeng Zhang

分類: cs.RO

原文アブストラクト

Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.

関連論文

PR本紙発行元 EmplifAI