RoboSPA: VLAモデルは単純なシーンと短期的なタスクを超えられるか?
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
VLAモデルの空間的・手続き的複雑さに対する推論能力を診断する大規模ロボット操作ベンチマークRoboSPAを提案し、既存モデルの限界を明らかにした。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
分類: cs.RO, cs.AI, cs.CV
原文アブストラクト
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.
関連論文
- Behavior-Skill: 長期的タスクにおける視覚-言語-行動ポリシー評価のための細粒度ベンチマークVLA/ベンチマーク
- InstructMove: 指示追従操作のためのテキスト必須ベンチマークVLA/ベンチマーク
- NumerosityVLM: 視覚言語モデルにおける数量表現の解釈のための認知に着想を得たベンチマークVLA/ベンチマーク
- SO-101における視覚言語行動モデルのベンチマーク:失敗と回復の分析VLA/ベンチマーク
- RoboProcessBench: 視覚言語ロボット操作におけるプロセス認識理解のベンチマークVLA/ベンチマーク
- EPIC-Bench: 視覚言語モデルにおける細粒度な身体化視覚グラウンディングのための知覚中心ベンチマークVLA/ベンチマーク