時間に根ざした操作:ロボットマニピュレーションにおける時間的グラウンディングのためのマルチソースデータセットとベンチマーク
Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation
ロボット操作において過去の出来事に基づく意思決定を評価するためのデータセットとベンチマークGiTを提案し、実世界とシミュレーションの両方で履歴依存タスクにおけるVLAモデルの性能を評価した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yi Wang, Yang Yang, Guangqi Xu, Sumin Lin, Ning Kang, Pengxiang Lu, Xiaotong Chen, Zeyu Xue, Ping Deng, Xing Liu, Chenguang Yang, Zhenyu Lu
分類: cs.RO, cs.CV
原文アブストラクト
Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.