日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04255

時間に根ざした操作:ロボットマニピュレーションにおける時間的グラウンディングのためのマルチソースデータセットとベンチマーク

Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作において過去の出来事に基づく意思決定を評価するためのデータセットとベンチマークGiTを提案し、実世界とシミュレーションの両方で履歴依存タスクにおけるVLAモデルの性能を評価した。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションにおける時間的groundingのためのデータセットとベンチマーク - 名称: GiT (Grounded in Time) - 対象: biolaboratory, household, industrialシナリオ - 実ロボットとUniversal Manipulation Interface (UMI)スタイルのデモを含む - 18のbimanualタスクをカバー - シミュレーションデータとManiSkillベースの評価スイート(9タスク)も含む - 細かいサブタスクアノテーションとcounterfactualタスクペアを提供

2. 先行研究と比べてどこがすごい?

- 従来のmemory-augmented vision-language-action (VLA)モデルのベンチマークでは、履歴依存の意味推論を要する応用指向タスクが不足していた - GiTはそのようなタスクを明示的に含む点で先行研究と異なる - 実世界とシミュレーションの両方で評価可能な点が特徴 - 要旨からは他の具体的な先行研究との比較は不明

3. 技術・手法の肝は?

- データセット構築: 実ロボットとUMIスタイルのデモ、シミュレーションデータを収集 - タスク設計: 現在の観測だけでは不十分で過去のイベントに基づく推論が必要なタスク - アノテーション: 細かいサブタスクアノテーションとcounterfactualタスクペア(類似の現在観測でも過去のイベントにより異なる行動を要求) - 評価: ManiSkillベースの評価スイートを提供

4. どうやって有効だと検証した?

- シミュレーションと選択された実世界タスクで代表的なend-to-end VLAモデルを評価 - 履歴依存のマニピュレーションにおいて改善の余地が大きいことを示した - 具体的な評価指標や結果の詳細は要旨からは不明

5. 議論はある?

- 履歴依存のマニピュレーションにおける課題を提起 - 代表的なVLAモデルが十分な性能を発揮できないことを示唆 - データセットとベンチマークの公開により今後の研究を促進 - 具体的な議論や限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてmemory-augmented vision-language-action (VLA)モデル、Universal Manipulation Interface (UMI)、ManiSkillが挙げられる - 同分野の定番としてvision-language-action (VLA)モデル全般が次に読むべき候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yi Wang, Yang Yang, Guangqi Xu, Sumin Lin, Ning Kang, Pengxiang Lu, Xiaotong Chen, Zeyu Xue, Ping Deng, Xing Liu, Chenguang Yang, Zhenyu Lu

分類: cs.RO, cs.CV

原文アブストラクト

Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.

関連論文

PR本紙発行元 EmplifAI