日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24124

ActiveArena:ロボットマニピュレーションにおける能動的知覚のベンチマークと理解

ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

能動的知覚と操作を評価するためのシミュレータと35タスクのベンチマーク、および13種のVLA構成を提案し、記憶管理やOOD汎化の要因を分析した。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションにおけるactive perceptionを評価するためのベンチマークとシミュレータ - ActiveArena-Sim: 視点制御可能で大規模ワークスペースを持つactive-perceptionシミュレータ - ActiveArena-Bench: 5つの細粒度カテゴリにわたる35タスク。視覚探索と対話的情報獲得をカバー - 各タスクは受動的観察だけでは難しく、多ラウンドの証拠獲得と記憶ベースの推論を要求 - 豊富なmemory annotations、標準化訓練データ、ID/OODプロトコル(disjoint scenes、unseen distractor configurations、novel backgrounds)を提供 - ActiveArena-VLA: memory writing、memory capacity、proprioceptive state、subtask supervision、high-level planningを制御研究するための13のvision-language-action構成のモジュールスイート

2. 先行研究と比べてどこがすごい?

- 既存ベンチマークは、ロボットが能動的に情報を獲得し記憶に保持する能力を評価するのが困難 - ActiveArenaは、制御可能な視点と大規模ワークスペースを備えたシミュレータを基盤とし、active perceptionに特化 - 受動的観察では解けない多ラウンド証拠獲得と記憶ベース推論を要求するタスク設計 - ID/OODプロトコルを備え、OOD汎化を評価可能 - 13のVLA構成を提供し、記憶書き込み・容量・固有受容感覚・サブタスク監督・高レベル計画を制御研究可能

3. 技術・手法の肝は?

- ActiveArena-Sim上に構築されたActiveArena-Bench: 35タスク、5カテゴリ - 各タスクは多ラウンドの証拠獲得と記憶ベース推論を必要とする - 豊富なmemory annotationsと標準化訓練データを提供 - ID/OODプロトコル: disjoint scenes、unseen distractor configurations、novel backgrounds - ActiveArena-VLA: 13のvision-language-action構成で、memory writing、memory capacity、proprioceptive state、subtask supervision、high-level planningを制御 - ベンチマーク結果から、uniform memory sampling、信頼できるwrite policies下でのmemory capacity増加、proprioceptive inputs、subtask supervisionがOOD汎化を改善 - planner-guided memory…

4. どうやって有効だと検証した?

- ActiveArena-Benchを用いたベンチマーク評価を実施 - 結果、ID-OODギャップが大きいことを明らかにした - uniform memory sampling、信頼できるwrite policies下でのmemory capacity増加、proprioceptive inputs、subtask supervisionがOOD汎化を改善することを示した - planner-guided memory managementとdecision-makingが、sparse memoryのみで最高性能構成に近い性能を達成することを示した

5. 議論はある?

- ベンチマーク結果からID-OODギャップが大きいことが判明 - uniform memory sampling、memory capacity増加(信頼できるwrite policies下)、proprioceptive inputs、subtask supervisionがOOD汎化を改善 - planner-guided memory managementとdecision-makingはsparse memoryで高性能 - これらの知見はactive perceptionとmanipulationのモデル開発と診断に有用 - 具体的な議論や限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の関連手法として、vision-language-action models、active perception、memory-based reasoning、OOD generalization in roboticsが挙げられる - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng

分類: cs.RO, cs.AI

原文アブストラクト

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.

関連論文

PR本紙発行元 EmplifAI