日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
評価ベンチマークarXiv:2607.24481

ArmnetBench v0.1: 低コストアーム群での並列実世界評価による操作ポリシーのベンチマーク

ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

シェア:XThreadsFacebookLINEはてブBluesky

低コストなロボットアーム群を用いて、複数の操作ポリシーを実世界で並列評価するベンチマークを構築し、7つのポリシーを12タスクで比較した。

著者: Praveen Selvaraj, Lorenzo Uttini, Ville Kuosmanen

分類: cs.RO

原文アブストラクト

Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.

関連論文