日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマノイド/ツール使用arXiv:2610.02089

ヒューマノイドツールベンチ:ツール選択から移動を伴う実行までを評価

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

シェア:XThreadsFacebookLINEはてブBluesky

ヒューマノイドが道具を選び、操作と移動を組み合わせてタスクを遂行する能力を評価する18タスクのベンチマークと3.1kデモのデータセットを提案し、実機・シミュレーションで既存手法の課題を明らかにした。

詳しい要約

1. どんなもの?

- ヒューマノイドによる tool use を評価する benchmark - 名称は HumanoidToolBench - 18 タスク、3 シナリオ、3 実行レベル、2 tool-set モード - 付随 dataset は ToolBook - simulation と real Unitree G1 で収集した 3.1k demonstrations - 対象は tool の選択から mobile execution まで

2. 先行研究と比べてどこがすごい?

- 既存 benchmark は humanoid 上の tool use 能力を同時評価しない - tool 選択と manipulation、必要時の locomotion の協調をまとめて見ない - 本 benchmark はこれらを統合的に評価 - 18 タスク、3 シナリオ、3 実行レベル、2 tool-set モードで構成 - simulation と real robot の両方を含む点が特徴

3. 技術・手法の肝は?

- benchmark 設計 - 3 シナリオ、3 実行レベル、2 tool-set モードの 18 タスク - dataset 構築 - ToolBook として simulation と real Unitree G1 で 3.1k demonstrations を収集 - 評価 - simulation で 7 policies、real robot で 3 policies を評価 - GR00T N1.7 の focused probes も実施

4. どうやって有効だと検証した?

- simulation で 7 policies を評価 - real robot で 3 policies を評価 - 結果として tool 選択とタスク完了の間に大きな gap を確認 - GR00T N1.7 probes では - unseen tools で選択精度が低下 - 無関係な指示でもタスク実行が継続

5. 議論はある?

- tool 選択とタスク完了の間に substantial gaps がある - GR00T N1.7 は unseen tools で選択精度が低下 - 無関係な指示でも実行を続ける傾向 - その他の議論や限界は要旨からは不明

6. 次に読むべき論文は?

- GR00T N1.7 - Unitree G1 - その他は要旨で参照/比較されている研究が明示されていないため、同分野の定番として humanoid manipulation や tool use の benchmark 研究を挙げる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, Idan Szpektor, Avi Caciularu, Jongmin Park, Youngjae Yu

分類: cs.RO, cs.AI

原文アブストラクト

As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.

PR本紙発行元 EmplifAI