日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマノイド操作arXiv:2610.07594

BiGym 2.0:ヒューマノイド家庭内操作のための学習済み・エージェント開発ポリシーのベンチマーク

BiGym 2.0: Benchmarking Learned and Agent-Developed Policies for Humanoid Household Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

Unitree G1ヒューマノイドによる20の家庭内タスクを対象に、全身制御器を用いたデモと評価を統合したベンチマークBiGym 2.0を構築し、VLA微調整・模倣学習・デモ駆動強化学習・コーディングエージェントを比較評価した。

詳しい要約

1. どんなもの?

- BiGym 2.0は、Unitree G1ヒューマノイド向けにBiGymを適応したベンチマーク。 - 20のhouseholdタスクを、統一whole-body controllerで実演・評価する。 - 各タスク60件のnative human virtual-reality demonstrationsを提供。 - 同期multi-camera viewsとfull-body execution recordsを含む。 - vision-language-action fine-tuning、imitation learning、demo-driven reinforcement learning、cold-start coding agentsを比較する。

2. 先行研究と比べてどこがすごい?

- 従来のBiGymをUnitree G1に拡張し、20 household tasksを統一whole-body controllerで扱う。 - 全手法に同じonboard views、proprioception、whole-body controllerを提供。 - 60 native human VR demonstrations per taskと同期multi-camera views、full-body execution recordsを備える。 - vision-language-action fine-tuning、imitation learning、demo-driven reinforcement learning、cold-start coding agentsを同一条件で比較。

3. 技術・手法の肝は?

- Unitree G1向けにBiGymを適応し、20 household tasksを設定。 - 実演と評価にunified whole-body controllerを使用。 - 各タスク60件のnative human virtual-reality demonstrationsを収集。 - 同期multi-camera viewsとfull-body execution recordsを提供。 - vision-language-action fine-tuning、imitation learning、demo-driven reinforcement learning、cold-start coding agentsをベンチマーク。

4. どうやって有効だと検証した?

- 20 household tasksで各手法を評価。 - 同一のonboard views、proprioception、whole-body controllerを全手法に適用。 - vision-language-action fine-tuningがnine-task meanで最高。 - agent-developed programsがdemo-driven reinforcement learning baselineを上回り、bimanual reachingでリード。 - cross-workspace stackingは未解決、π_{0.5}はpick-boxで低い、multi-object transportはimitation learning、demo-driven reinforcement learning、coding agentsで困難。

5. 議論はある?

- cross-workspace stackingは依然としてopen problem。 - π_{0.5}はpick-boxで低い性能にとどまる。 - multi-object transportはimitation learning、demo-driven reinforcement learning、coding agentsにとって難しい。 - これらの困難さが議論の焦点。 - 要旨からはその他の議論は不明。

6. 次に読むべき論文は?

- BiGym(元のベンチマーク)。 - vision-language-action fine-tuning。 - imitation learning。 - demo-driven reinforcement learning。 - cold-start coding agents。 - π_{0.5}。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zexi Zhang, Zecheng Zhu, Zidong Chen, Zulkhuu Tuya, Stephen James

分類: cs.RO, cs.LG

原文アブストラクト

Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, $π_{0.5}$ stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at https://github.com/swirl-uk/BiGym2.

関連論文

PR本紙発行元 EmplifAI