日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.28175

DAVIS: ヒューマノイドサッカースキルのための深度のみのエンドツーエンド能動視覚フレームワーク

DAVIS: A Depth-Only End-to-End Active-Vision Framework for Humanoid Soccer Skills

シェア:XThreadsFacebookLINEはてブBluesky

頭部搭載の深度画像と固有感覚履歴のみから、ヒューマノイドロボットがサッカーのシュートやドリブルを直接25自由度の関節PD目標として学習するエンドツーエンドフレームワークを提案。

詳しい要約

1. どんなもの?

- 人型ロボットのサッカー接触スキルを、頭部搭載のdepth画像のみから学習するend-to-endフレームワークDAVISを提案。 - 入力はdepth画像、proprioceptive history、任意の低次元task command。 - 出力は25-DoFのjoint PD targetsを直接生成し、実行時の追加知覚・計画モジュールを不要とする。 - 対象スキルはgoal-directed shootingとdirectional dribbling。

2. 先行研究と比べてどこがすごい?

- 従来のサッカー接触スキルは高インパクトのキック生成に焦点が当たりがち。 - 本手法はperception, approach, alignment, impact, recoveryのループを、自己運動による視点変化やボール見失い、接触結果の不確実性下で扱う。 - 実行時に追加の知覚・計画モジュールを必要とせず、depthのみで25-DoF関節目標を直接出力する点が特徴。

3. 技術・手法の肝は?

- 訓練中にvisibility-aware auxiliary geometryを学習。 - GT-to-prediction annealing、task curricula、AMP-style motion priorsを組み合わせ、privileged supervisionから実機展開へ滑らかに橋渡し。 - タスクごとにobjects, commands, rewards, curriculaを定義し、goal-directed shootingとdirectional dribblingを実装。

4. どうやって有効だと検証した?

- simulation、Noetix E1 real-robot experiments、ablationsを通じて検証。 - 具体的な評価指標や成功率は要旨からは不明。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- AMP-style motion priorsに関連するAdversarial Motion Priors (AMP) の原論文。 - 人型ロボットのサッカーに関する先行研究(要旨で具体的に参照されていないため、同分野の定番としてhumanoid soccer skillsに関する研究)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiakang Jin, Yixiao Huo, Pengyuan Wang, Yinan Han, Tingxuan Zhang, Zhuobing Zhao, Xuanxin Zhou, Zhangchen Ye, Enxuan Ruan, Yifei Bao, Jiankun Yang, Chenghao Sun, Wenhao Cui, Xiaoyu Tian, Yiming Li

分類: cs.RO

原文アブストラクト

Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this work, we ask a compact yet stricter question: can a humanoid learn soccer contact skills using only a head-mounted depth image, proprioceptive history, and an optional low-dimensional task command, and directly output 25-DoF joint PD targets without extra runtime perception or planning modules? To this end, we propose DAVIS, a depth-only end-to-end framework for humanoid soccer skills that learns visibility-aware auxiliary geometry during training, and combines GT-to-prediction annealing, task curricula, and AMP-style motion priors to smoothly bridge privileged supervision and real deployment. Built on this framework, we instantiate representative soccer contact skills, including goal-directed shooting and directional dribbling, through task-specific definitions of objects, commands, rewards, and curricula, and validate them through simulation, Noetix E1 real-robot experiments, and ablations.

関連論文

PR本紙発行元 EmplifAI