日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38087

CrossBFM: ヒューマノイド間で共有される潜在行動空間の蒸留

CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments

シェア:XThreadsFacebookLINEはてブBluesky

リターゲティングを利用して、複数のヒューマノイドに共通の潜在行動空間を1GPU時間未満で蒸留し、10GPU時間の追従学習で全身制御を実現する手法を提案。

詳しい要約

1. どんなもの?

ヒューマノイドの Behavior Foundation Model (BFM) を、複数の機体間で共有できる潜在行動空間として蒸留する手法 CrossBFM を提案する。従来の Forward-Backward 表現は 1 機体あたり数百 GPU 時間を要し、機体ごとに無関係な潜在空間が生成されていた。CrossBFM は retargeting によるフレーム単位の対応を利用し、機体固有パラメータを持たない統一 encoder で複数機体を同時に蒸留し、1 GPU 時間未満で学習する。その後 latent-conditioned tracker を PPO で 10 GPU 時間追加学習し、全身制御を実現する。

2. 先行研究と比べてどこがすごい?

従来の Forward-Backward 表現は 1 機体で数百 GPU 時間を要し、2 機体目を学習すると第 1 の空間と無関係な第 2 の空間が生成され、機体固有の latent となり統一・転移できなかった。CrossBFM は潜在空間を機体間で転移可能な資産として扱い、機体固有パラメータなしの統一 encoder で複数機体を同時に 1 GPU 時間未満で蒸留する。さらに latent-conditioned tracker を 10 GPU 時間で学習し、3 機体すべてで 3 つの prompting モードが転移することを示した。

3. 技術・手法の肝は?

retargeting がフレーム単位の機体間対応を与えることを利用し、機体固有パラメータを持たない統一 encoder アーキテクチャで行動空間を蒸留する。この encoder により複数の学習機体を同時に 1 GPU 時間未満で処理する。続いて latent-conditioned tracker が蒸留された latent を従来の PPO 学習で全身制御に変換し、追加 10 GPU 時間で完了する。

4. どうやって有効だと検証した?

3 機体の蒸留ヒューマノイドで、motion tracking は latent-conditioned policy が joint-conditioned 版に比べ 0.025 rad の損失で済み、pose 間の smooth goal reaching は転倒なし、41 の reward prompt すべてで reward optimization が成功した。追加実験で、encoder を motion corpus の 1/4 で回帰すると tracking 性能の 5% のみ低下し、一部機体で encoder を学習し未見機体で評価すると既見機体の tracking 性能の最大 89% を回復した。実機でも 3 つの prompting モードと flow-based 生成 latent でパイプラインを検証した。

5. 議論はある?

実験から、encoder を motion corpus の 1/4 で回帰すると tracking 性能の 5% のみ低下すること、一部機体で encoder を学習し未見機体で評価すると既見機体の tracking 性能の最大 89% を回復することが明らかになり、形態的に類似した機体への cross-embodiment 汎化を示す。ただし、形態的に類似した機体への汎化に留まる点や、実機検証の詳細な限界については要旨からは不明。

6. 次に読むべき論文は?

Forward-Backward representations、PPO、retargeting、flow-based 生成 latent が参照・比較されている。関連手法として Behavior Foundation Models (BFMs)、latent-conditioned policy、joint-conditioned policy が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger, Cuc T. Trinh, Siwei Ju, Vien Anh Ngo, Jan Peters, Xinchao Wang, An T. Le

分類: cs.RO

原文アブストラクト

Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only $0.025$ rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all $41$ reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only $5\%$ of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to $89\%$ of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: https://dotandung.github.io/crossbfm/

関連論文

PR本紙発行元 EmplifAI