日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12435

VioLA: 人間のデータから汎用人型制御ポリシーを学習

VioLA: Learning Generalist Humanoid Control Policies from Human Data

シェア:XThreadsFacebookLINEはてブBluesky

関節コマンドではなく身体・手の動作潜在表現を予測する汎用人型ポリシーを提案し、人間の動作記録をそのまま学習に活用することで、タスク固有の微調整なしに実機で指示追従を実現した。

詳しい要約

1. どんなもの?

ヒューマノイドロボットが全身で指示に従うための汎用制御ポリシー VioLA を提案する研究。 - 従来の関節コマンド予測ではなく、身体と手の motion latent を予測する。 - 事前学習済み body/hand controller が latent を実機で実行する。 - motion encoder が人間とロボットの動きを同一 latent 空間に写像する。 - これにより人間の記録がポリシーの行動空間でラベル付けされる。 - 学習デモプールは 1億4,060万フレーム、うち 93.2% が人間データ。

2. 先行研究と比べてどこがすごい?

従来のヒューマノイド汎用ポリシーは新規指示にそのまま従えず、タスクごとの teleoperated demonstration で fine-tuning が必要だった。 - 人間のデモは大量にあるが、人間の動きはそのままロボットコマンドにならないという課題があった。 - VioLA は行動空間の再定義により、タスク固有 fine-tuning なしで locomotion 指示に zero-shot で従う。 - 実機 locomotion で 100% 成功、GR00T N1.7 は 16.7%、Ψ_0 は 0%。 - manipulation も fine-tuning なしで 88.6% 成功。

3. 技術・手法の肝は?

ポリシーが予測する対象を関節コマンドから body/hand motion latent に変更する点が肝。 - 事前学習済み body- and hand-controller が latent を実機で実行する。 - 対応する motion encoder が人間とロボットの動きを同じ latent 空間へ写像する。 - その結果、人間の記録がポリシーの行動空間でラベル付け可能になる。 - 大規模デモプール(1億4,060万フレーム、93.2% 人間)で学習する。 - 同一手法が 2 つの VLA と 1 つの world-action model backbone で機能する。

4. どうやって有効だと検証した?

実機ヒューマノイドで zero-shot 評価を実施。 - locomotion 指示追従で 100% 成功。 - 比較対象の GR00T N1.7 は 16.7%、Ψ_0 は 0%。 - manipulation はタスク固有 fine-tuning なしで 88.6% 成功。 - 2 つの VLA と 1 つの world-action model backbone で同手法が機能することを確認。 - 人間デモのみで学習した汎用ポリシーが実機 locomotion を zero-shot で実行することも示す。

5. 議論はある?

要旨からは不明。 - 限界、失敗事例、安全性、計算コスト、latent 空間の汎化性などは記述されていない。 - コードとチェックポイントは公開予定と述べられている。

6. 次に読むべき論文は?

要旨で参照・比較されている研究を挙げる。 - GR00T N1.7 - Ψ_0 - VLA (vision-language-action) モデル - world-action model - 関連するヒューマノイド汎用ポリシー研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius, Martin Riedmiller

分類: cs.RO, cs.LG

原文アブストラクト

Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.

関連論文

PR本紙発行元 EmplifAI