日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35709

離散VLAモデルによるヒューマノイドの移動操作

Humanoid Loco-Manipulation With Discrete VLA Model

シェア:XThreadsFacebookLINEはてブBluesky

ヒューマノイドの全身動作を部位ごとの離散トークンに分解し、言語モデルの語彙を拡張して学習・リアルタイム制御を可能にしたVLAモデルHolo-Mを提案。

詳しい要約

1. どんなもの?

- ヒューマノイドのloco-manipulationのための離散VLAモデル「Holo-M」を提案。 - 言語モデルの語彙にaction tokenを追加し、全身の行動空間(脚、胴、腕、手)を扱う。 - 4つの身体部位別tokenizer(end-effector, body, hand, kinematics)を統合。 - リアルタイム制御のためgrouped discrete diffusion decodingを採用。 - SIMPLEベンチマークで最高成功率を達成。

2. 先行研究と比べてどこがすごい?

- 従来の離散VLAモデルはロボットアームのmanipulationに限定され、ヒューマノイドの高次元・異種全身行動空間には未対応。 - 連続action expertを分離するモデルに生じるknowledge-insulation問題を回避。 - ヒューマノイドloco-manipulationのための初の離散VLAモデル。 - 異なるembodimentやデータソース(teleoperation, ego-centric human video, simulation)を横断して学習可能。

3. 技術・手法の肝は?

- 言語モデルの語彙をaction tokenで拡張し、actionを言語トークンとして扱う。 - 統一action tokenizerが全身行動空間を4つの身体部位別tokenizerに分解。 - 各身体部位のaction tokenをgrouped discrete diffusion decodingでデコードし、autoregressionを回避してリアルタイム性を確保。 - これにより異種データソースからの学習を可能にする。

4. どうやって有効だと検証した?

- SIMPLE humanoid loco-manipulationベンチマークで広範な実験を実施。 - generalist評価とspecialist評価の両方で最高成功率を達成し、2位に大きな差をつけた。 - コードとモデル重みを公開予定。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、離散VLAモデル(例: RT-2, OpenVLA)、連続action expertを用いるVLAモデル、humanoid loco-manipulationのベンチマーク(SIMPLE)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenxin Shao, Siqi Chai, Kun Li, Kerou Zhang, Xinzhou Jiang, Wei Xu, Qiang Liu

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.

関連論文

PR本紙発行元 EmplifAI