日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00899

TOAST: 自己回帰型視覚-言語-行動モデルのための確率的ロボット行動トークン化

TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

同じ行動列を複数のトークン列で表現できる冗長性を活かし、学習時に確率的にサンプリングするトークン化手法を提案。データが少ないほど精度が向上し、実機タスクでも有効性を示した。

詳しい要約

1. どんなもの?

- 連続ロボット行動を離散トークン列として扱う Autoregressive Vision-Language-Action モデル向けの新しい行動トークン化手法。 - 名称は TOkenization of Action sequences with STochastic sampling (TOAST)。 - 同一の量子化行動列に対して複数のトークン列が存在し得る冗長性に着目。 - 学習中に確率的サンプリングで代替トークン化を選ぶことで、離散教師信号を多様化する。 - 元のロボット行動は保持し、追加のデモンストレーションも不要。

2. 先行研究と比べてどこがすごい?

- 先行の FAST は多様な時間周波数を含む行動を少数トークンに圧縮し、自己回帰予測を効率化。 - しかし FAST は各量子化行動列に単一の決定論的トークン化を割り当てる。 - TOAST はこの表現冗長性を活用し、同じ行動を表す複数のトークン列を学習時にサンプリング。 - 圧縮によるトークン数削減ではなく、限られたデモからの方策学習効率を改善する点が新しい。 - 決定論的版と比べ、データが少ないほど改善幅が大きいと報告。

3. 技術・手法の肝は?

- 量子化された行動列に対し、等価に復号可能な複数のトークン化候補を確率的にサンプリング。 - 方策学習中に代替トークン化を教師として用い、離散監督を多様化。 - サンプリング後も復号されるロボット行動は元の量子化行動列と同一に保たれる。 - 追加デモンストレーションを必要とせず、既存の自己回帰 next-token 目的と組み合わせ可能。 - 詳細なサンプリング分布やアルゴリズムは要旨からは不明。

4. どうやって有効だと検証した?

- LIBERO ベンチマークで決定論的 counterpart と比較し、一貫した改善を確認。 - 訓練データが 1/16 のみの場合、成功率が 6.8 ポイント向上。 - 実ロボットの 4 つの manipulation タスクでも評価。 - 決定論的 counterpart に対し平均成功率が 15.8 ポイント向上。 - 訓練データが限られる状況で特に有効であることを示す。

5. 議論はある?

- 確率的行動トークン化が自己回帰ロボット方策学習に有効であると主張。 - 特に訓練データが限られる場合の有効性を強調。 - 表現冗長性を活用することが方策学習効率を高める可能性を示唆。 - 限界や失敗事例、計算コスト、他手法との詳細比較は要旨からは不明。

6. 次に読むべき論文は?

- FAST(本手法が改善対象とする決定論的行動トークン化) - Autoregressive Vision-Language-Action models(自己回帰 VLA モデル全般) - LIBERO(評価に用いられたベンチマーク) - 関連する行動トークン化・離散化手法(例:VQ-VAE 系、action chunking 系)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.

関連論文

PR本紙発行元 EmplifAI