日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09178

CAP: コードブック整合予測によるトークン化ロボットポリシー

CAP: Codebook-Aligned Prediction for Tokenized Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

アクショントークナイザのコードベクトルをポリシーのクラスプロトタイプとして再利用することで、トークン予測誤差に頑健なロボットポリシーを実現する手法CAPを提案。

詳しい要約

1. どんなもの?

- 連続ロボット行動を離散トークン化し自己回帰モデルで扱う手法。 - 既存のトークナイザベース方策はトークンを無関係なクラス指数として扱い、学習済みコード構造を無視。 - 本研究はその構造を再利用するCodebook-Aligned Prediction (CAP)を提案。 - トークナイザのコードベクトルを方策のクラスプロトタイプとして直接使用。 - トークナイザと方策バックボーンは変更しない。

2. 先行研究と比べてどこがすごい?

- 従来のトークナイザベース方策はトークンを独立クラスとして扱い、新たな分類器をゼロから学習。 - CAPはトークナイザのコードブックを再利用し、学習済み潜在構造を保持。 - 4つの量子化器ファミリ、3つのシミュレーションベンチマーク、2つの実機タスクで標準トークン分類ヘッドより一貫して成功率向上。 - トークナイザの再構成品質を固定したまま改善。

3. 技術・手法の肝は?

- トークナイザのコードベクトルを方策のクラスプロトタイプとして直接再利用。 - トークナイザと方策バックボーンはそのまま。 - 標準的なトークン分類ヘッドをCAPに置き換える。 - コードブックの潜在構造を方策に提供し、トークン予測誤差を行動空間でより無害化。 - 方策バックボーンの表現学習も改善。

4. どうやって有効だと検証した?

- 4つの量子化器ファミリ、3つのシミュレーションベンチマーク、2つの実機タスクで評価。 - 標準トークン分類ヘッドと比較し、タスク成功率が一貫して向上。 - トークナイザの再構成品質は固定。 - 分析により、向上はトークン精度の向上や方策ヘッドの変更だけでは説明できないと示す。

5. 議論はある?

- 行動トークナイザは離散ターゲットを超えた有用な行動認識潜在構造を学習。 - 下流方策訓練時にその構造を保持すべき。 - 利得はトークン精度やヘッド変更では説明できず、コードブック再利用がトークン間の潜在構造情報を提供。 - トークン予測誤差が行動空間でより無害になり、方策バックボーンの表現も改善。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、行動トークン化手法(例:VQ-VAE、VQ-BeT、Action Chunking with Transformers (ACT))や自己回帰方策(例:Decision Transformer、Trajectory Transformer)が関連。 - トークナイザのコードブック構造を活用する研究や、ロボット方策のトークン化に関する最新論文を読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoran Chen, Jingtian Ji, Samuel Wheeler, Kaylene Caswell Stocking, Matthew Walter

分類: cs.RO

原文アブストラクト

Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively. However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch. We show that this discarded structure is valuable. We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged. Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed. Our analysis further shows that these gains are not explained by higher token accuracy or changes in the policy head alone. Instead, reusing the tokenizer codebook provides the policy with valuable information about the tokenizer's learned latent structure across tokens, making token prediction errors more benign in action space and improving the representations learned by the policy backbone. These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.

関連論文

PR本紙発行元 EmplifAI