日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09016

PAIR: 視覚言語行動モデルにおける知覚と行動の橋渡し

PAIR: Bridging Perception and Action in Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語行動モデルにおいて、知覚表現と行動生成の間を埋める共有表現を学習するフレームワークPAIRを提案し、LIBEROやCALVINなどで成功率を向上させた。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルは視覚観察と言語指示を連続ロボット行動に写像する。 - この課題はシーンと指示を記述する表現から行動生成を支える表現への遷移を要する。 - 多くの連続行動 VLA はこの遷移を暗黙のままにし、主に最終の行動予測損失で監督する。 - 本研究は PAIR を導入する。これはこれら二空間間で共有された知覚-行動表現を学習する枠組みである。

2. 先行研究と比べてどこがすごい?

- 多くの連続行動 VLA は知覚から行動への遷移を暗黙にし、最終行動予測損失のみで監督する。 - PAIR は知覚と言語の表現と行動潜在トークンの間の共有中間表現を明示的に学習する。 - これにより行動専門家の精緻化前に連続行動情報へアクセス可能となる。 - 評価された OpenVLA-OFT と VLA-Adapter モデルで利得を示す。

3. 技術・手法の肝は?

- 訓練時、Masked Action Autoencoder が専門家行動チャンクを horizon-aligned Action Latent Tokens に符号化する。 - Bridge Module が現在の視覚言語表現からタスク関連特徴を抽出する。 - PAIR はこれらの特徴を Action Latent Tokens と整列し、タスク情報を保持し専門家行動の構造を捉える Bridge Tokens を形成する。 - Bridge Tokens は行動トークン空間に射影され、初期 Action Tokens に注入され、Action Expert 精緻化のための行動準備済み出発点を提供する。 - 推論時、autoencoder は除去され、Bridge Tokens は現在の観察と指示のみから生成される。

4. どうやって有効だと検証した?

- LIBERO、LIBERO-Plus、CALVIN ABC-D での実験で評価された OpenVLA-OFT と VLA-Adapter モデルに利得。 - LIBERO-Plus で PAIR は VLA-Adapter の成功率を 59.1% から 64.2% に上げる。 - CALVIN で VLA-Adapter の平均完了シーケンス長を 4.42 から 4.53 に増加させる。 - 7 つの実世界タスクで PAIR は OpenVLA-OFT の成功率を 51.4% から 65.0% に上げる。 - 表現分析は Bridge Tokens がタスク情報を保持しつつ Action Expert 精緻化前に連続行動情報をアクセス可能にすることを示す。

5. 議論はある?

- 結果は共有中間表現が連続行動 VLA における知覚と行動の有用なインターフェースであることを支持する。 - 要旨からは限界や失敗事例、計算コスト、一般化の議論は不明。

6. 次に読むべき論文は?

- OpenVLA-OFT - VLA-Adapter - Masked Action Autoencoder - LIBERO - LIBERO-Plus - CALVIN ABC-D

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kaixi Feng, Guoheng Sun, Ang li

分類: cs.AI

原文アブストラクト

Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.

関連論文

PR本紙発行元 EmplifAI