日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18663

VLA-ULAP: クラウドVLA呼び出しと超軽量ローカル行動予測をエッジで交互実行

VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

シェア:XThreadsFacebookLINEはてブBluesky

クラウド上の大規模VLAモデルの呼び出しと、エッジ上の超軽量ローカル行動予測器(ULAP)を交互に実行することで、成功率をほぼ維持しながら推論時間と消費電力を大幅に削減する手法を提案。

詳しい要約

1. どんなもの?

- 提案手法はVLA-ULAP。 - クラウド上のVLAポリシー呼び出しと、エッジ上の超軽量ローカル行動予測器(ULAP)を交互に実行する。 - ULAPは視覚エンコーダを含め約7.4Mパラメータ。 - 現在の視覚入力、固有感覚、実行済み行動履歴から行動チャンクを1パスで予測。 - 独立に学習され、VLAの隠れ状態やオンライン検証、サーバ往復を必要としない。 - Jetson Orin Nanoで19.9ms/0.183J、RTX A6000上のGR00Tは284.3ms/50.55J。

2. 先行研究と比べてどこがすごい?

- 従来のVLAは数十億パラメータでオンボード電力と通信遅延が問題。 - ローカルVLA高速化手法(ACT, SP-VLA)と比較し、VLA-JEPAで推定49.2%の推論時間削減、51.0%のGPUエネルギー削減(ACT比、同等成功率)。 - SP-VLA比では77.1%時間削減、79.9%エネルギー削減(同等成功率)。 - シミュレーションでVLA呼び出しを48.8-76.7%削減しつつ、ベースライン成功率の95.0-97.5%を維持。 - 実機SO-101でもベースライン成功率の95.2-100%を維持し、推論時間47.9-58.0%削減、エネルギー52.1-62.5%削減。

3. 技術・手法の肝は?

- クラウドVLA呼び出しとULAPによるローカル予測を交互に実行。 - ULAPは現在の視覚入力、固有感覚、実行済み行動履歴を入力とし、行動チャンクを1パスで予測。 - 約7.4Mパラメータ(凍結視覚エンコーダ含む)。 - 独立学習でVLAの隠れ状態やオンライン検証、サーバ往復不要。 - エッジデバイス(Jetson Orin Nano)で動作。

4. どうやって有効だと検証した?

- 3つのシミュレーションのベースポリシー/ベンチマークペアで評価。 - VLA呼び出しを48.8-76.7%削減し、成功率95.0-97.5%維持。 - VLA-JEPAでACT比49.2%時間削減、51.0%エネルギー削減、SP-VLA比77.1%時間削減、79.9%エネルギー削減(同等成功率)。 - 実機SO-101で見た配置と未見配置で成功率95.2-100%維持、推論時間47.9-58.0%削減、エネルギー52.1-62.5%削減。 - 遅延対応LIBERO-Safetyシミュレーションでπ0.5を11.0および15.5ポイント上回り、VLA呼び出しを約半減。

5. 議論はある?

- 要旨からは不明。 - 限界や議論については記述なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: GR00T, ACT, SP-VLA, VLA-JEPA, π0.5, LIBERO-Safety。 - 関連手法としてVLAポリシー、ローカル行動予測、エッジ推論の研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Deyu Cao, Ryuji Oi, Kosuke Matsushima, Yuxuan Pan, Ziheng Wang, Daichi Fujiki, Atsutake Kosuge

分類: cs.RO, cs.LG

原文アブストラクト

Billion-parameter vision--language--action (VLA) policies demand substantial onboard power, while communication delays in remote inference hinder timely responses. We propose VLA-ULAP, which interleaves remote VLA calls with an Ultra-Lightweight Local Action Predictor (ULAP). With approximately 7.4M parameters including the frozen vision encoder, ULAP combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 19.9 ms and 0.183 J per inference, compared with 284.3 ms and 50.55 J for GR00T on RTX A6000. Across three simulated base-policy/benchmark pairs, selected operating points remove 48.8--76.7\% of VLA calls while retaining 95.0--97.5\% of the baseline success rate. Against local VLA-acceleration alternatives on VLA-JEPA, ULAP uses an estimated 49.2\% less inference time and 51.0\% less GPU energy per successful episode than ACT at comparable success rates, and 77.1\% less time and 79.9\% less energy than SP-VLA at equal success rates. Physical SO-101 experiments retain 95.2--100\% of the baseline success rate across seen and held-out placements while reducing inference time by an estimated 47.9--58.0\% and inference-device energy by 52.1--62.5\%, based on successful-episode call counts and measured device costs. Faster responses also improve dynamic-task success rates: in latency-aware LIBERO-Safety simulation, VLA-ULAP exceeds $π_{0.5}$ by 11.0 and 15.5 percentage points on two tasks while approximately halving VLA calls.

関連論文

PR本紙発行元 EmplifAI