日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02832

FastOPD:軽量VLA展開のためのオンポリシー蒸留

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

シェア:XThreadsFacebookLINEはてブBluesky

大規模VLAモデルをオンポリシー蒸留で軽量な生徒モデルに圧縮し、少ない推論ステップで実機展開を可能にするフレームワークを提案。

詳しい要約

1. どんなもの?

- 大規模VLAの実用的展開を可能にする、foundation-to-lightweight VLAフレームワークFastOPDを提案。 - on-policy distillationにより、大規模VLAから軽量なstudentモデルを構築。 - flow mapとself-consistency objectiveを組み合わせ、少ない推論ステップで高性能を維持。 - シミュレーションと実世界実験で評価。

2. 先行研究と比べてどこがすごい?

- 既存の軽量化手法は小規模アーキテクチャ設計やflow-based policyのdenoisingステップ削減が主流。 - FastOPDはon-policy distillationにより、大規模VLAの性能を保ちつつ推論遅延を大幅削減。 - LIBEROでπ_{0.5}の性能の84%を2ステップで維持し、推論遅延を78.1%削減。 - 既存のfew-step distillationベースラインを平均成功率で上回る。

3. 技術・手法の肝は?

- flow mapを単一状態のteacher supervisionに適応。 - self-consistency objectiveと組み合わせ、teacherのダイナミクスを学習するコンパクトなstudentを構築。 - 理論的に、この目的関数の最小化により、理想的なfew-step teacherモデルが誘導する分布と同等の分布をstudentが回復可能であることを示す。

4. どうやって有効だと検証した?

- 多様なfoundation policyを用いてシミュレーションと実世界実験で評価。 - LIBEROでπ_{0.5}をteacherとし、2推論ステップで84%の性能維持、推論遅延78.1%削減、既存few-step distillationベースラインを上回る。 - LingBot-VLAをteacherとし、RoboTwin 2.0で単一ステップ成功率をベースstudentより15.9ポイント改善。 - World Action Model (WAM)への適用性も実証。 - MolmoAct2から蒸留したコンパクトstudentを実ロボットに展開。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- π_{0.5}、LingBot-VLA、MolmoAct2、World Action Model (WAM)、LIBERO、RoboTwin 2.0。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $π_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.

関連論文

PR本紙発行元 EmplifAI