日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34982

ActionUNet: 効率的なマルチスケール微調整によるVLAモデルのロバスト性向上

ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みVLAモデルに軽量な時間的U-Netと連続行動デコーダを組み込み、少ない計算コストでマルチスケールの時間構造を融合し、動作の滑らかさとロバスト性を高める微調整フレームワークを提案。

詳しい要約

1. どんなもの?

- 目的 - Vision-Language-Action (VLA) モデルのロボット操作における汎化とロバスト性向上 - 課題 - 粗い semantics と細かい temporal execution の整合が困難 - 混雑環境で汎化・ロバスト性が不十分 - 提案 - ActionUNet: 効率的な multi-scale fine-tuning フレームワーク - 事前学習済み VLA モデルを低計算コストで強化 - 構成 - temporal-aligned action feature space 内に軽量 temporal U-Net を構築 - conditional SIREN を連続 action decoder として採用 - 効果 - 時間的連続性を保証し高周波の動きのジッタを低減 - 基盤 VLA モデルの汎化と操作ロバスト性を維持

2. 先行研究と比べてどこがすごい?

- 従来の課題 - VLA モデルは粗い semantics と細かい temporal execution の整合が困難 - 混雑環境での汎化・ロバスト性が限定的 - 提案の優位性 - 事前学習済み VLA モデルを最小限の計算コストで fine-tuning - multi-scale 構造事前分布を融合し semantics と temporal execution のスケールギャップを橋渡し - multi-scale モデリングによる微視的 temporal 連続性の破壊や機械的振動を抑制 - conditional SIREN と二次平滑性制約で時間的連続性を保証し高周波ジッタを低減 - 基盤 VLA モデルの汎化と操作ロバスト性を保持 - 定量的優位 - π0.5 の成功率を絶対値で 9.8%, 6.1%, 11.4% 改善 - regression-based OpenVLA-OFT backbone にも汎化

3. 技術・手法の肝は?

- 中核 - 効率的な multi-scale fine-tuning フレームワーク ActionUNet - 手順 - temporal-aligned action feature space 内に軽量 temporal U-Net を構築 - 階層的構造事前分布を融合し semantics と temporal execution のスケールギャップを橋渡し - 連続 action decoder - conditional SIREN を採用 - 明示的な二次平滑性制約を付与 - 時間的連続性を保証し高周波運動ジッタを低減 - 設計意図 - multi-scale 融合による時間的不連続を平滑化 - 機械的実行失敗を低減しつつ基盤 VLA モデルの汎化と操作ロバスト性を維持

4. どうやって有効だと検証した?

- ベンチマーク - RoboTwin 2.0 - LIBERO-Plus - 実世界評価 - real-world hard evaluations - 結果 - π0.5 の成功率を絶対値で 9.8%, 6.1%, 11.4% 改善 - regression-based OpenVLA-OFT backbone にも汎化 - 結論 - fine-tuning 戦略としての有効性と効率性を実証 - 補足 - コードと実装詳細は https://github.com/Di-Zhu123/ActionUNet で公開

5. 議論はある?

- 課題認識 - VLA モデルは粗い semantics と細かい temporal execution の整合が困難 - 混雑環境で汎化・ロバスト性が不十分 - 提案の議論 - multi-scale モデリングは微視的 temporal 連続性を破壊し機械的振動を引き起こす可能性 - conditional SIREN と二次平滑性制約で時間的連続性を保証し高周波ジッタを低減 - 時間的不連続の平滑化により機械的実行失敗を低減 - 基盤 VLA モデルの汎化と操作ロバスト性を保持 - 限界 - 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究 - π0.5 - OpenVLA-OFT - 関連手法 - Vision-Language-Action (VLA) モデル - temporal U-Net - conditional SIREN - ベンチマーク - RoboTwin 2.0 - LIBERO-Plus - 同分野の定番 - Vision-Language-Action (VLA) モデル全般 - robotic manipulation における fine-tuning 手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Di Zhu, Ziheng Yan, Fang Wan

分類: cs.CV, cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves π0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.

関連論文

PR本紙発行元 EmplifAI