日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37307

行動履歴メモリと二重エキスパートノイズ除去による長期的視覚言語行動ポリシー

Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

シェア:XThreadsFacebookLINEはてブBluesky

凍結したVLAにMambaベースのメモリと軽量PreAction Expertを追加し、行動履歴を活用して長期的タスクの知覚的曖昧性を解消する手法を提案。

詳しい要約

1. どんなもの?

- 長期的なロボット操作タスクにおける Vision-Language-Action (VLA) モデルの性能向上を目指す研究。 - 既存 VLA は相互作用履歴を明示的に利用しないため、知覚的曖昧さ (perceptual aliasing) に弱い。 - 提案手法 ActMem-VLA は、凍結した fine-tuned VLA に memory plugin を追加する dual-expert handover アーキテクチャ。 - memory plugin は Mamba-based memory module と lightweight PreAction Expert (PAE) から構成。 - 実行された行動履歴を Mamba でエンコードし、PAE を条件付け、初期の高ノイズ denoising でタスク進行を誘導。 - その後、部分的に denoise された行動を凍結した Action Expert (AE) に渡し、低ノイズステップで詳細を洗練。

2. 先行研究と比べてどこがすごい?

- 既存手法は temporal/progress cues を feature conditioning、action-prior modification、sampling guidance で組み込むが、memory module と base VLA を joint fine-tune するため追加の policy-training コストがかかる。 - 提案手法は trainable history-conditioned steering と frozen base-policy refinement を分離。 - 学習中は fine-tuned base VLA を凍結し、Mamba module と PAE のみを jointly optimize するため、追加パラメータは 3.45% のみ。 - LIBERO-Mem で平均成功率 80.8% を達成し、π_{0.5} の 65.2%、MemoryVLA の 49.5% を上回る。 - 実世界 4 タスクで π_{0.5} より平均成功率が 28.8% 向上。

3. 技術・手法の肝は?

- 凍結した fine-tuned VLA に memory plugin を追加する dual-expert handover アーキテクチャ。 - Mamba-based memory module が実行行動履歴をエンコードし、メモリを生成。 - メモリは現在のコンテキストとともに lightweight PreAction Expert (PAE) を条件付け。 - PAE は初期の高ノイズ denoising 段階でタスク進行を誘導。 - 部分的に denoise された行動を凍結した Action Expert (AE) に渡し、残りの低ノイズステップで行動詳細を洗練。 - 学習中は Mamba module と PAE のみを jointly optimize し、base VLA は凍結。

4. どうやって有効だと検証した?

- LIBERO-Mem ベンチマークで評価。全 10 タスクの平均成功率 80.8% を達成。 - 比較対象: π_{0.5} は 65.2%、MemoryVLA は 49.5%。 - 実世界 4 タスクで評価し、π_{0.5} より平均成功率が 28.8% 向上。 - 追加パラメータは 3.45% のみ。

5. 議論はある?

- 既存手法は memory module と base VLA の joint fine-tune により追加の policy-training コストがかかる点が課題。 - 提案手法は trainable history-conditioned steering と frozen base-policy refinement を分離することでこの問題に対処。 - 知覚的曖昧さ (perceptual aliasing) が長期的タスクの成功率低下を招くことが背景。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- π_{0.5} (比較対象の VLA モデル) - MemoryVLA (比較対象の memory 拡張 VLA) - Mamba (memory module の基盤) - LIBERO-Mem (評価ベンチマーク) - Vision-Language-Action (VLA) モデル全般

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yaxin Zhao, Dianye Huang, Chenwei Wang, Chenguang Yang, Zhongliang Jiang

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8\% average success across all ten tasks, compared with 65.2\% for $π_{0.5}$ and 49.5\% for MemoryVLA, while introducing only 3.45\% additional parameters. Across four real-world tasks, it improves the average success rate over $π_{0.5}$ by 28.8\%.

関連論文

PR本紙発行元 EmplifAI