日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34792

D$^2$-VLA:長期動的マニピュレーションのための二重記憶・二重周波数視覚言語行動モデル

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みVLAのKVキャッシュに二重記憶と二重周波数制御を組み込み、動く物体を扱う長期タスクで過去の視覚手がかりを保持しつつ応答できるようにした。

詳しい要約

1. どんなもの?

- 長期的な操作タスクにおいて、視界から消えた手がかりを記憶しつつ動く物体に反応する必要がある。 - 既存の vision-language-action (VLA) ポリシーは最新の観測に依存し、視覚コンテキストの更新には高コストな vision-language model (VLM) の再実行が必要。 - 本研究では、事前学習済み VLA の KV-cache インターフェースで dual memory と dual-frequency control を組み合わせた D$^2$-VLA を提案。 - また、移動物体を操作する際に以前の視覚的手がかりを使うことを要求する 10 タスクのベンチマーク DOMINO-Long を導入。

2. 先行研究と比べてどこがすごい?

- 従来の VLA は最新観測のみに依存し、視覚コンテキスト更新に別の VLM パスが必要だった。 - D$^2$-VLA は dual memory と dual-frequency control により、高コストな VLM 再実行を減らしつつ長期動的操作を可能にする。 - 実験では DOMINO で完全タスク成功率 29.3%($\pi_{0.5}$ は 9.6%、PUMA は 17.2%)、DOMINO-Long で 60.0%(それぞれ 35.4%、20.6%)を達成。 - 実ロボット 8 タスクで成功率を改善し、LIBERO-Long で 97.5%、RoboTwin 2.0 で 74.3% を達成。

3. 技術・手法の肝は?

- 事前学習済み VLA の KV-cache インターフェースを活用。 - block-wise causal KV caching により観測を逐次エンコード。 - 異なる temporal attention patterns に基づき、VLM と action expert 用に別々の historical KV read views を構築。 - 定期的な VLM 更新の間に、gated adapter が新しい視覚特徴を最新の history-conditioned KV block に組み込む。 - 短い fast-memory queue が action replanning を支援。

4. どうやって有効だと検証した?

- DOMINO および DOMINO-Long ベンチマークで評価。 - DOMINO で完全タスク成功率 29.3%、DOMINO-Long で 60.0% を達成。 - 実ロボット 8 タスクで成功率を改善。 - LIBERO-Long で 97.5%、RoboTwin 2.0 で 74.3% を達成。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- $\pi_{0.5}$ - PUMA - LIBERO-Long - RoboTwin 2.0 - DOMINO - DOMINO-Long

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi

分類: cs.CV

原文アブストラクト

Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D$^2$-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D$^2$-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D$^2$-VLA achieves complete-task success rates of 29.3\% on DOMINO, compared with 9.6\% for $π_{0.5}$ and 17.2\% for PUMA, and 60.0\% on DOMINO-Long, compared with 35.4\% and 20.6\%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5\% on LIBERO-Long and 74.3\% on RoboTwin 2.0.

関連論文

PR本紙発行元 EmplifAI