日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12090

VLAモデルのための潜在推論フローの再構成と洗練

Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルが過去の成功した潜在推論を記憶し、動的に再利用・洗練することで、ロボット制御の成功率を向上させる手法FLOWMEMを提案。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデル向けの新手法 FLOWMEM を提案。 - 実行後に破棄されていた成功した latent reasoning を再利用可能な経験として蓄積。 - 身体化コンテキストの変化に応じて latent fragment を動的に検索・再構成し、reasoning route を形成。 - 現在の視覚・固有受容感覚の証拠で route を refine してから行動生成の条件付けに使う。 - 閉ループ VLA 制御のための統合モデル。

2. 先行研究と比べてどこがすごい?

- 既存手法は各 policy query ごとに latent state を生成・refine するが、実行後の成功した reasoning を破棄し、同様の計算をゼロから再構築していた。 - FLOWMEM は成功した latent computation を再利用可能な reasoning 経験に変換する点が新しい。 - 固定の retrieved context を追加するのではなく、身体化コンテキストの進展に合わせて動的に検索・再構成する。 - 時間構造と進捗に沿った reasoning route を形成できる。

3. 技術・手法の肝は?

- 成功した latent computation を記憶として保持し、再利用する仕組み。 - 身体化コンテキストの変化に応じて互換性のある latent fragment を動的に検索・再構成。 - 成功した計算の時間構造と進捗に従う reasoning route を形成。 - 現在の視覚・固有受容感覚の証拠で route を refine。 - refine 後の route が行動生成を条件付ける。

4. どうやって有効だと検証した?

- RoboMME と LIBERO-Plus で評価。 - FLOWMEM はそれぞれ 48.0% と 77.3% の成功率を達成。 - memory-free な policy をそれぞれ 1.7 ポイント、4.1 ポイント上回った。 - 成功した latent computation の再利用が閉ループ VLA 制御に有効であることを示す。

5. 議論はある?

- 成功した latent computation を再利用することが閉ループ VLA 制御に価値があると主張。 - 具体的な限界や失敗ケース、計算コスト、記憶の一般化性などは要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法として memory-free policies、latent reasoning を用いる VLA モデル、retrieval-augmented な制御手法が挙げられる。 - 同分野の定番として RoboMME、LIBERO-Plus ベンチマークや VLA モデル全般を参照するとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hongyu Shi, Sen Zhao, Zuyu Zhang, Lifeng Shen, Ding Zou, Xinyu He, Xu Zhang, Qinghua Zhang

分類: cs.AI

原文アブストラクト

Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.

関連論文

PR本紙発行元 EmplifAI