PonderPounce: 事前学習済みMLLMをロボット制御のエピソードコンテキストエンジンとして活用
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
PonderPounceは、マルチモーダル大規模言語モデル(MLLM)のネイティブな因果コンテキストをロボットの記憶として再利用し、エピソード観察やデモンストレーションを蓄積してサブゴール生成を行うSystem2と、現在の観察から直接アクションを出力するSystem1のVLAを非同期に連携させる新しいアーキテクチャを提案する。専用の記憶モジュールやブリッジ事前学習なしでエンドツーエンドに訓練され、低遅延で高周波のアクション生成を実現する。
著者: Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu
分類: cs.RO, cs.AI
原文アブストラクト
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without using this contextual capacity as episode memory. Memory-dependent policies address this gap through purpose-built history mechanisms. PonderPounce instead reuses an MLLM's native causal context as robot memory. Ponder, a System2 MLLM, accumulates episode observations, demonstrations, and prior cognition in its native causal context and can generate subgoal text and demonstration reasoning for internal use. Pounce, a System1 VLA, receives the current observation, instruction, and proprioception directly; through the Ponder--Pounce interface, it asynchronously receives only the newest continuous cognition token and its age. Both are jointly trained end to end without a purpose-built memory module or separate bridge pretraining. Optimized serving achieves p50 latencies of 78ms for cognition refresh and 25ms for action-model invocation, supporting 20Hz action playback. On RoboMME with base-scale training data, PonderPounce reaches 60.83% with 9B and 50.04% with 0.8B under the same Pounce architecture and interface, versus 44.51% for FrameSamp+Modul and 17.93% for the current-observation π_{0.5}. With 9x data, it reaches 75.54% versus 57.88% for FrameSamp+Modul. On RoboCasa-DC, the same interface learns from action supervision alone and reaches 12.5% versus 11.6% for the strongest published demonstration-conditioned baseline, falling to 8.6% when cognition is replaced by a learned null state.