日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成arXiv:2610.03664

ProAR: 自己回帰型動画モデルによる目標指向推論の学習

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

シェア:XThreadsFacebookLINEはてブBluesky

自己回帰型動画生成に目標フレーム予測と未来表現の自己整合を組み込み、最終的な結果から逆算する目標指向の推論プロセスとして動画生成を再構成するフレームワークを提案。

詳しい要約

1. どんなもの?

- 自己回帰(AR)ビデオモデルを目標指向の推論プロセスに変換するフレームワーク「ProAR」を提案。 - 次チャンク予測に依存する短絡的・反応的なパラダイムを克服し、長期的な結果に基づく推論を可能にする。 - 視覚的推論ベンチマークでの性能向上と、embodied reasoningタスクへの適用可能性を示す。

2. 先行研究と比べてどこがすごい?

- 従来のARビデオモデルは次チャンク予測に限定され、局所的な視覚的妥当性に囚われる。 - ProARは目標フレーム予測と未来表現自己整合により、長距離の結果と短距離の遷移を同時にガイド。 - 標準ARベースラインを25%の訓練ステップで上回る訓練効率を実現。

3. 技術・手法の肝は?

- 目標フレーム予測を非対称アテンションマスクを介してARループに統合し、予測目標フレームが中間状態生成を導く。 - 未来表現自己整合を導入し、現在の隠れ状態が将来の時間的ダイナミクスを予期するよう促す。 - teacher-forcingをAR訓練で活用し、単一フォワードパスでクリーンな未来表現を抽出。軽量な訓練時のみのpredictorで現在表現を整合。 - 明示的・スパースな目標監督と暗黙的・密なステップワイズガイダンスを組み合わせ。

4. どうやって有効だと検証した?

- 多様な視覚的推論ベンチマークで実験を行い、ProARの補完的コンポーネントが一貫して性能を向上させることを確認。 - 訓練効率が高く、標準ARベースラインを25%の訓練ステップで上回ることを示す。 - embodied reasoningタスクへの適用可能性も実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法として、自己回帰ビデオモデル、teacher-forcing、非対称アテンションマスク、未来表現自己整合、embodied reasoningに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

分類: cs.CV

原文アブストラクト

Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them. (2) To guide short-range transitions, we introduce future representation self-alignment to encourage current hidden states to anticipate upcoming temporal dynamics. By leveraging teacher-forcing in AR training, we extract clean future representations in a single forward pass and align current representations with them using a lightweight, training-only predictor. Together, these two mechanisms seamlessly combine explicit, sparse target supervision with implicit, dense step-wise guidance, promoting coherent, goal-directed reasoning progress with modest computational cost. Experiments show that ProAR's complementary components consistently improve performance across diverse visual reasoning benchmarks. The framework proves highly training-efficient, surpassing fully trained standard AR baselines using only 25% of the training steps. This paradigm also demonstrates promising applicability to embodied reasoning tasks.

関連論文

PR本紙発行元 EmplifAI