日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.07355

追跡は永続性ではない:動画世界モデルは隠れた物体をどこまで保持するか

Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object

シェア:XThreadsFacebookLINEはてブBluesky

凍結したV-JEPA 2予測器に物体を隠して、予測器が隠れた物体の情報をどれだけ保持するかを調べた研究。物体の永続性は予測器側に欠けており、訓練で安価に導入できることを示した。

詳しい要約

1. どんなもの?

ビデオ世界モデルが「見えない物体」をどれだけ保持するかを調べる研究。frozen V-JEPA 2 predictor で物体を隠し、隠れた領域の予測と、その領域だけが異なる2つの世界の encoder 表現を比較。予測器は静止物体を部分的に保持、容器内の物体は全く保持せず、動く物体は0.3秒で失う(V-JEPAのtube maskで0.5秒、ViT-Hは1.1秒)。投影では痕跡が残るが中点以下。encoderは物体の存在を1.00で読み、閉じた容器の中身を3.5秒間デコード可能。予測器の出力は箱を閉じて0.5秒後には2%のシーンでしかボールを含まない。

2. 先行研究と比べてどこがすごい?

従来のビデオ世界モデルは可視物体の追跡に注目。本研究は不可視物体の保持(permanence)を定量化。V-JEPA 2 predictor は静止物体を部分的に保持するが、容器内の物体は全く保持せず、動く物体も短時間で失う。一方、encoderは物体の存在を高精度で読み、閉じた容器の中身も長時間デコード可能。VideoMAEはほとんど保持せず、Cosmosのnext-token predictionは静止隠れ物体は保持するが動く容器内の物体は保持しない。

3. 技術・手法の肝は?

frozen V-JEPA 2 predictor を使用し、物体を隠したビデオを入力。隠れた領域の予測と、その領域だけが異なる2つの世界の encoder 表現を比較。予測器の決定が物体を保持するかを評価。さらに、合成容器での予測器のみの3000ステップの訓練で、permanenceの信念を0.05から1.00に向上。IntPhys-2019を84.2%から93.3%に改善。tube maskを用いた継続訓練で、操作およびインターネットスタイルビデオで1.1-1.6秒の動く物体のキャリーオーバーを実現。

4. どうやって有効だと検証した?

隠れた領域の予測とencoder表現の比較、投影での痕跡の定量化、encoderのprobeによる物体存在の読み取り(1.00)、閉じた容器の中身のデコード可能性(3.5秒)、予測器出力のボール含有率(2%)。合成容器での訓練によるpermanence信念の変化(0.05→1.00)と2つのマッチしたコントロール。IntPhys-2019のスコア改善(84.2%→93.3%)。tube mask継続訓練によるキャリーオーバー時間の測定。

5. 議論はある?

予測器側にpermanenceが欠如しているが、訓練で安価にpriorとしてインストール可能。IntPhys-2019の改善は容器なしのカリキュラムでも同様に達成され、ベンチマークが評価する訓練習慣はスコアリングルールによって変わる。tube maskを用いた継続訓練で動く物体のキャリーオーバーが1.1-1.6秒得られるため、欠陥はlatent predictionに本質的ではない。VideoMAEはほとんど保持せず、Cosmosのnext-token predictionは静止隠れ物体は保持するが動く容器内の物体は保持しない。

6. 次に読むべき論文は?

V-JEPA 2, ViT-H, VideoMAE, Cosmos, IntPhys-2019

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Peng Xie, Amr Alanwar

分類: cs.CV, cs.CL

原文アブストラクト

Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.

関連論文

PR本紙発行元 EmplifAI