日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ビデオ世界モデルarXiv:2609.35052

OPIS: ビデオ世界モデルにおけるマルチオブジェクト記憶のための入力接地ベンチマーク

OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

シェア:XThreadsFacebookLINEはてブBluesky

ビデオ世界モデルが初期観測の物体をどれだけ記憶できるかを評価するベンチマークOPISを提案し、8モデルで物体の存在・同一性・構造を測定した。

詳しい要約

1. どんなもの?

- ビデオ世界モデルのマルチオブジェクト記憶を評価するベンチマーク「OPIS」を提案。 - 初期観察の固定オブジェクト集合に厳密に基づき、生成履歴や参照ビデオに依存しない。 - 実世界、embodied-robotic、ゲーム世界の3ドメインから500ケース、12,672インスタンス(rigid, articulated, deformable)の密な注釈を提供。 - オブジェクト中心の評価器でObject Presence, Identity, Structureを階層的に測定。

2. 先行研究と比べてどこがすごい?

- 既存評価は生成履歴、ビデオ参照、選択された再訪視点に依存し、真の記憶能力評価が混乱する問題があった。 - OPISは入力を厳密にアンカーとし、固定オブジェクト集合に基づくため、より信頼性の高い評価が可能。 - 8つのimage-to-videoまたはcamera-conditioned世界モデルでスコア48.65~56.01と、記憶保持の難しさを定量的に示した。

3. 技術・手法の肝は?

- 入力グラウンディング:初期観察の固定オブジェクトインスタンスを基準に評価。 - オブジェクト中心評価器:associationと明示的visibility reasoningを組み合わせ、Presence, Identity, Structureを階層的に測定。 - オブジェクトの運動学に基づき静的または動的評価トラックを利用。 - データセットは500ケース、12,672インスタンスに密なオブジェクトレベル注釈を付与。

4. どうやって有効だと検証した?

- 8つのimage-to-videoまたはcamera-conditioned世界モデルを評価。 - OPISスコアは48.65から56.01の範囲。 - 参照インベントリが20未満から40超に増えると、Presence, Identity, Structureスコアが全体的に低下。 - 平均Identityスコアは40.22から23.11に低下。 - 入力の特定オブジェクトインスタンス保持が、もっともらしい視覚要素生成より難しいことを実証。

5. 議論はある?

- 結果は、入力の特定オブジェクトインスタンスを保持することが、もっともらしい視覚要素を生成するより considerably harder であることを示す。 - 参照インベントリの増加に伴うスコア低下が観察され、マルチオブジェクト記憶の課題を浮き彫りに。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、image-to-video世界モデル、camera-conditioned世界モデル、オブジェクト中心評価、video world modelsの記憶評価に関する研究が挙げられる。 - 同分野の定番として、video prediction、object-centric learning、benchmark for video generationの論文を読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hao Wang, Tao Yu, Liuzhou Zhang, HeXin Wang, Haopeng Jin, Yuxuan Zhou, Xinming Wang, Hongzhu Yi, Xinye Li, Yuanlei Wang, Ping Nie, Yan Huang, Yuxuan Zhang, Pengfei Zhou, Yanyan Zou, Wei Yang

分類: cs.CV, cs.AI

原文アブストラクト

Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.

関連論文

PR本紙発行元 EmplifAI