日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成/世界モデルarXiv:2610.03154

物理は活性化に宿るか?動画拡散モデルにおける物理量の局在化

Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

シェア:XThreadsFacebookLINEはてブBluesky

動画拡散Transformerの内部表現を調べ、運動や剛体力学の物理量がノイズ除去過程の早い段階で線形に読み出せ、物体トークンに局在することを示した研究。

詳しい要約

1. どんなもの?

- ビデオ生成モデルが物理原理を内部化しているかを探る研究。 - Video Diffusion Transformers (DiTs) の内部表現をプローブし、シミュレータ由来の物理量(運動学、重力・接触下の剛体力学)を線形デコード可能か検証。 - 物理情報がデノイジング過程で構築され、トークン列内で局所化されることを示す。

2. 先行研究と比べてどこがすごい?

- 従来のベンチマークは物理推論の欠陥を指摘するが、内部表現の分析は不十分。 - 本研究は、物理量がモデルの活性化から高精度に線形デコード可能であり、入力ノイズから直接デコードするベースラインを大幅に上回ることを示す。 - 物理情報がデノイジング中に能動的に構築されることを明らかにした点が新しい。

3. 技術・手法の肝は?

- Video DiTs の内部活性化をプローブし、シミュレータ由来の物理量を線形デコード。 - デノイジング過程の早期段階で高精度にデコード可能。 - 物体トークン上の活性化が物理情報を保持し、複数フレームにわたる量が単一潜在フレームから読み取れる。 - 情報はトークン列内で鋭く局所化され、大域的に計算されるが局所的に保存される。 - プローブは部分的な外挿を示し、訓練範囲外のシーン変化や物体配置に転移。 - フル解像度活性化空間でフィットしたプローブ方向はステアリングベクトルとしてモデル出力を変更可能。

4. どうやって有効だと検証した?

- シミュレータ由来の物理量をグラウンドトゥルースとして使用。 - 線形デコーダの精度を評価し、モデル自身のノイズ付き潜在から直接デコードするベースラインと比較。 - 物体トークン上の活性化と単一潜在フレームからの読み取りを検証。 - 訓練範囲外のシーン変化や物体配置への転移(部分的外挿)を確認。 - ステアリングベクトルとしての有効性を検証。

5. 議論はある?

- 物理情報がデノイジング中に能動的に構築されることを示唆。 - 情報がトークン列内で局所化される一方、大域的に計算されるという特性。 - プローブの部分的外挿は、単なる相関ではないことを示す。 - ステアリングベクトルとしての利用可能性。 - 限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:Video Diffusion Transformers (DiTs)、物理推論ベンチマーク。 - 関連手法:線形プローブ、ステアリングベクトル。 - 同分野の定番:物理推論を評価するベンチマーク(例:Physics-IQ、IntPhysなど)や世界モデルとしてのビデオ生成モデルに関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jonas Kneifl, Jakub Skalski, Bartłomiej Twardowski, Kamil Deja

分類: cs.CV, cs.LG

原文アブストラクト

Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We address this by probing internal representations of video Diffusion Transformers (DiTs) for simulator-derived ground-truth physical quantities spanning kinematic motion and rigid-body dynamics under gravity and contact. We find that these quantities are linearly decodable with high accuracy early in the denoising process, substantially outperforming a baseline decoded directly from the model's own noised latents, indicating that the relevant physical information is actively constructed during denoising rather than already present in the input. Additionally, we show that activations at on-object tokens carry the relevant physical information and that quantities defined over multiple frames are readable from single latent frames. Hence, information is sharply localized within the token sequence and is computed globally but stored locally. The probes further show partial extrapolation, transferring to scene variations and object configurations outside their training regime, so what they read is not simply a correlate of the scenes they were fit on. When fitted directly in the full-resolution activation space, the probing directions can serve as steering vectors to change the model's output.

関連論文

PR本紙発行元 EmplifAI