日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2608.01049

FactorJEPA:混雑した混沌とした南半球の都市世界におけるレイアウト・エージェント・相互作用チャネルへのモノリシックな未来の因子分解

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

シェア:XThreadsFacebookLINEはてブBluesky

南半球の混雑した都市環境(DENSEWORLD)向けに、世界モデルの予測をレイアウト・エンティティ・相互作用に分解するFactorJEPAを提案し、1000時間の大規模データセットで精度とロバスト性を向上させた。

詳しい要約

1. どんなもの?

FactorJEPAは、混雑した混沌としたGlobal Southの都市環境(DENSEWORLD)向けに設計された世界モデルである。既存のJEPA(Joint Embedding Predictive Architecture)が単一の潜在表現で未来を予測するのに対し、FactorJEPAは未来をレイアウト、エージェント、相互作用のチャネルに分解し、可視性ゲートと分離された部分空間を用いて部分的に観測されたエージェントを保持し、クロスファクターの近道を防ぐ。また、22都市にわたる1,000時間のドライブスルー、ウォークスルー、航空ビデオからなる大規模データセットDENSEWORLD-115kを導入する。

2. 先行研究と比べてどこがすごい?

既存のJEPAは低密度でレーン構造のある環境で評価されており、混雑したGlobal Southの都市環境では、ソフトな空間境界、極端なエージェントの不均一性、持続的なオクルージョン、混合交通下での迅速な社会的交渉といった特性を扱うのが難しい。FactorJEPAは、世界構造を第一級の予測プリミティブとして扱い、モノリシックな潜在表現ではなく、レイアウト、エンティティ、相互作用を構成することで、これらの課題に対処する。

3. 技術・手法の肝は?

手法の肝は、未来の潜在表現をレイアウト、エージェント、相互作用のチャネルに分解することである。可視性ゲートを使用して部分的に観測されたエージェントを保持し、分離された部分空間を導入してクロスファクターの近道(ショートカット)を防ぐ。これにより、不均一性と部分的な観測可能性の下での密集した相互作用ダイナミクスを保存する。

4. どうやって有効だと検証した?

有効性は、Future-frame L1、Causal L1、Mask-ratio slope、Motion cosineの4つの指標で検証された。Future-frame L1は将来の潜在精度、Causal L1は介入感度予測、Mask-ratio slopeは視覚的証拠の減少に対するロバスト性、Motion cosineは再現可能なモーション情報のトレードオフを示す。方法のランキングは2Bおよび1BのV-JEPA 2.1バックボーンで再現され、rho = 0.895から0.978の相関が得られた。

5. 議論はある?

要旨からは、議論の余地がある点として、モーション情報のトレードオフ(Motion cosine)が挙げられるが、詳細な議論は不明。また、DENSEWORLDデータセットのバイアスや、FactorJEPAの一般化可能性についての議論は要旨には含まれていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、Joint Embedding Predictive Architectures (JEPA) と V-JEPA 2.1 が挙げられる。次に読むべき論文としては、JEPAの基礎論文やV-JEPA 2.1の論文が適切である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das

分類: cs.AI, cs.CV, cs.LG

原文アブストラクト

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).

関連論文