潜在計画は点群でも生き残るか?幾何学的観測のための行動条件付きJEPAワールドモデル
Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations
画像ベースのJEPAワールドモデルを点群観測に拡張し、潜在空間での計画が機能することを示した。特に、分布事前モデルが画像ベースと同等の性能を達成し、行動感応モデルが最も動きの多いシーンで最強の結果を得た。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Fabio F. Oberweger, Michael Schwingshackl
分類: cs.LG, cs.AI, cs.CV
原文アブストラクト
JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, and with 0.3-15% of scene points moving, the slow-feature optimum of latent prediction compounds with the geometric shortcut of 3D self-supervision. We lift three canonical JEPA designs to point clouds, frozen-encoder, distribution-prior, and action-sensitive, and re-sense the stable-worldmodel benchmark so that only the observation differs from the image baselines. All three plan without collapse: the distribution-prior model is statistically equivalent to its re-evaluated image counterpart on every benchmark, and the action-sensitive model attains the strongest result in our controlled comparison where the most geometry moves. Probing explains why: object positions are almost perfectly linearly decodable and attention falls on the few moving points. Planning withstands heavy dropout never seen in training, though range noise defeats the thinnest scene. Geometry finally makes a commanded 3D target a natural goal interface: we construct the goal latent from the target and the current latent, at no cost in success rate, without a goal observation.