日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
表現学習arXiv:2610.09457

DSReg: 再構成なしで個々の世界潜在変数を証明可能に復元

DSReg: Provably Recovering Individual World Latents without Reconstruction

シェア:XThreadsFacebookLINEはてブBluesky

再構成やデコーダ、ラベルなしで、構造的多様性を条件に依存性スパース正則化により世界の個々の潜在変数を符号付き置換まで復元する手法を提案。

詳しい要約

1. どんなもの?

- 非線形ICAや辞書学習、因果表現学習などの手法は、再構成や補助監督、非ガウス性などの分布的非対称性を利用して世界の個々の潜在変数を復元する。 - これらのアンカーがない手法(JEPAなど)は、潜在状態を線形変換までしか同定できず、個々の潜在変数は混合したままである。 - 本研究は、再構成もデコーダもラベルもなしに個々の世界潜在変数を証明可能に復元するDSRegを提案する。 - 鍵となる条件はStructural Diversity(構造的多様性)で、異なる潜在変数が観測に異なる依存性の足跡を残すことである。 - LeJEPAが提供する線形同定性に基づき、Structural Diversityの下でDSRegが符号付き置換まで個々の潜在変数を復元することを証明する。 - 任意の線形同定表現に事後適用でき、訓練済みチェックポイントを再利用可能で、共同訓練に対する損失がない。 - 初の完全に同定可能なJEPAを確立し、すべての世界潜在変数を復元する。

2. 先行研究と比べてどこがすごい?

- 従来の同定可能な潜在変数モデルは、再構成、補助監督、非ガウス性などの分布的非対称性を必要とした。 - JEPAなどのアンカーなし手法は線形変換までしか同定できず、個々の潜在変数は混合したままであった。 - DSRegは再構成もデコーダもラベルもなしに個々の潜在変数を復元する点で画期的である。 - Structural Diversityは依存性の足跡に関する条件であり、従来の同定可能な潜在変数モデルのすべての構造条件よりも厳密に弱い。 - これにより、より広いクラスのモデルに適用可能となる。

3. 技術・手法の肝は?

- 鍵となる条件はStructural Diversity:異なる潜在変数が観測に異なる依存性の足跡を残すこと。 - LeJEPAが提供する線形同定性を基盤とし、その上でDSReg(Dependency-Sparsity Regularization)を適用する。 - DSRegは依存性のスパース性を正則化することで、線形同定された表現から個々の潜在変数を符号付き置換まで復元する。 - 再構成やデコーダを必要とせず、任意の線形同定表現に事後適用可能。 - 訓練済みチェックポイントを再利用でき、共同訓練と比べて損失がない。

4. どうやって有効だと検証した?

- 合成レジーム、世界モデルプローブ、学習済み視覚エンコーダ、外部レンダラーにわたって検証。 - DSRegは密な予測を保持しつつ、個々の潜在変数の復元と下流利用をスケールとともに改善する。 - 具体的な評価指標やベースラインとの比較数値は要旨からは不明。

5. 議論はある?

- Structural Diversityは依存性の足跡に関する条件であり、従来の構造条件よりも厳密に弱い。 - これにより、より一般的な設定での同定性が可能になる。 - 再構成やデコーダなしで個々の潜在変数を復元できることの理論的保証が示された。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- LeJEPA(線形同定性を提供する手法) - 非線形ICA、辞書学習、因果表現学習などの従来の同定可能潜在変数モデル - JEPA(Joint-Embedding Predictive Architectures) - これらの関連手法を比較・参照している。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf

分類: cs.LG, cs.AI, cs.RO, stat.ML

原文アブストラクト

Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.

関連論文

PR本紙発行元 EmplifAI