日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈性arXiv:2610.05017

ニューラルネットワーク概念の位相的表現による操作的抽象化

Operational Abstractions of Neural Network Concepts via Topological Representations

シェア:XThreadsFacebookLINEはてブBluesky

学習済み表現に含まれる概念とその関係を位相的に捉える抽象表現TCRを提案し、概念編集や転移を統一的に行えるようにした。

詳しい要約

1. どんなもの?

- 学習済み表現に含まれる概念とその関係を事後的に抽象化する Topological Concept Representations (TCR) を提案。 - 概念の recoverability と interaction score から中間概念空間を構築し、その組織を topology でコンパクトに符号化。 - 介入は抽象化の修正として表現され、元の学習表現へ伝播される。 - 概念レベルの異なる操作が同一の最適化枠組みを共有でき、望む概念組織と実現機構を分離可能。

2. 先行研究と比べてどこがすごい?

- 既存の concept-based 手法は特定の介入に特化し、概念組織の共通かつ編集可能な表現を提供しない。 - TCR は概念とその関係を同時に特徴づける共通の編集可能表現を提供。 - 異なる概念操作を同一最適化枠組みで扱え、概念組織と実現機構を分離。 - 既存の concept-editing 定式化との接続も示す。

3. 技術・手法の肝は?

- concept recoverability と interaction score から中間概念空間を構築。 - その組織を topology でコンパクトに符号化。 - 介入を抽象化の修正として表現し、学習表現へ逆伝播。 - TCR の stability と reparameterization-invariance を確立。 - 既存 concept-editing 定式化との接続を示す。

4. どうやって有効だと検証した?

- TCR を概念の disentangle の前処理として既存 erasure 手法に適用。 - 同程度の concept leakage で worst-group accuracy を平均 21.89 改善。 - teacher から student への概念転移に TCR を使用。 - concept recoverability を最大 5.54 改善し、test top-1 accuracy も改善または維持。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている既存の concept-editing 手法や erasure 手法、concept-based 手法。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mathieu Pont, Christoph Garth

分類: cs.LG, cs.AI

原文アブストラクト

Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept organization. To achieve this, we introduce Topological Concept Representations (TCR), a post-hoc operational abstraction that jointly characterizes the concepts encoded in a learned representation and their relationships. TCR constructs an intermediate concept space from concept recoverability and interaction scores, and compactly encodes its organization through topology. Interventions are expressed through modifications of this abstraction that are propagated back to the underlying learned representation. This allows different concept-level operations to share the same optimization framework and separates the desired concept organization from the mechanism to achieve it. We establish stability and reparameterization-invariance properties of TCR and its connections to existing concept-editing formulations. We use TCR to disentangle concepts as a preprocessing step for existing erasure methods, improving worst-group accuracy by 21.89 on average at comparable concept leakage. We further use TCR to transfer concepts from teacher to student models, improving concept recoverability by up to 5.54 while also improving or maintaining competitive test top-1 accuracy.

関連論文

PR本紙発行元 EmplifAI