日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
顔認識/解釈可能性arXiv:2608.16251v1

SCOUT: 顔認識テンプレートのオープンボキャブラリ編集のための意味概念発見

SCOUT: Semantic Concept Discovery for Open-Vocabulary Editing of face Recognition Templates

シェア:XThreadsFacebookLINEはてブBluesky

顔認識テンプレート内の意味概念を直接発見・操作するフレームワークを提案し、自然言語記述から潜在特徴の意味仮説を生成して安定性を検証することで、編集可能な意味方向を獲得する。

著者: Leon Todorov, Peter Rot, Peter Peer, Vitomir Štruc, Klemen Grm

分類: cs.CV

原文アブストラクト

Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely unexplored. Existing interpretability methods for face recognition often rely on manual neuron inspection or predefined attribute labels, limiting scalability and semantic flexibility. To address this gap, we propose SCOUT (Semantic Concept Discovery for Open-VocabUlary Editing of Face Recognition Templates), an end-to-end framework for discovering and directly manipulating semantic concepts in face recognition templates using mechanistic interpretability. SCOUT learns sparse template representations, generates semantic hypotheses for latent features from natural-language descriptions, and validates their stability. The resulting features act as controllable semantic directions for direct editing, avoiding costly edit--re-encode pipelines. Experiments with face recognition models using CNN, ViT, and Swin backbones show that SCOUT discovers interpretable concepts beyond standard attribute labels and enables controllable, identity-aware template manipulation with negligible impact on identity matching. We further show that edited templates can subsequently be decoded with independent inversion models for visualization and evaluation.