日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
物体検出arXiv:2609.36426

名前を失ってから箱を失う:狭いファインチューニングが検出器の展開語彙外で犠牲にするものの測定と修復

Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary

シェア:XThreadsFacebookLINEはてブBluesky

広範な事前学習済み検出器を狭い領域でファインチューニングすると、領域内精度は上がる一方で語彙にない物体へのカバレッジが低下することを長期追跡で示し、事前学習状態を一部混ぜる訓練不要の修復法を提案した。

詳しい要約

1. どんなもの?

事前学習済みdetectorを狭いdomainでfine-tuningすると、in-domain精度は上がる一方、vocabularyが名前を持たない物体へのcoverageが失われる現象を扱う。in-domain test setにはその例が無いため、同一pretrained checkpointとそのfine-tuned descendantsを縦断比較するprotocolを提案。held-out top-K proposal coverage C_τ(pretrainingがカバーしvocabularyが省いたカテゴリのboxをtop K領域がどれだけ覆うか)を追跡する。

2. 先行研究と比べてどこがすごい?

C_τ自体はopen-world proposal literatureの量だが、縦断的に読む点が新規。in-domain accuracyが上昇する間にC_τが低下し、4 architectures・3 domainsで1024 px²超のboxに対し5.12〜63.35ポイント低下。in-domainの数値でもdetection APでもこの低下は識別できない(APはmissedとmisnamedを同様に扱う)。

3. 技術・手法の肝は?

1つのpretrained checkpointとそのfine-tuned descendantsを比較するlongitudinal protocol。held-out top-K proposal coverage C_τを指標化。freeze ladderの6 depthsでnamingが最初に失われることを観察。さらにtraining不要のrepairとして、pretrained stateの1/4をnormalisation statistics込みで混ぜ戻す手法を提示。

4. どうやって有効だと検証した?

4 architectures・3 domainsでC_τ低下を確認。1024 px²超のboxで5.12〜63.35ポイント低下。両方を採点した1 architectureではadaptationがAPの87%を犠牲にする一方coverageは1/5。freeze ladder全6 depths・全runでnamingが先に失われる。pretrainingを共有しない3 architecturesがcoverageを失うカテゴリで一致し、モデルが学習しなかったカテゴリは失われない。repairは全cellでcoverageを上げ、in-domain accuracyの犠牲は最大2.47ポイント。追加評価1回で検出可能。

5. 議論はある?

失われるものは構造的で、pretrainingを共有しない3 architecturesが同じカテゴリでcoverageを失い、未学習カテゴリは失われない。in-domain数値やAPではこのリスクを捉えられない。repairはtraining不要でnormalisation statistics込みのpretrained state混合が有効だが、要旨からは限界や一般化可能性の詳細は不明。

6. 次に読むべき論文は?

要旨で参照/比較されているのはopen-world proposal literature、detection average precision、freeze ladder、pretrained state mixingによるrepair。関連手法としてopen-world object detection、incremental learning、catastrophic forgetting、normalisation statisticsの転移に関する研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Trung Minh Bui, Jongsul Moon, YoungOuk Kim, Jung-Hoon Hwang, Dongin Shin

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-$K$ proposal coverage $C_τ$: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top $K$ regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. $C_τ$ falls while in-domain accuracy rises, on four architectures and three domains, by $5.12$ to $63.35$ points on boxes above $1024$ px$^2$. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs $87\%$ of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most $2.47$ points of in-domain accuracy. Seeing it costs one extra evaluation pass.

関連論文

PR本紙発行元 EmplifAI