日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
画像編集arXiv:2609.34237

DecFlowEdit: ガイダンス分離による自己局所化フローベース画像編集

DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling

シェア:XThreadsFacebookLINEはてブBluesky

フローベース画像編集において、局所化と編集性のためのガイダンススケールを分離することで、背景を保ちながら編集精度を向上させる学習不要・逆変換不要の手法を提案。

詳しい要約

1. どんなもの?

- 本論文は、FlowEdit を基盤とした画像編集手法 DecFlowEdit を提案する。 - FlowEdit は source と target の velocity 差を利用し、inversion-free な semantic change を実現する。 - しかし、デフォルトの classifier-free guidance (CFG) 設定では背景漏れ (background leakage) が生じる。 - DecFlowEdit は、局所化と編集性のための最適な guidance scale を分離 (decouple) する。 - training-free かつ inversion-free で、外部 spatial mask や attention 操作を必要としない。

2. 先行研究と比べてどこがすごい?

- FlowEdit のデフォルト CFG は非対称な source/target scale により背景漏れを引き起こす。 - guidance scale を一致させると (例: CFG 除去)、編集関連の局所化は改善するが編集性が大幅に低下する。 - DecFlowEdit は両者の利点を両立し、局所化と編集性を同時に向上させる。 - 従来の FlowEdit と比較して、PIE-Bench 上で structure distance を約 61-73% 削減、background LPIPS を 68-80% 削減。 - 編集忠実度は同等に保たれる。

3. 技術・手法の肝は?

- まず、CFG なしで評価した velocity 差を時間的に集約し、編集関連の prior を抽出する。 - 次に、この prior を用いて、デフォルト CFG 下での元の更新を再重み付け (reweight) する。 - これにより、局所化のための guidance scale と編集のための guidance scale を分離する。 - training-free かつ inversion-free で、外部 spatial mask や attention 操作を必要としない。

4. どうやって有効だと検証した?

- PIE-Bench 上で FLUX, SD3, SD3.5 を用いて実験。 - FlowEdit と比較して、structure distance を約 61-73% 削減、background LPIPS を 68-80% 削減。 - 編集忠実度は同等であることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- FlowEdit (比較対象として参照) - classifier-free guidance (CFG) に関する研究 - PIE-Bench (評価ベンチマーク) - FLUX, SD3, SD3.5 (基盤モデル)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zheyuan Zhan, Can Wang, Jiawei Chen, Chun Chen, Siwei Lyu, Zeyu Zheng, Defang Chen

分類: cs.CV

原文アブストラクト

Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matching these guidance scales, for example by removing CFG, improves edit-relevant localization but severely degrades editability. To get the best of both worlds, we propose DecFlowEdit, which decouples the optimal guidance scales for localization and for editing in flow-based generative models. In particular, DecFlowEdit first extracts an edit-relevant prior by temporally aggregating velocity differences evaluated without CFG, and then uses this prior to reweight the original updates under default CFG. Our method remains training-free and inversion-free, requiring neither external spatial masks nor attention manipulation. Experiments on PIE-Bench across FLUX, SD3, and SD3.5 show that DecFlowEdit improves background preservation, reducing structure distance by approximately 61 to 73 percent and background LPIPS by 68 to 80 percent relative to FlowEdit at comparable editing fidelity.

関連論文

PR本紙発行元 EmplifAI