日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
画像編集arXiv:2609.36755

ドラッグ編集のための運動に基づく潜在再構成手法MoRe-Drag

Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing

シェア:XThreadsFacebookLINEはてブBluesky

ドラッグ操作による画像編集において、ピクセル空間のワーピングを運動の手がかりとして生成過程に注入し、領域ごとの潜在再構成と段階的適応条件付けで精度と自然さを両立させる手法を提案。

詳しい要約

1. どんなもの?

- ドラッグベース画像編集の新手法 MoRe-Drag を提案。 - ピクセル空間のワーピングを粗い動きの証拠とみなし、生成サンプリング軌道に注入。 - region-aware latent recomposition を refinement, inpainting, anchor 領域に適用。 - stage-adaptive conditioning で motion-grounded 構造形成から意味的洗練へ段階的に移行。 - MLLM-based text encoder を drag-aware 指示推論に適応し、指示不要インターフェースを実現。

2. 先行研究と比べてどこがすごい?

- 既存の drag-based 手法はドラッグ精度と自然で意図に沿う生成のバランスに苦戦。 - MoRe-Drag は DragBench-SR と DragBench-DR で強い base editor よりドラッグ精度を大幅改善。 - SOTA の drag-based 手法の中で優れたドラッグ精度を達成。 - 強い意味的一貫性と視覚的にリアルな結果を両立。

3. 技術・手法の肝は?

- ピクセル空間ワーピングを粗い動き証拠として扱い、生成サンプリング軌道に注入。 - region-aware latent recomposition を refinement, inpainting, anchor 領域に対して実施。 - stage-adaptive conditioning により motion-grounded 構造形成から意味的洗練へ段階的に条件を変化。 - MLLM-based text encoder を drag-aware 指示推論に適応し、指示不要インターフェースをサポート。

4. どうやって有効だと検証した?

- DragBench-SR と DragBench-DR で実験。 - 強い base editor よりドラッグ精度を大幅改善。 - SOTA の drag-based 手法の中で優れたドラッグ精度を達成。 - 強い意味的一貫性と視覚的にリアルな結果を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- DragBench-SR, DragBench-DR に関連する drag-based 編集の先行研究。 - MLLM-based text encoder を利用した指示推論手法。 - 同分野の定番として DragGAN などの drag-based 編集手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou

分類: cs.CV

原文アブストラクト

Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.

関連論文

PR本紙発行元 EmplifAI