日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダル生成arXiv:2610.10400

インターリーブ型マルチモーダル生成のための自己修正最適化

Self-correction Optimization for Interleaved Multimodal Generation

シェア:XThreadsFacebookLINEはてブBluesky

マルチモーダル大規模言語モデルにおいて、追加学習なしで画像とテキストが交互に並ぶ生成の一貫性を高める自己修正最適化手法を提案し、時間的一貫性や視覚的主題の保持を改善した。

詳しい要約

1. どんなもの?

- 自己補正最適化(SCO)を提案 - 訓練不要で一貫したインターリーブ生成を実現 - マルチモーダル大規模言語モデル(MLLM)の課題に対処 - 画像とテキストの交互生成における時間的一貫性と視覚的主題の保持を向上

2. 先行研究と比べてどこがすごい?

- 既存のMLLMは追加訓練と拡張データに依存し計算コストが高い - 視覚的主題、時間的一貫性、物理的妥当性の保持に限界 - SCOは訓練不要でこれらの問題を改善 - 時間的一貫性と視覚的主題の保持で大幅な改善を実証

3. 技術・手法の肝は?

- classifier-free guidance updateを参照として利用 - 最小限の自己補正を2つの制約下で実施 - 新イベント制約:画像-テキストシーケンス間の時間的一貫性を促進 - 状態保持制約:後続生成ステップを通じて視覚的主題のコヒーレンスを維持

4. どうやって有効だと検証した?

- 挑戦的なインターリーブマルチモーダル生成ベンチマークで実験 - 時間的一貫性と視覚的主題の保持で大幅な改善を確認 - ビデオ生成への拡張も検証 - ロボット操作や長期的な手作業を含む物理的プロセスのモデリングを改善

5. 議論はある?

- 訓練不要で計算コストを削減 - 時間的一貫性と視覚的主題の保持を向上 - ビデオ生成や物理的プロセスへの拡張可能性 - 具体的な限界や議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてclassifier-free guidance、MLLM、インターリーブ生成の既存研究 - 同分野の定番としてマルチモーダル生成、ビデオ生成、ロボット操作の論文

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xin You, Zhiwei Ning, Zukai Chen, Minghui Zhang, Xuanke Shi, Hanxiao Zhang, Jingsong Liu, Jie Yang, Quan Wang, Yun Gu

分類: cs.CV

原文アブストラクト

Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.

関連論文

PR本紙発行元 EmplifAI