手術映像生成:拡散モデルからワールドモデルへの調査
Surgical Video Generation From Diffusion to World Models: A Survey
手術映像生成の研究を、無条件生成・条件付き生成・ワールドモデル生成の3分類で整理し、ピクセル忠実度と臨床的妥当性のギャップや課題を明らかにしたサーベイ論文。
著者: Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang
分類: cs.CV
原文アブストラクト
Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.