日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
手術AI/視覚言語モデルarXiv:2607.19889

潜在アクション誘導による手術インタラクション認識のための視覚ファインチューニング

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

シェア:XThreadsFacebookLINEはてブBluesky

手術中の器具と組織の相互作用を認識するため、潜在アクションで誘導する視覚言語モデルのファインチューニング手法を提案。逆動力学モデルと前方世界モデルを用いてアクション関連領域の表現を強化し、パッチレベルの正則化で局所特徴の崩壊を防ぐ。

著者: Jiajun Cheng, Subarna Tripathi, Sainan Liu, Xiaofan Yu, Shan Lin

分類: cs.CV

原文アブストラクト

Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.