日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
長文脈動画理解arXiv:2610.10156

HeiCo-FOCUS: 外科手術の長文脈動画理解のための臨床データセット

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

シェア:XThreadsFacebookLINEはてブBluesky

大腸手術の長時間動画で物体の挿入・操作・遮蔽・除去を追跡するVQAデータセットを構築し、最先端VLMの長文脈理解能力を評価した。

詳しい要約

1. どんなもの?

- 臨床に基盤を置いた長文脈ビデオ理解のためのデータセット HeiCo-FOCUS を提案。 - 手術中の Foreign Object Contextual Understanding をタスクとし、Heidelberg Colorectal surgeries のデータセット上に構築。 - 最大数時間に及ぶ手技中に、複数物体の挿入・操作・遮蔽・除去を継続的に追跡する必要がある。 - 30,000 の VQA ペアを含み、object recognition、temporal grounding、aggregation、event and procedural understanding、complex reasoning の5能力をカバー。 - 大規模クラウドソーシングと39名の外科領域専門家による多段階アノテーションパイプラインで構築。 - 単一フレームから全手技まで段階的に時間的・文脈的負荷を高める multi-track 評価フレームワークを導入。

2. 先行研究と比べてどこがすごい?

- 既存のビデオ理解ベンチマークは短期的推論に焦点を当て、長時間にわたる累積的な時間的一貫性の評価が不足。 - HeiCo-FOCUS は数時間に及ぶ手術ビデオで物体の挿入・操作・遮蔽・除去を継続追跡する長文脈理解を評価。 - 臨床的に grounded なデータセットで、39名の外科専門家が関与し臨床関連性と高品質を確保。 - 単一フレームから全手技まで段階的に難易度を上げる multi-track 評価を導入。 - 10の最先端 VLM の実験で、タスクは未解決であり、約半数が text-only baseline を明確に上回るのみ。

3. 技術・手法の肝は?

- Heidelberg Colorectal surgeries のデータセットを基盤に、Foreign Object Contextual Understanding タスクを設定。 - 30,000 の VQA ペアを作成し、object recognition、temporal grounding、aggregation、event and procedural understanding、complex reasoning の5能力を評価。 - 大規模クラウドソーシングと39名の外科領域専門家による多段階アノテーションパイプラインでデータ構築。 - multi-track 評価フレームワークで、単一フレームから全手技まで時間的・文脈的負荷を段階的に増加。 - 10の frontier VLMs を評価し、text-only baseline と比較。

4. どうやって有効だと検証した?

- 10の frontier VLMs を用いた実験を実施。 - HeiCo-FOCUS のタスクは far from solved であり、約半数のモデルのみが text-only baseline を明確に上回る。 - ビデオトラック全体で、event and procedural understanding が最も良好(全モデル平均 Accuracy: 56.5%)。 - temporal grounding は全評価モデルにとって特に困難(平均 Accuracy: 19.7%)。 - これにより長文脈ビデオ理解の評価ギャップを示し、データセットの有効性を検証。

5. 議論はある?

- 既存評価は短期的推論に偏り、長時間の累積的時間一貫性を評価できていないという問題意識。 - HeiCo-FOCUS により、数時間ビデオでの信頼性の高い時間的一貫推論の必要性を提起。 - 実験結果から、現在の VLM は長文脈手術ビデオ理解において未解決であり、特に temporal grounding が課題。 - データセットが、時間的一貫性を備えたモデル開発の触媒となることを期待。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、Vision-Language Models (VLMs) のビデオ理解ベンチマーク、長文脈ビデオ理解、手術ビデオ理解、VQA データセットが挙げられる。 - 同分野の定番として、VideoQA、Temporal Grounding、Surgical Video Understanding の論文を読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Leon Mayer, Lucas Luttner, Patrick Godau, Kai Fritzsche, Annika Reinke, Leonie Boland, Jule Brandt, Janne Heinecke, Chloe K. Nobuhara, Niklas Holzwarth, Evangelia Christodoulou, Marcel Knopp, Dominik Michael, Pascale Piermarco, Saliq Neyaz, Korhan Derin Özarslan, Jakob Hennighausen, Carlos Aumente-Maestro, Tim Rädsch, Dheeraj Baji, Peter Maximilian Full, Finn Aichholz, Justus Veit Erpenbeck, Linus Finn Schott, Bastian Winkelhausen, Claas de Boer, Bianca Güttner, Anneli Hummel, Gregor Just, Max Kirchner, Chenyang Li, Rozenn Raffaut, Ariel Rodriguez, Danush Kumar Venkatesh, Kevin Wang, Jinjing Xu, Mona Sheikh Zeinoddin, Salman Khan, Thomas M. Pausch, Stefanie Speidel, Danail Stoyanov, Daniel A. Hashimoto, Fiona R. Kolbinger, Thomas G. Weiser, Lena Maier-Hein

分類: cs.CV

原文アブストラクト

Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.

PR本紙発行元 EmplifAI