日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
コンピュテーショナルイメージングarXiv:2606.19985v2

視覚推論ガイドによるライトフィールドからの遮蔽除去

Vision-Reasoning-Guided Occlusion Removal from Light Fields

シェア:XThreadsFacebookLINEはてブBluesky

ライトフィールド統合と視覚言語モデルの推論を組み合わせ、遮蔽物を除去してシーンを復元するフレームワークを提案。合成・実データで最高性能を達成した。

著者: Mohamed Youssef, Oliver Bimber

分類: cs.CV

原文アブストラクト

Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly where dense vegetation severely limits visibility. We propose a visionreasoning-guided light field occlusion removal framework combining light field integration (LFI) with vision-language model (VLM) semantic reasoning. Multi-view observations are first integrated via LFI to suppress foreground occlusions, producing an initial visibility-enhanced representation, a VLM then acts as a conditional semantic prior to restore degraded structures and fine details. A multi-sample fusion strategy aggregates multiple generated hypotheses to improve consistency and reduce hallucination. Experimental results on synthetic and real-world datasets show state-of-the-art performance, achieving the highest average SSIM across four synthetic benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings, with applicability to search-and-rescue and exploratory robotic navigation.