産業用物体検出における人間可読なXAIのための視覚言語モデル適応
Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection
産業製造ラインの物体検出において、非専門家向けに直感的な説明を生成する視覚言語モデルを微調整し、XAIインターフェースを構築した。
著者: Sarvenaz Sardari, Freddy Fernandes, Samarth Yelvande, Jose Moises Araya-Martinez, Alina Roitberg
分類: cs.CV, cs.AI
原文アブストラクト
Explainable Artificial Intelligence (XAI) solutions are essential for building trust in AI technologies and their integration in real manufacturing lines. However, most existing methods are tailored to technical experts, limiting their accessibility to diverse user groups such as blue-collar workers in manufacturing lines who use AI for quality control. In this work, we introduce an XAI interface for object detection in industrial manufacturing based on a fine-tuned vision-language model, designed to generate intuitive explanations for non-expert users. We benchmark existing vision-language models and demonstrate that out-of-the-box models often fall short in delivering clear, context-relevant explanations for non-expert users. To address this, we fine-tune a vision-language model and integrate it into our interface, enabling contextualized, accessible explanations for non-expert users. We demonstrate improvements in explanation clarity, instruction adherence, image groundedness, and contextual awareness over GPT 4o-mini on proprietary and public robotics dataset. This approach advances the accessibility and usability of AI explanations, making them more intuitive and applicable in manufacturing domain.