身体化された能動視覚における視空間複雑性の人間工学的認知モデル
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
視覚・聴覚・空間刺激を含むマルチモーダルデータの複雑性を、身体化認知と能動視覚の理論に基づいて分析する枠組みを提案し、運転などの日常環境での応用を示した論文。
著者: Vasiliki Kondyli, Jakob Suchan, Mehul Bhatt
分類: q-bio.NC, cs.AI, cs.CV
原文アブストラクト
We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of complexity in embodied perception and interaction in dynamic, naturalistic settings. Grounded in theories of embodied cognition and active vision, we argue that embodied perceptual complexity emerges from an agent's dynamic engagement with the environment and must be analyzed holistically, as a combination of qualitative and quantitative attributes pertaining to, for instance, visuospatial and auditory features. Building on previous work on visual complexity, we expand this into a categorization of diverse complexity attributes -- quantitative, structural, dynamic, auditory, and interactional -- that together characterize multimodal complexity. We demonstrate how this model provides a theoretical framework for characterizing aspects of visuospatial complexity and their interactions, specifically in the context of everyday driving. We also discuss practical applications of the proposed model for creating and evaluating benchmark datasets (e.g., in driving) that centralize cognitive human factors, as well as applications aimed at systematically investigating the effect of visuospatial complexity on human active vision from the viewpoint of visual perception research. The proposed framework lays the foundation for automated methods that interpret complexity in 3D dynamic environments from a human-centered perspective, serving as a semantic template for explainable computational analysis of visuospatial complexity with a categorical focus on cognitive human factors.