日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
感情理解arXiv:2604.15823

人間のように映画を観る:身体化されたコンパニオンのための自己中心視点の感情理解

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

シェア:XThreadsFacebookLINEはてブBluesky

映画を自己中心的なスクリーン視点で観るロボットの感情理解を扱い、新しいデータセットとマルチモーダル推論フレームワークを提案した論文。

著者: Ze Dong, Hao Shi, Zejia Gao, Zhonghua Yi, Kaiwei Wang, Lin Wang

分類: cs.CV

原文アブストラクト

Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shifts such as viewpoint distortion, scale variation, illumination changes, and environmental interference. However, existing research on movie emotion understanding is almost exclusively conducted on cinematic footage, limiting cross-domain generalization to real-world viewing scenarios. To bridge this gap, we introduce EgoScreen-Emotion (ESE), the first benchmark dataset for egocentric screen-view movie emotion understanding. ESE contains 224 movie trailers captured under controlled egocentric screen-view conditions, producing 28,667 temporally aligned key-frames annotated by multiple raters with a confidence-aware multi-label protocol to address emotional ambiguity. We further build a multimodal long-context emotion reasoning framework that models temporal visual evidence, narrative summaries, compressed historical context, and audio cues. Cross-domain experiments reveal a severe domain gap: models trained on cinematic footage drop from 27.99 to 16.69 Macro-F1 when evaluated on realistic egocentric screen-view observations. Training on ESE substantially improves robustness under realistic viewing conditions. Our approach achieves competitive performance compared with strong closed-source multimodal models, highlighting the importance of domain-specific data and long-context multimodal reasoning.