日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
社会的前向き知能/動画ベンチマークarXiv:2609.21371

RobotEQ-Video:世界状態タクソノミーを用いた社会的前向き知能の動画中心ベンチマーク

RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy

シェア:XThreadsFacebookLINEはてブBluesky

動画から人間の状態やニーズを推論し、社会的に適切な先回り支援を行う「社会的前向き知能」を評価するため、4階層の世界状態タクソノミーに基づく2K以上の動画ベンチマークを構築した。

詳しい要約

1. どんなもの?

- 社会 proactive intelligence (SPI) を評価するための video-centric benchmark である RobotEQ-Video を提案。 - SPI は embodied scenarios での社会的適切性を考慮した proactive assistance を指す。 - 従来の image-centric な研究を video-centric に移行し、動画の手がかりを活用。 - 階層的な world-state taxonomy を構築し、6 domains, 20 dimensions, 142 level-1 attributes, 816 level-2 attributes からなる。 - 2K+ の動画、100K+ の human annotations、16K+ の behavior properness ラベルを含む。

2. 先行研究と比べてどこがすごい?

- 先行研究は static images に焦点を当てていたが、本手法は dynamic videos を扱う。 - 従来の free-form data collection ではカバレッジが不十分だったが、階層的 taxonomy で多様なシナリオを網羅。 - これにより SPI 研究を静的画像から動画へと進展させ、ベンチマークのシナリオカバレッジを向上。

3. 技術・手法の肝は?

- video-centric な分析を可能にするため、階層的 world-state taxonomy を構築。 - taxonomy は coarse-to-fine の4レベル構造 (6 domains, 20 dimensions, 142 level-1 attributes, 816 level-2 attributes)。 - これに基づき 2K+ 動画を収集し、100K+ human annotations と 16K+ behavior properness ラベルを付与。 - 評価では現在のシステムの信頼性を検証し、world models の活用可能性も探索。

4. どうやって有効だと検証した?

- 構築した benchmark を用いて評価を実施。 - 現在のシステムは unreliable であり、human performance に及ばないことを示した。 - world models がこのタスクにどのように役立つかを探索。

5. 議論はある?

- 現在のシステムは信頼性が低く、人間のパフォーマンスに達していない。 - world models の活用が有望である可能性を示唆。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として world models が挙げられている。 - 同分野の定番として social proactive intelligence (SPI) や embodied AI に関する研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xinyi Che, Zheng Lian, Kuofei Fang, Xuehao Wang, Xinghai Gao, Junqing Wu, Chuyu Wu, Liyi Liu, Yanhan Huang, Keyi Xie, Haomin Ouyang, Jinyang Wu, Fan Zhang, Runhao Zeng, Xun Yang, Bin He

分類: cs.CV, cs.HC

原文アブストラクト

Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.

PR本紙発行元 EmplifAI