DeepSpeak-Agenticデータセット:人間と具現化AIエージェントの対話記録
The DeepSpeak-Agentic Dataset
人間と具現化AIエージェントの半構造化対話を37時間以上収録したビデオデータセットを構築し、AIエージェントの自動フォレンジック識別や人間とエージェントの相互作用の分析に活用する。
著者: Sarah Barrington, Maty Bohacek, Hany Farid
分類: cs.AI
原文アブストラクト
We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We use this dataset to evaluate the automatic forensic identification (audio, video, or text) of AI agents, study the nature of human-agent interactions, and provide a benchmark for future advances in the large-language models and AI-generated voices and faces that power embodied AI agents. We also contribute a scalable data-capture system that creates agents, automatically pairs them with human crowd workers, records audiovisual conversations across specified scenarios, and identifies and separates the human and agent in the combined stream.
関連論文
- DARP: 多視点ロボット知覚のための校正済み双腕RGB-D-IRデータセットデータセット
- uScenes: 水中ロボット知覚のためのマルチモーダルRGB・3Dソナー画像データセットデータセット
- PRISM:マルチモーダルセンシングを備えた精密で接触豊富な実世界産業スキルデータセットデータセット
- NARRATE: 自動運転における人間中心の説明のためのマルチモーダル実世界オーストラリア運転データセットデータセット
- 衛星画像の改ざんとディープフェイク位置特定のためのベンチマークデータセット構築に向けてデータセット
- InteracVid: ライブチャット動画から構築した実インタラクティブ音声視覚応答データセットデータセット