日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11310

FloorSAV: 2Dフロアマップで空間的な音声視覚コンテキストを解明するAV-LLM

FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs

シェア:XThreadsFacebookLINEはてブBluesky

動的な一人称視点環境において、3D点群やカメラ軌跡、空間音声、物体ランドマークから2Dフロアマップを生成し、AV-LLMに同期入力することで空間推論能力を向上させるフレームワークを提案。

詳しい要約

1. どんなもの?

- 動的なegocentric環境における3D空間推論のためのAV-LLMフレームワーク。 - 2D floormapを動的にレンダリングし、空間音声視覚コンテキストを明示的に接地。 - 3D point clouds、camera trajectories、spatial audio cues、semantically grounded object landmarksを統合。 - floormapをegocentric videoと同期したストリームとしてAV-LLMに注入。 - AV-LLMのマルチモーダル能力を活用し、単一推論で視覚・聴覚・幾何的手がかりを統合推論。 - SAVED-Benchを導入し、動的相対性、地域、経路推論QAを構築。

2. 先行研究と比べてどこがすごい?

- 既存手法は高コストなfine-tuningを必要とするか、モデルのcross-modal reasoning能力を十分に活用していない。 - FloorSAVはfine-tuning不要で、floormapを介して空間情報を明示的に注入。 - AV-LLMのマルチモーダル推論能力を最大限に引き出し、空間推論を改善。 - SAVED-BenchとSAVVY-Benchの両方で空間推論性能を向上。

3. 技術・手法の肝は?

- 3D point clouds、camera trajectories、spatial audio cues、semantically grounded object landmarksを統合し、動的2D floormapをレンダリング。 - floormapをegocentric videoと同期したストリームとしてAV-LLMに注入。 - floormap解釈ガイダンスを提供し、AV-LLMが視覚・聴覚・幾何学的手がかりを単一推論で統合推論。 - これにより、生のセンサストリームからグローバルジオメトリを直接処理・内部化。

4. どうやって有効だと検証した?

- SAVED-Bench(動的相対性、地域、経路推論QA)とSAVVY-Benchで評価。 - FloorSAVがAV-LLMの空間推論を様々なタスクで改善することを示す。 - ground-truth floormapを用いた研究により、正確な空間情報があればFloorSAVが大きな可能性を持つことを実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- SAVVY-Bench(関連ベンチマーク) - 3D spatial reasoning in dynamic egocentric environments(関連分野) - AV-LLMs(関連モデル)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kyeong-Rae Kim, Sungnyun Kim, Tae-Hyun Oh

分類: cs.CV, cs.LG

原文アブストラクト

While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.

関連論文

PR本紙発行元 EmplifAI