日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D視覚arXiv:2610.07982

M3SunAgent: 単眼3D空間理解のためのエージェント

M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

シェア:XThreadsFacebookLINEはてブBluesky

LLMをタスクプランナーとして用い、単眼深度推定と3D視覚グラウンディングを統合的に行うエージェントを提案し、ベンチマークデータセットも構築した。

詳しい要約

1. どんなもの?

- 単眼3D空間理解(M3Sun)のための統合エージェントM3SunAgentを提案 - 対象はmonocular metric depth estimationと3D visual groundingの2タスク - LLMをtask plannerとして使い、structured programを生成しツールを調整 - instance-level depth estimationと3D bounding box予測を単一枠組みで実行 - 評価用ベンチマークM3Sun Instance(M3SI)を構築(2,910サンプル)

2. 先行研究と比べてどこがすごい?

- 従来は2つの相補的タスクを別々のframeworkで扱い、空間情報が不柔軟・不整合 - 本研究は単一agentで両タスクを統合し、embodied intelligence向けの一貫した空間情報取得を狙う - instance-level depth estimationで比較モデル中最高性能(δ<0.25が52.61%) - 3D visual groundingでmIoU 41.73%、SOTAのMonoVLMを3.62%上回る

3. 技術・手法の肝は?

- LLMをtask plannerとしてspatial visual programmingを実行 - 構造化programを柔軟に生成し、複数ツールをcoordinatedに呼び出す - instance-level depth: object detectorで対象を定位→depth estimation toolで選択点の深度推定→集約 - 3D visual grounding: VLM toolで対象定位と基本空間属性出力→back-projection toolとdimension-lifting toolで3D bounding box予測

4. どうやって有効だと検証した?

- M3Sun Instance(M3SI)ベンチマーク2,910サンプルを構築し評価 - instance-level monocular metric depth estimationで全比較モデル中最高、δ<0.25が52.61% - monocular 3D visual groundingでmIoU 41.73%、MonoVLMを3.62%上回る - visionモデルおよびVLMモデルと比較し総合的に競争力のある性能を確認

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例、計算コスト、ツール依存性などの議論は要旨に記載なし

6. 次に読むべき論文は?

- MonoVLM(3D visual groundingのSOTA比較対象) - monocular metric depth estimationの代表的手法 - 3D visual groundingのvision/VLMベース手法 - LLMをtask plannerとするvisual programming関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinsong Zhang, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, Zhengguo Li

分類: cs.CV

原文アブストラクト

Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($δ< 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.

PR本紙発行元 EmplifAI