日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02021

2D VLMを3Dプログラマに変えるタスク適応型グラウンディング

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

シェア:XThreadsFacebookLINEはてブBluesky

2Dの視覚言語モデルに座標系の統一とタスク適応フィードバックを組み合わせ、再学習なしで3D理解・操作・生成を可能にするフレームワーク3D-Progを提案。

詳しい要約

1. どんなもの?

- 2D VLMを再学習なしで3Dタスクに適応させる枠組み。 - 新概念Canonical Coordinate Framing (CCF)とTask-Adaptive Feedback (TAF)を導入。 - 3D-Progという3D理解・推論・生成フレームワークを提案。 - オープンボキャブラリの3D理解、操作、生成を物体レベルとシーンレベルで実行。

2. 先行研究と比べてどこがすごい?

- 従来はVLMを3Dに拡張する際、データ規模・学習多様性・推論能力の制約があった。 - 本研究はモデルを3D化せず、2D VLMを3Dで信頼できるよう動作させる点が異なる。 - CCFとTAFの併用により、再学習不要で幾何認識型3Dプログラマに変換。 - 一貫性・解釈性・高品質な結果を多様な3Dタスクで実現。

3. 技術・手法の肝は?

- CCF: 入力と出力を共有ユークリッド座標系に固定する統一視覚表現。 - 軸の曖昧さ、メートルスケールの不一致、浮動参照を解決。 - TAF: タスク適応型動的フィードバックで推論ループを閉じる。 - 2D VLMがネイティブ視覚文脈内で多様なオープンボキャブラリタスクを実行可能。 - 3D-Prog: CCFとTAFを強力なVLMと組み合わせた3D理解・推論・生成フレームワーク。

4. どうやって有効だと検証した?

- 物体レベルとシーンレベルの多様な3Dタスクで実験。 - CCFとTAFの併用が2D VLMを幾何認識型3Dプログラマに変換することを示す。 - 一貫性、解釈性、高品質な結果が得られることを確認。 - 具体的なデータセット名や評価指標は要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、スケーラビリティに関する議論は記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として2D VLM (Vision-Language Models) や3D grounding、open-vocabulary 3D understandingの定番研究が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel

分類: cs.CV, cs.AI

原文アブストラクト

Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.

関連論文

PR本紙発行元 EmplifAI