日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37655

Exemplar2VQA:マルチエージェントコーディングによる模範駆動型視覚質問応答生成フレームワーク

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

シェア:XThreadsFacebookLINEはてブBluesky

マルチエージェントコーディングにより、幾何学的ユーティリティを活用して大規模な空間QAデータを合成し、MLLMの空間知能を向上させるフレームワークを提案。

詳しい要約

1. どんなもの?

- 複雑でスケーラブルな3D質問応答(QA)データの不足を解決するため、模擬環境で大規模な空間QAペアを迅速に合成するフレームワーク。 - マルチエージェントコーディングを活用し、例示駆動型で視覚質問応答を生成。 - 静的なオブジェクト中心の空間クエリテンプレートを例示として与え、自律的に大規模で高忠実度の合成データセットに拡張。

2. 先行研究と比べてどこがすごい?

- 手動アノテーションは労力が大きく、LLMを直接使ったQA合成は空間・幾何計算の欠陥により失敗しがち。 - 本手法は決定論的コード実行を通じてLLMの空間推論の欠陥を回避。 - 多様な静的な例示テンプレートから自律的にスケール可能で、屋内だけでなく屋外や混合シーンベンチマークにも効果が拡張。

3. 技術・手法の肝は?

- 協調エージェントに幾何ユーティリティのライブラリを装備。 - 決定論的コード実行により、LLMの空間推論の欠陥をバイパス。 - 例示駆動型で、静的なオブジェクト中心の空間クエリテンプレートを大規模な合成データセットに自律的に拡張。

4. どうやって有効だと検証した?

- Qwen2.5-VL (3B/7B)をExemplar2VQA生成の合成屋内データのみでファインチューニング。 - 多様なベンチマークで大幅な性能向上を確認。 - 屋内データセットに限らず、屋外や混合シーンのベンチマークにも効果が頑健に拡張。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。同分野の定番として、3D視覚質問応答(VQA)やEmbodied AIにおけるsim-to-realギャップに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin Tan

分類: cs.CV

原文アブストラクト

Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA

関連論文

PR本紙発行元 EmplifAI