日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12266

DVD: 動的ベクトルデコーディングによる効率的なMLLMベース知覚

DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

シェア:XThreadsFacebookLINEはてブBluesky

2D/3Dの知覚表現を1次元ベクトル列に変換し、コンパクトな離散トークンとしてMLLMで扱うことで、トークン数と推論遅延を削減しつつ高精度な知覚を実現する手法を提案。

詳しい要約

1. どんなもの?

- MLLMベースの知覚タスク(2D/3D)向けの新しいデコーディング手法「DVD (Dynamic Vector Decoding)」を提案。 - 2D bounding boxes, 2D masks, 3D bounding boxes を1Dベクトル系列に変換し、高次元空間でコンパクトな離散トークンにマッピング。 - 軽量な de-tokenizer により、MLLMの出力トークンを元の2D/3D知覚表現にデコード。 - 2Dと3Dの知覚タスクを統一的に表現し、トークンオーバーヘッドと推論レイテンシを大幅に削減。

2. 先行研究と比べてどこがすごい?

- 既存のMLLMベース知覚手法は、テキストベースの座標表現(トークンオーバーヘッド大)や固定範囲量子化(範囲と精度の制約)に依存。 - 特に3D領域では空間範囲が非有界で高い位置精度が要求されるため、これらの手法は限界がある。 - DVDはこれらの問題を克服し、2D/3Dタスクで優れた性能を達成しつつ、トークンオーバーヘッドと推論レイテンシを大幅に削減。

3. 技術・手法の肝は?

- 多様な知覚表現(2D bounding boxes, 2D masks, 3D bounding boxes)を1Dベクトル系列に変換。 - そのベクトル系列を高次元空間でコンパクトな離散トークンにマッピング。 - 軽量な de-tokenizer がMLLMの出力トークンを元の2D/3D知覚表現にデコード。 - これによりMLLMとのシームレスな統合を実現。

4. どうやって有効だと検証した?

- 2Dおよび3D知覚ベンチマーク(RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes)で広範な実験を実施。 - DVDは2D/3Dタスクで優れた性能を達成し、トークンオーバーヘッドと推論レイテンシを大幅に削減することを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていないが、関連手法としてテキストベース座標表現や固定範囲量子化を用いたMLLMベース知覚手法が挙げられる。 - 同分野の定番として、MLLMベースの2D/3D知覚(例:RefCOCO, SUN-RGBD, KITTI, Hypersim, nuScenes を用いた研究)が次に読むべき論文として考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinghua Hou, Zhe Liu, Hengshuang Zhao

分類: cs.CV, cs.AI

原文アブストラクト

Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.

関連論文

PR本紙発行元 EmplifAI