日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
能動知覚arXiv:2609.23974

LEAP-NBV: 基盤モデルによる次善視点計画のための軽量エッジ能動知覚

LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning

シェア:XThreadsFacebookLINEはてブBluesky

大規模なHuman Mesh Recoveryモデルを蒸留・量子化してエッジデバイス上で動作させ、オクルージョン対応の能動知覚ループで次善視点計画をリアルタイム実行する軽量フレームワークを提案。

詳しい要約

1. どんなもの?

- 基盤モデル駆動のNext-Best-View (NBV) 計画をエッジデバイス上で実行する軽量アクティブ知覚フレームワークLEAP-NBVを提案。 - Human Mesh Recovery (HMR) を対象に、大規模教師モデルを32Mの学生モデルに蒸留し、視覚エンコーダをFP16量子化。 - オクルージョン対応のアクティブ知覚ループ内で評価し、NVIDIA Jetson Xavier NX上でエンドツーエンドのパイプラインを展開。 - オンデバイスのレイテンシとエネルギーを測定し、リアルタイム性能と省エネを実現。

2. 先行研究と比べてどこがすごい?

- 従来の基盤モデルはサイズと電力要件が大きく、エッジプラットフォームでの実行が困難でリアルタイム性能が制限されていた。 - 特に競合通信下で計算をオフロードできない戦術エッジ展開の要件を満たせなかった。 - LEAP-NBVは蒸留と量子化により、非圧縮モデルと比較して2.0倍の高速化と3.0倍のエネルギー削減を達成し、下流タスク品質をほぼ維持。 - エッジデバイス上で完全な閉ループを3.6 FPS、1フレームあたり2.6 Jで実行可能にした点が革新的。

3. 技術・手法の肝は?

- 大規模HMR教師モデル群を、オフラインメッシュ目的関数を用いてコンパクトな32M学生モデルに蒸留。 - 視覚エンコーダをFP16に量子化し、オンデバイスでの精度とレイテンシを特性評価。 - オクルージョン対応のアクティブ知覚ループ内で、全ての構成を同一のホールドアウトベンチマークで評価。 - エッジ最適な圧縮モデルを選択し、HMRエンジンを約12 msで動作させ、完全な閉ループを実現。

4. どうやって有効だと検証した?

- 同一のホールドアウトベンチマーク上で全ての構成を評価。 - 蒸留により、未蒸留学生モデルと比較してProcrustes-aligned mean per-vertex position error (PA-MPVPE) が6-7 mm改善。 - NVIDIA Jetson Xavier NX上でエンドツーエンドパイプラインを展開し、オンデバイスレイテンシとエネルギーを測定。 - エッジ最適圧縮モデルが約12 msで動作し、完全閉ループを3.6 FPS、2.6 J/フレームで実行、非圧縮モデル比2.0倍高速化、3.0倍省エネを達成し、下流タスク品質をほぼ維持。

5. 議論はある?

- 蒸留と量子化により精度を小さなコストで大幅な高速化と省エネを実現。 - エッジデバイス上でのリアルタイムアクティブ知覚の実現可能性を示す。 - 下流タスク品質をほぼ維持しつつ、戦術エッジ展開の要件を満たす。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- Human Mesh Recovery (HMR) に関する研究。 - Next-Best-View (NBV) 計画に関する研究。 - 知識蒸留 (knowledge distillation) やモデル量子化 (quantization) に関する研究。 - エッジデバイス上での基盤モデル展開に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Boxun Hu, Jiawei Ge, Axel Krieger, Peng Wang, Tinoosh Mohsenin

分類: cs.AI

原文アブストラクト

Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human Mesh Recovery (HMR), which provides useful estimates of a target's 3D pose and shape that can benefit tactical missions. However, the size and power demands of such models make them difficult to run on edge platforms and limit their real-time performance, undermining the requirements of tactical edge deployment - especially for active perception, where a mobile robot must plan its next-best view on-board and cannot offload computation under contested communications. We present LEAP-NBV, a lightweight active-perception framework that runs foundation-model-driven Next-Best-View (NBV) planning on-board an edge device. To this end, we distill a family of large HMR teachers, each into a compact 32M student, with an offline mesh objective, then quantize the vision encoder to FP16 and characterize its on-device accuracy and latency. Within an occlusion-aware active perception loop, we evaluate all configurations on the same held-out benchmark and deploy the end-to-end pipeline on an NVIDIA Jetson Xavier NX, reporting measured on-device latency and energy. Distillation recovers 6-7 mm of Procrustes-aligned mean per-vertex position error (PA-MPVPE) over the undistilled student on the test set. Selecting the edge-optimal compression model brings the HMR engine to ~12 ms at a small accuracy cost and runs the full closed loop at 3.6 FPS and 2.6 J per frame, achieving a 2.0x speedup and 3.0x lower energy than the uncompressed model while nearly matching downstream task quality.

関連論文

PR本紙発行元 EmplifAI