日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLMarXiv:2609.30629

FreshLatent: 資源制約下の身体性VLM知覚のためのチャネル適応型潜在アダプテーション

FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception

シェア:XThreadsFacebookLINEはてブBluesky

無線通信で劣化した中間特徴に適応する軽量なチャネル認識型潜在アダプタを提案し、VLM本体を凍結したまま省資源でロバストなUAV知覚を実現した。

詳しい要約

1. どんなもの?

- 資源制約下のUAV向けsplit VLM知覚のための軽量なchannel-aware latent adapter「FreshLatent」を提案。 - 伝送される中間特徴の劣化に対し、周囲のVLMを凍結したままpower-normalized encoder-decoderを無線劣化下で訓練。 - ミッション条件付き知覚要件と組込みインタフェースコストに基づき、知覚が使用可能な動作条件を定式化。

2. 先行研究と比べてどこがすごい?

- clean-trained split interfaceは伝送特徴の劣化でdeployment mismatchを生じ、channel-aware codecは組込みコストが大きい。 - FreshLatentは軽量ながら、0 dB・最小通信予算でclean split compression比gIoU +20.79、cIoU +20.87点。 - 0 dBで重いrange-trained feature-JSCC codecのgIoU改善の63.5-69.1%を回復。 - Jetson AGX Xavier 10-Wでencoderパラメータ37-40x少、edge-interface遅延7.7-9.9x低、エネルギー8.8-10.0x低。

3. 技術・手法の肝は?

- 周囲のVLMを凍結し、power-normalized encoder-decoderのみを無線corruption下で訓練するchannel-aware latent adapter。 - ミッション条件付き知覚要件とembedded interface costを組み合わせ、channel qualityとcommunication budgetを使用可能条件に結びつける定式化。 - 詳細なネットワーク構成や訓練損失は要旨からは不明。

4. どうやって有効だと検証した?

- 0 dBおよび最も厳しい通信予算でgIoUとcIoUをclean split compressionと比較。 - 最も不利なSNR 0 dBで3つの通信予算すべてにおいて、重いrange-trained feature-JSCC codecのgIoU改善に対する回復率を評価。 - NVIDIA Jetson AGX Xavier 10-Wモードでencoderパラメータ数、edge-interface遅延、edge-interfaceエネルギーを比較。

5. 議論はある?

- 軽量なchannel-aware adaptationが大規模通信インタフェースの堅牢性の相当部分を回復し、制約無線下で品質有効動作を広げることを示す。 - 限界や失敗条件、他のタスク・モデルへの一般化については要旨からは不明。

6. 次に読むべき論文は?

- 比較対象のrange-trained feature-JSCC codec。 - clean split compression。 - split VLM perception、feature-JSCC、channel-aware codecの関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Rajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu, Eli Bozorgzadeh, Marco Levorato, Nikil Dutt

分類: eess.SP, cs.CV, cs.DC, cs.LG, cs.RO

原文アブストラクト

Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while keeping the surrounding VLM frozen. We formulate deployment around a mission-conditioned perception requirement and embedded interface cost, linking channel quality and communication budget to the operating conditions under which perception remains usable. At 0 dB and the tightest communication budget, FreshLatent improves gIoU and cIoU over clean split compression by 20.79 and 20.87 points, respectively. At the most adverse evaluated SNR (0 dB), across all three communication budgets, FreshLatent recovers 63.5-69.1% of the gIoU improvement achieved by a much heavier, range-trained feature-JSCC codec. On an NVIDIA Jetson AGX Xavier in 10-W mode, FreshLatent uses 37-40x fewer encoder parameters, 7.7-9.9x lower edge-interface latency, and 8.8-10.0x lower edge-interface energy than the heavier codec. Together, these results show that lightweight channel-aware adaptation can recover a substantial fraction of the robustness of a much larger communication interface while broadening quality-valid operation under constrained wireless conditions.

関連論文

PR本紙発行元 EmplifAI