日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04727

衝突予測のためのビデオトランスフォーマーにおける時空間冗長性の調査

Investigating Spatiotemporal Redundancy in Video Transformer for Collision Anticipation

シェア:XThreadsFacebookLINEはてブBluesky

衝突予測に用いるVideoMAEの計算冗長性を解析し、時間方向のトークン統合で精度を保ちつつ1.76倍高速化した研究。

詳しい要約

1. どんなもの?

ビデオトランスフォーマー(VideoMAEv2-Base)を用いた衝突予測において、計算の冗長性を調査し、性能を大きく損なわずに削減する方法を検討した研究。Nexar Collision Prediction datasetを使用し、MLPの構造的容量冗長性とトークン列の時空間冗長性の2種類を分析。

2. 先行研究と比べてどこがすごい?

先行研究と比べて、衝突予測におけるビデオトランスフォーマーの計算冗長性を詳細に分析し、特に時間方向のトークン冗長性が空間方向より大きいことを定量的に示した点が新しい。また、時間トークンマージにより1.76倍の高速化を達成しつつmAPの低下を0.0035に抑えた。

3. 技術・手法の肝は?

VideoMAEv2-Baseの各層で線形プローブを用いて衝突関連情報の進化を調査。MLPの構造的容量冗長性とトークン列の時空間冗長性を分析。時間トークンの類似度が後半層で0.970に達する一方、空間類似度は0.297に低下する非対称性を利用し、時間トークンマージを適用。また、重要度に基づくMLPユニットの50%保持を実施。

4. どうやって有効だと検証した?

Nexar Collision Prediction datasetで評価。時間トークンマージによりバックボーン計算が356.99から178.50 GFLOPsに、レイテンシが12.69から7.21 ms/clipに削減され、mAPは0.7478から0.7443にわずかに変化。MLPユニットの重要度保持50%でAUC 0.753を維持(ランダム保持では0.529)。

5. 議論はある?

時間トークンマージとニューロンプルーニングが有効であることを示し、より高速なアルゴリズムへの新たな道筋を提示。MLPプルーニングはスコア較正を変化させるが識別ランキングが崩壊する前に影響を与えることを明らかにした。

6. 次に読むべき論文は?

VideoMAEv2-Base、Nexar Collision Prediction dataset、および関連するビデオトランスフォーマーの効率化手法(例:Token Merging, MLP pruning)に関する論文。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaoshan Zhou

分類: cs.CV

原文アブストラクト

In worker-equipment proximity monitoring, video transformers are widely used for collision anticipation and have demonstrated strong performance. However, their accuracy comes with substantial computational demands, creating a tension with the need for low-latency inference on mobile robots and the pursuit of lower-carbon computation in construction. To address this, this study investigates where computation within an established video transformer is redundant and whether that redundancy can be removed without materially degrading predictive performance. Using VideoMAEv2-Base on the Nexar Collision Prediction dataset, we first examine how collision-relevant information evolves across network depth and then investigate two complementary forms of redundancy: structured capacity redundancy in multilayer perceptrons (MLPs) and spatiotemporal redundancy in the token stream. Linear probes show that interpretable motion cues, including flow magnitude, looming, and approach versus retreat, are most accessible at intermediate layers, whereas collision-label discrimination strengthens toward the final layer. Token redundancy is axis-specific: adjacent temporal-token similarity reaches 0.970 in later layers, while spatial similarity falls to 0.297, indicating substantially greater redundancy across time than across space. Exploiting this asymmetry, temporal token merging reduces backbone computation from 356.99 to 178.50 GFLOPs and latency from 12.69 to 7.21 ms per clip, a 1.76x speedup, while mean average precision changes only from 0.7478 to 0.7443. Importance-guided retention of 50% of MLP units preserves an AUC of 0.753, compared with 0.529 under matched random retention, and reveals that pruning alters score calibration before discriminative ranking collapses. These findings establish a new pathway for pursuing faster algorithms through targeted temporal token compression and neuron pruning.

関連論文

PR本紙発行元 EmplifAI