日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08713

SpaTime: 時空間推論のためのストリーミング視覚言語モデル

SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

動画が到着する途中でも因果的幾何トークンを融合し、応答時間損失で適切な応答タイミングを学習するストリーミングVLMを提案。

詳しい要約

1. どんなもの?

- ストリーミングVision-Language Model (VLM) であるSpaTimeを提案。 - ビデオが到着している最中に3D空間推論を行い、十分観察した時点で回答する。 - 各フレームでcausal geometry tokensを言語モデルに融合。 - 応答タイミングを学習するresponse-time lossを導入。 - StreamVSTI-BenchとStreamVSI-Benchを構築し評価。

2. 先行研究と比べてどこがすごい?

- 3D幾何事前知識を持つVLMはオフラインで全ビデオが必要。 - ストリーミングVLMは因果的処理と応答決定が可能だが3D表現が不足。 - SpaTimeは両者を統合し、ストリーミング設定で3D空間推論を実現。 - 最強のストリーミングベースラインと比較し、平均応答時間誤差を66%削減。 - StreamVSTI-Benchで49.2%の総合精度を達成。

3. 技術・手法の肝は?

- 各フレームでcausal geometry tokensを抽出し、言語モデルに融合。 - 観測済みフレームのみを使用する因果的処理。 - response-time loss: フレームごとの応答確率を微分可能な期待応答時間に写像し、正解フレームとの距離をペナルティ。 - これによりモデルが自ら応答タイミングを学習。

4. どうやって有効だと検証した?

- StreamVSTI-BenchとStreamVSI-Benchを構築。 - これらはVSTI-BenchとVSI-Benchのストリーミング適応版。 - StreamVSTI-Benchで49.2%の総合精度を達成。 - 最強のストリーミングベースラインと比較し、平均応答時間誤差を66%削減。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- VSTI-Bench, VSI-Bench, およびストリーミングVLMの関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hairong Yin, Huangying Zhan, Shin-Fang Chng, Yi Xu, Raymond A. Yeh

分類: cs.CV

原文アブストラクト

Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.

関連論文

PR本紙発行元 EmplifAI