日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37519

Video2STL: VLM生成の時間仕様をロボット学習に接地するフレームワーク

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

シェア:XThreadsFacebookLINEはてブBluesky

観測のみの動画からVLMで時間論理仕様(STL)を生成し、ロボット軌道で閾値を接地して強化学習の報酬に用いる手法。4つのマニピュレーションタスクで平均85.8%の成功率を達成。

詳しい要約

1. どんなもの?

- 観測のみの動画をパラメトリックな Signal Temporal Logic (STL) 仕様に変換し、ロボット学習に用いる Video2STL を提案。 - VLM が embodiment 非依存の意味イベントトレースを抽出し、記号的時間仕様のバンクを構築。 - タスク構造はモデルが決定し、述語の閾値と時間境界は成功したロボット軌道から接地。 - 短時間仕様は rolling-window の定量的 robustness で dense reward を、長時間仕様は causal monitor で一度きりの progress reward を与える。 - 人間や動物の動画からロボット制御への cross-embodiment 転移も可能。

2. 先行研究と比べてどこがすごい?

- 既存手法は視覚観測をスカラー類似度や価値信号に変換するか、foundation model に直接 reward code を生成させるものが多い。 - それらはタスクの時間構造を検査・接地・再利用しにくい。 - Video2STL は形式的な STL 表現を用いることで、時間構造を明示的かつ再利用可能にする。 - 4つの manipulation タスクで平均 success-once 85.8%、success-at-end 67.0% を達成。 - 比較対象の native dense PPO は 81.5%/59.5%、Text2Reward は 65.0%/42.3%。 - quadruped locomotion では Qwen-3.8 と GPT-5.6 ベースの Video2STL が 0.3〜2.1 m/s で 100% 成功。

3. 技術・手法の肝は?

- VLM が動画から embodiment 非依存の意味イベントトレースを抽出。 - そのトレースから記号的時間仕様のバンクを構築し、タスク構造を決定。 - 述語の数値閾値と時間境界は成功したロボット軌道から接地。 - 短時間仕様は rolling-window の定量的 robustness により dense reward を提供。 - 長時間仕様は causal monitor により有効な時間プレフィックスに対する一度きりの progress reward を提供。 - 同一表現が cross-embodiment 転移を支援。

4. どうやって有効だと検証した?

- 4つの manipulation タスクで評価し、平均 success-once 85.8%、success-at-end 67.0% を達成。 - native dense PPO (81.5%/59.5%) および Text2Reward (65.0%/42.3%) と比較。 - quadruped locomotion では Qwen-3.8 と GPT-5.6 ベースの Video2STL が 0.3〜2.1 m/s の速度で 100% 成功。 - 高速域のエネルギー効率でも競争力を維持。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Text2Reward - native dense PPO - Signal Temporal Logic (STL) 関連研究 - vision-language model を用いた reward 生成研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh

分類: cs.RO, cs.AI

原文アブストラクト

Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.

関連論文

PR本紙発行元 EmplifAI