日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30715

RoboMonitor: 予測表現学習によるラベル効率的なロボットタスク実行モニタリング

RoboMonitor: Label-Efficient Runtime Monitoring of Robot Task Execution via Predictive Representation Learning

シェア:XThreadsFacebookLINEはてブBluesky

ロボット政策学習用データセットを活用し、行動条件付き未来特徴予測などで事前学習した視覚言語モデルを因果モニタに転移することで、少数のラベル付きエピソードで実行フェーズ・失敗・完了を高精度に監視できる手法を提案。

詳しい要約

1. どんなもの?

- ロボットのタスク実行を監視する vision-language モニタ「RoboMonitor」 - 実行中の観測から現在の実行フェーズ特定、失敗検出、タスク完了認識を行う - 学習には robot-policy learning 用データセットを活用し、監視用アノテーションを最小化 - 展開時は task instruction と camera observations のみを必要とする

2. 先行研究と比べてどこがすごい?

- 同じ監視 supervision で学習した Qwen3-VL と Robometer を上回る - 52 labeled episodes で 93.1% mean phase accuracy、85.9% macro recall を達成 - この phase accuracy は 100 episodes で学習した Qwen3-VL と Robometer も上回る - Temporal SFT により mean spurious phase switching を 15.23% から 4.95% に低減

3. 技術・手法の肝は?

- 25 時間・12 manipulation tasks・2 robot embodiments の multi-camera trajectories で事前学習 - action-conditioned future-feature prediction、inverse dynamics、masked-present prediction を使用 - 学習した visual/context encoders を causal monitor に転移 - Temporal SFT で各 observation window 全体の supervision と、window 内・window 間の consistency objectives を組み合わせる

4. どうやって有効だと検証した?

- 4 タスクの monitoring benchmark で評価 - 2 つの fine-tuning seeds で 52 labeled episodes を用いて学習し、mean phase accuracy 93.1%、macro recall 85.9% - Qwen3-VL ablation で Temporal SFT が spurious phase switching を 15.23% から 4.95% に削減 - closed-loop deployment で simulated Toolbox Sorting は 39/40、real-world Reel Packing は 35/40 を完了し、false recovery triggers は観測されず

5. 議論はある?

- 要旨からは不明 - 限界、失敗事例、計算コスト、一般化可能性についての議論は要旨に記載なし

6. 次に読むべき論文は?

- Qwen3-VL - Robometer - robot-policy learning 用データセット - action-conditioned future-feature prediction、inverse dynamics、masked-present prediction に関連する手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Abhiroop Ajith, Gokul Narayanan, Kyle Coelho, Tingji Zhao, Yash Shahapurkar, Brian Zhu, Melih Erdogan, Ted Krubasik, Constantinos Chamzas, Eugen Solowjow

分類: cs.RO

原文アブストラクト

Learned robot policies produce actions, but their outputs alone do not establish whether execution is progressing as intended. Robot execution monitoring requires identifying the current execution phase, detecting failures, and recognizing task completion from observations available during execution. Training such monitors requires annotations that are scarce in datasets collected for robot-policy learning. We present RoboMonitor, a label-efficient vision--language execution monitor that learns from these datasets before introducing monitoring supervision. We pre-train on 25 hours of multi-camera trajectories spanning 12 manipulation tasks and two robot embodiments, using action-conditioned future-feature prediction, inverse dynamics, and masked-present prediction. We then transfer the learned visual and context encoders to a causal monitor and apply temporal supervised fine-tuning (Temporal SFT), which combines supervision throughout each observation window with consistency objectives within and across overlapping windows. At deployment, RoboMonitor requires only the task instruction and camera observations. On a four-task monitoring benchmark, RoboMonitor trained with 52 labeled episodes achieves 93.1% mean phase accuracy and 85.9% macro recall over two fine-tuning seeds, exceeding Qwen3-VL and Robometer trained with the same monitoring supervision. Its phase accuracy also exceeds that of both Qwen3-VL and Robometer trained with 100 episodes. A Qwen3-VL ablation shows that Temporal SFT reduces mean spurious phase switching from 15.23% to 4.95%. In closed-loop deployment, the integrated system completes 39 of 40 simulated Toolbox Sorting trials and 35 of 40 real-world Reel Packing trials, with no false recovery triggers observed.

関連論文

PR本紙発行元 EmplifAI