日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.21369

ProTracer: 固有感覚誘導によるロボットマニピュレーションの失敗診断

ProTracer: Proprioception-Guided Failure Diagnosis in Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

固有感覚信号と視覚言語モデルを組み合わせ、訓練なしでロボット操作の失敗検出・分類・説明・発生時刻特定を行うフレームワークを提案。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションの失敗分析フレームワーク - 二値失敗検出、失敗カテゴリ分類、説明生成、失敗開始位置特定 - 失敗開始位置特定は、有効なタスク完了軌道から逸脱し最終的に失敗に至る最早の時点を同定 - ProTracerは訓練不要のフレームワークで、既存のVision-Language Models (VLMs)とproprioceptive signalsを活用 - FailTimeベンチマークを導入し、視覚とproprioceptiveの同期観測を提供

2. 先行研究と比べてどこがすごい?

- 従来の失敗診断は視覚のみに依存し、時間的精度が不足 - ProTracerはproprioceptive dynamicsを活用し、時間的に情報量の多いaction boundariesを特定 - 訓練不要で既存VLMを利用し、追加のモデル訓練を必要としない - 失敗開始位置特定という新タスクを導入し、細粒度の時間的失敗分析を可能に - 視覚とproprioceptiveのマルチモーダル推論を組み合わせ、従来手法より高い性能を実現

3. 技術・手法の肝は?

- proprioceptive dynamicsを用いてaction boundariesを特定 - 豊富なrobot-state signalsを構造化自然言語記述に変換 - VLMが視覚観測とproprioceptive記述を統合分析 - 訓練不要で、既存VLMのマルチモーダル推論能力を活用 - 時間的精度(proprioceptive)とマルチモーダル推論(VLM)を組み合わせ

4. どうやって有効だと検証した?

- FailTimeベンチマークを導入 - 視覚とproprioceptiveの同期観測を含む - 従来の失敗診断タスクと新規の失敗開始位置特定タスクで評価 - ProTracerが両タスクで強い性能を達成 - proprioceptive reasoningの重要性を実証

5. 議論はある?

- proprioceptive reasoningが細粒度の時間的失敗分析に重要 - 訓練不要フレームワークの有効性を示す - 失敗開始位置特定の新タスクの意義 - 限界や課題については要旨からは不明

6. 次に読むべき論文は?

- Vision-Language Models (VLMs) を用いたロボット失敗分析の関連研究 - proprioceptive signalsを活用したロボット学習 - 失敗検出・診断のベンチマーク(例: FailTime) - マルチモーダル推論によるロボットマニピュレーション - 具体的な参照論文は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chang Dong, Mehdi Hosseinzadeh, King Hang Wong, Lingqiao Liu, Francois Fraysse, Feras Dayoub, Minh Hoai Nguyen

分類: cs.RO, cs.CV

原文アブストラクト

This paper presents a comprehensive framework for robot manipulation failure analysis that includes binary failure detection, failure categorization, explanation generation, and the additional capability of failure onset localization, which aims to identify the earliest moment at which a robot execution deviates from a valid task-completion trajectory and is ultimately followed by task failure. To address these tasks, we propose ProTracer, a training-free framework that leverages existing Vision-Language Models (VLMs) together with proprioceptive signals for failure analysis. Our method uses proprioceptive dynamics to identify temporally informative action boundaries and converts richer robot-state signals into structured natural-language descriptions that can be jointly analyzed together with visual observations by the VLM. This design combines the temporal precision of proprioceptive signals with the multimodal reasoning capabilities of modern VLMs without requiring additional model training. We further introduce FailTime, a benchmark with synchronized visual and proprioceptive observations for evaluating conventional failure diagnosis tasks as well as failure onset localization. Experiments demonstrate that ProTracer achieves strong performance across both conventional failure diagnosis tasks and the newly introduced failure onset localization task, highlighting the importance of proprioceptive reasoning for fine-grained temporal failure analysis.

関連論文

PR本紙発行元 EmplifAI