日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.21982

CARF: 失敗誘導フローマッチングの対比的引力-斥力

CARF: Contrastive Attraction-Repulsion of Failure-Guided Flow Matching

シェア:XThreadsFacebookLINEはてブBluesky

失敗軌道から「進展する部分」は模倣し「失敗を招く部分」は回避する非対称な学習を、フローマッチングで統一的に実現する手法を提案。

詳しい要約

1. どんなもの?

- ロボットのデモンストレーション収集では、成功軌道だけでなく失敗軌道も生じる。 - 既存手法は失敗軌道から進捗のあるセグメントを模倣するが、失敗に直結する failure-critical behaviors を見落としている。 - 本論文は、progressive segments は模倣し、failure-critical segments は明示的に回避すべきという非対称な監督の考えを提案。 - これを実現する CARF (Contrastive Attraction-Repulsion of Failure-guided framework) を提案。 - 不完全なロボットデータから学習する枠組み。

2. 先行研究と比べてどこがすごい?

- 既存手法は失敗軌道から進捗セグメントのみを利用し、failure-critical behaviors を無視していた。 - CARF は progressive と failure-critical を非対称に扱い、模倣と回避を同時に行う。 - 曖昧なセグメントを除外することで、信頼できない監督を避ける。 - 不完全データをより包括的に活用できる点が先行研究と異なる。 - シミュレーションと実世界で一貫した改善を示す。

3. 技術・手法の肝は?

- progress-based importance scorer を導入。 - 成功した expert demonstrations とその perturbation 結果のみで scorer を訓練。 - 各ステップのタスク完了への寄与を推定し、失敗軌道中の informative regions を特定。 - スコアが unified flow-matching objective を導き、progressive behaviors へ attract、failure-critical から repel。 - 曖昧なセグメントは除外。

4. どうやって有効だと検証した?

- シミュレーションと実世界での広範な実験を実施。 - 多様な failure scenarios で競合ベースラインに対し一貫した改善を確認。 - ablation により scoring と attraction-repulsion メカニズムの有効性を検証。 - 詳細は要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - ただし、progressive と failure-critical の非対称監督の有効性、曖昧セグメント除外の影響、scorer の汎化性などが議論され得る。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として flow matching、imitation learning、learning from imperfect demonstrations、contrastive learning が挙げられる。 - 同分野の定番として Behavior Cloning、DAgger、GAIL などが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuqi Zhao, Bang Du, Cheng-En Wu, Yichen Xie, Yixiao Wang, Masayoshi Tomizuka

分類: cs.RO

原文アブストラクト

Robot demonstration collection often produces imperfect or failed trajectories in addition to successful demonstrations. Existing methods typically exploit failed trajectories by identifying segments that still make progress toward task completion, but largely overlook \textit{failure-critical behaviors} that directly lead to task failure. Here we argue that these two types of segments provide fundamentally asymmetric supervision: progressive segments should be imitated, whereas failure-critical segments should be explicitly avoided. Based on this observation, we propose CARF, a Contrastive Attraction-Repulsion of Failure-guided framework for learning from imperfect robot data. CARF introduces a progress-based importance scorer, trained solely on successful expert demonstrations and its perturbation results, to estimate step-wise contributions toward task completion and identify informative regions in failed trajectories. These scores guide a unified flow-matching objective that attracts the policy toward progressive behaviors and repels it from failure-critical ones, while excluding ambiguous segments. This enables more comprehensive utilization of imperfect data and avoids unreliable supervision from ambiguous failure segments. Extensive experiments in simulation and the real world demonstrate consistent improvements over competing baselines across diverse failure scenarios, with ablations further validating the effectiveness of the proposed scoring and attraction-repulsion mechanisms. Our website is https://zhao-sq.github.io/carf/#.

関連論文

PR本紙発行元 EmplifAI