日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン模倣学習arXiv:2609.38225

SynIL: 不完全なデモンストレーションデータセットからのオフライン模倣学習のためのシナジー活用

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

シェア:XThreadsFacebookLINEはてブBluesky

運動シナジーを自己教師ありで定量化し、遷移ごとの報酬を生成することで、質の低いデモを含むデータからでも高精度なオフライン模倣学習を可能にするフレームワークを提案。

著者: Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe

分類: cs.RO, cs.AI

原文アブストラクト

Imitation learning enables robots to acquire complex skills directly from massive demonstration datasets, but its performance degrades severely when datasets are contaminated with suboptimal or noisy demonstrations. While prior quality-assessment methods attempt to filter or reweight data, they typically rely on manual pre-selection of expert reference data or task-specific heuristics, limiting scalability. To address this challenge, we introduce SynIL (Synergy-based Imitation Learning), a novel framework for automated, label-free demonstration quality assessment in offline reinforcement learning. Grounded in neuroscientific evidence that motor synergy, a low-dimensional coordinated structure in movement, correlates directly with motor proficiency, SynIL algorithmically quantifies synergy manifestation to generate dense, transition-level reward signals via self-supervised reward regression. Comprehensive evaluations on D4RL locomotion benchmarks and multi-human Robomimic manipulation datasets demonstrate that synergy-derived rewards correlate strongly with ground-truth rewards. Furthermore, SynIL substantially outperforms Behavior Cloning (BC) and achieves performance comparable to, and in sparse-reward human teleoperation scenarios, superior to, offline reinforcement learning trained on true environment rewards.

PR本紙発行元 EmplifAI