日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.18004

欠けた橋を探せ:構成を考慮した能動的模倣学習

Missing Bridges: Composition-Aware Active Imitation Learning

シェア:XThreadsFacebookLINEはてブBluesky

潜在トポロジーを構築し、多数のタスクを一度に解ける「橋渡し」デモを能動的に要求することで、わずかな実演でロボットの多タスク成功率を大幅に向上させる手法を提案。

詳しい要約

1. どんなもの?

- 能動的模倣学習(active imitation learning)の新手法 AALT(Adaptive Agents via Latent Topologies)を提案。 - 構造化されたマルチタスク領域で、開始-目標タスクが組合せ的に増える問題に対処。 - 既存のデモンストレーションを潜在ハブ状態のトポロジーに組織し、学習済み行動で接続。 - 多くのタスクを一度に可能にする高価値なブリッジデモを特定し、専門家に要求。 - 推論時はトポロジーを通じて計画し、各ハブ遷移で diffusion policy を条件付け。

2. 先行研究と比べてどこがすごい?

- 既存手法は専門家政策に関する期待情報利得で要求を選択するが、AALT は開始-目標接続性の期待利得を最大化するように要求。 - この目的がタスク到達可能性に関する情報利得と形式的に関連することを示す。 - シミュレーション UR5e ロボットの ordered-retrieval 領域(72タスク)で、AALT は初期データセットに加え3デモ(計5遷移)のみで42/72から72/72(100%)成功。 - 最強ベースラインは20デモ(98遷移)後でも平均88.6%成功にとどまる。

3. 技術・手法の肝は?

- 既存デモを潜在ハブ状態のトポロジーに組織し、学習済み行動で接続。 - 高価値なブリッジデモを特定し、専門家クエリに基づいて要求。 - 目的関数は開始-目標接続性の期待利得を最大化し、タスク到達可能性の情報利得と形式的に関連。 - 推論時はトポロジーを通じて計画し、各ハブ遷移で diffusion policy を条件付け。

4. どうやって有効だと検証した?

- シミュレーション UR5e ロボットの ordered-retrieval 領域で72タスクを対象に検証。 - AALT は初期データセットに加え3デモ(計5遷移)のみで42/72から72/72(100%)成功を達成。 - 最強ベースラインは20デモ(98遷移)後でも平均88.6%成功。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として active imitation learning、diffusion policy、multi-task learning に関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Maxwell J. Jacobson, Ahmed H Qureshi, Yexiang Xue

分類: cs.AI, cs.RO

原文アブストラクト

Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs. Existing methods typically select these requests for their expected information gain about the expert policy. In structured multi-task domains, however, the number of start-goal tasks may grow combinatorially despite their solutions sharing reusable behavior. This makes composable behaviors especially valuable, since a single demonstration may help solve many tasks at once. Prior methods do not explicitly account for this value when selecting which demonstration to request. We introduce Adaptive Agents via Latent Topologies (AALT), which requests demonstrations that maximize expected gains in start-goal connectivity. We further show that this objective is formally tied to information gain about task reachability. AALT organizes existing demonstrations into a topology of latent hub states connected by learned behaviors, identifies high-value bridge demonstrations that are likely to enable many tasks at once, and grounds each to an expert query. At inference, it plans through the resulting topology and conditions a diffusion policy on each successive hub transition. In a simulated UR5e robot ordered-retrieval domain with 72 tasks, AALT improved from 42/72 to 72/72 (100%) successful tasks consistently using only 3 demonstrations totaling 5 transitions beyond the initial dataset. After 20 demonstrations, the strongest baseline averaged 88.6% success using 98 transitions.

関連論文

PR本紙発行元 EmplifAI