日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.15716

接触の多い模倣学習のための連続多様体分解インピーダンス再ターゲティング

Continuous Manifold-Decomposed Impedance Retargeting for Contact-Rich Imitation Learning

シェア:XThreadsFacebookLINEはてブBluesky

固定インピーダンスの実演を連続可変インピーダンス制御器に変換し、模倣学習の構造的教師信号として利用する手法を提案。実タスクで力変動とピーク力を低減し、学習可能性を示した。

詳しい要約

1. どんなもの?

- 固定インピーダンスのデモを連続可変インピーダンス制御器へ変換する CMDIR を提案。 - 変換結果は imitation learning の構造的 supervision としても利用可能。 - Continuous Task-Manifold Impedance Representation (TMIR) を導入。 - 接触の多い実タスク3種で検証。

2. 先行研究と比べてどこがすごい?

- 先行の Manifold-Decomposed Impedance Retargeting (MDIR) は離散的。 - CMDIR は連続可変インピーダンスへ拡張。 - 225 trials で mean task-proxy retention を改善。 - mean pose deviation、force fluctuation、peak force を低減。 - FastMPO は C-MPO 比 5.8–9.4× 高速化。

3. 技術・手法の肝は?

- Continuous Task-Manifold Impedance Representation (TMIR): 進化する task frame と controller instructions を対にする。 - Demo-relative Compromise dynamics: moving-basis transport と control/physical metric mismatch を保持。 - 変位、reaction-impulse、perturbation-sensitivity の基準を導出。 - Quality-to-Fast: 開発経路から solver 構造を自動コンパイルし、各デモに再適用、multi-resolution evaluation で候補を認証。

4. どうやって有効だと検証した?

- 3つの実接触タスクで 225 retargeted-controller trials を実施。 - full CMDIR は discrete MDIR に対し mean task-proxy retention を改善。 - mean pose deviation、force fluctuation、peak force を低減。 - FastMPO は C-MPO 比 5.8–9.4× 高速化し、closed-loop outcomes は同等。 - 下流実験で TMIR supervision interface の learnability を確認。成功実行で force fluctuation と peak force が低い。

5. 議論はある?

- 完了信頼性はタスクや環境間で不均一。 - 成功実行では force fluctuation と peak force が低い傾向。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- Manifold-Decomposed Impedance Retargeting (MDIR) - C-MPO - FastMPO - imitation learning における impedance control 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiahao Liu, Kento Kawaharazuka, Tasuku Makabe, Kei Okada

分類: cs.RO

原文アブストラクト

CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.

関連論文