日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.31439

コンプライアンスを活用する学習:ロボット挿入のための方策-アドミタンス学習フレームワーク

Learning to Leverage Compliance: A Policy-Admittance Learning Framework for Robotic Insertion

シェア:XThreadsFacebookLINEはてブBluesky

固定アドミタンス制御下で視覚方策を学習させ、方策と制御器の対立を報酬で抑えることで、コネクタ組立タスクの成功率94%と接触力ピークの大幅低減を達成した。

詳しい要約

1. どんなもの?

本論文は、ロボット組立における挿入タスクのための新しい学習フレームワーク「LeCo (Leverage Compliance)」を提案する。これは、visual policyと固定admittance制御を組み合わせ、実行時の相互作用を通じてpolicyを導くpolicy-admittance学習フレームワークである。pose errorsやcontact uncertainty下での信頼性の高い自律組立を目指す。

2. 先行研究と比べてどこがすごい?

従来、policy learningとcompliant controlの組み合わせは調整が不十分で、policyが接触に逆らい続け、controllerが譲歩することで持続的な負荷と進展の欠如が生じる問題があった。LeCoは、この問題に対処し、complianceを活用しつつ無駄な負荷を減らすようにpolicyを学習させる点が優れている。

3. 技術・手法の肝は?

LeCoの肝は、multirate feedback mechanismにより高レートのcontact-interaction記録をpolicy-transition rewardsに集約する点、integrated conflict costで持続的なpolicy-loading/controller-unloadingの対立を特徴づける点、directional high-force tail costでtransition内の持続的負荷イベントを捉える点である。これらのコストとタスク完了を組み合わせ、complianceを活用しつつ非生産的な負荷を減らす。

4. どうやって有効だと検証した?

4つの実connector-assemblyタスクで評価し、集計成功率94%を達成。比較ベースラインに対し、成功試行の平均resultant-forceピークとtorqueピークがそれぞれ約30%と64%減少。reward ablationにより、conflict shapingの追加が成功試行の中央値contact-conditioned conflict densityを約53%減少させることを示した。

5. 議論はある?

本論文では、固定complianceを活用する学習を、multirate policy-admittance相互作用を補完的なreward信号に変えることで支援できると結論づけている。有効性と低負荷挿入の実現が示唆されるが、限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として、policy learning、compliant control、admittance control、visual policy、connector-assemblyタスクに関する研究が挙げられる。同分野の定番としては、reinforcement learning for assembly、impedance control、learning from demonstrationなどが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chongren Wang, Minghe Li, Honghua Dai, Zhicheng Lin, Shiyang Wei, Xiaokui Yue

分類: cs.RO

原文アブストラクト

Policy learning and compliant control offer a promising route to reliable autonomous assembly under pose errors and contact uncertainty. However, combining them does not ensure coordination: the policy may continue pushing against contact while the controller yields, producing sustained loading with limited progress. To address this problem, we propose LeCo (Leverage Compliance), a policy-admittance learning framework that guides a visual policy through execution-time interaction under fixed admittance. A multirate feedback mechanism aggregates high-rate contact-interaction records into policy-transition rewards. An integrated conflict cost then characterizes sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition. Together with task completion, these costs encourage the policy to leverage compliance with less unproductive loading. We evaluate LeCo on four real connector-assembly tasks, obtaining an aggregate success rate of 94%. Across tasks, mean successful-trial resultant-force and torque peaks decrease by approximately 30% and 64% relative to the comparison baseline. Reward ablation further shows that adding conflict shaping reduces median successful-trial contact-conditioned conflict density by approximately 53%. These results support learning to leverage fixed compliance by turning multirate policy-admittance interaction into complementary reward signals for effective, lower-load insertion.

関連論文

PR本紙発行元 EmplifAI