日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.37552

Wrench-ACT: 直接的な力制御による接触の多い動作のためのロボットポリシー強化

Wrench-ACT: Enhancing Robot Policies for Contact Rich Behavior Using Direct Wrench Control

シェア:XThreadsFacebookLINEはてブBluesky

接触の多いマニピュレーションタスクにおいて、力(レンチ)を唯一の行動出力として予測する模倣学習ポリシーを提案し、位置ベースのベースラインを上回る性能を示した。

詳しい要約

1. どんなもの?

接触を伴うマニピュレーションでは力の調整が重要だが、近年のロボットマニピュレーション学習は行動を目標位置や姿勢で表すことが主流である。本論文は、wrench(力とトルク)を唯一の行動出力として予測し、pure force controllerで直接使う模倣学習ポリシーを提案する。ベースアーキテクチャにAction Chunking with Transformers (ACT)を用い、単一タスクモデルをbilateral wrenchのデモンストレーションで訓練し、5つの接触リッチなタスクで評価する。

2. 先行研究と比べてどこがすごい?

従来は力覚を観測としてのみ使うか、力を出力の一部として予測する場合でもhybrid force controllerに依存していた。本手法はwrenchを唯一の行動出力とし、pure force controllerで直接利用する点が異なる。位置ベースのベースラインと比較して全タスクで同等以上、意図的な力調整の度合いに応じて性能向上が変動する。また、bilateralデータ収集インターフェースとwrench行動空間がそれぞれ独立に性能に寄与することをクロス条件アブレーションで示した。

3. 技術・手法の肝は?

Action Chunking with Transformers (ACT)をベースに、単一タスクモデルをbilateral wrenchのデモンストレーションで訓練する。行動出力はwrenchのみとし、pure force controllerで直接使用する。力フィードバックテレオペレーションにより、操作者の意図的な力調整を捉えたデータ収集が重要であることを示す。

4. どうやって有効だと検証した?

5つの接触リッチなマニピュレーションタスクで評価し、wrenchポリシーが位置ベースのベースラインと同等以上であることを確認。性能向上は各タスクが要求する意図的な力調整の度合いに応じて変動した。クロス条件アブレーションにより、bilateralデータ収集インターフェースとwrench行動空間がそれぞれ独立に性能に寄与することを検証した。

5. 議論はある?

力領域の模倣学習はデータ収集に決定的に依存し、力フィードバックテレオペレーションが操作者の意図的な力調整を捉えることでポリシー性能を向上させる。bilateralデータ収集インターフェースとwrench行動空間の寄与は独立している。その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究として、Action Chunking with Transformers (ACT)、hybrid force controller、pure force controller、位置ベースのベースラインが挙げられる。関連手法として、力覚を観測として用いる模倣学習や、力を出力の一部として予測する手法が想定される。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Johannes Hechtl, Yannik Blei, Simon Ball, Reihaneh Mirjalili, Michael Krawez, Seongjin Bien, Philipp Schmitt, Wolfram Burgard

分類: cs.RO

原文アブストラクト

While contact-rich manipulation requires deliberate regulation of interaction forces, recent approaches to robot manipulation learning predominantly represent actions as target positions or poses. Even methods that incorporate force sensing either use it solely as an observation or, when predicting forces as part of the output, rely on a hybrid force controller. In this paper, we propose an imitation learning policy that predicts wrenches as its sole action output for direct use by a pure force controller. Our studies suggest that force-domain imitation learning depends critically on data collection, with force-feedback teleoperation improving policy performance by capturing the operator's deliberate force regulation. Using Action Chunking with Transformers (ACT) as the base architecture, we train single-task models on bilateral wrench demonstrations and evaluate them on five contact-rich manipulation tasks. The wrench policy matches or outperforms position-based baselines across all tasks, with gains varying according to the degree of deliberate force regulation each task requires. Cross-condition ablations show that the bilateral data collection interface and the wrench action space each contribute independently to performance. To support further research, we will release over 1000 wrench-action demonstrations spanning these tasks on a companion website upon publication.

関連論文

PR本紙発行元 EmplifAI