日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.25630

PAKT: 強化学習のための物理的整合性を備えたキネステティック教示

PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

人間がロボットを直接動かして教示する際に、ロボットや方策が物理的に再現できない軌道を防ぐアドミッタンス制御ベースの枠組みを提案し、低頻度のRL行動を高頻度トルク指令に変換する制御スタックを構築した。

詳しい要約

1. どんなもの?

- 接触の多い産業マニピュレーション向けの実世界RLの枠組み - マイクロメートル精度・99%超の成功率・人間並みのサイクルタイムが要求される - 実演と介入を活用するoff-policyアルゴリズムの性能を引き出す - 物理系とpolicyの制約に従いつつ直感的にガイダンスを集めるインタフェースが課題 - PAKTはkinesthetic teachingのためのフレームワーク - teleoperationではなく産業で広く使われるkinesthetic guidanceを採用

2. 先行研究と比べてどこがすごい?

- teleoperationベースの手法と異なりkinesthetic guidanceを採用 - kinesthetic guidanceの弱点は、操作者がrobotやpolicyが物理的に再現できない軌道(速度・加速度・jerk)に動かせる点 - PAKTはadmittance controlで人間の力を運動に写像 - 下流のreference generatorがpolicy実行時と同じkinematic limitsを適用し軌道を制約内に保つ - HIL-SERLベースラインに対しサイクルタイムを23%-48%、累積介入回数を62%-86%削減

3. 技術・手法の肝は?

- 操作者はadmittance controlを通じてrobotをガイド - admittance controlは人間が加えた力を運動にマッピング - reference generatorがpolicy実行時と同じkinematic limitsを適用 - 収集軌道をこれらの制約内に維持 - 低頻度のRL actionを高頻度のtorque commandに写像する高性能制御スタックを追加 - 構成はreference generatorと後段のimpedance controller - reference generatorはimpedance controllerの追従性能を保ちつつ接触処理を改善し、より滑らかなpolicy actionを生成

4. どうやって有効だと検証した?

- 4つのinsertionおよび産業assemblyベンチマークで報告された実行により検証 - ベンチマークにはdata center compute trayを含む - エンドツーエンドシステムがHIL-SERLベースライン比でサイクルタイムを23%-48%削減 - 累積介入回数を62%-86%削減 - 詳細な評価プロトコルは要旨からは不明

5. 議論はある?

- kinesthetic guidanceの弱点として、操作者がrobotやpolicyに物理的に再現不能な軌道を与え得る点を指摘 - PAKTはadmittance controlとreference generatorのkinematic limits適用でこの問題に対処 - 制御スタックが接触処理とpolicy actionの滑らかさを改善 - 限界や失敗事例、一般化可能性に関する議論は要旨からは不明

6. 次に読むべき論文は?

- HIL-SERL(要旨で比較ベースラインとして参照) - off-policy RL with demonstrations and interventions - teleoperationベースの実演収集手法 - admittance controlおよびimpedance control - kinesthetic teaching - 接触の多い産業マニピュレーション向け実世界RL

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lars Johannsmeier, Yashraj Narang

分類: cs.RO

原文アブストラクト

Real-world reinforcement learning (RL) systems still struggle with the demands of contact-rich industrial manipulation, including micrometer-level precision, success rates above 99%, and human-level cycle times. Although off-policy algorithms can improve performance by leveraging demonstrations and interventions, a key bottleneck is the lack of an intuitive interface for collecting such guidance while complying with constraints of the physical system and the policy. We propose PAKT, a framework for kinesthetic teaching in RL. As opposed to teleoperation approaches, PAKT relies on kinesthetic guidance, which is widely used in industry. However, a critical weakness of kinesthetic guidance is the possibility for the operator to move the robot along trajectories (e.g., velocities, accelerations, jerk) that the robot and/or policy cannot physically reproduce. Using PAKT, operators guide the robot through admittance control, which maps human-applied forces to motion. The downstream reference generator applies the same kinematic limits used during policy execution, keeping the collected trajectories within these limits. To support this teaching interface with an appropriate execution layer, PAKT adds a high-performance control stack that maps low-frequency RL actions to high-frequency torque commands. It consists of a reference generator and subsequent impedance controller, where the reference generator preserves the tracking performance of the impedance controller while improving contact handling and producing smoother policy actions. Across the reported runs on four insertion and industrial assembly benchmarks, including a data center compute tray, the end-to-end system reduces cycle time by 23%-48% and cumulative intervention count by 62%-86% relative to the HIL-SERL baseline. Project website: https://pakt-website.github.io/pakt-website}{https://pakt-website.github.io/pakt-website

関連論文

PR本紙発行元 EmplifAI