日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.29031

直接駆動グリッパのゼロショットsim-to-real把持のための簡易トルク観測アライメント

Simple Torque-Observation Alignment for Zero-Shot Sim-to-Real Grasping with a Direct-Drive Gripper

シェア:XThreadsFacebookLINEはてブBluesky

直接駆動アクチュエータのトルク定数校正とトルク差分観測により、シミュレーションと実機のトルクのスケール・オフセット・ノイズのずれを解消し、ゼロショットで把持方策を実機に転移させる手法を提案した。

詳しい要約

1. どんなもの?

本論文は、direct-drive (DD) actuator を備えたロボットにおける torque observation の sim-to-real ギャップを解消するための簡便な alignment 手法を提案する。シミュレーションと実機で torque の scale・offset・noise が異なる問題に対し、dynamometer calibration による torque constant K_tau の同定、delta_tau(t) = tau(t) - tau(t-1) の使用、測定データ由来の Gaussian noise 注入を組み合わせる。teacher-student grasping policy を完全にシミュレーションで学習し、distilled student を multifingered DD gripper に zero-shot 展開する。

2. 先行研究と比べてどこがすごい?

従来の sim-to-real では torque observation の scale・offset・noise の不一致が強化学習の転移を難しくしてきた。本手法は DD actuator の motor current が motor-type-specific torque constant K_tau を介して joint torque に線形対応する点に着目し、dynamometer calibration で K_tau* を同定して scale を補正する。さらに delta_tau(t) を用いてドメイン依存の定数 offset を除去し、測定由来の Gaussian noise を学習時に注入する。これにより、実世界の torque-observation mismatch に対する zero-shot policy transfer の robustness を向上させる点が先行研究と異なる。

3. 技術・手法の肝は?

技術の肝は次の3点である。第一に、dynamometer calibration により motor-type-specific torque constant K_tau* を同定し、シミュレーションと実機の torque scale の不一致を補正する。第二に、直接 torque tau(t) ではなく delta_tau(t) = tau(t) - tau(t-1) を両ドメインの observation として用い、ドメイン依存の定数 offset を除去する。第三に、dynamometer measurement data から得た Gaussian noise を学習過程で注入する。これらを組み合わせ、joint positions と torque differences のみを用いる proprioceptive grasping policy を teacher-student で学習・蒸留する。

4. どうやって有効だと検証した?

提案手法の検証として、teacher-student grasping policy を完全にシミュレーションで学習し、distilled student を multifingered DD gripper に zero-shot 展開した。展開された policy は joint positions と torque differences のみを用いた proprioceptive grasping を実行する。9つの in-distribution (ID) objects に対し、提案手法と代替 alignment variants を比較する ablation study を行い、提案手法が 100% grasp success を達成した。これにより、実世界の torque-observation mismatch に対する zero-shot policy transfer の robustness 向上が示された。

5. 議論はある?

本論文では、DD actuator における torque observation alignment が zero-shot sim-to-real grasping の robustness を改善することが示唆される。一方、要旨からは、out-of-distribution (OOD) objects や他の actuator タイプへの一般化、noise 注入の感度、K_tau 同定の誤差影響、実機での長期的な安定性などについては不明である。また、提案手法が 100% grasp success を達成したのは 9つの ID objects に限られており、より多様な条件での検証は今後の課題と考えられる。

6. 次に読むべき論文は?

要旨で参照・比較されている研究は明示されていない。関連手法として、sim-to-real transfer、domain randomization、teacher-student distillation、torque observation alignment、direct-drive actuator の制御、proprioceptive grasping に関する研究が次に読むべき候補として挙げられる。具体的な論文名は要旨からは不明である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Doyoung Kim, Edgar Lee, Hyeonsun Park, Chunghyeon Lee, Chihyun Han, Uisu Hwang, Seokhwan Jeong

分類: cs.RO

原文アブストラクト

Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.

関連論文

PR本紙発行元 EmplifAI