日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.07525

ReDex: 指先の柔軟なインタラクションによるSim-to-Real巧みな操作ポリシーの修復

ReDex: Repairing Sim-to-Real Dexterous Policies by Finger-Level Compliant Interaction

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションで訓練した多指ハンドポリシーに対し、実機で人間が指先の接触失敗のみを修正し、その際の力覚情報を模倣学習で取り込むことで、触覚シミュレーションなしに実世界へ適応させる手法。

詳しい要約

1. どんなもの?

- シミュレーションで訓練したdexterous manipulation policyを実世界に適応させるフレームワークReDexを提案。 - 接触タイミングや力調整の誤りを修正し、tactile feedbackを導入する。 - proprioception-onlyのbase policyから出発し、人間が選択した指の接触失敗をcompliant control下で物理的に修正。 - そのrolloutからforce-conditioned policyをbehavior cloningで学習する。

2. 先行研究と比べてどこがすごい?

- 従来のsim-to-real転移では接触タイミングと力調整の誤りで失敗しやすい。 - ReDexはbase policyの多指協調を保持しつつ、局所的な接触失敗のみを修正する。 - tactile simulationや複雑なfull-hand teleoperationを必要とせず、実世界のinteractionからcontact regulationを学習できる。 - 人間の修正負担を軽減し、proprioception-only policyに力フィードバックを導入する点が新しい。

3. 技術・手法の肝は?

- proprioception-onlyのbase policyを凍結し、残りの指を制御。 - 人間オペレータが選択した指の接触失敗をcompliant control下で物理的に修正。 - rolloutはbase policy実行、人間修正指運動、fingertip force観測を組み合わせる。 - これらのrolloutからforce-informed targetsを再構成し、behavior cloningでstandalone force-conditioned policyを訓練。

4. どうやって有効だと検証した?

- 実機上でcontact-richな2つのdexterous manipulationタスクで評価。 - Object Flippingの成功率がsim-to-real転移base policyの14%から86%に向上(2物体)。 - Screwdriver Rotationの平均進捗が26.0%から95.3%に向上(3物体)。

5. 議論はある?

- 要旨からは、人間の修正負担の定量的評価や、他のタスクへの汎化性、失敗事例に関する議論は不明。 - 提案手法の限界や、force-conditioned policyの学習に必要なデータ量などは要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、sim-to-real transfer、behavior cloning、compliant control、tactile feedbackを用いたdexterous manipulationの定番研究(例: domain randomization, teacher-student distillation, residual policy learning)を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinzhou Li, Hadi Tabatabaee, Kelin Yu, Yuyin Sun, Cheng-Hao Kuo, Roberto Martín-Martín, Nima Fazeli, X. Alice Wu, Xianyi Cheng

分類: cs.RO

原文アブストラクト

Dexterous manipulation policies trained in simulation often fail to transfer to the real world because of errors in contact timing and force regulation. Yet these policies can retain useful multi-finger coordination for task progression. We propose ReDex, a framework for adapting a simulation-trained base policy to the real world by correcting local contact failures and incorporating tactile feedback. Starting from a proprioception-only base policy, ReDex allows a human operator to physically correct contact failures at selected fingers under compliant control during real-world rollouts, while the frozen base policy continues to control the remaining fingers. These rollouts combine base policy execution, human-corrected finger motion, and fingertip force observations. We reconstruct force-informed targets from these rollouts to train a standalone force-conditioned policy via behavior cloning. This design reduces human correction effort, enables learning of contact regulation from real-world interaction, and introduces force feedback into a proprioception-only policy without tactile simulation or complex full-hand teleoperation. We evaluate ReDex on two challenging, contact-rich dexterous manipulation tasks on real hardware. Compared with sim-to-real transferred base policies, ReDex increases Object Flipping success rate from 14\% to 86\% across two objects and average Screwdriver Rotation progress from 26.0% to 95.3% across three objects.

関連論文

PR本紙発行元 EmplifAI