日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18174

TacBPM: 触覚条件付き行動事前モデルによる巧みな再配向

TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation

シェア:XThreadsFacebookLINEはてブBluesky

触覚と固有感覚の履歴を条件とする潜在行動事前モデルを学習し、残差潜在行動で下流方策を効率的に訓練することで、多様な物体の巧みな手内再配向と把持・運搬タスクを実現した。

詳しい要約

1. どんなもの?

- 器用なin-hand manipulationのためのtactile-conditioned behavior prior modelであるTacBPMを提案。 - 多尺度のsphere specialistsをlatent controllerに蒸留し、下流のpolicyが固定されたtactile priorをresidual latent actionsを通じて再利用する。 - 対象物の再配向を目的とし、接触を伴う高自由度ハンド関節の協調を実現する。

2. 先行研究と比べてどこがすごい?

- 従来のraw joint commandsからの探索と比較して、再探索を削減し、学習を加速する。 - マッチしたraw-action PPOがほぼ失敗するのに対し、安定したpolicyを実現する。 - 異方性物体への任意姿勢転送、指令軸回転、arm-hand Grasp-to-AnyPoseタスクで有効性を示す。

3. 技術・手法の肝は?

- 多尺度のsphere specialistsをlatent controllerに蒸留する。 - 下流のpolicyは固定されたtactile priorをresidual latent actionsで再利用する。 - priorはtactile-proprioceptive historyに条件付けられ、現在のhand-object interactionを反映する。

4. どうやって有効だと検証した?

- 異方性物体への任意姿勢転送、指令軸回転、arm-hand Grasp-to-AnyPoseタスクで評価。 - 新規工具形状と一般化配置に対するgrasp, lift, transport, reach goal posesを含む。 - 学習加速と安定policyを実証し、in-handおよびarm-handタスクでsim-to-real転送に成功。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。同分野の定番として、PPO、sphere specialists、tactile-conditioned behavior prior model、residual latent actions、sim-to-real transferに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jie Yin, Wanli Xing, Zeyuan Zhao, Xuezhou Zhu, Zhijie Deng, Kaifeng Zhang

分類: cs.RO

原文アブストラクト

Dexterous in-hand manipulation requires policies that coordinate high-DoF hand joints through intermittent, contact-rich interaction. Beyond target-orientation tracking, such policies must discover finger gaits that preserve object stability while adapting to geometry, anisotropy, pose, contact, and sensing changes. We propose \method, a tactile-conditioned behavior prior model for dexterous reorientation. \method distills multi-scale sphere specialists into a latent controller and lets downstream policies reuse the fixed tactile prior through residual latent actions, reducing renewed exploration from raw joint commands. The prior conditions on tactile-proprioceptive history so latent behavior reflects the current hand-object interaction. We evaluate arbitrary-pose transfer across anisotropic objects, commanded-axis rotation, and an arm-hand Grasp-to-AnyPose task in which the robot must grasp, lift, transport, and reach goal poses for novel tool geometries and generalized placements. Extensive experiments demonstrate that the proposed method accelerates training and enables stable policies where matched raw-action PPO remains near failure, with successful sim-to-real transfer in in-hand and arm-hand tasks.

関連論文

PR本紙発行元 EmplifAI