デモンストレーションで調整するポート・ハミルトン型マニピュレーション方策の再チューニング
Demonstration-Calibrated Port-Hamiltonian Retuning for Manipulation Policies
デモからポート・ハミルトンモデルを学習し、凍結した方策の予測行動に対する努力とエネルギーを評価して、下流のインピーダンス制御器のゲインを閉形式で一度に調整するオフライン手法を提案。
著者: Yulong Yang, Fan Wu, Christine Allen-Blanchette, Amit Chakraborty
分類: cs.RO, eess.SY
原文アブストラクト
Diffusion and VLA policies for manipulation are often deployed through downstream impedance controllers. The stiffness and damping gains of these controllers affect task success, yet are commonly inherited from data collection rather than selected for the deployed policy. Although empirical gain sweeps can improve performance, they require repeated evaluation rollouts. We introduce PHRetune, an offline method that derives controller gains for a frozen policy without evaluation rollouts or gain search. Our approach learns a port-Hamiltonian model from demonstrations to estimate the effort and energy associated with the policy's predicted actions. The policy is applied to recorded demonstration observations, and its predictions are assessed against demonstration-derived effort and energy budgets. From this comparison, we derive a single gain scale in closed form, adjusting the downstream controller while preserving the policy and its action representation. The gains are fixed before evaluation, without requiring a prior manipulator model, task rewards, policy retraining, or additional runtime computation. Across LIBERO suites, PHRetune improves Diffusion Policy success by up to 9.4 percentage points, with the derived gains achieving the highest observed success rates in empirical gain sweeps. On all four real-world manipulation tasks, PHRetuned Diffusion Policy outperforms the nominal policy, alternative gain-tuning methods, and a policy-retraining baseline. The same procedure improves success with SmolVLA and OpenVLA-OFT on every task, while reducing acceleration and jerk for both VLA backbones.
関連論文
- 不確実環境下での柔軟なリーチングのための仮想モデル制御マニピュレーション
- オープンエンド環境におけるロバストな把持マニピュレーションに向けてマニピュレーション
- DexForge: 高忠実度な物理情報に基づく巧みなリターゲティングマニピュレーション
- STC-MPM:軟組織切断における変形・損傷進展・切開形成の連成マニピュレーション
- 能動推論制御のための生成軌道モデルのベンチマークマニピュレーション
- 物理残差ダイナミクスと低次元全身計画による障害物回避型人間-ロボット協調布運搬マニピュレーション