日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モデルベース強化学習arXiv:2609.31025

速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習

Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control

シェア:XThreadsFacebookLINEはてブBluesky

油圧ショベルの制御において、確率的ダイナミクスアンサンブルモデルをオンラインで学習し、モデル予測制御に用いるフレームワークを提案。実機で20分の対話後、従来の100-150分学習した制御器と同等の追従精度を達成し、40分後には高速動作でサブセンチメートル精度を維持した。

詳しい要約

1. どんなもの?

- 油圧ショベル制御のためのオンラインmodel-based reinforcement learning (MBRL) フレームワーク。 - probabilistic dynamics ensemble modelをゼロから学習し、sampling-based model predictive control (MPC) に用いる。 - precision-gated contouring objectiveにより、progress rewardをpath accuracyで条件付けし、速度より精度を優先。 - 11.5トンのMenzi Muck M445油圧ショベルに直接適用し、demonstrationやsimulation pretrainingなしで学習。

2. 先行研究と比べてどこがすごい?

- 評価したMBRLベースラインより高いsample efficiencyをdata-driven excavator simulatorで達成。 - 実機で20分のinteraction後、100-150分のデータで訓練されたprior learned controllersと同等のtracking accuracyに到達。 - 40分後には高速動作でsub-centimeter mean path errorを維持。 - 実世界interactionのコスト制約下で、demonstrationやsimulation pretrainingなしに直接学習できる点が優位。

3. 技術・手法の肝は?

- probabilistic dynamics ensemble modelをオンラインでゼロから学習。 - sampling-based model predictive control (MPC) に組み込む。 - precision-gated contouring objective: progress rewardをpath accuracyでゲートし、精度を優先。 - 実機の油圧ショベル上で直接学習するonline MBRLフレームワーク。

4. どうやって有効だと検証した?

- data-driven excavator simulatorでMBRLベースラインと比較し、sample efficiencyを評価。 - 11.5トンのMenzi Muck M445油圧ショベルに直接適用し、demonstrationやsimulation pretrainingなしで検証。 - 20分interaction後のtracking accuracyを、100-150分データで訓練されたprior learned controllersと比較。 - 40分後に高速動作でsub-centimeter mean path errorを維持することを確認。

5. 議論はある?

- 要旨からは不明。 - 実機適用時の安全性、汎化性、他の油圧機械への転移、計算コストなどの議論は要旨に記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: prior learned controllers、evaluated model-based reinforcement learning baselines。 - 関連手法: model-based reinforcement learning (MBRL)、sampling-based model predictive control (MPC)、probabilistic dynamics ensemble model。 - 同分野の定番: 油圧ショベル制御、online learning、sim-to-real transferに関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar

分類: cs.RO, cs.LG, eess.SY

原文アブストラクト

Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.

関連論文

PR本紙発行元 EmplifAI