日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.19962

接触の多いロボット書架挿入のためのハイブリッド残差強化学習

Hybrid Residual Reinforcement Learning for Contact-Rich Robotic Book Insertion

シェア:XThreadsFacebookLINEはてブBluesky

把持後の本を狭い棚に挿入する接触の多い制御問題に対し、公称タスク空間制御器と残差PPOを組み合わせ、局所補正と解放判断を学習させて成功率を大幅に改善した。

詳しい要約

1. どんなもの?

- 把持済みの本を狭い棚に挿入する contact-rich 制御問題を扱う。 - 把持獲得と global approach の後の最終フェーズを対象。 - 既知の geometry と学習挙動の間で制御権限をどう配分するかを問う。 - nominal task-space controller を保持し、residual PPO が bounded local correction と release 判断を担う。 - open-retreat-reclose 遷移のみ scripted。

2. 先行研究と比べてどこがすごい?

- 従来の nominal control と比較して大幅に成功率を改善。 - deployment-matched simulation で nominal 37.89% に対し 98.50% mean success。 - 実機 xArm7 でも nominal 26.7% から 63.3% へ向上。 - 失敗を 22 から 11 に削減。 - 差が出た 15 条件中 13 条件で residual control が勝利。

3. 技術・手法の肝は?

- nominal task-space controller で構造化された insertion と seating を実行。 - residual PPO が bounded local correction を供給。 - residual PPO が release タイミングも決定。 - open-retreat-reclose 遷移のみ scripted。 - geometry で信頼できる task structure を保持し、学習を contact-sensitive 挙動に集中。

4. どうやって有効だと検証した?

- deployment-matched simulation で 512 fixed conditions を評価。 - 3 独立 training runs で 98.50% mean success、0.23 percentage-point sample SD。 - 実機 xArm7 で 30 matched conditions、60 trials を実施。 - 成功率 26.7% から 63.3%、失敗 22 から 11、15 条件中 13 条件で勝利。 - initialization perturbations 1.5x まで 87% 超を維持。 - 非常に tight な clearance では local correction の幾何限界が露呈。

5. 議論はある?

- geometry が reliable task structure を保持し、学習を contact-sensitive 挙動に集中する hybrid design を支持。 - 非常に tight な clearance では local correction の限界が示される。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として residual reinforcement learning、PPO、task-space control、contact-rich manipulation の定番文献を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianyuan Liu, Rutherford Agbeshi Patamia, Benjamin Champion, Akansel Cosgun, Richard Dazeley

分類: cs.RO

原文アブストラクト

Placing a grasped book into a tight shelf is a compact but difficult contact-rich control problem: millimetre-scale pose error can turn a geometrically valid approach into jamming, failed release, or incomplete seating. We study this final phase after grasp acquisition and global approach, and ask how control authority should be divided between known geometry and learned behaviour. Our method retains a nominal task-space controller for structured insertion and seating, while residual PPO supplies bounded local corrections and decides when to release. Only the brief open-retreat-reclose transition is scripted. For the final policy used on hardware, a deployment-matched simulation evaluation over 512 fixed conditions yields 98.50 percent mean success (0.23 percentage-point sample SD) across three independent training runs, compared with 37.89 percent for nominal control. On the physical xArm7, 60 trials over 30 matched conditions show the same qualitative advantage: residual control raises success from 26.7 percent to 63.3 percent, reduces failures from 22 to 11, and wins 13 of the 15 matched conditions in which the two controllers differ. Robustness tests show that performance remains above 87 percent under initialization perturbations up to 1.5x, while very tight clearances expose the geometric limit of local correction. These results support a hybrid design in which geometry preserves reliable task structure and learning is concentrated on the contact-sensitive behaviour that fixed rules handle poorly.

関連論文

PR本紙発行元 EmplifAI