日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.33053

SwingRL: ワールドモデル予測を用いた適応的観測強化学習によるケーブル吊り下げホイスティング制御

SwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting Control

シェア:XThreadsFacebookLINEはてブBluesky

遅延・欠損した視覚観測と不確実なダイナミクスに対処するため、経過時間を考慮したワールドモデルと古典的制振制御を組み合わせた残差強化学習フレームワークを提案し、建設現場での吊り荷の精密挿入タスクにおける有効性を示した。

著者: Guangming Wang, Xiaoyu Zhang, Yucheng Xin, Wanli Ma, Jiucai Liu, Yunxiang Ma, Joe Ingham, Haibing Wu, Yixiong Jing, Olaf Wysocki, Brian Sheil

分類: cs.RO

原文アブストラクト

Cable-suspended hoisting is widely used to move heavy or bulky payloads that cannot be handled conveniently by rigid pick-and-place systems, for example in crane-assisted construction. Robotic hoisting using flexible cables is challenging because payload motion is underactuated, external disturbances vary, and delayed or lost visual observations can make the perceived payload state stale at control execution. These effects are particularly critical during precise insertion of a suspended payload's sockets onto rebar pins, which is a very common task in construction environments. We present SwingRL, a residual reinforcement-learning (RL) framework that combines an age-aware world model, a classical anti-swing prior, and a recurrent residual policy to address two coupled problems: stale feedback and uncertain dynamics. The world model propagates the newest received payload observation to the current control step using the executed commands, providing a time-aligned state estimate under delayed and lossy sensing. The prior supplies nominal tracking and swing damping. The residual policy learns bounded corrections to the prior rather than the complete control law, compensating for system-parameter variation, external disturbances, and remaining state-estimation errors. We evaluate SwingRL against classical and learning-based baselines across a cumulative difficulty ladder covering system-parameter variation, wind disturbance, degraded sensing, and strong gusts. Under the most difficult setting, SwingRL achieves 69.5% strict and 77.3% broad success, exceeding all baselines by at least 60 percentage points, respectively. World-model ablations support the role of time-aligned state estimation in maintaining insertion success as observation loss increases. Finally, without real-robot fine-tuning, SwingRL achieves 90% success on the physical rig.

関連論文

PR本紙発行元 EmplifAI