日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/手術操作arXiv:2609.07211

疎な報酬強化学習のための位相・初到達VLMフィードバック:手術操作への応用

Phase-and-First-Arrival VLM Feedback for Sparse-Reward Reinforcement Learning in Surgical Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

手術操作の疎な報酬強化学習において、失敗した試行から部分的な成功段階を再利用できるよう、VLMを用いて各エピソードの到達段階とその初到達時刻を特定するフィードバック手法を提案し、シミュレーションと実機で有効性を示した。

著者: Wanli Liuchen, Fangyuan Wang, Bin Li, Anqing Duan, Yunhui Liu, Peng Zhou, David Navarro-Alarcon

分類: cs.RO

原文アブストラクト

Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Failed multi-stage surgical attempts can contain grasps, lifts, or transfers worth reusing. In sparse-reward reinforcement learning, terminal rewards collapse such attempts to the same outcome, while scalar vision-language model (VLM) ratings reveal neither what progress merits credit nor when it occurred. We introduce phase-and-first-arrival feedback: one VLM query per recorded episode identifies the furthest visually verified task phase and when that phase is first reached, allowing the learner to reuse partial behavior and localize credit. We instantiate it in SurgPhaseBench, a phase-structured suite spanning rigid and deformable tasks, and evaluate it in simulation and hardware. Across five simulated tasks, our method reaches 75.2% mean success, compared with 52.1% for a reward based on Contrastive Language-Image Pre-training (CLIP) using the same visual input; the advantage persists when only the feedback representation changes. On hardware, the same record supports autonomous block picking and slip recovery. Together, these results show that trajectory-level visual supervision can preserve partial progress while providing the temporal credit needed for sparse-reward control.