日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
報酬モデリングarXiv:2609.34484

ARS: ロボット学習のためのエージェント型報酬システム

ARS: Agentic Reward System for Robot Learning

シェア:XThreadsFacebookLINEはてブBluesky

汎用VLMを推論に用いて、追加学習なしにロボット軌跡の進捗報酬を推定するエージェント型フレームワークを提案し、実機の長期タスク学習で有効性を示した。

著者: Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang, Kai Xie, Xiaofeng Mou, Yi Xu

分類: cs.RO, cs.LG

原文アブストラクト

Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars

関連論文

PR本紙発行元 EmplifAI