報酬中心型ReST-MCTS:高不確実環境におけるロボットマニピュレーションのための堅牢な意思決定フレームワーク
Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments
不確実な環境でのロボット操作において、中間フィードバックを複数のチャネルに分解し、タスク文脈に照らして検索をバイアス・修正する報酬中心型MCTSフレームワークを提案。VLA行動修復にも応用可能。
著者: Xibai Wang
分類: cs.RO, cs.AI
原文アブストラクト
Monte Carlo tree search is attractive for robotic manipulation because it can improve action selection through simulation without requiring a fully differentiable policy. In uncertain domains, however, sparse terminal rewards and noisy transitions can make shallow search brittle: many candidate branches remain indistinguishable until late rollouts, and small simulation budgets amplify this ambiguity. This paper presents Reward-Centered ReST-MCTS, a decision-making framework that decomposes intermediate feedback into rule, heuristic, optional neural, and value-estimation channels, centers the resulting process signal against matched task contexts, and uses it to bias or repair search while preserving terminal-task evaluation. The primary evidence is intentionally tiered. Local tasks and matched ManiSkill diagnostics isolate reward-center mechanisms and ablations; matched option-level ManiSkill sweeps test robustness under primitive failure, observation noise, and initial-pose shifts while not claiming standard benchmark superiority; and an official same-backbone OpenVLA-OFT/LIBERO bridge tests bounded VLA action repair. The OpenVLA-OFT clean reproduction reaches 10/10 LIBERO-Spatial successes both with and without RCRM-Guard. A single-suite same-backbone action-channel stress artifact over ten paired LIBERO-Spatial action-channel stress episodes records 0/10 unguarded successes and 9/10 guarded successes. Additional observation-noise, language-perturbation, and visual-distractor probes are reported as coverage and negative-result context rather than superiority evidence. The resulting claim is bounded: Reward-Centered ReST-MCTS is an inspectable test-time verifier for same-backbone high-uncertainty manipulation, not a replacement VLA policy or a broad standard-benchmark superiority claim.
関連論文
- エンドタスク成功を超えて:ロボティクスにおける視覚経験検索の監査手法マニピュレーション
- 複数把持点における非無視可能な物理応答を伴う線形変形物体の安定性保証付きマニピュレーションマニピュレーション
- PAKT: 強化学習のための物理的整合性を備えたキネステティック教示マニピュレーション
- 不完全データを活用した高精度ロボットマニピュレーションマニピュレーション
- SafeLoop: 視覚言語行動マニピュレーションのためのリスク認識ロールバックマニピュレーション
- ローカルコーディングエージェントによるマニピュレーションスキルの汎化マニピュレーション