日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
タスク計画arXiv:2610.11223

SafeInferCom: 検証器ガイドによる生成途中介入で安全な推論時計算を実現するロボットタスク計画

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

シェア:XThreadsFacebookLINEはてブBluesky

大規模推論モデルの生成途中で中間計画を検証・修正するフレームワークを提案し、ロボットタスク計画の成功率向上とトークン削減を実現した。

詳しい要約

1. どんなもの?

- ロボットタスク計画のための推論時安全性フレームワーク - LRLMs (Large Reasoning Language Models) を用いた多段階推論 - 推論中に中間計画を検証し、エラー修正を促す - 推論時計算の無駄を削減し、計画信頼性を向上 - 形式検証器 (formal verifier) をガイドとして利用 - 有効な中間計画を保持しつつ、制約違反を解消 - 実験で計画成功率向上とエラー修正加速を確認 - VirtualHomeと実機ロボットアームで評価

2. 先行研究と比べてどこがすごい?

- 従来のone-shot推論では、推論が進むにつれて有効な中間計画が上書きされたり、制約違反が未解決のままになる問題があった - 既存の反復改善 (iterative refinement) 手法と比較して、SafeInferComはトークン使用量を削減しつつ成功率を向上 - 推論時モニタを導入し、デコード軌跡を乱さずに中間計画を検証・介入できる点が新しい - 形式検証器をガイドとして用いることで、自己修正能力の限界を補う

3. 技術・手法の肝は?

- 推論時モニタ (inference-time monitor) を開発 - 中間計画を露出・検証し、元のデコード軌跡を妨げない - SafeInferComフレームワーク - 形式検証器 (formal verifier) が中間計画を評価 - 有効な計画を保持し、エラー修正を生成中に指示 - 検証結果に基づき、推論をガイド - 制約違反を検出したら修正を促す - 反復改善と組み合わせ可能

4. どうやって有効だと検証した?

- 複数のLRLMsと計画ドメインで実験 - one-shot推論における推論と応答の不一致、自己修正の限界を確認 - SafeInferComが計画成功率を向上させ、エラー修正を加速することを示した - 反復改善と組み合わせると、成功率がさらに向上し、トークン使用量も削減 - VirtualHomeでの評価と実世界のロボットアーム実演を実施

5. 議論はある?

- 推論と応答の不一致 (reasoning-response inconsistency) が問題として指摘 - one-shot推論では自己修正が限定的 - SafeInferComは推論時計算の効率化と信頼性向上を両立 - 反復改善との併用で更なる効果 - 実世界適用の可能性を示すが、詳細な議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: one-shot inference, iterative refinement - 関連手法: formal verifier, inference-time monitor - 同分野の定番: LRLMs (Large Reasoning Language Models), VirtualHome - 具体的な論文名は要旨に記載なし

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Weizhe Xu, Jialiang Fan, Mengyu Liu, Fanxin Kong

分類: cs.RO, cs.AI, cs.CL, cs.LO

原文アブストラクト

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.

関連論文

PR本紙発行元 EmplifAI