日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
LLMエージェント評価arXiv:2606.05558v2

LLMエージェントのオフ方策評価のための自己回帰拡散ワールドモデル

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents

シェア:XThreadsFacebookLINEはてブBluesky

事前収集した軌跡のみから、新しいLLMエージェント方策の性能を推定する評価フレームワークADWMを提案。各遷移を独立した拡散過程としてモデル化し、方策条件付きスコア関数でエージェントの意思決定を反映したシミュレーションを実現する。

著者: Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang

分類: cs.LG

原文アブストラクト

Evaluating large language model (LLM) agents in multi-turn interactive environments is expensive and risky, as it requires online environment interaction. We propose ADWM (Autoregressive Diffusion World Model), an evaluation framework that estimates the performance of a new LLM agent policy purely from pre-collected trajectories. The core idea is to learn a latent diffusion world model that simulates how the environment responds to the evaluation policy, without ever executing it in the real environment. Existing diffusion-based OPE methods guide full trajectories in a single pass by jointly diffusing states and actions, an assumption that breaks down for LLM agents whose actions are discrete text that must be sampled from the policy after observing the environment. Unlike autoregressive world models that suffer from compounding errors, ADWM models each transition as an independent denoising process, enabling reliable step-by-step rollouts where the world model and agent alternate in causal order. Crucially, the LLM agent under evaluation directly guides the diffusion generation at each step via a policy-conditioned score function, ensuring that simulated trajectories accurately reflect its decision-making patterns. Empirically, ADWM achieves accurate value estimates and evaluation reliability across diverse multi-turn agent tasks, demonstrating its promise as a practical framework for offline LLM agent evaluation.