LLMエージェントのマルチターン一貫性評価:生存分析と失敗理由の分類体系
Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy
20ステップのマルチエージェント実験で8モデル・約8.5万軌跡を分析し、報酬を先延ばしにするか即時獲得するかの一貫性を生存分析で評価。失敗理由を7分類し、時間・文脈による変化や熟考の矛盾傾向を明らかにした。
著者: Igor Bogdanov, Olga Manakina, Chung-Horng Lung
分類: cs.AI, cs.CL, cs.LG, cs.MA
原文アブストラクト
Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories spanning 8 model families. Treating the first reward-claim as a time-to-event outcome, we estimate Kaplan-Meier survival curves and fit discrete-time hazard regression to quantify how experimental factors shift failure risk over time. Then, to analyze rationales and language patterns associated with failure, we build a seven-category taxonomy from 13,780 deliberation traces from agents who choose to terminate the episode, using an LLM-assisted labeling paired with human audit ($κ=0.83$). Rationale profiles change systematically with time and context: early failures are more impulse-driven, later failures more fatigue- and cost-benefit-framed, while public settings increase norm-oriented justifications. We also find a deliberation-inconsistency association: among failures, longer deliberation correlates with higher rates of intra-rationale contradiction (simultaneous pro-delay and pro-claim statements), challenging the assumption that more reasoning text implies greater consistency. Together, the survival and rationale analyses reveal distinct temporal reliability regimes and model-specific "failure fingerprints", offering an evaluation lens for diagnosing inconsistency in multi-turn agent behavior.
関連論文
- LLMエージェントのための世界モデル再考:エージェント編集型世界モデルLLMエージェント
- 世界を巻き戻し、反省を残す:長期LLMエージェントのためのロールバック誘導リフレクションLLMエージェント
- SkillGLoW: 手続き的ファミリーのスキル統合による長期的タスクストリーム上の自己改善エージェントLLMエージェント
- LLMエージェントにおける世界モデルと方策の合成:スペクトル解析と行動解析による統一的考察LLMエージェント
- 実行可能な幻覚検出:潜在的不確実性をエージェント的批判へ変換するLLMエージェント
- State2State: 環境から導出された中間学習によるLLMエージェントの訓練LLMエージェント