日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/オフラインRLarXiv:2609.31586

信頼誘導型Decision Transformer

Trust Guided Decision Transformer

シェア:XThreadsFacebookLINEはてブBluesky

Decision Transformerの長いロールアウトでの性能劣化を、次状態予測誤差に基づく信頼性評価で検出し、信頼できる文脈のみを用いて価値誘導を行う手法を提案。

詳しい要約

1. どんなもの?

- Decision Transformer (DT) の長いロールアウトにおける性能劣化を解決する Trust Guided Decision Transformer (TGDT) を提案。 - DT の性能劣化は、条件付けコンテキストが訓練分布から外れる (context drift) ことに起因。 - このドリフトはモデル自身の次状態予測誤差 (next state prediction error) の上昇として現れ、ロールアウト中に増加し高いままになる。 - TGDT はコンテキスト選択を価値ガイダンスの前に行うことで、信頼できるコンテキストのみを使用する。

2. 先行研究と比べてどこがすごい?

- 従来の value only elastic selection では、critic がモデル自身が信頼できないと判断したコンテキストから生成された行動を選択する可能性があった。 - TGDT はコンテキスト選択を先に行い、信頼できるコンテキストの中から critic で最高価値の行動を選ぶため、この問題を回避。 - 実験で、state prediction、critic guidance、hard context reset はそれぞれ問題の一部しか解決しないことを示し、TGDT はそれらを上回る性能を達成。 - 具体的には、vanilla Decision Transformer、reset based context control、value only context selection と比較してリターンを改善し、持続的な高エラー実行を減少させる。

3. 技術・手法の肝は?

- 各ステップで、最近の複数のコンテキストサフィックスを評価。 - 評価には rolling next state prediction error を使用し、held out offline data を用いた split conformal prediction でキャリブレーションされた閾値と比較。 - 誤差がキャリブレーションされた閾値内に収まるサフィックスのみを保持 (trusted suffixes)。 - その後、frozen critic を用いて信頼できるサフィックスの中から最高価値の行動を選択。 - これにより、コンテキスト選択と価値ガイダンスの順序を逆転させ、信頼性を優先。

4. どうやって有効だと検証した?

- D4RL の navigation および locomotion タスクで実験を実施。 - 比較対象: vanilla Decision Transformer、reset based context control、value only context selection。 - 結果: TGDT は持続的な高エラー実行を減少させ、リターンを改善。 - また、state prediction、critic guidance、hard context reset を個別に適用しても問題の一部しか解決できないことを示した。

5. 議論はある?

- 要旨からは、TGDT の限界や失敗ケース、計算コスト、ハイパーパラメータ感度についての議論は明示されていない。 - ただし、context drift が next state prediction error で可視化できること、およびその誤差を conformal prediction でキャリブレーションするアプローチの有効性が示唆されている。 - より詳細な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: Decision Transformer、value only elastic selection、reset based context control、value only context selection。 - 関連手法: split conformal prediction、D4RL ベンチマーク。 - 次に読むべき論文として、これらの手法の原論文や、Decision Transformer のロバスト性に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath

分類: cs.LG

原文アブストラクト

Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.

PR本紙発行元 EmplifAI