日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能な強化学習arXiv:2610.10367

時間的に解釈可能な微分可能決定木

Temporally Interpretable Differentiable Decision Trees

シェア:XThreadsFacebookLINEはてブBluesky

行動チャンキングと情報理論的な木再構築により、逐次意思決定タスクにおける微分可能決定木の時間的解釈性を向上させる手法を提案した。

詳しい要約

1. どんなもの?

本論文は、sequential-decision making タスクにおける differentiable decision trees (DDTs) の解釈性を高めるため、時間軸に着目した temporal interpretability を導入する。action chunking による temporal abstraction が解釈性を向上させることを示し、action chunking を組み込んだ2つの新しい policy gradient アルゴリズムと、パラメータ効率を保つ情報理論的 tree restructuring アルゴリズムを提案する。4つの simulation 環境で評価し、distilled action chunked policy から warm-start した action chunked DDT が、neural network policy と3/4の環境で同等性能を最大80%少ないパラメータで達成することを示した。

2. 先行研究と比べてどこがすごい?

従来の DDTs は single-timestep の挙動と人間の multi-timestep planning の間に本質的なミスマッチがあり、sequential-decision making に適していなかった。本研究は時間を新たな解釈性の次元として導入し、action chunking による temporal abstraction でこのミスマッチを解消する点が新しい。また、パラメータ効率を維持する tree restructuring を開発し、最大80%のパラメータ削減を実現した。

3. 技術・手法の肝は?

- temporal interpretability の導入 - action chunking を組み込んだ2つの新規 policy gradient アルゴリズム - 情報理論的 tree restructuring アルゴリズムによるパラメータ効率の維持 - distilled action chunked policy からの warm-start が有効

4. どうやって有効だと検証した?

4つの simulation 環境で評価。warm-start した action chunked DDTs が neural network policies と3/4の環境で同等性能を示し、最大80%少ないパラメータで実現することを確認。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法は明示されていない。同分野の定番として differentiable decision trees (DDTs)、policy gradient、action chunking、distillation などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Eisuke Hirota, Aarav Sane, Rohan Paleja

分類: cs.LG, cs.RO

原文アブストラクト

Interpretability offers a solution to safe autonomy by providing transparency into an agent's underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans with a discrete tree-based visualization. Nonetheless, current implementations of DDTs are not well-suited for sequential-decision making domains, as there exists an inherent mismatch between a tree's single-timestep behavior and a human's multi-timestep planning. Our work thus introduces time as a new dimension of interpretability, coined as temporal interpretability, and demonstrates how temporal abstractions via action chunking improve it. We achieve this by first introducing two novel policy gradient algorithms that incorporate action chunking. Additionally, to maintain parameter-efficient trees, we develop an information-theoretic tree restructuring algorithm that modifies the tree during training. Across four simulation environments, we find that warm-starting action chunked DDTs from a distilled action chunked policy is the most effective way to obtain temporally interpretable trees: they match neural network policies in three of the four domains while using up to 80$\%$ fewer parameters. Our code is available at https://github.com/ei5uke/temp-interp.

PR本紙発行元 EmplifAI