世界モデルが強化学習エージェントの最適採餌戦略を実現する
World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents
学習した世界モデルを持つ人工採餌エージェントが、限界価値定理に沿ったパッチ離脱行動を自然に獲得することを示し、予測的世界モデルが生物学的に妥当な意思決定の基盤となる可能性を提示した。
著者: Yesid Fonseca, Manuel S. Ríos, Nicanor Quijano, Luis F. Giraldo
分類: cs.AI, cs.LG
原文アブストラクト
Patch foraging involves the deliberate and planned process of determining the optimal time to depart from a resource-rich region and investigate potentially more beneficial alternatives. The Marginal Value Theorem (MVT) is frequently used to characterize this process, offering an optimality model for such foraging behaviors. Although this model has been widely used to make predictions in behavioral ecology, discovering the computational mechanisms that facilitate the emergence of optimal patch-foraging decisions in biological foragers remains under investigation. Here, we show that artificial foragers equipped with learned world models naturally converge to MVT-aligned strategies. Using a model-based reinforcement learning agent that acquires a parsimonious predictive representation of its environment, we demonstrate that anticipatory capabilities, rather than reward maximization alone, drive efficient patch-leaving behavior. Compared with standard model-free RL agents, these model-based agents exhibit decision patterns similar to many of their biological counterparts, suggesting that predictive world models can serve as a foundation for more explainable and biologically grounded decision-making in AI systems. Overall, our findings highlight the value of ecological optimality principles for advancing interpretable and adaptive AI.