Proceedings of 2006 International Conference on Artificial Intelligence: 50 YEARS' ACHIEVEMENTS, FUTURE DIRECTIONS AND SOCIAL IMPACTS(2006)
Tongji Univ
被引用6708|浏览58
摘要
Reinforcement Learning (RL) is developed from control theories, statistics, psychology etc. It is much more focused on goal-directed learning from interaction. It addresses the problem of learning optimal policies for sequential decision-making problems involving stochastic operators and numerical reward functions rather than the more traditional deterministic operators and logical goal predicates. Our goal in writing this paper was to provide a clear and simple account of the key ideas and algorithms of reinforcement learning. It presents a conceptual framework that might serve as an introduction to a more rigorous study of RL.
更多
查看译文
关键词
reinforcement learning, MDPs, reward function, value function, MAXQ, HEXQ