Policy Evaluation Using the Ω-Return

Philip S. Thomas,Scott Niekum,Georgios Theocharous,George Konidaris

Annual Conference on Neural Information Processing Systems（2015）

引用 24|浏览41

暂无评分

摘要

We propose the Ω-return as an alternative to the λ-return currently used by the TD(λ) family of algorithms. The benefit of the Ω-return is that it accounts for the correlation of different length returns. Because it is difficult to compute exactly, we suggest one way of approximating the Ω-return. We provide empirical studies that suggest that it is superior to the λ-return and γ-return for a variety of problems.

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要