Is the Policy Gradient a Gradient?

Chris Nota
Chris Nota

AAMAS '19: International Conference on Autonomous Agents and Multiagent Systems Auckland New Zealand May, 2020, pp. 939-947, 2019.

Cited by: 4|Bibtex|Views13|DOI:https://doi.org/10.5555/3398761.3398871
EI
Other Links: dblp.uni-trier.de|dl.acm.org|arxiv.org

Abstract:

The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the discount factor from the state distribution and therefore do not optimize the discounted objective. What do they optimize instead? This has been an open question ...More

Code:

Data:

Full Text
Your rating :
0

 

Tags
Comments