We study reinforcement learning with linear function approximation and finite-memory approximations for partially observed Markov decision processes (POMDPs). We first present an algorithm for the value evaluation of finite-memory feedback policies. We provide error bounds derived from filter stability and projection errors. We then study the learning of finite-memory based near-optimal Q values. Convergence in this case requires further assumptions on the exploration policy when using general basis functions. We then show that these assumptions can be relaxed for specific models such as those with perfectly linear cost and dynamics, or when using discretization based basis functions.
更多
查看译文
关键词
Linear Approximation,Function Approximation,Linear Function Approximation,Discretion,Policy Evaluation,Markov Decision Process,Error Bounds,Convergence In Case,Learning Rate,Value Function,Cost Function,State Space,Fixed Point,Vector Function,Convergence Of Algorithm,Measurement Invariance,Mapping Project,Restrictive Assumptions,Maximum Norm,Admission Policies,Composition Operator,Markov Kernel,Transition Kernel,Greedy Policy,Borel Measurable,Borel Set,Invariant Distribution,Greedy Selection,Joint Processing,Proof Of Theorem