Fitted Q-iteration in Continuous Action-space MDPs

Fitted Q-iteration in Continuous Action-space MDPs. Antos, A., Munos, R., & Szepesvári, C. In Advances in Neural Information Processing Systems, pages 9–16, 2007.

Paper abstract bibtex 3 downloads

We consider continuous state, continuous action batch reinforcement learning where the goal is to learn a good policy from a sufficiently rich trajectory generated by some policy. We study a variant of fitted Q-iteration, where the greedy action selection is replaced by searching for a policy in a restricted set of candidate policies by maximizing the average action values. We provide a rigorous analysis of this algorithm, proving what we believe is the first finite-time bound for value-function based algorithms for continuous state and action problems. Note: In retrospect, it would have been better to call this algorithm an actor-critic algorithm. The algorithm that we considers updates a policy and a value function (action-value function in this case).

@inproceedings{antos2007,
	abstract = {We consider continuous state, continuous action batch reinforcement learning where the goal is to learn a good policy from a sufficiently rich trajectory generated by some policy. We study a variant of fitted Q-iteration, where the greedy action selection is replaced by searching for a policy in a restricted set of candidate policies by maximizing the average action values. We provide a rigorous analysis of this algorithm, proving what we believe is the first finite-time bound for value-function based algorithms for continuous state and action problems.

		  Note: In retrospect, it would have been better to call this algorithm an actor-critic algorithm. The algorithm that we considers updates a policy and a value function (action-value function in this case).},
	author = {Antos, A. and Munos, R. and Szepesv{\'a}ri, Cs.},
	booktitle = {Advances in Neural Information Processing Systems},
	keywords = {batch learning, reinforcement learning, function approximation, performance bounds, actor-critic methods, nonparametrics},
	pages = {9--16},
	title = {Fitted Q-iteration in Continuous Action-space MDPs},
	url_paper = {rlca.pdf},
	year = {2007}}

Downloads: 3