Fairbank, Michael and Alonso, Eduardo (2012) The divergence of reinforcement learning algorithms with value-iteration and function approximation. In: 2012 International Joint Conference on Neural Networks (IJCNN 2012 - Brisbane), 2012-06-10 - 2012-06-15, Brisbane, QLD, Australia.
Fairbank, Michael and Alonso, Eduardo (2012) The divergence of reinforcement learning algorithms with value-iteration and function approximation. In: 2012 International Joint Conference on Neural Networks (IJCNN 2012 - Brisbane), 2012-06-10 - 2012-06-15, Brisbane, QLD, Australia.
Fairbank, Michael and Alonso, Eduardo (2012) The divergence of reinforcement learning algorithms with value-iteration and function approximation. In: 2012 International Joint Conference on Neural Networks (IJCNN 2012 - Brisbane), 2012-06-10 - 2012-06-15, Brisbane, QLD, Australia.
Abstract
This paper gives specific divergence examples of value-iteration for several major Reinforcement Learning and Adaptive Dynamic Programming algorithms, when using a function approximator for the value function. These divergence examples differ from previous divergence examples in the literature, in that they are applicable for a greedy policy, i.e. in a “value iteration” scenario. Perhaps surprisingly, with a greedy policy, it is also possible to get divergence for the algorithms TD(1) and Sarsa(1). In addition to these divergences, we also achieve divergence for the Adaptive Dynamic Programming algorithms HDP, DHP and GDHP.
| Item Type: | Conference or Workshop Item (Paper) |
|---|---|
| Uncontrolled Keywords: | Trajectory, Approximation algorithms, Vectors, Heuristic algorithms, Equations, Function approximation |
| Divisions: | Faculty of Science and Health Faculty of Science and Health > Computer Science and Electronic Engineering, School of |
| SWORD Depositor: | Unnamed user with email elements@essex.ac.uk |
| Depositing User: | Unnamed user with email elements@essex.ac.uk |
| Date Deposited: | 02 Sep 2026 13:49 |
| Last Modified: | 02 Sep 2026 13:49 |
| URI: | http://repository.essex.ac.uk/id/eprint/43787 |
Available files
Filename: The divergence of reinforcement learning algorithms with value-iteration and function approximation.pdf