Graph explorer

Quantile Reinforcement Learning

In reinforcement learning, the standard criterion to evaluate policies in a state is the expectation of (discounted) sum of rewards. However, this criterion may not always be suitable, we consider an alternative criterion based on the notion of quantiles. In the case of episodic reinforcement learning problems, we propose an algorithm based on stochastic approximation with two timescales. We evaluate our proposition on a simple model of the TV show, Who wants to be a millionaire.

5 nodes5 linksoverview mapQuantile Reinforcement Learning
5 nodes5 links
Quantile Reinforcement Learning5 visible / 5 total nodes / 6 links
Related contextCo-authorshipAuthorshipAuthorshipTopic signalTopic signalWQuantile Reinforcement Learningpreprint / 2016AHugo GilbertResearcherAPaul WengResearcherTMachine Learning49008 worksTArtificial Intelligence22915 works
PaperSignal 104 links

Quantile Reinforcement Learning

preprint / 2016

Open