Source author record

Paul Reverdy

Paul Reverdy appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

4works
5topics
3close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

4 published item(s)

preprint2020arXiv

Dynamical, value-based decision making among $N$ options

Decision making is a fundamental capability of autonomous systems. As decision making is a process which happens over time, it can be well modeled by dynamical systems. Often, decisions are made on the basis of perceived values of the underlying options and the desired outcome is to select the option with the highest value. This can be encoded as a bifurcation which produces a stable equilibrium corresponding to the high-value option. When some options have identical values, it is natural to design the decision-making model to be indifferent among the equally-valued options, leading to symmetries in the underlying dynamical system. For example, when all $N$ options have identical values, the dynamical system should have $S_N$ symmetry. Unfortunately, constructing a dynamical system that unfolds the $S_N$-symmetric pitchfork bifurcation is non-trivial. In this paper, we develop a method to construct an unfolding of the pitchfork bifurcation with a symmetry group that is a significant subgroup of $S_N$. The construction begins by parsing the decision among $N$ options into a hierarchical set of $N-1$ binary decisions encoded in a binary tree. By associating the unfolding of a standard $S_2$-symmetric pitchfork bifurcation with each of these binary decisions, we develop an unfolding of the pitchfork bifurcation with symmetries corresponding to isomorphisms of the underlying binary tree.

preprint2016arXiv

Satisficing in multi-armed bandit problems

Satisficing is a relaxation of maximizing and allows for less risky decision making in the face of uncertainty. We propose two sets of satisficing objectives for the multi-armed bandit problem, where the objective is to achieve reward-based decision-making performance above a given threshold. We show that these new problems are equivalent to various standard multi-armed bandit problems with maximizing objectives and use the equivalence to find bounds on performance. The different objectives can result in qualitatively different behavior; for example, agents explore their options continually in one case and only a finite number of times in another. For the case of Gaussian rewards we show an additional equivalence between the two sets of satisficing objectives that allows algorithms developed for one set to be applied to the other. We then develop variants of the Upper Credible Limit (UCL) algorithm that solve the problems with satisficing objectives and show that these modified UCL algorithms achieve efficient satisficing performance.

preprint2015arXiv

Correlated Multiarmed Bandit Problem: Bayesian Algorithms and Regret Analysis

We consider the correlated multiarmed bandit (MAB) problem in which the rewards associated with each arm are modeled by a multivariate Gaussian random variable, and we investigate the influence of the assumptions in the Bayesian prior on the performance of the upper credible limit (UCL) algorithm and a new correlated UCL algorithm. We rigorously characterize the influence of accuracy, confidence, and correlation scale in the prior on the decision-making performance of the algorithms. Our results show how priors and correlation structure can be leveraged to improve performance.

preprint2015arXiv

Parameter estimation in softmax decision-making models with linear objective functions

With an eye towards human-centered automation, we contribute to the development of a systematic means to infer features of human decision-making from behavioral data. Motivated by the common use of softmax selection in models of human decision-making, we study the maximum likelihood parameter estimation problem for softmax decision-making models with linear objective functions. We present conditions under which the likelihood function is convex. These allow us to provide sufficient conditions for convergence of the resulting maximum likelihood estimator and to construct its asymptotic distribution. In the case of models with nonlinear objective functions, we show how the estimator can be applied by linearizing about a nominal parameter value. We apply the estimator to fit the stochastic UCL (Upper Credible Limit) model of human decision-making to human subject data. We show statistically significant differences in behavior across related, but distinct, tasks.