Source author record

Yanling Chang

Yanling Chang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
3topics
2close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2020arXiv

Worst-Case Analysis for a Leader-follower Partially Observable Stochastic Game

Partially observable stochastic games provide a rich mathematical paradigm for modeling multi-agent dynamic decision making under uncertainty and partial information. However, they generally do not admit closed-form solutions and are notoriously difficult to solve. Also, in reality, each agent often does not have complete knowledge of the other agent. This paper studies a leader-follower partially observable stochastic game where the leader has little knowledge of the adversarial follower's reward structure, level of rationality, and process for gathering and transmitting data relevant for decision making. We introduce the worst-case analysis to the partially observable stochastic game to cope with this lack of knowledge and determine the best worst-case value function of the leader. The resulting problem from the leader's perspective has a simple sufficient statistic; however, different from a classical partially observable Markov decision process, the value function of the resulting problem may not be convex. We design a viable and computationally attractive solution procedure for computing a lower bound of the leader's value function as well as its associated control policy in the finite planning horizon. We illustrate the use of the proposed approach in a liquid egg production security problem.

preprint2014arXiv

Partially Observed, Multi-objective Markov Games

The intent of this research is to generate a set of non-dominated policies from which one of two agents (the leader) can select a most preferred policy to control a dynamic system that is also affected by the control decisions of the other agent (the follower). The problem is described by an infinite horizon, partially observed Markov game (POMG). At each decision epoch, each agent knows: its past and present states, its past actions, and noise corrupted observations of the other agent's past and present states. The actions of each agent are determined at each decision epoch based on these data. The leader considers multiple objectives in selecting its policy. The follower considers a single objective in selecting its policy with complete knowledge of and in response to the policy selected by the leader. This leader-follower assumption allows the POMG to be transformed into a specially structured, partially observed Markov decision process (POMDP). This POMDP is used to determine the follower's best response policy. A multi-objective genetic algorithm (MOGA) is used to create the next generation of leader policies based on the fitness measures of each leader policy in the current generation. Computing a fitness measure for a leader policy requires a value determination calculation, given the leader policy and the follower's best response policy. The policies from which the leader can select a most preferred policy are the non-dominated policies of the final generation of leader policies created by the MOGA. An example is presented that illustrates how these results can be used to support a manager of a liquid egg production process (the leader) in selecting a sequence of actions to best control this process over time, given that there is an attacker (the follower) who seeks to contaminate the liquid egg production process with a chemical or biological toxin.