Source author record

Jalal Arabneydi

Jalal Arabneydi appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
2topics
2close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2021arXiv

Reinforcement Learning in Deep Structured Teams: Initial Results with Finite and Infinite Valued Features

In this paper, we consider Markov chain and linear quadratic models for deep structured teams with discounted and time-average cost functions under two non-classical information structures, namely, deep state sharing and no sharing. In deep structured teams, agents are coupled in dynamics and cost functions through deep state, where deep state refers to a set of orthogonal linear regressions of the states. In this article, we consider a homogeneous linear regression for Markov chain models (i.e., empirical distribution of states) and a few orthonormal linear regressions for linear quadratic models (i.e., weighted average of states). Some planning algorithms are developed for the case when the model is known, and some reinforcement learning algorithms are proposed for the case when the model is not known completely. The convergence of two model-free (reinforcement learning) algorithms, one for Markov chain models and one for linear quadratic models, is established. The results are then applied to a smart grid.

preprint2020arXiv

Deep Structured Teams with Linear Quadratic Model: Partial Equivariance and Gauge Transformation

Motivated by the recent developments in artificial intelligence, we introduce linear quadratic deep structured teams in this paper. Two notions of equivariant and partially equivariant systems are defined, and it is shown that such systems can be partitioned into a few sub-populations of decision makers, where every decision maker in each sub-population is coupled in both dynamics and cost function through a set of linear regressions of the states and actions of all decision makers. Two non-classical information structures are considered: deep-state sharing and partial deep-state sharing, where deep state refers to the linear regression of the states of the decision makers in each sub-population. For a risk-sensitive cost function with deep-state sharing structure, a closed-form low-complexity representation of the globally optimal strategy is obtained, whose computational complexity is independent of the number of decision makers in each sub-population. In addition, it is shown that the risk-sensitive solution converges to the risk-neutral one as the number of decision makers increases to infinity. Moreover, two sub-optimal sequential strategies under partial deep-state sharing information structure are proposed by introducing two Kalman-like filters, one based on the finite-population model and the other one based on the infinite-population model. It is proved that the prices of information associated with the above sub-optimal solutions converge to zero as the number of decision makers goes to infinity. Furthermore, a class of feed-forward deep neural networks with multiple layers of weighted sums and products is introduced wherein the optimal weights and biases are explicitly obtained. A supply-chain management example is presented to demonstrate the efficacy of the obtained results.

preprint2020arXiv

Deep Teams: Decentralized Decision Making with Finite and Infinite Number of Agents

Inspired by the concepts of deep learning in artificial intelligence and fairness in behavioural economics, we introduce deep teams in this paper. In such systems, agents are partitioned into a few sub-populations so that the dynamics and cost of agents in each sub-population is invariant to the indexing of agents. The goal of agents is to minimize a common cost function in such a manner that the agents in each sub-population are not discriminated or privileged by the way they are indexed. Two non-classical information structures are studied. In the first one, each agent observes its local state as well as the empirical distribution of the states of agents in each sub-population, called deep state, whereas in the second one, the deep states of a subset (possibly all) of sub-populations are not observed. Novel dynamic programs are developed to identify globally optimal and sub-optimal solutions for the first and second information structures, respectively. The computational complexity of finding the optimal solution in both space and time is polynomial (rather than exponential) with respect to the number of agents in each sub-population and is linear (rather than exponential) with respect to the control horizon. This complexity is further reduced in time by introducing a forward equation, that we call deep Chapman-Kolmogorov equation, described by multiple convolutional layers of Binomial probability distributions. Two different prices are defined for computation and communication, and it is shown that under mild assumptions they converge to zero as the quantization level and the number of agents tend to infinity. In addition, the main results are extended to the infinite-horizon discounted model and arbitrarily asymmetric cost function. Finally, a service-management example with 200 users is presented.