Source author record

Arnab Chatterjee

Arnab Chatterjee appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

30works
14topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

30 published item(s)

preprint2021arXiv

CAMTA: Causal Attention Model for Multi-touch Attribution

Advertising channels have evolved from conventional print media, billboards and radio advertising to online digital advertising (ad), where the users are exposed to a sequence of ad campaigns via social networks, display ads, search etc. While advertisers revisit the design of ad campaigns to concurrently serve the requirements emerging out of new ad channels, it is also critical for advertisers to estimate the contribution from touch-points (view, clicks, converts) on different channels, based on the sequence of customer actions. This process of contribution measurement is often referred to as multi-touch attribution (MTA). In this work, we propose CAMTA, a novel deep recurrent neural network architecture which is a casual attribution mechanism for user-personalised MTA in the context of observational data. CAMTA minimizes the selection bias in channel assignment across time-steps and touchpoints. Furthermore, it utilizes the users' pre-conversion actions in a principled way in order to predict pre-channel attribution. To quantitatively benchmark the proposed MTA model, we employ the real world Criteo dataset and demonstrate the superior performance of CAMTA with respect to prediction accuracy as compared to several baselines. In addition, we provide results for budget allocation and user-behaviour modelling on the predicted channel attribution.

preprint2021arXiv

MetaCI: Meta-Learning for Causal Inference in a Heterogeneous Population

Performing inference on data obtained through observational studies is becoming extremely relevant due to the widespread availability of data in fields such as healthcare, education, retail, etc. Furthermore, this data is accrued from multiple homogeneous subgroups of a heterogeneous population, and hence, generalizing the inference mechanism over such data is essential. We propose the MetaCI framework with the goal of answering counterfactual questions in the context of causal inference (CI), where the factual observations are obtained from several homogeneous subgroups. While the CI network is designed to generalize from factual to counterfactual distribution in order to tackle covariate shift, MetaCI employs the meta-learning paradigm to tackle the shift in data distributions between training and test phase due to the presence of heterogeneity in the population, and due to drifts in the target distribution, also known as concept shift. We benchmark the performance of the MetaCI algorithm using the mean absolute percentage error over the average treatment effect as the metric, and demonstrate that meta initialization has significant gains compared to randomly initialized networks, and other methods.

preprint2020arXiv

Effect of lockdown interventions to control the COVID-19 epidemic in India

The pandemic caused by the novel Coronavirus SARS-CoV2 has been responsible for life threatening health complications, and extreme pressure on healthcare systems. While preventive and definite curative medical interventions are yet to arrive, Non-Pharmaceutical Interventions (NPIs) like physical isolation, quarantine and drastic social measures imposed by governing agencies are effective in arresting the spread of infections in a population. In densely populated countries like India, lockdown interventions are partially effective due to social and administrative complexities. Using detailed demographic data, we present an agent based model to imitate the behavior of the population and its mobility features, even under intervention. We demonstrate the effectiveness of contact tracing policies and how our model efficiently relates to empirical findings on testing efficiency. We also present various lockdown intervention strategies for mitigation - using the bare number of infections, the effective reproduction rate, as well as using reinforcement learning. Our analysis can help assess the socio-economic consequences of such interventions, and provide useful ideas and insights to policy makers for better decision making.

preprint2020arXiv

The Ising universality class of kinetic exchange models of opinion dynamics

We show using scaling arguments and Monte Carlo simulations that a class of binary interacting models of opinion evolution belong to the Ising universality class in presence of an annealed noise term of finite amplitude. While the zero noise limit is known to show an active-absorbing transition, addition of annealed noise induces a continuous order-disorder transition with Ising universality class in the infinite-range (mean field) limit of the models.

preprint2019arXiv

Characterizing behavioral trends in a community driven discussion platform

This article presents a systematic analysis of the patterns of behavior of individuals as well as groups observed in community-driven platforms for discussion like Reddit, where users usually exchange information and viewpoints on their topics of interest. We perform a statistical analysis of the behavior of posts and model the users' interactions around them. A platform like Reddit which has grown exponentially, starting from a very small community to one of the largest social networks, with its large user base and popularity harboring a variety of behavior of users in terms of their activity. Our work provides interesting insights about a huge number of inactive posts which fail to attract attention despite their authors exhibiting Cyborg-like behavior to attract attention. We also observe short-lived yet extremely active posts emulate a phenomenon like Mayfly Buzz. A method is presented, to study the activity around posts which are highly active, to determine the presence of Limelight hogging activity. We also present a systematic analysis to study the presence of controversies in posts. We analyzed data from two periods of one-year duration but separated by few years in time, to understand how social media has evolved through the years.

preprint2016arXiv

Disorder induced phase transition in an opinion dynamics model: results in 2 and 3 dimensions

We study a model of continuous opinion dynamics with both positive and negative mutual interaction. The model shows a continuous phase transition between a phase with consensus (order) and a phase having no consensus (disorder). The mean field version of the model was already studied. Using extensive numerical simulations, we study the same model in $2$ and $3$ dimensions. The critical points of the phase transitions for various cases and the associated critical exponents have been estimated. The universality class of the phase transitions in the model is found to be same as Ising model in the respective dimensions.

preprint2016arXiv

Inequality measures in kinetic exchange models of wealth distributions

In this paper, we study the inequality indices for some models of wealth exchange. We calculated Gini index and newly introduced k-index and compare the results with reported empirical data available for different countries. We have found lower and upper bounds for the indices and discuss the efficiencies of the models. Some exact analytical calculations are given for a few cases. We also exactly compute the quantities for Gamma and double Gamma distributions.

preprint2016arXiv

Socio-economic inequality and prospects of institutional Econophysics

Socio-economic inequality is measured using various indices. The Gini ($g$) index, giving the overall inequality is the most commonly used, while the recently introduced Kolkata ($k$) index gives a measure of $1-k$ fraction of population who possess top $k$ fraction of wealth in the society. This article reviews the character of such inequalities, as seen from a variety of data sources, the apparent relationship between the two indices, and what toy models tell us. These socio-economic inequalities are also investigated in the context of man-made social conflicts or wars, as well as in natural disasters. Finally, we forward a proposal for an international institution with sufficient fund for visitors, where natural and social scientists from various institutions of the world can come to discuss, debate and formulate further developments.

preprint2016arXiv

Socio-economic inequality: Relationship between Gini and Kolkata indices

Socio-economic inequality is characterized from data using various indices. The Gini ($g$) index, giving the overall inequality is the most common one, while the recently introduced Kolkata ($k$) index gives a measure of $1-k$ fraction of population who possess top $k$ fraction of wealth in the society. Here, we show the relationship between the two indices, using both empirical data and analytical estimates. The significance of their relationship has been discussed.

preprint2015arXiv

A theoretical model for the associative nature of conference participation

Participation in conferences is an important part of every scientific career. Conferences provide an opportunity for a fast dissemination of latest results, discussion and exchange of ideas, and broadening of scientists' collaboration network. The decision to participate in a conference depends on several factors like the location, cost, popularity of keynote speakers, and the scientists' association with the community. Here we discuss and formulate the problem of discovering how a scientists' previous participation affects her/his future participations in the same conference series. We develop a stochastic model to examine scientists' participation patterns in conferences and compare our model with data from six conferences across various scientific fields and communities. Our model shows that the probability for a scientist to participate in a given conference series strongly depends on the balance between the number of participations and non-participations during his/her early connections with the community. An active participation in a conference series strengthens the scientists' association with that particular conference community and thus increases the probability of future participations.

preprint2015arXiv

Invariant features of spatial inequality in consumption: the case of India

We study the distributional features and inequality of consumption expenditure across India, for different states, castes, religion and urban-rural divide. We find that even though the aggregate measures of inequality are fairly diversified across states, the consumption distributions show near identical statistics, once properly normalized. This feature is seen to be robust with respect to variations in sociological and economic factors. We also show that state-wise inequality seems to be positively correlated with growth which is in accord with the traditional idea of Kuznets' curve. We present a brief model to account for the invariance found empirically and show that better but riskier technology draws can create a positive correlation between inequality and growth.

preprint2015arXiv

Social inequality: from data to statistical physics modeling

Social inequality is a topic of interest since ages, and has attracted researchers across disciplines to ponder over it origin, manifestation, characteristics, consequences, and finally, the question of how to cope with it. It is manifested across different strata of human existence, and is quantified in several ways. In this review we discuss the origins of social inequality, the historical and commonly used non-entropic measures such as Lorenz curve, Gini index and the recently introduced $k$ index. We also discuss some analytical tools that aid in understanding and characterizing them. Finally, we argue how statistical physics modeling helps in reproducing the results and interpreting them.

preprint2015arXiv

Universality of citation distributions for academic institutions and journals

Citations measure the importance of a publication, and may serve as a proxy for its popularity and quality of its contents. Here we study the distributions of citations to publications from individual academic institutions for a single year. The average number of citations have large variations between different institutions across the world, but the probability distributions of citations for individual institutions can be rescaled to a common form by scaling the citations by the average number of citations for that institution. We find this feature seem to be universal for a broad selection of institutions irrespective of the average number of citations per article. A similar analysis for citations to publications in a particular journal in a single year reveals similar results. We find high absolute inequality for both these sets, Gini coefficients being around $0.66$ and $0.58$ for institutions and journals respectively. We also find that the top $25$% of the articles hold about $75$% of the total citations for institutions and the top $29$% of the articles hold about $71$% of the total citations for journals.

preprint2014arXiv

Measuring social inequality with quantitative methodology: analytical estimates and empirical data analysis by Gini and $k$ indices

Social inequality manifested across different strata of human existence can be quantified in several ways. Here we compute non-entropic measures of inequality such as Lorenz curve, Gini index and the recently introduced $k$ index analytically from known distribution functions. We characterize the distribution functions of different quantities such as votes, journal citations, city size, etc. with suitable fits, compute their inequality measures and compare with the analytical results. A single analytic function is often not sufficient to fit the entire range of the probability distribution of the empirical data, and fit better to two distinct functions with a single crossover point. Here we provide general formulas to calculate these inequality measures for the above cases. We attempt to specify the crossover point by minimizing the gap between empirical and analytical evaluations of measures. Regarding the $k$ index as an `extra dimension', both the lower and upper bounds of the Gini index are obtained as a function of the $k$ index. This type of inequality relations among inequality indices might help us to check the validity of empirical and analytical evaluations of those indices.

preprint2014arXiv

On the evolution and utility of annual citation indices

We study the statistics of citations made to the top ranked indexed journals for Science and Social Science databases in the Journal Citation Reports using different measures. Total annual citation and impact factor, as well as a third measure called the annual citation rate are used to make the detailed analysis. We observe that the distribution of the annual citation rate has an universal feature - it shows a maximum at the rate scaled by half the average, irrespective of how the journals are ranked, and even across Science and Social Science journals, and fits well to log-Gumbel distribution. Correlations between different quantities are studied and a comparative analysis of the three measures is presented. The newly introduced annual citation rate factor helps in understanding the effect of scaling the number of citation by the total number of publications. The effect of the impact factor on authors contributing to the journals as well as on editorial policies is also discussed.

preprint2014arXiv

Socio-economic inequalities: a statistical physics perspective

Socio-economic inequalities are manifested in different aspects of our social life. We discuss various aspects, beginning with the evolutionary and historical origins, and discussing the major issues from the social and economic point of view. The subject has attracted scholars from across various disciplines, including physicists, who bring in a unique perspective to the field. The major attempts to analyze the results, address the causes, and understand the origins using statistical tools and statistical physics concepts are discussed.

preprint2014arXiv

Statistical Mechanics of Competitive Resource Allocation using Agent-based Models

Demand outstrips available resources in most situations, which gives rise to competition, interaction and learning. In this article, we review a broad spectrum of multi-agent models of competition (El Farol Bar problem, Minority Game, Kolkata Paise Restaurant problem, Stable marriage problem, Parking space problem and others) and the methods used to understand them analytically. We emphasize the power of concepts and tools from statistical mechanics to understand and explain fully collective phenomena such as phase transitions and long memory, and the mapping between agent heterogeneity and physical disorder. As these methods can be applied to any large-scale model of competitive resource allocation made up of heterogeneous adaptive agent with non-linear interaction, they provide a prospective unifying paradigm for many scientific disciplines.

preprint2014arXiv

Zipf's law in city size from a resource utilization model

We study a resource utilization scenario characterized by intrinsic fitness. To describe the growth and organization of different cities, we consider a model for resource utilization where many restaurants compete, as in a game, to attract customers using an iterative learning process. Results for the case of restaurants with uniform fitness are reported. When fitness is uniformly distributed, it gives rise to a Zipf law for the number of customers. We perform an exact calculation for the utilization fraction for the case when choices are made independent of fitness. A variant of the model is also introduced where the fitness can be treated as an ability to stay in the business. When a restaurant loses customers, its fitness is replaced by a random fitness. The steady state fitness distribution is characterized by a power law, while the distribution of the number of customers still follows the Zipf law, implying the robustness of the model. Our model serves as a paradigm for the emergence of Zipf law in city size distribution.

preprint2013arXiv

Universality in voting behavior: an empirical analysis

Election data represent a precious source of information to study human behavior at a large scale. In proportional elections with open lists, the number of votes received by a candidate, rescaled by the average performance of all competitors in the same party list, has the same distribution regardless of the country and the year of the election. Here we provide the first thorough assessment of this claim. We analyzed election datasets of 15 countries with proportional systems. We confirm that a class of nations with similar election rules fulfill the universality claim. Discrepancies from this trend in other countries with open-lists elections are always associated with peculiar differences in the election rules, which matter more than differences between countries and historical periods. Our analysis shows that the role of parties in the electoral performance of candidates is crucial: alternative scalings not taking into account party affiliations lead to poor results.

preprint2012arXiv

Continuous transition of social efficiencies in the stochastic strategy Minority Game

We show that in a variant of the Minority Game problem, the agents can reach a state of maximum social efficiency, where the fluctuation between the two choices is minimum, by following a simple stochastic strategy. By imagining a social scenario where the agents can only guess about the number of excess people in the majority, we show that as long as the guess value is sufficiently close to the reality, the system can reach a state of full efficiency or minimum fluctuation. A continuous transition to less efficient condition is observed when the guess value becomes worse. Hence, people can optimize their guess value for excess population to optimize the period of being in the majority state. We also consider the situation where a finite fraction of agents always decide completely randomly (random trader) as opposed to the rest of the population that follow a certain strategy (chartist). For a single random trader the system becomes fully efficient with majority-minority crossover occurring every two-days interval on average. For just two random traders, all the agents have equal gain with arbitrarily small fluctuations.

preprint2012arXiv

Disorder induced phase transition in kinetic models of opinion dynamics

We propose a model of continuous opinion dynamics, where mutual interactions can be both positive and negative. Different types of distributions for the interactions, all characterized by a single parameter $p$ denoting the fraction of negative interactions, are considered. Results from exact calculation of a discrete version and numerical simulations of the continuous version of the model indicate the existence of a universal continuous phase transition at p=p_c below which a consensus is reached. Although the order-disorder transition is analogous to a ferromagnetic-paramagnetic phase transition with comparable critical exponents, the model is characterized by some distinctive features relevant to a social system.

preprint2012arXiv

Phase transitions in crowd dynamics of resource allocation

We define and study a class of resources allocation processes where $gN$ agents, by repeatedly visiting $N$ resources, try to converge to optimal configuration where each resource is occupied by at most one agent. The process exhibits a phase transition, as the density $g$ of agents grows, from an absorbing to an active phase. In the latter, even if the number of resources is in principle enough for all agents ($g<1$), the system never settles to a frozen configuration. We recast these processes in terms of zero-range interacting particles, studying analytically the mean field dynamics and investigating numerically the phase transition in finite dimensions. We find a good agreement with the critical exponents of the stochastic fixed-energy sandpile. The lack of coordination in the active phase also leads to a non-trivial faster-is-slower effect.

preprint2011arXiv

Antipersistent dynamics in kinetic models of wealth exchange

We investigate the detailed dynamics of gains and losses made by agents in some kinetic models of wealth exchange. The concept of a walk in an abstract gain-loss space for the agents had been introduced in an earlier work. For models in which agents do not save, or save with uniform saving propensity, this walk has diffusive behavior. In case the saving propensity $λ$ is distributed randomly ($0 \leq λ< 1$), the resultant walk showed a ballistic nature (except at a particular value of $λ^* \approx 0.47$). Here we consider several other features of the walk with random $λ$. While some macroscopic properties of this walk are comparable to a biased random walk, at microscopic level, there are gross differences. The difference turns out to be due to an antipersistent tendency towards making a gain (loss) immediately after making a loss (gain). This correlation is in fact present in kinetic models without saving or with uniform saving as well, such that the corresponding walks are not identical to ordinary random walks. In the distributed saving case, antipersistence occurs with a simultaneous overall bias.

preprint2011arXiv

Phase transitions and non-equilibrium relaxation in kinetic models of opinion formation

We review in details some recently proposed kinetic models of opinion dynamics. We discuss the several variants including a generalised model. We provide mean field estimates for the critical points, which are numerically supported with reasonable accuracy. Using non-equilibrium relaxation techniques, we also investigate the nature of phase transitions observed in these models. We study the nature of correlations as the critical points are approached, and comment on the universality of the phase transitions observed.

preprint2010arXiv

A new route to Explosive Percolation

The biased link occupation rule in the Achlioptas process (AP) discourages the large clusters to grow much ahead of others and encourages faster growth of clusters which lag behind. In this paper we propose a model where this tendency is sharply reflected in the Gamma distribution of the cluster sizes, unlike the power law distribution in AP. In this model single edges between pairs of clusters of sizes $s_i$ and $s_j$ are occupied with a probability $\propto (s_is_j)^α$. The parameter $α$ is continuously tunable over the entire real axis. Numerical studies indicate that for $α< α_c$ the transition is first order, $α_c=0$ for square lattice and $α_c=-1/2$ for random graphs. In the limits of $α= -\infty, +\infty$ this model coincides with models well established in the literature.

preprint2010arXiv

Agent dynamics in kinetic models of wealth exchange

We study the dynamics of individual agents in some kinetic models of wealth exchange, particularly, the models with savings. For the model with uniform savings, agents perform simple random walks in the "wealth space". On the other hand, we observe ballistic diffusion in the model with distributed savings. There is an associated skewness in the gain-loss distribution which explains the steady state behavior in such models. We find that in general an agent gains while interacting with an agent with a larger saving propensity.

preprint2010arXiv

Statistics of the Kolkata Paise Restaurant Problem

We study the dynamics of a few stochastic learning strategies for the 'Kolkata Paise Restaurant' problem, where N agents choose among N equally priced but differently ranked restaurants every evening such that each agent tries get to dinner in the best restaurant (each serving only one customer and the rest arriving there going without dinner that evening). We consider the learning strategies to be similar for all the agents and assume that each follow the same probabilistic or stochastic strategy dependent on the information of the past successes in the game. We show that some 'naive' strategies lead to much better utilization of the services than some relatively 'smarter' strategies. We also show that the service utilization fraction as high as 0.80 can result for a stochastic strategy, where each agent sticks to his past choice (independent of success achieved or not; with probability decreasing inversely in the past crowd size). The numerical results for utilization fraction of the services in some limiting cases are analytically examined.

preprint2009arXiv

The Kolkata Paise Restaurant Problem and Resource Utilization

We study the dynamics of the "Kolkata Paise Restaurant problem". The problem is the following: In each period, N agents have to choose between N restaurants. Agents have a common ranking of the restaurants. Restaurants can only serve one customer. When more than one customer arrives at the same restaurant, one customer is chosen at random and is served; the others do not get the service. We first introduce the one-shot versions of the Kolkata Paise Restaurant problem which we call one-shot KPR games. We then study the dynamics of the Kolkata Paise Restaurant problem (which is a repeated game version of any given one shot KPR game) for large N. For statistical analysis, we explore the long time steady state behavior. In many such models with myopic agents we get under-utilization of resources, that is, we get a lower aggregate payoff compared to the social optimum. We study a number of myopic strategies, focusing on the average occupation fraction of restaurants.