Source author record

Antonio Lijoi

Antonio Lijoi appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

15works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

15 published item(s)

preprint2022arXiv

Flexible clustering via hidden hierarchical Dirichlet priors

The Bayesian approach to inference stands out for naturally allowing borrowing information across heterogeneous populations, with different samples possibly sharing the same distribution. A popular Bayesian nonparametric model for clustering probability distributions is the nested Dirichlet process, which however has the drawback of grouping distributions in a single cluster when ties are observed across samples. With the goal of achieving a flexible and effective clustering method for both samples and observations, we investigate a nonparametric prior that arises as the composition of two different discrete random structures and derive a closed-form expression for the induced distribution of the random partition, the fundamental tool regulating the clustering behavior of the model. On the one hand, this allows to gain a deeper insight into the theoretical properties of the model and, on the other hand, it yields an MCMC algorithm for evaluating Bayesian inferences of interest. Moreover, we single out limitations of this algorithm when working with more than two populations and, consequently, devise an alternative more efficient sampling scheme, which as a by-product, allows testing homogeneity between different populations. Finally, we perform a comparison with the nested Dirichlet process and provide illustrative examples of both synthetic and real data.

preprint2022arXiv

Smoothing distributions for conditional Fleming-Viot and Dawson-Watanabe diffusions

We study the distribution of the unobserved states of two measure-valued diffusions of Fleming-Viot and Dawson-Watanabe type, conditional on observations from the underlying populations collected at past, present and future times. If seen as nonparametric hidden Markov models, this amounts to finding the smoothing distributions of these processes, which we show can be explicitly described in recursive form as finite mixtures of laws of Dirichlet and gamma random measures respectively. We characterize the time-dependent weights of these mixtures, accounting for potentially different time intervals between data collection times, and fully describe the implications of assuming a discrete or a nonatomic distribution for the underlying process that drives mutations. In particular, we show that with a nonatomic mutation offspring distribution, the inference automatically upweights mixture components that carry, as atoms, observed types shared at different collection times. The predictive distributions for further samples from the population conditional on the data are also identified and shown to be mixtures of generalized Polya urns, conditionally on a latent variable in the Dawson-Watanabe case.

preprint2020arXiv

Predictive inference with Fleming--Viot-driven dependent Dirichlet processes

We consider predictive inference using a class of temporally dependent Dirichlet processes driven by Fleming--Viot diffusions, which have a natural bearing in Bayesian nonparametrics and lend the resulting family of random probability measures to analytical posterior analysis. Formulating the implied statistical model as a hidden Markov model, we fully describe the predictive distribution induced by these Fleming--Viot-driven dependent Dirichlet processes, for a sequence of observations collected at a certain time given another set of draws collected at several previous times. This is identified as a mixture of Pólya urns, whereby the observations can be values from the baseline distribution or copies of previous draws collected at the same time as in the usual Pòlya urn, or can be sampled from a random subset of the data collected at previous times. We characterise the time-dependent weights of the mixture which select such subsets and discuss the asymptotic regimes. We describe the induced partition by means of a Chinese restaurant process metaphor with a conveyor belt, whereby new customers who do not sit at an occupied table open a new table by picking a dish either from the baseline distribution or from a time-varying offer available on the conveyor belt. We lay out explicit algorithms for exact and approximate posterior sampling of both observations and partitions, and illustrate our results on predictive problems with synthetic and real data.

preprint2016arXiv

Canonical correlations for dependent gamma processes

The present paper provides a characterisation of exchangeable pairs of random measures $(\widetildeμ_1,\widetildeμ_2)$ whose identical margins are fixed to coincide with the distribution of a gamma completely random measure, and whose dependence structure is given in terms of canonical correlations. It is first shown that canonical correlation sequences for the finite-dimensional distributions of $(\widetildeμ_1,\widetildeμ_2)$ are moments of means of a Dirichlet process having random base measure. Necessary and sufficient conditions are further given for canonically correlated gamma completely random measures to have independent joint increments. Finally, time-homogeneous Feller processes with gamma reversible measure and canonical autocorrelations are characterised as Dawson--Watanabe diffusions with independent homogeneous immigration, time-changed via an independent subordinator. It is thus shown that Dawson--Watanabe diffusions subordinated by pure drift are the only processes in this class whose time-finite-dimensional distributions have, jointly, independent increments.

preprint2015arXiv

Bayesian Survival Model based on Moment Characterization

Bayesian nonparametric marginal methods are very popular since they lead to fairly easy implementation due to the formal marginalization of the infinite-dimensional parameter of the model. However, the straightforwardness of these methods also entails some limitations: they typically yield point estimates in the form of posterior expectations, but cannot be used to estimate non-linear functionals of the posterior distribution, such as median, mode or credible intervals. This is particularly relevant in survival analysis where non-linear functionals such as e.g. the median survival time, play a central role for clinicians and practitioners. The main goal of this paper is to summarize the methodology introduced in [Arbel et al., Comput. Stat. Data. An., 2015] for hazard mixture models in order to draw approximate Bayesian inference on survival functions that is not limited to the posterior mean. In addition, we propose a practical implementation of an R package called momentify designed for moment-based density approximation, and, by means of an extensive simulation study, we thoroughly compare the introduced methodology with standard marginal methods and empirical estimation.

preprint2014arXiv

Bayesian inference with dependent normalized completely random measures

The proposal and study of dependent prior processes has been a major research focus in the recent Bayesian nonparametric literature. In this paper, we introduce a flexible class of dependent nonparametric priors, investigate their properties and derive a suitable sampling scheme which allows their concrete implementation. The proposed class is obtained by normalizing dependent completely random measures, where the dependence arises by virtue of a suitable construction of the Poisson random measures underlying the completely random measures. We first provide general distributional results for the whole class of dependent completely random measures and then we specialize them to two specific priors, which represent the natural candidates for concrete implementation due to their analytic tractability: the bivariate Dirichlet and normalized $σ$-stable processes. Our analytical results, and in particular the partially exchangeable partition probability function, form also the basis for the determination of a Markov Chain Monte Carlo algorithm for drawing posterior inferences, which reduces to the well-known Blackwell--MacQueen Pólya urn scheme in the univariate case. Such an algorithm can be used for density estimation and for analyzing the clustering structure of the data and is illustrated through a real two-sample dataset example.

preprint2014arXiv

Full Bayesian inference with hazard mixture models

Bayesian nonparametric inferential procedures based on Markov chain Monte Carlo marginal methods typically yield point estimates in the form of posterior expectations. Though very useful and easy to implement in a variety of statistical problems, these methods may suffer from some limitations if used to estimate non-linear functionals of the posterior distribution. The main goal of the present paper is to develop a novel methodology that extends a well-established marginal procedure designed for hazard mixture models, in order to draw approximate inference on survival functions that is not limited to the posterior mean but includes, as remarkable examples, credible intervals and median survival time. Our approach relies on a characterization of the posterior moments that, in turn, is used to approximate the posterior distribution by means of a technique based on Jacobi polynomials. The inferential performance of our methodology is analyzed by means of an extensive study of simulated data and real data consisting of leukemia remission times. Although tailored to the survival analysis context, the procedure we introduce can be adapted to a range of other models for which moments of the posterior distribution can be estimated.

preprint2013arXiv

Conditional formulae for Gibbs-type exchangeable random partitions

Gibbs-type random probability measures and the exchangeable random partitions they induce represent an important framework both from a theoretical and applied point of view. In the present paper, motivated by species sampling problems, we investigate some properties concerning the conditional distribution of the number of blocks with a certain frequency generated by Gibbs-type random partitions. The general results are then specialized to three noteworthy examples yielding completely explicit expressions of their distributions, moments and asymptotic behaviors. Such expressions can be interpreted as Bayesian nonparametric estimators of the rare species variety and their performance is tested on some real genomic data.

preprint2013arXiv

Modeling with Normalized Random Measure Mixture Models

The Dirichlet process mixture model and more general mixtures based on discrete random probability measures have been shown to be flexible and accurate models for density estimation and clustering. The goal of this paper is to illustrate the use of normalized random measures as mixing measures in nonparametric hierarchical mixture models and point out how possible computational issues can be successfully addressed. To this end, we first provide a concise and accessible introduction to normalized random measures with independent increments. Then, we explain in detail a particular way of sampling from the posterior using the Ferguson-Klass representation. We develop a thorough comparative analysis for location-scale mixtures that considers a set of alternatives for the mixture kernel and for the nonparametric component. Simulation results indicate that normalized random measure mixtures potentially represent a valid default choice for density estimation problems. As a byproduct of this study an R package to fit these models was produced and is available in the Comprehensive R Archive Network (CRAN).

preprint2012arXiv

A Conversation with Eugenio Regazzini

Eugenio Regazzini was born on August 12, 1946 in Cremona (Italy), and took his degree in 1969 at the University "L. Bocconi" of Milano. He has held positions at the universities of Torino, Bologna and Milano, and at the University "L. Bocconi" as assistant professor and lecturer from 1974 to 1980, and then professor since 1980. He is currently professor in probability and mathematical statistics at the University of Pavia. In the periods 1989-2001 and 2006-2009 he was head of the Institute for Applications of Mathematics and Computer Science of the Italian National Research Council (C.N.R.) in Milano and head of the Department of Mathematics at the University of Pavia, respectively. For twelve years between 1989 and 2006, he served as a member of the Scientific Board of the Italian Mathematical Union (U.M.I.). In 2007, he was elected Fellow of the IMS and, in 2001, Fellow of the "Istituto Lombardo---Accademia di Scienze e Lettere." His research activity in probability and statistics has covered a wide spectrum of topics, including finitely additive probabilities, foundations of the Bayesian paradigm, exchangeability and partial exchangeability, distribution of functionals of random probability measures, stochastic integration, history of probability and statistics. Overall, he has been one of the most authoritative developers of de Finetti's legacy. In the last five years, he has extended his scientific interests to probabilistic methods in mathematical physics; in particular, he has studied the asymptotic behavior of the solutions of equations, which are of interest for the kinetic theory of gases. The present interview was taken in occasion of his 65th birthday.

preprint2012arXiv

Asymptotics for a Bayesian nonparametric estimator of species variety

In Bayesian nonparametric inference, random discrete probability measures are commonly used as priors within hierarchical mixture models for density estimation and for inference on the clustering of the data. Recently, it has been shown that they can also be exploited in species sampling problems: indeed they are natural tools for modeling the random proportions of species within a population thus allowing for inference on various quantities of statistical interest. For applications that involve large samples, the exact evaluation of the corresponding estimators becomes impracticable and, therefore, asymptotic approximations are sought. In the present paper, we study the limiting behaviour of the number of new species to be observed from further sampling, conditional on observed data, assuming the observations are exchangeable and directed by a normalized generalized gamma process prior. Such an asymptotic study highlights a connection between the normalized generalized gamma process and the two-parameter Poisson-Dirichlet process that was previously known only in the unconditional case.

preprint2010arXiv

Limiting behavior of the search cost distribution for the move-to-front rule in the stable case

Move-to-front rule is a heuristic updating a list of n items according to requests. Items are required with unknown probabilities (or popularities). The induced Markov chain is known to be ergodic. One main problem is the study of the distribution of the search cost dened as the position of the required item. Here we first establish the link between two recent papers that both extend results proved by Kingman on the expected stationary search cost. Combining results contained in these papers, we obtain the limiting behavior for any moments of the stationary seach cost as n tends to innity.

preprint2010arXiv

On the posterior distribution of classes of random means

The study of properties of mean functionals of random probability measures is an important area of research in the theory of Bayesian nonparametric statistics. Many results are now known for random Dirichlet means, but little is known, especially in terms of posterior distributions, for classes of priors beyond the Dirichlet process. In this paper, we consider normalized random measures with independent increments (NRMI's) and mixtures of NRMI. In both cases, we are able to provide exact expressions for the posterior distribution of their means. These general results are then specialized, leading to distributional results for means of two important particular cases of NRMI's and also of the two-parameter Poisson--Dirichlet process.

preprint2004arXiv

Means of a Dirichlet process and multiple hypergeometric functions

The Lauricella theory of multiple hypergeometric functions is used to shed some light on certain distributional properties of the mean of a Dirichlet process. This approach leads to several results, which are illustrated here. Among these are a new and more direct procedure for determining the exact form of the distribution of the mean, a correspondence between the distribution of the mean and the parameter of a Dirichlet process, a characterization of the family of Cauchy distributions as the set of the fixed points of this correspondence, and an extension of the Markov-Krein identity. Moreover, an expression of the characteristic function of the mean of a Dirichlet process is obtained by resorting to an integral representation of a confluent form of the fourth Lauricella function. This expression is then employed to prove that the distribution of the mean of a Dirichlet process is symmetric if and only if the parameter of the process is symmetric, and to provide a new expression of the moment generating function of the variance of a Dirichlet process.