Source author record

Peter Neal

Peter Neal appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
3topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2016arXiv

Optimal scaling of the independence sampler: Theory and Practice

The independence sampler is one of the most commonly used MCMC algorithms usually as a component of a Metropolis-within-Gibbs algorithm. The common focus for the independence sampler is on the choice of proposal distribution to obtain an as high as possible acceptance rate. In this paper we have a somewhat different focus concentrating on the use of the independence sampler for updating augmented data in a Bayesian framework where a natural proposal distribution for the independence sampler exists. Thus we concentrate on the proportion of the augmented data to update to optimise the independence sampler. Generic guidelines for optimising the independence sampler are obtained for independent and identically distributed product densities mirroring findings for the random walk Metropolis algorithm. The generic guidelines are shown to be informative beyond the narrow confines of idealised product densities in two epidemic examples.

preprint2015arXiv

Model comparison with missing data using MCMC and importance sampling

Selecting between competing statistical models is a challenging problem especially when the competing models are non-nested. In this paper we offer a simple solution by devising an algorithm which combines MCMC and importance sampling to obtain computationally efficient estimates of the marginal likelihood which can then be used to compare the models. The algorithm is successfully applied to longitudinal epidemic and time series data sets and shown to outperform existing methods for computing the marginal likelihood.

preprint2014arXiv

On expected durations of birth-death processes, with applications to branching processes and SIS epidemics

We study continuous-time birth-death type processes, where individuals have independent and identically distributed lifetimes, according to a random variable Q, with E[Q]=1, and where the birth rate if the population is currently in state (has size) n is α(n). We focus on two important examples, namely α(n)=λn being a branching process, and α(n)=λn(N-n)/N which corresponds to an SIS epidemic model in a homogeneously mixing community of fixed size N. The processes are assumed to start with a single individual, i.e. in state 1. Let T, A_n, C and S denote the (random) time to extinction, the total time spent in state $n$, the total number of individuals ever alive and the sum of the lifetimes of all individuals in the birth-death process, respectively. The main results of the paper give expressions for the expectation of all these quantities, and shows that these expectations are insensitive to the distribution of Q. We also derive an asymptotic expression for the expected time to extinction of the SIS epidemic, but now starting at the endemic state, which is not independent of the distribution of Q. The results are also applied to the household SIS epidemic, showing that its threshold parameter R_* is insensitive to the distribution of Q, contrary to the household SIR epidemic, for which R_* does depend on Q.

preprint2014arXiv

Simulation based sequential Monte Carlo methods for discretely observed Markov processes

Parameter estimation for discretely observed Markov processes is a challenging problem. However, simulation of Markov processes is straightforward using the Gillespie algorithm. We exploit this ease of simulation to develop an effective sequential Monte Carlo (SMC) algorithm for obtaining samples from the posterior distribution of the parameters. In particular, we introduce two key innovations, coupled simulations, which allow us to study multiple parameter values on the basis of a single simulation, and a simple, yet effective, importance sampling scheme for steering simulations towards the observed data. These innovations substantially improve the efficiency of the SMC algorithm with minimal effect on the speed of the simulation process. The SMC algorithm is successfully applied to two examples, a Lotka-Volterra model and a Repressilator model.

preprint2013arXiv

On the expected time a branching process has K individuals alive

Consider a homogeneous time-continuous branching process where individuals have constant birth rate $δ$, and life length distribution $Q$ having mean $E(Q)=1$. Let $X(u)$ denote the number of individuals alive at time $u$, and assume that $X(0)=1$. Let $K$ be a positive integer and define $A_K:=\int_0^\infty 1_{\{X(u)=K\}}du$, the accumulated time that the branching process has exactly $K$ individuals alive. In this paper we prove that $E(A_K)=δ^{K-1}/\left(k(1\veeδ)^K\right)$, irrespective of the life length distribution $Q$, subject to the normalizing condition $E(Q)=1$.

preprint2013arXiv

Robust estimation of microbial diversity in theory and in practice

Quantifying diversity is of central importance for the study of structure, function and evolution of microbial communities. The estimation of microbial diversity has received renewed attention with the advent of large-scale metagenomic studies. Here, we consider what the diversity observed in a sample tells us about the diversity of the community being sampled. First, we argue that one cannot reliably estimate the absolute and relative number of microbial species present in a community without making unsupported assumptions about species abundance distributions. The reason for this is that sample data do not contain information about the number of rare species in the tail of species abundance distributions. We illustrate the difficulty in comparing species richness estimates by applying Chao's estimator of species richness to a set of in silico communities: they are ranked incorrectly in the presence of large numbers of rare species. Next, we extend our analysis to a general family of diversity metrics ("Hill diversities"), and construct lower and upper estimates of diversity values consistent with the sample data. The theory generalizes Chao's estimator, which we retrieve as the lower estimate of species richness. We show that Shannon and Simpson diversity can be robustly estimated for the in silico communities. We analyze nine metagenomic data sets from a wide range of environments, and show that our findings are relevant for empirically-sampled communities. Hence, we recommend the use of Shannon and Simpson diversity rather than species richness in efforts to quantify and compare microbial diversity.

preprint2012arXiv

Optimal scaling of random walk Metropolis algorithms with discontinuous target densities

We consider the optimal scaling problem for high-dimensional random walk Metropolis (RWM) algorithms where the target distribution has a discontinuous probability density function. Almost all previous analysis has focused upon continuous target densities. The main result is a weak convergence result as the dimensionality d of the target densities converges to infinity. In particular, when the proposal variance is scaled by $d^{-2}$, the sequence of stochastic processes formed by the first component of each Markov chain converges to an appropriate Langevin diffusion process. Therefore optimizing the efficiency of the RWM algorithm is equivalent to maximizing the speed of the limiting diffusion. This leads to an asymptotic optimal acceptance rate of $e^{-2}$ (=0.1353) under quite general conditions. The results have major practical implications for the implementation of RWM algorithms by highlighting the detrimental effect of choosing RWM algorithms over Metropolis-within-Gibbs algorithms.