Source author record

Peter Pfaffelhuber

Peter Pfaffelhuber appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

20works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

20 published item(s)

preprint2022arXiv

A central limit theorem concerning uncertainty in estimates of individual admixture

The concept of individual admixture (IA) assumes that the genome of individuals is composed of alleles inherited from $K$ ancestral populations. Each copy of each allele has the same chance $q_k$ to originate from population $k$, and together with the allele frequencies $p$ in all populations at all $M$ markers, comprises the admixture model. Here, we assume a supervised scheme, i.e.\ allele frequencies $p$ are given through a reference database of size $N$, and $q$ is estimated via maximum likelihood for a single sample. We study laws of large numbers and central limit theorems describing effects of finiteness of both, $M$ and $N$, on the estimate of $q$. We recall results for the effect of finite $M$, and provide a central limit theorem for the effect of finite $N$, introduce a new way to express the uncertainty in estimates in standard barplots, give simulation results, and discuss applications in forensic genetics.

preprint2016arXiv

The fixation time of a strongly beneficial allele in a structured population

For a beneficial allele which enters a large unstructured population and eventually goes to fixation, it is known that the time to fixation is approximately $2\log(α)/α$ for a large selection coefficient $α$. For a population that is distributed over finitely many colonies, with migration between these colonies, we detect various regimes of the migration rate $μ$ for which the fixation times have different asymptotics as $α\to \infty$. If $μ$ is of order $α$, the allele fixes (as in the spatially unstructured case) in time $\sim 2\log(α)/α$. If $μ$ is of order $α^γ, 0\leq γ\leq 1$, the fixation time is $\sim (2 + (1-γ)Δ) \log(α)/α$, where $Δ$ is the number of migration steps that are needed to reach all other colonies starting from the colony where the beneficial allele appeared. If $μ= 1/\log(α)$, the fixation time is $\sim (2+S)\log(α)/α$, where $S$ is a random time in a simple epidemic model. The main idea for our analysis is to combine a new moment dual for the process conditioned to fixation with the time reversal in equilibrium of a spatial version of Neuhauser and Krone's ancestral selection graph.

preprint2015arXiv

A mixing tree-valued process arising under neutral evolution with recombination

The genealogy at a single locus of a constant size $N$ population in equilibrium is given by the well-known Kingman's coalescent. When considering multiple loci under recombination, the ancestral recombination graph encodes the genealogies at all loci in one graph. For a continuous genome $G$, we study the tree-valued process $(T^N_u)_{u\in G}$ of genealogies along the genome in the limit $N\to\infty$. Encoding trees as metric measure spaces, we show convergence to a tree-valued process with cadlag paths. In addition, we study mixing properties of the resulting process for loci which are far apart.

preprint2015arXiv

Scaling limits of spatial compartment models for chemical reaction networks

We study the effects of fast spatial movement of molecules on the dynamics of chemical species in a spatially heterogeneous chemical reaction network using a compartment model. The reaction networks we consider are either single- or multi-scale. When reaction dynamics is on a single-scale, fast spatial movement has a simple effect of averaging reactions over the distribution of all the species. When reaction dynamics is on multiple scales, we show that spatial movement of molecules has different effects depending on whether the movement of each type of species is faster or slower than the effective reaction dynamics on this molecular type. We obtain results for both when the system is without and with conserved quantities, which are linear combinations of species evolving only on the slower time scale.

preprint2015arXiv

The stationary distribution of a Markov jump process glued together from two state spaces at two vertices

We compute the stationary distribution of a continuous-time Markov chain which is constructed by gluing together two finite, irreducible Markov chains by identifying a pair of states of one chain with a pair of states of the other and keeping all transition rates from either chain (the rates between the two shared states are summed). The result expresses the stationary distribution of the glued chain in terms of quantities of the two original chains. Some of the required terms are nonstandard but can be computed by solving systems of linear equations using the transition rate matrices of the two original chains. Special emphasis is given to the cases when the stationary distribution of the glued chain is a multiple of the equilibria of the original chains, and when not, for which bounds are derived.

preprint2014arXiv

Some large deviations in Kingman's coalescent

Kingman's coalescent is a random tree that arises from classical population genetic models such as the Moran model. The individuals alive in these models correspond to the leaves in the tree and the following two laws of large numbers concerning the structure of the tree-top are well-known: (i) The (shortest) distance, denoted by $T_n$, from the tree-top to the level when there are $n$ lines in the tree satisfies $nT_n \xrightarrow{n\to\infty} 2$ almost surely; (ii) At time $T_n$, the population is naturally partitioned in exactly $n$ families where individuals belong to the same family if they have a common ancestor at time $T_n$ in the past. If $F_{i,n}$ denotes the size of the $i$th family, then $n(F_{1,n}^2 + \cdots + F_{n,n}^2) \xrightarrow{n\to \infty}2$ almost surely. For both laws of large numbers we prove corresponding large deviations results. For (i), the rate of the large deviations is $n$ and we can give the rate function explicitly. For (ii), the rate is $n$ for downwards deviations and $\sqrt n$ for upwards deviations. For both cases we give the exact rate function.

preprint2014arXiv

Some limit results for Markov chains indexed by trees

We consider a sequence of Markov chains $(\mathcal X^n)_{n=1,2,...}$ with $\mathcal X^n = (X^n_σ)_{σ\in\mathcal T}$, indexed by the full binary tree $\mathcal T = \mathcal T_0 \cup \mathcal T_1 \cup ...$, where $\mathcal T_k$ is the $k$th generation of $\mathcal T$. In addition, let $(Σ_k)_{k=0,1,2,...}$ be a random walk on $\mathcal T$ with $Σ_k \in \mathcal T_k$ and $\widetilde{\mathcal R}^n = (\widetilde R_t^n)_{t\geq 0}$ with $\widetilde R_t^n := X_{Σ_{[tn]}}$, arising by observing the Markov chain $\mathcal X^n$ along the random walk. We present a law of large numbers concerning the empirical measure process $\widetilde{\mathcal Z}^n = (\widetilde Z_t^n)_{t\geq 0}$ where $\widetilde{Z}_t^n = \sum_{σ\in\mathcal T_{[tn]}} δ_{X_σ^n}$ as $n\to\infty$. Precisely, we show that if $\widetilde{\mathcal R}^n \to \mathcal R$ for some Feller process $\mathcal R = (R_t)_{t\geq 0}$ with deterministic initial condition, then $\widetilde{\mathcal Z}^n \to \mathcal Z$ with $Z_t = δ_{\mathcal L(R_t)}$.

preprint2014arXiv

Stochastic gene expression with delay

The expression of genes usually follows a two-step procedure. First, a gene (encoded in the genome) is transcribed resulting in a strand of (messenger) RNA. Afterwards, the RNA is translated into protein. Classically, this gene expression is modeled using a Markov jump process including activation and deactivation of the gene, transcription and translation rates together with degradation of RNA and protein. We extend this model by adding delays (with arbitrary distributions) to transcription and translation. Such delays can e.g.\ mean that RNA has to be transported to a different part of a cell before translation can be initiated. Already in the classical model, production of RNA and protein come in bursts by activation and deactivation of the gene, resulting in a large variance of the number of RNA and proteins in equilibrium. We derive precise formulas for this second-order structure with the model including delay in equilibrium. As a general fact, the delay decreases the variance of the number of RNA and proteins.

preprint2014arXiv

The infinitely many genes model with horizontal gene transfer

The genome of bacterial species is much more flexible than that of eukaryotes. Moreover, the distributed genome hypothesis for bacteria states that the total number of genes present in a bacterial population is greater than the genome of every single individual. The pangenome, i.e. the set of all genes of a bacterial species (or a sample), comprises the core genes which are present in all living individuals, and accessory genes, which are carried only by some individuals. In order to use accessory genes for adaptation to environmental forces, genes can be transferred horizontally between individuals. Here, we extend the infinitely many genes model from Baumdicker, Hess and Pfaffelhuber (2010) for horizontal gene transfer. We take a genealogical view and give a construction -- called the Ancestral Gene Transfer Graph -- of the joint genealogy of all genes in the pangenome. As application, we compute moments of several statistics (e.g. the number of differences between two individuals and the gene frequency spectrum) under the infinitely many genes model with horizontal gene transfer.

preprint2013arXiv

Path-properties of the tree-valued Fleming-Viot process

We consider the tree-valued Fleming-Viot process, $(\mathcal X_t)_{t\geq 0}$, with mutation and selection as studied in Depperschmidt, Greven, Pfaffelhuber (2012). This process models the stochastic evolution of the genealogies and (allelic) types under resampling, mutation and selection in the population currently alive in the limit of infinitely large populations. Genealogies and types are described by (isometry classes of) marked metric measure spaces. The long-time limit of the neutral tree-valued Fleming-Viot dynamics is an equilibrium given via the marked metric measure space associated with the Kingman coalescent. In the present paper we pursue two closely linked goals. First, we show that two well-known properties of the neutral Fleming-Viot genealogies at fixed time $t$ arising from the properties of the dual, namely the Kingman coalescent, hold for the whole path. These properties are related to the geometry of the family tree close to its leaves. In particular we consider the number and the size of subfamilies whose individuals are not further than $\ve$ apart in the limit $\ve\to 0$. Second, we answer two open questions about the sample paths of the tree-valued Fleming-Viot process. We show that for all $t>0$ almost surely the marked metric measure space $\mathcal X_t$ has no atoms and admits a mark function. The latter property means that all individuals in the tree-valued Fleming-Viot process can uniquely be assigned a type. All main results are proven for the neutral case and then carried over to selective cases via Girsanov's formula giving absolute continuity.

preprint2012arXiv

Finite populations with frequency-dependent selection: a genealogical approach

Evolutionary models for populations of constant size are frequently studied using the Moran model, the Wright-Fisher model, or their diffusion limits. When evolution is neutral, a random genealogy given through Kingman's coalescent is used in order to understand basic properties of such models. Here, we address the use of a genealogical perspective for models with weak frequency-dependent selection, i.e. N s =: α is small, and s is the fitness advantage of a fit individual and N is the population size. When computing fixation probabilities, this leads either to the approach proposed by Rousset (2003), who argues how to use the Kingman's coalescent for weak selection, or to extensions of the ancestral selection graph of Neuhauser and Krone (1997) and Neuhauser (1999). As an application, we re-derive the one-third law of evolutionary game theory (Nowak et al., 2004). In addition, we provide the approximate distribution of the genealogical distance of two randomly sampled individuals under linear frequency-dependence.

preprint2012arXiv

Tree-valued Fleming-Viot dynamics with mutation and selection

The Fleming-Viot measure-valued diffusion is a Markov process describing the evolution of (allelic) types under mutation, selection and random reproduction. We enrich this process by genealogical relations of individuals so that the random type distribution as well as the genealogical distances in the population evolve stochastically. The state space of this tree-valued enrichment of the Fleming-Viot dynamics with mutation and selection (TFVMS) consists of marked ultrametric measure spaces, equipped with the marked Gromov-weak topology and a suitable notion of polynomials as a separating algebra of test functions. The construction and study of the TFVMS is based on a well-posed martingale problem. For existence, we use approximating finite population models, the tree-valued Moran models, while uniqueness follows from duality to a function-valued process. Path properties of the resulting process carry over from the neutral case due to absolute continuity, given by a new Girsanov-type theorem on marked metric measure spaces. To study the long-time behavior of the process, we use a duality based on ideas from Dawson and Greven [On the effects of migration in spatial Fleming-Viot models with selection and mutation (2011c) Unpublished manuscript] and prove ergodicity of the TFVMS if the Fleming-Viot measure-valued diffusion is ergodic. As a further application, we consider the case of two allelic types and additive selection. For small selection strength, we give an expansion of the Laplace transform of genealogical distances in equilibrium, which is a first step in showing that distances are shorter in the selective case.

preprint2011arXiv

Compact metric measure spaces and Lambda-coalescents coming down from infinity

We study topological properties of random metric spaces which arise by Lambda-coalescents. These are stochastic processes, which start with an infinite number of lines and evolve through multiple mergers in an exchangeable setting. We show that the resulting Lambda-coalescent measure tree is compact iff the Lambda-coalescent comes down from infinity, i.e. only consists of finitely many lines at any positive time. If the Lambda-coalescent stays infinite, the resulting metric measure space is not even locally compact. Our results are based on general notions of compact and locally compact (isometry classes of) metric measure spaces. In particular, we give characterizations for general (random) metric measure spaces to be (locally) compact using the Gromov-weak topology.

preprint2011arXiv

Evolution of bacterial genomes under horizontal gene transfer

Unraveling the evolutionary forces shaping bacterial diversity can today be tackled using a growing amount of genomic data. While the genome of eukaryotes is highly stable, bacterial genomes from cells of the same species highly vary in gene content. This huge variation in gene content led to the concepts of the distributed genome of bacteria and their pangenome (Tettelin et al.,2005; Ehrlich et al.,2005). We present a population genetic model for gene content evolution which accounts for several mechanisms. Gene uptake from the environment is modeled by events of gene gain along the genealogical tree relating the population. Pseudogenization may lead to deletion of genes and is incoporated by gene loss. These two mechanisms were studied by Huson and Steel (2004) using a fixed phylogenetic tree. Taking the random genealogy given by the coalescent (Kingman, 1982; Hudson, 1983), we studied the resulting genomic diversity already in Baumdicker et al. (2010). In the present paper, we extend the model in order to incorporate events of interspecies horizontal gene transfer. Within this model, we derive expectations for the gene frequency spectrum and other quantities of interest.

preprint2011arXiv

Marked metric measure spaces

A marked metric measure space (mmm-space) is a triple (X,r,mu), where (X,r) is a complete and separable metric space and mu is a probability measure on XxI for some Polish space I of possible marks. We study the space of all (equivalence classes of) marked metric measure spaces for some fixed I. It arises as state space in the construction of Markov processes which take values in random graphs, e.g. tree-valued dynamics describing randomly evolving genealogical structures in population models. We derive here the topological properties of the space of mmm-spaces needed to study convergence in distribution of random mmm-spaces. Extending the notion of the Gromov-weak topology introduced in (Greven, Pfaffelhuber and Winter, 2009), we define the marked Gromov-weak topology, which turns the set of mmm-spaces into a Polish space. We give a characterization of tightness for families of distributions of random mmm- spaces and identify a convergence determining algebra of functions, called polynomials.

preprint2011arXiv

Sensitivity analysis of one parameter semigroups exemplified by the Wright--Fisher diffusion

We consider the sensitivity, with respect to a parameter θ, of parametric families of operators A_θ, vectors π_θ corresponding to the adjoints A_θ^{*} of A_θ via A_θ^{*}π_θ=0 and one parameter semigroups t\mapsto e^{tA_θ}. We display formulas relating weak differentiability of θ\mapsto π_θ (at θ=0) to weak differentiability of θ\mapsto A_θ^{*}π_{0} and [e^{A_θt}]^{*}π_{0}. We give two applications: The first one concerns the sensitivity of the Ornstein--Uhlenbeck process with respect to its location parameter. The second one provides new insights regarding the Wright--Fisher diffusion for small mutation parameter.

preprint2011arXiv

Tree-valued resampling dynamics: Martingale Problems and applications

The measure-valued Fleming-Viot process is a diffusion which models the evolution of allele frequencies in a multi-type population. In the neutral setting the Kingman coalescent is known to generate the genealogies of the "individuals" in the population at a fixed time. The goal of the present paper is to replace this static point of view on the genealogies by an analysis of the evolution of genealogies. We encode the genealogy of the population as an (isometry class of an) ultra-metric space which is equipped with a probability measure. The space of ultra-metric measure spaces together with the Gromov-weak topology serves as state space for tree-valued processes. We use well-posed martingale problems to construct the tree-valued resampling dynamics of the evolving genealogies for both the finite population Moran model and the infinite population Fleming-Viot diffusion. We show that sufficient information about any ultra-metric measure space is contained in the distribution of the vector of subtree lengths obtained by sequentially sampled "individuals". We give explicit formulas for the evolution of the Laplace transform of the distribution of finite subtrees under the tree-valued Fleming-Viot dynamics.

preprint2010arXiv

Selective sweeps for recessive alleles and for other modes of dominance

A selective sweep describes the reduction of linked genetic variation due to strong positive selection. If s is the fitness advantage of a homozygote for the beneficial allele and h its dominance coefficient, it is usually assumed that h=1/2, i.e. the beneficial allele is co-dominant. We complement existing theory for selective sweeps by assuming that h is any value in [0,1]. We show that genetic diversity patters under selective sweeps with strength s and dominance 0<h<1 are similar to co-dominant sweeps with selection strength 2hs. Moreover, we focus on the case h=0 of a completely recessive beneficial allele. We find that the length of the sweep, i.e. the time from occurrence until fixation of the beneficial allele, is of the order of sqrt(N/s) generations, if N is the population size. Simulations as well as our results show that genetic diversity patterns in the recessive case h=0 greatly differ from all other cases.

preprint2010arXiv

The Aldous-Shields model revisited (with application to cellular ageing)

In Aldous and Shields (1988), a model for a rooted, growing random binary tree was presented. For some c>0, an external vertex splits at rate c^(-i) (and becomes internal) if its distance from the root (depth) is i. For c>1, we reanalyse the tree profile, i.e. the numbers of external vertices in depth i=1,2,.... Our main result are concrete formulas for the expectation and covariance-structure of the profile. In addition, we present the application of the model to cellular ageing. Here, we assume that nodes in depth h+1 are senescent, i.e. do not split. We obtain a limit result for the proportion of non-senescent vertices for large h.

preprint2010arXiv

The tree length of an evolving coalescent

A well-established model for the genealogy of a large population in equilibrium is Kingman's coalescent. For the population together with its genealogy evolving in time, this gives rise to a time-stationary tree-valued process. We study the sum of the branch lengths, briefly denoted as tree length, and prove that the (suitably compensated) sequence of tree length processes converges, as the population size tends to infinity, to a limit process with cadlag paths, infinite infinitesimal variance, and a Gumbel distribution as its equilibrium.