Source author record

Jere Koskela

Jere Koskela appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2023arXiv

EWF : simulating exact paths of the Wright--Fisher diffusion

The Wright--Fisher diffusion is important in population genetics in modelling the evolution of allele frequencies over time subject to the influence of biological phenomena such as selection, mutation, and genetic drift. Simulating paths of the process is challenging due to the form of the transition density. We present EWF, a robust and efficient sampler which returns exact draws for the diffusion and diffusion bridge processes, accounting for general models of selection including those with frequency-dependence. Given a configuration of selection, mutation, and endpoints, EWF returns draws at the requested sampling times from the law of the corresponding Wright--Fisher process. Output was validated by comparison to approximations of the transition density via the Kolmogorov--Smirnov test and QQ plots. All software is available at https://github.com/JaroSant/EWF

preprint2020arXiv

Statistical tools for seed bank detection

In this article, we derive statistical tools to analyze and distinguish the patterns of genetic variability produced by classical and recent population genetic models related to seed banks. In particular, we are concerned with models described by the Kingman coalescent (K), models exhibiting so-called weak seed banks described by a time-changed Kingman coalescent (W), models with so-called strong seed bank described by the seed bank coalescent (S) and the classical two-island model by Wright, described by the structured coalescent (TI). As the presence of a (strong) seed bank should stratify a population, we expect it to produce a signal roughly comparable to the presence of population structure. We begin with a brief analysis of Wright's $F_{ST}$, which is a classical but crude measure for population structure, followed by a derivation of the expected site frequency spectrum (SFS) in the infinite sites model based on 'phase-type distribution calculus' as recently discussed by Hobolth et al. (2019). Both the $F_{ST}$ and the SFS can be readily computed under various population models, they discard statistical signal. Hence we also derive exact likelihoods for the full sampling probabilities, which can be achieved via recursions and a Monte Carlo scheme both in the infinite alleles and the infinite sites model. We employ a pseudo-marginal Metropolis-Hastings algorithm of Andrieu and Roberts (2009) to provide a method for simultaneous model selection and parameter inference under the so-called infinitely-many sites model, which is the most relevant in real applications. It turns out that this full likelihood method can reliably distinguish among the model classes (K, W), (S) and (TI) on the basis of simulated data even from moderate sample sizes. It is also possible to infer mutation rates, and in particular determine whether mutation is taking place in the (strong) seed bank.

preprint2015arXiv

Computational inference beyond Kingman's coalescent

Full likelihood inference under Kingman's coalescent is a computationally challenging problem to which importance sampling (IS) and the product of approximate conditionals (PAC) method have been applied successfully. Both methods can be expressed in terms of families of intractable conditional sampling distributions (CSDs), and rely on principled approximations for accurate inference. Recently, more general $Λ$- and $Ξ$-coalescents have been observed to provide better modelling fits to some genetic data sets. We derive families of approximate CSDs for finite sites $Λ$- and $Ξ$-coalescents, and use them to obtain "approximately optimal" IS and PAC algorithms for $Λ$-coalescents, yielding substantial gains in efficiency over existing methods.