Source author record

Ruth E. Baker

Ruth E. Baker appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

12works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

12 published item(s)

preprint2026arXiv

Structural identifiability of partially-observed stochastic processes: from single-particle trajectories to total particle density data

The increasing availability of experimental data has intensified interest in calibrating stochastic models, raising fundamental questions about parameter identifiability. Structural identifiability determines whether parameters can be uniquely recovered from idealised, noise-free data, a prerequisite to allow for parameter estimation. However, existing methods to assess structural identifiability are not generally applicable to stochastic processes. We develop a methodology to analyse structural identifiability for a class of spatio-temporal stochastic processes. We investigate how identifiability depends on the type of available data, distinguishing between single-particle trajectories and total particle density measurements. For trajectory data, we use the individual-based model description that explicitly represents single-particle dynamics. For population-level data, we derive a partial differential equation model representation, that describes the evolution of total particle density, and apply a differential algebra approach, common to ordinary differential equations analysis. We further introduce a novel method to study the initial condition, based on characteristic equations to construct a Taylor expansion of the density evolution, enabling identification of additional identifiable parameter combinations. We apply our methodology to a model, and show it is identifiable with trajectory data but only locally identifiable with density data, and demonstrate the critical role of initial conditions in the identifiability analysis.

preprint2026arXiv

The spontaneous emergence of leaders and followers in a mathematical model of cranial neural crest cell migration

Many agent-based mathematical models of cranial neural crest cell (CNCC) migration impose a binary phenotypic partition of cells into either leaders or followers. In such models, the movement of leader cells at the front of collectives is guided by local chemoattractant gradients, while follower cells behind leaders move according to local cell-cell guidance cues. Although such model formulations have yielded many insights into the mechanisms underpinning CNCC migration, they rely on fixed phenotypic traits that are difficult to reconcile with evidence of phenotypic plasticity in vivo. A later agent-based model of CNCC migration aimed to address this limitation by allowing cells to adaptively combine chemotactic and cell-cell guidance cues during migration. In this model, cell behaviour adapts instantaneously in response to environmental cues, which precludes the identification of a persistent subset of cells as leader-like over biologically relevant timescales, as observed in vivo. Here, we build on previous leader-follower and adaptive phenotype models to develop a polarity-based agent-based model of CNCC migration, in which all cells evolve according to identical rules, interact via a pairwise interaction potential, and carry polarity vectors that evolve according to a dynamical system driven by time-averaged exposure to chemoattractant gradients. Numerical simulations of this model show that a leader-follower phenotypic partition emerges spontaneously from the underlying collective dynamics of the model. Furthermore, the model reproduces behaviour that is consistent with experimental observations of CNCC migration in the chick embryo. Thus, we provide an experimentally consistent, mechanistically-grounded mathematical model that captures the emergence of leader and follower cell phenotypes without their imposition a priori.

preprint2022arXiv

Control of diffusion-driven pattern formation behind a wave of competency

In certain biological contexts, such as the plumage patterns of birds and stripes on certain species of fishes, pattern formation takes place behind a so-called "wave of competency". Currently, the effects of a wave of competency on the patterning outcome is not well-understood. In this study, we use Turing's diffusion-driven instability model to study pattern formation behind a wave of competency, under a range of wave speeds. Numerical simulations show that in one spatial dimension a slower wave speed drives a sequence of peak splittings in the pattern, whereas a higher wave speed leads to peak insertions. In two spatial dimensions, we observe stripes that are either perpendicular or parallel to the moving boundary under slow or fast wave speeds, respectively. We argue that there is a correspondence between the one- and two-dimensional phenomena, and that pattern formation behind a wave of competency can account for the pattern organization observed in many biological systems.

preprint2022arXiv

Multifidelity multilevel Monte Carlo to accelerate approximate Bayesian parameter inference for partially observed stochastic processes

Models of stochastic processes are widely used in almost all fields of science. Theory validation, parameter estimation, and prediction all require model calibration and statistical inference using data. However, data are almost always incomplete observations of reality. This leads to a great challenge for statistical inference because the likelihood function will be intractable for almost all partially observed stochastic processes. This renders many statistical methods, especially within a Bayesian framework, impossible to implement. Therefore, computationally expensive likelihood-free approaches are applied that replace likelihood evaluations with realisations of the model and observation process. For accurate inference, however, likelihood-free techniques may require millions of expensive stochastic simulations. To address this challenge, we develop a new method based on recent advances in multilevel and multifidelity. Our approach combines the multilevel Monte Carlo telescoping summation, applied to a sequence of approximate Bayesian posterior targets, with a multifidelity rejection sampler to minimise the number of computationally expensive exact simulations required for accurate inference. We present the derivation of our new algorithm for likelihood-free Bayesian inference, discuss practical implementation details, and demonstrate substantial performance improvements. Using examples from systems biology, we demonstrate improvements of more than two orders of magnitude over standard rejection sampling techniques. Our approach is generally applicable to accelerate other sampling schemes, such as sequential Monte Carlo, to enable feasible Bayesian analysis for realistic practical applications in physics, chemistry, biology, epidemiology, ecology and economics.

preprint2022arXiv

Symmetries of systems of first order ODEs: Symbolic symmetry computations, mechanistic model construction and applications in biology

We discuss the role and merits of symmetry methods for the analysis of biological systems. In particular, we consider systems of first order ordinary differential equations and provide a comprehensive review of the geometrical foundations pertinent to symmetries of such systems. Subsequently, we present an algorithm for finding infinitesimal generators of symmetries for systems with rational reaction terms, and an open-source implementation of the algorithm using symbolic computations. We discuss two complementary perspectives on symmetries in mechanistic modelling; as tools for the analysis of a given model or as a geometrical principle for incorporating biological properties in the construction of new models. Through numerous examples of relevance to modelling in biology we demonstrate the different uses of symmetry methods, and also discuss how to infer symmetries from experimental data.

preprint2020arXiv

Biologically-informed neural networks guide mechanistic modeling from sparse experimental data

Biologically-informed neural networks (BINNs), an extension of physics-informed neural networks [1], are introduced and used to discover the underlying dynamics of biological systems from sparse experimental data. In the present work, BINNs are trained in a supervised learning framework to approximate in vitro cell biology assay experiments while respecting a generalized form of the governing reaction-diffusion partial differential equation (PDE). By allowing the diffusion and reaction terms to be multilayer perceptrons (MLPs), the nonlinear forms of these terms can be learned while simultaneously converging to the solution of the governing PDE. Further, the trained MLPs are used to guide the selection of biologically interpretable mechanistic forms of the PDE terms which provides new insights into the biological and physical mechanisms that govern the dynamics of the observed system. The method is evaluated on sparse real-world data from wound healing assays with varying initial cell densities [2].

preprint2020arXiv

Efficiently simulating discrete-state models with binary decision trees

Stochastic simulation algorithms (SSAs) are widely used to numerically investigate the properties of stochastic, discrete-state models. The Gillespie Direct Method is the pre-eminent SSA, and is widely used to generate sample paths of so-called agent-based or individual-based models. However, the simplicity of the Gillespie Direct Method often renders it impractical where large-scale models are to be analysed in detail. In this work, we carefully modify the Gillespie Direct Method so that it uses a customised binary decision tree to trace out sample paths of the model of interest. We show that a decision tree can be constructed to exploit the specific features of the chosen model. Specifically, the events that underpin the model are placed in carefully-chosen leaves of the decision tree in order to minimise the work required to keep the tree up-to-date. The computational efficencies that we realise can provide the apparatus necessary for the investigation of large-scale, discrete-state models that would otherwise be intractable. Two case studies are presented to demonstrate the efficiency of the method.

preprint2016arXiv

Coupling volume-excluding compartment-based models of diffusion at different scales: Voronoi and pseudo-compartment approaches

Numerous processes across both the physical and biological sciences are driven by diffusion. Partial differential equations (PDEs) are a popular tool for modelling such phenomena deterministically, but it is often necessary to use stochastic models to accurately capture the behaviour of a system, especially when the number of diffusing particles is low. The stochastic models we consider in this paper are `compartment-based': the domain is discretized into compartments, and particles can jump between these compartments. Volume-excluding effects (crowding) can be incorporated by blocking movement with some probability. Recent work has established the connection between fine-grained models and coarse-grained models incorporating volume exclusion, but only for uniform lattices. In this paper we consider non-uniform, hybrid lattices that incorporate both fine- and coarse-grained regions, and present two different approaches to describing the interface of the regions. We test both techniques in a range of scenarios to establish their accuracy, benchmarking against fine-grained models, and show that the hybrid models developed in this paper can be significantly faster to simulate than the fine-grained models in certain situations, and are at least as fast otherwise.

preprint2016arXiv

Extending the multi-level method for the simulation of stochastic biological systems

The multi-level method for discrete state systems, first introduced by Anderson and Higham [Multiscale Model. Simul. 10:146--179, 2012], is a highly efficient simulation technique that can be used to elucidate statistical characteristics of biochemical reaction networks. A single point estimator is produced in a cost-effective manner by combining a number of estimators of differing accuracy in a telescoping sum, and, as such, the method has the potential to revolutionise the field of stochastic simulation. The first term in the sum is calculated using an approximate simulation algorithm, and can be calculated quickly but is of significant bias. Subsequent terms successively correct this bias by combining estimators from approximate stochastic simulations algorithms of increasing accuracy, until a desired level of accuracy is reached. In this paper we present several refinements of the multi-level method which render it easier to understand and implement, and also more efficient. Given the substantial and complex nature of the multi-level method, the first part of this work (Sections 2 - 5) is written as a tutorial, with the aim of providing a practical guide to its use. The second part (Sections 6 - 8) takes on a form akin to a research article, thereby providing the means for a deft implementation of the technique, and concludes with a discussion of a number of open problems.

preprint2016arXiv

Multi-level methods and approximating distribution functions

Biochemical reaction networks are often modelled using discrete-state, continuous-time Markov chains. System statistics of these Markov chains usually cannot be calculated analytically and therefore estimates must be generated via simulation techniques. There is a well documented class of simulation techniques known as exact stochastic simulation algorithms, an example of which is Gillespie's direct method. These algorithms often come with high computational costs, therefore approximate stochastic simulation algorithms such as the tau-leap method are used. However, in order to minimise the bias in the estimates generated using them, a relatively small value of tau is needed, rendering the computational costs comparable to Gillespie's direct method. The multi-level Monte Carlo method (Anderson and Higham, Multiscale Model. Simul. 10:146-179, 2012) provides a reduction in computational costs whilst minimising or even eliminating the bias in the estimates of system statistics. This is achieved by first crudely approximating required statistics with many sample paths of low accuracy. Then correction terms are added until a required level of accuracy is reached. Recent literature has primarily focussed on implementing the multi-level method efficiently to estimate a single system statistic. However, it is clearly also of interest to be able to approximate entire probability distributions of species counts. We present two novel methods that combine known techniques for distribution reconstruction with the multi-level method. We demonstrate the potential of our methods using a number of examples.

preprint2015arXiv

Reconciling transport models across scales: the role of volume exclusion

Diffusive transport is a universal phenomenon, throughout both biological and physical sciences, and models of diffusion are routinely used to interrogate diffusion-driven processes. However, most models neglect to take into account the role of volume exclusion, which can significantly alter diffusive transport, particularly within biological systems where the diffusing particles might occupy a significant fraction of the available space. In this work we use a random walk approach to provide a means to reconcile models that incorporate crowding effects on different spatial scales. Our work demonstrates that coarse-grained models incorporating simplified descriptions of excluded volume can be used in many circumstances, but that care must be taken in pushing the coarse-graining process too far.

preprint2014arXiv

An adaptive multi-level simulation algorithm for stochastic biological systems

Discrete-state, continuous-time Markov models are widely used in the modeling of biochemical reaction networks. Their complexity often precludes analytic solution, and we rely on stochastic simulation algorithms to estimate system statistics. The Gillespie algorithm is exact, but computationally costly as it simulates every single reaction. As such, approximate stochastic simulation algorithms such as the tau-leap algorithm are often used. Potentially computationally more efficient, the system statistics generated suffer from significant bias unless tau is relatively small, in which case the computational time can be comparable to that of the Gillespie algorithm. The multi-level method (Anderson and Higham, Multiscale Model. Simul. 2012) tackles this problem. A base estimator is computed using many (cheap) sample paths at low accuracy. The bias inherent in this estimator is then reduced using a number of corrections. Each correction term is estimated using a collection of paired sample paths where one path of each pair is generated at a higher accuracy compared to the other (and so more expensive). By sharing random variables between these paired paths the variance of each correction estimator can be reduced. This renders the multi-level method very efficient as only a relatively small number of paired paths are required to calculate each correction term. In the original multi-level method, each sample path is simulated using the tau-leap algorithm with a fixed value of $τ$. This approach can result in poor performance when the reaction activity of a system changes substantially over the timescale of interest. By introducing a novel, adaptive time-stepping approach where $τ$ is chosen according to the stochastic behaviour of each sample path we extend the applicability of the multi-level method to such cases. We demonstrate the efficiency of our method using a number of examples.