Catalog footprint

What is connected

26works
26topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

26 published item(s)

preprint2021arXiv

Modularity maximisation for graphons

Networks are a widely-used tool to investigate the large-scale connectivity structure in complex systems and graphons have been proposed as an infinite size limit of dense networks. The detection of communities or other meso-scale structures is a prominent topic in network science as it allows the identification of functional building blocks in complex systems. When such building blocks may be present in graphons is an open question. In this paper, we define a graphon-modularity and demonstrate that it can be maximised to detect communities in graphons. We then investigate specific synthetic graphons and show that they may show a wide range of different community structures. We also reformulate the graphon-modularity maximisation as a continuous optimisation problem and so prove the optimal community structure or lack thereof for some graphons, something that is usually not possible for networks. Furthermore, we demonstrate that estimating a graphon from network data as an intermediate step can improve the detection of communities, in comparison with exclusively maximising the modularity of the network. While the choice of graphon-estimator may strongly influence the accord between the community structure of a network and its estimated graphon, we find that there is a substantial overlap if an appropriate estimator is used. Our study demonstrates that community detection for graphons is possible and may serve as a privacy-preserving way to cluster network data.

preprint2021arXiv

What would it take to build a thermodynamically reversible Universal Turing machine? Computational and thermodynamic constraints in a molecular design

We outline the construction of a molecular system that could, in principle, implement a thermodynamically reversible Universal Turing Machine (UTM). By proposing a concrete-albeit idealised-design and operational protocol, we reveal fundamental challenges that arise when attempting to implement arbitrary computations reversibly. Firstly, the requirements of thermodynamic reversibility inevitably lead to an intricate design. Secondly, thermodynamically reversible UTMs, unlike simpler devices, must also be logically reversible. Finally, implementing multiple distinct computations in parallel is necessary to take the cost of external control per computation to zero, but this approach is complicated the distinct halting times of different computations.

preprint2020arXiv

Community detection in networks without observing edges

We develop a Bayesian hierarchical model to identify communities in networks for which we do not observe the edges directly, but instead observe a series of interdependent signals for each of the nodes. Fitting the model provides an end-to-end community detection algorithm that does not extract information as a sequence of point estimates but propagates uncertainties from the raw data to the community labels. Our approach naturally supports multiscale community detection as well as the selection of an optimal scale using model comparison. We study the properties of the algorithm using synthetic data and apply it to daily returns of constituents of the S&P100 index as well as climate data from US cities.

preprint2020arXiv

Democratizing University Research

We detail an experimental programme we have been testing in our university. Our Advanced Hackspace, attempts to give all members of the university, from students to technicians, free access to the means to develop their own interdisciplinary research ideas, with resources including access to specialized fellows and biological and chemical hacklabs. We assess the aspects of our programme that led to our community being one of the largest collectives in our university and critically examine the successes and failures of our trial programmes. We supply metrics for assessing progress and outline challenges. We conclude with future directions that advance interdisciplinary research empowerment for all university members.

preprint2020arXiv

Inference and Influence of Large-Scale Social Networks Using Snapshot Population Behaviour without Network Data

Population behaviours, such as voting and vaccination, depend on social networks. Social networks can differ depending on behaviour type and are typically hidden. However, we do often have large-scale behavioural data, albeit only snapshots taken at one timepoint. We present a method that jointly infers large-scale network structure and a networked model of human behaviour using only snapshot population behavioural data. This exploits the simplicity of a few parameter, geometric socio-demographic network model and a spin based model of behaviour. We illustrate, for the EU Referendum and two London Mayoral elections, how the model offers both prediction and the interpretation of our homophilic inclinations. Beyond offering the extraction of behaviour specific network structure from large-scale behavioural datasets, our approach yields a crude calculus linking inequalities and social preferences to behavioural outcomes. We give examples of potential network sensitive policies: how changes to income inequality, a social temperature and homophilic preferences might have reduced polarisation in a recent election.

preprint2020arXiv

Survival of the densest accounts for the expansion of mitochondrial mutations in ageing

The expansion of deleted mitochondrial DNA (mtDNA) molecules has been linked to ageing, particularly in skeletal muscle fibres; its mechanism has remained unclear for three decades. Previous accounts assigned a replicative advantage to the deletions, but there is evidence that cells can, instead, selectively remove defective mtDNA. We present a spatial model that, without a replicative advantage, but instead through a combination of enhanced density for mutants and noise, produces a wave of expanding mutations with wave speed consistent with experimental data, unlike a standard model based on replicative advantage. We provide a formula that predicts that the wave speed drops with copy number, in agreement with experimental data. Crucially, our model yields travelling waves of mutants even if mutants are preferentially eliminated. Justified by this exemplar of how noise, density and spatial structure affect muscle ageing, we introduce the mechanism of stochastic survival of the densest, an alternative to replicative advantage, that may underpin other phenomena, like the evolution of altruism.

preprint2017arXiv

Chemical Boltzmann Machines

How smart can a micron-sized bag of chemicals be? How can an artificial or real cell make inferences about its environment? From which kinds of probability distributions can chemical reaction networks sample? We begin tackling these questions by showing four ways in which a stochastic chemical reaction network can implement a Boltzmann machine, a stochastic neural network model that can generate a wide range of probability distributions and compute conditional probabilities. The resulting models, and the associated theorems, provide a road map for constructing chemical reaction networks that exploit their native stochasticity as a computational resource. Finally, to show the potential of our models, we simulate a chemical Boltzmann machine to classify and generate MNIST digits in-silico.

preprint2015arXiv

Closed-form stochastic solutions for non-equilibrium dynamics and inheritance of cellular components over many cell divisions

Stochastic dynamics govern many important processes in cellular biology, and an underlying theoretical approach describing these dynamics is desirable to address a wealth of questions in biology and medicine. Mathematical tools exist for treating several important examples of these stochastic processes, most notably gene expression, and random partitioning at single cell divisions or after a steady state has been reached. Comparatively little work exists exploring different and specific ways that repeated cell divisions can lead to stochastic inheritance of unequilibrated cellular populations. Here we introduce a mathematical formalism to describe cellular agents that are subject to random creation, replication, and/or degradation, and are inherited according to a range of random dynamics at cell divisions. We obtain closed-form generating functions describing systems at any time after any number of cell divisions for binomial partitioning and divisions provoking a deterministic or random, subtractive or additive change in copy number, and show that these solutions agree exactly with stochastic simulation. We apply this general formalism to several example problems involving the dynamics of mitochondrial DNA (mtDNA) during development and organismal lifetimes.

preprint2015arXiv

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

Dangerous damage to mitochondrial DNA (mtDNA) can be ameliorated during mammalian development through a highly debated mechanism called the mtDNA bottleneck. Uncertainty surrounding this process limits our ability to address inherited mtDNA diseases. We produce a new, physically motivated, generalisable theoretical model for mtDNA populations during development, allowing the first statistical comparison of proposed bottleneck mechanisms. Using approximate Bayesian computation and mouse data, we find most statistical support for a combination of binomial partitioning of mtDNAs at cell divisions and random mtDNA turnover, meaning that the debated exact magnitude of mtDNA copy number depletion is flexible. New experimental measurements from a wild-derived mtDNA pairing in mice confirm the theoretical predictions of this model. We analytically solve a mathematical description of this mechanism, computing probabilities of mtDNA disease onset, efficacy of clinical sampling strategies, and effects of potential dynamic interventions, thus developing a quantitative and experimentally-supported stochastic theory of the bottleneck.

preprint2015arXiv

Unified treatment of the asymptotics of asymmetric kernel density estimators

We extend balloon and sample-smoothing estimators, two types of variable-bandwidth kernel density estimators, by a shift parameter and derive their asymptotic properties. Our approach facilitates the unified study of a wide range of density estimators which are subsumed under these two general classes of kernel density estimators. We demonstrate our method by deriving the asymptotic bias, variance, and mean (integrated) squared error of density estimators with gamma, log-normal, Birnbaum-Saunders, inverse Gaussian and reciprocal inverse Gaussian kernels. We propose two new density estimators for positive random variables that yield properly-normalised density estimates. Plugin expressions for bandwidth estimation are provided to facilitate easy exploratory data analysis.

preprint2014arXiv

Explicit tracking of uncertainty increases the power of quantitative rule-of-thumb reasoning in cell biology

"Back-of-the-envelope" or "rule-of-thumb" calculations involving rough estimates of quantities play a central scientific role in developing intuition about the structure and behaviour of physical systems, for example in so-called `Fermi problems' in the physical sciences. Such calculations can be used to powerfully and quantitatively reason about biological systems, particularly at the interface between physics and biology. However, substantial uncertainties are often associated with values in cell biology, and performing calculations without taking this uncertainty into account may limit the extent to which results can be interpreted for a given problem. We present a means to facilitate such calculations where uncertainties are explicitly tracked through the line of reasoning, and introduce a `probabilistic calculator' called Caladis, a web tool freely available at www.caladis.org, designed to perform this tracking. This approach allows users to perform more statistically robust calculations in cell biology despite having uncertain values, and to identify which quantities need to be measured more precisely in order to make confident statements, facilitating efficient experimental design. We illustrate the use of our tool for tracking uncertainty in several example biological calculations, showing that the results yield powerful and interpretable statistics on the quantities of interest. We also demonstrate that the outcomes of calculations may differ from point estimates when uncertainty is accurately tracked. An integral link between Caladis and the Bionumbers repository of biological quantities further facilitates the straightforward location, selection, and use of a wealth of experimental data in cell biological calculations.

preprint2014arXiv

Highly comparative fetal heart rate analysis

A database of fetal heart rate (FHR) time series measured from 7221 patients during labor is analyzed with the aim of learning the types of features of these recordings that are informative of low cord pH. Our 'highly comparative' analysis involves extracting over 9000 time-series analysis features from each FHR time series, including measures of autocorrelation, entropy, distribution, and various model fits. This diverse collection of features was developed in previous work, and is publicly available. We describe five features that most accurately classify a balanced training set of 59 'low pH' and 59 'normal pH' FHR recordings. We then describe five of the features with the strongest linear correlation to cord pH across the full dataset of FHR time series. The features identified in this work may be used as part of a system for guiding intervention during labor in future. This work successfully demonstrates the utility of comparing across a large, interdisciplinary literature on time-series analysis to automatically contribute new scientific results for specific biomedical signal processing challenges.

preprint2013arXiv

Highly comparative time-series analysis: The empirical structure of time series and their methods

The process of collecting and organizing sets of observations represents a common theme throughout the history of science. However, despite the ubiquity of scientists measuring, recording, and analyzing the dynamics of different processes, an extensive organization of scientific time-series data and analysis methods has never been performed. Addressing this, annotated collections of over 35 000 real-world and model-generated time series and over 9000 time-series analysis algorithms are analyzed in this work. We introduce reduced representations of both time series, in terms of their properties measured by diverse scientific methods, and of time-series analysis methods, in terms of their behaviour on empirical time series, and use them to organize these interdisciplinary resources. This new approach to comparing across diverse scientific data and methods allows us to organize time-series datasets automatically according to their properties, retrieve alternatives to particular analysis methods developed in other scientific disciplines, and automate the selection of useful methods for time-series classification and regression tasks. The broad scientific utility of these tools is demonstrated on datasets of electroencephalograms, self-affine time series, heart beat intervals, speech signals, and others, in each case contributing novel analysis techniques to the existing literature. Highly comparative techniques that compare across an interdisciplinary literature can thus be used to guide more focused research in time-series analysis for applications across the scientific disciplines.

preprint2013arXiv

How modular structure can simplify tasks on networks

By considering the task of finding the shortest walk through a network we find an algorithm for which the run time is not as O(2^n), with n being the number of nodes, but instead scales with the number of nodes in a coarsened network. This coarsened network has a number of nodes related to the number of dense regions in the original graph. Since we exploit a form of local community detection as a preprocessing, this work gives support to the project of developing heuristic algorithms for detecting dense regions in networks: preprocessing of this kind can accelerate optimization tasks on networks. Our work also suggests a class of empirical conjectures for how structural features of efficient networked systems might scale with system size.

preprint2012arXiv

Ancestral Inference from Functional Data: Statistical Methods and Numerical Examples

Many biological characteristics of evolutionary interest are not scalar variables but continuous functions. Here we use phylogenetic Gaussian process regression to model the evolution of simulated function-valued traits. Given function-valued data only from the tips of an evolutionary tree and utilising independent principal component analysis (IPCA) as a method for dimension reduction, we construct distributional estimates of ancestral function-valued traits, and estimate parameters describing their evolutionary dynamics.

preprint2012arXiv

Evolutionary Inference for Function-valued Traits: Gaussian Process Regression on Phylogenies

Biological data objects often have both of the following features: (i) they are functions rather than single numbers or vectors, and (ii) they are correlated due to phylogenetic relationships. In this paper we give a flexible statistical model for such data, by combining assumptions from phylogenetics with Gaussian processes. We describe its use as a nonparametric Bayesian prior distribution, both for prediction (placing posterior distributions on ancestral functions) and model selection (comparing rates of evolution across a phylogeny, or identifying the most likely phylogenies consistent with the observed data). Our work is integrative, extending the popular phylogenetic Brownian Motion and Ornstein-Uhlenbeck models to functional data and Bayesian inference, and extending Gaussian Process regression to phylogenies. We provide a brief illustration of the application of our method.

preprint2012arXiv

Function-Valued Traits in Evolution

Many biological characteristics of evolutionary interest are not scalar variables but continuous functions. Given a dataset of function-valued traits generated by evolution, we develop a practical statistical approach to infer ancestral function-valued traits, and estimate the generative evolutionary process. We do this by combining dimension reduction and phylogenetic Gaussian process regression, a nonparametric procedure which explicitly accounts for known phylogenetic relationships. We test the methods' performance on simulated function-valued data generated from a stochastic evolutionary model. The methods are applied assuming that only the phylogeny and the function-valued traits of taxa at its tips are known. Our method is robust and applicable to a wide range of function-valued data, and also offers a phylogenetically aware method for estimating the autocorrelation of function-valued traits.

preprint2012arXiv

Taxonomies of Networks

The study of networks has grown into a substantial interdisciplinary endeavour that encompasses myriad disciplines in the natural, social, and information sciences. Here we introduce a framework for constructing taxonomies of networks based on their structural similarities. These networks can arise from any of numerous sources: they can be empirical or synthetic, they can arise from multiple realizations of a single process, empirical or synthetic, or they can represent entirely different systems in different disciplines. Since the mesoscopic properties of networks are hypothesized to be important for network function, we base our comparisons on summaries of network community structures. While we use a specific method for uncovering network communities, much of the introduced framework is independent of that choice. After introducing the framework, we apply it to construct a taxonomy for 746 individual networks and demonstrate that our approach usefully identifies similar networks. We also construct taxonomies within individual categories of networks, and in each case we expose non-trivial structure. For example we create taxonomies for similarity networks constructed from both political voting data and financial data. We also construct network taxonomies to compare the social structures of 100 Facebook networks and the growth structures produced by different types of fungi.

preprint2011arXiv

Advection, diffusion and delivery over a network

Many biological, geophysical and technological systems involve the transport of resource over a network. In this paper we present an algorithm for calculating the exact concentration of resource at any point in space or time, given that the resource in the network is lost or delivered out of the network at a given rate, while being subject to advection and diffusion. We consider the implications of advection, diffusion and delivery for simple models of glucose delivery through a vascular network, and conclude that in certain circumstances, increasing the volume of blood and the number of glucose transporters can actually decrease the total rate of glucose delivery. We also consider the case of empirically determined fungal networks, and analyze the distribution of resource that emerges as such networks grow over time. Fungal growth involves the expansion of fluid filled vessels, which necessarily involves the movement of fluid. In three empirically determined fungal networks we found that the minimum currents consistent with the observed growth would effectively transport resource throughout the network over the time-scale of growth. This suggests that in foraging fungi, the active transport mechanisms observed in the growing tips may not be required for long range transport.

preprint2011arXiv

Mitochondrial Variability as a Source of Extrinsic Cellular Noise

We present a study investigating the role of mitochondrial variability in generating noise in eukaryotic cells. Noise in cellular physiology plays an important role in many fundamental cellular processes, including transcription, translation, stem cell differentiation and response to medication, but the specific random influences that affect these processes have yet to be clearly elucidated. Here we present a mechanism by which variability in mitochondrial volume and functionality, along with cell cycle dynamics, is linked to variability in transcription rate and hence has a profound effect on downstream cellular processes. Our model mechanism is supported by an appreciable volume of recent experimental evidence, and we present the results of several new experiments with which our model is also consistent. We find that noise due to mitochondrial variability can sometimes dominate over other extrinsic noise sources (such as cell cycle asynchronicity) and can significantly affect large-scale observable properties such as cell cycle length and gene expression levels. We also explore two recent regulatory network-based models for stem cell differentiation, and find that extrinsic noise in transcription rate causes appreciable variability in the behaviour of these model systems. These results suggest that mitochondrial and transcriptional variability may be an important mechanism influencing a large variety of cellular processes and properties.

preprint2011arXiv

Temporal Evolution of Financial Market Correlations

We investigate financial market correlations using random matrix theory and principal component analysis. We use random matrix theory to demonstrate that correlation matrices of asset price changes contain structure that is incompatible with uncorrelated random price changes. We then identify the principal components of these correlation matrices and demonstrate that a small number of components accounts for a large proportion of the variability of the markets that we consider. We then characterize the time-evolving relationships between the different assets by investigating the correlations between the asset price time series and principal components. Using this approach, we uncover notable changes that occurred in financial markets and identify the assets that were significantly affected by these changes. We show in particular that there was an increase in the strength of the relationships between several different markets following the 2007--2008 credit and liquidity crisis.

preprint2010arXiv

Dynamical Clustering of Exchange Rates

We use techniques from network science to study correlations in the foreign exchange (FX) market over the period 1991--2008. We consider an FX market network in which each node represents an exchange rate and each weighted edge represents a time-dependent correlation between the rates. To provide insights into the clustering of the exchange rate time series, we investigate dynamic communities in the network. We show that there is a relationship between an exchange rate's functional role within the market and its position within its community and use a node-centric community analysis to track the time dynamics of this role. This reveals which exchange rates dominate the market at particular times and also identifies exchange rates that experienced significant changes in market role. We also use the community dynamics to uncover major structural changes that occurred in the FX market. Our techniques are general and will be similarly useful for investigating correlations in other markets.

preprint2010arXiv

Revisiting Date and Party Hubs: Novel Approaches to Role Assignment in Protein Interaction Networks

The idea of 'date' and 'party' hubs has been influential in the study of protein-protein interaction networks. Date hubs display low co-expression with their partners, whilst party hubs have high co-expression. It was proposed that party hubs are local coordinators whereas date hubs are global connectors. Here we show that the reported importance of date hubs to network connectivity can in fact be attributed to a tiny subset of them. Crucially, these few, extremely central, hubs do not display particularly low expression correlation, undermining the idea of a link between this quantity and hub function. The date/party distinction was originally motivated by an approximately bimodal distribution of hub co-expression; we show that this feature is not always robust to methodological changes. Additionally, topological properties of hubs do not in general correlate with co-expression. Thus, we suggest that a date/party dichotomy is not meaningful and it might be more useful to conceive of roles for protein-protein interactions rather than individual proteins. We find significant correlations between interaction centrality and the functional similarity of the interacting proteins.

preprint2010arXiv

Sparse bayesian step-filtering for high-throughput analysis of molecular machine dynamics

Nature has evolved many molecular machines such as kinesin, myosin, and the rotary flagellar motor powered by an ion current from the mitochondria. Direct observation of the step-like motion of these machines with time series from novel experimental assays has recently become possible. These time series are corrupted by molecular and experimental noise that requires removal, but classical signal processing is of limited use for recovering such step-like dynamics. This paper reports simple, novel Bayesian filters that are robust to step-like dynamics in noise, and introduce an L1-regularized, global filter whose sparse solution can be rapidly obtained by standard convex optimization methods. We show these techniques outperforming classical filters on simulated time series in terms of their ability to accurately recover the underlying step dynamics. To show the techniques in action, we extract step-like speed transitions from Rhodobacter sphaeroides flagellar motor time series. Code implementing these algorithms available from http://www.eng.ox.ac.uk/samp/members/max/software/.

preprint2010arXiv

The Function of Communities in Protein Interaction Networks at Multiple Scales

Background: If biology is modular then clusters, or communities, of proteins derived using only protein interaction network structure should define protein modules with similar biological roles. We investigate the link between biological modules and network communities in yeast and its relationship to the scale at which we probe the network. Results: Our results demonstrate that the functional homogeneity of communities depends on the scale selected, and that almost all proteins lie in a functionally homogeneous community at some scale. We judge functional homogeneity using a novel test and three independent characterizations of protein function, and find a high degree of overlap between these measures. We show that a high mean clustering coefficient of a community can be used to identify those that are functionally homogeneous. By tracing the community membership of a protein through multiple scales we demonstrate how our approach could be useful to biologists focusing on a particular protein. Conclusions: We show that there is no one scale of interest in the community structure of the yeast protein interaction network, but we can identify the range of resolution parameters that yield the most functionally coherent communities, and predict which communities are most likely to be functionally homogeneous.

preprint2009arXiv

Dynamic communities in multichannel data: An application to the foreign exchange market during the 2007--2008 credit crisis

We study the cluster dynamics of multichannel (multivariate) time series by representing their correlations as time-dependent networks and investigating the evolution of network communities. We employ a node-centric approach that allows us to track the effects of the community evolution on the functional roles of individual nodes without having to track entire communities. As an example, we consider a foreign exchange market network in which each node represents an exchange rate and each edge represents a time-dependent correlation between the rates. We study the period 2005-2008, which includes the recent credit and liquidity crisis. Using dynamical community detection, we find that exchange rates that are strongly attached to their community are persistently grouped with the same set of rates, whereas exchange rates that are important for the transfer of information tend to be positioned on the edges of communities. Our analysis successfully uncovers major trading changes that occurred in the market during the credit crisis.