Catalog footprint

What is connected

36works
23topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

36 published item(s)

preprint2022arXiv

Ballistic deposition with memory: a new universality class of surface growth with a new scaling law

Motivated by recent experimental studies in microbiology, we suggest a modification of the classic ballistic deposition model of surface growth, where the memory of a deposition at a site induces more depositions at that site or its neighbors. By studying the statistics of surfaces in this model, we obtain three independent critical exponents: the growth exponent $β=5/4$, the roughening exponent $α= 2$, and the new (size) exponent $γ= 1/2$. The model requires a modification to the Family-Vicsek scaling, resulting in the dynamical exponent $z = \frac{α+γ}β = 2$. This modified scaling collapses the surface width vs time curves for various lattice sizes. This is a previously unobserved universality class of surface growth that could describe surface properties of a wide range of natural systems.

preprint2022arXiv

Intrinsic Motivation in Dynamical Control Systems

Biological systems often choose actions without an explicit reward signal, a phenomenon known as intrinsic motivation. The computational principles underlying this behavior remain poorly understood. In this study, we investigate an information-theoretic approach to intrinsic motivation, based on maximizing an agent's empowerment (the mutual information between its past actions and future states). We show that this approach generalizes previous attempts to formalize intrinsic motivation, and we provide a computationally efficient algorithm for computing the necessary quantities. We test our approach on several benchmark control problems, and we explain its success in guiding intrinsically motivated behaviors by relating our information-theoretic control function to fundamental properties of the dynamical system representing the combined agent-environment system. This opens the door for designing practical artificial, intrinsically motivated controllers and for linking animal behaviors to their dynamical properties.

preprint2022arXiv

Multi-dimensional structure of C. elegans thermal learning

Quantitative models of associative learning that explain the behavior of real animals with high precision have turned out very difficult to construct. We do this in the context of the dynamics of the thermal preference of C. elegans. For this, we quantify C. elegans thermotaxis in response to various conditioning parameters, genetic perturbations, and operant behavior using a fast, high-throughput microfluidic droplet assay. We then model this data comprehensively, within a new, biologically interpretable, multi-modal framework. We discover that the dynamics of thermal preference are described by two independent contributions and require a model with at least four dynamical variables. One pathway positively associates the experienced temperature independently of food and the other negatively associates to the temperature when food is absent.

preprint2021arXiv

Statistical properties of large data sets with linear latent features

Analytical understanding of how low-dimensional latent features reveal themselves in large-dimensional data is still lacking. We study this by defining a linear latent feature model with additive noise constructed from probabilistic matrices, and analytically and numerically computing the statistical distributions of pairwise correlations and eigenvalues of the correlation matrix. This allows us to resolve the latent feature structure across a wide range of data regimes set by the number of recorded variables, observations, latent features and the signal-to-noise ratio. We find a characteristic imprint of latent features in the distribution of correlations and eigenvalues and provide an analytic estimate for the boundary between signal and noise even in the absence of a clear spectral gap.

preprint2019arXiv

Precise Spatial Memory in Local Random Networks

Self-sustained, elevated neuronal activity persisting on time scales of ten seconds or longer is thought to be vital for aspects of working memory, including brain representations of real space. Continuous-attractor neural networks, one of the most well-known modeling frameworks for persistent activity, have been able to model crucial aspects of such spatial memory. These models tend to require highly structured or regular synaptic architectures. In contrast, we elaborate a geometrically-embedded model with a local but otherwise random connectivity profile which, combined with a global regulation of the mean firing rate, produces localized, finely spaced discrete attractors that effectively span a 2D manifold. We demonstrate how the set of attracting states can reliably encode a representation of the spatial locations at which the system receives external input, thereby accomplishing spatial memory via attractor dynamics without synaptic fine-tuning or regular structure. We measure the network's storage capacity and find that the statistics of retrievable positions are also equivalent to a full tiling of the plane, something hitherto achievable only with (approximately) translationally invariant synapses, and which may be of interest in modeling such biological phenomena as visuospatial working memory in two dimensions.

preprint2019arXiv

Randomly connected networks generate emergent selectivity and predict decoding properties of large populations of neurons

Advances in neural recording methods enable sampling from populations of thousands of neurons during the performance of behavioral tasks, raising the question of how recorded activity relates to the theoretical models of computations underlying performance. In the context of decision making in rodents, patterns of functional connectivity between choice-selective cortical neurons, as well as broadly distributed choice information in both excitatory and inhibitory populations, were recently reported [1]. The straightforward interpretation of these data suggests a mechanism relying on specific patterns of anatomical connectivity to achieve selective pools of inhibitory as well as excitatory neurons. We investigate an alternative mechanism for the emergence of these experimental observations using a computational approach. We find that a randomly connected network of excitatory and inhibitory neurons generates single-cell selectivity, patterns of pairwise correlations, and indistinguishable excitatory and inhibitory readout weight distributions, as observed in recorded neural populations. Further, we make the readily verifiable experimental predictions that, for this type of evidence accumulation task, there are no anatomically defined sub-populations of neurons representing choice, and that choice preference of a particular neuron changes with the details of the task. This work suggests that distributed stimulus selectivity and patterns of functional organization in population codes could be emergent properties of randomly connected networks.

preprint2019arXiv

Universal properties of concentration sensing in large ligand-receptor networks

Cells estimate concentrations of chemical ligands in their environment using a limited set of receptors. Recent work has shown that the temporal sequence of binding and unbinding events on just a single receptor can be used to estimate the concentrations of multiple ligands. Here, for a network of many ligands and many receptors, we show that such temporal sequences can be used to estimate the concentration of a few times as many ligand species as there are receptors. Crucially, we show that the spectrum of the inverse covariance matrix of these estimates has several universal properties, which we trace to properties of Vandermonde matrices. We argue that this can be used by cells in realistic biochemical decoding networks.

preprint2017arXiv

Chance, long tails, and inference: a non-Gaussian, Bayesian theory of vocal learning in songbirds

Traditional theories of sensorimotor learning posit that animals use sensory error signals to find the optimal motor command in the face of Gaussian sensory and motor noise. However, most such theories cannot explain common behavioral observations, for example that smaller sensory errors are more readily corrected than larger errors and that large abrupt (but not gradually introduced) errors lead to weak learning. Here we propose a new theory of sensorimotor learning that explains these observations. The theory posits that the animal learns an entire probability distribution of motor commands rather than trying to arrive at a single optimal command, and that learning arises via Bayesian inference when new sensory information becomes available. We test this theory using data from a songbird, the Bengalese finch, that is adapting the pitch (fundamental frequency) of its song following perturbations of auditory feedback using miniature headphones. We observe the distribution of the sung pitches to have long, non-Gaussian tails, which, within our theory, explains the observed dynamics of learning. Further, the theory makes surprising predictions about the dynamics of the shape of the pitch distribution, which we confirm experimentally.

preprint2016arXiv

Extrinsic and intrinsic correlations in molecular information transmission

Cells measure concentrations of external ligands by capturing ligand molecules with cell surface receptors. The numbers of molecules captured by different receptors co-vary because they depend on the same extrinsic ligand fluctuations. However, these numbers also counter-vary due to the intrinsic stochasticity of chemical processes because a single molecule randomly captured by a receptor cannot be captured by another. Such structure of receptor correlations is generally believed to lead to an increase in information about the external signal compared to the case of independent receptors. We analyze a solvable model of two molecular receptors and show that, contrary to this widespread expectation, the correlations have a small and negative effect on the information about the ligand concentration. Further, we show that measurements that average over multiple receptors are almost as informative as those that track the states of every individual one.

preprint2015arXiv

Accurate sensing of multiple ligands with a single receptor

Cells use surface receptors to estimate the concentration of external ligands. Limits on the accuracy of such estimations have been well studied for pairs of ligand and receptor species. However, the environment typically contains many ligands, which can bind to the same receptors with different affinities, resulting in cross-talk. In traditional rate models, such cross-talk prevents accurate inference of individual ligand concentrations. In contrast, here we show that knowing the precise timing sequence of stochastic binding and unbinding events allows one receptor to provide information about multiple ligands simultaneously and with a high accuracy. We argue that such high-accuracy estimation of multiple concentrations can be realized by the familiar kinetic proofreading mechanism.

preprint2015arXiv

Cell-cell communication enhances the capacity of cell ensembles to sense shallow gradients during morphogenesis

Collective cell responses to exogenous cues depend on cell-cell interactions. In principle, these can result in enhanced sensitivity to weak and noisy stimuli. However, this has not yet been shown experimentally, and, little is known about how multicellular signal processing modulates single cell sensitivity to extracellular signaling inputs, including those guiding complex changes in the tissue form and function. Here we explored if cell-cell communication can enhance the ability of cell ensembles to sense and respond to weak gradients of chemotactic cues. Using a combination of experiments with mammary epithelial cells and mathematical modeling, we find that multicellular sensing enables detection of and response to shallow Epidermal Growth Factor (EGF) gradients that are undetectable by single cells. However, the advantage of this type of gradient sensing is limited by the noisiness of the signaling relay, necessary to integrate spatially distributed ligand concentration information. We calculate the fundamental sensory limits imposed by this communication noise and combine them with the experimental data to estimate the effective size of multicellular sensory groups involved in gradient sensing. Functional experiments strongly implicated intercellular communication through gap junctions and calcium release from intracellular stores as mediators of collective gradient sensing. The resulting integrative analysis provides a framework for understanding the advantages and limitations of sensory information processing by relays of chemically coupled cells.

preprint2015arXiv

Limits to the precision of gradient sensing with spatial communication and temporal integration

Gradient sensing requires at least two measurements at different points in space. These measurements must then be communicated to a common location to be compared, which is unavoidably noisy. While much is known about the limits of measurement precision by cells, the limits placed by the communication are not understood. Motivated by recent experiments, we derive the fundamental limits to the precision of gradient sensing in a multicellular system, accounting for communication and temporal integration. The gradient is estimated by comparing a "local" and a "global" molecular reporter of the external concentration, where the global reporter is exchanged between neighboring cells. Using the fluctuation-dissipation framework, we find, in contrast to the case when communication is ignored, that precision saturates with the number of cells independently of the measurement time duration, since communication establishes a maximum lengthscale over which sensory information can be reliably conveyed. Surprisingly, we also find that precision is improved if the local reporter is exchanged between cells as well, albeit more slowly than the global reporter. The reason is that while exchange of the local reporter weakens the comparison, it decreases the measurement noise. We term such a model "regional excitation--global inhibition" (REGI). Our results demonstrate that fundamental sensing limits are necessarily sharpened when the need to communicate information is taken into account.

preprint2015arXiv

Of fishes and birthdays: Efficient estimation of polymer configurational entropies

We present an algorithm to estimate the configurational entropy $S$ of a polymer. The algorithm uses the statistics of coincidences among random samples of configurations and is related to the catch-tag-release method for estimation of population sizes, and to the classic "birthday paradox". Bias in the entropy estimation is decreased by grouping configurations in nearly equiprobable partitions based on their energies, and estimating entropies separately within each partition. Whereas most entropy estimation algorithms require $N\sim 2^{S}$ samples to achieve small bias, our approach typically needs only $N\sim \sqrt{2^{S}}$. Thus the algorithm can be applied to estimate protein free energies with increased accuracy and decreased computational cost.

preprint2015arXiv

On the sufficiency of pairwise interactions in maximum entropy models of biological networks

Biological information processing networks consist of many components, which are coupled by an even larger number of complex multivariate interactions. However, analyses of data sets from fields as diverse as neuroscience, molecular biology, and behavior have reported that observed statistics of states of some biological networks can be approximated well by maximum entropy models with only pairwise interactions among the components. Based on simulations of random Ising spin networks with $p$-spin ($p>2$) interactions, here we argue that this reduction in complexity can be thought of as a natural property of densely interacting networks in certain regimes, and not necessarily as a special property of living systems. By connecting our analysis to the theory of random constraint satisfaction problems, we suggest a reason for why some biological systems may operate in this regime.

preprint2015arXiv

Role of spatial averaging in multicellular gradient sensing

Gradient sensing underlies important biological processes including morphogenesis, polarization, and cell migration. The precision of gradient sensing increases with the length of a detector (a cell or group of cells) in the gradient direction, since a longer detector spans a larger range of concentration values. Intuition from analyses of concentration sensing suggests that precision should also increase with detector length in the direction transverse to the gradient, since then spatial averaging should reduce the noise. However, here we show that, unlike for concentration sensing, the precision of gradient sensing decreases with transverse length for the simplest gradient sensing model, local excitation--global inhibition (LEGI). The reason is that gradient sensing ultimately relies on a subtraction of measured concentration values. While spatial averaging indeed reduces the noise in these measurements, which increases precision, it also reduces the covariance between the measurements, which results in the net decrease in precision. We demonstrate how a recently introduced gradient sensing mechanism, regional excitation--global inhibition (REGI), overcomes this effect and recovers the benefit of transverse averaging. Using a REGI-based model, we compute the optimal two- and three-dimensional detector shapes, and argue that they are consistent with the shapes of naturally occurring gradient-sensing cell populations.

preprint2015arXiv

Time to Quantify Falsifiability

Here we argue that the notion of falsifiability, a key concept in defining a valid scientific theory, can be quantified using Bayesian Model Selection, which is a standard tool in modern statistics. This relates falsifiability to the quantitative version of the Occam's razor, and allows transforming some long-running arguments about validity of certain scientific theories from philosophical discussions to mathematical calculations. This is a Letter to the editor.

preprint2014arXiv

Automated adaptive inference of coarse-grained dynamical models in systems biology

Cellular regulatory dynamics is driven by large and intricate networks of interactions at the molecular scale, whose sheer size obfuscates understanding. In light of limited experimental data, many parameters of such dynamics are unknown, and thus models built on the detailed, mechanistic viewpoint overfit and are not predictive. At the other extreme, simple ad hoc models of complex processes often miss defining features of the underlying systems. Here we propose an approach that instead constructs phenomenological, coarse-grained models of network dynamics that automatically adapt their complexity to the amount of available data. Such adaptive models lead to accurate predictions even when microscopic details of the studied systems are unknown due to insufficient data. The approach is computationally tractable, even for a relatively large number of dynamical variables, allowing its software realization, named Sir Isaac, to make successful predictions even when important dynamic variables are unobserved. For example, it matches the known phase space structure for simulated planetary motion data, avoids overfitting in a complex biological signaling system, and produces accurate predictions for a yeast glycolysis model with only tens of data points and over half of the interacting species unobserved.

preprint2014arXiv

Director Field Model of the Primary Visual Cortex for Contour Detection

We aim to build the simplest possible model capable of detecting long, noisy contours in a cluttered visual scene. For this, we model the neural dynamics in the primate primary visual cortex in terms of a continuous director field that describes the average rate and the average orientational preference of active neurons at a particular point in the cortex. We then use a linear-nonlinear dynamical model with long range connectivity patterns to enforce long-range statistical context present in the analyzed images. The resulting model has substantially fewer degrees of freedom than traditional models, and yet it can distinguish large contiguous objects from the background clutter by suppressing the clutter and by filling-in occluded elements of object contours. This results in high-precision, high-recall detection of large objects in cluttered scenes. Parenthetically, our model has a direct correspondence with the Landau - de Gennes theory of nematic liquid crystal in two dimensions.

preprint2014arXiv

Efficient inference of parsimonious phenomenological models of cellular dynamics using S-systems and alternating regression

The nonlinearity of dynamics in systems biology makes it hard to infer them from experimental data. Simple linear models are computationally efficient, but cannot incorporate these important nonlinearities. An adaptive method based on the S-system formalism, which is a sensible representation of nonlinear mass-action kinetics typically found in cellular dynamics, maintains the efficiency of linear regression. We combine this approach with adaptive model selection to obtain efficient and parsimonious representations of cellular dynamics. The approach is tested by inferring the dynamics of yeast glycolysis from simulated data. With little computing time, it produces dynamical models with high predictive power and with structural complexity adapted to the difficulty of the inference problem.

preprint2014arXiv

Millisecond-scale motor encoding in a cortical vocal area

Studies of motor control have almost universally examined firing rates to investigate how the brain shapes behavior. In principle, however, neurons could encode information through the precise temporal patterning of their spike trains as well as (or instead of) through their firing rates. Although the importance of spike timing has been demonstrated in sensory systems, it is largely unknown whether timing differences in motor areas could affect behavior. We tested the hypothesis that significant information about trial-by-trial variations in behavior is represented by spike timing in the songbird vocal motor system. We found that premotor neurons convey information via spike timing far more often than via spike rate and that the amount of information conveyed at the millisecond timescale greatly exceeds the information available from spike counts. These results demonstrate that information can be represented by spike timing in motor circuits and suggest that timing variations evoke differences in behavior.

preprint2014arXiv

Zipf's law and criticality in multivariate data without fine-tuning

The joint probability distribution of many degrees of freedom in biological systems, such as firing patterns in neural networks or antibody sequence composition in zebrafish, often follow Zipf's law, where a power law is observed on a rank-frequency plot. This behavior has recently been shown to imply that these systems reside near to a unique critical point where the extensive parts of the entropy and energy are exactly equal. Here we show analytically, and via numerical simulations, that Zipf-like probability distributions arise naturally if there is an unobserved variable (or variables) that affects the system, e. g. for neural networks an input stimulus that causes individual neurons in the network to fire at time-varying rates. In statistics and machine learning, these models are called latent-variable or mixture models. Our model shows that no fine-tuning is required, i.e. Zipf's law arises generically without tuning parameters to a point, and gives insight into the ubiquity of Zipf's law in a wide range of systems.

preprint2013arXiv

Genotype to phenotype mapping and the fitness landscape of the E. coli lac promoter

Genotype-to-phenotype maps and the related fitness landscapes that include epistatic interactions are difficult to measure because of their high dimensional structure. Here we construct such a map using the recently collected corpora of high-throughput sequence data from the 75 base pairs long mutagenized E. coli lac promoter region, where each sequence is associated with its phenotype, the induced transcriptional activity measured by a fluorescent reporter. We find that the additive (non-epistatic) contributions of individual mutations account for about two-thirds of the explainable phenotype variance, while pairwise epistasis explains about 7% of the variance for the full mutagenized sequence and about 15% for the subsequence associated with protein binding sites. Surprisingly, there is no evidence for third order epistatic contributions, and our inferred fitness landscape is essentially single peaked, with a small amount of antagonistic epistasis. There is a significant selective pressure on the wild type, which we deduce to be multi-objective optimal for gene expression in environments with different nutrient sources. We identify transcription factor (CRP) and RNA polymerase binding sites in the promotor region and their interactions without difficult optimization steps. In particular, we observe evidence for previously unexplored genetic regulatory mechanisms, possibly kinetic in nature. We conclude with a cautionary note that inferred properties of fitness landscapes may be severely influenced by biases in the sequence data.

preprint2012arXiv

Large number of receptors may reduce cellular response time variation

Cells often have tens of thousands of receptors, even though only a few activated receptors can trigger full cellular responses. Reasons for the overabundance of receptors remain unclear. We suggest that, in certain conditions, the large number of receptors results in a competition among receptors to be the first to activate the cell. The competition decreases the variability of the time to cellular activation, and hence results in a more synchronous activation of cells. We argue that, in simple models, this variability reduction does not necessarily interfere with the receptor specificity to ligands achieved by the kinetic proofreading mechanism. Thus cells can be activated accurately in time and specifically to certain signals. We predict the minimum number of receptors needed to reduce the coefficient of variation for the time to activation following binding of a specific ligand. Further, we predict the maximum number of receptors so that the kinetic proofreading mechanism still can improve the specificity of the activation. These predictions fall in line with experimentally reported receptor numbers for multiple systems.

preprint2012arXiv

Population-expression models of immune response

The immune response to a pathogen has two basic features. The first is the expansion of a few pathogen-specific cells to form a population large enough to control the pathogen. The second is the process of differentiation of cells from an initial naive phenotype to an effector phenotype which controls the pathogen, and subsequently to a memory phenotype that is maintained and responsible for long-term protection. The expansion and the differentiation have been considered largely independently. Changes in cell populations are typically described using ecologically based ordinary differential equation models. In contrast, differentiation of single cells is studied within systems biology and is frequently modeled by considering changes in gene and protein expression in individual cells. Recent advances in experimental systems biology make available for the first time data to allow the coupling of population and high dimensional expression data of immune cells during infections. Here we describe and develop population-expression models which integrate these two processes into systems biology on the multicellular level. When translated into mathematical equations, these models result in non-conservative, non-local advection-diffusion equations. We describe situations where the population-expression approach can make correct inference from data while previous modeling approaches based on common simplifying assumptions would fail. We also explore how model reduction techniques can be used to build population-expression models, minimizing the complexity of the model while keeping the essential features of the system. While we consider problems in immunology in this paper, we expect population-expression models to be more broadly applicable.

preprint2012arXiv

Predictive information in a nonequilibrium critical model

We propose predictive information, that is information between a long past of duration T and the entire infinitely long future of a time series, as a universal order parameter to study phase transitions in physical systems. It can be used, in particular, to study nonequlibrium transitions and other exotic transitions, where a simpler order parameter cannot be identifies using traditional symmetry arguments. As an example, we calculate predictive information for a stochastic nonequilibrium dynamics problem that forms an absorbing state under a continuous change of a parameter. The information at the transition point diverges as log(T), and a smooth crossover to constant away from the transition is observed.

preprint2011arXiv

Fitness in time-dependent environments includes a geometric contribution

Phenotypic evolution implies sequential fixations of new genomic sequences. The speed at which these mutations fixate depends, in part, on the relative fitness (selection coefficient) of the mutant vs. the ancestor. Using a simple population dynamics model we show that the relative fitness in dynamical environments is not equal to the fitness averaged over individual environments. Instead it includes a term that explicitly depends on the sequence of the environments. This term is geometric in nature and depends only on the oriented area enclosed by the trajectory taken by the system in the environment state space. It is related to the well-studied geometric phases in classical and quantum physical systems. We discuss possible biological implications of these observations, focusing on evolution of novel metabolic or stress-resistant functions.

preprint2011arXiv

Gain control in molecular information processing: Lessons from neuroscience

Statistical properties of environments experienced by biological signaling systems in the real world change, which necessitate adaptive responses to achieve high fidelity information transmission. One form of such adaptive response is gain control. Here we argue that a certain simple mechanism of gain control, understood well in the context of systems neuroscience, also works for molecular signaling. The mechanism allows to transmit more than one bit (on or off) of information about the signal independently of the signal variance. It does not require additional molecular circuitry beyond that already present in many molecular systems, and, in particular, it does not depend on existence of feedback loops. The mechanism provides a potential explanation for abundance of ultrasensitive response curves in biological regulatory networks.

preprint2011arXiv

Model Cortical Association Fields Account for the Time Course and Dependence on Target Complexity of Human Contour Perception

Can lateral connectivity in the primary visual cortex account for the time dependence and intrinsic task difficulty of human contour detection? To answer this question, we created a synthetic image set that prevents sole reliance on either low-level visual features or high-level context for the detection of target objects. Rendered images consist of smoothly varying, globally aligned contour fragments (amoebas) distributed among groups of randomly rotated fragments (clutter). The time course and accuracy of amoeba detection by humans was measured using a two-alternative forced choice protocol with self-reported confidence and variable image presentation time (20-200 ms), followed by an image mask optimized so as to interrupt visual processing. Measured psychometric functions were well fit by sigmoidal functions with exponential time constants of 30-91 ms, depending on amoeba complexity. Key aspects of the psychophysical experiments were accounted for by a computational network model, in which simulated responses across retinotopic arrays of orientation-selective elements were modulated by cortical association fields, represented as multiplicative kernels computed from the differences in pairwise edge statistics between target and distractor images. Comparing the experimental and the computational results suggests that each iteration of the lateral interactions takes at least 37.5 ms of cortical processing time. Our results provide evidence that cortical association fields between orientation selective elements in early visual areas can account for important temporal and task-dependent aspects of the psychometric curves characterizing human contour perception, with the remaining discrepancies postulated to arise from the influence of higher cortical areas.

preprint2010arXiv

Information theory and adaptation

In this Chapter, we ask questions (1) What is the right way to measure the quality of information processing in a biological system? and (2) What can real-life organisms do in order to improve their performance in information-processing tasks? We then review the body of work that investigates these questions experimentally, computationally, and theoretically in biological domains as diverse as cell biology, population biology, and computational neuroscience

preprint2010arXiv

Mass Conservation And Inference of Metabolic Networks from High-throughput Mass Spectrometry Data

We present a step towards the metabolome-wide computational inference of cellular metabolic reaction networks from metabolic profiling data, such as mass spectrometry. The reconstruction is based on identification of irreducible statistical interactions among the metabolite activities using the ARACNE reverse-engineering algorithm and on constraining possible metabolic transformations to satisfy the conservation of mass. The resulting algorithms are validated on synthetic data from an abridged computational model of Escherichia coli metabolism. Precision rates upwards of 50% are routinely observed for identification of full metabolic reactions, and recalls upwards of 20% are also seen.

preprint2010arXiv

Multivariate dependence and genetic networks inference

A critical task in systems biology is the identification of genes that interact to control cellular processes by transcriptional activation of a set of target genes. Many methods have been developed to use statistical correlations in high-throughput datasets to infer such interactions. However, cellular pathways are highly cooperative, often requiring the joint effect of many molecules, and few methods have been proposed to explicitly identify such higher-order interactions, partially due to the fact that the notion of multivariate statistical dependency itself remains imprecisely defined. We define the concept of dependence among multiple variables using maximum entropy techniques and introduce computational tests for their identification. Synthetic network results reveal that this procedure uncovers dependencies even in undersampled regimes, when the joint probability distribution cannot be reliably estimated. Analysis of microarray data from human B cells reveals that third-order statistics, but not second-order ones, uncover relationships between genes that interact in a pathway to cooperatively regulate a common set of targets.

preprint2010arXiv

Speeding up evolutionary search by small fitness fluctuations

We consider a fixed size population that undergoes an evolutionary adaptation in the weak mutuation rate limit, which we model as a biased Langevin process in the genotype space. We show analytically and numerically that, if the fitness landscape has a small highly epistatic (rough) and time-varying component, then the population genotype exhibits a high effective diffusion in the genotype space and is able to escape local fitness minima with a large probability. We argue that our principal finding that even very small time-dependent fluctuations of fitness can substantially speed up evolution is valid for a wide class of models.

preprint2008arXiv

Statistical properties of multistep enzyme-mediated reactions

Enzyme-mediated reactions may proceed through multiple intermediate conformational states before creating a final product molecule, and one often wishes to identify such intermediate structures from observations of the product creation. In this paper, we address this problem by solving the chemical master equations for various enzymatic reactions. We devise a perturbation theory analogous to that used in quantum mechanics that allows us to determine the first (<n>) and the second (variance) cumulants of the distribution of created product molecules as a function of the substrate concentration and the kinetic rates of the intermediate processes. The mean product flux V=d<n>/dt (or "dose-response" curve) and the Fano factor F=variance/<n> are both realistically measurable quantities, and while the mean flux can often appear the same for different reaction types, the Fano factor can be quite different. This suggests both qualitative and quantitative ways to discriminate between different reaction schemes, and we explore this possibility in the context of four sample multistep enzymatic reactions. We argue that measuring both the mean flux and the Fano factor can not only discriminate between reaction types, but can also provide some detailed information about the internal, unobserved kinetic rates, and this can be done without measuring single-molecule transition events.

preprint2006arXiv

Optimal signal processing in small stochastic biochemical networks

We quantify the influence of the topology of a transcriptional regulatory network on its ability to process environmental signals. By posing the problem in terms of information theory, we may do this without specifying the function performed by the network. Specifically, we study the maximum mutual information between the input (chemical) signal and the output (genetic) response attainable by the network in the context of an analytic model of particle number fluctuations. We perform this analysis for all biochemical circuits, including various feedback loops, that can be built out of 3 chemical species, each under the control of one regulator. We find that a generic network, constrained to low molecule numbers and reasonable response times, can transduce more information than a simple binary switch and, in fact, manages to achieve close to the optimal information transmission fidelity. These high-information solutions are robust to tenfold changes in most of the networks' biochemical parameters; moreover they are easier to achieve in networks containing cycles with an odd number of negative regulators (overall negative feedback) due to their decreased molecular noise (a result which we derive analytically). Finally, we demonstrate that a single circuit can support multiple high-information solutions. These findings suggest a potential resolution of the "cross-talk" dilemma as well as the previously unexplained observation that transcription factors which undergo proteolysis are more likely to be auto-repressive.

preprint2002arXiv

Inference of entropies of discrete random variables with unknown cardinalities

We examine the recently introduced NSB estimator of entropies of severely undersampled discrete variables and devise a procedure for calculating the involved integrals. We discover that the output of the estimator has a well defined limit for large cardinalities of the variables being studied. Thus one can estimate entropies with no a priori assumptions about these cardinalities, and a closed form solution for such estimates is given.

preprint2001arXiv

Predictability, complexity and learning

We define {\em predictive information} $I_{\rm pred} (T)$ as the mutual information between the past and the future of a time series. Three qualitatively different behaviors are found in the limit of large observation times $T$: $I_{\rm pred} (T)$ can remain finite, grow logarithmically, or grow as a fractional power law. If the time series allows us to learn a model with a finite number of parameters, then $I_{\rm pred} (T)$ grows logarithmically with a coefficient that counts the dimensionality of the model space. In contrast, power--law growth is associated, for example, with the learning of infinite parameter (or nonparametric) models such as continuous functions with smoothness constraints. There are connections between the predictive information and measures of complexity that have been defined both in learning theory and in the analysis of physical systems through statistical mechanics and dynamical systems theory. Further, in the same way that entropy provides the unique measure of available information consistent with some simple and plausible conditions, we argue that the divergent part of $I_{\rm pred} (T)$ provides the unique measure for the complexity of dynamics underlying a time series. Finally, we discuss how these ideas may be useful in different problems in physics, statistics, and biology.