Topic overview

Quantitative Methods

1848 works8425 researchers

Map preview

Start with the graph, then narrow the list

1848works
8425researchers

Next steps

Use the topic as a working map

Open the full map for clusters, then return here to scan ranked papers and people.

Topic graph

See the topic as a live network

Open full explorer

Inspect nearby papers, researchers, institutions and communities without opening a separate graph page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Papers in this area

24 paper(s) to start with

preprint2017arXiv

Inverse Protein Folding Problem via Quadratic Programming

This paper presents a method of reconstruction a primary structure of a protein that folds into a given geometrical shape. This method predicts the primary structure of a protein and restores its linear sequence of amino acids in the polypeptide chain using the tertiary structure of a molecule. Unknown amino acids are determined according to the principle of energy minimization. This study represents inverse folding problem as a quadratic optimization problem and uses different relaxation techniques to reduce it to the problem of convex optimizations. Computational experiment compares the quality of these approaches on real protein structures.

preprint2016arXiv

The impact of Gene Ontology evolution on GO-Term Information Content

The Gene Ontology (GO) is a major bioinformatics ontology that provides structured controlled vocabularies to classify gene and proteins function and role. The GO and its annotations to gene products are now an integral part of functional analysis. Recently, the evaluation of similarity among gene products starting from their annotations (also referred to as semantic similarities) has become an increasing area in bioinformatics. While many research on updates to the structure of GO and on the annotation corpora have been made, the impact of GO evolution on semantic similarities is quite unobserved. Here we extensively analyze how GO changes that should be carefully considered by all users of semantic similarities. GO changes in particular have a big impact on information content (IC) of GO terms. Since many semantic similarities rely on calculation of IC it is obvious that the study of these changes should be deeply investigated. Here we consider GO versions from 2005 to 2014 and we calculate IC of all GO Terms considering five different formulation. Then we compare these results. Analysis confirm that there exists a statistically significant difference among different calculation

preprint2016arXiv

A computational investigation of the relationships between single-neuron and network dynamics in the cerebral cortex

Functions of brain areas in complex animals are believed to rely on the dynamics of networks of neurons rather than on single neurons. On the other hand, the network dynamics reflect and arise from the integration and coordination of the activity of populations of single neurons. Understanding how single-neurons and neural-circuits dynamics complement each other to produce brain functions is thus of paramount importance. LFPs and EEGs are good indicators of the dynamics of mesoscopic and macroscopic populations of neurons, while microscopic-level activities can be documented by measuring the membrane potential, the synaptic currents or the spiking activity of individual neurons. In this thesis we develop mathematical modelling and mathematical analysis tools that can help the interpretation of joint measures of neural activity at microscopic and mesoscopic or macroscopic scales. In particular, we develop network models of recurrent cortical circuits that can clarify the impact of several aspects of single-neuron (i.e., microscopic-level) dynamics on the activity of the whole neural population (as measured by LFP). We then develop statistical tools to characterize the relationship b

preprint2016arXiv

Unsupervised cryo-EM data clustering through adaptively constrained K-means algorithm

In single-particle cryo-electron microscopy (cryo-EM), K-means clustering algorithm is widely used in unsupervised 2D classification of projection images of biological macromolecules. 3D ab initio reconstruction requires accurate unsupervised classification in order to separate molecular projections of distinct orientations. Due to background noise in single-particle images and uncertainty of molecular orientations, traditional K-means clustering algorithm may classify images into wrong classes and produce classes with a large variation in membership. Overcoming these limitations requires further development on clustering algorithms for cryo-EM data analysis. We propose a novel unsupervised data clustering method building upon the traditional K-means algorithm. By introducing an adaptive constraint term in the objective function, our algorithm not only avoids a large variation in class sizes but also produces more accurate data clustering. Applications of this approach to both simulated and experimental cryo-EM data demonstrate that our algorithm is a significantly improved alterative to the traditional K-means algorithm in single-particle cryo-EM analysis.

preprint2016arXiv

Learning Weighted Association Rules in Human Phenotype Ontology

The Human Phenotype Ontology (HPO) is a structured repository of concepts (HPO Terms) that are associated to one or more diseases. The process of association is referred to as annotation. The relevance and the specificity of both HPO terms and annotations are evaluated by a measure defined as Information Content (IC). The analysis of annotated data is thus an important challenge for bioinformatics. There exist different approaches of analysis. From those, the use of Association Rules (AR) may provide useful knowledge, and it has been used in some applications, e.g. improving the quality of annotations. Nevertheless classical association rules algorithms do not take into account the source of annotation nor the importance yielding to the generation of candidate rules with low IC. This paper presents HPO-Miner (Human Phenotype Ontology-based Weighted Association Rules) a methodology for extracting Weighted Association Rules. HPO-Miner can extract relevant rules from a biological point of view. A case study on using of HPO-Miner on publicly available HPO annotation datasets is used to demonstrate the effectiveness of our methodology.

preprint2017arXiv

Challenges ahead Electron Microscopy for Structural Biology from the Image Processing point of view

Since the introduction of Direct Electron Detectors (DEDs), the resolution and range of macromolecules amenable to this technique has significantly widened, generating a broad interest that explains the well over a dozen reviews in top journal in the last two years. Similarly, the number of job offers to lead EM groups and/or coordinate EM facilities has exploded, and FEI (the main microscope manufacturer for Life Sciences) has received more than 100 orders of high-end electron microscopes by summer 2016. Strategic corporate movements are also happening, with very big players entering the market through key acquisitions (Thermo Fisher has recently bought FEI for \$4.2B), partly attracted by new Pharma interest in the field, now perceived to be in a position to impact structure-based drug design. The scientific perspectives are indeed extremely positive but, in these moments of well-founded generalized optimists, we want to make a reflection on some of the hurdles ahead us, since they certainly exist and they indeed limit the informational content of cryoEM projects. Here we focus on image processing aspects, particularly in the so-called area of Single Particle Analysis, discussing

preprint2016arXiv

Inference of Causal Information Flow in Collective Animal Behavior

Understanding and even defining what constitutes animal interactions remains a challenging problem. Correlational tools may be inappropriate for detecting communication between a set of many agents exhibiting nonlinear behavior. A different approach is to define coordinated motions in terms of an information theoretic channel of direct causal information flow. In this work, we consider time series data obtained by an experimental protocol of optical tracking of the insect species Chironomus riparius. The data constitute reconstructed 3-D spatial trajectories of the insects' flight trajectories and kinematics. We present an application of the optimal causation entropy (oCSE) principle to identify direct causal relationships or information channels among the insects. The collection of channels inferred by oCSE describes a network of information flow within the swarm. We find that information channels with a long spatial range are more common than expected under the assumption that causal information flows should be spatially localized. The tools developed herein are general and applicable to the inference and study of intercommunication networks in a wide variety of natural setti

preprint2016arXiv

Bounds on stationary moments in stochastic chemical kinetics

In the stochastic formulation of chemical kinetics, the stationary moments of the population count of species can be described via a set of linear equations. However, except for some specific cases such as systems with linear reaction propensities, the moment equations are underdetermined as a lower order moment might depend upon a higher order moment. Here, we propose a method to find lower, and upper bounds on stationary moments of molecular counts in a chemical reaction system. The method exploits the fact that statistical moments of any positive-valued random variable must satisfy some constraints. Such constraints can be expressed as nonlinear inequalities on moments in terms of their lower order moments, and solving them in conjugation with the stationary moment equations results in bounds on the moments. Using two examples of biochemical systems, we illustrate that not only one obtains upper and lower bounds on a given stationary moment, but these bounds also improve as one uses more moment equations and utilizes the inequalities for the corresponding higher order moments. Our results provide avenues for development of moment approximations that provide explicit bounds on mo

preprint2016arXiv

Data-Driven Forecast of Dengue Outbreaks in Brazil: A Critical Assessment of Climate Conditions for Different Capitals

Local climate conditions play a major role in the development of the mosquito population responsible for transmitting Dengue Fever. Since the {\em Aedes Aegypti} mosquito is also a primary vector for the recent Zika and Chikungunya epidemics across the Americas, a detailed monitoring of periods with favorable climate conditions for mosquito profusion may improve the timing of vector-control efforts and other urgent public health strategies. We apply dimensionality reduction techniques and machine-learning algorithms to climate time series data and analyze their connection to the occurrence of Dengue outbreaks for seven major cities in Brazil. Specifically, we have identified two key variables and a period during the annual cycle that are highly predictive of epidemic outbreaks. The key variables are the frequency of precipitation and temperature during an approximately two month window of the winter season preceding the outbreak. Thus simple climate signatures may be influencing Dengue outbreaks even months before their occurrence. Some of the more challenging datasets required usage of compressive-sensing procedures to estimate missing entries for temperature and precipitation rec

preprint2016arXiv

Drug response prediction by inferring pathway-response associations with Kernelized Bayesian Matrix Factorization

A key goal of computational personalized medicine is to systematically utilize genomic and other molecular features of samples to predict drug responses for a previously unseen sample. Such predictions are valuable for developing hypotheses for selecting therapies tailored for individual patients. This is especially valuable in oncology, where molecular and genetic heterogeneity of the cells has a major impact on the response. However, the prediction task is extremely challenging, raising the need for methods that can effectively model and predict drug responses. In this study, we propose a novel formulation of multi-task matrix factorization that allows selective data integration for predicting drug responses. To solve the modeling task, we extend the state-of-the-art kernelized Bayesian matrix factorization (KBMF) method with component-wise multiple kernel learning. In addition, our approach exploits the known pathway information in a novel and biologically meaningful fashion to learn the drug response associations. Our method quantitatively outperforms the state of the art on predicting drug responses in two publicly available cancer data sets as well as on a synthetic data set.

preprint2016arXiv

Bone fusion in normal and pathological development is constrained by the network architecture of the human skull

The premature fusion of cranial bones, craniosynostosis, affects the correct development of the skull producing morphological malformations in newborns. To assess the susceptibility of each craniofacial articulation to close prematurely, we used a network model of the skull to quantify the link reliability (an index based on stochastic block modeling and Bayesian inference) of each articulation. We show that, of the 93 human skull articulations at birth, the few articulations that are associated with nonsyndromic craniosynostosis conditions have statistically significant lower reliability scores than the others. In a similar way, articulations that close during the normal postnatal development of the skull have also lower reliability scores than those articulations that persist through adult live. These results indicate a relationship between the architecture of the skull network and the specific articulations that close during normal development and in pathological conditions. Our findings suggest that the topological arrangement of skull bones might act as an epigenetic factor, predisposing some articulations to closure, both in normal and pathological development, and also affec

preprint2016arXiv

A Method for Massively Parallel Analysis of Time Series

Quantification of system-wide perturbations from time series -omic data (i.e. a large number of variables with multiple measures in time) provides the basis for many downstream hypothesis generating tools. Here we propose a method, Massively Parallel Analysis of Time Series (MPATS) that can be applied to quantify transcriptome-wide perturbations. The proposed method characterizes each individual time series through its $\ell_1$ distance to every other time series. Application of MPATS to compare biological conditions produces a ranked list of time series based on their magnitude of differences in their $\ell_1$ representation, which then can be further interpreted through enrichment analysis. The performance of MPATS was validated through its application to a study of IFN$α$ dendritic cell responses to viral and bacterial infection. In conjunction with Gene Set Enrichment Analysis (GSEA), MPATS produced consistently identified signature gene sets of anti-bacterial and anti-viral response. Traditional methods such as EDGE and GSEA Time Series (GSEA-TS) failed to identify the relevant signature gene sets. Furthermore, the results of MPATS highlighted the crucial functional difference

preprint2016arXiv

Rate-Equation Modelling and Ensemble Approach to Extraction of Parameters for Viral Infection-Induced Cell Apoptosis and Necrosis

We develop a theoretical approach that uses physiochemical kinetics modelling to describe cell population dynamics upon progression of viral infection in cell culture, which results in cell apoptosis (programmed cell death) and necrosis (direct cell death). Several model parameters necessary for computer simulation were determined by reviewing and analyzing available published experimental data. By comparing experimental data to computer modelling results, we identify the parameters that are the most sensitive to the measured system properties and allow for the best data fitting. Our model allows extraction of parameters from experimental data and also has predictive power. Using the model we describe interesting time-dependent quantities that were not directly measured in the experiment, and identify correlations among the fitted parameter values. Numerical simulation of viral infection progression is done by a rate-equation approach resulting in a system of "stiff" equations, which are solved by using a novel variant of the stochastic ensemble modelling approach. The latter was originally developed for coupled chemical reactions.

preprint2016arXiv

Functional Hypergraph Uncovers Novel Covariant Structures over Neurodevelopment

Brain development during adolescence is marked by substantial changes in brain structure and function, leading to a stable network topology in adulthood. However, most prior work has examined the data through the lens of brain areas connected to one another in large-scale functional networks. Here, we apply a recently-developed hypergraph approach that treats network connections (edges) rather than brain regions as the unit of interest, allowing us to describe functional network topology from a fundamentally different perspective. Capitalizing on a sample of 780 youth imaged as part of the Philadelphia Neurodevelopmental Cohort, this hypergraph representation of resting-state functional MRI data reveals three distinct classes of sub-networks (hyperedges): clusters, bridges, and stars, which represent spatially distributed, bipartite, and focal architectures, respectively. Cluster hyperedges show a strong resemblance to the functional modules of the brain including somatomotor, visual, default mode, and salience systems. In contrast, star hyperedges represent highly localized subnetworks centered on a small set of regions, and are distributed across the entire cortex. Finally, bridg

preprint2016arXiv

BaTFLED: Bayesian Tensor Factorization Linked to External Data

The vast majority of current machine learning algorithms are designed to predict single responses or a vector of responses, yet many types of response are more naturally organized as matrices or higher-order tensor objects where characteristics are shared across modes. We present a new machine learning algorithm BaTFLED (Bayesian Tensor Factorization Linked to External Data) that predicts values in a three-dimensional response tensor using input features for each of the dimensions. BaTFLED uses a probabilistic Bayesian framework to learn projection matrices mapping input features for each mode into latent representations that multiply to form the response tensor. By utilizing a Tucker decomposition, the model can capture weights for interactions between latent factors for each mode in a small core tensor. Priors that encourage sparsity in the projection matrices and core tensor allow for feature selection and model regularization. This method is shown to far outperform elastic net and neural net models on 'cold start' tasks from data simulated in a three-mode structure. Additionally, we apply the model to predict dose-response curves in a panel of breast cancer cell lines t

preprint2016arXiv

Spatial probabilistic pulsatility model for enhancing photoplethysmographic imaging systems

Photolethysmographic imaging (PPGI) is a widefield non-contact biophotonic technology able to remotely monitor cardiovascular function over anatomical areas. Though spatial context can provide increased physiological insight, existing PPGI systems rely on coarse spatial averaging with no anatomical priors for assessing arterial pulsatility. Here, we developed a continuous probabilistic pulsatility model for importance-weighted blood pulse waveform extraction. Using a data-driven approach, the model was constructed using a 23 participant sample with large demographic variation (11/12 female/male, age 11-60 years, BMI 16.4-35.1 kg$\cdot$m$^{-2}$). Using time-synchronized ground-truth waveforms, spatial correlation priors were computed and projected into a co-aligned importance-weighted Cartesian space. A modified Parzen-Rosenblatt kernel density estimation method was used to compute the continuous resolution-agnostic probabilistic pulsatility model. The model identified locations that consistently exhibited pulsatility across the sample. Blood pulse waveform signals extracted with the model exhibited significantly stronger temporal correlation ($W=35,p<0.01$) and spectral SNR ($W=31,

preprint2016arXiv

Embedding of biological distribution networks with differing environmental constraints

Distribution networks -- from vasculature to urban transportation systems -- are prevalent in both the natural and consumer worlds. These systems are intrinsically physical in composition and are embedded into real space, properties that lead to constraints on their topological organization. In this study, we compare and contrast two types of biological distribution networks: mycelial fungi and the vasculature system on the surface of rodent brains. Both systems are alike in that they must route resources efficiently, but they are also inherently distinct in terms of their growth mechanisms, and in that fungi are not attached to a larger organism and must often function in unregulated and varied environments. We begin by uncovering a common organizational principle -- Rentian scaling -- that manifests as hierarchical network layout in both physical and topological space. Simulated models of distribution networks optimized for transport in the presence of fluctuations are also shown to exhibit this feature in their embedding, with similar scaling exponents. However, we also find clear differences in how the fungi and vasculature balance tradeoffs in material cost, efficiency, and ro

preprint2016arXiv

Local equilibrium in bird flocks

The correlated motion of flocks is an instance of global order emerging from local interactions. An essential difference with analogous ferromagnetic systems is that flocks are active: animals move relative to each other, dynamically rearranging their interaction network. The effect of this off-equilibrium element is well studied theoretically, but its impact on actual biological groups deserves more experimental attention. Here, we introduce a novel dynamical inference technique, based on the principle of maximum entropy, which accodomates network rearrangements and overcomes the problem of slow experimental sampling rates. We use this method to infer the strength and range of alignment forces from data of starling flocks. We find that local bird alignment happens on a much faster timescale than neighbour rearrangement. Accordingly, equilibrium inference, which assumes a fixed interaction network, gives results consistent with dynamical inference. We conclude that bird orientations are in a state of local quasi-equilibrium over the interaction length scale, providing firm ground for the applicability of statistical physics in certain active systems.

preprint2016arXiv

A Hidden Markov Movement Model for rapidly identifying behavioral states from animal tracks

1. Electronic telemetry is frequently used to document animal movement through time. Methods that can identify underlying behaviors driving specific movement patterns can help us understand how and why animals use available space, thereby aiding conservation and management efforts. For aquatic animal tracking data with significant measurement error, a Bayesian state-space model called the first-Difference Correlated Random Walk with Switching (DCRWS) has often been used for this purpose. However, for aquatic animals, highly accurate tracking data of animal movement are now becoming more common. 2. We developed a new Hidden Markov Model (HMM) for identifying behavioral states from animal tracks with negligible error, which we called the Hidden Markov Movement Model (HMMM). We implemented as the basis for the HMMM the process equation of the DCRWS, but we used the method of maximum likelihood and the R package TMB for rapid model fitting. 3. We compared the HMMM to a modified version of the DCRWS for highly accurate tracks, the DCRWSnome, and to a common HMM for animal tracks fitted with the R package moveHMM. We show that the HMMM is both accurate and suitable for multiple species b

preprint2016arXiv

Partially blind domain adaptation for age prediction from DNA methylation data

Over the last years, huge resources of biological and medical data have become available for research. This data offers great chances for machine learning applications in health care, e.g. for precision medicine, but is also challenging to analyze. Typical challenges include a large number of possibly correlated features and heterogeneity in the data. One flourishing field of biological research in which this is relevant is epigenetics. Here, especially large amounts of DNA methylation data have emerged. This epigenetic mark has been used to predict a donor's 'epigenetic age' and increased epigenetic aging has been linked to lifestyle and disease history. In this paper we propose an adaptive model which performs feature selection for each test sample individually based on the distribution of the input data. The method can be seen as partially blind domain adaptation. We apply the model to the problem of age prediction based on DNA methylation data from a variety of tissues, and compare it to a standard model, which does not take heterogeneity into account. The standard approach has particularly bad performance on one tissue type on which we show substantial improvement

preprint2016arXiv

Systems-level approach to uncovering diffusive states and their transitions from single particle trajectories

The stochastic motions of a diffusing particle contain information concerning the particle's interactions with binding partners and with its local environment. However, accurate determination of the underlying diffusive properties, beyond normal diffusion, has remained challenging when analyzing particle trajectories on an individual basis. Here, we introduce the maximum likelihood estimator (MLE) for confined diffusion and fractional Brownian motion. We demonstrate that this MLE yields improved estimation over traditional mean square displacement analyses. We also introduce a model selection scheme (that we call mleBIC) that classifies individual trajectories to a given diffusion mode. We demonstrate the statistical limitations of classification via mleBIC using simulated data. To overcome these limitations, we introduce a new version of perturbation expectation-maximization (pEMv2), which simultaneously analyzes a collection of particle trajectories to uncover the system of interactions which give rise to unique normal and/or non-normal diffusive states within the population. We test and evaluate the performance of pEMv2 on various sets of simulated particle trajectories, whi

preprint2016arXiv

RIDS: Robust Identification of Sparse Gene Regulatory Networks from Perturbation Experiments

Reconstructing the causal network in a complex dynamical system plays a crucial role in many applications, from sub-cellular biology to economic systems. Here we focus on inferring gene regulation networks (GRNs) from perturbation or gene deletion experiments. Despite their scientific merit, such perturbation experiments are not often used for such inference due to their costly experimental procedure, requiring significant resources to complete the measurement of every single experiment. To overcome this challenge, we develop the Robust IDentification of Sparse networks (RIDS) method that reconstructs the GRN from a small number of perturbation experiments. Our method uses the gene expression data observed in each experiment and translates that into a steady state condition of the system's nonlinear interaction dynamics. Applying a sparse optimization criterion, we are able to extract the parameters of the underlying weighted network, even from very few experiments. In fact, we demonstrate analytically that, under certain conditions, the GRN can be perfectly reconstructed using $K = Ω(d_{max})$ perturbation experiments, where $d_{max}$ is the maximum in-degree of the GRN, a sma

preprint2016arXiv

Effects of cell cycle noise on excitable gene circuits

We assess the impact of cell cycle noise on gene circuit dynamics. For bistable genetic switches and excitable circuits, we find that transitions between metastable states most likely occur just after cell division and that this concentration effect intensifies in the presence of transcriptional delay. We explain this concentration effect with a 3-states stochastic model. For genetic oscillators, we quantify the temporal correlations between daughter cells induced by cell division. Temporal correlations must be captured properly in order to accurately quantify noise sources within gene networks.

preprint2016arXiv

Generation and precise control of dynamic biochemical gradients for cellular assays

Spatial gradients of diffusible signalling molecules play crucial roles in controlling diverse cellular behaviour such as cell differentiation, tissue patterning and chemotaxis. In this paper, we report the design and testing of a microfluidic device for diffusion-based gradient generation for cellular assays. A unique channel design of the device eliminates cross-flow between the source and sink channels, thereby stabilising gradients by passive diffusion. The platform also enables quick and flexible control of chemical concentration that makes highly dynamic gradients in diffusion chambers. A model with the first approximation of diffusion and surface adsorption of molecules recapitulates the experimentally observed gradients. Budding yeast cells cultured in a gradient of a chemical inducer expressed a reporter fluorescence protein in a concentration-dependent manner. This microfluidic platform serves as a versatile prototype applicable to a broad range of biomedical investigations.

People in this topic

12 visible researcher(s)