Source author record

Ernst Wit

Ernst Wit appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

12works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

12 published item(s)

preprint2022arXiv

Generative Models for Periodicity Detection in Noisy Signals

We introduce a new periodicity detection algorithm for binary time series of event onsets, the Gaussian Mixture Periodicity Detection Algorithm (GMPDA). The algorithm approaches the periodicity detection problem to infer the parameters of a generative model. We specified two models - the Clock and Random Walk - which describe two different periodic phenomena and provide a generative framework. The algorithm achieved strong results on test cases for single and multiple periodicity detection and varying noise levels. The performance of GMPDA was also evaluated on real data, recorded leg movements during sleep, where GMPDA was able to identify the expected periodicities despite high noise levels. The paper's key contributions are two new models for generating periodic event behavior and the GMPDA algorithm for multiple periodicity detection, which is highly accurate under noise.

preprint2020arXiv

Model-based clustering for populations of networks

Until recently obtaining data on populations of networks was typically rare. However, with the advancement of automatic monitoring devices and the growing social and scientific interest in networks, such data has become more widely available. From sociological experiments involving cognitive social structures to fMRI scans revealing large-scale brain networks of groups of patients, there is a growing awareness that we urgently need tools to analyse populations of networks and particularly to model the variation between networks due to covariates. We propose a model-based clustering method based on mixtures of generalized linear (mixed) models that can be employed to describe the joint distribution of a populations of networks in a parsimonious manner and to identify subpopulations of networks that share certain topological properties of interest (degree distribution, community structure, effect of covariates on the presence of an edge, etc.). Maximum likelihood estimation for the proposed model can be efficiently carried out with an implementation of the EM algorithm. We assess the performance of this method on simulated data and conclude with an example application on advice networks in a small business.

preprint2016arXiv

Estimating Causal Effects From Nonparanormal Observational Data

One of the basic aims in science is to unravel the chain of cause and effect of particular systems. Especially for large systems this can be a daunting task. Detailed interventional and randomized data sampling approaches can be used to resolve the causality question, but for many systems such interventions are impossible or too costly to obtain. Recently, Maathuis et al. (2010), following ideas from Spirtes et al. (2000), introduced a framework to estimate causal effects in large scale Gaussian systems. By describing the causal network as a directed acyclic graph it is a possible to estimate a class of Markov equivalent systems that describe the underlying causal interactions consistently, even for non-Gaussian systems. In these systems, causal effects stop being linear and cannot be described any more by a single coefficient. In this paper, we derive the general functional form of causal effect in a large subclass of non-Gaussian distributions, called the non- paranormal. We also derive a convenient approximation, which can be used effectively in estimation. We apply the method to an observational gene expression dataset.

preprint2016arXiv

Mixture model with multiple allocations for clustering spatially correlated observations in the analysis of ChIP-Seq data

Model-based clustering is a technique widely used to group a collection of units into mutually exclusive groups. There are, however, situations in which an observation could in principle belong to more than one cluster. In the context of Next-Generation Sequencing (NGS) experiments, for example, the signal observed in the data might be produced by two (or more) different biological processes operating together and a gene could participate in both (or all) of them. We propose a novel approach to cluster NGS discrete data, coming from a ChIP-Seq experiment, with a mixture model, allowing each unit to belong potentially to more than one group: these multiple allocation clusters can be flexibly defined via a function combining the features of the original groups without introducing new parameters. The formulation naturally gives rise to a `zero-inflation group' in which values close to zero can be allocated, acting as a correction for the abundance of zeros that manifest in this type of data. We take into account the spatial dependency between observations, which is described through a latent Conditional Auto-Regressive process that can reflect different dependency patterns. We assess the performance of our model within a simulation environment and then we apply it to ChIP-seq real data.

preprint2014arXiv

A computationally fast alternative to cross-validation in penalized Gaussian graphical models

We study the problem of selection of regularization parameter in penalized Gaussian graphical models. When the goal is to obtain the model with good predicting power, cross validation is the gold standard. We present a new estimator of Kullback-Leibler loss in Gaussian Graphical model which provides a computationally fast alternative to cross-validation. The estimator is obtained by approximating leave-one-out-cross validation. Our approach is demonstrated on simulated data sets for various types of graphs. The proposed formula exhibits superior performance, especially in the typical small sample size scenario, compared to other available alternatives to cross validation, such as Akaike's information criterion and Generalized approximate cross validation. We also show that the estimator can be used to improve the performance of the BIC when the sample size is small.

preprint2014arXiv

Generalized information criterion for model selection in penalized graphical models

This paper introduces an estimator of the relative directed distance between an estimated model and the true model, based on the Kulback-Leibler divergence and is motivated by the generalized information criterion proposed by Konishi and Kitagawa. This estimator can be used to select model in penalized Gaussian copula graphical models. The use of this estimator is not feasible for high-dimensional cases. However, we derive an efficient way to compute this estimator which is feasible for the latter class of problems. Moreover, this estimator is, generally, appropriate for several penalties such as lasso, adaptive lasso and smoothly clipped absolute deviation penalty. Simulations show that the method performs similarly to KL oracle estimator and it also improves BIC performance in terms of support recovery of the graph. Specifically, we compare our method with Akaike information criterion, Bayesian information criterion and cross validation for band, sparse and dense network structures.

preprint2014arXiv

Penalized EM algorithm and copula skeptic graphical models for inferring networks for mixed variables

In this article, we consider the problem of reconstructing networks for continuous, binary, count and discrete ordinal variables by estimating sparse precision matrix in Gaussian copula graphical models. We propose two approaches: $\ell_1$ penalized extended rank likelihood with Monte Carlo Expectation-Maximization algorithm (copula EM glasso) and copula skeptic with pair-wise copula estimation for copula Gaussian graphical models. The proposed approaches help to infer networks arising from nonnormal and mixed variables. We demonstrate the performance of our methods through simulation studies and analysis of breast cancer genomic and clinical data and maize genetics data.

preprint2014arXiv

Reproducing kernel Hilbert space based estimation of systems of ordinary differential equations

Non-linear systems of differential equations have attracted the interest in fields like system biology, ecology or biochemistry, due to their flexibility and their ability to describe dynamical systems. Despite the importance of such models in many branches of science they have not been the focus of systematic statistical analysis until recently. In this work we propose a general approach to estimate the parameters of systems of differential equations measured with noise. Our methodology is based on the maximization of the penalized likelihood where the system of differential equations is used as a penalty. To do so, we use a Reproducing Kernel Hilbert Space approach that allows to formulate the estimation problem as an unconstrained numeric maximization problem easy to solve. The proposed method is tested with synthetically simulated data and it is used to estimate the unobserved transcription factor CdaR in Steptomyes coelicolor using gene expression data of the genes it regulates.

preprint2013arXiv

High dimensional Sparse Gaussian Graphical Mixture Model

This paper considers the problem of networks reconstruction from heterogeneous data using a Gaussian Graphical Mixture Model (GGMM). It is well known that parameter estimation in this context is challenging due to large numbers of variables coupled with the degeneracy of the likelihood. We propose as a solution a penalized maximum likelihood technique by imposing an $l_{1}$ penalty on the precision matrix. Our approach shrinks the parameters thereby resulting in better identifiability and variable selection. We use the Expectation Maximization (EM) algorithm which involves the graphical LASSO to estimate the mixing coefficients and the precision matrices. We show that under certain regularity conditions the Penalized Maximum Likelihood (PML) estimates are consistent. We demonstrate the performance of the PML estimator through simulations and we show the utility of our method for high dimensional data analysis in a genomic application.

preprint2013arXiv

Network estimation in State Space Model with L1-regularization constraint

Biological networks have arisen as an attractive paradigm of genomic science ever since the introduction of large scale genomic technologies which carried the promise of elucidating the relationship in functional genomics. Microarray technologies coupled with appropriate mathematical or statistical models have made it possible to identify dynamic regulatory networks or to measure time course of the expression level of many genes simultaneously. However one of the few limitations fall on the high-dimensional nature of such data coupled with the fact that these gene expression data are known to include some hidden process. In that regards, we are concerned with deriving a method for inferring a sparse dynamic network in a high dimensional data setting. We assume that the observations are noisy measurements of gene expression in the form of mRNAs, whose dynamics can be described by some unknown or hidden process. We build an input-dependent linear state space model from these hidden states and demonstrate how an incorporated $L_{1}$ regularization constraint in an Expectation-Maximization (EM) algorithm can be used to reverse engineer transcriptional networks from gene expression profiling data. This corresponds to estimating the model interaction parameters. The proposed method is illustrated on time-course microarray data obtained from a well established T-cell data. At the optimum tuning parameters we found genes TRAF5, JUND, CDK4, CASP4, CD69, and C3X1 to have higher number of inwards directed connections and FYB, CCNA2, AKT1 and CASP8 to be genes with higher number of outwards directed connections. We recommend these genes to be object for further investigation. Caspase 4 is also found to activate the expression of JunD which in turn represses the cell cycle regulator CDC2.

preprint2013arXiv

State-space modeling of dynamic genetic networks

The genomic reality is a highly complex and dynamic system. The recent development of high-throughput technologies has enabled researchers to measure the abundance of many genes (in the order of thousands) simultaneously. The challenge is to unravel from such measurements, gene/protein or gene/gene or protein/ protein interactions and key biological features of cellular systems. Our goal is to devise a method for inferring transcriptional or gene regulatory networks from high-throughput data sources such as gene expression microarrays with potentially hidden states, such as unmeasured transcription factors (TFs), which satisfies certain Markov properties. We propose a dynamic state space representation. Our method is based on an EM algorithm with an incorporated Kalman smoothing algorithm in the E-step, a bootstrap for confidence intervals to infer the networks and the AIC for model selection. The state space model is an approach with proven effectiveness to reverse engineer transcriptional networks. The proposed method is applied to time course microarray data obtained from well established T-cell. When we applied the method to the T-cell data, we obtained 4, as the optimum number of hidden states. Our results support interesting biological properties in the family of Jun genes. The following genes were mostly seen as regulatory genes. These genes includes FYB, CCNA2, AKT1, TRAF5, CASP4, and CTNNB1. We found interaction between Jun-B and SMN1, and CDC2 activates Jun-D. We found few significant interactions or one-to-one correspondence among the 4 putative transcription factors. Among the important key genes in terms of outward-directed edges, we found genes such as CCNA2, JUNB, CDC2, CASP4, JUND to have a high degree of connectivity. R Computer source code is made available at our website at http://www.math.rug.nl/stat/Main/Software.

preprint2012arXiv

Modelling slowly changing dynamic gene-regulatory networks

Dynamic gene-regulatory networks are complex since the number of potential components involved in the system is very large. Estimating dynamic networks is an important task because they compromise valuable information about interactions among genes. Graphical models are a powerful class of models to estimate conditional independence among random variables, e.g. interactions in dynamic systems. Indeed, these interactions tend to vary over time. However, the literature has been focused on static networks, which can only reveal overall structures. Time-course experiments are performed in order to tease out significant changes in networks. It is typically reasonable to assume that changes in genomic networks are few because systems in biology tend to be stable. We introduce a new model for estimating slowly changes in dynamic gene-regulatory networks which is suitable for a high-dimensional dataset, e.g. time-course genomic data. Our method is based on i) the penalized likelihood with $\ell_1$-norm, ii) the penalized differences between conditional independence elements across time points and iii) the heuristic search strategy to find optimal smoothing parameters. We implement a set of linear constraints necessary to estimate sparse graphs and penalized changing in dynamic networks. These constraints are not in the linear form. For this reason, we introduce slack variables to re-write our problem into a standard convex optimization problem subject to equality linear constraints. We show that GL$_Δ$ performs well in a simulation study. Finally, we apply the proposed model to a time-course genetic dataset T-cell.