Source author record

Dimitris Kugiumtzis

Dimitris Kugiumtzis appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

13works
14topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

13 published item(s)

preprint2019arXiv

Evaluation of Granger causality measures for constructing networks from multivariate time series

Granger causality and variants of this concept allow the study of complex dynamical systems as networks constructed from multivariate time series. In this work, a large number of Granger causality measures used to form causality networks from multivariate time series are assessed. These measures are in the time domain, such as model-based and information measures, the frequency domain and the phase domain. The study aims also to compare bivariate and multivariate measures, linear and nonlinear measures, as well as the use of dimension reduction in linear model-based measures and information measures. The latter is particular relevant in the study of high-dimensional time series. For the performance of the multivariate causality measures, low and high dimensional coupled dynamical systems are considered in discrete and continuous time, as well as deterministic and stochastic. The measures are evaluated and ranked according to their ability to provide causality networks that match the original coupling structure. The simulation study concludes that the Granger causality measures using dimension reduction are superior and should be preferred particularly in studies involving many observed variables, such as multi-channel electroencephalograms and financial markets.

preprint2015arXiv

Granger Causality in Multi-variate Time Series using a Time Ordered Restricted Vector Autoregressive Model

Granger causality has been used for the investigation of the inter-dependence structure of the underlying systems of multi-variate time series. In particular, the direct causal effects are commonly estimated by the conditional Granger causality index (CGCI). In the presence of many observed variables and relatively short time series, CGCI may fail because it is based on vector autoregressive models (VAR) involving a large number of coefficients to be estimated. In this work, the VAR is restricted by a scheme that modifies the recently developed method of backward-in-time selection (BTS) of the lagged variables and the CGCI is combined with BTS. Further, the proposed approach is compared favorably to other restricted VAR representations, such as the top-down strategy, the bottom-up strategy, and the least absolute shrinkage and selection operator (LASSO), in terms of sensitivity and specificity of CGCI. This is shown by using simulations of linear and nonlinear, low and high-dimensional systems and different time series lengths. For nonlinear systems, CGCI from the restricted VAR representations are compared with analogous nonlinear causality indices. Further, CGCI in conjunction with BTS and other restricted VAR representations is applied to multi-channel scalp electroencephalogram (EEG) recordings of epileptic patients containing epileptiform discharges. CGCI on the restricted VAR, and BTS in particular, could track the changes in brain connectivity before, during and after epileptiform discharges, which was not possible using the full VAR representation.

preprint2015arXiv

Markov chain order estimation with parametric significance tests of conditional mutual information

Besides the different approaches suggested in the literature, accurate estimation of the order of a Markov chain from a given symbol sequence is an open issue, especially when the order is moderately large. Here, parametric significance tests of conditional mutual information (CMI) of increasing order $m$, $I_c(m)$, on a symbol sequence are conducted for increasing orders $m$ in order to estimate the true order $L$ of the underlying Markov chain. CMI of order $m$ is the mutual information of two variables in the Markov chain being $m$ time steps apart, conditioning on the intermediate variables of the chain. The null distribution of CMI is approximated with a normal and gamma distribution deriving analytic expressions of their parameters, and a gamma distribution deriving its parameters from the mean and variance of the normal distribution. The accuracy of order estimation is assessed with the three parametric tests, and the parametric tests are compared to the randomization significance test and other known order estimation criteria using Monte Carlo simulations of Markov chains with different order $L$, length of symbol sequence $N$ and number of symbols $K$. The parametric test using the gamma distribution (with directly defined parameters) is consistently better than the other two parametric tests and matches well the performance of the randomization test. The tests are applied to genes and intergenic regions of DNA sequences, and the estimated orders are interpreted in view of the results from the simulation study. The application shows the usefulness of the parametric gamma test for long symbol sequences where the randomization test becomes prohibitively slow to compute.

preprint2014arXiv

Simulation of Multivariate Non-Gaussian Autoregressive Time Series with Given Autocovariance and Marginals

A semi-analytic method is proposed for the generation of realizations of a multivariate process of a given linear correlation structure and marginal distribution. This is an extension of a similar method for univariate processes, transforming the autocorrelation of the non-Gaussian process to that of a Gaussian process based on a piece-wise linear marginal transform from non-Gaussian to Gaussian marginal. The extension to multivariate processes involves the derivation of the autocorrelation matrix from the marginal transforms, which determines the generating vector autoregressive process. The effectiveness of the approach is demonstrated on systems designed under different scenarios of autocovariance and marginals.

preprint2013arXiv

Backward-in-Time Selection of the Order of Dynamic Regression Prediction Model

We investigate the optimal structure of dynamic regression models used in multivariate time series prediction and propose a scheme to form the lagged variable structure called Backward-in-Time Selection (BTS) that takes into account feedback and multi-collinearity, often present in multivariate time series. We compare BTS to other known methods, also in conjunction with regularization techniques used for the estimation of model parameters, namely principal components, partial least squares and ridge regression estimation. The predictive efficiency of the different models is assessed by means of Monte Carlo simulations for different settings of feedback and multi-collinearity. The results show that BTS has consistently good prediction performance while other popular methods have varying and often inferior performance. The prediction performance of BTS was also found the best when tested on human electroencephalograms of an epileptic seizure, and to the prediction of returns of indices of world financial markets.

preprint2013arXiv

Direct coupling information measure from non-uniform embedding

A measure to estimate the direct and directional coupling in multivariate time series is proposed. The measure is an extension of a recently published measure of conditional Mutual Information from Mixed Embedding (MIME) for bivariate time series. In the proposed measure of Partial MIME (PMIME), the embedding is on all observed variables, and it is optimized in explaining the response variable. It is shown that PMIME detects correctly direct coupling, and outperforms the (linear) conditional Granger causality and the partial transfer entropy. We demonstrate that PMIME does not rely on significance test and embedding parameters, and the number of observed variables has no effect on its statistical accuracy, it may only slow the computations. The importance of these points is shown in simulations and in an application to epileptic multi-channel scalp EEG.

preprint2013arXiv

Markov Chain Order estimation with Conditional Mutual Information

We introduce the Conditional Mutual Information (CMI) for the estimation of the Markov chain order. For a Markov chain of $K$ symbols, we define CMI of order $m$, $I_c(m)$, as the mutual information of two variables in the chain being $m$ time steps apart, conditioning on the intermediate variables of the chain. We find approximate analytic significance limits based on the estimation bias of CMI and develop a randomization significance test of $I_c(m)$, where the randomized symbol sequences are formed by random permutation of the components of the original symbol sequence. The significance test is applied for increasing $m$ and the Markov chain order is estimated by the last order for which the null hypothesis is rejected. We present the appropriateness of CMI-testing on Monte Carlo simulations and compare it to the Akaike and Bayesian information criteria, the maximal fluctuation method (Peres-Shields estimator) and a likelihood ratio test for increasing orders using $ϕ$-divergence. The order criterion of CMI-testing turns out to be superior for orders larger than one, but its effectiveness for large orders depends on data availability. In view of the results from the simulations, we interpret the estimated orders by the CMI-testing and the other criteria on genes and intergenic regions of DNA chains.

preprint2013arXiv

Partial Transfer Entropy on Rank Vectors

For the evaluation of information flow in bivariate time series, information measures have been employed, such as the transfer entropy (TE), the symbolic transfer entropy (STE), defined similarly to TE but on the ranks of the components of the reconstructed vectors, and the transfer entropy on rank vectors (TERV), similar to STE but forming the ranks for the future samples of the response system with regard to the current reconstructed vector. Here we extend TERV for multivariate time series, and account for the presence of confounding variables, called partial transfer entropy on ranks (PTERV). We investigate the asymptotic properties of PTERV, and also partial STE (PSTE), construct parametric significance tests under approximations with Gaussian and gamma null distributions, and show that the parametric tests cannot achieve the power of the randomization test using time-shifted surrogates. Using simulations on known coupled dynamical systems and applying parametric and randomization significance tests, we show that PTERV performs better than PSTE but worse than the partial transfer entropy (PTE). However, PTERV, unlike PTE, is robust to the presence of drifts in the time series and it is also not affected by the level of detrending.

preprint2010arXiv

Measures of Analysis of Time Series (MATS): A MATLAB Toolkit for Computation of Multiple Measures on Time Series Data Bases

In many applications, such as physiology and finance, large time series data bases are to be analyzed requiring the computation of linear, nonlinear and other measures. Such measures have been developed and implemented in commercial and freeware softwares rather selectively and independently. The Measures of Analysis of Time Series ({\tt MATS}) {\tt MATLAB} toolkit is designed to handle an arbitrary large set of scalar time series and compute a large variety of measures on them, allowing for the specification of varying measure parameters as well. The variety of options with added facilities for visualization of the results support different settings of time series analysis, such as the detection of dynamics changes in long data records, resampling (surrogate or bootstrap) tests for independence and linearity with various test statistics, and discrimination power of different measures and for different combinations of their parameters. The basic features of {\tt MATS} are presented and the implemented measures are briefly described. The usefulness of {\tt MATS} is illustrated on some empirical examples along with screenshots.

preprint2010arXiv

Non-uniform state space reconstruction and coupling detection

We investigate the state space reconstruction from multiple time series derived from continuous and discrete systems and propose a method for building embedding vectors progressively using information measure criteria regarding past, current and future states. The embedding scheme can be adapted for different purposes, such as mixed modelling, cross-prediction and Granger causality. In particular we apply this method in order to detect and evaluate information transfer in coupled systems. As a practical application, we investigate in records of scalp epileptic EEG the information flow across brain areas.

preprint2010arXiv

Transfer Entropy on Rank Vectors

Transfer entropy (TE) is a popular measure of information flow found to perform consistently well in different settings. Symbolic transfer entropy (STE) is defined similarly to TE but on the ranks of the components of the reconstructed vectors rather than the reconstructed vectors themselves. First, we correct STE by forming the ranks for the future samples of the response system with regard to the current reconstructed vector. We give the grounds for this modified version of STE, which we call Transfer Entropy on Rank Vectors (TERV). Then we propose to use more than one step ahead in the formation of the future of the response in order to capture the information flow from the driving system over a longer time horizon. To assess the performance of STE, TE and TERV in detecting correctly the information flow we use receiver operating characteristic (ROC) curves formed by the measure values in the two coupling directions computed on a number of realizations of known weakly coupled systems. We also consider different settings of state space reconstruction, time series length and observational noise. The results show that TERV indeed improves STE and in some cases performs better than TE, particularly in the presence of noise, but overall TE gives more consistent results. The use of multiple steps ahead improves the accuracy of TE and TERV.

preprint2009arXiv

Evaluation of Mutual Information Estimators for Time Series

We study some of the most commonly used mutual information estimators, based on histograms of fixed or adaptive bin size, $k$-nearest neighbors and kernels, and focus on optimal selection of their free parameters. We examine the consistency of the estimators (convergence to a stable value with the increase of time series length) and the degree of deviation among the estimators. The optimization of parameters is assessed by quantifying the deviation of the estimated mutual information from its true or asymptotic value as a function of the free parameter. Moreover, some common-used criteria for parameter selection are evaluated for each estimator. The comparative study is based on Monte Carlo simulations on time series from several linear and nonlinear systems of different lengths and noise levels. The results show that the $k$-nearest neighbor is the most stable and less affected by the method-specific parameter. A data adaptive criterion for optimal binning is suggested for linear systems but it is found to be rather conservative for nonlinear systems. It turns out that the binning and kernel estimators give the least deviation in identifying the lag of the first minimum of mutual information from nonlinear systems, and are stable in the presence of noise.

preprint1996arXiv

State Space Reconstruction Parameters in the Analysis of Chaotic Time Series - the Role of the Time Window Length

The most common state space reconstruction method in the analysis of chaotic time series is the Method of Delays (MOD). Many techniques have been suggested to estimate the parameters of MOD, i.e. the time delay $τ$ and the embedding dimension $m$. We discuss the applicability of these techniques with a critical view as to their validity, and point out the necessity of determining the overall time window length, $τ_w$, for successful embedding. Emphasis is put on the relation between $τ_w$ and the dynamics of the underlying chaotic system, and we suggest to set $τ_w \geq τ_p$, the mean orbital period; $τ_p$ is approximated from the oscillations of the time series. The procedure is assessed using the correlation dimension for both synthetic and real data. For clean synthetic data, values of $τ_w$ larger than $τ_p$ always give good results given enough data and thus $τ_p$ can be considered as a lower limit ($τ_w \geq τ_p$). For noisy synthetic data and real data, an upper limit is reached for $τ_w$ which approaches $τ_p$ for increasing noise amplitude.