Source author record

Christian Bongiorno

Christian Bongiorno appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Applications Computation Genomics math.SP math.ST Methodology physics.data-an q-fin.RM q-fin.ST q-fin.TR Statistics Theory

Catalog footprint

What is connected

3works

11topics

3close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2022arXiv

Statistical inference of lead-lag at various timescales between asynchronous time series from p-values of transfer entropy

Symbolic transfer entropy is a powerful non-parametric tool to detect lead-lag between time series. Because a closed expression of the distribution of Transfer Entropy is not known for finite-size samples, statistical testing is often performed with bootstraps whose slowness prevents the inference of large lead-lag networks between long time series. On the other hand, the asymptotic distribution of Transfer Entropy between two time series is known. In this work, we derive the asymptotic distribution of the test for one time series having a larger Transfer Entropy than another one on a target time series. We then measure the convergence speed of both tests in the small sample size limits via benchmarks. We then introduce Transfer Entropy between time-shifted time series, which allows to measure the timescale at which information transfer is maximal and vanishes. We finally apply these methods to tick-by-tick price changes of several hundreds of stocks, yielding non-trivial statistically validated networks.

preprint2020arXiv

Bootstraps Regularize Singular Correlation Matrices

I show analytically that the average of $k$ bootstrapped correlation matrices rapidly becomes positive-definite as $k$ increases, which provides a simple approach to regularize singular Pearson correlation matrices. If $n$ is the number of objects and $t$ the number of features, the averaged correlation matrix is almost surely positive-definite if $k> \frac{e}{e-1}\frac{n}{t}\simeq 1.58\frac{n}{t}$ in the limit of large $t$ and $n$. The probability of obtaining a positive-definite correlation matrix with $k$ bootstraps is also derived for finite $n$ and $t$. Finally, I demonstrate that the number of required bootstraps is always smaller than $n$. This method is particularly relevant in fields where $n$ is orders of magnitude larger than the size of data points $t$, e.g., in finance, genetics, social science, or image processing.

preprint2019arXiv

Nested partitions from hierarchical clustering statistical validation

We develop a greedy algorithm that is fast and scalable in the detection of a nested partition extracted from a dendrogram obtained from hierarchical clustering of a multivariate series. Our algorithm provides a $p$-value for each clade observed in the hierarchical tree. The $p$-value is obtained by computing a number of bootstrap replicas of the dissimilarity matrix and by performing a statistical test on each difference between the dissimilarity associated with a given clade and the dissimilarity of the clade of its parent node. We prove the efficacy of our algorithm with a set of benchmarks generated by using a hierarchical factor model. We compare the results obtained by our algorithm with those of Pvclust. Pvclust is a widely used algorithm developed with a global approach originally motivated by phylogenetic studies. In our numerical experiments we focus on the role of multiple hypothesis test correction and on the robustness of the algorithms to inaccuracy and errors of datasets. We also apply our algorithm to a reference empirical dataset. We verify that our algorithm is much faster than Pvclust algorithm and has a better scalability both in the number of elements and in the number of records of the investigated multivariate set. Our algorithm provides a hierarchically nested partition in much shorter time than currently widely used algorithms allowing to perform a statistically validated cluster analysis detection in very large systems.