Source author record

Sylvain Delattre

Sylvain Delattre appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

14works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

14 published item(s)

preprint2020arXiv

Rate of Estimation for the Stationary Distribution of Stochastic Damping Hamiltonian Systems with Continuous Observations

We study the problem of the non-parametric estimation for the density $π$ of the stationary distribution of a stochastic two-dimensional damping Hamiltonian system $(Z_t)_{t\in[0,T]}=(X_t,Y_t)_{t \in [0,T]}$. From the continuous observation of the sampling path on $[0,T]$, we study the rate of estimation for $π(x_0,y_0)$ as $T \to \infty$. We show that kernel based estimators can achieve the rate $T^{-v}$ for some explicit exponent $v \in (0,1/2)$. One finding is that the rate of estimation depends on the smoothness of $π$ and is completely different with the rate appearing in the standard i.i.d.\ setting or in the case of two-dimensional non degenerate diffusion processes. Especially, this rate depends also on $y_0$. Moreover, we obtain a minimax lower bound on the $L^2$-risk for pointwise estimation, with the same rate $T^{-v}$, up to $\log(T)$ terms.

preprint2016arXiv

A note on dynamical models on random graphs and Fokker-Planck equations

We address the issue of the proximity of interacting diffusion models on large graphs with a uniform degree property and a corresponding mean field model, i.e. a model on the complete graph with a suitably renormalized interaction parameter. Examples include Erdős-Rényi graphs with edge probability $p_n$, $n$ is the number of vertices, such that $\lim_{n \to \infty}p_n n= \infty$. The purpose of this note it twofold: (1) to establish this proximity on finite time horizon, by exploiting the fact that both systems are accurately described by a Fokker-Planck PDE (or, equivalently, by a nonlinear diffusion process) in the $n=\infty$ limit; (2) to remark that in reality this result is unsatisfactory when it comes to applying it to systems with $N$ large but finite, for example the values of $N$ that can be reached in simulations or that correspond to the typical number of interacting units in a biological system.

preprint2016arXiv

On the Kozachenko-Leonenko entropy estimator

We study in details the bias and variance of the entropy estimator proposed by Kozachenko and Leonenko for a large class of densities on $\mathbb{R}^d$. We then use the work of Bickel and Breiman to prove a central limit theorem in dimensions $1$ and $2$. In higher dimensions, we provide a development of the bias in terms of powers of $N^{-2/d}$. This allows us to use a Richardson extrapolation to build, in any dimension, an estimator satisfying a central limit theorem and for which we can give some some explicit (asymptotic) confidence intervals.

preprint2016arXiv

Statistical inference versus mean field limit for Hawkes processes

We consider a population of $N$ individuals, of which we observe the number of actions as time evolves. For each couple of individuals $(i,j)$, $j$ may or not influence $i$, which we model by i.i.d. Bernoulli$(p)$-random variables, for some unknown parameter $p\in (0,1]$. Each individual acts autonomously at some unknown rate $μ>0$ and acts by mimetism at some rate depending on the number of recent actions of the individuals which influence him, the age of these actions being taken into account through an unknown function $φ$ (roughly, decreasing and with fast decay). The goal of this paper is to estimate $p$, which is the main charateristic of the graph of interactions, in the asymptotic $N\to\infty$, $t\to\infty$. The main issue is that the mean field limit (as $N \to \infty$) of this model is unidentifiable, in that it only depends on the parameters $μ$ and $pφ$. Fortunately, this mean field limit is not valid for large times. We distinguish the subcritical case, where, roughly, the mean number $m_t$ of actions per individual increases linearly and the supercritical case, where $m_t$ increases exponentially. Although the nuisance parameter $φ$ is non-parametric, we are able, in both cases, to estimate $p$ without estimating $φ$ in a nonparametric way, with a precision of order $N^{-1/2}+N^{1/2}m_t^{-1}$, up to some arbitrarily small loss. We explain, using a Gaussian toy model, the reason why this rate of convergence might be (almost) optimal.

preprint2015arXiv

New procedures controlling the false discovery proportion via Romano-Wolf's heuristic

The false discovery proportion (FDP) is a convenient way to account for false positives when a large number $m$ of tests are performed simultaneously. Romano and Wolf [Ann. Statist. 35 (2007) 1378-1408] have proposed a general principle that builds FDP controlling procedures from $k$-family-wise error rate controlling procedures while incorporating dependencies in an appropriate manner; see Korn et al. [J. Statist. Plann. Inference 124 (2004) 379-398]; Romano and Wolf (2007). However, the theoretical validity of the latter is still largely unknown. This paper provides a careful study of this heuristic: first, we extend this approach by using a notion of "bounding device" that allows us to cover a wide range of critical values, including those that adapt to $m\_0$, the number of true null hypotheses. Second, the theoretical validity of the latter is investigated both nonasymptotically and asymptotically. Third, we introduce suitable modifications of this heuristic that provide new methods, overcoming the existing procedures with a proven FDP control.

preprint2014arXiv

Asymptotic lower bounds in estimating jumps

We study the problem of the efficient estimation of the jumps for stochastic processes. We assume that the stochastic jump process $(X_t)_{t\in[0,1]}$ is observed discretely, with a sampling step of size $1/n$. In the spirit of Hajek's convolution theorem, we show some lower bounds for the estimation error of the sequence of the jumps $(ΔX_{T_k})_k$. As an intermediate result, we prove a LAMN property, with rate $\sqrt{n}$, when the marks of the underlying jump component are deterministic. We deduce then a convolution theorem, with an explicit asymptotic minimal variance, in the case where the marks of the jump component are random. To prove that this lower bound is optimal, we show that a threshold estimator of the sequence of jumps $(ΔX_{T_k})_k$ based on the discrete observations, reaches the minimal variance of the previous convolution theorem.

preprint2014arXiv

High dimensional Hawkes processes

We generalise the construction of multivariate Hawkes processes to a possibly infinite network of counting processes on a directed graph $\mathbb G$. The process is constructed as the solution to a system of Poisson driven stochastic differential equations, for which we prove pathwise existence and uniqueness under some reasonable conditions. We next investigate how to approximate a standard $N$-dimensional Hawkes process by a simple inhomogeneous Poisson process in the mean-field framework where each pair of individuals interact in the same way, in the limit $N \rightarrow \infty$. In the so-called linear case for the interaction, we further investigate the large time behaviour of the process. We study in particular the stability of the central limit theorem when exchanging the limits $N, T\rightarrow \infty$ and exhibit different possible behaviours. We finally consider the case $\mathbb G = \mathbb Z^d$ with nearest neighbour interactions. In the linear case, we prove some (large time) laws of large numbers and exhibit different behaviours, reminiscent of the infinite setting. Finally we study the propagation of a {\it single impulsion} started at a given point of $\zz^d$ at time $0$. We compute the probability of extinction of such an impulsion and, in some particular cases, we can accurately describe how it propagates to the whole space.

preprint2014arXiv

Testing over a continuum of null hypotheses with False Discovery Rate control

We consider statistical hypothesis testing simultaneously over a fairly general, possibly uncountably infinite, set of null hypotheses, under the assumption that a suitable single test (and corresponding $p$-value) is known for each individual hypothesis. We extend to this setting the notion of false discovery rate (FDR) as a measure of type I error. Our main result studies specific procedures based on the observation of the $p$-value process. Control of the FDR at a nominal level is ensured either under arbitrary dependence of $p$-values, or under the assumption that the finite dimensional distributions of the $p$-value process have positive correlations of a specific type (weak PRDS). Both cases generalize existing results established in the finite setting. Its interest is demonstrated in several non-parametric examples: testing the mean/signal in a Gaussian white noise model, testing the intensity of a Poisson process and testing the c.d.f. of i.i.d. random variables.

preprint2013arXiv

Estimating the efficient price from the order flow: a Brownian Cox process approach

At the ultra high frequency level, the notion of price of an asset is very ambiguous. Indeed, many different prices can be defined (last traded price, best bid price, mid price,...). Thus, in practice, market participants face the problem of choosing a price when implementing their strategies. In this work, we propose a notion of efficient price which seems relevant in practice. Furthermore, we provide a statistical methodology enabling to estimate this price form the order flow.

preprint2013arXiv

On empirical distribution function of high-dimensional Gaussian vector components with an application to multiple testing

This paper introduces a new framework to study the asymptotical behavior of the empirical distribution function (e.d.f.) of Gaussian vector components, whose correlation matrix $Γ^{(m)}$ is dimension-dependent. Hence, by contrast with the existing literature, the vector is not assumed to be stationary. Rather, we make a "vanishing second order" assumption ensuring that the covariance matrix $Γ^{(m)}$ is not too far from the identity matrix, while the behavior of the e.d.f. is affected by $Γ^{(m)}$ only through the sequence $γ_m=m^{-2} \sum_{i\neq j} Γ_{i,j}^{(m)}$, as $m$ grows to infinity. This result recovers some of the previous results for stationary long-range dependencies while it also applies to various, high-dimensional, non-stationary frameworks, for which the most correlated variables are not necessarily next to each other. Finally, we present an application of this work to the multiple testing problem, which was the initial statistical motivation for developing such a methodology.

preprint2012arXiv

Scaling limits for Hawkes processes and application to financial statistics

We prove a law of large numbers and a functional central limit theorem for multivariate Hawkes processes observed over a time interval $[0,T]$ in the limit $T \rightarrow \infty$. We further exhibit the asymptotic behaviour of the covariation of the increments of the components of a multivariate Hawkes process, when the observations are imposed by a discrete scheme with mesh $Δ$ over $[0,T]$ up to some further time shift $τ$. The behaviour of this functional depends on the relative size of $Δ$ and $τ$ with respect to $T$ and enables to give a full account of the second-order structure. As an application, we develop our results in the context of financial statistics. We introduced in a previous work a microscopic stochastic model for the variations of a multivariate financial asset, based on Hawkes processes and that is confined to live on a tick grid. We derive and characterise the exact macroscopic diffusion limit of this model and show in particular its ability to reproduce important empirical stylised fact such as the Epps effect and the lead-lag effect. Moreover, our approach enable to track these effects across scales in rigorous mathematical terms.

preprint2012arXiv

Testing the finiteness of the support of a distribution: a statistical look at Tsirelson's equation

We consider the following statistical problem: based on an i.i.d.sample of size n of integer valued random variables with common law m, is it possible to test whether or not the support of m is finite as n goes to infinity? This question is in particular connected to a simple case of Tsirelson's equation, for which it is natural to distinguish between two main configurations, the first one leading only to laws with finite support, and the second one including laws with infinite support. We show that it is in fact not possible to discriminate between the two situations, even using a very weak notion of statistical test.

preprint2010arXiv

Nonparametric regression with martingale increment errors

We consider the problem of adaptive estimation of the regression function in a framework where we replace ergodicity assumptions (such as independence or mixing) by another structural assumption on the model. Namely, we propose adaptive upper bounds for kernel estimators with data-driven bandwidth (Lepski's selection rule) in a regression model where the noise is an increment of martingale. It includes, as very particular cases, the usual i.i.d. regression and auto-regressive models. The cornerstone tool for this study is a new result for self-normalized martingales, called ``stability'', which is of independent interest. In a first part, we only use the martingale increment structure of the noise. We give an adaptive upper bound using a random rate, that involves the occupation time near the estimation point. Thanks to this approach, the theoretical study of the statistical procedure is disconnected from usual ergodicity properties like mixing. Then, in a second part, we make a link with the usual minimax theory of deterministic rates. Under a beta-mixing assumption on the covariates process, we prove that the random rate considered in the first part is equivalent, with large probability, to a deterministic rate which is the usual minimax adaptive one.

preprint2010arXiv

On the false discovery proportion convergence under Gaussian equi-correlation

We study the convergence of the false discovery proportion (FDP) of the Benjamini-Hochberg procedure in the Gaussian equi-correlated model, when the correlation $ρ_m$ converges to zero as the hypothesis number $m$ grows to infinity. By contrast with the standard convergence rate $m^{1/2}$ holding under independence, this study shows that the FDP converges to the false discovery rate (FDR) at rate $\{\min(m,1/ρ_m)\}^{1/2}$ in this equi-correlated model.