Source author record

François Roueff

François Roueff appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

19works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

19 published item(s)

preprint2021arXiv

Nonlinear Functional Output Regression: a Dictionary Approach

To address functional-output regression, we introduce projection learning (PL), a novel dictionary-based approach that learns to predict a function that is expanded on a dictionary while minimizing an empirical risk based on a functional loss. PL makes it possible to use non orthogonal dictionaries and can then be combined with dictionary learning; it is thus much more flexible than expansion-based approaches relying on vectorial losses. This general method is instantiated with reproducing kernel Hilbert spaces of vector-valued functions as kernel-based projection learning (KPL). For the functional square loss, two closed-form estimators are proposed, one for fully observed output functions and the other for partially observed ones. Both are backed theoretically by an excess risk analysis. Then, in the more general setting of integral losses based on differentiable ground losses, KPL is implemented using first-order optimization for both fully and partially observed output functions. Eventually, several robustness aspects of the proposed algorithms are highlighted on a toy dataset; and a study on two real datasets shows that they are competitive compared to other nonlinear approaches. Notably, using the square loss and a learnt dictionary, KPL enjoys a particularily attractive trade-off between computational cost and performances.

preprint2020arXiv

Necessary and sufficient conditions for the identifiability of observation-driven models

In this contribution we are interested in proving that a given observation-driven model is identifiable. In the case of a GARCH(p, q) model, a simple sufficient condition has been established in [1] for showing the consistency of the quasi-maximum likelihood estimator. It turns out that this condition applies for a much larger class of observation-driven models, that we call the class of linearly observation-driven models. This class includes standard integer valued observation-driven time series, such as the log-linear Poisson GARCH or the NBIN-GARCH models.

preprint2020arXiv

Spectral estimation for non-linear long range dependent discrete time trawl processes

Discrete time trawl processes constitute a large class of time series parameterized by a trawl sequence (a j) j$\in$N and defined though a sequence of independent and identically distributed (i.i.d.) copies of a continuous time process ($γ$(t)) t$\in$R called the seed process. They provide a general framework for modeling linear or non-linear long range dependent time series. We investigate the spectral estimation, either pointwise or broadband, of long range dependent discrete-time trawl processes. The difficulty arising from the variety of seed processes and of trawl sequences is twofold. First, the spectral density may take different forms, often including smooth additive correction terms. Second, trawl processes with similar spectral densities may exhibit very different statistical behaviors. We prove the consistency of our estimators under very general conditions and we show that a wide class of trawl processes satisfy them. This is done in particular by introducing a weighted weak dependence index that can be of independent interest. The broadband spectral estimator includes an estimator of the long memory parameter. We complete this work with numerical experiments to evaluate the finite sample size performance of this estimator for various integer valued discrete time trawl processes.

preprint2016arXiv

Anomaly Detection and Localisation using Mixed Graphical Models

We propose a method that performs anomaly detection and localisation within heterogeneous data using a pairwise undirected mixed graphical model. The data are a mixture of categorical and quantitative variables, and the model is learned over a dataset that is supposed not to contain any anomaly. We then use the model over temporal data, potentially a data stream, using a version of the two-sided CUSUM algorithm. The proposed decision statistic is based on a conditional likelihood ratio computed for each variable given the others. Our results show that this function allows to detect anomalies variable by variable, and thus to localise the variables involved in the anomalies more precisely than univariate methods based on simple marginals.

preprint2016arXiv

Nonparametric estimation of mark's distribution of an exponential Shot-noise process

In this paper, we consider a nonlinear inverse problem occurring in nuclear science. Gamma rays randomly hit a semiconductor detector which produces an impulse response of electric current. Because the sampling period of the measured current is larger than the mean inter arrival time of photons, the impulse responses associated to different gamma rays can overlap: this phenomenon is known as pileup. In this work, it is assumed that the impulse response is an exponentially decaying function. We propose a novel method to infer the distribution of gamma photon energies from the indirect measurements obtained from the detector. This technique is based on a formula linking the characteristic function of the photon density to a function involving the characteristic function and its derivative of the observations. We establish that our estimator converges to the mark density in uniform norm at a logarithmic rate. A limited Monte-Carlo experiment is provided to support our findings.

preprint2015arXiv

Aggregation of predictors for nonstationary sub-linear processes and online adaptive forecasting of time varying autoregressive processes

In this work, we study the problem of aggregating a finite number of predictors for nonstationary sub-linear processes. We provide oracle inequalities relying essentially on three ingredients: (1) a uniform bound of the $\ell^1$ norm of the time varying sub-linear coefficients, (2) a Lipschitz assumption on the predictors and (3) moment conditions on the noise appearing in the linear representation. Two kinds of aggregations are considered giving rise to different moment conditions on the noise and more or less sharp oracle inequalities. We apply this approach for deriving an adaptive predictor for locally stationary time varying autoregressive (TVAR) processes. It is obtained by aggregating a finite number of well chosen predictors, each of them enjoying an optimal minimax convergence rate under specific smoothness conditions on the TVAR coefficients. We show that the obtained aggregated predictor achieves a minimax rate while adapting to the unknown smoothness. To prove this result, a lower bound is established for the minimax rate of the prediction risk for the TVAR process. Numerical experiments complete this study. An important feature of this approach is that the aggregated predictor can be computed recursively and is thus applicable in an online prediction context.

preprint2015arXiv

Convergence to stable laws in the space $D$

We study the convergence of centered and normalized sums of i.i.d. random elements of the space $\mathcal{D}$ of c{á}dl{á}g functions endowed with Skorohod's $J\_1$ topology, to stable distributions in $\mathcal D$. Our results are based on the concept of regular variation on metric spaces and on point process convergence. We provide some applications, in particular to the empirical process of the renewal-reward process.

preprint2015arXiv

Handy sufficient conditions for the convergence of the maximum likelihood estimator in observation-driven models

This paper generalizes asymptotic properties obtained in the observation-driven times series models considered by \cite{dou:kou:mou:2013} in the sense that the conditional law of each observation is also permitted to depend on the parameter. The existence of ergodic solutions and the consistency of the Maximum Likelihood Estimator (MLE) are derived under easy-to-check conditions. The obtained conditions appear to apply for a wide class of models. We illustrate our results with specific observation-driven times series, including the recently introduced NBIN-GARCH and NM-GARCH models, demonstrating the consistency of the MLE for these two models.

preprint2015arXiv

Nonparametric estimation of the mixing density using polynomials

We consider the problem of estimating the mixing density $f$ from $n$ i.i.d. observations distributed according to a mixture density with unknown mixing distribution. In contrast with finite mixtures models, here the distribution of the hidden variable is not bounded to a finite set but is spread out over a given interval. We propose an approach to construct an orthogonal series estimator of the mixing density $f$ involving Legendre polynomials. The construction of the orthonormal sequence varies from one mixture model to another. Minimax upper and lower bounds of the mean integrated squared error are provided which apply in various contexts. In the specific case of exponential mixtures, it is shown that the estimator is adaptive over a collection of specific smoothness classes, more precisely, there exists a constant $A\textgreater{}0$ such that, when the order $m$ of the projection estimator verifies $m\sim A \log(n)$, the estimator achieves the minimax rate over this collection. Other cases are investigated such as Gamma shape mixtures and scale mixtures of compactly supported densities including Beta mixtures. Finally, a consistent estimator of the support of the mixing density $f$ is provided.

preprint2014arXiv

Asymptotic behavior of the quadratic variation of the sum of two Hermite processes of consecutive orders

Hermite processes are self--similar processes with stationary increments which appear as limits of normalized sums of random variables with long range dependence. The Hermite process of order $1$ is fractional Brownian motion and the Hermite process of order $2$ is the Rosenblatt process. We consider here the sum of two Hermite processes of order $q\geq 1$ and $q+1$ and of different Hurst parameters. We then study its quadratic variations at different scales. This is akin to a wavelet decomposition. We study both the cases where the Hermite processes are dependent and where they are independent. In the dependent case, we show that the quadratic variation, suitably normalized, converges either to a normal or to a Rosenblatt distribution, whatever the order of the original Hermite processes.

preprint2014arXiv

Ergodicity and scaling limit of a constrained multivariate Hawkes process

We introduce a multivariate Hawkes process with constraints on its conditional density. It is a multivariate point process with conditional intensity similar to that of a multivariate Hawkes process but certain events are forbidden with respect to boundary conditions on a multidimensional constraint variable, whose evolution is driven by the point process. We study this process in the special case where the fertility function is exponential so that the process is entirely described by an underlying Markov chain, which includes the constraint variable. Some conditions on the parameters are established to ensure the ergodicity of the chain. Moreover, scaling limits are derived for the integrated point process. This study is primarily motivated by the stochastic modelling of a limit order book for high frequency financial data analysis.

preprint2014arXiv

Large scale reduction principle and application to hypothesis testing

Consider a non--linear function $G(X_t)$ where $X_t$ is a stationary Gaussian sequence with long--range dependence. The usual reduction principle states that the partial sums of $G(X_t)$ behave asymptotically like the partial sums of the first term in the expansion of $G$ in Hermite polynomials. In the context of the wavelet estimation of the long--range dependence parameter, one replaces the partial sums of $G(X_t)$ by the wavelet scalogram, namely the partial sum of squares of the wavelet coefficients. Is there a reduction principle in the wavelet setting, namely is the asymptotic behavior of the scalogram for $G(X_t)$ the same as that for the first term in the expansion of $G$ in Hermite polynomial? The answer is negative in general. This paper provides a minimal growth condition on the scales of the wavelet coefficients which ensures that the reduction principle also holds for the scalogram. The results are applied to testing the hypothesis that the long-range dependence parameter takes a specific value.

preprint2013arXiv

High order chaotic limits of wavelet scalograms under long--range dependence

Let $G$ be a non--linear function of a Gaussian process $\{X_t\}_{t\in\mathbb{Z}}$ with long--range dependence. The resulting process $\{G(X_t)\}_{t\in\mathbb{Z}}$ is not Gaussian when $G$ is not linear. We consider random wavelet coefficients associated with $\{G(X_t)\}_{t\in\mathbb{Z}}$ and the corresponding wavelet scalogram which is the average of squares of wavelet coefficients over locations. We obtain the asymptotic behavior of the scalogram as the number of observations and scales tend to infinity. It is known that when $G$ is a Hermite polynomial of any order, then the limit is either the Gaussian or the Rosenblatt distribution, that is, the limit can be represented by a multiple Wiener-Itô integral of order one or two. We show, however, that there are large classes of functions $G$ which yield a higher order Hermite distribution, that is, the limit can be represented by a a multiple Wiener-Itô integral of order greater than two.

preprint2013arXiv

Wavelet estimation of the long memory parameter for Hermite polynomial of Gaussian processes

We consider stationary processes with long memory which are non-Gaussian and represented as Hermite polynomials of a Gaussian process. We focus on the corresponding wavelet coefficients and study the asymptotic behavior of the sum of their squares since this sum is often used for estimating the long-memory parameter. We show that the limit is not Gaussian but can be expressed using the non-Gaussian Rosenblatt process defined as a Wiener Itô integral of order 2. This happens even if the original process is defined through a Hermite polynomial of order higher than 2.

preprint2012arXiv

Function-indexed empirical processes based on an infinite source Poisson transmission stream

We study the asymptotic behavior of empirical processes generated by measurable bounded functions of an infinite source Poisson transmission process when the session length have infinite variance. In spite of the boundedness of the function, the normalized fluctuations of such an empirical process converge to a non-Gaussian stable process. This phenomenon can be viewed as caused by the long-range dependence in the transmission process. Completing previous results on the empirical mean of similar types of processes, our results on non-linear bounded functions exhibit the influence of the limit transmission rate distribution at high session lengths on the asymptotic behavior of the empirical process. As an illustration, we apply the main result to estimation of the distribution function of the steady state value of the transmission process.

preprint2011arXiv

Testing for homogeneity of variance in the wavelet domain

The danger of confusing long-range dependence with non-stationarity has been pointed out by many authors. Finding an answer to this difficult question is of importance to model time-series showing trend-like behavior, such as river run-off in hydrology, historical temperatures in the study of climates changes, or packet counts in network traffic engineering. The main goal of this paper is to develop a test procedure to detect the presence of non-stationarity for a class of processes whose $K$-th order difference is stationary. Contrary to most of the proposed methods, the test procedure has the same distribution for short-range and long-range dependence covariance stationary processes, which means that this test is able to detect the presence of non-stationarity for processes showing long-range dependence or which are unit root. The proposed test is formulated in the wavelet domain, where a change in the generalized spectral density results in a change in the variance of wavelet coefficients at one or several scales. Such tests have been already proposed in \cite{whitcher:2001}, but these authors do not have taken into account the dependence of the wavelet coefficients within scales and between scales. Therefore, the asymptotic distribution of the test they have proposed was erroneous; as a consequence, the level of the test under the null hypothesis of stationarity was wrong. In this contribution, we introduce two test procedures, both using an estimator of the variance of the scalogram at one or several scales. The asymptotic distribution of the test under the null is rigorously justified. The pointwise consistency of the test in the presence of a single jump in the general spectral density is also be presented. A limited Monte-Carlo experiment is performed to illustrate our findings.

preprint2010arXiv

Large scale behavior of wavelet coefficients of non-linear subordinated processes with long memory

We study the asymptotic behavior of wavelet coefficients of random processes with long memory. These processes may be stationary or not and are obtained as the output of non--linear filter with Gaussian input. The wavelet coefficients that appear in the limit are random, typically non--Gaussian and belong to a Wiener chaos. They can be interpreted as wavelet coefficients of a generalized self-similar process.

preprint2010arXiv

Locally stationary long memory estimation

There exists a wide literature on modelling strongly dependent time series using a longmemory parameter d, including more recent work on semiparametric wavelet estimation. As a generalization of these latter approaches, in this work we allow the long-memory parameter d to be varying over time. We embed our approach into the framework of locally stationary processes. We show weak consistency and a central limit theorem for our log-regression wavelet estimator of the time-dependent d in a Gaussian context. Both simulations and a real data example complete our work on providing a fairly general approach.

preprint2006arXiv

Nonparametric estimation of mixing densities for discrete distributions

By a mixture density is meant a density of the form $π_μ(\cdot)=\intπ_θ(\cdot)\timesμ(dθ)$, where $(π_θ)_{θ\inΘ}$ is a family of probability densities and $μ$ is a probability measure on $Θ$. We consider the problem of identifying the unknown part of this model, the mixing distribution $μ$, from a finite sample of independent observations from $π_μ$. Assuming that the mixing distribution has a density function, we wish to estimate this density within appropriate function classes. A general approach is proposed and its scope of application is investigated in the case of discrete distributions. Mixtures of power series distributions are more specifically studied. Standard methods for density estimation, such as kernel estimators, are available in this context, and it has been shown that these methods are rate optimal or almost rate optimal in balls of various smoothness spaces. For instance, these results apply to mixtures of the Poisson distribution parameterized by its mean. Estimators based on orthogonal polynomial sequences have also been proposed and shown to achieve similar rates. The general approach of this paper extends and simplifies such results. For instance, it allows us to prove asymptotic minimax efficiency over certain smoothness classes of the above-mentioned polynomial estimator in the Poisson case. We also study discrete location mixtures, or discrete deconvolution, and mixtures of discrete uniform distributions.