Source author record

Marc Hallin

Marc Hallin appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

26works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

26 published item(s)

preprint2024arXiv

Dynamic Factor Models: a Genealogy

Dynamic factor models have been developed out of the need of analyzing and forecasting time series in increasingly high dimensions. While mathematical statisticians faced with inference problems in high-dimensional observation spaces were focusing on the so-called spiked-model-asymptotics, econometricians adopted an entirely and considerably more effective asymptotic approach, rooted in the factor models originally considered in psychometrics. The so-called dynamic factor model methods, in two decades, has grown into a wide and successful body of techniques that are widely used in central banks, financial institutions, economic and statistical institutes. The objective of this chapter is not an extensive survey of the topic but a sketch of its historical growth, with emphasis on the various assumptions and interpretations, and a family tree of its main variants.

preprint2024arXiv

Multivariate Quantiles: Geometric and Measure-Transportation-Based Contours

Quantiles are a fundamental concept in probability and theoretical statistics and a daily tool in their applications. While the univariate concept of quantiles is quite clear and well understood, its multivariate extension is more problematic. After half a century of continued efforts and many proposals, two concepts, essentially, are emerging: the so-called (relabeled) geometric quantiles, extending the characterization of univariate quantiles as minimizers of an L1 loss function involving the check functions, and the more recent center-outward quantiles based on measure transportation ideas. These two concepts yield distinct families of quantile regions and quantile contours. Our objective here is to present a comparison of their main theoretical properties and a numerical investigation of their differences.

preprint2022arXiv

M-estimation in GARCH Models in the Absence of Higher-Order Moments

We consider a class of M-estimators of the parameters of a GARCH (p,q) model. These estimators involve score functions and, for adequate choices of the score functions, are asymptotically normal under milder moment assumptions than the usual quasi maximum likelihood, which makes them more reliable in the presence of heavy tails. We also consider weighted bootstrap approximations of the distributions of these M-estimators and establish their validity. Through extensive simulations, we demonstrate the robustness of these M-estimators under heavy tails and conduct a comparative study of the performance (bias and mean squared errors) of various score functions and the accuracy (confidence interval coverage rates) of their bootstrap approximations. In addition to the GARCH (1, 1) model, our simulations also involve higher-order models such as GARCH~(2, 1) and GARCH~(1,~\!2) which so far have received relatively little attention in the literature. We also consider the case of order-misspecified models. Finally, we use our M-estimators in the analysis of two real financial time series fitted with GARCH (1, 1) or GARCH (2, 1) models.

preprint2022arXiv

Nonparametric Multiple-Output Center-Outward Quantile Regression

Based on the novel concept of multivariate center-outward quantiles introduced recently in Chernozhukov et al. (2017) and Hallin et al. (2021), we are considering the problem of nonparametric multiple-output quantile regression. Our approach defines nested conditional center-outward quantile regression contours and regions with given conditional probability content irrespective of the underlying distribution; their graphs constitute nested center-outward quantile regression tubes. Empirical counterparts of these concepts are constructed, yielding interpretable empirical regions and contours which are shown to consistently reconstruct their population versions in the Pompeiu-Hausdorff topology. Our method is entirely non-parametric and performs well in simulations including heteroskedasticity and nonlinear trends; its power as a data-analytic tool is illustrated on some real datasets.

preprint2021arXiv

Multivariate goodness-of-Fit tests based on Wasserstein distance

Goodness-of-fit tests based on the empirical Wasserstein distance are proposed for simple and composite null hypotheses involving general multivariate distributions. For group families, the procedure is to be implemented after preliminary reduction of the data via invariance.This property allows for calculation of exact critical values and p-values at finite sample sizes. Applications include testing for location--scale families and testing for families arising from affine transformations, such as elliptical distributions with given standard radial density and unspecified location vector and scatter matrix. A novel test for multivariate normality with unspecified mean vector and covariance matrix arises as a special case. For more general parametric families, we propose a parametric bootstrap procedure to calculate critical values. The lack of asymptotic distribution theory for the empirical Wasserstein distance means that the validity of the parametric bootstrap under the null hypothesis remains a conjecture. Nevertheless, we show that the test is consistent against fixed alternatives. To this end, we prove a uniform law of large numbers for the empirical distribution in Wasserstein distance, where the uniformity is over any class of underlying distributions satisfying a uniform integrability condition but no additional moment assumptions. The calculation of test statistics boils down to solving the well-studied semi-discrete optimal transport problem. Extensive numerical experiments demonstrate the practical feasibility and the excellent performance of the proposed tests for the Wasserstein distance of order p = 1 and p = 2 and for dimensions at least up to d = 5. The simulations also lend support to the conjecture of the asymptotic validity of the parametric bootstrap.

preprint2020arXiv

Center-Outward Distribution Functions, Quantiles, Ranks, and Signs in $\mathbb{R}^d$

Univariate concepts as quantile and distribution functions involving ranks and signs, do not canonically extend to $\mathbb{R}^d, d\geq 2$. Palliating that has generated an abundant literature. Chapter 1 shows that, unlike the many definitions that have been proposed so far, the measure transportation-based ones introduced in Chernozhukov et al. (2017) enjoy all the properties that make univariate quantiles and ranks successful tools for semiparametric statistical inference. We therefore propose a new center-outward definition of multivariate distribution and quantile functions, along with their empirical counterparts, for which we obtain a Glivenko-Cantelli result. Our approach is geometric and, contrary to the Monge-Kantorovich one in Chernozhukov et al. (2017), does not require any moment assumptions. The resulting ranks and signs are strictly distribution-free, and maximal invariant under the action of a data-driven class of (order-preserving) transformations generating the family of absolutely continuous distributions; that property is the theoretical foundation of the semiparametric efficiency preservation property of ranks. The corresponding quantiles are equivariant under the same transformations. The empirical proposed distribution functions are defined at observed values only. A continuous extension to the entire $\mathbb{R}^d$, yielding continuous empirical quantile contours while preserving the monotonicity and Glivenko-Cantelli features is desirable. Such extension requires solving a nontrivial problem of smooth interpolation under cyclical monotonicity constraints. A complete solution of that problem is given in Chapter 2; we show that the resulting distribution and quantile functions are Lipschitz, and provide a sharp lower bound for the Lipschitz constants. A numerical study of empirical center-outward quantile contours and their consistency is conducted.

preprint2020arXiv

Center-Outward R-Estimation for Semiparametric VARMA Models

We propose a new class of R-estimators for semiparametric VARMA models in which the innovation density plays the role of the nuisance parameter. Our estimators are based on the novel concepts of multivariate center-outward ranks and signs. We show that these concepts, combined with Le Cam's asymptotic theory of statistical experiments, yield a class of semiparametric estimation procedures, which are efficient (at a given reference density), root-$n$ consistent, and asymptotically normal under a broad class of (possibly non elliptical) actual innovation densities. No kernel density estimation is required to implement our procedures. A Monte Carlo comparative study of our R-estimators and other routinely-applied competitors demonstrates the benefits of the novel methodology, in large and small sample. Proofs, computational aspects, and further numerical results are available in the supplementary material.

preprint2019arXiv

Generalized Dynamic Factor Models and Volatilities: Consistency, rates, and prediction intervals

Volatilities, in high-dimensional panels of economic time series with a dynamic factor structure on the levels or returns, typically also admit a dynamic factor decomposition. We consider a two-stage dynamic factor model method recovering the common and idiosyncratic components of both levels and log-volatilities. Specifically, in a first estimation step, we extract the common and idiosyncratic shocks for the levels, from which a log-volatility proxy is computed. In a second step, we estimate a dynamic factor model, which is equivalent to a multiplicative factor structure for volatilities, for the log-volatility panel. By exploiting this two-stage factor approach, we build one-step-ahead conditional prediction intervals for large $n \times T$ panels of returns. Those intervals are based on empirical quantiles, not on conditional variances; they can be either equal- or unequal- tailed. We provide uniform consistency and consistency rates results for the proposed estimators as both $n$ and $T$ tend to infinity. We study the finite-sample properties of our estimators by means of Monte Carlo simulations. Finally, we apply our methodology to a panel of asset returns belonging to the S&P100 index in order to compute one-step-ahead conditional prediction intervals for the period 2006-2013. A comparison with the componentwise GARCH benchmark (which does not take advantage of cross-sectional information) demonstrates the superiority of our approach, which is genuinely multivariate (and high-dimensional), nonparametric, and model-free.

preprint2016arXiv

Quantile Spectral Analysis for Locally Stationary Time Series

Classical spectral methods are subject to two fundamental limitations: they only can account for covariance-related serial dependencies, and they require second-order stationarity. Much attention has been devoted lately to quantile-based spectral methods that go beyond covariance-based serial dependence features. At the same time, covariance-based methods relaxing stationarity into much weaker {\it local stationarity} conditions have been developed for a variety of time-series models. Here, we are combining those two approaches by proposing quantile-based spectral methods for locally stationary processes. We therefore introduce a time-varying version of the copula spectra that have been recently proposed in the literature, along with a suitable local lag-window estimator. We propose a new definition of local {\it strict} stationarity that allows us to handle completely general non-linear processes without any moment assumptions, thus accommodating our quantile-based concepts and methods. We establish a central limit theorem for the new estimators, and illustrate the power of the proposed methodology by means of a simulation study. Moreover, in two empirical studies (namely of the Standard \& Poor's 500 series and a temperature dataset recorded in Hohenpeissenberg) we demonstrate that the new approach detects important variations in serial dependence structures both across time and across quantiles. Such variations remain completely undetected, and are actually undetectable, via classical covariance-based spectral methods.

preprint2016arXiv

Quantile spectral processes: Asymptotic analysis and inference

Quantile- and copula-related spectral concepts recently have been considered by various authors. Those spectra, in their most general form, provide a full characterization of the copulas associated with the pairs $(X_t,X_{t-k})$ in a process $(X_t)_{t\in\mathbb{Z}}$, and account for important dynamic features, such as changes in the conditional shape (skewness, kurtosis), time-irreversibility, or dependence in the extremes that their traditional counterparts cannot capture. Despite various proposals for estimation strategies, only quite incomplete asymptotic distributional results are available so far for the proposed estimators, which constitutes an important obstacle for their practical application. In this paper, we provide a detailed asymptotic analysis of a class of smoothed rank-based cross-periodograms associated with the copula spectral density kernels introduced in Dette et al. [Bernoulli 21 (2015) 781-831]. We show that, for a very general class of (possibly nonlinear) processes, properly scaled and centered smoothed versions of those cross-periodograms, indexed by couples of quantile levels, converge weakly, as stochastic processes, to Gaussian processes. A first application of those results is the construction of asymptotic confidence intervals for copula spectral density kernels. The same convergence results also provide asymptotic distributions (under serially dependent observations) for a new class of rank-based spectral methods involving the Fourier transforms of rank-based serial statistics such as the Spearman, Blomqvist or Gini autocovariance coefficients.

preprint2015arXiv

Dynamic Functional Principal Component

In this paper, we address the problem of dimension reduction for time series of functional data $(X_t\colon t\in\mathbb{Z})$. Such {\it functional time series} frequently arise, e.g., when a continuous-time process is segmented into some smaller natural units, such as days. Then each~$X_t$ represents one intraday curve. We argue that functional principal component analysis (FPCA), though a key technique in the field and a benchmark for any competitor, does not provide an adequate dimension reduction in a time-series setting. FPCA indeed is a {\it static} procedure which ignores the essential information provided by the serial dependence structure of the functional data under study. Therefore, inspired by Brillinger's theory of {\it dynamic principal components}, we propose a {\it dynamic} version of FPCA, which is based on a frequency-domain approach. By means of a simulation study and an empirical illustration, we show the considerable improvement the dynamic approach entails when compared to the usual static procedure.

preprint2015arXiv

Local bilinear multiple-output quantile/depth regression

A new quantile regression concept, based on a directional version of Koenker and Bassett's traditional single-output one, has been introduced in [Ann. Statist. (2010) 38 635-669] for multiple-output location/linear regression problems. The polyhedral contours provided by the empirical counterpart of that concept, however, cannot adapt to unknown nonlinear and/or heteroskedastic dependencies. This paper therefore introduces local constant and local linear (actually, bilinear) versions of those contours, which both allow to asymptotically recover the conditional halfspace depth contours that completely characterize the response's conditional distributions. Bahadur representation and asymptotic normality results are established. Illustrations are provided both on simulated and real data.

preprint2015arXiv

Of copulas, quantiles, ranks and spectra: An $L_1$-approach to spectral analysis

In this paper, we present an alternative method for the spectral analysis of a univariate, strictly stationary time series $\{Y_t\}_{t\in \mathbb {Z}}$. We define a "new" spectrum as the Fourier transform of the differences between copulas of the pairs $(Y_t,Y_{t-k})$ and the independence copula. This object is called a copula spectral density kernel and allows to separate the marginal and serial aspects of a time series. We show that this spectrum is closely related to the concept of quantile regression. Like quantile regression, which provides much more information about conditional distributions than classical location-scale regression models, copula spectral density kernels are more informative than traditional spectral densities obtained from classical autocovariances. In particular, copula spectral density kernels, in their population versions, provide (asymptotically provide, in their sample versions) a complete description of the copulas of all pairs $(Y_t,Y_{t-k})$. Moreover, they inherit the robustness properties of classical quantile regression, and do not require any distributional assumptions such as the existence of finite moments. In order to estimate the copula spectral density kernel, we introduce rank-based Laplace periodograms which are calculated as bilinear forms of weighted $L_1$-projections of the ranks of the observed time series onto a harmonic regression model. We establish the asymptotic distribution of those periodograms, and the consistency of adequately smoothed versions. The finite-sample properties of the new methodology, and its potential for applications are briefly investigated by simulations and a short empirical example.

preprint2014arXiv

Skew-symmetric distributions and Fisher information: The double sin of the skew-normal

Hallin and Ley [Bernoulli 18 (2012) 747-763] investigate and fully characterize the Fisher singularity phenomenon in univariate and multivariate families of skew-symmetric distributions. This paper proposes a refined analysis of the (univariate) problem, showing that singularity can be more or less severe, inducing $n^{1/4}$ ("simple singularity"), $n^{1/6}$ ("double singularity"), or $n^{1/8}$ ("triple singularity") consistency rates for the skewness parameter. We show, however, that simple singularity (yielding $n^{1/4}$ consistency rates), if any singularity at all, is the rule, in the sense that double and triple singularities are possible for generalized skew-normal families only. We also show that higher-order singularities, leading to worse-than-$n^{1/8}$ rates, cannot occur. Depending on the degree of the singularity, our analysis also suggests a simple reparametrization that offers an alternative to the so-called centred parametrization proposed, in the particular case of skew-normal and skew-$t$ families, by Azzalini [Scand. J. Stat. 12 (1985) 171-178], Arellano-Valle and Azzalini [J. Multivariate Anal. 113 (2013) 73-90], and DiCiccio and Monti [Quaderni di Statistica 13 (2011) 1-21], respectively.

preprint2013arXiv

Asymptotic power of sphericity tests for high-dimensional data

This paper studies the asymptotic power of tests of sphericity against perturbations in a single unknown direction as both the dimensionality of the data and the number of observations go to infinity. We establish the convergence, under the null hypothesis and contiguous alternatives, of the log ratio of the joint densities of the sample covariance eigenvalues to a Gaussian process indexed by the norm of the perturbation. When the perturbation norm is larger than the phase transition threshold studied in Baik, Ben Arous and Peche [Ann. Probab. 33 (2005) 1643-1697] the limiting process is degenerate, and discrimination between the null and the alternative is asymptotically certain. When the norm is below the threshold, the limiting process is nondegenerate, and the joint eigenvalue densities under the null and alternative hypotheses are mutually contiguous. Using the asymptotic theory of statistical experiments, we obtain asymptotic power envelopes and derive the asymptotic power for various sphericity tests in the contiguity region. In particular, we show that the asymptotic power of the Tracy-Widom-type tests is trivial (i.e., equals the asymptotic size), whereas that of the eigenvalue-based likelihood ratio test is strictly larger than the size, and close to the power envelope.

preprint2013arXiv

On Hodges and Lehmann's "$6/π$ result"

While the asymptotic relative efficiency (ARE) of Wilcoxon rank-based tests for location and regression with respect to their parametric Student competitors can be arbitrarily large, Hodges and Lehmann (1961) have shown that the ARE of the same Wilcoxon tests with respect to their van der Waerden or normal-score counterparts is bounded from above by $6/π\approx 1.910$. In this paper, we revisit that result, and investigate similar bounds for statistics based on Student scores. We also consider the serial version of this ARE. More precisely, we study the ARE, under various densities, of the Spearman-Wald-Wolfowitz and Kendall rank-based autocorrelations with respect to the van der Waerden or normal-score ones used to test (ARMA) serial dependence alternatives.

preprint2013arXiv

R-Estimation for Asymmetric Independent Component Analysis

Independent Component Analysis (ICA) recently has attracted attention in the statistical literature as an alternative to elliptical models. Whereas k-dimensional elliptical densities depend on one single unspecified radial density, however, k-dimensional independent component distributions involve k unspecified component densities that for given sample size n and dimension k making statistical analysis harder. We focus here on estimating the model's mixing matrix. Traditional methods (FOBI, Kernel-ICA, FastICA) originating from the engineering literature have consistency that requires moment conditions without achieving any type of asymptotic efficiency. When based on robust scatter matrices, the two-scatter methods developed by Oja, et al. (2006) and Nordhausen, et al. (2008) enjoy better robustness features but have unclear optimality properties. The semiparametric approach by Chen and Bickel (2006) achieves semiparametric efficiency but requires estimating the k unobserved independent component densities. As a reaction, an efficient (signed-)rank-based approach has been proposed by Ilmonen and Paindaveine (2011) for the case of symmetric component densities that fail to be root-n consistent as soon as one of the component densities is asymmetric. In this paper, using ranks rather than signed ranks, we extend their approach to the asymmetric case and propose a one-step R-estimator for ICA mixing matrices. The finite-sample performances of those estimators are investigated and compared to those of existing methods under moderately large sample sizes. Particularly good performances are obtained when using data-driven scores taking into account the skewness and kurtosis of residuals. Finally, we show, by an empirical exercise, that our methods also may provide excellent results in a context such as image analysis, where the basic assumptions of ICA are quite unlikely to hold.

preprint2012arXiv

One-Step R-Estimation in Linear Models with Stable Errors

Classical estimation techniques for linear models either are inconsistent, or perform rather poorly, under $α$-stable error densities; most of them are not even rate-optimal. In this paper, we propose an original one-step R-estimation method and investigate its asymptotic performances under stable densities. Contrary to traditional least squares, the proposed R-estimators remain root-$n$ consistent (the optimal rate) under the whole family of stable distributions, irrespective of their asymmetry and tail index. While parametric stable-likelihood estimation, due to the absence of a closed form for stable densities, is quite cumbersome, our method allows us to construct estimators reaching the parametric efficiency bounds associated with any prescribed values $(α_0, \ b_0)$ of the tail index $α$ and skewness parameter $b$, while preserving root-$n$ consistency under any $(α, \ b)$ as well as under usual light-tailed densities. The method furthermore avoids all forms of multidimensional argmin computation. Simulations confirm its excellent finite-sample performances.

preprint2012arXiv

Optimal rank-based testing for principal components

This paper provides parametric and rank-based optimal tests for eigenvectors and eigenvalues of covariance or scatter matrices in elliptical families. The parametric tests extend the Gaussian likelihood ratio tests of Anderson (1963) and their pseudo-Gaussian robustifications by Davis (1977) and Tyler (1981, 1983). The rank-based tests address a much broader class of problems, where covariance matrices need not exist and principal components are associated with more general scatter matrices. The proposed tests are shown to outperform daily practice both from the point of view of validity as from the point of view of efficiency. This is achieved by utilizing the Le Cam theory of locally asymptotically normal experiments, in the nonstandard context, however, of a curved parametrization. The results we derive for curved experiments are of independent interest, and likely to apply in other contexts.

preprint2012arXiv

Signal Detection in High Dimension: The Multispiked Case

This paper deals with the local asymptotic structure, in the sense of Le Cam's asymptotic theory of statistical experiments, of the signal detection problem in high dimension. More precisely, we consider the problem of testing the null hypothesis of sphericity of a high-dimensional covariance matrix against an alternative of (unspecified) multiple symmetry-breaking directions (\textit{multispiked} alternatives). Simple analytical expressions for the asymptotic power envelope and the asymptotic powers of previously proposed tests are derived. These asymptotic powers are shown to lie very substantially below the envelope, at least for relatively small values of the number of symmetry-breaking directions under the alternative. In contrast, the asymptotic power of the likelihood ratio test based on the eigenvalues of the sample covariance matrix is shown to be close to that envelope. These results extend to the case of multispiked alternatives the findings of an earlier study (Onatski, Moreira and Hallin, 2011) of the single-spiked case. The methods we are using here, however, are entirely new, as the Laplace approximations considered in the single-spiked context do not extend to the multispiked case.

preprint2012arXiv

Skew-symmetric distributions and Fisher information -- a tale of two densities

Skew-symmetric densities recently received much attention in the literature, giving rise to increasingly general families of univariate and multivariate skewed densities. Most of those families, however, suffer from the inferential drawback of a potentially singular Fisher information in the vicinity of symmetry. All existing results indicate that Gaussian densities (possibly after restriction to some linear subspace) play a special and somewhat intriguing role in that context. We dispel that widespread opinion by providing a full characterization, in a general multivariate context, of the information singularity phenomenon, highlighting its relation to a possible link between symmetric kernels and skewing functions -- a link that can be interpreted as the mismatch of two densities.

preprint2011arXiv

A class of optimal tests for symmetry based on local Edgeworth approximations

The objective of this paper is to provide, for the problem of univariate symmetry (with respect to specified or unspecified location), a concept of optimality, and to construct tests achieving such optimality. This requires embedding symmetry into adequate families of asymmetric (local) alternatives. We construct such families by considering non-Gaussian generalizations of classical first-order Edgeworth expansions indexed by a measure of skewness such that (i) location, scale and skewness play well-separated roles (diagonality of the corresponding information matrices) and (ii) the classical tests based on the Pearson--Fisher coefficient of skewness are optimal in the vicinity of Gaussian densities.

preprint2010arXiv

Multivariate quantiles and multiple-output regression quantiles: From $L_1$ optimization to halfspace depth

A new multivariate concept of quantile, based on a directional version of Koenker and Bassett's traditional regression quantiles, is introduced for multivariate location and multiple-output regression problems. In their empirical version, those quantiles can be computed efficiently via linear programming techniques. Consistency, Bahadur representation and asymptotic normality results are established. Most importantly, the contours generated by those quantiles are shown to coincide with the classical halfspace depth contours associated with the name of Tukey. This relation does not only allow for efficient depth contour computations by means of parametric linear programming, but also for transferring from the quantile to the depth universe such asymptotic results as Bahadur representations. Finally, linear programming duality opens the way to promising developments in depth-related multivariate rank-based inference.

preprint2007arXiv

Semiparametrically efficient rank-based inference for shape II. Optimal R-estimation of shape

A class of R-estimators based on the concepts of multivariate signed ranks and the optimal rank-based tests developed in Hallin and Paindaveine [Ann. Statist. 34 (2006)] is proposed for the estimation of the shape matrix of an elliptical distribution. These R-estimators are root-n consistent under any radial density g, without any moment assumptions, and semiparametrically efficient at some prespecified density f. When based on normal scores, they are uniformly more efficient than the traditional normal-theory estimator based on empirical covariance matrices (the asymptotic normality of which, moreover, requires finite moments of order four), irrespective of the actual underlying elliptical density. They rely on an original rank-based version of Le Cam's one-step methodology which avoids the unpleasant nonparametric estimation of cross-information quantities that is generally required in the context of R-estimation. Although they are not strictly affine-equivariant, they are shown to be equivariant in a weak asymptotic sense. Simulations confirm their feasibility and excellent finite-sample performances.