Source author record

Joni Virta

Joni Virta appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

10works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

10 published item(s)

preprint2022arXiv

Order Determination for Tensor-valued Observations Using Data Augmentation

Tensor-valued data benefits greatly from dimension reduction as the reduction in size is exponential in the number of modes. To achieve maximal reduction without loss in information, our objective in this work is to give an automated procedure for the optimal selection of the reduced dimensionality. Our approach combines a recently proposed data augmentation procedure with the higher-order singular value decomposition (HOSVD) in a tensorially natural way. We give theoretical guidelines on how to choose the tuning parameters and further inspect their influence in a simulation study. As our primary result, we show that the procedure consistently estimates the true latent dimensions under a noisy tensor model, both at the population and sample levels. Additionally, we propose a bootstrap-based alternative to the augmentation estimator. Simulations are used to demonstrate the estimation accuracy of the two methods under various settings.

preprint2022arXiv

Robust signal dimension estimation via SURE

The estimation of signal dimension under heavy-tailed latent factor models is studied. As a primary contribution, robust extensions of an earlier estimator based on Gaussian Stein's unbiased risk estimation are proposed. These novel extensions are based on the framework of elliptical distributions and robust scatter matrices. Extensive simulation studies are conducted in order to compare the novel methods with several well-known competitors in both estimation accuracy and computational speed. The novel methods are applied to a financial asset return data set.

preprint2022arXiv

Sliced Inverse Regression in Metric Spaces

In this article, we propose a general nonlinear sufficient dimension reduction (SDR) framework when both the predictor and response lie in some general metric spaces. We construct reproducing kernel Hilbert spaces whose kernels are fully determined by the distance functions of the metric spaces, then leverage the inherent structures of these spaces to define a nonlinear SDR framework. We adapt the classical sliced inverse regression of \citet{Li:1991} within this framework for the metric space data. We build the estimator based on the corresponding linear operators, and show it recovers the regression information unbiasedly. We derive the estimator at both the operator level and under a coordinate system, and also establish its convergence rate. We illustrate the proposed method with both synthetic and real datasets exhibiting non-Euclidean geometry.

preprint2020arXiv

Latent Model Extreme Value Index Estimation

We propose a novel strategy for multivariate extreme value index estimation. In applications such as finance, volatility and risk present in the components of a multivariate time series are often driven by the same underlying factors, such as the subprime crisis in the US. To estimate the latent risk, we apply a two-stage procedure. First, a set of independent latent series is estimated using a method of latent variable analysis. Then, univariate risk measures are estimated individually for the latent series to assess their contribution to the overall risk. As our main theoretical contribution, we derive conditions under which the effect of the first step to the asymptotic behavior of the risk estimators is negligible. Simulations demonstrate the theory under both i.i.d. and dependent data, and an application into financial data illustrates the usefulness of the method in extracting joint sources of risk in practice.

preprint2020arXiv

On the behavior of extreme $d$-dimensional spatial quantiles under minimal assumptions

"Spatial" or "geometric" quantiles are the only multivariate quantiles coping with both high-dimensional data and functional data, also in the framework of multiple-output quantile regression. This work studies spatial quantiles in the finite-dimensional case, where the spatial quantile $μ_{α,u}(P)$ of the distribution $P$ taking values in $\mathbb{R}^d $ is a point in $\mathbb{R}^d$ indexed by an order $α\in[0,1)$ and a direction $u$ in the unit sphere $\mathcal{S}^{d-1}$ of $\mathbb{R}^d$ --- or equivalently by a vector $αu$ in the open unit ball of $\mathbb{R}^d$. Recently, Girard and Stupfler (2017) proved that (i) the extreme quantiles $μ_{α,u}(P)$ obtained as $α\to 1$ exit all compact sets of $\mathbb{R}^d$ and that (ii) they do so in a direction converging to $u$. These results help understanding the nature of these quantiles: the first result is particularly striking as it holds even if $P$ has a bounded support, whereas the second one clarifies the delicate dependence of spatial quantiles on $u$. However, they were established under assumptions imposing that $P$ is non-atomic, so that it is unclear whether they hold for empirical probability measures. We improve on this by proving these results under much milder conditions, allowing for the sample case. This prevents using gradient condition arguments, which makes the proofs very challenging. We also weaken the well-known sufficient condition for uniqueness of finite-dimensional spatial quantiles.

preprint2019arXiv

Spatial Blind Source Separation

Recently a blind source separation model was suggested for spatial data together with an estimator based on the simultaneous diagonalisation of two scatter matrices. The asymptotic properties of this estimator are derived here and a new estimator, based on the joint diagonalisation of more than two scatter matrices, is proposed. The asymptotic properties and merits of the novel estimator are verified in simulation studies. A real data example illustrates the method.

preprint2017arXiv

Independent component analysis for multivariate functional data

We extend two methods of independent component analysis, fourth order blind identification and joint approximate diagonalization of eigen-matrices, to vector-valued functional data. Multivariate functional data occur naturally and frequently in modern applications, and extending independent component analysis to this setting allows us to distill important information from this type of data, going a step further than the functional principal component analysis. To allow the inversion of the covariance operator we make the assumption that the dependency between the component functions lies in a finite-dimensional subspace. In this subspace we define fourth cross-cumulant operators and use them to construct the two novel, Fisher consistent methods for solving the independent component problem for vector-valued functions. Both simulations and an application on a hand gesture data set show the usefulness and advantages of the proposed methods over functional principal component analysis.

preprint2016arXiv

Projection Pursuit for non-Gaussian Independent Components

In independent component analysis it is assumed that the observed random variables are linear combinations of latent, mutually independent random variables called the independent components. Our model further assumes that only the non-Gaussian independent components are of interest, the Gaussian components being treated as noise. In this paper projection pursuit is used to extract the non-Gaussian components and to separate the corresponding signal and noise subspaces. Our choice for the projection index is a convex combination of squared third and fourth cumulants and we estimate the non-Gaussian components either one-by-one (deflation-based approach) or simultaneously (symmetric approach). The properties of both estimates are considered in detail through the corresponding optimization problems, estimating equations, algorithms and asymptotic properties. Various comparisons of the estimates show that the two approaches separate the signal and noise subspaces equally well but the symmetric one is generally better in extracting the individual non-Gaussian components.

preprint2015arXiv

Joint Use of Third and Fourth Cumulants in Independent Component Analysis

The independent component model is a latent variable model where the components of the observed random vector are linear combinations of latent independent variables. The aim is to find an estimate for a transformation matrix back to independent components. In moment-based approaches third cumulants are often neglected in favor of fourth cumulants, even though both approaches have similar appealing properties. This paper considers the joint use of third and fourth cumulants in finding independent components. First, univariate cumulants are used as projection indices in search for independent components (projection pursuit). Second, multivariate cumulant matrices are jointly used to solve the problem. The properties of the estimates are considered in detail through corresponding optimization problems, estimating equations, algorithms and asymptotic statistical properties. Comparisons of the asymptotic variances of different estimates in wide independent component models show that in most cases symmetric projection pursuit approach using both third and fourth squared cumulants is a safe choice.

preprint2015arXiv

The squared symmetric FastICA estimator

In this paper we study the theoretical properties of the deflation-based FastICA method, the original symmetric FastICA method, and a modified symmetric FastICA method, here called the squared symmetric FastICA. This modification is obtained by replacing the absolute values in the FastICA objective function by their squares. In the deflation-based case this replacement has no effect on the estimate since the maximization problem stays the same. However, in the symmetric case a novel estimate with unknown properties is obtained. In the paper we review the classic deflation-based and symmetric FastICA approaches and contrast these with the new squared symmetric version of FastICA. We find the estimating equations and derive the asymptotical properties of the squared symmetric FastICA estimator with an arbitrary choice of nonlinearity. Asymptotic variances of the unmixing matrix estimates are then used to compare their efficiencies for large sample sizes showing that the squared symmetric FastICA estimator outperforms the other two estimators in a wide variety of situations.