Source author record

Antonio Cuevas

Antonio Cuevas appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

9works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

9 published item(s)

preprint2020arXiv

Set Estimation Under Biconvexity Restrictions

A set in the Euclidean plane is said to be biconvex if, for some angle $θ\in[0,π/2)$, all its sections along straight lines with inclination angles $θ$ and $θ+π/2$ are convex sets (i.e, empty sets or segments). Biconvexity is a natural notion with some useful applications in optimization theory. It has also be independently used, under the name of "rectilinear convexity", in computational geometry. We are concerned here with the problem of asymptotically reconstructing (or estimating) a biconvex set $S$ from a random sample of points drawn on $S$. By analogy with the classical convex case, one would like to define the "biconvex hull" of the sample points as a natural estimator for $S$. However, as previously pointed out by several authors, the notion of "hull" for a given set $A$ (understood as the "minimal" set including $A$ and having the required property) has no obvious, useful translation to the biconvex case. This is in sharp contrast with the well-known elementary definition of convex hull. Thus, we have selected the most commonly accepted notion of "biconvex hull" (often called "rectilinear convex hull"): we first provide additional motivations for this definition, proving some useful relations with other convexity-related notions. Then, we prove some results concerning the consistent approximation of a biconvex set $S$ and and the corresponding biconvex hull. An analogous result is also provided for the boundaries. A method to approximate, from a sample of points on $S$, the biconvexity angle $θ$ is also given.

preprint2016arXiv

On the use of reproducing kernel Hilbert spaces in functional classification

The Hájek-Feldman dichotomy establishes that two Gaussian measures are either mutually absolutely continuous with respect to each other (and hence there is a Radon-Nikodym density for each measure with respect to the other one) or mutually singular. Unlike the case of finite dimensional Gaussian measures, there are non-trivial examples of both situations when dealing with Gaussian stochastic processes. This paper provides: (a) Explicit expressions for the optimal (Bayes) rule and the minimal classification error probability in several relevant problems of supervised binary classification of mutually absolutely continuous Gaussian processes. The approach relies on some classical results in the theory of Reproducing Kernel Hilbert Spaces (RKHS). (b) An interpretation, in terms of mutual singularity, for the "near perfect classification" phenomenon described by Delaigle and Hall (2012). We show that the asymptotically optimal rule proposed by these authors can be identified with the sequence of optimal rules for an approximating sequence of classification problems in the absolutely continuous case. (c) A new model-based method for variable selection in binary classification problems, which arises in a very natural way from the explicit knowledge of the RN-derivatives and the underlying RKHS structure. Different classifiers might be used from the selected variables. In particular, the classical, linear finite-dimensional Fisher rule turns out to be consistent under some standard conditions on the underlying functional model.

preprint2015arXiv

On visual distances for spectrum-type functional data

A functional distance ${\mathbb H}$, based on the Hausdorff metric between the function hypographs, is proposed for the space ${\mathcal E}$ of non-negative real upper semicontinuous functions on a compact interval. The main goal of the paper is to show that the space $({\mathcal E},{\mathbb H})$ is particularly suitable in some statistical problems with functional data which involve functions with very wiggly graphs and narrow, sharp peaks. A typical example is given by spectrograms, either obtained by magnetic resonance or by mass spectrometry. On the theoretical side, we show that $({\mathcal E},{\mathbb H})$ is a complete, separable locally compact space and that the ${\mathbb H}$-convergence of a sequence of functions implies the convergence of the respective maximum values of these functions. The probabilistic and statistical implications of these results are discussed in particular, regarding the consistency of $k$-NN classifiers for supervised classification problems with functional data in ${\mathbb H}$. On the practical side, we provide the results of a small simulation study and check also the performance of our method in two real data problems of supervised classification involving mass spectra.

preprint2015arXiv

The mRMR variable selection method: a comparative study for functional data

The use of variable selection methods is particularly appealing in statistical problems with functional data. The obvious general criterion for variable selection is to choose the `most representative' or `most relevant' variables. However, it is also clear that a purely relevance-oriented criterion could lead to select many redundant variables. The mRMR (minimum Redundance Maximum Relevance) procedure, proposed by Ding and Peng (2005) and Peng et al. (2005) is an algorithm to systematically perform variable selection, achieving a reasonable trade-off between relevance and redundancy. In its original form, this procedure is based on the use of the so-called mutual information criterion to assess relevance and redundancy. Keeping the focus on functional data problems, we propose here a modified version of the mRMR method, obtained by replacing the mutual information by the new association measure (called distance correlation) suggested by Székely et al. (2007). We have also performed an extensive simulation study, including 1600 functional experiments (100 functional models $\times$ 4 sample sizes $\times$ 4 classifiers) and three real-data examples aimed at comparing the different versions of the mRMR methodology. The results are quite conclusive in favor of the new proposed alternative.

preprint2015arXiv

Variable selection in functional data classification: a maxima-hunting proposal

Variable selection is considered in the setting of supervised binary classification with functional data $\{X(t),\ t\in[0,1]\}$. By "variable selection" we mean any dimension-reduction method which leads to replace the whole trajectory $\{X(t),\ t\in[0,1]\}$, with a low-dimensional vector $(X(t_1),\ldots,X(t_k))$ still keeping a similar classification error. Our proposal for variable selection is based on the idea of selecting the local maxima $(t_1,\ldots,t_k)$ of the function ${\mathcal V}_X^2(t)={\mathcal V}^2(X(t),Y)$, where ${\mathcal V}$ denotes the "distance covariance" association measure for random variables due to Székely, Rizzo and Bakirov (2007). This method provides a simple natural way to deal with the relevance vs. redundancy trade-off which typically appears in variable selection. This paper includes (a) Some theoretical motivation: a result of consistent estimation on the maxima of ${\mathcal V}_X^2$ is shown. We also show different theoretical models for the underlying process $X(t)$ under which the relevant information in concentrated in the maxima of ${\mathcal V}_X^2$. (b) An extensive empirical study, including about 400 simulated models and real data examples, aimed at comparing our variable selection method with other standard proposals for dimension reduction.

preprint2014arXiv

A geometrically motivated parametric model in manifold estimation,

The general aim of manifold estimation is reconstructing, by statistical methods, an $m$-dimensional compact manifold $S$ on ${\mathbb R}^d$ (with $m\leq d$) or estimating some relevant quantities related to the geometric properties of $S$. We will assume that the sample data are given by the distances to the $(d-1)$-dimensional manifold $S$ from points randomly chosen on a band surrounding $S$, with $d=2$ and $d=3$. The point in this paper is to show that, if $S$ belongs to a wide class of compact sets (which we call \it sets with polynomial volume\rm), the proposed statistical model leads to a relatively simple parametric formulation. In this setup, standard methodologies (method of moments, maximum likelihood) can be used to estimate some interesting geometric parameters, including curvatures and Euler characteristic. We will particularly focus on the estimation of the $(d-1)$-dimensional boundary measure (in Minkowski's sense) of $S$. It turns out, however, that the estimation problem is not straightforward since the standard estimators show a remarkably pathological behavior: while they are consistent and asymptotically normal, their expectations are infinite. The theoretical and practical consequences of this fact are discussed in some detail.

preprint2014arXiv

On Poincaré cone property

A domain $S\subset{\mathbb{R}}^d$ is said to fulfill the Poincaré cone property if any point in the boundary of $S$ is the vertex of a (finite) cone which does not otherwise intersects the closure $\bar{S}$. For more than a century, this condition has played a relevant role in the theory of partial differential equations, as a shape assumption aimed to ensure the existence of a solution for the classical Dirichlet problem on $S$. In a completely different setting, this paper is devoted to analyze some statistical applications of the Poincaré cone property (when defined in a slightly stronger version). First, we show that this condition can be seen as a sort of generalized convexity: while it is considerably less restrictive than convexity, it still retains some ``convex flavour.'' In particular, when imposed to a probability support $S$, this property allows the estimation of $S$ from a random sample of points, using the ``hull principle'' much in the same way as a convex support is estimated using the convex hull of the sample points. The statistical properties of such hull estimator (consistency, convergence rates, boundary estimation) are considered in detail. Second, it is shown that the class of sets fulfilling the Poincaré property is a $P$-Glivenko-Cantelli class for any absolutely continuous distribution $P$ on $\mathbb{R}^d$. This has some independent interest in the theory of empirical processes, since it extends the classical analogous result, established for convex sets, to a much larger class. Third, an algorithm to approximate the cone-convex hull of a finite sample of points is proposed and some practical illustrations are given.

preprint2011arXiv

On statistical properties of sets fulfilling rolling-type conditions

Motivated by set estimation problems, we consider three closely related shape conditions for compact sets: positive reach, r-convexity and rolling condition. First, the relations between these shape conditions are analyzed. Second, we obtain for the estimation of sets fulfilling a rolling condition a result of "full consistency" (i.e., consistency with respect to the Hausdorff metric for the target set and for its boundary). Third, the class of uniformly bounded compact sets whose reach is not smaller than a given constant r is shown to be a P-uniformity class (in Billingsley and Topsoe's (1967) sense) and, in particular, a Glivenko-Cantelli class. Fourth, under broad conditions, the r-convex hull of the sample is proved to be a fully consistent estimator of an r-convex support in the two-dimensional case. Moreover, its boundary length is shown to converge (a.s.) to that of the underlying support. Fifth, the above results are applied to get new consistency statements for level set estimators based on the excess mass methodology (Polonik, 1995).

preprint2010arXiv

Supervised classification for a family of Gaussian functional models

In the framework of supervised classification (discrimination) for functional data, it is shown that the optimal classification rule can be explicitly obtained for a class of Gaussian processes with "triangular" covariance functions. This explicit knowledge has two practical consequences. First, the consistency of the well-known nearest neighbors classifier (which is not guaranteed in the problems with functional data) is established for the indicated class of processes. Second, and more important, parametric and nonparametric plug-in classifiers can be obtained by estimating the unknown elements in the optimal rule. The performance of these new plug-in classifiers is checked, with positive results, through a simulation study and a real data example.