Source author record

Tatjana Pavlenko

Tatjana Pavlenko appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
3topics
3close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2016arXiv

A $U$-classifier for high-dimensional data under non-normality

A classifier for two or more samples is proposed when the data are high-dimensional and the underlying distributions may be non-normal. The classifier is constructed as a linear combination of two easily computable and interpretable components, the $U$-component and the $P$-component. The $U$-component is a linear combination of $U$-statistics which are averages of bilinear forms of pairwise distinct vectors from two independent samples. The $P$-component is the discriminant score and is a function of the projection of the $U$-component on the observation to be classified. Combined, the two components constitute an inherently bias-adjusted classifier valid for high-dimensional data. The simplicity of the classifier helps conveniently study its properties, including its asymptotic normal limit, and extend it to multi-sample case. The classifier is linear but its linearity does not rest on the assumption of homoscedasticity. Probabilities of misclassification and asymptotic properties of their empirical versions are discussed in detail. Simulation results are used to show the accuracy of the proposed classifier for sample sizes as small as 5 or 7 and any large dimensions. Applications on real data sets are also demonstrated.

preprint2016arXiv

Goodness-of-fit tests based on sup-functionals of weighted empirical processes

A large class of goodness-of-fit test statistics based on sup-functionals of weighted empirical processes is proposed and studied. The weight functions employed are Erdős-Feller-Kolmogorov-Petrovski upper-class functions of a Brownian bridge. Based on the result of M. Csörgő, S. Csörgő, Horváth, and Mason obtained for this type of test statistics, we provide the asymptotic null distribution theory for the class of tests in hand, and present an algorithm for tabulating the limit distribution functions under the null hypothesis. A new family of nonparametric confidence bands is constructed for the true distribution function and it is found to perform very well. The results obtained, together with a new result on the convergence in distribution of the higher criticism statistic, introduced by Donoho and Jin, demonstrate the advantage of our approach over a common approach that utilizes a family of regularly varying weight functions. Furthermore, we show that, in various subtle problems of detecting sparse heterogeneous mixtures, the proposed test statistics achieve the detection boundary found by Ingster and, when distinguishing between the null and alternative hypotheses, perform optimally adaptively to unknown sparsity and size of the non-null effects.

preprint2012arXiv

Bayesian Network Classifiers in a High Dimensional Framework

We present a growing dimension asymptotic formalism. The perspective in this paper is classification theory and we show that it can accommodate probabilistic networks classifiers, including naive Bayes model and its augmented version. When represented as a Bayesian network these classifiers have an important advantage: The corresponding discriminant function turns out to be a specialized case of a generalized additive model, which makes it possible to get closed form expressions for the asymptotic misclassification probabilities used here as a measure of classification accuracy. Moreover, in this paper we propose a new quantity for assessing the discriminative power of a set of features which is then used to elaborate the augmented naive Bayes classifier. The result is a weighted form of the augmented naive Bayes that distributes weights among the sets of features according to their discriminative power. We derive the asymptotic distribution of the sample based discriminative power and show that it is seriously overestimated in a high dimensional case. We then apply this result to find the optimal, in a sense of minimum misclassification probability, type of weighting.