Source author record

Sihai D. Zhao

Sihai D. Zhao appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
1topics
2close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2014arXiv

Nonparametric empirical Bayes and maximum likelihood estimation for high-dimensional data analysis

Nonparametric empirical Bayes methods provide a flexible and attractive approach to high-dimensional data analysis. One particularly elegant empirical Bayes methodology, involving the Kiefer-Wolfowitz nonparametric maximum likelihood estimator (NPMLE) for mixture models, has been known for decades. However, implementation and theoretical analysis of the Kiefer-Wolfowitz NPMLE are notoriously difficult. A fast algorithm was recently proposed that makes NPMLE-based procedures feasible for use in large-scale problems, but the algorithm calculates only an approximation to the NPMLE. In this paper we make two contributions. First, we provide upper bounds on the convergence rate of the approximate NPMLE's statistical error, which have the same order as the best known bounds for the true NPMLE. This suggests that the approximate NPMLE is just as effective as the true NPMLE for statistical applications. Second, we illustrate the promise of NPMLE procedures in a high-dimensional binary classification problem. We propose a new procedure and show that it vastly outperforms existing methods in experiments with simulated data. In real data analyses involving cancer survival and gene expression data, we show that it is very competitive with several recently proposed methods for regularized linear discriminant analysis, another popular approach to high-dimensional classification.

preprint2012arXiv

Sure screening for estimating equations in ultra-high dimensions

As the number of possible predictors generated by high-throughput experiments continues to increase, methods are needed to quickly screen out unimportant covariates. Model-based screening methods have been proposed and theoretically justified, but only for a few specific models. Model-free screening methods have also recently been studied, but can have lower power to detect important covariates. In this paper we propose EEScreen, a screening procedure that can be used with any model that can be fit using estimating equations, and provide unified results on its finite-sample screening performance. EEScreen thus generalizes many recently proposed model-based and model-free screening procedures. We also propose iEEScreen, an iterative version of EEScreen, and show that it is closely related to a recently studied boosting method for estimating equations. We show via simulations for two different estimating equations that EEScreen and iEEScreen are useful and flexible screening procedures, and demonstrate our methods on data from a multiple myeloma study.