Source author record

Yi-Hui Zhou

Yi-Hui Zhou appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2026arXiv

Generative Conditional Missing Imputation Networks

In this study, we introduce a sophisticated generative conditional strategy designed to impute missing values within datasets, an area of considerable importance in statistical analysis. Specifically, we initially elucidate the theoretical underpinnings of the Generative Conditional Missing Imputation Networks (GCMI), demonstrating its robust properties in the context of the Missing Completely at Random (MCAR) and the Missing at Random (MAR) mechanisms. Subsequently, we enhance the robustness and accuracy of GCMI by integrating a multiple imputation framework using a chained equations approach. This innovation serves to bolster model stability and improve imputation performance significantly. Finally, through a series of meticulous simulations and empirical assessments utilizing benchmark datasets, we establish the superior efficacy of our proposed methods when juxtaposed with other leading imputation techniques currently available. This comprehensive evaluation not only underscores the practicality of GCMI but also affirms its potential as a leading-edge tool in the field of statistical data analysis.

preprint2020arXiv

Fast Multivariate Probit Estimation via a Two-Stage Composite Likelihood

The multivariate probit is popular for modeling correlated binary data, with an attractive balance of flexibility and simplicity. However, considerable challenges remain in computation and in devising a clear statistical framework. Interest in the multivariate probit has increased in recent years. Current applications include genomics and precision medicine, where simultaneous modeling of multiple traits may be of interest, and computational efficiency is an important consideration. We propose a fast method for multivariate probit estimation via a two-stage composite likelihood. We explore computational and statistical efficiency, and note that the approach sets the stage for extensions beyond the purely binary setting.

preprint2020arXiv

Reduced Rank Multivariate Kernel Ridge Regression

In the multivariate regression, also referred to as multi-task learning in machine learning, the goal is to recover a vector-valued function based on noisy observations. The vector-valued function is often assumed to be of low rank. Although the multivariate linear regression is extensively studied in the literature, a theoretical study on the multivariate nonlinear regression is lacking. In this paper, we study reduced rank multivariate kernel ridge regression, proposed by \cite{mukherjee2011reduced}. We prove the consistency of the function predictor and provide the convergence rate. An algorithm based on nuclear norm relaxation is proposed. A few numerical examples are presented to show the smaller mean squared prediction error comparing with the elementwise univariate kernel ridge regression.

preprint2016arXiv

A robust covariance testing approach for high-throughput data

The problem of testing changes in covariance has received increasing attention in recent years, especially in the context of high-dimensional testing. A number of approaches have been proposed, all limited to the two-sample problem and involving varying statistics and assumptions on the number of features $p$ vs. the sample size $n$. There are no general approaches to test association of covariances with a continuous outcome. We propose a uniform framework for testing association of covariances with an experimental variable, whether discrete or continuous. The approach is not limited by the data dimensions. Our test procedure (i) does not rely on parametric assumptions, (ii) works well for a range of $p$ and $n$ (e.g., does not require $n > p$), (iii) provides correct type I error control, and (iv) includes four different statistics, to ensure power and flexibility under various settings, including a new "connectivity" statistic that is sensitive to changes in overall covariance magnitude. We demonstrate that, for the two-sample special case, the proposed statistics are permutationally equivalent or similar to existing proposed statistics. We demonstrate the power and utility of our approaches via simulation and analysis of real data. The approach is implemented in an $R$ package.

preprint2016arXiv

Computation of ancestry scores with mixed families and unrelated individuals

The issue of robustness to family relationships in computing genotype ancestry scores such as eigenvector projections has received increased attention in genetic association, as the scores are widely used to control spurious association. We use a motivational example from the North American Cystic Fibrosis (CF) Consortium genetic association study with 3444 individuals and 898 family members to illustrate the challenge of computing ancestry scores when sets of both unrelated individuals and closely-related family members are included. We propose novel methods to obtain ancestry scores and demonstrate that the proposed methods outperform existing methods. The current standard is to compute loadings (left singular vectors) using unrelated individuals and to compute projected scores for remaining family members. However, projected ancestry scores from this approach suffer from shrinkage toward zero. We consider in turn alternate strategies: (i) within-family data orthogonalization, (ii) matrix substitution based on decomposition of a target family-orthogonalized covariance matrix, (iii) covariance-preserving whitening, retaining covariances between unrelated pairs while orthogonalizing family members, and (iv) using family-averaged data to obtain loadings. Except for within-family orthogonalization, our proposed approaches offer similar performance and are superior to the standard approaches. We illustrate the performance via simulation and analysis of the CF dataset.

preprint2014arXiv

Hypothesis testing at the extremes: fast and robust association for high-throughput data

A number of biomedical problems require performing many hypothesis tests, with an attendant need to apply stringent thresholds. Often the data take the form of a series of predictor vectors, each of which must be compared with a single response vector, perhaps with nuisance covariates. Parametric tests of association are often used, but can result in inaccurate type I error at the extreme thresholds, even for large sample sizes. Furthermore, standard two-sided testing can reduce power compared to the doubled $p$-value, due to asymmetry in the null distribution. Exact (permutation) testing is attractive, but can be computationally intensive and cumbersome. We present an approximation to exact association tests of trend that is accurate and fast enough for standard use in high-throughput settings, and can easily provide standard two-sided or doubled $p$-values. The approach is shown to be equivalent under permutation to likelihood ratio tests for the most commonly used generalized linear models. For linear regression, covariates are handled by working with covariate-residualized responses and predictors. For generalized linear models, stratified covariates can be handled in a manner similar to exact conditional testing. Simulations and examples illustrate the wide applicability of the approach.