Source author record

Xiongzhi Chen

Xiongzhi Chen appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

8works
7topics
3close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

8 published item(s)

preprint2022arXiv

Consistent estimation of the proportion of false nulls and FDR for adaptive multiple testing Normal means under weak dependence

We consider multiple testing means of many dependent Normal random variables that do not necessarily follow a joint Normal distribution. Under weak dependence, we show the uniform consistency of proportion estimators that are constructed as solutions to Lebesgue-Stieltjes equations for the setting of a point, bounded and one-sided null, respectively, and characterize via the index of weak dependence the sparsest proportion these estimators can consistently estimate. On the other hand, under a principal correlation structure and employing a suitable definition of p-value for composite null hypotheses, we show that three key empirical processes induced by a single-step multiple testing procedure (MTP) satisfy the strong law of large numbers for testing each of the three types of nulls. Further, under this structure and for testing a point null and a one-sided null respectively, we construct an adaptive single-step MTP that employs a proportion estimator mentioned earlier, and show that the false discovery proportion of this procedure satisfies the weak law of large numbers and hence consistently estimates the false discovery rate of the procedure. In addition, we report some findings on the estimators of Jin and of Meinshausen and Rice of the proportion of false nulls in the critically and very sparse regimes under weak dependence and model misspecifications, respectively.

preprint2019arXiv

A strong law of large numbers related to multiple testing Normal means

Assessing the stability of a multiple testing procedure under dependence is important but very challenging. Even for multiple testing which among a set of Normal random variables have mean zero, which we refer to as the "Normal means problem", to date there lacks a classification of the type of dependence under which the strong law of large numbers (SLLN) holds for the numbers of rejections and false rejections. We introduce the concept of "principal correlation structure (PCS)" that characterizes the type of dependence for which such SLLN holds, and establish the law. Further, we show that PCS ensures the SLLN for the false discover proportion when there is always a positive proportion of zero Normal means. We also investigate the stability of two conditional multiple testing procedures for the Normal means problem, and show that the associated SLLN holds when in addition the decomposition of the covariance matrix of the Normal random variables that induces PCS is homogeneous in certain sense. Our results also provide a formal way to check if the "weak dependence" assumption, a widely used assumption in the multiple testing literature, holds for the Normal means problem. As by-products, we establish a universal bound on Hermite polynomials and a universal comparison result on the covariance of the indicator functions of the two p-values of testing the marginal means of a bivariate Normal random vector and the correlation between the two components of the vector. These are of their own interests.

preprint2019arXiv

Uniformly consistently estimating the proportion of false null hypotheses via Lebesgue-Stieltjes integral equations

The proportion of false null hypotheses is a very important quantity in statistical modelling and inference based on the two-component mixture model and its extensions, and in control and estimation of the false discovery rate and false non-discovery rate. Most existing estimators of this proportion threshold p-values, deconvolve the mixture model under constraints on its components, or depend heavily on the location-shift property of distributions. Hence, they usually are not consistent, applicable to non-location-shift distributions, or applicable to discrete statistics or p-values. To eliminate these shortcomings, we construct uniformly consistent estimators of the proportion as solutions to Lebesgue-Stieltjes integral equations. In particular, we provide such estimators respectively for random variables whose distributions have Riemann-Lebesgue type characteristic functions, form discrete natural exponential families with infinite supports, and form natural exponential families with separable moment sequences. We provide the speed of convergence and uniform consistency class for each such estimator under independence. In addition, we provide example distribution families for which a consistent estimator of the proportion cannot be constructed using our techniques.

preprint2018arXiv

False discovery rate control for multiple testing based on p-values with càdlàg distribution functions

For multiple testing based on p-values with càdlàg distribution functions, we propose an FDR procedure "BH+" with proven conservativeness. BH+ is at least as powerful as the BH procedure when they are applied to super-uniform p-values. Further, when applied to mid p-values, BH+ is more powerful than it is applied to conventional p-values. An easily verifiable necessary and sufficient condition for this is provided. BH+ is perhaps the first conservative FDR procedure applicable to mid p-values. BH+ is applied to multiple testing based on discrete p-values in a methylation study, an HIV study and a clinical safety study, where it makes considerably more discoveries than the BH procedure.

preprint2016arXiv

Stopping time property of thresholds of Storey-type FDR procedures

For multiple testing, we introduce Storey-type FDR procedures and the concept of "regular estimator of the proportion of true nulls". We show that the rejection threshold of a Storey-type FDR procedure is a stopping time with respect to the backward filtration generated by the p-values and that a Storey-type FDR estimator at this rejection threshold equals the pre-specified FDR level, when the estimator of the proportion of true nulls is regular. These results hold regardless of the dependence among or the types of distributions of the p-values. They directly imply that a Storey-type FDR procedure is conservative when the null p-values are independent and uniformly distributed.

preprint2015arXiv

Consistent Estimation of Low-Dimensional Latent Structure in High-Dimensional Data

We consider the problem of extracting a low-dimensional, linear latent variable structure from high-dimensional random variables. Specifically, we show that under mild conditions and when this structure manifests itself as a linear space that spans the conditional means, it is possible to consistently recover the structure using only information up to the second moments of these random variables. This finding, specialized to one-parameter exponential families whose variance function is quadratic in their means, allows for the derivation of an explicit estimator of such latent structure. This approach serves as a latent variable model estimator and as a tool for dimension reduction for a high-dimensional matrix of data composed of many related variables. Our theoretical results are verified by simulation studies and an application to genomic data.

preprint2015arXiv

Explicit solutions to a vector time series model and its induced model for business cycles

This article gives the explicit solution to a general vector time series model that describes interacting, heterogeneous agents that operate under uncertainties but according to Keynesian principles, from which a model for business cycle is induced by a weighted average of the growth rates of the agents in the model. The explicit solution enables a direct simulation of the time series defined by the model and better understanding of the joint behavior of the growth rates. In addition, the induced model for business cycles and its solutions are explicitly given and analyzed. The explicit solutions provide a better understanding of the mathematics of these models and the econometric properties they try to incorporate.