Source author record

Dennis Leung

Dennis Leung appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2026arXiv

Berry-Esseen theorems for the asymptotic normality of incomplete U-statistics with Bernoulli sampling

There has been a resurgence of interest in incomplete U-statistics that only sum over a subset of kernel evaluations, due to their computational efficiency and asymptotic normality which can be leveraged to quantify the uncertainty of ensemble predictions in machine learning. In this paper, we study the weak convergences to normality of one such construction, the incomplete U-statistic with Bernoulli sampling, under three different regimes on the relative sizes of the raw sample and the computational budget. Under minimalistic moment assumptions, we establish accompanying Berry-Esseen bounds with the natural rates that characterize the accuracy of these normal approximations. The key ingredients in our proofs include a variable censoring technique and a methodology for establishing Berry-Esseen bounds for the so-called Studentized nonlinear statistics recently formalized in the Stein's method literature, as well as an exponential lower tail bound for non-negative kernel U-statistics.

preprint2024arXiv

Singularity-agnostic incomplete U-statistics for testing polynomial constraints in Gaussian covariance matrices

Testing the goodness-of-fit of a model with its defining functional constraints in the parameters could date back to Spearman (1927), who analyzed the famous "tetrad" polynomial in the covariance matrix of the observed variables in a single-factor model. Despite its long history, the Wald test typically employed to operationalize this approach could produce very inaccurate test sizes in many situations, even when the regular conditions for the classical normal asymptotics are met and a very large sample is available. Focusing on testing a polynomial constraint in a Gaussian covariance matrix, we obtained a new understanding of this baffling phenomenon: When the null hypothesis is true but "near-singular", the standardized Wald test exhibits slow weak convergence, owing to the sophisticated dependency structure inherent to the underlying U-statistic that ultimately drives its limiting distribution; this can also be rigorously explained by a key ratio of moments encoded in the Berry-Esseen bound quantifying the normal approximation error involved. As an alternative, we advocate the use of an incomplete U-statistic to mildly tone down the dependence thereof and render the speed of convergence agnostic to the singularity status of the hypothesis. In parallel, we develop a Berry-Esseen bound that is mathematically descriptive of the singularity-agnostic nature of our standardized incomplete U-statistic, using some of the finest exponential-type inequalities in the literature.

preprint2016arXiv

Testing independence in high dimensions with sums of rank correlations

We treat the problem of testing independence between m continuous variables when m can be larger than the available sample size n. We consider three types of test statistics that are constructed as sums or sums of squares of pairwise rank correlations. In the asymptotic regime where both m and n tend to infinity, a martingale central limit theorem is applied to show that the null distributions of these statistics converge to Gaussian limits, which are valid with no specific distributional or moment assumptions on the data. Using the framework of U-statistics, our result covers a variety of rank correlations including Kendall's tau and a dominating term of Spearman's rank correlation coefficient (rho), but also degenerate U-statistics such as Hoeffding's $D$, or the $τ^*$ of Bergsma and Dassios (2014). As in the classical theory for U-statistics, the test statistics need to be scaled differently when the rank correlations used to construct them are degenerate U-statistics. The power of the considered tests is explored in rate-optimality theory under Gaussian equicorrelation alternatives as well as in numerical experiments for specific cases of more general alternatives.

preprint2015arXiv

Efficient Computation of the Bergsma-Dassios Sign Covariance

In an extension of Kendall's $τ$, Bergsma and Dassios (2014) introduced a covariance measure $τ^*$ for two ordinal random variables that vanishes if and only if the two variables are independent. For a sample of size $n$, a direct computation of $t^*$, the empirical version of $τ^*$, requires $O(n^4)$ operations. We derive an algorithm that computes the statistic using only $O(n^2\log(n))$ operations.

preprint2015arXiv

Identifiability of directed Gaussian graphical models with one latent source

We study parameter identifiability of directed Gaussian graphical models with one latent variable. In the scenario we consider, the latent variable is a confounder that forms a source node of the graph and is a parent to all other nodes, which correspond to the observed variables. We give a graphical condition that is sufficient for the Jacobian matrix of the parametrization map to be full rank, which entails that the parametrization is generically finite-to-one, a fact that is sometimes also referred to as local identifiability. We also derive a graphical condition that is necessary for such identifiability. Finally, we give a condition under which generic parameter identifiability can be determined from identifiability of a model associated with a subgraph. The power of these criteria is assessed via an exhaustive algebraic computational study on models with 4, 5, and 6 observable variables.

preprint2014arXiv

Order-invariant prior specification in Bayesian factor analysis

In (exploratory) factor analysis, the loading matrix is identified only up to orthogonal rotation. For identifiability, one thus often takes the loading matrix to be lower triangular with positive diagonal entries. In Bayesian inference, a standard practice is then to specify a prior under which the loadings are independent, the off-diagonal loadings are normally distributed, and the diagonal loadings follow a truncated normal distribution. This prior specification, however, depends in an important way on how the variables and associated rows of the loading matrix are ordered. We show how a minor modification of the approach allows one to compute with the identifiable lower triangular loading matrix but maintain invariance properties under reordering of the variables.