Source author record

Houying Zhu

Houying Zhu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2026arXiv

2D Stability Selection: Design Jittering for Doubly Stable Feature Selection

We study feature selection in high-dimensional regression under two distinct sources of instability: sampling variability and measurement error in the design matrix. Stability Selection addresses the former through sub-sampling and aggregation, but does not explicitly stress-test robustness to noisy predictors. We introduce doubly stable feature selection, a perturb-and-aggregate framework that targets features whose inclusion is stable both across randomization and across increasing levels of design noise. The method injects controlled additive noise into the design matrix, fits a fixed base selector such as the Lasso on the perturbed data, and aggregates selection frequencies. Sweeping over a grid of noise levels yields a stability path that summarizes robustness to measurement error while using the full sample size and isolating the effect of design perturbations. On the theory side, we show that classical model-selection conditions are preserved under sufficiently small perturbations, with a high-probability extension for Gaussian noise. Empirically, experiments on synthetic and real datasets show improved robustness compared with Stability Selection and standard base selectors.

preprint2016arXiv

A Discrepancy Bound for Deterministic Acceptance-Rejection Samplers Beyond $N^{-1/2}$ in Dimension 1

In this paper we consider an acceptance-rejection (AR) sampler based on deterministic driver sequences. We prove that the discrepancy of an $N$ element sample set generated in this way is bounded by $\mathcal{O} (N^{-2/3}\log N)$, provided that the target density is twice continuously differentiable with non-vanishing curvature and the AR sampler uses the driver sequence $$\mathcal{K}_M= \{( j α, j β) ~~ mod~~1 \mid j = 1,\ldots,M\}, $$ where $α,β$ are real algebraic numbers such that $1,α,β$ is a basis of a number field over $\mathbb{Q}$ of degree $3$. For the driver sequence $$\mathcal{F}_k= \{ ({j}/{F_k}, \{jF_{k-1}/{F_k}\} ) \mid j=1,\ldots, F_k\},$$ where $F_k$ is the $k$-th Fibonacci number and $\{x\}=x-\lfloor x \rfloor$ is the fractional part of a non-negative real number $x$, we can remove the $\log$ factor to improve the convergence rate to $\mathcal{O}(N^{-2/3})$, where again $N$ is the number of samples we accepted. We also introduce a criterion for measuring the goodness of driver sequences. The proposed approach is numerically tested by calculating the star-discrepancy of samples generated for some target densities using $\mathcal{K}_M$ and $\mathcal{F}_k$ as driver sequences. These results confirm that achieving a convergence rate beyond $N^{-1/2}$ is possible in practice using $\mathcal{K}_M$ and $\mathcal{F}_k$ as driver sequences in the acceptance-rejection sampler.

preprint2016arXiv

Discrepancy bounds for uniformly ergodic Markov chain quasi-Monte Carlo

Markov chains can be used to generate samples whose distribution approximates a given target distribution. The quality of the samples of such Markov chains can be measured by the discrepancy between the empirical distribution of the samples and the target distribution. We prove upper bounds on this discrepancy under the assumption that the Markov chain is uniformly ergodic and the driver sequence is deterministic rather than independent $U(0,1)$ random variables. In particular, we show the existence of driver sequences for which the discrepancy of the Markov chain from the target distribution with respect to certain test sets converges with (almost) the usual Monte Carlo rate of $n^{-1/2}$.

preprint2014arXiv

A Discrepancy Bound for a Deterministic Acceptance-Rejection Sampler

We consider an acceptance-rejection sampler based on a deterministic driver sequence. The deterministic sequence is chosen such that the discrepancy between the empirical target distribution and the target distribution is small. We use quasi-Monte Carlo (QMC) point sets for this purpose. The empirical evidence shows convergence rates beyond the crude Monte Carlo rate of $N^{-1/2}$. We prove that the discrepancy of samples generated by the QMC acceptance-rejection sampler is bounded from above by $N^{-1/s}$. A lower bound shows that for any given driver sequence, there always exists a target density such that the star discrepancy is at most $N^{-2/(s+1)}$. For a general density, whose domain is the real state space $\mathbb{R}^{s-1}$, the inverse Rosenblatt transformation can be used to convert samples from the $(s-1)-$dimensional cube to $\mathbb{R}^{s-1}$. We show that this transformation is measure preserving. This way, under certain conditions, we obtain the same convergence rate for a general target density defined in $\mathbb{R}^{s-1}$. Moreover, we also consider a deterministic reduced acceptance-rejection algorithm recently introduced by Barekat and Caflisch [F. Barekat and R.Caflisch. Simulation with Fluctuation and Singular Rates. ArXiv:1310.4555[math.NA], 2013.]

preprint2014arXiv

Discrepancy Estimates for Acceptance-Rejection Samplers Using Stratified Inputs

In this paper we propose an acceptance-rejection sampler using stratified inputs as diver sequence. We estimate the discrepancy of the points generated by this algorithm. First we show an upper bound on the star discrepancy of order $N^{-1/2-1/(2s)}$. Further we prove an upper bound on the $q$-th moment of the $L_q$-discrepancy $(\mathbb{E}[N^{q}L^{q}_{q,N}])^{1/q}$ for $2\le q\le \infty$, which is of order $N^{(1-1/s)(1-1/q)}$. We also present an improved convergence rate for a deterministic acceptance-rejection algorithm using $(t,m,s)-$nets as driver sequence.