Source author record

Juntong Chen

Juntong Chen appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
3topics
1close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2022arXiv

Estimating a regression function in exponential families by model selection

Let $X_{1}=(W_{1},Y_{1}),\ldots,X_{n}=(W_{n},Y_{n})$ be $n$ pairs of independent random variables. We assume that, for each $i\in\{1,\ldots,n\}$, the conditional distribution of $Y_{i}$ given $W_{i}$ belongs to a one-parameter exponential family with parameter ${\boldsymbolγ}^{\star}(W_{i})\in{\mathbb{R}}$, or at least, is close enough to a distribution of this form. The objective of the present paper is to estimate these conditional distributions on the basis of the observation ${\boldsymbol{X}}=(X_{1},\ldots,X_{n})$ and to do so, we propose a model selection procedure together with a non-asymptotic risk bound for the resulted estimator with respect to a Hellinger-type distance. When ${\boldsymbolγ}^{\star}$ does exist, the procedure allows to obtain an estimator $\widehat{\boldsymbolγ}$ of ${\boldsymbolγ}^{\star}$ adapted to a wide range of the anisotropic Besov spaces. When ${\boldsymbolγ}^{\star}$ has a general additive or multiple index structure, we construct suitable models and show the resulted estimators by our procedure based on such models can circumvent the curse of dimensionality. Moreover, we consider model selection problems for ReLU neural networks and provide an example where estimation based on neural networks enjoys a much faster converge rate than the classical models. Finally, we apply this procedure to solve variable selection problem in exponential families. The proofs in the paper rely on bounding the VC dimensions of several collections of functions, which can be of independent interest.

preprint2022arXiv

Robust estimation of a regression function in exponential families

We observe $n$ pairs of independent (but not necessarily i.i.d.) random variables $X_{1}=(W_{1},Y_{1}),\ldots,X_{n}=(W_{n},Y_{n})$ and tackle the problem of estimating the conditional distributions $Q_{i}^{\star}(w_{i})$ of $Y_{i}$ given $W_{i}=w_{i}$ for all $i\in\{1,\ldots,n\}$. Even though these might not be true, we base our estimator on the assumptions that the data are i.i.d.\ and the conditional distributions of $Y_{i}$ given $W_{i}=w_{i}$ belong to a one parameter exponential family $\bar{\mathscr{Q}}$ with parameter space given by an interval $I$. More precisely, we pretend that these conditional distributions take the form $Q_{{\boldsymbolθ}(w_{i})}\in \bar{\mathscr{Q}}$ for some ${\boldsymbolθ}$ that belongs to a VC-class $\bar{\boldsymbolΘ}$ of functions with values in $I$. For each $i\in\{1,\ldots,n\}$, we estimate $Q_{i}^{\star}(w_{i})$ by a distribution of the same form, i.e.\ $Q_{\hat{\boldsymbolθ}(w_{i})}\in \bar{\mathscr{Q}}$, where $\hat {\boldsymbolθ}=\hat {\boldsymbolθ}(X_{1},\ldots,X_{n})$ is a well-chosen estimator with values in $\bar{\boldsymbolΘ}$. We show that our estimation strategy is robust to model misspecification, contamination and the presence of outliers. Besides, we provide an algorithm for calculating $\hat{\boldsymbolθ}$ when $\bar{\boldsymbolΘ}$ is a VC-class of functions of low or moderate dimension and we carry out a simulation study to compare the performance of $\hat{\boldsymbolθ}$ to that of the MLE and median-based estimators.