Source author record

Xuan Liang

Xuan Liang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2021arXiv

On the Subbagging Estimation for Massive Data

This article introduces subbagging (subsample aggregating) estimation approaches for big data analysis with memory constraints of computers. Specifically, for the whole dataset with size $N$, $m_N$ subsamples are randomly drawn, and each subsample with a subsample size $k_N\ll N$ to meet the memory constraint is sampled uniformly without replacement. Aggregating the estimators of $m_N$ subsamples can lead to subbagging estimation. To analyze the theoretical properties of the subbagging estimator, we adapt the incomplete $U$-statistics theory with an infinite order kernel to allow overlapping drawn subsamples in the sampling procedure. Utilizing this novel theoretical framework, we demonstrate that via a proper hyperparameter selection of $k_N$ and $m_N$, the subbagging estimator can achieve $\sqrt{N}$-consistency and asymptotic normality under the condition $(k_Nm_N)/N\to α\in (0,\infty]$. Compared to the full sample estimator, we theoretically show that the $\sqrt{N}$-consistent subbagging estimator has an inflation rate of $1/α$ in its asymptotic variance. Simulation experiments are presented to demonstrate the finite sample performances. An American airline dataset is analyzed to illustrate that the subbagging estimate is numerically close to the full sample estimate, and can be computationally fast under the memory constraint.

preprint2014arXiv

Metric Entropy and the Optimal Prediction of Chaotic Signals

Suppose we are given a time series or a signal $x(t)$ for $0\leq t\leq T$. We consider the problem of predicting the signal in the interval $T<t\leq T+t_{f}$ from a knowledge of its history and nothing more. We ask the following question: what is the largest value of $t_{f}$ for which a prediction can be made? We show that the answer to this question is contained in a fundamental result of information theory due to Wyner, Ziv, Ornstein, and Weiss. In particular, for the class of chaotic signals, the upper bound is $t_{f}\leq\log_{2}T/H$ in the limit $T\rightarrow\infty$, with $H$ being entropy in a sense that is explained in the text. If $\bigl|x(T-s)-x(t^{\ast}-s)\bigr|$ is small for $0\leq s\leqτ$, where $τ$ is of the order of a characteristic time scale, the pattern of events leading up to $t=T$ is similar to the pattern of events leading up to $t=t^{\ast}$. It is reasonable to expect $x(t^{\ast}+t_{f})$ to be a good predictor of $x(T+t_{f}).$ All existing methods for prediction use this idea in some way or the other. Unfortunately, this intuitively reasonable idea is fundamentally deficient and all existing methods fall well short of the Wyner-Ziv entropy bound on $t_{f}$. An optimal predictor should decompose the distance between the pattern of events leading up to $t=T$ and the pattern leading up to $t=t^{\ast}$ into stable and unstable components. A good match should have suitably small unstable components but will in general allow stable components which are as large as the tolerance for correct prediction. For the special case of hyperbolic toral automorphisms, we derive an optimal predictor using Pade approximation.