Source author record

Guangming Pan

Guangming Pan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

24works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

24 published item(s)

preprint2026arXiv

High-Dimensional Precision Matrix Quadratic Forms: Estimation Framework for $p > n$

We propose a novel estimation framework for quadratic functionals of precision matrices in high-dimensional settings, particularly in regimes where the feature dimension $p$ exceeds the sample size $n$. Traditional moment-based estimators with bias correction remain consistent when $p<n$ (i.e., $p/n \to c <1$). However, they break down entirely once $p>n$, highlighting a fundamental distinction between the two regimes due to rank deficiency and high-dimensional complexity. Our approach resolves these issues by combining a spectral-moment representation with constrained optimization, resulting in consistent estimation under mild moment conditions. The proposed framework provides a unified approach for inference on a broad class of high-dimensional statistical measures. We illustrate its utility through two representative examples: the optimal Sharpe ratio in portfolio optimization and the multiple correlation coefficient in regression analysis. Simulation studies demonstrate that the proposed estimator effectively overcomes the fundamental $p>n$ barrier where conventional methods fail.

preprint2022arXiv

A new model for preferential attachment scheme with time-varying parameters

We propose an extension of the preferential attachment scheme by allowing the connecting probability to depend on time t. We estimate the parameters involved in the model by minimizing the expected squared difference between the number of vertices of degree one and its conditional expectation. The asymptotic properties of the estimators are also investigated when the parameters are time-varying by establishing the central limit theorem (CLT) of the number of vertices of degree one. We propose a new statistic to test whether the parameters have change points. We also offer some methods to estimate the number of change points and detect the locations of change points. Simulations are conducted to illustrate the performances of the above results.

preprint2022arXiv

Factor Modelling for Clustering High-dimensional Time Series

We propose a new unsupervised learning method for clustering a large number of time series based on a latent factor structure. Each cluster is characterized by its own cluster-specific factors in addition to some common factors which impact on all the time series concerned. Our setting also offers the flexibility that some time series may not belong to any clusters. The consistency with explicit convergence rates is established for the estimation of the common factors, the cluster-specific factors, the latent clusters. Numerical illustration with both simulated data as well as a real data example is also reported. As a spin-off, the proposed new approach also advances significantly the statistical inference for the factor model of Lam and Yao (2012).

preprint2016arXiv

A unified matrix model including both CCA and F matrices in multivariate analysis: the largest eigenvalue and its applications

Let $\bbZ_{M_1\times N}=\bbT^{\frac{1}{2}}\bbX$ where $(\bbT^{\frac{1}{2}})^2=\bbT$ is a positive definite matrix and $\bbX$ consists of independent random variables with mean zero and variance one. This paper proposes a unified matrix model $$\bold{\bbom}=(\bbZ\bbU_2\bbU_2^T\bbZ^T)^{-1}\bbZ\bbU_1\bbU_1^T\bbZ^T,$$ where $\bbU_1$ and $\bbU_2$ are isometric with dimensions $N\times N_1$ and $N\times (N-N_2)$ respectively such that $\bbU_1^T\bbU_1=\bbI_{N_1}$, $\bbU_2^T\bbU_2=\bbI_{N-N_2}$ and $\bbU_1^T\bbU_2=0$. Moreover, $\bbU_1$ and $\bbU_2$ (random or non-random) are independent of $\bbZ_{M_1\times N}$ and with probability tending to one, $rank(\bbU_1)=N_1$ and $rank(\bbU_2)=N-N_2$. We establish the asymptotic Tracy-Widom distribution for its largest eigenvalue under moment assumptions on $\bbX$ when $N_1,N_2$ and $M_1$ are comparable. By selecting appropriate matrices $\bbU_1$ and $\bbU_2$, the asymptotic distributions of the maximum eigenvalues of the matrices used in Canonical Correlation Analysis (CCA) and of F matrices (including centered and non-centered versions) can be both obtained from that of $\bold{\bbom}$. %In particular, $\bbom$ can also cover nonzero mean by appropriate matrices $\bbU_1$ and $\bbU_2$. %relax the zero mean value restriction for F matrix in \cite{WY} to allow for any nonzero mean vetors. %thus a direct application of our proposed Tracy-Widom distribution is the independence testing via CCA. Moreover, via appropriate matrices $\bbU_1$ and $\bbU_2$, this matrix $\bold{\bbom}$ can be applied to some multivariate testing problems that cannot be done by the traditional CCA matrix.

preprint2015arXiv

CLT for linear spectral statistics of normalized sample covariance matrices with the dimension much larger than the sample size

Let $\mathbf{A}=\frac{1}{\sqrt{np}}(\mathbf{X}^T\mathbf{X}-p\mathbf {I}_n)$ where $\mathbf{X}$ is a $p\times n$ matrix, consisting of independent and identically distributed (i.i.d.) real random variables $X_{ij}$ with mean zero and variance one. When $p/n\to\infty$, under fourth moment conditions a central limit theorem (CLT) for linear spectral statistics (LSS) of $\mathbf{A}$ defined by the eigenvalues is established. We also explore its applications in testing whether a population covariance matrix is an identity matrix.

preprint2015arXiv

Convergence of the empirical spectral distribution function of Beta matrices

Let $\mathbf{B}_n=\mathbf {S}_n(\mathbf {S}_n+α_n\mathbf {T}_N)^{-1}$, where $\mathbf {S}_n$ and $\mathbf {T}_N$ are two independent sample covariance matrices with dimension $p$ and sample sizes $n$ and $N$, respectively. This is the so-called Beta matrix. In this paper, we focus on the limiting spectral distribution function and the central limit theorem of linear spectral statistics of $\mathbf {B}_n$. Especially, we do not require $\mathbf {S}_n$ or $\mathbf {T}_N$ to be invertible. Namely, we can deal with the case where $p>\max\{n,N\}$ and $p<n+N$. Therefore, our results cover many important applications which cannot be simply deduced from the corresponding results for multivariate $F$ matrices.

preprint2015arXiv

Independence test for high dimensional data based on regularized canonical correlation coefficients

This paper proposes a new statistic to test independence between two high dimensional random vectors ${\mathbf{X}}:p_1\times1$ and ${\mathbf{Y}}:p_2\times1$. The proposed statistic is based on the sum of regularized sample canonical correlation coefficients of ${\mathbf{X}}$ and ${\mathbf{Y}}$. The asymptotic distribution of the statistic under the null hypothesis is established as a corollary of general central limit theorems (CLT) for the linear statistics of classical and regularized sample canonical correlation coefficients when $p_1$ and $p_2$ are both comparable to the sample size $n$. As applications of the developed independence test, various types of dependent structures, such as factor models, ARCH models and a general uncorrelated but dependent case, etc., are investigated by simulations. As an empirical application, cross-sectional dependence of daily stock returns of companies between different sections in the New York Stock Exchange (NYSE) is detected by the proposed test.

preprint2015arXiv

Spectral statistics of large dimensional Spearman's rank correlation matrix and its application

Let $\mathbf{Q}=(Q_1,\ldots,Q_n)$ be a random vector drawn from the uniform distribution on the set of all $n!$ permutations of $\{1,2,\ldots,n\}$. Let $\mathbf{Z}=(Z_1,\ldots,Z_n)$, where $Z_j$ is the mean zero variance one random variable obtained by centralizing and normalizing $Q_j$, $j=1,\ldots,n$. Assume that $\mathbf {X}_i,i=1,\ldots ,p$ are i.i.d. copies of $\frac{1}{\sqrt{p}}\mathbf{Z}$ and $X=X_{p,n}$ is the $p\times n$ random matrix with $\mathbf{X}_i$ as its $i$th row. Then $S_n=XX^*$ is called the $p\times n$ Spearman's rank correlation matrix which can be regarded as a high dimensional extension of the classical nonparametric statistic Spearman's rank correlation coefficient between two independent random variables. In this paper, we establish a CLT for the linear spectral statistics of this nonparametric random matrix model in the scenario of high dimension, namely, $p=p(n)$ and $p/n\to c\in(0,\infty)$ as $n\to\infty$. We propose a novel evaluation scheme to estimate the core quantity in Anderson and Zeitouni's cumulant method in [Ann. Statist. 36 (2008) 2553-2576] to bypass the so-called joint cumulant summability. In addition, we raise a two-step comparison approach to obtain the explicit formulae for the mean and covariance functions in the CLT. Relying on this CLT, we then construct a distribution-free statistic to test complete independence for components of random vectors. Owing to the nonparametric property, we can use this test on generally distributed random variables including the heavy-tailed ones.

preprint2015arXiv

The logarithmic law of random determinant

Consider the square random matrix $A_n=(a_{ij})_{n,n}$, where $\{a_{ij}:=a_{ij}^{(n)},i,j=1,\ldots,n\}$ is a collection of independent real random variables with means zero and variances one. Under the additional moment condition \[\sup_n\max_{1\leq i,j\leq n}\mathbb{E}a_{ij}^4<\infty,\] we prove Girko's logarithmic law of $\det A_n$ in the sense that as $n\rightarrow\infty$ \begin{eqnarray*}\frac{\log|\det A_n|-(1/2)\log(n-1)!}{\sqrt{(1/2)\log n}}\stackrel{d}{ \longrightarrow}N(0,1).\end{eqnarray*}

preprint2015arXiv

Universality for the largest eigenvalue of sample covariance matrices with general population

This paper is aimed at deriving the universality of the largest eigenvalue of a class of high-dimensional real or complex sample covariance matrices of the form $\mathcal{W}_N=Σ^{1/2}XX^*Σ^{1/2}$. Here, $X=(x_{ij})_{M,N}$ is an $M\times N$ random matrix with independent entries $x_{ij},1\leq i\leq M,1\leq j\leq N$ such that $\mathbb{E}x_{ij}=0$, $\mathbb{E}|x_{ij}|^2=1/N$. On dimensionality, we assume that $M=M(N)$ and $N/M\rightarrow d\in(0,\infty)$ as $N\rightarrow\infty$. For a class of general deterministic positive-definite $M\times M$ matrices $Σ$, under some additional assumptions on the distribution of $x_{ij}$'s, we show that the limiting behavior of the largest eigenvalue of $\mathcal{W}_N$ is universal, via pursuing a Green function comparison strategy raised in [Probab. Theory Related Fields 154 (2012) 341-407, Adv. Math. 229 (2012) 1435-1515] by Erdős, Yau and Yin for Wigner matrices and extended by Pillai and Yin [Ann. Appl. Probab. 24 (2014) 935-1001] to sample covariance matrices in the null case ($Σ=I$). Consequently, in the standard complex case ($\mathbb{E}x_{ij}^2=0$), combing this universality property and the results known for Gaussian matrices obtained by El Karoui in [Ann. Probab. 35 (2007) 663-714] (nonsingular case) and Onatski in [Ann. Appl. Probab. 18 (2008) 470-490] (singular case), we show that after an appropriate normalization the largest eigenvalue of $\mathcal{W}_N$ converges weakly to the type 2 Tracy-Widom distribution $\mathrm{TW}_2$. Moreover, in the real case, we show that when $Σ$ is spiked with a fixed number of subcritical spikes, the type 1 Tracy-Widom limit $\mathrm{TW}_1$ holds for the normalized largest eigenvalue of $\mathcal {W}_N$, which extends a result of Féral and Péché in [J. Math. Phys. 50 (2009) 073302] to the scenario of nondiagonal $Σ$ and more generally distributed $X$.

preprint2014arXiv

Canonical correlation coefficients of high-dimensional normal vectors: finite rank case

Consider a normal vector $\mathbf{z}=(\mathbf{x}',\mathbf{y}')'$, consisting of two sub-vectors $\mathbf{x}$ and $\mathbf{y}$ with dimensions $p$ and $q$ respectively. With $n$ independent observations of $\mathbf{z}$ at hand, we study the correlation between $\mathbf{x}$ and $\mathbf{y}$, from the perspective of the Canonical Correlation Analysis, under the high-dimensional setting: both $p$ and $q$ are proportional to the sample size $n$. In this paper, we focus on the case that $Σ_{\mathbf{x}\mathbf{y}}$ is of finite rank $k$, i.e. there are $k$ nonzero canonical correlation coefficients, whose squares are denoted by $r_1\geq\cdots\geq r_k>0$. Under the additional assumptions $(p+q)/n\to y\in (0,1)$ and $p/q\not\to 1$, we study the sample counterparts of $r_i,i=1,\ldots,k$, i.e. the largest k eigenvalues of the sample canonical correlation matrix $S_{\mathbf{x}\mathbf{x}}^{-1}S_{\mathbf{x}\mathbf{y}}S_{\mathbf{y}\mathbf{y}}^{-1}S_{\mathbf{y}\mathbf{x}}$, namely $λ_1\geq\cdots\geq λ_k$. We show that there exists a threshold $r_c\in(0,1)$, such that for each $i\in\{1,\ldots,k\}$, when $r_i\leq r_c$, $λ_i$ converges almost surely to the right edge of the limiting spectral distribution of the sample canonical correlation matrix, denoted by $d_r$. When $r_i>r_c$, $λ_i$ possesses an almost sure limit in $(d_r,1]$, from which we can recover $r_i$ in turn, thus provide an estimate of the latter in the high-dimensional scenario.

preprint2014arXiv

High Dimensional Correlation Matrices: CLT and Its Applications

Statistical inferences for sample correlation matrices are important in high dimensional data analysis. Motivated by this, this paper establishes a new central limit theorem (CLT) for a linear spectral statistic (LSS) of high dimensional sample correlation matrices for the case where the dimension p and the sample size $n$ are comparable. This result is of independent interest in large dimensional random matrix theory. Meanwhile, we apply the linear spectral statistic to an independence test for $p$ random variables, and then an equivalence test for p factor loadings and $n$ factors in a factor model. The finite sample performance of the proposed test shows its applicability and effectiveness in practice. An empirical application to test the independence of household incomes from different cities in China is also conducted.

preprint2014arXiv

On singular value distribution of large dimensional auto-covariance matrices

Let $(\varepsilon_j)_{j\geq 0}$ be a sequence of independent $p-$dimensional random vectors and $τ\geq1$ a given integer. From a sample $\varepsilon_1,\cdots,\varepsilon_{T+τ-1},\varepsilon_{T+τ}$ of the sequence, the so-called lag $-τ$ auto-covariance matrix is $C_τ=T^{-1}\sum_{j=1}^T\varepsilon_{τ+j}\varepsilon_{j}^t$. When the dimension $p$ is large compared to the sample size $T$, this paper establishes the limit of the singular value distribution of $C_τ$ assuming that $p$ and $T$ grow to infinity proportionally and the sequence satisfies a Lindeberg condition on fourth order moments. Compared to existing asymptotic results on sample covariance matrices developed in random matrix theory, the case of an auto-covariance matrix is much more involved due to the fact that the summands are dependent and the matrix $C_τ$ is not symmetric. Several new techniques are introduced for the derivation of the main theorem.

preprint2014arXiv

Test of Independence for High-dimensional Random Vectors Based on Block Correlation Matrices

In this paper, we are concerned with the independence test for $k$ high-dimensional sub-vectors of a normal vector, with fixed positive integer $k$. A natural high-dimensional extension of the classical sample correlation matrix, namely block correlation matrix, is raised for this purpose. We then construct the so-called Schott type statistic as our test statistic, which turns out to be a particular linear spectral statistic of the block correlation matrix. Interestingly, the limiting behavior of the Schott type statistic can be figured out with the aid of the Free Probability Theory and the Random Matrix Theory. Specifically, we will bring the so-called real second order freeness for Haar distributed orthogonal matrices, derived in \cite{MP2013}, into the framework of this high-dimensional testing problem. Our test does not require the sample size to be larger than the total or any partial sum of the dimensions of the $k$ sub-vectors. Simulated results show the effect of the Schott type statistic, in contrast to those of the statistics proposed in \cite{JY2013} and \cite{JBZ2013}, is satisfactory. Real data analysis is also used to illustrate our method.

preprint2013arXiv

Universality for a global property of the eigenvectors of Wigner matrices

Let $M_n$ be an $n\times n$ real (resp. complex) Wigner matrix and $U_nΛ_n U_n^*$ be its spectral decomposition. Set $(y_1,y_2...,y_n)^T=U_n^*x$, where $x=(x_1,x_2,...,$ $x_n)^T$ is a real (resp. complex) unit vector. Under the assumption that the elements of $M_n$ have 4 matching moments with those of GOE (resp. GUE), we show that the process $X_n(t)=\sqrt{\frac{βn}{2}}\sum_{i=1}^{\lfloor nt\rfloor}(|y_i|^2-\frac1n)$ converges weakly to the Brownian bridge for any $\mathbf{x}$ such that $||x||_\infty\rightarrow 0$ as $n\rightarrow \infty$, where $β=1$ for the real case and $β=2$ for the complex case. Such a result indicates that the othorgonal (resp. unitary) matrices with columns being the eigenvectors of Wigner matrices are asymptotically Haar distributed on the orthorgonal (resp. unitary) group from a certain perspective.

preprint2012arXiv

Central limit theorem for partial linear eigenvalue statistics of Wigner matrices

In this paper, we study the complex Wigner matrices $M_n=\frac{1}{\sqrt{n}}W_n$ whose eigenvalues are typically in the interval $[-2,2]$. Let $λ_1\leq λ_2...\leqλ_n$ be the ordered eigenvalues of $M_n$. Under the assumption of four matching moments with the Gaussian Unitary Ensemble(GUE), for test function $f$ 4-times continuously differentiable on an open interval including $[-2,2]$, we establish central limit theorems for two types of partial linear statistics of the eigenvalues. The first type is defined with a threshold $u$ in the bulk of the Wigner semicircle law as $\mathcal{A}_n[f; u]=\sum_{l=1}^nf(λ_l)\mathbf{1}_{\{λ_l\leq u\}}$. And the second one is $\mathcal{B}_n[f; k]=\sum_{l=1}^{k}f(λ_l)$ with positive integer $k=k_n$ such that $k/n\rightarrow y\in (0,1)$ as $n$ tends to infinity. Moreover, we derive a weak convergence result for a partial sum process constructed from $\mathcal{B}_n[f; \lfloor nt\rfloor]$.

preprint2011arXiv

A Deterministic Equivalent for the Analysis of Non-Gaussian Correlated MIMO Multiple Access Channels

Large dimensional random matrix theory (RMT) has provided an efficient analytical tool to understand multiple-input multiple-output (MIMO) channels and to aid the design of MIMO wireless communication systems. However, previous studies based on large dimensional RMT rely on the assumption that the transmit correlation matrix is diagonal or the propagation channel matrix is Gaussian. There is an increasing interest in the channels where the transmit correlation matrices are generally nonnegative definite and the channel entries are non-Gaussian. This class of channel models appears in several applications in MIMO multiple access systems, such as small cell networks (SCNs). To address these problems, we use the generalized Lindeberg principle to show that the Stieltjes transforms of this class of random matrices with Gaussian or non-Gaussian independent entries coincide in the large dimensional regime. This result permits to derive the deterministic equivalents (e.g., the Stieltjes transform and the ergodic mutual information) for non-Gaussian MIMO channels from the known results developed for Gaussian MIMO channels, and is of great importance in characterizing the spectral efficiency of SCNs.

preprint2011arXiv

Tracy-Widom law for the extreme eigenvalues of sample correlation matrices

Let the sample correlation matrix be $W=YY^T$, where $Y=(y_{ij})_{p,n}$ with $y_{ij}=x_{ij}/\sqrt{\sum_{j=1}^nx_{ij}^2}$. We assume $\{x_{ij}: 1\leq i\leq p, 1\leq j\leq n\}$ to be a collection of independent symmetric distributed random variables with sub-exponential tails. Moreover, for any $i$, we assume $x_{ij}, 1\leq j\leq n$ to be identically distributed. We assume $0<p<n$ and $p/n\rightarrow y$ with some $y\in(0,1)$ as $p,n\rightarrow\infty$. In this paper, we provide the Tracy-Widom law ($TW_1$) for both the largest and smallest eigenvalues of $W$. If $x_{ij}$ are i.i.d. standard normal, we can derive the $TW_1$ for both the largest and smallest eigenvalues of the matrix $\mathcal{R}=RR^T$, where $R=(r_{ij})_{p,n}$ with $r_{ij}=(x_{ij}-\bar x_i)/\sqrt{\sum_{j=1}^n(x_{ij}-\bar x_i)^2}$, $\bar x_i=n^{-1}\sum_{j=1}^nx_{ij}$.

preprint2010arXiv

On the Performance of Spectrum Sensing Algorithms using Multiple Antennas

In recent years, some spectrum sensing algorithms using multiple antennas, such as the eigenvalue based detection (EBD), have attracted a lot of attention. In this paper, we are interested in deriving the asymptotic distributions of the test statistics of the EBD algorithms. Two EBD algorithms using sample covariance matrices are considered: maximum eigenvalue detection (MED) and condition number detection (CND). The earlier studies usually assume that the number of antennas (K) and the number of samples (N) are both large, thus random matrix theory (RMT) can be used to derive the asymptotic distributions of the maximum and minimum eigenvalues of the sample covariance matrices. While assuming the number of antennas being large simplifies the derivations, in practice, the number of antennas equipped at a single secondary user is usually small, say 2 or 3, and once designed, this antenna number is fixed. Thus in this paper, our objective is to derive the asymptotic distributions of the eigenvalues and condition numbers of the sample covariance matrices for any fixed K but large N, from which the probability of detection and probability of false alarm can be obtained. The proposed methodology can also be used to analyze the performance of other EBD algorithms. Finally, computer simulations are presented to validate the accuracy of the derived results.