Source author record

Debashis Paul

Debashis Paul appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

17works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

17 published item(s)

preprint2021arXiv

Variance Estimation and Confidence Intervals from High-dimensional Genome-wide Association Studies Through Misspecified Mixed Model Analysis

We study variance estimation and associated confidence intervals for parameters characterizing genetic effects from genome-wide association studies (GWAS) misspecified mixed model analysis. Previous studies have shown that, in spite of the model misspecification, certain quantities of genetic interests are estimable, and consistent estimators of these quantities can be obtained using the restricted maximum likelihood (REML) method under a misspecified linear mixed model. However, the asymptotic variance of such a REML estimator is complicated and not ready to be implemented for practical use. In this paper, we develop practical and computationally convenient methods for estimating such asymptotic variances and constructing the associated confidence intervals. Performance of the proposed methods is evaluated empirically based on Monte-Carlo simulations and real-data application.

preprint2016arXiv

Estimating fiber orientation distribution from diffusion MRI with spherical needlets

We present a novel method for estimation of the fiber orientation distribution (FOD) function based on diffusion-weighted Magnetic Resonance Imaging (D-MRI) data. We formulate the problem of FOD estimation as a regression problem through spherical deconvolution and a sparse representation of the FOD by a spherical needlets basis that form a multi-resolution tight frame for spherical functions. This sparse representation allows us to estimate FOD by an $l_1$-penalized regression under a non-negativity constraint. The resulting convex optimization problem is solved by an alternating direction method of multipliers (ADMM) algorithm. The proposed method leads to a reconstruction of the FODs that is accurate, has low variability and preserves sharp features. Through extensive experiments, we demonstrate the effectiveness and favorable performance of the proposed method compared with two existing methods. Particularly, we show the ability of the proposed method in successfully resolving fiber crossing at small angles and in automatically identifying isotropic diffusion. We also apply the proposed method to real 3T D-MRI data sets of healthy elderly individuals. The results show realistic descriptions of crossing fibers that are more accurate and less noisy than competing methods even with a relatively small number of gradient directions.

preprint2016arXiv

Modeling Tangential Vector Fields on a Sphere

Physical processes that manifest as tangential vector fields on a sphere are common in geophysical and environmental sciences. These naturally occurring vector fields are often subject to physical constraints, such as being curl-free or divergence-free. We construct a new class of parametric models for cross-covariance functions of curl-free and divergence-free vector fields that are tangential to the unit sphere. These models are constructed by applying the surface gradient or the surface curl operator to scalar random potential fields defined on the unit sphere. We propose a likelihood-based estimation procedure for the model parameters and show that fast computation is possible even for large data sets when the observations are on a regular latitude-longitude grid. Characteristics and utility of the proposed methodology are illustrated through simulation studies and by applying it to an ocean surface wind velocity data set collected through satellite-based scatterometry remote sensing. We also compare the performance of the proposed model with a class of bivariate Matérn models in terms of estimation and prediction, and demonstrate that the proposed model is superior in capturing certain physical characteristics of the wind fields.

preprint2015arXiv

Fiber Direction Estimation, Smoothing and Tracking in Diffusion MRI

Diffusion magnetic resonance imaging is an imaging technology designed to probe anatomical architectures of biological samples in an in vivo and non-invasive manner through measuring water diffusion. The contribution of this paper is threefold. First it proposes a new method to identify and estimate multiple diffusion directions within a voxel through a new and identifiable parametrization of the widely used multi-tensor model. Unlike many existing methods, this method focuses on the estimation of diffusion directions rather than the diffusion tensors. Second, this paper proposes a novel direction smoothing method which greatly improves direction estimation in regions with crossing fibers. This smoothing method is shown to have excellent theoretical and empirical properties. Lastly, this paper develops a fiber tracking algorithm that can handle multiple directions within a voxel. The overall methodology is illustrated with simulated data and a data set collected for the study of Alzheimer's disease by the Alzheimer's Disease Neuroimaging Initiative (ADNI).

preprint2015arXiv

On the Marčenko-Pastur law for linear time series

This paper is concerned with extensions of the classical Marčenko-Pastur law to time series. Specifically, $p$-dimensional linear processes are considered which are built from innovation vectors with independent, identically distributed (real- or complex-valued) entries possessing zero mean, unit variance and finite fourth moments. The coefficient matrices of the linear process are assumed to be simultaneously diagonalizable. In this setting, the limiting behavior of the empirical spectral distribution of both sample covariance and symmetrized sample autocovariance matrices is determined in the high-dimensional setting $p/n\to c\in (0,\infty)$ for which dimension $p$ and sample size $n$ diverge to infinity at the same rate. The results extend existing contributions available in the literature for the covariance case and are one of the first of their kind for the autocovariance case.

preprint2015arXiv

Spectral analysis of linear time series in moderately high dimensions

This article is concerned with the spectral behavior of $p$-dimensional linear processes in the moderately high-dimensional case when both dimensionality $p$ and sample size $n$ tend to infinity so that $p/n\to0$. It is shown that, under an appropriate set of assumptions, the empirical spectral distributions of the renormalized and symmetrized sample autocovariance matrices converge almost surely to a nonrandom limit distribution supported on the real line. The key assumption is that the linear process is driven by a sequence of $p$-dimensional real or complex random vectors with i.i.d. entries possessing zero mean, unit variance and finite fourth moments, and that the $p\times p$ linear process coefficient matrices are Hermitian and simultaneously diagonalizable. Several relaxations of these assumptions are discussed. The results put forth in this paper can help facilitate inference on model parameters, model diagnostics and prediction of future values of the linear process.

preprint2014arXiv

Adaptation in a class of linear inverse problems

We consider the linear inverse problem of estimating an unknown signal $f$ from noisy measurements on $Kf$ where the linear operator $K$ admits a wavelet-vaguelette decomposition (WVD). We formulate the problem in the Gaussian sequence model and propose estimation based on complexity penalized regression on a level-by-level basis. We adopt squared error loss and show that the estimator achieves exact rate-adaptive optimality as $f$ varies over a wide range of Besov function classes.

preprint2014arXiv

High-dimensional genome-wide association study and misspecified mixed model analysis

We study behavior of the restricted maximum likelihood (REML) estimator under a misspecified linear mixed model (LMM) that has received much attention in recent gnome-wide association studies. The asymptotic analysis establishes consistency of the REML estimator of the variance of the errors in the LMM, and convergence in probability of the REML estimator of the variance of the random effects in the LMM to a certain limit, which is equal to the true variance of the random effects multiplied by the limiting proportion of the nonzero random effects present in the LMM. The aymptotic results also establish convergence rate (in probability) of the REML estimators as well as a result regarding convergence of the asymptotic conditional variance of the REML estimator. The asymptotic results are fully supported by the results of empirical studies, which include extensive simulation studies that compare the performance of the REML estimator (under the misspecified LMM) with other existing methods.

preprint2014arXiv

Nonparametric estimation of dynamics of monotone trajectories

We study a class of nonlinear nonparametric inverse problems. Specifically, we propose a nonparametric estimator of the dynamics of a monotonically increasing trajectory defined on a finite time interval. Under suitable regularity conditions, we prove consistency of the proposed estimator and show that in terms of $L^2$-loss, the optimal rate of convergence for the proposed estimator is the same as that for the estimation of the derivative of a trajectory. This is a new contribution to the area of nonlinear nonparametric inverse problems. We conduct a simulation study to examine the finite sample behavior of the proposed estimator and apply it to the Berkeley growth data.

preprint2014arXiv

Time-warped growth processes, with applications to the modeling of boom-bust cycles in house prices

House price increases have been steady over much of the last 40 years, but there have been occasional declines, most notably in the recent housing bust that started around 2007, on the heels of the preceding housing bubble. We introduce a novel growth model that is motivated by time-warping models in functional data analysis and includes a nonmonotone time-warping component that allows the inclusion and description of boom-bust cycles and facilitates insights into the dynamics of asset bubbles. The underlying idea is to model longitudinal growth trajectories for house prices and other phenomena, where temporal setbacks and deflation may be encountered, by decomposing such trajectories into two components. A first component corresponds to underlying steady growth driven by inflation that anchors the observed trajectories on a simple first order linear differential equation, while a second boom-bust component is implemented as time warping. Time warping is a commonly encountered phenomenon and reflects random variation along the time axis. Our approach to time warping is more general than previous approaches by admitting the inclusion of nonmonotone warping functions. The anchoring of the trajectories on an underlying linear dynamic system also makes the time-warping component identifiable and enables straightforward estimation procedures for all model components. The application to the dynamics of housing prices as observed for 19 metropolitan areas in the U.S. from December 1998 to July 2013 reveals that the time setbacks corresponding to nonmonotone time warping vary substantially across markets and we find indications that they are related to market-specific growth rates.

preprint2013arXiv

Limiting spectral distribution of renormalized separable sample covariance matrices when $p/n\to 0$

We are concerned with the behavior of the eigenvalues of renormalized sample covariance matrices of the form C_n=\sqrt{\frac{n}{p}}\left(\frac{1}{n}A_{p}^{1/2}X_{n}B_{n}X_{n}^{*}A_{p}^{1/2}-\frac{1}{n}\tr(B_{n})A_{p}\right) as $p,n\to \infty$ and $p/n\to 0$, where $X_{n}$ is a $p\times n$ matrix with i.i.d. real or complex valued entries $X_{ij}$ satisfying $E(X_{ij})=0$, $E|X_{ij}|^2=1$ and having finite fourth moment. $A_{p}^{1/2}$ is a square-root of the nonnegative definite Hermitian matrix $A_{p}$, and $B_{n}$ is an $n\times n$ nonnegative definite Hermitian matrix. We show that the empirical spectral distribution (ESD) of $C_n$ converges a.s. to a nonrandom limiting distribution under some assumptions. The probability density function of the LSD of $C_{n}$ is derived and it is shown that it depends on the LSD of $A_{p}$ and the limiting value of $n^{-1}\tr(B_{n}^2)$. We propose a computational algorithm for evaluating this limiting density when the LSD of $A_{p}$ is a mixture of point masses. In addition, when the entries of $X_{n}$ are sub-Gaussian, we derive the limiting empirical distribution of $\{\sqrt{n/p}(λ_j(S_n) - n^{-1}\tr(B_n) λ_j(A_{p}))\}_{j=1}^p$ where $S_n := n^{-1} A_{p}^{1/2}X_{n}B_{n}X_{n}^{*}A_{p}^{1/2}$ is the sample covariance matrix and $λ_j$ denotes the $j$-th largest eigenvalue, when $F^A$ is a finite mixture of point masses. These results are utilized to propose a test for the covariance structure of the data where the null hypothesis is that the joint covariance matrix is of the form $A_{p} \otimes B_n$ for $\otimes$ denoting the Kronecker product, as well as $A_{p}$ and the first two spectral moments of $B_n$ are specified. The performance of this test is illustrated through a simulation study.

preprint2012arXiv

Augmented sparse principal component analysis for high dimensional data

We study the problem of estimating the leading eigenvectors of a high-dimensional population covariance matrix based on independent Gaussian observations. We establish lower bounds on the rates of convergence of the estimators of the leading eigenvectors under $l^q$-sparsity constraints when an $l^2$ loss function is used. We also propose an estimator of the leading eigenvectors based on a coordinate selection scheme combined with PCA and show that the proposed estimator achieves the optimal rate of convergence under a sparsity regime. Moreover, we establish that under certain scenarios, the usual PCA achieves the minimax convergence rate.

preprint2012arXiv

Minimax bounds for sparse PCA with noisy high-dimensional data

We study the problem of estimating the leading eigenvectors of a high-dimensional population covariance matrix based on independent Gaussian observations. We establish a lower bound on the minimax risk of estimators under the $l_2$ loss, in the joint limit as dimension and sample size increase to infinity, under various models of sparsity for the population eigenvectors. The lower bound on the risk points to the existence of different regimes of sparsity of the eigenvectors. We also propose a new method for estimating the eigenvectors by a two-stage coordinate selection scheme.

preprint2011arXiv

Semiparametric modeling of autonomous nonlinear dynamical systems with application to plant growth

We propose a semiparametric model for autonomous nonlinear dynamical systems and devise an estimation procedure for model fitting. This model incorporates subject-specific effects and can be viewed as a nonlinear semiparametric mixed effects model. We also propose a computationally efficient model selection procedure. We show by simulation studies that the proposed estimation as well as model selection procedures can efficiently handle sparse and noisy measurements. Finally, we apply the proposed method to a plant growth data used to study growth displacement rates within meristems of maize roots under two different experimental conditions.

preprint2011arXiv

Shrinking the Quadratic Estimator

We study a regression characterization for the quadratic estimator of weak lensing, developed by Hu and Okamoto (2001,2002), for cosmic microwave background observations. This characterization motivates a modification of the quadratic estimator by an adaptive Wiener filter which uses the robust Bayesian techniques described in Strawderman (1971) and Berger (1980). This technique requires the user to propose a fiducial model for the spectral density of the unknown lensing potential but the resulting estimator is developed to be robust to misspecification of this model. The role of the fiducial spectral density is to give the estimator superior statistical performance in a "neighborhood of the fiducial model" while controlling the statistical errors when the fiducial spectral density is drastically wrong. Our estimate also highlights some advantages provided by a Bayesian analysis of the quadratic estimator.

preprint2010arXiv

Geometric kernel smoothing of tensor fields

In this paper, we study a kernel smoothing approach for denoising a tensor field. Particularly, both simulation studies and theoretical analysis are conducted to understand the effects of the noise structure and the structure of the tensor field on the performance of different smoothers arising from using different metrics, viz., Euclidean, log-Euclidean and affine invariant metrics. We also study the Rician noise model and compare two regression estimators of diffusion tensors based on raw diffusion weighted imaging data at each voxel.

preprint2007arXiv

"Pre-conditioning" for feature selection and regression in high-dimensional problems

We consider regression problems where the number of predictors greatly exceeds the number of observations. We propose a method for variable selection that first estimates the regression function, yielding a "pre-conditioned" response variable. The primary method used for this initial regression is supervised principal components. Then we apply a standard procedure such as forward stepwise selection or the LASSO to the pre-conditioned response variable. In a number of simulated and real data examples, this two-step procedure outperforms forward stepwise selection or the usual LASSO (applied directly to the raw outcome). We also show that under a certain Gaussian latent variable model, application of the LASSO to the pre-conditioned response variable is consistent as the number of predictors and observations increases. Moreover, when the observational noise is rather large, the suggested procedure can give a more accurate estimate than LASSO. We illustrate our method on some real problems, including survival analysis with microarray data.