Source author record

Hua Zhou

Hua Zhou appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

38works
16topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

38 published item(s)

preprint2023arXiv

High-throughput combinatorial approach expedites the synthesis of a lead-free relaxor ferroelectric system

Developing novel lead-free ferroelectric materials is crucial for next-generation microelectronic technologies that are energy efficient and environment friendly. However, materials discovery and property optimization are typically time-consuming due to the limited throughput of traditional synthesis methods. In this work, we use a high-throughput combinatorial synthesis approach to fabricate lead-free ferroelectric superlattices and solid solutions of (Ba0.7Ca0.3)TiO3 (BCT) and Ba(Zr0.2Ti0.8)O3 (BZT) phases with continuous variation of composition and layer thickness. High-resolution X-ray diffraction (XRD) and analytical scanning transmission electron microscopy (STEM) demonstrate high film quality and well-controlled compositional gradients. Ferroelectric and dielectric property measurements identify the optimal property point achieved at the morphotropic phase boundary (MPB) with a composition of 48BZT-52BCT. Displacement vector maps reveal that ferroelectric domain sizes are tunable by varying {BCT-BZT}N superlattice geometry. This high-throughput synthesis approach can be applied to many other material systems to expedite new materials discovery and properties optimization, allowing for the exploration of a large area of phase space within a single growth.

preprint2022arXiv

Efficient Algorithms and Implementation of a Semiparametric Joint Model for Longitudinal and Competing Risks Data: With Applications to Massive Biobank Data

Semiparametric joint models of longitudinal and competing risks data are computationally costly and their current implementations do not scale well to massive biobank data. This paper identifies and addresses some key computational barriers in a semiparametric joint model for longitudinal and competing risks survival data. By developing and implementing customized linear scan algorithms, we reduce the computational complexities from $O(n^2)$ or $O(n^3)$ to $O(n)$ in various components including numerical integration, risk set calculation, and standard error estimation, where $n$ is the number of subjects. Using both simulated and real world biobank data, we demonstrate that these linear scan algorithms generate drastic speed-up of up to hundreds of thousands fold when $n>10^4$, sometimes reducing the run-time from days to minutes. We have developed an R-package, FastJM, based on the proposed algorithms for joint modeling of longitudinal and time-to-event data with and without competing risks, and made it publicly available on the Comprehensive R Archive Network (CRAN).

preprint2022arXiv

Extensions to the Proximal Distance Method of Constrained Optimization

The current paper studies the problem of minimizing a loss $f(\boldsymbol{x})$ subject to constraints of the form $\boldsymbol{D}\boldsymbol{x} \in S$, where $S$ is a closed set, convex or not, and $\boldsymbol{D}$ is a matrix that fuses parameters. Fusion constraints can capture smoothness, sparsity, or more general constraint patterns. To tackle this generic class of problems, we combine the Beltrami-Courant penalty method with the proximal distance principle. The latter is driven by minimization of penalized objectives $f(\boldsymbol{x})+\fracρ{2}\text{dist}(\boldsymbol{D}\boldsymbol{x},S)^2$ involving large tuning constants $ρ$ and the squared Euclidean distance of $\boldsymbol{D}\boldsymbol{x}$ from $S$. The next iterate $\boldsymbol{x}_{n+1}$ of the corresponding proximal distance algorithm is constructed from the current iterate $\boldsymbol{x}_n$ by minimizing the majorizing surrogate function $f(\boldsymbol{x})+\fracρ{2}\|\boldsymbol{D}\boldsymbol{x}-\mathcal{P}_{S}(\boldsymbol{D}\boldsymbol{x}_n)\|^2$. For fixed $ρ$ and a subanalytic loss $f(\boldsymbol{x})$ and a subanalytic constraint set $S$, we prove convergence to a stationary point. Under stronger assumptions, we provide convergence rates and demonstrate linear local convergence. We also construct a steepest descent (SD) variant to avoid costly linear system solves. To benchmark our algorithms, we compare against the alternating direction method of multipliers (ADMM). Our extensive numerical tests include problems on metric projection, convex regression, convex clustering, total variation image denoising, and projection of a matrix to a good condition number. These experiments demonstrate the superior speed and acceptable accuracy of our steepest variant on high-dimensional problems.

preprint2022arXiv

Novel and self-consistency analysis of the QCD running coupling $α_s(Q)$ in both the perturbative and nonperturbative domains

The QCD coupling $α_s$ is the most important parameter for achieving precise QCD predictions. By using the well measured effective coupling $α^{g_1}_{s}(Q)$ defined from the Bjorken sum rules as a basis, we suggest a novel and self-consistency way to fix the $α_s$ at all scales: The QCD light-front holographic model is adopted for its infrared behavior, and the fixed-order pQCD prediction under the principle of maximum conformality (PMC) is used for its high-energy behavior. Using the PMC scheme-and-scale independent perturbative series, and by transforming it into the one under the physical $V$-scheme, we observe that a precise $α_s$ running behavior in both the perturbative and nonperturbative domains with a smooth transition from small to large scales can be achieved.

preprint2022arXiv

Reheating constraints on modified single-field Natural Inflation models

In this paper, we discuss three modified single-field natural inflation models in detail, including Special generalized Natural Inflation model(SNI), Extended Natural Inflation model(ENI) and Natural Inflation inspired model(NII). We derive the analytical expression of the tensor-to-scalar ratio $r$ and the spectral index $n_s$ for those models. Then the reheating temperature $T_{re}$ and reheating duration $N_{re}$ are analytically derived. Moreover, considering the CMB constraints, the feasible space of the SNI model in $(n_s, r)$ plane is almost covered by that of the NII, which means the NII is more general than the SNI. In addition, there is no overlapping space between the ENI and the other two models in $(n_s, r)$ plane, which indicates that the ENI and the other two models exclude each other, and more accurate experiments can verify them. Furthermore, the reheating brings tighter constraints to the inflation models, but they still work for a different reheating universe. Considering the constraints of $n_s$, $r$, $N_k$ and choosing $T_{re}$ near the electroweak energy scale, one can find that the decay constants of the three models have no overlapping area and the effective equations of state $ω_{re}$ should be within $\frac{1}{4}\lesssim ω_{re} \lesssim \frac{4}{5}$ for the three models.

preprint2022arXiv

Room-Temperature Valence Transition in a Strain-Tuned Perovskite Oxide

Cobalt oxides have long been understood to display intriguing phenomena known as spin-state crossovers, where the cobalt ion spin changes vs. temperature, pressure, etc. A very different situation was recently uncovered in praseodymium-containing cobalt oxides, where a first-order coupled spin-state/structural/metal-insulator transition occurs, driven by a remarkable praseodymium valence transition. Such valence transitions, particularly when triggering spin-state and metal-insulator transitions, offer highly appealing functionality, but have thus far been confined to cryogenic temperatures in bulk materials (e.g., 90 K in Pr1-xCaxCoO3). Here, we show that in thin films of the complex perovskite (Pr1-yYy)1-xCaxCoO3-δ, heteroepitaxial strain tuning enables stabilization of valence-driven spin-state/structural/metal-insulator transitions to at least 291 K, i.e., around room temperature. The technological implications of this result are accompanied by fundamental prospects, as complete strain control of the electronic ground state is demonstrated, from ferromagnetic metal under tension to nonmagnetic insulator under compression, thereby exposing a potential novel quantum critical point.

preprint2022arXiv

Thermodynamics and quark condensates of three-flavor QCD at low temperature

We use three-flavor chiral perturbation theory ($χ$PT) to calculate the pressure, light and $s$-quark condensates of QCD in the confined phase at finite temperature to ${\cal O}(p^6)$ in the low-energy expansion. We also include electromagnetic effects to order $e^2$, where the electromagnetic coupling $e$ counts as order $p$. Our results for the pressure and the condensates suggest that $χ$PT converges very well for temperatures up to approximately 150 MeV. We combine $χ$PT and the Hadron Resonance Gas (HRG) model by adding heavier baryons and mesons. Our results are compared with lattice simulations an d the agreement is very good for temperatures below {170} MeV, in contrast to the results from $χ$PT which agree with the lattice only up to $T\approx120$ MeV. Our value for the chiral crossover temperature is 160.1 MeV, which compares favorably to the lattice result of $157.3$ MeV.

preprint2021arXiv

Orthogonal Trace-Sum Maximization: Applications, Local Algorithms, and Global Optimality

This paper studies the problem of maximizing the sum of traces of matrix quadratic forms on a product of Stiefel manifolds. This orthogonal trace-sum maximization (OTSM) problem generalizes many interesting problems such as generalized canonical correlation analysis (CCA), Procrustes analysis, and cryo-electron microscopy of the Nobel prize fame. For these applications finding global solutions is highly desirable but it has been unclear how to find even a stationary point, let alone testing its global optimality. Through a close inspection of Ky Fan's classical result (1949) on the variational formulation of the sum of largest eigenvalues of a symmetric matrix, and a semidefinite programming (SDP) relaxation of the latter, we first provide a simple method to certify global optimality of a given stationary point of OTSM. This method only requires testing whether a symmetric matrix is positive semidefinite. A by-product of this analysis is an unexpected strong duality between Shapiro-Botha (1988) and Zhang-Singer (2017). After showing that a popular algorithm for generalized CCA and Procrustes analysis may generate oscillating iterates, we propose a simple fix that provably guarantees convergence to a stationary point. The combination of our algorithm and certificate reveals novel global optima of various instances of OTSM.

preprint2021arXiv

Unconventional hysteretic transition in a charge density wave

Hysteresis underlies a large number of phase transitions in solids, giving rise to exotic metastable states that are otherwise inaccessible. Here, we report an unconventional hysteretic transition in a quasi-2D material, EuTe4. By combining transport, photoemission, diffraction, and x-ray absorption measurements, we observed that the hysteresis loop has a temperature width of more than 400 K, setting a record among crystalline solids. The transition has an origin distinct from known mechanisms, lying entirely within the incommensurate charge-density-wave (CDW) phase of EuTe4 with no change in the CDW modulation periodicity. We interpret the hysteresis as an unusual switching of the relative CDW phases in different layers, a phenomenon unique to quasi-2D compounds that is not present in either purely 2D or strongly-coupled 3D systems. Our findings challenge the established theories on metastable states in density wave systems, pushing the boundary of understanding hysteretic transitions in a broken-symmetry state.

preprint2020arXiv

Correlation-driven eightfold magnetic anisotropy in a two-dimensional oxide monolayer

Engineering magnetic anisotropy in two-dimensional systems has enormous scientific and technological implications. The uniaxial anisotropy universally exhibited by two-dimensional magnets has only two stable spin directions, demanding 180 degrees spin switching between states. We demonstrate a novel eightfold anisotropy in magnetic SrRuO3 monolayers by inducing a spin reorientation in (SrRuO3)1/(SrTiO3)N superlattices, in which the magnetic easy axis of Ru spins is transformed from uniaxial <001> direction (N = 1 and 2) to eightfold <111> directions (N = 3, 4 and 5). This eightfold anisotropy enables 71 and 109 degrees spin switching in SrRuO3 monolayers, analogous to 71 and 109 degrees polarization switching in ferroelectric BiFeO3. First-principle calculations reveal that increasing the SrTiO3 layer thickness induces an emergent correlation-driven orbital ordering, tuning spin-orbit interactions and reorienting the SrRuO3 monolayer easy axis. Our work demonstrates that correlation effects can be exploited to substantially change spin-orbit interactions, stabilizing unprecedented properties in two-dimensional magnets and opening rich opportunities for low-power, multi-state device applications.

preprint2018arXiv

Electric-field Control of Magnetism with Emergent Topological Hall Effect in SrRuO3 through Proton Evolution

Ionic substitution forms an essential pathway to manipulate the carrier density and crystalline symmetry of materials via ion-lattice-electron coupling, leading to a rich spectrum of electronic states in strongly correlated systems. Using the ferromagnetic metal SrRuO3 as a model system, we demonstrate an efficient and reversible control of both carrier density and crystalline symmetry through the ionic liquid gating induced protonation. The insertion of protons electron-dopes SrRuO3, leading to an exotic ferromagnetic to paramagnetic phase transition along with the increase of proton concentration. Intriguingly, we observe an emergent topological Hall effect at the boundary of the phase transition as the consequence of the newly-established Dzyaloshinskii-Moriya interaction owing to the breaking of inversion symmetry in protonated SrRuO3 with the proton compositional film-depth gradient. We envision that electric-field controlled protonation opens a novel strategy to design material functionalities.

preprint2018arXiv

Persistence of Island Arrangements During Layer-by-Layer Growth Revealed Using Coherent X-rays

Understanding surface dynamics during epitaxial film growth is key to growing high quality materials with controllable properties. X-ray photon correlation spectroscopy (XPCS) using coherent x-rays opens new opportunities for in situ observation of atomic-scale fluctuation dynamics during crystal growth. Here, we present the first XPCS measurements of 2D island dynamics during homoepitaxial growth in the layer-by-layer mode. Analysis of the results using two-time correlations reveals a new phenomenon - a memory effect in island nucleation sites on successive crystal layers. Simulations indicate that this persistence in the island arrangements arises from communication between islands on different layers via adatoms. With the worldwide advent of new coherent x-ray sources, the XPCS methods pioneered here will be widely applicable to atomic-scale processes on surfaces.

preprint2016arXiv

Algorithms for Fitting the Constrained Lasso

We compare alternative computing strategies for solving the constrained lasso problem. As its name suggests, the constrained lasso extends the widely-used lasso to handle linear constraints, which allow the user to incorporate prior information into the model. In addition to quadratic programming, we employ the alternating direction method of multipliers (ADMM) and also derive an efficient solution path algorithm. Through both simulations and real data examples, we compare the different algorithms and provide practical recommendations in terms of efficiency and accuracy for various sizes of data. We also show that, for an arbitrary penalty matrix, the generalized lasso can be transformed to a constrained lasso, while the converse is not true. Thus, our methods can also be used for estimating a generalized lasso, which has wide-ranging applications. Code for implementing the algorithms is freely available in the Matlab toolbox SparseReg.

preprint2015arXiv

Co-GISAXS as a New Technique to Investigate Surface Growth Dynamics

Detailed quantitative measurement of surface dynamics during thin film growth is a major experimental challenge. Here X-ray Photon Correlation Spectroscopy with coherent hard X-rays is used in a Grazing-Incidence Small-Angle X-ray Scattering (i.e. Co-GISAXS) geometry as a new tool to investigate nanoscale surface dynamics during sputter deposition of a-Si and a-WSi$_2$ thin films. For both films, kinetic roughening during surface growth reaches a dynamic steady state at late times in which the intensity autocorrelation function $g_2$(q,t) becomes stationary. The $g_2$(q,t) functions exhibit compressed exponential behavior at all wavenumbers studied. The overall dynamics are complex, but the most surface sensitive sections of the structure factor and correlation time exhibit power law behaviors consistent with dynamical scaling.

preprint2015arXiv

MM Algorithms for Variance Components Models

Variance components estimation and mixed model analysis are central themes in statistics with applications in numerous scientific disciplines. Despite the best efforts of generations of statisticians and numerical analysts, maximum likelihood estimation and restricted maximum likelihood estimation of variance component models remain numerically challenging. Building on the minorization-maximization (MM) principle, this paper presents a novel iterative algorithm for variance components estimation. MM algorithm is trivial to implement and competitive on large data problems. The algorithm readily extends to more complicated problems such as linear mixed models, multivariate response models possibly with missing data, maximum a posteriori estimation, penalized estimation, and generalized estimating equations (GEE). We establish the global convergence of the MM algorithm to a KKT point and demonstrate, both numerically and theoretically, that it converges faster than the classical EM algorithm when the number of variance components is greater than two and all covariance matrices are positive definite.

preprint2015arXiv

Using coherent X-rays to directly measure the propagation velocity of defects during thin film deposition

The properties of artificially grown thin films are often strongly affected by the dynamic relationship between surface growth processes and subsurface structure. Coherent mixing of X-ray signals promises to provide an approach to better understand such processes. Here, we demonstrate the continuously variable mixing of surface and bulk scattering signals during real-time studies of sputter deposition of a-Si and a-WiS2 films by controlling the X-ray penetration and escape depths in coherent grazing incidence small angle X-ray scattering (Co-GISAXS). Under conditions where the X-ray signal comes from both the growth surface and the thin film bulk, oscillations in temporal correlations arise from coherent interference between scattering from stationary bulk features and from the advancing surface. We also observe evidence that elongated bulk features propagate upward at the same velocity as the surface. Additionally, a highly surface sensitive mode is demonstrated that can access the surface dynamics independently of the subsurface structure.

preprint2014arXiv

Evolution Process of Wurtzite ZnO Films on Cubic MgO (001) Substrates: a Structural, Optical and Electronic Investigation of the Misfit Structures

The interface between hexagonal ZnO films and cubic MgO (001) substrates, fabricated through molecular beam epitaxy, are thoroughly investigated. X-ray diffraction and (scanning) transmission electron microscopy reveal that, at the substrate temperature above 200 degree C, the growth follows the single [0001] direction; while at the substrate below 150 degree C, the growth is initially along [0001] and then mainly changes to [0-332] variants beyond the thickness of about 10 nm. Interestingly, a double-domain feature with a rotational angle of 30 degree appears for the growth along [0001] regardless of the growth temperature, experimentally demonstrated the theoretical predictions for occurrence of double rotational domains in such a heteroepitaxy [Grundmann et al, Phys. Rev. Lett. 105, 146102 (2010)]. It is also found that, the optical transmissivity of the ZnO film is greatly influenced by the mutation of growth directions, stimulated by the bond-length modulations, as further determined by X-ray absorption Spectra (XAS) at Zn K edge. The XAS results also show the evolution of 4pxy and 4pz states in the conduction band as the growth temperature increases. The results obtained from this work can hopefully promote the applications of ZnO in advanced optoelectronics for which its integration with other materials of different phases is desirable.

preprint2014arXiv

Fast Genome-Wide QTL Analysis Using Mendel

Pedigree GWAS (Option 29) in the current version of the Mendel software is an optimized subroutine for performing large scale genome-wide QTL analysis. This analysis (a) works for random sample data, pedigree data, or a mix of both, (b) is highly efficient in both run time and memory requirement, (c) accommodates both univariate and multivariate traits, (d) works for autosomal and x-linked loci, (e) correctly deals with missing data in traits, covariates, and genotypes, (f) allows for covariate adjustment and constraints among parameters, (g) uses either theoretical or SNP-based empirical kinship matrix for additive polygenic effects, (h) allows extra variance components such as dominant polygenic effects and household effects, (i) detects and reports outlier individuals and pedigrees, and (j) allows for robust estimation via the $t$-distribution. The current paper assesses these capabilities on the genetics analysis workshop 19 (GAW19) sequencing data. We analyzed simulated and real phenotypes for both family and random sample data sets. For instance, when jointly testing the 8 longitudinally measured systolic blood pressure (SBP) and diastolic blood pressure (DBP) traits, it takes Mendel 78 minutes on a standard laptop computer to read, quality check, and analyze a data set with 849 individuals and 8.3 million SNPs. Genome-wide eQTL analysis of 20,643 expression traits on 641 individuals with 8.3 million SNPs takes 30 hours using 20 parallel runs on a cluster. Mendel is freely available at \url{http://www.genetics.ucla.edu/software}.

preprint2014arXiv

Fast Genome-Wide QTL Association Mapping on Pedigree and Population Data

Since most analysis software for genome-wide association studies (GWAS) currently exploit only unrelated individuals, there is a need for efficient applications that can handle general pedigree data or mixtures of both population and pedigree data. Even data sets thought to consist of only unrelated individuals may include cryptic relationships that can lead to false positives if not discovered and controlled for. In addition, family designs possess compelling advantages. They are better equipped to detect rare variants, control for population stratification, and facilitate the study of parent-of-origin effects. Pedigrees selected for extreme trait values often segregate a single gene with strong effect. Finally, many pedigrees are available as an important legacy from the era of linkage analysis. Unfortunately, pedigree likelihoods are notoriously hard to compute. In this paper we re-examine the computational bottlenecks and implement ultra-fast pedigree-based GWAS analysis. Kinship coefficients can either be based on explicitly provided pedigrees or automatically estimated from dense markers. Our strategy (a) works for random sample data, pedigree data, or a mix of both; (b) entails no loss of power; (c) allows for any number of covariate adjustments, including correction for population stratification; (d) allows for testing SNPs under additive, dominant, and recessive models; and (e) accommodates both univariate and multivariate quantitative traits. On a typical personal computer (6 CPU cores at 2.67 GHz), analyzing a univariate HDL (high-density lipoprotein) trait from the San Antonio Family Heart Study (935,392 SNPs on 1357 individuals in 124 pedigrees) takes less than 2 minutes and 1.5 GB of memory. Complete multivariate QTL analysis of the three time-points of the longitudinal HDL multivariate trait takes less than 5 minutes and 1.5 GB of memory.

preprint2014arXiv

Interfacial Ionic Liquids: Connecting Static and Dynamic Structures

It is well-known that room temperature ionic liquids (RTILs) often adopt a charge-separated layered structure, i.e., with alternating cation- and anion-rich layers, at electrified interfaces. However, the dynamic response of the layered structure to temporal variations in applied potential is not well understood. We used in situ, real-time X-ray reflectivity (XR) to study the potential-dependent electric double layer (EDL) structure of an imidazolium-based RTIL on charged epitaxial graphene during potential cycling as a function of temperature. The results suggest that the graphene-RTIL interfacial structure is bistable in which the EDL structure at any intermediate potential can be described by the combination of two extreme-potential structures whose proportions vary depending on the polarity and magnitude of the applied potential. This picture is supported by the EDL structures obtained by fully atomistic molecular dynamics (MD) simulations at various static potentials. The potential-driven transition between the two structures is characterized by an increasing width but with an approximately fixed hysteresis magnitude as a function of temperature. The results are consistent with the coexistence of distinct anion and cation adsorbed structures separated by an energy barrier (~0.15 eV).

preprint2014arXiv

IsoDOT Detects Differential RNA-isoform Expression/Usage with respect to a Categorical or Continuous Covariate with High Sensitivity and Specificity

We have developed a statistical method named IsoDOT to assess differential isoform expression (DIE) and differential isoform usage (DIU) using RNA-seq data. Here isoform usage refers to relative isoform expression given the total expression of the corresponding gene. IsoDOT performs two tasks that cannot be accomplished by existing methods: to test DIE/DIU with respect to a continuous covariate, and to test DIE/DIU for one case versus one control. The latter task is not an uncommon situation in practice, e.g., comparing paternal and maternal allele of one individual or comparing tumor and normal sample of one cancer patient. Simulation studies demonstrate the high sensitivity and specificity of IsoDOT. We apply IsoDOT to study the effects of haloperidol treatment on mouse transcriptome and identify a group of genes whose isoform usages respond to haloperidol treatment.

preprint2014arXiv

Tensor Generalized Estimating Equations for Longitudinal Imaging Analysis

In an increasing number of neuroimaging studies, brain images, which are in the form of multidimensional arrays (tensors), have been collected on multiple subjects at multiple time points. Of scientific interest is to analyze such massive and complex longitudinal images to diagnose neurodegenerative disorders and to identify disease relevant brain regions. In this article, we treat those problems in a unifying regression framework with image predictors, and propose tensor generalized estimating equations (GEE) for longitudinal imaging analysis. The GEE approach takes into account intra-subject correlation of responses, whereas a low rank tensor decomposition of the coefficient array enables effective estimation and prediction with limited sample size. We propose an efficient estimation algorithm, study the asymptotics in both fixed $p$ and diverging $p$ regimes, and also investigate tensor GEE with regularization that is particularly useful for region selection. The efficacy of the proposed tensor GEE is demonstrated on both simulated data and a real data set from the Alzheimer's Disease Neuroimaging Initiative (ADNI).

preprint2013arXiv

Distance Majorization and Its Applications

The problem of minimizing a continuously differentiable convex function over an intersection of closed convex sets is ubiquitous in applied mathematics. It is particularly interesting when it is easy to project onto each separate set, but nontrivial to project onto their intersection. Algorithms based on Newton's method such as the interior point method are viable for small to medium-scale problems. However, modern applications in statistics, engineering, and machine learning are posing problems with potentially tens of thousands of parameters or more. We revisit this convex programming problem and propose an algorithm that scales well with dimensionality. Our proposal is an instance of a sequential unconstrained minimization technique and revolves around three ideas: the majorization-minimization (MM) principle, the classical penalty method for constrained optimization, and quasi-Newton acceleration of fixed-point algorithms. The performance of our distance majorization algorithms is illustrated in several applications.

preprint2013arXiv

Tucker Tensor Regression and Neuroimaging Analysis

Large-scale neuroimaging studies have been collecting brain images of study individuals, which take the form of two-dimensional, three-dimensional, or higher dimensional arrays, also known as tensors. Addressing scientific questions arising from such data demands new regression models that take multidimensional arrays as covariates. Simply turning an image array into a long vector causes extremely high dimensionality that compromises classical regression methods, and, more seriously, destroys the inherent spatial structure of array data that possesses wealth of information. In this article, we propose a family of generalized linear tensor regression models based upon the Tucker decomposition of regression coefficient arrays. Effectively exploiting the low rank structure of tensor covariates brings the ultrahigh dimensionality to a manageable level that leads to efficient estimation. We demonstrate, both numerically that the new model could provide a sound recovery of even high rank signals, and asymptotically that the model is consistently estimating the best Tucker structure approximation to the full array model in the sense of Kullback-Liebler distance. The new model is also compared to a recently proposed tensor regression model that relies upon an alternative CANDECOMP/PARAFAC (CP) decomposition.

preprint2012arXiv

A Generic Path Algorithm for Regularized Statistical Estimation

Regularization is widely used in statistics and machine learning to prevent overfitting and gear solution towards prior information. In general, a regularized estimation problem minimizes the sum of a loss function and a penalty term. The penalty term is usually weighted by a tuning parameter and encourages certain constraints on the parameters to be estimated. Particular choices of constraints lead to the popular lasso, fused-lasso, and other generalized $l_1$ penalized regression methods. Although there has been a lot of research in this area, developing efficient optimization methods for many nonseparable penalties remains a challenge. In this article we propose an exact path solver based on ordinary differential equations (EPSODE) that works for any convex loss function and can deal with generalized $l_1$ penalties as well as more complicated regularization such as inequality constraints encountered in shape-restricted regressions and nonparametric density estimation. In the path following process, the solution path hits, exits, and slides along the various constraints and vividly illustrates the tradeoffs between goodness of fit and model parsimony. In practice, the EPSODE can be coupled with AIC, BIC, $C_p$ or cross-validation to select an optimal tuning parameter. Our applications to generalized $l_1$ regularized generalized linear models, shape-restricted regressions, Gaussian graphical models, and nonparametric density estimation showcase the potential of the EPSODE algorithm.

preprint2012arXiv

Atomic layer engineering of perovskite oxides for chemically sharp heterointerfaces

Advances in synthesis techniques and materials understanding have given rise to oxide heterostructures with intriguing physical phenomena that cannot be found in their constituents. In these structures, precise control of interface quality, including oxygen stoichiometry, is critical for unambiguous tailoring of the interfacial properties, with deposition of the first monolayer being the most important step in shaping a well-defined functional interface. Here, we studied interface formation and strain evolution during the initial growth of LaAlO3 on SrTiO3 by pulsed laser deposition, in search of a means for controlling the atomic-sharpness of the interfaces. Our experimental results show that growth of LaAlO3 at a high oxygen pressure dramatically enhances interface abruptness. As a consequence, the critical thickness for strain relaxation was increased, facilitating coherent epitaxy of perovskite oxides. This provides a clear understanding of the role of oxygen pressure during the interface formation, and enables the synthesis of oxide heterostructures with chemically-sharper interfaces.

preprint2012arXiv

Path Following and Empirical Bayes Model Selection for Sparse Regression

In recent years, a rich variety of regularization procedures have been proposed for high dimensional regression problems. However, tuning parameter choice and computational efficiency in ultra-high dimensional problems remain vexing issues. The routine use of $\ell_1$ regularization is largely attributable to the computational efficiency of the LARS algorithm, but similar efficiency for better behaved penalties has remained elusive. In this article, we propose a highly efficient path following procedure for combination of any convex loss function and a broad class of penalties. From a Bayesian perspective, this algorithm rapidly yields maximum a posteriori estimates at different hyper-parameter values. To bypass the inefficiency and potential instability of cross validation, we propose an empirical Bayes procedure for rapidly choosing the optimal model and corresponding hyper-parameter value. This approach applies to any penalty that corresponds to a proper prior distribution on the regression coefficients. While we mainly focus on sparse estimation of generalized linear models, the method extends to more general regularizations such as polynomial trend filtering after reparameterization. The proposed algorithm scales efficiently to large $p$ and/or $n$. Solution paths of 10,000 dimensional examples are computed within one minute on a laptop for various generalized linear models (GLM). Operating characteristics are assessed through simulation studies and the methods are applied to several real data sets.

preprint2012arXiv

Path Following in the Exact Penalty Method of Convex Programming

Classical penalty methods solve a sequence of unconstrained problems that put greater and greater stress on meeting the constraints. In the limit as the penalty constant tends to $\infty$, one recovers the constrained solution. In the exact penalty method, squared penalties are replaced by absolute value penalties, and the solution is recovered for a finite value of the penalty constant. In practice, the kinks in the penalty and the unknown magnitude of the penalty constant prevent wide application of the exact penalty method in nonlinear programming. In this article, we examine a strategy of path following consistent with the exact penalty method. Instead of performing optimization at a single penalty constant, we trace the solution as a continuous function of the penalty constant. Thus, path following starts at the unconstrained solution and follows the solution path as the penalty constant increases. In the process, the solution path hits, slides along, and exits from the various constraints. For quadratic programming, the solution path is piecewise linear and takes large jumps from constraint to constraint. For a general convex program, the solution path is piecewise smooth, and path following operates by numerically solving an ordinary differential equation segment by segment. Our diverse applications to a) projection onto a convex set, b) nonnegative least squares, c) quadratically constrained quadratic programming, d) geometric programming, and e) semidefinite programming illustrate the mechanics and potential of path following. The final detour to image denoising demonstrates the relevance of path following to regularized estimation in inverse problems. In regularized estimation, one follows the solution path as the penalty constant decreases from a large value.

preprint2012arXiv

Regularized Matrix Regression

Modern technologies are producing a wealth of data with complex structures. For instance, in two-dimensional digital imaging, flow cytometry, and electroencephalography, matrix type covariates frequently arise when measurements are obtained for each combination of two underlying variables. To address scientific questions arising from those data, new regression methods that take matrices as covariates are needed, and sparsity or other forms of regularization are crucial due to the ultrahigh dimensionality and complex structure of the matrix data. The popular lasso and related regularization methods hinge upon the sparsity of the true signal in terms of the number of its nonzero coefficients. However, for the matrix data, the true signal is often of, or can be well approximated by, a low rank structure. As such, the sparsity is frequently in the form of low rank of the matrix parameters, which may seriously violate the assumption of the classical lasso. In this article, we propose a class of regularized matrix regression methods based on spectral regularization. Highly efficient and scalable estimation algorithm is developed, and a degrees of freedom formula is derived to facilitate model selection along the regularization path. Superior performance of the proposed method is demonstrated on both synthetic and real examples.

preprint2012arXiv

Tensor Regression with Applications in Neuroimaging Data Analysis

Classical regression methods treat covariates as a vector and estimate a corresponding vector of regression coefficients. Modern applications in medical imaging generate covariates of more complex form such as multidimensional arrays (tensors). Traditional statistical and computational methods are proving insufficient for analysis of these high-throughput data due to their ultrahigh dimensionality as well as complex structure. In this article, we propose a new family of tensor regression models that efficiently exploit the special structure of tensor covariates. Under this framework, ultrahigh dimensionality is reduced to a manageable level, resulting in efficient estimation and prediction. A fast and highly scalable estimation algorithm is proposed for maximum likelihood estimation and its associated asymptotic properties are studied. Effectiveness of the new methods is demonstrated on both synthetic and real MRI imaging data.

preprint2012arXiv

Understanding Controls on Interfacial Wetting at Epitaxial Graphene: Experiment and Theory

The interaction of interfacial water with graphitic carbon at the atomic scale is studied as a function of the hydrophobicity of epitaxial graphene. High resolution X-ray reflectivity shows that the graphene-water contact angle is controlled by the average graphene thickness, due to the fraction of the film surface expressed as the epitaxial buffer layer whose contact angle (contact angle θ_c = 73°) is substantially smaller than that of multilayer graphene (θ_c = 93°). Classical and ab initio molecular dynamics simulations show that the reduced contact angle of the buffer layer is due to both its epitaxy with the SiC substrate and the presence of interfacial defects. This insight clarifies the relationship between interfacial water structure and hydrophobicity, in general, and suggests new routes to control interface properties of epitaxial graphene.

preprint2011arXiv

A Path Algorithm for Constrained Estimation

Many least squares problems involve affine equality and inequality constraints. Although there are variety of methods for solving such problems, most statisticians find constrained estimation challenging. The current paper proposes a new path following algorithm for quadratic programming based on exact penalization. Similar penalties arise in $l_1$ regularization in model selection. Classical penalty methods solve a sequence of unconstrained problems that put greater and greater stress on meeting the constraints. In the limit as the penalty constant tends to $\infty$, one recovers the constrained solution. In the exact penalty method, squared penalties are replaced by absolute value penalties, and the solution is recovered for a finite value of the penalty constant. The exact path following method starts at the unconstrained solution and follows the solution path as the penalty constant increases. In the process, the solution path hits, slides along, and exits from the various constraints. Path following in lasso penalized regression, in contrast, starts with a large value of the penalty constant and works its way downward. In both settings, inspection of the entire solution path is revealing. Just as with the lasso and generalized lasso, it is possible to plot the effective degrees of freedom along the solution path. For a strictly convex quadratic program, the exact penalty algorithm can be framed entirely in terms of the sweep operator of regression analysis. A few well chosen examples illustrate the mechanics and potential of path following.

preprint2011arXiv

Growth of macroscopic-area single crystal polyacene thin films on arbitrary substrates

Organic electronic materials have potential applications in a number of low-cost, large area electronic devices such as flat panel displays and inexpensive solar panels. Small molecules in the series Anthracene, Tetracene, Pentacene, are model molecules for organic semiconductor thin films to be used as the active layers in such devices. This has motivated a number of studies of polyacene thin film growth and structure. Although the majority of these studies rely on vapor-deposited films, solvent-based deposition of films with improved properties onto non-crystalline substrates is desired for industrial production of devices. Improved ordering in thin films will have a large impact on their electronic properties, since grain boundaries and other defects are detrimental to carrier mobilities and lifetimes. Thus, a long-standing challenge in this field is to prepare largearea single crystal films on arbitrary substrates. Here we demonstrate a solvent-based method to deposit thin films of organic semiconductors, which is very general. Anthracene thin films with single-crystal domain sizes exceeding 1x1 cm2 can be prepared on various substrates by the technique. Films of 6,13-bis(triisopropylsilylethynyl)pentacene are also demonstrated to have grain sizes larger than 2x2 mm2. In contrast, films produced by conventional means such as vapor deposition or spin coating are polycrystalline with micron-scale grain sizes. The general propensity of these small molecules towards crystalline order in an optimized solvent deposition process shows that there is great potential for thin films with improved properties. Films prepared by these methods will also be useful in exploring the limits of performance in organic thin film devices.

preprint2010arXiv

Graphics Processing Units and High-Dimensional Optimization

This paper discusses the potential of graphics processing units (GPUs) in high-dimensional optimization problems. A single GPU card with hundreds of arithmetic cores can be inserted in a personal computer and dramatically accelerates many statistical algorithms. To exploit these devices fully, optimization algorithms should reduce to multiple parallel tasks, each accessing a limited amount of data. These criteria favor EM and MM algorithms that separate parameters and data. To a lesser extent block relaxation and coordinate descent and ascent also qualify. We demonstrate the utility of GPUs in nonnegative matrix factorization, PET image reconstruction, and multidimensional scaling. Speedups of 100 fold can easily be attained. Over the next decade, GPUs will fundamentally alter the landscape of computational statistics. It is time for more statisticians to get on-board.

preprint2010arXiv

MM Algorithms for Geometric and Signomial Programming

This paper derives new algorithms for signomial programming, a generalization of geometric programming. The algorithms are based on a generic principle for optimization called the MM algorithm. In this setting, one can apply the geometric-arithmetic mean inequality and a supporting hyperplane inequality to create a surrogate function with parameters separated. Thus, unconstrained signomial programming reduces to a sequence of one-dimensional minimization problems. Simple examples demonstrate that the MM algorithm derived can converge to a boundary point or to one point of a continuum of minimum points. Conditions under which the minimum point is unique or occurs in the interior of parameter space are proved for geometric programming. Convergence to an interior point occurs at a linear rate. Finally, the MM framework easily accommodates equality and inequality constraints of signomial type. For the most important special case, constrained quadratic programming, the MM algorithm involves very simple updates.

preprint2010arXiv

Pressure-dependent transition from atoms to nanoparticles in magnetron sputtering: Effect on WSi2 film roughness and stress

We report on the transition between two regimes from several-atom clusters to much larger nanoparticles in Ar magnetron sputter deposition of WSi2, and the effect of nanoparticles on the properties of amorphous thin films and multilayers. Sputter deposition of thin films is monitored by in situ x-ray scattering, including x-ray reflectivity and grazing incidence small angle x-ray scattering. The results show an abrupt transition at an Ar background pressure Pc; the transition is associated with the threshold for energetic particle thermalization, which is known to scale as the product of the Ar pressure and the working distance between the magnetron source and the substrate surface. Below Pc smooth films are produced, while above Pc roughness increases abruptly, consistent with a model in which particles aggregate in the deposition flux before reaching the growth surface. The results from WSi2 films are correlated with in situ measurement of stress in WSi2/Si multilayers, which exhibits a corresponding transition from compressive to tensile stress at Pc. The tensile stress is attributed to coalescence of nanoparticles and the elimination of nano-voids.

preprint2009arXiv

Anomalous Expansion of the Copper-Apical Oxygen Distance in Superconducting La$_{2}$CuO$_{4}$ - La$_{1.55}$Sr$_{0.45}$CuO$_{4}$ Bilayers

We have introduced an improved X-ray phase-retrieval method with unprecedented speed of convergence and precision, and used it to determine with sub-Ångstrom resolution the complete atomic structure of an ultrathin superconducting bilayer film, composed of La$_{1.55}$Sr$_{0.45}$CuO$_{4}$ and La$_{2}$CuO$_{4}$ neither of which is superconducting by itself. The results show that phase-retrieval diffraction techniques enable accurate measurement of structural modifications in near-surface layers, which may be critically important for elucidation of surface-sensitive experiments. Specifically we find that close to the sample surface the unit cell size remains constant while the copper-apical oxygen distance shows a dramatic increase, by as much as 0.45 Å. The apical oxygen displacement is known to have a profound effect on the superconducting transition temperature.

preprint2006arXiv

Wavelength Tunability of Ion-bombardment Induced Ripples on Sapphire

A study of ripple formation on sapphire surfaces by 300-2000 eV Ar+ ion bombardment is presented. Surface characterization by in-situ synchrotron grazing incidence small angle x-ray scattering and ex-situ atomic force microscopy is performed in order to study the wavelength of ripples formed on sapphire (0001) surfaces. We find that the wavelength can be varied over a remarkably wide range-nearly two orders of magnitude-by changing the ion incidence angle. Within the linear theory regime, the ion induced viscous flow smoothing mechanism explains the general trends of the ripple wavelength at low temperature and incidence angles larger than 30. In this model, relaxation is confined to a few-nm thick damaged surface layer. The behavior at high temperature suggests relaxation by surface diffusion. However, strong smoothing is inferred from the observed ripple wavelength near normal incidence, which is not consistent with either surface diffusion or viscous flow relaxation.