Source author record

Peter Hoff

Peter Hoff appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
3topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2022arXiv

Core Shrinkage Covariance Estimation for Matrix-variate Data

A separable covariance model for a random matrix provides a parsimonious description of the covariances among the rows and among the columns of the matrix, and permits likelihood-based inference with a very small sample size. However, in many applications the assumption of exact separability is unlikely to be met, and data analysis with a separable model may overlook or misrepresent important dependence patterns in the data. In this article, we propose a compromise between separable and unstructured covariance estimation. We show how the set of covariance matrices may be uniquely parametrized in terms of the set of separable covariance matrices and a complementary set of "core" covariance matrices, where the core of a separable covariance matrix is the identity matrix. This parametrization defines a Kronecker-core decomposition of a covariance matrix. By shrinking the core of the sample covariance matrix with an empirical Bayes procedure, we obtain an estimator that can adapt to the degree of separability of the population covariance matrix.

preprint2022arXiv

Coverage Properties of Empirical Bayes Intervals

This note is an invited discussion of the article "Confidence Intervals for Nonparametric Empirical Bayes Analysis" by Ignatiadis and Wager. In this discussion, I review some goals of empirical Bayes data analysis and the contribution of Ignatiadis and Wager. Differences between across-group inference and group-specific inference are discussed. Standard empirical Bayes interval procedures focus on controlling the across-group average coverage rate. However, if group-specific inferences are of primary interest, confidence intervals with group-specific coverage control may be preferable.

preprint2022arXiv

Tests of Linear Hypotheses using Indirect Information

In multigroup data settings with small within-group sample sizes, standard $F$-tests of group-specific linear hypotheses can have low power, particularly if the within-group sample sizes are not large relative to the number of explanatory variables. To remedy this situation, in this article we derive alternative test statistics based on information-sharing across groups. Each group-specific test has potentially much larger power than the standard $F$-test, while still exactly maintaining a target type I error rate if the hypothesis for the group is true. The proposed test for a given group uses a statistic that has optimal marginal power under a prior distribution derived from the data of the other groups. This statistic approaches the usual $F$-statistic as the prior distribution becomes more diffuse, but approaches a limiting "cone" test statistic as the prior distribution becomes extremely concentrated. We compare the power and $p$-values of the cone test to that of the $F$-test in some high-dimensional asymptotic scenarios. An analysis of educational outcome data is provided, demonstrating empirically that the proposed test is more powerful than the $F$-test.

preprint2021arXiv

Existence and Uniqueness of the Kronecker Covariance MLE

In matrix-valued datasets the sampled matrices often exhibit correlations among both their rows and their columns. A useful and parsimonious model of such dependence is the matrix normal model, in which the covariances among the elements of a random matrix are parameterized in terms of the Kronecker product of two covariance matrices, one representing row covariances and one representing column covariance. An appealing feature of such a matrix normal model is that the Kronecker covariance structure allows for standard likelihood inference even when only a very small number of data matrices is available. For instance, in some cases a likelihood ratio test of dependence may be performed with a sample size of one. However, more generally the sample size required to ensure boundedness of the matrix normal likelihood or the existence of a unique maximizer depends in a complicated way on the matrix dimensions. This motivates the study of how large a sample size is needed to ensure that maximum likelihood estimators exist, and exist uniquely with probability one. Our main result gives precise sample size thresholds in the paradigm where the number of rows and the number of columns of the data matrices differ by at most a factor of two. Our proof uses invariance properties that allow us to consider data matrices in canonical form, as obtained from the Kronecker canonical form for matrix pencils.

preprint2020arXiv

The Stein Effect for Frechet Means

The Frechet mean is a useful description of location for a probability distribution on a metric space that is not necessarily a vector space. This article considers simultaneous estimation of multiple Frechet means from a decision-theoretic perspective, and in particular, the extent to which the unbiased estimator of a Frechet mean can be dominated by a generalization of the James-Stein shrinkage estimator. It is shown that if the metric space satisfies a non-positive curvature condition, then this generalized James-Stein estimator asymptotically dominates the unbiased estimator as the dimension of the space grows. These results hold for a large class of distributions on a variety of spaces - including Hilbert spaces - and therefore partially extend known results on the applicability of the James-Stein estimator to non-normal distributions on Euclidean spaces. Simulation studies on metric trees and symmetric-positive-definite matrices are presented, numerically demonstrating the efficacy of this generalized James-Stein estimator.

preprint2012arXiv

Bayesian sandwich posteriors for pseudo-true parameters

Under model misspecification, the MLE generally converges to the pseudo-true parameter, the parameter corresponding to the distribution within the model that is closest to the distribution from which the data are sampled. In many problems, the pseudo-true parameter corresponds to a population parameter of interest, and so a misspecified model can provide consistent estimation for this parameter. Furthermore, the well-known sandwich variance formula of Huber(1967) provides an asymptotically accurate sampling distribution for the MLE, even under model misspecification. However, confidence intervals based on a sandwich variance estimate may behave poorly for low sample sizes, partly due to the use of a plug-in estimate of the variance. From a Bayesian perspective, plug-in estimates of nuisance parameters generally underrepresent uncertainty in the unknown parameters, and averaging over such parameters is expected to give better performance. With this in mind, we present a Bayesian sandwich posterior distribution, whose likelihood is based on the sandwich sampling distribution of the MLE. This Bayesian approach allows for the incorporation of prior information about the parameter of interest, averages over uncertainty in the nuisance parameter and is asymptotically robust to model misspecification. In a small simulation study on estimating a regression parameter under heteroscedasticity, the addition of accurate prior information and the averaging over the nuisance parameter are both seen to improve the accuracy and calibration of confidence intervals for the parameter of interest.

preprint2012arXiv

Likelihoods for fixed rank nomination networks

Many studies that gather social network data use survey methods that lead to censored, missing or otherwise incomplete information. For example, the popular fixed rank nomination (FRN) scheme, often used in studies of schools and businesses, asks study participants to nominate and rank at most a small number of contacts or friends, leaving the existence other relations uncertain. However, most statistical models are formulated in terms of completely observed binary networks. Statistical analyses of FRN data with such models ignore the censored and ranked nature of the data and could potentially result in misleading statistical inference. To investigate this possibility, we compare parameter estimates obtained from a likelihood for complete binary networks to those from a likelihood that is derived from the FRN scheme, and therefore recognizes the ranked and censored nature of the data. We show analytically and via simulation that the binary likelihood can provide misleading inference, at least for certain model parameters that relate network ties to characteristics of individuals and pairs of individuals. We also compare these different likelihoods in a data analysis of several adolescent social networks. For some of these networks, the parameter estimates from the binary and FRN likelihoods lead to different conclusions, indicating the importance of analyzing FRN data with a method that accounts for the FRN survey design.