Source author record

Brian J. Reich

Brian J. Reich appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

27works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

27 published item(s)

preprint2026arXiv

Demonstrating the power and flexibility of variational assumptions for amortized neural posterior estimation in environmental applications

Classic Bayesian methods with complex models are frequently infeasible due to an intractable likelihood. Simulation-based inference methods, such as Approximate Bayesian Computing (ABC), calculate posteriors without accessing a likelihood function by leveraging the fact that data can be quickly simulated from the model, but converge slowly and/or poorly in high-dimensional settings. In this paper, we propose a framework for Bayesian posterior estimation by mapping data to posteriors of parameters using a neural network trained on data simulated from the complex model. Posterior distributions of model parameters are efficiently obtained by feeding observed data into the trained neural network. We show theoretically that our posteriors converge to the true posteriors in Kullback-Leibler divergence. Our approach yields computationally efficient and theoretically justified uncertainty quantification, which is lacking in existing simulation-based neural network approaches. Comprehensive simulation studies highlight our method's robustness and accuracy.

preprint2022arXiv

A Gaussian-process approximation to a spatial SIR process using moment closures and emulators

The dynamics that govern disease spread are hard to model because infections are functions of both the underlying pathogen as well as human or animal behavior. This challenge is increased when modeling how diseases spread between different spatial locations. Many proposed spatial epidemiological models require trade-offs to fit, either by abstracting away theoretical spread dynamics, fitting a deterministic model, or by requiring large computational resources for many simulations. We propose an approach that approximates the complex spatial spread dynamics with a Gaussian process. We first propose a flexible spatial extension to the well-known SIR stochastic process, and then we derive a moment-closure approximation to this stochastic process. This moment-closure approximation yields ordinary differential equations for the evolution of the means and covariances of the susceptibles and infectious through time. Because these ODEs are a bottleneck to fitting our model by MCMC, we approximate them using a low-rank emulator. This approximation serves as the basis for our hierarchical model for noisy, underreported counts of new infections by spatial location and time. We demonstrate using our model to conduct inference on simulated infections from the underlying, true spatial SIR jump process. We then apply our method to model counts of new Zika infections in Brazil from late 2015 through early 2016.

preprint2022arXiv

Distributed Inference for Spatial Extremes Modeling in High Dimensions

Extreme environmental events frequently exhibit spatial and temporal dependence. These data are often modeled using max stable processes (MSPs). MSPs are computationally prohibitive to fit for as few as a dozen observations, with supposed computationally-efficient approaches like the composite likelihood remaining computationally burdensome with a few hundred observations. In this paper, we propose a spatial partitioning approach based on local modeling of subsets of the spatial domain that delivers computationally and statistically efficient inference. Marginal and dependence parameters of the MSP are estimated locally on subsets of observations using censored pairwise composite likelihood, and combined using a modified generalized method of moments procedure. The proposed distributed approach is extended to estimate spatially varying coefficient models to deliver computationally efficient modeling of spatial variation in marginal parameters. We demonstrate consistency and asymptotic normality of estimators, and show empirically that our approach leads to a surprising reduction in bias of parameter estimates over a full data approach. We illustrate the flexibility and practicability of our approach through simulations and the analysis of streamflow data from the U.S. Geological Survey.

preprint2021arXiv

Estimating intervention effects on infectious disease control: the effect of community mobility reduction on Coronavirus spread

Understanding the effects of interventions, such as restrictions on community and large group gatherings, is critical to controlling the spread of COVID-19. Susceptible-Infectious-Recovered (SIR) models are traditionally used to forecast the infection rates but do not provide insights into the causal effects of interventions. We propose a spatiotemporal model that estimates the causal effect of changes in community mobility (intervention) on infection rates. Using an approximation to the SIR model and incorporating spatiotemporal dependence, the proposed model estimates a direct and indirect (spillover) effect of intervention. Under an interference and treatment ignorability assumption, this model is able to estimate causal intervention effects, and additionally allows for spatial interference between locations. Reductions in community mobility were measured by cell phone movement data. The results suggest that the reductions in mobility decrease Coronavirus cases 4 to 7 weeks after the intervention.

preprint2021arXiv

Instrumental variables, spatial confounding and interference

Unobserved spatial confounding variables are prevalent in environmental and ecological applications where the system under study is complex and the data are often observational. Instrumental variables (IVs) are a common way to address unobserved confounding; however, the efficacy of using IVs on spatial confounding is largely unknown. This paper explores the effectiveness of IVs in this situation -- with particular attention paid to the spatial scale of the instrument. We show that, in case of spatially-dependent treatments, IVs are most effective when they vary at a finer spatial resolution than the treatment. We investigate IV performance in extensive simulations and apply the model in the example of long term trends in the air pollution and cardiovascular mortality in the United States over 1990-2010. Finally, the IV approach is also extended to the spatial interference setting, in which treatments can affect nearby responses.

preprint2020arXiv

A Bayesian semi-parametric hybrid model for spatial extremes with unknown dependence structure

The max-stable process is an asymptotically justified model for spatial extremes. In particular, we focus on the hierarchical extreme-value process (HEVP), which is a particular max-stable process that is conducive to Bayesian computing. The HEVP and all max-stable process models are parametric and impose strong assumptions including that all marginal distributions belong to the generalized extreme value family and that nearby sites are asymptotically dependent. We generalize the HEVP by relaxing these assumptions to provide a wider class of marginal distributions via a Dirichlet process prior for the spatial random effects distribution. In addition, we present a hybrid max-mixture model that combines the strengths of the parametric and semi-parametric models. We show that this versatile max-mixture model accommodates both asymptotic independence and dependence and can be fit using standard Markov chain Monte Carlo algorithms. The utility of our model is evaluated in Monte Carlo simulation studies and application to Netherlands wind gust data.

preprint2020arXiv

A spatial causal analysis of wildland fire-contributed PM2.5 using numerical model output

Wildland fire smoke contains hazardous levels of fine particulate matter PM2.5, a pollutant shown to adversely effect health. Estimating fire attributable PM2.5 concentrations is key to quantifying the impact on air quality and subsequent health burden. This is a challenging problem since only total PM2.5 is measured at monitoring stations and both fire-attributable PM2.5 and PM2.5 from all other sources are correlated in space and time. We propose a framework for estimating fire-contributed PM2.5 and PM2.5 from all other sources using a novel causal inference framework and bias-adjusted chemical model representations of PM2.5 under counterfactual scenarios. The chemical model representation of PM2.5 for this analysis is simulated using Community Multi-Scale Air Quality Modeling System (CMAQ), run with and without fire emissions across the contiguous U.S. for the 2008-2012 wildfire seasons. The CMAQ output is calibrated with observations from monitoring sites for the same spatial domain and time period. We use a Bayesian model that accounts for spatial variation to estimate the effect of wildland fires on PM2.5 and state assumptions under which the estimate has a valid causal interpretation. Our results include estimates of absolute, relative and cumulative contributions of wildfire smoke to PM2.5 for the contiguous U.S. Additionally, we compute the health burden associated with the PM2.5 attributable to wildfire smoke.

preprint2020arXiv

A spatiotemporal recommendation engine for malaria control

Malaria is an infectious disease affecting a large population across the world, and interventions need to be efficiently applied to reduce the burden of malaria. We develop a framework to help policy-makers decide how to allocate limited resources in realtime for malaria control. We formalize a policy for the resource allocation as a sequence of decisions, one per intervention decision, that map up-to-date disease related information to a resource allocation. An optimal policy must control the spread of the disease while being interpretable and viewed as equitable to stakeholders. We construct an interpretable class of resource allocation policies that can accommodate allocation of resources residing in a continuous domain, and combine a hierarchical Bayesian spatiotemporal model for disease transmission with a policy-search algorithm to estimate an optimal policy for resource allocation within the pre-specified class. The estimated optimal policy under the proposed framework improves the cumulative long-term outcome compared with naive approaches in both simulation experiments and application to malaria interventions in the Democratic Republic of the Congo.

preprint2020arXiv

Accounting for Location Measurement Error in Imaging Data with Application to Atomic Resolution Images of Crystalline Materials

Scientists use imaging to identify objects of interest and infer properties of these objects. The locations of these objects are often measured with error, which when ignored leads to biased parameter estimates and inflated variance. Current measurement error methods require an estimate or knowledge of the measurement error variance to correct these estimates, which may not be available. Instead, we create a spatial Bayesian hierarchical model that treats the locations as parameters, it using the image itself to incorporate positional uncertainty. We lower the computational burden by approximating the likelihood using a non-contiguous block design around the object locations. We apply this model in a materials science setting to study the relationship between the chemistry and displacement of hundreds of atom columns in crystal structures directly imaged via scanning transmission electron microscopy. Greater knowledge of this relationship can lead to engineering materials with improved properties of interest. We find strong evidence of a negative relationship between atom column displacement and the intensity of neighboring atom columns, which is related to the local chemistry. A simulation study shows our method corrects the bias in the parameter of interest and drastically improves coverage in high noise scenarios compared to non-measurement error models.

preprint2020arXiv

Bayesian Regression Using a Prior on the Model Fit: The R2-D2 Shrinkage Prior

Prior distributions for high-dimensional linear regression require specifying a joint distribution for the unobserved regression coefficients, which is inherently difficult. We instead propose a new class of shrinkage priors for linear regression via specifying a prior first on the model fit, in particular, the coefficient of determination, and then distributing through to the coefficients in a novel way. The proposed method compares favourably to previous approaches in terms of both concentration around the origin and tail behavior, which leads to improved performance both in posterior contraction and in empirical performance. The limiting behavior of the proposed prior is $1/x$, both around the origin and in the tails. This behavior is optimal in the sense that it simultaneously lies on the boundary of being an improper prior both in the tails and around the origin. None of the existing shrinkage priors obtain this behavior in both regions simultaneously. We also demonstrate that our proposed prior leads to the same near-minimax posterior contraction rate as the spike-and-slab prior.

preprint2020arXiv

Constrained Bayesian Nonparametric Regression for Grain Boundary Energy Predictions

Grain boundary (GB) energy is a fundamental property that affects the form of grain boundary and plays an important role to unveil the behavior of polycrystalline materials. With a better understanding of grain boundary energy distribution (GBED), we can produce more durable and efficient materials that will further improve productivity and reduce loss. The lack of robust GB structure-property relationships still remains one of the biggest obstacles towards developing true bottom-up models for the behavior of polycrystalline materials. Progress has been slow because of the inherent complexity associated with the structure of interfaces and the vast five-dimensional configurational space in which they reside. Estimating the GBED is challenging from a statistical perspective because there are not direct measurements on the grain boundary energy. We only have indirect information in the form of an unidentifiable homogeneous set of linear equations. In this paper, we propose a new statistical model to determine the GBED from the microstructures of polycrystalline materials. We apply spline-based regression with constraints to successfully recover the GB energy surface. Hamiltonian Monte Carlo and Gibbs sampling are used for computation and model fitting. Compared with conventional methods, our method not only gives more accurate predictions but also provides prediction uncertainties.

preprint2020arXiv

Nonparametric Conditional Density Estimation In A Deep Learning Framework For Short-Term Forecasting

Short-term forecasting is an important tool in understanding environmental processes. In this paper, we incorporate machine learning algorithms into a conditional distribution estimator for the purposes of forecasting tropical cyclone intensity. Many machine learning techniques give a single-point prediction of the conditional distribution of the target variable, which does not give a full accounting of the prediction variability. Conditional distribution estimation can provide extra insight on predicted response behavior, which could influence decision-making and policy. We propose a technique that simultaneously estimates the entire conditional distribution and flexibly allows for machine learning techniques to be incorporated. A smooth model is fit over both the target variable and covariates, and a logistic transformation is applied on the model output layer to produce an expression of the conditional density function. We provide two examples of machine learning models that can be used, polynomial regression and deep learning models. To achieve computational efficiency we propose a case-control sampling approximation to the conditional distribution. A simulation study for four different data distributions highlights the effectiveness of our method compared to other machine learning-based conditional distribution estimation techniques. We then demonstrate the utility of our approach for forecasting purposes using tropical cyclone data from the Atlantic Seaboard. This paper gives a proof of concept for the promise of our method, further computational developments can fully unlock its insights in more complex forecasting and other applications.

preprint2020arXiv

Spatial shrinkage via the product independent Gaussian process prior

We study the problem of sparse signal detection on a spatial domain. We propose a novel approach to model continuous signals that are sparse and piecewise smooth as product of independent Gaussian processes (PING) with a smooth covariance kernel. The smoothness of the PING process is ensured by the smoothness of the covariance kernels of Gaussian components in the product, and sparsity is controlled by the number of components. The bivariate kurtosis of the PING process shows more components in the product results in thicker tail and sharper peak at zero. The simulation results demonstrate the improvement in estimation using the PING prior over Gaussian process (GP) prior for different image regressions. We apply our method to a longitudinal MRI dataset to detect the regions that are affected by multiple sclerosis (MS) in the greatest magnitude through an image-on-scalar regression model. Due to huge dimensionality of these images, we transform the data into the spectral domain and develop methods to conduct computation in this domain. In our MS imaging study, the estimates from the PING model are more informative than those from the GP model.

preprint2017arXiv

A Test for Isotropy on a Sphere using Spherical Harmonic Functions

Analysis of geostatistical data is often based on the assumption that the spatial random field is isotropic. This assumption, if erroneous, can adversely affect model predictions and statistical inference. Nowadays many applications consider data over the entire globe and hence it is necessary to check the assumption of isotropy on a sphere. In this paper, a test for spatial isotropy on a sphere is proposed. The data are first projected onto the set of spherical harmonic functions. Under isotropy, the spherical harmonic coefficients are uncorrelated whereas they are correlated if the underlying fields are not isotropic. This motivates a test based on the sample correlation matrix of the spherical harmonic coefficients. In particular, we use the largest eigenvalue of the sample correlation matrix as the test statistic. Extensive simulations are conducted to assess the Type I errors of the test under different scenarios. We show how temporal correlation affects the test and provide a method for handling temporal correlation. We also gauge the power of the test as we move away from isotropy. The method is applied to the near-surface air temperature data which is part of the HadCM3 model output. Although we do not expect global temperature fields to be isotropic, we propose several anisotropic models with increasing complexity, each of which has an isotropic process as model component and we apply the test to the isotropic component in a sequence of such models as a method of determining how well the models capture the anisotropy in the fields.

preprint2016arXiv

A Markov-switching model for heat waves

Heat waves merit careful study because they inflict severe economic and societal damage. We use an intuitive, informal working definition of a heat wave-a persistent event in the tail of the temperature distribution-to motivate an interpretable latent state extreme value model. A latent variable with dependence in time indicates membership in the heat wave state. The strength of the temporal dependence of the latent variable controls the frequency and persistence of heat waves. Within each heat wave, temperatures are modeled using extreme value distributions, with extremal dependence across time accomplished through an extreme value Markov model. One important virtue of interpretability is that model parameters directly translate into quantities of interest for risk management, so that questions like whether heat waves are becoming longer, more severe or more frequent are easily answered by querying an appropriate fitted model. We demonstrate the latent state model on two recent, calamitous, examples: the European heat wave of 2003 and the Russian heat wave of 2010.

preprint2016arXiv

Data Mining to Investigate the Meteorological Drivers for Extreme Ground Level Ozone Events

This project aims to explore which combinations of meteorological conditions are associated with extreme ground level ozone conditions. Our approach focuses only on the tail by optimizing the tail dependence between the ozone response and functions of meteorological covariates. Since there is a long list of possible meteorological covariates, the space of possible models cannot be explored completely. Consequently, we perform data mining within the model selection context, employing an automated model search procedure. Our study is unique among extremes applications as optimizing tail dependence has not previously been attempted, and it presents new challenges, such as requiring a smooth threshold. We present a simulation study which shows that the method can detect complicated conditions leading to extreme responses and resists overfitting. We apply the method to ozone data for Atlanta and Charlotte and find similar meteorological drivers for these two Southeastern US cities. We identify several covariates which help to differentiate the meteorological conditions which lead to extreme ozone levels from those which lead to merely high levels.

preprint2016arXiv

Modeling Multivariate Mixed-Response Functional Data

We propose a Bayesian modeling framework for jointly analyzing multiple functional responses of different types (e.g. binary and continuous data). Our approach is based on a multivariate latent Gaussian process and models the dependence among the functional responses through the dependence of the latent process. Our framework easily accommodates additional covariates. We offer a way to estimate the multivariate latent covariance, allowing for implementation of multivariate functional principal components analysis (FPCA) to specify basis expansions and simplify computation. We demonstrate our method through both simulation studies and an application to real data from a periodontal study.

preprint2016arXiv

Scalar-on-Image Regression via the Soft-Thresholded Gaussian Process

The focus of this work is on spatial variable selection for scalar-on-image regression. We propose a new class of Bayesian nonparametric models, soft-thresholded Gaussian processes and develop the efficient posterior computation algorithms. Theoretically, soft-thresholded Gaussian processes provide large prior support for the spatially varying coefficients that enjoy piecewise smoothness, sparsity and continuity, characterizing the important features of imaging data. Also, under some mild regularity conditions, the soft-thresholded Gaussian process leads to the posterior consistency for both parameter estimation and variable selection for scalar-on-image regression, even when the number of true predictors is larger than the sample size. The proposed method is illustrated via simulations, compared numerically with existing alternatives and applied to Electroencephalography (EEG) study of alcoholism.

preprint2015arXiv

Quantile regression for mixed models with an application to examine blood pressure trends in China

Cardiometabolic diseases have substantially increased in China in the past 20 years and blood pressure is a primary modifiable risk factor. Using data from the China Health and Nutrition Survey, we examine blood pressure trends in China from 1991 to 2009, with a concentration on age cohorts and urbanicity. Very large values of blood pressure are of interest, so we model the conditional quantile functions of systolic and diastolic blood pressure. This allows the covariate effects in the middle of the distribution to vary from those in the upper tail, the focal point of our analysis. We join the distributions of systolic and diastolic blood pressure using a copula, which permits the relationships between the covariates and the two responses to share information and enables probabilistic statements about systolic and diastolic blood pressure jointly. Our copula maintains the marginal distributions of the group quantile effects while accounting for within-subject dependence, enabling inference at the population and subject levels. Our population-level regression effects change across quantile level, year and blood pressure type, providing a rich environment for inference. To our knowledge, this is the first quantile function model to explicitly model within-subject autocorrelation and is the first quantile function approach that simultaneously models multivariate conditional response. We find that the association between high blood pressure and living in an urban area has evolved from positive to negative, with the strongest changes occurring in the upper tail. The increase in urbanization over the last twenty years coupled with the transition from the positive association between urbanization and blood pressure in earlier years to a more uniform association with urbanization suggests increasing blood pressure over time throughout China, even in less urbanized areas. Our methods are available in the R package BSquare.

preprint2014arXiv

A spatial capture-recapture model for territorial species

Advances in field techniques have lead to an increase in spatially-referenced capture-recapture data to estimate a species' population size as well as other demographic parameters and patterns of space usage. Statistical models for these data have assumed that the number of individuals in the population and their spatial locations follow a homogeneous Poisson point process model, which implies that the individuals are uniformly and independently distributed over the spatial domain of interest. In many applications there is reason to question independence, for example when species display territorial behavior. In this paper, we propose a new statistical model which allows for dependence between locations to account for avoidance or territorial behavior. We show via a simulation study that accounting for this can improve population size estimates. The method is illustrated using a case study of small mammal trapping data to estimate avoidance and population density of adult female field voles (Microtus agrestis) in northern England.

preprint2014arXiv

Modeling the effect of temperature on ozone-related mortality

Climate change is expected to alter the distribution of ambient ozone levels and temperatures which, in turn, may impact public health. Much research has focused on the effect of short-term ozone exposures on mortality and morbidity while controlling for temperature as a confounder, but less is known about the joint effects of ozone and temperature. The extent of the health effects of changing ozone levels and temperatures will depend on whether these effects are additive or synergistic. In this paper we propose a spatial, semi-parametric model to estimate the joint ozone-temperature risk surfaces in 95 US urban areas. Our methodology restricts the ozone-temperature risk surfaces to be monotone in ozone and allows for both nonadditive and nonlinear effects of ozone and temperature. We use data from the National Mortality and Morbidity Air Pollution Study (NMMAPS) and show that the proposed model fits the data better than additive linear and nonlinear models. We then examine the synergistic effect of ozone and temperature both nationally and locally and find evidence of a nonlinear ozone effect and an ozone-temperature interaction at higher temperatures and ozone concentrations.

preprint2013arXiv

A hierarchical max-stable spatial model for extreme precipitation

Extreme environmental phenomena such as major precipitation events manifestly exhibit spatial dependence. Max-stable processes are a class of asymptotically-justified models that are capable of representing spatial dependence among extreme values. While these models satisfy modeling requirements, they are limited in their utility because their corresponding joint likelihoods are unknown for more than a trivial number of spatial locations, preventing, in particular, Bayesian analyses. In this paper, we propose a new random effects model to account for spatial dependence. We show that our specification of the random effect distribution leads to a max-stable process that has the popular Gaussian extreme value process (GEVP) as a limiting case. The proposed model is used to analyze the yearly maximum precipitation from a regional climate model.

preprint2012arXiv

A class of covariate-dependent spatiotemporal covariance functions for the analysis of daily ozone concentration

In geostatistics, it is common to model spatially distributed phenomena through an underlying stationary and isotropic spatial process. However, these assumptions are often untenable in practice because of the influence of local effects in the correlation structure. Therefore, it has been of prolonged interest in the literature to provide flexible and effective ways to model nonstationarity in the spatial effects. Arguably, due to the local nature of the problem, we might envision that the correlation structure would be highly dependent on local characteristics of the domain of study, namely, the latitude, longitude and altitude of the observation sites, as well as other locally defined covariate information. In this work, we provide a flexible and computationally feasible way for allowing the correlation structure of the underlying processes to depend on local covariate information. We discuss the properties of the induced covariance functions and methods to assess its dependence on local covariate information. The proposed method is used to analyze daily ozone in the southeast United States.

preprint2012arXiv

Mixture Likelihood Ratio Scan Statistic for Disease Outbreak Detection

Early detection of disease outbreaks is of paramount importance to implementing intervention strategies to mitigate the severity and duration of the outbreak. We build methodology that utilizes the characteristic profile of disease outbreaks to reduce the time to detection and false positive rate. We model daily counts through a Poisson distribution with additive background plus outbreak components. The outbreak component has a parametric form with unknown underlying parameters. A mixture likelihood ratio scan statistic is developed to maximize parameters over a window in time. This provides an alert statistic with early time to detection and low false positive rate. The methodology is demonstrated on three simulated data sets meant to represent E. coli, Cryptosporidium, and Influenza outbreaks.

preprint2010arXiv

A latent factor model for spatial data with informative missingness

A large amount of data is typically collected during a periodontal exam. Analyzing these data poses several challenges. Several types of measurements are taken at many locations throughout the mouth. These spatially-referenced data are a mix of binary and continuous responses, making joint modeling difficult. Also, most patients have missing teeth. Periodontal disease is a leading cause of tooth loss, so it is likely that the number and location of missing teeth informs about the patient's periodontal health. In this paper we develop a multivariate spatial framework for these data which jointly models the binary and continuous responses as a function of a single latent spatial process representing general periodontal health. We also use the latent spatial process to model the location of missing teeth. We show using simulated and real data that exploiting spatial associations and jointly modeling the responses and locations of missing teeth mitigates the problems presented by these data.