Source author record

Reinhard Furrer

Reinhard Furrer appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

13works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

13 published item(s)

preprint2026arXiv

Iterative Methods for Full-Scale Gaussian Process Approximations for Large Spatial Data

Gaussian processes are flexible probabilistic regression models which are widely used in statistics and machine learning. However, a drawback is their limited scalability to large data sets. To alleviate this, full-scale approximations (FSAs) combine predictive process methods and covariance tapering, thus approximating both global and local structures. We show how iterative methods can be used to reduce computational costs in calculating likelihoods, gradients, and predictive distributions with FSAs. In particular, we introduce a novel preconditioner and show theoretically and empirically that it accelerates the conjugate gradient method's convergence speed and mitigates its sensitivity with respect to the FSA parameters and the eigenvalue structure of the original covariance matrix, and we demonstrate empirically that it outperforms a state-of-the-art pivoted Cholesky preconditioner. Furthermore, we introduce an accurate and fast way to calculate predictive variances using stochastic simulation and iterative methods. In addition, we show how our newly proposed fully independent training conditional (FITC) preconditioner can also be used in iterative methods for Vecchia approximations. In our experiments, it outperforms existing state-of-the-art preconditioners for Vecchia approximations. All methods are implemented in a free C++ software library with high-level Python and R packages.

preprint2022arXiv

Dominant-feature identification in data from Gaussian processes applied to Finnish forest inventory records

In spatial data, location-dependent variation leads to connected structures known as features. Variations occur at different spatial scales and possibly originate from distinct underlying processes. Each of these scales is characterized by its own dominant features. Here we introduce a statistical method for identifying these scales and their dominant features in data from Gaussian processes. This identification involves credibly recognizing the dominant features by scale-space decomposition and assessing feature attributes by estimating covariance function parameters of the underlying processes and their associations to potential drivers. We analyze Finnish forest inventory data from the 1920s using this dominant-feature identification method and identify the scales of variation in basal area estimates of most common Finnish trees, including Scots pine, Norway spruce, birch, and other native deciduous trees. Comparing the resulting scale-dependent features and their attributes in these tree species, we identify the different effects of edaphic and anthropogenic drivers on the spatial distribution of their basal areas. These data are analyzed for the first time in terms of their scale of variation, and the resulting scale-dependent maps and estimates are an essential contribution to the historical forest ecology of Fennoscandia. Until now, this analysis was not possible with conventional methods.

preprint2022arXiv

EggCounts: a Bayesian hierarchical toolkit to model faecal egg count reductions

This is a vignette for the R package eggCounts version 2.0. The package implements a suite of Bayesian hierarchical models dealing with faecal egg count reductions. The models are designed for a variety of practical situations, including individual treatment efficacy, zero inflation, small sample size (less than 10) and potential outliers. The functions are intuitive to use and their output are easy to interpret, such that users are protected from being exposed to complex Bayesian hierarchical modelling tasks. In addition, the package includes plotting functions to display data and results in a visually appealing manner. The models are implemented in Stan modelling language, which provides efficient sampling technique to obtain posterior samples. This vignette briefly introduces different models, and provides a short walk-through analysis with example data.

preprint2021arXiv

Joint Variable Selection of both Fixed and Random Effects for Gaussian Process-based Spatially Varying Coefficient Models

Spatially varying coefficient (SVC) models are a type of regression model for spatial data where covariate effects vary over space. If there are several covariates, a natural question is which covariates have a spatially varying effect and which not. We present a new variable selection approach for Gaussian process-based SVC models. It relies on a penalized maximum likelihood estimation (PMLE) and allows variable selection both with respect to fixed effects and Gaussian process random effects. We validate our approach both in a simulation study as well as a real world data set. Our novel approach shows good selection performance in the simulation study. In the real data application, our proposed PMLE yields sparser SVC models and achieves a smaller information criterion than classical MLE. In a cross-validation applied on the real data, we show that sparser PML estimated SVC models are on par with ML estimated SVC models with respect to predictive performance.

preprint2020arXiv

Asymptotically Equivalent Prediction in Multivariate Geostatistics

Cokriging is the common method of spatial interpolation (best linear unbiased prediction) in multivariate geostatistics. While best linear prediction has been well understood in univariate spatial statistics, the literature for the multivariate case has been elusive so far. The new challenges provided by modern spatial datasets, being typically multivariate, call for a deeper study of cokriging. In particular, we deal with the problem of misspecified cokriging prediction within the framework of fixed domain asymptotics. Specifically, we provide conditions for equivalence of measures associated with multivariate Gaussian random fields, with index set in a compact set of a d-dimensional Euclidean space. Such conditions have been elusive for over about 50 years of spatial statistics. We then focus on the multivariate Matérn and Generalized Wendland classes of matrix valued covariance functions, that have been very popular for having parameters that are crucial to spatial interpolation, and that control the mean square differentiability of the associated Gaussian process. We provide sufficient conditions, for equivalence of Gaussian measures, relying on the covariance parameters of these two classes. This enables to identify the parameters that are crucial to asymptotically equivalent interpolation in multivariate geostatistics. Our findings are then illustrated through simulation studies.

preprint2016arXiv

On the smallest eigenvalues of covariance matrices of multivariate spatial processes

There has been a growing interest in providing models for multivariate spatial processes. A majority of these models specify a parametric matrix covariance function. Based on observations, the parameters are estimated by maximum likelihood or variants thereof. While the asymptotic properties of maximum likelihood estimators for univariate spatial processes have been analyzed in detail, maximum likelihood estimators for multivariate spatial processes have not received their deserved attention yet. In this article we consider the classical increasing-domain asymptotic setting restricting the minimum distance between the locations. Then, one of the main components to be studied from a theoretical point of view is the asymptotic positive definiteness of the underlying covariance matrix. Based on very weak assumptions on the matrix covariance function we show that the smallest eigenvalue of the covariance matrix is asymptotically bounded away from zero. Several practical implications are discussed as well.

preprint2016arXiv

Valid parameter space of a bivariate Gaussian Markov random field with a generalized block-Toeplitz precision matrix

Gaussian Markov random fields (GMRFs) are extensively used in statistics to model area-based data and usually depend on several parameters in order to capture complex spatial correlations. In this context, it is important to determine the valid parameter space, namely the domain ensuring (semi) positive-definiteness of the precision matrix. Depending on the structure of the latter, this task can be challenging. While univari- ate GMRFs with block-Toeplitz precision are well studied in the literature, not much is analytically known about bivariate GMRFs. So far, only restrictive sufficient conditions and brute-force approaches were proposed, which are computationally expensive for the size of modern datasets. In this paper, we consider a bivariate GMRF, which is part of a hierarchical model used in spatial statistics to analyze data coming from projec- tions of regional climate change. By extending classical convergence results of univariate fields with toroidal boundary conditions to fields without boundary conditions, we pro- vide asymptotically closed-form expressions of the valid parameter space. We develop a general methodology that can be used to determine the valid parameter space of bivariate GMRFs whose precision matrix has a generalized block-Toeplitz structure and for which classical convergence results are not directly applicable. Finally, we quantify the rate of convergence of our approach through a numerical study in R.

preprint2014arXiv

Hierarchical modelling of faecal egg counts to assess anthelmintic efficacy

Counting the number of parasite eggs in faecal samples is a widely used diagnostic method to evaluate parasite burden. Typically a sub-sample of the diluted faeces is examined for eggs. The resulting egg counts are multiplied by a specific correction factor to estimate the mean parasite burden. To detect anthelmintic resistance, the mean parasite burden from treated and untreated animals are compared. However, this standard method has some limitations. In particular, the analysis of repeated samples may produce quite variable results as the sampling variability due to the counting technique is ignored. We propose a hierarchical model that takes this sampling variability as well as between-animal variation into account. Bayesian inference is done via Markov chain Monte Carlo sampling. The performance of the hierarchical model is illustrated by a re-analysis of faecal egg count data from a Swedish study assessing the anthelmintic resistance of nematode parasite in sheep. A simulation study shows that the hierarchical model provides better classification of anthelmintic resistance compared to the standard method.

preprint2013arXiv

Conjugate distributions in hierarchical Bayesian ANOVA for computational efficiency and assessments of both practical and statistical significance

Assessing variability according to distinct factors in data is a fundamental technique of statistics. The method commonly regarded to as analysis of variance (ANOVA) is, however, typically confined to the case where all levels of a factor are present in the data (i.e. the population of factor levels has been exhausted). Random and mixed effects models are used for more elaborate cases, but require distinct nomenclature, concepts and theory, as well as distinct inferential procedures. Following a hierarchical Bayesian approach, a comprehensive ANOVA framework is shown, which unifies the above statistical models, emphasizes practical rather than statistical significance, addresses issues of parameter identifiability for random effects, and provides straightforward computational procedures for inferential steps. Although this is done in a rigorous manner the contents herein can be seen as ideological in supporting a shift in the approach taken towards analysis of variance.

preprint2013arXiv

Intelligent Compaction and Quality Assurance of Roller Measurement Values utilizing Backfitting and Multiresolution Scale Space Analysis

Modern earthwork compaction rollers collect location and compaction information as they traverse a compaction site. These roller measurement values present a challenging spatio-temporal statistical problem that requires careful implementation of a proper stochastic model and estimation procedure. Heersink and Furrer (2013) proposed a sequential, spatial mixed-effects model and a sequential, spatial backfitting routine for estimation of the modeling terms for such data. The estimated fields produced from this backfitting procedure are analyzed using a multiresolution scale space analysis developed by Holmstrom et al. (2011). This image analysis is proposed as a viable solution to improved intelligent compaction and quality assurance of the compaction process.

preprint2013arXiv

Spatial Backfitting of Roller Measurement Values from a Florida Test Bed

Modern earthwork compaction rollers collect location and compaction information as they traverse a compaction site. These data are indirectly observed through non-linear measurement operators, inherently multivariate with complex correlation structures, and collected in huge quantities. The nature of such data was investigated at a large, atypically compacted test bed in Florida, USA. Exploratory analysis of this data through detrending and empirical semivariogram estimation is performed. A second analysis using a sequential, spatial backfitting algorithm is used to investigate the importance of driving direction of the roller.

preprint2012arXiv

MMANOVA: A general multilevel framework for multivariate analysis of variance

Classical analysis of variance requires that model terms be labeled as fixed or random and typically culminate by comparing variability from each batch (factor) to variability from errors; without a standard methodology to assess the magnitude of a batch's variability, to compare variability between batches, nor to consider the uncertainty in this assessment. In this paper we support recent work, placing ANOVA into a general multilevel framework, then refine this through batch level model specifications, and develop it further by extension to the multivariate case. Adopting a Bayesian multilevel model parametrization, with improper batch level prior densities, we derive a method that facilitates comparison across all sources of variability. Whereas classical multivariate ANOVA often utilizes a single covariance criterion, e.g. determinant for Wilks' lambda distribution, the method allows arbitrary covariance criteria to be employed. The proposed method also addresses computation. By introducing implicit batch level constraints, which yield improper priors, the full posterior is efficiently factored, thus alleviating computational demands. For a large class of models, the partitioning mitigates, or even obviates the need for methods such as MCMC. The method is illustrated with simulated examples and an application focusing on climate projections with global climate models.

preprint2011arXiv

A spatial analysis of multivariate output from regional climate models

Climate models have become an important tool in the study of climate and climate change, and ensemble experiments consisting of multiple climate-model runs are used in studying and quantifying the uncertainty in climate-model output. However, there are often only a limited number of model runs available for a particular experiment, and one of the statistical challenges is to characterize the distribution of the model output. To that end, we have developed a multivariate hierarchical approach, at the heart of which is a new representation of a multivariate Markov random field. This approach allows for flexible modeling of the multivariate spatial dependencies, including the cross-dependencies between variables. We demonstrate this statistical model on an ensemble arising from a regional-climate-model experiment over the western United States, and we focus on the projected change in seasonal temperature and precipitation over the next 50 years.