Source author record

Tilmann Gneiting

Tilmann Gneiting appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

22works
13topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

22 published item(s)

preprint2024arXiv

Physics-based vs. data-driven 24-hour probabilistic forecasts of precipitation for northern tropical Africa

Numerical weather prediction (NWP) models struggle to skillfully predict tropical precipitation occurrence and amount, calling for alternative approaches. For instance, it has been shown that fairly simple, purely data-driven logistic regression models for 24-hour precipitation occurrence outperform both climatological and NWP forecasts for the West African summer monsoon. More complex neural network based approaches, however, remain underdeveloped due to the non-Gaussian character of precipitation. In this study, we develop, apply and evaluate a novel two-stage approach, where we train a U-Net convolutional neural network (CNN) model on gridded rainfall data to obtain a deterministic forecast and then apply the recently developed, nonparametric Easy Uncertainty Quantification (EasyUQ) approach to convert it into a probabilistic forecast. We evaluate CNN+EasyUQ for one-day ahead 24-hour accumulated precipitation forecasts over northern tropical Africa for 2011--2019, with the Integrated Multi-satellitE Retrievals for GPM (IMERG) data serving as ground truth. In the most comprehensive assessment to date we compare CNN+EasyUQ to state-of-the-art physics-based and data-driven approaches such as a monthly probabilistic climatology, raw and postprocessed ensemble forecasts from the European Centre for Medium-Range Weather Forecasts (ECMWF), and traditional statistical approaches that use up to 25 predictor variables from IMERG and the ERA5 reanalysis.Generally, statistical approaches perform about en par with post-processed ECMWF ensemble forecasts. The CNN+EasyUQ approach, however, clearly outperforms all competitors for both occurrence and amount. Hybrid methods that merge CNN+EasyUQ and physics-based forecasts show slight further improvement. Thus, the CNN+EasyUQ approach can likely improve operational probabilistic forecasts of rainfall in the tropics, and potentially even beyond.

preprint2020arXiv

Predictive Inference Based on Markov Chain Monte Carlo Output

In Bayesian inference, predictive distributions are typically in the form of samples generated via Markov chain Monte Carlo (MCMC) or related algorithms. In this paper, we conduct a systematic analysis of how to make and evaluate probabilistic forecasts from such simulation output. Based on proper scoring rules, we develop a notion of consistency that allows to assess the adequacy of methods for estimating the stationary distribution underlying the simulation output. We then provide asymptotic results that account for the salient features of Bayesian posterior simulators, and derive conditions under which choices from the literature satisfy our notion of consistency. Importantly, these conditions depend on the scoring rule being used, such that the choices of approximation method and scoring rule are intertwined. While the logarithmic rule requires fairly stringent conditions, the continuous ranked probability score (CRPS) yields consistent approximations under minimal assumptions. These results are illustrated in a simulation study and an economic data example. Overall, mixture-of-parameters approximations which exploit the parametric structure of Bayesian models perform particularly well. Under the CRPS, the empirical distribution function is a simple and appealing alternative option.

preprint2018arXiv

Properization: Constructing Proper Scoring Rules via Bayes Acts

Scoring rules serve to quantify predictive performance. A scoring rule is proper if truth telling is an optimal strategy in expectation. Subject to customary regularity conditions, every scoring rule can be made proper, by applying a special case of the Bayes act construction studied by Grünwald and Dawid (2004) and Dawid (2007), to which we refer as properization. We discuss examples from the recent literature and apply the construction to create new types, and reinterpret existing forms, of proper scoring rules and consistent scoring functions. In an abstract setting, we formulate sufficient conditions under which Bayes acts exist and scoring rules can be made proper.

preprint2016arXiv

Spatially adaptive, Bayesian estimation for probabilistic temperature forecasts

Uncertainty in the prediction of future weather is commonly assessed through the use of forecast ensembles that employ a numerical weather prediction model in distinct variants. Statistical postprocessing can correct for biases in the numerical model and improves calibration. We propose a Bayesian version of the standard ensemble model output statistics (EMOS) postprocessing method, in which spatially varying bias coefficients are interpreted as realizations of Gaussian Markov random fields. Our Markovian EMOS (MEMOS) technique utilizes the recently developed stochastic partial differential equation (SPDE) and integrated nested Laplace approximation (INLA) methods for computationally efficient inference. The MEMOS approach shows good predictive performance in a comparative study of 24-hour ahead temperature forecasts over Germany based on the 50-member ensemble of the European Centre for Medium-Range Weather Forecasting (ECMWF).

preprint2015arXiv

Forecaster's Dilemma: Extreme Events and Forecast Evaluation

In public discussions of the quality of forecasts, attention typically focuses on the predictive performance in cases of extreme events. However, the restriction of conventional forecast evaluation methods to subsets of extreme observations has unexpected and undesired effects, and is bound to discredit skillful forecasts when the signal-to-noise ratio in the data generating process is low. Conditioning on outcomes is incompatible with the theoretical assumptions of established forecast evaluation methods, thereby confronting forecasters with what we refer to as the forecaster's dilemma. For probabilistic forecasts, proper weighted scoring rules have been proposed as decision theoretically justifiable alternatives for forecast evaluation with an emphasis on extreme events. Using theoretical arguments, simulation experiments, and a real data study on probabilistic forecasts of U.S. inflation and gross domestic product growth, we illustrate and discuss the forecaster's dilemma along with potential remedies.

preprint2015arXiv

Gaussian Random Particles with Flexible Hausdorff Dimension

Gaussian particles provide a flexible framework for modelling and simulating three-dimensional star-shaped random sets. In our framework, the radial function of the particle arises from a kernel smoothing, and is associated with an isotropic random field on the sphere. If the kernel is a von Mises--Fisher density, or uniform on a spherical cap, the correlation function of the associated random field admits a closed form expression. The Hausdorff dimension of the surface of the Gaussian particle reflects the decay of the correlation function at the origin, as quantified by the fractal index. Under power kernels we obtain particles with boundaries of any Hausdorff dimension between 2 and 3.

preprint2015arXiv

Of Quantiles and Expectiles: Consistent Scoring Functions, Choquet Representations, and Forecast Rankings

In the practice of point prediction, it is desirable that forecasters receive a directive in the form of a statistical functional, such as the mean or a quantile of the predictive distribution. When evaluating and comparing competing forecasts, it is then critical that the scoring function used for these purposes be consistent for the functional at hand, in the sense that the expected score is minimized when following the directive. We show that any scoring function that is consistent for a quantile or an expectile functional, respectively, can be represented as a mixture of extremal scoring functions that form a linearly parameterized family. Scoring functions for the mean value and probability forecasts of binary events constitute important examples. The quantile and expectile functionals along with the respective extremal scoring functions admit appealing economic interpretations in terms of thresholds in decision making. The Choquet type mixture representations give rise to simple checks of whether a forecast dominates another in the sense that it is preferable under any consistent scoring function. In empirical settings it suffices to compare the average scores for only a finite number of extremal elements. Plots of the average scores with respect to the extremal scoring functions, which we call Murphy diagrams, permit detailed comparisons of the relative merits of competing forecasts.

preprint2014arXiv

The affinely invariant distance correlation

Székely, Rizzo and Bakirov (Ann. Statist. 35 (2007) 2769-2794) and Székely and Rizzo (Ann. Appl. Statist. 3 (2009) 1236-1265), in two seminal papers, introduced the powerful concept of distance correlation as a measure of dependence between sets of random variables. We study in this paper an affinely invariant version of the distance correlation and an empirical version of that distance correlation, and we establish the consistency of the empirical quantity. In the case of subvectors of a multivariate normally distributed random vector, we provide exact expressions for the affinely invariant distance correlation in both finite-dimensional and asymptotic settings, and in the finite-dimensional case we find that the affinely invariant distance correlation is a function of the canonical correlation coefficients. To illustrate our results, we consider time series of wind vectors at the Stateline wind energy center in Oregon and Washington, and we derive the empirical auto and cross distance correlation functions between wind vectors at distinct meteorological stations.

preprint2013arXiv

Copula Calibration

We propose notions of calibration for probabilistic forecasts of general multivariate quantities. Probabilistic copula calibration is a natural analogue of probabilistic calibration in the univariate setting. It can be assessed empirically by checking for the uniformity of the copula probability integral transform (CopPIT), which is invariant under coordinate permutations and coordinatewise strictly monotone transformations of the predictive distribution and the outcome. The CopPIT histogram can be interpreted as a generalization and variant of the multivariate rank histogram, which has been used to check the calibration of ensemble forecasts. Climatological copula calibration is an analogue of marginal calibration in the univariate setting. Methods and tools are illustrated in a simulation study and applied to compare raw numerical model and statistically postprocessed ensemble forecasts of bivariate wind vectors.

preprint2013arXiv

Section on the special year for mathematics of planet earth (MPE 2013)

Dozens of research centers, foundations, international organizations and scientific societies, including the Institute of Mathematical Statistics, have joined forces to celebrate 2013 as a special year for the Mathematics of Planet Earth. In its five-year history, the Annals of Applied Statistics has been publishing cutting edge research in this area, including geophysical, biological and socio-economic aspects of planet Earth, with the special section on statistics in the atmospheric sciences edited by Fuentes, Guttorp and Stein (2008) and the discussion paper by McShane and Wyner (2011) on paleoclimate reconstructions [Stein (2011)] having been highlights.

preprint2013arXiv

Strictly and non-strictly positive definite functions on spheres

Isotropic positive definite functions on spheres play important roles in spatial statistics, where they occur as the correlation functions of homogeneous random fields and star-shaped random particles. In approximation theory, strictly positive definite functions serve as radial basis functions for interpolating scattered data on spherical domains. We review characterizations of positive definite functions on spheres in terms of Gegenbauer expansions and apply them to dimension walks, where monotonicity properties of the Gegenbauer coefficients guarantee positive definiteness in higher dimensions. Subject to a natural support condition, isotropic positive definite functions on the Euclidean space $\mathbb{R}^3$, such as Askey's and Wendland's functions, allow for the direct substitution of the Euclidean distance by the great circle distance on a one-, two- or three-dimensional sphere, as opposed to the traditional approach, where the distances are transformed into each other. Completely monotone functions are positive definite on spheres of any dimension and provide rich parametric classes of such functions, including members of the powered exponential, Matérn, generalized Cauchy and Dagum families. The sine power family permits a continuous parameterization of the roughness of the sample paths of a Gaussian process. A collection of research problems provides challenges for future work in mathematical analysis, probability theory and spatial statistics.

preprint2013arXiv

Uncertainty Quantification in Complex Simulation Models Using Ensemble Copula Coupling

Critical decisions frequently rely on high-dimensional output from complex computer simulation models that show intricate cross-variable, spatial and temporal dependence structures, with weather and climate predictions being key examples. There is a strongly increasing recognition of the need for uncertainty quantification in such settings, for which we propose and review a general multi-stage procedure called ensemble copula coupling (ECC), proceeding as follows: 1. Generate a raw ensemble, consisting of multiple runs of the computer model that differ in the inputs or model parameters in suitable ways. 2. Apply statistical postprocessing techniques, such as Bayesian model averaging or nonhomogeneous regression, to correct for systematic errors in the raw ensemble, to obtain calibrated and sharp predictive distributions for each univariate output variable individually. 3. Draw a sample from each postprocessed predictive distribution. 4. Rearrange the sampled values in the rank order structure of the raw ensemble to obtain the ECC postprocessed ensemble. The use of ensembles and statistical postprocessing have become routine in weather forecasting over the past decade. We show that seemingly unrelated, recent advances can be interpreted, fused and consolidated within the framework of ECC, the common thread being the adoption of the empirical copula of the raw ensemble. Depending on the use of Quantiles, Random draws or Transformations at the sampling stage, we distinguish the ECC-Q, ECC-R and ECC-T variants, respectively. We also describe relations to the Schaake shuffle and extant copula-based techniques. In a case study, the ECC approach is applied to predictions of temperature, pressure, precipitation and wind over Germany, based on the 50-member European Centre for Medium-Range Weather Forecasts (ECMWF) ensemble.

preprint2013arXiv

Using proper divergence functions to evaluate climate models

It has been argued persuasively that, in order to evaluate climate models, the probability distributions of model output need to be compared to the corresponding empirical distributions of observed data. Distance measures between probability distributions, also called divergence functions, can be used for this purpose. We contend that divergence functions ought to be proper, in the sense that acting on modelers' true beliefs is an optimal strategy. Score divergences that derive from proper scoring rules are proper, with the integrated quadratic distance and the Kullback-Leibler divergence being particularly attractive choices. Other commonly used divergences fail to be proper. In an illustration, we evaluate and rank simulations from fifteen climate models for temperature extremes in a comparison to re-analysis data.

preprint2012arXiv

Ensemble model output statistics for wind vectors

A bivariate ensemble model output statistics (EMOS) technique for the postprocessing of ensemble forecasts of two-dimensional wind vectors is proposed, where the postprocessed probabilistic forecast takes the form of a bivariate normal probability density function. The postprocessed means and variances of the wind vector components are linearly bias-corrected versions of the ensemble means and ensemble variances, respectively, and the conditional correlation between the wind components is represented by a trigonometric function of the ensemble mean wind direction. In a case study on 48-hour forecasts of wind vectors over the North American Pacific Northwest with the University of Washington Mesoscale Ensemble, the bivariate EMOS density forecasts were calibrated and sharp, and showed considerable improvement over the raw ensemble and reference forecasts, including ensemble copula coupling.

preprint2012arXiv

Estimators of Fractal Dimension: Assessing the Roughness of Time Series and Spatial Data

The fractal or Hausdorff dimension is a measure of roughness (or smoothness) for time series and spatial data. The graph of a smooth, differentiable surface indexed in $\mathbb{R}^d$ has topological and fractal dimension $d$. If the surface is nondifferentiable and rough, the fractal dimension takes values between the topological dimension, $d$, and $d+1$. We review and assess estimators of fractal dimension by their large sample behavior under infill asymptotics, in extensive finite sample simulation studies, and in a data example on arctic sea-ice profiles. For time series or line transect data, box-count, Hall--Wood, semi-periodogram, discrete cosine transform and wavelet estimators are studied along with variation estimators with power indices 2 (variogram) and 1 (madogram), all implemented in the R package fractaldim. Considering both efficiency and robustness, we recommend the use of the madogram estimator, which can be interpreted as a statistically more efficient version of the Hall--Wood estimator. For two-dimensional lattice data, we propose robust transect estimators that use the median of variation estimates along rows and columns. Generally, the link between power variations of index $p>0$ for stochastic processes, and the Hausdorff dimension of their sample paths, appears to be particularly robust and inclusive when $p=1$.

preprint2012arXiv

Local proper scoring rules of order two

Scoring rules assess the quality of probabilistic forecasts, by assigning a numerical score based on the predictive distribution and on the event or value that materializes. A scoring rule is proper if it encourages truthful reporting. It is local of order $k$ if the score depends on the predictive density only through its value and the values of its derivatives of order up to $k$ at the realizing event. Complementing fundamental recent work by Parry, Dawid and Lauritzen, we characterize the local proper scoring rules of order 2 relative to a broad class of Lebesgue densities on the real line, using a different approach. In a data example, we use local and nonlocal proper scoring rules to assess statistically postprocessed ensemble weather forecasts.

preprint2012arXiv

On the Cover-Hart Inequality: What's a Sample of Size One Worth?

Bob predicts a future observation based on a sample of size one. Alice can draw a sample of any size before issuing her prediction. How much better can she do than Bob? Perhaps surprisingly, under a large class of loss functions, which we refer to as the Cover-Hart family, the best Alice can do is to halve Bob's risk. In this sense, half the information in an infinite sample is contained in a sample of size one. The Cover-Hart family is a convex cone that includes metrics and negative definite functions, subject to slight regularity conditions. These results may help explain the small relative differences in empirical performance measures in applied classification and forecasting problems, as well as the success of reasoning and learning by analogy in general, and nearest neighbor techniques in particular.

preprint2011arXiv

Combining Predictive Distributions

Predictive distributions need to be aggregated when probabilistic forecasts are merged, or when expert opinions expressed in terms of probability distributions are fused. We take a prediction space approach that applies to discrete, mixed discrete-continuous and continuous predictive distributions alike, and study combination formulas for cumulative distribution functions from the perspectives of coherence, probabilistic and conditional calibration, and dispersion. Both linear and non-linear aggregation methods are investigated, including generalized, spread-adjusted and beta-transformed linear pools. The effects and techniques are demonstrated theoretically, in simulation examples, and in case studies on density forecasts for S&P 500 returns and daily maximum temperature at Seattle-Tacoma Airport.

preprint2010arXiv

Making and Evaluating Point Forecasts

Typically, point forecasting methods are compared and assessed by means of an error measure or scoring function, such as the absolute error or the squared error. The individual scores are then averaged over forecast cases, to result in a summary measure of the predictive performance, such as the mean absolute error or the (root) mean squared error. I demonstrate that this common practice can lead to grossly misguided inferences, unless the scoring function and the forecasting task are carefully matched. Effective point forecasting requires that the scoring function be specified ex ante, or that the forecaster receives a directive in the form of a statistical functional, such as the mean or a quantile of the predictive distribution. If the scoring function is specified ex ante, the forecaster can issue the optimal point forecast, namely, the Bayes rule. If the forecaster receives a directive in the form of a functional, it is critical that the scoring function be consistent for it, in the sense that the expected score is minimized when following the directive. A functional is elicitable if there exists a scoring function that is strictly consistent for it. Expectations, ratios of expectations and quantiles are elicitable. For example, a scoring function is consistent for the mean functional if and only if it is a Bregman function. It is consistent for a quantile if and only if it is generalized piecewise linear. Similar characterizations apply to ratios of expectations and to expectiles. Weighted scoring functions are consistent for functionals that adapt to the weighting in peculiar ways. Not all functionals are elicitable; for instance, conditional value-at-risk is not, despite its popularity in quantitative finance.

preprint2010arXiv

Predicting Inflation: Professional Experts Versus No-Change Forecasts

We compare forecasts of United States inflation from the Survey of Professional Forecasters (SPF) to predictions made by simple statistical techniques. In nowcasting, economic expertise is persuasive. When projecting beyond the current quarter, novel yet simplistic probabilistic no-change forecasts are equally competitive. We further interpret surveys as ensembles of forecasts, and show that they can be used similarly to the ways in which ensemble prediction systems have transformed weather forecasting. Then we borrow another idea from weather forecasting, in that we apply statistical techniques to postprocess the SPF forecast, based on experience from the recent past. The foregoing conclusions remain unchanged after survey postprocessing.

preprint2001arXiv

Stochastic models which separate fractal dimension and Hurst effect

Fractal behavior and long-range dependence have been observed in an astonishing number of physical systems. Either phenomenon has been modeled by self-similar random functions, thereby implying a linear relationship between fractal dimension, a measure of roughness, and Hurst coefficient, a measure of long-memory dependence. This letter introduces simple stochastic models which allow for any combination of fractal dimension and Hurst exponent. We synthesize images from these models, with arbitrary fractal properties and power-law correlations, and propose a test for self-similarity.