Source author record

Alex Lenkoski

Alex Lenkoski appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

15works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

15 published item(s)

preprint2019arXiv

Rapid adjustment and post-processing of temperature forecast trajectories

Modern weather forecasts are commonly issued as consistent multi-day forecast trajectories with a time resolution of 1-3 hours. Prior to issuing, statistical post-processing is routinely used to correct systematic errors and misrepresentations of the forecast uncertainty. However, once the forecast has been issued, it is rarely updated before it is replaced in the next forecast cycle of the numerical weather prediction (NWP) model. This paper shows that the error correlation structure within the forecast trajectory can be utilized to substantially improve the forecast between the NWP forecast cycles by applying additional post-processing steps each time new observations become available. The proposed rapid adjustment is applied to temperature forecast trajectories from the UK Met Office's convective-scale ensemble MOGREPS-UK. MOGREPS-UK is run four times daily and produces hourly forecasts for up to 36 hours ahead. Our results indicate that the rapidly adjusted forecast from the previous NWP forecast cycle outperforms the new forecast for the first few hours of the next cycle, or until the new forecast itself can be rapidly adjusted, suggesting a new strategy for updating the forecast cycle.

preprint2016arXiv

Exact formulas for the normalizing constants of Wishart distributions for graphical models

Gaussian graphical models have received considerable attention during the past four decades from the statistical and machine learning communities. In Bayesian treatments of this model, the G-Wishart distribution serves as the conjugate prior for inverse covariance matrices satisfying graphical constraints. While it is straightforward to posit the unnormalized densities, the normalizing constants of these distributions have been known only for graphs that are chordal, or decomposable. Up until now, it was unknown whether the normalizing constant for a general graph could be represented explicitly, and a considerable body of computational literature emerged that attempted to avoid this apparent intractability. We close this question by providing an explicit representation of the G-Wishart normalizing constant for general graphs.

preprint2016arXiv

Spatially adaptive, Bayesian estimation for probabilistic temperature forecasts

Uncertainty in the prediction of future weather is commonly assessed through the use of forecast ensembles that employ a numerical weather prediction model in distinct variants. Statistical postprocessing can correct for biases in the numerical model and improves calibration. We propose a Bayesian version of the standard ensemble model output statistics (EMOS) postprocessing method, in which spatially varying bias coefficients are interpreted as realizations of Gaussian Markov random fields. Our Markovian EMOS (MEMOS) technique utilizes the recently developed stochastic partial differential equation (SPDE) and integrated nested Laplace approximation (INLA) methods for computationally efficient inference. The MEMOS approach shows good predictive performance in a comparative study of 24-hour ahead temperature forecasts over Germany based on the 50-member ensemble of the European Centre for Medium-Range Weather Forecasting (ECMWF).

preprint2014arXiv

Bayesian hierarchical modeling of extreme hourly precipitation in Norway

Spatial maps of extreme precipitation are a critical component of flood estimation in hydrological modeling, as well as in the planning and design of important infrastructure. This is particularly relevant in countries such as Norway that have a high density of hydrological power generating facilities and are exposed to significant risk of infrastructure damage due to flooding. In this work, we estimate a spatially coherent map of the distribution of extreme hourly precipitation in Norway, in terms of return levels, by linking generalized extreme value (GEV) distributions with latent Gaussian fields in a Bayesian hierarchical model. Generalized linear models on the parameters of the GEV distribution are able to incorporate location-specific geographic and meteorological information and thereby accommodate these effects on extreme precipitation. A Gaussian field on the GEV parameters captures additional unexplained spatial heterogeneity and overcomes the sparse grid on which observations are collected. We conduct an extensive analysis of the factors that affect the GEV parameters and show that our combination is able to appropriately characterize both the spatial variability of the distribution of extreme hourly precipitation in Norway, and the associated uncertainty in these estimates.

preprint2014arXiv

Efficient sampling of Gaussian graphical models using conditional Bayes factors

Bayesian estimation of Gaussian graphical models has proven to be challenging because the conjugate prior distribution on the Gaussian precision matrix, the G-Wishart distribution, has a doubly intractable partition function. Recent developments provide a direct way to sample from the G-Wishart distribution, which allows for more efficient algorithms for model selection than previously possible. Still, estimating Gaussian graphical models with more than a handful of variables remains a nearly infeasible task. Here, we propose two novel algorithms that use the direct sampler to more efficiently approximate the posterior distribution of the Gaussian graphical model. The first algorithm uses conditional Bayes factors to compare models in a Metropolis-Hastings framework. The second algorithm is based on a continuous time Markov process. We show that both algorithms are substantially faster than state-of-the-art alternatives. Finally, we show how the algorithms may be used to simultaneously estimate both structural and functional connectivity between subcortical brain regions using resting-state fMRI.

preprint2013arXiv

A Direct Sampler for G-Wishart Variates

The G-Wishart distribution is the conjugate prior for precision matrices that encode the conditional independencies of a Gaussian graphical model. While the distribution has received considerable attention, posterior inference has proven computationally challenging, in part due to the lack of a direct sampler. In this note, we rectify this situation. The existence of a direct sampler offers a host of new possibilities for the use of G-Wishart variates. We discuss one such development by outlining a new transdimensional model search algorithm--which we term double reversible jump--that leverages this sampler to avoid normalizing constant calculation when comparing graphical models. We conclude with two short studies meant to investigate our algorithm's validity.

preprint2013arXiv

Bayesian Motion Estimation for Dust Aerosols

Dust storms in the earth's major desert regions significantly influence microphysical weather processes, the CO$_2$-cycle and the global climate in general. Recent increases in the spatio-temporal resolution of remote sensing instruments have created new opportunities to understand these phenomena. However, the scale of the data collected and the inherent stochasticity of the underlying process pose significant challenges, requiring a careful combination of image processing and statistical techniques. In particular, using satellite imagery data, we develop a statistical model of atmospheric transport that relies on a latent Gaussian Markov random field (GMRF) for inference. In doing so, we make a link between the optical flow method of Horn and Schunck and the formulation of the transport process as a latent field in a generalized linear model, which enables the use of the integrated nested Laplace approximation for inference. This framework is specified such that it satisfies the so-called integrated continuity equation, thereby intrinsically expressing the divergence of the field as a multiplicative factor covering air compressibility and satellite column projection. The importance of this step -- as well as treating the problem in a fully statistical manner -- is emphasized by a simulation study where inference based on this latent GMRF clearly reduces errors of the estimated flow field. We conclude with a study of the dynamics of dust storms formed over Saharan Africa and show that our methodology is able to accurately and coherently track the storm movement, a critical problem in this field.

preprint2013arXiv

Shape from Texture using Locally Scaled Point Processes

Shape from texture refers to the extraction of 3D information from 2D images with irregular texture. This paper introduces a statistical framework to learn shape from texture where convex texture elements in a 2D image are represented through a point process. In a first step, the 2D image is preprocessed to generate a probability map corresponding to an estimate of the unnormalized intensity of the latent point process underlying the texture elements. The latent point process is subsequently inferred from the probability map in a non-parametric, model free manner. Finally, the 3D information is extracted from the point pattern by applying a locally scaled point process model where the local scaling function represents the deformation caused by the projection of a 3D surface onto a 2D image.

preprint2012arXiv

A Multivariate Graphical Stochastic Volatility Model

The Gaussian Graphical Model (GGM) is a popular tool for incorporating sparsity into joint multivariate distributions. The G-Wishart distribution, a conjugate prior for precision matrices satisfying general GGM constraints, has now been in existence for over a decade. However, due to the lack of a direct sampler, its use has been limited in hierarchical Bayesian contexts, relegating mixing over the class of GGMs mostly to situations involving standard Gaussian likelihoods. Recent work, however, has developed methods that couple model and parameter moves, first through reversible jump methods and later by direct evaluation of conditional Bayes factors and subsequent resampling. Further, methods for avoiding prior normalizing constant calculations--a serious bottleneck and source of numerical instability--have been proposed. We review and clarify these developments and then propose a new methodology for GGM comparison that blends many recent themes. Theoretical developments and computational timing experiments reveal an algorithm that has limited computational demands and dramatically improves on computing times of existing methods. We conclude by developing a parsimonious multivariate stochastic volatility model that embeds GGM uncertainty in a larger hierarchical framework. The method is shown to be capable of adapting to the extreme swings in market volatility experienced in 2008 after the collapse of Lehman Brothers, offering considerable improvement in posterior predictive distribution calibration.

preprint2012arXiv

Instrumental Variable Bayesian Model Averaging via Conditional Bayes Factors

We develop a method to perform model averaging in two-stage linear regression systems subject to endogeneity. Our method extends an existing Gibbs sampler for instrumental variables to incorporate a component of model uncertainty. Direct evaluation of model probabilities is intractable in this setting. We show that by nesting model moves inside the Gibbs sampler, model comparison can be performed via conditional Bayes factors, leading to straightforward calculations. This new Gibbs sampler is only slightly more involved than the original algorithm and exhibits no evidence of mixing difficulties. We conclude with a study of two different modeling challenges: incorporating uncertainty into the determinants of macroeconomic growth, and estimating a demand function by instrumenting wholesale on retail prices.

preprint2012arXiv

Multivariate probabilistic forecasting using Bayesian model averaging and copulas

We propose a method for post-processing an ensemble of multivariate forecasts in order to obtain a joint predictive distribution of weather. Our method utilizes existing univariate post-processing techniques, in this case ensemble Bayesian model averaging (BMA), to obtain estimated marginal distributions. However, implementing these methods individually offers no information regarding the joint distribution. To correct this, we propose the use of a Gaussian copula, which offers a simple procedure for recovering the dependence that is lost in the estimation of the ensemble BMA marginals. Our method is applied to 48-h forecasts of a set of five weather quantities using the 8-member University of Washington mesoscale ensemble. We show that our method recovers many well-understood dependencies between weather quantities and subsequently improves calibration and sharpness over both the raw ensemble and a method which does not incorporate joint distributional information.

preprint2012arXiv

Tobit Bayesian Model Averaging and the Determinants of Foreign Direct Investment

We develop a fully Bayesian, computationally efficient framework for incorporating model uncertainty into Type II Tobit models and apply this to the investigation of the determinants of Foreign Direct Investment (FDI). While direct evaluation of modelprobabilities is intractable in this setting, we show that by using conditional Bayes Factors, which nest model moves inside a Gibbs sampler, we are able to incorporate model uncertainty in a straight-forward fashion. We conclude with a study of global FDI flows between 1988-2000.

preprint2011arXiv

Copula Gaussian graphical models and their application to modeling functional disability data

We propose a comprehensive Bayesian approach for graphical model determination in observational studies that can accommodate binary, ordinal or continuous variables simultaneously. Our new models are called copula Gaussian graphical models (CGGMs) and embed graphical model selection inside a semiparametric Gaussian copula. The domain of applicability of our methods is very broad and encompasses many studies from social science and economics. We illustrate the use of the copula Gaussian graphical models in the analysis of a 16-dimensional functional disability contingency table.

preprint2010arXiv

Bayesian inference for general Gaussian graphical models with application to multivariate lattice data

We introduce efficient Markov chain Monte Carlo methods for inference and model determination in multivariate and matrix-variate Gaussian graphical models. Our framework is based on the G-Wishart prior for the precision matrix associated with graphs that can be decomposable or non-decomposable. We extend our sampling algorithms to a novel class of conditionally autoregressive models for sparse estimation in multivariate lattice data, with a special emphasis on the analysis of spatial data. These models embed a great deal of flexibility in estimating both the correlation structure across outcomes and the spatial correlation structure, thereby allowing for adaptive smoothing and spatial autocorrelation parameters. Our methods are illustrated using simulated and real-world examples, including an application to cancer mortality surveillance.

preprint2010arXiv

Sparse covariance estimation in heterogeneous samples

Standard Gaussian graphical models (GGMs) implicitly assume that the conditional independence among variables is common to all observations in the sample. However, in practice, observations are usually collected form heterogeneous populations where such assumption is not satisfied, leading in turn to nonlinear relationships among variables. To tackle these problems we explore mixtures of GGMs; in particular, we consider both infinite mixture models of GGMs and infinite hidden Markov models with GGM emission distributions. Such models allow us to divide a heterogeneous population into homogenous groups, with each cluster having its own conditional independence structure. The main advantage of considering infinite mixtures is that they allow us easily to estimate the number of number of subpopulations in the sample. As an illustration, we study the trends in exchange rate fluctuations in the pre-Euro era. This example demonstrates that the models are very flexible while providing extremely interesting interesting insights into real-life applications.