Source author record

Axel Gandy

Axel Gandy appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

16works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

16 published item(s)

preprint2022arXiv

Age-stratified epidemic model using a latent marked Hawkes process

We extend the unstructured homogeneously mixing epidemic model introduced by Lamprinakou et al. [arXiv:2208.07340] considering a finite population stratified by age bands. We model the actual unobserved infections using a latent marked Hawkes process and the reported aggregated infections as random quantities driven by the underlying Hawkes process. We apply a Kernel Density Particle Filter (KDPF) to infer the marked counting process, the instantaneous reproduction number for each age group and forecast the epidemic's future trajectory in the near future; considering the age bands and the population size does not increase the computational effort. We demonstrate the performance of the proposed inference algorithm on synthetic data sets and COVID-19 reported cases in various local authorities in the UK. We illustrate that taking into account the individual heterogeneity in age decreases the uncertainty of estimates and provides a real-time measurement of interventions and behavioural changes.

preprint2022arXiv

Can a latent Hawkes process be used for epidemiological modelling?

Understanding the spread of COVID-19 has been the subject of numerous studies, highlighting the significance of reliable epidemic models. Here, we introduce a novel epidemic model using a latent Hawkes process with temporal covariates for modelling the infections. Unlike other models, we model the reported cases via a probability distribution driven by the underlying Hawkes process. Modelling the infections via a Hawkes process allows us to estimate by whom an infected individual was infected. We propose a Kernel Density Particle Filter (KDPF) for inference of both latent cases and reproduction number and for predicting the new cases in the near future. The computational effort is proportional to the number of infections making it possible to use particle filter type algorithms, such as the KDPF. We demonstrate the performance of the proposed algorithm on synthetic data sets and COVID-19 reported cases in various local authorities in the UK, and benchmark our model to alternative approaches.

preprint2020arXiv

Estimating the number of infections and the impact of non-pharmaceutical interventions on COVID-19 in European countries: technical description update

Following the emergence of a novel coronavirus (SARS-CoV-2) and its spread outside of China, Europe has experienced large epidemics. In response, many European countries have implemented unprecedented non-pharmaceutical interventions including case isolation, the closure of schools and universities, banning of mass gatherings and/or public events, and most recently, wide-scale social distancing including local and national lockdowns. In this technical update, we extend a semi-mechanistic Bayesian hierarchical model that infers the impact of these interventions and estimates the number of infections over time. Our methods assume that changes in the reproductive number - a measure of transmission - are an immediate response to these interventions being implemented rather than broader gradual changes in behaviour. Our model estimates these changes by calculating backwards from temporal data on observed to estimate the number of infections and rate of transmission that occurred several weeks prior, allowing for a probabilistic time lag between infection and death. In this update we extend our original model [Flaxman, Mishra, Gandy et al 2020, Report #13, Imperial College London] to include (a) population saturation effects, (b) prior uncertainty on the infection fatality ratio, (c) a more balanced prior on intervention effects and (d) partial pooling of the lockdown intervention covariate. We also (e) included another 3 countries (Greece, the Netherlands and Portugal). The model code is available at https://github.com/ImperialCollegeLondon/covid19model/ We are now reporting the results of our updated model online at https://mrc-ide.github.io/covid19estimates/ We estimated parameters jointly for all M=14 countries in a single hierarchical model. Inference is performed in the probabilistic programming language Stan using an adaptive Hamiltonian Monte Carlo (HMC) sampler.

preprint2020arXiv

Importance subsampling for power system planning under multi-year demand and weather uncertainty

This paper introduces a generalised version of importance subsampling for time series reduction/aggregation in optimisation-based power system planning models. Recent studies indicate that reliably determining optimal electricity (investment) strategy under climate variability requires the consideration of multiple years of demand and weather data. However, solving planning models over long simulation lengths is typically computationally unfeasible, and established time series reduction approaches induce significant errors. The importance subsampling method reliably estimates long-term planning model outputs at greatly reduced computational cost, allowing the consideration of multi-decadal samples. The key innovation is a systematic identification and preservation of relevant extreme events in modeling subsamples. Simulation studies on generation and transmission expansion planning models illustrate the method's enhanced performance over established "representative days" clustering approaches. The models, data and sample code are made available as open-source software.

preprint2020arXiv

Inference of COVID-19 epidemiological distributions from Brazilian hospital data

Knowing COVID-19 epidemiological distributions, such as the time from patient admission to death, is directly relevant to effective primary and secondary care planning, and moreover, the mathematical modelling of the pandemic generally. We determine epidemiological distributions for patients hospitalised with COVID-19 using a large dataset ($N=21{,}000-157{,}000$) from the Brazilian Sistema de Informação de Vigilância Epidemiológica da Gripe database. A joint Bayesian subnational model with partial pooling is used to simultaneously describe the 26 states and one federal district of Brazil, and shows significant variation in the mean of the symptom-onset-to-death time, with ranges between 11.2-17.8 days across the different states, and a mean of 15.2 days for Brazil. We find strong evidence in favour of specific probability density function choices: for example, the gamma distribution gives the best fit for onset-to-death and the generalised log-normal for onset-to-hospital-admission. Our results show that epidemiological distributions have considerable geographical variation, and provide the first estimates of these distributions in a low and middle-income setting. At the subnational level, variation in COVID-19 outcome timings are found to be correlated with poverty, deprivation and segregation levels, and weaker correlation is observed for mean age, wealth and urbanicity.

preprint2020arXiv

On the derivation of the renewal equation from an age-dependent branching process: an epidemic modelling perspective

Renewal processes are a popular approach used in modelling infectious disease outbreaks. In a renewal process, previous infections give rise to future infections. However, while this formulation seems sensible, its application to infectious disease can be difficult to justify from first principles. It has been shown from the seminal work of Bellman and Harris that the renewal equation arises as the expectation of an age-dependent branching process. In this paper we provide a detailed derivation of the original Bellman Harris process. We introduce generalisations, that allow for time-varying reproduction numbers and the accounting of exogenous events, such as importations. We show how inference on the renewal equation is easy to accomplish within a Bayesian hierarchical framework. Using off the shelf MCMC packages, we fit to South Korea COVID-19 case data to estimate reproduction numbers and importations. Our derivation provides the mathematical fundamentals and assumptions underpinning the use of the renewal equation for modelling outbreaks.

preprint2020arXiv

Semi-Mechanistic Bayesian Modeling of COVID-19 with Renewal Processes

We propose a general Bayesian approach to modeling epidemics such as COVID-19. The approach grew out of specific analyses conducted during the pandemic, in particular an analysis concerning the effects of non-pharmaceutical interventions (NPIs) in reducing COVID-19 transmission in 11 European countries. The model parameterizes the time varying reproduction number $R_t$ through a regression framework in which covariates can e.g be governmental interventions or changes in mobility patterns. This allows a joint fit across regions and partial pooling to share strength. This innovation was critical to our timely estimates of the impact of lockdown and other NPIs in the European epidemics, whose validity was borne out by the subsequent course of the epidemic. Our framework provides a fully generative model for latent infections and observations deriving from them, including deaths, cases, hospitalizations, ICU admissions and seroprevalence surveys. One issue surrounding our model's use during the COVID-19 pandemic is the confounded nature of NPIs and mobility. We use our framework to explore this issue. We have open sourced an R package epidemia implementing our approach in Stan. Versions of the model are used by New York State, Tennessee and Scotland to estimate the current situation and make policy decisions.

preprint2016arXiv

The chopthin algorithm for resampling

Resampling is a standard step in particle filters and more generally sequential Monte Carlo methods. We present an algorithm, called chopthin, for resampling weighted particles. In contrast to standard resampling methods the algorithm does not produce a set of equally weighted particles; instead it merely enforces an upper bound on the ratio between the weights. Simulation studies show that the chopthin algorithm consistently outperforms standard resampling methods. The algorithms chops up particles with large weight and thins out particles with low weight, hence its name. It implicitly guarantees a lower bound on the effective sample size. The algorithm can be implemented efficiently, making it practically useful. We show that the expected computational effort is linear in the number of particles. Implementations for C++, R (on CRAN), Python and Matlab are available.

preprint2015arXiv

A Lévy-driven rainfall model with applications to futures pricing

We propose a parsimonious stochastic model for characterising the distributional and temporal properties of rainfall. The model is based on an integrated Ornstein-Uhlenbeck process driven by the Hougaard Lévy process. We derive properties of this process and propose an extended model which generalises the Ornstein-Uhlenbeck process to the class of continuous-time ARMA (CARMA) processes. The model is illustrated by fitting it to empirical rainfall data on both daily and hourly time scales. It is shown that the model is sufficiently flexible to capture important features of the rainfall process across locations and time scales. Finally we study an application to the pricing of rainfall derivatives which introduces the market price of risk via the Esscher transform. We first give a result specifying the risk-neutral expectation of a general moving average process. Then we illustrate the pricing method by calculating futures prices based on empirical daily rainfall data, where the rainfall process is specified by our model.

preprint2014arXiv

RMCMC: A System for Updating Bayesian Models

A system to update estimates from a sequence of probability distributions is presented. The aim of the system is to quickly produce estimates with a user-specified bound on the Monte Carlo error. The estimates are based upon weighted samples stored in a database. The stored samples are maintained such that the accuracy of the estimates and quality of the samples is satisfactory. This maintenance involves varying the number of samples in the database and updating their weights. New samples are generated, when required, by a Markov chain Monte Carlo algorithm. The system is demonstrated using a football league model that is used to predict the end of season table. Correctness of the estimates and their accuracy is shown in a simulation using a linear Gaussian model.

preprint2013arXiv

An algorithm to compute the power of Monte Carlo tests with guaranteed precision

This article presents an algorithm that generates a conservative confidence interval of a specified length and coverage probability for the power of a Monte Carlo test (such as a bootstrap or permutation test). It is the first method that achieves this aim for almost any Monte Carlo test. Previous research has focused on obtaining as accurate a result as possible for a fixed computational effort, without providing a guaranteed precision in the above sense. The algorithm we propose does not have a fixed effort and runs until a confidence interval with a user-specified length and coverage probability can be constructed. We show that the expected effort required by the algorithm is finite in most cases of practical interest, including situations where the distribution of the p-value is absolutely continuous or discrete with finite support. The algorithm is implemented in the R-package simctest, available on CRAN.

preprint2012arXiv

Non-Restarting CUSUM charts and Control of the False Discovery Rate

Cumulative sum (CUSUM) charts are typically used to detect changes in a stream of observations e.g. shifts in the mean. Usually, after signalling, the chart is restarted by setting it to some value below the signalling threshold. We propose a non-restarting CUSUM chart which is able to detect periods during which the stream is out of control. Further, we advocate an upper boundary to prevent the CUSUM chart rising too high, which helps detecting a change back into control. We present a novel algorithm to control the false discovery rate (FDR) pointwise in time when considering CUSUM charts based on multiple streams of data. We prove that the FDR is controlled under two definitions of a false discovery simultaneously. Simulations reveal the difference in FDR control when using these two definitions and other desirable definitions of a false discovery.

preprint2012arXiv

Optimality of Non-Restarting CUSUM charts

We show optimality, in a well-defined sense, using cumulative sum (CUSUM) charts for detecting changes in distributions. We consider a setting with multiple changes between two known distributions. This result advocates the use of non-restarting CUSUM charts with an upper boundary. Typically, after signalling, a CUSUM chart is restarted by setting it to some value below the threshold. A non-restarting CUSUM chart is not reset after signalling; thus is able to signal continuously. Imposing an upper boundary prevents the CUSUM chart rising too high, which facilitates detection in our setting. We discuss, via simulations, how the choice of the upper boundary changes the signals made by the non-restarting CUSUM charts.

preprint2011arXiv

Guaranteed Conditional Performance of Control Charts via Bootstrap Methods

To use control charts in practice, the in-control state usually has to be estimated. This estimation has a detrimental effect on the performance of control charts, which is often measured for example by the false alarm probability or the average run length. We suggest an adjustment of the monitoring schemes to overcome these problems. It guarantees, with a certain probability, a conditional performance given the estimated in-control state. The suggested method is based on bootstrapping the data used to estimate the in-control state. The method applies to different types of control charts, and also works with charts based on regression models, survival models, etc. If a nonparametric bootstrap is used, the method is robust to model errors. We show large sample properties of the adjustment. The usefulness of our approach is demonstrated through simulation studies.

preprint2010arXiv

A Bayesian approach to star-galaxy classification

Star-galaxy classification is one of the most fundamental data-processing tasks in survey astronomy, and a critical starting point for the scientific exploitation of survey data. For bright sources this classification can be done with almost complete reliability, but for the numerous sources close to a survey's detection limit each image encodes only limited morphological information. In this regime, from which many of the new scientific discoveries are likely to come, it is vital to utilise all the available information about a source, both from multiple measurements and also prior knowledge about the star and galaxy populations. It is also more useful and realistic to provide classification probabilities than decisive classifications. All these desiderata can be met by adopting a Bayesian approach to star-galaxy classification, and we develop a very general formalism for doing so. An immediate implication of applying Bayes's theorem to this problem is that it is formally impossible to combine morphological measurements in different bands without using colour information as well; however we develop several approximations that disregard colour information as much as possible. The resultant scheme is applied to data from the UKIRT Infrared Deep Sky Survey (UKIDSS), and tested by comparing the results to deep Sloan Digital Sky Survey (SDSS) Stripe 82 measurements of the same sources. The Bayesian classification probabilities obtained from the UKIDSS data agree well with the deep SDSS classifications both overall (a mismatch rate of 0.022, compared to 0.044 for the UKIDSS pipeline classifier) and close to the UKIDSS detection limit (a mismatch rate of 0.068 compared to 0.075 for the UKIDSS pipeline classifier). The Bayesian formalism developed here can be applied to improve the reliability of any star-galaxy classification schemes based on the measured values of morphology statistics alone.

preprint2008arXiv

Sequential Implementation of Monte Carlo Tests with Uniformly Bounded Resampling Risk

This paper introduces an open-ended sequential algorithm for computing the p-value of a test using Monte Carlo simulation. It guarantees that the resampling risk, the probability of a different decision than the one based on the theoretical p-value, is uniformly bounded by an arbitrarily small constant. Previously suggested sequential or non-sequential algorithms, using a bounded sample size, do not have this property. Although the algorithm is open-ended, the expected number of steps is finite, except when the p-value is on the threshold between rejecting and not rejecting. The algorithm is suitable as standard for implementing tests that require (re-)sampling. It can also be used in other situations: to check whether a test is conservative, iteratively to implement double bootstrap tests, and to determine the sample size required for a certain power.