Source author record

Ewan Cameron

Ewan Cameron appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

17works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

17 published item(s)

preprint2020arXiv

A simulation study of disaggregation regression for spatial disease mapping

Disaggregation regression has become an important tool in spatial disease mapping for making fine-scale predictions of disease risk from aggregated response data. By including high resolution covariate information and modelling the data generating process on a fine scale, it is hoped that these models can accurately learn the relationships between covariates and response at a fine spatial scale. However, validating these high resolution predictions can be a challenge, as often there is no data observed at this spatial scale. In this study, disaggregation regression was performed on simulated data in various settings and the resulting fine-scale predictions are compared to the simulated ground truth. Performance was investigated with varying numbers of data points, sizes of aggregated areas and levels of model misspecification. The effectiveness of cross validation on the aggregate level as a measure of fine-scale predictive performance was also investigated. Predictive performance improved as the number of observations increased and as the size of the aggregated areas decreased. When the model was well-specified, fine-scale predictions were accurate even with small numbers of observations and large aggregated areas. Under model misspecification predictive performance was significantly worse for large aggregated areas but remained high when response data was aggregated over smaller regions. Cross-validation correlation on the aggregate level was a moderately good predictor of fine-scale predictive performance. While the simulations are unlikely to capture the nuances of real-life response data, this study gives insight into the effectiveness of disaggregation regression in different contexts.

preprint2020arXiv

Spatiotemporal mapping of malaria prevalence in Madagascar using routine surveillance and health survey data

Malaria transmission in Madagascar is highly heterogeneous, exhibiting spatial, seasonal and long-term trends. Previous efforts to map malaria risk in Madagascar used prevalence data from Malaria Indicator Surveys. These cross-sectional surveys, conducted during the high transmission season most recently in 2013 and 2016, provide nationally representative prevalence data but cover relatively short time frames. Conversely, monthly case data are collected at health facilities but suffer from biases, including incomplete reporting. We combined survey and case data to make monthly maps of prevalence between 2013 and 2016. Health facility catchments were estimated and incidence surfaces, environmental and socioeconomic covariates, and survey data informed a Bayesian prevalence model. Prevalence estimates were consistently high in the coastal regions and low in the highlands. Prevalence was lowest in 2014 and peaked in 2015, highlighting the importance of estimates between survey years. Seasonality was widely observed. Similar multi-metric approaches may be applicable across sub-Saharan Africa.

preprint2019arXiv

Mapping malaria seasonality: a case study from Madagascar

Many malaria-endemic areas experience seasonal fluctuations in case incidence as Anopheles mosquito and Plasmodium parasite life cycles respond to changing environmental conditions. While most existing maps of malaria seasonality use fixed thresholds of rainfall, temperature, and/or vegetation indices to identify suitable transmission months, we develop a statistical modelling framework for characterising the seasonal patterns derived directly from case data. The procedure involves a spatiotemporal regression model for estimating the monthly proportions of total annual cases and an algorithm to identify operationally relevant characteristics such as the transmission start and peak months. A seasonality index combines the monthly proportion estimates and existing estimates of annual case incidence to provide a summary of "how seasonal" locations are relative to their surroundings. An advancement upon past seasonality mapping endeavours is the presentation of the uncertainty associated with each map, which will enable policymakers to make more statistically sound decisions. The methodology is illustrated using health facility data from Madagascar.

preprint2015arXiv

The star cluster mass--galactocentric radius relation: Implications for cluster formation

Whether or not the initial star cluster mass function is established through a universal, galactocentric-distance-independent stochastic process, on the scales of individual galaxies, remains an unsolved problem. This debate has recently gained new impetus through the publication of a study that concluded that the maximum cluster mass in a given population is not solely determined by size-of-sample effects. Here, we revisit the evidence in favor and against stochastic cluster formation by examining the young ($\lesssim$ a few $\times 10^8$ yr-old) star cluster mass--galactocentric radius relation in M33, M51, M83, and the Large Magellanic Cloud. To eliminate size-of-sample effects, we first adopt radial bin sizes containing constant numbers of clusters, which we use to quantify the radial distribution of the first- to fifth-ranked most massive clusters using ordinary least-squares fitting. We supplement this analysis with an application of quantile regression, a binless approach to rank-based regression taking an absolute-value-distance penalty. Both methods yield, within the $1σ$ to $3σ$ uncertainties, near-zero slopes in the diagnostic plane, largely irrespective of the maximum age or minimum mass imposed on our sample selection, or of the radial bin size adopted. We conclude that, at least in our four well-studied sample galaxies, star cluster formation does not necessarily require an environment-dependent cluster formation scenario, which thus supports the notion of stochastic star cluster formation as the dominant star cluster-formation process within a given galaxy.

preprint2014arXiv

Recursive Pathways to Marginal Likelihood Estimation with Prior-Sensitivity Analysis

We investigate the utility to computational Bayesian analyses of a particular family of recursive marginal likelihood estimators characterized by the (equivalent) algorithms known as "biased sampling" or "reverse logistic regression" in the statistics literature and "the density of states" in physics. Through a pair of numerical examples (including mixture modeling of the well-known galaxy data set) we highlight the remarkable diversity of sampling schemes amenable to such recursive normalization, as well as the notable efficiency of the resulting pseudo-mixture distributions for gauging prior sensitivity in the Bayesian model selection context. Our key theoretical contributions are to introduce a novel heuristic ("thermodynamic integration via importance sampling") for qualifying the role of the bridging sequence in this procedure and to reveal various connections between these recursive estimators and the nested sampling technique.

preprint2014arXiv

What we talk about when we talk about fields

In astronomical and cosmological studies one often wishes to infer some properties of an infinite-dimensional field indexed within a finite-dimensional metric space given only a finite collection of noisy observational data. Bayesian inference offers an increasingly-popular strategy to overcome the inherent ill-posedness of this signal reconstruction challenge. However, there remains a great deal of confusion within the astronomical community regarding the appropriate mathematical devices for framing such analyses and the diversity of available computational procedures for recovering posterior functionals. In this brief research note I will attempt to clarify both these issues from an "applied statistics" perpective, with insights garnered from my post-astronomy experiences as a computational Bayesian / epidemiological geostatistician.

preprint2013arXiv

A Generalized Savage-Dickey Ratio

In this brief research note I present a generalized version of the Savage-Dickey Density Ratio for representation of the Bayes factor (or marginal likelihood ratio) of nested statistical models; the new version takes the form of a Radon-Nikodym derivative and is thus applicable to a wider family of probability spaces than the original (restricted to those admitting an ordinary Lebesgue density). A derivation is given following the measure-theoretic construction of Marin & Robert (2010), and the equivalent estimator is demonstrated in application to a distributional modeling problem.

preprint2013arXiv

On the Evidence for Cosmic Variation of the Fine Structure Constant (I): A Parametric Bayesian Model Selection Analysis of the Quasar Dataset

We review the evidence behind recent claims of spatial variation in the fine structure constant deriving from observations of ionic absorption lines in the light from distant quasars. To this end we expand upon previous non-Bayesian analyses limited by the assumptions of an unbiased and strictly Normal distribution for the "unexplained errors" of the benchmark quasar dataset. Through the technique of reverse logistic regression we estimate and compare marginal likelihoods for three competing hypotheses---(i) the null hypothesis (no cosmic variation), (ii) the monopole hypothesis (a constant Earth-to-quasar offset), and (iii) the monopole+dipole hypothesis (a cosmic variation manifest to the Earth-bound observer as a North-South divergence)---under a variety of candidate parametric forms for the unexplained error term. Our analysis reveals weak support for a skeptical interpretation in which the apparent dipole effect is driven solely by systematic errors of opposing sign inherent in measurements from the two telescopes employed to obtain these observations. Throughout we seek to exemplify a 'best practice' approach to Bayesian model selection with prior-sensitivity analysis; in a companion paper we extend this methodology to a semi-parametric framework using the infinite-dimensional Dirichlet process.

preprint2013arXiv

On the Evidence for Cosmic Variation of the Fine Structure Constant (II): A Semi-Parametric Bayesian Model Selection Analysis of the Quasar Dataset

In the second paper of this series we extend our Bayesian reanalysis of the evidence for a cosmic variation of the fine structure constant to the semi-parametric modelling regime. By adopting a mixture of Dirichlet processes prior for the unexplained errors in each instrumental subgroup of the benchmark quasar dataset we go some way towards freeing our model selection procedure from the apparent subjectivity of a fixed distributional form. Despite the infinite-dimensional domain of the error hierarchy so constructed we are able to demonstrate a recursive scheme for marginal likelihood estimation with prior-sensitivity analysis directly analogous to that presented in Paper I, thereby allowing the robustness of our posterior Bayes factors to hyper-parameter choice and model specification to be readily verified. In the course of this work we elucidate various similarities between unexplained error problems in the seemingly disparate fields of astronomy and clinical meta-analysis, and we highlight a number of sophisticated techniques for handling such problems made available by past research in the latter. It is our hope that the novel approach to semi-parametric model selection demonstrated herein may serve as a useful reference for others exploring this potentially difficult class of error model.

preprint2011arXiv

Galaxy And Mass Assembly: Stellar Mass Estimates

This paper describes the first catalogue of photometrically-derived stellar mass estimates for intermediate-redshift (z < 0.65) galaxies in the Galaxy And Mass Assembly (GAMA) spectroscopic redshift survey. These masses, as well as the full set of ancillary stellar population parameters, will be made public as part of GAMA data release 2. Although the GAMA database does include NIR photometry, we show that the quality of our stellar population synthesis fits is significantly poorer when these NIR data are included. Further, for a large fraction of galaxies, the stellar population parameters inferred from the optical-plus-NIR photometry are formally inconsistent with those inferred from the optical data alone. This may indicate problems in our stellar population library, or NIR data issues, or both; these issues will be addressed for future versions of the catalogue. For now, we have chosen to base our stellar mass estimates on optical photometry only. In light of our decision to ignore the available NIR data, we examine how well stellar mass can be constrained based on optical data alone. We use generic properties of stellar population synthesis models to demonstrate that restframe colour alone is in principle a very good estimator of stellar mass-to-light ratio, M*/Li. Further, we use the observed relation between restframe (g-i) and M*/Li for real GAMA galaxies to argue that, modulo uncertainties in the stellar evolution models themselves, (g-i) colour can in practice be used to estimate M*/Li to an accuracy of < ~0.1 dex. This 'empirically calibrated' (g-i)-M*/Li relation offers a simple and transparent means for estimating galaxies' stellar masses based on minimal data, and so provides a solid basis for other surveys to compare their results to z < ~0.4 measurements from GAMA.

preprint2011arXiv

On the Estimation of Confidence Intervals for Binomial Population Proportions in Astronomy: The Simplicity and Superiority of the Bayesian Approach

I present a critical review of techniques for estimating confidence intervals on binomial population proportions inferred from success counts in small-to-intermediate samples. Population proportions arise frequently as quantities of interest in astronomical research; for instance, in studies aiming to constrain the bar fraction, AGN fraction, SMBH fraction, merger fraction, or red sequence fraction from counts of galaxies exhibiting distinct morphological features or stellar populations. However, two of the most widely-used techniques for estimating binomial confidence intervals--the 'normal approximation' and the Clopper & Pearson approach--are liable to misrepresent the degree of statistical uncertainty present under sampling conditions routinely encountered in astronomical surveys, leading to an ineffective use of the experimental data (and, worse, an inefficient use of the resources expended in obtaining that data). Hence, I provide here an overview of the fundamentals of binomial statistics with two principal aims: (i) to reveal the ease with which (Bayesian) binomial confidence intervals with more satisfactory behaviour may be estimated from the quantiles of the beta distribution using modern mathematical software packages (e.g. R, matlab, mathematica, IDL, python); and (ii) to demonstrate convincingly the major flaws of both the 'normal approximation' and the Clopper & Pearson approach for error estimation.

preprint2011arXiv

The near-IR $M_{bh}$ - L and $M_{bh}$ - n relations

We present near-IR surface photometry (2D-profiling) for a sample of 29 nearby galaxies for which super-massive black hole (SMBH) masses are constrained. The data is derived from the UKIDSS-LASS survey representing a significant improvement in image quality and depth over previous studies based on 2MASS data. We derive the spheroid luminosity and spheroid Sérsic index for each galaxy with GALFIT3 and use these data to construct SMBH mass -bulge luminosity ($M_{\rm bh}$--$L$) and SMBH - Sérsic index ($M_{\rm bh}$--$n$) relations. The best fit K-band relation for elliptical and disk galaxies is $\log(M_{\rm bh}/M_{\odot})= -0.36(\pm 0.03) (M_{\rm K} + 18) + 6.17(\pm 0.16)$ with an intrinsic scatter of 0.4$^{+0.09}_{-0.06}$dex whilst for elliptical galaxies we find $\log(M_{\rm bh}/M_{\odot})= -0.42(\pm 0.06) (M_{\rm K} + 22) + 7.5(\pm 0.15)$ with an intrinsic scatter of 0.31$^{+0.087}_{-0.047}$dex. Our revised $M_{\rm bh}$--$L$ relation agrees closely with the previous near-IR constraint by \citet{tex:G07}. The lack of improvement in the intrinsic scatter in moving to higher quality near-IR data suggests that the SMBH relations are not currently limited by the quality of the imaging data but is either intrinsic or a result of uncertainty in the precise number of required components required in the profiling process. Contrary to expectation (see \citealt{tex:GD07a}) a relation between SMBH mass and the Sérsic index was not found at near-IR wavelengths. This latter outcome is believed to be explained by the generic inconsistencies between 1D and 2D galaxy profiling which are currently under further investigation.

preprint2010arXiv

Galaxy and Mass Assembly (GAMA): FUV, NUV, ugrizYJHK Petrosian, Kron and Sèrsic photometry

In order to generate credible 0.1-2 μm SEDs, the GAMA project requires many Gigabytes of imaging data from a number of instruments to be re-processed into a standard format. In this paper we discuss the software infrastructure we use, and create self-consistent ugrizYJHK photometry for all sources within the GAMA sample. Using UKIDSS and SDSS archive data, we outline the pre-processing necessary to standardise all images to a common zeropoint, the steps taken to correct for seeing bias across the dataset, and the creation of Gigapixel-scale mosaics of the three 4x12 deg GAMA regions in each filter. From these mosaics, we extract source catalogues for the GAMA regions using elliptical Kron and Petrosian matched apertures. We also calculate Sérsic magnitudes for all galaxies within the GAMA sample using SIGMA, a galaxy component modelling wrapper for GALFIT 3. We compare the resultant photometry directly, and also calculate the r band galaxy LF for all photometric datasets to highlight the uncertainty introduced by the photometric method. We find that (1) Changing the object detection threshold has a minor effect on the best-fitting Schechter parameters of the overall population (M* +/- 0.055mag, α +/- 0.014, Φ* +/- 0.0005 h^3 Mpc^{-3}). (2) An offset between datasets that use Kron or Petrosian photometry regardless of the filter. (3) The decision to use circular or elliptical apertures causes an offset in M* of 0.20mag. (4) The best-fitting Schechter parameters from total-magnitude photometric systems (such as SDSS modelmag or Sérsic magnitudes) have a steeper faint-end slope than photometry dependent on Kron or Petrosian magnitudes. (5) Our Universe's total luminosity density, when calculated using Kron or Petrosian r-band photometry, is underestimated by at least 15%.

preprint2010arXiv

The ugrizYJHK luminosity distributions and densities from the combined MGC, SDSS and UKIDSS LAS datasets

We combine data from the MGC, SDSS and UKIDSS LAS surveys to produce ugrizYJHK luminosity functions and densities from within a common, low redshift volume (z<0.1, ~71,000 h_1^-3 Mpc^3 for L* systems) with 100 per cent spectroscopic completeness. In the optical the fitted Schechter functions are comparable in shape to those previously reported values but with higher normalisations (typically 0, 30, 20, 15, 5 per cent higher phi*-values in u, g, r, i, z respectively over those reported by the SDSS team). We attribute these to differences in the redshift ranges probed, incompleteness, and adopted normalisation methods. In the NIR we find significantly different Schechter function parameters (mainly in the M* values) to those previously reported and attribute this to the improvement in the quality of the imaging data over previous studies. This is the first homogeneous measurement of the extragalactic luminosity density which fully samples both the optical and near-IR regimes. Unlike previous compilations that have noted a discontinuity between the optical and near-IR regimes our homogeneous dataset shows a smooth cosmic spectral energy distribution (CSED). After correcting for dust attenuation we compare our CSED to the expected values based on recent constraints on the cosmic star-formation history and the initial mass function.

preprint2009arXiv

GAMA: towards a physical understanding of galaxy formation

The Galaxy And Mass Assembly (GAMA) project is the latest in a tradition of large galaxy redshift surveys, and is now underway on the 3.9m Anglo-Australian Telescope at Siding Spring Observatory. GAMA is designed to map extragalactic structures on scales of 1kpc - 1Mpc in complete detail to a redshift of z~0.2, and to trace the distribution of luminous galaxies out to z~0.5. The principal science aim is to test the standard hierarchical structure formation paradigm of Cold Dark Matter (CDM) on scales of galaxy groups, pairs, discs, bulges and bars. We will measure (1) the Dark Matter Halo Mass Function (as inferred from galaxy group velocity dispersions); (2) baryonic processes, such as star formation and galaxy formation efficiency (as derived from Galaxy Stellar Mass Functions); and (3) the evolution of galaxy merger rates (via galaxy close pairs and galaxy asymmetries). Additionally, GAMA will form the central part of a new galaxy database, which aims to contain 275,000 galaxies with multi-wavelength coverage from coordinated observations with the latest international ground- and space-based facilities: GALEX, VST, VISTA, WISE, HERSCHEL, GMRT and ASKAP. Together, these data will provide increased depth (over 2 magnitudes), doubled spatial resolution (0.7"), and significantly extended wavelength coverage (UV through Far-IR to radio) over the main SDSS spectroscopic survey for five regions, each of around 50 deg^2. This database will permit detailed investigations of the structural, chemical, and dynamical properties of all galaxy types, across all environments, and over a 5 billion year timeline.

preprint2006arXiv

The Millennium Galaxy Catalogue: Bulge/Disc Decomposition of 10095 Nearby Galaxies

We have modelled the light distribution in 10095 galaxies from the Millennium Galaxy Catalogue (MGC), providing publically available structural catalogues for a large, representative sample of galaxies in the local Universe. Three different models were used: (1) a single Sersic function for the whole galaxy, (2) a bulge-disc decomposition model using a de Vaucouleurs (R^{1/4}) bulge plus exponential disc, (3) a bulge-disc decomposition model using a Sersic (R^{1/n}) bulge plus exponential disc. Repeat observations for 700 galaxies demonstrate that stable measurements can be obtained for object components with a half-light radius comparable to, or larger than, the seeing half-width at half maximum. We show that with careful quality control, robust measurements can be obtained for large samples such as the MGC. We use the catalogues to show that the galaxy colour bimodality is due to the two-component nature of galaxies (i.e. bulges and discs) and not to two distinct galaxy populations. We conclude that understanding galaxy evolution demands the routine bulge-disc decomposition of the giant galaxy population at all redshifts.

preprint2006arXiv

The Millennium Galaxy Catalogue: morphological classification and bimodality in the colour-concentration plane

Using 10 095 galaxies (B < 20 mag) from the Millennium Galaxy Catalogue, we derive B-band luminosity distributions and selected bivariate brightness distributions for the galaxy population. All subdivisions extract highly correlated sub-sets of the galaxy population which consistently point towards two overlapping distributions. A clear bimodality in the observed distribution is seen in both the rest-(u-r) colour and log(n) distributions. The rest-(u-r) colour bimodality becomes more pronounced when using the core colour as opposed to global colour. The two populations are extremely well separated in the colour-log(n) plane. Using our sample of 3 314 (B < 19 mag) eyeball classified galaxies, we show that the bulge-dominated, early-type galaxies populate one peak and the bulge-less, late-type galaxies occupy the second. The early- and mid-type spirals sprawl across and between the peaks. This constitutes extremely strong evidence that the fundamental way to divide the luminous galaxy population is into bulges and discs and that the galaxy bimodality reflects the two component nature of galaxies and not two distinct galaxy classes. We argue that these two-components require two independent formation mechanisms/processes and advocate early bulge formation through initial collapse and ongoing disc formation through splashback, infall and merging/accretion. We calculate the B-band luminosity-densities and stellar-mass densities within each subdivision and estimate that the z ~ 0 stellar mass content in spheroids, bulges and discs is 35 +/- 2 per cent, 18 +/- 7 and 47 +/- 7 per cent respectively. [Abridged]