Source author record

Martin Ingram

Martin Ingram appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Applications Computation Machine Learning

Catalog footprint

What is connected

2works

3topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2022arXiv

Scaling multi-species occupancy models to large citizen science datasets

Citizen science datasets can be very large and promise to improve species distribution modelling, but detection is imperfect, risking bias when fitting models. In particular, observers may not detect species that are actually present. Occupancy models can estimate and correct for this observation process, and multi-species occupancy models exploit similarities in the observation process, which can improve estimates for rare species. However, the computational methods currently used to fit these models do not scale to large datasets. We develop approximate Bayesian inference methods and use graphics processing units (GPUs) to scale multi-species occupancy models to very large citizen science data. We fit multi-species occupancy models to one month of data from the eBird project consisting of 186,811 checklist records comprising 430 bird species. We evaluate the predictions on a spatially separated test set of 59,338 records, comparing two different inference methods -- Markov chain Monte Carlo (MCMC) and variational inference (VI) -- to occupancy models fitted to each species separately using maximum likelihood. We fitted models to the entire dataset using VI, and up to 32,000 records with MCMC. VI fitted to the entire dataset performed best, outperforming single-species models on both AUC (90.4% compared to 88.7%) and on log likelihood (-0.080 compared to -0.085). We also evaluate how well range maps predicted by the model agree with expert maps. We find that modelling the detection process greatly improves agreement and that the resulting maps agree as closely with expert maps as ones estimated using high quality survey data. Our results demonstrate that multi-species occupancy models are a compelling approach to model large citizen science datasets, and that, once the observation process is taken into account, they can model species distributions accurately.

preprint2020arXiv

Space-Time VON CRAMM: Evaluating Decision-Making in Tennis with Variational generatiON of Complete Resolution Arcs via Mixture Modeling

Sports tracking data are the high-resolution spatiotemporal observations of a competitive event. The growing collection of these data in professional sport allows us to address a fundamental problem of modern sport: how to attribute value to individual actions? Taking advantage of the smoothness of ball and player movement in tennis, we present a functional data framework for estimating expected shot value (ESV) in continuous time. Our approach is a three-step recipe: 1) a generative model for a full-resolution functional representation of ball and player trajectories using an infinite Bayesian Gaussian mixture model (GMM), 2) conditioning of the GMM on observed positional data, and 3) the prediction of shot outcomes given the functional encoding of a shot event. From the ESV we derive three metrics of central interest: value added with shot taking (VAST), Shot IQ, and value added with court coverage (VACC), which respectively attribute value to shot execution, shot selection and movement around the court. We rate player performance at the 2019 US Open on these advanced metrics and show how each adds a novel perspective to performance evaluation in tennis that goes beyond simple counts of outcomes by quantitatively assessing the decisions players make throughout a point.