Source author record

Spencer Wheatley

Spencer Wheatley appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2020arXiv

Revisiting the predictability of the Haicheng and Tangshan earthquakes

We analyse the compiled set of precursory data that were reported to be available in real time before the Ms 7.5 Haicheng earthquake in Feb. 1975 and the Ms 7.6-7.8 Tangshan earthquake in July 1976. We propose a robust and simple coarse-graining method consisting in aggregating and counting how all the anomalies together (geodesy, levelling, geomagnetism, soil resistivity, Earth currents, gravity, Earth stress, well water radon, well water level) develop as a function of time. We demonstrate a strong evidence for the existence of an acceleration of the number of anomalies leading up to the major Haicheng and Tangshan earthquakes. In particular for the Tangshan earthquake, the frequency of occurrence of anomalies is found to be well described by the log-periodic power law singularity (LPPLS) model, previously proposed for the prediction of engineering failures and later adapted to the prediction of financial crashes. Based on a mock real-time prediction experiment, and simulation study, we show the potential for an early warning system with lead-time of a few days, based on this methodology of monitoring accelerated rates of anomalies.

preprint2016arXiv

The Extreme Risk of Personal Data Breaches & The Erosion of Privacy

Personal data breaches from organisations, enabling mass identity fraud, constitute an \emph{extreme risk}. This risk worsens daily as an ever-growing amount of personal data are stored by organisations and on-line, and the attack surface surrounding this data becomes larger and harder to secure. Further, breached information is distributed and accumulates in the hands of cyber criminals, thus driving a cumulative erosion of privacy. Statistical modeling of breach data from 2000 through 2015 provides insights into this risk: A current maximum breach size of about 200 million is detected, and is expected to grow by fifty percent over the next five years. The breach sizes are found to be well modeled by an \emph{extremely heavy tailed} truncated Pareto distribution, with tail exponent parameter decreasing linearly from 0.57 in 2007 to 0.37 in 2015. With this current model, given a breach contains above fifty thousand items, there is a ten percent probability of exceeding ten million. A size effect is unearthed where both the frequency and severity of breaches scale with organisation size like $s^{0.6}$. Projections indicate that the total amount of breached information is expected to double from two to four billion items within the next five years, eclipsing the population of users of the Internet. This massive and uncontrolled dissemination of personal identities raises fundamental concerns about privacy.

preprint2015arXiv

Of Disasters and Dragon Kings: A Statistical Analysis of Nuclear Power Incidents & Accidents

We provide, and perform a risk theoretic statistical analysis of, a dataset that is 75 percent larger than the previous best dataset on nuclear incidents and accidents, comparing three measures of severity: INES (International Nuclear Event Scale), radiation released, and damage dollar losses. The annual rate of nuclear accidents, with size above 20 Million US$, per plant, decreased from the 1950s until dropping significantly after Chernobyl (April, 1986). The rate is now roughly stable at 0.002 to 0.003, i.e., around 1 event per year across the current fleet. The distribution of damage values changed after Three Mile Island (TMI; March, 1979), where moderate damages were suppressed but the tail became very heavy, being described by a Pareto distribution with tail index 0.55. Further, there is a runaway disaster regime, associated with the "dragon-king" phenomenon, amplifying the risk of extreme damage. In fact, the damage of the largest event (Fukushima; March, 2011) is equal to 60 percent of the total damage of all 174 accidents in our database since 1946. In dollar losses we compute a 50% chance that (i) a Fukushima event (or larger) occurs in the next 50 years, (ii) a Chernobyl event (or larger) occurs in the next 27 years and (iii) a TMI event (or larger) occurs in the next 10 years. Finally, we find that the INES scale is inconsistent. To be consistent with damage, the Fukushima disaster would need to have an INES level of 11, rather than the maximum of 7.

preprint2014arXiv

Effective Measure of Endogeneity for the Autoregressive Conditional Duration Point Processes via Mapping to the Self-Excited Hawkes Process

In order to disentangle the internal dynamics from exogenous factors within the Autoregressive Conditional Duration (ACD) model, we present an effective measure of endogeneity. Inspired from the Hawkes model, this measure is defined as the average fraction of events that are triggered due to internal feedback mechanisms within the total population. We provide a direct comparison of the Hawkes and ACD models based on numerical simulations and show that our effective measure of endogeneity for the ACD can be mapped onto the "branching ratio" of the Hawkes model.

preprint2014arXiv

Estimation of the Hawkes Process With Renewal Immigration Using the EM Algorithm

We introduce the Hawkes process with renewal immigration and make its statistical estimation possible with two Expectation Maximization (EM) algorithms. The standard Hawkes process introduces immigrant points via a Poisson process, and each immigrant has a subsequent cluster of associated offspring of multiple generations. We generalize the immigration to come from a Renewal process; introducing dependence between neighbouring clusters, and allowing for over/under dispersion in cluster locations. This complicates evaluation of the likelihood since one needs to know which subset of the observed points are immigrants. Two EM algorithms enable estimation here: The first is an extension of an existing algorithm that treats the entire branching structure - which points are immigrants, and which point is the parent of each offspring - as missing data. The second considers only if a point is an immigrant or not as missing data and can be implemented with linear time complexity. Both algorithms are found to be consistent in simulation studies. Further, we show that misspecifying the immigration process introduces signficant bias into model estimation-- especially the branching ratio, which quantifies the strength of self excitation. Thus, this extended model provides a valuable alternative model in practice.