Source author record

Richard A. Blythe

Richard A. Blythe appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

11works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

11 published item(s)

preprint2020arXiv

Combinatorial mappings of exclusion processes

We review various combinatorial interpretations and mappings of stationary-state probabilities of the totally asymmetric, partially asymmetric and symmetric simple exclusion processes (TASEP, PASEP, SSEP respectively). In these steady states, the statistical weight of a configuration is determined from a matrix product, which can be written explicitly in terms of generalised ladder operators. This lends a natural association to the enumeration of random walks with certain properties. Specifically, there is a one-to-many mapping of steady-state configurations to a larger state space of discrete paths, which themselves map to an even larger state space of number permutations. It is often the case that the configuration weights in the extended space are of a relatively simple form (e.g., a Boltzmann-like distribution). Meanwhile, various physical properties of the nonequilibrium steady state - such as the entropy - can be interpreted in terms of how this larger state space has been partitioned. These mappings sometimes allow physical results to be derived very simply, and conversely the physical approach allows some new combinatorial problems to be solved. This work brings together results and observations scattered in the combinatorics and statistical physics literature, and also presents new results. The review is pitched at statistical physicists who, though not professional combinatorialists, are competent and enthusiastic amateurs.

preprint2020arXiv

Communicative need modulates competition in language change

All living languages change over time. The causes for this are many, one being the emergence and borrowing of new linguistic elements. Competition between the new elements and older ones with a similar semantic or grammatical function may lead to speakers preferring one of them, and leaving the other to go out of use. We introduce a general method for quantifying competition between linguistic elements in diachronic corpora which does not require language-specific resources other than a sufficiently large corpus. This approach is readily applicable to a wide range of languages and linguistic subsystems. Here, we apply it to lexical data in five corpora differing in language, type, genre, and time span. We find that changes in communicative need are consistently predictive of lexical competition dynamics. Near-synonymous words are more likely to directly compete if they belong to a topic of conversation whose importance to language users is constant over time, possibly leading to the extinction of one of the competing words. By contrast, in topics which are increasing in importance for language users, near-synonymous words tend not to compete directly and can coexist. This suggests that, in addition to direct competition between words, language change can be driven by competition between topics or semantic subspaces.

preprint2020arXiv

Coupled differentiation and division of embryonic stem cells inferred from clonal snapshots

The deluge of single-cell data obtained by sequencing, imaging and epigenetic markers has led to an increasingly detailed description of cell state. However, it remains challenging to identify how cells transition between different states, in part because data are typically limited to snapshots in time. A prerequisite for inferring cell state transitions from such snapshots is to distinguish whether transitions are coupled to cell divisions. To address this, we present two minimal branching process models of cell division and differentiation in a well-mixed population. These models describe dynamics where differentiation and division are coupled or uncoupled. For each model, we derive analytic expressions for each subpopulation's mean and variance and for the likelihood, allowing exact Bayesian parameter inference and model selection in the idealised case of fully observed trajectories of differentiation and division events. In the case of snapshots, we present a sample path algorithm and use this to predict optimal temporal spacing of measurements for experimental design. We then apply this methodology to an \textit{in vitro} dataset assaying the clonal growth of epiblast stem cells in culture conditions promoting self-renewal or differentiation. Here, the larger number of cell states necessitates approximate Bayesian computation. For both culture conditions, our inference supports the model where cell state transitions are coupled to division. For culture conditions promoting differentiation, our analysis indicates a possible shift in dynamics, with these processes becoming more coupled over time.

preprint2020arXiv

Inter-particle ratchet effect determines global current of heterogeneous particles diffusing in confinement

In a model of $N$ volume-excluding spheres in a $d$-dimensional tube, we consider how differences between particles in their drift velocities, diffusivities, and sizes influence the steady state distribution and axial particle current. We show that the model is exactly solvable when the geometrical constraints prevent any particle from overtaking every other -- a notion we term quasi-one-dimensionality. Then, due to a ratchet effect, the current is biased towards the velocities of the least diffusive particles. We consider special cases of this model in one dimension, and derive the exact joint gap distribution for driven tracers in a passive bath. We describe the relationship between phase space structure and irreversible drift that makes the quasi-one-dimensional supposition key to the model's solvability.

preprint2019arXiv

Challenges in detecting evolutionary forces in language change using diachronic corpora

Newberry et al. (Detecting evolutionary forces in language change, Nature 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having applied a test from population genetics (the Frequency Increment Test) to a number of relevant examples, they suggest stochasticity has a previously under-appreciated role in language evolution. We replicate their results and find that while the overall observation holds, results produced by this approach on individual time series can be sensitive to how the corpus is organized into temporal segments (binning). Furthermore, we use a large set of simulations in conjunction with binning to systematically explore the range of applicability of the Frequency Increment Test. We conclude that care should be exercised with interpreting results of tests like the Frequency Increment Test on individual series, given the researcher degrees of freedom available when applying the test to corpus data, and fundamental differences between genetic and linguistic data. Our findings have implications for selection testing and temporal binning in general, as well as demonstrating the usefulness of simulations for evaluating methods newly introduced to the field.

preprint2019arXiv

Quantifying the dynamics of topical fluctuations in language

The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements in texts, with changes in frequency taken to reflect the popularity or selective fitness of an element. However, corpus frequencies may change for a wide variety of reasons, including purely random sampling effects, or because corpora are composed of contemporary media and fiction texts within which the underlying topics ebb and flow with cultural and socio-political trends. In this work, we introduce a simple model for controlling for topical fluctuations in corpora - the topical-cultural advection model - and demonstrate how it provides a robust baseline of variability in word frequency changes over time. We validate the model on a diachronic corpus spanning two centuries, and a carefully-controlled artificial language change scenario, and then use it to correct for topical fluctuations in historical time series. Finally, we use the model to show that the emergence of new words typically corresponds with the rise of a trending topic. This suggests that some lexical innovations occur due to growing communicative need in a subspace of the lexicon, and that the topical-cultural advection model can be used to quantify this.

preprint2016arXiv

Word learning under infinite uncertainty

Language learners must learn the meanings of many thousands of words, despite those words occurring in complex environments in which infinitely many meanings might be inferred by the learner as a word's true meaning. This problem of infinite referential uncertainty is often attributed to Willard Van Orman Quine. We provide a mathematical formalisation of an ideal cross-situational learner attempting to learn under infinite referential uncertainty, and identify conditions under which word learning is possible. As Quine's intuitions suggest, learning under infinite uncertainty is in fact possible, provided that learners have some means of ranking candidate word meanings in terms of their plausibility; furthermore, our analysis shows that this ranking could in fact be exceedingly weak, implying that constraints which allow learners to infer the plausibility of candidate word meanings could themselves be weak. This approach lifts the burden of explanation from `smart' word learning constraints in learners, and suggests a programme of research into weak, unreliable, probabilistic constraints on the inference of word meaning in real word learners.

preprint2015arXiv

Hierarchy of Scales in Language Dynamics

Methods and insights from statistical physics are finding an increasing variety of applications where one seeks to understand the emergent properties of a complex interacting system. One such area concerns the dynamics of language at a variety of levels of description, from the behaviour of individual agents learning simple artificial languages from each other, up to changes in the structure of languages shared by large groups of speakers over historical timescales. In this Colloquium, we survey a hierarchy of scales at which language and linguistic behaviour can be described, along with the main progress in understanding that has been made at each of them---much of which has come from the statistical physics community. We argue that future developments may arise by linking the different levels of the hierarchy together in a more coherent fashion, in particular where this allows more effective use of rich empirical data sets.

preprint2013arXiv

Parasites on parasites: coupled fluctuations in stacked contact processes

We present a model for host-parasite dynamics which incorporates both vertical and horizontal transmission as well as spatial structure. Our model consists of stacked contact processes (CP), where the dynamics of the host is a simple CP on a lattice while the dynamics of the parasite is a secondary CP which sits on top of the host-occupied sites. In the simplest case, where infection does not incur any cost, we uncover a novel effect: a nonmonotonic dependence of parasite prevalence on host turnover. Inspired by natural examples of hyperparasitism, we extend our model to multiple levels of parasites and identify a transition between the maintenance of a finite and infinite number of levels, which we conjecture is connected to a roughening transition in models of surface-growth.

preprint2013arXiv

Stochastic dynamics of lexicon learning in an uncertain and nonuniform world

We study the time taken by a language learner to correctly identify the meaning of all words in a lexicon under conditions where many plausible meanings can be inferred whenever a word is uttered. We show that the most basic form of cross-situational learning - whereby information from multiple episodes is combined to eliminate incorrect meanings - can perform badly when words are learned independently and meanings are drawn from a nonuniform distribution. If learners further assume that no two words share a common meaning, we find a phase transition between a maximally-efficient learning regime, where the learning time is reduced to the shortest it can possibly be, and a partially-efficient regime where incorrect candidate meanings for words persist at late times. We obtain exact results for the word-learning process through an equivalence to a statistical mechanical problem of enumerating loops in the space of word-meaning mappings.

preprint2013arXiv

Why is combinatorial communication rare in the natural world, and why is language an exception to this trend?

In a combinatorial communication system, some signals consist of the combinations of other signals. Such systems are more efficient than equivalent, non-combinatorial systems, yet despite this they are rare in nature. Why? Previous explanations have focused on the adaptive limits of combinatorial communication, or on its purported cognitive difficulties, but neither of these explains the full distribution of combinatorial communication in the natural world. Here we present a nonlinear dynamical model of the emergence of combinatorial communication that, unlike previous models, considers how initially non-communicative behaviour evolves to take on a communicative function. We derive three basic principles about the emergence of combinatorial communication. We hence show that the interdependence of signals and responses places significant constraints on the historical pathways by which combinatorial signals might emerge, to the extent that anything other than the most simple form of combinatorial communication is extremely unlikely. We also argue that these constraints can be bypassed if individuals have the socio-cognitive capacity to engage in ostensive communication. Humans, but probably no other species, have this ability. This may explain why language, which is massively combinatorial, is such an extreme exception to nature's general trend for non-combinatorial communication.