Catalog footprint

What is connected

24works
34topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

24 published item(s)

preprint2025arXiv

Explaining Necessary Truths

Knowing the truth is rarely enough -- we also seek out reasons why the fact is true. While much is known about how we explain contingent truths, we understand less about how we explain facts, such as those in mathematics, that are true as a matter of logical necessity. We present a framework, based in computational complexity, where explanations for deductive truths co-emerge with discoveries of simplifying steps during the search process. When such structures are missing, we revert, in turn, to error-based reasons, where a (corrected) mistake can serve as fictitious, but explanatory, contingency-cause: not making the mistake serves as a reason why the truth takes the form it does. We simulate human subjects, using GPT-4o, presented with SAT puzzles of varying complexity and reasonableness, validating our theory and showing how its predictions can be tested in future human studies.

preprint2022arXiv

Epistemic Phase Transitions in Mathematical Proofs

Mathematical proofs are both paradigms of certainty and some of the most explicitly-justified arguments that we have in the cultural record. Their very explicitness, however, leads to a paradox, because the probability of error grows exponentially as the argument expands. When a mathematician encounters a proof, how does she come to believe it? Here we show that, under a cognitively-plausible belief formation mechanism combining deductive and abductive reasoning, belief in mathematical arguments can undergo what we call an epistemic phase transition: a dramatic and rapidly-propagating jump from uncertainty to near-complete confidence at reasonable levels of claim-to-claim error rates. To show this, we analyze an unusual dataset of forty-eight machine-aided proofs from the formalized reasoning system Coq, including major theorems ranging from ancient to 21st Century mathematics, along with five hand-constructed cases including Euclid, Apollonius, Hernstein's Topics in Algebra, and Andrew Wiles's proof of Fermat's Last Theorem. Our results bear both on recent work in the history and philosophy of mathematics on how we understand proofs, and on a question, basic to cognitive science, of how we justify complex beliefs.

preprint2020arXiv

When Science is a Game

What happens when scientists are, at certain points in a field's development, playing a game? I present a framework for such an analysis that draws on the theory of games provided by the historian Johan Huizinga. Huizinga gives five conditions for a social practice to become a game: free engagement, disconnection, boundedness in time and arena, the order-creation of rules, and the presence of tension. Application of this theory to scientific practice predicts patterns of behavior that can be tested by quantitative analysis: the emergence of hard boundaries between disciplines, the closure of loopholes in theory creation, resistance to certain innovations in journal publication, and the ways in which scientists fail to prosecute colleagues who engage in questionable research practices.

preprint2019arXiv

How we do things with words: Analyzing text as social and cultural data

In this article we describe our experiences with computational text analysis. We hope to achieve three primary goals. First, we aim to shed light on thorny issues not always at the forefront of discussions about computational text analysis methods. Second, we hope to provide a set of best practices for working with thick social and cultural concepts. Our guidance is based on our own experiences and is therefore inherently imperfect. Still, given our diversity of disciplinary backgrounds and research practices, we hope to capture a range of ideas and identify commonalities that will resonate for many. And this leads to our final goal: to help promote interdisciplinary collaborations. Interdisciplinary insights and partnerships are essential for realizing the full potential of any computational text analysis that involves social and cultural concepts, and the more we are able to bridge these divides, the more fruitful we believe our work will be.

preprint2016arXiv

A quantitative definition of organismality and its application to lichen

The organism is a fundamental concept in biology. However there is no universally accepted, formal, and yet broadly applicable definition of what an organism is. Here we introduce a candidate definition. We adopt the view that the "organism" is a functional concept, used by scientists to address particular questions concerning the future state of a biological system, rather than something wholly defined by that system. In this approach organisms are a coarse-graining of a fine-grained dynamical model of a biological system. Crucially, the coarse-graining of the system into organisms is chosen so that their dynamics can be used by scientists to make accurate predictions of those features of the biological system that interests them, and do so with minimal computational burden. To illustrate our framework we apply it to a dynamic model of lichen symbiosis---a system where either the lichen or its constituent fungi and algae could reasonably be considered "organisms." We find that the best choice for what organisms are in this scenario are complex mixtures of many entities that do not resemble standard notions of organisms. When we restrict our allowed coarse-grainings to more traditional types of organisms, we find that ecological conditions, such as niche competition and predation pressure, play a significant role in determining the best choice for organisms.

preprint2016arXiv

The Evolution of Wikipedia's Norm Network

Social norms have traditionally been difficult to quantify. In any particular society, their sheer number and complex interdependencies often limit a system-level analysis. One exception is that of the network of norms that sustain the online Wikipedia community. We study the fifteen-year evolution of this network using the interconnected set of pages that establish, describe, and interpret the community's norms. Despite Wikipedia's reputation for \textit{ad hoc} governance, we find that its normative evolution is highly conservative. The earliest users create norms that both dominate the network and persist over time. These core norms govern both content and interpersonal interactions using abstract principles such as neutrality, verifiability, and assume good faith. As the network grows, norm neighborhoods decouple topologically from each other, while increasing in semantic coherence. Taken together, these results suggest that the evolution of Wikipedia's norm network is akin to bureaucratic systems that predate the information age.

preprint2015arXiv

Common Knowledge on Networks

Common knowledge of intentions is crucial to basic social tasks ranging from cooperative hunting to oligopoly collusion, riots, revolutions, and the evolution of social norms and human culture. Yet little is known about how common knowledge leaves a trace on the dynamics of a social network. Here we show how an individual's network properties---primarily local clustering and betweenness centrality---provide strong signals of the ability to successfully participate in common knowledge tasks. These signals are distinct from those expected when practices are contagious, or when people use less-sophisticated heuristics that do not yield true coordination. This makes it possible to infer decision rules from observation. We also find that tasks that require common knowledge can yield significant inequalities in success, in contrast to the relative equality that results when practices spread by contagion alone.

preprint2015arXiv

Optimal high-level descriptions of dynamical systems

To analyze high-dimensional systems, many fields in science and engineering rely on high-level descriptions, sometimes called "macrostates," "coarse-grainings," or "effective theories". Examples of such descriptions include the thermodynamic properties of a large collection of point particles undergoing reversible dynamics, the variables in a macroeconomic model describing the individuals that participate in an economy, and the summary state of a cell composed of a large set of biochemical networks. Often these high-level descriptions are constructed without considering the ultimate reason for needing them in the first place. Here, we formalize and quantify one such purpose: the need to predict observables of interest concerning the high-dimensional system with as high accuracy as possible, while minimizing the computational cost of doing so. The resulting State Space Compression (SSC) framework provides a guide for how to solve for the {optimal} high-level description of a given dynamical system, rather than constructing it based on human intuition alone. In this preliminary report, we introduce SSC, and illustrate it with several information-theoretic quantifications of "accuracy", all with different implications for the optimal compression. We also discuss some other possible applications of SSC beyond the goal of accurate prediction. These include SSC as a measure of the complexity of a dynamical system, and as a way to quantify information flow between the scales of a system.

preprint2014arXiv

Demystifying Information-Theoretic Clustering

We propose a novel method for clustering data which is grounded in information-theoretic principles and requires no parametric assumptions. Previous attempts to use information theory to define clusters in an assumption-free way are based on maximizing mutual information between data and cluster labels. We demonstrate that this intuition suffers from a fundamental conceptual flaw that causes clustering performance to deteriorate as the amount of data increases. Instead, we return to the axiomatic foundations of information theory to define a meaningful clustering measure based on the notion of consistency under coarse-graining for finite data.

preprint2014arXiv

Group Minds and the Case of Wikipedia

Group-level cognitive states are widely observed in human social systems, but their discussion is often ruled out a priori in quantitative approaches. In this paper, we show how reference to the irreducible mental states and psychological dynamics of a group is necessary to make sense of large scale social phenomena. We introduce the problem of mental boundaries by reference to a classic problem in the evolution of cooperation. We then provide an explicit quantitative example drawn from ongoing work on cooperation and conflict among Wikipedia editors, showing how some, but not all, effects of individual experience persist in the aggregate. We show the limitations of methodological individualism, and the substantial benefits that come from being able to refer to collective intentions, and attributions of cognitive states of the form "what the group believes" and "what the group values".

preprint2013arXiv

Bootstrap Methods for the Empirical Study of Decision-Making and Information Flows in Social Systems

We characterize the statistical bootstrap for the estimation of information-theoretic quantities from data, with particular reference to its use in the study of large-scale social phenomena. Our methods allow one to preserve, approximately, the underlying axiomatic relationships of information theory---in particular, consistency under arbitrary coarse-graining---that motivate use of these quantities in the first place, while providing reliability comparable to the state of the art for Bayesian estimators. We show how information-theoretic quantities allow for rigorous empirical study of the decision-making capacities of rational agents and the time-asymmetric flows of information in distributed systems. We provide illustrative examples by reference to ongoing collaborative work on the semantic structure of the British Criminal Court system and the conflict dynamics of the contemporary Afghanistan insurgency.

preprint2013arXiv

Collective Phenomena and Non-Finite State Computation in a Human Social System

We investigate the computational structure of a paradigmatic example of distributed social interaction: that of the open-source Wikipedia community. We examine the statistical properties of its cooperative behavior, and perform model selection to determine whether this aspect of the system can be described by a finite-state process, or whether reference to an effectively unbounded resource allows for a more parsimonious description. We find strong evidence, in a majority of the most-edited pages, in favor of a collective-state model, where the probability of a "revert" action declines as the square root of the number of non-revert actions seen since the last revert. We provide evidence that the emergence of this social counter is driven by collective interaction effects, rather than properties of individual users.

preprint2013arXiv

Dynamical Structure of a Traditional Amazonian Social Network

Reciprocity is a vital feature of social networks, but relatively little is known about its temporal structure or the mechanisms underlying its persistence in real world behavior. In pursuit of these two questions, we study the stationary and dynamical signals of reciprocity in a network of manioc beer (Spanish: chicha; Tsimane': shocdye') drinking events in a Tsimane' village in lowland Bolivia. At the stationary level, our analysis reveals that social exchange within the community is heterogeneously patterned according to kinship and spatial proximity. A positive relationship between the frequencies at which two families host each other, controlling for kinship and proximity, provides evidence for stationary reciprocity. Our analysis of the dynamical structure of this network presents a novel method for the study of conditional, or non-stationary, reciprocity effects. We find evidence that short-timescale reciprocity (within three days) is present among non- and distant-kin pairs; conversely, we find that levels of cooperation among close kin can be accounted for on the stationary hypothesis alone.

preprint2013arXiv

Estimating Functions of Distributions Defined over Spaces of Unknown Size

We consider Bayesian estimation of information-theoretic quantities from data, using a Dirichlet prior. Acknowledging the uncertainty of the event space size $m$ and the Dirichlet prior's concentration parameter $c$, we treat both as random variables set by a hyperprior. We show that the associated hyperprior, $P(c, m)$, obeys a simple "Irrelevance of Unseen Variables" (IUV) desideratum iff $P(c, m) = P(c) P(m)$. Thus, requiring IUV greatly reduces the number of degrees of freedom of the hyperprior. Some information-theoretic quantities can be expressed multiple ways, in terms of different event spaces, e.g., mutual information. With all hyperpriors (implicitly) used in earlier work, different choices of this event space lead to different posterior expected values of these information-theoretic quantities. We show that there is no such dependence on the choice of event space for a hyperprior that obeys IUV. We also derive a result that allows us to exploit IUV to greatly simplify calculations, like the posterior expected mutual information or posterior expected multi-information. We also use computer experiments to favorably compare an IUV-based estimator of entropy to three alternative methods in common use. We end by discussing how seemingly innocuous changes to the formalization of an estimation problem can substantially affect the resultant estimates of posterior expectations.

preprint2013arXiv

Robust Compressed Sensing and Sparse Coding with the Difference Map

In compressed sensing, we wish to reconstruct a sparse signal $x$ from observed data $y$. In sparse coding, on the other hand, we wish to find a representation of an observed signal $y$ as a sparse linear combination, with coefficients $x$, of elements from an overcomplete dictionary. While many algorithms are competitive at both problems when $x$ is very sparse, it can be challenging to recover $x$ when it is less sparse. We present the Difference Map, which excels at sparse recovery when sparseness is lower and noise is higher. The Difference Map out-performs the state of the art with reconstruction from random measurements and natural image reconstruction via sparse coding.

preprint2012arXiv

Dynamics and Processing in Finite Self-Similar Networks

A common feature of biological networks is the geometric property of self-similarity. Molecular regulatory networks through to circulatory systems, nervous systems, social systems and ecological trophic networks, show self-similar connectivity at multiple scales. We analyze the relationship between topology and signaling in contrasting classes of such topologies. We find that networks differ in their ability to contain or propagate signals between arbitrary nodes in a network depending on whether they possess branching or loop-like features. Networks also differ in how they respond to noise, such that one allows for greater integration at high noise, and this performance is reversed at low noise. Surprisingly, small-world topologies, with diameters logarithmic in system size, have slower dynamical timescales, and may be less integrated (more modular) than networks with longer path lengths. All of these phenomena are essentially mesoscopic, vanishing in the infinite limit but producing strong effects at sizes and timescales relevant to biology.

preprint2012arXiv

Effective Theories for Circuits and Automata

Abstracting an effective theory from a complicated process is central to the study of complexity. Even when the underlying mechanisms are understood, or at least measurable, the presence of dissipation and irreversibility in biological, computational and social systems makes the problem harder. Here we demonstrate the construction of effective theories in the presence of both irreversibility and noise, in a dynamical model with underlying feedback. We use the Krohn-Rhodes theorem to show how the composition of underlying mechanisms can lead to innovations in the emergent effective theory. We show how dissipation and irreversibility fundamentally limit the lifetimes of these emergent structures, even though, on short timescales, the group properties may be enriched compared to their noiseless counterparts.

preprint2011arXiv

Evidence of strategic periodicities in collective conflict dynamics

We analyze the timescales of conflict decision-making in a primate society. We present evidence for multiple, periodic timescales associated with social decision-making and behavioral patterns. We demonstrate the existence of periodicities that are not directly coupled to environmental cycles or known ultraridian mechanisms. Among specific biological and socially-defined demographic classes, periodicities span timescales between hours and days, and many are not driven by exogenous or internal regularities. Our results indicate that they are instead driven by strategic responses to social interaction patterns. Analyses also reveal that a class of individuals, playing a critical functional role, policing, have a signature timescale on the order of one hour. We propose a classification of behavioral timescales analogous to those of the nervous system, with high-frequency, or $α$-scale, behavior occurring on hour-long scales, through to multi-hour, or $β$-scale, behavior, and, finally $γ$ periodicities observed on a timescale of days.

preprint2011arXiv

Parallel Complexity of Random Boolean Circuits

Random instances of feedforward Boolean circuits are studied both analytically and numerically. Evaluating these circuits is known to be a P-complete problem and thus, in the worst case, believed to be impossible to perform, even given a massively parallel computer, in time much less than the depth of the circuit. Nonetheless, it is found that for some ensembles of random circuits, saturation to a fixed truth value occurs rapidly so that evaluation of the circuit can be accomplished in much less parallel time than the depth of the circuit. For other ensembles saturation does not occur and circuit evaluation is apparently hard. In particular, for some random circuits composed of connectives with five or more inputs, the number of true outputs at each level is a chaotic sequence. Finally, while the average case complexity depends on the choice of ensemble, it is shown that for all ensembles it is possible to simultaneously construct a typical circuit together with its solution in polylogarithmic parallel time.

preprint2010arXiv

Inductive Game Theory and the Dynamics of Animal Conflict

Conflict destabilizes social interactions and impedes cooperation at multiple scales of biological organization. Of fundamental interest are the causes of turbulent periods of conflict. We analyze conflict dynamics in a monkey society model system. We develop a technique, Inductive Game Theory, to extract directly from time-series data the decision-making strategies used by individuals and groups. This technique uses Monte Carlo simulation to test alternative causal models of conflict dynamics. We find individuals base their decision to fight on memory of social factors, not on short timescale ecological resource competition. Furthermore, the social assessments on which these decisions are based are triadic (self in relation to another pair of individuals), not pairwise. We show that this triadic decision making causes long conflict cascades and that there is a high population cost of the large fights associated with these cascades. These results suggest that individual agency has been over-emphasized in the social evolution of complex aggregates, and that pair-wise formalisms are inadequate. An appreciation of the empirical foundations of the collective dynamics of conflict is a crucial step towards its effective management.

preprint2010arXiv

Neutron Stars in f(R) Gravity with Perturbative Constraints

We study the structure of neutron stars in f(R) gravity theories with perturbative constraints. We derive the modified Tolman-Oppenheimer-Volkov equations and solve them for a polytropic equation of state. We investigate the resulting modifications to the masses and radii of neutron stars and show that observations of surface phenomena alone cannot break the degeneracy between altering the theory of gravity versus choosing a different equation of state of neutron-star matter. On the other hand, observations of neutron-star cooling, which depends on the density of matter at the stellar interior, can place significant constraints on the parameters of the theory.

preprint2007arXiv

Cluster Mass Estimators from CMB Temperature and Polarization Lensing

Upcoming Sunyaev-Zel'dovich surveys are expected to return ~10^4 intermediate mass clusters at high redshift. Their average masses must be known to same accuracy as desired for the dark energy properties. Internal to the surveys, the CMB potentially provides a source for lensing mass measurements whose distance is precisely known and behind all clusters. We develop statistical mass estimators from 6 quadratic combinations of CMB temperature and polarization fields that can simultaneously recover large-scale structure and cluster mass profiles. The performance of these estimators on idealized NFW clusters suggests that surveys with a ~1' beam and 10uK' noise in uncontaminated temperature maps can make a ~10sigma detection, or equivalently a ~10% mass measurement for each 10^3 set of clusters. With internal or external acoustic scale E-polarization measurements, the ET cross correlation estimator can provide a stringent test for contaminants on a first detection at \~1/3 the significance. For surveys that reach below 3muK', the EB cross correlation estimator should provide the most precise measurements and potentially the strongest control over contaminants.

preprint2004arXiv

Effects of the Sound Speed of Quintessence on the Microwave Background and Large Scale Structure

We consider how quintessence models in which the sound speed differs from the speed of light and varies with time affect the cosmic microwave background and the fluctuation power spectrum. Significant modifications occur on length scales related to the Hubble radius during epochs in which the sound speed is near zero and the quintessence contributes a non-negligible fraction of the total energy density. For the microwave background, we find that the usual enhancement of the lowest multipole moments by the integrated Sachs-Wolfe effect can be modified, resulting in suppression or bumps instead. Also, the sound speed can produce oscillations and other effects at wavenumbers $k > 10^{-2}$ h/Mpc in the fluctuation power spectrum.

preprint2002arXiv

An Eternal Time Machine in 2+1 Dimensional anti-de Sitter Space

2+1 dimensional anti-de Sitter space has been the subject of much recent investigation. Studies of the behaviour of point particles in this space have given us a greater understanding of the BTZ black hole solutions produced by topological identification of adS isometries. In this paper, we present a new configuration of two orbiting massive point particles that leads to an ``eternal'' time machine, where closed timelike curves fill the entire space. In contrast to previous solutions, this configuration has no event or chronology horizons. Another interesting feature is that there is no lower bound on the relative velocities of the point masses used to construct the time machine; as long as the particles exceed a certain mass threshold, an eternal time machine will be produced.