Source author record

Joshua B. Plotkin

Joshua B. Plotkin appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

23works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

23 published item(s)

preprint2022arXiv

Selfish optimization and collective learning in populations

A selfish learner seeks to maximize their own success, disregarding others. When success is measured as payoff in a game played against another learner, mutual selfishness typically fails to produce the optimal outcome for a pair of individuals. However, learners often operate in populations, and each learner may have a limited duration of interaction with any other individual. Here, we compare selfish learning in stable pairs to selfish learning with stochastic encounters in a population. We study gradient-based optimization in repeated games like the prisoner's dilemma, which feature multiple Nash equilibria, many of which are suboptimal. We find that myopic, selfish learning, when distributed in a population via ephemeral encounters, can reverse the dynamics that occur in stable pairs. In particular, when there is flexibility in partner choice, selfish learning in large populations can produce optimal payoffs in repeated social dilemmas. This result holds for the entire population, not just for a small subset of individuals. Furthermore, as the population size grows, the timescale to reach the optimal population payoff remains finite in the number of learning steps per individual. While it is not universally true that interacting with many partners in a population improves outcomes, this form of collective learning achieves optimality for several important classes of social dilemmas. We conclude that naïve learning can be surprisingly effective in populations of individuals navigating conflicts of interest.

preprint2021arXiv

Evolution of prosocial behavior in multilayer populations

Human societies include diverse social relationships. Friends, family, business colleagues, and online contacts can all contribute to one's social life. Individuals may behave differently in different domains, but success in one domain may engender success in another. Here, we study this problem using multilayer networks to model multiple domains of social interactions, in which individuals experience different environments and may express different behaviors. We provide a mathematical analysis and find that coupling between layers tends to promote prosocial behavior. Even if prosociality is disfavored in each layer alone, multilayer coupling can promote its proliferation in all layers simultaneously. We apply this analysis to six real-world multilayer networks, ranging from the socio-emotional and professional relationships in a Zambian community, to the online and offline relationships within an academic University. We discuss the implications of our results, which suggest that small modifications to interactions in one domain may catalyze prosociality in a different domain.

preprint2020arXiv

The natural selection of good science

Scientists in some fields are concerned that many, or even most, published results are false. A high rate of false positives might arise accidentally, from shoddy research practices. Or it might be the inevitable result of institutional incentives that reward publication irrespective of veracity. Recent models and discussion of scientific culture predict selection for false-positive publications, as research labs that publish more positive findings out-compete more diligent labs. There is widespread debate about how scientific practices should be modified to avoid this degeneration. Some analyses suggest that "bad science" will persist even when labs are incentivized to undertake replication studies, and penalized for publications that later fail to replicate. Here we develop a framework for modelling the cultural evolution of research practices that allows labs to expend effort on theory - enabling them, at a cost, to focus on hypotheses that are more likely to be true on theoretical grounds. Theory restores the evolution of high effort in laboratory practice, and it suppresses false-positive publications to a technical minimum, even in the absence of replication. In fact, the mere ability choose between two sets of hypotheses - one with greater chance of being correct than the other - promotes better science than can be achieved by having effortless access to the better set of hypotheses. Combining theory and replication can have a synergistic effect in promoting good scientific methodology and reducing the rate of false-positive publications. Based on our analysis we propose four simple rules to promote good science in the face of pressure to publish.

preprint2016arXiv

Host-pathogen coevolution and the emergence of broadly neutralizing antibodies in chronic infections

The vertebrate adaptive immune system provides a flexible and diverse set of molecules to neutralize pathogens. Yet, viruses such as HIV can cause chronic infections by evolving as quickly as the adaptive immune system, forming an evolutionary arms race. Here we introduce a mathematical framework to study the coevolutionary dynamics of antibodies with antigens within a host. We focus on changes in the binding interactions between the antibody and antigen populations, which result from the underlying stochastic evolution of genotype frequencies driven by mutation, selection, and drift. We identify the critical viral and immune parameters that determine the distribution of antibody-antigen binding affinities. We also identify definitive signatures of coevolution that measure the reciprocal response between antibodies and viruses, and we introduce experimentally measurable quantities that quantify the extent of adaptation during continual coevolution of the two opposing populations. Using this analytical framework, we infer rates of viral and immune adaptation based on time-shifted neutralization assays in two HIV-infected patients. Finally, we analyze competition between clonal lineages of antibodies and characterize the fate of a given lineage in terms of the state of the antibody and viral populations. In particular, we derive the conditions that favor the emergence of broadly neutralizing antibodies, which may be useful in designing a vaccine against HIV.

preprint2016arXiv

Integrating Theory and Experiment to Explain the Breakdown of Population Synchrony in a Complex Microbial Community

We consider the extension of the `Moran effect', where correlated noise generates synchrony between isolated single species populations, to the study of synchrony between populations embedded in multi-species communities. In laboratory experiments on complex microbial communities, comprising both predators (protozoa) and prey (bacteria), we observe synchrony in abundances between isolated replicates. A breakdown in synchrony occurs for both predator and prey as the reactor dilution rate increases, which corresponds to both an increased rate of input of external resources and an increased effective mortality though washout. The breakdown is more rapid, however, for the lower trophic level. We can explain this phenomenon using a mathematical framework for determining synchrony between populations in multi-species communities at equilibrium. We assume that there are multiple sources of environmental noise with different degrees of correlation that affect the individual species population dynamics differently. The deterministic dynamics can then influence the degree of synchrony between species in different communities. In the case of a stable equilibrium community synchrony is controlled by the eigenvalue with smallest negative real part. Intuitively fluctuations are minimally damped in this direction. We show that the experimental observations are consistent with this framework but only for multiplicative noise.

preprint2015arXiv

Small games and long memories promote cooperation

Complex social behaviors lie at the heart of many of the challenges facing evolutionary biology, sociology, economics, and beyond. For evolutionary biologists in particular the question is often how such behaviors can arise \textit{de novo} in a simple evolving system. How can group behaviors such as collective action, or decision making that accounts for memories of past experience, emerge and persist? Evolutionary game theory provides a framework for formalizing these questions and admitting them to rigorous study. Here we develop such a framework to study the evolution of sustained collective action in multi-player public-goods games, in which players have arbitrarily long memories of prior rounds of play and can react to their experience in an arbitrary way. To study this problem we construct a coordinate system for memory-$m$ strategies in iterated $n$-player games that permits us to characterize all the cooperative strategies that resist invasion by any mutant strategy, and thus stabilize cooperative behavior. We show that while larger games inevitably make cooperation harder to evolve, there nevertheless always exists a positive volume of strategies that stabilize cooperation provided the population size is large enough. We also show that, when games are small, longer-memory strategies make cooperation easier to evolve, by increasing the number of ways to stabilize cooperation. Finally we explore the co-evolution of behavior and memory capacity, and we find that longer-memory strategies tend to evolve in small games, which in turn drives the evolution of cooperation even when the benefits for cooperation are low.

preprint2015arXiv

The diversity of evolutionary dynamics on epistatic versus non-epistatic fitness landscapes

The class of epistatic fitness landscapes is much more diverse than the class of non-epistatic landscapes, and so it stands to reason that there exist dynamical phenomena that can only be realized in the presence of epistasis. Here, we compare evolutionary dynamics on all finite epistatic landscapes versus all finite non-epistatic landscapes, under weak mutation. We first analyze the mean fitness trajectory - that is, the time course of the expected fitness of a population. We show that for any epistatic fitness landscape and starting genotype, there always exists a non-epistatic fitness landscape and starting genotype that produces the exact same mean fitness trajectory. Thus, surprisingly, the space of mean fitness trajectories that can be realized by epistatic landscapes is no more diverse than the space of mean fitness trajectories that can be realized by non-epistatic landscapes. On the other hand, we show that epistatic fitness landscapes can produce dynamics in the time-evolution of the variance in fitness across replicate populations and in the time-evolution of the expected number of substitutions that cannot be produced by any non-epistatic landscape. These results on identifiability have implications for efforts to infer epistasis from the types of data often measured in experimental populations.

preprint2014arXiv

Formal properties of the probability of fixation: identities, inequalities and approximations

The formula for the probability of fixation of a new mutation is widely used in theoretical population genetics and molecular evolution. Here we derive a series of identities, inequalities and approximations for the exact probability of fixation of a new mutation under the Moran process (equivalent results hold for the approximate probability of fixation for the Wright-Fisher process after an appropriate change of variables). We show that the behavior of the logarithm of the probability of fixation is particularly simple when the selection coefficient is measured as a difference of Malthusian fitnesses, and we exploit this simplicity to derive several inequalities and approximations. We also present a comprehensive comparison of both existing and new approximations for the probability of fixation, highlighting in particular approximations that result in a reversible Markov chain when used to model the dynamics of evolution under weak mutation.

preprint2014arXiv

Historical contingency and entrenchment in protein evolution under purifying selection

The fitness contribution of an allele at one genetic site may depend on alleles at other sites, a phenomenon known as epistasis. Epistasis can profoundly influence the process of evolution in populations under selection, and can shape the course of protein evolution across divergent species. Whereas epistasis between adaptive substitutions has been the subject of extensive study, relatively little is known about epistasis under purifying selection. Here we use mechanistic models of thermodynamic stability in a ligand-binding protein to explore the structure of epistatic interactions between substitutions that fix in protein sequences under purifying selection. We find that the selection coefficients of mutations that are nearly-neutral when they fix are highly contingent on the presence of preceding mutations. Conversely, mutations that are nearly-neutral when they fix are subsequently entrenched due to epistasis with later substitutions. Our evolutionary model includes insertions and deletions, as well as point mutations, and so it allows us to quantify epistasis within each of these classes of mutations, and also to study the evolution of protein length. We find that protein length remains largely constant over time, because indels are more deleterious than point mutations. Our results imply that, even under purifying selection, protein sequence evolution is highly contingent on history and so it cannot be predicted by the phenotypic effects of mutations assayed in the wild-type sequence.

preprint2014arXiv

Inferring fitness landscapes by regression produces biased estimates of epistasis

The genotype-fitness map plays a fundamental role in shaping the dynamics of evolution. However, it is difficult to directly measure a fitness landscape in practice, because the number of possible genotypes is astronomical. One approach is to sample as many genotypes as possible, measure their fitnesses, and fit a statistical model of the landscape that includes additive and pairwise interactive effects between loci. Here we elucidate the pitfalls of using such regressions, by studying artificial but mathematically convenient fitness landscapes. We identify two sources of bias inherent in these regression procedures that each tends to under-estimate high fitnesses and over-estimate low fitnesses. We characterize these biases for random sampling of genotypes, as well as for samples drawn from a population under selection in the Wright-Fisher model of evolutionary dynamics. We show that common measures of epistasis, such as the number of monotonically increasing paths between ancestral and derived genotypes, the prevalence of sign epistasis, and the number of local fitness maxima, are distorted in the inferred landscape. As a result, the inferred landscape will provide systematically biased predictions for the dynamics of adaptation. We identify the same biases in a computational RNA-folding landscape, as well as in regulatory sequence binding data, treated with the same fitting procedure. Finally, we present a method that may ameliorate these biases in some cases.

preprint2014arXiv

The collapse of cooperation in evolving games

Game theory provides a quantitative framework for analyzing the behavior of rational agents. The Iterated Prisoner's Dilemma in particular has become a standard model for studying cooperation and cheating, with cooperation often emerging as a robust outcome in evolving populations. Here we extend evolutionary game theory by allowing players' strategies as well as their payoffs to evolve in response to selection on heritable mutations. In nature, many organisms engage in mutually beneficial interactions, and individuals may seek to change the ratio of risk to reward for cooperation by altering the resources they commit to cooperative interactions. To study this, we construct a general framework for the co-evolution of strategies and payoffs in arbitrary iterated games. We show that, as payoffs evolve, a trade-off between the benefits and costs of cooperation precipitates a dramatic loss of cooperation under the Iterated Prisoner's Dilemma; and eventually to evolution away from the Prisoner's Dilemma altogether. The collapse of cooperation is so extreme that the average payoff in a population may decline, even as the potential payoff for mutual cooperation increases. Our work offers a new perspective on the Prisoner's Dilemma and its predictions for cooperation in natural populations; and it provides a general framework to understand the co-evolution of strategies and payoffs in iterated interactions.

preprint2013arXiv

From extortion to generosity, the evolution of zero-determinant strategies in the prisoner's dilemma

Recent work has revealed a new class of "zero-determinant" (ZD) strategies for iterated, two-player games. ZD strategies allow a player to unilaterally enforce a linear relationship between her score and her opponent's score, and thus achieve an unusual degree of control over both players' long-term payoffs. Although originally conceived in the context of classical, two-player game theory, ZD strategies also have consequences in evolving populations of players. Here we explore the evolutionary prospects for ZD strategies in the Iterated Prisoner's Dilemma (IPD). Several recent studies have focused on the evolution of "extortion strategies" - a subset of zero-determinant strategies - and found them to be unsuccessful in populations. Nevertheless, we identify a different subset of ZD strategies, called "generous ZD strategies", that forgive defecting opponents, but nonetheless dominate in evolving populations. For all but the smallest population sizes, generous ZD strategies are not only robust to being replaced by other strategies, but they also can selectively replace any non-cooperative ZD strategy. Generous strategies can be generalized beyond the space of ZD strategies, and they remain robust to invasion. When evolution occurs on the full set of all IPD strategies, selection disproportionately favors these generous strategies. In some regimes, generous strategies outperform even the most successful of the well-known Iterated Prisoner's Dilemma strategies, including win-stay-lose-shift.

preprint2013arXiv

Identifying Signatures of Selection in Genetic Time Series

Both genetic drift and natural selection cause the frequencies of alleles in a population to vary over time. Discriminating between these two evolutionary forces, based on a time series of samples from a population, remains an outstanding problem with increasing relevance to modern data sets. Even in the idealized situation when the sampled locus is independent of all other loci this problem is difficult to solve, especially when the size of the population from which the samples are drawn is unknown. A standard $χ^2$-based likelihood ratio test was previously proposed to address this problem. Here we show that the $χ^2$ test of selection substantially underestimates the probability of Type I error, leading to more false positives than indicated by its $P$-value, especially at stringent $P$-values. We introduce two methods to correct this bias. The empirical likelihood ratio test (ELRT) rejects neutrality when the likelihood ratio statistic falls in the tail of the empirical distribution obtained under the most likely neutral population size. The frequency increment test (FIT) rejects neutrality if the distribution of normalized allele frequency increments exhibits a mean that deviates significantly from zero. We characterize the statistical power of these two tests for selection, and we apply them to three experimental data sets. We demonstrate that both ELRT and FIT have power to detect selection in practical parameter regimes, such as those encountered in microbial evolution experiments. Our analysis applies to a single diallelic locus, assumed independent of all other loci, which is most relevant to full-genome selection scans in sexual organisms, and also to evolution experiments in asexual organisms as long as clonal interference is weak. Different techniques will be required to detect selection in time series of co-segregating linked loci.

preprint2013arXiv

The equilibrium allele frequency distribution for a population with reproductive skew

We study the population genetics of two neutral alleles under reversible mutation in the Λ-processes, a population model that features a skewed offspring distribution. We describe the shape of the equilibrium allele frequency distribution as a function of the model parameters. We show that the mutation rates can be uniquely identified from the equilibrium distribution, but that the form of the offspring distribution itself cannot be uniquely identified. We also introduce an infinite-sites version of the Λ-process, and we use it to study how reproductive skew influences standing genetic diversity in a population. We derive asymptotic formulae for the expected number of segregating sizes as a function of sample size. We find that the Wright-Fisher model minimizes the equilibrium genetic diversity, for a given mutation rate and variance effective population size, compared to all other Λ-processes.

preprint2013arXiv

The evolution of genetic architectures underlying quantitative traits

In the classic view introduced by R. A. Fisher, a quantitative trait is encoded by many loci with small, additive effects. Recent advances in QTL mapping have begun to elucidate the genetic architectures underlying vast numbers of phenotypes across diverse taxa, producing observations that sometimes contrast with Fisher's blueprint. Despite these considerable empirical efforts to map the genetic determinants of traits, it remains poorly understood how the genetic architecture of a trait should evolve, or how it depends on the selection pressures on the trait. Here we develop a simple, population-genetic model for the evolution of genetic architectures. Our model predicts that traits under moderate selection should be encoded by many loci with highly variable effects, whereas traits under either weak or strong selection should be encoded by relatively few loci. We compare these theoretical predictions to qualitative trends in the genetics of human traits, and to systematic data on the genetics of gene expression levels in yeast. Our analysis provides an evolutionary explanation for broad empirical patterns in the genetic basis of traits, and it introduces a single framework that unifies the diversity of observed genetic architectures, ranging from Mendelian to Fisherian.

preprint2013arXiv

The inevitability of unconditionally deleterious substitutions during adaptation

Studies on the genetics of adaptation typically neglect the possibility that a deleterious mutation might fix. Nonetheless, here we show that, in many regimes, the first substitution is most often deleterious, even when fitness is expected to increase in the long term. In particular, we prove that this phenomenon occurs under weak mutation for any house-of-cards model with an equilibrium distribution. We find that the same qualitative results hold under Fisher's geometric model. We also provide a simple intuition for the surprising prevalence of unconditionally deleterious substitutions during early adaptation. Importantly, the phenomenon we describe occurs on fitness landscapes without any local maxima and is therefore distinct from "valley-crossing". Our results imply that the common practice of ignoring deleterious substitutions leads to qualitatively incorrect predictions in many regimes. Our results also have implications for the substitution process at equilibrium and for the response to a sudden decrease in population size.

preprint2012arXiv

Epistasis not needed to explain low dN/dS

An important question in molecular evolution is whether an amino acid that occurs at a given position makes an independent contribution to fitness, or whether its effect depends on the state of other loci in the organism's genome, a phenomenon known as epistasis. In a recent letter to Nature, Breen et al. (2012) argued that epistasis must be "pervasive throughout protein evolution" because the observed ratio between the per-site rates of non-synonymous and synonymous substitutions (dN/dS) is much lower than would be expected in the absence of epistasis. However, when calculating the expected dN/dS ratio in the absence of epistasis, Breen et al. assumed that all amino acids observed in a protein alignment at any particular position have equal fitness. Here, we relax this unrealistic assumption and show that any dN/dS value can in principle be achieved at a site, without epistasis. Furthermore, for all nuclear and chloroplast genes in the Breen et al. dataset, we show that the observed dN/dS values and the observed patterns of amino acid diversity at each site are jointly consistent with a non-epistatic model of protein evolution.

preprint2012arXiv

Selection biases the prevalence and type of epistasis along adaptive trajectories

The contribution to an organism's phenotype from one genetic locus may depend upon the status of other loci. Such epistatic interactions among loci are now recognized as fundamental to shaping the process of adaptation in evolving populations. Although little is known about the structure of epistasis in most organisms, recent experiments with bacterial populations have concluded that antagonistic interactions abound and tend to de-accelerate the pace of adaptation over time. Here, we use a broad class of mathematical fitness landscapes to examine how natural selection biases the mutations that substitute during evolution based on their epistatic interactions. We find that, even when beneficial mutations are rare, these biases are strong and change substantially throughout the course of adaptation. In particular, epistasis is less prevalent than the neutral expectation early in adaptation and much more prevalent later, with a concomitant shift from predominantly antagonistic interactions early in adaptation to synergistic and sign epistasis later in adaptation. We observe the same patterns when re-analyzing data from a recent microbial evolution experiment. Since these biases depend on the population size and other parameters, they must be quantified before we can hope to use experimental data to infer an organism's underlying fitness landscape or to understand the role of epistasis in shaping its adaptation. In particular, we show that when the order of substitutions is not known to an experimentalist, then standard methods of analysis may suggest that epistasis retards adaptation when in fact it accelerates it.

preprint2012arXiv

The evolution of complex gene regulation by low specificity binding sites

Transcription factor binding sites vary in their specificity, both within and between species. Binding specificity has a strong impact on the evolution of gene expression, because it determines how easily regulatory interactions are gained and lost. Nevertheless, we have a relatively poor understanding of what evolutionary forces determine the specificity of binding sites. Here we address this question by studying regulatory modules composed of multiple binding sites. Using a population-genetic model, we show that more complex regulatory modules, composed of a greater number of binding sites, must employ binding sites that are individually less specific, compared to less complex regulatory modules. This effect is extremely general, and it hold regardless of the regulatory logic of a module. We attribute this phenomenon to the inability of stabilising selection to maintain highly specific sites in large regulatory modules. Our analysis helps to explain broad empirical trends in the yeast regulatory network: those genes with a greater number of transcriptional regulators feature by less specific binding sites, and there is less variance in their specificity, compared to genes with fewer regulators. Likewise, our results also help to explain the well-known trend towards lower specificity in the transcription factor binding sites of higher eukaryotes, which perform complex regulatory tasks, compared to prokaryotes.

preprint2011arXiv

The structure of allelic diversity in the presence of purifying selection

In the absence of selection, the structure of allelic diversity is described by the elegant sampling formula of Ewens. This formula has helped shape our expectations of empirical patterns of molecular variation. Along with coalescent theory, it provides statistical techniques for rejecting the null model of neutrality. However, we still do not fully understand the statistics of the allelic diversity we expect to see in the presence of natural selection. Earlier work has described the effects of strongly deleterious mutations linked to many neutral sites, and allelic variation in models where offspring fitness is unrelated to parental fitness, but it has proven difficult to understand allelic diversity in the presence of purifying selection at many linked sites. Here, we study the population genetics of infinitely many perfectly linked sites, some neutral and some deleterious. Our approach is based on studying the lineage structure within each class of individuals of similar fitness in the deleterious mutation-selection balance. Analogous to the Ewens sampling formula, we derive expressions for the likelihoods of any configuration of allelic types in a sample. We find that for moderate and weak selection pressures the patterns of allelic diversity cannot be described by a neutral model for any choice of the effective population size, indicating that there is power to detect selection from patterns of sampled allelic diversity.

preprint2011arXiv

The Structure of Genealogies in the Presence of Purifying Selection: A "Fitness-Class Coalescent"

Compared to a neutral model, purifying selection distorts the structure of genealogies and hence alters the patterns of sampled genetic variation. Although these distortions may be common in nature, our understanding of how we expect purifying selection to affect patterns of molecular variation remains incomplete. Genealogical approaches such as coalescent theory have proven difficult to generalize to situations involving selection at many linked sites, unless selection pressures are extremely strong. Here, we introduce an effective coalescent theory (a "fitness-class coalescent") to describe the structure of genealogies in the presence of purifying selection at many linked sites. We use this effective theory to calculate several simple statistics describing the expected patterns of variation in sequence data, both at the sites under selection and at linked neutral sites. Our analysis combines our earlier description of the allele frequency spectrum in the presence of purifying selection (Desai et al. 2010) with the structured coalescent approach of Nordborg (1997), to trace the ancestry of individuals through the distribution of fitnesses within the population. Alternatively, we can derive our results using an extension of the coalescent approach of Hudson and Kaplan (1994). We find that purifying selection leads to patterns of genetic variation that are related but not identical to a neutrally evolving population in which population size has varied in a specific way in the past.

preprint2009arXiv

On the accessibility of adaptive phenotypes of a bacterial metabolic network

The mechanisms by which adaptive phenotypes spread within an evolving population after their emergence are understood fairly well. Much less is known about the factors that influence the evolutionary accessibility of such phenotypes, a pre-requisite for their emergence in a population. Here, we investigate the influence of environmental quality on the accessibility of adaptive phenotypes of Escherichia coli's central metabolic network. We used an established flux-balance model of metabolism as the basis for a genotype-phenotype map (GPM). We quantified the effects of seven qualitatively different environments (corresponding to both carbohydrate and gluconeogenic metabolic substrates) on the structure of this GPM. We found that the GPM has a more rugged structure in qualitatively poorer environments, suggesting that adaptive phenotypes could be intrinsically less accessible in such environments. Nevertheless, on average ~74% of the genotype can be altered by neutral drift, in the environment where the GPM is most rugged; this could allow evolving populations to circumvent such ruggedness. Furthermore, we found that the normalized mutual information (NMI) of genotype differences relative to phenotype differences, which measures the GPM's capacity to transmit information about phenotype differences, is positively correlated with (simulation-based) estimates of the accessibility of adaptive phenotypes in different environments. These results are consistent with the predictions of a simple analytic theory and they suggest an intuitive information-theoretic principle for evolutionary adaptation; adaptation could be faster in environments where the GPM has a greater capacity to transmit information about phenotype differences.

preprint2004arXiv

Synonymous codon usage and selection on proteins

Selection pressures on proteins are usually measured by comparing homologous nucleotide sequences (Zuckerkandl and Pauling 1965). Recently we introduced a novel method, termed `volatility', to estimate selection pressures on protein sequences from their synonymous codon usage (Plotkin and Dushoff 2003, Plotkin et al 2004a). Here we provide a theoretical foundation for this approach. We derive the expected frequencies of synonymous codons as a function of the strength of selection, the mutation rate, and the effective population size. We analyze the conditions under which we can expect to draw inferences from biased codon usage, and we estimate the time scales required to establish and maintain such a signal. Our results indicate that, over a broad range of parameters, synonymous codon usage can reliably distinguish between negative selection, positive selection, and neutrality. While the power of volatility to detect negative selection depends on the population size, there is no such dependence for the detection of positive selection. Furthermore, we show that phenomena such as transient hyper-mutators in microbes can improve the power of volatility to detect negative selection, even when the typical observed neutral site heterozygosity is low.