We study evolutionary dynamics on graphs in which each step consists of one birth and one death, referred to generally as Moran processes. In standard simplified models, there are two types of individuals: residents, who have a fitness of 1, and mutants, who have a fitness of r. Two standard update rules are used in the literature. In Birth-death (Bd), a vertex is chosen to reproduce proportional to fitness, and one of its neighbors is selected uniformly at random to die and be replaced by the offspring. In death-Birth (dB), a vertex is chosen uniformly to die, and then one of its neighbors is chosen, proportional to fitness, to place an offspring into the vacancy. Two crucial quantities are: the unconditional absorption time, which is the expected time until only residents or only mutants remain, and the fixation probability of the mutant, which is the probability that at some time the mutants occupy the whole graph. Birth-death and death-Birth rules can yield significantly different outcomes for these quantities on the same graph, rendering conclusions dependent on the update rule. We formalize and study a unified model, the lambda-mixed Moran process, in which each step is independently a Bd step with probability lambda is an element of [0, 1] and a dB step otherwise. We analyze this mixed process and establish a few results that form a starting point for its further study. All of our results are for undirected, connected graphs. As an interesting special case, we show at lambda = 1/2 for any graph that the fixation probability when r = 1 with a single mutant initially on the graph is exactly 1/n, and also at lambda = 1/2 that the absorption time for any r is O-r (n(4)) (that is, with an r-dependent constant). We also show results for graphs that are "almost regular," in a manner defined in the paper. We use this to show that for suitable random graphs from G similar to G( n, p) and fixed r > 1, with high probability over the choice of graph, the absorption time is O-r (n(4)), the fixation probability is Omega(r) (n(-2)), and we can approximate the fixation probability in polynomial time. Another special case is when the graph has only two possible values for the degree {d(1), d(2)} with d(1) <= d(2). For those graphs, we give exact formulas for fixation probabilities under r = 1 and any lambda, and establish O-r (n(4)alpha(4)) absorption time regardless of lambda, where alpha = d(2)/d(1). We also provide explicit formulas for the star and cycle under any r or lambda. 2012 ACM Subject Classification Theory of computation; Theory of computation -> Randomness, geometry and discrete structures; Theory of computation -> Graph algorithms analysis; Theory of computation -> Approximation algorithms analysis
Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leading models produce sensible strategies in the repeated prisoner's dilemma, which is the main metaphor of reciprocal cooperation. First, we measure the propensity of LLMs to cooperate in a neutral setting, without using language reminiscent of how this game is usually presented. We record to what extent LLMs implement Nash equilibria or other well-known strategy classes. Thereafter, we explore how LLMs adapt their strategies to changes in parameter values. We vary the game's stopping probability, the payoff values, and if the total number of rounds is commonly known. We also study the effect of different framings. In each case, we test whether the adaptations of the LLMs are in line with basic intuition, theoretical predictions of evolutionary game theory, and experimental evidence of human participants. While all LLMs perform well in many of the tasks, none of them exhibit full consistency over all tasks. We also conduct tournaments between the inferred LLM strategies and study direct interaction between LLMs in games over ten rounds with known or unknown last round. Our experiments shed light on how current LLMs instantiate reciprocal cooperation.
Repeated games and stochastic games are important frameworks to study direct reciprocity. Individuals react strategically to their coplayers’ previous behavior. While strategies in such games can be arbitrarily complex, explorations of evolutionary dynamics are often done in specific strategy spaces. Individuals may consider a fixed number of past rounds, or only some of the partner’s previous actions. Such restrictions can make the interpretation of the results difficult. Strategies found to be superior within a restricted set may lose stability when more complex strategies are permitted. We discuss two notions of completeness that rule out this possibility. If a strategy space, S, is best-reply-complete, then any strategy in S is guaranteed to have a best reply in S. If a space, S, is payoff-complete, then any strategy playing against an opponent in S can be replaced by an equivalent strategy within S without affecting either player’s payoff. Sufficient conditions for best-reply-completeness have been given in a seminal paper by Levínský et al. Here, we show that for strategies of bounded memory, the same conditions are sufficient for payoff-completeness. Furthermore, using those conditions, we illustrate how to construct many complete spaces for simple games. Taken together, our findings highlight the importance of complete strategy spaces, which are particularly useful when interpreting evolutionary simulations and determining best responses.
Cancer-causing mutations have been identified primarily from positive selection signals in cancer genomes. However, positive selection is also a ubiquitous feature of normal tissue aging. Here we develop a statistical framework to disentangle selection in normal tissue and causation of carcinogenesis. By comparing cancer and normal tissue genomes, we estimate the effects of mutations on cancer risk in the blood, esophagus and colon. We determine that stronger cancer-causing mutations are enriched at younger patient ages. This enables cancer-causing mutations to be identified from patient age distributions, even without normal tissue data. Moreover, we show for acute myeloid leukemia that the age-dependence of purported causal mutations can be explained largely by normal blood evolution, challenging the long-standing notion that childhood cancers require distinct mutations. Broadly, our framework delineates carcinogenesis from normal tissue aging, improving the assessment of cancer risk conferred by mutations.
Direct reciprocity is a mechanism for evolution of cooperation based on repeated interactions between the same individuals. Direct reciprocity can help natural selection to favor cooperators over defectors, but whether or not cooperation prevails depends on the details of the evolutionary dynamics. We describe simple processes of evolution that have the astonishing ability of driving direct reciprocity to maximum payoff in all social dilemmas which we study. The basic process is based on mutation and pairwise comparison. Mutation samples strategies near the boundary of the strategy space. Pairwise comparison includes a parameter for intensity of selection. For large population sizes, intermediate to high mutation rates and intermediate to strong intensities of selection, we find that the process leads to communities of strategies that reach maximum payoff in Prisoner's Dilemma, Snowdrift, Stag Hunt and Harmony games. Maximum payoff in all four games is consistently achieved if players have access to memory-2 strategies. Memory-1 strategies have the capacity to resolve all four social dilemmas, but they are usually defeated by a “Hold-trap” in Snowdrift games.
Understanding lung cancer evolution can identify tools for intercepting its growth. In a landscape analysis of 1024 lung adenocarcinomas (LUAD) with deep whole-genome sequencing integrated with multiomic data, we identified 542 LUAD that displayed diverse clonal architecture. In this group, we observed an interplay between mobile elements, endogenous and exogenous mutational processes, distinct driver genes, and epidemiological features. Our results revealed divergent evolutionary trajectories based on tobacco smoking exposure, ancestry, and sex. LUAD from smokers showed an abundance of tobacco-related C:G>A:T driver mutations in KRAS plus short subclonal diversification. LUAD in never smokers showed early occurrence of copy number alterations and EGFR mutations associated with SBS5 and SBS40a mutational signatures. Tumors harboring EGFR mutations exhibited long latency, particularly in females of European-ancestry (EU_N). In EU_N, EGFR mutations preceded the occurrence of other driver genes, including TP53 and RBM10. Tumors from Asian never smokers showed a short clonal evolution and presented with heterogeneous repetitive patterns for the inferred mutational order. Importantly, we found that the mutational signature ID2 is a marker of a previously unrecognized mechanism for LUAD evolution. Tumors with ID2 showed short latency and high L1 retrotransposon activity linked to L1 promoter demethylation. These tumors exhibited an aggressive phenotype, characterized by increased genomic instability, elevated hypoxia scores, low burden of neoantigens, propensity to develop metastasis, and poor overall survival. Reactivated L1 retrotransposition-induced mutagenesis can contribute to the origin of the mutational signature ID2, including through the regulation of the transcriptional factor ZNF695, a member of the KZFP family. The complex nature of LUAD evolution creates both challenges and opportunities for screening and treatment plans.
Evolution occurs in populations of reproducing individuals. In stochastic descriptions of evolutionary dynamics, such as the Moran process, individuals are chosen randomly for birth and for death. If the same type is chosen for both steps, then the reproductive event is wasted, because the composition of the population remains unchanged. Here we introduce a new phenotype, which we call a replacer. Replacers are efficient competitors. When a replacer is chosen for reproduction, the offspring will always replace an individual of another type (if available). We determine the selective advantage of replacers in well-mixed populations and on one-dimensional lattices. We find that being a replacer substantially boosts the fixation probability of neutral and deleterious mutants. In particular, fixation probability of a single neutral replacer who invades a well-mixed population of size N is of the order of 1/√(N) rather than the standard 1/N. Even more importantly, replacers are much better protected against invasions once they have reached fixation. Therefore, replacers dominate the mutation selection equilibrium even if the phenotype of being a replacer comes at a substantial cost: curiously, for large population size and small mutation rate the relative reproductive rate of a successful replacer can be as low as 1/e.
We examine population structures for their ability to maintain diversity in neutral evolution. We use the general framework of evolutionary graph theory and consider birth-death (bd) and death-birth (db) updating. The population is of size N. Initially all individuals represent different types. The basic question is: what is the time T_N until one type takes over the population? This time is known as consensus time in computer science and as total coalescent time in evolutionary biology. For the complete graph, it is known that T_N is quadratic in N for db and bd. For the cycle, we prove that T_N is cubic in N for db and bd. For the star, we prove that T_N is cubic for bd and quasilinear (Nlog N) for db. For the double star, we show that T_N is quartic for bd. We derive upper and lower bounds for all undirected graphs for bd and db. We also show the Pareto front of graphs (of size N=8) that maintain diversity the longest for bd and db. Further, we show that some graphs that quickly homogenize can maintain high levels of diversity longer than graphs that slowly homogenize. For directed graphs, we give simple contracting star-like structures that have superexponential time scales for maintaining diversity.
Elucidating the evolution of cancers allows us to understand their key events, and the order in which they occur. To chart and interpret these evolutionary trajectories, we leverage whole-genome sequencing of lung tumours, including those from the largest cohort to date of lung cancers in subjects who have never smoked. Through ordering frequent genomic alterations, we discover three distinct evolutionary paths taken by lung adenocarcinomas; two dominated by tumours from people who have never smoked (NS-LUAD), and one followed by the vast majority of those who have smoked (S-LUAD). However, one in six NS-LUAD follow the smoking-dominant trajectory. These tumours, surprisingly, have fewer somatic alterations than the other NS-LUAD, and have shorter latency. They are strongly enriched for KRAS mutations. Our results suggest that gaining KRAS mutations allows these tumours to evolve more rapidly, acquiring a set of smoking-associated key alterations, with less need for genomic instability to progress. These tumours are three times more frequent in subjects of European vs. East Asian ancestry. These findings could shape clinical management strategies for lung adenocarcinoma patients, particularly for tumours driven by smoking-like evolutionary trajectories.
Lung cancer in never smokers (LCINS) accounts for around 25% of all lung cancers1,2 and has been associated with exposure to second-hand tobacco smoke and air pollution in observational studies3-5. Here we use data from the Sherlock-Lung study to evaluate mutagenic exposures in LCINS by examining the cancer genomes of 871 treatment-naive individuals with lung cancer who had never smoked, from 28 geographical locations. KRAS mutations were 3.8 times more common in adenocarcinomas of never smokers from North America and Europe than in those from East Asia, whereas a higher prevalence of EGFR and TP53 mutations was observed in adenocarcinomas of never smokers from East Asia. Signature SBS40a, with unknown cause6, contributed the largest proportion of single base substitutions in adenocarcinomas, and was enriched in cases with EGFR mutations. Signature SBS22a, which is associated with exposure to aristolochic acid7,8, was observed almost exclusively in patients from Taiwan. Exposure to secondhand smoke was not associated with individual driver mutations or mutational signatures. By contrast, patients from regions with high levels of air pollution were more likely to have TP53 mutations and shorter telomeres. They also exhibited an increase in most types of mutations, including a 3.9-fold increase in signature SBS4, which has previously been linked with tobacco smoking9, and a 76% increase in the clock-like10 signature SBS5. A positive dose-response effect was observed with air-pollution levels, correlating with both a decrease in telomere length and an increase in somatic mutations, mainly attributed to signatures SBS4 and SBS5. Our results elucidate the diversity of mutational processes shaping the genomic landscape of lung cancer in never smokers.
In spite of the growing interest in the microbiome in human cancer, there are currently only small-scale lung cancer microbiome studies conducted directly on tissue. As part of the Sherlock-Lung study, we studied the microbiomes of 940 lung cancers (4090 samples) in never smokers (LCINS) directly from lung tissue using three data types: 16S rRNA gene sequencing (16S), whole-genome sequencing (WGS) with paired blood, and RNA-seq. We observe very low biomass and few microbiome associations in LCINS using 16S and WGS tissue. Using RNA-seq, we observe more total microbial reads, and decreased relative abundance of several commensal bacteria at the genus and species levels in tumors relative to paired normal lung tissue. Among all datasets, we see no consistent associations between the lung tissue microbiome, or circulating bacterial DNA, and any available demographic and clinical features, including age, sex, genetic ancestry, second-hand tobacco smoking exposure, LCINS histology, stage, and overall survival. We also observe no microbiome associations with any human genomic alterations within the same samples. Every null result should be interpreted with caution given the possibility of future methodological breakthroughs. However, all together, using multiple data types in nearly 1000 patients, we find no substantive role for the lung cancer microbiome in treatment-naïve LCINS.
Understanding lung cancer evolution can identify tools for intercepting its growth(1,2). Here, in a landscape analysis of 1,024 lung adenocarcinomas (LUADs) with deep whole-genome sequencing integrated with multiomic data, we identified 542 LUADs with a diverse clonal architecture. In this group, we observed divergent evolutionary trajectories based on tobacco smoking exposure, ancestry and sex. LUAD from smokers showed an abundance of tobacco-related C:G>A:T driver mutations(3) in KRAS and short subclonal diversification. LUAD in people who have never smoked (hereafter, never-smokers) showed early occurrence of copy-number alterations and EGFR mutations associated with SBS5 and SBS40a mutational signatures. Tumours containing EGFR mutations exhibited long latency, particularly in female individuals of European-ancestry. Tumours from Asian never-smokers showed a short clonal evolution. Importantly, we found that the mutational signature ID2(4) is a marker of a previously unrecognized mechanism for LUAD evolution. Tumours with ID2 showed short latency and high long interspersed nuclear element-1 (LINE-1, hereafter L1) retrotransposon activity linked to L1 promoter demethylation. These tumours exhibited an aggressive phenotype with genomic instability, elevated hypoxia scores, low neoantigen burden, metastasis propensity and poor overall survival. Reactivated L1-retrotransposition-induced mutagenesis probably contributes to the mutational signature ID2, including through the regulation of the transcriptional factor ZNF695, a member of the KZFP family(5). The complex nature of LUAD evolution creates both challenges and opportunities for screening and treatment plans.
Population-suppressing gene drives may be capable of extinguishing wild populations, with proposed applications in conservation, agriculture, and public health. However, unintended and potentially disastrous consequences of release of drive-engineered individuals are extremely difficult to predict. We propose a model for the dynamics of a sex ratio-biasing drive, and using simulations, we show that failure of the suppression drive is often a natural outcome due to stochastic and spatial effects. We further demonstrate rock-paper-scissors dynamics among wild-type, drive-infected, and extinct populations that can persist for arbitrarily long times. Gene drive-mediated extinction of wild populations entails critical complications that lurk far beyond the reach of laboratory-based studies. Our findings help in addressing these challenges.
ABSTRACTLung cancer in never smokers (LCINS) accounts for up to 25% of all lung cancers and has been associated with exposure to secondhand tobacco smoke and air pollution in observational studies. Here, we evaluate the mutagenic exposures in LCINS by examining deep whole-genome sequencing data from a large international cohort of 871 treatment-naïve LCINS recruited from 28 geographical locations within the Sherlock-Lungstudy.KRASmutations were 3.8-fold more common in adenocarcinomas of never smokers from North America and Europe, while a 1.6-fold higher prevalence ofEGFRandTP53mutations was observed in adenocarcinomas from East Asia. Signature SBS40a, with unknown cause, was found in most samples and accounted for the largest proportion of single base substitutions in adenocarcinomas, being enriched inEGFR-mutated cases. Conversely, the aristolochic acid signature SBS22a was almost exclusively observed in patients from Taipei. Even though LCINS exposed to secondhand smoke had an 8.3% higher mutational burden and 5.4% shorter telomeres, passive smoking was not associated with driver mutations in cancer driver genes or the activities of individual mutational signatures. In contrast, patients from regions with high levels of air pollution were more likely to haveTP53mutations while exhibiting shorter telomeres and an increase in most types of somatic mutations, including a 3.9-fold elevation of signature SBS4 (q-value=3.1 × 10−5), previously linked mainly to tobacco smoking, and a 76% increase of clock-like signature SBS5 (q-value=5.0 × 10−5). A positive dose-response effect was observed with air pollution levels, which correlated with both a decrease in telomere length and an elevation in somatic mutations, notably attributed to signatures SBS4 and SBS5. Our results elucidate the diversity of mutational processes shaping the genomic landscape of lung cancer in never smokers.
Some drugs increase the mutation rate of their target pathogen, a potentially concerning mechanism as the pathogen might evolve faster toward an undesired phenotype. We suggest a four-step assessment of evolutionary safety for the approval of such treatments.
Computing the rate of evolution in spatially structured populations is difficult. A key quantity is the fixation time of a single mutant with relative reproduction rate r which invades a population of residents. We say that the fixation time is "fast" if it is at most a polynomial function in terms of the population size N. Here we study fixation times of advantageous mutants (r>1) and neutral mutants (r=1) on directed graphs, which are those graphs that have at least some one-way connections. We obtain three main results. First, we prove that for any directed graph the fixation time is fast, provided that r is sufficiently large. Second, we construct an efficient algorithm that gives an upper bound for the fixation time for any graph and any r≥ 1. Third, we identify a broad class of directed graphs with fast fixation times for any r≥ 1. This class includes previously studied amplifiers of selection, such as Superstars and Metafunnels. We also show that on some graphs the fixation time is not a monotonically declining function of r; in particular, neutral fixation can occur faster than fixation for small selective advantages.
Abstract Lung cancer is the leading cause of cancer-related mortality worldwide. Understanding lung cancer evolutionary dynamics can help identify tools to intercept its growth and suggest strategies for treatment. Multiple factors can impact the tumors’ natural history and distinctly affect growth rate. However, research on the evolutionary trajectories of lung cancer across demographic or exposure scenarios remains inadequate. Additionally, the roles of mutational processes and complex genomic alterations on the evolution of lung cancer are still largely unexplored. In the largest genomic study of lung cancer to date, we analyzed deep whole-genome sequencing (~ 81x) and other omics data of 1217 lung cancers from the Sherlock-Lung study. To ensure adequate statistical power for identifying subclone architectures and constructing lung cancer evolutionary histories, we utilized a metric known as NRPCC (number of reads per tumor chromosomal copy) to select 542 lung adenocarcinoma (LUAD) samples for clonal evolution analyses, including 186 and 181 samples from never-smoker subjects of European and Asian ancestry, respectively, and 121 samples from smokers of European ancestry. We found that major driver genes and exogenous mutations contribute to tumor initiation, while copy number gains and endogenous processes appear later in tumor evolution. Tumors harboring EGFR mutations in never-smoker females of European descent show long latency, while tumors with KRAS mutations have shorter latency regardless of ancestry and sex. Notably, tumors harboring the mutational signature ID2 have short latency and aggressive phenotype, accompanied by increased genomic instability, elevated hypoxia scores, high CpG methylator phenotype, low neoantigen burden, and propensity to develop metastasis. We show that LINE-1 retrotransposition-induced mutagenesis contributes to the origin of ID2 mutations. The transcriptional factor ZNF695, a member of the KZFP family, up-regulated in LUAD, appears to contribute to LINE-1 retrotransposition through a dominant-negative effect and LINE-1 promoter demethylation. In a multivariate analysis of genomics, exposures and demographic factors, LUAD latency was most significantly associated with ID2, followed by EGFR mutations, KRAS mutations, and sex, suggesting an independent impact of these factors on LUAD evolution. Our findings underscore the complex interplay of ancestry, sex, exogenous mutagenesis, epigenetic regulation, and LINE-1 retrotransposition in shaping LUAD evolutionary trajectories, paving the way for potential targeted therapeutic interventions. Citation Format: Tongwu Zhang, Wei Zhao, Christopher Wirth, Marcos Díaz-Gay, Jinhu Yin, Phuc H. Hoang, Jian Sang, John McElderry, Alyssa Klein, Azhar Khandekar, Caleb Hartman, Jennifer Rosenbaum, Frank Colon-Matos, Kristine M. Jones, Neil E. Caporaso, Robert Homer, Angela C. Pesatori, Dario Consonni, Lixing Yang, Bin Zhu, Jianxin Shi, Kevin Brown, Nathaniel Rothman, Stephen J. Chanock, Ludmil B. Alexandrov, Jiyeon Choi, Maurizio Cardelli, Qing Lan, Martin A. Nowak, David C. Wedge, Maria Teresa Landi. Deciphering lung adenocarcinoma evolution and the role of LINE-1 retrotransposition [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(7_Suppl):Abstract nr LB226.