Abstract Cancer treatment failure is often attributed to two processes: selection for pre-existing resistant cells and cellular plasticity that allows drug-sensitive cells to become resistant. Plasticity is a hallmark of cancer, but we still lack a clear definition and a mechanistic understanding of how it is controlled in tumors. These gaps are especially problematic in pancreatic ductal adenocarcinoma (PDAC), where epithelial-to-mesenchymal plasticity (EMP) is thought to drive metastasis, tumor initiation, and resistance to therapy. Causally dissecting plasticity requires large-scale perturbation studies that systematically target many genes and measure how they reshape cell-state dynamics at single-cell resolution over weeks to months. Such studies must follow enough cells to observe rare state transitions and clonal heterogeneity, yet current methods force a tradeoff between (i) sufficient molecular phenotyping, (ii) temporal duration, and (iii) cellular throughput, limiting our ability to map regulators of plasticity at scale. We set out to address both issues by (1) developing a screening platform that overcomes major limitations in existing paradigms to (2) enable discovery of the mechanistic underpinnings of cellular plasticity in pancreatic cancer. Using an engineered patient-derived PDAC model, we performed a paired single-cell CRISPRi transcriptomic screen (Perturb-seq) and high-throughput imaging-based optical pooled screen (OPS) of EMP across a panel of chromatin and transcriptional regulators. We identified factors responsible for distinct plastic behaviors by quantifying state transitions from a FACS-defined initial state linked to high-content imaging phenotypes for 150,000 clones, across 1000 genetic perturbations and 22 million cells after two weeks of sustained perturbation in culture. Crucially, this framework sorted regulators into distinct functional classes of plasticity control, including “maintenance” genes stabilizing the starting state, “transition” genes that biased state switching in one direction, and “catalytic” genes that regulated switching bidirectionally. Our methodology recapitulated many genes previously implicated in EMP and uncovered novel candidate “catalytic” targets that converge on H3K9me3-associated epigenetic reprogramming. This work clarifies how cell state is regulated at clonal resolution, identifies potential vulnerabilities in plastic tumor cells, and establishes a platform that can be extended to studies of cellular plasticity in other biological contexts. Citation Format: Mikolaj Godzik, Russell Walton, Michael Bogaev, Lynn Bi, Julien Dilly, Martin Jankowiak, Elisa Donnard, David T. Ting, Nir Hacohen, Eric S. Lander, Paul Blainey, Arnav Mehta. Mechanistic dissection of regulators of cancer plasticity using high-content optical pooled screens at clonal resolution [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 6798.
Ancient DNA has transformed our understanding of population history1, but its potential to reveal as much about human evolutionary biology has not been realized because of limited sample sizes and the difficulty of distinguishing sustained rises in allele frequency increasing fitness—directional selection—from shifts due to migrations, population structure, or non-adaptive purifying or stabilizing selection2–7. Here we present a method for detecting directional selection in ancient DNA time-series data that tests for consistent trends in allele frequency change over time, and apply it to 15,836 West Eurasians (10,016 with new data). Previous work has shown that classic hard sweeps driving advantageous mutations to fixation have been rare over the broad span of human evolution8,9. By contrast, in the past ten millennia, we find that many hundreds of alleles have been affected by strong directional selection. We also document one-standard-deviation changes on the scale of modern variation in combinations of alleles that today predict complex traits. This includes decreases in predicted body fat and schizophrenia, and increases in measures of cognitive performance. These effects were measured in industrialized societies, and it remains unclear how these relate to phenotypes that were adaptive in the past. We estimate selection coefficients at 9.7 million variants, enabling study of how Darwinian forces couple to allelic effects and shape the genetic architecture of complex traits. Analysis of 15,836 ancient West Eurasian genomes reveals hundreds of instances of directional selection, showing that sustained changes in allele frequency were widespread, rather than being rare over this period as previously assumed.
Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics1-5. Most of these variants individually have weak effects6 and lie in non-coding gene-regulatory elements7-10, for which we lack a complete understanding of how single-nucleotide alterations modulate transcriptional processes to affect human phenotypes5,11-15. To address this problem, we measured the activity of 221,412 fine-mapped trait-associated variants using a massively parallel reporter assay16-20 in 5 diverse cell types. We show that this assay effectively discriminates between likely causal variants and controls, and identified 13,121 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% of them can plausibly be explained by the disruption of a known transcription factor binding motif. We investigated the mechanisms of 136 variants using saturation mutagenesis and assigned affected transcription factors for 91% of variants without a clear canonical mechanism. Finally, we detected regulatory epistasis at 11% of tested regulatory variants in close proximity and identified multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants that underlie complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.
Severe proteinopathies-such as retinitis pigmentosa, a form of inherited blindness-are driven by genetic mutations that overwhelm the quality control of the post-endoplasmic reticulum (post-ER) secretory pathway, causing toxic protein accumulation. Here, we identify a therapeutic node defined by a hetero-oligomeric cargo receptor complex consisting of TMED7, 2, 9, and 10. This "entrapment complex" anchors structurally and functionally diverse mutant clients within the early secretory pathway via TMED7 binding to the integral Golgi protein GRASP55. Disruption of the entrapment complex results in the clearance of accumulated protein cargoes. In vivo ablation of the entrapment node via inducible genetic deletion or via the small molecule BRD7635 reverses histopathological hallmarks and rescues functional deficits in clinically distinct proteinopathies of the kidney and the eye, including mitigating vision loss in a mouse model of retinitis pigmentosa.
KRAS is a major oncogenic driver in pancreatic ductal adenocarcinoma (PDAC), mutationally activated in approximately 90% of cases. Mutations in this oncogene have been associated with more aggressive disease, poorer outcomes, and have remained hard to drug for over three decades. Recently developed small molecule inhibitors of KRAS (KRASi) have shown promising efficacy in advanced, previously treated PDAC patients but ultimately, most patients develop acquired resistance. Approximately half of these resistance cases do not present a putative genomic driver. In preclinical studies using genetically engineered mouse models of PDAC, we and others have demonstrated that cell state identity along the classical/epithelial-mesenchymal axis is a key determinant of response and resistance to KRASi. Notably, acute KRASi treatment initially induces a strong and selective bottlenecking of the malignant population in vivo, resulting in tumors that are predominantly classical with depleted mesenchymal characteristics. While these findings highlight a potent cell state-selective effect of acute KRASi treatment on PDAC cells, the underlying mechanisms driving the drug-induced cell state selection and plasticity remain poorly characterized. These mechanisms may reveal novel therapeutic targets for developing combination therapies that improve therapeutic responses. To comprehensively map and modulate epithelial-mesenchymal (E-M) cell state plasticity in response to KRASi treatment, we performed CRISPRi Perturb-seq (single-cell gene expression readout) with lineage tracing on a patient-derived PDAC cell line treated with the RAS(ON) multi-selective inhibitor RMC-7977. We captured 734,092 high-quality single-cell transcriptomic profiles and 156,119 unique clones across 60 genetic perturbations, including E- and M-specific transcription factors (TF) inferred to influence E-M plasticity from previous single-cell lineage tracing experiments. To optimally model clonal growth and transition rates, we profiled cells across a time series at day 6, day 15 and day 22, after 1 week of DMSO or KRASi treatment. We developed a hierarchical generative probabilistic model to jointly infer E↔M transition rates and cell state-specific growth rates, while capturing the effects of perturbations and KRASi treatment on state dynamics. Analysis of longitudinally tracked clones revealed that KRASi treatment specifically reduced the M-state growth rate while increasing the probability of M→E transitions compared to DMSO. Notably, we identified several TFs, displaying KRASi-dependent changes in state transition dynamics, either promoting (ELF3, MEIS2 among others) or decreasing (ZNF281, IRF9, among others) probabilities of M→E transitions when compared to non-targeting controls. These preclinical findings suggest that perturbation of these factors may be used as a paradigm to prevent cell state plasticity upon acute KRASi treatment and could be used to homogenize tumor populations towards a treatment-sensitive state. Julien Dilly, Mike Bogaev, Lynn Bi, Abigail Collins, Martin Jankowiak, Aziz Al'Khafaji, Kyle E. Evans, Mehrtash Babadi, Thouis R. Jones, Elisa Donnard, David T. Ting, Nir Hacohen, Dana Pe'er, Eric S. Lander, Andrew J. Aguirre, Arnav Mehta. Mapping and modulating epithelial-mesenchymal plasticity under RAS(ON) multi-selective inhibition in PDAC through lineage tracing and Perturb-seq [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Pancreatic Cancer Research—Emerging Science Driving Transformative Solutions; Boston, MA; 2025 Sep 28-Oct 1; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(18_Suppl_3):Abstract nr B120.
Regulatory DNA provides a platform for transcription factor binding to encode cell-type-specific patterns of gene expression. However, the effects and programmability of regulatory DNA sequences remain difficult to map or predict. Here, we develop variant effects from flow-sorting experiments with CRISPR targeting screens (Variant-EFFECTS) to introduce hundreds of designed edits to endogenous regulatory DNA and quantify their effects on gene expression. We systematically dissect and reprogram 3 regulatory elements for 2 genes in 2 cell types. These data reveal endogenous binding sites with effects specific to genomic context, transcription factor motifs with cell-type-specific activities, and limitations of computational models for predicting the effect sizes of variants. We identify small edits that can tune gene expression over a large dynamic range, suggesting new possibilities for prime-editing-based therapeutics targeting regulatory DNA. Variant-EFFECTS provides a generalizable tool to dissect regulatory DNA and to identify genome editing reagents that tune gene expression in an endogenous context.
Analyses of ancient DNA typically involve sequencing the surviving short oligonucleotides and aligning to genome assemblies from related, modern species. Here, we report that skin from a female woolly mammoth (†Mammuthus primigenius) that died 52,000 years ago retained its ancient genome architecture. We use PaleoHi-C to map chromatin contacts and assemble its genome, yielding 28 chromosome-length scaffolds. Chromosome territories, compartments, loops, Barr bodies, and inactive X chromosome (Xi) superdomains persist. The active and inactive genome compartments in mammoth skin more closely resemble Asian elephant skin than other elephant tissues. Our analyses uncover new biology. Differences in compartmentalization reveal genes whose transcription was potentially altered in mammoths vs. elephants. Mammoth Xi has a tetradic architecture, not bipartite like human and mouse. We hypothesize that, shortly after this mammoth’s death, the sample spontaneously freeze-dried in the Siberian cold, leading to a glass transition that preserved subfossils of ancient chromosomes at nanometer scale.
Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics. The majority of these variants have individually weak effects and lie in non-coding gene-regulatory elements where we lack a complete understanding of how single nucleotide alterations modulate transcriptional processes to affect human phenotypes. To address this, we measured the activity of 221,412 trait-associated variants that had been statistically fine-mapped using a Massively Parallel Reporter Assay (MPRA) in 5 diverse cell-types. We show that MPRA is able to discriminate between likely causal variants and controls, identifying 12,025 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% can plausibly be explained by the disruption of a known transcription factor (TF) binding motif. We dissect the mechanisms of 136 variants using saturation mutagenesis and assign impacted TFs for 91% of variants without a clear canonical mechanism. Finally, we provide evidence that epistasis is prevalent for variants in close proximity and identify multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants underlying complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.
The time required to conduct clinical trials limits the rate at which we can evaluate and deliver new treatment options to patients with cancer. New approaches to increase trial efficiency while maintaining rigor would benefit patients, especially in oncology, in which adjuvant trials hold promise for intercepting metastatic disease, but typically require large numbers of patients and many years to complete. We envision a standing platform - an infrastructure to support ongoing identification and trial enrolment of patients with cancer with early molecular evidence of disease (MED) after curative-intent therapy for early-stage cancer, based on the presence of circulating tumour DNA. MED strongly predicts subsequent recurrence, with the vast majority of patients showing radiographic evidence of disease within 18 months. Such a platform would allow efficient testing of many treatments, from small exploratory studies to larger pivotal trials. Trials enrolling patients with MED but without radiographic evidence of disease have the potential to advance drug evaluation because they can be smaller (given high probability of recurrence) and faster (given short time to recurrence) than conventional adjuvant trials. Circulating tumour DNA may also provide a valuable early biomarker of treatment effect, which would allow small signal-finding trials. In this Perspective, we discuss how such a platform could be established.
Infection with Lassa virus (LASV) can cause Lassa fever, a haemorrhagic illness with an estimated fatality rate of 29.7%, but causes no or mild symptoms in many individuals. Here, to investigate whether human genetic variation underlies the heterogeneity of LASV infection, we carried out genome-wide association studies (GWAS) as well as seroprevalence surveys, human leukocyte antigen typing and high-throughput variant functional characterization assays. We analysed Lassa fever susceptibility and fatal outcomes in 533 cases of Lassa fever and 1,986 population controls recruited over a 7 year period in Nigeria and Sierra Leone. We detected genome-wide significant variant associations with Lassa fever fatal outcomes near GRM7 and LIF in the Nigerian cohort. We also show that a haplotype bearing signatures of positive selection and overlapping LARGE1, a required LASV entry factor, is associated with decreased risk of Lassa fever in the Nigerian cohort but not in the Sierra Leone cohort. Overall, we identified variants and genes that may impact the risk of severe Lassa fever, demonstrating how GWAS can provide insight into viral pathogenesis. GWAS in difficult-to-recruit populations identifies variants associated with Lassa fever outcome and susceptibility at loci proximal to LIF, GRM7 and LARGE1.
Most phenotype-associated genetic variants map to noncoding regulatory regions of the human genome, but their mechanisms remain elusive in most cases. We developed a highly efficient strategy, Perturb-multiome, to simultaneously profile chromatin accessibility and gene expression in single cells with CRISPR-mediated perturbation of master transcription factors (TFs). We examined the connection between TFs, accessible regions, and gene expression across the genome throughout hematopoietic differentiation. We discovered that variants within TF-sensitive accessible chromatin regions in erythroid differentiation, although representing <0.3% of the genome, show a ~100-fold enrichment for blood cell phenotype heritability, which is substantially higher than that for other accessible chromatin regions. Our approach facilitates large-scale mechanistic understanding of phenotype-associated genetic variants by connecting key cis-regulatory elements and their target genes within gene regulatory networks.
Linking variants from genome-wide association studies (GWAS) to underlying mechanisms of disease remains a challenge1,4,6. For some diseases, a successful strategy has been to look for cases where multiple GWAS loci contain genes that act in the same biological pathway1–6. However, our knowledge of which genes act in which pathways is incomplete, particularly for cell-type specific pathways or understudied genes. Here we introduce a method to connect GWAS variants to functions, which links variants to genes using epigenomic data, links genes to pathways de novo using Perturb-seq, and integrates these data to identify convergence of GWAS loci onto pathways. We apply this approach to study the role of endothelial cells in genetic risk for coronary artery disease (CAD), and discover that 43 CAD GWAS signals converge on the cerebral cavernous malformations (CCM) signaling pathway. Two regulators of this pathway, CCM2 and TLNRD1,are each linked to a CAD risk variant, regulate other CAD risk genes, and affect atheroprotective processes in endothelial cells. These results suggest a model where CAD risk is driven in part by the convergence of causal genes onto a particular transcriptional pathway in endothelial cells, highlight shared genes between common and rare vascular diseases (CAD and CCM), and identify TLNRD1 as a new, previously uncharacterized member of the CCM signaling pathway. This approach will be widely useful for linking variants to functions for other common polygenic diseases. Note: The list of authors for this protocol does not include all authors of the accompanying manuscript, only those who played a role in developing and executing the Perturb-seq method. For a complete list of manuscript authors, see the Manuscript Citation. Notes are provided to indicate which authors are best to contact for questions regarding specific methods.
Enhancers are key drivers of gene regulation thought to act via 3D physical interactions with the promoters of their target genes. However, genome-wide depletions of architectural proteins such as cohesin result in only limited changes in gene expression, despite a loss of contact domains and loops. Consequently, the role of cohesin and 3D contacts in enhancer function remains debated. Here, we developed CRISPRi of regulatory elements upon degron operation (CRUDO), a novel approach to measure how changes in contact frequency impact enhancer effects on target genes by perturbing enhancers with CRISPRi and measuring gene expression in the presence or absence of cohesin. We systematically perturbed all 1,039 candidate enhancers near five cohesin-dependent genes and identified 34 enhancer-gene regulatory interactions. Of 26 regulatory interactions with sufficient statistical power to evaluate cohesin dependence, 18 show cohesin-dependent effects. A decrease in enhancer-promoter contact frequency upon removal of cohesin is frequently accompanied by a decrease in the regulatory effect of the enhancer on gene expression, consistent with a contact-based model for enhancer function. However, changes in contact frequency and regulatory effects on gene expression vary as a function of distance, with distal enhancers (e.g., >50Kb) experiencing much larger changes than proximal ones (e.g., <50Kb). Because most enhancers are located close to their target genes, these observations can explain how only a small subset of genes - those with strong distal enhancers - are sensitive to cohesin. Together, our results illuminate how 3D contacts, influenced by both cohesin and genomic distance, tune enhancer effects on gene expression.
Linking variants from genome-wide association studies (GWAS) to underlying mechanisms of disease remains a challenge1-3. For some diseases, a successful strategy has been to look for cases in which multiple GWAS loci contain genes that act in the same biological pathway1-6. However, our knowledge of which genes act in which pathways is incomplete, particularly for cell-type-specific pathways or understudied genes. Here we introduce a method to connect GWAS variants to functions. This method links variants to genes using epigenomics data, links genes to pathways de novo using Perturb-seq and integrates these data to identify convergence of GWAS loci onto pathways. We apply this approach to study the role of endothelial cells in genetic risk for coronary artery disease (CAD), and discover 43 CAD GWAS signals that converge on the cerebral cavernous malformation (CCM) signalling pathway. Two regulators of this pathway, CCM2 and TLNRD1, are each linked to a CAD risk variant, regulate other CAD risk genes and affect atheroprotective processes in endothelial cells. These results suggest a model whereby CAD risk is driven in part by the convergence of causal genes onto a particular transcriptional pathway in endothelial cells. They highlight shared genes between common and rare vascular diseases (CAD and CCM), and identify TLNRD1 as a new, previously uncharacterized member of the CCM signalling pathway. This approach will be widely useful for linking variants to functions for other common polygenic diseases.
Abstract Epithelial-mesenchymal transition (EMT) is a hallmark of pancreatic ductal adenocarcinoma (PDAC) invasion. In general, cellular plasticity is a major contributor to both tumor progression and therapy resistance. The major goals of our study are to 1) directly and quantitatively measure plasticity in PDAC3 cell lines, and 2) identify candidate genes that could alter the plasticity of cancer cells. By understanding the regulation of tumor cell plasticity, we may find ways to perturb cells and prevent the transition towards treatment resistant cell states. A strict definition of plasticity in cancer has yet to be established, however, and is imperative to our understanding of cell state regulation and for identifying vulnerabilities in plastic cells. We argue that in order to prove plasticity occurs within a population of cells, one must either 1) demonstrate that a new cell state has emerged that was previously not present, or 2) trace the lineage of a cell to show it can adopt multiple cell states. For our study, we profile 12 patient-derived PDAC cell lines to find convergent EMT programs and perform lineage tracing on these cell lines using a modified CROP-seq vector, called ClonMapper, to quantitatively measure plasticity. Towards this goal, we perform single-cell RNA-seq (scRNA) and single-cell multiomic measurements of 1000 uniquely barcoded cells expanded over a time course of 2, 3 and 4 weeks. We refer to cells that have the same lineage barcode - originating from the same initial cell - as families. We found that ~50% of families are composed of cells that are epithelial and mesenchymal, thus proving plasticity must exist, as it shows that a given parental cell of unknown state can give rise to progeny that are in multiple states. We then develop mathematical models to show that families have different levels of plasticity, defined by the rates of transition between distinct cell states. We next leverage our barcode-resolution time course data to identify factors at early time points that may influence plasticity at later time points. We show there is differential accessibility of several of these factors across families with different plasticity properties, thus proposing a role for epigenetic regulation of EMT. Importantly, we find that specific factors drive gene regulatory networks important for epithelial and mesenchymal programs, and therefore propose candidates for the modulation of EMT for therapeutic benefit. Citation Format: Deepika Yeramosu, Lynn Bi, Abigail Collins, Mike Bogaev, Martin Jankowiak, Aziz Al’Khafaji, Milan Parikh, Mehrtash Babadi, Alex Bloemendal, Surya Nagaraja, Ray Jones, Jay Shendure, David T. Ting, Andrew Aguirre, Nir Hacohen, Dana Pe’er, Eric S. Lander, Arnav Mehta. The transcriptional and epigenetic regulation of epithelial-mesenchymal plasticity in patient-derived pancreatic cancer cell lines [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 6940.
Abstract Pancreatic cancer (PDAC) is a lethal disease in part because tumor cells exist in distinct transcriptional states (e.g. basal/mesenchymal v.s. classical/epithelial) with unique phenotypic properties that contribute to tumor growth and treatment resistance. Two major mechanisms have been suggested for treatment evasion: (1) the intrinsic resistance of an existing state to a therapy regimen and (2) plasticity of therapy-sensitive states to adopt more resistant states. The relative contribution of these mechanisms to treatment resistance is still poorly understood. Historically, measurement of plasticity in both human patients and mouse models has involved one of three principles: (1) observing a redistribution of cell states in tissue across timepoints or conditions; (2) identifying cells that have genomic, epigenetic or proteomic features of more than one state (mixed states); and (3) performing single-cell cloning of cells and observing the cell states adopted by clonal progeny. While these approaches are observationally consistent with the notion of plasticity, they either fail to definitively prove the existence of plasticity, are restricted in measurements of plasticity outside of native tissues or are unable to quantify the role of plasticity in treatment resistance. Amongst the most well described forms of plasticity in human development and cancer is epithelial-mesenchymal plasticity (EMP), which includes epithelial to mesenchymal transition (EMT) and mesenchymal to epithelial transition (MET). To better understand and quantify the role of EMP in driving treatment resistance of human PDAC, we have developed single-cell multiomic, functional genomic and computational methods applied to patient-derived models and clinical biopsies. We first profiled twelve patient-derived PDAC cell lines by single-cell RNA-seq (scSeq) and learned convergent epithelial and mesenchymal gene programs that were consistent with programs observed in patient samples. We next performed lineage tracing experiments in three PDAC cell lines using an expressed lentiviral barcoding system (ClonMapper). By performing scSeq on these barcoded lines at weekly timepoints over four weeks, we proved the presence of EMP by showing a single cell can produce progeny in both epithelial and mesenchymal states. We next developed a generative probabilistic model of our lineage tracing data. This demonstrated that clones (cells sharing a barcode) had different transition matrices (different EMT and MET rates), thus suggesting each clone has a distinct level of plasticity. Having established this, we focused on identifying genes that might explain the differing plasticity properties of clones. Using elastic net regression we identified 50 transcription factors (TFs) whose expression significantly explained the propensity for EMP over time across clonal populations. Among these were were several known EMP TFs (Zeb1 and Gata6), understudied TFs (Elf3, Sox2, Sox4, Klf3, Klf5 and Atf4) and novel TFs (Meis2, Meis3, FoxA1 and the interferon regulatory factors Irf6, Irf7 and Irf9). Using single-cell multiomics (paired scSeq and single-cell ATAC-seq) on our barcoded population, we found that 9 of the 50 predicted TFs, including Elf3, had differential accessibility between clones with different plasticity properties, suggesting a role for epigenetic regulation of these TFs in facilitating EMP. Importantly, we leveraged our multiomic data to infer gene regulatory networks influenced by these TFs and found an enrichment of binding motifs for these TFs in enhancer regions of genes in epithelial and mesenchymal programs. To study the role of these predicted TFs in modulating EMP, we developed a CRISPRi system that enabled gene perturbations alongside lineage tracing. We performed a CRISPRi perturb-seq experiment (CRISPR perturbation with scSeq readouts), perturbing the 50 predicted TFs above and 10 control genes, and collected 1.5 million single-cell transcriptomic profiles, in addition to two other CRISPR KO perturb-seq experiments. We performed a negative binomial regression to estimate effect sizes of guide RNAs on all genes. 60% of our guides had significant perturbation effects on their target gene. We subsequently found Klf3 as an important regulator of PDAC proliferation independent of cell state. Importantly, we found that knockdown (KD) of several factors, such has Grhl2 and FoxA1, bias towards mesenchymal cell states, whereas KD of others such as Batf2, Snai1, Rel, Zeb1, Nr2f1 and Sox4 led to a bias towards epithelial cell states. Interestingly, KD of several TFs influenced transition properties of cells by decreasing rates of EMT and MET across barcodes without biasing the overall clonal distribution towards a single cell state. This suggests a role for these TFs in enabling plasticity and facilitating state transitions. To study the effect of plasticity in treatment resistance, we treated four barcoded cell lines with the first-line chemotherapy combination FOLFIRINOX (5-fluorouracil, oxaliplatin and SN-38, the active metabolite of irinotecan) or targeted therapy and performed scSeq yielding over 600,00 single-cell transcriptomic profiles. We found an enrichment after treatment with FOLFIRINOX of barcodes that were biased for cells in mesenchymal states, consistent with selection against epithelial cells, but also of those barcodes with the highest inferred state transition rates. With targeted therapies, we found selective depletion of mesenchymal states. We next treated our CRISPR perturbed cell lines and found overall a significantly more restricted barcode diversity in cells containing guide RNAs targeting plasticity factors compared to non-targeting controls, suggestive of the role of plasticity in facilitation resistance. To validate the role of these proposed plasticity factors in human patients, we collected paired biopsy samples from 23 patients in a phase 2 clinical trial of metastatic PDAC patients being treated with radiation therapy and dual checkpoint blockade (NCT03104439). We performed scSeq on these samples, and used a supervised Bayesian matrix factorization approach (Spectra) to learn epithelial and mesenchymal gene programs within tumor cells. We subsequently classified cells as epithelial, mesenchymal and intermediate cell types using a gaussian mixture model on gene expression features. We found the intermediate states were enriched in expression of our proposed plasticity factors, and importantly high expression of these factors in baseline samples correlated with a redistribution of states in follow-up biopsies. Our efforts define a robust experimental and quantitative framework for studying tumor cell plasticity in patient-derived model systems with validation in human patient samples using single-cell and spatial transcriptomics. Collectively, we nominate several regulators that alter the propensity of EMP in PDAC, thus posing a paradigm whereby perturbations may be used to homogenize tumor populations towards treatment-sensitive phenotypes for combination therapy. Citation Format: Arnav Mehta, Lynn Bi, Deepika Yeramosu, Michael Bogaev, Martin Jankowiak, Abigail Collins, Aziz Al'Khafaji, Milan Parikh, Mehrtash Babadi, Kyle Evans, Alex Bloemendal, Russell Kunnes, Marc Schwartz, Glen Munson, Elisa Donnard, Thouis R. Jones, Ben Z. Stanger, Jay Shendure, Jonathan Weissman, David T. Ting, Andrew Aguirre, Nir Hacohen, Dana Pe'er, Eric S. Lander. Dissecting and quantifying pancreatic cancer plasticity using single-cell multiomics, lineage tracing and functional genomics reveals novel mediators of therapy resistance [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(7_Suppl):Abstract nr NG08.
We present a method for detecting evidence of natural selection in ancient DNA time-series data that leverages an opportunity not utilized in previous scans: testing for a consistent trend in allele frequency change over time. By applying this to 8433 West Eurasians who lived over the past 14000 years and 6510 contemporary people, we find an order of magnitude more genome-wide significant signals than previous studies: 347 independent loci with >99% probability of selection. Previous work showed that classic hard sweeps driving advantageous mutations to fixation have been rare over the broad span of human evolution, but in the last ten millennia, many hundreds of alleles have been affected by strong directional selection. Discoveries include an increase from ~0% to ~20% in 4000 years for the major risk factor for celiac disease at HLA-DQB1; a rise from ~0% to ~8% in 6000 years of blood type B; and fluctuating selection at the TYK2 tuberculosis risk allele rising from ~2% to ~9% from ~5500 to ~3000 years ago before dropping to ~3%. We identify instances of coordinated selection on alleles affecting the same trait, with the polygenic score today predictive of body fat percentage decreasing by around a standard deviation over ten millennia, consistent with the "Thrifty Gene" hypothesis that a genetic predisposition to store energy during food scarcity became disadvantageous after farming. We also identify selection for combinations of alleles that are today associated with lighter skin color, lower risk for schizophrenia and bipolar disease, slower health decline, and increased measures related to cognitive performance (scores on intelligence tests, household income, and years of schooling). These traits are measured in modern industrialized societies, so what phenotypes were adaptive in the past is unclear. We estimate selection coefficients at 9.9 million variants, enabling study of how Darwinian forces couple to allelic effects and shape the genetic architecture of complex traits.
Capturing the full complexity of the clinical experiences of metastatic breast cancer (MBC) patients treated in a variety of settings is needed to better understand this disease and develop new treatment modalities. Yet, challenges exist to establish and share a large MBC dataset that integrates genomic, clinical, and patient-reported data as it requires collecting information and samples from many geographically dispersed patients and institutions. We explored whether a patient-partnered research approach that uses online engagement could enable patients living across the United States and Canada to accelerate cancer research by sharing their samples, clinical information, and experiences. In collaboration with patients and patient advocates, the Metastatic Breast Cancer Project (MBCproject; [www.mbcproject.org][1]) was developed and launched in October 2015. As of March 2020, 3,246 MBC patients who received treatment at ∼1,700 institutions had consented for the MBCproject, providing patient-reported information via surveys, as well as access to medical records and biological samples. Through the collection and analysis of tumor and germline samples, medical records, and patient-reported data, the MBCproject generates and publicly releases clinically-annotated genomic data on primary and metastatic tumor specimens on a recurring basis. Herein we describe the MBCproject cohort in detail and describe the clinico-genomic landscape of the MBCproject dataset. The complete dataset consists of whole exome sequencing (WES) for 379 tumors with matching germline from 301 patients, WES on germline samples from 377 patients, and transcriptome sequencing (RNA-seq) for 200 tumors from 141 patients, with clinical data from medical records and patient-reported information. A comparison of various clinical fields (diagnostic dates, tumor histology, tumor sites, treatments received) obtained from patient-reported data and the abstracted from medical records found a high degree of concordance, with multiple fields having over 90% concordance. Analysis of the somatic alterations in the 249 tumors taken after metastatic diagnosis found a significant enrichment of mutations in the cancer genes TP53 , PIK3CA , CDH1 , PTEN, AKT1, NF1 , and ESR1 , among others. Tumor evolutionary analysis of 14 patients with 3 or more samples identified oncogenic mutations in ESR1 , NF1 , and TP53 , genes associated with MBC and/or resistance to endocrine therapy. Analysis of germline samples identified pathogenic variants in the cancer-associated genes BRCA1, BRCA2 , ATM, and PALB2 . Comparing the frequency of pathogenic variants in patients diagnosed before/at or after the age of 40 years old, we found that the presence of these variants in BRCA1 or BRCA2 was enriched in the younger group compared to the older group (9.2% vs 2.5%, p=0.0089; two-sided Fisher exact test). Transcriptome sequencing identified putatively oncogenic in-frame fusions in cancer genes such as FANCD2 , FGFR3 , ESR1 , BRAF and NCOR1 . Analysis of tumor’s intrinsic molecular subtype (research-based PAM50) found a depletion of the Luminal A subtype in MBCproject compared to The Cancer Genome Atlas, and a switch in molecular subtype in 15 out of 35 patients with 2 or more samples. A case study of a patient with sequencing data from 4 tumor biopsies obtained during the course of their metastatic disease is presented. An integrated analysis of the clinical and multi-omic data from this patient identified distinct drivers of resistance to endocrine therapy in each of these tumors. The MBCproject clinico-genomic dataset is one of the largest available MBC patient cohorts This integrated dataset is poised for studying several understudied clinical cohorts (young women with breast cancer, de novo MBC), rare disease subtypes (e.g. lobular, metaplastic, extraordinary responders), biomarkers of response/resistance (e.g. CDK4/6 inhibitors), and real world patterns, among others, and will serve as an invaluable resource to accelerate discoveries. ### Competing Interest Statement JG owns stocks in the biotechnology exchange-traded funds CNCR, IDNA, IBB, and XBI, and owned stocks in Adaptive Biotechnologies, 2seventy bio, and bluebird bio. EJ is a current employee of Repare Therapeutics. SB is a current employee of GRAIL Inc.. JEB-B is a current employee of Cellarity. EMV reports advisory/consulting from Tango Therapeutics, Genome Medical, Genomic Life, Enara Bio, Manifold Bio, Monte Rosa, Novartis Institute for Biomedical Research, Riva Therapeutics, and Serinus Bio; research support from Novartis, BMS, and Sanof; equity from Tango Therapeutics, Genome Medical, Genomic Life, Syapse, Enara Bio, Manifold Bio, Microsoft, Monte Rosa, Riva Therapeutics, and Serinus Bio; institutional patents on chromatin mutations and immunotherapy response, and methods for clinical interpretation; intermittent legal consulting on patents for Foaley & Hoag; being on the Editorial Board of JCO Precision Oncology and Science Advances. TRG is a co-founder, holds equity, and was previously a scientific advisor in Sherlock Biosciences, Inc.; receives compensation from Anji Oncology (cash and equity), Braidwell (cash), Dewpoint Therapeutics (cash and equity); received compensation from GlaxoSmithCline (unpaid as of January 2021). CAP is a current employee at Precede Biosciences. NW is an employee of Genentech as of Feb 13, 2023 and has equity in Roche; holds equity in Relay Therapeutics and Flare Therapeutics; is a consultant for Flare Therapeutics; prior to Jan 31 2023 was a scientific advisory board member of Relay Therapeutics, an advisory board member for Eli Lilly, and received research support from Astra-Zeneca and Puma Biotechnologies. The remaining authors declare no conflicts of interest. ### Funding Statement This research was supported by the non-profit organization Count Me In (joincountmein.org) and by anonymous philanthropic support to the Broad Institute. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Institutional review board (IRB) of Dana-Farber/Harvard Cancer Center (DF/HCC) gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The clinically annotated genomic dataset of the MBCproject is shared publicly in order for all researchers to be able to utilize this data to better understand metastatic breast cancer. De-identified data has been shared on a recurring basis as the data is generated. Data from The Metastatic Breast Cancer Project is available at cBioPortal (https://www.cbioportal.org/study/summary?id=brca\_mbcproject\_2022), National Cancer Institute's Genomic Data Commons (https://portal.gdc.cancer.gov/projects/CMI-MBC) and dbGaP (Study Accession phs001709). cBioPortal has the clinical and genomic data for the MBCproject clinico-genomic dataset (N=379 tumor, 301 patients). As of May 1st 2023, updating the Genomic Data Commons and dbGaP data repositories with the genomic data for all the tumor samples used in this study is in progress. The clinically annotated genomic dataset of the MBCproject is shared publicly in order for all researchers to be able to utilize this data to better understand metastatic breast cancer. De-identified data has been shared on a recurring basis as the data is generated. Data from The Metastatic Breast Cancer Project is available at cBioPortal ([https://www.cbioportal.org/study/summary?id=brca\_mbcproject\_2022][2]), National Cancer Institute's Genomic Data Commons () and dbGaP (Study Accession phs001709). cBioPortal has the clinical and genomic data for the MBCproject clinico-genomic dataset (N=379 tumor, 301 patients). As of May 1st 2023, updating the Genomic Data Commons and dbGaP data repositories with the genomic data for all the tumor samples used in this study is in progress. [1]: http://www.mbcproject.org [2]: https://www.cbioportal.org/study/summary?id=brca_mbcproject_2022
Supplementary Figure Legends 1-6 from Loss of E-Cadherin Promotes Metastasis via Multiple Downstream Transcriptional Pathways