Estrogen receptor α (ER) mutations occur in up to 30% of metastatic ER-positive breast cancers. Recent data has shown that ER mutations impact the expression of thousands of genes not typically regulated by wildtype ER. While the majority of these altered genes can be explained by constant activity of mutant ER or genomic changes such as altered ER binding and chromatin accessibility, as much as 33% remain unexplained, indicating the potential for post-transcriptional effects. Here we explored the role of microRNAs in mutant ER-driven gene regulation and identified several microRNAs that are dysregulated in ER mutant cells. These differentially regulated microRNAs target a significant portion of mutant-specific genes involved in key cellular processes. When the activity of microRNAs is altered using mimics or inhibitors, significant changes are observed in gene expression and cellular proliferation related to mutant ER. An in-depth evaluation of miR-301b led us to discover an important role for PRKD3 in the proliferation of ER mutant cells. Our findings show that microRNAs contribute to mutant ER gene regulation and cellular effects in breast cancer cells.
Many cancers carry recurrent, change-of-function mutations affecting RNA splicing factors. Here, we describe a method to harness this abnormal splicing activity to drive splicing factor mutation-dependent gene expression to selectively eliminate tumor cells. We engineered synthetic introns that were efficiently spliced in cancer cells bearing SF3B1 mutations, but unspliced in otherwise isogenic wild-type cells, to yield mutation-dependent protein production. A massively parallel screen of 8,878 introns delineated ideal intronic size and mapped elements underlying mutation-dependent splicing. Synthetic introns enabled mutation-dependent expression of herpes simplex virus–thymidine kinase (HSV–TK) and subsequent ganciclovir (GCV)-mediated killing of SF3B1-mutant leukemia, breast cancer, uveal melanoma and pancreatic cancer cells in vitro, while leaving wild-type cells unaffected. Delivery of synthetic intron-containing HSV–TK constructs to leukemia, breast cancer and uveal melanoma cells and GCV treatment in vivo significantly suppressed the growth of these otherwise lethal xenografts and improved mouse host survival. Synthetic introns provide a means to exploit tumor-specific changes in RNA splicing for cancer gene therapy. Synthetic introns tailored for specific splice-factor mutations enable targeted cancer gene therapy.
Most endometrial cancers express the hormone receptor estrogen receptor alpha (ER) and are driven by excess estrogen signaling. However, evaluation of the estrogen response in endometrial cancer cells has been limited by the availability of hormonally responsive in vitro models, with one cell line, Ishikawa, being used in most studies. Here, we describe a novel, adherent endometrioid endometrial cancer (EEC) cell line model, HCI-EC-23. We show that HCI-EC-23 retains ER expression and that ER functionally responds to estrogen induction over a range of passages. We also demonstrate that this cell line retains paradoxical activation of ER by tamoxifen, which is also observed in Ishikawa and is consistent with clinical data. The mutational landscape shows that HCI-EC-23 is mutated at many of the commonly altered genes in EEC, has relatively few copy-number alterations, and is microsatellite instable high (MSI-high). In vitro proliferation of HCI-EC-23 is strongly reduced upon combination estrogen and progesterone treatment. HCI-EC-23 exhibits strong estrogen dependence for tumor growth in vivo and tumor size is reduced by combination estrogen and progesterone treatment. Molecular characterization of estrogen induction in HCI-EC-23 revealed hundreds of estrogen-responsive genes that significantly overlapped with those regulated in Ishikawa. Analysis of ER genome binding identified similar patterns in HCI-EC-23 and Ishikawa, although ER exhibited more bound sites in Ishikawa. This study demonstrates that HCI-EC-23 is an estrogen- and progesterone-responsive cell line model that can be used to study the hormonal aspects of endometrial cancer.
Pancreatic adenosquamous carcinoma (PASC) is an aggressive cancer whose mutational origins are poorly understood. An early study reported high-frequency somatic mutations affecting UPF1, a nonsense-mediated mRNA decay (NMD) factor, in PASC, but subsequent studies did not observe these lesions. The corresponding controversy about whether UPF1 mutations are important contributors to PASC has been exacerbated by a paucity of functional studies. Here, we modeled two UPF1 mutations in human and mouse cells to find no significant effects on pancreatic cancer growth, acquisition of adenosquamous features, UPF1 splicing, UPF1 protein, or NMD efficiency. We subsequently discovered that 45% of UPF1 mutations reportedly present in PASCs are identical to standing genetic variants in the human population, suggesting that they may be non-pathogenic inherited variants rather than pathogenic mutations. Our data suggest that UPF1 is not a common functional driver of PASC and motivate further attempts to understand the genetic origins of these malignancies.
Most eukaryotes harbor two distinct pre-mRNA splicing machineries: the major spliceosome, which removes >99% of introns, and the minor spliceosome, which removes rare, evolutionarily conserved introns. Although hypothesized to serve important regulatory functions, physiologic roles of the minor spliceosome are not well understood. For example, the minor spliceosome component ZRSR2 is subject to recurrent, leukemia-associated mutations, yet functional connections among minor introns, hematopoiesis and cancers are unclear. Here, we identify that impaired minor intron excision via ZRSR2 loss enhances hematopoietic stem cell self-renewal. CRISPR screens mimicking nonsense-mediated decay of minor intron-containing mRNA species converged on LZTR1, a regulator of RAS-related GTPases. LZTR1 minor intron retention was also discovered in the RASopathy Noonan syndrome, due to intronic mutations disrupting splicing and diverse solid tumors. These data uncover minor intron recognition as a regulator of hematopoiesis, noncoding mutations within minor introns as potential cancer drivers and links among ZRSR2 mutations, LZTR1 regulation and leukemias.
Mutations in RNA splicing factors are amongst the most common genetic alterations in myeloid malignancies. Mutations in the splicing factors SF3B1, SRSF2, and U2AF1 occur as heterozygous, missense mutations and have been shown to confer a change-of-function. In contrast, the X chromosome encoded ZRSR2 is enriched in nonsense/frameshift mutations in males, consistent with loss of function. To date however, we do not understand the basis for enrichment of ZRSR2 mutations in leukemia. Moreover, ZRSR2 is the only one of these factors that primarily functions in the minor spliceosome. While most introns are spliced by the major spliceosome, a small subset (<1%) of introns are recognized by a separate complex, the minor spliceosome. Although minor (or "U12") introns are present in only ~800 genes in humans, their sequences and positions are highly evolutionarily conserved - more so than their U2 counterparts. The high conservation of minor introns suggests key regulatory roles yet few functional roles for the minor spliceosome in regulating biological phenotypes are known. The rarity and conservation of minor introns offered a unique opportunity to investigate splicing factor mutations and identify potential tissue-specific roles of the minor spliceosome. Modeling loss-of-function mutations in ZRSR2 via a mouse model for induced deletion of Zrsr2 revealed strikingly enhanced self-renewal of Zrsr2-deficient male and female hematopoietic cells (Fig. A-C). This was in stark contrast to the effects of hotspot mutations in Sf3b1and Srsf2 and similar to those of Tet2 loss on increasing self-renewal and numbers of HSCs. Zrsr2 loss was also associated increased myeloid cells in the blood and long-term hematopoietic stem cells (HSCs) in the marrow (Fig. C). To understand the mechanistic basis by which ZRSR2 loss causes aberrant HSC self-renewal, we quantified transcriptome-wide splicing patterns in MDS patients. ZRSR2-mutant samples had widespread, dysfunctional recognition of minor introns- 48% of minor introns exhibiting significantly increased retention (Fig. D). We next systematically mimicked the effects of nonsense-mediated decay caused by minor intron retention in ZRSR2-mutants. Every gene containing a ZRSR2-regulated minor intron was targeted by 4 sgRNAs via a positive-enrichment CRISPR screen using pools of lentiviral sgRNAs in cytokine-dependent human and mouse hematopoietic cell lines. This identified several minor intron-containing genes whose downregulation conferred cytokine independence. Strikingly, just one gene was enriched in all lines (Fig. E): LZTR1, a cullin-3 adaptor for ubiquitin-mediated suppression of RAS-related GTPases which is subject to loss-of-function mutations in several cancers and the RASopathy Noonan Syndrome. Minor intron retention in LZTR1 correlated with reduced LZTR1 protein in MDS patients (Fig. F-G). Inducing mutations in either the protein-coding region of LZTR1 or its minor intron resulted in cytokine independence (Fig. H), reduced LZTR1, and dramatic accumulation of RIT1, a RAS GTPase substrate of LZTR1. In a Noonan Syndrome family wherein one child died of AML, the mother and all children carried an intronic mutation within LZTR1's minor intron (Fig. I-J). Fibroblasts from each family member revealed clear LZTR1 minor intron retention with impaired LZTR1 protein expression and RIT1 accumulation in subjects bearing the LZTR1 minor intron mutation (Fig. J). We next interrogated LZTR1 minor intron splicing across all cancers in the TCGA. While LZTR1's minor intron was efficiently excised in normal samples, a notable subset of tumors in almost all cancer types exhibited significantly increased retention within LZTR1's minor intron. These data indicate LZTR1 is frequently dysregulated via perturbed minor intron splicing - much more so than by protein-coding mutations alone. Here we uncover a heretofore unrecognized role of minor intron excision in regulating HSC self-renewal, a molecular link between ZRSR2 mutations and aberrant LZTR1 splicing and expression, and frequent LZTR1 minor intron retention in diverse cancers and cancer predisposition syndromes. Given frequent post-transcriptional disruption of LZTR1 in the absence of protein-coding mutations, our data additionally motivate study of other cancer-associated minor intron-containing genes which may be dysregulated via similar, and as-yet-undetected, aberrant splicing. Figure Disclosures Abdel-Wahab: Merck: Consultancy; Envisagenics Inc.: Current equity holder in private company; H3 Biomedicine Inc.: Consultancy, Research Funding; Janssen: Consultancy.
Article Figures and data Abstract eLife digest Introduction Results Discussion Materials and methods Appendix 1 Data availability References Decision letter Author response Article and author information Metrics Abstract Pancreatic adenosquamous carcinoma (PASC) is an aggressive cancer whose mutational origins are poorly understood. An early study reported high-frequency somatic mutations affecting UPF1, a nonsense-mediated mRNA decay (NMD) factor, in PASC, but subsequent studies did not observe these lesions. The corresponding controversy about whether UPF1 mutations are important contributors to PASC has been exacerbated by a paucity of functional studies. Here, we modeled two UPF1 mutations in human and mouse cells to find no significant effects on pancreatic cancer growth, acquisition of adenosquamous features, UPF1 splicing, UPF1 protein, or NMD efficiency. We subsequently discovered that 45% of UPF1 mutations reportedly present in PASCs are identical to standing genetic variants in the human population, suggesting that they may be non-pathogenic inherited variants rather than pathogenic mutations. Our data suggest that UPF1 is not a common functional driver of PASC and motivate further attempts to understand the genetic origins of these malignancies. eLife digest Cancer is a group of complex diseases in which cells grow uncontrollably and spread into surrounding tissues and other parts of the body. All types of cancers develop from changes – or mutations – in the genes that affect the pathways involved in controlling the growth of cells. Different cancers possess unique sets of mutations that affect specific genes, and often, it is difficult to determine which of them play the most important role in a particular type of cancer. For example, pancreatic adenosquamous carcinoma, a rare and aggressive form of pancreatic cancer, is a devastating disease with a poor chance of survival – patients rarely live longer than one year after diagnosis. While the cells of this particular cancer display distinct features that separate them from other forms of pancreatic cancer, the genetic causes of these features are unclear. Using new technologies, some researchers have reported mutations in a ‘quality control’ gene called ‘UPF1’, which is responsible for destroying faulty forms of genetic material. However, subsequent studies did not find such mutations. To clarify the role of UPF1 in pancreatic adenosquamous carcinoma, Polaski et al. used mouse and human cancer cells with UPF1 mutations and monitored their effects on tumour growth and the development of features unique to this disease. Polaski et al. first injected mice with mouse pancreatic cancer cells containing mutations in UPF1 (mutated cells) and cancer cells without. Both groups of mice developed pancreatic tumours but there was no difference in tumour growth between the mutated and non-mutated cells, and neither cell type displayed distinct features. The researchers then generated human mutated cells, which were also found to lack any specific characteristics. Further analysis showed that the mutations did not stop UPF1 from working, in fact, over 40% of these mutations occurred naturally in humans without causing cancer. This suggests that UPF1 does not seem to be involved in pancreatic adenosquamous carcinoma. Further investigation is needed to illuminate key genetic players in the development of this type of cancer, which will be vital for improving treatments and outcomes for patients suffering from this disease. Introduction Pancreatic adenosquamous carcinoma (PASC) is a rare and aggressive disease that constitutes 1–4% of pancreatic exocrine tumors (Madura et al., 1999). Patient prognosis is extremely poor, with a median survival of 8 months (Simone et al., 2013). Although PASC is clinically and histologically distinct from the more common disease pancreatic adenocarcinoma, the genetic and molecular origins of PASC’s unique features are unknown. A recent study reported a potential breakthrough in our understanding of PASC etiology. Liu et al., 2014 reported high-frequency mutations affecting UPF1, which encodes a core component of the nonsense-mediated mRNA decay (NMD) pathway, in 78% (18 of 23) of PASC patients. These mutations were absent from patient-matched normal pancreatic tissue (0 of 18) and from non-PASC tumors (0 of 29 non-adenosquamous pancreatic carcinomas and 0 of 21 lung squamous cell carcinomas). The authors used a combination of molecular and histological assays to find that the UPF1 mutations caused UPF1 mis-splicing, loss of UPF1 protein, and impaired NMD, resulting in stable expression of aberrant mRNAs containing premature termination codons that would normally be degraded by NMD. The recurrent, PASC-specific, and focal nature of the reported UPF1 mutations, together with their dramatic effects on NMD activity, suggested that UPF1 mutations are a key feature of PASC biology. Three subsequent studies of distinct PASC cohorts, however, did not report somatic mutations in UPF1 (Fang et al., 2017; Hayashi et al., 2020; Witkiewicz et al., 2015). This absence of UPF1 mutations is significantly different from the high rate reported by Liu et al. (0 of 34 total PASC samples from three cohorts [Fang et al., 2017; Hayashi et al., 2020; Witkiewicz et al., 2015] vs. 18 of 23 PASC samples from Liu et al; p<10−8 by the two-sided binomial proportion test). Although these other studies relied on whole-exome and/or genome sequencing instead of targeted UPF1 gene sequencing, those technologies yield good coverage of the relevant UPF1 gene regions because the affected introns are very short. Given this discrepancy, we sought to directly assess the functional contribution of UPF1 mutations to PASC using a combination of biological and molecular assays. Results We first tested the role of the reported UPF1 mutations during tumorigenesis in vivo. Liu et al. reported that the majority of UPF1 mutations caused skipping of UPF1 exons 10 and 11, disrupting UPF1’s RNA helicase domain that is essential for its NMD activity (Lee et al., 2015). We therefore modeled UPF1 mutation-induced exon skipping by designing paired guide RNAs flanking Upf1 exons 10 and 11, such that these exons would be deleted upon Cas9 expression (Figure 1A). We chose mouse pancreatic cancer cells (KPC cells: KrasG12D; Trp53R172H/null; Pdx1-Cre) as a model system. KPC cells are defined by mutations affecting KRAS and p53 (encoded by Kras and Trp53 in mouse) that also occur in the vast majority of PASC cases (Borazanci et al., 2015; Fang et al., 2017; Hayashi et al., 2020), making them a genetically appropriate system. We delivered Upf1-targeting paired guide RNAs to KPC cells using recombinant adenoviral vectors and confirmed that guide delivery resulted in the production of UPF1 mRNA lacking exons 10 and 11 and a corresponding reduction in full-length UPF1 protein levels (Figure 1—figure supplement 1A–G). We injected subcloned control and Upf1-targeted KPC cells into the tails of the pancreata of B6 albino mice (n = 10 mice per treatment) and monitored tumor growth and animal survival. We detected no significant differences in tumor volume or survival in mice implanted with control or Upf1-targeted KPC cells (Figure 1B–D and Figure 1—source datas 1 and 2). Tumors derived from control as well as Upf1-targeted cells displayed similar histopathological features characteristic of moderately to poorly differentiated pancreatic ductal adenocarcinomas (Figure 1E–G). Moderately differentiated areas were composed of medium to small duct-like structures or tubules with lower mucin production, while poorly differentiated components were characterized by solid sheets or nests of tumor cells with large eosinophilic cytoplasms and large pleomorphic nuclei. No squamous differentiation was identified by histomorphologic evaluation and no expression of the squamous marker p40 (ΔNp63) was detected (Figure 1H and Supplementary file 1a). We concluded that inducing the reported Upf1 exon skipping in vivo had no detectable effects on pancreatic cancer growth or acquisition of adenosquamous features in the KPC model. However, there are several important caveats to our data. First, we cannot rule out the possibility that inducing Upf1 exon skipping in a different model system or cell type could influence tumorigenesis. Second, as our assays were performed in the complex setting of in vivo tumorigenesis, we cannot infer how inducing the reported Upf1 exon skipping might affect tumor cell proliferation in the controlled setting of in vitro growth. Third, as accurate measurement of Upf1 spicing and UPF1 protein isoforms was only possible for KPC cells prior to orthotopic injection, we cannot infer how the relative frequencies of mis-spliced UPF1 mRNA and the resulting truncated proteins may have changed during tumorigenesis. Figure 1 with 1 supplement see all Download asset Open asset UPF1 mutations do not result in the acquisition of squamous histological features or confer a growth advantage to mutant cells in vivo. (A) Schematic of UPF1 gene structure and corresponding encoded protein domains. Intron 10 (I10) contains the bulk of the mutations reported by Liu et al. Scissors indicate the sites targeted by the paired guide RNAs used to excise exons 10 and 11 (E10 and E11). Red nucleotides represent positions subject to point mutations reported in Liu et al. Arrows indicate specific mutations that we modeled in 293 T cells. The horizontal black line indicates the nucleotide within the protospacer adjacent motif (PAM) site that we mutated to prevent repeated cutting by Cas9 in 293 T cells. (B) Top, experimental strategy for testing whether mimicking UPF1 mis-splicing by deleting exons 10 and 11 promoted pancreatic cancer growth. Mice were orthotopically injected with mouse pancreatic cancer cells (KPC cells: KrasG12D; Trp53R172H/null; Pdx1-Cre) lacking Upf1 exons 10 and 11. Bottom, hematoxylin and eosin (H and E) stain of pancreatic tumor tissue harvested from the mice. (C) Line graph comparing tumor volume between mice injected with control (AdCas9; Cas9 only) or treatment (AdUpf1; Cas9 with Upf1-targeting guide RNAs) KPC cells. Tumor volume measured by ultrasound imaging. Error bars, standard deviation computed over surviving animals (n = 10 at first time point). n.s., not significant (p>0.05). p-values at each timepoint were calculated relative to the control group with an unpaired, two-tailed t-test. (D) Survival curves for the control (AdCas9) or treatment (AdUpf1) cohorts. Error bars, standard deviation computed over biological replicates (n = 10, each group). p-value was calculated relative to the control group by a logrank test. (E) Representative hematoxylin and eosin (H and E) staining of a pancreatic tumor resulting from orthotopic injection of control KPC cells displaying features of a moderately to poorly differentiated pancreatic ductal adenocarcinoma. Tumors were composed of medium-size duct-like structures and small tubular glands with lower mucin production. (F) Representative H and E image illustrating a moderately to poorly differentiated pancreatic ductal adenocarcinoma resulting from orthotopic injection of Upf1-targeted KPC cells. Depicted here is a section of the poorly differentiated component (arrow), which was characterized by solid sheets of tumor cells with large eosinophilic cytoplasms and marked nuclear polymorphism. (G) Representative H and E image of a pancreatic tumor resulting from orthotopic injection of Upf1-targeted KPC cells. The dashed circle marks a moderately differentiated component; the remainder is poorly differentiated. (H) Representative IHC image of a pancreatic tumor resulting from orthotopic injection of Upf1-targeted KPC cells for the squamous marker p40 (ΔNp63). No expression of the marker was observed in tumor cells. Figure 1—source data 1 Source data for mouse tumor volume (Figure 1C). https://cdn.elifesciences.org/articles/62209/elife-62209-fig1-data1-v2.xlsx Download elife-62209-fig1-data1-v2.xlsx Figure 1—source data 2 Source data for mouse survival (Figure 1D). https://cdn.elifesciences.org/articles/62209/elife-62209-fig1-data2-v2.xlsx Download elife-62209-fig1-data2-v2.xlsx We next assessed molecular phenotypes induced with UPF1 mutations. Liu et al. measured the effects of each mutation on UPF1 splicing using a minigene assay, in which each mutation was introduced into a plasmid containing a small fragment of the UPF1 gene that was subsequently transfected into 293 T cells. Liu et al. concluded that all reported UPF1 mutations caused dramatic UPF1 mis-splicing that disrupted key protein domains that are essential for UPF1 function in NMD. Minigenes are common tools for studying splicing, but they are frequently spliced less efficiently than endogenous genes, presumably because they are gene fragments that lack potentially important sequence features that promote splicing and incompletely capture the close relationship between chromatin and splicing (Luco et al., 2011; Naftelberg et al., 2015). We modeled UPF1 mutations in 293 T cells in order to mimic Liu et al.’s experimental strategy, but introduced mutations into their endogenous genomic contexts rather than using minigenes. We selected two distinct UPF1 mutations in intron 10 for these studies. We selected IVS10+31G>A (patient 1; P1) because it was reportedly recurrent across three different patients (making it equally or more common than any other mutation) and induced strong mis-splicing on its own (36% mis-spliced mRNA, versus 0% for wild-type UPF1); we selected IVS10-17G>A (patient 9; P9) because it had one of the strongest effects on splicing (90% mis-spliced mRNA). IVS10+31G>A was present in a homozygous state in two of the three patients carrying it, while IVS10-17G>A was present in a heterozygous state. We introduced each mutation into its endogenous context by transiently transfecting a plasmid expressing Cas9 and a single guide RNA (sgRNA) targeting UPF1 intron 10 as well as appropriate donor DNA for homology-directed repair, screened the resulting cells for the desired genotypes, and established clonal lines. The resulting cell lines contained the desired mutations in the correct copy numbers as well as a point mutation disrupting the protospacer adjacent motif (PAM) site (Figure 2—figure supplement 1A–C). As neither the PAM site itself nor nearby positions were reported as mutated in Liu et al., we additionally established a cell line in which only the PAM site was mutated as a wild-type control. We systematically tested the functional consequences of UPF1 mutations for NMD efficiency, UPF1 protein levels, and UPF1 splicing. We measured NMD efficiency in our engineered cells using the well-established beta-globin reporter system, which permits controlled measurement of the relative levels of mRNAs that do or do not contain an NMD-inducing premature termination codon, but which are otherwise identical (Zhang et al., 1998). We did not observe decreased NMD efficiency in UPF1-mutant versus wild-type cells; instead, UPF1-mutant cells exhibited evidence of modestly more efficient NMD, although these differences were not statistically significant (Figure 2A and Figure 2—source data 1). To confirm these results from reporter experiments, we then queried levels of endogenous NMD substrates across the transcriptome. We performed high-coverage RNA-seq on each of the three 293 T cell lines that we engineered to lack or contain defined UPF1 mutations in biological triplicate, quantified transcript expression, and identified differentially expressed transcripts. We focused on NMD substrates arising from alternative splicing, as these are abundant and sensitive biomarkers of NMD efficiency that are internally controlled for gene expression variation (Feng et al., 2015). These analyses revealed that neither UPF1-mutant cell line exhibited global increases in the expression of NMD substrates relative to wild-type cells. Instead, both UPF1-mutant cell lines exhibited modestly lower global levels of endogenous NMD substrates than did wild-type cells, mimicking the trend observed with our NMD reporter experiments. Together, these data confirm that the tested mutations in UPF1 intron 10 do not affect NMD activity (Figure 2B–D and Supplementary file 1b and c). Figure 2 with 1 supplement see all Download asset Open asset Mutations in UPF1 intron 10 do not inhibit nonsense-mediated mRNA decay (NMD) or cause exon skipping. (A) Box plot of NMD efficiency in 293 T cells engineered to contain wild-type (WT) or mutant (P1, P9) UPF1. P1 and P9 correspond to the IVS10+31G>A and IVS10-17G>A mutations reported by Liu et al. All cells have the protospacer adjacent motif (PAM) site mutation illustrated in Figure 1A. NMD efficiency estimated via the beta-globin reporter assay11. Middle line, notches, and whiskers indicate median, first and third quartiles, and range of data. Each point corresponds to a single biological replicate. n.s., not significant (p>0.05). p-values were calculated for each variant relative to the control by a two-sided Mann–Whitney U test (p=0.40 for P1, 0.30 for P9). (B) Scatter plot showing transcriptome-wide quantification of transcripts containing NMD-promoting features in 293 T cells carrying the UPF1 mutation that was reportedly observed in patient 1 relative to control, wild-type cells. Each point corresponds to a single isoform that is a predicted NMD substrate (NMD(+)). Purple points represent NMD substrates that are significantly increased in UPF1-mutant cells relative to wild-type cells; black points represent NMD substrates that exhibit the opposite behavior. Plot is restricted to NMD substrates arising from differential inclusion of cassette exons. Significantly increased/decreased NMD substrates were defined as transcripts that displayed either an absolute increase/decrease in isoform ratio of ≥10% or an absolute log fold-change in expression of ≥2 with associated p≤0.05 (two-sided t-test). (C) As (B), but for 293 T cells carrying the UPF1 mutation that was reportedly observed in patient 9. Gold points represent NMD substrates that are significantly increased in UPF1-mutant cells relative to wild-type cells. (D) Summary of the numbers of NMD substrates arising from differential alternative splicing that exhibit significantly higher or lower levels in UPF1-mutant cells relative to wild-type cells. Analysis is identical to (B) and (C), but extended to the illustrated different types of alternative splicing. (E) Left, immunoblot of full-length UPF1 protein for the 293 T cell lines. Each lane represents a single biological replicate with the indicated genotype. GAPDH serves as a loading control. Equal amounts of protein were loaded in each lane (measured by fluorescence). Right, box plot illustrating UPF1 protein levels relative to GAPDH for each genotype. Middle line, notches, and whiskers indicate median, first and third quartiles, and range of data. Each point corresponds to a single biological replicate. Data was quantified with Fiji (v2.0.0). A.U., arbitrary units. n.s., not significant (p>0.05). p-values were calculated for each variant relative to the control by a two-sided Mann–Whitney U test (p=0.10 for P1, 1.0 for P9). (F) PCR using primers that amplify both full-length UPF1 mRNA (FL) and mRNA lacking exons 10 and 11 (ΔE10-11). UPF1 mRNA lacking exons 10 and 11 was only detected in the positive control lanes (ΔE10-11 spike in), in which DNA corresponding to UPF1 cDNA lacking exons 10 and 11 was synthesized and added to cDNA libraries created from WT cells prior to PCR. Numbers above each lane indicate biological replicates. Numbers below each lane represent the abundance of the lower band as a percentage of total intensity (see Materials and methods). Data was quantified with Fiji (v2.0.0). (G) RNA-seq read coverage across the genomic locus containing UPF1 exons 9–12 in the indicated 293 T cell lines. Each sample corresponds to a distinct biological replicate. Numbers represent read counts that supported each indicated splice junction (Katz et al., 2015). Figure 2—source data 1 Source data for qPCR in HEK 293 T cell lines (Figure 2A). https://cdn.elifesciences.org/articles/62209/elife-62209-fig2-data1-v2.xlsx Download elife-62209-fig2-data1-v2.xlsx Figure 2—source data 2 Source data for western blot in HEK 293 T cell lines (Figure 2E). https://cdn.elifesciences.org/articles/62209/elife-62209-fig2-data2-v2.xlsx Download elife-62209-fig2-data2-v2.xlsx Figure 2—source data 3 Source data for RT-PCR in HEK 293 T cell lines (Figure 2F). https://cdn.elifesciences.org/articles/62209/elife-62209-fig2-data3-v2.xlsx Download elife-62209-fig2-data3-v2.xlsx Consistent with similar NMD activity independent of UPF1 mutational status, UPF1 mutations did not cause loss of full-length UPF1 protein (Figure 2E, Figure 2—figure supplement 1D, and Figure 2—source data 2). Although UPF1 protein levels varied between the individual cell lines, this variation in UPF1 protein levels was not associated with variation in NMD efficiency and did not segregate with UPF1 mutational status. We therefore measured the levels of normally spliced and mis-spliced UPF1 mRNA. We readily detected normally spliced UPF1 mRNA in all samples by RT-PCR, but found no evidence of mis-spliced UPF1 mRNA, except in positive control samples in which we spiked in synthesized DNA corresponding to the exon skipping isoform reported in Liu et al. (Figure 2F, Figure 2—figure supplement 1E, and Figure 2—source data 3). We confirmed these results with our RNA-seq data by mapping all reads against all possible splice junctions connecting exons 9, 10, 11, and 12. These analyses revealed no evidence of splice junctions consistent with the reported exon 10 and 11 skipping or other abnormal exon skipping isoforms (Figure 2G). Given the differences between Liu et al.’s findings of common UPF1 mutations and their absence from subsequent studies of PASC, we wondered whether some of the UPF1 mutations reported by Liu et al. might correspond to inherited genetic variation rather than somatically acquired mutations. We searched for each mutation reported by Liu et al. within databases compiled by the 1000 Genomes Project, NHLBI Exome Sequencing Project, Exome Aggregation Consortium (ExAC), and the genome aggregation database (gnomAD) (Auton et al., 2015; Karczewski et al., 2020; Exome Aggregation Consortium et al., 2016; Server EV, 2016). These databases were constructed from a mix of whole-genome and whole-exome sequencing, both of which are effective for discovering variants within the relevant regions of UPF1 (because UPF1 introns 10, 21, and 22 are very short, they are well covered by exon-capture technologies). We found genetic variants identical to 45% (18 of 40) of the reported UPF1 mutations, one of which is present in the reference human genome. Eighty-nine percent (16 of 18) of UPF1-mutant patients had one or more reported mutations that corresponded to standing genetic variation (Figure 3A–F and Supplementary file 1d). The distribution of overlaps between reported UPF1 mutations and standing genetic variation depended strongly upon genic context. A large fraction of reported intronic UPF1 mutations were identical to standing genetic variation, while the majority of reported exonic UPF1 mutations were not (Supplementary file 1d). Figure 3 Download asset Open asset Many reported UPF1 mutations are identical to genetic variants. (A) Illustration of the mutations in UPF1 intron 10 (I10) reported by Liu et al. Each row indicates the wild-type (WT) sequence from the reference human genome or mutations reported by Liu et al. (P1, patient 1). Purple and gold arrows indicate the mutations that we modeled with genome engineering in 293 T cells for patient 1 and patient 9, respectively. Red nucleotides represent positions subject to point mutations reported in Liu et al. The horizontal black line indicates the nucleotide within the protospacer adjacent motif (PAM) site that we mutated to prevent repeated cutting by Cas9. Parentheses indicate where we found genetic variation at a reported mutation position that differed from the specific mutated nucleotide reported by Liu et al. (B) As (A), but for UPF1 exon 10 (E10). (C) As (A), but for UPF1 exon 11 (E11). (D) As (A), but for UPF1 exon 21 (E21). (E) As (A), but for UPF1 intron 22 (I22). (F) As (A), but for UPF1 exon 23 (E23). Our discovery that a large fraction of the reported UPF1 mutations are present in databases of germline genetic variation was surprising for two reasons. First, when strongly cancer-linked mutations occur as germline variants, they frequently manifest as cancer predisposition syndromes. However, no such relationship is known for UPF1 genetic variants, despite their reportedly high prevalence as identical somatic mutations in PASC. Second, UPF1 is essential for embryonic viability and development in mammals (Medghalchi et al., 2001), zebrafish (Wittkopp et al., 2009), and Drosophila (Avery et al., 2011). As Liu et al. reported that all UPF1 mutations caused mis-splicing that is expected to disable UPF1 protein function (Liu et al., 2014), then those mutations should be incompatible with life when present as inherited genetic variants. Our finding that two reported mutations had no effect on UPF1 splicing when introduced into their endogenous genomic contexts offers a way to explain this incongruity, at least for the two reported lesions that we studied. Given these discrepancies, we next sought to verify the somatic nature of the UPF1 mutations described in Liu et al., which was reportedly determined by sequencing both tumors and patient-matched controls. The GenBank accession codes reported in Liu et al. corresponded to short nucleotide sequences containing UPF1 mutations, without corresponding data for patient-matched controls. We contacted the senior author (Dr. YanJun Lu) to request primary sequencing data from patient-matched tumor and normal samples, but neither primary sequencing data from matched samples nor the samples themselves were available. To further explore whether UPF1 is recurrently mutated in PASC, we reanalyzed sequencing data from Fang et al., 2017 to manually search for UPF1 mutations (Supplementary file 1e-f). We focused on the two loci that contained all mutations reported by Liu et al. (UPF1 exons 10-11 and exons 21–23). Because the relevant introns are very short, they were well covered by both the whole-exome and whole-genome sequencing used by Fang et al. Using relaxed mutation-calling criteria to maximize sensitivity (details in Materials and methods), we identified somatic UPF1 mutations in samples from 6 of 17 PASC patients. However, those mutations exhibited genetic characteristics expected of passenger, not driver, mutations. None of those UPF1 mutations matched the UPF1 mutations reported by Liu et al., and only one was present at an allelic frequency equal to the allelic frequency of mutant KRAS, which is a known driver and which we detected in samples from all PASC patients (median allelic frequencies of 12% versus 34% for UPF1 versus KRAS mutations). Furthermore, we also identified UPF1 mutations in samples from patients with non-adenosquamous tumors (3 of 34 pancreatic ductal adenocarcinomas), whereas Liu et al. reported finding no UPF1 mutations in non-adenosquamous pancreatic cancers (0 of 29). In concert with the reports of Witkiewicz et al., 2015 and Hayashi et al., 2020 of finding no UPF1 mutations in their PASC samples, these analyses suggest that UPF1 is not a frequent or adenosquamous-specific mutational target in most PASC cohorts. Discussion UPF1’s role in the pathogenesis of PASC has been unclear and controversial given the seeming discrepancies between its mutational spectrum in different PASC cohorts. Although it is difficult to conclusively prove that a specific genetic change does not promote cancer, we were unable to detect biological or molecular changes arising from two mutations reported by Liu et al. UPF1’s status as an essential gene and our discovery that many reported UPF1 mutations occur as germline genetic variants of no known pathogenicity together suggest that other UPF1 mutations reported by Liu et al. could similarly represent genetic differences that do not functionally contribute to PASC. Our study highlights the need for continued study of the PASC mutational spectrum in order to understand the molecular basis of this disease. Materials and methods Construction of mouse KPC cells carrying a deletion of Upf1 exons 10 and 11 Request a detailed protocol Mouse KPC cells (KrasG12D; Trp53R172H/null; Pdx1-Cre) were obtained from Dr. Robert Vonderheide and were cultured in DMEM (GIBCO) supplemented with 10% fetal bovine serum (FBS) and 1% Penicillin/Streptomycin (GIBCO). All cell lines were incubated at 37°C and 5% CO2. Guide RNAs targeting mouse Upf1 introns 9 and 11 were cloned into a paired guide expression vector (px333) as previously described (Maddalo et al., 2014). An EcoRI-XhoI fragment containing the double U6-sgRNA cassette and Flag-tagged Cas9 was then ligated into the EcoRI-XhoI-digested pacAd5 shuttle vector. Recombinant adenoviruses were generated by Viraquest (Ad-Upf1 and Ad-Cas9) or purchased from the University of Iowa (Ad-Cre). KPC cells were infected with (5 × 106 PFU) of Ad-Cas9 or Ad-Upf1 in each well of a 6-well plate. Genomic DNA was extracted 48 hr post infection to confirm excision of Upf1 exons 10 and 11. For PCR analysis of genomic DNA, cells were collected in lysis buffer (100 nM Tris-HCl at pH 8.5, 5 mM EDTA, 0.2% SDS, 200 mM NaCl supplemented with fresh proteinase K at a final concentration of 100 ng/mL). Genomic DNA was extracted
Genes encoding the RNA splicing factors SF3B1, SRSF2, and U2AF1 are subject to frequent missense mutations in clonal hematopoiesis and diverse neoplastic diseases. Most "spliceosomal" mutations affect specific hotspot residues, resulting in splicing changes that promote disease pathophysiology. However, a subset of patients carry spliceosomal mutations that affect non-hotspot residues, whose potential functional contributions to disease are unstudied. Here, we undertook a systematic characterization of diverse rare and private spliceosomal mutations to infer their likely disease relevance. We utilized isogenic cell lines and primary patient materials to discover that 11 of 14 studied rare and private mutations in SRSF2 and U2AF1 induced distinct splicing alterations, including partially or completely phenocopying the alterations in exon and splice site recognition induced by hotspot mutations or driving "dual" phenocopies that mimicked two co-occurring hotspot mutations. Our data suggest that many rare and private spliceosomal mutations contribute to disease pathogenesis and illustrate the utility of molecular assays to inform precision medicine by inferring the potential disease relevance of newly discovered mutations.
While RNA-seq has enabled comprehensive quantification of alternative splicing, no correspondingly high-throughput assay exists for functionally interrogating individual isoforms. We describe pgFARM (paired guide RNAs for alternative exon removal), a CRISPR–Cas9-based method to manipulate isoforms independent of gene inactivation. This approach enabled rapid suppression of exon recognition in polyclonal settings to identify functional roles for individual exons, such as an SMNDC1 cassette exon that regulates pan-cancer intron retention. We generalized this method to a pooled screen to measure the functional relevance of ‘poison’ cassette exons, which disrupt their host genes’ reading frames yet are frequently ultraconserved. Many poison exons were essential for the growth of both cultured cells and lung adenocarcinoma xenografts, while a subset had clinically relevant tumor-suppressor activity. The essentiality and cancer relevance of poison exons are likely to contribute to their unusually high conservation and contrast with the dispensability of other ultraconserved elements for viability. pgFARM (paired guide RNAs for alternative exon removal) is a CRISPR–Cas9-based approach to manipulate alternative splicing and identify functional roles for individual exons, including poison exons with essential and tumor-suppressor roles.
RNAs directly regulate a vast array of cellular processes, emphasizing the need for robust approaches to fluorescently label and track RNAs in living cells. Here, we develop an RNA imaging platform using the cobalamin riboswitch as an RNA tag and a series of probes containing cobalamin as a fluorescence quencher. This highly modular 'Riboglow' platform leverages different colored fluorescent dyes, linkers and riboswitch RNA tags to elicit fluorescence turn-on upon binding RNA. We demonstrate the ability of two different Riboglow probes to track mRNA and small noncoding RNA in live mammalian cells. A side-by-side comparison revealed that Riboglow outperformed the dye-binding aptamer Broccoli and performed on par with the gold standard RNA imaging system, the MS2-fluorescent protein system, while featuring a much smaller RNA tag. Together, the versatility of the Riboglow platform and ability to track diverse RNAs suggest broad applicability for a variety of imaging approaches.
Riboswitches are structured mRNA sequences that regulate gene expression by directly binding intracellular metabolites. Generating the appropriate regulatory response requires the RNA rapidly and stably acquire higher-order structure to form the binding pocket, bind the appropriate effector molecule and undergo a structural transition to inform the expression machinery. These requirements place riboswitches under strong kinetic constraints, likely restricting the sequence space accessible by recurrent structural modules such as the kink turn and the T-loop. Class-II cobalamin riboswitches contain two T-loop modules: one directing global folding of the RNA and another buttressing the ligand binding pocket. While the T-loop module directing folding is highly conserved, the T-loop associated with binding is substantially less so, with no clear consensus sequence. To further understand the functional role of the binding-associated module, a functional genetic screen of a library of riboswitches with the T-loop and its interacting nucleotides was used to build an experimental phylogeny comprised of sequences that possess a wide range of cobalamin-dependent regulatory activity. Our results reveal conservation patterns of the T-loop and its interaction with the binding core that allow for rapid tertiary structure formation and demonstrate its importance for generating strong ligand-dependent repression of mRNA expression.
Riboswitches are a widely distributed class of regulatory RNAs in bacteria that modulate gene expression via small-molecule-induced conformational changes. Generally, these RNA elements are grouped into classes based upon conserved primary and secondary structure and their cognate effector molecule. Although this approach has been very successful in identifying new riboswitch families and defining their distributions, small sequence differences between structurally related RNAs can alter their ligand selectivity and regulatory behavior. Herein, we use a structure-based mutagenic approach to demonstrate that cobalamin riboswitches have a broad spectrum of preference for the two biological forms of cobalamin in vitro using isothermal titration calorimetry. This selectivity is primarily mediated by the interaction between a peripheral element of the RNA that forms a T-loop module and a subset of nucleotides in the cobalamin-binding pocket. Cell-based fluorescence reporter assays in Escherichia coli revealed that mutations that switch effector preference in vitro lead to differential regulatory responses in a biological context. These data demonstrate that a more comprehensive analysis of representative sequences of both previously and newly discovered classes of riboswitches might reveal subgroups of RNAs that respond to different effectors. Furthermore, this study demonstrates a second distinct means by which tertiary structural interactions in cobalamin riboswitches dictate ligand selectivity.
Allosteric RNA devices are increasingly being viewed as important tools capable of monitoring enzyme evolution, optimizing engineered metabolic pathways, facilitating gene discovery and regulators of nucleic acid-based therapeutics. A key bottleneck in the development of these platforms is the availability of small-molecule-binding RNA aptamers that robustly function in the cellular environment. Although aptamers can be raised against nearly any desired target through in vitro selection, many cannot easily be integrated into devices or do not reliably function in a cellular context. Here, we describe a new approach using secondary- and tertiary-structural scaffolds derived from biologically active riboswitches and small ribozymes. When applied to the neurotransmitter precursors 5-hydroxytryptophan and 3,4-dihydroxyphenylalanine, this approach yielded easily identifiable and characterizable aptamers predisposed for coupling to readout domains to allow engineering of nucleic acid-sensory devices that function in vitro and in the cellular context.
Riboswitches are mRNA elements regulating gene expression in response to direct binding of a metabolite. While these RNAs are increasingly well understood with respect to interactions between receptor domains and their cognate effector molecules, little is known about the specific mechanistic relationship between metabolite binding and gene regulation by the downstream regulatory domain. Using a combination of cell-based, biochemical, and biophysical techniques, we reveal the specific RNA architectural features enabling a cobalamin-dependent hairpin loop docking interaction between receptor and regulatory domains. Furthermore, these data demonstrate that docking kinetics dictate a regulatory response involving the coupling of translation initiation to general mechanisms that control mRNA abundance. These results yield a comprehensive picture of how RNA structure in the riboswitch regulatory domain enables kinetically constrained ligand-dependent regulation of gene expression.
RNA folding in vivo is significantly influenced by transcription, which is not necessarily recapitulated by Mg2+-induced folding of the corresponding full-length RNA in vitro. Riboswitches that regulate gene expression at the transcriptional level are an ideal system for investigating this aspect of RNA folding as ligand-dependent termination is obligatorily co-transcriptional, providing a clear readout of the folding outcome. The folding of representative members of the SAM-I family of riboswitches has been extensively analyzed using approaches focusing almost exclusively upon Mg2+ and/or S-adenosylmethionine (SAM)-induced folding of full-length transcripts of the ligand binding domain. To relate these findings to co-transcriptional regulatory activity, we have investigated a set of structure-guided mutations of conserved tertiary architectural elements of the ligand binding domain using an in vitro single-turnover transcriptional termination assay, complemented with phylogenetic analysis and isothermal titration calorimetry data. This analysis revealed a conserved internal loop adjacent to the SAM binding site that significantly affects ligand binding and regulatory activity. Conversely, most single point mutations throughout key conserved features in peripheral tertiary architecture supporting the SAM binding pocket have relatively little impact on riboswitch activity. Instead, a secondary structural element in the peripheral subdomain appears to be the key determinant in observed differences in regulatory properties across the SAM-I family. These data reveal a highly coupled network of tertiary interactions that promote high-fidelity co-transcriptional folding of the riboswitch but are only indirectly linked to regulatory tuning.
Riboswitches represent a family of highly structured regulatory elements found primarily in the leader sequences of bacterial mRNAs. They function as molecular switches capable of altering gene expression; commonly, this occurs via a conformational change in a regulatory element of a riboswitch that results from ligand binding in the aptamer domain. Numerous studies have investigated the ligand binding process, but little is known about the structural changes in the regulatory element. A mechanistic description of both processes is essential for deeply understanding how riboswitches modulate gene expression. This task is greatly facilitated by studying all aspects of riboswitch structure/dynamics/function in the same model system. To this end, single-molecule fluorescence resonance energy transfer (smFRET) techniques have been used to directly observe the conformational dynamics of a hydroxocobalamin (HyCbl) binding riboswitch (env8HyCbl) with a known crystallographic structure.1 The single-molecule RNA construct studied in this work is unique in that it contains all of the structural elements both necessary and sufficient for regulation of gene expression in a biological context. The results of this investigation reveal that the undocking rate constant associated with the disruption of a long-range kissing-loop (KL) interaction is substantially decreased when the ligand is bound to the RNA, resulting in a preferential stabilization of the docked conformation. Notably, the formation of this tertiary KL interaction directly sequesters the Shine-Dalgarno sequence (i.e., the ribosome binding site) via base-pairing, thus preventing translation initiation. These results reveal that the conformational dynamics of this regulatory switch are quantitatively described by a four-state kinetic model, whereby ligand binding promotes formation of the KL interaction. The results of complementary cell-based gene expression experiments conducted in Escherichia coli are highly correlated with the smFRET results, suggesting that KL formation is directly responsible for regulating gene expression.
The crystal structures of two different cobalamin (vitamin B12)-binding riboswitches are determined; the structures reveal how cobalamin facilitates interdomain interactions to regulate gene expression. Small metabolites and ligands can affect gene expression by binding to a structured part of an RNA known as a riboswitch. Although the structures of many riboswitch receptor domains have been solved, the complete riboswitch structure with regulatory domain had not been determined. Robert Batey and colleagues have now solved the structure of two different cobalamin (vitamin B12) riboswitches that include the downstream regulatory domain. Ligand recognition occurs largely as a result of shape complementarity, rather than the more typical hydrogen bonding. Structures of riboswitch receptor domains bound to their effector have shown how messenger RNAs recognize diverse small molecules, but mechanistic details linking the structures to the regulation of gene expression remain elusive1,2. To address this, here we solve crystal structures of two different classes of cobalamin (vitamin B12)-binding riboswitches that include the structural switch of the downstream regulatory domain. These classes share a common cobalamin-binding core, but use distinct peripheral extensions to recognize different B12 derivatives. In each case, recognition is accomplished through shape complementarity between the RNA and cobalamin, with relatively few hydrogen bonding interactions that typically govern RNA–small molecule recognition. We show that a composite cobalamin–RNA scaffold stabilizes an unusual long-range intramolecular kissing-loop interaction that controls mRNA expression. This is the first, to our knowledge, riboswitch crystal structure detailing how the receptor and regulatory domains communicate in a ligand-dependent fashion to regulate mRNA expression.