Diffuse midline gliomas (DMG) are deadly pediatric brain cancers with limited treatment options. These tumors likely arise from oligodendrocyte precursor cells (OPC) that acquire a driver histone mutation, leading to an aberrant epigenome. RNA N6-methyladenosine (m6A) is a vital epi-transcriptomic modification that regulates RNA processes and plays a significant role in OPC development through its regulation of transcripts involved in histone modification processes. Despite this pivotal role in OPC biology, the epi-transcriptome has not yet been investigated in DMG, and its interrogation may uncover new therapeutic options and understanding of this disease. Therefore, for the first time, we generated base-resolution m6A landscapes for patient-derived DMG cultures and found that DMG exhibits elevated m6A levels compared to non-neoplastic patient cells, with particularly strong enrichment on transcripts involved in cell motility and migration. In contrast, the minority of transcripts that have lower levels of m6A in DMG were associated with cell cycle regulation, especially components of chromosome segregation machinery. We also demonstrate that DMG is sensitive to inhibition of the m6A demethylase FTO, with FB23-2 treatment resulting in decreased proliferation, reduced survival, and pronounced S-phase arrest/stress, accompanied by robust induction of CDKN1A, GADD45B, and TFRC. Furthermore, FTO inhibition led to significant downregulation of key cell cycle regulators at both the transcriptomic and proteomic levels. Collectively, these findings highlight RNA methylation as a critical regulator of DMG tumorigenicity and identify FTO as a promising therapeutic target for this currently incurable disease.
While MYCN-amplified neuroblastoma has been the focus of neuroblastoma research in the past three decades, most human neuroblastomas do not harbour MYCN oncogene amplification, and their tumorigenic factors are unknown. Long non-coding RNAs (lncRNAs) regulate tumorigenesis by modulating the expression of molecular targets, however, there is limited literature on therapeutic targeting of lncRNAs with small molecule compounds. To determine the oncogenic mechanism though which the lncRNA lncNB promotes neuroblastoma cell proliferation, survival and tumour progression; and to identify a small molecule compound that inhibits the interaction between lncNB and its binding protein MSI2 as an effective anticancer strategy. Kaplan Meier analysis showed that high levels of lncNB expression in neuroblastoma tissues correlated with poor prognosis in 476 patients. We identified lncNB as the lncRNA most over-expressed in MYCN non-amplified, compared with MYCN-amplified, neuroblastoma cell lines. lncNB expression was controlled by super-enhancers, and lncNB RNA bound to MSI2 protein. RNA immunoprecipitation and sequencing identified BMX mRNA as the transcript most significantly disrupted from binding to MSI2 protein, after lncNB knockdown. lncNB or MSI2 knockdown reduced, while their over-expression enhanced, BMX mRNA stability and expression, ERK protein phosphorylation and MYCN non-amplified neuroblastoma cell proliferation. lncNB knockdown significantly suppressed neuroblastoma progression in mice. AlphaScreen of a compound library identified NSC617570 as an efficient inhibitor of lncNB RNA and MSI2 protein interaction, and NSC617570 reduced BMX expression, ERK protein phosphorylation, neuroblastoma cell proliferation in vitro and tumor progression in mice. Our study demonstrates that lncNB RNA interacts with MSI2 protein to induce neuroblastoma tumorigenesis, and that targeting lncNB and MSI2 interaction with small molecule compounds is an effective anticancer strategy. Sujanna Mondal, Pei Y. Liu, Janith Seneviratne, Antoine De Weck, Pooja Venkat, Chelsea Mayoh, Jing Wu, Jesper Maag, Jingwei Chen, Matthew Wong, Nenad Bartonicek, Poh Khoo, Lei Jin, Louise E. Ludlow, David S. Ziegler, Toby Trahair, Pieter Mestdagh, Belamy B. Cheung, Jinyan Li, Marcel E. Dinger, Ian Street, Xu D. Zhang, Glenn M. Marshall, Tao Liu. The super enhancer-driven long noncoding RNA lncNB promotes neuroblastoma tumorigenesis by interacting with MSI2 protein and is targetable by small molecule compounds [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2608.
Tumorigenic drivers of MYCN gene nonamplified neuroblastoma remain largely uncharacterized. Long noncoding RNAs (lncRNAs) regulate tumorigenesis, however, there is little literature on therapeutic targeting of lncRNAs with small molecule compounds. Here PRKCQ-AS1 is identified as the lncRNA most overexpressed in MYCN nonamplified, compared with MYCN-amplified, neuroblastoma cell lines. PRKCQ-AS1 expression is controlled by super-enhancers, and PRKCQ-AS1 RNA bound to MSI2 protein. RNA immunoprecipitation and sequencing identified BMX mRNA as the transcript most significantly disrupted from binding to MSI2 protein, after PRKCQ-AS1 knockdown. PRKCQ-AS1 or MSI2 knockdown reduces, while its overexpression enhances, BMX mRNA stability and expression, ERK protein phosphorylation and MYCN nonamplified neuroblastoma cell proliferation. PRKCQ-AS1 knockdown significantly suppresses neuroblastoma progression in mice. In human neuroblastoma tissues, high levels of PRKCQ-AS1 and MSI2 expression correlate with poor patient outcomes, independent of current prognostic markers. AlphaScreen of a compound library identifies NSC617570 as an efficient inhibitor of PRKCQ-AS1 RNA and MSI2 protein interaction, and NSC617570 reduces BMX expression, ERK protein phosphorylation, neuroblastoma cell proliferation in vitro and tumor progression in mice. The study demonstrates that PRKCQ-AS1 RNA interacts with MSI2 protein to induce neuroblastoma tumorigenesis, and that targeting PRKCQ-AS1 and MSI2 interaction with small molecule compounds is an effective anticancer strategy.
We develop GeneRAIN, a suite of Transformer-based models that learn gene expression relationships from 410 K human bulk RNA-seq samples. Featuring a novel Binning-By-Gene normalization technique, our models capture diverse biological information beyond expression. We introduce GeneRAIN-vec, a multifaceted vectorized gene representation that outperforms those from existing models. We demonstrate knowledge transfer from protein-coding genes to Make 62.5 million biological attribute predictions for 13,030 long noncoding RNAs. This work advances Transformer and self-supervised deep learning applications to expression data, enhancing biological exploration.
Abstract Diffuse Midline Glioma (DMG) is an incurable pediatric brain tumor thought to originate from oligodendrocyte precursor cells in midline brain structures. The RNA modification N6-methyladenosine (m6A) plays an important role in RNA stability and is critical to neuronal stem-cell self-renewal and differentiation. We therefore sought to investigate m6A as a therapeutic target in DMG. Moreover, targeting the epitranscriptome has shown promise in the treatment of other cancers, and several small-molecule inhibitors of m6A writers and erasers have been recently developed. To that end, we tested the sensitivity of a panel of patient-derived DMG cell lines to FB23-2, an inhibitor of the m6A eraser FTO, and STM2457, an inhibitor of the mRNA m6A writer METTL3. In order to interrogate the therapeutic mechanisms and identify predictive biomarkers for response, we then performed RNA-seq to measure gene expression changes, and native RNA-seq to measure RNA m6A levels. We found that DMG cell lines were more sensitive to FB23-2 (IC50 ~10μM) than STM2457 (IC50 ~100μM), indicating that m6A gain rather than loss may be a potential therapeutic strategy. Furthermore, we observed variation in FB23-2 sensitivity between cell lines with RNA-sequencing identifying marked changes in the expression of cell cycle, cell stress response, and differentiation associated genes in the cell lines which responded best to FB23-2. Additionally, these analyses also identified several potential biomarkers for response to FB23-2. Finally, we generated the first m6A transcriptome maps for DMG using native RNA sequencing and demonstrated that m6A is abundant in DMG cell lines and is particularly enriched on cell cycle, cell stress response, and metabolic pathway transcripts. Overall, our work identified the FTO m6A demethylase as a potential therapeutic target in DMG, whereby its inhibition with FB23-2 increases m6A abundance, altering the stability of critical transcripts and pathways. Citation Format: Samuel E. Ross, Holly Holliday, Maria Tsoli, David S. Ziegler, Marcel E. Dinger. RNA N6-methyladenosine (m6A) as a therapeutic target in Diffuse Midline Glioma (DMG) [abstract]. In: Proceedings of the AACR Special Conference on Brain Cancer; 2023 Oct 19-22; Minneapolis, Minnesota. Philadelphia (PA): AACR; Cancer Res 2024;84(5 Suppl_1):Abstract nr B015.
Motivation Mitochondrial diseases (MDs) are the most common group of inherited metabolic disorders and are often challenging to diagnose due to extensive genotype-phenotype heterogeneity. MDs are caused by mutations in the nuclear or mitochondrial genome, where pathogenic mitochondrial variants are usually heteroplasmic and typically at much lower allelic fraction in the blood than affected tissues. Both genomes can now be readily analysed using unbiased whole genome sequencing (WGS), but most nuclear variant detection methods fail to detect low heteroplasmy variants in the mitochondrial genome. Results We present mity , a bioinformatics pipeline for detecting and interpreting heteroplasmic SNVs and INDELs in the mitochondrial genome using WGS data. In 2,980 healthy controls, we observed on average 3,166× coverage in the mitochondrial genome using WGS from blood. mity utilises this high depth to detect pathogenic mitochondrial variants, even at low heteroplasmy. mity enables easy interpretation of mitochondrial variants and can be incorporated into existing diagnostic WGS pipelines. This could simplify the diagnostic pathway, avoid invasive tissue biopsies and increase the diagnostic rate for MDs and other conditions caused by impaired mitochondrial function. Availability mity is available from https://github.com/KCCG/mity under an MIT license. Contact clare.puttick@crick.ac.uk , carolyn.sue@sydney.edu.au , MCowley@ccia.org.au
Mice are widely used as animal models in biomedical research, favored for their small size, ease of breeding, and anatomical and physiological similarities to humans[1][1],[2][2]. However, discrepancies between mouse gene experimental results and the actual behavior of human genes are not uncommon, despite their shared DNA sequence similarity[3][3]–[8][4]. This suggests that DNA sequence similarity does not always reliably predict functional similarity. On the other hand, RNA-level gene expression could offer additional information about gene function[9][5],[10][6]. In this study, we undertook characterization and inter-species comparison of human and mouse genes by applying innovative deep learning methodologies to a large dataset of 410K human and 366K mouse bulk RNA-seq samples. This was achieved by using gene representations from our Transformer-based GeneRAIN model[11][7],[12][8]. These gene representations aggregate information from large gene expression datasets, and provide insights beyond DNA sequence similarity. We identified 2,407 human-mouse homologous genes with high DNA similarity but distinct RNA characteristics, and showed that these genes are more likely to have differing disease/phenotype associations between the two species. Additionally, we found 3,070 homologous genes with low similarity at both the DNA and RNA levels, suggesting the highest risk of discrepancies in study results between the two species. We propose that this approach will support future decision making around whether the mouse will be an appropriate model for studying specific human genes, and whether the results of specific mouse gene studies are likely to be recapitulated in humans. Our methodological innovations offer valuable lessons for future deep learning applications in cross-species omics data. The interspecies gene relationship findings from our study also contribute valuable insights into the gene biology and evolution of the two species. ### Competing Interest Statement F.V. declares commercial association with OmniOmics.AI Pty Ltd. [1]: #ref-1 [2]: #ref-2 [3]: #ref-3 [4]: #ref-8 [5]: #ref-9 [6]: #ref-10 [7]: #ref-11 [8]: #ref-12
[This corrects the article DOI: 10.3389/fmolb.2021.665199.].
Gene expression regulation is a sophisticated, multi-stage process, and its robustness is critical to normal cell function and the survival of an organism. Previous studies indicate that differential gene expression at the RNA level is typically attenuated at the protein level through translational regulation. However, how post-transcriptional regulation (PTR) influences expression change during the RNA maturation process remains unclear. In this study, we investigated this by quantifying the magnitude of expression change in precursor RNA and mature RNA across a vast range of different biological conditions. We analyzed bulk tissue RNA sequencing data from 4689 samples, including healthy and diseased tissues from human, chimpanzee, rhesus macaque, and murine sources. We demonstrated that PTR tends to support homeostatic expression of mature RNA by amplifying normal tissue-specific expression of precursor RNA, while reducing expression change of precursor RNA in disease contexts. Our study provides insight into the general influence of PTR on gene expression homeostasis. Our analysis also suggests that intronic reads in RNA-seq studies may contain under-utilized information about disease associations. Additionally, our findings may assist in identifying new disease biomarkers and more effective ways of altering gene expression as a therapeutic strategy.
DNA i-motif structures are formed in the nuclei of human cells and are believed to provide critical genomic regulation. While the existence, abundance, and distribution of i-motif structures in human cells has been demonstrated and studied by immunofluorescent staining, and more recently NMR and CUT&Tag, the abundance and distribution of such structures in human genomic DNA have remained unclear. Here we utilise high-affinity i-motif immunoprecipitation followed by sequencing to map i-motifs in the purified genomic DNA of human MCF7, U2OS and HEK293T cells. Validated by biolayer interferometry and circular dichroism spectroscopy, our approach aimed to identify DNA sequences capable of i-motif formation on a genome-wide scale, revealing that such sequences are widely distributed throughout the human genome and are common in genes upregulated in G0/G1 cell cycle phases. Our findings provide experimental evidence for the widespread formation of i-motif structures in human genomic DNA and a foundational resource for future studies of their genomic, structural, and molecular roles.
Genes specifying long non-coding RNAs (lncRNAs) occupy a large fraction of the genomes of complex organisms. The term 'lncRNAs' encompasses RNA polymerase I (Pol I), Pol II and Pol III transcribed RNAs, and RNAs from processed introns. The various functions of lncRNAs and their many isoforms and interleaved relationships with other genes make lncRNA classification and annotation difficult. Most lncRNAs evolve more rapidly than protein-coding sequences, are cell type specific and regulate many aspects of cell differentiation and development and other physiological processes. Many lncRNAs associate with chromatin-modifying complexes, are transcribed from enhancers and nucleate phase separation of nuclear condensates and domains, indicating an intimate link between lncRNA expression and the spatial control of gene expression during development. lncRNAs also have important roles in the cytoplasm and beyond, including in the regulation of translation, metabolism and signalling. lncRNAs often have a modular structure and are rich in repeats, which are increasingly being shown to be relevant to their function. In this Consensus Statement, we address the definition and nomenclature of lncRNAs and their conservation, expression, phenotypic visibility, structure and functions. We also discuss research challenges and provide recommendations to advance the understanding of the roles of lncRNAs in development, cell biology and disease.
MOTIVATION:Methods for concept recognition (CR) in clinical texts have largely been tested on abstracts or articles from the medical literature. However, texts from electronic health records (EHRs) frequently contain spelling errors, abbreviations, and other nonstandard ways of representing clinical concepts. RESULTS:Here, we present a method inspired by the BLAST algorithm for biosequence alignment that screens texts for potential matches on the basis of matching k-mer counts and scores candidates based on conformance to typical patterns of spelling errors derived from 2.9 million clinical notes. Our method, the Term-BLAST-like alignment tool (TBLAT) leverages a gold standard corpus for typographical errors to implement a sequence alignment-inspired method for efficient entity linkage. We present a comprehensive experimental comparison of TBLAT with five widely used tools. Experimental results show an increase of 10% in recall on scientific publications and 20% increase in recall on EHR records (when compared against the next best method), hence supporting a significant enhancement of the entity linking task. The method can be used stand-alone or as a complement to existing approaches. AVAILABILITY AND IMPLEMENTATION:Fenominal is a Java library that implements TBLAT for named CR of Human Phenotype Ontology terms and is available at https://github.com/monarch-initiative/fenominal under the GNU General Public License v3.0.
Supplementary Methods from The Melanoma‐Upregulated Long Noncoding RNA SPRY4-IT1 Modulates Apoptosis and Invasion
Intercalated motifs or i-Motifs (iMs) are nucleic acid structures formed by cytosine-rich sequences, which may regulate cellular processes and have broad applications in nanotechnology due to their pH-dependent nature. We have developed an iM-specific nanobody (iMbody) that can recognize iM DNA structures regardless of their sequences, making it a versatile research tool for studying iMs in various contexts. Here, we provide a protocol for the bacterial expression and His-tag purification of iMbody. We then describe procedures for performing ELISA and immunostaining using iMbody.
Supplementary figure 4. Esophageal adenocarcinoma cell line viability changes after shRNA knockdown of 21 genes identified through network and machine learning analysis, and Representative IHC staining slides of the 4-gene signature. (a) Two esophageal cancer cell lines, JHESOAD1 and OE33 were interrogated through the Broad Institute Project Achilles for changes in viability after pooled shRNA mediated gene knockdown. Y-axis measures relative change in cell viability. (b) Comparison of the effect of knockdowns of genes of interest on different GI cell types from the Achilles data base. (c) Representative slides from the IHC staining between NE, NDBE, and EAC, NDBE and EAC slides from COL17A1 and from E2F3 are the same as in Figure 3f.
Supplementary Video 2 from The Melanoma‐Upregulated Long Noncoding RNA SPRY4-IT1 Modulates Apoptosis and Invasion
Supplementary figure 2. Differentially expressed transcriptions factors between LGD and NDBE, and examples of EAC specific genes and eigen-modules from WGCNA. (a) Box plots of the top transcription factors (FOSB, FOSB, EGR1, EGR3, NR4A1, ATF3) differentially expressed between LGD and BE and the corresponding expression in Nsq and EAC. y-axis is log2(cpm 1). (b) Boxplot (top) and IGV coverage plot (bottom) of two examples of highly connected, and also differentially expressed gens between both EAC and NDBE/LGD in the brown eigenmodule. (c, d ) (left) Heatmap (top), median eigenexpression (bottom), (right) protein-protein-interaction network (top), and network gene ontology enrichment (bottom) of eigenmodules found to be either over- (c) or underexpressed (d) in EAC. Genes are coloured based on differential expression, of either EAC and NDBE/LGD (red), EAC vs NDBE (blue), EAC vs LGD (green).
Tissue purity, RNA-seq based correlation analysis, and gene ontology enrichment. ESTIMATE (see methods) was used for in silico assessment of the (a) tissue purity, (c) stromal contamination, and (b) immune infiltration of each patient from the Normal esophagus (NE), Non-dysplastic Barrett's esophagus (NDBE), Low-grade dysplasia (LGD) and Esophageal adenocarcinoma (EAC) group. (d) Scatterplot comparing the mean expression (log2(cpm 1)) of all Gencode genes between each condition. Spearman correlation (R) shows the correlation score. (e) Venn diagram of the differentially expressed (absolute log2FC{greater than or equal to}1, FDR
Predicting the impact of coding and noncoding variants on splicing is challenging, particularly in non-canonical splice sites, leading to missed diagnoses in patients. Existing splice prediction tools are complementary but knowing which to use for each splicing context remains difficult. Here, we describe Introme, which uses machine learning to integrate predictions from several splice detection tools, additional splicing rules, and gene architecture features to comprehensively evaluate the likelihood of a variant impacting splicing. Through extensive benchmarking across 21,000 splice-altering variants, Introme outperformed all tools (auPRC: 0.98) for the detection of clinically significant splice variants. Introme is available at https://github.com/CCICB/introme .
Supplementary Figures 1-5 from The Melanoma‐Upregulated Long Noncoding RNA SPRY4-IT1 Modulates Apoptosis and Invasion