Kawasaki disease (KD) and viral infection (VD) share similar clinical features but require distinct treatments. A practical biomarker to distinguish them is therefore clinically important. Previous blood transcriptome studies identified many differentially expressed genes, but large gene panels are impractical for routine laboratories. A ratio-based Direct Leukocyte Single-cell-type Transcript Abundance (DIRECT LS-TA) assay was recently developed to quantify monocyte gene expression directly in whole blood using a monocyte-specific target gene relative to monocyte reference genes (PSAP or CTSS). Interferon-stimulated genes IFI27, IFI44L, and SIGLEC1 can be measured by this approach. In this study, three ratio biomarkers (IFI27/PSAP, IFI44L/PSAP, SIGLEC1/PSAP) and a conventional interferon (IFN) score derived from eight genes were calculated from public blood transcriptome datasets (GSE73461 and GSE68004) and compared between KD and VD. VD patients showed markedly elevated IFN-related biomarkers, with all three ratios significantly higher in VD and IFI27/PSAP giving the largest increase. IFI27/PSAP achieved the highest diagnostic performance (AUC 0.90), slightly exceeding the conventional IFN score (AUC 0.89). These findings suggest that absent or minimal IFN activation argues against VD and supports KD, and that this simple ratio assay could serve as a clinically useful exclusion test.
A genome sequence is not made up of random nucleotides. Instead, it has distinctive features for evolutionary adaptation. Herein, we identified a higher frequency of continuous guanines (G-runs) located downstream of the transcription start site (TSS) in the non-template strand than that of the template strand by analyzing the genomic region around TSS (TSS ± 1 kb) across different species. G-runs are known to have the propensity to form G-quadruplex structures (G4). By integrative analysis of large-scale multi-omic datasets, predicted G4 structures in TSS downstream region (TSS-to-+300 bp) of the non-template strand were found to be associated with lower promoter DNA methylation, more accessible chromatin, and higher transcript levels. Compared to non-cancer genes, tumor-suppressor genes (TSGs) exhibited a higher G-run frequency in TSS downstream region (TSS-to-+300 bp) of the non-template strand, contributing to both their high transcript levels and resistance to tumor-specific downregulation. These results were successfully validated with independent datasets. Taken together, our study reveals an evolutionarily conserved higher G-run frequency in TSS downstream region (TSS-to-+300 bp) of the non-template strand as a beneficial genetic feature of TSGs for optimizing their function in tumor suppression.
The intestinal epithelium maintains host-microbiota homeostasis, while inflammatory conditions, such as inflammatory bowel disease (IBD), induce pathological shifts in intestinal epithelial cell (IEC) subtypes. We unveil METTL3, an RNA m6A methyltransferase, as a pivotal regulator of this balance. METTL3 is enriched in intestinal stem cells and transit-amplifying cells (TACs), and upregulated in patients with IBD and a mouse model of IBD. DSS-challenged, intestine-specific Mettl3 knockout mice exhibited exacerbated colitis as exemplified by more weight loss, elevated disease activity index (DAI), and higher extent of colon shortening. Single-cell transcriptomics of colonic tissues from DSS-challenged intestine-specific Mettl3 knockout mice revealed that Mettl3 ablation depleted epithelial lineages (TACs, goblet cells, enterocytes) but amplified immune infiltration (macrophages, neutrophils, T cells) within the intestinal mucosa. Crucially, METTL3 loss impaired TAC multipotency and increased epithelial-neutrophil crosstalk mediated by the TNF pathway. Mechanistically, METTL3-mediated m6A modification increases Slc39a8 expression, whose knockdown in colon organoids phenocopied METTL3 deficiency in impairing self-renewal. Our work establishes METTL3 as a dual guardian of intestinal homeostasis-preserving epithelial regeneration and restraining inflammation by calibrating epithelial-immune dialogue. The former is mediated at least in part by regulating Slc39a8 expression through m6A modification.
A rapid method for triaging febrile patients by aetiology (e.g., viral or bacterial infection) using gene expression in peripheral blood (PB) is an intensively researched area. However, gene expression in blood represents a composite sum of gene expression of all the component cell types present in the sample. As a result, numerous genes are measured in most proposed signatures. Herein, we propose a simple ratio-based biomarker (RBB) called direct leukocyte subpopulation-transcript abundance assay (DIRECT LS-TA) that recapitulates gene expressions of a single cell type in PB (i.e., monocytes). Based on single-cell RNA sequencing (scRNAseq) data and bulk expression data, IFI27 and SIGLEC1 are found as interferon-stimulated genes (ISGs) predominantly expressed by monocytes. The DIRECT LS-TA method can use a simple ratio of two genes measured in PB as an RBB to represent the target gene expression in monocytes without the need for monocyte purification. Both scRNAseq and bulk RNA sequencing datasets were used to evaluate the correlation between ISG expression in monocytes and PB, with a particular focus on monocyte expression of IFI27. An iceberg plot of bulk transcriptome data was used to identify genes that were predominantly expressed by monocytes in PB. DIRECT LS-TA RBBs of the three genes (IFI27, IFI44L and SIGLEC1) were evaluated by group-wise comparison, receiver operating characteristic and meta-analysis. In addition, the conventional interferon (IFN) score was evaluated for comparison of diagnostic performance. In viral infection datasets, DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) was most intensely activated (p value by t test <1e-9) and had the best area under the curve (0.94) among the three potential monocyte ISGs analysed. DIRECT LS-TA SIGLEC1 was also another monocyte biomarker but showed a lower activation (p<9e-5). IFI27/PSAP showed better diagnostic performance than the conventional IFN score. On the other hand, IFI44L was not a predominant monocyte expression gene. DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) measured in PB was the best biomarker of viral infection and IFN activation among ISGs predominantly expressed by monocytes. It performed even better than the conventional IFN score which required quantification of eight genes. The results suggest that DIRECT LS-TA of IFI27 is a monocyte-informative biomarker which is easy to determine in PB without the need for cell sorting.
BACKGROUND:The role of N1-methyladenosine (m1A) in cancer is poorly understood. Here we explored the function of RNA methyltransferase TRNA methyltransferase 61A (TRMT61A) in colorectal cancer (CRC) and its potential as a therapeutic target. METHODS:RNA m1A levels were assessed through liquid chromatography-mass spectrometry. The expression and clinical significance of TRMT61A were investigated across five human CRC cohorts. The function of TRMT61A was elucidated using CRC cell lines, patient-derived organoids, xenografts, and transgenic mouse models. Integrated analyses of m1A-sequencing and RNA-sequencing revealed the underlying mechanisms of TRMT61A. A nanoparticle-based small interfering RNA (siRNA) delivery system and a specific inhibitor were developed to target TRMT61A. The efficacy and safety of targeting TRMT61A were assessed. RESULTS:Our research revealed a consistent increase in TRMT61A expression and total RNA m1A levels within primary CRCs. High TRMT61A expression was associated with poor prognosis of CRC patients. Through CRISPR/Cas9 screenings, we identified TRMT61A as the most essential gene among m1A regulators. Furthermore, we established that TRMT61A promoted CRC tumorigenesis and progression by enhancing the mRNA stability of critical targets in an m1A-dependent manner. In particular, TRMT61A boosted the mRNA stability of one cut homeobox 2 (ONECUT2), which in turn triggered son of sevenless homolog 1 (SOS1) transcription, leading to the induction of mitogen-activated protein kinase (MAPK)/extracellular signal-regulated kinase (ERK) signaling in CRC. Notably, our study underscored the safety and substantial anti-CRC effects achievable by inhibiting TRMT61A using nanoparticle-encapsulated siTRMT61A or our newly discovered small molecule compound, pentagalloylglucose. CONCLUSIONS:Our study unveiled the tumor-promoting role of TRMT61A in CRC via the m1A-ONECUT2-SOS1-MAPK/ERK pathway. Targeting TRMT61A showed promise as a therapeutic strategy for treating CRC.
Spearman's correlation between evolutionary distance from humans and concordance scores of SD-associated TSGs across 32 non-human species.
Size-adjusted comparison of deletion, structural variation, and mutation frequencies in TSGs with SDs versus TSGs without SDs using multivariate linear regression.
Positive correlation between gene size and missense mutation frequency in oncogenes with SDs across 33 TCGA cancers.
Positive correlation between gene size and nonsynonymous/frameshift mutation frequency in TSGs with SDs across 33 cancers.
Gene size comparisons between with SDs and without SDs genes for TSGs and oncogenes.
Results of Fisher's exact tests comparing the number of SDs association with oncogenes and tumor suppressor genes (TSGs) against non-cancer genes across multiple species.
Positive correlation between SD repeat number and homozygous deletion frequency in SD-containing TSGs across 33 TCGA cancers.
A rapid method to triage febrile patients into different categories of etiologies remains a significant challenge even nowadays, when many molecular tests for pathogens are available. Routine serum protein tests like C-reactive protein and procalcitonin have limited specificity. Host response gene signatures are promising biomarkers but they usually require assaying many genes, e.g. 7 genes are commonly used to calculate the interferon (IFN) score. However, these gene panels fail to capture cell-type-specific host responses. Measuring gene expression of a specified single cell population, like monocytes, offers enhanced biological insight. However, it currently requires laborious cell sorting or costly single-cell sequencing techniques, limiting its clinical applicability. This study aims to develop a simple ratio-based biomarker (RBB) representing monocyte-specific host response to viral infection called DIRECT LS-TA method. A simple ratio of 2 genes (both are shortlist monocyte informative genes) quantified in peripheral blood (PB) samples correlated with gene expression in purified monocytes in the corresponding individual. These RBBs cover 3 interferon-stimulated genes (ISGs): IFI27/PSAP , IFI44L/PSAP and SIGLEC1/PSAP . They are compared to the conventional multi-gene IFN score in the differentiation of viral infection. Public gene expression datasets from NCBI GEO were used to shortlist monocyte-informative genes that can be used as the RBB in PB. The DIRECT LS-TA RBB was calculated as the ratio of the target ISG transcript abundance (TA) to that of another reference gene ( PSAP or CTSS) directly quantified from bulk PB data (e.g., Log( IFI27/PSAP ) in WB). The correlation (expressed by coefficient of determination, R²) between these DIRECT LS-TA RBBs and the gold-standard target gene TA measured in purified monocytes was assessed. The diagnostic performance of selected RBBs ( IFI27/PSAP , IFI44L/PSAP , SIGLEC1/PSAP) was compared against the conventional 8-gene IFN score for differentiating viral infections from controls. Direct LS-TA RBBs measured in PB showed strong correlation with gold-standard gene expression measured in purified monocytes (R2 ranged from 0.53 for the target gene IFI27 to >0.9 for the target gene IFI44L ). This high level of correlation supports that this simple RBB (DIRECT LS-TA) method can replace the tedious cell sorting approach to obtain single-cell-type gene expression data. All DIRECT LS-TA results of ISGs were raised during viral infection. The best clinical performance in triaging viral infection patients was achieved by IFI27/PSAP or IFI27/CTSS across all datasets. For example, in the GSE111368 dataset, IFI27/PSAP achieved an AUC of 0.94 (95% CI 0.90-0.97) with 88% sensitivity and 95% specificity, surpassing the IFN score’s AUC of 0.90 (95% CI 0.85-0.94) with 79% sensitivity and 93% specificity. Conclusion The DIRECT LS-TA method, utilizing the format of simple two-gene ratio-based biomarkers like IFI27/PSAP , provides a robust and accurate measure of monocyte-specific interferon pathway activation directly from peripheral blood. The superior performance of the DIRECT LS-TA method makes it a promising, readily implementable tool for clinical triage. Its ability to provide single-cell-type specific information, rapid turnaround using standard qPCR/dPCR technology, and enhanced biological specificity make it a valuable molecular host response assessment. ### Competing Interest Statement The authors declare the following potential conflict of interest. Nelson LS Tang is the inventor of the patent Determination of gene expression levels of a cell type which has been assigned to The Chinese University of Hong Kong. K.S. Leung and Nelson LS Tang are share-holders of Cytomics Ltd. Cytomics Ltd. holds a license to use a patent related to DIRECT LS-TA assay. Patent application pending
Odds ratios comparing SD association between oncogenes (including orthologs in 32 species) and non-cancer genes.
Background Non-translated transcripts (nt-RNAs) with frame-shifts or premature termination codons resulting from alternative splicing events (ASE), have been recently found at unexpectedly abundant in transcriptomes of cancer tissue. However, their full genomic spectrum has not yet been fully elucidated. This study comprehensively characterised the expression of signature junctions of these nt-RNA (termed “toxic junctions” here) of both known and novel nt-RNA across multiple cancer types and investigated their potential as biomarkers. Methods RNA-seq data of ∼6,000 samples, including the tumor and normal samples for 13 cancer types were retrieved from The Cancer Genome Atlas database (TCGA) together with data from Cancer Cell Line Encyclopedia (CCLE) project. Due to the difficulty in quantifying the entire transcript isoform of nt-RNA, we pioneered an algorithm to focus exclusively on the expression of junctional reads, which also circumvented the limitation of non-directional RNA- seq of TCGA data. We showed that the majority of nt-RNA is associated with at least one toxic junction. We built a comprehensive catalogue of known nt-RNA toxic junctions from genome databases. And novel toxic junctions were also identified by a new junction-focused algorithm from the higher quality discovery subsets of TCGA data. Splicing in Ratio (SiR) was used to quantify ASE leading to nt-RNA, enabling: Differential expression analysis between cancer and normal tissue and across cancer types. Identification of different profiles of nt-RNA abundance and various factor which may be the causes of differential nt-RNA abundance and SiR results Identification of specific nt-RNA and toxic junctions that were expressed in various cancer (and/or normal tissue) types. Assessment of nt-RNA and their toxic junction expression as biomarkers or prognosis indicators. Results We profiled the expressed known nt-RNA (toxic) junctions of known transcripts and discovered ∼22,000 novel toxic junctions out of ∼250,000 novel junctions found in the transcriptome data. The expression of nt-RNA was as high as 10% of all transcripts of the corresponding gene in cancer transcriptomes. Interestingly, some signature toxic junctions of nt-RNA are expressed in even higher quantities, e.g. up to 50% or more, which is reminiscent of a heterozygous mutation. We identified distinct patterns between cancer and normal samples, including example of nt-RNA expressing toxic junctions exclusively in normal or tumor samples. Clinically relevant examples included ANXA6 in breast cancer, where the nt-RNA isoform showed significantly higher expression in tumors (p=1.8e-15). In kidney renal clear cell carcinoma (KIRC), a significant isoform switch of ESYT2 based on the RNA-seq data was confirmed. The Kaplan-Meier survival curves showed that samples with the higher expression ratio of ESYT2-L are associated with better survival (p=2.0e-06). Unsupervised clustering showed that SiR results of 150 toxic signatures defined 4 subgroups of patients with different prognosis. Through principal component analysis (PCA), PC1 and PC2 can be used as an independent prognosis biomarkers. nt-RNA accounting for these PCs included splicing factors SRSF3 and CLK1, where CLK1 phosphorylates SRSF3 to promote exon 4 inclusion in both genes. Conclusions In summary, the expression profiles of all known and novel toxic junctions were explored using pan-cancer RNA-seq data. A dual 10% rule emerged from this study: ∼10% of novel junctions were toxic junctions associated with nt-RNA, and up to 10% of RNA transcripts inside a cell were also nt-RNA. The SiR metric enables accurate quantification of unproductive splicing and identification of cancer biomarkers. Our findings reveal that unproductive splicing represents functionally important post-transcriptional regulation in cancer. These expression profiles allow researchers to study the expression of nt-RNA signature junctions or novel signature junctions in or near the genes they are interested in, which could provide a new direction for their research. The SRSF3-CLK1 regulatory mechanism provides insights into splicing dysregulation. Our comprehensive toxic junction catalogue serves as a valuable resource, suggesting that targeting unproductive splicing pathways may offer novel therapeutic strategies for cancer treatment. Data availability The catalogue is available on GitHub and UCSC browser. for GitHub overview [https://genome.ucsc.edu/s/dandan\_0909/hg38\_all\_new\_nr][1] for genome browsing of all novel (unannotated) toxic junctions [https://genome.ucsc.edu/s/dandan\_0909/hg38\_5_26][2] for toxic junctions in known (annotated) nt-RNA. ### Competing Interest Statement The authors have declared no competing interest. [1]: https://genome.ucsc.edu/s/dandan_0909/hg38_all_new_nr [2]: https://genome.ucsc.edu/s/dandan_0909/hg38_5_26
Results from multivariate linear regression of somatic non-CNA SV mutation frequency of TSGs with SDs versus TSGs without SDs.
No significant difference in ionizing radiation-induced SV between with SDs and without SDs genes across normal human cell colonies.
Boxplot showing the distance distribution of non-cancer genes, oncogenes, and TSGs to their nearest SDs, with individual genes represented as dots and significant differences marked (p < 0.05).