Small RNAs play essential roles in gene regulation across diverse biological processes. Crosslinking, ligation, and sequencing of hybrids (CLASH) experiments have revealed that PIWI and Argonaute proteins can each bind a wide range of mRNA targets with distinct base-pairing rules, raising questions about the flexibility and functional relevance of these interactions. Given that crosslinking-induced mutations (CIMs) provide single-nucleotide resolution molecular footprints of RNA-binding proteins, we developed MUTACLASH, a bioinformatics tool for systematically analyzing CIMs in CLASH data sets. Our analyses indicate that CIMs function as molecular footprints of Argonaute binding on target mRNAs. Specifically, for Caenorhabditis elegans miRNA and piRNA CLASH data, CIMs are enriched at the center of small RNA binding sites, as well as at nucleotides within mRNA target sites that exhibit local mismatches in piRNA interactions. Furthermore, we show that mRNAs with noncanonical miRNA and piRNA binding sites and/or low hybrid abundance marked by CIMs exhibit stronger regulatory effects than those without CIMs, demonstrating the utility of CIM analysis in identifying functional small RNA binding sites, including those that are otherwise likely overlooked with current analysis tools.
Abstracts Membraneless organelles (MLOs) often exhibit internal architecture, yet whether the local transcriptome differentially partitions across MLO subdomains remains largely uncharacterized. Here we combine super-resolution imaging with in situ reverse transcription-based sequencing to profile transcriptomes within MLO subdomains. Using the human tripartite nucleolus as a model system, we identify distinct RNA populations in the fibrillar center (FC), dense fibrillar component (DFC), and granular component (GC). Pre-rRNA processing intermediates demonstrate a layered progression across nucleolar subdomains, reflecting the temporal order of the processing steps. Processing steps involved in large-small subunit separation show increased retention in the DFC in highly differentiated cells. Mature small nucleolar RNAs (snoRNAs) are preferentially enriched in the DFC and spatially segregated from their precursor transcripts. Many non-snoRNA-related transcripts, often derived from nucleolus-proximal genes, show modest enrichment in the GC. These results illustrate functional RNA organization across nucleolar subdomains and provide a framework for nanoscale transcriptome mapping of biomolecular condensates.
Per- and polyfluoroalkyl substances (PFAS) represent a critical class of persistent environmental contaminants with significant ecological and human health implications. However, the rapid emergence of novel PFAS has far outpaced the development of reference mass spectral databases. Here, Neural Per- and Polyfluoroalkyl Substances Mass Spectrometry (NPFAS-MS), a transfer learning-based neural network model, was developed to predict PFAS-specific high-resolution mass spectra. NPFAS-MS was fine-tuned from a pretrained model using PFAS tandem mass (MS/MS) spectra. NPFAS-MS outperformed other in silico spectral prediction models for PFAS spectra prediction across multiple spectral similarity metrics. In library searching tasks, libraries generated by other spectral prediction models showed top-1 recall between 42.1% and 55.4%, while NPFAS-MS demonstrated 71.1%. Applying the virtual PFAS mass spectral library generated with NPFAS-MS using 10,553 PFAS structures from the U.S. EPA and NORMAN databases to groundwater and aqueous film-forming foam (AFFF) samples revealed more potential PFAS than other mass spectral databases. Specifically, 38 potential PFAS were annotated in AFFF products and 40 in groundwater samples. NPFAS-MS enabled characterization of emerging PFAS, including ultrashort-chain, unsaturated, and substituted derivatives in environmental matrices. This advancement enables comprehensive environmental monitoring of rapidly evolving PFAS contamination. NPFAS-MS and associated resources were deployed as a web-based tool at https://cosbi10.ee.ncku.edu.tw/NPFAS_MS/, enabling both structure-to-spectrum prediction and library searching against 31,659 predicted PFAS spectra.
The cost of sequencing a genome has become affordable for many research groups. However, with the growing number of sequenced genomes from non-model organisms, manually building functional genome annotation knowledge databases for each species is no longer feasible. To address this, we developed NoAC (Non-model Organism Atlas Constructor), a web tool that automatically constructs knowledge bases and query interfaces for non-model organism genomes without programming skills. In NoAC, users simply upload the gene or transcript information of a given non-model organism genome and select an appropriate reference model organism. NoAC then identifies orthologous genes, infers functional annotations, and sets up a searchable knowledge base. Functional annotations for the non-model organism such as gene ontology (GO) terms, protein domains, pathways, and physical/genetic interactors are predicted and transferred from the reference organism to the target genome. In an example non-model organism Phalaenopsis equestris, NoAC associates functional annotations for more than half of its 21,938 genes. Through case studies of the non-model organism Phalaenopsis equestris, we demonstrated that the knowledge base constructed by NoAC can reveal key functional aspects of PeSEP2 and PaMLS, supporting the study of novel genes involved in flower development. Another case study on the gene Wnt-1 in Bicyclus anynana further illustrates the applicability of NoAC in investigating insect segmentation and morphogen activity, highlighting its broader utility across diverse taxonomic genomes. In summary, NoAC allows general researchers to study non-model organisms with minimal in silico barriers. NoAC and its user tutorial are freely available at https://github.com/cosbi-nckuee/NoAC/.
BACKGROUND:Orchids are well-known for their rich diversity of species as well as wide range habitats. Their floral structures are so unique in angiosperms that many of orchids are economically and culturally important in human society. Orchids pollination strategy and evolutionary trajectory are also fantastic human for centuries. Previously, OrchidBase was created not only for storage and management of orchid genomic and transcriptomic information including Apostasia shenzhenica, Dendrobium catenatum, Phalaenopsis equestris, and two species of Platanthera that belong to three different subfamilies of Orchidaceae, but explored orchid genetic sequences for their function. The OrchidBase offers an opportunity for the plant science community to compare orchid genomes and transcriptomes, and retrieve orchid sequences for further study. DESCRIPTION:Recently, three whole-genome sequences of the Epidendroideae species, Cymbidium sinense, C. ensifolium and C. goeringii, were sequenced de novo, assembled, and analyzed. In addition, the systemic transcriptomes of these three species have been established. We included these datasets to develop a new version of OrchidBase 6.0. Furthermore, four new analytical methods, namely regulation, updated transcriptome, advanced BLAST, and domain search, were developed for orchid genome analyses. CONCLUSION:OrchidBase 6.0 extended genetic information to that of eight orchid species and created new tools for an expanded community curation in response to the ever-increasing volume and complexity of data.
Allele-specific expression (ASE) analyses from RNA-Seq data provide quantitative insights into genomic imprinting and the genetic variants that affect transcription. Robust ASE analysis requires the integration of multiple computational steps, including read alignment, read counting, data visualization, and statistical testing—this complexity creates challenges for reproducibility, scalability, and ease of use. Here, we present ASE Toolkit (ASET), an end-to-end pipeline that streamlines SNP-level ASE data generation, visualization, and testing for parent-of-origin (PofO) effect. ASET includes a modular pipeline built with Nextflow for ASE quantification from short-read transcriptome sequencing reads, an R library for data visualization, and a Julia script for PofO testing. ASET performs comprehensive read quality control, SNP-tolerant alignment to reference genomes, read counting with allele and strand resolution, annotation with genes and exons, and estimation of contamination. In sum, ASET provides a complete and easy-to-use solution for molecular and biomedical scientists to identify and interpret patterns of ASE from RNA-Seq data.
In germ cells, small RNAs function as a defense system to silence invading RNAs like viruses and transposons to protect genome integrity. The ability of small RNAs to robustly silence diverse RNA sequences prompts the question of how endogenous mRNAs avoid this silencing. In C. elegans, small RNAs bound by the Argonaute CSR-1 protect endogenous mRNAs from silencing, while also fine-tuning a subset of these mRNAs. Here, we identify RNA Helicase A (RHA-1) as a key regulator of CSR-1 small RNA biogenesis and function in mRNA fine-tuning. RHA-1 localizes to germ granules dependent on EGO-1, which synthesizes CSR-1 small RNAs. We find RHA-1 promotes small RNA production from the 5' regions of mRNAs and small RNA sorting to CSR-1. Loss of RHA-1 leads to elevated CSR-1 target mRNA levels and compromised fertility. Our study highlights the importance of small RNA regulation, mediated by RHA-1, to protect endogenous gene expression programs and germ cell function.
MicroRNAs (miRNAs) can target messenger RNAs to control their degradation or translation repression effects. Therefore, identifying the target and binding sites of different miRNAs is essential for understanding miRNA functions. To investigate these interactions, researchers have employed the cross-linking, ligation, and sequencing of hybrids (CLASH-seq) and similar CLASH-like approaches to generate chimeric reads formed by miRNAs and their targeting segments. These chimeric reads allow for the direct extraction of both the miRNA-target gene pairs and their corresponding binding sites. Nevertheless, these studies lack user-friendly platforms for researchers to investigate these interactions efficiently, thus hindering scientists' ability to explore miRNA functions. To address this gap, we developed mirTarCLASH, a comprehensive database that deposits 502 061/322 707/224 452 unique hybrid reads from human/mouse/worm miRNA chimeric read-based experiments. In mirTarCLASH, the chimera analysis algorithm ChiRA and two distinct binding site inference tools, RNAup and miRanda, were adopted to facilitate the exploration of miRNA-target pairs derived from CLASH-like experiments. Compared with existing similar repositories, mirTarCLASH further enables several confidence evaluation filters with visualization functions for the extracted results. The results can be further refined based on the key properties of the miRNA targeting sites, including read depths, numbers of supporting algorithms, and cross-linking-induced mutations, to enhance confidence levels. In addition, these miRNA-binding sites are visually represented through an integrated transcript atlas. Finally, we demonstrated the biological applicability of mirTarCLASH via the well-characterized example interaction between cel-let-7-5p and lin-41 in Caenorhabditis elegans, showcasing the potential of mirTarCLASH to provide novel insights for subsequent experimental research designs. The constructed mirTarCLASH database is freely available at https://cosbi.ee.ncku.edu.tw/MirTarClash. Database URL: https://cosbi.ee.ncku.edu.tw/MirTarClash.
Small RNAs play critical roles in gene regulation in diverse processes across organisms. Crosslinking, ligation, and analyses of sequence hybrid (CLASH) experiments have shown PIWI and Argonaute proteins bind to diverse mRNA targets, raising questions about their functional relevance and the degree of flexibility in target recognition. As crosslinking-induced mutations (CIMs) provides nucleotide-resolution of RNA binding sites, we developed MUTACLASH to systematically analyze CIMs in piRNA and miRNA CLASH data in C. elegans . We found CIMs are enriched at the nucleotide positions of mRNA corresponding to the center of targeting piRNAs and miRNAs. Notably, CIMs are also enriched at nucleotides with local pairing mismatches to piRNA. In addition, distinct patterns of CIMs are observed between canonical and non-canonical base pairing interactions, suggesting that the worm PIWI Argonaute PRG-1 adopts distinct conformations for canonical vs. non-canonical interactions. Critically, non-canonical miRNA or piRNA binding sites with CIMs exhibit more regulatory effects than those without CIMs, demonstrating CIM analysis as a valuable approach in assessing functional significance of small RNA targeting sites in CLASH data. Together, our analyses reveal the landscapes of Argonaute crosslinking sites on mRNAs and highlight MUTACLASH as an advanced tool in analyzing CLASH data.
Wolfberry, also known as goji berry or Lycium barbarum, is a highly valued fruit with significant health benefits and nutritional value. For more efficient and comprehensive usage of published L. barbarum genomic data, we established the Wolfberry database. The utility of the Wolfberry Genome Database (WGDB) is highlighted through the Genome browser, which enables the user to explore the L. barbarum genome, browse specific chromosomes, and access gene sequences. Gene annotation features provide comprehensive information about gene functions, locations, expression profiles, pathway involvement, protein domains, and regulatory transcription factors. The transcriptome feature allows the user to explore gene expression patterns using transcripts per kilobase million (TPM) and fragments per kilobase per million mapped reads (FPKM) metrics. The Metabolism pathway page provides insights into metabolic pathways and the involvement of the selected genes. In addition to the database content, we also introduce six analysis tools developed for the WGDB. These tools offer functionalities for gene function prediction, nucleotide and amino acid BLAST analysis, protein domain analysis, GO annotation, and gene expression pattern analysis. The WGDB is freely accessible at https://cosbi7.ee.ncku.edu.tw/Wolfberry/. Overall, WGDB serves as a valuable resource for researchers interested in the genomics and transcriptomics of L. barbarum. Its user-friendly web interface and comprehensive data facilitate the exploration of gene functions, regulatory mechanisms, and metabolic pathways, ultimately contributing to a deeper understanding of wolfberry and its potential applications in agronomy and nutrition.
Background and Objective:Proteome microarrays are one of the popular high-throughput screening methods for large-scale investigation of protein interactions in cells. These interactions can be measured on protein chips when coupled with fluorescence-labeled probes, helping indicate potential biomarkers or discover drugs. Several computational tools were developed to help analyze the protein chip results. However, existing tools fail to provide a user-friendly interface for biologists and present only one or two data analysis methods suitable for limited experimental designs, restricting the use cases.Methods:In order to facilitate the biomarker examination using protein chips, we implemented a user-friendly and comprehensive web tool called BAPCP (Biomarker Analysis tool for Protein Chip Platforms) in this research to deal with diverse chip data distributions.Results:BAPCP is well integrated with standard chip result files and includes 7 data normalization methods and 7 custom-designed quality control/differential analysis filters for biomarker extraction among experiment groups. Moreover, it can handle cost-efficient chip designs that repeat several blocks/samples within one single slide. Using experiments of the human coronavirus (HCoV) protein microarray and the E. coli proteome chip that helps study the immune response of Kawasaki disease as examples, we demonstrated that BAPCP can accelerate the time-consuming week-long manual biomarker identification process to merely 3 min.Conclusions:The developed BAPCP tool provides substantial analysis support for protein interaction studies and conforms to the necessity of expanding computer usage and exchanging information in bioscience and medicine. The web service of BAPCP is available at https://cosbi.ee.ncku.edu.tw/BAPCP/.
Circular RNAs (circRNAs) are RNA molecules with a continuous loop structure characterized by back-splice junctions (BSJs). While analyses of short-read RNA sequencing have identified millions of BSJ events, it is inherently challenging to determine exact full-length sequences and alternatively spliced (AS) isoforms of circRNAs. Recent advances in nanopore long-read sequencing with circRNA enrichment bring an unprecedented opportunity for investigating the issues. Here, we developed FL-circAS (https://cosbi.ee.ncku.edu.tw/FL-circAS/), which collected such long-read sequencing data of 20 cell lines/tissues and thereby identified 884 636 BSJs with 1 853 692 full-length circRNA isoforms in human and 115 173 BSJs with 135 617 full-length circRNA isoforms in mouse. FL-circAS also provides multiple circRNA features. For circRNA expression, FL-circAS calculates expression levels for each circRNA isoform, cell line/tissue specificity at both the BSJ and isoform levels, and AS entropy for each BSJ across samples. For circRNA biogenesis, FL-circAS identifies reverse complementary sequences and RNA binding protein (RBP) binding sites residing in flanking sequences of BSJs. For functional patterns, FL-circAS identifies potential microRNA/RBP binding sites and several types of evidence for circRNA translation on each full-length circRNA isoform. FL-circAS provides user-friendly interfaces for browsing, searching, analyzing, and downloading data, serving as the first resource for discovering full-length circRNAs at the isoform level.
Gram-negative bacteremia is a major cause of global morbidity involving three phases of pathogenesis: initial site infection, dissemination, and survival in the blood and filtering organs. Klebsiella pneumoniae is a leading cause of bacteremia and pneumonia is often the initial infection. In the lung, K. pneumoniae relies on many factors like capsular polysaccharide and branched chain amino acid biosynthesis for virulence and fitness. However, mechanisms directly enabling bloodstream fitness are unclear. Here, we performed transposon insertion sequencing (TnSeq) in a tail-vein injection model of bacteremia and identified 58 K. pneumoniae bloodstream fitness genes. These factors are diverse and represent a variety of cellular processes. In vivo validation revealed tissue-specific mechanisms by which distinct factors support bacteremia. ArnD, involved in Lipid A modification, was required across blood filtering organs and supported resistance to soluble splenic factors. The purine biosynthesis enzyme PurD supported liver fitness in vivo and was required for replication in serum. PdxA, a member of the endogenous vitamin B6 biosynthesis pathway, optimized replication in serum and lung fitness. The stringent response regulator SspA was required for splenic fitness yet was dispensable in the liver. In a bacteremic pneumonia model that incorporates initial site infection and dissemination, splenic fitness defects were enhanced. ArnD, PurD, DsbA, SspA, and PdxA increased fitness across bacteremia phases and each demonstrated unique fitness dynamics within compartments in this model. SspA and PdxA enhanced K. pnuemoniae resistance to oxidative stress. SspA, but not PdxA, specifically resists oxidative stress produced by NADPH oxidase Nox2 in the lung, spleen, and liver, as it was a fitness factor in wild-type but not Nox2-deficient (Cybb-/-) mice. These results identify site-specific fitness factors that act during the progression of Gram-negative bacteremia. Defining K. pneumoniae fitness strategies across bacteremia phases could illuminate therapeutic targets that prevent infection and sepsis.
Biomolecular condensates have been shown to interact in vivo, yet it is unclear whether these interactions are functionally meaningful. Here, we demonstrate that cooperativity between two distinct condensates-germ granules and P bodies-is required for transgenerational gene silencing in C. elegans. We find that P bodies form a coating around perinuclear germ granules and that P body components CGH-1/DDX6 and CAR-1/ LSM14 are required for germ granules to organize into sub-compartments and concentrate small RNA silencing factors. Functionally, while the P body mutant cgh-1 is competent to initially trigger gene silencing, it is unable to propagate the silencing to subsequent generations. Mechanistically, we trace this loss of transgenerational silencing to defects in amplifying secondary small RNAs and the stability of WAGO-4 Argonaute, both known carriers of gene silencing memories. Together, these data reveal that cooperation between condensates results in an emergent capability of germ cells to establish heritable memory.
The identification of xenobiotic biotransformation products is crucial for delineating toxicity and carcinogenicity that might be caused by xenobiotic exposures and for establishing monitoring systems for public health. However, the lack of available reference standards and spectral data leads to the generation of multiple candidate structures during identification and reduces the confidence in identification. Here, a UHPLC-HRMS-based metabolomics strategy integrated with a metabolite structure elucidation approach, namely, FragAssembler, was proposed to reduce the number of false-positive structure candidates. biotransformation product candidates were filtered by mass defect filtering (MDF) and multiple-group comparison. FragAssembler assembled fragment signatures from the MS/MS spectra and generated the modified moieties corresponding to the identified biotransformation products. The feasibility of this approach was demonstrated by the three biotransformation products of di(2-ethylhexyl)phthalate (DEHP). Comprehensive identification was carried out, and 24 and 13 biotransformation products of two xenobiotics, DEHP and 4'-Methoxy-α-pyrrolidinopentiophenone (4-MeO-α-PVP), were annotated, respectively. The number of 4-MeO-α-PVP biotransformation product candidates in the FragAssembler calculation results was approximately 2.1 times lower than that generated by BioTransformer 3.0. Our study indicates that the proposed approach has great potential for efficiently and reliably identifying xenobiotic biotransformation products, which is attributed to the fact that FragAssembler eliminates false-positive reactions and chemical structures and distinguishes modified moieties on isomeric biotransformation products. The FragAssembler software and associated tutorial are freely available at https://cosbi.ee.ncku.edu.tw/FragAssembler/ and the source code can be found at https://github.com/YuanChihChen/FragAssembler.
Coronavirus-associated coagulopathy (CAC) is a morbid and lethal sequela of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection. CAC results from a perturbed balance between coagulation and fibrinolysis and occurs in conjunction with exaggerated activation of monocytes/macrophages (MO/Mφs), and the mechanisms that collectively govern this phenotype seen in CAC remain unclear. Here, using experimental models that use the murine betacoronavirus MHVA59, a well-established model of SARS-CoV-2 infection, we identify that the histone methyltransferase mixed lineage leukemia 1 (MLL1/KMT2A) is an important regulator of MO/Mφ expression of procoagulant and profibrinolytic factors such as tissue factor (F3; TF), urokinase (PLAU), and urokinase receptor (PLAUR) (herein, "coagulopathy-related factors") in noninfected and infected cells. We show that MLL1 concurrently promotes the expression of the proinflammatory cytokines while suppressing the expression of interferon alfa (IFN-α), a well-known inducer of TF and PLAUR. Using in vitro models, we identify MLL1-dependent NF-κB/RelA-mediated transcription of these coagulation-related factors and identify a context-dependent, MLL1-independent role for RelA in the expression of these factors in vivo. As functional correlates for these findings, we demonstrate that the inflammatory, procoagulant, and profibrinolytic phenotypes seen in vivo after coronavirus infection were MLL1-dependent despite blunted Ifna induction in MO/Mφs. Finally, in an analysis of SARS-CoV-2 positive human samples, we identify differential upregulation of MLL1 and coagulopathy-related factor expression and activity in CD14+ MO/Mφs relative to noninfected and healthy controls. We also observed elevated plasma PLAU and TF activity in COVID-positive samples. Collectively, these findings highlight an important role for MO/Mφ MLL1 in promoting CAC and inflammation.
PIWI-interacting RNAs (piRNAs) protect genome integrity by silencing transposon mRNAs and some endogenous mRNAs in various animals. However, C. elegans piRNAs only trigger gene silencing at select predicted targeting sites, suggesting additional cellular mechanisms regulate piRNA silencing. To gain insight into possible mechanisms, we compared the transcriptome-wide predicted piRNA targeting sites to the in vivo piRNA binding sites. Surprisingly, while sequence-based predicted piRNA targeting sites are enriched in 3' UTRs, we found that C. elegans piRNAs preferentially bind to coding regions (CDS) of target mRNAs, leading to preferential production of secondary silencing small RNAs in the CDS. However, our analyses suggest that this CDS binding preference cannot be explained by the action of anti-silencing Argonaute CSR-1. Instead, our analyses imply that CSR-1 protects mRNAs from piRNA silencing through two distinct mechanisms - by inhibiting piRNA binding across the entire CSR-1 targeted transcript, and by inhibiting secondary silencing small RNA production locally at CSR-1 bound sites. Together, our work identifies the CDS as the critical region that is uniquely competent for piRNA binding in C. elegans. We speculate the CDS binding preference may have evolved to allow the piRNA pathway to maintain robust recognition of RNA targets in spite of genetic drift. Together, our analyses revealed that distinct mechanisms are responsible for restricting piRNA binding and silencing to achieve proper transcriptome surveillance.