Accurate protein domain annotation is essential for inferring protein function, and databases such as Pfam provide sequence-derived signatures for thousands of domain families. Because protein structure is more evolutionarily conserved than sequence, structure-based searches can detect homologous relationships even at low sequence identity (typically below 30%), where pairwise sequence aligners often lose sensitivity. Here, we leverage AlphaFold-derived structures of Pfam domain instances to systematically evaluate structure-based versus sequence-based methods for Pfam annotation. We benchmarked three structural aligners (Reseek, Foldseek, TM-align) against sequence-based methods (MMseqs, HMMER) using both exhaustive all-against-all searches and a split-family design that enables direct comparison of pairwise and profile-based ranking performance. We also evaluated residue-level alignment accuracy using Pfam multiple sequence alignments as reference and investigated whether profile-derived information can improve structural hit ranking. In all-against-all searches, Reseek achieved the highest sensitivity up to the first false positive (AUC = 0.85), outperforming Foldseek (0.81), TM-align (0.76), and MMseqs (0.46). In split-family evaluation, HMMER remained superior (maximum F1 = 0.991), highlighting the continued strength of sequence-profile approaches for family-level annotation. Performance varied substantially across domain families, with average sequence identity emerging as the strongest predictor of success. Structural aligners consistently produced more accurate residue-level mappings than pairwise sequence methods. Finally, incorporating profile-derived information via rescoring improved structural annotation performance for short domains, suggesting a path toward profile-informed structure-based domain annotation.
Independent protein-protein interaction networks in kinetoplastid parasites show little overlap, often interpreted as biological divergence. We argue that this pattern largely reflects fragmented sampling. Integrating interactomes from Trypanosoma brucei, Trypanosoma cruzi, and Leishmania donovani improves functional coverage and interpretability while preserving lineage-specific assemblies, providing a framework for hypothesis generation across species.
RNA metabolism in kinetoplastid protists (Kinetoplastea), including trypanosomes and Leishmania, involves unique post-transcriptional mitochondrial RNA editing that creates translatable mRNAs through uridine (U) insertions and deletions (U-indels) directed by antisense guide RNAs (gRNAs). Like other biological processes that require specific RNA targeting, this system faces several challenges beyond coordinating its many components: assembling mRNA-gRNA hybrids, recognizing hundreds of sites, and accurately distinguishing pre-edited, partially edited, and fully edited transcripts in the mitochondrial environment. In parasites such as Trypanosoma brucei, significant energetic adaptations to different host environments also involve critical editing changes during development. The editing holoenzyme includes three molecular complexes and isoforms that carry most proteins: RNA Editing Catalytic Complexes (RECCs), which catalyze U-indel cycles; RNA Editing Substrate Complexes (RESCs), which serve as scaffolds to coordinate the editing components; and the RNA Editing Helicase 2 Complex (REH2C), which contains key proteins involved in developmental editing regulation. However, more proteins and functions are being discovered. The editing system, best understood in T. brucei, shows considerable evolutionary conservation in its core machinery; however, it varies in the extent of RNA editing and the organization of mitochondrial mRNA and gRNA genes across different species. Here we explore recent progress in our understanding of RNA editing and the growing use of modern computational tools, including artificial intelligence (AI) and structural methods, to examine function, organization, developmental regulation, and evolutionary aspects of this amazing system. This article is categorized under: RNA Interactions with Proteins and Other Molecules > RNA-Protein Complexes RNA Processing > RNA Editing and Modification.
Leishmania donovani is the causative agent of visceral leishmaniasis, a tropical disease affecting millions worldwide. While proteomic studies of Leishmania species have been conducted, the organization of protein-protein interaction (PPI) networks in L. donovani remains largely unexplored. Here, we present a protein interaction network for L. donovani generated through size-exclusion chromatography coupled with mass spectrometry (SEC-MS) and computational analysis. We quantified 3468 proteins with high confidence, of which approximately 70 % are conserved across the Tritryps (Trypanosoma brucei, T. cruzi, and L. donovani). The resulting network contains 1509 nodes and 16,095 interactions, exhibiting scale-free topology and covering key cellular machineries such as the proteasome, ribosome, translation initiation complexes, and BBSome. Remarkably, most annotated Leishmania complexes remained intact within the network, highlighting its high quality. Paralogs within L. donovani frequently interacted with each other, a phenomenon observed at a higher rate than reported in different organisms. Beyond structural organization, the network also provided interaction-based evidence that functionally contextualizes previously uncharacterized or poorly annotated proteins. Complexes involved in mRNA metabolism and flagellar assembly revealed novel components supported by conserved interaction patterns, underscoring the biological utility of the network for functional inference. Our study provides the first experimentally derived, large-scale interaction network specific to L. donovani, offering critical insights into the parasite's molecular architecture. All interaction data are available through our dedicated database at https://2025.trypsnetdb.org.
Trypanosoma cruzi, the causative agent of Chagas disease, poses a significant health challenge due to limited therapeutic options and an incomplete understanding of its biology. Approximately half of the genome encodes hypothetical proteins with unknown functions, underscoring the need for systematic functional annotation. Protein-protein interactions (PPIs) underpin essential cellular processes, yet no large-scale PPI map has been developed for T. cruzi ─a critical gap that impedes both functional annotation of its proteome and drug discovery. This study presents the first comprehensive PPI network for T. cruzi, constructed using quantitative mass spectrometry-based cofractionation data. The network includes 1319 proteins and more than 16,000 predicted interactions, with 47% of the proteins classified as hypothetical, consistent with the 49% hypothetical annotation rate in the proteome. Their placement within functionally enriched network modules provides unprecedented insights into their potential biological roles. Network analysis revealed densely interconnected cores enriched with essential cellular functions. This PPI network exhibits small-world properties, with conserved proteins showing higher connectivity, reinforcing their central roles in the parasite's biology. This resource, publicly available at https://2025.trypsnetdb.org/, offers a powerful platform for exploring T. cruzi biology and prioritizing novel therapeutic targets, revealing central hubs of protein organization, resolving ribosomal and proteasomal complexes, and enabling functional predictions for numerous hypothetical proteins through integrative structural modeling.
Significant variations in the abundance of mitochondrial RNA processing proteins and their target RNAs across trypanosome life stages present an opportunity to explore the regulatory mechanisms that drive these changes. Utilizing omics approaches can uncover unconventional targets, aiding our understanding of the parasites’ adaptation and enabling targeted interventions for differentiation.
Traditional automated in silico functional annotation uses tools like Pfam that rely on sequence similarities for domain annotation. However, structural conservation often exceeds sequence conservation, suggesting an untapped potential for improved annotation through structural similarity. This approach was previously overlooked before the AlphaFold2 introduction due to the need for more high-quality protein structures. Leveraging structural information especially holds significant promise to enhance accurate annotation in diverse proteins across phylogenetic distances. In our study, we evaluated the feasibility of annotating Pfam domains based on structural similarity. To this end, we created a database from segmented full-length protein structures at their domain boundaries, representing the structure of Pfam seeds. We used Trypanosomabrucei , a phylogenetically distant protozoan parasite as our model organism. Its structome was aligned with our database using Foldseek, the ultra-fast structural alignment tool, and the top non-overlapping hits were annotated as domains. Our method identified over 400 new domains in the T. brucei proteome, surpassing the benchmark set by sequence-based tools, Pfam and Pfam-N, with some predictions validated manually. We have also addressed limitations and suggested avenues for further enhancing structure-based domain annotation.
Background Trypanosoma brucei is the causative agent for trypanosomiasis in humans and livestock, which presents a growing challenge due to drug resistance. While identifying novel drug targets is vital, the process is delayed due to a lack of functional information on many of the pathogen’s proteins. Accordingly, this paper presents a computational framework for prioritizing drug targets within the editosome, a vital molecular machinery responsible for mitochondrial RNA processing in T. brucei . Importantly, this framework may eliminate the need for prior gene or protein characterization, potentially accelerating drug discovery efforts. Results By integrating protein-protein interaction (PPI) network analysis, PPI structural modeling, and residue interaction network (RIN) analysis, we quantitatively ranked and identified top hub editosome proteins, their key interaction interfaces, and hotspot residues. Our findings were cross-validated and further prioritized by incorporating them into gene set analysis and differential expression analysis of existing quantitative proteomics data across various life stages of T. brucei . In doing so, we highlighted PPIs such as KREL2-KREPA1, RESC2-RESC1, RESC12A-RESC13, and RESC10-RESC6 as top candidates for further investigation. This includes examining their interfaces and hotspot residues, which could guide drug candidate selection and functional studies. Conclusion RNA editing offers promise for target-based drug discovery, particularly with proteins and interfaces that play central roles in the pathogen’s life cycle. This study introduces an integrative drug target identification workflow combining information from the PPI network, PPI 3D structure, and reside-level information of their interface which can be applicable to diverse pathogens. In the case of T. brucei , via this pipeline, the present study suggested potential drug targets with residue-resolution from RNA editing machinery. However, experimental validation is needed to fully realize its potential in advancing urgently needed antiparasitic drug development.
Trypanosomatids are the causative agents of deadly diseases in humans and livestock. Given the high phylogenetic distance of trypanosomatids from model organisms, these organisms have ample unannotated genes. Manual functional annotation is time-consuming, highlighting the importance of automated functional annotation tools. The development of automated functional tools is a hot research topic, and multiple tools have been developed for the task. PANNZER2 is an automated functional annotation tool that merely relies on the sequence similarity of the query to the annotated proteins. We tried PANNZER2 on Trypanosoma brucei, the most studied organism among trypanosomatids, to see if it could improve our knowledge of the functions of the genes.Even with the availability of automated annotation tools like InterPro2GO in databases such as TriTrypDB, PANNZER2 has made surprisingly confident predictions for some hypothetical proteins in T. brucei. In this study, we identify gaps in such annotations because of not employing pairwise sequence alignment tools in TriTrypDB's automated annotation process. Our findings demonstrate that even the use of stringent cutoffs can successfully annotate a significant number of proteins. Additionally, we discovered that adjusting the open reading frames in certain genes leads to sequences with increased sequence signature coverage—characterized by the length covered by at least one sequence signature—compared to the original sequences. This enhanced sequence signature coverage suggests these genomic fragments could be pseudogenes. To facilitate further exploration, we developed a script to help identify potential pseudogenes within an organism's genome, offering researchers a new tool for genomic analysis and understanding. We extended all our analysis to Trypanosoma cruzi and Leishmania major to assess the impact of this approach across different species.Our study demonstrates that by utilizing pairwise sequence similarity alignment, even with stringent cutoffs, we can attribute 2986, 3953, and 3798 new GO terms to the genomes of T. brucei, T. cruzi, and L. major. Additionally, we found that 210, 239, and 29 genes exhibit increased sequence signature coverage following frame correction, suggesting the presence of pseudogenes.
RNA-specific nucleotidyltransferases (rNTrs) add nontemplated nucleotides to the 3′ end of RNA. Two noncanonical rNTRs that are thought to be poly(A) polymerases (PAPs) have been identified in the mitochondria of trypanosomes – KPAP1 and KPAP2. KPAP1 is the primary polymerase that adds adenines (As) to trypanosome mitochondrial mRNA 3′ tails, while KPAP2 is a non-essential putative polymerase whose role in the mitochondria is ambiguous. Here, we elucidate the effects of manipulations of KPAP1 and KPAP2 on the 5′ and 3′ termini of transcripts and their 3′ tails. Using glycerol gradients followed by immunoblotting, we present evidence that KPAP2 is found in protein complexes of up to about 1600 kDa. High-throughput sequencing of mRNA termini showed that KPAP2 overexpression subtly changes an edited transcript’s 3′ tails, though not in a way consistent with general PAP activity. Next, to identify possible roles of posttranslational modifications on KPAP1 regulation, we mutated two KPAP1 arginine methylation sites to either mimic methylation or hypomethylation. We assessed their effect on 3′ mRNA tail characteristics and found that the two mutants generally had opposing effects, though some of these were transcript-specific. We present results suggesting that while methylation increases KPAP1 substrate binding and/or initial nucleotide additions, unmethylated KPAP1is more processive. We also present a comprehensive review of UTR termini, and evidence that tail addition activity may change as mRNA editing is initiated. Together, this work furthers our understanding of the role of KPAP1 and KPAP2 on trypanosome mitochondrial mRNA 3′ tail addition, as well as provides more information on mRNA termini processing in general.
RNA editing pathway is a validated target in kinetoplastid parasites (Trypanosoma brucei, Trypanosoma cruzi, and Leishmania spp.) that cause severe diseases in humans and livestock. An essential large protein complex, the editosome, mediates uridine insertion and deletion in RNA editing through a stepwise process. This study details the discovery of editosome inhibitors by screening a library of widely used human drugs using our previously developed in vitro biochemical Ribozyme Insertion Deletion Editing (RIDE) assay. Subsequent studies on the mode of action of the identified hits and hit expansion efforts unveiled compounds that interfere with RNA-editosome interactions and novel ligase inhibitors with IC50 values in the low micromolar range. Docking studies on the ligase demonstrated similar binding characteristics for ATP and our novel epigallocatechin gallate inhibitor. The inhibitors demonstrated potent trypanocidal activity and are promising candidates for drug repurposing due to their lack of cytotoxic effects. Further studies are necessary to validate these targets using more definitive gene-editing techniques and to enhance the safety profile.
The wide applications of liquid chromatography - mass spectrometry (LC-MS) in untargeted metabolomics demand an easy-to-use, comprehensive computational workflow to support efficient and reproducible data analysis. However, current tools were primarily developed to perform specific tasks in LC-MS based metabolomics data analysis. Here we introduce MetaboAnalystR 4.0 as a streamlined pipeline covering raw spectra processing, compound identification, statistical analysis, and functional interpretation. The key features of MetaboAnalystR 4.0 includes an auto-optimized feature detection and quantification algorithm for LC-MS1 spectra processing, efficient MS2 spectra deconvolution and compound identification for data-dependent or data-independent acquisition, and more accurate functional interpretation through integrated spectral annotation. Comprehensive validation studies using LC-MS1 and MS2 spectra obtained from standards mixtures, dilution series and clinical metabolomics samples have shown its excellent performance across a wide range of common tasks such as peak picking, spectral deconvolution, and compound identification with good computing efficiency. Together with its existing statistical analysis utilities, MetaboAnalystR 4.0 represents a significant step toward a unified, end-to-end workflow for LC-MS based global metabolomics in the open-source R environment.
Mitochondrial uridine insertion/deletion RNA editing, catalyzed by a multiprotein complex (editosome), is essential for gene expression in trypanosomes and Leishmania parasites. As this process is absent in the human host, a drug targeting this mechanism promises high selectivity and reduced toxicity. Here, we successfully miniaturized our FRET-based full-round RNA editing assay, which replicates the complete RNA editing process, adapting it into a 1536-well format. Leveraging this assay, we screened over 100,000 compounds against purified editosomes derived from Trypanosoma brucei, identifying seven confirmed primary hits. We sourced and evaluated various analogs to enhance the inhibitory and parasiticidal effects of these primary hits. In combination with secondary assays, our compounds marked inhibition of essential catalytic activities, including the RNA editing ligase and interactions of editosome proteins. Although the primary hits did not exhibit any growth inhibitory effect on parasites, we describe eight analog compounds capable of effectively killing T. brucei and/or Leishmania donovani parasites within a low micromolar concentration. Whether parasite killing is - at least in part - due to inhibition of RNA editing in vivo remains to be assessed. Our findings introduce novel molecular scaffolds with the potential for broad antitrypanosomal effects.
Since the first identification of circular RNA (circRNA) in viral-like systems, reports of circRNAs and their functions in various organisms, cell types, and organelles have greatly expanded. Here, we report the first evidence, to our knowledge, of circular mRNA in the mitochondrion of the eukaryotic parasite, Trypanosoma brucei . While using a circular RT-PCR technique developed to sequence mRNA tails of mitochondrial transcripts, we found that some mRNAs are circularized without an in vitro circularization step normally required to produce PCR products. Starting from total in vitro circularized RNA and in vivo circRNA, we high-throughput sequenced three transcripts from the 3′ end of the coding region, through the 3′ tail, to the 5′ start of the coding region. We found that fewer reads in the circRNA libraries contained tails than in the total RNA libraries. When tails were present on circRNAs, they were shorter and less adenine-rich than the total population of RNA tails of the same transcript. Additionally, using hidden Markov modelling we determined that enzymatic activity during tail addition is different for circRNAs than for total RNA. Lastly, circRNA UTRs tended to be shorter and more variable than those of the same transcript sequenced from total RNA. We propose a revised model of Trypanosome mitochondrial tail addition, in which a fraction of mRNAs is circularized prior to the addition of adenine-rich tails and may act as a new regulatory molecule or in a degradation pathway.
RNA editing, a unique post-transcriptional modification, is observed in trypanosomatid parasites as a crucial procedure for the maturation of mitochondrial mRNAs. The editosome protein complex, involving multiple protein components, plays a key role in this process. In Trypanosoma brucei, a putative Z-DNA binding protein known as RBP7910 is associated with the editosome. However, the specific Z-DNA/Z-RNA binding activity and the interacting interface of RBP7910 have yet to be determined. In this study, we conducted a comparative analysis of the binding behavior of RBP7910 with different potential ligands using microscale thermophoresis (MST). Additionally, we generated a 3D model of the protein, revealing potential Z-α and Z-β nucleic acid-binding domains of RBP7910. RBP7910 belongs to the winged-helix–turn–helix (HTH) superfamily of proteins with an α1α2α3β1β2 topology. Finally, using docking techniques, potential interacting surface regions of RBP7910 with notable oligonucleotide ligands were identified. Our findings indicate that RBP7910 exhibits a notable affinity for (CG)n Z-DNA, both in single-stranded and double-stranded forms. Moreover, we observed a broader interacting interface across its Z-α domain when bound to Z-DNA/Z-RNA compared to when bound to non-Z-form nucleic acid ligands.
Untranslatable mitochondrial transcripts in kinetoplastids are decrypted post-transcriptionally through an RNA editing process that entails uridine insertion/deletion. This unique stepwise process is mediated by the editosome, a multiprotein complex that is a validated drug target of considerable interest in addressing the unmet medical needs for kinetoplastid diseases. With that objective, several in vitro RNA editing assays have been developed, albeit with limited success in discovering potent inhibitors. This manuscript describes the development of three hammerhead ribozyme (HHR) FRET reporter-based RNA editing assays for precleaved deletion, insertion, and ligation assays that bypass the rate-limiting endonucleolytic cleavage step, providing information on U-deletion, U-insertion, and ligation activities. These assays exhibit higher editing efficiencies in shorter incubation times while requiring significantly less purified editosome and 10,000-fold less ATP than the previously published full round of in vitro RNA editing assay. Moreover, modifications in the reporter ribozyme sequence enable the feasibility of multiplexing a ribozyme-based insertion/deletion editing (RIDE) assay that simultaneously surveils U-insertion and deletion editing suitable for HTS. These assays can be used to find novel chemical compounds with chemotherapeutic applications or as probes for studying the editosome machinery.
Parasitic protozoans of theTrypanosomaandLeishmaniaspecies have a uniquely organized mitochondrial genome, the kinetoplast. Most kinetoplast-transcribed mRNAs are cryptic and encode multiple subunits for the electron transport chain following maturation through a uridine insertion/deletion process called RNA editing. This process is achieved through an enzyme cascade by an RNA editing catalytic complex (RECC), where the final ligation step is catalyzed by the kinetoplastid RNA editing ligases, KREL1 and KREL2. While the amino-terminal domain (NTD) of these proteins is highly conserved with other DNA ligases and mRNA capping enzymes, with five recognizable motifs, the functional role of their diverged carboxy-terminal domain (CTD) has remained elusive. In this manuscript, we assayed recombinant KREL1 in vitro to unveil critical residues from its CTD to be involved in protein–protein interaction and dsRNA ligation activity. Our data show that the α-helix (H)3 of KREL1 CTD interacts with the αH1 of its editosome protein partner KREPA2. Intriguingly, the OB-fold domain and the zinc fingers on KREPA2 do not appear to influence the RNA ligation activity of KREL1. Moreover, a specific KWKE motif on the αH4 of KREL1 CTD is found to be implicated in ligase auto-adenylylation analogous to motif VI in DNA ligases. In summary, we present in the KREL1 CTD a motif VI for auto-adenylylation and a KREPA2 binding motif for RECC integration.
The RNA editing core complex (RECC) catalyzes mitochondrial U-insertion/deletion mRNA editing in trypanosomatid flagellates. Some naphthalene-based sulfonated compounds, such as C35 and MrB, competitively inhibit the auto-adenylylation activity of an essential RECC enzyme, kinetoplastid RNA editing ligase 1 (KREL1), required for the final step in editing. Previous studies revealed the ability of these compounds to interfere with the interaction between the editosome and its RNA substrates, consequently affecting all catalytic activities that comprise RNA editing. This observation implicates a critical function for the affected RNA binding proteins in RNA editing. In this study, using the inhibitory compounds, we analyzed the composition and editing activities of functional editosomes and identified the mitochondrial RNA binding proteins 1 and 2 (MRP1/2) as their preferred targets. While the MRP1/2 heterotetramer complex is known to bind guide RNA and promote annealing to its cognate pre-edited mRNA, its role in RNA editing remained enigmatic. We show that the compounds affect the association between the RECC and MRP1/2 heterotetramer. Furthermore, RECC purified post-treatment with these compounds exhibit compromised in vitro RNA editing activity that, remarkably, recovers upon the addition of recombinant MRP1/2 proteins. This work provides experimental evidence that the MRP1/2 heterotetramer is required for in vitro RNA editing activity and substantiates the hypothesized role of these proteins in presenting the RNA duplex to the catalytic complex in the initial steps of RNA editing.
Trypanosoma brucei spp. cause African human and animal trypanosomiasis, a burden on health and economy in Africa. These hemoflagellates are distinguished by a kinetoplast nucleoid containing mitochondrial DNAs of two kinds: maxicircles encoding ribosomal RNAs (rRNAs) and proteins and minicircles bearing guide RNAs (gRNAs) for mRNA editing. All RNAs are produced by a phage-type RNA polymerase as 3' extended precursors, which undergo exonucleolytic trimming. Most pre-mRNAs proceed through 3' adenylation, uridine insertion/deletion editing, and 3' A/U-tailing. The rRNAs and gRNAs are 3' uridylated. Historically, RNA editing has attracted major research effort, and recently essential pre- and postediting processing events have been discovered. Here, we classify the key players that transform primary transcripts into mature molecules and regulate their function and turnover.