The 3'-5' exoribonuclease EXOSC10 degrades aberrant mRNAs and noncoding RNAs in cooperation with the nuclear RNA exosome. EXOSC10's localization and stability are regulated by sumoylation and proteasomal degradation in response to stress, and the protein is essential for cell growth and proliferation, fertility, hematopoiesis, and brain development. EXOSC10 is a cancer biomarker; its activity is inhibited by the widely used anticancer drug 5-fluorouracil (5-FU) and the protein's depletion sensitizes cells to 5-FU. We employed mass spectrometry to reveal EXOSC10's post-translational modifications, such as phosphorylation, acetylation and ubiquitination, and to explore its protein interaction network, which includes RNA exosome subunits and enzymes involved in protein degradation. Furthermore, we find that the EXOSC10S402T allele identified in colon cancer and located within a motif for targeted proteolysis is stable, nuclear but nonfunctional in vivo, since homozygous Exosc10S402T mice exhibit early embryonic lethality. We identified equivalent S402P/S402A variants and heterozygous loss-of-function (LoF) alleles in cancers and healthy individuals using public genomics data. Our findings suggest that recessive EXOSC10 LoF alleles may cause increased 5-FU sensitivity in tumors bearing de novo mutations and hypertoxicity in heterozygous carriers.
Previous RNA profiling studies revealed co-expression of overlapping sense/antisense (s/a) transcripts in pro- and eukaryotic organisms. Functional analyses in yeast have shown that certain s/a mRNA/mRNA and mRNA/lncRNA pairs form stable double-stranded RNAs (dsRNAs) that affect transcript stability. Little is known, however, about the genome-wide prevalence of dsRNA formation and its potential functional implications during growth and development in diploid budding yeast. To address this question, we monitored dsRNAs in a Saccharomyces cerevisiae strain expressing the ribonuclease DCR1 and the RNA binding protein AGO1 from Naumovozyma castellii. We identify dsRNAs at 347 s/a loci that express partially or completely overlapping transcripts during mitosis, meiosis or both stages of the diploid life cycle. We associate dsRNAs with s/a loci previously thought to be exclusively regulated by antisense interference, and others that encode antisense RNAs, which downregulate sense mRNA-encoded protein levels. To facilitate hypothesis building we developed the Sense/Antisense double-stranded RNA (SensR) expression viewer. Users are able to retrieve different graphical displays of dsRNA and RNA expression data using genome coordinates and systematic or standard names for mRNAs and different types of stable or cryptic long non-coding RNAs (lncRNAs). Our data are a useful resource for improving yeast genome annotation and for work on RNA-based regulatory mechanisms controlling transcript and protein levels. The data are also interesting from an evolutionary perspective, since natural antisense transcripts that form stable dsRNAs have been detected in many species from bacteria to humans. The SensR viewer is freely accessible at https://sensr.genouest.org.
The expansion of multi-omics datasets raises significant challenges for data integration and querying. To overcome these challenges, we developed a generic RDF-based integration schema that connects various types of differential -omics data, epigenomics, and regulatory information. This schema employs the FALDO ontology to enable querying based on genomic locations. It is designed to be fully or partially populated, providing both flexibility and extensibility while supporting complex queries. We validated the schema by reproducing two recently published studies, one in biomedicine and the other in environmental science, proving its genericity and its ability to integrate data efficiently. This schema serves as an effective tool for managing and querying a wide range of multi-omics datasets.
In systems biology, the study of biological pathways plays a central role in understanding the complexity of biological systems. The massification of pathway data made available by numerous online databases in recent years has given rise to an important need for standardization of this data. The BioPAX format (Biological Pathway Exchange) emerged in 2010 as a solution for standardizing and exchanging pathway data across databases. BioPAX is a Semantic Web format associated to an ontology. It is highly expressive, allowing to finely describe biological pathways at the molecular and cellular levels, but the associated intrinsic complexity may be an obstacle to its widespread adoption. Here, we report on the use of the BioPAX format in 2024. We compare how the different pathway databases use BioPAX to standardize their data and point out possible avenues for improvement to make full use of its potential. We also report on the various tools and software that have been developed to work with BioPAX data. Finally, we present a new concept of abstraction on BioPAX graphs that would allow to specifically target areas in a BioPAX graph needed for a specific analysis, thus differentiating the format suited for representation and the abstraction suited for contextual analysis.
Motivation Biological Pathway Exchange (BioPAX) is a standard language, represented in OWL, that aims to enable the integration, exchange, visualization and analysis of biological pathway data. While public databanks increasingly provide datasets in BioPAX format, their use remains below potential. Users may encounter challenges in harnessing the data due to the BioPAX intricately detailed underlying model. Moreover, extracting data demands specific technical skills, posing a barrier for many potential users. Results To address these obstacles, we developped BioPAX-Explorer. This toolis designed to facilitate the adoption and usage of BioPAX for extracting data or build algorithms and models, within the Python community. BioPAX-Explorer is a Python package that provides an object-oriented data model automatically generated from the BioPAX OWL specification. Moreover, it offers expressive query capabilities that shield users from BioPAX inner complexity. BioPAX-Explorer supports dataset building features, validation facilities and pre-build queries. It simplifies the extraction and processing of data from BioPAX sources by automatically generating SPARQL queries. BioPAX-Explorer also offers a user-friendly interface for Python users, allowing exhaustive exploration of large datasets through features such as memory-efficient query execution, entity-oriented queries without the need for SPARQL knowledge. It also allows to learn and reuse complex SPARQL queries for biological network analysis. Additionally, BioPAX-Explorer can accelerate the development of Python-based network analysis software, since it generates graph data structures from BioPAX queries and facilitates the creation of transparent, reproducible workflows based on the BioPAX OWL standard. Availability and implementation BioPAX-Explorer is freely available. We provide the source code, documentation, installation instructions and a Jupyter notebook with tutorial at ### Competing Interest Statement The authors have declared no competing interest.
Abstract Motivation Molecular complexes play a major role in the regulation of biological pathways. The Biological Pathway Exchange format (BioPAX) facilitates the integration of data sources describing interactions some of which involving complexes. The BioPAX specification explicitly prevents complexes to have any component that is another complex (unless this component is a black-box complex whose composition is unknown). However, we observed that the well-curated Reactome pathway database contains such recursive complexes of complexes. We propose reproductible and semantically rich SPARQL queries for identifying and fixing invalid complexes in BioPAX databases, and evaluate the consequences of fixing these nonconformities in the Reactome database. Results For the Homo sapiens version of Reactome, we identify 5833 recursively defined complexes out of the 14 987 complexes (39%). This situation is not specific to the Human dataset, as all tested species of Reactome exhibit between 30% (Plasmodium falciparum) and 40% (Sus scrofa, Bos taurus, Canis familiaris, and Gallus gallus) of recursive complexes. As an additional consequence, the procedure also allows the detection of complex redundancies. Overall, this method improves the conformity and the automated analysis of the graph by repairing the topology of the complexes in the graph. This will allow to apply further reasoning methods on better consistent data. Availability and implementation We provide a Jupyter notebook detailing the analysis https://github.com/cjuigne/non_conformities_detection_biopax.
Feed efficiency is a research priority to support a sustainable meat production. It is recognized as a complex trait that integrates multiple biological pathways orchestrated in and by various tissues. This study aims to determine networks between biological entities to explain inter-individual variation of feed efficiency in growing pigs. The feed conversion ratio (FCR), a measure of feed efficiency, and its two component traits, average daily gain and average daily feed intake, were obtained from 47 growing pigs from a divergent selection for residual feed intake and fed high-starch or high-fat high-fiber diets during 58 days. Datasets of transcriptomics (60 k porcine microarray) in the whole blood and metabolomics (1H-NMR analysis and target gas chromatography) in plasma were available for all pigs at the end of the trial. A weighted gene co-expression network was built from the transcriptomics dataset, resulting in 33 modules of co-expressed molecular probes. The eigengenes of eight of these modules were significantly ( P ≤ 0.05 ) or tended to be ( 0.05 < P ≤ 0.10 ) correlated to FCR. Great homogeneity in the enriched biological pathways was observed in these modules, suggesting co-expressed and co-regulated constitutive genes. They were mainly enriched in genes participating to immune and defense-related processes, and to a lesser extent, to translation, cell development or learning. They were also generally associated with growth rate and percentage of lean mass. In the whole network, only one module composed of genes participating to the response to substances, was significantly associated with daily feed intake and body adiposity. The plasma profiles in circulating metabolites and in fatty acids were summarized by weighted linear combinations using a dimensionality reduction method. Close association was thus found between a module composed of co-expressed genes participating to T cell receptor signaling and cell development process in the whole blood and related to FCR, and the circulating concentrations of polyunsaturated fatty acids in plasma. These systemic approaches have highlighted networks of entities driving key biological processes involved in the phenotypic difference in feed efficiency between animals. Connecting transcriptomics and metabolic levels together had some additional benefits.
OBJECTIVE:MicroRNAs are promising biomarkers of frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS), but discrepant results between studies have so far hampered their use in clinical trials. We aim to assess all previously identified circulating microRNA signatures as potential biomarkers of genetic FTD and/or ALS, using homogeneous, independent validation cohorts of C9orf72 and GRN mutation carriers. METHODS:104 individuals carrying a C9orf72 or a GRN mutation, along with 31 controls, were recruited through the French research network on FTD/ALS. All subjects underwent blood sampling, from which circulating microRNAs were extracted. We measured differences in the expression levels of 65 microRNAs, selected from 15 published studies about FTD or ALS, between 31 controls, 17 C9orf72 presymptomatic subjects, and 29 C9orf72 patients. We also assessed differences in the expression levels of 30 microRNAs, selected from five studies about FTD, between 31 controls, 30 GRN presymptomatic subjects, and 28 GRN patients. RESULTS:More than half (35/65) of the selected microRNAs were differentially expressed in the C9orf72 cohort, while only a small proportion (5/30) of microRNAs were differentially expressed in the GRN cohort. In multivariate analyses, only individuals in the C9orf72 cohort could be adequately classified (ROC AUC up to 0.98 for controls versus presymptomatic subjects, 0.94 for controls versus patients, and 0.77 for presymptomatic subjects versus patients) with some of the signatures. INTERPRETATION:Our results suggest that previously identified microRNAs using sporadic or mixed cohorts of FTD and ALS patients could potentially serve as biomarkers of C9orf72-associated disease, but not GRN-associated disease.
Frontotemporal dementia and amyotrophic lateral sclerosis are rare neurodegenerative diseases with no effective treatment. The development of biomarkers allowing an accurate assessment of disease progression is crucial for evaluating new therapies. Concretely, neuroimaging and transcriptomic (microRNA) data have been shown useful in tracking their progression. However, no single biomarker can accurately measure progression in these complex diseases. Additionally, large samples are not available for such rare disorders. It is thus essential to develop methods that can model disease progression by combining multiple biomarkers from small samples. In this paper, we propose a new framework for computing a disease progression score (DPS) from cross-sectional multimodal data. Specifically, we introduce a supervised multimodal variational autoencoder that can infer a meaningful latent space, where latent representations are placed along a disease trajectory. A score is computed by orthogonal projections onto this path. We evaluate our framework with multiple synthetic datasets and with a real dataset containing 14 patients, 40 presymptomatic genetic mutation carriers and 37 controls from the PREV-DEMALS study. There is no ground truth for the DPS in real-world scenarios, therefore we use the area under the ROC curve (AUC) as a proxy metric. Results with the synthetic datasets support this choice, since the higher the AUC, the more accurate the predicted simulated DPS. Experiments with the real dataset demonstrate better performance in comparison with state-of-the-art approaches. The proposed framework thus leverages cross-sectional multimodal datasets with small sample sizes to objectively measure disease progression, with potential application in clinical trials.
Protein-protein interactions (PPIs) play an ubiquitous and fundamental role in all biological processes. Information on PPIs described in the literature is annotated and made available by several protein-interaction databases. Because most databases have their own curation rules and priorities, they often annotate overlapping sets of publications, which leads to redundancies. We developed a semantic-based approach which enables to accurately detect redundancies within PPI datasets from multiple databases. We applied this approach to assemble a "reproducible interactome", with PPIs supported by at least two methods or publications.
Acetaminophen, aspirin, and ibuprofen are mild analgesics commonly used by pregnant women, the sole current recommendation being to avoid ibuprofen from the fifth month of gestation. The nephrotoxicity of these three analgesics is well documented in adults, as is their interference with prostaglandins biosynthesis. Here we investigated the effect of these analgesics on human first trimester kidneys ex vivo. We first evaluated prostaglandins biosynthesis functionality by performing a wide screening of prostaglandin expression patterns in first trimester human kidneys. We demonstrated that prostaglandins biosynthesis machinery is functional during early nephrogenesis. Human fetal kidney explants aged 7-12 developmental weeks were exposed ex vivo to ibuprofen, aspirin or acetaminophen for 7 days, and analyzed by histology, immunohistochemistry, and flow cytometry. This study has revealed that these analgesics induced a spectrum of abnormalities within early developing structures, ranging from cell death to a decline in differentiating glomeruli density. These results warrant caution for the use of these medicines during the first trimester of pregnancy.
The conserved 3'-5' exoribonuclease EXOSC10/Rrp6 is required for gametogenesis, brain development, erythropoiesis and blood cell enhancer function. The human ortholog is essential for mitosis in cultured cancer cells. Little is known, however, about the role of Exosc10 during embryo development and organogenesis. We generated an Exosc10 knockout model and find that Exosc10-/- mice show an embryonic lethal phenotype. We demonstrate that Exosc10 maternal wild type mRNA is present in mutant oocytes and that the gene is expressed during all stages of early embryogenesis. Furthermore, we observe that EXOSC10 early on localizes to the periphery of nucleolus precursor bodies in blastomeres, which is in keeping with the protein's role in rRNA processing and may indicate a function in the establishment of chromatin domains during initial stages of embryogenesis. Finally, we infer from genotyping data for embryonic days e7.5, e6.5 and e4.5 and embryos cultured in vitro that Exosc10-/- mutants arrest at the eight-cell embryo/morula transition. Our results demonstrate a novel essential role for Exosc10 during early embryogenesis, and they are consistent with earlier work showing that impaired ribosome biogenesis causes a developmental arrest at the morula stage.
Abstract Summary PAX2GRAPHML is an open-source Python library that allows to easily manipulate BioPAX source files as regulated reaction graphs described in.graphml format. The concept of regulated reactions, which allows connecting regulatory, signaling and metabolic levels, has been used. Biochemical reactions and regulatory interactions are homogeneously described by regulated reactions involving substrates, products, activators and inhibitors as elements. PAX2GRAPHML is highly flexible and allows generating graphs of regulated reactions from a single BioPAX source or by combining and filtering BioPAX sources. Supported by the graph exchange format .graphml, the large-scale graphs produced from one or more data sources can be further analyzed with PAX2GRAPHML or standard Python and R graph libraries. Availability and implementation https://pax2graphml.genouest.org.
Objective To identify potential biomarkers of preclinical and clinical progression in chromosome 9 open reading frame 72 gene (C9orf72)-associated disease by assessing the expression levels of plasma microRNAs (miRNAs) in C9orf72 patients and presymptomatic carriers. Methods The PREV-DEMALS study is a prospective study including 22 C9orf72 patients, 45 presymptomatic C9orf72 mutation carriers and 43 controls. We assessed the expression levels of 2576 miRNAs, among which 589 were above noise level, in plasma samples of all participants using RNA sequencing. The expression levels of the differentially expressed miRNAs between patients, presymptomatic carriers and controls were further used to build logistic regression classifiers. Results Four miRNAs were differentially expressed between patients and controls: miR-34a-5p and miR-345-5p were overexpressed, while miR-200c-3p and miR-10a-3p were underexpressed in patients. MiR-34a-5p was also overexpressed in presymptomatic carriers compared with healthy controls, suggesting that miR-34a-5p expression is deregulated in cases with C9orf72 mutation. Moreover, miR-345-5p was also overexpressed in patients compared with presymptomatic carriers, which supports the correlation of miR-345-5p expression with the progression of C9orf72-associated disease. Together, miR-200c-3p and miR-10a-3p underexpression might be associated with full-blown disease. Four presymptomatic subjects in transitional/prodromal stage, close to the disease conversion, exhibited a stronger similarity with the expression levels of patients. Conclusions We identified a signature of four miRNAs differentially expressed in plasma between clinical conditions that have potential to represent progression biomarkers for C9orf72-associated frontotemporal dementia and amyotrophic lateral sclerosis. This study suggests that dysregulation of miRNAs is dynamically altered throughout neurodegenerative diseases progression, and can be detectable even long before clinical onset. Trial registration number NCT02590276.
5-fluorouracil (5-FU) was isolated as an inhibitor of thymidylate synthase, which is important for DNA synthesis. The drug was later found to also affect the conserved 3'-5' exoribonuclease EXOSC10/Rrp6, a catalytic subunit of the RNA exosome that degrades and processes protein-coding and non-coding transcripts. Work on 5-FU's cytotoxicity has been focused on mRNAs and non-coding transcripts such as rRNAs, tRNAs and snoRNAs. However, the effect of 5-FU on long non-coding RNAs (lncRNAs), which include regulatory transcripts important for cell growth and differentiation, is poorly understood. RNA profiling of synchronized 5-FU treated yeast cells and protein assays reveal that the drug specifically inhibits a set of cell cycle regulated genes involved in mitotic division, by decreasing levels of the paralogous Swi5 and Ace2 transcriptional activators. We also observe widespread accumulation of different lncRNA types in treated cells, which are typically present at high levels in a strain lacking EXOSC10/Rrp6. 5-FU responsive lncRNAs include potential regulatory antisense transcripts that form double-stranded RNAs (dsRNAs) with overlapping sense mRNAs. Some of these transcripts encode proteins important for cell growth and division, such as the transcription factor Ace2, and the RNA exosome subunit EXOSC6/Mtr3. In addition to revealing a transcriptional effect of 5-FU action via DNA binding regulators involved in cell cycle progression, our results have implications for the function of putative regulatory lncRNAs in 5-FU mediated cytotoxicity. The data raise the intriguing possibility that the drug deregulates lncRNAs/dsRNAs involved in controlling eukaryotic cell division, thereby highlighting a new class of promising therapeutical targets.
ABSTRACT The sexual transmission of viruses is responsible for the spread of multiple infectious diseases. Although the human immunodeficiency virus (HIV)/AIDS pandemic remains fueled by sexual contacts with infected semen, the origin of virus in semen is still unknown. In a substantial number of HIV-infected men, viral strains present in semen differ from the ones in blood, suggesting that HIV is locally produced within the genital tract. Such local production may be responsible for the persistence of HIV in semen despite effective antiretroviral therapy. In this study, we used single-genome amplification, amplicon sequencing (env gene), and phylogenetic analyses to compare the genetic structures of simian immunodeficiency virus (SIV) populations across all the male genital organs and blood in intravenously inoculated cynomolgus macaques in the chronic stage of infection. Examination of the virus populations present in the male genital tissues of the macaques revealed compartmentalized SIV populations in testis, epididymis, vas deferens, seminal vesicles, and urethra. We found genetic similarities between the viral strains present in semen and those in epididymis, vas deferens, and seminal vesicles. The contribution of male genital organs to virus shedding in semen varied among individuals and could not be predicted based on their infection or proinflammatory cytokine mRNA levels. These data indicate that rather than a single source, multiple genital organs are involved in the release of free virus and infected cells into semen. These findings have important implications for our understanding of systemic virus shedding and persistence in semen and for the design of eradication strategies to access viral reservoirs. IMPORTANCE Semen is instrumental for the dissemination of viruses through sexual contacts. Worryingly, a number of systemic viruses, such as HIV, can persist in this body fluid in the absence of viremia. The local source(s) of virus in semen, however, remains unknown. To elucidate the anatomic origin(s) of the virus released in semen, we compared viral populations present in semen with those in the male genital organs and blood of the Asian macaque model, using single-genome amplification, amplicon sequencing (env gene), and phylogenetic analysis. Our results show that multiple genital tissues harbor compartmentalized strains, some of them (i.e., from epididymis, vas deferens, and seminal vesicles) displaying genetic similarities with the viral populations present in semen. This study is the first to uncover local genital sources of viral populations in semen, providing a new basis for innovative targeted strategies to prevent and eradicate HIV in the male genital tract.
Motivation: At the same time that toxicologists express increasing concern about reproducibility in this field, the development of dedicated databases has already smoothed the path toward improving the storage and exchange of raw toxicogenomic data. Nevertheless, none provides access to analyzed and interpreted data as originally reported in scientific publications. Given the increasing demand for access to this information, we developed TOXsIgN, a repository for TOXicogenomic sIgNatures. Results: The TOXsIgN repository provides a flexible environment that facilitates online submission, storage and retrieval of toxicogenomic signatures by the scientific community. It currently hosts 754 projects that describe more than 450 distinct chemicals and their 8491 associated signatures. It also provides users with a working environment containing a powerful search engine as well as bioinformatics/biostatistics modules that enable signature comparisons or enrichment analyses. Availability and implementation: The TOXsIgN repository is freely accessible at http://toxsign.genouest.org. Website implemented in Python, JavaScript and MongoDB, with all major browsers supported. Supplementary information: Supplementary data are available at Bioinformatics online.
STUDY QUESTION:Does ibuprofen use during the first trimester of pregnancy interfere with the development of the human fetal ovary?SUMMARY ANSWER:In human fetuses, ibuprofen exposure is deleterious for ovarian germ cells.WHAT IS KNOWN ALREADY:In utero stages of ovarian development define the future reproductive capacity of a woman. In rodents, analgesics can impair the development of the fetal ovary leading to early onset of fertility failure. Ibuprofen, which is available over-the-counter, has been reported as a frequently consumed medication during pregnancy, especially during the first trimester when the ovarian germ cells undergo crucial steps of proliferation and differentiation.STUDY DESIGN, SIZE, DURATION:Organotypic cultures of human ovaries obtained from 7 to 12 developmental week (DW) fetuses were exposed to ibuprofen at 1-100 μM for 2, 4 or 7 days. For each individual, a control culture (vehicle) was included and compared to its treated counterpart. A total of 185 individual samples were included.PARTICIPANTS/MATERIALS, SETTING, METHODS:Ovarian explants were analyzed by flow cytometry, immunohistochemistry and quantitative PCR. Endpoints focused on ovarian cell number, cell death, proliferation and germ cell complement. To analyze the possible range of exposure, ibuprofen was measured in the umbilical cord blood from the women exposed or not to ibuprofen prior to termination of pregnancy.MAIN RESULTS AND THE ROLE OF CHANCE:Human ovarian explants exposed to 10 and 100 μM ibuprofen showed reduced cell number, less proliferating cells, increased apoptosis and a dramatic loss of germ cell number, regardless of the gestational age of the fetus. Significant effects were observed after 7 days of exposure to 10 μM ibuprofen. At this concentration, apoptosis was observed as early as 2 days of treatment, along with a decrease in M2A-positive germ cell number. These deleterious effects of ibuprofen were not fully rescued after 5 days of drug withdrawal.LARGE SCALE DATA:N/A.LIMITATIONS, REASONS FOR CAUTION:This study was performed in an experimental setting of human ovaries explants exposed to the drug in culture, which may not fully recapitulate the complexity of in vivo exposure and organ development. Inter-individual variability is also to be taken into account.WIDER IMPLICATIONS OF THE FINDINGS:Whereas ibuprofen is currently only contra-indicated after 24 weeks of pregnancy, our results points to a deleterious effect of this drug on first trimester fetal ovaries ex vivo. These findings deserve to be considered in light of the present recommendations about ibuprofen consumption pregnancy, and reveal the urgent need for further investigations on the cellular and molecular mechanisms that underlie the effect of ibuprofen on fetal ovary development.
(Abstracted from Hum Reprod 2018;33(3):482–493) Nonsteroidal anti-inflammatory drugs (NSAIDs) are one of the most commonly used over-the-counter medications for treatment of pain, inflammation, and fever. In 2013, up to 28.3% of pregnant women reported the use of ibuprofen at some stage during their pregnancy.