Collecting the relevant list of patient phenotypes, known as deep phenotyping, can significantly improve the final diagnosis. As textual clinical reports are the richest source of phenotypes information, their automatic extraction is a critical task. The main challenges of this Information Extraction (IE) task are to identify precisely the text spans related to a phenotype and to link them unequivocally to referenced entities from a source such as the Human Phenotype Ontology (HPO). Recently, Language Models (LMs) have been the most suc-cessful approach for extracting phenotypes from clinical reports. Solutions such as PhenoBERT, relying on BERT or GPT, have shown promising results when applied to datasets built on the hypothesis that most phenotypes are explicitly mentioned in the text. However, this assumption is not always true in medical genetics. Hence, although the LMs carry powerful semantic abilities, their contributions are not clear compared to syntactic string-matching steps that are used within the current pipelines. The goal of this study is to improve phenotype extraction from clinical notes related to genetic diseases. Our contributions are threefold: First, we provide a clear definition of the phenotype extraction task from free text, along with a high-level overview of the involved functions. Second, we conduct an in-depth analysis of PhenoBERT, one of the best existing solutions, to evaluate the proportion of phenotypes predicted with simple string-matching. Third, we demonstrate how utilizing and incorporating large language models (LLMs) for span detection step can improve performance especially with implicit phenotypes. In addition, this experiment revealed that the annotations of existing dataset are not exhaustive, and that LLM can identify relevant spans missed by human labelers.
The expansion of multi-omics datasets raises significant challenges for data integration and querying. To overcome these challenges, we developed a generic RDF-based integration schema that connects various types of differential -omics data, epigenomics, and regulatory information. This schema employs the FALDO ontology to enable querying based on genomic locations. It is designed to be fully or partially populated, providing both flexibility and extensibility while supporting complex queries. We validated the schema by reproducing two recently published studies, one in biomedicine and the other in environmental science, proving its genericity and its ability to integrate data efficiently. This schema serves as an effective tool for managing and querying a wide range of multi-omics datasets.
In systems biology, the study of biological pathways plays a central role in understanding the complexity of biological systems. The massification of pathway data made available by numerous online databases in recent years has given rise to an important need for standardization of this data. The BioPAX format (Biological Pathway Exchange) emerged in 2010 as a solution for standardizing and exchanging pathway data across databases. BioPAX is a Semantic Web format associated to an ontology. It is highly expressive, allowing to finely describe biological pathways at the molecular and cellular levels, but the associated intrinsic complexity may be an obstacle to its widespread adoption. Here, we report on the use of the BioPAX format in 2024. We compare how the different pathway databases use BioPAX to standardize their data and point out possible avenues for improvement to make full use of its potential. We also report on the various tools and software that have been developed to work with BioPAX data. Finally, we present a new concept of abstraction on BioPAX graphs that would allow to specifically target areas in a BioPAX graph needed for a specific analysis, thus differentiating the format suited for representation and the abstraction suited for contextual analysis.
Motivation Biological Pathway Exchange (BioPAX) is a standard language, represented in OWL, that aims to enable the integration, exchange, visualization and analysis of biological pathway data. While public databanks increasingly provide datasets in BioPAX format, their use remains below potential. Users may encounter challenges in harnessing the data due to the BioPAX intricately detailed underlying model. Moreover, extracting data demands specific technical skills, posing a barrier for many potential users. Results To address these obstacles, we developped BioPAX-Explorer. This toolis designed to facilitate the adoption and usage of BioPAX for extracting data or build algorithms and models, within the Python community. BioPAX-Explorer is a Python package that provides an object-oriented data model automatically generated from the BioPAX OWL specification. Moreover, it offers expressive query capabilities that shield users from BioPAX inner complexity. BioPAX-Explorer supports dataset building features, validation facilities and pre-build queries. It simplifies the extraction and processing of data from BioPAX sources by automatically generating SPARQL queries. BioPAX-Explorer also offers a user-friendly interface for Python users, allowing exhaustive exploration of large datasets through features such as memory-efficient query execution, entity-oriented queries without the need for SPARQL knowledge. It also allows to learn and reuse complex SPARQL queries for biological network analysis. Additionally, BioPAX-Explorer can accelerate the development of Python-based network analysis software, since it generates graph data structures from BioPAX queries and facilitates the creation of transparent, reproducible workflows based on the BioPAX OWL standard. Availability and implementation BioPAX-Explorer is freely available. We provide the source code, documentation, installation instructions and a Jupyter notebook with tutorial at ### Competing Interest Statement The authors have declared no competing interest.
MotivationTranscriptional regulation is performed by transcription factors (TF) binding to DNA in context-dependent regulatory regions and determines the activation or inhibition of gene expression. Current methods of transcriptional regulatory circuits inference, based on one or all of TF, regions and genes activity measurements require a large number of samples for ranking the candidate TF-gene regulation relations and rarely predict whether they are activations or inhibitions. We hypothesize that transcriptional regulatory circuits can be inferred from fewer samples by (1) fully integrating information on TF binding, gene expression and regulatory regions accessibility, (2) reducing data complexity and (3) using biology-based likelihood constraints to determine the global consistency between a candidate TF-gene relation and patterns of genes expressions and region activations, as well as qualify regulations as activations or inhibitions.ResultsWe introduce Regulus, a method which computes TF-gene relations from gene expressions, regulatory region activities and TF binding sites data, together with the genomic locations of all entities. After aggregating gene expressions and region activities into patterns, data are integrated into a RDF (Resource Description Framework) endpoint. A dedicated SPARQL (SPARQL Protocol and RDF Query Language) query retrieves all potential relations between expressed TF and genes involving active regulatory regions. These TF-region-gene relations are then filtered using biological likelihood constraints allowing to qualify them as activation or inhibition. Regulus provides signed relations consistent with public databases and, when applied to biological data, identifies both known and potential new regulators. Regulus is devoted to context-specific transcriptional circuits inference in human settings where samples are scarce and cell populations are closely related, using discretization into patterns and likelihood reasoning to decipher the most robust regulatory relations.
Abstract Motivation Molecular complexes play a major role in the regulation of biological pathways. The Biological Pathway Exchange format (BioPAX) facilitates the integration of data sources describing interactions some of which involving complexes. The BioPAX specification explicitly prevents complexes to have any component that is another complex (unless this component is a black-box complex whose composition is unknown). However, we observed that the well-curated Reactome pathway database contains such recursive complexes of complexes. We propose reproductible and semantically rich SPARQL queries for identifying and fixing invalid complexes in BioPAX databases, and evaluate the consequences of fixing these nonconformities in the Reactome database. Results For the Homo sapiens version of Reactome, we identify 5833 recursively defined complexes out of the 14 987 complexes (39%). This situation is not specific to the Human dataset, as all tested species of Reactome exhibit between 30% (Plasmodium falciparum) and 40% (Sus scrofa, Bos taurus, Canis familiaris, and Gallus gallus) of recursive complexes. As an additional consequence, the procedure also allows the detection of complex redundancies. Overall, this method improves the conformity and the automated analysis of the graph by repairing the topology of the complexes in the graph. This will allow to apply further reasoning methods on better consistent data. Availability and implementation We provide a Jupyter notebook detailing the analysis https://github.com/cjuigne/non_conformities_detection_biopax.
Les données en sciences de la vie sont massives, hétérogènes, compliquées et complexes. L'enjeu est d'en automatiser le traitement afin de le rendre systématique, ce qui nécessite à la fois intégration (data engineering) et méthodes d'analyse (data science). Ce chapitre montre comment le Web Sémantique offre une solution générique adoptée à large échelle par la communauté bioinformatique.
Omics technologies offer great promises for improving our understanding of diseases. The integration and interpretation of such data pose major challenges, calling for adequate knowledge models. Disease maps provide curated knowledge about disorders' pathophysiology at the molecular level adapted to omics measurements. However, the expressiveness of disease maps could be increased to help in avoiding ambiguities and misinterpretations and to reinforce their interoperability with other knowledge resources. Ontology is an adequate framework to overcome this limitation, through their axiomatic definitions and logical reasoning properties. We introduce the Disease Map Ontology (DMO), an ontological upper model based on systems biology terms. We then propose to apply DMO to Alzheimer's disease (AD). Specifically, we use it to drive the conversion of AlzPathway, a disease map devoted to AD, into a formal ontology: Alzheimer DMO. We demonstrate that it allows one to deal with issues related to redundancy, naming, consistency, process classification and pathway relationships. Furthermore, we show that it can store and manage multi-omics data. Finally, we expand the model using elements from other resources, such as clinical features contained in the AD Ontology, resulting in an enriched model called ADMO-plus. The current versions of DMO, ADMO and ADMO-plus are freely available at http://bioportal.bioontology.org/ontologies/ADMO.
Protein-protein interactions (PPIs) play an ubiquitous and fundamental role in all biological processes. Information on PPIs described in the literature is annotated and made available by several protein-interaction databases. Because most databases have their own curation rules and priorities, they often annotate overlapping sets of publications, which leads to redundancies. We developed a semantic-based approach which enables to accurately detect redundancies within PPI datasets from multiple databases. We applied this approach to assemble a "reproducible interactome", with PPIs supported by at least two methods or publications.
The Regulatory Circuits project is among the most recent and the most complete attempts to identify cell-type specific regulatory networks in Human. It is one of the largest efforts of public genomics data integration, based on data from the major consortia FANTOM5, ENCODE and Roadmap Epigenomics. This project is a main provider of biological data, cited more than 224 times (Google Scholar) and its resulting networks were used in at least 42 other articles. For such a general resource, reproducibility of both the outputs (regulation networks) and methods (data integration pipeline) is a major issue, since biological data are updated regularly. In addition, users may want to introduce new data into the Regulatory Circuits framework to provide networks about previously uncharacterized cell types or to add information about specific regulators, which require to re-execute the whole pipeline on the new data. In this article, we analyze the various factors limiting reproducibility of the Regulatory Circuits data and methods. Starting from a factual description of our understanding of the methods used in Regulatory Circuits, our contribution is two-fold: we propose (1) a characterization of the different levels of reusability, reproducibility and conceptual issues in the original workflow and (2) a new implementation of the workflow ensuring its consistency with the published description and allowing for an easier reuse and reproduction of the published outputs. Both are applicable beyond the case of Regulatory Circuits.
Medico-administrative databases contain information about patients’ medical events, i.e. their care trajectories. Semantic Web technologies are used by epidemiologists to query these databases in order to identify patients whose care trajectories conform to some criteria. In this article we are interested in care trajectories involving temporal constraints. In such cases, Semantic Web tools lack computational efficiency while temporal pattern matching algorithms are efficient but lack of expressiveness. We propose to use a temporal pattern called chronicles to represent temporal constraints on care trajectories. We also propose an hybrid approach, combining the expressiveness of SPARQL and the efficiency of chronicle recognition to query care trajectories. We evaluate our approach on synthetic data and real large data. The results show that the hybrid approach is more efficient than pure SPARQL, and validate the interest of our tool to detect patients having venous thromboembolism disease in the French medico-administrative database.
MotivationTranscriptional regulation is performed by transcription factors (TF) binding to DNA in context-dependent regulatory regions and determines the activation or inhibition of gene expression. Current methods of transcriptional regulatory networks inference, based on one or all of TF, regions and genes activity measurements require a large number of samples for ranking the candidate TF-gene regulation relations and rarely predict whether they are activations or inhibitions. We hypothesize that transcriptional regulatory networks can be inferred from fewer samples by (1) fully integrating information on TF binding, gene expression and regulatory regions accessibility, (2) reducing data complexity and (3) using biology-based logical constraints to determine the global consistency of the candidate TF-gene relations and qualify them as activations or inhibitions.ResultsWe introduce Regulus, a method which computes TF-gene relations from gene expressions, regulatory region activities and TF binding sites data, together with the genomic locations of all entities. After aggregating gene expressions and region activities into patterns, data are integrated into a RDF endpoint. A dedicated SPARQL query retrieves all potential relations between expressed TF and genes involving active regulatory regions. These TF-region-gene relations are then filtered using a logical consistency check translated from biological knowledge, also allowing to qualify them as activation or inhibition. Regulus compares favorably to the closest network inference method, provides signed relations consistent with public databases and, when applied to biological data, identifies both known and potential new regulators. Altogether, Regulus is devoted to transcriptional network inference in settings where samples are scarce and cell populations are closely related. Regulus is available at https://gitlab.com/teamDyliss/regulus
The recent development of data analysis provides opportunities for improving healthcare systems through analysis of health databases. However the thirst for data is conflicting with preserving the privacy of individuals. The generation of synthetic datasets may foster research on healthcare data analytics. It is mostly based on generative statistical models fitted on real data. Thus it still requires access to sensitive data. This article proposes a probabilistic relational model fitted on publicly available datasets. Public healthcare statistics provide valuable information to mimic statistical distributions and do not hold sensitive personal data. More specifically, we propose to generate a synthetic version of the national database of French insured patients. We do not only provide synthetic datasets, but a generator of datasets that can be used without any data access request. Experiments compare official statistics with those computed on synthetic datasets.
Transcriptional regulation -a major field of investigation in life science-is performed by binding of specialized proteins called transcription factors (TF) to DNA in specific, context-dependent regulatory regions, leading to either activation or inhibition of gene expression. Relations between TF, regions and genes can be described as regulatory networks, which are basically knowledge graphs containing the relationships between the different entities. Current methods of transcriptional regulatory networks inference rarely use information about TF binding or regulatory regions, often require a large number of samples and most of time do not indicate if the TF-gene relation is an activation or an inhibition. The resulting networks may then contain inconsistent relations and the methods are not applicable for common experimental or clinical settings, where the number of samples is limited. Therefore, based on our previous experience of formalizing the Regulatory Circuits data-sets with Semantic Web Technologies, we decided to create a new tool for transcriptional networks inference, that could solve these issues. Results: Our tool, Regulus, provides candidate signed TF-gene relations computed from gene expressions, regulatory region activities and TF binding sites data, together with the genomic location of all entities. After creating expressions and activities patterns, data are integrated into a RDF endpoint. A dedicated SPARQL query retrieves all potential TF-region relations for a given gene expression pattern. These ternary TF-region-gene pattern relations are then filtered and signed using a logical consistency check translated from biological knowledge. Regulus compares favorably to its closest network inference method, provides signs which are consistent with public databases and, when applied to real biological data, identifies both known and potential new regulators. We also provide several means to more stringently filter the output regulators. Altogether, we propose a new tool devoted to transcriptional network inference in settings where samples are scarce and cell populations may be closely related.The Regulus package is available at https://gitlab.com/gcollet/regulus