
Abstract Background The extensive volume of biomedical scientific literature requires efficient methods for retrieving relevant documents based on semantic technologies and biomedical concepts. While embedding-based methods have shown improvements over traditional keyword-based methods, the integration of domain-specific terminologies like Medical Subject Headings (MeSH) into these models remain underexplored. Methods This study compares three hybrid methods that integrate MeSH-based annotations with document embeddings (called “pre-annotation”, “post-annotation” and“post-reduction”). We benchmark these hybrid methods against the following traditional methods: TF-IDF, standard neural embeddings (Word2Vec, fastText, Doc2Vec) and publicly available transformer-based models (BioBERT, SciBERT, SPECTER, SapBERT), using cosine similarity and Word Mover’s Distance (WMD) as evaluation metrics. The benchmark experiments are based on the RELISH corpus, a manually curated dataset of PubMed articles where experts have labeled pairs of documents with regards to their relevance to each other, providing a 2-class (relevant vs. non-relevant) as well as a 3-class (relevant, partially relevant, non-relevant) judgment. Results Transformer-based models, particularly fine-tuned BioBERT and SciBERT align best with the expert judgements after fine-tuning. Among non-transformer methods, Doc2Vec and MeSH-based hybrid methods also perform well, demonstrating the benefits from combining structured biomedical vocabularies with embedding methods. Our experiments deliver extensive results showing that the baseline performance of 76–78% precision at position 5 can be achieved through almost all approaches, improvements of 2–4% with MeSH concepts can be achieved, but performances up to 90% is left to the fine-tuned large-scale public models. Conclusion The performance gains from the integration of concepts may be underwhelming, however the benefits lie in the successful integration and benchmarking of structured vocabularies with embedding methods, the applicability of these techniques to aligning literature with other data sources via a controlled vocabulary and the potential for stronger performance on tasks and corpora where concept-based resources are better suited. All experiments have been conserved as a Dockerized pipeline, making the full benchmarking workflow reproducible and supporting future research in biomedical document retrieval.
Abstract Background Safe and accurate sharing of healthcare information is essential for good health and well-being. Information must be exchanged among multiple parties, making the transformation between different information structures unavoidable. It has been suggested that such a transformation could be automated if information were sufficiently structured. This study aims to explore whether terminology binding to SNOMED CT can act as a bridge between different information structures. Methods This study employed a mixed-methods approach using two information structures from the breast cancer pathology domain. Pairs of semantically equivalent user interface terms from these structures were identified by a clinician and verified by two informaticians. The terms were considered semantically equivalent when a clinician would have used them to convey the same information. The pairs of SNOMED CT concepts bound to each of the interface terms were then compared and categorised as transformable, partially transformable, or non-transformable. Non-transformable concept pairs were analysed qualitatively, using an inductive content analysis approach, to identify underlying causes. Results 73 term pairs were identified. In 70% of these, the user interface terms were terminology bound to the same SNOMED CT concept. Among term pairs with terms bound to different concepts, 9% were fully transformable, 27% partially transformable, and 64% non-transformable. All is-a-related concept pairs were either fully or partially transformable, whereas concept pairs from different hierarchies were never transformable. Major obstacles within the same hierarchy included the use of primitive concepts and nearly identical sibling concepts. For concept pairs across different hierarchies, divergent terminology-binding strategies—such as handling of integers and “other” values—hindered transformability. Additional issues were observed with local user interface terms and SNOMED CT translations. Conclusions Although concordance between the two structures was relatively high (70%), transformability among semantically equivalent term pairs bound to different SNOMED CT concepts was low (9%). This can be improved on the terminology-binding side through guidelines and shared terminology bindings, and on the ontology side by increasing the proportion of sufficiently defined concepts and ensuring the quality of descriptions and translations. These findings are significant for practitioners seeking to enhance interoperability, as well as for SNOMED International and researchers focusing on terminology quality and transformation approaches.
Unstructured clinical narratives in electronic health records contain essential information for healthcare delivery and research. However, the presence of personally identifiable information poses significant privacy risks, which limit the secondary use of data. Therefore, reliable automated de-identification is a prerequisite for the reuse of clinical texts. This study aims to evaluate different strategies for Spanish clinical text de-identification by assessing the performance of language models on synthetic and real-world datasets. This study presents a comparative evaluation of de-identification approaches using two datasets in Spanish: MEDDOCAN, a publicly available synthetic corpus, and ObstEHR, real-world clinical narratives from an obstetrics department. Three de-identification strategies were assessed: (i) inference with a pre-trained task-specific model used as-is, (ii) prompt-based inference using locally deployed large language models (LLMs), and (iii) fine-tuned task-specific models. Model performance was evaluated using precision, recall, and F1 score on test sets. In addition, a text preservation metric was introduced to assess prompt-based LLMs, and the impact of training set size was analyzed using progressively larger subsets of annotated ObstEHR data. Models without fine-tuning and prompt-based LLMs showed limited performance, with macro-averaged F1 scores ranging from 0.14 to 0.5 on ObstEHR and from 0.19 to 0.5 on MEDDOCAN. Fine-tuned models achieved higher performance, reaching macro-averaged F1 scores of up to 0.956. Learning curve analyses showed consistently high precision and gradual improvements in recall as the amount of training data increased, with high performance achieved with moderate amounts of annotated data. Fine-tuning-based approaches outperform models without task-specific adaptation and prompt-based strategies. Despite their flexibility in generative settings, LLM-based prompt strategies show limited reliability and information preservation in clinical de-identification. Prompt-based LLMs make slight textual modifications that cause token misalignment, leading to a subsequent decrease in evaluation metrics. Therefore, domain-specific fine-tuning remains the most effective strategy for real-world clinical text de-identification.
Behavioral and social science research (BSSR) is essential for understanding human behavior and informing interventions, policy, and health outcomes; however, such research faces persistent challenges related to fragmented knowledge, imprecise and inconsistent terminology, and limited interoperability across studies. Ontologies provide a promising solution by enabling standardized, machine-readable representations of concepts and relationships to support data integration, knowledge synthesis, and reproducibility. We present an initial overview of the ACCELERATE-BASSO Research Network, an NIH-supported consortium established to advance BSSR through ontology-driven approaches. Specifically, we describe the role of the Dissemination and Coordination Center (DCC) and the Best Practices Working Group (BPWG) in supporting a network of projects focusing on accelerating BSSR through ontology development and use. Based on the consortium’s early experiences, we summarize emerging ontology-driven practices, current interoperability efforts, and common methodological considerations identified across the participating projects. Rather than proposing a formal consensus framework, this work provides an initial consortium perspective on current ontology development activities and shared lessons learned. We also discuss the potential role of ontologies in supporting AI-related applications in BSSR. This paper provides an initial consortium perspective on ontology-driven approaches for advancing BSSR. By synthesizing early experiences, emerging practices, and shared challenges across multiple projects, it offers insights to inform future ontology development, interoperability, and community consensus while laying the foundation for more systematic evaluation and broader adoption within the BSSR community.
Ontologies support data integration and knowledge discovery across multiple life science disciplines through the application of formal semantics that define hierarchical and cross-hierarchical relationships which enable automated reasoning capabilities. At the same time, tabular data is by far the most commonly produced format of data in most disciplines. This is especially true in biodiversity dataset production. Creating rich semantic data from tables can be complex because of a mismatch between flat tabular data and hierarchical ontology models. Here, we present a practical approach for closing this gap using Linked Data Modeling Language (LinkML) in combination with OWL punning. LinkML allows tabular schemas to be formally defined, validated, and automatically documented in combination. OWL punning allows schema elements to function simultaneously as both classes and data properties. This approach has many practical advantages that we showcase using mammal trait data assembled from the Ranges digitization network. Our work demonstrates how using LinkML + punning simplifies alignment of trait data in tabular format with ontologies such as the FuTRES Ontology of Vertebrate Traits (FOVT). Additionally, we show multiple advantages of our practical approach, including: maintaining synchronization between schema, ontology, and documentation, reducing modeling overhead, and preserving compatibility with conventional relational data workflows. On the downside, full logical reasoning is limited relative to more verbose RDF translations. Often this is a perfectly acceptable tradeoff for many use cases focused on data discovery and integration. The LinkML + punning framework shown here provides a scalable and pragmatic strategy for building semantically rich, ontology-aligned data resources that lower the barrier to semantic interoperability in biodiversity and trait informatics.
This research addresses the development and validation of an information model for pregnancy monitoring that, through the Fast Healthcare Interoperability Resources (FHIR) specification, seeks to promote interoperability for the quality of prenatal care by multidisciplinary teams in Brazilian primary healthcare contexts. The Design Science Research (DSR) approach is applied to propose and evaluate the model, through quality strategies, where iterations increase the mode maturity in accordance with healthcare scenarios. A set of use cases is constructed from user stories in primary care that abstract health concepts for pregnancy monitoring. The information model is presented in a hierarchical structure, with the categorization of related concepts grouped into pillars, in the context of the integration of health professionals who work in a complementary manner. The traceability of the use cases in relation to the semantic pillars of the information model, the implementation of the model from a FHIR perspective, and how the model complies with the ISO 13972:2022 standard (Health informatics – Clinical information models - Characteristics, structures and requirements) are analyzed. This work makes several key contributions, starting with the development of an information model for pregnancy monitoring. This model integrates different areas of primary care to promote a holistic and personalized approach to prenatal care. To ensure data can be shared effectively, it establishes health information interoperability using the HL7 FHIR R4 standard, which involved mapping resources and creating profiles as specified in an implementation guide. The study rigorously applies the Design Science Research method, utilizing multidisciplinary consensus strategies to scientifically improve and evolve the information model.
Background Semantic clarity and standardization in representing dental restoration materials are essential for ensuring interoperability across research and clinical settings. However, existing ontologies, including those in the Open Biological and Biomedical Ontology (OBO) Foundry, provide minimal coverage of this clinically relevant domain. Methods Guided by OBO Foundry principles, developed an ontological extension of the Oral Health and Disease (OHD) ontology to formally represent dental restoration materials. The modeling process included iterative expert input to ensure clinical accuracy and precise terminology. Logical definitions and subclass hierarchies were created to classify materials by composition and microstructure. Results The extended ontology introduces a logically structured taxonomy of dental restoration materials using genus-differentia definitions. Its class hierarchy aligns with domain-specific classification systems and captures current and emerging material types. Reuse of class from the Chemical Entities of Biological Interest (ChEBI) and the Environment Ontology (ENVO) as well as relations from the Relation Ontology (RO) supports semantic integration with related biomedical ontologies. Conclusion This work highlights how domain-informed ontology design can effectively capture complex material knowledge. The extensive collaboration integrates expertise from ontology engineers, academic researchers, and practicing clinicians. This approach enables the capture of intricate, often-disregarded real clinical factors, resulting in a model that reflects both the structural properties of materials and their practical use. The extended OHD ontology improves the semantic representation of dental restoration materials and provides a scalable foundation for advancing oral health informatics.
BackgroundData-driven AI models in clinical care increasingly rely on integrating data from heterogeneous sources to improve predictive performance. Traditional Electronic Health Record (EHR) systems however pose significant challenges for data integration due to the use of diverse standards and inconsistent semantics. Personalized Health Knowledge Graphs (PHKG), which leverage biomedical ontologies to encode and interconnect clinical knowledge, have emerged as a promising solution for unifying heterogeneous health data. Yet the process of PHKG integration can result in incomplete and ambiguous representations.ObjectiveOur primary objective is to apply advanced embedding techniques to mitigate the incompleteness problem in PHKGs. We propose a framework that combines schema-based semantic integration, domain-specific ontology alignment, and patient context representations through embedding techniques to improve incomplete and ambiguous PHKG representations. We further demonstrate the framework's capabilities for producing enriched representations in a use case concerning cardiovascular outcome prediction by training machine learning models which utilize the learned embeddings.ResultsWe embed EHR from a critical care unit as structured PHKGs mapped with two different schemas for comparison across knowledge completion and alignment tasks following a modular design. Each module is individually optimized using three baseline embedding methods with enhanced loss functions. We evaluate each task for both accuracy and semantic consistency and show the contribution of each module to the overall performance. We find and report settings for each module in which the proposed framework outperforms baseline results. The learned representations are subsequently used to generate patient contexts for the task of Heart Failure diagnosis as a use case. Our experiments demonstrate that semantically enhanced PHKG embeddings achieve better precision and recall scores compared to baseline models.ConclusionOur proposed method addresses the challenges in generating heterogenuous Personalized Health Knowledge Graphs (PHKG) through a modular framework that integrates schema mapping, ontology alignment, and contextual patient embeddings. Results on real-world patient records support the potential for improved performance in clinical decision making. Particularly in scenarios where sparse and fragmented health records prove problematic for data driven applications, our method can provide a robust approach for the disambiguation and completion of coded information. Our implementation is available at: https://github.com/AIDAVA-DEV/kge-framework.
The latest release of the BioAssay Ontology (BAO), version 2.8.16, introduces major updates that expand its ability to describe and categorize assays related to pharmacokinetics, pharmacodynamics, and safety pharmacology. These refinements, driven by collaboration with the Semantic Enrichment of Electronic Laboratory Notebook Data project, an industry initiative led by the Pistoia Alliance, address previously identified gaps in ontology coverage. New terms were added for disease models, preclinical study parameters, toxicological measurements, and detailed classifications of pharmacokinetics and pharmacodynamics assays. The update also incorporates biologically relevant target classes, including cytochrome P450 enzymes, solute carrier transporters, and uridine diphosphate-glucuronosyltransferase enzymes. Beyond content expansion, structural improvements include reassignment of terms to more specific and semantically appropriate parent classes, refinement of class dependencies, and enhanced alignment with external ontologies. Anatomical terms were reorganized to follow the Uber Anatomy Ontology hierarchy, and new chemical classes were incorporated to improve compatibility with the Chemical Entities of Biological Interest Ontology. Hundreds of additional axioms were added using existing object properties to capture assay formats, endpoints, detection methods, substrates, design strategies, and biological context. These refinements improve BAO’s semantic precision, interoperability, and reasoning capabilities. As a demonstration of these capabilities, we present a reasoning-based use case in which BAO’s equivalent class axioms enable automated classification of passive cell permeability and active efflux substrate assays. The ontology infers broader mechanistic categories from shared modeled characteristics, grouping transporter-specific and cell-based permeability assays under their mechanistic parents, while excluding assays that do not meet all restrictions. This example demonstrates its ability to support precise, inference-driven retrieval and integration of assay classes. By extending its scope and improving the clarity, consistency, and semantic depth of its classifications, BAO continues to serve as a vital resource for organizing pharmacological data and advancing research in both academic and industrial settings.
Background Entity linking maps textual mentions with entities in vocabularies. While accuracy is the primary metric for entity linking evaluation, it fails to capture the complexity of model behaviour. Results We propose an entity linking evaluation framework that clarifies the target objects of metrics calculation, incorporates term hierarchy, generates performance profiles, and summarises them as model characteristics. The framework emphasises hierarchical vocabulary structures, shifting the focus from text-level label matching to semantic comparison between concepts. We illustrate the competence and utility of this framework through a case study on disease entity linking. Conclusion Our results highlight the importance of aligning evaluation metrics with application-specific requirements, and provide structured hierarchical error analysis for entity linking, paving the way for more nuanced and practical assessments of entity linking systems.
Background The geriatric population is a vulnerable population that has higher rate of chronic disease and represents the largest portion of healthcare delivered. This population is vulnerable to medication non-adherence. Patient medication non-adherence is a problem that can lead to increase morbidity and mortality and waste of resources. Reasons for this phenomenon are multifactorial and include poor health literacy. Results To address this issue, we created a patient-directed drug information knowledge graph using both patient-directed resources and the Vaccine Information Statement Ontology (VISO), and attempt to validate the knowledge graph with common patient questions. This knowledge graph model (Patient-centric Drug Knowledge Graph) includes a terminological size of 577 term nodes and 113 links. We also created five knowledge graphs using the Patient-centric Drug Knowledge Graph (PcDKG) framework that represent five top medications (atorvastatin, levothyroxine, lisinopril, metformin, and amlodipine) used by the geriatric population. The common patient questions were converted to SPARQL queries to assess the coverage the model. Conclusion This initial development of the PcDKG is first knowledge graph that synthesizes concepts relating to drug information needs of the consumer population, specifically the geriatric population. This work also evolves the previous work of the VISO knowledge graph to cover wider range of medication knowledge for patients. PcDKG aims to be integrated in patient-directed tools to leverage its knowledge base, and future direction is to integrate PcDKG for digital health tools directed to the geriatric population. Our work is publicly available on our GitHub repository along with the five instance knowledge graph models.
BACKGROUND: Cholangiocarcinoma (CCA) is a critical public health problem in Thailand. The prevalence is much higher than other areas in the world. Data about CCA are stored in different data sources and standards in both research data sets and electronic health records (EHR). OBJECTIVE: This study aims to integrate and analyze CCA data from various sources to investigate risk factors and develop prediction models using the Cholangiocarcinoma Ontology (CCAO). METHODS: Datasets from Thailand were annotated with CCAO and analyzed using ontology-based term enrichment methods. We applied ontology term enrichment analysis, similar to that used with the Gene Ontology, for identifying significant risk factors for suspected CCA and patients with CCA. Our program provided a list of significant terms associated with CCA and a visualization of the ontology hierarchy with significant terms highlighted. The outputs of the term enrichment analyses have been used as the inputs to machine learning classification tasks. RESULTS: The results confirmed that indicators for CCA include dilated bile ducts, periductal fibrosis, and hepatic mass, based on ultrasound findings from several years prior. Our analysis also revealed demographic and lifestyle risk factors such as male gender, having no education, alcohol consumption, smoking, being a farmer, and having diabetes. We seeded a random forest classifier with the term enrichment results and predicted CCA patients with average 0.92 precision-recall curve score (0.023 standard deviation) with age, dilated bile ducts, periductal fibrosis, suspected CCA, and hepatic mass as the top five important features. CONCLUSIONS: These findings can be used to focus and monitor populations at risk for CCA. Expanding CCAO with molecular data related to CCA using ontology-driven term enrichment analysis and machine learning will help us to discover new hypotheses to decrease the morbidity and mortality of CCA in Thailand.
BACKGROUND:The Expression Constraint Language (ECL) is a powerful query language for SNOMED CT, enabling precise semantic queries across clinical concepts. However, its complex syntax and reliance on the SNOMED CT Concept Model make it difficult for non-experts to use, limiting its broader adoption in clinical research and healthcare analytics. OBJECTIVE:This work presents ECLed, a web-based tool designed to simplify access to ECL queries by abstracting the complexity of ECL syntax and the SNOMED CT Concept Model. ECLed is aimed at non-technical users, enabling the creation and modification of ECL queries and facilitating the querying of patient data coded with SNOMED CT. METHODS:ECLed was developed following a detailed requirements analysis, addressing both functional and non-functional needs. The tool supports the creation and editing of SNOMED CT ECL queries, integrates a processed Concept Model, and uses FHIR terminology services for semantic validation. Its modular architecture, with a frontend based on Angular and a backend on Spring Boot, ensures seamless communication through RESTful interfaces. RESULT:ECLed demonstrated high usability in a user survey. Technical validation confirmed that it reliably generates and edits complex ECL queries. The tool was successfully integrated into the DaWiMed research platform, enhancing clinical analysis workflows. It also worked effectively with clinical data in FHIR format, although scalability with larger datasets remains to be tested. DISCUSSION:ECLed overcomes the limitations of existing ECL tools by abstracting the complexity of both the syntax and the SNOMED CT Concept Model. It provides a user-friendly solution that enables both technical and non-technical users to easily create and edit ECL queries. CONCLUSION:ECLed offers a practical, user-friendly solution for creating SNOMED CT ECL queries, effectively hiding the underlying complexity while optimizing clinical research and data analysis workflows. It holds significant potential for further development and integration into additional research platforms.
The rapid expansion of biomedical literature has made comprehensive manual synthesis increasingly difficult to perform effectively, creating a pressing need for AI systems capable of reasoning across verified evidence rather than merely retrieving it. However, existing retrieval-augmented generation (RAG) methods often fall short when faced with complex biomedical questions that require iterative reasoning and multi-step synthesis. Here, we developed Queryome, a deep research system consisting of specialized large language model (LLM) agents that can adapt their orchestration dynamically to a wide range of queries. Using a hybrid semantic–lexical retrieval engine spanning 28.3 million PubMed abstracts, it performs iterative, evidence-grounded synthesis. On the MIRAGE benchmark, Queryome achieved 88.98 % accuracy, surpassing prior systems by up to 14 points, and improved reasoning accuracy on the biomedical Human’s Last Exam (HLE) subset from 15.8% to 19.3%. Moreover, in a task for constructing a review article, it earned the highest composite score in comparison with Deep Research from OpenAI, Google, Perplexity, and Scite.AI, reflecting its strong literature retrieval and synthesis capabilities.
Background. Ontology development is a complex, iterative process that traditionally requires extensive collaboration between ontology developers and subject matter experts (SMEs). While effective, this manual approach is time-consuming, labor-intensive, and prone to cognitive bias. To streamline early-stage ontology development and uncover concepts that might be overlooked through manual review alone, we applied automated topic modeling with BERTopic to extract topics, keywords, topic labels, and summaries from The Handbook of Solitude: Psychological Perspectives on Social Isolation, Social Withdrawal, and Being Alone and Gerotranscendence: A Developmental Theory of Positive Aging . The extracted topic labels were used as candidate concepts for the Promoting Healthy Aging through Semantic Enrichment of Solitude Research (PHASES) Ontology. Methods. We implemented and compared two BERTopic pipelines: (1) the default configuration and (2) a custom preprocessing pipeline incorporating part-of-speech filtering and n-gram tuning.The pipeline is customized to flexibly extract any specified number of topics and keywords based on user-defined parameters. To compare and merge topic modeling outputs across solitude and gerotranscendence, we used semantic embeddings of topic labels, keywords, and summaries from the custom pipeline. Cosine similarity identified semantically matched topic pairs above a set threshold, enabling categorization and integration into a merged conceptual framework that bridges both domains. Results. From the solitude corpus, BERTopic generated 244 initial topics, which SME review refined to 32 high-quality topics with the custom pipeline and 46 with the default pipeline. For the gerotranscendence corpus, the pipeline produced 172 initial topics, refined to 33 (custom) and 32 (default) high-quality topics. Across both corpora, BERTopic contributed 90 ontology terms, 52 from the solitude corpus and 38 from the gerotranscendence corpus. Visual evaluations, including keyword score bar charts, hierarchical clustering dendrograms, and BART-generated summaries, revealed that the custom pipeline produced more fine-grained, domain-specific topics, while the default pipeline offered broader thematic coverage and clearer labels. Certain theory-laden concepts, however, required SME interpretive input. Conclusions. BERTopic provided an efficient, semi-automated approach for identifying candidate ontology terms from domain literature, supporting both breadth and specificity in concept capture. Integrating semantic similarity analysis across thematic domains revealed conceptual intersections and overlaps, enhancing the semantic foundation of the PHASES Ontology and offering a replicable method for cross-domain ontology development.
Background. Dental caries is an oral health condition in which cariogenic bacteria demineralize and decay teeth. It arises due to interaction between the host, environment, and oral microbiome. Current terminologies and ontologies, however, do not accurately represent the important role that the microbiome has in the formation of carious lesions. Rather, they focus on the anatomical features of carious lesions and often obfuscate the distinctions between dental caries as a disease affecting a tooth, as lesions that are produced because of the disease, and as lesions produced as a result of dysbiosis in the oral microbiome. To capture the current state of evidence and provide flexibility for evolving literature on host-environment-microbiome interactions, there is a need to revise and expand the ontological framework for dental caries. Results. Several established terminologies and ontologies were reviewed for terms used to represent dental caries and the oral microbiome. We found that they either did not represent or misrepresented the current scientific understanding of caries and its relation to the microbial dysbiosis. As a result of these deficiencies, we added terms and relations to the Oral Health and Disease Ontology (OHD) that more accurately represent how oral microbial dysbiosis influences the development of dental caries. Conclusions. The Oral Health and Disease Ontology is an advance over existing ontologies for representing the impact of oral microbial dysbiosis on dental caries. It provides a semantic framework that better serves the needs of cariology researchers and can more easily incorporate new oral microbiome findings.
BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., ‘CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W’ that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (relative to that event); and (ii) a compact version named ‘stub’ (e.g., ‘CT01001LTR0N401T1W’) optimized for filenames, pipelines, and labeling. ClarID is supported by an open-source reference implementation, ClarID-Tools, a command-line tool that processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID’s utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID is designed to complement, not replace, persistent identifiers such as UUIDs, by providing a human-readable layer that enhances interpretability and facilitates metadata curation. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools .
Background:Vaccines have the ability to induce a range of immune responses under different conditions, for example stimulating the production of neutralizing antibodies to block pathogen entry or activating cytotoxic T-cells to eliminate infected cells. Many such immune responses have not been thoroughly examined and classified. The Vaccine Ontology (VO) is a community-based ontology in the domain of vaccinology. We here describe how VO is used to represent the variety of immune responses associated with vaccines, together with associated biomarkers and profiles. Results:The VO differentiates 'vaccination' and 'vaccine immunization.' The former is a process of administering a vaccine in vivo; the latter is the outcome of vaccine induction of immune response. This distinction is critical for understanding both the procedure of vaccination and the resulting immune effects. VO also models and represents various vaccine-induced responses at multiple biological levels, including population, organism, organ/tissue, cell, and gene/protein levels. Such an approach captures the complexity of vaccine-induced immunity, from population-wide trend (for example: herd immunity) to molecular mechanisms. VO defines immune biomarkers as material entities such as neutralizing antibodies that signify humoral immune response, and IFN-gamma that is indicative of cell-mediated responses. Such biomarkers provide measurable indicators of the immune system's functional state post vaccination, enabling robust evaluation of vaccine efficacy. VO classifies 'immune response profile' and 'correlated profile (or correlate) of immune protection' as 'process profiles,' a class in the Basic Formal Ontology (BFO 2.0). Immune response profiles, such as 'Th1 (or Th2)-biased profile,' can be induced by various vaccines and vaccine adjuvants. Different types of 'correlated profile of immune protection' are also identified, such as mechanistic and non-mechanistic correlates of immune protection. Such distinctions help us to quickly identify biomarkers and associated prediction and measurement of different kinds of vaccine protection. Conclusion:The important immune-related terms for immune biomarkers, profiles, and responses are modeled ontologically in VO together with their interrelations. The results support enhanced classification and analysis of vaccine-induced immune responses and related biomarkers and immune profiles, leading to further understanding of the vaccine immune mechanisms and enhanced vaccine research and development.
Around 30 million people in Europe are affected by a rare (or orphan) disease, defined as a condition occurring in fewer than 1 in 2,000 individuals. The primary challenge is to automatically and efficiently identify scientific articles and guidelines that address a particular rare disease. We present a novel methodology to annotate and index scientific text with taxonomical concepts describing rare diseases from the OrphaNet taxonomy. This task is complicated by several technical challenges, including the lack of sufficiently large, human-annotated datasets for supervised training and the polysemy/synonymy and surface-form variation of rare disease names, which can hinder any annotation engine. We introduce a framework that operationalizes OrphaNet for large-scale literature annotation by integrating the TERMite engine with curated synonym expansion, label normalization (including deprecated/renamed concepts), and fuzzy matching. On benchmark datasets, the approach achieves precision = 92
BACKGROUND:HL7 FHIR terminological services (TS) are a valuable tool towards better healthcare interoperability, but require representations of terminologies using FHIR resources to provide their services. As most terminologies are not natively distributed using FHIR resources, converters are needed. Large-scale FHIR projects, especially those with a national or even an international scope, define enormous numbers of value sets and reference many large and complex code systems, which must be regularly updated in TS and other systems. This necessitates a flexible, scalable and efficient provision of these artifacts. This work aims to develop a comprehensive, extensible and accessible toolkit for FHIR terminology conversion, making it possible for terminology authors, FHIR profilers and other actors to provide standardized TS for large-scale terminological artifacts. IMPLEMENTATION:Based on the prevalent HL7 FHIR Shorthand (FSH) specification, a converter toolkit, called BabelFSH, was created that utilizes an adaptable plugin architecture to separate the definition of content from that of the needed declarative metadata. The development process was guided by formalized design goals. RESULTS:All eight design goals were addressed by BabelFSH. Validation of the systems' performance and completeness was exemplarily demonstrated using Alpha-ID-SE, an important terminology used for diagnosis coding especially of rare diseases within Germany. The tool is now used extensively within the content delivery pipeline for a central FHIR TS with a national scope within the German Medical Informatics Initiative and Network University Medicine and demonstrates adequate usability for FHIR developers. DISCUSSION:The first development focus was geared towards the requirements of the central research FHIR TS for the federated FHIR infrastructure in Germany, and has proven to be very useful towards that goal. Opportunities for further improvement were identified in the validation process especially, as the validation messages are currently imprecise at times. The design of the application lends itself to the implementation of further use cases, such as direct connectivity to legacy systems for catalog conversion to FHIR. CONCLUSIONS:The developed BabelFSH tool is a novel, powerful and open-source approach to making heterogenous sources of terminological knowledge accessible as FHIR resources, thus aiding semantic interoperability in healthcare in general.