Hepatocellular carcinoma (HCC) is steadily increasing in incidence worldwide and requires data-driven approaches to improve diagnosis, prognosis, and therapeutic decisions. We describe the harmonization of IMALIVE -a real-world HCC cohort- with the common data model of the European Cancer Imaging Initiative (EUCAIM). IMALIVE integrates demographic, clinical, and imaging-related variables, which were aligned with core variables of the EUCAIM common data model through a robust mapping process and expert validation. The process achieved full coverage of the core dataset and added HCC-specific variables, including tumor staging and liver function scores. In addition, metadata from digital pathology was standardized using the international MIABIS/BBMRI model, extending interoperability across radiology and histology. This work demonstrates the feasibility of harmonizing local cohorts enabling their integration in the data catalogue of large scale federated platforms It highlights how medical experts' contributions to this harmonization process can enrich common models with clinically relevant variables for AI development in their domain of expertise.
Building federated infrastructure for European ICU data requires establishing standards for semantic interoperability across heterogeneous clinical settings. Even with standardized terminologies like LOINC and SNOMED CT, a single clinical concept may correspond to hundreds of alternative codes, precluding interoperability without explicit harmonization guidance. We developed the INDICATE Minimal Data Dictionary through an iterative, consensus-based process involving clinical experts and data scientists. The dictionary defines standardized OHDSI concept sets across nine clinical domains and supports six concrete use cases. Implementation is supported by a web-based application that provides expert-curated guidance for terminology selection and mapping. This approach reduces terminology ambiguity while remaining adaptable to local ICU practices, enabling scalable data harmonization across 15 European data providers from 12 countries.
Background:Semantic interoperability in health care, essential for seamless integration of information systems, is partially achieved through the use of terminologies and common data standards that define the semantic structure of data. Various complexities arise when using real-world health care data, including different interpretations of terms and concepts and gaps in domain coverage in standard terminologies. However, ensuring compatibility becomes increasingly challenging when big data are distributed across diverse repositories that use heterogeneous health care standards and overlapping terminologies. Ontologies are key solutions to bridge these gaps, enabling consistent semantic interoperability and data harmonization. Objective:We aim to develop and validate a hyperontology within the EUCAIM (Cancer Image Europe) project to semantically integrate and harmonize clinical, biological, and imaging metadata, along with associated data from heterogeneous, disparate cancer image data models, to achieve semantic interoperability in oncology and medical imaging. The hyperontology will be used to support several EUCAIM components, including the extract, transform, and load process; federated query; image annotation and segmentation; and ultimately, AI-federated processing. Methods:The ontology development process combines real-world data from a network of European projects on cancer imaging (AI for Health Imaging) with their semantic mappings, as well as conceptual unpacking and modeling of the Minimal Common Oncology Data Elements (mCODE) specifications. The mCODE is a core set of structured data elements for oncology electronic health records. The building process is supported by ontology grounding, layering, and modularization. We adopted this hybrid approach to simplify ontology design, semantically reflect oncology's essential entities and their interactions, and enhance the extensibility and reusability of the hyperontology. We initiated ontology development with a set of competency questions derived from the provided knowledge, which helped clarify the ontology's scope and requirements and identify inconsistencies or incomplete information. We also assessed whether the requirements were fulfilled by formalizing the competency questions using SPARQL. Results:We developed a FAIR hyperontology that semantically integrates and harmonizes clinical, biological, and imaging metadata and data spread across disparate sources. The ontology also captures and accurately represents oncology and medical imaging. The hyperontology, which covers various cancer types, is rich in axiomatizations and patterns, supporting the semantic understanding and harmonization of heterogeneous data. Additionally, semantic mappings are established across data models and standards, ensuring the efficient and meaningful sharing and integration of health care data. Finally, we evaluated the ontology model and demonstrated its applicability using real-world prostate and breast cancer use cases. Conclusions:EUCAIM's hyperontology is a valuable effort that provides a unifying framework for the essentials of oncology and medical imaging, facilitating communication among disparate and heterogeneous cancer image data models. The ontology model is evaluated and validated using multiple methods, demonstrating compliance with the specified ontological requirements. Challenges include ensuring that the ontology is scalable, extensible, and applicable, given the complexity and dynamic nature of the application domain.
Implementing patient-trial matching during Clinical Trials (CT) execution requires a time consuming and usually manual processing of eligibility criteria (EC) focusing on those that are both the most relevant at the pre-inclusion step and the most computable using Electronic Health Records data. The aim of this study is to design an ontology-based data dictionary for standardizing EC to support patient eligibility determination. We previously decomposed manually free text EC from 51 CT into data elements to be considered at the prescreening step to build the PENELOPE data dictionary. We aligned this dictionary to an existing common data model and ontology developed within the EUCAIM (Europe CAncer IMage) project. Results: A total of 100 data elements DE were captured in the alignment framework including 10 DE with direct match between PENELOPE and EUCAIM data models and 74 with partial match with syntactic or semantic gaps. Three DE were present only in PENELOPE and 13 only in EUCAIM. The structure of the PENELOPE data dictionary has been revised and its semantic content enriched with 292 concepts of the EUCAIM ontology that were relevant for prescreening. This study provides an ontology-based enrichment of the PENELOPE data dictionary used to represent prescreening-oriented EC from CT protocols ensuring the scalability of the PENELOPE prescreening use case beyond the 51 CT initially considered.
INTRODUCTION:L'utilisation des Systèmes d'Aide au Recrutement dans les Essais Cliniques (SAREC) appliqués aux données massives de santé pourrait faciliter l'inclusion des patients dans des essais cliniques. La revue exploratoire de la littérature a porté sur l'intégration en vie courante de SARECs appliqués aux dossiers patients informatisés (DPI) et ayant bénéficié d'une évaluation scientifique. MéTHODES: Dans cette revue, les publications ont été extraites de PubMed avec les mots-clés « essais cliniques comme sujet », « sélection de patients » et « entrepôt de données de santé ». Deux investigateurs en aveugle ont extrait les données d'intérêt : intégration des SAREC dans le processus de soin, contextes et solutions techniques, évaluation scientifique de leur utilisabilité et leur plus-value clinique. RéSULTATS: A partir de 801 articles, 41 ont été retenus. La majorité décrivait l'intégration de l'outil dans le parcours patient (27/41), les méthodes de transformation des critères d'éligibilité (34/41) et les défis liés à la qualité des données des DPIs. Près de la moitié (17/41) rapportait une évaluation prospective de la mise en correspondance patient-essai clinique et/ou l'amélioration des performances d'inclusion. DISCUSSION/CONCLUSION:Malgré la diversité des systèmes et des outils, des métriques communes émergent pour évaluer les SAREC selon leur utilisabilité, leurs caractéristiques techniques et leur impact. Des collaborations multidisciplinaires sont nécessaires pour promouvoir l'adoption des SAREC à grande échelle et améliorer la qualité et transformation des données de vie réelle et des critères d'éligibilité.
OBJECTIVE:Electronic Health Record (EHR) data is increasingly used to support patient recruitment in clinical trials (CTs). This scoping review focused on articles reporting an evaluation of EHR-based CT recruitment support systems (CTRSS). METHODS:Querying 3 biomedical literature databases using keywords as "clinical trial as topic", "patient recruitment", and "EHR", research articles describing evaluations of EHR-based CTRSS and published in English from 2013 to 2024 were included. Abstract screening was performed by 5 reviewers. Two reviewers performed full-text screening, data extraction and synthesis following 3 major themes: clinical workflow integration, technical aspects, and evaluation of EHR-based CTRSS. RESULTS:Of 927 screened articles, 44 qualified in scope, providing an evaluation of EHR-based CTRSS conducted within a diverse set of use cases, settings, and technologies. Results showed that 75% of the systems were CT rather than patient centric, 20% provided patient eligibility information at clinically relevant timepoints in the patient journey and 9% sent notifications to the targeted users. EHR data was transformed according to standard models in 32% of the studies. Eligibility criteria transformation was computer assisted in 48% of the studies. The main patient-trial matching technique was querying structured data in 55% of the cases. Finally, the majority (61%) of the evaluations were done retrospectively and led to a complete system evaluation from data transformation to patient - trial matching in 23%. CONCLUSION:Despite the diversity of system scopes and designs, metrics for effective evaluation throughout the entire cycle of system implementation across user, technology, and society levels are emerging. Research effort is still needed to extend the scope, quality, and computability of both the source EHR-data and EC and to conduct harmonized evaluation of these systems to demonstrate their impact on recruitment rates, delays, and costs.
This work aims to identify the Key Research Areas for building and deploying the semantic interoperability framework for the Intensive Medicine Data Space in Europe. A set of European experts defined four research areas and associated challenges: i) to characterize the value of Common Data Models, vocabularies and standardized data elements for creating the foundation for the standardization of Intensive Care Unit (ICU) data sets, ii) to derive the Common Data Model ensuring seamless data interoperability supporting the querying and access to high-quality, multi-modal data located on federated nodes according to the FAIR principles, iii) to define and support the data standardization process enabling efficient secondary use of ICU data set during the execution of key decision support use cases through the Intensive Medicine Data Space in Europe and beyond, and iv) to define requirements for anonymization and pseudonymization for the good balance between data protection and innovation.
Objectives. - Hospital-acquired infections (HAIs) during community disease outbreaks threaten vulnerable hospitalized patients. This study compares the outcomes of hospitalized patients who had COVID-19 as either a HAI or a community-acquired infection (CAI). Methods. - We conducted a retrospective cohort study involving adult patients hospitalized across 39 greater Paris University hospitals between January 27th, 2020, and April 21st, 2021, who tested positive for SARS-CoV-2 PCR during their stay. Patients were classified as CAI if they tested positive within 72 hours of admission and HAI if they tested negative within 72 hours but later positive. HAI was subclassified as possible (first positive test between days 4-7), probable (days 8-13), or definite (day 14 onward). Patients with probable or definite HAI were matched 1:3 to CAI patients for age, sex, and comorbidities, to compare intensive care unit (ICU) transfer and in-hospital death between both groups. Results. - Of 10,831 patients, 506 (4.7%) were classified as HAI. They were older and had more comorbidities. After matching, the 333 patients with probable or definite HAI were less likely to be transferred to the ICU (hazard ratio [HR] 0.57, 95% CI 0.38-0.85) compared to their 999 CAI controls and had a higher risk for in-hospital death (HR 1.58, 95% CI 1.16-2.14). Conclusion. - Patients with COVID-19 as a HAI face a higher risk of death compared to patients hospitalized with COVID-19 acquired in the community and are less likely to be admitted to the ICU. Strict infection control measures are needed during community outbreaks to protect hospitalized patients. (c) 2025 The Authors. Published by Elsevier Masson SAS on behalf of Societe Nationale Franc, aise de Medecine Interne (SNFMI). This is an open access article under the CC BY license (http:// creativecommons.org/licenses/by/4.0/).
Clinical Trial (CT) Recruitment Support Systems (CTRSS) querying Electronic Health Records (EHR) for patient-trial matching during CT execution have been expanding. Since free text CT eligibility criteria (EC) are not readily suitable for the automation of the EHR querying, the configuration of EHR-based CTRSS requires a time-consuming and usually manual processing of EC focusing on those that are the most relevant at the pre-inclusion (prescreening) step. The aim of this study is to provide a methodological approach to semi-automatically detect Prescreening-Oriented Eligibility Criteria (POEC) and build a library of POEC usable in the context of the development and evaluation of EHR-based Clinical Trial Recruitment Support Systems (CTRSS). We proposed an approach for decomposing free text EC into standardized elements and developing a rule-based algorithm to semi-automatically detect POEC. In addition, this paper describes the characteristics of a publicly available POEC library usable for CTRSS evaluation. An annotation framework consisting in 96 patterns of elementary EC categorized in 17 domains was used to annotate 381 free text EC from 20 CT dedicated to various cancer types. This training dataset was used to develop a rule-based algorithm detecting POEC. This study provides a methodological approach to (semi)-automatically spot POEC and store them in a library considering advances in the field of CTRSS. The PENELOPE-C2Q pipeline is designed to feed the PENELOPE POEC library, both having the potential to facilitate the reuse of EHR data for better participation of patients to research.
BACKGROUND:Clinical Research Informatics (CRI) is a subspeciality of biomedical informatics that has substantially matured during the last decade. Advances in CRI have transformed the way clinical research is conducted. In recent years, there has been growing interest in CRI, as reflected by a vast and expanding scientific literature focused on the topic. The main objectives of this review are: 1) to provide an overview of the evolving definition and scope of this biomedical informatics subspecialty over the past 10 years; 2) to highlight major contributions to the field during the past decade; and 3) to provide insights about more recent CRI research trends and perspectives. METHODS:We adopted a modified thematic review approach focused on understanding the evolution and current status of the CRI field based on literature sources identified through two complementary review processes (AMIA CRI year-in-review/IMIA Yearbook of Medical Informatics) conducted annually during the last decade. RESULTS:More than 1,500 potentially relevant publications were considered, and 205 sources were included in the final review. The review identified key publications defining the scope of CRI and/or capturing its evolution over time as illustrated by impactful tools and methods in different categories of CRI focus. The review also revealed current topics of interest in CRI and prevailing research trends. CONCLUSION:This scoping review provides an overview of a decade of research in CRI, highlighting major changes in the core CRI discoveries as well as increasingly impactful methods and tools that have bridged the principles-to-practice gap. Practical CRI solutions as well as examples of CRI-enabled large-scale, multi-organizational and/or multi-national research projects demonstrate the maturity of the field. Despite the progress demonstrated, some topics remain challenging, highlighting the need for ongoing CRI development and research, including the need of more rigorous evaluations of CRI solutions and further formalization and maturation of CRI services and capabilities across the research enterprise.
To develop and validate algorithms to automate the calculation of Healthcare Quality and Safety Indicators (HQSIs) from electronic health records (EHR) multi-sources data. Three HQSIs of interest were identified by experts. The required variables were identified and algorithms for extracting these variables from various sources of EHR data, prioritized based on their quality, were performed and evaluated on the Assistance Publique - Hôpitaux de Paris (AP-HP) and Bordeaux University Hospital clinical data warehouses (CDWs). Algorithm performances (positive predictive value (PPV), accuracy, f1-score) were assessed to a manual review of 100 EHRs randomly selected among newly referred patients with head and neck cancer (HNC) in 2023. Three HQSI were computed: number of newly referred HNC, number of HNC cases treated by front line anticancer surgery and the number of resected HNC cases requiring a post-operative surgery based on ICD-10 and their French Association pour le Développement de l’Informatique en Cytologie et en Anatomie Pathologique (ADICAP) pathology codes, 6,943 and 7,112 HNC patients were identified as referred to AP-HP and Bordeaux University Hospitals between 1998 and 2023, respectively. The algorithm related to the number of newly referred HNC diagnoses in 2023 had the following performances: PPV of 37
Introduction L'inclusion de patients dans les essais cliniques (EC) en oncologie est insuffisante. L'objectif de cette étude était de comprendre les modalités d'identification des patients éligibles à l'inclusion dans les EC, et les perspectives possibles de support informatique fondé sur les bases de données de santé. Méthodes Etude qualitative par entretiens semi-structurés d'acteurs impliquées dans la recherche clinique en oncologie depuis au moins cinq ans. Recrutement par échantillonnage délibéré pour intégrer divers profils (investigateurs, attachés de recherche clinique (ARC), support IT, promoteurs, patient) et institutions (public/privé, entreprises, tutelles). Entretiens enregistrés et retranscrits. Analyse qualitative par codage inductif des transcrits. Résultats Nous avons interrogé 24 participants, dont 10 oncologues et 5 ARCs. Nous mettons en évidence de grandes étapes stables entre institutions, avec une étape de préscreening très dépendante de l'oncologue référent du patient et de l'ARC, puis une étape de vérification manuelle chronophage. Le processus est très dépendant de la qualité des relations entre acteurs, et de la possibilité d'envisager l'inclusion dans un EC à chaque étape de soin du patient. La performance de ce processus est modérée par la culture de recherche clinique des centres : représentation des EC comme option de soin, satisfaction quant aux taux d'inclusion actuels, incitation à inclure pour les médecins. Par ailleurs, le préscreening est affecté par l'infrastructure en données mise à disposition des investigateurs. D'abord, l'absence de registre en temps réel des ECs ouverts est préjudiciable. Ensuite, l'automatisation même partielle du préscreening, pour aligner les critères d'inclusion aux caractéristiques des patients, supposerait à la fois une structuration standardisée de ces informations suivant une norme interopérable, et la compatibilité avec le dossier patient. Automatiser demanderait de redistribuer les rôles dans le processus, avec implication des spécialistes des données, et réflexion sur l'utilisateur principal d'une solution. La RCP apparaît comme un temps stratégique pour renforcer l'inclusion des patients dans les ECs. Conclusion L'inclusion dans les essais reste un processus manuel. Un support informatique serait utile mais les prérequis techniques et organisationnels ne sont pas encore remplis.
Objective To access Electronic Health Record (EHR) data, hospitals have implemented Clinical Data Warehouses (CDWs) using Extract Transform and Load (ETL) processes. While ETL performances are typically evaluated individually, our study examines the cumulative impact of ETLs on data availability. Methods Using a real multi-hospital CDW as a case study, we modeled EHR data processing from the software sources to the CDW's data store. We simulated a scenario where researchers aimed to reconstruct breast cancer care trajectories using EHR data. We calculated the size and characteristics of the data store population, and compared them to the original population. Results EHR data are recorded in various software depending on data category, hospital, and year, each requiring specific series of ETLs for integration in the CDW. Despite acceptable transfer rates for each ETL (range 73 %-100 %), cumulative losses led to study populations in the data store being up to 90 % smaller than anticipated when researchers required data exhaustivity for patients. Population size decreased steeply with the more data categories required. No difference was found in population characteristics between the data store and the original cohorts. Discussion & Conclusion Researchers should scrutinize data availability in CDWs as missing data could result from outsourced care, incomplete input, or underperforming ETLs. Integrating more data sources in CDWs increases the number of data routes, necessitating time for ETL implementation and maintenance, and increases data loss risks. Though commonly perceived as a “black box”, data transformation can significantly influence the reliability of populations studied in CDWs. Public interest Summary To access data generated during care, researchers build Clinical Data Warehouses (CDWs). CDWs are infrastructures composed of a series of processing steps to extract the data from the data source, transform it according to the needs and load it into a data store. Usually, the performances of these processing steps are evaluated one a time. However, each data point goes through a series of processing steps before being made available for research. In this study, we aim to evaluate the impact of the entire data processing pipeline on the availability of data points in a CDW by simulating a study on breast cancer and evaluating the impact on the size and the characteristics of the final cohort. The cumulative losses of the processing steps resulted in a population 90 % smaller than anticipated. The characteristics of the final population showed no difference to those of the original cohort.
Semantic interoperability is a growing and challenging subject in the healthcare domain. It aims to ensure a coherent and unambiguous exchange, use, and reuse of health information among different systems and applications. In the context of the EUCAIM (Cancer Image Europe) project, semantic interoperability among various heterogeneous cancer image data models is required to support the communication, integration, and sharing of data in a standardized and structured way. For this purpose, hyper-ontology is developed as a common semantic meta-model that bridges the disparate imaging and clinical knowledge of the various repositories in EUCAIM and supports their integration. EUCAIM’s hyper-ontology is also an application-based ontology targeted for federated semantic querying and image annotation. To facilitate the hyper-ontology building process and ensure the extensibility of the ontology model, an iterative hybrid well-founded approach that divides the ontology structure into layers and modules is established.
Interoperability is crucial to overcoming various challenges of data integration in the healthcare domain. While OMOP and FHIR data standards handle syntactic heterogeneity among heterogeneous data sources, ontologies support semantic interoperability to overcome the complexity and disparity of healthcare data. This study proposes an ontological approach in the context of the EUCAIM project to support semantic interoperability among distributed big data repositories that have applied heterogeneous cancer image data models using a semantically well-founded Hyperontology for the oncology domain.
Key Research Areas (KRAs) were identified to establish a semantic interoperability framework for intensive medicine data in Europe. These include assessing common data model value, ensuring smooth data interoperability, supporting data standardization for efficient dataset use, and defining anonymization requirements to balance data protection and innovation.
BACKGROUND:The SARS CoV-2 pandemic disrupted healthcare systems. We compared the cancer stage for new breast cancers (BCs) before and during the pandemic.METHODS:We performed a retrospective multicenter cohort study on the data warehouse of Greater Paris University Hospitals (AP-HP). We identified all female patients newly referred with a BC in 2019 and 2020. We assessed the timeline of their care trajectories, initial tumor stage, and treatment received: BC resection, exclusive systemic therapy, exclusive radiation therapy, or exclusive best supportive care (BSC). We calculated patients' 1-year overall survival (OS) and compared indicators in 2019 and 2020.RESULTS:In 2019 and 2020, 2055 and 1988, new BC patients underwent cancer treatment, and during the two lockdowns, the BC diagnoses varied by -18% and by +23% compared to 2019. De novo metastatic tumors (15% and 15%, p = 0.95), pTNM and ypTNM distributions of 1332 cases with upfront resection and of 296 cases with neoadjuvant therapy did not differ (p = 0.37, p = 0.3). The median times from first multidisciplinary meeting and from diagnosis to treatment of 19 days (interquartile 11-39 days) and 35 days (interquartile 22-65 days) did not differ. Access to plastic surgery (15% and 17%, p = 0.08) and to treatment categories did not vary: tumor resection (73% and 72%), exclusive systemic therapy (13% and 14%), exclusive radiation therapy (9% and 9%), exclusive BSC (5% and 5%) (p = 0.8). Among resected patients, the neoadjuvant therapy rate was lower in 2019 (16%) versus 2020 (20%) (p = 0.02). One-year OS rates were 99.3% versus 98.9% (HR = 0.96; 95% CI, 0.77-1.2), 72.6% versus 76.6% (HR = 1.28; 95% CI, 0.95-1.72), 96.6% versus 97.8% (HR = 1.09; 95% CI, 0.61-1.94), and 15.5% versus 15.1% (HR = 0.99; 95% CI, 0.72-1.37), in the treatment groups.CONCLUSIONS:Despite a decrease in the number of new BCs, there was no tumor stage shift, and OS did not vary.
In recent years, the development of clinical data warehouses (CDW) has put Electronic Health Records (EHR) data in the spotlight. More and more innovative technologies for healthcare are based on these EHR data. However, quality assessments on EHR data are fundamental to gain confidence in the performances of new technologies. The infrastructure developed to access EHR data - CDW - can affect EHR data quality but its impact is difficult to measure. We conducted a simulation on the Assistance Publique - Hôpitaux de Paris (AP-HP) infrastructure to assess how a study on breast cancer care pathways could be affected by the complexity of the data flows between the AP-HP Hospital Information System, the CDW, and the analysis platform. A model of the data flows was developed. We retraced the flows of specific data elements for a simulated cohort of 1,000 patients. We estimated that 756 [743;770] and 423 [367;483] patients had all the data elements necessary to reconstruct the care pathway in the analysis platform in the "best case" scenarios (losses affect the same patients) and in a random distribution scenario (losses affect patients at random), respectively.