Routinely collected EHR data are increasingly repurposed for secondary use. As data move through multi-stage pipelines, data quality losses can be introduced at each transformation step. Yet most assessment approaches treat data quality as a static, post-hoc property of a final dataset, making transformation induced losses invisible. Emerging regulatory frameworks demand transparency and traceability across the full data lifecycle but provide limited operational guidance on how to achieve this. To introduce a lifecycle framework for assessing data quality across sequential ETL transformations and validate it using the Intego-II database. The framework comprised two components: information loss and selection bias. It was applied to Intego-II database by comparing three dataset representations within ETL Stage 2: pseudonymised source data, transformed data, and standardised data. Four indicators were evaluated across the condition, procedure, and drug exposure domains: patient coverage, domain coverage, temporal coverage, and mapping coverage. TPQAs evaluated selection bias using Cohen’s d and standardised mean differences. Patient attrition of 31,6% (1,421,150 to 971,762) arose from SSN-based deduplication of legacy EHR records and eligibility-driven exclusion. ICD-10 to SNOMED CT mapping achieved 58,3% coverage, attributable to a decimal separator inconsistency affecting 3,871 codes and approximately 326,800 condition records (1,7%). Drug concept granularity was compressed by 99,9% due to three co-occurring failures: non-standard source codes, incomplete national mapping coverage, and structural ATC-to-RxNorm vocabulary gaps. TPQA-1 showed eligibility-driven attrition was demographically neutral (Cohen’s d = -0,003; SMD = 0,031). TPQA-2 identified vocabulary-driven drug domain loss disproportionately affecting older patients (Cohen’s d = -0,49). Sequential ETL processing introduces heterogeneous, mechanism-specific data quality losses invisible to post-hoc assessment. The framework is reproducible, generalisable, and supports the transparency and traceability requirements of the European Health Data Space and EU AI Act.
Background:The secondary use of health data is essential for advancing medical research and improving clinical practice. The Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) enables large-scale, multicenter studies but faces challenges related to consistency, completeness, and transparency during data mapping from original data sources. Objective:This study aimed to evaluate the quality of the mapping process for lung cancer data within the Federated Health Innovation Network project, with a focus on consistency, completeness, and challenges encountered throughout the process. Methods:Clinical data from Ghent University Hospital were mapped to the OMOP CDM using a reference data dictionary. Consistency was assessed using Cohen kappa (κ) scores, while completeness was evaluated by comparing patient and record counts before and after mapping. Challenges, including unstructured data and an evolving reference standard, were documented and analyzed. Results:High consistency was observed for structured variables, while some unstructured variables, such as "Smoking status," were excluded due to their free-text format and the lack of suitable OMOP concepts. The completeness analysis showed minimal data loss for most structured variables but highlighted substantial challenges associated with unstructured data. Persistent issues included evolving data dictionary versions and mismatches in diagnostic code granularity between institutions, underscoring structural challenges in standardization. Conclusions:The transformation of lung cancer data to the OMOP CDM highlighted both technical and systemic challenges, including the handling of unstructured data and the resolution of granularity discrepancies. A multidisciplinary approach involving clinical and technical expertise is crucial for ensuring reliable, high-quality datasets for multicenter research.
eSource - particularly EHR-to-EDC - is an emerging paradigm in clinical research that enables automated transfer of electronic health record (EHR) data into electronic data capture (EDC) systems, with the potential to reduce site burden, improve data quality and accelerate oncology clinical trial workflows. However, widespread implementation remains limited due to technical, regulatory and operational barriers. To address these challenges, the European Institute for Innovation through Health Data (i~HD) launched the eSource Scale-Up Task Force in 2024. This multi-stakeholder initiative brings together leading oncology centres and pharmaceutical sponsors to establish a consensus-driven roadmap for eSource adoption. Central to this effort are three foundational resources: readiness criteria for early adopters, a performance indicator framework for monitoring success and an operational playbook to guide implementation. This article provides a structured overview of the Task Force's objectives, collaborative model and outputs, with specific attention to its focus on interoperability, regulatory alignment and real-world validation. While initially developed for oncology, the Task Force's framework is applicable across therapeutic areas characterized by data-intensive workflows.
IntroductionThe General Data Protection Regulation (“GDPR”) legal basis for obtaining consent for the processing of personal data for research purposes, where those purposes cannot be fully specified in advance, is provided for in Articles 6, 7 and Recital 33. However, GDPR’s requirements for obtaining consent, as to the secondary use and sharing of data in research, have been argued to have generated confusion, whilst the conflicts between the Regulation itself, its practical application and research ethics are well-documented (1). The requirements for “informed consent”, as defined within the GDPR, have not been well defined in the context of genome research or clinical trials (2), which has in turn led to the implementation and interpretation of the lawful basis to span into different idiosyncratic models. This naturally has fed into the uncertainty of how the legal basis can be applied in practice and calls for an investigation into the requirements for consent to be “informed” in the context of health research. This work aims to provide a scoping review and analysis of relevant publications with ultimate purpose to examine whether the concept of ‘data altruism’, as stipulated within Article 2 (10) of the Data Governance Act (“DGA”), addresses the gaps left behind by the application of the legal basis of ‘consent’, under the GDPR (Art. 6 (1) and 7), in so far as the secondary uses of data for research are concerned. In this light the article, by exploring available solutions found in relevant literature and used in practice in national and European projects, examines how ‘data altruism’ can add any value and work as a cohesive solution that the research community can use.ObjectivesThe article, through its research, intends to answer the following questions:What gaps has the GDPR left when it comes to the interpretation and practical application of “consent” towards the secondary use of health data;Can the DGA, through the mechanism of ‘data altruism’, address these issues and provide a solution;What solutions have been used so far in practice to address this issue.MethodologyTo address the above-mentioned questions, the Arskey and O’Malley scoping review methodology and best practice, as outlined in the Joanna Briggs scoping review guidelines, have been applied. The research questions have been identified through an extensive literature review and consultation with subject matter experts. The search was conducted using six search engines and utilising a tailored search strategy, with the application of both MESH and non-MESH based search terms. From the identified relevant publications, 148 abstracts were kept to be read and 60 of those publications were kept as relevant. A PRISMA chart showcases the process in which the publications were reviewed and the process which led to the final papers kept as relevant. The title-abstract and full text screening and charting the data were concluded independently by two reviewers. Discrepancies were then resolved by a third reviewer. Results are summarised in both chart and narrative form below.ResultsThe final 60 publications were then split into three subcategories: (i) GDPR critique (23 publications listed); (ii) iterations of consent and data altruism (21 publications listed); and (iii) proposed solutions and current practices (31 publications). Certain of the publications fell into more than one of the above subcategories, given the interdisciplinary element of the subject and theme of each paper. Throughout the research, 5 of the publications discuss the Data Governance Act and data altruism, with 4 of those providing a critique over the text used in the DGA and the concept of ‘data altruism’ in relation to ‘consent’ as defined within the GDPR and the overall legislative framework for the secondary uses of data.
The secondary use of health data is essential for advancing medical research and improving clinical practices. The Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) enables large-scale, multi-center studies but faces challenges in consistency, completeness, and transparency during data mapping from the original data sources. This study aimed to evaluate the quality of the mapping process for lung cancer data within the Federated Health Innovation Network (FHIN) project, focusing on consistency, completeness, and challenges encountered throughout the process. Clinical data from Ghent University Hospital was mapped to the OMOP CDM using a reference data dictionary. Consistency was assessed through Cohen’s kappa scores, while completeness was evaluated by comparing patient and record counts pre- and post-mapping. Challenges, including unstructured data and evolving reference standards, were documented and analysed. High consistency was observed for structured variables, while some unstructured variables like “Smoking status” were excluded due to free-text format and a lack of suitable OMOP concepts. Completeness analysis showed minimal data loss for most structured variables but significant challenges for unstructured data. Persistent issues included evolving data dictionary versions and diagnostic code granularity mismatches between institutions, underscoring structural challenges in standardization. The transformation of lung cancer data to the OMOP CDM highlights both technical and systemic challenges, including handling unstructured data and addressing granularity discrepancies. A multidisciplinary approach involving clinical and technical expertise is crucial to ensure reliable, high-quality datasets for multi-center research.
Background: Data quality is fundamental to maintaining the trust and reliability of health data for both primary and secondary purposes. However, before the secondary use of health data, it is essential to assess the quality at the source and to develop systematic methods for the assessment of important data quality dimensions. Objective: This case study aims to offer a dual aim-to assess the data quality of height and weight measurements across 7 Belgian hospitals, focusing on the dimensions of completeness and consistency, and to outline the obstacles these hospitals face in sharing and improving data quality standards. Methods: Focusing on data quality dimensions completeness and consistency, this study examined height and weight data collected from 2021 to 2022 within 3 distinct departments-surgical, geriatrics, and pediatrics-in each of the 7 hospitals. Results: Variability was observed in the completeness scores for height across hospitals and departments, especially within surgical and geriatric wards. In contrast, weight data uniformly achieved high completeness scores. Notably, the consistency of height and weight data recording was uniformly high across all departments. Conclusions: A collective collaboration among Belgian hospitals, transcending network affiliations, was formed to conduct this data quality assessment. This study demonstrates the potential for improving data quality across health care organizations by sharing knowledge and good practices, establishing a foundation for future, similar research.
BackgroundHealth care has not reached the full potential of the secondary use of health data because of—among other issues—concerns about the quality of the data being used. The shift toward digital health has led to an increase in the volume of health data. However, this increase in quantity has not been matched by a proportional improvement in the quality of health data. ObjectiveThis review aims to offer a comprehensive overview of the existing frameworks for data quality dimensions and assessment methods for the secondary use of health data. In addition, it aims to consolidate the results into a unified framework. MethodsA review of reviews was conducted including reviews describing frameworks of data quality dimensions and their assessment methods, specifically from a secondary use perspective. Reviews were excluded if they were not related to the health care ecosystem, lacked relevant information related to our research objective, and were published in languages other than English. ResultsA total of 22 reviews were included, comprising 22 frameworks, with 23 different terms for dimensions, and 62 definitions of dimensions. All dimensions were mapped toward the data quality framework of the European Institute for Innovation through Health Data. In total, 8 reviews mentioned 38 different assessment methods, pertaining to 31 definitions of the dimensions. ConclusionsThe findings in this review revealed a lack of consensus in the literature regarding the terminology, definitions, and assessment methods for data quality dimensions. This creates ambiguity and difficulties in developing specific assessment methods. This study goes a step further by assigning all observed definitions to a consolidated framework of 9 data quality dimensions.
reports of data from original research.• Review Papers -comprehensive, authoritative, reviews within the journal's scope.• Short Reports -brief reports of data from original research.• Methodology Papers -Papers that present different methodological approaches that can be used to investigate problems in a relevant scientific field and to encourage innovation.• Policy Case Studies -brief articles on policy development at a regional or national level.• Study Protocols -articles describing a research protocol of a study.
Background There has been an exponential growth in the availability of apps, resulting in increased use of pregnancy apps. However, information on resources and use of apps among pregnant women is relatively limited. Objective The aim of this study is to map the current information resources and the use of pregnancy apps among pregnant women in Flanders. Methods A cross-sectional study was conducted, using a semistructured survey (April-June 2019) consisting of four different domains: (1) demographics; (2) use of devices; (3) sources of information; and (4) use of pregnancy apps. Women were recruited by social media, flyers, and paper questionnaires at prenatal consultations. Statistical analysis was mainly focused on descriptive statistics. Differences in continuous and categorical variables were tested using independent Student t tests and chi-square tests. Correlations were investigated between maternal characteristics and the women’s responses. Results In total, 311 women completed the entire questionnaire. Obstetricians were the primary source of information (268/311, 86.2%) for pregnant women, followed by websites/internet (267/311, 85.9%) and apps (233/311, 74.9%). The information that was most searched for was information about the development of the baby (275/311, 88.5%), discomfort/complaints (251/311, 80.7%) and health during pregnancy (248/311, 79.7%), administrative/practical issues (233/311, 74.9%), and breastfeeding (176/311, 56.6%). About half of the women (172/311, 55.3%) downloaded a pregnancy app, and primarily searched app stores (133/311, 43.0%). Pregnant women who are single asked their mothers (22/30, 73.3%) or other family members (13/30, 43.3%) for significantly more information than did married women (mother [in law]: 82/160, 51.3%, P=.02; family members: 35/160, 21.9%, P=.01). Pregnant women with lower education were significantly more likely to have a PC or laptop than those with higher education (72/73, 98.6% vs 203/237, 85.5%; P=.008), and to consult other family members for pregnancy information (30/73, 41.1% vs 55/237, 23.1%; P<.001), but were less likely to consult a gynecologist (70/73, 95.9% vs 198/237, 83.5%; P=.001). They also followed more prenatal sessions (59/73, 80.8% vs 77/237, 32.5%; P=.04) and were more likely to search for information regarding discomfort/complaints during pregnancy (65/73, 89% vs 188/237, 79.5%; P=.02). Compared to multigravida, primigravida were more likely to solicit advice about their pregnancy from other women in their social networks (family members: primigravida 44/109, 40.4% vs multigravida 40/199, 20.1%; P<.001; other pregnant women: primigravida 58/109, 53.2% vs multigravida 80/199, 40.2%; P<.03). Conclusions Health care professionals need to be aware that apps are important and are a growing source of information for pregnant women. Concerns rise about the quality and safety of those apps, as only a limited number of apps are subjected to an external quality check. Therefore, it is important that health care providers refer to high-quality digital resources and take the opportunity to discuss digital information with pregnant women.
There is an exponential growth in the availability of mobile Health (mHealth) applications, resulting in increased use of pregnancy apps. However, the actual use, experience, and characteristics of women using pregnancy apps are relatively unknown. To map the current use of the Internet and mobile applications and the needs and expectations among pregnant women in Flanders. A cross-sectional study was conducted, using a semi-structured survey (April - June 2019) consisting of four different domains: (1) demographics; (2) use of multimedia; (3) sources of information; and (4) use of pregnancy apps. Women were recruited by social media, flyers, and paper questionnaires at prenatal consultations. Statistical analysis was mainly focused on descriptive statistics. Differences in continuous and categorical variables were tested using Independent Student’s t-tests and Chi-square tests. Correlations were investigated between maternal characteristics and the women’s responses. In total, 311 women fulfilled the questionnaire completely. The majority of multimedia were daily used by the women (computer/laptop 40,84%; GSM: 80,71%; and smartphone/iPhone: 97,43%). The obstetrician was their prior source of information (86.17%), followed by ‘websites/Internet’ (85.85%) and ‘apps’ (74.92%). Information was mostly searched about the development of the baby (88.45%), discomfort/complaints (80.71%) and health during pregnancy (79.74%), administrative/practical issues (74.92%) and breastfeeding (56.59%). About half of the women (55.31%) downloaded a pregnancy app (172/311), mostly searched app stores (43.02%; 74/172). Singleton pregnancies asked significantly more information to their mother (73.33%) or other family members (43.33%) than married women (mother (in law): 51.26% (p= 0.02); family members: 21.88% (p = 0.01)) or cohabiting women (mother (in law): 50.00% (p = 0.02)). Pregnant women with lower education had significantly more a pc or laptop than those with higher education (98.63% vs. 85.47%; p = 0.008), consulted more other family members for pregnancy information (41.10% vs. 23.08%; p <0.01) but less a gynaecologist (95.89% vs. 83.54%; p = 0.001), followed more prenatal sessions (80.77% vs. 32.48%; p = 0.04) and searched more for information on discomfort/complains during pregnancy (89.04% vs. 79.49%; p = 0.02). Compared to multigravida, primigravida asked more advice about their pregnancy to their environment (family members: primigravida: 40.37% vs. multigravida 20.10%; p < 0.001; or other pregnant women: primigravida: 53.21% vs. multigravida 40.20%; p < 0.03). Healthcare professionals need to be aware that mHealth apps are important and are a growing source of information for pregnant women. Concerns rise about the quality and safety of those apps, as only a limited amount of apps are subjected to an external quality check. Therefore, it is important that caregivers refer to high quality digital resources and take the opportunity to discuss digital information with pregnant women. The study was approved by the Medical Ethics Committees of the hospital Oost-Limburg (no. 19/0026U, B-no. B371201939699) and Ghent University hospital (EC 2018/0120, B-no. B670201835156).
There is increasing recognition that healthcare providers need to focus attention, and be judged against, the impact they have on the health outcomes experienced by patients. The measurement of health outcomes as a routine part of clinical documentation is probably the only scalable way of collecting outcomes evidence, since secondary data collection is expensive and error prone. However, there is uncertainty about whether routinely collected clinical data within EHR systems includes the data most relevant to measuring and comparing outcomes, and if those items are collected to a good enough data quality to be relied upon for outcomes assessment, since several studies have pointed out significant issues regarding EHR data availability and quality. In this paper, we first describe a practical approach to data quality assessment of health outcomes, based on a literature review of existing frameworks for quality assessment of health data and multi-stakeholder consultation. Adopting this approach, we perform a pilot study on a subset of 21 International Consortium for Health Outcomes Measurement (ICHOM) outcomes data items from patients with congestive heart failure. To this end, all available registries compatible with the diagnosis of heart failure within the IMASIS-2 data repository connected to the Hospital del Mar network (142,345 visits of 12,503 patients) were extracted and mapped to the ICHOM format. We focus our pilot assessment on five commonly used data quality dimensions: completeness, correctness, consistency, uniqueness and temporal stability. Overall, this pilot study reveals high scores on the consistency, completeness and uniqueness dimensions. Temporal stability analyses show some changes over time in the reported use of medication to treat heart failure, as well as in the recording of past medical conditions. Finally, investigation of data correctness suggests several issues concerning the proper characterization of missing data values. Many of these issues appear to be introduced while mapping the IMASIS-2 relational database contents to the ICHOM format, as the latter requires a level of detail which is not explicitly available in the coded data of an EHR. To truly examine to what extent hospitals today are able to routinely collect the evidence of their success in achieving good health outcomes, future research would benefit from performing more extensive data quality assessments, including all data items from the ICHOM heart failure standard set, across multiple hospitals.
Summary Objective: This survey article presents a literature review of relevant publications aiming to explore whether the EU's General Data Protection Regulation (GDPR) has held true during a time of crisis and the implications that arose during the COVID-19 outbreak. Method and Results: Based on the approach taken and the screening of the relevant articles, the results focus on three themes: a critique on GDPR; the ethics surrounding the use of digital health technologies, namely in the form of mobile applications; and the possibility of cross border transfers of said data outside of Europe. Within this context, the article reviews the arising themes, considers the use of data through mobile health applications, and discusses whether data protection may require a revision when balancing societal and personal interests. Conclusions: In summary, although it is clear that the GDPR has been applied through a mixed and complex experience with data handling during the pandemic, the COVID-19 pandemic has indeed shown that it was a test the GDPR was designed and prepared to undertake. The article suggests that further review and research is needed to first ensure that an understanding of the state of the art in data protection during the pandemic is maintained and second to subsequently explore and carefully create a specific framework for the ethical considerations involved. The paper echoes the literature reviewed and calls for the creation of a unified and harmonised network or database to enable the secure data sharing across borders.
Background There is increasing recognition that health care providers need to focus attention, and be judged against, the impact they have on the health outcomes experienced by patients. The measurement of health outcomes as a routine part of clinical documentation is probably the only scalable way of collecting outcomes evidence, since secondary data collection is expensive and error-prone. However, there is uncertainty about whether routinely collected clinical data within electronic health record (EHR) systems includes the data most relevant to measuring and comparing outcomes and if those items are collected to a good enough data quality to be relied upon for outcomes assessment, since several studies have pointed out significant issues regarding EHR data availability and quality. Objective In this paper, we first describe a practical approach to data quality assessment of health outcomes, based on a literature review of existing frameworks for quality assessment of health data and multistakeholder consultation. Adopting this approach, we performed a pilot study on a subset of 21 International Consortium for Health Outcomes Measurement (ICHOM) outcomes data items from patients with congestive heart failure. Methods All available registries compatible with the diagnosis of heart failure within an EHR data repository of a general hospital (142,345 visits and 12,503 patients) were extracted and mapped to the ICHOM format. We focused our pilot assessment on 5 commonly used data quality dimensions: completeness, correctness, consistency, uniqueness, and temporal stability. Results We found high scores (>95%) for the consistency, completeness, and uniqueness dimensions. Temporal stability analyses showed some changes over time in the reported use of medication to treat heart failure, as well as in the recording of past medical conditions. Finally, the investigation of data correctness suggested several issues concerning the characterization of missing data values. Many of these issues appear to be introduced while mapping the IMASIS-2 relational database contents to the ICHOM format, as the latter requires a level of detail that is not explicitly available in the coded data of an EHR. Conclusions Overall, results of this pilot study revealed good data quality for the subset of heart failure outcomes collected at the Hospital del Mar. Nevertheless, some important data errors were identified that were caused by fundamentally different data collection practices in routine clinical care versus research, for which the ICHOM standard set was originally developed. To truly examine to what extent hospitals today are able to routinely collect the evidence of their success in achieving good health outcomes, future research would benefit from performing more extensive data quality assessments, including all data items from the ICHOM standards set and across multiple hospitals.
Les dossiers de santé électroniques hospitaliers contribuent à l’amélioration de la qualité des soins en permettant une meilleure gestion des informations cliniques. Les bases de données numériques ainsi constituées facilitent l’échange des informations de santé avec les prestataires de soins et optimisent la coordination multidisciplinaire pour de meilleurs résultats thérapeutiques. Le projet européen EHR4CR (electronic health records for clinical research) a développé une plateforme pilote innovante permettant de réutiliser ces données numériques pour la recherche clinique. En améliorant et en accélérant les procédures de recherche clinique, cette approche permet d’envisager la réalisation d’études cliniques de manière plus efficiente, plus rapide et plus économique.
Electronic health records in hospitals contribute to improving the quality of care by enabling better management of clinical information. The databases thus constituted facilitate the exchange of health information with healthcare providers and optimize multidisciplinary coordination for better therapeutic results. The EHR4CR (Electronic Health Records for Clinical Research) European project has developed an innovative pilot platform enabling the reuse of this digital information for clinical research. By enhancing and speeding up clinical research procedures, this innovative approach makes it possible to conduct clinical trials more efficiently, faster, and more economically.
Electronic health records in hospitals contribute to improving the quality of care by enabling better management of clinical information. The databases thus constituted facilitate the exchange of health information with healthcare providers and optimize multidisciplinary coordination for better therapeutic results. The EHR4CR (Electronic Health Records for Clinical Research) European project has developed an innovative pilot platform enabling the reuse of this digital information for clinical research. By enhancing and speeding up clinical research procedures, this innovative approach makes it possible to conduct clinical trials more efficiently, faster, and more economically.
The European Institute for Innovation through Health Data (i~HD) has been formed as one of the sustainable entities arising from the Electronic Health Records for Clinical Research (EHR4CR) and SemanticHealthNet projects, in collaboration with other European Commission projects and initiatives. The vision of i~HD is to become the European organisation of reference for guiding and catalysing the best, most efficient and trustworthy uses of health data and interoperability, for optimizing health and knowledge discovery. i~HD has been established in recognition that there is a need to tackle areas of challenge in the successful scaling up of innovations that rely on high-quality and interoperable health data, to sustain and propagate the results of eHealth research, and to address current-day obstacles to using health data. i~HD was launched at an inaugural conference in Paris, in March 2016. This was attended by over 200 European clinicians, healthcare providers and researchers, representatives of the pharma industry, patient associations, health professional associations, the health ICT industry and standards bodies. The event showcased issues and approaches, that are presented in this paper to highlight the activities that i~HD intends to pursue as enablers of the better uses of health data, for care and research.