AIM:To estimate the prevalence of medication use in nonhospitalized pregnant women with COVID-19. METHODS:A prospective two-stage individual patient meta-analysis across 10 data sources in Europe and North America studied medication use among nonhospitalized pregnant women with COVID-19 between January 2020 and December 2022. Comparisons were made between medication use within 30 days pre- and post-COVID-19 diagnosis in this cohort and two comparator groups: pregnant women without COVID-19 and nonpregnant women with COVID-19. Prevalence estimates were pooled using a random-effects model stratified by trimester. RESULTS:50 335 nonhospitalized pregnant women with COVID-19 were identified. The pooled prevalence of antibacterial use in the third trimester was higher post-COVID-19 diagnosis (6.8%, 95% confidence interval [CI] = 5.5-8.4, I2 = 94%) compared with the same women pre-COVID-19 (3.9%, 95% CI = 3.1-4.9, I2 = 89%). Overall, pregnant women with COVID-19 had higher medication use compared to pregnant women without COVID-19, although the CIs of the prevalence overlapped. Post-COVID-19, antithrombotic prevalence was 4.5% (95% CI = 1.1-16.5, I2 = 100%) among pregnant women with COVID-19 in the third trimester, compared to 2.1% (95% CI = 1.2-3.6, I2 = 99%) among those without COVID-19 in the third trimester. Compared to nonpregnant women with COVID-19, pregnant women with COVID-19 were less likely to be prescribed analgesics, antiprotozoals, corticosteroids, psychoanaleptics and psycholeptics, and more likely to be prescribed antithrombotics, cough and cold and nasal preparations, and drugs used in diabetes across all trimesters. High heterogeneity existed in nearly all analyses. CONCLUSION:This international meta-analysis reveals low medication use and country-specific variations, enhancing insight into the management of COVID-19 in nonhospitalized pregnant women. Higher antithrombotic use post-COVID-19 suggests prophylactic treatment in this population, but variation between countries emphasizes the challenges of combining multinational data.
PURPOSE:In 2019, the Innovative Medicines Initiative funded the ConcePTION project to enhance monitoring of medication safety in pregnancy and breastfeeding. This paper describes how the ConcePTION Pregnancy Algorithm (PA) identified pregnancies in 10 diverse European electronic healthcare data sources and estimated their duration. METHODS:Data sources from six European countries were mapped to the ConcePTION Common Data Model. Any pregnancy-related record was retrieved from various available data banks, including birth register, primary care records, and hospital records, and reconciled into episodes of pregnancy (starting between 01/2015 and 12/2019), each with start date, end date, and type of end. A random forest model was used to estimate missing gestational ages for incomplete records. Parameters were tailored to data sources to address local variations in data availability, collection, and governance. Model performance was evaluated using cross-validated Root Mean Squared Error (RMSE). RESULTS:The PA identified ~2.7 million pregnancies, in over 2.2 million individuals. Most ended in live births (50%-83%), 1%-15% in elective terminations, and 4%-10% in spontaneous abortions, depending on data sources. Pregnancies with unknown type of end were also retrieved (2%-34%). Gestational age was predicted for 6%-89% of records (RMSE: 17-50 days). The median gestational age at first identified pregnancy record ranged from 47 to 280 days. CONCLUSIONS:We developed an open-source algorithm to identify and date pregnancies, including early-stage pregnancies with unknown end and/or ongoing at the time of data extraction. This algorithm may facilitate multinational studies, improving generation of timely real-world evidence about use and safety of medicinal products in pregnancy.
ObjectiveThe HDRUK Phenotype Library (phenotypes.healthdatagateway.org) shares definitions used to measure concepts of interest (such as diagnoses or treatments) within health datasets. It already holds more than 1000 phenotypes, with researchers able to contribute their work via an API. We aimed to create a more user-friendly method of contributing to the Library. ApproachWe designed a phenotype creation workflow enabling users to create and submit new content via web interface. Goals included clarity and ease of use for a broad range of users. ResultsA home page shows researchers’ own content. Researchers can create new phenotypes using a web form, entering metadata such as name, authors, description, and publications. Code lists are defined via one or more rules, including search terms or referring to another phenotype, or by CSV upload. Users can make their phenotypes accessible to a research group or to all authenticated users, as well as publish content on the web. Publication requests are reviewed to ensure content is complete and appropriate. Editing, with full history and version control, is also supported. ConclusionsWe implemented and released the new features. We are currently engaging researchers to get feedback and invite content submission. ImplicationsThe benefit of tools to support research transparency and repeatability is only realized when they are adopted. We hope that a GUI to support phenotype creation will broaden the Library’s user base help it serve as an enabler of higher-quality, more efficient research across the worldwide health data research community.
Purpose The COVID-19 pandemic has impacted medication needs and prescribing practices, including those affecting pregnant women. Our goal was to investigate patterns of medication use among pregnant women with COVID-19, focusing on variations by trimester of infection and location. Methods We conducted an observational study using six electronic healthcare databases from six European regions (Aragon/Spain; France; Norway; Tuscany, Italy; Valencia/Spain; and Wales/UK). The prevalence of primary care prescribing or dispensing was compared in the 30-day periods before and after a positive COVID-19 test or diagnosis. Results The study included 294,126 pregnant women, of whom 8943 (3.0%) tested positive for, or were diagnosed with, COVID-19 during their pregnancy. A significantly higher use of antithrombotic medications was observed particularly after COVID-19 infection in the second and third trimesters. The highest increase was observed in the Valencia region where use of antithrombotic medications in the third trimester increased from 3.8% before COVID-19 to 61.9% after the infection. Increases in other countries were lower; for example, in Norway, the prevalence of antithrombotic medication use changed from around 1–2% before to around 6% after COVID-19 in the third trimester. Smaller and less consistent increases were observed in the use of other drug classes, such as antimicrobials and systemic corticosteroids. Conclusion Our findings highlight the substantial impact of COVID-19 on primary care medication use among pregnant women, with a marked increase in the use of antithrombotic medications post-COVID-19. These results underscore the need for further research to understand the broader implications of these patterns on maternal and neonatal/fetal health outcomes.
ObjectiveThe Welsh Longitudinal General Practice (WLGP) dataset contains over 4 billion records. Due to the size and the need to filter the dataset to get the necessary general practice interaction results, query performance is poor. To overcome this, a ‘cleaned’ version of the dataset needed to be created. ApproachAn R package was created that would be run when a new refresh of the WLGP data is provided. Two new tables would be created based on the original WLGP Events table. ResultsA WLGP Cleaned Dataset R Package was produced and creates a reformatted events table and events look up table. The reformatted GP events table reduces the original events table from 14 columns to 8 key columns. The events look up table takes distinct events from the original WLGP events table and produces a new table including columns such as the event description, type and hierarchy levels. Both tables are then linked using an event code ID. ConclusionThe R package now creates a new reformatted events table as well as an events look up table which is run after a WLGP refresh is provided. The new tables are then provisioned to any projects that have access to the dataset. ImplicationsThe new tables has improved query performance for our researchers, providing them with a table with all the necessary information and an easy way to decipher events codes. This has improved the amount of time researchers have spent manipulating and querying the tables.
ObjectiveWe’re introducing a comprehensive approach to enhance sensitive data flagging within our Data Quality (DQ) tool. Within de-identified health records, sensitive data such as user ids, post codes, phone numbers, and addresses pose significant privacy risks if exposed in their raw form. To address this challenge, we propose using all necessary regex patterns and free text field checks for sensitive data flagging with the objective of efficient detection within large datasets and ultimately reporting such data. ApproachA regex map has been developed along free text field thresholds, distribution counts and percentage checks to allow a generic approach. This ensures that any new regexes can be added to the flagging process without significant code changes. Once identified, sensitive data instances, including fields that contain free text are flagged and displayed in an html report that is human readable by DQ reviewers. ResultsThrough a combination of SQL and R regex based text processing, our approach allows a seamless identification and flagging of sensitive data within datasets. Sensitive data for large tables of +25 million records with over 50 columns are getting flagged in less than 150 seconds (approx. 2 minutes). ConclusionOur development offers a practical solution for sensitive data flagging in complex datasets where team members can implement robust sensitive data flagging by adding more regex formulas in the codebase. ImplicationsDQ reviewers can inspect the sensitive flagged fields in a human readable way, providing an extra measure of confidence when determining the quality of de-identified health records.
ObjectiveConstruct an innovative open-learning solution that provides comprehensive training specific to Trusted Research Environments (TRE) and the broader research community of administrative data users, irrespective of their proficiency levels. DOL offers training opportunities tailored to meet each user's unique learning needs, enabling them to utilise complex, linked administrative datasets confidently and effectively for meaningful research outcomes, thereby building capacity for sustainable national data infrastructure. ApproachDOL’s innovative open-learning solution offers two learning formats: adaptive and experiential. Adaptive learning provides registered users with bitesize self-paced training based on Administrative Data Research UK's priorities. Experiential learning involves online workgroups with real-world context and practical application. They meet twice a year and are designed around specific topics with frequent guest speakers who are experts in their fields. ConclusionDOL’s innovative open-learning solution empowers TREs, such as SAIL Databank, to provide well-rounded learning that fosters community support, knowledge-sharing, and networking opportunities for its users, while gathering valuable user feedback. Users can personalise learning and test their knowledge in a flexible training environment, allowing them to take charge of their learning journey. ImplicationIn response to the increasing demand for training services from novice to advanced users, SAIL Databank adopted DOL's dual learning approach by (1) developing training courses that cover access, process, methodology, integration, and analytical tools for SAIL TRE users, (2) engaging users in a series of workgroups focused on themed datasets including Justice, Environmental, Maternal, Education, and Core Health.
ObjectiveThe Cohort Builder package aims to streamline the cohort building process using R, with simplicity and efficiency as the primary focus. Researchers can generate cohorts from TRE (Trusted Research Environment) datasets by specifying simple cohort definitions as YAML (YAML Ain't a Markup Language) parameters using this package, eliminating the need to write complex queries. ApproachCohort Builder offers researchers a user-friendly interface for defining cohort criteria by abstracting away the complexities of database queries using YAML configuration files. The package enables researchers to focus on defining the desired cohort characteristics rather than navigating the database and tables. Users can easily specify inclusion and exclusion criteria, demographic constraints, temporal constraints, and other dataset parameters to create cohorts tailored to their research goals and reproduce the results using the same configuration file. ResultUsing the Cohort Builder, users may quickly create cohorts from TRE datasets with minimum code. The package facilitates reproducibility, transparency, and increased productivity. Researchers gain from having to rely less on specialised database expertise and having more flexibility when defining and exploring cohorts. ConclusionCohort Builder offers a straightforward yet effective method for defining cohorts. The package frees researchers from having to worry about intricate technological details in favour of their analytical objectives. ImplicationsMoving forward, Cohort Builder will support creative research projects across a range of disciplines and promote collaboration among researchers by allowing them to share configuration files for their cohorts.
ObjectiveThe volume and frequency of refreshed data within the [organisation] has increased significantly since the beginning of the COVID-19 pandemic. Therefore, a more efficient data quality (DQ) checking process was necessary. Approach Having previously developed an automated DQ checking tool, the focus was on re-engineering the process of DQ task allocation and communication of results. Results 5 analysts were trained in DQ checking. A JIRA workflow tracks the management of data loading. When a dataset is ready for DQ, the Data Manager allocates a ticket to the DQ Lead who then allocates it onto one of the 5 analysts. Via a DQ Slack channel, the analyst is informed and acknowledges receipt of the task. On completion of DQ, the analyst updates the ticket and transfers it to the appropriate workflow stage. Passed DQ tickets are transferred to the Data Manager for data release, whereas failed ones are placed “On Hold”. The DQ Lead triages the issues and liaises with relevant parties for resolution, which may require data amendments. On receipt of amended data, the DQ ticket is transferred back to the queue and the analyst is notified to re-check the data. ConclusionsThe team now complete a high volume of DQ checks efficiently. In 2023, 405 datasets, containing 1716 tables, were quality checked, with the initial DQ taking, on average, 2.6 days. Implications The improved speed of DQ checking ensures projects access the latest available data whilst maintaining the expected DQ levels, integrity and reputation of the TRE.
Objective:To enable reproducible research at scale by creating a platform that enables health data users to find, access, curate, and re-use electronic health record phenotyping algorithms.Materials and Methods:We undertook a structured approach to identifying requirements for a phenotype algorithm platform by engaging with key stakeholders. User experience analysis was used to inform the design, which we implemented as a web application featuring a novel metadata standard for defining phenotyping algorithms, access via Application Programming Interface (API), support for computable data flows, and version control. The application has creation and editing functionality, enabling researchers to submit phenotypes directly.Results:We created and launched the Phenotype Library in October 2021. The platform currently hosts 1049 phenotype definitions defined against 40 health data sources and >200K terms across 16 medical ontologies. We present several case studies demonstrating its utility for supporting and enabling research: the library hosts curated phenotype collections for the BREATHE respiratory health research hub and the Adolescent Mental Health Data Platform, and it is supporting the development of an informatics tool to generate clinical evidence for clinical guideline development groups.Discussion:This platform makes an impact by being open to all health data users and accepting all appropriate content, as well as implementing key features that have not been widely available, including managing structured metadata, access via an API, and support for computable phenotypes.Conclusions:We have created the first openly available, programmatically accessible resource enabling the global health research community to store and manage phenotyping algorithms. Removing barriers to describing, sharing, and computing phenotypes will help unleash the potential benefit of health data for patients and the public.
Aims In patients with non-valvular atrial fibrillation (NVAF) prescribed warfarin, the association between guideline defined international normalised ratio (INR) control and adverse outcomes in unknown. We aimed to (i) determine stroke and systemic embolism (SSE) and bleeding events in NVAF patients prescribed warfarin; and (ii) estimate the increased risk of these adverse events associated with poor INR control in this population. Methods and results Individual-level population-scale linked patient data were used to investigate the association between INR control and both SSE and bleeding events using (i) the National Institute for Health and Care Excellence (NICE) criteria of poor INR control [time in therapeutic range (TTR) <65%, two INRs <1.5 or two INRs >5 in a 6-month period or any INR >8]. A total of 35 891 patients were included for SSE and 35 035 for bleeding outcome analyses. Mean CHA(2)DS(2)-VASc score was 3.5 (SD = 1.7), and the mean follow up was 4.3 years for both analyses. Mean TTR was 71.9%, with 34% of time spent in poor INR control according to NICE criteria. SSE and bleeding event rates (per 100 patient years) were 1.01 (95%CI 0.95-1.08) and 3.4 (95%CI 3.3-3.5), respectively, during adequate INR control, rising to 1.82 (95%CI 1.70-1.94) and 4.8 (95% CI 4.6-5.0) during poor INR control. Poor INR control was independently associated with increased risk of both SSE [HR = 1.69 (95%CI = 1.54-1.86), P < 0.001] and bleeding [HR = 1.40 (95%CI 1.33-1.48), P < 0.001] in Cox-multivariable models. Conclusion Guideline-defined poor INR control is associated with significantly higher SSE and bleeding event rates, independent of recognised risk factors for stroke or bleeding.
Background Congenital anomalies (CAs) increase the risk of death during infancy and childhood. This study aimed to evaluate the accuracy of using death certificates to estimate the burden of CAs on mortality for children under 10 years old. Methods Children born alive with a major CA between 1 January 1995 and 31 December 2014, from 13 population-based European CA registries were linked to mortality records up to their 10th birthday or 31 December 2015, whichever was earlier. Results In total 4199 neonatal, 2100 postneonatal and 1087 deaths in children aged 1–9 years were reported. The underlying cause of death was a CA in 71% (95% CI 64% to 78%) of neonatal and 68% (95% CI 61% to 74%) of postneonatal infant deaths. For neonatal deaths the proportions varied by registry from 45% to 89% and by anomaly from 53% for Down syndrome to 94% for tetralogy of Fallot. In children aged 1–9, 49% (95% CI 42% to 57%) were attributed to a CA. Comparing mortality in children with anomalies to population mortality predicts that over 90% of all deaths at all ages are attributable to the anomalies. The specific CA was often not reported on the death certificate, even for lethal anomalies such as trisomy 13 (only 80% included the code for trisomy 13). Conclusions Data on the underlying cause of death from death certificates alone are not sufficient to evaluate the burden of CAs on infant and childhood mortality across countries and over time. Linked data from CA registries and death certificates are necessary for obtaining accurate estimates.
OBJECTIVES:To investigate the survival up to age 10 for children born alive with a major congenital anomaly (CA).METHODS:This population-based linked cohort study (EUROlinkCAT) linked data on live births from 2005 to 2014 from 13 European CA registries with mortality data. Pooled Kaplan-Meier survival estimates up to age 10 were calculated for these children (77 054 children with isolated structural anomalies and 4011 children with Down syndrome).RESULTS:The highest mortality of children with isolated structural CAs was within infancy, with survival of 97.3% (95% confidence interval [CI]: 96.6%-98.1%) and 96.9% (95% CI: 96.0%-97.7%) at age 1 and 10, respectively. The 10-year survival exceeded 90% for the majority of specific CAs (27 of 32), with considerable variations between CAs of different severity. Survival of children with a specific isolated anomaly was higher than in all children with the same anomaly when those with associated anomalies were included. For children with Down syndrome, the 10-year survival was significantly higher for those without associated cardiac or digestive system anomalies (97.6%; 95% CI: 96.5%-98.7%) compared with children with Down syndrome associated with a cardiac anomaly (92.3%; 95% CI: 89.4%-95.3%), digestive system anomaly (92.8%; 95% CI: 87.7%-98.2%), or both (88.6%; 95% CI: 83.2%-94.3%).CONCLUSIONS:Ten-year survival of children born with congenital anomalies in Western Europe from 2005 to 2014 was relatively high. Reliable information on long-term survival of children born with specific CAs is of major importance for parents of these children and for the health care professionals involved in their care.
In 2019, the Innovative Medicines Initiative (IMI) funded the ConcePTION project—Building an ecosystem for better monitoring and communicating safety of medicines use in pregnancy and breastfeeding: validated and regulatory endorsed workflows for fast, optimised evidence generation—with the vision that there is a societal obligation to rapidly reduce uncertainty about the safety of medication use in pregnancy and breastfeeding. The present paper introduces the set of concepts used to describe the European data sources involved in the ConcePTION project and illustrates the ConcePTION Common Data Model (CDM), which serves as the keystone of the federated ConcePTION network. Based on data availability and content analysis of 21 European data sources, the ConcePTION CDM has been structured with six tables designed to capture data from routine healthcare, three tables for data from public health surveillance activities, three curated tables for derived data on population (e.g., observation time and mother‐child linkage), plus four metadata tables. By its first anniversary, the ConcePTION CDM has enabled 13 data sources to run common scripts to contribute to major European projects, demonstrating its capacity to facilitate effective and transparent deployment of distributed analytics, and its potential to address questions about utilization, effectiveness, and safety of medicines in special populations, including during pregnancy and breastfeeding, and, more broadly, in the general population.
Aims In patients with non-valvular atrial fibrillation prescribed warfarin, the UK National Institute of Health and Care Excellence (NICE) defines poor anticoagulation as a time in therapeutic range (TTR) of <65%, any two international normalized ratios (INRs) within a 6-month period of <= 1.5 ('low'), two INRs >= 5 within 6 months, or any INR >= 8 ('high'). Our objectives were to (i) quantify the number of patients with poor INR control and (ii) describe the demographic and clinical characteristics associated with poor INR control. Method and results Linked anonymized health record data for Wales, UK (2006-2017) was used to evaluate patients prescribed warfarin who had at least 6 months of INR data. 32 380 patients were included. In total, 13 913 (43.0%) patients had at least one of the NICE markers of poor INR control. Importantly, in the 24 123 (74.6%) of the cohort with an acceptable TTR (>= 65%), 5676 (23.5%) had either low or high INR readings at some point in their history. In a multivariable regression female gender, age (>= 75 years), excess alcohol, diabetes heart failure, ischaemic heart disease, and respiratory disease were independently associated with all markers of poor INR control. Conclusion Acceptable INR control according to NICE standards is poor. Of those with an acceptable TTR (>65%), one-quarter still had unacceptably low or high INR levels according to NICE criteria. Thus, only using TTR to assess effectiveness with warfarin has the potential to miss a large number of patients with non-therapeutic INRs who are likely to be at increased risk.