Health research increasingly relies on observational datasets that capture systematically different patient populations, affecting the representativeness, comparability, and transportability of study findings. Yet these differences are rarely quantified across clinical domains or explored interactively, and the workflow for pairwise dataset comparison remains fragmented. We present Syrona, a visual analytics workflow for systematic pairwise comparison of datasets and subcohorts on the OMOP Common Data Model. Syrona extracts annual prevalence for conditions, procedures, and drugs from any OMOP CDM database, computes prevalence ratios across demographic strata, and synthesizes them via multilevel meta-analysis. A coordinated dashboard - distributional overviews, stratified heatmaps, forest plots, and absolute prevalence comparisons - enables interactive exploration driven by domain-adaptive SNOMED CT and ATC filtering. Three case studies on Estonian national health data (495,000 persons, 2012-2024) demonstrate how the workflow supports representativeness assessment, institutional practice comparison, and coding artifact detection. Syrona is open-source and applicable to any OMOP CDM database. Explore the demo at: http://omop-apps.cloud.ut.ee/ShinyApps/Syrona/ ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by the Estonian Research Council (PRG1844, PRG2078, PSG809, PSG1216, PRG1414, PRG1291). The study was funded by the European Union and co-funded by the Ministry of Education and Research (TEM-TA72). The European Union funded the project under its Horizon Europe research and innovation programme (grant agreement No 101060011, TeamPerMed) and co-funded the research through the European Regional Development Fund (Project No. 2021-2027.1.01.24-0444). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. This work was also supported by the Estonian Centre of Excellence in Artificial Intelligence (EXAI), the Estonian Centre of Excellence in Personalised Medicine (CEPM), and the Center of Excellence for Well-Being Sciences (EstWell) funded by the Estonian Ministry of Education and Research grants TK213, TK214, and TK218 respectively. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This work was approved by the Estonian Bioethics and Human Research Council (1.1-12/653, 1.1-12/1039) and the Ethics Committee of University of Tartu (401/T-34). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced are available online at https://github.com/MaarjaPajusalu/Syrona/
Abstract Objective To address the unresolved bottleneck of selecting cohort-relevant clinical concepts for treatment trajectory analysis in observational health data, we introduce CohortContrast, an OMOP-compatible R package for enrichment-based concept identification, temporal and semantic noise reduction, and concept aggregation, enabling cohort-level characterization and downstream trajectory analysis. Materials and Methods We developed CohortContrast and applied it to OMOP-mapped observational data from the Estonian nationwide OPTIMA database, which includes all cases of lung, breast, and prostate cancer, focusing here on lung and prostate cancer cohorts. The workflow combines target-control statistical enrichment, temporal/global noise filtering, hierarchical concept aggregation and correlation-based merging, with optional patient clustering for downstream trajectory exploration. We validated the approach with a clinician-based plausibility assessment of extracted diagnosis-concept pairs and evaluated a large language model (LLM) as an auxiliary filtering step. Results We analyzed 7,579 lung cancer and 11,547 prostate cancer patients. The workflow reduced concept dimensionality from 5,793 to 296 concepts (94.9%) in lung cancer and from 5,759 to 170 concepts (97.0%) in prostate cancer, and identified three exploratory patient subgroups in both cohorts. In a plausibility assessment of 466 diagnosis-concept pairs, validators rated 31.3% as directly linked and 57.5% as indirectly linked. Discussion CohortContrast reduces manual concept curation by prioritizing and aggregating cohort-relevant concepts while preserving clinically interpretable treatment patterns in OMOP-based real-world data. Conclusion CohortContrast enables scalable reduction of broad OMOP concept spaces into clinically interpretable, cohort-specific representations for exploratory trajectory analysis and real-world evidence research.
Objective Selecting cohort-relevant clinical concepts from high-dimensional observational health data remains a persistent bottleneck in trajectory analysis, phenotyping, and real-world evidence studies. We developed an OMOP-compatible workflow for deriving compact, clinically interpretable concept sets from broad real-world data concept spaces. Methods We applied the workflow to OMOP-mapped observational data from the Estonian nationwide OPTIMA database using lung cancer and prostate cancer cohorts. Candidate concepts observed during cohort-specific windows were contrasted against control periods or cohorts using statistical enrichment tests, filtered against time-matched background populations, aggregated through OMOP vocabulary hierarchies, and merged using correlation-based redundancy reduction. We evaluated the resulting concept sets by quantifying dimensionality reduction, examining downstream exploratory patient clustering, and conducting clinician-based relevance assessment of extracted diagnosis-concept pairs. Results The study included 7,579 patients with lung cancer and 11,547 patients with prostate cancer. The workflow reduced the concept space from 5,793 to 296 concepts in lung cancer and from 5,759 to 170 concepts in prostate cancer, corresponding to 94.9% and 97.0% reductions, respectively. Retained concepts spanned diagnostic, treatment, monitoring, supportive-care, and outcome domains. Exploratory clustering identified three clinically interpretable pathway-pattern groupings in each cohort. Across two clinician raters, 85.4–88.0% of the 466 extracted diagnosis–concept pairs were assessed as directly or indirectly related to the target condition. Conclusion Enrichment-based cohort contrast can transform broad OMOP concept spaces into compact, clinically interpretable, cohort-specific representations for exploratory trajectory analysis and real-world evidence research. The workflow provides a reproducible concept-prioritization step for OMOP-based cohort characterization, with an open-source implementation available for reuse.
Purpose To develop and describe DrugUtilisation, an open-source R package that facilitates drug utilisation studies using data mapped to the OMOP Common Data Model (CDM).Methods Core functionalities include creating drug user cohorts, identifying and summarising indications, describing the duration and dose of medication/s and assessing treatment adherence. The package works with packages developed within the DARWIN EU initiative to support study-specific workflows. We show the package's workflow by analysing the use of simvastatin in three European real-world databases.Results This paper outlines the DrugUtilisation package's functions and demonstrates their application with a clinical example of simvastatin use in databases from the United Kingdom, Estonia and the Netherlands. We generated results including cohort counts, indication summaries, measures of dose and duration, as well as publication-ready tables and figures. We implemented comprehensive unit tests and standardised output format, which ensured consistency across databases and minimised coding errors.Conclusion The development of this software allows for researchers to quickly perform common drug utilisation analyses, while also providing the foundation for additional, bespoke study-specific analyses.
Abstract Patients with earlier SARS-CoV-2 variants are at increased risk of venous and arterial thromboembolic (VTE, ATE) events. Here we aimed to contextualise the incidence of thromboembolic events among patients with COVID-19 during the Omicron period. We conducted a population-based cohort study using electronic health records from the UK (CPRD GOLD), the Netherlands (IPCI), and Spain (SIDIAP) within the DARWIN EU ® network. Two cohorts were included: a pre-pandemic population (2017–2019) and individuals infected with SARS-CoV-2 during the Omicron-dominant period. We estimated incidence rates (IRs) of VTE, ATE, and other cardiovascular events at 30-, 60-, 90-, and 180-days post-infection. Crude incidence rate ratios (IRRs) and age-sex standardized incidence ratios (SIRs) were calculated relative to the pre-pandemic cohort. Analyses were stratified by prior infection, vaccination status, and immunocompromised status. In total, we included over 7.6 million individuals (CPRD GOLD: 5.28 M; IPCI: 1.59 M; SIDIAP: 0.75 M) in the general population cohort, and about 0.8 million individuals (CPRD GOLD: 248,847; IPCI: 330,200; SIDIAP: 200,563) in the COVID-19 Omicron cohort. Crude IRs varied by outcome and data source. For VTE, IRs per 100,000 person-years were 136 [95%CI 131–141] in SIDIAP, 167 [164–169] in CPRD GOLD, and 264 [259–270] in IPCI. Elevated SIRs for VTE and ATE were observed following SARS-CoV-2 infection, highest within 30 days and persisting up to 180 days. In CPRD GOLD, the VTE SIR was 3.61 [2.45–5.53] at 30 days, decreasing to 1.88 [1.52–2.34] at 180 days. Higher SIRs were observed among immunocompromised individuals and those without prior infection. Our findings indicate that among individuals diagnosed with SARS-CoV-2 infection during the Omicron-dominant period, observed rates of thromboembolic events exceeded expected background incidence, particularly in the early post-infection period.
BackgroundDrug adherence is crucial for chronic disease management, yet treatment discontinuation remains common due to factors such as side effects, inefficacy, or cost. These reasons are often recorded only in free-text clinical notes, making large-scale analysis difficult. While large language models (LLMs) can interpret such unstructured data more effectively than traditional natural language processing methods, few studies have systematically categorized reasons for discontinuation or identified whether the decision was initiated by the patient or the clinician, especially in low-resource languages such as Estonian. ObjectiveThis study aimed to assess the ability of LLMs to extract and classify reasons for drug discontinuation and identify who initiated it using Estonian electronic health records and characterize the observed discontinuation patterns and initiators for statins and antidiabetic medications. MethodsWe combined prescription data with free-text anamneses from a 10% sample of the Estonian population (2012-2019). LLMs (Llama 3.1-70B and GPT-4o) were applied to extract discontinuation phrases and reasons, classify them into a clinician-developed taxonomy, and identify who discontinued the treatment. Performance was evaluated on 100 randomly chosen cases per drug group. ResultsExtraction yielded 625 antidiabetic drug and 233 statin discontinuation cases. Validation confirmed a precision of 0.93 to 0.98 for extracting phrases and 0.95 to 0.96 for extracting reasons. Classification of discontinuation reasons achieved weighted F1-scores of 0.81 to 0.84, whereas classification of who initiated discontinuation achieved weighted F1-scores of 0.64 to 0.78. Adverse reactions were the most frequent reason overall, accounting for 70% (163/233) of statin discontinuations and 44.8% (280/625) of antidiabetic drug discontinuations. Regarding antidiabetic drugs, treatment inefficacy and contraindications were more common. Patients more often stopped due to adverse reactions or nonmedical reasons, whereas physicians more often initiated discontinuation for contraindications. ConclusionsLLMs can accurately extract and classify medication discontinuation reasons and show variable performance in identifying discontinuation initiators in Estonian clinical narratives. Both local and proprietary models showed promising results, enabling scalable analyses that complement structured health records. This demonstrates the potential of LLMs to unlock information from clinical notes, turning this underused electronic health record component into a valuable resource for monitoring treatment patterns and detecting adverse event signals.
Background: The increasing availability of routinely collected health data offers new opportunities for population-level research, yet access to comprehensive, linked, and standardised datasets remains limited. We describe EST-Health-30, a large-scale, population-representative health data resource from Estonia. Methods: EST-Health-30 comprises a random 30% sample of the Estonian population (~500,000 individuals), with longitudinal data from 2012 to 2024 and annual updates planned through 2026. Individual-level records are linked across five nationwide databases, including electronic health records, health insurance claims, prescription data, cancer registry, and cause of death records. A privacy-preserving hashing approach ensures consistent cohort inclusion over time while maintaining pseudonymisation. All data are harmonised to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (version 5.4) using international standard vocabularies. Data quality was assessed using established OMOP-based validation frameworks. Results: The dataset contains rich multimodal information on diagnoses, procedures, laboratory measurements, prescriptions, free-text clinical notes, healthcare utilisation, and costs, with high population coverage and longitudinal depth. Data quality assessment showed high completeness and consistency, with 99.2% of applicable checks passing. The age-sex distribution closely reflects the national population, supporting representativeness, though coverage is marginally below the target 30% (29.2%), primarily attributable to recent immigrants without health system contact. The dataset enables construction of detailed clinical cohorts, analysis of disease trajectories, and evaluation of healthcare utilisation and outcomes across the life course. Conclusions: EST-Health-30 is a comprehensive, standardised, and population-representative real-world data resource that supports epidemiological, clinical, and methodological research. Its alignment with the OMOP CDM facilitates reproducible analytics and participation in international federated research networks, while secure access infrastructure ensures compliance with data protection regulations. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by the Estonian Research Council (PRG1844). The study was funded by the European Union and co-funded by the Ministry of Education and Research (TEM-TA72). The European Union funded the project under its Horizon Europe research and innovation programme (grant agreement No 101060011, TeamPerMed) and co-funded the research through the European Regional Development Fund (Project No. 2021-2027.1.01.24-0444). Views and opinions expressed are, however, those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. This work was also supported by the Estonian Centre of Excellence in Artificial Intelligence (EXAI), funded by the Estonian Ministry of Education and Research grant TK213. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The compilation and use of the EST-Health-30 dataset was approved by the Estonian Bioethics and Human Research Council (no. 1.1-12/817) on 3 March 2025. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Access is governed through a multi-step process requiring a cooperation agreement with the Health Informatics Research Group at the University of Tartu, ethics approval from the national scientific ethics committee coordinated by the Estonian Research Council, and subsequent authorization from the Ministry of Social Affairs.
BackgroundFor accurate medication usage statistics and medication adherence calculations, we need to have an accurate days’ supply (DS) for each prescription. Unfortunately, often the DS or the information needed for calculating the DS is not provided. Therefore, other methods need to be applied to acquire missing values or substitute incorrect values. ObjectiveThis study aims to apply a variety of methods for managing incomplete and missing data to enhance the accuracy of calculating DS for all medications and drug forms alike. Furthermore, to describe the effect of applied methods on the medication adherence calculated on real-world data. MethodsA dataset comprising prescription records from a 10% (150,824 patients) random sample of the Estonian population between 2012 and 2019 was used. The workflow consisted of 3 steps: data cleaning, imputation, and calculation of DS. For imputation, different methods were combined, such as calculating mode-based daily dose, or using usage guidelines from the Summary of Product Characteristics or legislation. DS was calculated based on the provided daily dose or imputed value. To evaluate the impact of data cleaning, medication adherence for the baseline dataset and corrected dataset for 2 time periods, 2012-2015 and 2017-2019, was calculated and compared. ResultsThe drug forms with the lowest proportion of correct DS provided were insulin injections (2601/82,867, 3.1%) and intravaginal contraceptives (1692/21,145, 8%) while the highest proportion of DS was provided for inhalation medication (78,541/126,588, 62%), oral drops (52,085/98,221, 53%) and tablets, capsules, suppositories (2,828,617/6,176,585, 45.8%). As a result of applying different imputation approaches, we successfully found the DS for 98.3% (7,415,347/7,544,892) of dispensed prescriptions. For the remaining 1.7% (129,545/7,544,892) of prescriptions, DS could not be imputed nor calculated with these methods. As for the medication adherence, the distinction between 2 observed time periods was more distinct in the baseline dataset compared with the corrected dataset for most of the drug groups, indicating that the applied correction methods had lessened the stark contrast. ConclusionsIn summary, our study demonstrated that with a carefully designed imputation pipeline where data-driven imputation is combined with domain knowledge and literature information, it is possible to meaningfully improve the quality of prescription datasets and generate more accurate and consistent adherence metrics across various drug forms. Nonetheless, future efforts should continue to refine imputation techniques, incorporate machine learning approaches where appropriate, and expand validation efforts using external benchmarks or clinical outcomes.
Abstract Background The current knowledge about medication adherence is based on studies focusing only on few health conditions and little is known about how strongly adherence is shaped by person-specific behaviour. The aim of the cohort study is to 1) evaluate the effect of multiple factors affecting medication adherence in a consistent manner across 137 active substances, and 2) calculate individual medication adherence score (IMAS), evaluate its predictive power, stability over time, and impact on health outcomes. In essence, IMAS describes persons’ medication-taking “baseline”. Methods We utilised a representative dataset with electronic health records, claims, and dispensed medications across 137 active substances and applied continuous multiple interval measures of medication availability (CMA). To assess the effect of various demographic, health, and medication-related variables on CMA, we employed linear mixed models. Results Here we show that the medication adherence ranged from 0.423 (albuterol, 95% CI 0.414–0.432) to 0.922 (warfarin, 95% CI 0.917–0.926). The demographic, health- and medication-related factors explained 11.6% and IMAS 22.0% of the variation in adherence. IMAS predicted adherence across medication classes, reduced the risk of overall hospitalisation (hazard ratio = 0.76, 95% CI 0.60–0.97, p < 0.05) and cause-specific incidence for 17 conditions. Conclusions Thus, IMAS represents a person-level metric that captures baseline medication-taking behaviour across therapeutic classes and predicts both medication adherence as well as health outcomes. Our analysis suggests that medication-taking behaviour represents a broader patient-level phenomenon manifesting consistently across medications, suggesting its potential for personalised interventions in clinical practice and more efficient public health strategies and policies.
Heart failure (HF) prevalence is increasing and is relatively well described. Less is known about using tests in diagnosing HF in clinical practice. The aim of this article was to describe the diagnostic practices among incident HF patients in Estonia. Electronic health records and healthcare provision claims from 1st January 2012 to 31st December 2019 for a random sample of 10
Background:Real-world evidence provides valuable insights into cancer burden, presentation, and care variations. Through a large-scale federated approach, this study aims to explore patient characteristics and overall survival for eight cancers using data from 11 electronic health records and cancer registries from eight European countries, mapped to the Observational Medical Outcomes Partnership Common Data Model (OMOP-CDM). Methods:Patients aged 18 years or older with a primary cancer diagnosis between 2000 and 2019 were included. Patients were followed from cancer diagnosis until death, database exit, or study end. Mortality data was sourced from linked national or subnational death registries for most databases. Patient characteristics, including comorbidities, and medication use, were summarised. Age-standardised overall survival (OS) at one, five, and ten years were calculated using the Kaplan-Meier method and stratified by cancer type, age group and sex. Findings:There were 1,796,278 eligible cancer patients included with most diagnoses in individuals aged 60-79 years. Top comorbidities and medications were relatively consistent across databases, with certain variations observed by cancer type, possibly indicative of early cancer signs and risk factors. For instance, anaemia was frequent in colorectal (9% [HUS]-23% [IMASIS]; 791/8395-730/3141 individuals) and stomach cancers (10% [HUS]-34% [IMASIS]; 130/1277-225/670), while chronic obstructive pulmonary disease (18% [SIDIAP]-34% [HUVM], 5310/29,009-1039/3063) and pneumonia (5% [CPRD GOLD]-33% [UTARTU], 1904/34,990-1001/3063) were common in lung cancer patients. Breast and prostate cancers had the highest one, five and ten-year overall survival, with 5-year OS ranging from 76% [ECi]-85% [IMASIS] and 75% [HUVM]-83% [SIDIAP], respectively. Pancreatic cancer showed the lowest survival ranging from 3% [NCR]-25% [IMASIS] 5-year OS. Variations in cancer survival estimates were observed across data sources and countries. Interpretation:Federated analysis of diverse European real-world databases, standardised to OMOP-CDM, offer a valuable benchmark for future cancer research, particularly in understanding prodromes and risk factors, often recorded in routinely collected healthcare data prior to cancer onset. Funding:The European Health Data & Evidence Network has received funding from the Innovative Medicines Initiative 2 Joint Undertaking (JU) under grant agreement No 806968. The JU receives support from the European Union's Horizon 2020 research and innovation programme and the European Federation of Pharmaceutical Industries and Associations partners.
Abstract Background: For accurate medication usage statistics and medication adherence calculations, we need to have an accurate days' supply (DS) for each prescription. Unfortunately, often the DS or information needed for calculating the DS is not provided. Therefore, other methods need to be applied to acquire missing values or substituting incorrect values. Objective: The aim of this study is to apply a variety of methods for managing incomplete and missing data to enhance the accuracy of calculating DS for all medications and drug forms alike. Furthermore, to describe the effect of applied methods on the medication adherence calculated on real-world data. Methods: A dataset comprising prescription records from a 10% random sample of the Estonian population between 2012 and 2019 was used. The workflow consisted of three steps - data cleaning, imputation and calculation of DS. For imputation, different methods were combined, such as calculating mode-based daily dose, or using usage guidelines from Summary of Product Characteristics (SPCs) or legislation. DS was calculated based on provided daily dose or imputed value. To evaluate the impact of data cleaning, medication adherence for baseline dataset and corrected dataset for two time periods 2012-2015 and 2017-2019 was calculated and compared. Results: The drug forms with the lowest proportion of correct DS provided were insulin injections (3.1%) and intravaginal contraceptives (8.0%) while the highest proportion of DS was provided for inhalation medication (57.5%), oral drops (53.0%) and tablets, capsules, suppositories (45.8%). As a result of applying different imputation approaches, we successfully found the DS for 98.3% (N=7,415,347) of dispensed prescriptions. For the remaining 1.7% (N=129,545) of prescriptions DS could not be imputed nor calculated with these methods. As for the medication adherence, the distinction between two observed time periods was more distinct in the baseline dataset compared with the corrected dataset for most of the drug groups, indicating that the applied correction methods had lessened the stark contrast. Conclusions: In summary, our study demonstrated that with a carefully designed imputation pipeline where data-driven imputation is combined with domain knowledge and literature information, it is possible to meaningfully improve the quality of prescription datasets and generate more accurate and consistent adherence metrics across various drug form. Nonetheless, future efforts should continue to refine imputation techniques, incorporate machine learning approaches where appropriate, and expand validation efforts using external benchmarks or clinical outcomes. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study was co-funded by the European Union and Estonian Ministry of Education and Research via project TEM-TA72 and Estonian Research Council grant PRG1844. This project has received funding from the European Union's Horizon Europe research and innovation programme under grant agreement No 101060011. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. This research was co-funded by the European Union through the European Regional Development Fund (Project No. 2021-2027.1.01.24-0444). This research was funded by the Estonian Ministry of Education and Research (Teaming for Excellence). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study was approved by the Research Ethics Committee of the University of Tartu (300/T-23) and the Estonian Committee on Bioethics and Human Research (1.1-12/653), and the requirement for informed consent was waived. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The datasets generated and analysed during this study are not publicly available due to legal restrictions on sharing de-identified data. According to legislative regulation and data protection law in Estonia, the authors cannot publicly release the data received from the health data registries in Estonia. However, the data can be requested by completing necessary applications in order to carry out research or an evaluation of public interest and acquiring the permission of the controller of the databases.
BIG-HEART cohort was established to study and lower health inequalities of cardiovascular disease, by linking a multidimensional and rich collection of Estonia's electronic health and social data. The dataset includes all individuals aged 36 and above who resided in Estonia in 2012 (N=770,323). Its complete population representation minimises sampling and healthy volunteer bias. Existing funding and permits will follow up for health outcomes annually until at least 2026, with future extensions possible. The dataset integrates all routinely collected individual-level primary and secondary care health data (including inpatient and outpatient attendance, diagnoses made, prescriptions issued and purchased), along with mortality data, plus comprehensive social data (including ethnicity, education, marital status, social benefits, unemployment history, land and business ownership) from eight national registries. This enables researchers to explore novel dimensions of social epidemiology, including unbiased measures of wealth, medication adherence and care quality, as well as the derivation and validation of equity-enhancing clinical risk prediction algorithms and large language models. Health and social data are linked using pseudonymised identifiers derived from national personal identification numbers, ensuring accuracy and privacy. The data are stored in the OMOP common data model, facilitating international collaboration. Collaboration inquiries can be directed to the BIG-HEART team at taavi.tillmann{at}ut.ee to explore potential opportunities. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by the Estonian Research Council grants PSG809 and PRG1844. This research was co-funded by the European Union through the European Regional Development Fund (Project No. 2021-2027.1.01.24-0444) and through Estonian Ministry of Education and Research (project TEM-TA72; and grant TK218 establishing the Estonian Centre of Excellence for Well-Being Sciences (EstWell)). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The BIG-HEART project has received approval from the Research Ethics Committee of the University of Tartu [no. 384/T-8, 20.11.2023]. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Background:The opioid crisis has been a serious public health challenge in North America for decades, despite numerous efforts to mitigate its devastating consequences. As concerns grow about a similar situation developing in Europe, we evaluated the trends in opioid use and characterized prescribing indications across seven European countries. Methods:We conducted a multinational cohort study using electronic health records from various healthcare settings: primary care [Clinical Practice Research Datalink (CPRD) GOLD (United Kingdom), Sistema d'Informació per al Desenvolupament de la Investigació en Atenció Primària (SIDIAP, Spain), and Integrated Primary Care Information Project (IPCI, the Netherlands)]; primary and outpatient specialist care [IQVIA Disease Analyzer (DA) Germany and IQVIA Longitudinal Patient Database (LPD) Belgium]; hospital care [Clinical Data Warehouse of Bordeaux University Hospital (CHUBX, France)]; and the Estonian Biobank (EBB). All data were mapped to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM). All people registered in a contributing database for ≥365 days between 2012 and 2022 were included. Annual period prevalence and incidence rates of opioid prescriptions were estimated, and long-term trends were quantified as the percent change from 2012 to 2019. New opioid users were characterized, including potential prescribing indications. Results:Between 2012 and 2019, the incidence of opioid prescriptions in primary care decreased by -50·7% (CPRD GOLD) and -2·0% (SIDIAP), while it increased in EBB (+52·8%) and CHUBX (+25·3%) data. The incidence of codeine and tramadol use decreased in most databases. However, the prevalence of oxycodone, morphine, and fentanyl increased. Opioid use was highest among older age groups, and the majority of prescriptions were for oral formulations. Respiratory and pain-related conditions were the most common indications for new opioid users in outpatient settings. Conclusion:Despite a decrease in new opioid prescriptions in many European countries, the prevalence of opioid use remained largely stable over the last decade. More data are needed to monitor evolving opioid prescription patterns in Europe, particularly in the post-pandemic era.
Large biobanks have set a new standard for research and innovation in human genomics and implementation of personalized medicine. The Estonian Biobank was founded a quarter of a century ago, and its biological specimens, clinical, health, omics, and lifestyle data have been included in over 800 publications to date. What makes the biobank unique internationally is its translational focus, with active efforts to conduct clinical studies based on genetic findings, and to explore the effects of return of results on participants. In this review, we provide an overview of the Estonian Biobank, highlight its strengths for studying the effects of genetic variation and quantitative phenotypes on health-related traits, development of methods and frameworks for bringing genomics into the clinic, and its role as a driving force for implementing personalized medicine on a national level and beyond.
Objective This study aims to address the gap in the literature on converting real-world Clinical Document Architecture (CDA) data into the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), focusing on the initial steps preceding the mapping phase. We highlight the importance of a repeatable Extract-Transform-Load (ETL) pipeline for health data extraction from HL7 CDA documents in Estonia for research purposes. Methods We developed a repeatable ETL pipeline to facilitate the extraction, cleaning, and restructuring of health data from CDA documents to OMOP CDM, ensuring a high-quality and structured data format. This pipeline was designed to adapt to continuously updated data exchange format changes and handle various CDA document subsets for different scientific studies. Results We demonstrated via selected use cases that our pipeline successfully transformed a significant portion of diagnosis codes, body weight and eGFR measurements, and PAP test results from CDA documents into OMOP CDM, showing the ease of extracting structured data. However, challenges such as harmonising diverse coding systems and extracting lab results from free-text sections were encountered. The iterative development of the pipeline facilitated swift error detection and correction, enhancing the process’s efficiency. Conclusion After a decade of focused work, our research has led to the development of an ETL pipeline that effectively transforms HL7 CDA documents into OMOP CDM in Estonia, addressing key data extraction and transformation challenges. The pipeline’s repeatability and adaptability to various data subsets make it a valuable resource for researchers dealing with health data. While tested on Estonian data, the principles outlined are broadly applicable, potentially aiding in handling health data standards that vary by country. Despite newer health data standards emerging, the relevance of CDA for retrospective health studies ensures the continuing importance of this work.
Sven Laur合作论文数Institute of Computer Science, Faculty of Mathematics and Computer Science, University of Tartu8