A post-authorisation safety study was conducted for the AZD1222 COVID-19 vaccine. This paper presents one study outcome, thrombosis with thrombocytopenia syndrome (TTS), and estimates TTS risk in subjects administered ≥1 AZD1222 dose versus concurrent unvaccinated, pre-pandemic historical, or mRNA-vaccinated subjects. The cohort study used data from CPRD Aurum (UK), VID (Spain), SIDIAP (Spain) and PHARMO-GP database (PHARMO) (the Netherlands). AZD1222-vaccinated subjects were matched on age, sex, region, prior COVID-19, and special population status. Incident venous TTS was defined as a thromboembolic event and thrombocytopaenia within ±10 days and no TTS within the prior year. 5,321,930 subjects were matched with concurrent unvaccinated comparators, 4,831,010 with historical comparators, and 4,028,091 with mRNA active comparators (CPRD only). In CPRD, 83% of subjects were vaccinated in Q1 2021; 64% were < 60 years. In VID, SIDIAP, and PHARMO, >59% were vaccinated after Q1 2021; most subjects were ≥ 60 years. Propensity score-weighted incidence rate ratios (IRRs) (95% confidence intervals) for TTS were CPRD, 1.14 (0.60-2.17); VID, 0.34 (0.10-1.18); SIDIAP, 0.66 (0.33-1.34); and zero events in PHARMO. Incidence rates (IRs) and IRRs for TTS, where available, were higher in AZD1222-vaccinated versus concurrent unvaccinated subjects <60 years or during shorter risk windows. After case validation, positive predictive value-adjusted IRRs were < 1. For historical comparators, meta-analysis resulted in an IRR of 1.78 (95% CI,1.12-2.82; I2 = 0%). For mRNA active comparators, the IRR was 1.12 (95% CI,0.61-2.05). Considering the magnitude, precision, and potential biases-such as selection bias due to informative censoring and potential outcome misclassification-the totality of evidence suggests a possible increased risk of TTS with post-AZD1222 vaccination that may be higher among subjects <60 years and 1-42 days after first AZD1222 dose, in line with the literature. Differential age distributions resulting from country-level differences in the risk minimisation measures may explain IRR disparities across data sources.
Synthetic health data enables privacy-safe data sharing, development of code, and mitigation of data scarcity and bias. While most research focuses on generation from electronic health records (EHR), this scoping review focuses on generation with access to only descriptive metadata information (metadata-based methods). Metadata-based methods are suited to privacy sensitive-contexts because they do not require access to individual EHR.This study aimed to identify current methods and research gaps, as well as synthesise relevant evaluation, ethical, and legal frameworks within the context of code development for federated analysis. We searched PubMed, Ovid Emabse, and Web of Science through March 2025 to identify peer-reviewed articles and preprints which focused on metadata-based methods, or which reviewed synthetic data evaluation and/or legal and ethical requirements for health research from all years. After deduplication, abstract and full-text screening were conducted by two researchers for each record. Data extraction and synthesis were conducted by one researcher and validated by a second. Twelve original articles and eleven reviews were included. The twelve articles covered six techniques: four stepwise, one LLM-based, and one syntactical. The methods covered in the primary sources were not benchmarked according to existing evaluation frameworks to identify their strengths and weaknesses; only fidelity metrics were commonly used. Ten of eleven reviews proposed evaluation frameworks and three addressed ethical or legal considerations. We observed inconsistent evaluation definitions and limited exploration of ethical, legal, and governance issues. No reviews specifically discussed code development. In conclusion, in this scoping review we identified several metadata-based methods but limited evaluations for these methods. We found a strong body of literature regarding evaluation but limited specific context-specific guidance. There is a need for consistent evaluation of metadata-based methods, and for specific guidelines aligned with generation method and intended use to ensure quality and compliance with emerging AI requirements.
Objective: To identify current methods and research gaps for synthetic health data generation with access to only descriptive metadata information (metadata-based methods), and to synthesise current research on evaluation, ethical and legal frameworks which may apply to synthetic health data generation for code development in federated analysis. Design: Scoping review Data Sources: PubMed, Ovid Embase, Web of Science through March 18, 2025 Eligibility Criteria: Peer-reviewed articles and preprints which focus on metadata-based methods, or which review synthetic data evaluation and/or legal and ethical requirements for health research from all years. Review Methods: After deduplication, abstract and full-text screening were conducted by two researchers for each record. Data extraction and synthesis were conducted by one researcher. Results: Twelve articles regarding metadata-based methods and eleven literature reviews regarding evaluation and legal or ethical considerations were included. The twelve articles covered six techniques, including Synthea, OSIM-I/II, and SASC. Ten out of eleven literature reviews proposed evaluation metric guidelines or frameworks and three reported on ethical or legal considerations. The methods covered in the primary sources were not benchmarked according to existing evaluation frameworks to identify their strengths and weaknesses; only fidelity metrics were commonly used. Additionally, we found a lack of a standardised evaluation definitions and a limited exploration of ethical, legal, and governance issues in the literature reviews. No literature reviews specifically discussed code development as a use case. Conclusion: We found several methods for metadata-based generation, but limited comparison and evaluation of these methods. Although we found an overall strong body of literature regarding synthetic health data evaluation, we identified a lack of specific guidance for the federated analysis context. There is a need for future research to consistently evaluate metadata-based methods, and for context-specific guidelines aligned with both generation method and intended use to ensure overall quality and compliance with emerging AI requirements. Scoping Review Registration: https://osf.io/9chu6
Methods Several statistical analysis plans (SAP) from the Vaccine Monitoring Collaboration for Europe (VAC4EU) were analyzed to identify the study design sections and specifications for programming RWE studies based on multi-databases standardized to common data models. We envisioned a metadata schema that transforms the epidemiologist's knowledge into a machine-readable format. This machine-readable metadata schema must also contain the different study sections, code lists, and time anchoring specified in the SAPs. Further desired attributes are adaptability and user-friendliness. Results We developed RWE-BRIDGE, a metadata schema with a star-schema model divided into four study design sections with 12 tables: Study Variable Definition with two tables, Cohort Definition with two tables, Post-Exposure Outcome Analysis with seven tables, and Data Retrieval with one table. We provide examples and a step-by-step guide to populate this metadata schema. In addition, we provide a Shiny app that checks the several tables proposed in this metadata strategy. RWE-BRIDGE is available at https://github.com/UMC-Utrecht-RWE/RWE-BRIDGE. Discussion The RWE-BRIDGE has been designed to support the translation of study design sections from statistical analysis plans into analytical pipelines, facilitating collaboration and transparency between lead researchers and scientific programmers and reducing hard coding and repetition. This metadata schema strategy is flexible by supporting different common data models and programming languages, and it is adaptable to the specific needs of each SAP by adding further tables or fields, if necessary. Modified versions of the RWE-BRIGE have been applied in several RWE studies within the VAC4EU ecosystem. Conclusion The RWE-BRIDGE offers a systematic approach to detailing what type of variables, time anchoring, and algorithms are required for a specific RWE study. Applying this metadata schema can facilitate the communication between epidemiologists and programmers in a transparent manner.### Competing Interest StatementACR, DW, VH, MS, TAV, and CLAN are currently salaried employees at University Medical Center Utrecht, which receives institutional research funding from pharmaceutical companies and regulatory agencies and is administered by University Medical Center Utrecht. RE and ZK were salaried employees of University Medical Center Utrecht at the time this project was performed. ### Funding StatementThis study did not receive any funding### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesNo data was used in the present study
OBJECTIVE:To enhance documentation on programming decisions in Real World Evidence (RWE) studies. MATERIALS AND METHODS:We analyzed several statistical analysis plans (SAP) within the Vaccine Monitoring Collaboration for Europe (VAC4EU) to identify study design sections and specifications for programming RWE studies. We designed a machine-readable metadata schema containing study sections, codelists, and time anchoring definitions specified in the SAPs with adaptability and user-friendliness. RESULTS:We developed the RWE-BRIDGE, a metadata schema in form of relational database divided into four study design sections with 12 tables: Study Variable Definition (two tables), Cohort Definition (two tables), Post-Exposure Outcome Analysis (one table), and Data Retrieval (seven tables). We provide a guide to populate this metadata schema and a Shiny app that checks the tables. RWE-BRIDGE is available on GitHub (github.com/UMC-Utrecht-RWE/RWE-BRIDGE). DISCUSSION:The RWE-BRIDGE has been designed to support the translation of study design sections from statistical analysis plans into analytical pipelines and to adhere to the FAIR principles, facilitating collaboration and transparency between researcher and programmers. This metadata schema strategy is flexible as it can support different common data models and programming languages, and it is adaptable to the specific needs of each SAP by adding further tables or fields, if necessary. Modified versions of the RWE-BRIGE have been applied in several RWE studies within VAC4EU. CONCLUSION:RWE-BRIDGE offers a systematic approach to detailing variables, time anchoring, and algorithms for RWE studies. This metadata schema facilitates communication between researcher and programmers.