In Deutschland sind gegenwärtig mindestens 1,8 Millionen von Demenz betroffen; bis 2060 droht ein Anstieg auf 2,1 Millionen Betroffene. Prävention bietet die derzeit beste Möglichkeit zur effektiven Linderung der Krankheitslast. Das Krankheitsrisiko und der Krankheitsverlauf sowie die Effektivität von Präventionsmaßnahmen werden von diversen Faktoren beeinflusst; in Kombination ergeben diese potenziell individuelle Risikoprofile als Grundlage für eine effektive Demenzprävention. Die Hebung dieser Potenziale erfordert kurzfristig insbesondere eine bessere Nutzung bestehender Daten inklusive einer erweiterten Datenerfassung und Nutzbarmachung für die Forschung, wodurch bevölkerungsbasierte Screening- und Interventionsmaßnahmen entwickelt und eine Strategie zur personalisierten Demenzprävention realisiert werden können.
BACKGROUND:Food allergy (FA) arises from a complex interplay between an individual's genetic predisposition and environmental factors, and its prevalence is increasing. Genome-wide association studies to date have been hindered by small sample sizes and varying FA definitions. OBJECTIVE:We sought to identify novel FA risk loci by conducting a genome-wide association study meta-analysis in children and adults by using a multiphenotype approach to ensure a good trade-off between sufficient sample size and valid FA definitions. METHODS:Analyses were conducted separately in children and adults on the basis of the following FA phenotypes: self-report, doctor diagnosis, food-specific sensitization, and doctor diagnosis plus food-specific sensitization. A meta-analysis was performed of genome-wide association studies from up to 16 cohorts of people of European ancestry including 229,426 adults and 14,234 children. Models were adjusted for sex, age, principal components, and, if applicable, further study-specific confounders. Sensitivity models were additionally adjusted for hay fever. Replication was conducted in additional external cohorts and a validation in oral food challenge-defined FA cases. RESULTS:Thirty-seven single nucleotide polymorphisms met suggestive significance (P < 1 × 10-6), with two reaching genome-wide significance: rs116936231 (FGL1) in adult doctor-diagnosed FA plus food-specific sensitization phenotype (stable after additional hay fever adjustment) and rs8022829 (AKAP6-NPAS3), which was significant only in the hay fever-adjusted model in adults. However, neither variant was validated. Further, we identified 3 single nucleotide polymorphisms previously reported for FA and atopic disease. CONCLUSION:This study identified 37 single nucleotide polymorphisms suggestively associated with FA and demonstrated genetic differences across phenotypes. It highlights the need for a unified FA definition and sheds light on FA's shared genetic architecture with allergies.
Data harmonization is a prerequisite for joint cohort analyses. In this review, we aim to identify and contrast statistical methods for retrospective harmonization of longitudinal data. We performed a scoping review following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews guidelines. Studies were included if they described statistical methods for retrospectively harmonizing longitudinal data at the participant level. From 35 included papers out of 1,234 hits, we identified three types of statistical methods applicable to tabular data commonly collected in longitudinal epidemiological studies (e.g., questionnaires): (1) distribution-based methods, (2) the proportion score model, and (3) latent variable models. Our results suggest that the suitability of a statistical harmonization method mainly depends on the measurement scales of the original variables as well as on the type of target variable (directly measurable vs. latent). The chosen harmonization method influences how missing subsets of variables are addressed. None of the included studies applied more automated approaches such as machine learning-based procedures for deriving a harmonized dataset. Based on our findings, we present a roadmap that can guide researchers in selecting the most appropriate statistical method for a specific harmonization task and in handling variables collected only in a subset of studies. Data harmonization is still a demanding task that requires the development and application of novel tools for automating the procedures.
BACKGROUND:The German National Cohort (NAKO Gesundheitsstudie) is a prospective cohort study with 205 053 participants. Its goal is to identify risk factors for chronic diseases, including cancer. METHODS:We describe the methods of ascertaining cancer cases in NAKO (i.e., linkage with cancer registries and self-reporting), the incidence and prevalence figures obtained so far, and the ratios of observed to expected (O/E) case numbers. RESULTS:Case ascertainment identified 2774 existing cancer cases diagnosed within the five years prior to study enrollment and 4295 new cancer cases diagnosed up to five years after enrollment. Cancers of the breast, prostate, lung, and colorectum made up 55% of the incident cases. Fewer incident cases were found than would have been expected on the basis of incidences in the general population (O/E ratio 0.80, 95% confidence interval [0.76;0.83] for the first two-years of prospective follow-up); the O/E ratio varied across tumor sites (breast 1.02, prostate 1.12, lung 0.38, colorectum 0.62). Possible explanations include healthy volunteer bias, delayed reporting, as well as incomplete data collection from registries and self-reporting. CONCLUSION:Rising case numbers for the most common types of cancer are expected in the next few years and will enable comprehensive epidemiological analysis of the risk factors of cancer.
Background: Access to high-quality, FAIR health data is essential for advancing epidemiological, public health and clinical research. NFDI4Health, one of the consortia of Germany’s National Research Data Infrastructure (NFDI), was established in 2020 to address this need by improving data FAIRness for the scientific community. Methods: During the initial funding phase, NFDI4Health developed key infrastructure components and services focusing on interoperability, data sharing and research support. Our approach was guided by user needs and real-world use cases in alignment with (inter)national FAIR standards and infrastructures. Results: We advanced findability of health data by establishing a central Health Study Hub and connecting it to local infrastructures, e.g., through the Local Data Hub software, to facilitate transfer of metadata from the local infrastructures to the Health Study Hub. Accessibility was improved by expanding the German Research Data Portal for Health (FDPG) to provide central access to study data and by providing tools to support anonymisation and synthetic data generation. We also developed an interoperable metadata schema for publishing study data and implemented it in the Health Study Hub. To further improve interoperability, a NFDI4Health FAIR sharing collection and AI support for metadata annotation and harmonisation workflows were established. To enhance reusability, data quality assessment tools were further developed, and two frameworks for federated analysis of sensitive health data were piloted. Intensive engagement with our communities through training, advisory services, and collaborative development ensured user relevance. Conclusion/outlook: NFDI4Health has established a scalable, interoperable, and user-centred infrastructure to support FAIR data sharing in health research. Future work will focus on further developing and consolidating the infrastructure, expanding cross-domain data integration, fostering broader adoption within our research community, and strengthening national and international collaboration.
BACKGROUND:Synthetic data hold substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. METHODS:We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF's performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalization, and runtime. RESULTS:Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalization relative to other synthesizers and superior computational efficiency. CONCLUSIONS:In summary, ARF reliably generates high-quality synthetic data that replicate diverse epidemiological analyses while offering a competitive privacy-utility trade-off.
Linking different health data at the personal level (record linkage, RL) allows answering scientific questions that could otherwise not be answered by a single data source. Linked data therefore offer great potential for health research to improve prevention, treatment, and care at the population level. Personal health data are protected by strict legal regulations. Its use requires balancing legitimate interests in protecting personal data and health benefits. However, current laws and their interpretations in Germany place severe restrictions on health data RL such that its potential for improving health outcomes is still to be leveraged. In Germany, RL is also hindered by the lack of a unique identifier that enables error-free merging across different data sources. Overall, there is a lack of interoperable solutions to perform comprehensive RL across studies and data sources in a secure environment.In this article, we propose solutions for the linkage of personal health data from different sources based on the White Paper - Verbesserung des Record Linkage für die Gesundheitsforschung in Deutschland. Our proposed solutions include, among others, the establishment of a health ID and the creation of a decentralized federated research data infrastructure with central components. Although these proposals are in line with the General Data Protection Regulation, there is a need for further legal regulation in specific cases.
Academic performance in children is associated with a range of health-related factors, including physical fitness, mental well-being, sleep, and behavioral patterns. While previous studies have examined these factors individually, fewer have assessed their independent associations with academic achievement while accounting for other relevant health indicators. This study uses data from the I.Family study to explore how physical, mental, sleep-related, and behavioral health indicators relate to academic achievement among European adolescents, considering each factor’s contribution while adjusting for the others. We used data from the 2013–2014 wave of the I.Family study to investigate eight health indicators: health related quality of life (HRQoL), body mass index (BMI), diet, media use, physical activity, sleep duration and quality, and stressful life events. Their associations with self-reported academic performance in mathematics and language were analyzed using binary logistic regression models, adjusting for confounders such as parents’ education, income, survey country and child’s age. We conducted separate analyses for girls and boys to capture associations that are specific to academic subject and sex. A number of significant associations were found between several health indicators and academic performance. Higher HRQoL scores, reduced media time, and increased physical activity were linked to better academic performance in both mathematics and language for both boys and girls. Variation by sex and academic subjects were observed, with lower BMI, higher healthy diet scores and better sleep quality associated with better academic performance in language among girls. For mathematics, emotional, self-esteem, and family-related HRQoL were all significantly associated with higher performance for both boys and girls. In contrast, for language achievement, only family-related HRQoL was significant for both sexes. Our study underscores the need to consider both the importance of accounting for heterogeneity in sex and the differences between math and language academic subjects when investigating determinants of academic performance, setting the stage for further research on this topic to explore potential competing, synergistic, or time-dependent effects among these different health dimensions.
The global increase of overweight and obesity in children and adults is one of the most prominent public health threats, often accompanied by insulin resistance, hypertension, and dyslipidemia. The simultaneous occurrence of these health problems is referred to as metabolic syndrome. Various criteria have been proposed to define this syndrome, but no general consensus on the specific markers and the respective cut-offs has been achieved yet. As a consequence, it is difficult to assess regional variations and temporal trends and to obtain a comprehensive picture of the global burden of this major health threat. This limitation is most striking in childhood and adolescence, when metabolic parameters change with developmental stage. Obesity and related metabolic disorders develop early in life and then track into adulthood, i.e., the metabolic syndrome seems to originate in the early life course. Thus, it would be important to monitor the trajectories of cardio-metabolic parameters from early on. We will summarize selected key studies to provide a narrative overview of the global epidemiology of the metabolic syndrome while considering the limitations that hinder us to provide a comprehensive full picture of the problem. A particular focus will be given to the situation in children and adolescents and the risk factors impacting on their cardio-metabolic health. This summary will be complemented by key findings of a pan-European children cohort and first results of a large German adult cohort.
Mit der Verknüpfung unterschiedlicher Gesundheitsdaten auf Personenebene (Record Linkage, RL) lassen sich wissenschaftliche Fragestellungen beantworten, die mit einer Datenquelle alleine nicht zu beantworten wären. Die verknüpften Daten entfalten daher ein großes Potenzial für die Gesundheitsforschung, um Prävention, Therapie und Versorgung auf Bevölkerungsebene zu verbessern. Personenbezogene Gesundheitsdaten sind durch strenge Rechtsvorschriften geschützt. Ihre Nutzung bedarf einer Abwägung zwischen dem berechtigten Interesse am Schutz der persönlichen Daten und dem gesundheitlichen Nutzen. Allerdings schränken derzeit Gesetze und ihre Auslegung das RL von Gesundheitsdaten in Deutschland so stark ein, dass ihr Potenzial zur Verbesserung der Gesundheit bisher nur unzureichend genutzt werden kann. Das RL wird in Deutschland insbesondere durch das Fehlen eines Unique Identifier erschwert, der eine fehlerfreie Zusammenführung über verschiedene Datenkörper hinweg ermöglicht. Insgesamt fehlen interoperable Lösungen, um ein umfassendes studien- und datenkörperübergreifendes RL in einer gesicherten Umgebung durchführen zu können. In diesem Artikel schlagen wir basierend auf dem „White Paper – Verbesserung des Record Linkage für die Gesundheitsforschung in Deutschland“ Lösungen für die personenbezogene Datensatzverknüpfung von unterschiedlichen Datenquellen vor. Unsere Lösungsvorschläge beinhalten u. a. die Etablierung einer Gesundheits-ID und die Schaffung einer dezentral-föderierten Forschungsdateninfrastruktur mit zentralen Komponenten. Auch wenn diese Vorschläge im Einklang mit der Datenschutzgrundverordnung stehen, gibt es im Einzelfall weiteren gesetzlichen Regelungsbedarf.
According to the World Health Organization, public health surveillance is a continuous process of collecting, analyzing, and interpreting health data such that the results can be directly disseminated to responsible actors to take immediate action if necessary. Although disease surveillance has its roots in managing infectious diseases, it has also become essential for tracking non-communicable diseases (NCDs) as well as health-related behaviors and contextual factors. Due to the increasing digitalization of the health space, traditional methods of data collection are being complemented by digital information such as Electronic Health Records or self-generated internet data, e.g., on social media. In this chapter, we will first define the term surveillance before introducing different approaches that are used depending on whether the focus is on infectious diseases or NCDs. Here, we will distinguish two major threads that underline the huge impact of digitalization on surveillance systems, i.e., the usage of digital technologies for primary data collection and the usage of routinely collected available digital data. We will then provide some prominent examples of surveillance systems before we conclude by highlighting the importance of the FAIR guiding principles for data sharing and of record linkage for disease surveillance.
Digital public health has received a significant boost in recent years, especially due to the demands associated with the COVID-19 pandemic. In this report, we provide an overview of the developments in digitalization in the field of public health in Germany since 2020 and illustrate these with examples from the Leibniz ScienceCampus Digital Public Health Bremen (LSC DiPH).The following topics are central: How do digital survey methods as well as digital biomarkers and artificial intelligence methods shape modern epidemiology and prevention research? What is the status of digitalization in public health offices? Which approaches to health economics evaluation of digital public health interventions have been utilized so far? What is the status of training and further education in digital public health?The first years of the Leibniz ScienceCampus Digital Public Health Bremen (LSC DiPH) were also strongly influenced by the COVID-19 pandemic. Repeated population-based digital surveys of the LSC indicated an increase in use of health apps in the population, for example, in applications to support physical activity. The COVID-19-pandemic has also shown that the digitalization of public health enhances the risk of misinformation and disinformation.
IntroductionPharmacovigilance is vital for drug safety. The process typically involves two key steps: initial signal generation from spontaneous reporting systems (SRSs) and subsequent expert review to assess the signals’ (potential) causality and decide on the appropriate action.MethodsWe propose a novel discovery and verification approach to pharmacovigilance based on electronic healthcare data. We enhance the signal detection phase by introducing an ensemble of methods which generated signals are combined using Borda count ranking; a method designed to emphasize consensus. Ensemble methods tend to perform better when data is noisy and leverage the strengths of individual classifiers, while trying to mitigate some of their limitations. Additionally, we offer the committee of medical experts with the option to perform an in-depth investigation of selected signals through tailored pharmacoepidemiological studies to evaluate their plausibility or spuriousness. To illustrate our approach, we utilize data from the German Pharmacoepidemiological Research Database, focusing on drug reactions to the direct oral anticoagulant rivaroxaban.ResultsIn this example, the ensemble method is built upon the Bayesian confidence propagation neural network, longitudinal Gamma Poisson shrinker, penalized regression and random forests. We also conduct a pharmacoepidemiological verification study in the form of a nested active comparator case-control study, involving patients diagnosed with atrial fibrillation who initiated anticoagulant treatment between 2011 and 2017.DiscussionThe case study reveals our ability to detect known adverse drug reactions and discover new signals. Importantly, the ensemble method is computationally efficient. Hasty false conclusions can be avoided by a verification study, which is, however, time-consuming to carry out. We provide an online tool for easy application: https://borda.bips.eu.
Childhood obesity is a complex disorder that appears to be influenced by an interacting system of many factors. Taking this complexity into account, we aim to investigate the causal structure underlying childhood obesity. Our focus is on identifying potential early, direct or indirect, causes of obesity which may be promising targets for prevention strategies. Using a causal discovery algorithm, we estimate a cohort causal graph (CCG) over the life course from childhood to adolescence. We adapt a popular method, the so-called PC-algorithm, to deal with missing values by multiple imputation, with mixed discrete and continuous variables, and that takes background knowledge such as the time-structure of cohort data into account. The algorithm is then applied to learn the causal structure among 51 variables including obesity, early life factors, diet, lifestyle, insulin resistance, puberty stage and cultural background of 5112 children from the European IDEFICS/I.Family cohort across three waves (2007–2014). The robustness of the learned causal structure is addressed in a series of alternative and sensitivity analyses; in particular, we use bootstrap resamples to assess the stability of aspects of the learned CCG. Our results suggest some but only indirect possible causal paths from early modifiable risk factors, such as audio-visual media consumption and physical activity, to obesity (measured by age- and sex-adjusted BMI z-scores) 6 years later.
Background In Germany, all citizens must purchase health insurance, in either statutory (SHI) or private health insurance (PHI). Because of the division into SHI and PHI, person insurance's status is an important variable for studies in the context of public health research. In the German National Cohort (NAKO), the variable on self-reported health insurance status of the participants has a high proportion of missing values (55.4%). The aim of our study was to develop and internally validate models to predict the health insurance status of NAKO baseline survey participants in order to replace missing values. In this respect, our research interest was focused on the question to which extent socio-demographic characteristics are suitable for predicting health insurance status. Methods We developed two prediction models including 53,796 participants to estimate the probability that a participant is either member of a SHI (model 1) or PHI (model 2). We identified eight predictors by literature research: occupation, income, education, sex, age, employment status, residential area, and marital status. The predictive performance was determined in the internal validation considering discrimination and calibration. Discrimination was assessed based on the Area Under the Curve (AUC) and the Receiver Operating Characteristic (ROC) curve and calibration was assessed based on the calibration slope and calibration plot. Results In model 1, the AUC was 0.91 (95% CI: 0.91-0.92) and the calibration slope was 0.97 (95% CI: 0.97-0.97). Model 2 had an AUC of 0.91 (95% CI: 0.90-0.91) and a calibration slope of 0.97 (95% CI: 0.97-0.97). Based on the calculated performance parameters both models turned out to show an almost ideal discrimination and calibration. Employment status and household income and to a lesser extent educational level, age, sex, marital status, and residential area are suitable for predicting health insurance status. Conclusions Socio-demographic characteristics especially employment status and household income assessed at NAKO's baseline were suitable for predicting the statutory and private health insurance status. However, before applying the prediction models in other studies, an external validation in population-based studies is recommended.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis project was conducted with data from the German National Cohort (NAKO) (www.nako.de). The NAKO is funded by the Federal Ministry of Education and Research (BMBF) [project funding reference numbers: 01ER1301A/B/C and 01ER1511D], federal states and the Helmholtz Association with additional financial support by the participating universities and the institutes of the Leibniz Association. We thank all participants who took part in the German National Cohort and the staff in this research program.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:The study protocol of the NAKO was approved by the ethics committee of the Bavarian State Medical Association (13023 and 13031) and by the locally responsible ethics committees of the institutions of the 18 study centres. All the described investigations were conducted in compliance with national law and in accordance with the declaration of Helsinki (in the latest revised version). All participants have been fully informed and have given their written informed consent to participate in the study.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesThe data that support the findings of this study are available from the German National Cohort but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available.
FAIRification of personal health data is of utmost importance to improve health research and political as well as medical decision-making, which ultimately contributes to a better health of the general population. Despite the many advances in information technology, several obstacles such as interoperability problems remain and relevant research on the health topic of interest is likely to be missed out due to time-consuming search and access processes. A recent example is the COVID-19 pandemic, where a better understanding of the virus’ transmission dynamics as well as preventive and therapeutic options would have improved public health and medical decision-making. Consequently, the NFDI4Health Task Force COVID-19 was established to foster the FAIRification of German COVID-19 studies. This paper describes the various steps that have been taken to create low barrier workflows for scientists in finding and accessing German COVID-19 research. It provides an overview on the building blocks for FAIR health research within the Task Force COVID-19 and how this initial work was subsequently expanded by the German consortium National Research Data Infrastructure for Personal Health Data (NFDI4Health) to cover a wider range of studies and research areas in epidemiological, public health and clinical research. Lessons learned from the Task Force helped to improve the respective tasks of NFDI4Health.
Zusammenfassung Digital Public Health hat in den vergangenen Jahren insbesondere durch die mit der COVID-19-Pandemie verbundenen Anforderungen einen erheblichen Schub erfahren. In diesem Bericht geben wir einen Überblick über die Entwicklungen in der Digitalisierung im Bereich Public Health in Deutschland seit 2020 und illustrieren diese mit Beispielen aus dem Leibniz-WissenschaftsCampus Digital Public Health Bremen (LWC DiPH). Zentral sind dabei folgende Themen: Wie prägen digitale Erhebungsmethoden sowie digitale Biomarker und Methoden der künstlichen Intelligenz die moderne epidemiologische und Präventionsforschung? Wie steht es um die Digitalisierung im öffentlichen Gesundheitsdienst? Welche Ansätze der gesundheitsökonomischen Evaluation von digitalen Public-Health-Interventionen wurden bisher eingesetzt? Wie steht es um die Aus- und Weiterbildung in diesem Bereich? Auch die Arbeit des LWC DiPH war zunächst stark durch die COVID-19-Pandemie geprägt. Wiederholte populationsbezogene digitale Surveys des LWC DiPH ergaben Hinweise auf eine häufigere Nutzung von Gesundheitsapps in der Bevölkerung in Deutschland, z. B. bei den Anwendungen zur Unterstützung der körperlichen Aktivität. Dass die Digitalisierung von Public Health das Risiko von gezielten Fehl- und Desinformationen mit sich bringt, hat die COVID-19-Pandemie ebenfalls gezeigt.
To investigate the reliability of parental recall of birth weight, birth length and gestational age several years after birth. Parentally recalled birth parameters were obtained from the European multicentric cohort study IDEFICS (Identification and prevention of dietary- and lifestyle-induced health effects in children and infants) and compared to the corresponding data externally recorded in the child’s medical check-up booklet. The agreement between the two sources was examined using Bland–Altman plots, intraclass correlation coefficients and Cohen’s kappa for clinically relevant categories. Additionally, logistic regression models were used to identify factors related to parental recall accuracy. A total of 4930 children aged 2 to 11 years were included. Accuracy of birth weight within 100 g was 88