Chronic kidney disease (CKD) is a critical, progressive condition associated with high mortality and substantial healthcare costs. Early detection is essential, as it can slow disease progression and improve patient outcomes. With the increasing availability of large-scale electronic health records (EHRs), the question arises to what extent these data, when combined with machine-learning algorithms specifically tailored to EHR characteristics, can enhance personalized CKD risk prediction. We developed a transformer model adapted from BEHRT (Bidirectional Encoder Representations from Transformers for EHRs) and specifically tailored for German (GER) EHRs, which we refer to as GERBEHRT. GERBEHRT was pre-trained on outpatient claims data from more than 9 million statutorily insured patients and fine-tuned with nearly 1 million additional patients to predict CKD. The model incorporates EHR features not previously explored in BERT-based approaches and introduces an efficient method to represent multiple attributes per medical concept, such as diagnoses and medications. GERBEHRT was compared with more traditional models and predictions restricted to established risk factors, and the importance of its input features was assessed through an ablation study. In a test cohort of 3.7 million patients with 1.5
Abstract Background Digital decision-support tools such as triage systems and symptom checkers support millions of health-related decisions each year. Their quality and safety are commonly evaluated using textual patient cases, known as case vignettes. However, existing vignette sets written by medical experts cover only a limited spectrum of real-world patient presentations and lack population weights, which would allow extrapolating evaluation results to the underlying patient population. Objective This study aims to develop a data-driven framework for automatically generating a human-manageable set of case vignettes from nationwide triage data that captures broad presentation diversity and links each vignette to a quantitative weight reflecting the number of underlying patient assessments. Methods From 3.2 million triage assessments conducted over one year using structured triage software in the German medical on-call service (telephone triage and online self-triage) and at the joint contact points of the outpatient emergency care service and hospital emergency departments, we randomly sampled 50,000 cases. Triage questionnaires were converted into semantic embeddings using a German Sentence Transformer Model and grouped by agglomerative clustering. For clusters containing sufficient assessments, we generated one representative assessment using a two-phase simulated-annealing optimization. The optimization minimized the distance to the cluster centroid while maximizing the number of answered triage questions, aiming for high representativeness and information content. Each representative assessment was assigned the size of its source cluster as its sample-based weight. A similarity-based sensitivity analysis was performed to examine whether these weights were preserved in the full 1-year population. Finally, the question-answer pairs of the representative assessments were converted into structured textual case vignettes using controlled prompting of a large language model. Results The cluster analysis yielded 514 included clusters covering 96.8% of the sampled 50,000 assessments. The generated representatives showed strong agreement with the majority treatment-urgency recommendation of their source cluster (Spearman’s ρ=0.78, p<0.001) and contained on average 4.3 more answered triage questions than the original assessments within their clusters. When weighted by cluster size, the representatives approximated the sample distributions of treatment urgency, demographics, and symptoms, although some systematic deviations remained, most notably an overrepresentation of female cases (+13.5%), patients aged 14-49 years (+8.0%), and the urgency category “As soon as possible” (+6.6%). Of 121 recorded symptoms, 101 (83.5%) were covered by the representatives; the rest each occurred in <0.5% of the sample. In a sensitivity analysis, cluster-based vignette weights were strongly correlated with similarity-based population weights (Spearman’s ρ=0.77, p<0.001), and 90.1% of assessments in the full 1-year population were matched to at least one vignette. Conclusions We present a data-driven framework for deriving a manageable set of population-weighted case vignettes from nationwide triage data. The resulting vignettes captured broad presentation diversity, approximated key sample characteristics, and provided an explicit quantitative link to the number of underlying patient assessments. After medical expert review and refinement, the vignettes may support more population-aware evaluation and quality assurance of digital decision-support tools.
BackgroundIn health care, diagnosis codes in claims data and electronic health records (EHRs) play an important role in data-driven decision making. Any analysis that uses a patient’s diagnosis codes to predict future outcomes or describe morbidity requires a numerical representation of this diagnosis profile made up of string-based diagnosis codes. These numerical representations are especially important for machine learning models. Most commonly, binary-encoded representations have been used, usually for a subset of diagnoses. In real-world health care applications, several issues arise: patient profiles show high variability even when the underlying diseases are the same, they may have gaps and not contain all available information, and a large number of appropriate diagnoses must be considered.ObjectiveWe herein present Pat2Vec, a self-supervised machine learning framework inspired by neural network–based natural language processing that embeds complete diagnosis profiles into a small real-valued numerical vector.MethodsBased on German outpatient claims data with diagnosis codes according to the International Statistical Classification of Diseases and Related Health Problems, 10th Revision (ICD-10), we discovered an optimal vectorization embedding model for patient diagnosis profiles with Bayesian optimization for the hyperparameters. The calibration process ensured a robust embedding model for health care–relevant tasks by aggregating the metrics of different regression and classification tasks using different machine learning algorithms (linear and logistic regression as well as gradient-boosted trees). The models were tested against a baseline model that binary encodes the most common diagnoses. The study used diagnosis profiles and supplementary data from more than 10 million patients from 2016 to 2019 and was based on the largest German ambulatory claims data set. To describe subpopulations in health care, we identified clusters (via density-based clustering) and visualized patient vectors in 2D (via dimensionality reduction with uniform manifold approximation). Furthermore, we applied our vectorization model to predict prospective drug prescription costs based on patients’ diagnoses.ResultsOur final models outperform the baseline model (binary encoding) with equal dimensions. They are more robust to missing data and show large performance gains, particularly in lower dimensions, demonstrating the embedding model’s compression of nonlinear information. In the future, other sources of health care data can be integrated into the current diagnosis-based framework. Other researchers can apply our publicly shared embedding model to their own diagnosis data.ConclusionsWe envision a wide range of applications for Pat2Vec that will improve health care quality, including personalized prevention and signal detection in patient surveillance as well as health care resource planning based on subcohorts identified by our data-driven machine learning framework.
Background:Regional deprivation indices enable researchers to analyse associations between socioeconomic disadvantages and health outcomes even if the health data of interest does not include information on the individuals' socioeconomic position. This article introduces the recent revision of the German Index of Socioeconomic Deprivation (GISD) and presents associations with life expectancy as well as age-standardised cardiovascular mortality rates and cancer incidences as applications. Methods:The GISD measures the level of socioeconomic deprivation using administrative data of education, employment, and income situations at the district and municipality level from the INKAR database. The indicators are weighted via principal component analyses. The regional distribution is depicted cartographically, regional level associations with health outcomes are presented. Results:The principal component analysis indicates medium to high correlations of the indicators with the index subdimensions. Correlation analyses show that in districts with the lowest deprivation, the average life expectancy of men is approximately six years longer (up to three years longer for women) than for those from districts with the highest deprivation. A similar social gradient is observed for cardiovascular mortality and lung cancer incidence. Conclusions:The GISD provides a valuable tool to analyse socioeconomic inequalities in health conditions, diseases, and their determinants at the regional level.
Several determinants are suspected to be causal drivers for new cases of COVID-19 infection. Correcting for possible confounders, we estimated the effects of the most prominent determining factors on reported case numbers. To this end, we used a directed acyclic graph (DAG) as a graphical representation of the hypothesized causal effects of the determinants on new reported cases of COVID-19. Based on this, we computed valid adjustment sets of the possible confounding factors. We collected data for Germany from publicly available sources (e.g. Robert Koch Institute, Germany’s National Meteorological Service, Google) for 401 German districts over the period of 15 February to 8 July 2020, and estimated total causal effects based on our DAG analysis by negative binomial regression. Our analysis revealed favorable effects of increasing temperature, increased public mobility for essential shopping (grocery and pharmacy) or within residential areas, and awareness measured by COVID-19 burden, all of them reducing the outcome of newly reported COVID-19 cases. Conversely, we saw adverse effects leading to an increase in new COVID-19 cases for public mobility in retail and recreational areas or workplaces, awareness measured by searches for “corona” in Google, higher rainfall, and some socio-demographic factors. Non-pharmaceutical interventions were found to be effective in reducing case numbers. This comprehensive causal graph analysis of a variety of determinants affecting COVID-19 progression gives strong evidence for the driving forces of mobility, public awareness, and temperature, whose implications need to be taken into account for future decisions regarding pandemic management.
Mobility, awareness, and weather are suspected to be causal drivers for new cases of COVID-19 infection. Correcting for possible confounders, we estimated their causal effects on reported case numbers. To this end, we used a directed acyclic graph (DAG) as a graphical representation of the hypothesized causal effects of the aforementioned determinants on new reported cases of COVID-19. Based on this, we computed valid adjustment sets of the possible confounding factors. We collected data for Germany from publicly available sources (e.g. Robert Koch Institute, Germany’s National Meteorological Service, Google) for 401 German districts over the period of 15 February to 8 July 2020, and estimated total causal effects based on our DAG analysis by negative binomial regression. Our analysis revealed favorable causal effects of increasing temperature, increased public mobility for essential shopping (grocery and pharmacy), and awareness measured by COVID-19 burden, all of them reducing the outcome of newly reported COVID-19 cases. Conversely, we saw adverse effects of public mobility in retail and recreational areas, awareness measured by searches for “corona” in Google, and higher rainfall, leading to an increase in new COVID-19 cases. This comprehensive causal analysis of a variety of determinants affecting COVID-19 progression gives strong evidence for the driving forces of mobility, public awareness, and temperature, whose implications need to be taken into account for future decisions regarding pandemic management.
Der Verlust des Arbeitsplatzes geht mit erheblichen gesundheitlichen Folgen einher, Arbeitslose sind von Depressionen häufiger betroffen als Erwerbstätige. Der Beitrag geht der Frage nach, inwiefern der Zusammenhang zwischen Arbeitslosigkeitserfahrung und Depression durch soziale Unterstützung vermittelt wird. Dazu werden bevölkerungsweite Querschnittsdaten des Zusatzmoduls «Psychische Gesundheit» der Studie zur Gesundheit Erwachsener in Deutschland (DEGS1-MH, 2008–2011) verwendet und depressive Störungen anhand der DSM-IV-Kriterien des psychiatrischen Diagnoseinterviews «Composite International Diagnostic Interview» (DIA-X/M-CIDI) gemessen. Die Fallzahl für multivariate Analysen beträgt n=2.806 im Alter zwischen 18 und 64 Jahren. Frauen und Männer mit Arbeitslosigkeitserfahrung sind etwa doppelt so oft von Depressionen betroffen wie Erwerbstätige ohne Arbeitslosigkeitserfahrung in den letzten fünf Jahren. Der Erklärungsanteil sozialer Unterstützung am Zusammenhang zwischen Arbeitslosigkeitserfahrungen und Depression liegt bei Frauen bei 20,8% (p=0.008), bei Männern bei 15,7% (p=0.140) Die Analysen betonen die Bedeutung sozialer Ressourcen für den Zusammenhang zwischen Arbeitslosigkeit und Depressionen.
Although there is extensive research on health inequalities in childhood and adolescence, it remains unclear to what extent different aspects of social inequality (family-based and non-family, objective and subjective) are linked and shape mental health problems over time. By using directed acyclic graphs, a structured causal model was developed that takes these interrelations into account. Prospective data from the KiGGS cohort (735 boys; 830 girls) were analysed. Mental health problems were recorded with the Strengths and Difficulties Questionnaire. Mental health problems in childhood show strongest effects on those in adolescence. Young people who attend the highest German type of school and who rate their social status higher show fewer mental health problems. Boys show slightly more mental health problems than girls. Overall, the analysis indicates the simultaneous occurrence of causation (social position in childhood influences mental health problems in adolescence) and selection (mental health problems in childhood are relevant for social position in adolescence).
Kapitel des Online Lehrbuch der Medizinischen Psychologie und Medizinischen Soziologie
ObjectivesThis study aimed to investigate associations between occupational physical activity patterns (physical work demands linked to job title) and leisure time physical activity (assessed by questionnaire) with cardiorespiratory fitness (assessed by exercise test) among men and women in the German working population.DesignPopulation-based cross-sectional study.SettingTwo-stage cluster-randomised general population sample selected from population registries of 180 nationally distributed sample points. Information was collected from 2008 to 2011.Participants1296 women and 1199 men aged 18–64 from the resident working population.Outcome measureEstimated low maximal oxygen consumption (V˙O2max), defined as first and second sex-specific quintile, assessed by a standardised, submaximal cycle ergometer test.ResultsLow estimatedV˙O2maxwas strongly linked to low leisure time physical activity, but not occupational physical activity. The association of domain-specific physical activity patterns with lowV˙O2maxvaried by sex: women doing no leisure time physical activity with high occupational physical activity levels were more likely to have lowV˙O2max(OR 6.54; 95% CI 2.98 to 14.3) compared with women with ≥2 hours of leisure time physical activity and high occupational physical activity. Men with no leisure time physical activity and low occupational physical activity had the highest odds of lowV˙O2max(OR 4.37; 95% CI 2.02 to 9.47).ConclusionThere was a strong association between patterns of leisure time and occupational physical activity and cardiorespiratory fitness within the adult working population in Germany. Women doing no leisure time physical activity were likely to have poor cardiorespiratory fitness, especially if they worked in physically demanding jobs. However, further investigation is needed to understand the relationships between activity and fitness in different domains. Current guidelines do not distinguish between activity during work and leisure time, so specifying leisure time recommendations by occupational physical activity level should be considered.
Background High-sensitivity C-reactive protein (hsCRP) is a sensitive biomarker of systemic inflammation and is related to the development and progression of cardiometabolic diseases. Beyond individual-level determinants, characteristics of the residential physical and social environment are increasingly recognized as contextual determinants of systemic inflammation and cardiometabolic risks. Based on a large nationwide sample of adults in Germany, we analyzed the cross-sectional association of hsCRP with residential environment characteristics. We specifically asked whether these associations are observed independent of determinants at the individual level. Methods Data on serum hsCRP levels and individual sociodemographic, behavioral, and anthropometric characteristics were available from the German Health Interview and Examination Survey for Adults (2008-2011). Area-level variables included, firstly, the predefined German Index of Socioeconomic Deprivation (GISD) derived from the INKAR (indicators and maps on spatial and urban development in Germany and Europe) database and, secondly, population- weighted annual average concentration of particulate matter (PM10) in ambient air provided by the German Environment Agency. Associations with log-transformed hsCRP levels were analyzed using random-intercept multi-level linear regression models including 6,768 participants aged 18-79 years nested in 162 municipalities. Results No statistically significant association of PM10 exposure with hsCRP was observed. However, adults residing in municipalities with high compared to those with low social deprivation showed significantly elevated hsCRP levels (change in geometric mean 13.5%, 95% CI 3.2%-24.7%) after adjusting for age and sex. The observed relationship was independent of individual-level educational status. Further adjustment for smoking, sports activity, and abdominal obesity appeared to markedly reduce the association between area-level social deprivation and hsCRP, whereas all individual-level variables contributed significantly to the model. Conclusions Area-level social deprivation is associated with higher systemic inflammation and the potentially mediating role of modifiable risk factors needs further elucidation. Identifying and assessing the source-specific harmful components of ambient air pollution in populationbased studies remains challenging.
Social differences in mortality and life expectancy are a clear demonstration of the social and health-related inequalities that exist within a particular population. According to data from the Socio-Economic Panel (SOEP) for the period ranging from 1992 to 2016, 13% of women and 27% of men in the lowest income group died before the age of 65; the same can be said for just 8% of women and 14% of men in the highest income group. The difference between mean life expectancy at birth among the lowest and highest income groups is 4.4 years for women and 8.6 years for men. Substantial differences also exist between income groups regarding further life expectancy at the age of 65: women in the lowest income group have a 3.7-year shorter life expectancy than women in the highest income group. Similarly, men in the lowest income group have a 6.6-year shorter life expectancy than men in the highest income group. Finally, results from the trend analyses suggest that social differences in life expectancy have remained relatively stable over the last 25 years.
Abstract Background Studies show that occupational physical activity (OPA) has less health-enhancing effects than leisure-time physical activity (LTPA). The spare data available suggests that OPA rarely includes aerobic PAs with little or no enhancing effects on cardiorespiratory fitness (CRF) as a possible explanation. This study aims to investigate the associations between patterns of OPA and LTPA and CRF among adults in Germany. Methods 1,204 men and 1,303 women (18-64 years), who participated in the German Health Interview and Examination Survey 2008-2011, completed a standardized sub-maximal cycle ergometer test to estimate maximal oxygen consumption (VO2max). Job positions were coded according to the level of physical effort to construct an occupational PA index and categorized as low vs. high OPA. LTPA was assessed via questionnaires and dichotomized in no vs. any LTPA participation. A combined LTPA/OPA variable was used (high OPA/ LTPA, low OPA/LTPA, high OPA/no LTPA, low OPA/no LTPA). Information on potential confounders was obtained via questionnaires (e.g., smoking and education) or physical measurements (e.g., waist circumference). Multi-variable logistic regression was used to analyze associations between OPA/LTPA patterns and VO2max. Results Preliminary analyses showed that less-active men were more likely to have a low VO2max with odds ratios (ORs) of 0.80 for low OPA/LTPA, 1.84 for high OPA/no LTPA and 3.46 for low OPA/no LTPA compared to high OPA/LTPA. The corresponding ORs for women were 1.11 for low OPA/LTPA, 3.99 for high OPA/no LTPA and 2.44 for low OPA/no LTPA, indicating the highest likelihood of low fitness for women working in physically demanding jobs and not engaging in LTPA. Conclusions Findings confirm a strong association between LTPA and CRF and suggest an interaction between OPA and LTPA patterns on CRF within the workforce in Germany. Women without LTPA are at high risk of having a low CRF, especially if they work in physically demanding jobs. Key messages Women not practicing leisure-time physical activity are at risk of having a low cardiorespiratory fitness, especially if they work in physically demanding jobs. Different impact of domains of physical activity should be considered when planning interventions to enhance fitness among the adult population.
Ziel des Beitrags ist es, sozioökonomische Ungleichheiten in der Lungenkrebsinzidenz in Deutschland zu analysieren und abzuschätzen, welchen Erklärungsbeitrag das Tabakrauchen leistet.
Objective: Despite extensive study of the obesity epidemic, research on whether obesity has risen faster in lower or in higher socioeconomic groups is inconsistent. This study examined secular trends in obesity prevalence by socioeconomic position and the resulting obesity inequalities in the German adult population. Methods: Data were drawn from three national examination surveys conducted in 1990–1992, 1997–1999 and 2008–2011 (n = 18,541; age range: 25–69 years). Obesity was defined by a body mass index ≥30 kg/m2 using standardised measurements of body height and weight. Education and equivalised household disposable income were used as indicators of socioeconomic position. Time trends in socioeconomic inequalities in obesity were examined using linear probability and log-binomial regression models. Results: In each survey period, the highest socioeconomic groups had the lowest prevalence of obesity. The low and medium socioeconomic groups showed increases in obesity prevalence, whereas no such trend was observed in the high socioeconomic groups. Absolute inequalities in obesity by income increased by an average of 0.53 percentage points per year (95% confidence interval [CI] 0.01–1.05, p = 0.047) among men and 0.47 percentage points per year (95% CI 0.05–0.90, p = 0.029) among women. Absolute inequalities in obesity by education increased on average by 0.64 percentage points per year (95% CI 0.19–1.08, p = 0.005) among women but not among men (0.33 percentage points, 95% CI –0.27 to 0.92, p = 0.283). Conclusions: These findings suggest a widening obesity gap between the top and the bottom of the socioeconomic spectrum. This has the potential to have adverse consequences for population health and health inequalities in coming decades. Interventions that are effective in preventing and reducing obesity in socially disadvantaged groups are needed.
This study aimed at estimating the prevalence in adults of complying with the aerobic physical activity (PA) recommendation through transportation-related walking and cycling. Furthermore, potential determinants of transportation-related PA recommendation compliance were investigated. 10,872 men and 13,144 women aged 18 years or older participated in the cross-sectional ‘German Health Update 2014/15 – EHIS’ in Germany. Transportation-related walking and cycling were assessed using the European Health Interview Survey-Physical Activity Questionnaire. Three outcome indicators were constructed: walking, cycling, and total active transportation (≥600 metabolic equivalent, MET-min/week). Associations were analyzed using multilevel regression analysis. Forty-two percent of men and 39% of women achieved ≥600 MET-min/week with total active transportation. The corresponding percentages for walking were 27% and 28% and for cycling 17% and 13%, respectively. Higher population density, older age, lower income, higher work-related and leisure-time PA, not being obese, and better self-perceived health were positively associated with transportation-related walking and cycling and total active transportation among both men and women. The promotion of walking and cycling among inactive people has great potential to increase PA in the general adult population and to comply with PA recommendations. Several correlates of active transportation were identified which should be considered when planning public health policies and interventions.
[This corrects the article on p. 98-114 in vol. 2.].