DNA foundation models have become transformative tools in bioinformatics and healthcare applications. Trained on vast genomic datasets, these models can be used to generate sequence embeddings, dense vector representations that capture complex genomic information. These embeddings are increasingly being shared via Embeddings-as-a-Service (EaaS) frameworks to facilitate downstream tasks, while supposedly protecting the privacy of the underlying raw sequences. However, as this practice becomes more prevalent, the security of these representations is being called into question. This study evaluates the resilience of DNA foundation models to model inversion attacks, whereby adversaries attempt to reconstruct sensitive training data from model outputs. In our study, the model's output for reconstructing the DNA sequence is a zero-shot embedding, which is then fed to a decoder. We evaluated the privacy of three DNA foundation models: DNABERT-2, Evo 2, and Nucleotide Transformer v2 (NTv2). Our results show that per-token embeddings allow near-perfect sequence reconstruction across all models. For mean-pooled embeddings, reconstruction quality degrades as sequence length increases, though it remains substantially above random baselines. Evo 2 and NTv2 prove to be most vulnerable, especially for shorter sequences with reconstruction similarities > 90
BACKGROUND:Missense variants in genes encoding GABAA receptors are involved in the pathophysiology of common and rare epilepsies. Variant effects on channel biophysical function are associated with key clinical characteristics and treatment response. Predicting variant effects is therefore key to improving care for individuals with GABAA receptor-related disorders. METHODS:We collected data from 505 affected individuals with 272 (likely) pathogenic GABAA receptor missense variants (GABRA1, GABRB2, GABRB3, GABRG2). All variants were evaluated with in-vitro electrophysiology. Variants were annotated with features based on sequence, structure, and phenotype. Model performance was estimated using cross-validation and external validation on a further 197 individuals with 138 (likely) pathogenic variants. FINDINGS:Our models enable highly accurate prediction of missense variant effects in GABAA (AU-ROC 0.862-0.946), outperforming state-of-the-art models (AU-ROC 0.495-0.756) and clinical decision-making. Model scores correlated with GABA sensitivity and were consistent with expert-based structure-function hypotheses, supporting plausibility. Predictions on population variants were similar to functionally neutral variants, while cases from ClinVar were similar to GOF/LOF variants. Our model may provide additional evidence for 5-25% of variants in ClinVar. Lastly, we show that we can predict likely clinical characteristics from variant information alone (median Lin similarity 0.754 IQR 0.161). INTERPRETATION:We demonstrate accurate missense variant effect prediction in GABAA receptors with rigorous validation across a large dataset of functionally tested variants. These predictions may facilitate timely diagnosis and precision treatment of individuals with GABAA receptor-related disorders, pending prospective clinical validation. A web interface, precomputed scores, and calibrated score thresholds for all possible variants are openly available. FUNDING:Else Kröner-Fresenius-Stiftung; German Federal Ministry of Research, Technology and Space; German Research Foundation; Medical Faculty University of Tübingen; Lundbeck Foundation; Novo Nordisk Foundation.
Abstract Motivation Artificial Neural Networks (ANNs) hold promise in predicting disease severity from viral protein sequences. To gain clinical insights, the models must adhere to two key factors: first, they need to be interpretable, as using black-box models within a healthcare setting poses risks, and second, they should integrate viral and clinical features to correct for biases within the data. We propose ClinIAN (Clinically Informed Attention Network), an inherently interpretable end-to-end learning model that effectively combines clinical and sequence parameters. We evaluate ClinIAN within the context of the challenging task of predicting the severity of SARS-CoV-2 infection, but it could be applied to other medical and biological contexts, with sequence and tabular features, such as bacterial infections or cancer research. Results We evaluate the performance capabilities of ClinIAN in predicting SARS-CoV-2 severity using a subset of the EuCARE hospitalized cohort. We show ClinIAN’s ability to capture biologically relevant features across multiple layers of resolution while retaining stable state-of-the-art performance. Availability and implementation The code is available on Zenodo and on GitHub. Supplementary Information Supplementary materials are published at Bioinformatics Advances online.
Meta-research and Trustworthy AI (TAI) share common goals, namely improving evidence, robustness, and transparency, yet there is very little interplay between the two fields. To investigate the potential benefits of closer collaboration between the domains of TAI in healthcare and meta-research, we convened an interdisciplinary workshop funded by the Volkswagen Foundation in February 2025. The workshop aimed to collaboratively examine key challenges in translating AI ethics principles into practice and to identify potential solutions informed by meta-research approaches. A Design Thinking-informed co-creation approach was followed by an inductive descriptive analysis of the outputs. Our results demonstrate how meta-research can offer concrete contributions to address pressing challenges of TAI in healthcare. These challenges include the dynamic and complex nature of TAI ethical requirements and principles, common terminology and understanding of TAI, ensuring robustness, replicability, and reproducibility, choosing adequate evaluation metrics, lack of transparency, advancing preclinical biomedical research, and validation in real-world clinical environments. We present a catalog of ideas and a roadmap for future research, which synthesize existing interconnections and identify concrete next steps and open research gaps, thereby serving as a foundation for future interdisciplinary efforts.
Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN v2.5 has the highest point estimate, followed by Logistic Regression and TabDPT, whose point estimates are nearly identical. A paired bootstrap over the target pool separates RealTabPFN v2.5 from Logistic Regression by 145 Elo (95 We invite the community to contribute: TabBench-Bio is designed to grow, and we welcome submissions of new biomedical tabular datasets, particularly from underrepresented assays and clinical endpoints, for inclusion in future releases. The interactive leaderboard is available at: https://tabbench-bio.eu
Revealing novel insights from the relationship between molecular measurements and pathology remains a very impactful application of machine learning in biomedicine. Data in this domain typically contain only a few observations but thousands of potentially noisy features, posing challenges for conventional tabular machine learning approaches. While prior-data fitted networks emerge as foundation models for predictive tabular data tasks, they are currently not suited to handle large feature counts (>500). Although feature reduction enables their application, it hinders feature importance analysis. We propose a strategy that extends existing models through continued pre-training on synthetic data sampled from a customized prior. The resulting model, TabPFN-Wide, matches or exceeds its base model's performance, while exhibiting improved robustness to noise. It seamlessly scales beyond 30,000 categorical and continuous features, regardless of noise levels, while maintaining inherent interpretability, which is critical for biomedical applications. Our results demonstrate that prior-informed adaptation is suitable to enhance the capability of foundation models for high-dimensional data. On real-world omics datasets, we show that many of the most relevant features identified by the model overlap with previous biological findings, while others propose potential starting points for future studies.
Motivation:The emergence of multidrug class resistance (MDR) in Human Immunodeficiency Virus (HIV) is a rare but significant challenge in antiretroviral therapy (ART). MDR, which may arise from prolonged drug exposure, treatment failures, or transmission of resistant strains, accelerates disease progression and poses particular challenges in resource-limited settings with restricted access to resistance testing and advanced therapies. Early prediction of future MDR development is important to inform therapeutic decisions and mitigate its occurrence. Results:In this study, we employ various machine learning classifiers to predict future resistance to all four major antiretroviral drug classes using features extracted from clinical HIV sequence data. We systematically explore several variations of the problem that differ in the pre-existing resistance level and the temporal gap between sample collection and observed MDR occurrence. Our models show the ability to predict multidrug class resistance even in the most challenging variations, albeit at a reduced accuracy. Feature importance analysis reveals that our models primarily utilize known drug resistance mutations for easier classification tasks, but rely on new mutations for the difficult task of distinguishing four class drug resistance from three class drug resistance. Availability and implementation:All analysis was performed using the Euresist Integrated DataBase (EIDB). Researchers wishing to reproduce, validate or extend these findings can request access to the latest EIDB release via the Euresist Network.
BACKGROUND:Respiratory viral diseases are one of the greatest challenges facing our healthcare system, with them being one of the main causes of death. This has been demonstrated once again by the impact of the SARS-CoV-2 pandemic in recent years. We study the impact of the SARS-CoV-2 pandemic on the prevalence of respiratory viruses by analysing a subset of the Clinical Virology network database, covering 2,216,198 samples tested for 18 different viral pathogens in the time span from 2010 to 2024. METHODS:We calculated the prevalence of 17 respiratory viruses before and after onset of the SARS-CoV-2 pandemic and compared the degree of seasonality shift with a newly developed a metric dubbed seasonal disruption index. In addition, we compared coinfection statistics prior to and after the pandemic onset, and also studied the correlation of infection counts with non-pharmaceutical interventions in the time frame from early 2020 to end of 2022. RESULTS:We found that the viral pathogens show a varying degree of seasonality disruption. It is largest among those that are known to show a highly seasonal behavior, namely Influenza and RSV, the latter having the highest seasonal disruption index. Most perennial viruses continued to appear throughout the year. Coinfections occurred before and after the pandemic; patterns before and after pandemic onset are surprisingly similar. The occurrence of most viruses is nonlinearly correlated with the degree of non-pharmaceutical interventions. CONCLUSION:The SARS-CoV-2 pandemic had a considerable impact on the occurrence and seasonality of other respiratory viruses. While nearly all seasonality patterns were initially disrupted due to the heavy non-pharmaceutical interventions, viruses are regaining their pre-pandemic seasonality.
Ensuring privacy in distributed machine learning while computing the Area Under the Curve (AUC) is a significant challenge because pooling sensitive test data is often not allowed. Although cryptographic methods can address some of these concerns, they may compromise either scalability or accuracy. In this paper, we present two privacy-preserving solutions for secure AUC computation across multiple institutions: (1) an exact global AUC method that handles ties in prediction scores and scales linearly with the number of samples, and (2) an approximation method that substantially reduces runtime while maintaining acceptable accuracy. Our protocols leverage a combination of homomorphic encryption (modified Paillier), symmetric and asymmetric cryptography, and randomized encoding to preserve the confidentiality of true labels and model predictions. We integrate these methods into the Personal Health Train (PHT)-meDIC platform, a distributed machine learning environment designed for healthcare, to demonstrate their correctness and feasibility. Results using both real-world and synthetic datasets confirm the accuracy of our approach: the exact method computes the true AUC without revealing private inputs, and the approximation provides a balanced trade-off between computational efficiency and precision. All relevant code and data is publicly available at https://github.com/PHT-meDIC/PP-AUC, facilitating straightforward adoption and further development within broader distributed learning ecosystems.
Platelet reconstitution after allogeneic hematopoietic cell transplantation (allo-HCT) is heterogeneous and influenced by various patient- and transplantation-related factors, associated with poor prognoses for poor graft function (PGF) and isolated thrombocytopenia. Tailored interventions could improve the outcome of patients with PGF and post-HCT thrombocytopenia. To provide individual predictions of 180-day platelet counts from early phase data, we developed a model of long-term platelet reconstitution after allo-HCT. A large cohort (n = 1949) of adult patients undergoing their first allo-HCT was included. Real-world data from 1,048 retrospective patients were used for non-linear mixed-effects model development. Bayesian forecasting was used to predict platelet-time profiles for 518 retrospective and 383 prospective patients during internal and external model validation, respectively. Thrombocytopenia was defined as mean platelet count < 75 × 109/L, derived from the last 12 platelet measurements within the first 180 days post-HCT. Thrombocytopenia affected 37% of all patients and was associated with significantly reduced overall survival (P-value < 0.0001). On days +7, +14, +21, and +28, the developed model achieved areas under the receiver-operating characteristic of ≥ 0.68, ≥ 0.75, ≥ 0.78, and 0.81 for the prediction of post-HCT thrombocytopenia, respectively, with anti-thymocyte globulin, donor relation, and total protein measurements representing prognostic markers for post-HCT platelet kinetics. A publicly accessible web-based demonstrator of the model was established (https://hsct.precisiondosing.de). In summary, the developed model predicts individual platelet counts from day +28 post-HCT adequately, utilizing internal and external datasets. The web-based demonstrator provides a basis to implement model-based predictions in clinical practice and to confirm these findings in future clinical studies.
Several domains increasingly rely on machine learning in their applications. The resulting heavy dependence on data has led to the emergence of various laws and regulations around data ethics and privacy and growing awareness of the need for privacy-preserving machine learning (ppML). Current ppML techniques utilize methods that are either purely based on cryptography, such as homomorphic encryption, or that introduce noise into the input, such as differential privacy. The main criticism given to those techniques is the fact that they either are too slow or they trade off a model's performance for improved confidentiality. To address this performance reduction, we aim to leverage robust representation learning as a way of encoding our data while optimising the privacy-utility trade-off. Our method centers on training autoencoders in a multi-objective manner and then concatenating the latent and learned features from the encoding part as the encoded form of our data. Such a deep learning-powered encoding can then safely be sent to a third party for intensive training and hyperparameter tuning. With our proposed framework, we can share our data and use third party tools without being under the threat of revealing its original form. We empirically validate our results on unimodal and multimodal settings, the latter following a vertical splitting system and show improved performance over state-of-the-art.
MOTIVATION:Generalizing machine learning models across small, high-dimensional, and heterogeneous biological datasets remains a critical challenge due to domain shifts caused by variations in data collection, population differences, and privacy constraints that restrict data sharing. Existing federated domain adaptation (FDA) approaches primarily rely on deep learning and focus on classification tasks, making them unsuitable for privacy-sensitive, small-scale regression problems in biomedical research. We introduce a privacy-preserving federated method for unsupervised domain adaptation in regression, enabling robust learning across distributed, high-dimensional datasets while maintaining full data privacy. RESULTS:Our method is the first to enable distributed training of Gaussian processes for domain adaptation, ensuring complete privacy through randomized encoding and secure aggregation. Unlike deep learning-based FDA approaches, our method is specifically designed for small-scale, high-dimensional biological data, overcoming prior limitations in scalability and generalization. We evaluate our approach on age prediction from DNA methylation data, demonstrating that it achieves performance comparable to non-private state-of-the-art methods while fully preserving data privacy. This work enables secure and effective cross-institutional collaboration in biomedical research without requiring raw data sharing. AVAILABILITY AND IMPLEMENTATION:The source code for our method is available at https://github.com/mdppml/FREDA.
ABSTRACTHepatitis C virus infection is a significant global health concern, affecting millions worldwide. Although direct‐acting antivirals achieve over 90% success rate, treatment failures still occur, particularly when pan‐genotypic DAAs are unavailable, and drugs need to be chosen based on the present HCV genotype. Genotyping tests can be misleading, especially in cases involving the 2k/1b recombinant variant. The 2k/1b variant was first discovered in Saint Petersburg in 2002 and is most commonly observed in Eastern European countries, including Russia, Georgia, and Ukraine. Due to migration, the 2k/1b variant has spread to Western Europe and other regions, potentially increasing HCV transmission and changing the virus's epidemiological landscape. The situation highlights the importance of molecular epidemiology in monitoring the spread of the 2k/1b variant. Accurate detection and characterization of the 2k/1b variant are crucial for an effective treatment if no pan‐genotypic DAAs are available. To address this need, machine learning models were developed to predict the 2k/1b variant based on 1b and 2k/1b sequence data from nonstructural proteins. They were integrated into the geno2pheno[HCV] tool, providing physicians and researchers with an open‐access resource for determining HCV genotypes, including the 2k/1b variant.
BACKGROUND:Limited evidence exists on how bacterial and viral coinfections have developed since the SARS-CoV-2 Omicron variant emerged. We investigated whether community-onset coinfections in adult patients hospitalized with COVID-19 differed during the wild type, Alpha, Delta, and Omicron periods and whether such coinfections were associated with an increased risk of mortality. METHODS:We conducted a multinational cohort study including COVID-19 hospitalizations until 30 April 2023 in 5 European countries. The outcome was bacterial and viral coinfections based on 5 test modalities. Variant periods were compared with regard to occurrences of coinfections and risk ratios for coinfections (Omicron vs pre-Omicron), as well as association with in-hospital mortality (Omicron vs pre-Omicron). RESULTS:A total of 29 564 cases were included: 12 601 wild type, 5256 Alpha, 2433 Delta, and 9274 Omicron. The coinfection rate was 2.6% (327/12 601) for wild type, 2.0% (105/5256) for Alpha, 3.2% (77/2433) for Delta, and 7.9% (737/9274) for Omicron. Omicron had a significantly increased risk ratio of coinfection when compared with preceding variants (1.88 [95% CI, 1.53-2.32], P < .001). These results were consistent across several subgroup analyses. An increased occurrence (19% [232/1246] vs 11% [3042/28 318]) and adjusted risk (1.69 [95% CI, 1.49-1.91], P < .001) of in-hospital mortality were observed in patients with a verified coinfection as compared with patients without a coinfection. CONCLUSIONS:Bacterial and viral coinfections were more prevalent during the Omicron period as compared with preceding variants. Such coinfections were associated with an increased risk of in-hospital mortality, calling for sustained monitoring and clinical vigilance.
Methods: Patients with breast cancer were eligible to participate in the study before neoadjuvant, adjuvant, postneoadjuvant, or palliative systemic therapy against breast cancer was initiated at the Heidelberg, Mannheim, and Tuebingen, Germany, university hospitals. After 1:1 randomization into an intervention and a control group, HRQoL assessments were performed at six fixed time points during the therapy using validated questionnaires. In the intervention group, HRQoL was also assessed briefly every week using a visual analog scale (EQ-VAS). In cases of significant deterioration, therapy-associated side effects were assessed in a graduated manner, recommendations were sent to the patient, and the treatment team was informed. Additionally, the app served as an "eHealth companion" for education, training, and organizational support during therapy. Results: Recruitment started in March 2021; follow-up was completed in February 2024. In total, 606 patients were enrolled, and 592 patients participated in the study. Enrollment was completed in September 2023, and the last visit was in February 2024. The first results are expected to be published in Q2 2025. Conclusions: Participation in the intervention group is expected to improve treatment satisfaction, adherence, detection, and timely treatment of critical AEs. The close-meshed, weekly, brief HRQoL assessment will also be tested as a screening tool to detect relevant side effects during therapy. The study offers a more objective HRQoL assessment across treatment strategies. Trial Registration: Deutsches Register Klinischer Studien DRKS00025611; https://drks.de/search/en/trial/DRKS00025611 International Registered Report Identifier (IRRID): DERR1-10.2196/69855
The Omicron (B.1.1.529) variant of SARS-CoV-2 emerged in November 2021 and has since evolved into multiple lineages. Understanding its transmission, vaccine efficacy, and potential for reinfection is crucial. This study examines the dynamics of Omicron in Germany, France, and Italy by employing Physics-Informed Neural Networks to estimate the temporal parameters influencing its spread. We validated the performance of our model using the Root Mean Squared Percent Error (RMSPE). Our analysis revealed significant correlations between specific viral mutations-S371F, T376A, D405N, and R408S-and increased transmission rates in all three countries. These mutations, prevalent in the Omicron BA.2 and BA.3 sublineages, are linked to immune evasion and heightened transmissibility.
Genome-wide association studies help uncover genetic influences on complex traits and diseases. Importantly, multi-site data collaborations enhance the statistical power of these studies but pose challenges due to the sensitivity of genomic data. Existing privacy-preserving approaches to performing multi-site genome-wide association studies rely on computationally expensive cryptographic techniques, which limit applicability. To address this, we present PP-GWAS, a privacy-preserving algorithm that improves efficiency and scalability while maintaining data privacy. Our method leverages randomized encoding within a distributed framework to perform stacked ridge regression on a linear mixed model, enabling robust analysis of quantitative phenotypes. We show experimentally using real-world and synthetic data that our approach achieves twice the computational speed of comparable methods while reducing resource consumption.
GABA A receptors are critical for inhibitory neurotransmission. Variants in genes encoding these receptors are involved in the pathophysiology of both common and rare epilepsy syndromes. Variant effects on channel biophysical function, broadly classified as gain-of-function (GOF) or loss-of-function (LOF), are associated with key clinical characteristics and treatment response. Understanding and predicting variant effects is therefore essential to improve care for individuals with GABA A -related disorders. Here, we present GABA A receptor functional variant effect prediction using multi-task phenotypic learning (GENTLY). We collected clinical data from 505 affected individuals with 272 (likely) pathogenic GABA A receptor variants across GABRA1, GABRB2, GABRB3 , and GABRG2 . All variants were evaluated with in-vitro electrophysiology using receptor assemblies that reflect heteropentamer composition in heterozygous carriers. Variants were annotated with features based on sequence (e.g. physicochemical properties, conservation), structure (e.g. binding sites, domains), and phenotypes represented by 8185 HPO terms. We trained separate models on all features and without clinical features. Model performance was estimated using ablation, cross-validation, and external validation on a further 197 individuals with 138 (likely) pathogenic GABA A receptors variants. Our models enable highly accurate prediction of GOF/LOF in GABA A (AU-ROC 0.863-0.946), outperforming state-of-the-art genome-wide predictors (LoGoFunc: AU-ROC 0.495; evo2: AU-ROC 0.559-0.755) and clinical decision-making (decision tree: AU-ROC 0.823). Model scores correlated strongly with GABA sensitivity ( r = -0.77, p < 0.001). Predictions were consistent with expert-based structure-function hypotheses: variants located in transmembrane domains were more likely GOF ( p < 0.001), and variants in GABA binding sites were more likely LOF ( p < 0.001). Predictions on variants from population databases behaved as expected: 13,389 population variants were similar to functionally neutral variants, and predictions from (likely) pathogenic ClinVar variants were similar to GOF/LOF variants. Our model may provide additional evidence for 10-29% of 2,295 variants in ClinVar. Lastly, we show that a simple k -nearest neighbour algorithm can predict likely clinical characteristics only from variant information (median Lin similarity 0.754 IQR 0.161). We demonstrate accurate variant effect prediction in GABA A receptors with rigorous validation across the largest dataset of functionally tested variants to date. Our predictions correlate with continuous electrophysiological measurements not directly used during training and conform to known structure-function relationships, supporting their biological plausibility. These predictions may facilitate timely diagnosis, precision treatment, and prognosis of individuals with GABA A receptor related disorders. A web interface, precomputed scores, and ACMG-calibrated score thresholds for all possible variants are openly available.
Thomas Lengauer合作论文数Max-Planck-Institut fur Informatik26