Lung transplantation is a life-saving therapy for end-stage lung disease but has the poorest survival among solid organ transplants. We analyzed standardized electronic health record (EHR) data from the United Network for Organ Sharing (UNOS) to predict one-, three-, and five-year survival and favorable long-term outcomes post-lung transplant. We applied two multivariate machine learning approaches, XGBoost or a tabular BERT model called EHRFormer, to data from 43,869 first-time lung transplant recipients (1987–2022). XGBoost and EHRFormer identified features that align closely with established risk factors for worse outcomes such as length of index stay, recipient age, and creatinine at the time of transplant. We developed a simple perturbation method with EHRFormer to probe in silico multivariate interactions between features that influence model prediction. Despite their attention to known risk factors, machine learning applied to EHR data collected by UNOS poorly predict one-, three-, and five-year mortality after lung transplant.
Abstract Patient cohort profiling increasingly includes structured views for multiple modalities, such as single-cell RNA sequencing, spatial transcriptomics or proteomics, and histology, each providing multiple subobservations per patient, including single cells, spatial spots or patches. To model such data along with simple patient-level views, current multimodal integration methods typically rely on separately precomputed summaries and fail to fully leverage information in structured views. Here we present FACTMx, a variational framework that jointly models structured and simple views to learn interpretable patient-level representations. FACTMx couples latent patient factors with subobservation clustering and per-patient component proportions, enabling direct interpretation and downstream association analyses. The framework supports different structured-view mixture assumptions, including topic- and Gaussian-structured data, while retaining modular encoder-decoder parameterisations. In simulations spanning sparse and dense dependencies and multiple noise regimes, FACTMx improved reconstruction, integration and recovery of structured components relative to previous methods. Applied to non-small cell lung cancer cohorts, FACTMx captured survival-associated latent signals linked to immune microenvironments, gene expression pathways and spatially coherent histological patterns. In a longitudinal coronary syndrome cohort, FACTMx highlighted an outcome-associated axis connected to ejection-fraction change, immune cell states, soluble mediators and cardiac injury markers. These results support joint structured-simple modelling for interpretable multimodal patient stratification.
We present ImmuVis, a family of efficient foundation models for imaging mass cytometry (IMC), a high-throughput multiplex imaging technology that handles molecular marker measurements as image channels and enables large-scale spatial tissue profiling. Unlike natural images, multiplex imaging lacks a fixed channel space, as real-world marker sets vary across studies, violating a core assumption of standard vision backbones. To address this, ImmuVis introduces marker-adaptive hyperconvolutions that generate convolutional kernels from learned marker embeddings, enabling a single model to operate on arbitrary measured marker subsets without retraining. We pretrain ImmuVis on the largest dataset to date, IMC17M (28 cohorts, 24,405 images, 265 markers, over 17M patches), using self-supervised masked reconstruction. ImmuVis outperforms state-of-the-art baselines and ablations in virtual staining and downstream classification tasks at substantially lower compute cost than transformer-based alternatives, and is the sole model that provides calibrated uncertainty via a heteroscedastic likelihood objective. These results position ImmuVis as a practical framework for real-world IMC modeling.
Abstract The human gut microbiome is a powerful indicator of host health, yet its compositional nature, high sparsity, and inter-individual variability complicate downstream analysis. Here, we introduce two complementary approaches to characterize gut microbiome structure at population scale. First, we define eight functional signatures of the human gut microbiome using Non-negative Matrix Factorization, revealing coordinated metabolic patterns that partially decouple from taxonomic composition. Second, we present GUT-FORMer, a transformer-based autoencoder that jointly models taxonomic and functional metagenomic profiles from close to 21,000 publicly available samples. The learned latent representations capture biologically meaningful structure, reflect geographic and disease-associated variation, and enable accurate classification of 25 diseases in both binary and multiclass settings, as well as regression of host age and BMI. GUT-FORMer outperforms existing microbiome indices and deep learning methods across all tasks, establishing a generalizable framework for microbiome-based precision medicine.
In cancer research, multiplexed imaging allows detailed characterization of the tumor microenvironment (TME) and its link to patient prognosis. The integrated immunoprofiling of large adaptive cancer patient cohorts (IMMUcan) consortium collects multi-modal imaging data from thousands of patients with cancer to perform broad molecular and cellular spatial profiling. Here, we describe and compare two workflows for multiplexed immunofluorescence (mIF) and imaging mass cytometry (IMC) developed within IMMUcan to enable the generation of standardized data for cancer tissue analysis. The IFQuant software supports web-based, user-friendly, and reproducible analysis of mIF data. High sample throughput for IMC is achieved by optimizing experimental protocols, developing a robotic arm for automated slide loading, and classification-based cell typing. Using our manually labeled single-cell data, we show that tree-based methods outperform other cell-phenotyping tools. These pipelines form the basis for multiplexed image analysis within IMMUcan, and we summarize our learnings from 5 years of development and optimization.
Discovering cellular neighborhoods and their roles in disease requires computational methods that consider the full breadth of data and offer multi-level granularity and interpretability. Here, we introduce Cellohood, a permutation-invariant set transformer auto-encoder equipped with a clinical association pipeline. Cellohood encodes full readouts for bags of spatially co-localized cells and supports multi-level analyses, providing interpretability by mapping latent dimensions to spatial tissue features. Our model surpasses current methods in accuracy of cellular neighborhood detection for spatial transcriptomics of the human cortex and CODEX spleen data measuring lupus progression. Applied to cancer data across three granularity levels, our method recovers neighborhoods along the expected immune-cold to immune-hot spectrum and further refines them into biologically and clinically meaningful subclasses, revealing spatial patterns linked to prognosis, histology, stage, and tumor mutational burden, and uncovering subgroups that transcend standard classifications. Overall, Cellohood enables in-depth analysis of complex tissues, revealing clinically informative spatial neighborhoods. ### Competing Interest Statement Projects at Szczurek lab at the University of Warsaw are co-founded by Merck Healthcare. Innovative Medicines Initiative 2 Joint Undertaking, 821558 Polish National Science Center, 2020/38/E/NZ2/00305
Non-small cell lung cancer (NSCLC) is one of the leading causes of cancer-related death despite the availability of therapies targeting the tumor microenvironment (TME) and/or the tumor. Immunotherapy and novel targeted therapy have improved the management of NSCLC, but the heterogeneity of the disease still hampers treatment efficacy as well as further therapeutic advances. Here, we used spatial single-cell data complemented with bulk RNAseq and whole exome sequencing to investigate the TME of 192 resected, mostly early-stage NSCLC tumors and to characterize associations of genetic, clinical and lifestyle factors with cellular and histological properties of the TME. We found that the TME of squamous cell carcinoma (LUSC) harbored stronger signs of inflammation and immune exhaustion than adenocarcinoma (LUAD), but interestingly that elevated PD-L1 expression was associated with inflammation and markers of lymphocyte activation/exhaustion only in LUAD. Smoking correlated with T cell infiltration and TP53 mutation in LUAD, and TP53 mutation was associated with proliferation of tumor cells. In both histologies, naive and BCL2+ T cells decreased with higher clinical stages. EGFR-driven tumors showed fewer proliferating, activated and exhausted T cells, but more CD4 T cells and HLA-DR+ tumor cells, than tumors with wild type EGFR. Multi-modal integration of our four data types showed that most variation in our cohort was along the axes of histology, smoking, TP53 mutation, and inflammation. Our integration identified histology-specific prognostic signatures, with a LUSC-enriched profile of proliferation, inflammation, and smoking associated with poor prognosis in LUAD. In summary, we provide a rich, multimodal, single-cell characterization of a large NSCLC cohort as a resource and suggest that future investigations of biomarkers for ICI in NSCLC will benefit from stratification by histology. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement The project was funded by the Innovative Medicines Initiative 2 Joint Undertaking under grant agreement No 821558. This Joint Undertaking receives support from the European Union's Horizon 2020 research and innovation program and EFPIA https://IMI.europa.eu. The SPECTA Platform is supported by Alliance Healthcare. Alliance Healthcare will become Cencora ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The SPECTAlung study was approved by two Ethische Commissie Onderzoek UZ/KU Leuven, in Belgium (S57513) and the Comite de protection des personnes "Ile-de-France VII", in France (15-030 (PP 15-001)) I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data will be available upon peer-review publication
AbstractBackgroundLung transplantation is the only life-saving therapy for end-stage lung disease. However, lung transplantation has the worst survival among all solid organ transplants.1We applied machine learning to a large standardized electronic health record (EHR) dataset from the United Network for Organ Sharing (UNOS) to test whether pre- transplant and peri-transplant donor and recipient features can predict one-, three- and five-year survival, or favorable long-term outcomes in lung transplant.MethodsWe used data from 43,869 first time lung transplant recipients >18 years old from 1987 to November 2022 for whom one-, three-, and five-year survival outcomes were available. We applied XGBoost or a tabular BERT model called EHRFormer to the UNOS EHR dataset.ResultsUsing pre-transplant features XGBoost predicted one year mortality with a test AUC = 0.6 [0.57, 0.64] 95% CI. Addition of peri-transplant features only modestly improved AUC for one-year mortality prediction (test AUC = 0.63 [0.60, 0.67] 95% CI and 0.64 [0.63, 0.66] 95% CI for XGBoost and EHRFormer, respectively). Top predictive features of one year mortality using peri-transplant features from each model were length of index stay, transplant type, recipient age, ventilation status during the index stay, and creatinine at the time of transplant. Both XGBoost and EHRFormer performed better when predicting lung function at one-year post-transplant (XGBoost test AUC = 0.74; EHRFormer test AUC = 0.76). Both models identified and used features previously associated with transplant outcomes to inform predictions.ConclusionsDespite machine learning approaches identifying known risk factors for transplant outcomes, EHR data collected by UNOS poorly predict one-, three-, and five-year mortality outcomes of lung transplantation. These results suggest caution when using pre-transplant EHR features to predict lung transplant outcomes.
Antimicrobial peptides emerge as compounds that can alleviate the global health hazard of antimicrobial resistance, prompting a need for novel computational approaches to peptide generation. Here, we propose HydrAMP, a conditional variational autoencoder that learns lower-dimensional, continuous representation of peptides and captures their antimicrobial properties. The model disentangles the learnt representation of a peptide from its antimicrobial conditions and leverages parameter-controlled creativity. HydrAMP is the first model that is directly optimized for diverse tasks, including unconstrained and analogue generation and outperforms other approaches in these tasks. An additional preselection procedure based on ranking of generated peptides and molecular dynamics simulations increases experimental validation rate. Wet-lab experiments on five bacterial strains confirm high activity of nine peptides generated as analogues of clinically relevant prototypes, as well as six analogues of an inactive peptide. HydrAMP enables generation of diverse and potent peptides, making a step towards resolving the antimicrobial resistance crisis.
Recent emergence of high-throughput drug screening assays sparkled an intensive development of machine learning methods, including models for prediction of sensitivity of cancer cell lines to anti-cancer drugs, as well as methods for generation of potential drug candidates. However, a concept of generation of compounds with specific properties and simultaneous modeling of their efficacy against cancer cell lines has not been comprehensively explored. To address this need, we present VADEERS, a Variational Autoencoder-based Drug Efficacy Estimation Recommender System. The generation of compounds is performed by a novel variational autoencoder with a semi-supervised Gaussian Mixture Model (GMM) prior. The prior defines a clustering in the latent space, where the clusters are associated with specific drug properties. In addition, VADEERS is equipped with a cell line autoencoder and a sensitivity prediction network. The model combines data for SMILES string representations of anti-cancer drugs, their inhibition profiles against a panel of protein kinases, cell lines biological features and measurements of the sensitivity of the cell lines to the drugs. The evaluated variants of VADEERS achieve a high r=0.87 Pearson correlation between true and predicted drug sensitivity estimates. We train the GMM prior in such a way that the clusters in the latent space correspond to a pre-computed clustering of the drugs by their inhibitory profiles. We show that the learned latent representations and new generated data points accurately reflect the given clustering. In summary, VADEERS offers a comprehensive model of drugs and cell lines properties and relationships between them, as well as a guided generation of novel compounds.
We investigate the performance of sparsely connected neural networks, with connectivity determined by road network graphs, for solving the Traffic Signal Setting optimization problem. We conducted experiments on three realistic road network topologies and found these types of graph neural networks superior to fully connected ones, both in terms of generalization properties on fixed test sets and more importantly near target function minima obtained in the gradient descent optimization process. We additionally confirm the soundness of our method by showing that random perturbations of the actual graph lead to consistent deterioration of model performance.
This paper presents our contribution to PolEval 2019 Task 6: Hate speech and bullying detection. We describe three parallel approaches that we followed: fine-tuning a pre-trained ULMFiT model to our classification task, fine-tuning a pre-trained BERT model to our classification task, and using the TPOT library to find the optimal pipeline. We present results achieved by these three tools and review their advantages and disadvantages in terms of user experience. Our team placed second in subtask 2 with a shallow model found by TPOT: a~logistic regression classifier with non-trivial feature engineering.
We present a method for optimizing traffic signal settings which can be used for offline planning and realtime adaptive traffic management. The method is based on metaheuristics efficiently exploring space of possible settings and evaluating candidate solutions using a microscopic traffic simulation or metamodels of simulations built using machine learning algorithms (e.g., neural networks, LightGBM). We present results of extensive experiments and compare different algorithms and their configurations in order to find the best approach in our use case. Experiments were carried out on a realistic road network of Warsaw (maps originated from the OpenStreetMap service) and showed that LightGBM may outperform neural networks in terms of accuracy of approximations, time efficiency and optimality of traffic signal settings, which is a new and important result. We also show that in terms of traffic optimization genetic algorithms give the best results comparing to other metaheuristics.
Machine learning algorithms hold the promise to effectively automate the analysis of histopathological images that are routinely generated in clinical practice. Any machine learning method used in the clinical diagnostic process has to be extremely accurate and, ideally, provide a measure of uncertainty for its predictions. Such accurate and reliable classifiers need enough labelled data for training, which requires time-consuming and costly manual annotation by pathologists. Thus, it is critical to minimise the amount of data needed to reach the desired accuracy by maximising the efficiency of training. We propose an accurate, reliable and active (ARA) image classification framework and introduce a new Bayesian Convolutional Neural Network (ARA-CNN) for classifying histopathological images of colorectal cancer. The model achieves exceptional classification accuracy, outperforming other models trained on the same dataset. The network outputs an uncertainty measurement for each tested image. We show that uncertainty measures can be used to detect mislabelled training samples and can be employed in an efficient active learning workflow. Using a variational dropout-based entropy measure of uncertainty in the workflow speeds up the learning process by roughly 45%. Finally, we utilise our model to segment whole-slide images of colorectal tissue and compute segmentation-based spatial statistics.
We present a new method for uncertainty estimation and out-of-distribution detection in neural networks with softmax output. We extend softmax layer with an additional constant input. The corresponding additional output is able to represent the uncertainty of the network. The proposed method requires neither additional parameters nor multiple forward passes nor input preprocessing nor out-of-distribution datasets. We show that our method performs comparably to more computationally expensive methods and outperforms baselines on our experiments from image recognition and sentiment analysis domains.
We analyze the accuracy of traffic simulations metamodels based on neural networks and gradient boosting models (LightGBM), applied to traffic optimization as fitness functions of genetic algorithms. Our metamodels approximate outcomes of traffic simulations (the total time of waiting on a red signal) taking as an input different traffic signal settings, in order to efficiently find (sub)optimal settings. Their accuracy was proven to be very good on randomly selected test sets, but it turned out that the accuracy may drop in case of settings expected (according to genetic algorithms) to be close to local optima, which makes the traffic optimization process more difficult. In this work, we investigate 16 different metamodels and 20 settings of genetic algorithms, in order to understand what are the reasons of this phenomenon, what is its scale, how it can be mitigated and what can be potentially done to design better real-time traffic optimization methods.
Schedae Informaticae » 2018 » Volume 27 » Traffic Signal Settings Optimization Using Gradient Descent A A A