
Protein tertiary structures capture the complex three-dimensional (3D) arrangements of the constituent atoms that define how proteins function. Different strategies have been applied in the form of representation learning to interpret these high-dimensional structure spaces. Factorization-based techniques, such as matrix and tensor factorization, have also been proven to be an effective set of approaches due to their utility in capturing latent organizations from complicated structural spaces and have already been applied in the context of tasks like tracking conformational changes across structure ensembles of biomolecules, identifying the biologically-active tertiary structure(s) from the given computed protein models, analyzing biomolecular dynamics simulations, etc. This paper exhibits a comprehensive review of factorization-driven methods employed to learn representations of tertiary structures for the protein structure prediction task. It covers how protein structures are encoded, represented, analyzed, and evaluated via different factorization-based frameworks. Furthermore, it outlines major open challenges and offers prospective research directions.
Since its introduction, the double cut and join (DCJ) operation has gained popularity in Computational Biology as a convenient abstraction that generalizes various genome rearrangement operations – mutation events that affect large genome segments. The central problem involving this operation is the DCJ distance problem, which seeks to minimize the number of DCJ operations required to transform one genome into another. In particular, recent studies incorporate intergenic region sizes to obtain more realistic rearrangement distance estimates. When the genomes differ in gene content, indel (insertion or deletion) events must also be considered. We present the first algorithm for the DCJ distance that incorporates intergenic information and supports indels, a 2-approximation applicable to single-copy, co-tailed genomes.
We present a dynamic functional connectivity (DFC)-based classification analysis of functional magnetic resonance imaging (fMRI) data from veterans with a type of post-traumatic stress disorder (PTSD) and from matched normal control (NC) veterans. Whole-brain resting-state fMRI (rsfMRI) data which were scanned from 23 PTSD (mean age 49) and 30 NC (mean age 50) veterans were used for analyses. A computational method using statistics of DFC and support-vector machine (SVM) classifier were used to correctly classify PTSD vs NC with up to 98
Representation learning is central to structural bioinformatics, enabling shared embeddings for protein fold recognition, structural similarity search, and function prediction. However, the organization and interpretability of latent representations learned from protein distance or contact maps remain poorly understood. We present a representation-focused comparative analysis of two of our previously developed geometric autoencoders, SuperFoldAE and its contractive variant ConSOLAE, trained on C_α distance maps. To characterize latent representations, we examine latent organization using visualizations, clustering alignment, perturbation sensitivity, resolution studies, and transfer tasks. While both models capture coarse fold structure, their latent spaces differ markedly in geometry and robustness. Representations learned with contractive regularization exhibit smoother, more compact, and more stable latent geometry, with stronger unsupervised alignment to fold structure (ARI = 0.87, NMI = 0.91). These geometric advantages are reflected in downstream behavior, including consistent Top-1 and Top-5 accuracy and support for unsupervised protein length prediction ( R^2 = 0.64 ). Overall, this study provides a concise comparative characterization of latent spaces learned by distance-map autoencoders, offering a geometric perspective that links latent geometry and stability to robust protein fold representations.
Spatial domain identification is a fundamental task in spatial transcriptomics, aiming to partition tissues into coherent regions that reflect both transcriptional similarity and spatial organization. Existing methods typically generate low-dimensional embeddings through predefined transformations or neural network encoders, and treat these representations as fixed inputs for downstream clustering. This limits the model’s ability to adapt to multiple spatial and expression-derived constraints. To address this limitation, we propose Graph-Regularized Embedding Refinement for Spatial Transcriptomics (GRER-ST), a refinement framework that enhances spatial domain identification by explicitly optimizing the latent embedding. GRER-ST first integrates spatial coordinates and gene expression profiles to construct an initial low-dimensional representation that captures local spatial context. A trainable latent embedding, initialized from this representation, is then refined under a unified objective that incorporates spatial smoothness, expression similarity, and neighborhood-based contrastive regularization. By refining the embedding rather than relying on a fixed representation, GRER-ST produces spatial domains that are both locally coherent and discriminative. Extensive experiments on multiple spatial transcriptomics datasets demonstrate that GRER-ST consistently outperforms baseline methods in spatial domain identification, yielding more spatially coherent domain structures. These findings suggest that embedding refinement is an effective and flexible strategy for integrating heterogeneous spatial and transcriptional information in spatial transcriptomics.
Single-cell sequencing has transformed the study of cellular heterogeneity and its role in health and disease. Current models predominantly adopt expression-driven learning frameworks that encode cellular states through the joint modeling of genes and their associated expressions. Current methods treat gene expression in various ways, such as using a fixed gene order, ranking genes based on expression levels, or combining gene IDs with expression values through simple addition. However, these approaches overlook the biological context of genes and rely on static representations that fail to capture gene activity’s functional dynamics. Here, we present scFiLM, a dynamic fusion framework built upon the Feature-wise Linear Modulation (FiLM) paradigm, designed to integrate prior biological knowledge encoded in pretrained semantic gene embeddings derived from NCBI gene descriptions with expression-driven feature modulation. Specifically, for each gene–expression pair, scFiLM adaptively modulates the corresponding semantic embedding through FiLM layers, utilizing expression-derived scaling and shifting parameters. The resulting modulated embeddings are subsequently fed into a bidirectional Mamba encoder to capture long-range dependencies, thereby generating informative cell-level representations for downstream analyses. Comprehensive cell-level evaluations demonstrate that scFiLM not only effectively mitigates batch effects, but also achieves state-of-the-art performance in B cell type classification. At gene level, scFiLM retains the intrinsic functionality of genes and enables effective categorization of gene functions, thereby enhancing the interpretability of learned representations. The results highlight the potential of scFiLM as a robust and generalizable framework for single-cell transcriptomic analysis.
Spatial multi-omics technologies enable joint profiling of molecular modalities within native tissue contexts, but integrating data across heterogeneous tissue sections and experimental batches remains a fundamental challenge. A key limitation of existing methods is that spatial graphs are constructed using manually defined radii or fixed neighbor sizes prior to training, preventing models from adapting spatial interaction scales across tissues with variable architectures and cell densities. To address this limitation, we propose SpaHMS, a fully dynamic framework for spatial multi-omics integration that learns hierarchical, multi-scale spatial receptive fields in an end-to-end manner. SpaHMS introduces a Dynamic Hierarchical Multi-Scale Graph Neural Network that jointly optimizes spatial graph structure and multimodal representations, enabling adaptive spatial modeling across heterogeneous tissues. Each modality is encoded using a dedicated network with cross-section parameter sharing, preserving modality-specific structure while mitigating batch effects through explicit mutual nearest neighbor (MNN) alignment. Across datasets involving RNA, ADT, and ATAC, SpaHMS produces robustly aligned, batch-corrected, and biologically coherent embeddings that preserve meaningful spatial organization.
Protein language models (PLMs) are increasingly used for tasks such as structure prediction, function annotation, and protein–protein interaction modeling. However, the rapid expansion of available models including sequence-only, structure-informed, and multimodal architectures makes it challenging for researchers to determine which model best fits a specific biological task, computational budget, or deployment setting. Existing surveys often emphasize architectural or benchmark comparisons but provide limited practical guidance for real-world usability and integration. To address this gap, we introduce the Adaptability Bioinformatics Usability Score (ABUS), a structured evaluation framework that characterizes PLMs across five key dimensions: adaptability, bioinformatics relevance, usability, computational efficiency, and output suitability. Each dimension is decomposed into specific rubric-based sub-features scored on a 0/1/2 scale based on objective evidence. These sub-category values are then aggregated and transformed into a normalized overall ABUS score ranging from 0 to 100, providing both granular qualitative insights and a comparative quantitative metric. ABUS is implemented as an evidence-driven scoring engine supported by human-in-the-loop validation. The framework includes a curated database of over 20 models, an interactive web interface for exploring detailed model profiles, and a prototype AI agent designed to extract evidence and recommend scores directly from research papers and repositories.
Chronic wound assessment requires accurate segmentation and interpretable analysis for effective treatment planning. Although deep learning models show promise in automated wound segmentation, clinical adoption is hampered by limited interpretability. In this paper, we present a framework that integrates advanced wound segmentation with conversational interfaces powered by a large language model (LLM). Our EfficientNet-Attention U-Net trained on 2,760 wound images achieves state-of-the-art performance with an accuracy of 99.76
Open-TGGATEs is an important toxicogenomics database that has been used and referenced in many studies. Unfortunately, Open-TGGATEs were curated decades ago on the Affymetrix platform. This creates barriers for current studies that often use profiles from modern platforms such as TempO-Seq. In our study we present an AI-based methodology for modernizing Open-TGGATEs. We demonstrate the effectiveness of our approach using the measured liver profiles. Our results show that the modernized profiles maintain important relationships among treatment profiles in the original measurements.
We investigate eco-epidemic models to analyze the averaged dynamics of spatio-temporal systems. Specifically, we study a coupled Lotka–Volterra–Susceptible–Infected–Susceptible model with diffusion, posed on various graph structures. This eco-epidemic framework, formulated as ordinary differential equations on graphs (Graph ODEs), consists of four equations describing the evolution of susceptible and infected prey and predator populations, capturing ecological interactions, disease transmission, and spatial dispersal on a graph. Our primary goal is to understand how spatial coupling influence the system’s averaged temporal dynamics. Numerical experiments demonstrate that diffusion strength has a significant impact on the temporal behavior, underscoring the crucial role of spatial processes in shaping eco-epidemic dynamics.
Background: Unplanned Intensive Care Unit (ICU) readmission after breast cancer surgery is associated with increased morbidity, mortality, and healthcare utilization [1, 2]. Existing ICU readmission models are typically developed for heterogeneous populations and may not generalize to narrowly defined postoperative oncology cohorts [3–5]. Objective: This study presents a pilot feasibility analysis of an interpretable, cohort-specific clinical decision support framework for identifying postoperative breast cancer patients at elevated risk of ICU readmission using routinely collected electronic health record (EHR) data. Methods: A retrospective cohort was curated from the MIMIC-IV database [6] using reproducible phenotyping and transfer-based ICU readmission logic. Thirteen physiologic, laboratory, and medication-related variables recorded within 24 h after surgery or index ICU discharge were extracted in raw clinical units. Exploratory modeling using logistic regression, random forest, and k-nearest neighbors models was performed solely to assess feature stability and relevance rather than predictive performance. In parallel, a transparent rule-based alert module categorized abnormalities using established clinical thresholds [7–9] and generated color-coded (Green/ Yellow/ Red) severity alerts. Results: The final analytic cohort consisted of 51 postoperative breast cancer patients of whom 40 (78.4
A combination of biological delays, tumor-microenvironment transport barriers, and dosing cadence determines the success of oncolytic virotherapy. In this study, we develop a delay-differential model incorporating virus-tumor-immune dynamics with explicit infection, lysis, and immune-priming delays, as well as a tunable transport-impedance field that modulates viral contact and diffusion. A tractable viral invasion metric relates barrier severity to eradication thresholds. Meanwhile, stability and Hopf bifurcation analyses demonstrate that cumulative delays can induce oscillations that promote relapse. A safety-aware optimal control formulation, combined with learning-based model predictive control (MPC), enables adaptive, biomarker-guided dosing under toxicity constraints. A spatial reaction-diffusion delay extension reveals stacked traveling fronts under severe transport impedance, which motivates the normalization of transport and optimized therapy sequencing. Based on this framework, personalized oncolytic virotherapy can be designed.
We investigate the use of supervised machine learning models to predict equilibrium outcomes in a predator–prey system with disease. Epidemics in this system are described by a coupled model that integrates Lotka–Volterra (LV) predator–prey dynamics with a Susceptible–Infected–Susceptible (SIS) epidemic framework. A synthetic dataset is generated by randomly varying model parameters within prescribed ranges and recording the final susceptible and infected populations as equilibrium states. This dataset is then used to train a wide suite of supervised models, including a linear model, tree ensembles, boosting methods, and neural networks. The results demonstrate that machine learning provides fast and reliable equilibrium prediction, enabling large-scale scenario exploration without repeatedly integrating the underlying system of ordinary differential equations.
Matrix-assisted laser desorption ionization–time of flight (MALDI-TOF) mass spectrometry (MS) is a rapid and widely utilized method for microbial identification and detecting their antibiotic-resistance profiles. This method extends to the detection of clinically significant organisms, such as vancomycin-resistant Enterococcus faecium (VRE), an opportunistic pathogen of the gut microbiome. Machine learning is well-suited and has been increasingly applied for analyzing MALDI-TOF MS data. However, the MALDI-TOF studies that focus on VRE are limited. In binary resistant vs. susceptible E. faecium classification using a feature matrix with 36,000 features created from preprocessed spectra without signal-to-noise ratio subtraction (SNR), random forest (RF) achieves the highest accuracy (78.7
Protein structure provides the foundation for a protein’s function. Therefore, incorporating structural information is essential in methods for protein function prediction. The attention layer in transformer models has proven to be a powerful tool for encoding rich sequential information in a wide range of machine learning tasks. Although originally developed for sequence inputs, the attention mechanism can be extended to integrate structural information. In this paper, we present a structure-aware attention mechanism that incorporates protein structural information for protein function prediction. We demonstrate that integrating structure-derived features significantly improves the performance of predicting protein functions. Our method outperforms the current state-of-the-art in comprehensive benchmark evaluations, with gains up to 29.6
Protein sequence generation has traditionally relied on autoregressive and diffusion based models, while Fourier-based approaches remain largely unexplored in this domain. We apply FNet, a Fourier transform based architecture originally developed for natural language processing, to protein sequence generation and compare its performance against ESM2’s direct masked sampling strategy and the autoregressive model ProGen. All three models were trained and evaluated on a curated kinase dataset under identical experimental conditions. FNet achieves substantial computational gains over ProGen, reducing generation time by 95 https://github.com/bashirgit/Fourier_vs._Attention .
Pregnancy care often involves simultaneous obstetric and other medical conditions, but their co-occurrence patterns are rarely modeled explicitly in a systematic, network-based approach. In this work, we formulate obstetric and non-obstetric diagnoses co-occurrences as a link prediction problem on a diagnosis-level homogeneous graph constructed from pregnancy encounters. Diagnoses are represented as nodes connected by co-occurrence edges, with node features capturing graph structure and demographic statistics. We address this challenge by leveraging collected electronic health records data and study several standalone and hybrid graph neural network (GNN) architectures, including GCN, GAT, GraphSAGE, and three hybrid encoders that combine complementary aggregation mechanisms, namely GCN+GraphSAGE, GCN+GAT, and GAT+GraphSAGE. All models used consistent train-validation-test splits and are evaluated on 5-fold cross-validation sets. Among standalone models, GraphSAGE achieved the strongest performance, whereas hybrid GraphSAGE-based models (GCN+GraphSAGE and GAT+GraphSAGE) are best performers. The GCN+GraphSAGE hybrid, reaching an AUROC and AUPRC of approximately 0.90, consistently outperformed all other architectures. Further analysis of top-ranked predicted links revealed clinically plausible associations between pregnancy stage and risk-related diagnoses and common endocrine, metabolic, and hematological conditions. These findings indicate that graph-based link prediction may effectively prioritize obstetric and non-obstetric diagnosis pairs, providing a scalable framework for identifying clinically meaningful comorbidity patterns. They may further support hypothesis generation and downstream obstetric risk stratification efforts.
In this work, we investigate whether motif subsequence features can predict accessible chromatin regions (ACR) in a genome, i.e. regions that are accessible to regulatory proteins, thus enabling transcription of associated genes. We focus on plants, whose agricultural and ecological importance make them interesting and important organisms to study, and whose complex genomes provide important stress tests for our algorithm. We show that regulatory motif sequence similarity can be found efficiently using co-linear chaining. The similarity scores found are then used as features in machine learning models to explore the feasibility of effectively predicting ACRs in genome assemblies.
Parkinson’s disease (PD) is a chronic and complex neurodegenerative disorder influenced by genetic, clinical, and lifestyle factors. Predicting this disease early is challenging because it depends on traditional diagnostic methods that face issues of subjectivity, which commonly delay diagnosis. Several objective analyses are currently in practice to help overcome the challenges of subjectivity; however, a proper explanation of these analyses is still lacking. While machine learning (ML) has demonstrated potential in supporting PD diagnosis, existing approaches often rely on subjective reports only and lack interpretability for individualized risk estimation. This study proposed SCOPE-PD, an explainable AI-based prediction framework by integrating subjective and objective assessments to provide personalized health decisions. Subjective and objective clinical assessment data are collected from the Parkinson’s Progression Markers Initiative (PPMI) study to construct a multimodal prediction framework. Several ML techniques were applied to these data, and the best ML model was selected to interpret the results. Model interpretability was examined using SHAP-based analysis. The Random Forest algorithm achieved the highest accuracy of 98.66