Motivation The advent of single-cell RNA sequencing (scRNA-seq) has enhanced our ability to study cellular heterogeneity. Accurately identifying distinct subpopulations and their defining markers is critical for understanding tissue diversity. We introduce CORTADO, a hill-climbing optimization framework for marker discovery and clustering refinement.Results CORTADO maximizes differential expression, minimizes redundancy via cosine similarity, and enforces sparsity for interpretability. By using CORTADO-selected markers to inform the cell-type identification process, an iterative refinement approach markedly increases the Adjusted Rand Index (ARI), a metric that quantifies how well the clustering assignments align with gold-standard cell-type annotations. Benchmarking across brain, immune, spatial, and cancer datasets confirms that CORTADO delivers biologically relevant markers and consistently outperforms state-of-the-art methods in both marker discovery and clustering accuracy.
In the domain of network biology, the interactions among heterogeneous genomic and molecular entities are represented through networks. Link prediction (LP) methodologies are instrumental in inferring missing or prospective associations within these biological networks. In this review, we systematically dissect the attributes of local, centrality, and embedding-based LP approaches, applied to static and dynamic biological networks. We undertake an examination of the current applications of LP metrics for predicting links between diseases, genes, proteins, RNA, microbiomes, drugs, and neurons. We carry out comprehensive performance evaluations on established biological network datasets to show the practical applications of standard LP models. Moreover, we compare the similarity in prediction trends among the models and the specific network attributes that contribute to effective link prediction, before underscoring the role of LP in addressing the formidable challenges prevalent in biological systems, ranging from noise, bias, and data sparseness to interpretability. We conclude the review with an exploration of the essential characteristics expected from future LP models, poised to advance our comprehension of the intricate interactions governing biological systems.
Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance the prediction capability of the models. To fill this gap, we propose GramSeq-DTA, which integrates chemical perturbation information with the structural information of drugs and targets. We applied a Grammar Variational Autoencoder (GVAE) for drug feature extraction and utilized two different approaches for protein feature extraction as follows: a Convolutional Neural Network (CNN) and a Recurrent Neural Network (RNN). The chemical perturbation data are obtained from the L1000 project, which provides information on the up-regulation and down-regulation of genes caused by selected drugs. This chemical perturbation information is processed, and a compact dataset is prepared, serving as the functional feature set of the drugs. By integrating the drug, gene, and target features in the model, our approach outperforms the current state-of-the-art DTA prediction models when validated on widely used DTA datasets (BindingDB, Davis, and KIBA). This work provides a novel and practical approach to DTA prediction by merging the structural and functional aspects of biological entities, and it encourages further research in multi-modal DTA prediction.
Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug-target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug discovery research since researchers started using deep learning as a potent tool for DTB prediction. In particular, we examine how methodologies have evolved, starting with early heterogeneous network-based approaches, progressing to graph-based approaches that were widely accepted, followed by modern attention-based architectures, and finally, the most recent multimodal approaches. We also provide case studies utilizing an extensive compound library against specific protein targets implicated in critical cancer pathways to demonstrate the usefulness of these approaches. In addition to summarizing the latest developments in DTB prediction models, this review also identifies their drawbacks. It also highlights the outlook for the DTB prediction domain and future research directions. Combined, these studies present a more comprehensive view of how deep learning offers a quantitative framework for researching drug-target relationships, speeding up the identification of new drug candidates and making it easier to identify possible DTBs.
The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation-maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene analysis and pathway enrichment for BRCA patients. The analysis compared 302 tumor samples to 113 normal samples across 60,660 genes. The highly upregulated gene COL10A1, promoting breast cancer progression and poor prognosis, and the consistently downregulated gene CDG300LG, linked to brain metastatic cancer, were identified. Pathway enrichment analysis revealed similarities in cellular matrix organization pathways across subtypes, with notable differences in functions like cell proliferation regulation and endocytosis by host cells. GO Semantic Similarity analysis quantified gene relationships in each subtype, identifying potential biomarkers like MATN2, similar to COL10A1. These insights suggest deeper relationships within clusters and highlight personalized treatment potential based on subtypes.
The study explores the therapeutic relevance of Cardiac glycosides (CGs), including lanatoside C (LC), peruvoside (PS), and strophanthidin (STR) in treating breast cancer, using network pharmacology studies and bioinformatics approaches. Building on our prior in vitro studies and transcriptome profiling, we aimed to explore protein expression alterations influenced by the selected compounds in the present study. The methodology was structured and directed to delineate the active protein targets and their molecular mechanism of action in controlling cancer progression. Initially, we predicted the protein targets of individual compounds using SWISSTargetPrediction, and the results were compared with the differentially expressed genes from the transcriptome data acquired in the preliminary studies. The identified protein targets were further studied for their network relatedness and cross-verified by comparing their expression in cancer and normal patient data from TCGA using the UALCAN algorithm. Additionally, we aimed to identify the candidate biomarkers that potentially served as predictive or prognostic indicators in malignant breast cancer by conducting survival analysis of the crucial proteins using the GEPIA2 database. Overall, the analysis allowed us to understand the co-dependence expression between MAPK1 and EGR1 proteins, further emphasizing their clinical significance in cancer diagnosis and probable therapeutic outcomes. MD simulation studies further verified the most significant protein targets with their interaction scores and structural stability of the compounds, which showed higher structural stability around 300-ns trajectories for MAPK1 and EGR1 proteins. Finally, the pathway simulation studies, modeling on the MAPK/ERK signaling cascade, showed significant alteration in the biochemical parameters and stability of the system depending on the concentration of crucial proteins ERK2, also known as MAPK1 and EGR1 proteins in the pathway. These findings underscore the therapeutic potential of CGs and further highlight the significant role of identified proteins in targeting breast cancer.
This study examined the effects of 24R,25-dihydroxyvitamin D3 (24R,25(OH)2D3) in estrogen-responsive laryngeal cancer tumorigenesis in vivo, the mechanisms involved, and whether the ability of the tumor cells to produce 24R,25(OH)2D3 locally is estrogen-dependent. Estrogen receptor alpha-66 positive (ER+) UM-SCC-12 cells and ER- UM-SCC-11A cells responded differently to 24R,25(OH)2D3 in vivo; 24R,25(OH)2D3 enhanced tumorigenesis in ER+ tumors but inhibited tumorigenesis in ER- tumors. Treatment with 17β-estradiol (E2) for 24 h reduced levels of CYP24A1 protein but increased 24R,25(OH)2D3 production in ER+ cells; treatment with E2 for 9 min reduced CYP24A1 at 24 h and reduced 24R,25(OH)2D3 production in ER- cells. These findings suggest the involvement of E2 receptor(s) in addition to ERα66. To investigate if 24R,25(OH)2D3 can act locally, ER+ and ER- cells were treated with 24R,25(OH)2D3 after inhibiting putative 24R,25(OH)2D3 receptors, and the cells were assessed for effects on DNA synthesis (proliferation) and p53 production (apoptosis). Specific inhibitors were used to assess downstream secondary messenger signaling pathways and requirements for palmitoylation and caveolae in both cell lines. The results show that 24R,25(OH)2D3 binds to a complex of receptors, including TLCD3B2, VDR, and protein disulfide-isomerase A3 (PDIA3) in ER+ UM-SCC-12 cells. The mechanism requires palmitoylation, and PLD, PI3K, and LPAR are involved. The anti-tumorigenic effects of 24R,25(OH)2D3 in ER- UM-SCC-11A cells involve a membrane-receptor complex consisting of VDR, PDIA3, and ROR2 within caveolae to activate a yet-to-be-elucidated downstream signaling cascade. This work demonstrates a driving mechanism for the therapeutic agent 24R,25(OH)2D3 that may be used for laryngeal cancer patients.
During the COVID-19 pandemic, the prevalence of asymptomatic cases challenged the reliability of epidemiological statistics in policymaking. To address this, we introduced contagion potential (CP) as a continuous metric derived from sociodemographic and epidemiological data to quantify the infection risk posed by the asymptomatic within a region. However, CP estimation is hindered by incomplete or biased incidence data, where underreporting and testing constraints make direct estimation infeasible. To overcome this limitation, we employ a hypothesis-testing approach to infer CP from sampled data, allowing for robust estimation despite missing information. Even within the sample collected from spatial contact data, individuals possess partial knowledge of their neighborhoods, as their awareness is restricted to interactions captured by available tracking data. We introduce an adjustment factor that calibrates the sample CPs so that the sample is a reasonable estimate of the population CP. Further complicating estimation, biases in epidemiological and mobility data arise from heterogeneous reporting rates and sampling inconsistencies, which we address through inverse probability weighting to enhance reliability. Using a spatial model for infection spread through social mixing and an optimization framework based on the SIRS epidemic model, we analyze real infection datasets from Italy, Germany, and Austria. Our findings demonstrate that statistical methods can achieve high-confidence CP estimates while accounting for variations in sample size, confidence level, mobility models, and viral strains. By assessing the effects of bias, social mixing, and sampling frequency, we propose statistical corrections to improve CP prediction accuracy. Finally, we discuss how reliable CP estimates can inform outbreak mitigation strategies despite the inherent uncertainties in epidemiological data.
The surface topography and chemistry of titanium–aluminum–vanadium (Ti6Al4V) implants play critical roles in the osteoblast differentiation of human bone marrow stromal cells (MSCs) and the creation of an osteogenic microenvironment. To assess the effects of a microscale/nanoscale (MN) topography, this study compared the effects of MN-modified, anodized, and smooth Ti6Al4V surfaces on MSC response, and for the first time, directly contrasted MN-induced osteoblast differentiation with culture on tissue culture polystyrene (TCPS) in osteogenic medium (OM). Surface characterization revealed distinct differences in microroughness, composition, and topography among the Ti6Al4V substrates. MSCs on MN surfaces exhibited enhanced osteoblastic differentiation, evidenced by increased expression of RUNX2, SP7, BGLAP, BMP2, and BMPR1A (fold increases: 3.2, 1.8, 1.4, 1.3, and 1.2). The MN surface also induced a pro-healing inflammasome with upregulation of anti-inflammatory mediators (170–200% increase) and downregulation of pro-inflammatory factors (40–82% reduction). Integrin expression shifted towards osteoblast-associated integrins on MN surfaces. RNA-seq analysis revealed distinct gene expression profiles between MSCs on MN surfaces and those in OM, with only 199 shared genes out of over 1000 differentially expressed genes. Pathway analysis showed that MN surfaces promoted bone formation, maturation, and remodeling through non-canonical Wnt signaling, while OM stimulated endochondral bone development and mineralization via canonical Wnt3a signaling. These findings highlight the importance of Ti6Al4V surface properties in directing MSC differentiation and indicate that MN-modified surfaces act via signaling pathways that differ from OM culture methods, more accurately mimicking peri-implant osteogenesis in vivo.
Physics-informed machine learning bridges the gap between the high fidelity of mechanistic models and the adaptive insights of artificial intelligence. In chemical reaction network modeling, this synergy proves valuable, addressing the high computational costs of detailed mechanistic models while leveraging the predictive power of machine learning. This study applies this fusion to the biomedical challenge of A $$\beta$$ fibril aggregation, a key factor in Alzheimer’s disease. Central to the research is the introduction of an automatic reaction order model reduction framework, designed to optimize reduced-order kinetic models. This framework represents a shift in model construction, automatically determining the appropriate level of detail for reaction network modeling. The proposed approach significantly improves simulation efficiency and accuracy, particularly in systems like A $$\beta$$ aggregation, where precise modeling of nucleation and growth kinetics can reveal potential therapeutic targets. Additionally, the automatic model reduction technique has the potential to generalize to other network models. The methodology offers a scalable and adaptable tool for applications beyond biomedical research. Its ability to dynamically adjust model complexity based on system-specific needs ensures that models remain both computationally feasible and scientifically relevant, accommodating new data and evolving understandings of complex phenomena.
Accurate prediction of hospital length of stay (LoS) is a vital component in optimizing clinical workflows, resource allocation, and patient care. This study presents a comprehensive evaluation of machine learning models for both binary and multi-class LoS classification tasks using structured clinical variables, physiological measurements, and unstructured clinical notes. Seven data configurations were constructed from combinations of structured features (Z), including diagnoses, procedures, medications, laboratory tests, and microbiology results; MeSH-based symptoms (S); physiological signals (F); and textual representations (E): Z, F, E, ZS, ZSF, ZSE, and ZSEF. Five predictive models-Artificial Neural Networks (ANN), XGBoost, Logistic Regression (LR), Random Forest (RF), and Support Vector Machine (SVM)-were applied, with and without feature selection, where categorical features and Bag-of-Words representations were reduced to varied dimensions. Results indicate that the base structured feature set (Z) alone yields strong predictive performance across tasks. Moreover, the integration of additional data types-S, F, and E-either individually or in combination, consistently enhanced performance, with the ZSEF configuration achieving the highest F1-scores and AUC values in most cases. While the application of SMOTE did not yield substantial improvements in the global setting encompassing all hospital admissions, it demonstrated enhanced performance in disease-specific cohorts, particularly for patients admitted with lung cancer. Among the evaluated models, XGBoost and ANN demonstrated superior generalizability. These findings underscore the effectiveness of multimodal data integration and feature reduction techniques in advancing predictive modeling for hospital length of stay across diverse patient populations.
Wireless Body Area Networks (WBANs) are pivotal in health care and wearable technologies, enabling seamless communication between miniature sensors and devices on or within the human body. These biosensors capture critical physiological parameters, ranging from body temperature and blood oxygen levels to real-time electrocardiogram readings. However, WBANs face significant challenges during and after deployment, including energy conservation, security, reliability, and failure vulnerability. Sensor nodes, which are often battery-operated, expend considerable energy during sensing and transmission due to inherent spatiotemporal patterns in biomedical data streams. This paper provides a comprehensive survey of data-driven approaches that address these challenges, focusing on device placement and routing, sampling rate calibration, and the application of machine learning (ML) and statistical learning techniques to enhance network performance. Additionally, we validate three existing models (statistical, ML, and coding-based models) using two real datasets, namely the MIMIC clinical database and biomarkers collected from six subjects with a prototype biosensing device developed by our team. Our findings offer insights into strategies for optimizing energy efficiency while ensuring security and reliability in WBANs. We conclude by outlining future directions to leverage approaches to meet the evolving demands of healthcare applications.
The advent of single-cell RNA sequencing (scRNA-seq) has greatly enhanced our ability to explore cellular heterogeneity with high resolution. Identifying subpopulations of cells and their associated molecular markers is crucial in understanding their distinct roles in tissues. To address the challenges in marker gene selection, we introduce CORTADO, a computational framework based on hill-climbing optimization for the efficient discovery of cell-type-specific markers. CORTADO optimizes three critical properties: differential expression in the clusters of interest, distinctiveness in gene expression profiles to minimize redundancy, and sparseness to ensure a concise and biologically meaningful marker set. Unlike traditional methods that rely on ranking genes by p-values, CORTADO incorporates both differential expression metrics and penalties for overlapping expression profiles, ensuring that each selected marker uniquely represents its cluster while maintaining biological relevance. Its flexibility supports both constrained and unconstrained marker selection, allowing users to specify the number of markers to identify, making it adaptable to diverse analytical needs and scalable to datasets with varying complexities. To validate its performance, we apply CORTADO to several datasets, including the DLPFC 151507 dataset, the Zeisel mouse brain dataset, and a peripheral blood mononuclear cell dataset. Through enrichment analysis and examination of spatial localization-based expression, we demonstrate the robustness of CORTADO in identifying biologically relevant and non-redundant markers in complex datasets. CORTADO provides an efficient and scalable solution for cell-type marker discovery, offering improved sensitivity and specificity compared to existing methods.
Physics-informed machine learning emerges as a transformative approach, bridging the gap between the high fidelity of mechanistic models and the adaptive, data-driven insights afforded by artificial intelligence and machine learning. In the realm of chemical reaction network modeling, this synergy is particularly valuable. It offers a solution to the pro-hibitive computational costs associated with detailed mechanistic models, while also capitalizing on the predictive power and flexibility of machine learning algorithms. This study exemplifies this innovative fusion by applying it to the critical biomedical challenge of A β fibril aggregation, shedding light on the mechanisms underlying Alzheimer’s disease. A corner-stone of this research is the introduction of an automatic reaction order model reduction framework, tailored to optimize the scale of reduced order kinetic models. This framework is not merely a technical enhancement; it represents a paradigm shift in how models are constructed and refined. By automatically determining the most appropriate level of detail for modeling reaction networks, our proposed approach significantly enhances the efficiency and accuracy of simulations. This is particularly crucial for systems like A β aggregation, where the precise characterization of nucleation and growth kinetics can provide insights into potential therapeutic targets. The potential generalizability of this automatic model reduction technique to other network models is a key highlight of this study. The methodology developed here has far-reaching implications, offering a scalable and adaptable tool for a wide range of applications beyond biomedical research. The ability to dynamically adjust model complexity in response to the specific demands of the system under study is a powerful asset. This flexibility ensures that the models remain both computationally feasible and scientifically relevant, capable of accommodating new data and evolving understandings of complex phenomena. ### Competing Interest Statement The authors have declared no competing interest.
Several methods have been developed to computationally predict cell-types for single cell RNA sequencing (scRNAseq) data. As methods are developed, a common problem for investigators has been identifying the best method they should apply to their specific use-case. To address this challenge, we present CHAI (consensus Clustering tHrough similArIty matrix integratIon for single cell-type identification), a wisdom of crowds approach for scRNAseq clustering. CHAI presents two competing methods which aggregate the clustering results from seven state-of-the-art clustering methods: CHAI-AvgSim and CHAI-SNF. CHAI-AvgSim and CHAI-SNF demonstrate superior performance across several benchmarking datasets. Furthermore, both CHAI methods outperform the most recent consensus clustering method, SAME-clustering. We demonstrate CHAI's practical use case by identifying a leader tumor cell cluster enriched with CDH3. CHAI provides a platform for multiomic integration, and we demonstrate CHAI-SNF to have improved performance when including spatial transcriptomics data. CHAI overcomes previous limitations by incorporating the most recent and top performing scRNAseq clustering algorithms into the aggregation framework. It is also an intuitive and easily customizable R package where users may add their own clustering methods to the pipeline, or down-select just the ones they want to use for the clustering aggregation. This ensures that as more advanced clustering algorithms are developed, CHAI will remain useful to the community as a generalized framework. CHAI is available as an open source R package on GitHub: https://github.com/lodimk2/chai.
The pandemic caused by Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) has impacted the economy, health, and society. Emerging strains are making pandemic management challenging. There is an urge to collect epidemiological, clinical, and physiological data to make an informed decision on mitigation. Advances in the Internet of Things (IoT) and edge computing provide solutions for pandemic management through data collection and intelligent computation. While existing data-driven architectures operate on specific application domains and attempt to automate decision-making, they do not capture the multifaceted interaction among computational models, communication infrastructure, and data. In this article, we survey the existing approaches for pandemic management, including data repositories and contact-tracing applications. We envision a unified pandemic management architecture that leverages the IoT and edge computing paradigms to automate recommendations on vaccine distribution, dynamic lockdown, mobility scheduling, and pandemic trend prediction. We elucidate the data flow among the layers, namely, cloud, edge, and end device layers. Moreover, we address the privacy implications, threats, regulations, and solutions that may be adapted to optimize the utility of health data with security guarantees. The article ends with a discussion of the limitations of the architecture and research directions to enhance its practicality.
In the era of big data, it is necessary to provide novel and efficient platforms for training machine learning models over large volumes of data. The MapReduce approach and its Apache Spark implementation are among the most popular methods that provide high-performance computing for classification algorithms. However, they require dedicated implementations that will take advantage of such architectures. Additionally, many real-world big data problems are plagued by class imbalance, posing challenges to the classifier training step. Existing solutions for alleviating skewed distributions do not work well in the MapReduce environment. In this paper, we propose a novel KD-tree based classifier, together with a variation of the SMOTE algorithm dedicated to the Spark platform. Our algorithms offer excellent predictive power and can work simultaneously with binary and multi-class imbalanced data. Exhaustive experiments conducted using the Amazon Web Service platform showcase the high efficiency and flexibility of our proposed algorithms.
In areas affected by natural disasters, the functionality of communication networks is often compromised, resulting in partial or complete outages. Effective message sharing, crucial for facilitating prompt recovery efforts, is achieved by establishing mobile ad-hoc networks among user-owned devices. Network operations in post-disaster scenarios are inhibited by intermittent connectivity, delays, and energy constraints, necessitating routing strategies that ensure seamless communication amidst node failures and mobility challenges. To meet this challenge, we previously introduced an adaptive and distributed routing mechanism named ADRIN , capitalizing on the inherent periodicity in human mobility. The present work, ADRIN2.0 , extends the capabilities of ADRIN by incorporating a first-order Markov model-based mobility tracing approach to discern stable communication routes. It creates a spatiotemporal ad-hoc network to relay data multi-hop to the base station. Extensive simulations affirm that ADRIN2.0 approximates the underlying mobility distribution and facilitates data forwarding, even during node failures. ADRIN2.0 achieves a balance between data delivery rate and energy efficiency while minimizing latency, showing improvements over three centralized and two distributed routing benchmarks.