
The Sjogren syndrome (SS) is a chronic autoimmune disease that primarily affects the salivary glands. It has been associated with increased cancer risk, which implies that carcinogenic and autoimmune pathways may be interacting. Although being an integral component of mTOR complex 1 (mTORC1), regulatory-associated protein of mTOR (RAPTOR) is crucial for linking mTOR immune signaling to cellular metabolism. Signal transducer and activator of transcription 3 (STAT3) is the link between cancer-related and inflammation pathways. In this study, we used an integrative computational biology analysis, which focuses on non-experimental-based evidence, to investigate the potential of STAT3-RAPTOR crosstalk in the pathophysiology of Sjogren's syndrome and its related cancer. Protein-protein interaction (PPI) networks constructed in STRING and analysed and confirmed in Cytoscape network topology parameters demonstrated that STAT3 and RAPTOR are important interacting nodes in immune-metabolic signaling networks. Protein-protein interface characterisation of STAT3 by PDBsum revealed significant contact interfaces between the protein and RAPTOR, and molecular docking that ClusPro was performed and confirmed by HADDOCK. UniProt, EMBL-EBI ProtVar, and AlphaMissense helped in determining the human missense variants of STAT3 and RAPTOR to determine the functional relevance. GEPIA shows that RAPTOR and STAT3 are co-expressed in cancer data sets in terms of transcriptomic co-expression study. Importantly, the simultaneous expression was found to be preserved in autoimmune target tissue through disease-specific validation of salivary gland transcriptome data of Sjogren syndrome patients (GSE23117) using Autoimmune Disease Explorer (ADEx). All these findings indicate that there exists a bidirectional interaction between STAT3 and RAPTOR as a common immunometabolic signaling pathway that connects Sjogren's disease to the molecular signaling that follows the malignancy. Future wet-laboratory experiments will be required to experimentally validate the predicted interactions and their functional implications.
PURPOSE:To explore the relationship between surfactant protein B (SFTPB) gene expression and multiple biological processes in lung adenocarcinoma (LUAD) and construct a SFTPB-related gene signature for predicting LUAD prognosis. METHODS:Utilizing open-access datasets, we systematically investigated the association of SFTPB expression with patient prognosis, clinicopathological characteristics, immunoinfiltration, drug sensitivity, gene mutations, and methylation levels using the publicly available data. Pathway analysis revealed SFTPB-related signaling cascades, with distinct gene expression profiles being discerned when comparing groups exhibiting high versus low SFTPB expression levels. Additionally, we screened for optimal gene combinations for identifying prognostic signature and generated a nomogram using independent prognostic factors. RESULTS:SFTPB was significantly downregulated in LUAD samples, and patients with high SFTPB expression exhibited favorable survival. The high-SFTPB expression group exhibited lower tumor immune dysfunction and exclusion scores, IC50 values for eight chemotherapy drugs, and tumor mutation burden values. SFTPB expression was associated with multiple immune cells, including macrophages and activated CD4 T cells. Furthermore, we identified seven pathways, including such as PI3K-AKT-mTOR, associated with SFTPB. A risk model with high predictive value for the prognosis of patients with LUAD was constructed using ANLN, LYPD3, and IRX5. CONCLUSIONS:SFTPB expression was correlated with multiple biological processes in LUAD and may therefore serve as a valuable prognostic marker for LUAD.
Breast cancer (BC) is one of the most common malignancies in women globally, characterized by significant genetic and clinical heterogeneity. This complexity emphasises the need for reliable biomarkers and novel therapeutic strategies to improve patient outcomes.To address this, the present study introduces an integrated computational pipeline using multiple datasets to identify robust biomarkers and potential repurposed drugs in BC. From transcriptomic profiles across 12 GEO datasets, 143 differentially expressed genes (DEGs) were identified. Gene Ontology and functional enrichment analyses were performed, followed by the construction of a protein-protein interaction (PPI) network to pinpoint hub genes. These hubs were prioritised using multiple topological centrality measures and validated with independent datasets (cBioPortal, GEPIA, KM plotter) and machine-learning classification using five algorithms: Random Forest, XGBoost, Support vector machine, K-nearest neighbour, and Logistic regression. Machine-learning models achieved test accuracies >0.90 on TCGA-BRCA data (n=1,203), with K-nearest neighbour (class-weighted: accuracy 0.954) and XGBoost (SMOTE: accuracy 0.931) showing the strongest performance. Prioritised hub genes, TPX2 and BUB1B, were subjected to virtual-screening across five drug databases (DrugBank, DGIdb, OpenTargets, SwissTargetPrediction, and GSCALite), combined with molecular-docking, ADMET profiling, and drug-likeness evaluation, to identify promising repurposed candidates. Vorinostat exhibited the highest binding affinity to TPX2 (-29.32kcal/mol) and BUB1B (-23.71kcal/mol), followed by BRD-K90370028 (-18.40kcal/mol on BUB1B), NSC19630 and Dasatinib (with consistent dual-target binding), CD-437, and four additional prioritised compounds that exhibited favourable interactions. In conclusion, this coherent transcriptomics-to-therapeutics workflow establishes TPX2 and BUB1B as strong prognostic biomarkers in BC, with promising repurposed drugs targeting these mitotic regulators.
Background Glioblastoma (GBM) remains lethal due to high molecular heterogeneity and treatment resistance. While previous studies have proposed various biomarkers, a critical research gap exists: the lack of robust algorithmic validation and systematic linkage to drug discovery. Existing research predominantly relies on single machine learning models or traditional statistics, which often fail to provide stable results across diverse clinical datasets. To address this, we developed a high-dimensional pipeline that compares and ensembles 175 machine learning algorithm combinations. Unlike conventional single-model workflows, this approach ensures superior target identification stability and utilizes SHAP-based explainability and molecular dynamics to bridge the gap between biomarker discovery and precision therapeutics. Methods DEGs from GEO datasets were refined via PPI and functional analyses. The 175-algorithm ensemble identified core genes, with clinical utility validated via survival analysis. A drug discovery pipeline incorporating virtual screening, ADMET, and molecular dynamics (MD) was then implemented to evaluate compounds targeting the identified core genes. Results From 771 DEGs, 34 key genes were identified, with LOX validated as the core therapeutic target. The optimal predictive model achieved a robust AUC of 0.953, while survival analysis underscored the significant prognostic value of LOX. Following systematic screening, the most outstanding compound was prioritized via MD simulations, exhibiting exceptional binding stability, favorable pharmacokinetics, and minimal toxicity risk. Conclusion This integrated pipeline provides a robust framework for identifying precision targets and potent candidate compounds, offering a novel strategy for overcoming GBM treatment barriers.
Chromatin accessibility is generally associated with the binding of transcription factors and other regulatory proteins, which is fundamental to governing gene transcription. While the association between chromatin accessibility and gene expression levels is critical for transcriptional regulation, it remains incompletely characterized. Saccharomyces cerevisiae is a key eukaryotic model organism and a widely used chassis in synthetic biology, but studies on predicting gene expression from chromatin accessible regions are lacking. We developed Yeast-Gene, a supervised machine learning model that uses k-mer features from chromatin accessible regions to predict gene expression. Yeast-Gene focuses on local sequences of a few hundred base pairs within chromatin accessible regions. The model achieves an Area Under the Curve (AUC) of 0.90. The interpretability analysis identified AAGAA and CAAGA as highly influential motifs in the prediction of gene expression, and both motifs are potentially associated with mRNA splicing. These predictive features may contribute to the rational design of high-expression regulatory elements in synthetic biology.
Protein subcellular localization is key to understanding cellular function. Traditional methods are slow, prompting the use of machine learning and deep learning to enhance prediction accuracy. This study aims to leverage these approaches for more efficient and accurate localization prediction. This study presents a novel two-step approach for protein localization. First, feature extraction is performed using node embeddings derived from Gene Ontology graphs that capture hierarchical relationships between Gene Ontology terms representing molecular functions, biological processes, and cellular components. Embedding these Gene Ontology terms into a continuous vector space enhances protein representations within complex cellular environments. The second step involves multi-label classification using a dynamic thresholding-based Deep Neural Network to classify proteins into multiple subcellular locations. Unlike fixed thresholds, which typically yield single-label predictions, dynamic thresholding adapts to model outputs, enabling simultaneous prediction across multiple locations. The proposed model is evaluated on two benchmark datasets, namely DeepLoc 2.0 and a curated version of the Plant-mSubP dataset. The model achieves 44.69% Overall Actual Accuracy and 91.21% Relaxed Accuracy on the DeepLoc 2.0 dataset, and 82.43% Overall Actual Accuracy and 97.25% Relaxed Accuracy on the Plant-mSubP dataset. These results show significant improvements, with a minimum increase of 5.69% in Overall Actual Accuracy on DeepLoc 2.0 and 12.67% in Relaxed Accuracy on Plant-mSubP compared to state-of-the-art models. The proposed model significantly improves multi-label protein localization by combining Gene Ontology embeddings with dynamic thresholding, outperforming existing methods.
Allergic diseases, encompassing allergic asthma, atopic dermatitis, food allergies, have emerged as substantial global public health challenges, with IgE-stimulated mast cell and basophil playing pivotal roles in disease pathogenesis. Our prior research has indicated the potential therapeutic effects of Ginkgo biloba leaf (GBL) extracts in IgE-mediated mast cell degranulation. The present study aimed to identify key genes and pathways associated with IgE-mediated mast cell and basophil responses and explore the mechanisms of GBL components. Initially, DEGs and WGCNA analysis were employed to preliminarily identify candidate genes associated with IgE stimulation in GSE96696 dataset. Subsequently, transcriptomic analysis was conducted on an anti-DNP IgE-induced RBL-2H3 degranulation model to validate the key genes, including Egr1, Tnfaip3, Nfkbid, Il13, Nfkbia, Il3, Il4, Ccl7, Cdkn1a, Csf2, Rel, alongside crucial signaling pathways such as NF-κB, and PI3K-AKT pathways. Machine learning algorithms were then applied to confirm the predictive capacity of key genes. Building upon these findings, molecular biology experiments, molecular docking, and molecular dynamics simulations demonstrated that GBL flavonoids, including quercetin, amentoflavone, ginkgetin, and bilobetin, significantly inhibited the release of cytokines, blocked calcium ion influx, and suppressed degranulation potentially by targeting RELA and AKT1. Collectively, our findings not only identify key genes for IgE-mediated response but also provide preliminary mechanistic insights into GBL.
Dengue virus remains one of the most significant worldwide threats and has yet to yield to effective antiviral therapies. The viral NS2B-NS3 protease is a critical factor for dengue virus infection and is a validated target of therapeutic interest. In this work, we present a computational drug-design and similarity-based prioritization approach to evaluate triterpenoid, amidinium, and flavonoid inhibitors with their structural analogues. We have examined the binding modes, conformational stability, and energetic profiles of the molecules through molecular docking and explicit-solvent molecular dynamics simulations. While parent compounds (Glycyrrhizin, Benzamidine, and Myricetin; abbreviated as GLY, BENZ, and MYR) exhibited moderate to strong binding affinities, their derivatives demonstrated expanded interaction networks and enhanced occupancy within the deeper hydrophobic pockets of the active site. Free-energy landscapes unveiled thermodynamically favourable states with compact minima in GLY- and MYR-based systems. In addition, principal component analysis revealed ligand-dependent modulation of catalytic-site dynamics. MM/PBSA calculations yielded more favourable binding energies for derivatives, especially GLY-D and BENZ-D, driven by non-bonded interactions results identify tractable scaffold families whose members enhance protease stabilization and frame a rational basis for prioritizing candidates toward experimental evaluation for dengue antiviral development. The present study thus offered a thorough framework for finding lead compounds for additional drug optimization and design as NS2B-NS3 protease inhibitors, as well as insightful information about the interactions between active derivatives and DENV NS2B-NS3 protease.
Accurate drug-target affinity (DTA) prediction is essential for virtual screening and early-stage drug discovery. We propose TransDTAP, a multimodal Transformer-based architecture that integrates ligand SMILES sequences, protein amino-acid sequences, molecular physicochemical descriptors, and protein biochemical features within a unified regression framework. A curated dataset comprising 4793 protein targets, 23,531 ligands, and 33,457 experimentally measured IC50 interactions was constructed from ChEMBL and UniProt sources. Under a standard random split, TransDTAP achieved an R2 of 0.7794, MSE of 0.2142, and MAE of 0.2827 on the pIC50 scale (0.8556nM after inverse transformation). Robustness was further evaluated using Bemis-Murcko scaffold-based splitting, where the model maintained competitive performance (R2=0.689), indicating effective generalization to chemically novel scaffolds. Fair benchmarking against re-implemented baseline models demonstrated consistent performance advantages under both random and scaffold settings. Ablation analysis confirmed that predictive improvements arise from complementary integration of sequence-based encoders and descriptor-level features, while SHAP-based interpretability analysis revealed that influential features align with established determinants of ligand-protein binding, including hydrophobicity, polarity, and electrostatic properties. These results demonstrate that multimodal fusion of learned sequence representations and domain-informed descriptors provides a robust, interpretable, and scalable framework for drug-target affinity prediction, supporting its application in large-scale virtual screening pipelines.
This study develops and analyses a mathematical model that represents dengue transmission between human and mosquito populations. The human population is divided into susceptible, infected, and recovered classes, while the mosquito population is divided into susceptible and infected classes. This model captures interaction between humans and vectors with a number of significant epidemiological parameters governing the spread of disease. In devising effective intervention strategies, two optimal control techniques are considered. The first technique applies Pontryagin's Minimum Principle for obtaining analytical optimal control conditions, whereas the second one makes use of an Artificial Neural Network structure that approximates the dynamics of the system and optimizes the control process by a learning-based strategy. Numerical simulations indicate that both methods reduce significantly the infected populations and vectors, and the ANN-based method is more efficient and flexible when dealing with the nonlinear dynamics of the system. The results highlight the importance of integrating traditional mathematical analysis with state-of-the-art computational intelligence methods to facilitate dengue prevention and control.
Analyzing gene expression data is essential for predicting and detecting diseases, including cancer. The data is very repetitive and noisy, which makes it hard to find important information about illnesses. In the past decade, several traditional machine learning and feature selection models have been developed for cancer type classification from gene expression data. Rather than introducing new deep learning primitives, this work presents a principled integration framework that combines sparsity-aware representation learning, structure-inducing spatial embedding, and hierarchical multi-scale attention. This paper presents a method for cancer gene classification based on Hierarchical Attention Assisted Feature Pyramid Network (HA-FPN). The work involved two publicly available datasets. The proposed methodology starts with dimensionality reduction through a variational sparse autoencoder (VSAE), followed by an updated DeepInsight algorithm for image conversion of the input. Next, the classification technique is constructed using the proposed HA-FPN model. Moreover, the improved gradient descent optimization (IGDO) is utilized to change the hyperparameter of the classification model. In addition, the results demonstrate that the model combined with an IGDO outperforms the currently existing methods in terms of accuracy, precision, recall, and F1-score. The method that is presented efficiently brings out different aspects of the data through t-SNE computations. Moreover, the proposed approach is very robust, and hence, it can reach high performance levels on two different datasets.
High-throughput identification of potential ligand-binding pockets is essential for structure-based drug discovery and for anticipating off-target interactions. In this study, we developed an automated workflow to analyze 220,019 protein structures from the Protein Data Bank (PDB) using Fpocket, obtaining 219,762 valid outputs. Fpocket detected 12,360,014 pockets in total, from which 85,132 pockets with PocketScore and DrugScore ≥ 0.5 were retained as having predicted ligandability. All corresponding protein structures were converted into PDBQT format for immediate use with AutoDock Vina, and comprehensive pocket metadata were extracted, including pocket centers, box dimensions, and the identities and 3D coordinates of all amino acid residues lining each pocket. These annotations allow the dataset to be used not only with Vina but also with other docking engines and pocket-based analyses. The resulting PDBPockets resource provides a robust foundation for large-scale inverse docking, structure-based virtual screening, therapeutic repurposing opportunities, and global exploration of protein-ligand interaction space. PDBPockets is available at http://www.khashanlab.org/pdbpockets/.
Atherosclerosis (AS) is a chronic inflammatory disease of the vascular wall driven by a complex interplay between dysregulated immune responses and lipid retention. Its systemic complications remain the leading global cause of mortality. Despite current therapeutic strategies, substantial residual cardiovascular risk persists, largely due to insufficiently targeting of inflammatory and lipid-related pathways. Radix Puerariae (Gegen) is rich in isoflavonoids with reported cardioprotective properties; however, their coordinated anti-atherosclerotic mechanisms have not been fully characterized. Using an integrated network pharmacology framework, five active isoflavonoids (daidzein-4,7-diglucoside, formononetin, 3'-methoxydaidzein, puerarin, and 7,8,4'-trihydroxyisoflavonoid) were identified and predicted to target 79 AS-related proteins. Functional enrichment analyses indicated significant involvement in lipid metabolism, inflammatory response, apoptosis, and fluid shear stress pathways. Multi-algorithm machine learning (LASSO, SVM-RFE, RF, and NNET) consistently identified TNF and MMP9 as dominant diagnostic drivers, while RXRA emerged as a mechanistically relevant consensus target. Molecular docking and 100-ns molecular dynamics simulations indicated stable and favorable binding interactions between the isoflavonoids and their corresponding targets. Notably, daidzein-4,7-diglucoside consistently exhibited the strongest target-binding affinity, surpassing the well-studied puerarin, and is thus proposed as a potential novel lead candidate. Collectively, these findings suggest that the five Radix Puerariae isoflavonoids may exert coordinated anti-atherosclerotic effects through a multi-target, multi-pathway, and multicellular regulatory network, supporting their potential as a mechanism-driven adjunctive therapeutic strategy for AS.
Neurotrophic tyrosine receptor kinase 2 (NTRK2/TRKB) demonstrates oncogenic roles across cancers, with notable significance in gliomas where its overexpression is linked to aggressive clinical phenotypes. Lucitanib (AL3810), a multi-target tyrosine kinase inhibitor, shows unexplored potential for treating NTRK2-driven gliomas. This study employed an integrated approach combining pan-cancer analysis, computational drug screening, and experimental validation to systematically evaluate the oncogenic function of NTRK2 and the therapeutic efficacy of Lucitanib. Multi-omics analysis across 32 cancer types using TCGA/GTEx/CPTAC datasets revealed NTRK2 overexpression in gliomas (GBMLGG/LGG), correlating with poor prognosis (p < 0.01) and implicating AKT signaling and immune microenvironment modulation. Structure-based virtual screening of 26,996 compounds against NTRK2 identified Lucitanib as a high-affinity binder (ΔG < -8 kcal/mol), forming stable interactions with key residues (MET-636, PHE-633), further validated by 500 ns molecular dynamics simulations. In vitro experiments using U251MG glioblastoma cells and primary astrocytes demonstrated that Lucitanib significantly inhibited proliferation (p < 0.001), suppressed invasion and migration via MMP9 downregulation (p < 0.001), and induced apoptosis through Bcl-2/Bax modulation (p < 0.001). Its efficacy was intermediate between methotrexate and the selective NTRK2 inhibitor ana-12. Mechanistically, Lucitanib targeted the NTRK2-AKT-MMP9 axis while preserving immune effector functions. These findings establish NTRK2 as a viable therapeutic target in gliomas and highlight Lucitanib as a novel multi-mechanistic inhibitor with balanced efficacy and favorable pharmacokinetic properties, supporting its further development for clinical translation in NTRK2-overexpressing gliomas.
Recent evidence suggests that linagliptin and its structural derivatives exhibit biological activities beyond glycemic control, highlighting the linagliptin scaffold as a promising platform for multifunctional drug design. In this study, eight novel linagliptin-based azo-imine derivatives (3a-h) were designed, synthesized, and structurally characterized, with the stereochemistry of representative compound 3e confirmed by 2D NMR. Computational analyses (DFT, molecular docking, and molecular dynamics simulations) were employed to support electronic structure elucidation and ligand-protein interaction stability. Specifically, 250 ns molecular dynamics simulations confirmed that the 6XFP-3h complex maintains high structural stability and thermodynamic equilibration throughout the trajectory. All derivatives displayed moderate antioxidant activity (TEAC = 0.46-0.84). Antiproliferative effects were evaluated against A549 and DLD1 cancer cell lines, with WI-38 fibroblasts used to assess selectivity. Compound 3h exhibited the highest potency, with IC50 values of 0.66 µM (A549) and 0.29 µM (DLD1), while compounds 3d and 3e showed cytotoxicity comparable to cisplatin. Selectivity index analysis revealed moderate but discernible cancer cell preference, with compound 3h demonstrating the most favorable selectivity toward DLD1 cells (SI = 3.62). These results suggest that compound 3h represents a prioritized candidate within this series and support linagliptin-based azo-imine derivatives as a tractable framework for further structure-activity relationship optimization.
Human metapneumovirus (hMPV) is a serious global health threat because it causes human respiratory diseases in people of all ages. The complicated dynamics of this virus transmission exist in complications of waning immunity and reinfection that are extremely problematic for mathematical epidemiology. The aim of this research is to develop an integrated approach that combines a fractional-order mathematical model of hMPV transmission with a computational framework that forecasts outbreaks and optimal control, which applies effective control measures and captures complex immunological memory effects. This study presents a novel eight-compartment S0EIsIaHR1S1R2+ mathematical model that incorporates some complex transmission pathways, like naive susceptible, asymptomatic, initially recovered, susceptible with immunity, and recovered with immunity. We examine the model's mathematical validity by demonstrating its well-posedness, boundedness, and non-negativity. We use the next-generation matrix to estimate the basic reproduction number R0, as well as the hMPV-free and endemic equilibrium points. Moreover, we conduct a complete stability analysis of both disease-free and endemic equilibrium. According to a global sensitivity analysis, the transmission rate β (+0.788), symptomatic recovery rate γs (-0.394), and symptomatic proportion p (+0.234) have the greatest effects on R0. Also, inspired by physics-informed neural networks, a modified disease-informed neural network architecture is designed to overcome the shortcomings of traditional computational methods by including physical constraints into the training procedure. Using synthetic data for validation, our analysis demonstrates an excellent predictive performance with R2 values ranging from 0.8500 to 0.9800 across all test scenarios, and the average error accuracy of compartments stays extremely low with MSE≤0.0443 across multiple independent trials. Furthermore, our optimal control analysis validates that a synergy of intervention strategies achieves a 47.3% reduction in the peak hospitalizations of the model compared to the baseline model estimates. These results offer practical insights for public health policies, highlighting the significance of enhanced screening for symptoms and asymptomatic conditions, immunity maintenance initiatives, and effective control methods. This unified framework gives public health officials a useful tool to test intervention methods and forecast outbreaks using this innovative method that links mathematical models with real-world information.
Using two-sample Mendelian randomization with genome-wide association study (GWAS) data from European-ancestry populations, we identified a significant mediation pathway linking genus Akkermansia to type 2 diabetes (T2D) via phosphatidylcholine (O-18:1_20:3). GWAS data were obtained from the Dutch Microbiome Project (gut microbiota), Finnish cohorts (lipid metabolites), and FinnGen (T2D). Lower levels of phosphatidylcholine (O-18:1_20:3), a protective lipid, mediated the association. The indirect effect was IE = 0.0206 (SE = 0.00187; 95% CI 0.01694-0.02428; P = 3.06 ×10⁻²⁸; OR = e^IE = 1.0208), with a total effect TE = 0.0795 (SE = 0.00767; OR = 1.0827), corresponding to a mediation proportion of 25.9%. In contrast, higher-level Verrucomicrobiaceae taxa (phylum, class, order, and family aggregates) showed weaker, non-significant mediation (95% confidence intervals including zero). We also observed that the microbial pathway PWY.7221(de novo guanosine nucleotide biosynthesis) was inversely associated with T2D (OR = 0.903; 95% CI 0.852-0.957). ediation analyses further supported phosphatidylcholine (O-18:1_20:3) as a key lipid linking genus Akkermansia to T2D. Collectively, these findings support a causal gut microbiota-lipid-T2D axis and highlight phosphatidylcholine (O-18:1_20:3) as a potential target for therapeutic strategies that modulate lipid metabolism through microbiome interventions.
The MYB family genes encode transcription factors that orchestrate cellular proliferation, differentiation, and apoptosis; however, the regulatory architectures underlying their expression remain incompletely understood. Here, a genome-wide expression quantitative trait loci (eQTL) analysis of MYB family genes was conducted in a mixed-model framework using transcriptome and genotype data from lymphoblastoid cell lines (LCLs) of 373 European individuals. We identified 51 significant eQTLs (P<5 × 10⁻⁸), revealing a strikingly divergent regulatory architecture between MYB and MYBL2. Specifically, all 26 MYBL2-associated eQTLs were cis-acting, whereas MYB expression was regulated exclusively by 25 trans-eQTLs. These distinct cis- and trans-regulatory patterns were consistently replicated across multiple cohorts representing diverse ancestral backgrounds, supporting the robustness of the observed architecture. MYBL2 cis‑eQTLs demonstrated broad cross-tissue reproducibility, including in immune- and metabolism-related tissues, highlighting the biological coherence and broad regulatory activity of MYBL2. Transcriptome‑wide eGenes associated with MYB family eQTLs supported a ubiquitously acting regulatory profile for MYBL2, in contrast to a context‑dependent regulatory role for MYB. Integrative downstream analyses incorporating genome-wide association study data, including Bayesian colocalization and Mendelian randomization, provided evidence that regulatory variants modulating MYB family gene expression contribute to immunometabolic traits. Collectively, these findings elucidate distinct cis-trans regulatory divergence of MYB family genes and a causal role of their regulatory variants in immunometabolic phenotypes. This work establishes a conceptual and analytical framework for future functional investigations aimed at elucidating the mechanisms by which MYB family genes and their regulatory architectures shape complex human traits.
Metagenome-resolved metagenomics refers to the recovery of metagenome-assembled genomes from the metagenomic datasets. It is a multi-step and laborious process that requires substantial computational resources and technical expertise. Though various semi-automated pipelines have been developed to automate the recovery process, high computational requirements remain a major bottleneck. Since de novo assembly is the key step that consumes higher computational time and resources, optimizing this step can address the underlying challenges. Hence, to address these limitations, we introduce MetaBolt, an automated Nextflow-based pipeline designed for the rapid recovery of metagenome-assembled genomes from short-read metagenomic datasets. Based on an empirically optimized set of k-mers for MEGAHIT-based assembly, this pipeline offers a unique solution. When tested on both real and simulated metagenomic datasets, it consistently exhibited efficient performance within reduced computational time. From gut metagenomes, MetaBolt recovered MAGs at a 2.3 and 3.8-times faster rate than nf-core/mag and MetaWRAP, respectively, while recovering ∼5% more high-quality MAGs compared to the other two pipelines. Whereas, in the case of real metagenome samples, it reduced the computational times to 2-4%, particularly for low-biomass samples. By integrating optimized assembly parameters with automated workflow management, MetaBolt lowers computational barriers to genome-resolved metagenomics without compromising output quality. MetaBolt is available on the web at https://github.com/muneebdev7/metabolt.
One of the biggest causes of cancer-related mortality worldwide is still colon adenocarcinoma (COAD). Therefore, it is essential to investigate new therapy strategies. This study presents a multi-scale, hypothesis-generating framework that integrates network pharmacology with multi-omics data and quantum chemical analyses to systematically explore the repositioning potential of amodiaquine and desethylamodiaquine in colon adenocarcinoma. Eleven gene targets that were found to be shared across medication, gene expression, and illness databases included several computationally significant central hub genes with high network centrality, including BTK, SIGMAR1, SYK, and KCNH2. Amodiaquine and Desethylamodiaquine both showed high binding scores (-7.9 to -12.6 kcal·mol⁻¹) across all the selected proteins. The protein-ligand complexes' structural stability was confirmed by molecular dynamic simulations lasting more than 100 ns; RMSD, RMSF, and radius of gyration analyses collectively indicate that the protein-ligand complexes maintain structural stability, compactness, and limited conformational fluctuation throughout the simulation, reflecting preserved protein integrity rather than direct binding strength. While density functional theory investigations indicated stable geometries with high electronic polarizability, MM-PBSA calculations yielded binding free energies within a range of ΔG_bind = -59.0 to -67.0 kcal·mol⁻¹ . These results suggest that amodiaquine derivatives have anti-COAD actions due to their disruption of important immune-regulatory and apoptotic pathways. These findings computationally prioritize amodiaquine and desethylamodiaquine as candidate multi-target interactors in colon adenocarcinoma, warranting further experimental investigation rather than implying established therapeutic efficacy.