ABSTRACT Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA . KEY POINTS CoTRA is an open-source R/Shiny framework that provides comprehensive downstream bulk and single-cell RNA-seq analysis within a single graphical environment, reducing the need to move between separate analysis tools. CoTRA retains transparency and user control by exposing key analytical parameters and providing alternative established methods for differential expression, marker detection, trajectory inference, pathway activity analysis, and other analytical steps. CoTRA supports reproducible and data-controlled analysis through local, institutional server, or HPC execution without mandatory external data submission, together with extensive export, reporting, processed-object, and session-information capabilities. Comparative assessment showed full support for 46 of 49 predefined functionality criteria, while retinal case studies demonstrated strong bulk RNA-seq reproducibility, including 1,947 concordant DEGs and Pearson r = 0.988 for log2 fold changes, and advanced cell-type-resolved scRNA-seq analysis.
EcoDrug PLUS (EcoDrug+; https://ecodrugplus.helsinki.fi/) is a freely available and publicly accessible database, established to facilitate environmental risk assessment of pharmaceuticals and other bioactive substances, including veterinary medicines and pesticides. EcoDrug+ advancements on the original ECODrug database include a more extensive and intuitive graphical user interface for investigating the potential for chemicals to interact with protein targets based on their conservation with human drug, veterinary, and pesticide targets across 180 phylogenetically diverse wildlife taxa. EcoDrug+ integrates genomic and chemoinformatic data from open-access sources for ~7200 pharmaceuticals, 34 000 agrochemicals, 61 000 human metabolites, and 5800 other bioactive chemicals. Advanced search capabilities of EcoDrug+ include the ability to interrogate the database via text queries, chemical structure drawings, target protein sequence BLAST, and/or specified mechanisms of action. Chemical compound data are organized into clusters to facilitate the exploration of similar groups using interactive knowledge graphs. The integration of effects-based knowledge informs on appropriate endpoints and susceptible species for the testing of drugs (and other bioactive chemicals). Georeferenced measured environmental concentrations (for n = 266 chemicals) furthermore provide relevant exposure data for testing and environmental risk analysis.
Abstract Background Disease-centered knowledge graphs (KGs) support drug repurposing and precision medicine research, yet many remain static after release while primary databases and literature continue to expand. PrimeKG (Precision Medicine Knowledge Graph) is a widely used multimodal KG providing a holistic view of diseases. However, its public release reflects a June 2021 data cutoff, omitting several years of subsequent data growth. This lag is especially consequential for rare diseases, where mechanistic and therapeutic evidence often remains scattered across publications rather than structured resources. Finding We present PrimeKG-Plus, a refreshed, rare-disease–enriched release of PrimeKG, rebuilt from all twenty original data resources updated to their December 2025 releases and three additional resources: OpenTargets, RepurposeDrugs, and nSIDES. Beyond synchronizing biomedical databases, PrimeKG-Plus captures approximately five years of previously unavailable rare-disease knowledge from the biomedical literature, curated from 637 PubMed abstracts and PubMed Central full-text articles using a language-model-assisted workflow focused on four rare neurological disorders: Canavan disease, Niemann–Pick disease type C, Tay–Sachs disease, and Batten disease. Extracted relations were refined through entity normalization, UMLS synonym mapping, embedding-based similarity ranking, and human expert review. Network topology analysis showed improved indirect drug–disease connectivity across three to six hops and added 447,288 drug–protein–disease paths linking previously unreachable drug–disease pairs. Temporal validation using drug approval records identified 55 molecular entities approved after the original PrimeKG June 2021 cutoff, 46 of which were absent from the original graph. Conclusion PrimeKG-Plus restores the temporal relevance of PrimeKG, providing an updated resource for drug repurposing, rare-disease research, and downstream machine-learning applications.
Understanding the molecular vulnerabilities associated with Fanconi anemia (FA) is essential for identifying therapeutic opportunities and elucidating the mechanisms underlying disease progression and cancer predisposition. However, progress in this area remains constrained by limited availability of representative FA cellular models. To address this challenge, we defined an FA-like cellular state by identifying cancer cell lines exhibiting high-dependency on core FA pathway genes, and integrated CRISPR-Cas9 gene essentiality data at multiple molecular layers, including mutation, copy number alterations, mRNA expression, and independent patient-derived transcriptomic datasets. Functional enrichment analyses highlighted biological pathways previously implicated in FA pathogenesis, most notably aldehyde detoxification, cholesterol/fatty acid metabolism, and androgen signaling. Analysis of LINCS-L1000 perturbational transcriptomics resource identified compounds, capable of reversing the FA-associated transcriptional signature, further supporting the pharmacological tractability of the identified molecular vulnerabilities. In addition, drug-target affinity analysis prioritized aldehyde-metabolizing enzymes, including ALDH1A1 and ALDH2, as potentially druggable candidates. Notably, disulfiram demonstrated predicted high-affinity interactions with multiple proteins involved in aldehyde and lipid metabolism, including ALDH1A1, ALDH2, and MGLL, supporting its potential for further investigation in FA-related settings. Although additional validations are required, the identified vulnerabilities and candidate targets provide a foundation for future mechanistic and therapeutic investigations in FA and FA-associated malignancies.
Ocular diseases such as age-related macular degeneration, glaucoma, diabetic retinopathy, and inherited retinal dystrophies are leading causes of vision loss worldwide, yet existing databases often address only limited aspects of these disorders. To fill this gap, we developed the Ocular Disease Database (ODDB), a web-based resource that integrates genes, biomarkers, variants, and drugs associated to ocular diseases. Data were systematically collected through literature mining of PubMed-indexed journals, the NCBI Gene Expression Omnibus (GEO), and drug regulatory agency datasets. Multi-omics, experimental, and clinical information were harmonized using standardized integration workflows. The database is organized according to two complementary ontologies: one based on the anatomical site of pathology (cornea, retina, optic nerve) and another on gene inheritance pattern. ODDB currently covers over 170 ocular diseases, more than 1190 genes, 2400+ variants, and 386 drugs, including both approved and investigational compounds. Each record includes detailed annotations of associated genes, variants, therapeutic targets, and mechanisms of action. The platform supports interactive querying and network-based visualization of disease–gene–drug relationships. All data was internally validated for accuracy and are compliant with FAIR principles, ensuring accessibility and interoperability. ODDB () provides a comprehensive and standardized reference for exploring molecular mechanisms and therapeutic opportunities in ocular diseases. ### Competing Interest Statement The authors have declared no competing interest. RCF, 346295
Drug discovery is a complex, time-intensive, and costly process, often requiring more than a decade and substantial financial investment to bring a single therapeutic to market. Drug repurposing, the systematic identification of new indications for existing approved drugs, offers a cost-effective and expedited alternative to traditional pipelines, with the potential to address unmet clinical needs. In this study, we present a comparative analysis of drug-target interaction data from three extensively curated resources: ChEMBL, BindingDB, and GtoPdb, evaluating their release histories, curation methodologies, and coverage of approved and investigational compounds and targets. To facilitate therapeutic interpretation, we manually classified ChEMBL targets into 12 high-level biological families and mapped 817 clinically approved drug indications into 28 broader therapeutic groups. This structured framework enabled a systematic profiling of physicochemical properties among approved drugs across therapeutic categories. Our analyses revealed associations between physicochemical characteristics and therapeutic groups, providing practical guidance for indication-specific compound prioritization and refining the repurposing studies. We also examined cross-indication drug approvals to identify areas with high repurposing potential. Finally, we implemented a pathway-based computational pipeline to predict repositioning opportunities for FDA-approved drugs across 10 major cancer types, demonstrating its adaptability to other disease contexts. Overall, this work consolidates drug-target data and computational repurposing into a data-driven framework that advances drug discovery and translational applications.
Reliable and reproducible drug screening experiments are essential for drug discovery and personalized medicine. We demonstrate how systematic experimental errors in drug plates negatively impact data reproducibility, and that conventional quality control (QC) methods based on plate controls fail to detect these spatial errors. To address this limitation, we developed a control-independent QC approach that uses normalized residual fit error (NRFE) to identify systematic artifacts in drug screening experiments. Analysis of >100,000 duplicate measurements from the PRISM pharmacogenomic study revealed that NRFE-flagged experiments show 3-fold lower reproducibility among technical replicates. By integrating NRFE with QC methods to analyze 41,762 matched drug-cell line pairs between two datasets from the Genomics of Drug Sensitivity in Cancer project, we improved the cross-dataset correlation from 0.66 to 0.76. Available as an R package at https://github.com/IanevskiAleksandr/plateQC, plateQC provides a robust toolset for enhancing drug screening data reliability and consistency for basic research and translational applications.
In the rapidly advancing landscape of drug discovery and repurposing, efficient access and integration of chemical and bioactivity data from public repositories have become essential. To address this need, we developed two complementary annotation pipelines (KNIME- and Python-based) that automate the extraction and integration of curated chemical and bioactivity data from public repositories. These pipelines support any user-provided compound library, enabling reproducible workflows that integrate data from heterogeneous sources such as ChEMBL and PubChem. As part of the REMEDi4ALL project, with the aim of establishing a European platform for drug repurposing, we validated our framework using a harmonized subset of the Specs repurposing collection, which includes >5000 compounds available at the partner institutes. We also developed two interactive dashboards that support multilayered analyses and visualization by integrating chemical properties, bioactivity profiles, and relational data. Our results demonstrate that this framework streamlines the collection of harmonized data and facilitates analyses that are critical for drug repurposing efforts, while remaining versatile for broader applications in drug discovery. Moreover, the analysis of the annotations reveals that the Specs subset includes chemical scaffolds representative of a significant portion of approved drugs and compounds undergoing clinical evaluation, underscoring its potential as a rich source of drug repurposing candidates. Both pipeline protocols are publicly available online, and the dashboards are open access.
A key challenge in drug development is identification of druggable targets, the modulation of which attenuates disease progression, while avoiding inhibition of proteins that lead to dose-limiting toxicities. Here, we investigate a drug target Casein kinase 2 (CK2) - a serine/threonine kinase implicated in cancer, for which existing molecules have so far failed clinical trials. Using molecular and pharmacoepidemiology approaches, we show that molecules targeting CDK kinase family members CDK1/2/7/9 - such as the existing CK2 inhibitors - have a higher risk to induce adverse effects or fail in clinical trials. Based on this finding, we establish a machine learning assisted pipeline to redesign more specific and allosteric lead compounds against CK2, with more selective on-target binding and favourable off-target profile. Importantly, we show that such design is possible via machine learning powered, docking assisted discovery pipeline, when standard ML algorithms were combined with an error prediction model. In conclusion, our study reports a simple yet efficient machine learning-powered drug discovery pipeline and novel submicromolar CK2 inhibitors targeted. Importantly, our prediction pipeline was able to achieve a 90% hitrate, significantly reducing the need for subsequent wet-lab validation. ### Competing Interest Statement Jordi Mestres is the founder and research director of Chemotargets. The authors declare no other competing interests.
Repurposing of existing drugs for new indications has attracted substantial attention owing to its potential to accelerate drug development and reduce costs. Hundreds of computational resources such as databases and predictive platforms have been developed that can be applied for drug repurposing, making it challenging to select the right resource for a specific drug repurposing project. With the aim of helping to address this challenge, here we overview computational approaches to drug repurposing based on a comprehensive survey of available in silico resources using a purpose-built drug repurposing ontology that classifies the resources into hierarchical categories and provides application-specific information. We also present an expert evaluation of selected resources and three drug repurposing case studies implemented within the Horizon Europe REMEDi4ALL project to demonstrate the practical use of the resources. This comprehensive Review with expert evaluations and case studies provides guidelines and recommendations on the best use of various in silico resources for drug repurposing and establishes a basis for a sustainable and extendable drug repurposing web catalogue.
Aims/Purpose: To develop and validate a comprehensive web‐based interactive database of ocular diseases, including information of, for instance, associated genes, biomarkers, and medications.Methods: Initially, a thorough literature review was conducted to identify key ocular diseases and their associated genes and biomarkers. Data was then curated from multiple reputable sources, including peer‐reviewed journals, clinical trial repositories and drug regulatory agency databases. Advanced data integration techniques were employed to link various data types. The database interface was designed to allow intuitive navigation and robust querying capabilities, ensuring ease of access and data utilization. The database underwent internal testing to ensure user friendliness as well as data accuracy and reliability.Results: The resulting web‐based database encompasses detailed information on a wide array of ocular diseases. The diseases are categorized based on location of primary pathology in various parts of the eye, such as the cornea, retina, and optic nerve, among others. For each disease, the database provides a comprehensive list of biomarkers and associated genes, derived from scientific literature. Additionally, the database includes information of drug regulatory agency‐approved drugs and drugs currently undergoing clinical trials, with detailed descriptions of their therapeutic targets, mechanisms of action and clinical efficacy. The database also features tools for epidemiological analysis.Conclusions: This user‐friendly database of ocular diseases is a valuable resource for the scientific community, providing a wealth of information that can aid in the advancement of ophthalmology research. By consolidating data of biomarkers, genes, drugs and epidemiology, the database supports basic and clinical research, and facilitates the development of new therapeutic strategies. This initiative sets a precedent for future databases in other medical disciplines.
RNA-based therapies are a rapidly expanding field, offering treatments for a wide range of diseases, including many rare conditions. To date, 24 RNA therapeutics have received FDA approval, with 131 more in clinical trials, underscoring RNA’s growing role in modern medicine. In this context, Bidirectional Encoder Representations from Transformers (BERT) models provide a cost-effective and accurate virtual screening strategy for accelerating RNA-targeted drug discovery. These models take RNA FASTA sequences and compound SMILES strings as inputs and generate predicted binding affinities in nanomolar units. In this study, we introduce DLRNA-BERTa, a RoBERTa-based framework combining RNA-BERTa, pretrained on 9.76 million RNA sequences, with ChemBERTa-v2 for predicting small molecule–RNA interactions. The framework includes six class-specific models, aptamers, repeats, ribosomal RNAs, riboswitches, microRNAs (miRNAs), and viral RNAs, plus a general model for cases where the RNA class is unknown. Proposed DLRNA-BERTa consistently outperforms existing RNA–drug interaction prediction methods. Pearson correlation coefficients achieved are: 0.94 (aptamers), 0.95 (repeats), 0.93 (ribosomal RNAs), 0.94 (riboswitches), 0.95 (viral RNAs), 0.98 (miRNAs), and 0.92 (general model), demonstrating robust performance across RNA classes. Benchmarking against four independent datasets from the ROBIN repository further confirms generalizability. Application of DLRNA-BERTa to 3,492 approved drugs from the ChEMBL database identified 2,859 compounds with predicted affinities (pKd ≥ 6) across 294 RNA targets. As proof of concept, bleomycin is highlighted, supported by literature evidence of RNA-binding activity. A publicly accessible web application is available at , in alignment with FAIR principles. ### Competing Interest Statement The authors have declared no competing interest. Research Council of Finland, https://ror.org/05k73zm37
Motivation:Drug-target interactions (DTIs) play a pivotal role in drug discovery, as it aims to identify potential drug targets and elucidate their mechanism of action. In recent years, the application of natural language processing (NLP), particularly when combined with pre-trained language models, has gained considerable momentum in the biomedical domain, with the potential to mine vast amounts of texts to facilitate the efficient extraction of DTIs from the literature. Results:In this article, we approach the task of DTIs as an entity-relationship extraction problem, utilizing different pre-trained transformer language models, such as BERT, to extract DTIs. Our results indicate that an ensemble approach, by combining gene descriptions from the Entrez Gene database with chemical descriptions from the Comparative Toxicogenomics Database (CTD), is critical for achieving optimal performance. The proposed model achieves an F1 score of 80.6 on the hidden DrugProt test set, which is the top-ranked performance among all the submitted models in the official evaluation. Furthermore, we conduct a comparative analysis to evaluate the effectiveness of various gene textual descriptions sourced from Entrez Gene and UniProt databases to gain insights into their impact on the performance. Our findings highlight the potential of NLP-based text mining using gene and chemical descriptions to improve drug-target extraction tasks. Availability and implementation:Datasets utilized in this study are accessible at https://dtis.drugtargetcommons.org/.
RepurposeDrugs (https://repurposedrugs.org/) is a comprehensive web-portal that combines a unique drug indication database with a machine learning (ML) predictor to discover new drug-indication associations for approved as well as investigational mono and combination therapies. The platform provides detailed information on treatment status, disease indications and clinical trials across 25 indication categories, including neoplasms and cardiovascular conditions. The current version comprises 4314 compounds (approved, terminated or investigational) and 161 drug combinations linked to 1756 indications/conditions, totaling 28 148 drug-disease pairs. By leveraging data on both approved and failed indications, RepurposeDrugs provides ML-based predictions for the approval potential of new drug-disease indications, both for mono- and combinatorial therapies, demonstrating high predictive accuracy in cross-validation. The validity of the ML predictor is validated through a number of real-world case studies, demonstrating its predictive power to accurately identify repurposing candidates with a high likelihood of future approval. To our knowledge, RepurposeDrugs web-portal is the first integrative database and ML-based predictor for interactive exploration and prediction of both single-drug and combination approval likelihood across indications. Given its broad coverage of indication areas and therapeutic options, we expect it accelerates many future drug repurposing projects.
Advances in deep learning are re-defining how visual data is processed and understand by the machines. Vision Transformers (ViTs) have recently demonstrated prominent performance in computer vision related tasks. However, their performance improves with increasing numbers of labeled data, indicating reliance on labeled data. Humanly annotated data are difficult to acquire and thus shifted the focus from traditional annotations to unsupervised learning strategies that learn structures inside the data. In response to this challenge, self-supervised learning (SSL) has emerged as a promising technique. SSL utilize inherent relationships within the data as a form of supervision. This technique can reduce the dependence on manual annotations and offers a more scalable and resource-effective approach to training models. Taking these strengths into account, it is necessary to assess the combination of SSL methods with ViTs, especially in the cases of limited labeled data. Inspired by this evolving trend, this survey aims to systematically review SSL mechanisms tailored for ViTs. We propose a comprehensive taxonomy to classify SSL techniques based on their representations and pre-training tasks. Furthermore, we highlighted the motivations behind the study of SSL, reviewed prominent pre-training tasks, and highlight advancements and challenges in this field. Furthermore, we conduct a comparative analysis of various SSL methods designed for ViTs, evaluating their strengths, limitations, and applicability to different scenarios.
Pharmacogenomics, the study of how an individual's genetic makeup influences their response to medications, is a rapidly evolving field with significant implications for personalized medicine. As researchers and healthcare professionals face challenges in exploring the intricate relationships between genetic profiles and therapeutic outcomes, the demand for effective and user-friendly tools to access and analyze genetic data related to drug responses continues to grow. To address these challenges, we have developed PGxDB, an interactive, web-based platform specifically designed for comprehensive pharmacogenomics research. PGxDB enables the analysis across a wide range of genetic and drug response data types - informing cell-based validations and translational treatment strategies. We developed a pipeline that uniquely combines the relationship between medications indexed with Anatomical Therapeutic Chemical (ATC) codes with molecular target profiles with their genetic variability and predicted variant effects. This enables scientists from diverse backgrounds - including molecular scientists and clinicians - to link genetic variability to curated drug response variability and investigate indication or treatment associations in a single resource. With PGxDB, we aim to catalyze innovations in pharmacogenomics research, empower drug discovery, support clinical decision-making, and pave the way for more effective treatment regimens. PGxDB is a freely accessible database available at https://pgx-db.org/.
MOTIVATION:Drug-target interactions (DTIs) hold a pivotal role in drug repurposing and elucidation of drug mechanisms of action. While single-targeted drugs have demonstrated clinical success, they often exhibit limited efficacy against complex diseases, such as cancers, whose development and treatment is dependent on several biological processes. Therefore, a comprehensive understanding of primary, secondary and even inactive targets becomes essential in the quest for effective and safe treatments for cancer and other indications. The human proteome offers over a thousand druggable targets, yet most FDA-approved drugs bind to only a small fraction of these targets. RESULTS:This study introduces an attention-based method (called as MMAtt-DTA) to predict drug-target bioactivities across human proteins within seven superfamilies. We meticulously examined nine different descriptor sets to identify optimal signature descriptors for predicting novel DTIs. Our testing results demonstrated Spearman correlations exceeding 0.72 (P < 0.001) for six out of seven superfamilies. The proposed method outperformed fourteen state-of-the-art machine learning, deep learning and graph-based methods and maintained relatively high performance for most target superfamilies when tested with independent bioactivity data sources. We computationally validated 185 676 drug-target pairs from ChEMBL-V33 that were not available during model training, achieving a reasonable performance with Spearman correlation >0.57 (P < 0.001) for most superfamilies. This underscores the robustness of the proposed method for predicting novel DTIs. Finally, we applied our method to predict missing bioactivities among 3492 approved molecules in ChEMBL-V33, offering a valuable tool for advancing drug mechanism discovery and repurposing existing drugs for new indications. AVAILABILITY AND IMPLEMENTATION:https://github.com/AronSchulman/MMAtt-DTA.
INTRODUCTION:Mapping the interactions between pharmaceutical compounds and their molecular targets is a fundamental aspect of drug discovery and repurposing. Drug-target interactions are important for elucidating mechanisms of action and optimizing drug efficacy and safety profiles. Several computational methods have been developed to systematically predict drug-target interactions. However, computational and experimental validation of the drug-target predictions greatly vary across the studies. AREAS COVERED:Through a PubMed query, a corpus comprising 3,286 articles on drug-target interaction prediction published within the past decade was covered. Natural language processing was used for automated abstract classification to study the evolution of computational methods, validation strategies and performance assessment metrics in the 3,286 articles. Additionally, a manual analysis of 259 studies that performed experimental validation of computational predictions revealed prevalent experimental protocols. EXPERT OPINION:Starting from 2014, there has been a noticeable increase in articles focusing on drug-target interaction prediction. Docking and regression stands out as the most commonly used techniques among computational methods, and cross-validation is frequently employed as the computational validation strategy. Testing the predictions using multiple, orthogonal validation strategies is recommended and should be reported for the specific target prediction applications. Experimental validation remains relatively rare and should be performed more routinely to evaluate biological relevance of predictions.