The evolution of transcriptomic technologies requires effective translation between legacy and modern platforms to fully leverage historical data. We present GANomics, a generative adversarial network (GAN) framework that enables bidirectional translation between microarray and RNA-seq data using a small number of samples profiled by both technologies. By integrating paired and unpaired samples through a pair-aware feedback loss, GANomics enforces one-to-one transcript mappings while preserving global gene expression distributions. Applied to a neuroblastoma cohort (n = 498; 10,042 genes) and benchmarked against six alternative methods, GANomics achieved high per-sample correlations ( > 0.96) between real and synthetic data with as few as ten paired profiles. With fifty paired profiles, it accurately recapitulated differential expression, maintained pathway-level rankings, and enabled a cross-platform classifier to transfer comparable to real data. Consistent performance across five additional benchmark datasets confirmed its robustness. By bridging legacy and contemporary transcriptomic data while retaining biological consistency, GANomics facilitates scalable data integration for biomarker reuse and the expansion of transcriptomic resources in clinical applications.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) therapeutic monoclonal antibody bebtelovimab (BEB) was designed to target the spike protein of SARS-CoV-2 and block its entry into host cells via the angiotensin-converting enzyme 2 (ACE2) receptor. In this study, we used molecular modeling to predict interactions of six clinically observed viral S-protein mutants with human ACE2, as well as BEB. We found that some of these mutants could reduce the activity of BEB by changing the relative binding affinity of the viral S-protein with BEB and ACE2. This may contribute to the resistance in viral mutants to BEB. Additionally, we demonstrated that this effect is highly dependent on the dynamic behavior of the involved proteins and their structural variations. We further assessed BEB's activity against SARS-CoV-2 Omicron subvariants BQ.1 and BQ.1.1. Our study provides insights into how mutations may affect BEB's activity, offering valuable information for the development of future therapeutics against rapidly mutating pathogens.
Reference materials should be adopted as a common calibrator for multiomics measurement and co‑profiled with study samples. Multiomics results should be reported as sample‑to‑reference ratios so that they are reproducible and suitable for artificial intelligence tools.
Minimizing the use of animals as concurrent controls in in vivo studies directly supports the 3Rs principles of Replacement, Reduction, and Refinement. Current virtual control group (VCG) approaches primarily rely on historical control data, but their utility may be limited by cross-study variability from differences in study design, laboratory practices, and data heterogeneity. We propose GanCtrl, a generative AI approach to infer study-specific control data directly from time-matched treatment data. By generating synthetic controls analogous to concurrent controls, GanCtrl aims to mitigate biological, temporal, and technical biases inherent in VCG approaches based on historical control data. GanCtrl was applied to rat repeat-dose toxicity studies to simulate 38 clinical pathology endpoints under control conditions using corresponding treatment data. Synthetic controls closely approximated real concurrent controls, with differences smaller than intra- and inter-laboratory baseline variation and comparable to biological replicate variability, while also preserving the typical magnitude and distribution of control responses across studies. Importantly, synthetic controls enabled detection of drug-induced clinical pathology signals and maintained biologically relevant endpoint relationships, such as ALT-AST. For practical utility, toxicity assessments using GanCtrl-derived synthetic controls were compared with those using real concurrent controls and benchmarked against VCGs constructed from single-laboratory or combined multi-laboratory historical data. Although both approaches performed comparably in the single-laboratory setting, GanCtrl outperformed VCGs when data from multiple laboratories were combined. These findings suggest that GanCtrl offers a potential approach for generating VCGs that may reduce the use of concurrent control animals and advance the 3Rs.
Social media platforms host abundant and timely descriptions of medication experiences that can complement traditional pharmacovigilance systems. Yet the linguistic informality of these data presents a major challenge for mapping adverse drug event (ADE) expressions to standardized medical terminology. In this study, we developed BERT-based language models to classify ADE mentions from social media into MedDRA System Organ Classes (SOCs). Using the SMM4H and CADEC corpora, as well as their combination, we performed 20 iterations of 20% holdout validation for 3-, 6-, 22-, and 25-SOC classification tasks with a selected fixed training configuration (learning rate, batch size, and training epochs) based on training-loss convergence. The models achieved accuracies ranging from 75% to 94%, demonstrating strong performance for SOC-level classification of noisy and informal ADE expressions under the evaluated settings. These results are based on a controlled mention-level evaluation using deduplicated adverse drug event strings and do not establish document-level or real-world deployment generalization. This work provides a systematic evaluation of BERT-based models for SOC-level classification of ADEs and demonstrates consistent performance within the evaluated datasets and label granularities. While direct comparison with prior studies is limited by differences in datasets and evaluation protocols, the results demonstrate that transformer-based models can effectively classify ADEs into SOCs. These findings support the use of transformer-based normalization for SOC-level aggregation of user-reported adverse events and their integration into large-scale social media pharmacovigilance pipelines as a downstream component under controlled conditions.
Regulatory efforts to reduce animal testing have accelerated the adoption of new approach methodologies (NAMs). We present an integrated framework combining a published in silico model, DeepDILI (a deep learning model for predicting drug-induced liver injury [DILI]) with six independent human liver spheroid assays (in vitro NAMs) for predicting DILI. Integration was performed at the decision level using a Confidence-Guided Integration (CGI) rule, whereby DeepDILI predictions were applied only when confidence exceeded a predefined threshold. CGI outperformed spheroid assays alone in five of six datasets, improving accuracy by 9-55%. Drugs with concordant DILI classifications by both methods represented a high-confidence subset, achieving substantially high predictive performance (86-100% accuracy), while most of the gain in prediction with CGI was for discordant cases with up to 53% error reduction. CGI produced results comparable to a Liver-Chip system in an eight-drug evaluation. The study demonstrated the benefits of combining results from orthogonal NAMs, supporting the 3Rs (Replacement, Reduction and Refinement) and the FDA's 2025 road map for alternative methods.
Screening tests for disease have their performance measured through sensitivity and specificity, which inform how well the test can discriminate between those with and without the condition. Typically, high values for sensitivity and specificity are desired. These two measures of performance are unaffected by the outcome prevalence of the disease in the population. Research projects into the health of the American Indian frequently develop Machine learning algorithms as predictors of conditions in this population. In essence, these models serve as in silico screening tests for disease. A screening test’s sensitivity and specificity values, typically determined during the development of the test, inform on the performance at the population level and are not affected by the prevalence of disease. A screening test’s positive predictive value (PPV) is susceptible to the prevalence of the outcome. As the number of artificial intelligence and machine learning models flourish to predict disease outcomes, it is crucial to understand if the PPV values for these in silico methods suffer as traditional screening tests in a low prevalence outcome environment. The Strong Heart Study (SHS) is an epidemiological study of the American Indian and has been utilized in predictive models for health outcomes. We used data from the SHS focusing on the samples taken during Phases V and VI. Logistic Regression, Artificial Neural Network, and Random Forest were utilized as in silico screening tests within the SHS group. Their sensitivity, specificity, and PPV performance were assessed with health outcomes of varying prevalence within the SHS subjects. Although sensitivity and specificity remained high in these in silico screening tests, the PPVs’ values declined as the outcome’s prevalence became rare. Machine learning models used as in silico screening tests are subject to the same drawbacks as traditional screening tests when the outcome to be predicted is of low prevalence.
Microphysiological systems (MPS) is an emerging in vitro technology designed to recapitulate human physiology for applications in drug development and safety assessment. Compared to conventional in vitro systems, MPS may contain multiple types of cells and display dynamic and mechanical features of organ microenvironments. As part of New Approach Methodologies (NAMs), MPS systems are expected to contribute to reducing reliance on animal testing by providing human-relevant models that align with the principles of replacement, reduction, and refinement. This review discusses the advantages of MPS over conventional in vitro systems for drug absorption, distribution, metabolism, and excretion (ADME) and toxicity evaluation. We then revisit organ specific MPS platforms used in ADME and toxicity studies. Next, we briefly touched the reproducibility of MPS across different systems. Finally, we provide our perspectives and considerations on employing MPS in regulatory applications.
There is significant interest in combining adverse outcome pathways (AOPs) with Bayesian networks (BNs) because of their shared representation using directed acyclic graphs (DAGs). However, it has not been verified empirically whether AOP networks are mathematically congruent with BNs. Furthermore, important properties for BNs, such as Markov blankets, have not been emphasized, which is a missed opportunity for simplifying and optimizing the model. Here, we summarize the connection between AOP networks and BNs and explore the implications for toxicity modeling. We also present a case study in drug-related liver toxicity. Our results confirm that AOP networks are congruent mathematically with BNs, with incorporation of the mathematical properties of BN leading to significantly simplified and more efficient models.
Drug-induced cardiotoxicity (DICT) is a significant challenge in drug development and public health. DICT can arise from various mechanisms; New Approach Methods (NAMs), including quantitative structure-activity relationships (QSARs), have been extensively developed to predict DICT based solely on individual mechanisms (e.g., hERG-related cardiotoxicity) due to the availability of datasets limited to specific mechanisms. While these efforts have significantly contributed to our understanding of cardiotoxicity, DICT assessment remains challenging, suggesting that approaches focusing on isolated mechanisms may not provide a comprehensive evaluation. To address this, we previously developed DICTrank, the largest dataset for assessing overall cardiotoxicity liability in humans based on FDA drug labels. In this study, we evaluated the utility of DICTrank for QSAR modeling using five machine learning methods─Logistic Regression (LR), K-Nearest Neighbors, Support Vector Machines, Random Forest (RF), and extreme gradient boosting (XGBoost)─which vary in algorithmic complexity and explainability. To reflect real-world scenarios, models were trained on drugs approved before and within 2005 to predict the DICT risk of those approved thereafter. While we observed no clear association between prediction performance and model complexity, LR and XGBoost achieved the best results with DICTrank. Additionally, our significant-feature analyses with RF and XGBoost models provided novel insights into DICT mechanisms, revealing that drug properties associated with descriptors such as "structural and topological", "polarizability", and "electronegativity" contributed significantly to DICT. Moreover, we found that model performance varied by therapeutic category, suggesting the need to tailor models accordingly. In conclusion, our study demonstrated the robustness and reliability of DICTrank for cardiotoxicity prediction in humans using machine learning methods.
Trust is key in AI for regulatory science, but its definition is debated. If AI models use different features yet perform similarly, which should be trusted? If scientific theories must be testable, how critical is explainability? At the Global Summit on Regulatory Science (GSRS24), regulators agreed that successful AI adoption requires ongoing dialogue, adaptability, and AI-trained personnel to harness its potential for regulatory responsibilities in the evolving 21st-century landscape.
The FDA's 'Roadmap to Reducing Animal Testing in Preclinical Safety Studies' supports the regulatory advancement of new approach methodologies (NAMs). Focusing on overlapping drugs, this review compared the performance of in vitro NAMs, animal studies, and microphysiological systems (MPS) in predicting drug-induced liver injury (DILI). We observed considerable variability among in vitro NAMs and their potential advantages over animal models. Moreover, limited MPS data hindered meaningful comparison. To enable objective and systematic assessment of NAMs, we propose DILIference, a curated reference drug list compiled from the literature to guide the development and benchmarking of DILI-predictive NAMs.
Hepatotoxicity can lead to the discontinuation of approved or investigational drugs. The evaluation of the potential hepatoxicity of drugs in development is challenging because current models assessing this adverse effect are not always predictive of the outcome in human beings. Cell lines are routinely used for early hepatotoxicity screening, but to improve the detection of potential hepatotoxicity, in vitro models that better reflect liver morphology and function are needed. One such promising model is human liver microtissues. These are spheroids made of primary human parenchymal and nonparenchymal liver cells, which are amenable to high throughput screening. To test the predictivity of this model, the cytotoxicity of 152 FDA (US Food & Drug Administration)-approved small molecule drugs was measured as per changes in ATP content in human liver microtissues incubated in 384-well microplates. The results were analyzed with respect to drug label information, drug-induced liver injury (DILI) concern class, and drug class. The threshold IC50ATP-to-Cmax ratio of 176 was used to discriminate between safe and hepatotoxic drugs. "vMost-DILI-concern" drugs were detected with a sensitivity of 72% and a specificity of 89%, and "vMost-DILI-concern" drugs affecting the nervous system were detected with a sensitivity of 92% and a specificity of 91%. The robustness and relevance of this evaluation were assessed using a 5-fold cross-validation. The good predictivity, together with the in vivo-like morphology of the liver microtissues and scalability to a 384-well microplate, makes this method a promising and practical in vitro alternative to 2D cell line cultures for the early hepatotoxicity screening of drug candidates.
Drug-induced liver injury (DILI) is of great concern in drug development and public health. DILIrank 1.0, a widely used public dataset that ranks FDA-approved drugs by their potential to cause DILI, has significantly enabled the development of new methods for improved DILI assessment. Here, we introduce DILIrank 2.0, an essential update of DILIrank 1.0, to capture new non-biologics drug and liver-related adverse event data generated since its inception 15 years ago. Using DILIrank 2.0, we also observe changes in the DILI profiles of approved drugs following the introduction of different FDA regulatory programs. DILIrank 2.0 is an up-to-date resource that offers opportunities for supporting DILI safety assessment, predictive model development, and evaluation of new approach methods (NAMs).
Computational toxicology plays an important role in risk assessment and drug safety. The field has been traditionally dominated by Quantitative Structure-Activity Relationships (QSARs), which predict toxicological effects based solely on chemical structure. Although QSARs have achieved successes, their structure reliance limits drug toxicity predictions, where small structural modifications may cause major toxicity changes. Advances in artificial intelligence (AI), especially text embedding and generative AI, provide an opportunity to enhance toxicity predictions by leveraging broader chemical knowledge and its integration with structural data. In this study, we propose a novel framework, Quantitative Knowledge-Activity Relationships (QKARs), which predicts toxicity using domain-specific knowledge. We developed QKAR models for two drug toxicity endpoints, drug-induced liver injury (DILI) and drug-induced cardiotoxicity (DICT), using three different knowledge representations with varying levels of knowledge. The representations based on comprehensive knowledge of the drugs yielded better prediction than those with simpler knowledge. Five machine learning algorithms of distinct complexity were applied in QKAR models, and we observed little association between model complexity and performance. Further, we evaluated QKARs against QSARs on the same endpoints using identical datasets. We found that QKARs consistently outperformed QSARs for DILI and DICT. Notably, QKARs demonstrated better capability than QSARs in differentiating drugs with similar structures but different liver toxicity profiles. We also investigated integrating knowledge-based and structure-based representations, Q(K + S)ARs, for further enhanced prediction accuracy. Our findings demonstrate the potential of QKARs as a robust alternative to QSARs, offering additional opportunities in drug toxicity assessments by leveraging both domain-specific knowledge and structural data.
Drug-induced liver injury (DILI) is a significant concern with prescription medications and supplements. Accordingly, it is crucial to develop tools and approaches that can predict DILI likelihood of existing medications and supplements, as well as potential drug candidates under development. The complexity of liver injury mechanisms and the limited availability of DILI data hamper the development of robust predictive models. In order to overcome these challenges, this study investigated enriching machine learning/artificial intelligence (ML/AI) models that predict the risk of DILI using drug structural parameters along with rat liver transcriptomics data, quantum mechanics-derived features of the drug molecules, and metrics for interspecies variability of drug exposure. The enrichment of ML/AI models with such features dramatically improved ML/AI models' DILI predictive ability, even in a severely data-limited scenario. The approach used in the study, especially the incorporation of knowledge-based features to enrich AI models, holds tremendous promise for not only assessing safety and toxicity assessments of drug candidates but also in other aspects such as target engagement and efficacy of these candidates, early in the development phase.
Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.
Drug adverse events (AEs) represent a significant public health concern. US Food and Drug Administration (FDA) drug labeling documents are an essential resource for studying drug safety such as assessing a drug’s likelihood to cause certain organ toxicities; however, the manual extraction of AEs is labor-intensive, requires specialized expertise, and is challenging to maintain, due to frequent updates of the labeling documents. To automate the extraction of AE data from FDA drug labeling documents, we developed a workflow based on AskFDALabel, a large language model (LLM)-powered framework, and its demonstration in drug safety studies. This framework incorporates a retrieval-augmented generation (RAG) component based on FDALabel to enhance standard LLM inference. Key steps include (1) selection of a task-specific template, (2) FDALabel database querying, and (3) content preparation for LLM processing. We evaluated the performance of the framework in three benchmark experiments, including drug-induced liver injury (DILI) classification, drug-induced cardiotoxicity (DICT) classification, and AE term recognition. AskFDALabel achieved F1-scores of 0.978 for DILI, 0.931 for DICT, and 0.911 for AE annotation, outperforming other traditional methods. It also provided cited labeling content and detailed explanations, facilitating manual verification. AskFDALabel exhibited high consistency with human AE annotation, particularly in classifying and profiling DILI and DICT. Thus, it can significantly enhance the efficiency and accuracy of AE annotation, with promising potential for advanced AE surveillance and drug safety research.
Microphysiological systems (MPSs) are emerging in vitro technologies designed to recapitulate human physiology for applications in drug development and safety assessment. Compared with conventional in vitro systems, MPSs may contain multiple types of cells and display dynamic and mechanical features of organ microenvironments. As part of new approach methodologies, MPSs are expected to contribute to reducing reliance on animal testing by providing human-relevant models that align with the principles of replacement, reduction, and refinement. This review discusses the advantages of MPSs over conventional in vitro systems for drug absorption, distribution, metabolism, and excretion and toxicity evaluation. We then systematically examines organ-specific MPS platforms used in absorption, distribution, metabolism, and excretion and toxicity studies. Next, we briefly evaluated the reproducibility of MPSs across different systems. Finally, we provide our perspectives and considerations on employing MPSs in regulatory applications. SIGNIFICANCE STATEMENT: This minireview provides an overview of the current trends in the field of microphysiological systems within the framework of new approach methodologies. This review will give readers insights into key differences between microphysiological systems and other in vitro methods in drug absorption, distribution, metabolism, and excretion, followed by recent advances in several organ chips for drug evaluation.