BACKGROUND: Toxicology in the 21st Century (Tox21) assay data provide a valuable resource for the prediction of in vivo toxicity using machine learning models. However, the performances of these models previously developed using the pre-existing Tox21 assay data were less than ideal, likely due to insufficient coverage of the biological response space by the assay targets. OBJECTIVES: This study aimed to assess whether expanding the Tox21 portfolio with new assays that probe under-represented targets/pathways related to unanticipated adverse drug effects could improve the predictive capacity of in vitro assay data for in vivo toxicity such as drug-induced liver injury (DILI) and cardiotoxicity (DICT). METHODS: Models were constructed using data from the pre-existing panel of 36 assay targets and the expanded panel of 49 assay targets. A feature selection approach was used to determine the optimal number of assays needed for each model. The models were then applied to predict the potential hepatotoxicity and cardiotoxicity of compounds in the Tox21 10K compound library. RESULTS: For both DILI and DICT prediction, the best-performing models developed using the expanded assay panel required a smaller number of assays to achieve the same level of performance compared to those based on the pre-existing assays. Models constructed by combining both assay data (pre-existing + expanded) and chemical structure consistently outperformed those constructed based on assay data alone but showed similar performance to those constructed based on chemical structure. The compounds predicted to have the highest toxic potential were experimentally verified to demonstrate the effectiveness of our models in identifying new potentially toxic compounds. DISCUSSION: The expansion of the Tox21 assay panel has significantly enhanced the predictive capacity of assay data for predicting the DILI and DICT potential. This improvement underscores the importance of a diverse and comprehensive in vitro assay portfolio in advancing safety assessment.
Beta-1 adrenergic receptors (ADR beta 1) are critical regulators of cardiac function; however, the potential modulation of ADR beta 1 by environmental chemicals remains largely underexplored, raising concerns about unintended impacts on cardiovascular health. We applied a quantitative high-throughput screening (qHTS) approach to identify ADR beta 1 agonists within the Tox21 10K compound library, which includes environmental chemicals, pharmaceuticals, industrial agents, and consumer products, using an HTRF-based cAMP assay in ADR beta 1-overexpressing HEK293 cells. Primary screening of 8,947 unique compounds identified 118 potential ADR beta 1 agonists. Among these, 94 were confirmed and further evaluated for beta-adrenergic receptor subtype selectivity (ADR beta 2 and ADR beta 3) and hERG channel inhibition to assess potential cardiotoxicity liability. Known ADR beta 1 agonists, isoproterenol (EC50, 0.91 nM) and dobutamine (EC50, 10 nM), were identified, supporting the validity of the assay. In addition, several compounds with limited prior ADR beta 1-specific characterization, such as GR 103691 and N,N '-dibenzylethane-1,2-diamine, demonstrated subtype-selective or mixed agonist profiles, with some exhibiting minimal hERG inhibition. These findings expand the catalog of ADR beta 1 modulators and demonstrate the utility of qHTS for identifying chemicals that may affect cardiovascular signaling pathways.
Metals and metalloids are widely used in industrial applications, and increasing experimental and epidemiological evidence has linkded metal exposure to adverse health outcomes. However, the underlying mechanisms for these effects have not been fully understood. As part of the Toxicology in 21st century (Tox21) program, we have screened more than 150 metal-containing compounds and their salt forms across over 90 biological endpoints. In this study, we analyzed the comprehensive toxicity of metal compounds using Tox21 screening data to enhance the understanding of their mechanism and molecular pathways involved in molecular initiating events. Integrated data analysis and in vitro confirmation experiments identified three potential novel targets of metal compounds (i.e., sonic hedgehog pathway, thyroid-stimulating hormone receptor, and thyrotropin-releasing hormone receptor). We also found that mercury- and tin-containing substances were highly bioactive. Furthermore, cell painting analysis uncovered metal-induced bioactivity could be classified into two patterns depending on their respective associations to mitochondrial-related morphology changes. Our results provide a comprehensive analysis of metals-association bioactivity data within Tox21 assays, which can be applied to estimate the potency ranges for metal-induced bioactivity that support risk assessment efforts for metal and metalloid exposures. These findings identify previously undercharacterized molecular targets of metal compounds, offering new insights into mechanisms of metal toxicity and informing improved risk assessment methodologies.
G-protein-coupled receptors (GPCRs) are a diverse family of seven-transmembrane domain receptors that play pivotal roles in various physiological and neurological processes by mediating extracellular signals through G proteins. Notable GPCRs such as ADRB2, CHRM1, DRD2, and HTR2A are important therapeutic targets linked to conditions ranging from asthma to schizophrenia. The human ether-à-go-go-related gene (hERG), encoding the Kv11.1 potassium channel, is critical for cardiac repolarization, the inhibition of which can lead to prolonged QT intervals and an increased risk of arrhythmias. Consequently, assessing hERG-GPCR interactions is essential during drug development to enhance safety and ensure regulatory compliance. In this study, we utilized quantitative high-throughput screening (qHTS) to identify GPCR agonists and inhibitors in the Tox21 10K compound library. We applied machine-learning (ML)-based quantitative structure-activity relationship (QSAR) models to predict selective GPCR-targeting compounds with reduced hERG liability, employing different data processing sequences. Our models trained on the Tox21 10K library screening data were subsequently validated by using the Library of Pharmacologically Active Compounds (LOPAC). Furthermore, the models were applied to virtually screen approximately 360 K diverse compounds, with the top predictions experimentally validated, revealing new GPCR modulators with minimal hERG liability. The findings provide efficient strategies for the development of lead compounds targeting GPCRs while minimizing the cardiac risks associated with hERG inhibition.
Beta-1 adrenergic receptors (ADRβ1) are critical regulators of cardiac function; however, the potential modulation of ADRβ1 by environmental chemicals remains largely underexplored, raising concerns about unintended impacts on cardiovascular health. We applied a quantitative high-throughput screening (qHTS) approach to identify ADRβ1 agonists within the Tox21 10K compound library, which includes environmental chemicals, pharmaceuticals, industrial agents, and consumer products, using an HTRF-based cAMP assay in ADRβ1-overexpressing HEK293 cells. Primary screening of 8,947 unique compounds identified 118 potential ADRβ1 agonists. Among these, 94 were confirmed and further evaluated for β-adrenergic receptor subtype selectivity (ADRβ2 and ADRβ3) and hERG channel inhibition to assess potential cardiotoxicity liability. Known ADRβ1 agonists, isoproterenol (EC50, 0.91 nM) and dobutamine (EC50, 10 nM), were identified, supporting the validity of the assay. In addition, several compounds with limited prior ADRβ1-specific characterization, such as GR 103691 and N,N'-dibenzylethane-1,2-diamine, demonstrated subtype-selective or mixed agonist profiles, with some exhibiting minimal hERG inhibition. These findings expand the catalog of ADRβ1 modulators and demonstrate the utility of qHTS for identifying chemicals that may affect cardiovascular signaling pathways.
Abstract Background Cosmetics are defined by the U.S. Food and Drug Administration (FDA) as “articles intended to be rubbed, poured, sprinkled, or sprayed on, introduced into, or otherwise applied to the human body…for cleansing, beautifying, promoting attractiveness, or altering the appearance”. However, the safety of cosmetic ingredients is the responsibility of the manufacturer and exposure to some of these chemicals can result in unintentional harmful effects in humans. Thus, some states are banning specific chemicals, at the state legislation level, from being used in cosmetics, which includes known endocrine disruptors such as dibutyl phthalate, diethylhexyl phthalate, and per- and polyfluoroalkyl substances (PFAS). In this study, we aim to determine the toxicity pathways significantly affected by the cosmetic ingredients (both banned and non-banned), compared to the non-cosmetics compounds in the Tox21 10K library. Methods The Tox21 10K compound library (which includes 113 banned, 927 non-banned cosmetic ingredients and 8,185 non-cosmetic compounds) was previously screened against a panel of ~ 90 cell-based and biochemical assays. The hit rates (i.e., the percentage of compounds active in an assay) of the cosmetic and non-cosmetic compounds, as well as banned and non-banned compounds, in each Tox21 assay were compared using a Fisher’s exact test, and a p-value < 0.05 was considered statistically significant. Results We found that the evaluated cosmetic compounds were significantly more active in 11 antagonist mode assays and 23 agonist mode assays compared to the non-cosmetic compounds. The banned compounds were significantly more active in 7 antagonist mode and 6 agonist mode assays compared to the non-banned compounds. The targets of the antagonist assays included enzymes such as CYP2C19 and aromatase (CYP19A1), and the agonist assays included the Keap1/Nrf2 antioxidant response element pathway. Conclusions In the present study, we compared the bioactivity profiles of cosmetic ingredients and non-cosmetic compounds (not used in cosmetics), as well as the banned and non-banned cosmetic ingredients, across the Tox21 assays. The analyses revealed distinct molecular pathways for these chemicals which furthers our mechanistic understanding of cosmetics ingredient toxicity and may help direct the selection of cosmetic ingredients in the future with potentially less bioactivity. Clinical trial number Not applicable.
The Tox21 10K chemical library, an in vitro toxicology toolbox consisting of more than 8900 unique chemical entities including environmental chemicals and drugs, has undergone analytical quality control (QC) testing after storage at room temperature for 0 and 4 months (T0 and T4). Each chemical was previously assigned a QC grade based on purity, identity, and concentration. In parallel, the Tox21 10K library has been tested across approximately 90 in vitro assays in a quantitative high-throughput screening (qHTS) format, generating >120 M data points to date. These data were used to analyze the correlation between chemical quality and bioassay activity, as well as chemical structure. The chemical characteristics of poor-quality and unstable compounds were explored to identify structural features that should be avoided. In addition, one of the high-throughput assays measuring the induction of p53 activity by small molecules was used to test the Tox21 10K compound library at T0 and T4 due to its robust performance and reproducibility. Approximately 2% of compounds in the library showed a significant change in activity in the p53 assay between T0 and T4 (active to inactive or vice versa), which also correlated with chemical stability. Here, machine learning models were constructed using bioassay data or chemical structures to predict poor-quality (low QC grades at T0) and unstable (grade drop from T0 to T4) chemicals. Chemical structure was found to be highly predictive (0.75) of chemical quality and stability, whereas bioassay data was less predictive (0.66) but still showed better than random performance. Taken together, these findings provide valuable guidance for interpreting the Tox21 assay results and informing best practices for future chemical selection and handling.
BACKGROUND:Exposome-wide association studies (ExWAS) that include measures of social vulnerability may be sensitive to the geographic scale of exposure variables. While linking de-identified geomarkers to participant locations facilitates data sharing and improves privacy protection, the spatial scale and boundaries of geographic units used for analyses may yield varying estimates of epidemiologic associations. METHODS:We quantified the impact of the modifiable areal unit problem in the context of ExWAS designs via a simulation study and case study approach. We replicated the CDC/ATSDR Social Vulnerability Index (SVI) and its n = 16 components and n = 4 themes across North Carolina census tracts, county subdivisions, five-digit ZIP Code Tabulation Areas (ZCTA5), counties, and three-digit ZIP Code Tabulation Areas (ZCTA3). We fit 67,392 logistic regression models of theoretical associations using a simulated epidemiologic cohort of one million locations. We investigated exposure-phenotype relationships between SVI indicators and 11 disease outcomes reported by participants in the Personalized Environment and Genes Study cohort. RESULTS:With increased spatial aggregation encompassing larger administrative areas, the exposure distribution became less Gaussian, and there was less evidence of spatial clustering. Additionally, with coarser spatial aggregation, models lost precision: empirical coverage decreased, model estimates were more biased, and power was reduced. Twenty significant associations were found in exposure-phenotype relationships at the census-tract level. Fifteen were found at the ZCTA5 level and 12 at the ZCTA3 level. DISCUSSION:The efficacy of a simulated ExWAS to detect robust associations was diminished when using more coarsely aggregated geomarkers. SVI indicators became less informative when aggregated at spatially larger administrative units. CONCLUSION:Our results emphasize the importance of using data at spatial resolutions that align with hypothesized exposure-phenotype mechanisms and anticipating the potential loss of epidemiologic efficacy when participant data must be aggregated to large spatial scales to protect privacy.
The large and steadily increasing volume of scientific publications presents a challenge in accessing and utilizing data due to their unstructured nature. Toxicology, in particular, depends on structured data from diverse study types for study evaluation, weight-of-evidence chemical assessments, and validation of new approach methodologies (NAMs). Manual data extraction is time and labor-intensive. This work presents an automated data extraction workflow using large language models (LLMs) within the KNIME platform. The workflow integrates document parsing tools with LLMs to extract variables from scientific publications and general PDF files. Two execution modes are available: text mode and image mode. Text mode applies tools for extracting text and tables, while image mode uses multimodal LLMs to process non-linear layouts and graphical content. The workflow achieves 81.14% accuracy in text mode for scientific publications and up to 98.54% in image mode for general PDF files. The KNIME platform ensures accessibility through a user-friendly interface, allowing non-experts to use advanced data extraction methods. This automated approach facilitates toxicological research by improving the retrieval of structured data. By democratizing access to LLM-powered workflows, this approach paves the way for significant advancements in knowledge synthesis to support biomedical research. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning Data Science > Computer Algorithms and Programming Data Science > Databases and Expert Systems
Over the past decade, global contamination from per- and polyfluoroalkyl substances (PFAS) has become apparent due to their detection in countless matrices worldwide, from consumer products to human blood to drinking water. As researchers implement nontargeted analyses (NTA) to more fully understand the PFAS present in the environment and human bodies, clear guidance is needed for consistent and objective reporting of the identified molecules. Confidence levels for small molecules analyzed and identified with high-resolution mass spectrometry (HRMS) have existed since 2014; however, unification of currently used levels and improved guidance for their application is needed due to inconsistencies in reporting and continuing innovations in analytical methods. Here, we (i) investigate current practices for confidence level reporting of PFAS identified with liquid chromatography (LC), gas chromatography (GC), and/or ion mobility spectrometry (IMS) coupled with HRMS and (ii) propose a simple, unified confidence level guidance that incorporates both PFAS-specific attributes and IMS collision cross section (CCS) values.
Metabolically active compounds can cause toxicity which would otherwise be undetected using traditional in vitro assays with limited proficiency for xenobiotic metabolism. Introduction of liver microsomes to assay systems enables enhanced identification of compounds that require biotransformation to induce toxicity. Previously, metabolically active compounds from the Tox21 10 K compound library were identified using assays probing two targets, p53 and acetylcholinesterase (AChE), in the presence and absence of human or rat liver microsomes, due to the established roles of cytochrome P450 (CYP) enzymes in human drug metabolism. To further explore the role of metabolic activation, the activities of the identified metabolically active compounds were evaluated against five CYP enzymes: CYP1A2, CYP2C9, CYP2C19, CYP2D6, and CYP3A4. CYP bioactivities were found to be highly predictive (>80 % accuracy) of compounds that required metabolic activation in these assays. Chemical features significantly enriched in metabolically active compounds, as well as chemical features that were specific for each of the five CYPs, were identified. Product use exposures of the metabolically active compounds were examined in this study, with "pesticides" appearing to be the largest category that may produce harmful metabolites. Additionally, the compound interactions with different CYPs were assessed and frequencies for both classes of compounds, drugs and environmental chemicals, were found to be proportionally similar across the five CYP isoforms.
β-adrenergic receptors play important roles in heart failure and drug-induced cardiotoxicity (DICT). The Tox21 10 K library of drugs and environmental chemicals have been tested for their activity against β-adrenergic receptor subtypes 1 and 2 (ADRB1 and ADRB2), as well as inhibition of the human ether-à-go-go-related gene (hERG) in a quantitative high-throughput screening (qHTS) format. In this study, the Tox21 compound activity profiles in the ADRB1/2 and hERG assays were compared in relation to their DICT potential. The results showed that compounds that acted as ADRB1 agonists, ADRB2 antagonists, or hERG inhibitors were more likely to exhibit DICT. The ADRB1 and ADRB2 assays shared similar compound activity profiles, while the hERG inhibition assay identified a distinct set of active compounds. In addition, we identified structural features that may differentiate the cardiotoxic and non-toxic ADRB1 agonists. Finally, machine learning models were developed for ADRB1 activity prediction based on chemical structure. The models were used to virtually screen a collection of approximately 360 K diverse compounds, with the highest-ranked compounds selected for experimental validation. This work represents the first systematic study of drugs and environmental chemicals against ADRB1/2, providing important insights into β-adrenergic receptor-related cardiotoxicity mechanisms. By clarifying how specific pharmacological interactions contribute to cardiac risk, it provides a framework for early cardiotoxicity prediction and the design of safer therapeutics through integrated profiling and modeling.
Cytochrome P450 (CYP) enzymes are membrane-bound hemoproteins crucial for drug and xenobiotic metabolism. While more than 50 CYPs have been identified in humans, the isoforms from CYP1, 2, and 3 families contribute to the metabolism of about 80% of clinically approved drugs. To evaluate the effects of environmental chemicals on the activities of these important CYP enzyme families, we screened the Tox21 10K compound library to identify chemicals that inhibit CYP1A2, 2C9, 2C19, 2D6, and 3A4 enzymes. The data obtained from these five screenings were analyzed to reveal the structural classes responsible for inhibiting multiple and/or selective CYPs. Some known structural compound classes exhibiting pan-CYP inhibition, such as azole fungicides, along with established clinical inhibitors of CYPs, including erythromycin and verapamil inhibiting CYP3A4 and paroxetine and terbinafine inhibiting CYP2D6, were all confirmed in the current study. In addition, some selective CYP inhibitors, previously unknown but with potent activity (IC50 values < 1 µM), were identified. Examples included yohimbine, an indole alkaloid, and loteprednol, a corticosteroid, which showed inhibitory activity in CYP2D6 and 3A4 assays, respectively. These findings suggest that assessment of a candidate compound’s impact on CYP function may allow pre-emptive mitigation of potential adverse reactions and toxicity during drug development or toxicological characterization of environmental chemicals.
IntroductionData science training has the potential to propel environmental health research efforts into territories that remain untapped and holds immense promise to change our understanding of human health and the environment. Though data science training resources are expanding, they are still limited in terms of public accessibility, user friendliness, breadth of content, tangibility through real-world examples, and applicability to the field of environmental health science.MethodsTo fill this gap, we developed an environmental health data science training resource, the inTelligence And Machine lEarning (TAME) Toolkit, version 2.0 (TAME 2.0).ResultsTAME 2.0 is a publicly available website that includes training modules organized into seven chapters. Training topics were prioritized based upon ongoing engagement with trainees, professional colleague feedback, and emerging topics in the field of environmental health research (e.g., artificial intelligence and machine learning). TAME 2.0 is a significant expansion upon the original TAME training resource pilot. TAME 2.0 specifically includes training organized into the following chapters: (1) Data management to enable scientific collaborations; (2) Coding in R; (3) Basics of data analysis and visualizations; (4) Converting wet lab data into dry lab analyses; (5) Machine learning; (6) Applications in toxicology and exposure science; and (7) Environmental health database mining. Also new to TAME 2.0 are “Test Your Knowledge” activities at the end of each training module, in which participants are asked additional module-specific questions about the example datasets and apply skills introduced in the module to answer them. TAME 2.0 effectiveness was evaluated via participant surveys during graduate-level workshops and coursework, as well as undergraduate-level summer research training events, and suggested edits were incorporated while overall metrics of effectiveness were quantified.DiscussionCollectively, TAME 2.0 now serves as a valuable resource to address the growing demand of increased data science training in environmental health research. TAME 2.0 is publicly available at: https://uncsrp.github.io/TAME2/.
Per- and polyfluoroalkyl substances (PFAS) are a diverse class of anthropogenic chemicals; many are persistent, bioaccumulative, and mobile in the environment. Worldwide, PFAS bioaccumulation causes serious adverse health impacts, yet the physiochemical determinants of bioaccumulation and toxicity for most PFAS are not well understood, largely due to experimental data deficiencies. As most PFAS are proteinophilic, protein binding is a critical parameter for predicting PFAS bioaccumulation and toxicity. Among these proteins, human serum albumin (HSA) is the predominant blood transport protein for many PFAS. We previously demonstrated the utility of an in vitro differential scanning fluorimetry assay for determining relative HSA binding affinities for 24 PFAS. Here, we report HSA affinities for 65 structurally diverse PFAS from 20 chemical classes. We leverage these experimental data, and chemical/molecular descriptors of PFAS, to build 7 machine learning classifier algorithms and 9 regression algorithms, and evaluate their performance to identify the best predictive binding models. Evaluation of model accuracy revealed that the top-performing classifier model, logistic regression, had an AUROC (area under the receiver operating characteristic curve) statistic of 0.936. The top-performing regression model, support vector regression, had an R2 of 0.854. These top-performing models were then used to predict HSA-PFAS binding for chemicals in the EPAPFASINV list of 430 PFAS. These developed in vitro and in silico methodologies represent a high-throughput framework for predicting protein-PFAS binding based on empirical data, and generate directly comparable binding data of potential use in predictive modeling of PFAS bioaccumulation and other toxicokinetic endpoints.
BACKGROUND:Comprehensive environmental risk characterization, encompassing physical, chemical, social, ecological, and lifestyle stressors, necessitates innovative approaches to handle the escalating complexity. This is especially true when considering individual and population-level diversity, where the myriad combinations of real-world exposures magnify the combinatoric challenges. The GeoTox framework offers a tractable solution by integrating geospatial exposure data from source-to-outcome in a series of modular, interconnected steps. RESULTS:Here, we introduce the GeoTox open-source R software package for characterizing the risk of perturbing molecular targets involved in adverse human health outcomes based on exposure to spatially-referenced stressor mixtures. We demonstrate its usage in building computational workflows that incorporate individual and population-level diversity. Our results demonstrate the applicability of GeoTox for individual and population-level risk assessment, highlighting its capacity to capture the complex interplay of environmental stressors on human health. CONCLUSIONS:The GeoTox package represents a significant advancement in environmental risk characterization, providing modular software to facilitate the application and further development of the GeoTox framework for quantifying the relationship between environmental exposures and health outcomes. By integrating geospatial methods with cutting-edge exposure and toxicological frameworks, GeoTox offers a robust tool for assessing individual and population-level risks from environmental stressors. GeoTox is freely available at https://niehs.github.io/GeoTox/ .
This manuscript critically examines the landscape of public-facing web-based environmental health (EH) and environmental justice (EJ) screening tools aimed at mitigating environmental health crises that are involved in a substantial percentage of deaths globally. These EJ/EH screening tools have proliferated with the growth of publicly available data sources and computational advances that have fueled novel analytics and have made strides toward democratizing access to EJ/EH information impacting communities. The interactive, highly visual analytics offered by some of these EJ/EH screening tools could help address the role of environmental injustice in exacerbating environmental health-related causes of mortality and enable affected communities to take a more active role in EJ/EH efforts. Environmental injustice results from environmental conditions that affect communities differently based on residents’ race, income level, national origin, and level of participation in decision-making processes. We survey existing EJ/EH screening tools and evaluate selected examples based on parameters that include data availability, characterization of environmental burden and vulnerability, evaluation of stressor levels, and interpretability of environmental health and justice scores. This review highlights the unique capabilities and limitations of EJ/EH screening tools used at the local (US-Centric), national (US-Centric), and international levels. We then discuss unmet needs and thematic limitations apparent in this survey, related to data availability, relevancy of stressors, assignment of indicator weights, threshold values for action and intervention, modeling robustness, and appropriate community focus. The results underline the need for robust, accessible, and community-centric EJ/EH screening tools that can effectively address the unique environmental health burdens and vulnerabilities faced by communities. We conclude with proposed strategies to enhance EJ/EH screening tool development.
BACKGROUND:Geospatial methods are common in environmental exposure assessments and increasingly integrated with health data to generate comprehensive models of environmental impacts on public health. OBJECTIVE:Our objective is to review geospatial exposure models and approaches for health data integration in environmental health applications. METHODS:We conduct a literature review and synthesis. RESULTS:First, we discuss key concepts and terminology for geospatial exposure data and models. Second, we provide an overview of workflows in geospatial exposure model development and health data integration. Third, we review modeling approaches, including proximity-based, statistical, and mechanistic approaches, across diverse exposure types, such as air quality, water quality, climate, and socioeconomic factors. For each model type, we provide descriptions, general equations, and example applications for environmental exposure assessment. Fourth, we discuss the approaches used to integrate geospatial exposure data and health data, such as methods to link data sources with disparate spatial and temporal scales. Fifth, we describe the landscape of open-source tools supporting these workflows.