
The Toxicity Values Database, ToxValDB, was developed by the U.S. EPA Center for Computational Toxicology and Exposure as a resource to curate, store, standardize, and make accessible a wide range of human health-relevant toxicity information. The database originated in response to the need for harmonized and computationally accessible toxicology data. The scope and design of the database have evolved over time since its first release in 2016. Herein, the newly redesigned structure and development of ToxValDB v9.6.1 is described. The database is a compilation of three classes of summary-level values for chemical substances: in vivo toxicity study results (e.g., lowest-and no-observed adverse effect level), derived toxicity values (e.g., maximum acceptable oral dose), and media exposure guidelines (e.g., maximum contaminant level for drinking water). The current version of the database (9.6.1) contains 242,149 records covering 41,769 unique chemicals from 36 sources (55 source tables). With all records in a consistent structure normalized to a standardized vocabulary, the chemical and data landscape of ToxValDB v9.6.1 can be evaluated. To illustrate chemical coverage, the available data were mapped to chemical lists of regulatory importance. Further, the distribution of oral administered doses within in vivo toxicity studies was assessed by annotated chemical class. The harmonized in vivo data within ToxValDB have many applications including use in chemical screening and prioritization for human health assessment, modeling predictions, and benchmarking for New Approach Methods (NAMs), as well as to address a diverse range of novel research questions.
Adverse Outcome Pathways (AOPs) describe the mechanistic interactions of biological entities with a stressor (chemical, nanomaterial, radiation, virus, etc.) that produce an adverse response. How these interactions and associations are catalogued contributes to our ability to understand mechanistic effects and apply this knowledge to New Approach Methods (NAMs) that have the potential to reduce animal testing in chemical, biological, and material safety assessments. Making AOP data align with FAIR (Findable, Accessible, Interoperable, and Reusable) metadata standards relies on technical tools that implement and process AOP data and related metadata, and the establishment of coordinated and consensus computational bioinformatic methods. Herein current efforts in addressing the FAIRification of AOP mechanistic data and metadata, as well as the international, collaborative efforts to document, and improve the (re)-use and reliability of AOP information will be described. These coordinated efforts contribute to the establishment of a directive for the processing and storing of standardized AOP mechanistic data in the AOP-Wiki repository, and application of these data to next generation risk assessment.
Read-across is a data-gap filling technique used to predict the toxicity of a target chemical based on data from similar analogues. It is predominantly performed through expert-driven assessments which can limit reproducibility and broader acceptance. Data-driven approaches such as Generalised Read-Across (GenRA) offer the potential to generate more reproducible read-across predictions with quantified uncertainties and performance metrics. A key challenge is reconciling expert- and data-driven approaches particularly in how analogues are identified, evaluated and used to derive predictions. A critical aspect of analogue selection lies in understanding the relative contribution of different similarity contexts e.g. whether structural similarity plays a larger role than metabolism similarity. This study explored these considerations by compiling a compendium of expert-driven read-across assessments for repeated dose toxicity endpoints from peer reviewed and grey literature. Pairwise similarity was quantified across structural, physicochemical, metabolic and reactivity features within each case and a prediction model was developed to evaluate the contribution of each similarity context in analogue selection. Although the dataset comprised 157 read-across cases and 695 unique substances, it was limited in size, heterogeneous in origin and variable in analogue selection criteria and use contexts. These factors constrain generalisability of the findings and indicate that conclusions should be interpreted with caution. Nonetheless, the qualitative insight that structure and metabolism were influential led to a followup investigation using graphbased deep learning to explore whether embeddings derived from structure and/or metabolism information could improve read-across predictions, using repeated dose toxicity as a case study, relative to structural similarity baselines.
Read-across is a technique used to fill data gaps for substances lacking specific hazard data. The technique relies on identifying source analogues with relevant data that are 'similar' to the substance of interest (target). Typically, source analogues are identified on the basis of structural similarity but the evaluation of their suitability for read-across depends on other contexts of similarity. This manuscript aimed to review the ways in which source analogues are identified for read-across using chemical fingerprint/scaffold approaches before describing graphbased approaches including; graph kernel, graph embedding, and deep learning. To demonstrate how these could be practically used for analogue identification, five different toxicity datasets of varying size and diversity were selected that had been the subject of previous read-across or QSAR analyses. One dataset was an analogue set whereas the other four datasets comprised substances evaluated for their skin sensitisation, skin irritation, fathead minnow aquatic toxicity and genotoxicity potential. The analogues and their associated similarities using the different graph based approaches were compared with the outcomes from two chemical fingerprint approaches (ToxPrints and Morgan). The results for each dataset are briefly described. Based on the examples evaluated, graph kernel approaches were found to have some promise, in contrast unsupervised whole graph embedding approaches were ineffective for all the datasets evaluated. Graph convolutional networks produced meaningful embeddings for the genotoxicity dataset evaluated. Depending on use case, availability and size of training data, graph similarity approaches have the potential to play a larger role in analogue identification and evaluation for read-across.
New approach methods (NAMs) have been prioritized to reduce the use of animals for chemical safety assessment while continuing to protect human health and the environment. A key challenge of generating toxicity data is the implementation of a standardized analysis approach for transparent and reproducible benchmark concentration (BMC) estimation and uncertainty quantification for assay developers, regulators, and other stakeholders. In this study, we compared the bioactivity results of 321 chemical samples from four established BMC analysis pipelines used for evaluation of developmental neurotoxicity (DNT) NAMs data: the ToxCast pipeline (tcpl), CRStats, DNT DIVER (Curvep and Hill pipelines). We found an overall activity hit call concordance of 77.2% and highly correlated BMC estimations (r = 0.92 ± 0.02 SD), demonstrating generally good agreement across pipelines. Discordance appeared to be explained predominantly by noise within the data and borderline activity (activity occuring near the benchmark response level). Evaluation of the BMC confidence intervals indicated that pipeline selection may impact the estimation of the BMC lower bound. Consideration of biphasic models appeared important for capturing biologically-relevant changes in activity in the DNT battery. Lastly, different approaches to compute 'selective' bioactivity (activity below the threshold of cytotoxicity) were compared, identifying the CRstats classification model as more stringent for classifying selective activity. Overall, these findings indicated greater confidence in NAMs bioactivity results and emphasize the importance of understanding strengths and uncertainties of concentration-response modeling pipelines for informing biological interpretation and application decision making.
Per- and Polyfluoroalkyl substances (PFAS) are a class of manufactured chemicals that are in widespread use and many present concerns for persistence, bioaccumulation and toxicity. Whilst a handful of PFAS have been characterized for their hazard profiles, the vast majority have not been extensively studied. Herein, a chemical category approach was developed and applied to PFAS that could be readily characterized by a chemical structure. The PFAS definition as described in the Toxic Substances Control Act (TSCA) section 8(a)(7) rule was applied to the Distributed Structure-Searchable Toxicity (DSSTox) database to retrieve an initial list of 13,054 PFAS. Plausible degradation products from the 563 PFAS on the non-confidential TSCA Inventory were simulated using the Catalogic expert system, and the unique predicted PFAS degradants (2484) that conformed to the same PFAS definition were added to the list resulting in a set of 15,538 PFAS. Each PFAS was then assigned into a primary category using Organisation for Economic Co-operation and Development (OECD) structure-based classifications. The primary categories were subdivided into secondary categories based on a chain length threshold (>=7 =7 vs < 7). Secondary categories were subcategorized using chemical fingerprints to achieve a balance between total number of structural categories vs. level of structural similarity within a category based on the Jaccard index. A set of 128 terminal structural categories were derived from which a subset of representative candidates could be proposed for potential data collection, considering the sparsity of relevant toxicity data within each category, presence on environmental monitoring lists, and the ability to identify plausible manufacturers/importers. Refinements to the approach taking into consideration ways in which the categories could be updated by mechanistic data and physicochemical property information are also described. This categorization approach may be used to form the basis of identifying candidates for data collection with related applications in QSAR development, read-across and hazard assessment.
The advancement of protein structural prediction tools, exemplified by AlphaFold and Iterative Threading ASSEmbly Refinement, has enabled the prediction of protein structures across species based on available protein sequence and structural data. In this study, we introduce an innovative molecular docking method that capitalizes on this wealth of structural data to enhance predictions of chemical susceptibility across species. We demonstrated this method using the androgen receptor as a pertinent modulator of endocrine function. By using protein structures, this method contextualizes species susceptibility within a functional framework and helps to integrate molecular docking into the repertoire of New Approach Methodologies (NAMs) that support the Next-Generation Risk Assessment (NGRA) paradigm through the novel integration of various open-source tools.
The National Nanotechnology Initiative organized a Nanoinformatics Conference in the 2023 Biden-Harris Administration’s Year of Open Science, which included interested U.S. and EU stakeholders, and preceded the U.S.-EU COR meeting on November 15th, 2023 in Washington, D.C. Progress in the development of a common nanoinformatics infrastructure in the European Union and United States were discussed. Development of contributing, individual database projects, and their strengths and weaknesses, were highlighted. Recommendations and next steps for a U.S. nanoEHS common infrastructure were discussed in light of the pending update of the National Nanotechnology Initiative (NNI)’s Environmental, Health and Safety Research Strategy, and U.S. efforts to curate and house nano Environmental Health and Safety (nanoEHS) data from U.S. federal stakeholder groups. Improved data standards, for reporting and storage have been identified as areas where concerted efforts could most benefit initially. Areas that were not addressed at the conference, but that are critical to progress of the U.S. federal consortium effort are the evaluation of data formats according to use and sustainability measures; modeler and end user, including risk-assessor and regulator perspectives; a need for a community forum or shared data location that is not hosted by any individual U.S. federal agency, and is accessible to the public; as well as emerging needs for integration with new data types such as micro and nano plastics, and interoperability with other data and meta-data, such as adverse outcome pathway information. Future progress will depend on continued interaction of the U.S. and EU CORs, stakeholders and partners in the continued development goals for shared or interoperable infrastructure for nanoEHS.
Read-across is a well-established data-gap filling technique used within analogue or category approaches. Acceptance remains an issue, mainly due to the difficulties of addressing residual uncertainties associated with a read-across prediction and because assessments are expert-driven. Frameworks to develop, assess and document read-across may help reduce variability in read-across results. Data-driven read-across approaches such as Generalised Read-Across (GenRA) include quantification of uncertainties and performance. GenRA also offers opportunities on how New Approach Method (NAM) data can be systematically incorporated to support the read-across hypothesis. Herein, a systematic investigation of differences in expert-driven read-across with data-driven approaches was pursued in terms of building scientific confidence in the use of read-across. A dataset of expert-driven read-across assessments that made use of registration data as disseminated in the public International Uniform Chemical Information Database (IUCLID) (version 6) of Registration, Evaluation, Authorisation and Restriction of Chemicals (REACH) Study Results were compiled. A dataset of ∼5000 read-across cases pertaining to repeated dose and developmental toxicity was extracted and mapped to content within EPA’s Distributed Structure Searchable Toxicity database (DSSTox) to retrieve chemical name and structural identification information. Content could be mapped to ∼3600 cases which when filtered for unique cases with curated quantitative structure–activity relationship-ready SMILES resulted in 389 target-source analogue pairs. The similarity between target and the source analogues on the basis of different contexts – from structural similarity using chemical fingerprints to metabolic similarity using predicted metabolic information was evaluated. An attempt was also made to quantify the relative contribution each similarity context played relative to the target-source analogue pairs by deriving a model which predicted known analogue pairs. Finally, point of departure values (PODs) were predicted using the GenRA approach underpinned by data extracted from the EPA’s Toxicity Values Database (ToxValDB). The GenRA predicted PODs were compared with those reported within the REACH dossiers themselves. This study offers generalisable insights on how read-across is already applied for regulatory submissions and expectations on the levels of similarity necessary to make decisions.
Animal toxicity testing is time and resource intensive, making it difficult to keep pace with the number of substances requiring assessment. Machine learning (ML) models that use chemical structure information and high-throughput experimental data can be helpful in predicting potential toxicity . However, much of the toxicity data used to train ML models is biased with an unequal balance of positives and negatives primarily since substances selected for in vivo testing are expected to elicit some toxicity effect. To investigate the impact this bias had on predictive performance, various sampling approaches were used to balance in vivo toxicity data as part of a supervised ML workflow to predict hepatotoxicity outcomes from chemical structure and/or targeted transcriptomic data. From the chronic, subchronic, developmental, multigenerational reproductive, and subacute repeat-dose testing toxicity outcomes with a minimum of 50 positive and 50 negative substances, 18 different study-toxicity outcome combinations were evaluated in up to 7 ML models. These included Artificial Neural Networks, Random Forests, Bernouilli Naïve Bayes, Gradient Boosting, and Support Vector classification algorithms which were compared with a local approach, Generalised Read-Across (GenRA), a similarity-weighted k-Nearest Neighbour (k-NN) method. The mean CV F1 performance for unbalanced data across all classifiers and descriptors for chronic liver effects was 0.735 (0.0395 SD). Mean CV F1 performance dropped to 0.639 (0.073 SD) with over-sampling approaches though the poorer performance of KNN approaches in some cases contributed to the observed decrease (mean CV F1 performance excluding KNN was 0.697 (0.072 SD)). With under-sampling approaches, the mean CV F1 was 0.523 (0.083 SD). For developmental liver effects, the mean CV F1 performance was much lower with 0.089 (0.111 SD) for unbalanced approaches and 0.149 (0.084 SD) for under-sampling. Over-sampling approaches led to an increase in mean CV F1 performance (0.234, (0.107 SD)) for developmental liver toxicity. Model performance was found to be dependent on dataset, model type, balancing approach and feature selection. Accordingly tailoring ML workflows for predicting toxicity should consider class imbalance and rely on simpler classifiers first.
This paper presents an ordinary differential equation (ODE) model of endogenous H2O2 production and elimination in hepatocytes that is unique, at the time of writing, in its ability to accurately compute intracellular H2O2 concentration during incidents of oxidative stress and in its usefulness for constructing PBPK/PD models for ROS-generating xenobiotics. Versions of the model are presented for rat hepatocytes in vitro and mouse liver in vivo. A generic method is given for using the model to create PBPK/PD models which predict intracellular H2O2 concentration and oxidative-stress-induced hepatocyte death; these are identifiable from in vitro data sets reporting cell mortality following xenobiotic exposure at various levels. The procedure is demonstrated for the trivalent arsenical dimethylarsinous acid (DMAIII), which is produced in liver as part of the arsenic elimination pathway. This is the first model of H2O2 metabolism in hepatocytes to feature values for the endogenous rates of H2O2 production by mitochondria and other organelles which are inferred from the physiology literature, and to feature a detailed, realistic treatment of GSH metabolism; the latter is achieved by incorporating a minimal version of Reed and coworkers’ pioneering model of GSH metabolism in liver. Model simulations indicate that critical GSH depletion is the immediate trigger for intracellular H2O2 rising to concentrations associated with apoptosis (>1μM), that this may only occur hours after the xenobiotic concentration peaks (”delay effect”), that when critical GSH depletion does occur, H2O2 concentration rises rapidly in a sequence of two boundary layers, characterized by the kinetics of glutathione peroxidase (first boundary layer) and catalase (second boundary layer), and that intracellular H2O2 concentration >1μM implies critical GSH depletion. There has been speculation that ROS levels in the range associated with apoptosis simply indicate, rather than cause, an apoptotic milieu. Model simulations are consistent with this view. In a result of interest to the wider physiology community, the delay effect is shown to provide a GSH-based mechanism by which cells can distinguish transient elevations in H2O2 concentration, of use in intracellular signaling, from persistent ones indicative of either pathology or the presence of toxins, the second state of affairs eventually triggering apoptosis.
The Toxicological Prioritization Index (ToxPi) is a visual analysis and decision support tool for dimension reduction and visualization of high throughput, multi-dimensional feature data. ToxPi was originally developed for assessing the relative toxicity of multiple chemicals or stressors by synthesizing complex toxicological data to provide a single comprehensive view of the potential health effects. It continues to be used for profiling chemicals and has since been applied to other types of "sample" entities, including geospatial (e.g. county-level Covid-19 risk and sites of historical PFAS exposure) and other profiling applications. For any set of features (data collected on a set of sample entities), ToxPi integrates the data into a set of weighted slices that provide a visual profile and a score metric for comparison. This scoring system is highly dependent on user-provided feature weights, yet users often lack knowledge of how to define these feature weights. Common methods for predicting feature weights are generally unusable due to inappropriate statistical assumptions and lack of global distributional expectation. However, users often have an inherent understanding of expected results for a small subset of samples. For example, in chemical toxicity, prior knowledge can often place subsets of chemicals into categories of low, moderate or high toxicity (reference chemicals). Ordinal regression can be used to predict weights based on these response levels that are applicable to the entire feature set, analogous to using positive and negative controls to contextualize an empirical distribution. We propose a semi-supervised method utilizing ordinal regression to predict a set of feature weights that produces the best fit for the known response ("reference") data and subsequently fine-tunes the weights via a customized genetic algorithm. We conduct a simulation study to show when this method can improve the results of ordinal regression, allowing for accurate feature weight prediction and sample ranking in scenarios with minimal response data. To ground-truth the guided weight optimization, we test this method on published data to build a ToxPi model for comparison against expert-knowledge-driven weight assignments.
High-throughput screening (HTS) assays for bioactivity in the Tox21 program aim to evaluate an array of different biological targets and pathways, but a significant barrier to interpretation of these data is the lack of high-throughput screening (HTS) assays intended to identify non-specific reactive chemicals. This is an important aspect for prioritising chemicals to test in specific assays, identifying promiscuous chemicals based on their reactivity, as well as addressing hazards such as skin sensitisation which are not necessarily initiated by a receptor-mediated effect but act through a non-specific mechanism. Herein, a fluorescence-based HTS assay that allows the identification of thiol-reactive compounds was used to screen 7,872 unique chemicals in the Tox21 10K chemical library. Active chemicals were compared with profiling outcomes using structural alerts encoding electrophilic information. Random Forest classification models based on chemical fingerprints were developed to predict assay outcomes and evaluated through 10-fold stratified cross validation (CV). The mean CV Balanced Accuracy of the validation set was 0.648. The model developed shows promise as a tool to screen untested chemicals for their potential electrophilic reactivity based solely on chemical structural features.
This work estimates benchmarks for new approach method (NAM) performance in predicting organ-level effects in repeat dose studies of adult animals based on variability in replicate animal studies. Treatment-related effect values from the Toxicity Reference database (v2.1) for weight, gross, or histopathological changes in the adrenal gland, liver, kidney, spleen, stomach, and thyroid were used. Rates of chemical concordance among organ-level findings in replicate studies, defined by repeated chemical only, chemical and species, or chemical and study type, were calculated. Concordance was 39 - 88%, depending on organ, and was highest within species. Variance in treatment-related effect values, including lowest effect level (LEL) values and benchmark dose (BMD) values when available, was calculated by organ. Multilinear regression modeling, using study descriptors of organ-level effect values as covariates, was used to estimate total variance, mean square error (MSE), and root residual mean square error (RMSE). MSE values, interpreted as estimates of unexplained variance, suggest study descriptors accounted for 52-69% of total variance in organ-level LELs. RMSE ranged from 0.41 - 0.68 log10-mg/kg/day. Differences between organ-level effects from chronic (CHR) and subchronic (SUB) dosing regimens were also quantified. Odds ratios indicated CHR organ effects were unlikely if the SUB study was negative. Mean differences of CHR - SUB organ-level LELs ranged from -0.38 to -0.19 log10 mg/kg/day; the magnitudes of these mean differences were less than RMSE for replicate studies. Finally, in vitro to in vivo extrapolation (IVIVE) was employed to compare bioactive concentrations from in vitro NAMs for kidney and liver to LELs. The observed mean difference between LELs and mean IVIVE dose predictions approached 0.5 log10-mg/kg/day, but differences by chemical ranged widely. Overall, variability in repeat dose organ-level effects suggests expectations for quantitative accuracy of NAM prediction of LELs should be at least ± 1 log10-mg/kg/day, with qualitative accuracy not exceeding 70%.
The Analog Identification Methodology (AIM) was developed over 20 years ago to identify analogues to support read-across at the US Environmental Protection Agency. However, the current public version of the standalone tool, released in 2012, is no longer usable on Windows operating systems supported by Microsoft. Additionally, the structural logic for analogue selection is based on older, customised Simplified molecular-input-line-entry system (SMILES)-type features that are incompatible with modern cheminformatics tools. Given these limitations, a case study was undertaken to explore a more transparent, extensible method of implementing the AIM fragments using Chemical Subgraphs and Reactions Mark-up Language (CSRML). A CSRML file was developed to codify the original AIM fragments, and the extent to which AIM fragments were faithfully replicated was assessed using the AIM Database. The overall mean performance of the CSRML-AIM across all fragments in terms of sensitivity, specificity, and Jaccard similarity was 89.5%, 99.9%, and 82.2%, respectively. Comparing the AIM fragments with public ToxPrints using a large set of similar to 25,000 substances of regulatory interest to EPA found them to be dissimilar, with an average maximum Jaccard score of 0.24 for AIM and 0.29 for ToxPrint fingerprints. Both fragment sets were then used as inputs in the automated read-across approach, Generalised Read-Across (GenRA), to evaluate the quality of fit in predicting rat acute oral toxicity LD50 values with the coefficient of determination (R-2) and root mean squared error (RMSE). The performance of AIM fragments was R-2=0.434 and RMSE=0.663 whereas that of ToxPrints was R-2=0.477 and RMSE=0.638. A bootstrap resampling using 100 iterations found the mean and the 95th confidence interval of R-2 to be 0.349 [0.319, 0.379] for AIM fragments and 0.377 [0.338, 0.412] for ToxPrints. Although AIM and ToxPrints performed similarly in predicting LD50, they differed in their performance at a local level, revealing that their features can offer complementary insights.
Adverse outcome pathways provide a powerful tool for understanding the biological signaling cascades that lead to disease outcomes following toxicity. The framework outlines downstream responses known as key events, culminating in a clinically significant adverse outcome as a final result of the toxic exposure. Here we use the AOP framework combined with artificial intelligence methods to gain novel insights into genetic mechanisms that underlie toxicity-mediated adverse health outcomes. Specifically, we focus on liver cancer as a case study with diverse underlying mechanisms that are clinically significant. Our approach uses two complementary AI techniques: Generative modeling via automated machine learning and genetic algorithms, and graph machine learning. We used data from the US Environmental Protection Agency's Adverse Outcome Pathway Database (AOP-DB; aopdb.epa.gov) and the UK Biobank's genetic data repository. We use the AOP-DB to extract disease-specific AOPs and build graph neural networks used in our final analyses. We use the UK Biobank to retrieve real-world genotype and phenotype data, where genotypes are based on single nucleotide polymorphism data extracted from the AOP-DB, and phenotypes are case/control cohorts for the disease of interest (liver cancer) corresponding to those adverse outcome pathways. We also use propensity score matching to appropriately sample based on important covariates (demographics, comorbidities, and social deprivation indices) and to balance the case and control populations in our machine language training/testing datasets. Finally, we describe a novel putative risk factor for LC that depends on genetic variation in both the aryl-hydrocarbon receptor (AHR) and ATP binding cassette subfamily B member 11 (ABCB11) genes.
Read-across continues to be a popular data gap filling technique within category and analogue approaches. One of the main issues hindering read-across acceptance is the notion of addressing and reducing uncertainties. Frameworks and formats have been created to help facilitate read-across development, evaluation, and residual uncertainties. However, read-across remains an expert-driven approach with each assessment decided on its own merits with no objective means of evaluating performance or quantifying uncertainties. Here, the underlying motivation of creating an algorithmic approach to read-across, namely the Generalised Read-Across (GenRA) approach, is described. The overall objectives of the approach were to quantify performance and uncertainty. Progress made in quantifying the impact of each similarity context commonly relied upon as part of read-across assessment are discussed. The framework underpinning the approach, the software tools developed to date and how GenRA can be used to make and interpret predictions as part of a screening level hazard assessment decision context are illustrated. Future directions and some of the overarching issues still needed in this field and the extent to which GenRA might facilitate those needs are discussed.
The adverse outcome pathway (AOP) is a conceptual construct that facilitates organisation and interpretation of mechanistic data representing multiple biological levels and deriving from a range of methodological approaches including in silico, in vitro and in vivo assays. AOPs are playing an increasingly important role in the chemical safety assessment paradigm and quantification of AOPs is an important step towards a more reliable prediction of chemically induced adverse effects. Modelling methodologies require the identification, extraction and use of reliable data and information to support the inclusion of quantitative considerations in AOP development. An extensive and growing range of digital resources are available to support the modelling of quantitative AOPs, providing a wide range of information, but also requiring guidance for their practical application. A framework for qAOP development is proposed based on feedback from a group of experts and three qAOP case studies. The proposed framework provides a harmonised approach for both regulators and scientists working in this area.
Neurotoxicology is the study of adverse effects on the structure or function of the developing or mature adult nervous system following exposure to chemical, biological, or physical agents. The development of more informative alternative methods to assess developmental (DNT) and adult (NT) neurotoxicity induced by xenobiotics is critically needed. The use of such alternative methods including in silico approaches that predict DNT or NT from chemical structure (e.g., statistical-based and expert rule-based systems) is ideally based on a comprehensive understanding of the relevant biological mechanisms. This paper discusses known mechanisms alongside the current state of the art in DNT/NT testing. In silico approaches available today that support the assessment of neurotoxicity based on knowledge of chemical structure are reviewed, and a conceptual framework for the integration of in silico methods with experimental information is presented. Establishing this framework is essential for the development of protocols, namely standardized approaches, to ensure that assessments of NT and DNT based on chemical structures are generated in a transparent, consistent, and defendable manner.
Toxicology in the 21st Century has seen a shift from chemical risk assessment based on traditional animal tests, identifying apical endpoints and doses that are "safe", to the prospect of Next Generation Risk Assessment based on non-animal methods. Increasingly, large and high throughput in vitro datasets are being generated and exploited to develop computational models. This is accompanied by an increased use of machine learning approaches in the model building process. A potential problem, however, is that such models, while robust and predictive, may still lack credibility from the perspective of the end-user. In this commentary, we argue that the science of causal inference and reasoning, as proposed by Judea Pearl, will facilitate the development, use and acceptance of quantitative AOP models. Our hope is that by importing established concepts of causality from outside the field of toxicology, we can be "constructively disruptive" to the current toxicological paradigm, using the "Causal Revolution" to bring about a "Toxicological Revolution" more rapidly.