Since the initial release of the Chemical and Products Database (CPDat) in 2018, the United States Environmental Protection Agency has added a considerable amount of chemical exposure-related information to the database and has expanded its schema to accommodate new types of data. This data descriptor provides information regarding the structure and types of data contained within CPDat (both existing and new), new controlled vocabularies implemented to harmonize terminology across the different data types, application of a rigorous data curation and quality assurance tracking system, and various methods of accessing CPDat.
BACKGROUND:The number of chemicals present in the environment exceeds the capacity of government bodies to characterize risk. Therefore, data-informed and reproducible processes are needed for identifying chemicals for further assessment. The Minnesota Department of Health (MDH), under its Contaminants of Emerging Concern (CEC) initiative, uses a standardized process to screen potential drinking water contaminants based on toxicity and exposure potential.OBJECTIVE:Recently, MDH partnered with the U.S. Environmental Protection Agency (EPA) Office of Research and Development (ORD) to accelerate the screening process via development of an automated workflow accessing relevant exposure data, including exposure new approach methodologies (NAMs) from ORD's ExpoCast project.METHODS:The workflow incorporated information from 27 data sources related to persistence and fate, release potential, water occurrence, and exposure potential, making use of ORD tools for harmonization of chemical names and identifiers. The workflow also incorporated data and criteria specific to Minnesota and MDH's regulatory authority. The collected data were used to score chemicals using quantitative algorithms developed by MDH. The workflow was applied to 1867 case study chemicals, including 82 chemicals that were previously manually evaluated by MDH.RESULTS:Evaluation of the automated and manual results for these 82 chemicals indicated reasonable agreement between the scores although agreement depended on data availability; automated scores were lower than manual scores for chemicals with fewer available data. Case study chemicals with high exposure scores included disinfection by-products, pharmaceuticals, consumer product chemicals, per- and polyfluoroalkyl substances, pesticides, and metals. Scores were integrated with in vitro bioactivity data to assess the feasibility of using NAMs for further risk prioritization.SIGNIFICANCE:This workflow will allow MDH to accelerate exposure screening and expand the number of chemicals examined, freeing resources for in-depth assessments. The workflow will be useful in screening large libraries of chemicals for candidates for the CEC program.
Direct monitoring of chemical concentrations in different environmental and biological media is critical to understanding the mechanisms by which human and ecological receptors are exposed to exogenous chemicals. Monitoring data provides evidence of chemical occurrence in different media and can be used to inform exposure assessments. Monitoring data provide required information for parameterization and evaluation of predictive models based on chemical uses, fate and transport, and release or emission processes. Finally, these data are useful in supporting regulatory chemical assessment and decision-making. There are a wide variety of public monitoring data available from existing government programs, historical efforts, public data repositories, and peer-reviewed literature databases. However, these data are difficult to access and analyze in a coordinated manner. Here, data from 20 individual public monitoring data sources were extracted, curated for chemical and medium, and harmonized into a sustainable machine-readable data format for support of exposure assessments.
The rapid characterization of risk to humans and ecosystems from exogenous chemicals requires information on both hazard and exposure. The U.S. Environmental Protection Agency's ToxCast program and the interagency Tox21 initiative have screened thousands of chemicals in various high-throughput (HT) assay systems for in vitro bioactivity. EPA's ExpoCast program is developing complementary HT methods for characterizing the human and ecological exposures necessary to interpret HT hazard data in a real-world risk context. These new approach methodologies (NAMs) for exposure include computational and analytical tools for characterizing multiple components of the complex pathways chemicals take from their source to human and ecological receptors. Here, we analyze the landscape of exposure NAMs developed in ExpoCast in the context of various chemical lists of scientific and regulatory interest, including the ToxCast and Tox21 libraries and the Toxic Substances Control Act (TSCA) inventory. We examine the landscape of traditional and exposure NAM data covering chemical use, emission, environmental fate, toxicokinetics, and ultimately external and internal exposure. We consider new chemical descriptors, machine learning models that draw inferences from existing data, high-throughput exposure models, statistical frameworks that integrate multiple model predictions, and non-targeted analytical screening methods that generate new HT monitoring information. We demonstrate that exposure NAMs drastically improve the coverage of the chemical landscape compared to traditional approaches and recommend a set of research activities to further expand the development of HT exposure data for application to risk characterization. Continuing to develop exposure NAMs to fill priority data gaps identified here will improve the availability and defensibility of risk-based metrics for use in chemical prioritization and screening. IMPACT: This analysis describes the current state of exposure assessment-based new approach methodologies across varied chemical landscapes and provides recommendations for filling key data gaps.
Background Although evidence linking environmental chemicals to breast cancer is growing, mixtures-based exposure evaluations are lacking. Objective This study aimed to identify environmental chemicals in use inventories that co-occur and share properties with chemicals that have association with breast cancer, highlighting exposure combinations that may alter disease risk. Methods The occurrence of chemicals within chemical use categories was characterized using the Chemical and Products Database. Co-exposure patterns were evaluated for chemicals that have an association with breast cancer (BC), no known association (NBC), and understudied chemicals (UC) identified through query of the Silent Spring Institute's Mammary Carcinogens Review Database and the U.S. Environmental Protection Agency's Toxicity Reference Database. UCs were ranked based on structure and physicochemical similarities and co-occurrence patterns with BCs within environmentally relevant exposure sources. Results A total of 6793 chemicals had data available for exposure source occurrence analyses. 50 top-ranking UCs spanning five clusters of co-occurring chemicals were prioritized, based on shared properties with co-occuring BCs, including chemicals used in food production and consumer/personal care products, as well as potential endocrine system modulators. Significance Results highlight important co-exposure conditions that are likely prevalent within our everyday environments that warrant further evaluation for possible breast cancer risk. Impact statement Most environmental studies on breast cancer have focused on evaluating relationships between individual, well-known chemicals and breast cancer risk. This study set out to expand this research field by identifying understudied chemicals and mixtures that may occur in everyday environments due to their patterns of commercial use. Analyses focused on those that co-occur alongside chemicals associated with breast cancer, based upon in silico chemical database querying and analysis. Particularly in instances when understudied chemicals share physicochemical properties and structural features with carcinogens, these chemical mixtures represent conditions that should be studied in future clinical, epidemiological, and toxicological studies.
Exposure to chemicals is influenced by associations between the individual’s location and activities as well as demographic and physiological characteristics. Currently, many exposure models simulate individuals by drawing distributions from population-level data or use exposure factors for single individuals. The Residential Population Generator (RPGen) binds US surveys of individuals and households and combines the population with physiological characteristics to create a synthetic population. In general, the model must be supported by internal consistency; i.e., values that could have come from a single individual. In addition, intraindividual variation must be representative of the variation present in the modeled population. This is performed by linking individuals and similar households across income, location, family type, and house type. Physiological data are generated by linking census data to National Health and Nutrition Examination Survey data with a model of interindividual variation of parameters used in toxicokinetic modeling. The final modeled population data parameters include characteristics of the individual’s community (region, state, urban or rural), residence (size of property, size of home, number of rooms), demographics (age, ethnicity, income, gender), and physiology (body weight, skin surface area, breathing rate, cardiac output, blood volume, and volumes for body compartments and organs). RPGen output is used to support user-developed chemical exposure models that estimate intraindividual exposure in a desired population. By creating profiles and characteristics that determine exposure, synthetic populations produced by RPGen increases the ability of modelers to identify subgroups potentially vulnerable to chemical exposures. To demonstrate application, RPGen is used to estimate exposure to Toluene in an exposure modeling case example.
The use of consumer products presents a potential for chemical exposures to humans. Toxicity testing and exposure models are routinely employed to estimate risks from their use; however, a key challenge is the sparseness of information concerning who uses products and which products are used contemporaneously. Our goal was to demonstrate a method to infer use patterns by way of purchase data. We examined purchase patterns for three types of personal care products (cosmetics, hair care, and skin care) and two household care products (household cleaners and laundry supplies) using data from 60,000 households collected over a one-year period in 2012. The market basket analysis methodology frequent itemset mining (FIM) was used to identify co-occurring sets of product purchases for all households and demographic groups based on income, education, race/ethnicity, and family composition. Our methodology captured robust co-occurrence patterns for personal and household products, globally and for different demographic groups. FIM identified cosmetic co-occurrence patterns captured in prior surveys of cosmetic use, as well as a trend of increased diversity of cosmetic purchases as children mature to teenage years. We propose that consumer product purchase data can be mined to inform person-oriented use patterns for high-throughput chemical screening applications, for aggregate and combined chemical risk evaluations.
Background: Chemicals in consumer products are a major contributor to human chemical coexposures. Consumers purchase and use a wide variety of products containing potentially thousands of chemicals. There is a need to identify potential real-world chemical coexposures to prioritize in vitro toxicity screening. However, due to the vast number of potential chemical combinations, this identification has been a major challenge. Objectives: We aimed to develop and implement a data-driven procedure for identifying prevalent chemical combinations to which humans are exposed through purchase and use of consumer products. Methods: We applied frequent itemset mining to an integrated data set linking consumer product chemical ingredient data with product purchasing data from 60,000 households to identify chemical combinations resulting from co-use of consumer products. Results: We identified co-occurrence patterns of chemicals over all households as well as those specific to demographic groups based on race/ethnicity, income, education, and family composition. We also identified chemicals with the highest potential for aggregate exposure by identifying chemicals occurring in multiple products used by the same household. Last, a case study of chemicals active in estrogen and androgen receptor in silico models revealed priority chemical combinations co-targeting receptors involved in important biological signaling pathways. Discussion: Integration and comprehensive analysis of household purchasing data and product-chemical information provided a means to assess human near-field exposure and inform selection of chemical combinations for high-throughput screening in in vitro assays. https://doi.org/10.1289/EHP8610
Presentation to the International Society of Exposure Science (ISES) Meeting September 2020
Consumer product categorizations for use in predicting human chemical exposure provide a bridge between product composition data and consumer product use pattern information. Furthermore, the categories reflect other factors relevant to developing consumer product exposure scenarios, such as microenvironment of use (e.g., indoors or outdoors), method of application/form of release (e.g., spray versus liquid), release to various media, removal processes (e.g., rinse-off or wipe-off), and route-specific exposure factors (dermal surface areas of application, fraction of release in respirable form). While challenging, developing harmonized product categories can generalize the factors described above allowing for rapid parameterization of route-specific exposure scenario algorithms for new chemical/product applications and efficient utilization of new data on product use or composition. This can be accomplished via mapping product categories to likewise categorized release and use patterns or exposure factors. Here, hierarchical product use categories (PUCs) for consumer products that provide such mappings are presented and crosswalked with other internationally harmonized product categories for consumer exposure assessment. The PUCs were defined by applying use and exposure scenario information to the products in EPA's Chemical and Products Database (CPDat). This paper demonstrates how these PUCs are being used to rapidly parameterize algorithms for scenario-specific use, fate, and exposure in a probabilistic aggregate model of human exposure to chemicals used in consumer products. The PUCs provide a generic representation of consumer products for use in exposure assessment and provide an efficient framework for flexible and rapid data reporting and consumer exposure model parameterization.
The U.S. Environmental Protection Agency (EPA) is faced with the challenge of efficiently and credibly evaluating chemical safety often with limited or no available toxicity data. The expanding number of chemicals found in commerce and the environment, coupled with time and resource requirements for traditional toxicity testing and exposure characterization, continue to underscore the need for new approaches. In 2005, EPA charted a new course to address this challenge by embracing computational toxicology (CompTox) and investing in the technologies and capabilities to push the field forward. The return on this investment has been demonstrated through results and applications across a range of human and environmental health problems, as well as initial application to regulatory decision-making within programs such as the EPA's Endocrine Disruptor Screening Program. The CompTox initiative at EPA is more than a decade old. This manuscript presents a blueprint to guide the strategic and operational direction over the next 5 years. The primary goal is to obtain broader acceptance of the CompTox approaches for application to higher tier regulatory decisions, such as chemical assessments. To achieve this goal, the blueprint expands and refines the use of high-throughput and computational modeling approaches to transform the components in chemical risk assessment, while systematically addressing key challenges that have hindered progress. In addition, the blueprint outlines additional investments in cross-cutting efforts to characterize uncertainty and variability, develop software and information technology tools, provide outreach and training, and establish scientific confidence for application to different public health and environmental regulatory decisions.
Chemical risk assessment relies on knowledge of hazard, the dose-response relationship, and exposure to characterize potential risks to public health and the environment. A chemical with minimal toxicity might pose a risk if exposures are extensive, repeated, and/or occurring during critical windows across the human life span. Exposure assessment involves understanding human activity, and this activity is confounded by interindividual variability that is both biological and behavioral. Exposures further vary between the general population and susceptible or occupationally exposed populations. Recent computational exposure efforts have tackled these problems through the creation of new tools and predictive models. These tools include machine learning to draw inferences from existing data and computer-enhanced screening analyses to generate new data. Mathematical models provide frameworks describing chemical exposure processes. These models can be statistically evaluated to establish rigorous confidence in their predictions. The computational exposure tools reviewed here are oriented toward 'high-throughput' application, that is, they are suitable for dealing with the thousands of chemicals in commerce with limited sources of chemical exposure information. These new tools and models are moving chemical exposure and risk assessment forward in the 21st century.
Evaluation of biomonitoring data supports the importance of near-field chemical exposure pathways. High throughput (HT) exposure models for consumer products have been developed, but require large amounts of composition, product use, and exposure factor data. EPA’s ExpoCast program has worked to obtain and organize publicly-available information for thousands of chemicals to support these models. However, challenges arise in evaluating the applicability of these data to new models, chemicals, products, uses, or populations. Thoughtful model design and data organization can promote reuse of collected information, as can the definition of formal linkages among model inputs and model algorithms describing exposure processes. For example, efficiently modeling chemicals in consumer products is enabled by a fit-for-purpose system of consumer product categories. Categories can be linked to generic exposure scenarios that define indoor fate and transport, route-specific intake, and disposal, and to associated product-specific factors (e.g., use and release patterns), to create a library of fully parameterized algorithms. In this framework, obtaining results for a new chemical does not require de novo parameterization, only knowledge of its concentration in products. In addition, defining generic chemicals within categories (e.g., via Latin hypercube sampling of physicochemical space for chemicals in products) will allow for rapid read-across of exposure results. While near-field models may adequately capture chemical-to-chemical variability in population median exposure, new HT ambient exposure (e.g., associated with industrial releases) or occupational exposure models will be required to address populations having outlier exposure patterns. Obtaining and organizing available data and algorithms for these populations is a current focus in ExpoCast. This abstract may not reflect U.S. EPA policy.
Exposure to a chemical is a critical consideration in the assessment of risk, as it adds real-world context to toxicological information. Descriptions of where and how individuals spend their time are important for characterizing exposures to chemicals in consumer products and in indoor environments. Herein we create an agent-based model (ABM) that simulates longitudinal patterns in human behavior. By basing the ABM upon an artificial intelligence (AI) system, we create agents that mimic human decisions on performing behaviors relevant for determining exposures to chemicals and other stressors. We implement the ABM in a computer program called the Agent-Based Model of Human Activity Patterns (ABMHAP) that predicts the longitudinal patterns for sleeping, eating, commuting, and working. We then show that ABMHAP is capable of simulating behavior over extended periods of time. We propose that this framework, and models based on it, can generate longitudinal human behavior data for use in exposure assessments.
Within the US EPA Office of Research and Development's (ORD's) Chemicals for Safety and Sustainability (CSS) research program, efforts are underway to provide information that is useful for the risk-based prioritization of thousands of chemicals, particularly those found in consumer products. The information has been made publicly available through ORD's CompTox Chemistry Dashboard (https://comptox.epa.gov). In this presentation the research efforts that contribute to the collection and curation of consumer product formulations, article compositions and ingredient uses will be discussed. This information is obtained through both automated collection techniques from manufacturer-supplied documents and measurement of chemicals via suspect screening analysis. Finally, predictions from ORD-developed models for such values as physicochemical properties, functional use, product composition, and chemical exposure, which are provided via the Chemistry Dashboard, will be discussed. The views expressed in this abstract are those of the authors and do not necessarily reflect the views or policies of the US Environmental Agency.
Quantitative data on product chemical composition is a necessary parameter for characterizing near-field exposure. This data set comprises reported and predicted information on more than 75,000 chemicals and more than 15,000 consumer products. The data's primary intended use is for exposure, risk, and safety assessments. The data set includes specific products with quantitative or qualitative ingredient information, which has been publicly disclosed through material safety data sheets (MSDS) and ingredient lists. A single product category from a refined and harmonized set of categories has been assigned to each product. The data set also contains information on the functional role of chemicals in products, which can inform predictions of the concentrations in which they occur. These data will be useful to exposure and risk assessors evaluating chemical and product safety.
Multi-city population-based epidemiological studies of short-term fine particulate matter (PM2.5) exposures and mortality have observed heterogeneity in risk estimates between cities. Factors affecting exposures, such as pollutant infiltration, which are not captured by central-site monitoring data, can differ between communities potentially explaining some of this heterogeneity. This analysis evaluates exposure factors as potential determinants of the heterogeneity in 312 core-based statistical areas (CBSA)-specific associations between PM2.5 and mortality using inverse variance weighted linear regression. Exposure factor variables were created based on data on housing characteristics, commuting patterns, heating fuel usage, and climatic factors from national surveys. When survey data were not available, air conditioning (AC) prevalence was predicted utilizing machine learning techniques. Across all CBSAs, there was a 0.95% (Interquartile range (IQR) of 2.25) increase in non-accidental mortality per 10 µg/m3 increase in PM2.5 and significant heterogeneity between CBSAs. CBSAs with larger homes, more heating degree days, a higher percentage of home heating with oil had significantly (p < 0.05) higher health effect estimates, while cities with more gas heating had significantly lower health effect estimates. While univariate models did not explain much of heterogeneity in health effect estimates (R2 < 1%), multivariate models began to explain some of the observed heterogeneity (R2 = 13%).
Background/Aim: Air pollution levels in fast-growing sub-Saharan African cities are among the highest in the world, but human exposure studies are limited, especially, traffic-related exposures. Our aim was to measure personal fine particulate matter (PM2.5) exposure of commercial minibus and taxi drivers (the most popular means of public transportation in Ghana's capital), and street mobile vendors (hawkers) and street stationary vendors (vendors) in Accra, Ghana. Methods: We measured 24-hour personal PM2.5 exposure of 99 subjects, comprising 29 minibus drivers, 26 taxi drivers, 29 street vendors, and 15 street hawkers in the Accra metropolis. PM2.5 was measured both gravimetrically and continuously, with time-matched global positioning system coordinates. The instruments were placed in backpacks, which were worn by the vendors and hawkers during the 24-hour measurement period. Taxi drivers had the backpacks placed beside them near the front passenger's seat. For the mini bus drivers, field assistants carried the backpacks on their laps at the front passenger's seat and rode along the drivers. Results: Across all four occupational groups, average (SD) personal PM2.5 exposure was 56.4 (63.2) μg/m3; group means ranged from 26.0 μg/m3 for hawkers to 83.4 μg/m3 for taxi drivers. Individual exposure was > 100 μg/m3 for some drivers. Exposure was significantly higher among drivers than vendors (78 vs 29 μg/m3; 95% CI: 28, 70; p-value < 0.001). The highest exposure to the drivers occurred on major and secondary roads, although < 20% of their commute time was spent there compared to minor roads and alleys. Conclusions: Our results support the need for urban air quality management plans that will address the role traffic to reduce exposure to the thousands who commute daily by minibuses and taxis or work near roadways in Accra, and when combined with policies to reduce urban biomass use, will greatly reduce urban population exposure in Accra.