Environmental risk assessments often rely on measured concentrations in environmental matrices to characterize exposure of the population of interest-typically, humans, aquatic biota, or other wildlife. Yet, there is limited guidance available on how to report and evaluate exposure datasets for reliability and relevance, despite their importance to regulatory decision-making. This paper is the second of a four-paper series detailing the outcomes of a Society of Environmental Toxicology and Chemistry Technical Workshop that has developed Criteria for Reporting and Evaluating Exposure Datasets (CREED). It presents specific criteria to systematically evaluate the reliability of environmental exposure datasets. These criteria can help risk assessors understand and characterize uncertainties when existing data are used in various types of assessments and can serve as guidance on best practice for the reporting of data for data generators (to maximize utility of their datasets). Although most reliability criteria are universal, some practices may need to be evaluated considering the purpose of the assessment. Reliability refers to the inherent quality of the dataset and evaluation criteria address the identification of analytes, study sites, environmental matrices, sampling dates, sample collection methods, analytical method performance, data handling or aggregation, treatment of censored data, and generation of summary statistics. Each criterion is evaluated as "fully met," "partly met," "not met or inappropriate," "not reported," or "not applicable" for the dataset being reviewed. The evaluation concludes with a scheme for scoring the dataset as reliable with or without restrictions, not reliable, or not assignable, and is demonstrated with three case studies representing both organic and inorganic constituents, and different study designs and assessment purposes. Reliability evaluation can be used in conjunction with relevance evaluation (assessed separately) to determine the extent to which environmental monitoring datasets are "fit for purpose," that is, suitable for use in various types of assessments. Integr Environ Assess Manag 2024;20:981-1003. © 2024 Society of Environmental Toxicology & Chemistry (SETAC). This article has been contributed to by U.S. Government employees and their work is in the public domain in the USA.
Aim: Analysis of wastewater samples can be used to assess population drug use, but reporting and statistical issues have limited the utility of the approach for epidemiology due to analytical results that are below the limit of quantification or detection. Unobserved or non-quantifiable-censored-data are common and likely to persist as the methodology is applied to more municipalities and a broader array of substances. We demonstrate the use of censored data techniques and account for measurement errors to explore distributions and annual estimates of the daily mean level of drugs excreted per capita. Measurements: Daily 24-hour composite wastewater samples for 56 days in 2009 were obtained using a random sample stratified by day of week and season for 19 municipalities in the Northwest region of the U.S.Methods: Methamphetamine, benzoylecgonine (cocaine metabolite), 3,4 methylenedioxymethamphetamine (MDMA), methadone, oxycodone and hydrocodone were identified and quantified in wastewater samples. Four statistical approaches (reporting censoring, Maximum Likelihood Estimation, Kaplan-Meier estimates, or complete data calculations) were used to estimate an annual average, including confidence bounds where appropriate, dependent upon the amount of censoring in the data.Findings: The proportion of days within a year with censored data varied greatly by drug across the 19 municipalities, with MDMA varying the most (4% to 94% of observations censored). The different statistical approaches each needed to be used given the levels of censoring of measured drug concentrations. Figures incorporating confidence bounds allow visualization of the data that facilitates appropriate comparisons across municipalities. Conclusions: Results from wastewater sampling that are below detection or quantification limits contain important information and can be incorporated to create a more complete and valid estimate of drug excretion. (C) 2016 Published by Elsevier B.V.
Praise for the First Edition" . . . an excellent addition to an upper-level undergraduate course on environmental statistics, and . . . a 'must-have' desk reference for environmental practitioners dealing with censored datasets." Vadose Zone JournalStatistical Methods for Censored Environmental Data Using Minitab and R, Second Edition introduces and explains methods for analyzing and interpreting censored data in the environmental sciences. Adapting survival analysis techniques from other fields, the book translates well-established methods from other disciplines into new solutions for environmental studies. This new edition applies methods of survival analysis, including methods for interval-censored data to the interpretation of low-level contaminants in environmental sciences and occupational health. Now incorporating the freely available R software as well as Minitab into the discussed analyses, the book features newly developed and updated material including: A new chapter on multivariate methods for censored data Use of interval-censored methods for treating true nondetects as lower than and separate from values between the detection and quantitation limits ("remarked data") A section on summing data with nondetects A newly written introduction that discusses invasive data, showing why substitution methods fail Expanded coverage of graphical methods for censored data The author writes in a style that focuses on applications rather than derivations, with chapters organized by key objectives such as computing intervals, comparing groups, and correlation. Examples accompany each procedure, utilizing real-world data that can be analyzed using the Minitab and R software macros available on the book's related website, and extensive references direct readers to authoritative literature from the environmental sciences.Statistics for Censored Environmental Data Using Minitab and R, Second Edition is an excellent book for courses on environmental statistics at the upper-undergraduate and graduate levels. The book also serves as a valuable reference for?environmental professionals, biologists, and ecologists who focus on the water sciences, air quality, and soil science.
This chapter contains sections titled: Why Not Substitute? Nonparametric Methods After Censoring at the Highest Reporting Limit Maximum Likelihood Estimation Akritas–Theil–Sen Nonparametric Regression Additional Methods for Censored Regression Exercises
This chapter contains sections titled: A Brief Overview of R and the NADA Software Summary of the Commands Available in NADA
This chapter contains sections titled: Parametric Intervals Nonparametric Intervals Intervals for Censored Data by Substitution Intervals for Censored Data by Maximum Likelihood Intervals for the Lognormal Distribution Intervals Using “Robust” Parametric Methods Nonparametric Intervals for Censored Data Bootstrapped Intervals For Further Study Exercises
This chapter contains sections titled: Substitution Does Not Work—Invasive Data Nonparametric Methods after Censoring at the Highest Reporting Limit Maximum Likelihood Estimation Nonparametric Method—The Generalized Wilcoxon Test Summary Exercises
Previously reported dendrochemical data showed temporal variability in concentration of tungsten (W) and cobalt (Co) in tree rings of Fallon, Nevada, US. Criticism of this work questioned the use of the Mann–Whitney test for determining change in element concentrations. Here, we demonstrate that Mann–Whitney is appropriate for comparing background element concentrations to possibly elevated concentrations in environmental media. Given that Mann–Whitney tests for differences in shapes of distributions, inter-tree variability (e.g., “coefficient of median variation”) was calculated for each measured element across trees within subsites and time periods. For W and Co, the metals of highest interest in Fallon, inter-tree variability was always higher within versus outside of Fallon. For calibration purposes, this entire analysis was repeated at a different town, Sweet Home, Oregon, which has a known tungsten-powder facility, and inter-tree variability of W in tree rings confirmed the establishment date of that facility. Mann–Whitney testing of simulated data also confirmed its appropriateness for analysis of data affected by point-source contamination. This research adds important new dimensions to dendrochemistry of point-source contamination by adding analysis of inter-tree variability to analysis of central tendency. Fallon remains distinctive by a temporal increase in W beginning by the mid 1990s and by elevated Co since at least the early 1990s, as well as by high inter-tree variability for W and Co relative to comparison towns.
This chapter contains sections titled: Approach 1: Nonparametric Methods after Censoring at the Highest Reporting Limit Approach 2: Maximum Likelihood Estimation Approach 3: Nonparametric Survival Analysis Methods Application of Survival Analysis Methods to Environmental Data Parallels to Uncensored Methods
This chapter contains sections titled: Why Not Substitute—Missing the Signals that Are Present in the Data Why Not Substitute?—Finding Signals that Are Not There So Why Not Substitute? Other Common Misuses of Censored Data
Release 4.0 of the U.S. Geological Survey S-PLUS library supercedes release 2.1. It comprises functions, dialogs, and datasets used in the U.S. Geological Survey for the analysis of water-resources data. This version does not contain ESTREND, which was in version 2.1. See Release 2.1 for information and access to that version. This library requires Release 8.1 or later of S-PLUS for Windows. S-PLUS is a commercial statistical and graphical analysis software package produced by TIBCO corporation(http://www.tibco.com/). The USGS library is not supported by TIBCO or its technical support staff.
Debris-retention basins in Southern California are frequently used to protect communities and infrastructure from the hazards of flooding and debris flow. Empirical models that predict sediment yields are used to determine the size of the basins. Such models have been developed using analyses of records of the amount of material removed from debris retention basins, associated rainfall amounts, measures of watershed characteristics, and wildfire extent and history. In this study we used multiple linear regression methods to develop two updated empirical models to predict sediment yields for watersheds located in Southern California. The models are based on both new and existing measures of volume of sediment removed from debris retention basins, measures of watershed morphology, and characterization of burn severity distributions for watersheds located in Ventura, Los Angeles, and San Bernardino Counties. The first model presented reflects conditions in watersheds located throughout the Transverse Ranges of Southern California and is based on volumes of sediment measured following single storm events with known rainfall conditions. The second model presented is specific to conditions in Ventura County watersheds and was developed using volumes of sediment measured following multiple storm events. To relate sediment volumes to triggering storm rainfall, a rainfall threshold was developed to identify storms likely to have caused sediment deposition. A measured volume of sediment deposited by numerous storms was parsed among the threshold-exceeding storms based on relative storm rainfall totals. The predictive strength of the two models developed here, and of previously-published models, was evaluated using a test dataset consisting of 65 volumes of sediment yields measured in Southern California. The evaluation indicated that the model developed using information from single storm events in the Transverse Ranges best predicted sediment yields for watersheds in San Bernardino, Los Angeles, and Ventura Counties. This model predicts sediment yield as a function of the peak 1-hour rainfall, the watershed area burned by the most recent fire (at all severities), the time since the most recent fire, watershed area, average gradient, and relief ratio. The model that reflects conditions specific to Ventura County watersheds consistently under-predicted sediment yields and is not recommended for application. Some previously-published models performed reasonably well, while others either under-predicted sediment yields or had a larger range of errors in the predicted sediment yields.
Regional-scale variations in soil geochemistry were investigated in a 20,000-km2 study area in northern California that includes the western slope of the Sierra Nevada, the southern Sacramento Valley and the northern Coast Ranges. Over 1300 archival soil samples collected from the late 1970s to 1980 in El Dorado, Placer, Sutter, Sacramento, Yolo and Solano counties were analyzed for 42 elements by inductively coupled plasma-atomic emission spectrometry and inductively coupled plasma-mass spectrometry following a near-total dissolution. These data were supplemented by analysis of more than 500 stream-sediment samples from higher elevations in the Sierra Nevada from the same study site. The relatively high-density data (1 sample per 15km2 for much of the study area) allows the delineation of regional geochemical patterns and the identification of processes that produced these patterns. The geochemical results segregate broadly into distinct element groupings whose distribution reflects the interplay of geologic, hydrologic, geomorphic and anthropogenic factors. One such group includes elements associated with mafic and ultramafic rocks including Cr, Ni, V, Co, Cu and Mg. Using Cr as an example, elevated concentrations occur in soils overlying ultramafic rocks in the foothills of the Sierra Nevada (median Cr=160mg/kg) as well as in the northern Coast Ranges. Low concentrations of these elements occur in soils located further upslope in the Sierra Nevada overlying Tertiary volcanic, metasedimentary and plutonic rocks (granodiorite and diorite). Eastern Sacramento Valley soil samples, defined as those located east of the Sacramento River, are lower in Cr (median Cr=84mg/kg), and are systematically lower in this suite compared to soils from the west side of the Sacramento Valley (median Cr=130mg/kg). A second group of elements showing a coherent pattern, including Ca, K, Sr and REE, is derived from relatively silicic rocks types. This group occurs at elevated concentrations in soils overlying volcanic and plutonic rocks at higher elevations in the Sierras (e.g. median La=28mg/kg) and the east side of the Sacramento Valley (median 20mg/kg) compared to soils overlying ultramafic rocks in the Sierra Nevada foothills (median 15mg/kg) and the western Sacramento Valley (median 14mg/kg). The segregation of soil geochemistry into distinctive groupings across the Sacramento River arises from the former presence of a natural levee (now replaced by an artificial one) along the banks of the river. This levee has been a barrier to sediment transport. Sediment transport to the Valley by glacial outwash from higher elevations in the Sierra Nevada and, more recently, debris from placer Au mining has dominated sediment transport to the eastern Valley. High content of mafic elements (and low content of silicic elements) in surface soil in the west side of the valley is due to a combination of lack of silicic source rocks, transport of ultramafic rock material from the Coast Ranges, and input of sediment from the late Mesozoic Great Valley Group, which is itself enriched in mafic elements. A third group of elements (Zn, Cd, As and Cu) reflect the impact of mining activity. Soil with elevated content of these elements occurs along the Sacramento River in both levee and adjacent flood basin settings. It is interpreted that transport of sediment down the Sacramento River from massive sulfide mines in the Klamath Mountains to the north has caused this pattern. The Pb, and to some extent Zn, distribution patterns are strongly impacted by anthropogenic inputs. Elevated Pb content is localized in major cites and along major highways due to inputs from leaded gasoline. Zinc has a similar distribution pattern but the source is tire wear.