Many data collection efforts and modeling studies have focused on providing accurate estimates of streamflow while fewer efforts have sought to identify when and where surface water is present and the duration of surface water presence in stream channels, hereafter referred to as streamflow permanence. While physically-based hydrological models are frequently used to explore how water quantity may be influenced by various climatic and basin characteristics at local, regional, national, and global extents they are less often used to explore streamflow permanence. Herein, the Watershed Erosion Prediction Project (WEPP) hydrological model is applied to watersheds in the humid H. J. Andrews Experimental Forest (HJA) and watersheds of the arid Willow and Whitehorse creeks (WW), both in Oregon, to simulate daily (WW) and annual (HJA and WW) streamflow permanence. One thousand parameter combinations were tested to calibrate WEPP to observed streamflow in the HJA watersheds and one hundred parameter combinations were tested to calibrate WEPP to observed surface water presence time series data in WW watersheds. When calibrated to observed streamflow, WEPP correctly classified annual streamflow permanence for 83 % of HJA stream reaches. In the WW, WEPP simulations correctly classified 63-93 % of daily streamflow permanence observations and 59-87 % of annual streamflow permanence classifications. Inclusion of a dry-day threshold (the maximum number of days a stream reach could be modeled 'dry' but still classified as permanent for the year) improved annual accuracy in three WW watersheds from 2 to 10 %. Parameter sets that produced the best daily accuracies in WW resulted in poor annual accuracies. Results highlight the importance of evaluating physically-based streamflow permanence models on both permanent and nonpermanent streams at daily and annual time scales to ensure evaluation metrics are appropriate for interpretation purposes. Additionally, results suggest that strategic collection of surface water presence observations and streamflow observations may support robust calibration of physically based models to simulate streamflow permanence moving forward.
Agricultural crop insurance is an important component for mitigating farm risk, particularly given the potential for unexpected climatic events. Using a 2.8 million nationwide insurance claim dataset from the United States Department of Agriculture (USDA), this research study examines spatiotemporal variations of over 31,000 agricultural insurance loss claims across the 24-county region of the inland Pacific Northwest (iPNW) portion of the United States from 2001 to 2022. Wheat is the dominant insurance loss crop for the region, accounting for over USD 2.8 billion in indemnities, with over USD 1.5 billion resulting in claims due to drought (across the 22 year time period). While fruit production generates considerably lesser insurance losses (USD 400 million) as a primary result of freeze, frost, and hail, overall revenue ranks number one for the region, with USD 2 billion in sales, across the same time range. Principal components analysis of crop insurance claims showed distinct spatial and temporal differentiation in wheat and apples insurance losses using the range of damage causes as factor loadings. The first two factor loadings for wheat accounts for approximately 50 percent of total variance for the region, while a separate analysis of apples accounts for over 60 percent of total variance. These distinct orthogonal differences in losses by year and commodity in relationship to damage causes suggest that insurance loss analysis may serve as an effective barometer in gauging climatic influences.
Stream permanence classifications (i.e., perennial, intermittent, ephemeral) are a primary consideration to determine stream regulatory status in the United States (U.S.) and are an important indicator of environmental conditions and biodiversity. However, at present, no models or products adequately describe surface water presence for regulatory determinations. We modified the Thornthwaite monthly water balance model (MWBM) with a flow threshold parameter to estimate flow permanence and evaluated the model’s accuracy and precision for more than 1.3 million headwater stream reaches in the U.S. Pacific Northwest (PNW). Stream reaches were assigned to one of eight calibration groups by unsupervised classification based on sensitivity to MWBM parameters. Suitable MWBM parameter sets were identified by comparing modeled stream permanence estimates to surface water presence observations (SWPO). Parameter sets with accuracies > 65% were considered suitable. The MWBM estimated stream permanence with high precision at 40% of reaches, with poor precision at 20% of reaches, and no suitable parameter sets were identified for 40% of reaches. Results highlight the need for increased SWPO collection to improve calibration and assessment of stream permanence models. Additionally, implementation of the MWBM to estimate surface water presence indicates potential for process-based models to predict stream permanence with future development.
We compared climatic relationships to insurance loss across the inland Pacific Northwest region of the United States, using a design matrix methodology, to identify optimum temporal windows for climate variables by county in relationship to wheat insurance loss due to drought. The results of our temporal window construction for water availability variables (precipitation, temperature, evapotranspiration, and the Palmer drought severity index [PDSI]) identified spatial patterns across the study area that aligned with regional climate patterns, particularly with regards to drought-prone counties of eastern Washington. Using these optimum time-lagged correlational relationships between insurance loss and individual climate variables, along with commodity pricing, we constructed a regression-based random forest model for insurance loss prediction and evaluation of climatic feature importance. Our cross-validated model results indicated that PDSI was the most important factor in predicting total seasonal wheat/drought insurance loss, with wheat pricing and potential evapotranspiration having noted contributions. Our overall regional model had a $ {R}^2 $ of 0.49, and a RMSE of $30.8 million. Model performance typically underestimated annual losses, with moderate spatial variability in terms of performance between counties.
Forecasting the risk of pathogen spillover from reservoir populations of wild or domestic animals is essential for the effective deployment of interventions such as wildlife vaccination or culling. Due to the sporadic nature of spillover events and limited availability of data, developing and validating robust, spatially explicit, predictions is challenging. Recent efforts have begun to make progress in this direction by capitalizing on machine learning methodologies. An important weakness of existing approaches, however, is that they generally rely on combining human and reservoir infection data during the training process and thus conflate risk attributable to the prevalence of the pathogen in the reservoir population with the risk attributed to the realized rate of spillover into the human population. Because effective planning of interventions requires that these components of risk be disentangled, we developed a multi-layer machine learning framework that separates these processes. Our approach begins by training models to predict the geographic range of the primary reservoir and the subset of this range in which the pathogen occurs. The spillover risk predicted by the product of these reservoir specific models is then fit to data on realized patterns of historical spillover into the human population. The result is a geographically specific spillover risk forecast that can be easily decomposed and used to guide effective intervention. Applying our method to Lassa virus, a zoonotic pathogen that regularly spills over into the human population across West Africa, results in a model that explains a modest but statistically significant portion of geographic variation in historical patterns of spillover. When combined with a mechanistic mathematical model of infection dynamics, our spillover risk model predicts that 897,700 humans are infected by Lassa virus each year across West Africa, with Nigeria accounting for more than half of these human infections.
National Hydrography Dataset (NHD) stream permanence classifications (SPC; perennial, intermittent, and ephemeral) are widely used for data visualization and applied science, and have implications for resource policy and management. NHD SPC were assigned using a combination of topographic field surveys and interviews with local residents. However, previous studies indicate that non-NHD,in situstreamflow observations (NNO) frequently disagree with NHD SPC. We hypothesized that differences in annual climate conditions between map creation years and the years NNO were collected contributed to disagreement between NNO and NHD SPC. We compared NHD SPC to 10,055 NNO (classified as "wet" or "dry") collected in the Pacific Northwest between 1977 and 2015. Annual climate conditions were described with the Palmer Drought Severity Index (PDSI). Stream order was added as a covariate to account for different effects along the stream network. NHD SPC agreed with 80.5% of NNO. "Dry" NNO were five times more likely to disagree with NHD than "wet" NNO (p < 0.0001). Disagreement was greatest on first-order streams. When NHD SPC were collected during a wetter period than NNO the probability of disagreement increased by a factor of 1.17 (p < 0.0001) per unit difference in PDSI. The influence of climate on disagreements between NNO and NHD SPC provides support for the continued development of dynamic models representing SPC as opposed to static NHD classifications.
As funding agencies embrace open science principles that encourage sharing data and computer code developed to produce research outputs, we must respond with new modes of publication. Furthermore, as we address the expanding reproducibility crisis in the sciences, we must work to release research materials in ways that enable reproducibility-publishing data, computer code, and research products in addition to the traditional journal article. Toward addressing these needs, we present an example framework to model and map soil organic carbon (SOC) in the cereal grains production region of the northwestern United States. Primarily associated with soil organic matter, SOC relates to many soil properties that influence resiliency and soil health for agriculture. It is also critical for understanding soil-atmospheric C flux, a significant part of the overall C budget of the Earth. The technique for modeling soil properties uses seven categories of environmental input data to make predictions: known soil attributes, climatic values, organisms present, relief, parent material, age, and spatial location. We gather data representing these categories from various public sources. The map is produced using a random forest statistical model with inputs to predict SOC content on a 30-m spatial grid. All modeling components including input data, metadata, computer code, and output products are made freely available under an explicit open-source license. In this way, reproducibility is supported, the methods and code released are available to be reused by other researchers, and the research products are open to critical review and improvement.
The Consultative Committee for Space Data Systems (CCSDS), in 2002, released their first version of a Reference Model for an Open Archival Information System (OAIS). In 2003, the model was adopted by the International Standards Organization (ISO) as ISO 14721:2003. The CCSDS document was updated in 2012 with additional focus on verifying the authenticity of data and developing concepts of access rights and a security model. The OAIS model is the basis of research data management systems across institutions and disciplines around the world. The Organization for the Advancement of Structured Information Standards (OASIS), in 2006, released their first version of a Reference Model for Service Oriented Architecture (SOA). OASIS defines the SOA as “a paradigm for organizing and utilizing distributed capabilities that may be under the control of different ownership domains.” Systems designed around the SOA model benefit from improved scalability, flexibility, and agility. This paper applies the SOA model to the OAIS repository to describe how repositories can be implemented and extended through the use of services that may be internal or external to the host institution, including the consumption of network- or cloud-based services and resources. We use the Service Oriented Architecture (SOA) design paradigm to describe a set of potential extensions to OAIS Reference Model: purpose and justification for each extension, where and how each extension connects to the model, and an example of a specific service that meets the purpose.
Core IdeasPodzolization is a widespread pedogenic process in the Northern Rocky Mountains.Podzolization is strongly influenced by terrain attributes and environmental gradients.Northern Rocky Mountain landscapes are dominated by Spodosols, spodic intergrades, and Andisols.Although not recognized on land system inventory maps, Spodosols have been documented in the volcanic ash–mantled landscapes of the Northern Rocky Mountains (Major Land Resource Area 43A). This study focuses on terrain attributes and environmental gradients as potential predictors of the distribution and properties of podzolized soils in the region. Seventy‐two forested sites were sampled using hierarchical randomization. Podzolization, as indicated by the presence of a spodic horizon, was observed at 43 sites. Soils meeting all or most taxonomic criteria for Spodosols were present at 26 sites. These Spodosols are concentrated above elevations of approximately 1100 m and on northerly aspects. In this mountainous region, more northerly aspects and increasing elevation serve as proxies for lower temperatures, increased effective precipitation, and greater snow cover, all of which promote podzolization. Spodosol E horizons are very strongly acidic, with pH values as low as 3.1 and Al saturation as high as 89% of the effective cation exchange capacity. Other soils exhibit evidence of podzolization as interpreted from subsurface increases in pyrophosphate‐ and oxalate‐extractable Fe and Al. However, most of these soils have darker and/or higher‐chroma colors in the surface mineral horizon and lack the albic–spodic morphology of Spodosols. Results indicate that podzolization is a widespread pedogenic process in the Northern Rocky Mountains and occurs under different climatic conditions, topographic settings, and parent materials than those of the upper Midwest and northeastern United States.
Robust classification approaches are required for accurate classification of complex land-use/land-cover categories of desert landscapes using remotely sensed data. Machine-learning ensemble classifiers have proved to be powerful for the classification of remotely sensed data. However, they have not been evaluated for classifying land-cover categories in desert regions. In this study, the performance of two machine-learning ensemble classifiers – random forests (RF) and boosted artificial neural networks – is explored in the context of classification of land use/land cover of desert landscapes. The evaluation is based on the accuracy of classification of remotely sensed data, with and without integration of ancillary data. Landsat-5 Thematic Mapper data captured for a desert landscape in the north-western coastal desert of Egypt are used with ancillary variables derived from a digital terrain model to classify 13 different land-use/land-cover categories. Results show that the two ensemble methods produce accurate land-cover classifications, with and without integrating spectral data with ancillary data. In general, the overall accuracy exceeded 85% and the kappa coefficient (κ) attained values over 0.83. The integration of ancillary data improved the performance of the boosted artificial neural networks by approximately 5% and the random forests by 9%. The latter showed overall higher accuracy; however, boosted artificial neural networks showed better generalization ability and lower overfitting tendencies. The results reveal the merit of applying ensemble methods to integrated spectral and ancillary data of similar desert landscapes for achieving high classification accuracies.
Detecting land-use change has become of concern to environmentalists, conservationists and land use planners due to its impact on natural ecosystems. We studied land use/land cover (LULC) changes in part of the northwestern desert of Egypt and used the Markov-CA integrated approach to predict future changes. We mapped the LULC distribution of the desert landscape for 1988, 1999, and 2011. Landsat Thematic Mapper 5 data and ancillary data were classified using the random forests approach. The technique produced LULC maps with an overall accuracy of more than 90%. Analysis of LULC classes from the three dates revealed that the study area was subjected to three different stages of modification, each dominated by different land uses. The use of a spatially explicit land use change modeling approach, such as Markov-CA approach, provides ways for projecting different future scenarios. Markov-CA was used to predict land use change in 2011 and project changes in 2023 by extrapolating current trends. The technique was successful in predicting LULC distribution in 2011 and the results were comparable to the actual LULC for 2011. The projected LULC for 2023 revealed more urbanization of the landscape with potential expansion in the croplands westward and northward, an increase in quarries, and growth in residential centers. The outcomes can help management activities directed toward protection of wildlife in the area. The study can also be used as a guide to other studies aiming at projecting changes in arid areas experiencing similar land use changes.
The application of species distribution modeling in deserts is a useful tool for mapping species and assessing the impact of human induced changes on individual species. Such applications are still rare, and this may be attributed to the fact that much of the arid lands and deserts around the world are located in inaccessible areas. Few studies have conducted spatially explicit modeling of plant species distribution in Egypt. The random forest modeling approach was applied to climatic and land-surface parameters to predict the distribution of ten important plant species in an arid landscape in the northwestern coastal desert of Egypt. The impact of changes in land use and climate on the distribution of the plant species was assessed. The results indicate that the changes in land use in the area have resulted in habitat loss for all the modeled species. Projected future changes in land use reveals that all the modeled species will continue to suffer habitat loss. The projected impact of modeled climate scenarios (A1B, A2A and B2A) on the distribution of the modeled species by 2040 varied. Some of the species were projected to be adversely affected by the changes in climate, while other species are expected to benefit from these changes. The combined impact of the changes in land use and climate pose serious threats to most of the modeled species. The study found that all the species are expected to suffer loss in habitat, except Gymnocarpos decanderus. The study highlights the importance of assessing the impact of land use/climate change scenarios on other species of restricted distribution in the area and can help shape policy and mitigation measures directed toward biodiversity conservation in Egypt.
Sustainable forest management requires timely, detailed forest inventory data across large areas, which is difficult to obtain via traditional forest inventory techniques. This study evaluated k-nearest neighbor imputation models incorporating LiDAR data to predict tree-level inventory data (individual tree height, diameter at breast height, and species) across a 12 100 ha study area in northeastern Oregon, USA. The primary objective was to provide spatially explicit data to parameterize the Forest Vegetation Simulator, a tree-level forest growth model. The final imputation model utilized LiDAR-derived height measurements and topographic variables to spatially predict tree-level forest inventory data. When compared with an independent data set, the accuracy of forest inventory metrics was high; the root mean square difference of imputed basal area and stem volume estimates were 5 m2 ha–1 and 16 m3 ha–1, respectively. However, the error of imputed forest inventory metrics incorporating small trees (e.g., quadratic mean diameter, tree density) was considerably higher. Forest Vegetation Simulator growth projections based upon imputed forest inventory data follow trends similar to growth projections based upon independent inventory data. This study represents a significant improvement in our capabilities to predict detailed, tree-level forest inventory data across large areas, which could ultimately lead to more informed forest management practices and policies. Résumé : L’aménagement durable des forêts requiert des données appropriées et détaillées d’inventaire forestier sur de grandes superficies, ce qui est difficile à obtenir par le biais de techniques traditionnelles d’inventaire forestier. Cette étude évalue des modèles d’imputation basés sur les k plus proches voisins incorporant des données lidar pour prédire des mesures d’inventaire à l’échelle de l’arbre (hauteur, diamètre à hauteur de poitrine et espèce des arbres individuels) dans une aire d’étude de 12 100 ha du nord-est de l’Oregon, aux États-Unis. L’objectif premier est de fournir des données spatialement explicites pour paramétrer un modèle de croissance forestière à l’échelle de l’arbre, le «Forest Vegetation Simulator». Le modèle final d’imputation utilise des mesures de hauteur et des variables topographiques dérivées du lidar pour prédire spatialement des données d’inventaire forestier à l’échelle de l’arbre. Lorsqu’elles ont été comparées à un fichier indépendant de données, la précision des mesures d’inventaire forestier était élevée: l’erreur quadratique moyenne des estimations imputées de surface terrière et de volume étaient respectivement de 5 m2 ha–1 et 16 m3 ha–1. Cependant, l’erreur des mesures imputées d’inventaire forestier qui tiennent compte des petits arbres (p. ex. le diamètre moyen quadratique et la densité des arbres) était considérablement plus élevée. Les projections de croissance du «Forest Vegetation Simulator» basées sur des données imputées d’inventaire forestier suivent une tendance similaire aux projections basées sur des données indépendantes d’inventaire. Cette étude représente une amélioration importante de nos capacités à prédire des données détaillées d’inventaire forestier à l’échelle de l’arbre sur de grandes superficies, ce qui pourrait éventuellement mener à des pratiques et des politiques d’aménagement forestier mieux fondées. [Traduit par la Rédaction]
Douglas-fir (Pseudotsuga menziesii [Mirb.] Franco) in the Inland Northwest region of the USA are nitrogen (N) deficient; however stem growth responses to N fertilizers are unpredictable, which may be due to poor accounting of other limiting nutrients. Screening trial experiments, including potassium (K), sulfur (S), and boron (B) multiple nutrient treatments, have been conducted to learn about Douglas-fir nutritional status and fertilizer growth response. The data from the screening trial experiments were compiled to test whether the soil parent materials of the region could be used to predict nutritional status. Estimating effects of fertilizers and soil parent materials on Douglas-fir growth from compilations of such experiments, however, poses challenges and opportunity; experiments clustered in time and space introduce latent variables that drive between-site variation. We used a two-stage modeling approach to efficiently take advantage of the information in these data. First, we employed a mixed model approach to test the primary hypothesis of soil parent material influence upon stem growth response to fertilizer. As the second-stage to the analysis, the predicted random effects estimated from the mixed model were used as a response variable to test how strongly precipitation drives between-site variation. As expected, including the random site effect significantly improved the model fit of the growth model (Lambda = 436.5, P < 0.0001). The full mixed model accounted for 85% of the variation in the growth data (R-2 = 0.85) and revealed an interaction between fertilizer treatment and soil parent material class (P = 0.0179). Post hoc analysis suggested that Douglas-fir growing on loessal soils are not constrained by K, S, or B, but no general consistency was apparent with tephra or underlying geology. The second stage modeling suggested that winter precipitation explains variation in predicted random site effects (r(2) = 0.23), and hence the growth difference, better than total precipitation. Also, the annual lag precipitation explains variation in predicted random effects comparably well (r(2) = 0.22). (C) 2012 Elsevier B.V. All rights reserved.