National monitoring of forestlands and the processes causing canopy cover loss, be they abrupt or gradual, partial or stand clearing, temporary (disturbance) or persisting (deforestation), are necessary at fine scales to inform management, science and policy. This study utilizes the Landsat archive and an ensemble of disturbance algorithms to produce maps attributing event type and timing to > 258 million ha of contiguous Unites States forested ecosystems (1986–2010). Nationally, 75.95 million forest ha (759,531 km2) experienced change, with 80.6% attributed to removals, 12.4% to wildfire, 4.7% to stress and 2.2% to conversion. Between regions, the relative amounts and rates of removals, wildfire, stress and conversion varied substantially. The removal class had 82.3% (0.01 S.E.) user’s and 72.2% (0.02 S.E.) producer’s accuracy. A survey of available national attribution datasets, from the data user’s perspective, of scale, relevant processes and ecological depth suggests knowledge gaps remain.
In light of Earth's changing climate and growing human population, there is an urgent need to improve monitoring of natural and anthropogenic disturbances which effect forests' ability to sequester carbon and provide other ecosystem services. In this study, a two-step modeling approach was used to map the type and timing of forest disturbances occurring between 1984 and 2010 in ten Landsat scenes located in diverse forest systems of the conterminous U.S. In step one, Random Forest (RF) models were developed to predict the presence of five forest disturbance agents (conversion, fire, harvest, stress and wind) and stable (i.e. undisturbed) forest. Models were developed using a suite of predictors including spectral change metrics derived from a nonparametric shape-restricted spline fitting algorithm, as well as several topographic and biophysical variables which potentially influence the initiation and/or spread of forest disturbance agents. Step two involved applying a rule-based model to the spectrally-based shape parameters (e.g. shape type, year and duration) to assign a year to the disturbance types and locations predicted in step one. Out of bag (OOB) predictions from RF showed that across the ten scenes, overall agreement was highest when only causal agent was considered (avg=80%, min=69%, max=86%), and was lowest when both agent and year (within ±1 of the reference date) were required to be correct (avg=71%, min=56%, max=80%). Across scene omission and commission errors for fire and stable forest classes were mostly around 10% to 20%, respectively. Harvests were also modeled well, as five of nine test scenes had error rates <26%. Accuracy of the wind and stress classes were much more variable with model errors ranging from 24% to 88%. The years assigned by the rule-based model were reasonably accurate, as 88% of all disturbances were assigned a year that fell within ±2years of the reference date. Fire disturbances were assigned the correct year 78% of the time, followed by harvest (69%) and conversion (54%). Although 17% and 63% of wind and stress disturbances were under-estimated by 5 or more years, the impact on overall accuracy was nominal given these two classes only accounted for roughly 5% of all disturbances. Our results also revealed that causal agent models summarized to broader disturbed/not disturbed classes were as accurate as models specifically constructed to predict binary disturbance, thus there appears to be no advantage to modeling disturbance prior to assigning causality. A relative evaluation of mean decrease in accuracy from RF showed that although a wide range of predictor variables contributed to the successful modeling of causal agents and stable forest (e.g. patch metrics, forest occurrence, and topography), disturbance variables (e.g. MTBS) and spectral change metrics (e.g. absolute and relative magnitude) were by far the most important. Modeled causality maps and annual disturbance rates were examined and found to be in good agreement with existing literature and other published data sets. Lastly, results are used to make recommendations for mapping forest disturbance agents nationally across the U.S.
The ModelMap package (Freeman, 2009) for R (R Development Core Team, 2008) has added two additional variants of random forests: quantile regression forests and conditional inference forests. The quantregForest package (Meinshausen and Schiesser, 2015) is used for quantile regression forest (QRF) models. QRF models provide the ability to map the predicted median and individual quantiles. This makes it possible to map lower and upper bounds for the predictions without relying on the assumption that the predictions of individual trees in the model follow a normal distribution. The party package (Hothorn et al., 2006; Strobl et al., 2007, 2008) is used for conditional inference forest (CF) models. CF models offer two advantages over traditional RF models: they avoid RF’s bias towards predictor variables with higher numbers of categories; and, they provide a conditional importance measure, allowing a better understanding of the relative importance of correlated predictor variables.
We present a new methodology for fitting nonparametric shape-restricted regression splines to time series of Landsat imagery for the purpose of modeling, mapping, and monitoring annual forest disturbance dynamics over nearly three decades. For each pixel and spectral band or index of choice in temporal Landsat data, our method delivers a smoothed rendition of the trajectory constrained to behave in an ecologically sensible manner, reflecting one of seven possible 'shapes'. It also provides parameters summarizing the patterns of each change including year of onset, duration, magnitude, and pre- and postchange rates of growth or recovery. Through a case study featuring fire, harvest, and bark beetle outbreak, we illustrate how resultant fitted values and parameters can be fed into empirical models to map disturbance causal agent and tree canopy cover changes coincident with disturbance events through time. We provide our code in the r package ShapeSelectForest on the Comprehensive R Archival Network and describe our computational approaches for running the method over large geographic areas. We also discuss how this methodology is currently being used for forest disturbance and attribute mapping across the conterminous United States.
Questions regarding the impact of natural and anthropogenic forest change events (temporary and persisting) on energy, water and nutrient cycling, forest sustainability and resilience, and ecosystem services call for a full suite of information on the spatial and temporal trends of forest dynamics. Temporal and spatial patterns of change along with their magnitude and cause are all equally important when weaving together the full story of our forests’ history. National statistical estimation and mapping of land use and cover changes have been progressing for decades. However, especially in the case of forest cover changes, attributing the magnitude and underlying causal processes to areas of change are newly developing endeavors. The NASA/NACP funded North American Forest Dynamics (NAFD) project has conducted nearly a decade of research in mapping U.S. forest dynamics using Landsat imagery. One part of this research is an empirical and rule-based modeling approach to attribute the casual processes underlying temporary forest changes from wind, fire, insects/stress, harvest and persisting change from land cover conversion. In this presentation we address model matters including the utility of using the outputs (temporal, spatial and magnitude) from multiple forest disturbance algorithms as predictors to reduce commission and omission among response classes, insufficient and imbalanced training data, model and map accuracy, as well as initial results from these maps of CONUS forest change causal processes over two and a half decades.
Recently completing over a decade of research, the NASA/NACP funded North American Forest Dynamics (NAFD) project has led to several important advancements in the way U.S. forest disturbance dynamics are mapped at regional and continental scales. One major contribution has been the development of an empirical and rule-based modeling approach which addresses two of the major challenges associated with mapping forest disturbance. The first challenge is that no single spectral band or index responds consistently to all disturbance types. To overcome this challenge we use a new, non-parametric shape-fitting algorithm to derive pixel-level temporal change metrics (e.g. timing, magnitude, duration) from four different types of Landsat trajectories. By incorporating both shortwave-infrared data (e.g. Landsat band 5) and near-infrared-based vegetation indices (e.g. NDVI, NBR) we increase capture of subtle changes which alter forest structure and/or canopy leaf area. The second challenge is that certain types of disturbance are influenced by topographic and biophysical factors which are not inherently captured by optical remote sensing data. To overcome this challenge we use Random Forest models to integrate multiple spectral and non-spectral predictor variables to map fires, harvests, wind damage, as well as stress brought on by insect/disease outbreaks and land use conversion resulting in permanent forest cover loss. In this presentation we show results from 10 Landsat scenes representing a diverse array of causal agents, forest types, and forest prevalence levels found across the country. Using these example scenes we discuss the construction and importance of various predictor variables, as well as examine how model prediction accuracy varies as a function of geographic location, forest type and input training data. Lastly, we discuss how these initial results are being used to guide development of a nationwide map aimed at improving quantification of continental scale disturbance rates occurring over the last two decades.
We present a new R package called ShapeSelectForest recently posted to the Comprehensive R Archival Network. The package was developed to fit nonparametric shape-restricted regression splines to time series of Landsat imagery for the purpose of modeling, mapping, and monitoring annual forest disturbance dynamics over nearly three decades. For each pixel and spectral band or index of choice in temporal Landsat data, the package delivers an optimally smoothed rendition of the trajectory constrained to behave in an ecologically sensible manner, assuming one of seven possible “shapes”. It also provides parameters summarizing the temporal pattern including year(s) of inflection, magnitude of change, and pre- and post- inflection rates of growth or recovery. In addition, the package contains functions for deriving annual predictions of forest disturbance, as well as graphical displays of the shape fits.
FIESTA (Forest Inventory ESTimation for Analysis) is a user-friendly R package that was originally developed to support the production of estimates consistent with current tools available for the Forest Inventory and Analysis (FIA) National Program, such as FIDO (Forest Inventory Data Online) and EVALIDator. FIESTA provides an alternative data retrieval and reporting tool that is functional within the R environment, allowing customized applications and compatibility with other R-based analyses. Over the last few years, the tool has expanded to include new modules that accommodate nonresponse, photo-based estimators, two-phase regression estimators for inclusion of temporal remote sensing data, as well as small area estimates. Here, we describe these new modules and illustrate with FIA applications in the Interior West.
As part of the development of the 2011 National Land Cover Database (NLCD) tree canopy cover layer, a pilot project was launched to test the use of high-resolution photography coupled with extensive ancillary data to map the distribution of tree canopy cover over four study regions in the conterminous US. Two stochastic modeling techniques, random forests (RF) and stochastic gradient boosting (SGB), are compared. The objectives of this study were first to explore the sensitivity of RF and SGB to choices in tuning parameters and, second, to compare the performance of the two final models by assessing the importance of, and interaction between, predictor variables, the global accuracy metrics derived from an independent test set, as well as the visual quality of the resultant maps of tree canopy cover. The predictive accuracy of RF and SGB was remarkably similar on all four of our pilot regions. In all four study regions, the independent test set mean squared error (MSE) was identical to three decimal places, with the largest difference in Kansas where RF gave an MSE of 0.0113 and SGB gave an MSE of 0.0117. With correlated predictor variables, SGB had a tendency to concentrate variable importance in fewer variables, whereas RF tended to spread importance among more variables. RF is simpler to implement than SGB, as RF has fewer parameters needing tuning and also was less sensitive to these parameters. As stochastic techniques, both RF and SGB introduce a new component of uncertainty: repeated model runs will potentially result in different final predictions. We demonstrate how RF allows the production of a spatially explicit map of this stochastic uncertainty of the final model.
Lidar height data collected by the Geosciences Laser Altimeter System (GLAS) from 2002 to 2008 has the potential to form the basis of a globally consistent sample-based inventory of forest biomass. GLAS lidar return data were collected globally in spatially discrete full waveform “shots,” which have been shown to be strongly correlated with aboveground forest biomass. Relationships observed at spatially coincident field plots may be used to model biomass at all GLAS shots, and well-established methods of model-based inference may then be used to estimate biomass and variance for specific spatial domains. However, the spatial pattern of GLAS acquisition is neither random across the surface of the earth nor is it identifiable with any particular systematic design. Undefined sample properties therefore hinder the use of GLAS in global forest sampling.
Random Forests is frequently used to model species distributions over large geographic areas. Complications arise when data used to train the models have been collected in stratified designs that involve different sampling intensity per stratum. The modeling process is further complicated if some of the target species are relatively rare on the landscape leading to an unbalanced number of presences and absences in the training data. We explored means to accommodate unequal sampling intensity across strata as well as the unbalanced species prevalence in Random Forest models for tree and shrub species distributions in the state of Nevada. For the unequal sampling intensity issue, we tested three modeling strategies: fitting models using all the data, down-sampling the intensified stratum; and building separate models for each stratum. We explored unbalanced species prevalence by investigating the effects of down-sampling the more prevalent response (presence or absence), and by optimizing the cutoff thresholds for declaring a species present. When modeling species presence with stratified data that was collected with different sampling intensities per stratum, we found that neither down-sampling the intensified stratum, nor fitting individual strata models, improved model performance. We also found that balancing the number of presences and absences in a training data set by down-sampling did not improve predictive models of species distributions, and did not eliminate the need to optimize thresholds. We then apply our final choice of model to the full raster layers for Nevada to produce statewide species distribution maps.
Light Detection and Ranging (LiDAR) returns from the spaceborne Geoscience Laser Altimeter (GLAS) sensor may offer an alternative to solely field-based forest biomass sampling. Such an approach would rely upon model-based inference, which can account for the uncertainty associated with using modeled, instead of field- collected, measurements. Model-based methods have been thoroughly described in the statistical literature, and an increasing number of model-based forestry applications use tactically acquired airborne LiDAR. Adapting these methods to GLAS's irregular acquisition pattern requires a strategy for identifying a subset of GLAS "shots" that can be considered a simple random sample. We have developed a flexible method of dividing the landscape into equal-area polygons from which a GLAS shot can be chosen at random as a member of the sample. This process bears similarities to the approach used by the Forest Inventory and Analysis (FIA) Program as it moved toward its current hexagonal sample grid. Although the ultimate application of this approach would be production of consistent biomass estimates across different countries, well-calibrated FIA estimates over the United States provide a convenient testing ground. Applied to California, this approach produced almost exactly the same estimate of biomass density (Mg/ha) as the FIA sample. The GLAS-based estimate had a considerably higher standard error than FIA's estimate, but it comes at a much lower cost and is based upon globally available GLAS measurements.
Understanding the potential of forest ecosystems as global carbon sinks requires a thorough knowledge of forest carbon dynamics, including both sequestration and fluxes among multiple pools. The accurate quantification of biomass is important to better understand forest productivity and carbon cycling dynamics. Stand-based inventories (SBIs) are widely used for quantifying forest characteristics and for estimating biomass, but information may quickly become outdated in dynamic forest environments. Satellite remote sensing may provide a supplement or substitute. We tested the accuracy of aboveground biomass estimates modeled from a combination of Landsat Thematic Mapper (TM) imagery and topographic data, as well as SBI-derived variables in a Picea abies forest in the Western Carpathian Mountains. We employed Random Forests for non-parametric, regression tree-based modeling. Results indicated a difference in the importance of SBI-based and remote sensing-based predictors when estimating aboveground biomass. The most accurate models for biomass prediction ranged from a correlation coefficient of 0.52 for the TM- and topography-based model, to 0.98 for the inventory-based model. While Landsat-based biomass estimates were measurably less accurate than those derived from SBI, adding tree height or stand-volume as a field-based predictor to TM and topography-based models increased performance to 0.36 and 0.86, respectively. Our results illustrate the potential of spectral data to reveal spatial details in stand structure and ecological complexity.
The Forest Inventory and Analysis program (FIA) of the U.S. Forest Service monitors status and trends in forested ecoregions nationwide.The complex nature of this broad-scale, strategic-level inventory demands constant evolution and evaluation of methods to get the best information possible while continuously increasing efficiency.In 2004, the "Nevada Photo-Based Inventory Pilot" (NPIP) was launched and involved the acquisition and processing of large-scale aerial photography (LSP) throughout the State of Nevada.The over-arching goals of this pilot are to exceed information requirements, accelerate inventory timelines, and reduce inventory costs.Meeting these objectives requires the development of several complex and inter-related procedures, including photo-sampling protocol, statistical estimators, cover measurement techniques, and improved methods for mapping forest and nonforest attributes.This report documents the first of these procedures, the photo-sampling protocol for the NPIP project.
Modelling techniques used in binary classification problems often result in a predicted probability surface, which is then translated into a presence–absence classification map. However, this translation requires a (possibly subjective) choice of threshold above which the variable of interest is predicted to be present. The selection of this threshold value can have dramatic effects on model accuracy as well as the predicted prevalence for the variable (the overall proportion of locations where the variable is predicted to be present). The traditional default is to simply use a threshold of 0.5 as the cut-off, but this does not necessarily preserve the observed prevalence or result in the highest prediction accuracy, especially for data sets with very high or very low observed prevalence. Alternatively, the thresholds can be chosen to optimize map accuracy, as judged by various criteria. Here we examine the effect of 11 of these potential criteria on predicted prevalence, prediction accuracy, and the resulting map output. Comparisons are made using output from presence–absence models developed for 13 tree species in the northern mountains of Utah. We found that species with poor model quality or low prevalence were most sensitive to the choice of threshold. For these species, a 0.5 cut-off was unreliable, sometimes resulting in substantially lower kappa and underestimated prevalence, with possible detrimental effects on a management decision. If a management objective requires a map to portray unbiased estimates of species prevalence, then the best results were obtained from thresholds deliberately chosen so that the predicted prevalence equaled the observed prevalence, followed closely by thresholds chosen to maximize kappa. These were also the two criteria with the highest mean kappa from our independent test data. For particular management applications the special cases of user specified required accuracy may be most appropriate. Ultimately, maps will typically have multiple and somewhat conflicting management applications. Therefore, providing users with a continuous probability surface may be the most versatile and powerful method, allowing threshold choice to be matched with each maps intended use.
The PresenceAbsence package for R provides a set of functions useful when evaluating the results of presence-absence analysis, for example, models of species distribution or the analysis of diagnostic tests. The package provides a toolkit for selecting the optimal threshold for translating a probability surface into presence-absence maps specifically tailored to their intended use. The package includes functions for calculating threshold dependent measures such as confusion matrices, percent correctly classified (PCC), sensitivity, specificity, and Kappa, and produces plots of each measure as the threshold is varied. It also includes functions to plot the Receiver Operator Characteristic (ROC) curve and calculates the associated area under the curve (AUC), a threshold independent measure of model quality. Finally, the package computes optimal thresholds by multiple criteria, and plots these optimized thresholds on the graphs.
Andrew Lister合作论文数Department of Physical Sciences and Architecture, University of Queensland1