The fate of live forest biomass is largely controlled by growth and disturbance processes, both natural and anthropogenic. Thus, biomass monitoring strategies must characterize both the biomass of the forests at a given point in time and the dynamic processes that change it. Here, we describe and test an empirical monitoring system designed to meet those needs. Our system uses a mix of field data, statistical modeling, remotely-sensed time-series imagery, and small-footprint lidar data to build and evaluate maps of forest biomass. It ascribes biomass change to specific change agents, and attempts to capture the impact of uncertainty in methodology. We find that: A common image framework for biomass estimation and for change detection allows for consistent comparison of both state and change processes controlling biomass dynamics. Regional estimates of total biomass agree well with those from plot data alone. The system tracks biomass densities up to 450-500 Mg ha(-1) with little bias, but begins underestimating true biomass as densities increase further. Scale considerations are important. Estimates at the 30 m grain size are noisy, but agreement at broad scales is good. Further investigation to determine the appropriate scales is underway. Uncertainty from methodological choices is evident, but much smaller than uncertainty based on choice of allometric equation used to estimate biomass from tree data. In this forest-dominated study area, growth and loss processes largely balance in most years, with loss processes dominated by human removal through harvest. In years with substantial fire activity, however, overall biomass loss greatly outpaces growth. Taken together, our methods represent a unique combination of elements foundational to an operational landscape-scale forest biomass monitoring program.
Comparison of predictions from the Gradient Nearest Neighbor structure model with ground observations for several measures of forest composition and structure.
Accuracy assessments of remote sensing products are necessary for identifying map strengths and weaknesses in scientific and management applications. However, not all accuracy assessments are created equal. Motivated by a recent study published in Forest Ecology and Management (Volume 342, pages 8-20), we explored the potential limitations of accuracy assessments related to characteristics of the field data: sampling bias and spatial resolution. The authors of the previous paper used data from variable radius plots near northern spotted owl nest sites to assess the predictive accuracy of gradient nearest neighbor (GNN) maps in portions of Oregon and Washington, USA. The field plots used for accuracy assessment (1) potentially biased the accuracy assessment toward older forests and (2) examined accuracy at finer scales than the imputation map predictions under consideration. To examine both the impacts of bias and scale in accuracy assessment, we assessed the predictive accuracy of GNN maps in western and southern Oregon. We found correlation coefficients between predicted (900 m(2)) and observed forest attributes for small plots (506 m(2)) were consistently lower than accuracy assessments using larger plots (4048 m(2)). Similarly, correlation coefficients based only on field plots near nest sites were lower than correlations based on all field plots. These results imply that sampling bias and small plot areas result in accuracy assessments that underestimate map predictive performance. In particular, assessing accuracy at spatial scales below the resolution of the map products are overly pessimistic (i.e., low correlation coefficients). While accuracy assessment is important, care needs to be taken to ensure that the sampling design for field data does not limit inference on map accuracy. Published by Elsevier B.V.
Mapping vegetation and landscape change at fine spatial scales is needed to inform natural resource and conservation planning, but such maps are expensive and time-consuming to produce. For Landsat-based methodologies, mapping efforts are hampered by the daunting task of manipulating multivariate data for millions to billions of pixels. The advent of cloud-based geospatial computing platforms, such as the Google Earth Engine (GEE), enables a solution to big data problems by providing an environment for massively parallel processing of simple to complex algorithms. In addition to the obvious processing benefits, GEE supplies access to petabytes of remote sensing, topographic, and climatological data, including the entire Landsat archive. As a proof of concept, we will demonstrate the utility of GEE in vegetation change detection and mapping at both regional and national scales. We showcase two current projects utilizing GEE: 1) a random-forest based ensemble model incorporating information from leading change detection algorithms and 2) a nearest neighbors model combining forest inventory plots and spatial predictors to produce regional to national forest vegetation maps. Our early results suggest that this programming approach is ideal for rapid prototyping of change detection and forest vegetation modeling, including flexibility in specifying model forms and spatial covariates. We envision that this type of computing system could support many of FIA’s national data products.
Imputation provides a useful method for mapping forest attributes across broad geographic areas based on field plot measurements and Landsat multi-spectral data, but the resulting map products may be of limited use without corresponding analyses of uncertainties in predictions. In the case of k-nearest neighbor (kNN) imputation with k = 1, such as the Gradient Nearest Neighbor (GNN) approach, where the field plot with the most similar spectral signature is attributed to a given pixel, there has been limited guidance on methods of examining uncertainty. In this study, we use a bootstrapping method to assess the uncertainty associated with the imputation process on predictions of live tree structure (canopy cover, quadratic mean diameter, and aboveground biomass), dead tree structure (snag density and downed wood volume), and community composition (proportion hardwood) for a portion of the Cascade Mountains in Oregon, USA. We performed kNN with k = 1 imputation with 4000 bootstrap samples of the field plot data and examined three metrics of uncertainty: the width of 95% interpercentile ranges (IPR), the proportion of bootstrap samples with no tally (i.e., forest attribute was imputed as zero), and the imputation deviations (i.e., mean prediction from the bootstrap sample minus baseline GNN prediction [no bootstrapping]). Imputed values of dead tree components and species composition exhibited greater IPR, proportion no tally near 0.5, and greater magnitudes of imputation deviations compared to live tree components, indicating greater uncertainties. Our uncertainty metrics varied spatially with respect to environmental gradients and the variation was not consistent among metrics. Geographic patterns in prediction uncertainties implicated biogeography and disturbance as major factors influencing regional variation in imputation uncertainty. Spatial patterns differed not only by forest attribute, but by uncertainty metric, indicating that no single measure of uncertainty or forest structure provides a full description of imputation performance. Users of imputed map products need to consider the pattern of and the processes that contribute to uncertainty during the early stages of project development and execution. Published by Elsevier B.V.
Aim Landscape management and conservation planning require maps of vegetation composition and structure over large regions. Species distribution models (SDMs) are often used for individual species, but projects mapping multiple species are rarer. We compare maps of plant community composition assembled by stacking results from many SDMs with multivariate maps constructed using nearest-neighbor imputation. Location Western Cascades ecoregion, Oregon and California, USA. Methods We mapped distributions and abundances of 28 tree species over 4,007,110 ha at 30-m resolution using three approaches: SDMs using machine learning (random forest) to yield: (1) binary (RF_Bin); (2) basal area (abundance; RF_Abund) predictions; and (3) multi-species basal area predictions using a nearest-neighbor imputation variant based on random forest (RF_NN). We evaluated accuracy of binary predictions for all models, compared area mapped with plot-based areal estimates, assessed species abundance at two spatial scales and evaluated communities for species richness, problematic compositional errors and overall community composition. Results RF_Bin yielded the strongest binary predictions (median True Skill Statistics; RF_Bin: 0.57, RF_NN: 0.38, RF_Abund: 0.27). Plot-scale predictions of abundance were poor for RF_Abund and RF_NN (median Agreement Coefficient (AC): −1.77 and −2.28), but strong when summarized over 50-km radius tessellated hexagons (median AC for both: 0.79). RF_Abund's strength with abundance and weakness with binary predictions stems from predicting small values instead of zeros. The number of zero value predictions from RF_NN was closest to counts of zeros in the plot data. Correspondingly, RF_NN's map-based species area estimates closely matched plot-based area estimates. RF_NN also performed best for community-level accuracy metrics. Conclusions RF_NN was the best technique for building a broad-scale map of diversity and composition because the modelling framework maintained inter-species relationships from the input plot data. Re-assembling communities from single variable maps often yielded unrealistic communities. Although RF_NN rarely excelled at single species predictions of presence or abundance, it was often adequate to many (but not all) applications in both dimensions. We discuss our results in the context of map utility for applications in the fields of ecology, conservation and natural resource management planning. We highlight how RF_NN is well-suited for mapping current but not future vegetation.
This study investigated how lidar-derived vegetation indices, disturbance history from Landsat time series (LTS) imagery, plot location accuracy, and plot size influenced accuracy of statistical spatial models (nearest-neighbor imputation maps) of forest vegetation composition and structure. Nearest-neighbor (NN) imputation maps were developed for 539,000ha in the central Oregon Cascades, USA. Mapped explanatory data included tasseled-cap indices and disturbance history metrics (year, magnitude, and duration of disturbance) from LTS imagery, lidar-derived vegetation metrics, climate, topography, and soil parent material. Vegetation data from USDA Forest Service forest inventory plots was summarized at two plot sizes (plot and subplot) and geographically located with two levels of accuracy (standard and improved). Maps of vegetation composition and structure were developed with the Gradient Nearest Neighbor (GNN) method of NN imputation using different combinations of explanatory variables, plot spatial resolution, and plot positional accuracy. Lidar vegetation indices greatly improved predictions of live tree structure, moderately improved predictions of snag density and down wood volume, but did not consistently improve species predictions. LTS disturbance metrics improved predictions of forest structure, but not to the degree of lidar indices, while also improving predictions of many species. Absence of disturbance attribution (i.e. disturbance type such as fire or timber harvest) in LTS disturbance metrics may have limited our ability to predict forest structure. Absence of corrected lidar intensity values may also have lowered accuracy of snag and species predictions. However, LTS disturbance attribution and lidar corrected intensity values may not be able to overcome fundamental limitations of remote sensing for predicting snags and down wood that are obscured by the forest canopy. Improved GPS plot locations had little influence on map accuracy, and we suggest under what conditions improved GPS plot locations may or may not improve the accuracy of predictive maps that link remote sensing with forest inventory plots. Subplot NN imputation maps had much lower accuracy compared to maps generated using response variables from larger whole plots. No single map had optimal results for every mapped variable, suggesting map users and developers need to prioritize what forest vegetation attributes are most important for any given map application.
The integration of satellite image data with forest inventory plot data is a popular approach for mapping forest vegetation over large regions. Several methodological choices regarding spatial scale, mostly related to spatial resolution or grain, can profoundly influence forest maps developed from plot and imagery data. Yet often the consequences of scaling choices are not explicitly addressed. Our objective was to quantify the effects of several scale-related methods on map accuracy for multiple forest attributes, using a variety of diagnostics that address different map characteristics, to help guide map developers and users. We conducted nearest-neighbor imputation over a large region in the Pacific Northwest, USA, to investigate effects of imputation grain (single pixel or kernel); inclusion of heterogeneous plots; accuracy assessment grain and extent; and value of k (k=1 and k=5), where k is the number of nearest-neighbor plots. Spatial predictors were from rasters describing climate and topography and a time-series of Landsat imagery. Reference data were from regional forest inventory plots measured over two decades. All analyses were conducted at a spatial resolution of 30m×30m. Effects of imputation grain and heterogeneous plots on map accuracy were small. Excluding heterogeneous plots slightly improved map accuracy and did not lessen the systematic agreement (ACSYS) between our maps and observed plot data. Accuracy assessment grain strongly influenced map accuracy: maps assessed with a multi-pixel block were much more accurate than when assessed with a single pixel for almost all map diagnostics, but this was an artifact of methods rather than reflecting real differences among maps. Unsystematic agreement (ACUNS) between our maps and plots, or random error, improved notably with increasing accuracy assessment extent for all scaling methods, indicating that reliability of most map applications can be improved through coarsening the map grain. Value of k strongly influenced map diagnostics. The k=5 maps were better than k=1 maps for local-scale accuracy, but at the cost of reduced ACSYS, and loss of variability and poor areal representation of forest conditions over the study region. The k=1 maps produced notably better predictions of the least abundant forest conditions (early-successional, late-successional, and broadleaf). None of the scaling methods were optimal for all map diagnostics. Nevertheless, given a variety of diagnostics associated with a range of scaling options, map developers and map users can make informed choices about methods and resulting maps that best meet their particular objectives, and we present some general guidelines in this regard. Most of our findings are applicable to mapping with Landsat data in other forested regions with similar forest inventory data, and to other methods for spatial prediction.
Imputation is commonly used to assign reference stand observations to target stands based on covariate relationships to remotely sensed data to assign inventory attributes across the entire landscape. However, most remotely sensed data are collected at higher resolution than the stand inventory data often used by operational foresters. Our primary goal was to compare various aggregation strategies for modeling and mapping forest attributes summarized from stand inventory data, using predictor variables derived from either light detection and ranging (LiDAR) or Landsat and a US Geological Survey (USGS) digital terrain model (DTM). We found that LiDAR metrics produced more accurate models than models using Landsat/USGS-DTM predictors. Calculating stand-level means of all predictors or all responses proved most accurate for developing imputation models or validating imputed maps, respectively. Developing models or validating maps at the unaggregated scale of individual stand subplots proved to be very inaccurate, presumably due to poor geolocation accuracies. However, using a sample of pixels within stands proved only slightly less accurate than using all available pixels. Furthermore, bootstrap tests of similarity between imputations and observations showed no evidence of bias regardless of aggregation strategy. We conclude that pixels sampled from within stands provide sufficient information for modeling or validating stand attributes of interest to foresters.
Natural resource policy analysis and conservation planning are best served by broad-scale information about vegetation that is detailed, spatially complete, and consistent across land ownerships and allocations. In this paper I describe how a new generation of forest vegetation maps can be used to assess the distribution of vegetation biodiversity among land ownerships and land allocations at a regional scale. The vegetation maps contain detailed tree- and stand-level attributes of vegetation composition and structure for each forest land pixel in a regional landscape, which can be translated into vegetation biodiversity indicators at individual tree, species, community, and landscape levels. The new forest vegetation maps can be combined with models of stand and landscape dynamics to assess potential effects of alternative forest policies on biodiversity in future landscapes. Lastly, I discuss the importance of considering all land ownerships and allocations, including the matrix of semi-natural, managed forests, in regional biodiversity assessments.
The Northwest Forest Plan (NWFP), which aims to conserve late-successional and old-growth forests (older forests) and associated species, established new policies on federal lands in the Pacific Northwest USA. As part of monitoring for the NWFP, we tested nearest-neighbor imputation for mapping change in older forest, defined by threshold values for forest attributes that vary with forest succession. We mapped forest conditions on >19 million ha of forest for the beginning (Time 1) and end (Time 2) of a 13-year period using gradient nearest neighbor (GNN) imputation. Reference data were basal area by species and size class from 17,000 forest inventory plots measured from 1993 to 2008. Spatial predictors were from Landsat time-series and GIS data on climate, topography, parent material, and location. The Landsat data were temporally normalized at the pixel level using LandTrendr algorithms, which minimized year-to-year spectral variability and provided seamless multi-scene mosaics. We mapped older forest change by spatially differencing the Time 1 and Time 2 GNN maps for average tree size (MNDBH) and for old-growth structure index (OGSI), a composite index of stand age, large live trees and snags, down wood, and diversity of tree sizes. Forests with higher values of MNDBH and OGSI occurred disproportionately on federal lands. Estimates of older forest area and change varied with definition. About 10% of forest at Time 2 had OGS1 >= 50, with a net loss of about 4% over the period. Considered spatially, gross gain and gross loss of older forest were much greater than net change. As definition threshold value increased, absolute area of mapped change decreased, but increased as a percentage of older forest at Time 1. Pixel-level change was noisy, but change summarized to larger spatial units compared reasonably to known changes. Geographic patterns of older forest loss coincided with areas mapped as disturbed by LandTrendr, including large wildfires on federal lands and timber harvests on nonfederal lands. The GNN distribution of older forest attributes closely represented the range of variation observed from a systematic plot sample. Validation using expert image interpretation of an independent plot sample in TimeSync corroborated forest changes from GNN. An advantage of imputed maps is their flexibility for post-classification, summary, and rescaling to address a range of objectives. Our methods for characterizing forest conditions and dynamics over large regions, and for describing the reliability of the information, should help inform the debate over conservation and management of older forest. Published by Elsevier B.V.
Question: How can nearest-neighbour (NN) imputation be used to develop maps of multiple species and plant communities?Location: Western and central Oregon, USA, but methods are applicable anywhere.Methods: We demonstrate NN imputation by mapping woody plant communities for > 100 000 km(2) of diverse forests and woodlands. Species abundances on similar to 25 000 plots were related to spatial predictors (rasters) describing climate, topography, soil and geographic location using constrained ordination (CCA). Species data from the nearest plot in multi-dimensional CCA space were imputed to each map pixel. Maps of multiple individual species and community types were constructed from the single imputed surface. We computed a variety of diagnostics to characterize different qualities of the imputed (mapped) community data.Results: Community composition gradients were strongly associated with climate and elevation, and less so with topography and soil. Accuracy of the imputation model for presence/absence of 150 species varied widely (kappa 0.00 to 0.80). Omission error rates were higher than commission rates due to low species prevalence, and areal representation of species was only slightly inflated. A map of 78 community types was 41% correct and 78% fuzzy correct. Errors of omission and commission were balanced, and areal representation of both rare and abundant communities was accurate. Map accuracy may be lower for some species than with other methods, but areal representation of species and communities across large landscapes is preserved. Because imputed vegetation surfaces are developed for all species simultaneously, map units contain suites of species known to co-occur in nature. Maps of individual species, and of community types derived from them, will be internally consistent at map locations.Conclusions: NN imputation is a useful modelling approach where maps of multiple species and plant communities are needed, such as in natural resource management and conservation planning or models that project landscape change under alternative disturbance or climate scenarios. More research is needed to evaluate other ordination methods for NN imputation of plant communities.
Spatially and temporally explicit knowledge of biomass dynamics at broad scales is critical to understanding how forest disturbance and regrowth processes influence carbon dynamics. We modeled live, aboveground tree biomass using Forest Inventory and Analysis (FIA) field data and applied the models to 20+ year time-series of Landsat satellite imagery to derive trajectories of aboveground forest biomass for study locations in Arizona and Minnesota. We compared three statistical techniques (Reduced Major Axis regression, Gradient Nearest Neighbor imputation, and Random Forests regression trees) for modeling biomass to better understand how the choice of model type affected predictions of biomass dynamics. Models from each technique were applied across the 20+ year Landsat time-series to derive biomass trajectories, to which a curve-fitting algorithm was applied to leverage the temporal information contained within the time-series itself and to minimize error associated with exogenous effects such as biomass measurements, phenology, sun angle, and other sources. The effect of curve-fitting was an improvement in predictions of biomass change when validated against observed biomass change from repeat FIA inventories. Maps of biomass dynamics were integrated with maps depicting the location and timing of forest disturbance and regrowth to assess the biomass consequences of these processes over large areas and long time frames. The application of these techniques to a large sample of Landsat scenes across North America will facilitate spatial and temporal estimation of biomass dynamics associated with forest disturbance and regrowth, and aid in national-level estimates of biomass change in support of the North American Carbon Program.