National Forest Inventories (NFIs) provide valuable land cover (LC) information but often lack spatial continuity and an adequate update frequency. Satellite-based remote sensing offers a viable alternative, employing machine learning to extract thematic data. State-of-the-art methods such as convolutional neural networks rely on fully pixel-level annotated images, which are difficult to obtain. Although reference LC datasets have been widely used to derive annotations, NFIs consist of point-based data, providing only sparse annotations. Weakly supervised and self-supervised learning approaches help address this issue by reducing dependence on fully annotated images and leveraging unlabeled data. However, their potential for large-scale LC mapping needs further investigation. This study explored the use of NFI data with deep learning and weakly supervised and self-supervised methods. Using Sentinel-2 images and the Portuguese NFI, which covers other LC types beyond forest, as sparse labels, we performed weakly supervised semantic segmentation with a convolutional neural network to create an updated and spatially continuous national LC map. Additionally, we investigated the potential of self-supervised learning by pretraining a masked autoencoder on 65,000 Sentinel-2 image chips and then fine-tuning the model with NFI-derived sparse labels. The weakly supervised baseline achieved a validation accuracy of 69.60%, surpassing Random Forest (67.90%). The self-supervised model achieved 71.29%, performing on par with the baseline using half the training data. The results demonstrated that integrating both learning approaches enabled successful countrywide LC mapping with limited training data.
The current land cover (LC) mapping paradigm relies on automatic satellite imagery classification, predominantly through supervised methods, which depend on training data to calibrate classification algorithms. Hence, training data have a critical influence on classification accuracy. Although research on specific aspects of training data in the LC classification context exists, a study that organizes and synthetizes the multiplicity of aspects and findings of these researches is needed. In this article, we review the training data used for LC classification of satellite imagery. A protocol of identification and selection of relevant documents was followed, resulting in 114 peer-reviewed studies included. Main research topics were identified and documents were characterized according to their contribution to each topic, which allowed uncovering subtopics and categories and synthetizing the main findings regarding different aspects of the training dataset. The analysis found four research topics, namely construction of the training dataset, sample quality, sampling design and advanced learning techniques. Subtopics included sample collection method, sample cleaning procedures, sample size, sampling method, class balance and distribution, among others. A summary of the main findings and approaches provided an overview of the research in this area, which may serve as a starting point for new LC mapping initiatives.
Recent advances in satellite data availability, computing storage and processing power introduced a new land cover monitoring paradigm, settled on a continuous and timely identification of changes. The Continuous Change Detection and Classification (CCDC) algorithm has emerged as a powerful tool for continuous monitoring, being noteworthy for its ability to process high temporal frequency satellite data with components of seasonality, trend and break. Studies using CCDC were mostly limited to Landsat data, which offer lower spatial and temporal resolution in comparison to Sentinel-2 data. Therefore, our study aims to explore the potential of CCDC with Sentinel-2 data. For that purpose, an extensive reference dataset was developed for change detection accuracy assessment, comprising 290 sites of 200 m radius in a disturbance prone region in Central Portugal, ensuring an adequate representation of areas of vegetation loss. We focused on two specific forest species from this region, eucalyptus and maritime pine. Change date was determined through interpretation of orthophotos and satellite time series. We explored determinant aspects to CCDC performance, namely cloud and cloud shadow masking, algorithm parameterization, use of distinct vegetation indices and detection timeliness. Optimal accuracy was achieved with s2cloudless masking, lambda of 200, chi-square of 0.999, minYears of 1 and the Normalized Difference Vegetation Index. We computed the time lag vs omission error curve, showing comparable results (omission error rate close to 20 % was obtained with a time lag from 30 to 40 days) to methods designed to achieve near-real-time detection. Detections were spatially coherent, with patches of vegetation loss detected only with minor errors, mostly located in polygon borders. Disturbances in the first months resulted in poor model fitting, which undermined detection performance in some cases. Overall, results demonstrated how CCDC and Sentinel-2 data can be used to successfully monitor vegetation loss in a timely manner, especially as the satellite's time series grows.
The current land cover (LC) mapping paradigm relies on automatic satellite imagery classification, predominantly through supervised methods, which depend on training data to calibrate classification algorithms. Hence, training data have a critical influence on classification accuracy. Although research on specific aspects of training data in the LC classification context exists, a study that organizes and synthetizes the multiplicity of aspects and findings of these researches is needed. In this article, we review the training data used for LC classification of satellite imagery. A protocol of identification and selection of relevant documents was followed, resulting in 114 peer-reviewed studies included. Main research topics were identified and documents were characterized according to their contribution to each topic, which allowed uncovering subtopics and categories and synthetizing the main findings regarding different aspects of the training dataset. The analysis found four research topics, namely construction of the training dataset, sample quality, sampling design and advanced learning techniques. Subtopics included sample collection method, sample cleaning procedures, sample size, sampling method, class balance and distribution, among others. A summary of the main findings and approaches provided an overview of the research in this area, which may serve as a starting point for new LC mapping initiatives.
Land use/land cover (LULC) change detection and classification in maps based on automated data processing are becoming increasingly sophisticated in Earth Observation (EO).There is a growing number of annual maps available, with diverse but related production structures consisting primarily of classification and postclassification phases, the latter of which deals with inaccuracies of the first.The methodology production of the "Carta de Ocupação do Solo conjuntural" (COSc), a thematic land cover map of continental Portugal produced by the Directorate-General for Territory (DGT) mostly based on Sentinel-2 images classification, includes a semi-automatic phase of correction that combines expert knowledge and ancillary data in if-thenelse rules validated by photointerpretation.Although this approach reduces misclassifications from an initial Random Forest (RF) prediction map, improving consistency between years and compliance with ecological succession, requires a lot of time-consuming semi-automatic procedures.This work evaluates the relevance of exploring an additional set of variables for automatic classification over disturbance-prone areas.A multitemporal dataset with 124 variables was analysed using data dimensionality reduction techniques, resulting in the identification of 35 major explanatory indicators, which were then used as inputs for RF classification with cross-validation.The estimated importance of the explanatory variables shows that composites of spectral bands, which are already included in the current COSc workflow, in conjunction with the inclusion of additional data namely, historical land cover information and change detection coefficients, from the Continuous Change Detection and Classification (CCDC) algorithm, are relevant for predicting land cover classes after disturbance.Since map updating is a more challenging task for disturbed pixels, we focused our analysis on locations where COSc indicated potential land cover change.Nonetheless, the overall classification accuracy for our experiments was 72.34 % which is similar to the accuracy of COSc for this region of Portugal.The findings suggest new variables that could improve future COSc maps.
The current land cover mapping paradigm relies on automatic classification of satellite images, with supervised methods being the most used, implying training data to have a crucial role. Aspects such as training sample size and quality should be carefully considered. This paper proposes assessing the use of a detailed class nomenclature to reinforce class diversity in the training sample. A Random Forest (RF) classification of Sentinel-2 multi-temporal data was conducted. Additionally, the effect of sample size and class distribution were evaluated. The results indicate that the use of a detailed nomenclature provided better results in terms of classification accuracy. With respect to sample distribution, adopting class sizes proportional to their occurrence in a reference land cover map exhibited superior performance in comparison to an equal size approach. The effect of sample size on classification performance was limited, as previous studies with RF suggested.
Portugal is building a land cover monitoring system to deliver land cover products annually for its mainland territory. This paper presents the methodology developed to produce a prototype relative to 2018 as the first land cover map of the future annual map series (COSsim). A total of thirteen land cover classes are represented, including the most important tree species in Portugal. The mapping approach developed includes two levels of spatial stratification based on landscape dynamics. Strata are analysed independently at the higher level, while nested sublevels can share data and procedures. Multiple stages of analysis are implemented in which subsequent stages improve the outputs of precedent stages. The goal is to adjust mapping to the local landscape and tackle specific problems or divide complex mapping tasks in several parts. Supervised classification of Sentinel-2 time series and post-classification analysis with expert knowledge were performed throughout four stages. The overall accuracy of the map is estimated at 81.3% (±2.1) at the 95% confidence level. Higher thematic accuracy was achieved in southern Portugal, and expert knowledge significantly improved the quality of the map.
Abstract. Supervised classification of remotely sensed images has been widely used to map land cover and land use. Since the performance of supervised methods depends on the quality of the training data, it is essential to develop methods to generate an enhanced training dataset. Active learning represents an alternative for such purpose as it proposes to create a dataset of optimized samples, normally collected based on classification uncertainty. However, it is heavily dependent on human interaction, since the user has to label selected samples over a number of iterations. In this paper, we explore the use of uncertainty to improve classification accuracy through a single iteration. We conducted experiments in a region of Portugal (Trás-os-Montes), using multi-temporal Sentinel-2 images. The proposed approach consisted in computing the classification uncertainty of a Random Forest to collect additional training data from areas of high uncertainty and perform a new classification. An accuracy assessment was performed to compare the overall accuracy of the initial and new classifications. The results exhibited an increase in accuracy, though considered not statistically significant. Obstacles related to labelling additional sampling units resulted in a lack of additional training data for various classes, which might have limited the accuracy improvement. Additionally, an uneven proportion of additional training sampling units per class and the collection of new sample data from a limited number of uncertainty regions might also have prevented a higher increase in accuracy. Nevertheless, visual inspection of the maps revealed that the new classification reduced the confusion between some classes.
Experiments were carried out to investigate the use of Land Use and Coverage Area frame Survey (LUCAS) dataset and Sentinel-2 imagery to produce a land cover map in Portugal through automated supervised classification. LUCAS is a free land cover land use (LCLU) dataset based in Europe, while Sentinel-2 satellites provide also free images with short revisit frequency. The goal was to evaluate if LUCAS dataset from 2018 can be used as a single reference dataset for land cover classification at national level. The Random Forest (RF) algorithm was used. Some processing steps were undertaken to use LUCAS as reference dataset. The original LUCAS LCLU nomenclature was modified into a new nomenclature composed of 12 and 6 level-2 and level-1 map classes, respectively. Filtering was performed on LUCAS metadata, reducing the initial number of LUCAS points over Portugal from 7168 to 4910. Monthly composites of Sentinel-2 images acquired between October 2017 and September 2018 were used. To reduce the imbalance in LUCAS training points, an oversampling technique based on Synthetic Minority Over-Sampling Technique (SMOTE) was used. An independent validation dataset was produced with 600 points. RF shows an overall accuracy (OA) of 57% for level-2 and 72% for level-1 nomenclatures. When using the oversampling technique, the OA accuracy increases by 3% for level2 and 2% for level-1. The preliminary results of this experiment show that LUCAS dataset used in supervised machine learning classification has potential to produce a reliable land cover map at national scale.
Classification accuracy of remote sensing images with supervised learning depends on the quality and characteristics of training samples. Size is a key aspect of a sample and its impact on classification depends on several factors, including the classifier employed, dimension on the feature space and land cover characteristics. Random Forest classifier is considered to be of low sensitivity to variations in sample size. However, further investigation is required when feature spaces are large and training is performed with spectral subclasses of the land cover classes to be mapped. This paper proposes to assess the impact of sample size in the classification accuracy of Random Forest using multitemporal Sentinel-2 data and a detailed set of training subclasses to produce a map with general land cover classes. The results revealed similar classification accuracies after major reductions in sample size.
Southern Portugal is characterized by disperse tree cover of Cork and Holm oaks in an agro-forestry system known as montado. Mapping these trees has been historically very difficult as they occur in isolation or in groups with different understory vegetation, including grass and shrubland. Automatic classification for binary tree/non-tree map production has been used elsewhere, but with limited success in the context of montado. Here, the potential of Sentinel-2 data was explored to map oaks using pure and mixed pixels to train a random forest. The output depicts a gradient of tree cover that can be transformed into a crisp map. The accuracy assessment of the latter shows commission and omission errors of 17% and 18%.
This paper presents an experimental crop classification of the 10 most abundant annual crop types in Portugal, using a study area located in Alentejo region. This region has great diversity of land uses as well as multiple crop types. Sentinel-2 2018 intra-annual time-series imagery is considered in the experiment. The Portuguese Land Parcel Identification System (LPIS) is used to extract automatic training samples. LPIS information is automatically processed with the help of auxiliary datasets to filter out crop areas more likely to have been mislabeled. Classification is obtained using random forest. Validation is performed using an independent dataset also based on LPIS. A global accuracy of 76% is obtained. The novelty of the methodology here presented shows that LPIS can be used together with auxiliary data for crop type mapping, helping to characterize the agriculture land diversity in Portugal.
The performance of supervised classification depends on the size and quality of the training data. Multiple studies have used reference datasets to extract training data automatically in an efficient way. However, automatic extraction might be inappropriate for some classes. Furthermore, classes can have distinct spectral characteristics across large areas. Thus, dividing the study area into subregions can be beneficial. This study proposes to assess the impact of the introduction of spatial stratification and manually collected training data on classification performance. Two classifications were conducted with the Random Forest classifier and multi-temporal Sentinel-2 data. The classifications’ performance was evaluated by accuracy metrics and visual inspection of the maps. The results indicate that introducing spatial stratification and manual training yielded a higher overall accuracy (66.7%) when compared to the accuracy of a benchmark classification (60.2%) conducted without stratification and with training data collected exclusively by automatic methods. Visual inspection of the maps also revealed some advantages of the novel approach, namely constraining some land cover classes to be present only within specific strata, which avoids commission errors of the class to spread freely across the map. Most of the classification improvements were observed in subregions with specific landscapes and spectral patterns, although these strata represent a small fraction of the study area, which might have contributed to the small increase in accuracy.
Moraes, D., Ribeiro, S., & Costa, A. C. (2019). Modelling air temperature in Brazilian northeast to evaluate change patterns from 2000 to 2017. In 19th International Multidisciplinary Scientific Geoconference, SGEM 2019: Conference Proceedings. (2.2 ed., Vol. 19, pp. 915-922). (International Multidisciplinary Scientific GeoConference Surveying Geology and Mining Ecology Management, SGEM). https://doi.org/10.5593/sgem2019/2.2/S11.113