Timely and accurate crop-type mapping is fundamental for sustainable agricultural management and food security in semi-arid regions, where climate variability and fragmented landscapes present persistent challenges. This study develops and validates an operational, multi-temporal framework for classifying five key agricultural classes (i.e., soft wheat, durum wheat, barley, trees, and other crops) across diverse Moroccan agroecosystems. By integrating monthly Sentinel-1 Synthetic Aperture Radar and Sentinel-2 optical time series spanning six growing seasons (2018–2025), we extracted 156 features comprising 13 spectral indices across 12 monthly composites. Ground truth data from the national Al Moutmir database, strategically balanced to address natural class imbalances, supported comprehensive training and validation of six machine learning models (i.e., Random Forest, Extra Trees, XGBoost, LightGBM, Voting Ensemble, and Stacking Ensemble). Our findings showed that the LightGBM and Stacking Ensemble achieved the highest performance with 88.04% overall accuracy, followed closely by XGBoost (87.93%). Feature importance analysis revealed that monthly temporal resolution significantly outperformed traditional phenological-stage approaches, with March and April indices (particularly Normalized Difference Vegetation Index and Normalized Difference Red Edge) contributing most to class discrimination. Notably, early-season radar features (Vertical-Vertical polarization in September) provided valuable complementary information when optical data were limited. The framework demonstrated robust generalization through 10-fold cross-validation while explicitly quantifying a 12.56% overfitting gap (train-CV difference), acknowledging a non-negligible overfitting risk. Offering transparent performance assessment. Error analysis identified persistent confusion between spectrally similar cereals, particularly durum and soft wheat, highlighting priority areas for future sensor integration. This scalable, cloud-based pipeline directly supports Morocco’s Green Generation strategy by providing a reproducible, high-accuracy solution for annual crop inventories, with transferable applications across similar Mediterranean and semi-arid agricultural systems.
Operational crop mapping requires classifiers capable of robust generalization across years. While feature importance is routinely used for model optimization, its temporal stability has rarely been systematically investigated, creating a critical gap in deploying reliable monitoring systems. This study moves beyond identifying “most important” features to systematically evaluate and quantify their inter-annual stability for enabling automated classification. Using six agricultural years (2018, 2019, 2020, 2023, 2024 and 2025) of Sentinel-1 and Sentinel-2 data over Morocco, we extracted 156 multi-sensor features across 12 monthly composites and analyzed their importance stability through statistical metrics, clustering, and novel composite indices: the Reliability Index (RI) and Automatic Selection Score (AuSS). This framework automates feature selection by ranking features with RI and AuSS and then applying Pareto optimization to identify a minimal stable feature set—without requiring annual retraining or expert intervention. Our analysis confirms a fundamental tension: the most discriminative features (e.g., NDVI, VH, VV) are also the most volatile, while stable features (e.g., NDRE, MSI, NDMI) offer modest predictive power. Hierarchical clustering revealed four behavioral typologies (Dominant Stable, Performant Volatile, Stable Minor, and Noise), guiding strategic feature management. Crucially, a Pareto analysis demonstrated that a refined portfolio of 6 indices (VH, VV, NDVI, NDRE, GCVI, RVI) captures 57.2% of cumulative predictive importance, filtering out inter-annual noise while preserving discriminative signal. The Voting Ensemble leveraging this Stable Portfolio maintained consistent high accuracy (87.4% accuracy, 87.2% F1-score) with minimal performance degradation during temporal transfer, while models based on volatile top features exhibited significant drops. Entropy analysis confirmed that all features in the Stable Portfolio provide consistent informational certainty, indicating that stability-driven selection does not increase model uncertainty. We conclude that feature stability is not merely a diagnostic metric but a foundational criterion for operational design. We propose a practical, metrics-driven framework for constructing automated crop classification systems that are more resilient to inter-annual climate variability.
Predicting continuous soil properties from limited field observations remains a central challenge in precision agriculture, particularly when models must operate across multiple spatial scales. This study develops a multi-scale framework that combines multi-temporal Sentinel-2 imagery with legacy soil maps to estimate spatial variation in soil organic matter (SOM). Bare-soil pixels were extracted and spectral indices calculated after cloud and snow masking, and a LASSO-based procedure was used to select informative images before modelling. Two strategies were evaluated: one using only satellite-derived indices, and another integrating soil texture information. At the farm scale, twelve Sentinel-2 images yielded 422 field-image records. A simple field-level averaging baseline achieved RMSE = 0.14 log(%SOM) and R & sup2; = 0.74, while the hybrid model predicting at the zone level achieved RMSE = 0.16 log(%SOM) and R & sup2; = 0.83, capturing within-field variability despite slightly higher point-wise error. Remote-sensing-only models performed poorly (RMSE approximate to 0.35-0.37 log(%SOM) and R & sup2; < 0.10), demonstrating that spectral indices alone cannot represent subsurface conditions. The framework was then scaled to the province of Qu & eacute;bec. Multi-year images, soil texture, topography, and climate variables were combined, and Random Forest, LightGBM, and CatBoost were tested after feature screening. At the province scale, predictive performance decreased (R & sup2; = 0.287-0.364), reflecting the increased agroclimatic, edaphic, and management heterogeneity across Qu & eacute;bec. However, the corresponding RMSE values (approximate to 0.101-0.103 log(%SOM)) indicate that prediction errors remain quantitatively moderate after back-transformation. Therefore, model performance at this scale should be interpreted not only in terms of explained variance, but also in terms of operational prediction error and regional differentiation capability. At the regional level, the framework supports decision-oriented applications that rely on relative field differentiation rather than precise point estimation.
The study aims to improve the interpretability of habitat suitability models for beekeeping, which are important for ecosystem services and Sustainable Development Goals. The research proposes a framework using artificial intelligence, specifically expert systems, to create more easily interpretable models. Three models were developed: a basic flat-fuzzy inference system (FIS), an optimized version, and a simpli-FIS model. These models used spatio-temporal data considering weather, topography, proximity features, and landscape quality. The final Simpli-FIS model achieves comparable predictive performance and improves the Flat-FIS model by a 55% reduction in RMSE with less input variables and rules. The study highlights the benefits of combining expert knowledge with data-driven approaches to enhance habitat suitability models for pollinators.
Vegetated riparian buffers play a critical role in maintaining ecological health and water quality, yet efficient characterization and large-scale monitoring remain challenging due to the resource demands of traditional field campaigns. This study, conducted in an agricultural setting, introduces a straightforward, image-based methodology for riparian buffer characterization, exploiting advancements in deep convolutional neural networks (DCNN) and very high spatial resolution satellite imagery. Leveraging a large Riparian Strip Quality Index (RSQI) field dataset, the proposed approach adapts a Multi-View DCNN (MVDCNN) architecture, originally developed for 3D object recognition, to correlate satellite images of riparian strips with RSQI metrics. Of the seven spectral band combinations, multiple input views, and two training modes evaluated, the configuration using four views with RGB bands from a pretrained network achieved the best results. However, the alternative spectral band combinations produced similar levels of performance, suggesting that texture and shape information are key factors in the model’s effectiveness. Comparisons with a conventional workflow involving object-based land cover classification followed by RSQI calculation indicate that the trained MVDCNN achieves stronger correlations between imagery and RSQI scores (average RMSE = 7.35, R2 = 0.93 using RGB bands) compared to the object-based method (RMSE = 11.25, R2 = 0.87). To our knowledge, this is the first direct application of DCNNs to riparian buffer quality assessment. Requiring minimal preprocessing and no photo-interpretation expertise, the proposed approach leverages existing field data to facilitate more accessible, scalable, and adaptable riparian buffer monitoring, with potential for application in diverse environmental contexts.
The Bidirectional Reflectance Distribution Function (BRDF) describes surface reflectance anisotropy, introducing viewing-angle variability that affects time-series consistency and multi-sensor integration in agricultural remote sensing. Conventional BRDF correction methods rely either on semi empirical models requiring multi-day compositing windows or physically based models dependent on simulated datasets, both of which involve simplifying assumptions that limit their ability to represent real surface dynamics. In contrast, this study proposes a novel data-driven, Transformer-based deep learning framework for BRDF modeling and correction that operates at the parcel level and daily temporal scale. The approach uses real MODIS Terra (morning) and Aqua (afternoon) observations processed within the Google Earth Engine (GEE) platform, enabling efficient large scale data access and same-day multi-angular observations under different viewing geometries. The model takes as input sun–sensor geometry and auxiliary surface parameters (e.g., leaf area index and soil moisture) to predict directional reflectance across spectral bands. The model was trained and evaluated on 48,344 agricultural parcels across Quebec. Predictive performance, assessed on an independent dataset of 7,084 paired observations, is highest in the NIR (Band 2, R² = 0.758) and green band (Band 4, R² = 0.715), while it is moderate in the red (Band 1) and blue (Band 3) bands (R² ≈ 0.52), reflecting their lower dynamic range over croplands. Using the trained BRDF model, directional reflectance is normalized to a reference geometry to generate Nadir BRDF Adjusted Reflectance (NBAR). The effectiveness of this normalization was evaluated by comparing Terra–Aqua discrepancies before and after correction. Before normalization, Bands 2 and 4 exhibit greater divergence than Bands 1 and 3. After normalization, mean absolute differences decrease by 60–73% across bands, and distributional overlap increases—from 0.64 to 0.89 (Band 1), 0.25 to 0.88 (Band 2), 0.63 to 0.88 (Band 3), and 0.30 to 0.81 (Band 4). The results indicate that systematic bias and viewing-angle effects are largely removed after normalization. In contrast, NDVI shows limited improvement (14%), indicating BRDF effects primarily influence absolute reflectance rather than ratio-based indices. Overall, the proposed framework improves multi-sensor consistency for large-scale agricultural monitoring.
A robust decision-support infrastructure is fundamental to modern nitrogen management, where recommendations must be both site-specific and economically resilient under uncertain environmental and market conditions. This study presents a modular, pre-computed decision support system that delivers near real-time nitrogen rate recommendations by combining agronomic similarity, probabilistic yield modelling, and scenario-based economic optimization. Instead of fitting a single model on demand for each user query, the system pre-computes millions of combinations of soil, weather, and management conditions, storing yield response, error distributions, and expected profit metrics in a relational database. At runtime, user inputs are converted into binned feature keys, and the closest precomputed records are retrieved, allowing nitrogen rate profitability to be evaluated in seconds without additional regression or simulation. To illustrate the proposed concept, a hybrid soil module links legacy soil maps, soil organic matter (SOM) estimates, weather factors and similarity search to historical field trials, enabling the reuse of existing agronomic information at scale. The architecture is explicitly modular: each component, similarity search, yield response modelling, uncertainty treatment, and profit calculation, can be updated or replaced without altering the rest of the system. This design supports the integration of richer data sources (e.g., improved SOM models, climate scenarios, or alternative crops) and the extension of new analytical modules beyond nitrogen. The system was evaluated across nitrogen rate scenarios ranging from 0 to 250 kg N ha-1, achieving response times of under 30 s, thus enabling timely evaluation of management decisions under uncertainty. Overall, the proposed framework provides a practical path toward scalable, real-time, and uncertainty-aware decision support for precision agriculture.
This paper presents a methodology for predicting poverty using semi-supervised learning techniques, specifically pseudo-labeling, and deep learning algorithms. Standard poverty prediction models rely on limited household survey data, whereas our approach exploits large amounts of unlabeled census data to improve prediction accuracy. By applying pseudo-labeling, we improve key performance metrics across various African regions, where our models outperform conventional approaches to identifying poor individuals. Deep neural networks (DNNs) trained on pseudo-labeled data exhibited area under the curve (AUC) scores ranging from 0.8 to over 0.9, a notable improvement over previous machine learning survey-based methods. Furthermore, random undersampling was key to refining model performance, balancing higher coverage with some reduction in precision. These findings have significant implications for poverty targeting, enabling more accurate identification of poor individuals and supporting better resource allocation.
Accurate estimation of the spatial distribution of soil properties is essential for advancing precision agriculture. This study evaluates two modeling strategies for predicting a continuous soil attribute, log-transformed soil organic matter (SOM%), using zone-level categorical predictors such as soil type. The first approach employs a zonal regression model incorporating soil-type coefficients, while the second leverages normalized pairwise comparisons among intra-field zones. Both models are assessed against a baseline strategy reflecting composite sampling practice, in which a single field-wise average is assumed. Incorporating categorical structure increased the prediction accuracy; the zonal regression model achieved an RMSE of 0.097 (R2 = 0.89), and the normalized pairwise model reached an RMSE of 0.108 (R2 = 0.85), both improving upon the baseline field averaging method based on within-field samples (RMSE = 0.140, R2 = 0.75). The proposed framework is generalizable to other soil attributes (e.g., pH, CEC, texture fractions) where zone-level delineation is available. This work offers a scalable, interpretable, and field-deployable methodology, contributing to spatial soil inference and site-specific agricultural decision-making under sparse sampling conditions.
The study employs a predictive modelling approach using a fuzzy inference system to assess the beekeeping potential of a geographic area. Specifically, an adaptive neuro-fuzzy inference system with subtractive clustering (ANFIS-SC) was utilized, incorporating six input variables that influence Apis mellifera health and productivity, and field data as the output variable reflecting the state of a colony. The results demonstrate the model’s effectiveness in predicting the suitability of areas for beekeeping. Sensitivity analysis highlighted the significant effects of relative humidity on the model’s output. The research underscores the importance of data quality, particularly in determining the local land cover quality index (LLCQI), on the outcomes. This study highlights the role of data science in enhancing precision in beekeeping and proposes its integration into management practices to support honey bee health.
It is becoming increasingly accepted that beekeeping is declining due to the damaging effect of global changes such as climate and land-use change that directly and indirectly impact Apis Melliferas. Despite numerous investigations, a comprehensive study that incorporates both global and local knowledge has yet to be conducted. For a long time, researchers have suggested that expert knowledge should be taken into account when creating decision support tools for managing activities related to natural resources, such as beekeeping. Unlike previous studies, this research seeks to tackle these questions while also introducing the concept of ecosystem service in modelling, offering a fresh perspective on sustainable land use. To achieve this goal, we combined several methods, including using literature knowledge, beekeeper knowledge, and multi-source geospatial data. These data are employed in a hierarchical fuzzy inference system in a unified way. The proposed approach was applied in the Québec region and the suggested technique appears to be both reliable and effective. The validation step revealed that the landscape variable, particularly the area used for agriculture or grassland, had the greatest impact on changes in hive weight throughout the season. In addition, we demonstrated that meteorological factors such as rainfall and relative humidity are strongly correlated to beekeeping. We showed that access to data and knowledge can be a critical factor in decision-making in the beekeeping industry, and thus we suggest that wild-bees conservationists, decision-makers, farmers, beekeepers, and other stakeholders adopt a collaborative approach.
Accurate and up-to-date land cover (LC) maps are critical for informed decision-making and resource management. While global products may not include region-specific classes or have optimal accuracy, this paper presents a method to produce a 10 m resolution LC map for Southern Quebec, Canada with 10 classes related to soil artificialization. To overcome annotation effort, a DeepResUNet model is trained on 400 simulated Sentinel-2 scene mosaics instead of real images. Scene realism is enhanced through the use of data augmentation and linear mixing of pure samples. Preliminary results demonstrate improved spatial detail compared to other products, with better delineation of roads and urban areas. A preliminary quantitative analysis based on a visual interpretation of 400 locations across 4 areas indicates that our land cover map aligns with ground conditions three times more often than a recent global land cover product, where they disagree.
Digital twins are increasingly gaining popularity as a method for simulating intricate natural and urban environments, with the precise segmentation of 3D objects playing an important role. This study focuses on developing a methodology for extracting buildings from textured 3D meshes, employing the PicassoNet-II semantic segmentation architecture. Additionally, we integrate Markov field-based contextual analysis for post-segmentation assessment and cluster analysis algorithms for building instantiation. Training a model to adapt to diverse datasets necessitates a substantial volume of annotated data, encompassing both real data from Quebec City, Canada, and simulated data from Evermotion and Unreal Engine. The experimental results indicate that incorporating simulated data improves segmentation accuracy, especially for under-represented features, and the DBSCAN algorithm proves effective in extracting isolated buildings. We further show that the model is highly sensible for the method of creating 3D meshes.
Satellite observations provide critical data for a myriad of applications, but automated information extraction from such vast datasets remains challenging. While artificial intelligence (AI), particularly deep learning methods, offers promising solutions for land cover classification, it often requires massive amounts of accurate, error-free annotations. This paper introduces a novel approach to generate a segmentation task dataset with minimal human intervention, thus significantly reducing annotation time and potential human errors. ‘Samples’ extracted from actual imagery were utilized to construct synthetic composite images, representing 10 segmentation classes. A DeepResUNet was solely trained on this synthesized dataset, eliminating the need for further fine-tuning. Preliminary findings demonstrate impressive generalization abilities on real data across various regions of Quebec. We endeavored to conduct a quantitative assessment without reliance on manually annotated data, and the results appear to be comparable, if not superior, to models trained on genuine datasets.
Digital twins are becoming increasingly popular in society for performing simulations. However, to conduct simulations, it is necessary to extract information about the objects composing an urban environment. Recently, a new semantic segmentation model applied to textured meshes, named PicassoNet-II, has been developed. The architecture of this model will be modified to perform segmentation of building instances rather than semantic segmentation. Additionally, a contextual analysis based on Markov fields is integrated into the algorithm to perform a contextual analysis of the features following segmentation. To train a 3D city segmentation model that can be generalized to any dataset, a large amount of annotated data is required. The model is trained using real data from Quebec City, Canada, as well as simulated data from different platforms such as Unreal Engine and Evermotion. Experimental results on semantic segmentation demonstrate that both simulated data and a Markov based analysis improves segmentation results overall.
Knowledge of demographic data is valuable information for planning initiatives. Typically, census, survey, and population projection exercises provide this information. In some developing countries, these operations pose a variety of economic and logistical challenges, thereby depriving authorities of accurate and timely information on their populations. To provide approaches for solving this situation, our study evaluates a population estimation method that is based on detection of residential geo-objects (houses) on very-high-resolution (VHR) satellite images using convolutional neural networks (CNN). The approach would be applicable to countries where a complete census is difficult to perform due to resource constraints or political instability. A 2008 VHR satellite image of Sudan is annotated according to seven classes of buildings to create a dataset that was used to train an object detection model, faster region-based CNN, by transfer learning. The model obtained mean average precision of 79% and 99% during training and validation, respectively. This unusual difference is due to the dominance of well detected classes in the validation dataset. The model was fine-tuned to detect the same building classes on images in 2021. A link between residential geo-objects and population size was established using 2008 population data and available field data. Subsequent characterization of the current population should assist in preparation of the 2023 census. Limitations of this approach were raised, but it could be used to improve the framework for population data collection in developing countries.
The Bidirectional Reflectance Distribution Function (BRDF) defines the anisotropy of surface reflectance and plays a fundamental role in many remote sensing applications. This study proposes a new machine learning-based model for characterizing the BRDF. The model integrates the capability of Radiative Transfer Models (RTMs) to generate simulated remote sensing data with the power of deep neural networks to emulate, learn and approximate the complex pattern of physical RTMs for BRDF modeling. To implement this idea, we used a one-dimensional convolutional neural network (1D-CNN) trained with a dataset simulated using two widely used RTMs: PROSAIL and 6S. The proposed 1D-CNN consists of convolutional, max poling, and dropout layers that collaborate to establish a more efficient relationship between the input and output variables from the coupled PROSAIL and 6S yielding a robust, fast, and accurate BRDF model. We evaluated the proposed approach performance using a collection of an independent testing dataset. The results indicated that the proposed framework for BRDF modeling performed well at four simulated Sentinel-3 OLCI bands, including Oa04 (blue), Oa06 (green), Oa08 (red), and Oa17 (NIR), with a mean correlation coefficient of around 0.97, and RMSE around 0.003 and an average relative percentage error of under 4%. Furthermore, to assess the performance of the developed network in the real domain, a collection of multi-temporals OLCI real data was used. The results indicated that the proposed framework has a good performance in the real domain with a coefficient correlation (R2), 0.88, 0.76, 0.7527, and 0.7560 respectively for the blue, green, red, and NIR bands.