Earth observation foundation models (FMs) require vast amounts of high-quality training data, and the remote sensing community has responded by developing numerous benchmark datasets. However, these benchmarks are predominantly static and often lack rigorous quality assessments and difficulty measures. The convergence of petabyte-scale satellite archives and large FMs therefore necessitates rigorous validation frameworks for dataset quality and complexity. To address these limitations, we introduce a dual validation framework supported by a scalable, on-demand dataset curation pipeline. The framework integrates two complementary components. First, intrinsic validation characterizes dataset quality through radiometric fidelity checks and a composite dataset difficulty index (DI). This index synthesizes spatial heterogeneity, phenological variability, and data scarcity into a single normalized metric. Second, extrinsic validation utilizes the DI to enable stratified model performance analysis across distinct difficulty levels of the dataset, providing deeper insights than conventional aggregate evaluation. To operationalize this framework, the pipeline employs a cloud-native Zarr architecture with distributed computing to enable efficient parallel data access and automated quality supervision. We validate this framework on the cloud-gap imputation task, demonstrating that stratified analysis reveals severe performance degradation patterns completely obscured by aggregate metrics. Specifically, models evaluated on high-difficulty scenes exhibited over a 50% increase in error (RMSE) and a notable degradation in structural similarity, with SSIM falling from 0.83 to 0.77. This study provides a comprehensive foundation for transforming standard analysis-ready satellite data into rigorously validated, machine learning-ready benchmarks, advancing data-centric artificial intelligence in remote sensing through the quantitative assessment of dataset characteristics and model robustness.
In the field of early fire and smoke detection using unmanned aerial vehicle (UAV) remote sensing, existing research mainly focuses on single-target segmentation, and there are problems of neglecting the inherent correlation and coexistence between fire and smoke in real scenarios, as well as class imbalance. Furthermore, traditional convolutional neural networks (CNNs) suffer from fixed receptive fields, which make it difficult to balance the detection of large-scale smoke and small-scale flame plumes, as well as insufficient modelling of long-range dependencies, thus resulting in incomplete smoke segmentation. To address these challenges, we propose the Fire and Smoke Segmentation Network FSSNet incorporates an improved lightweight hybrid backbone, integrating ResNet and MobileNetV3, paired with a DeepLabHead-FCNHead decoder to enable high-resolution feature extraction and multi-scale segmentation. A weighted cross-entropy combined loss function is introduced to alleviate data imbalance and enhance performance on small targets. To support validation and address the scarcity of relevant datasets, we have constructed an annotated multi-target dataset encompassing fire, smoke, and background classes. Comprehensive experiments demonstrate that FSSNet outperforms baseline models, FCNHead, DeepLabHead, and LR-ASPPHead, across key metrics. FSSNet achieves a total recall of 84.8%, an MF1 score of 90.3%, and an MIoU of 75.4%, with fire and smoke recall rates exceeding 96.0% and F1 scores of at least 71.4%, indicating a balanced trade-off between precision and recall. This approach enhances the accuracy of collaborative fire and smoke segmentation in UAV remote sensing, and it offers a practical technical solution for early disaster warning. Its hybrid architecture further provides valuable insights for multi-target segmentation in UAV remote sensing. Furthermore, FSSNet maintains lightweight efficiency, thereby meeting the computational constraints of UAV edge devices.
Timely landslide detection and rapid qualitative assessment are fundamental to effective warning systems, hazard management, and risk mitigation. Yet, current practices that rely on on-site surveys and manual expert assessment remain risky, costly, and time-consuming. These limitations result in substantial delays between the event and the availability of actionable information. This study proposes a hybrid, multi-model framework that fuses RGB remote-sensing imagery with geospatial layers to enable timely landslide detection and actionable reporting. The pipeline couples an enhanced SegFormer (denoted as SDF-SegFormer-B2) model for landslide localization, a feature extraction technique for per-slide geo-attribute computation, and a lightweight instruction-tuned LLM (Mistral-7B-Instruct-v0.3) for structured, expert-style reporting. Although a few previous studies have explored landslide captioning, to our knowledge this is the first framework designed to generate structured technical reports enriched with terrain-context interpretation and qualitative intervention-priority indicators. Experiments use 26,758 georeferenced RGB tiles (64 × 64) with 3 m of spatial resolution from PlanetScope satellite imagery over Emilia–Romagna, Italy, with 68,592 annotated landslide boxes collected after the May 2023 rainfall events (~200 mm in 48 h on 1–3 May; 200–250 mm in 48 h on 16–17 May). The proposed SDF-SegFormer-B2 segmentation model achieved a precision of 85.54%, recall of 72.31%, and an F1-score of 78.39% on the unseen test dataset. To evaluate the quality of the generated landslide reports, 100 images were selected for domain-expert assessment. Among these, 58% of the reports were rated as “Very Good,” 30% as “Good,” 8% as “Acceptable,” and 4% as “Poor.” When considering only reports with complete and accurate inputs, 81.48% were rated “Very Good,” and 96.30% were rated either “Good” or “Very Good.” By integrating complementary models and modalities, the proposed approach automates localization-to-reporting and enables the generation of terrain-aware landslide summaries that may support preliminary decision-making and rapid post-disaster screening.
In this study, we demonstrate the central role of satellitederived geospatial datasets in mapping the suitability of bifacial photovoltaic (PV) systems in Algeria. The methodology integrates a suite of remote sensing inputs, including solar irradiation, sunshine duration, cloud climatology, surface albedo, and terrain parameters such as elevation, slope, and aspect. The resulting suitability maps reveal that approximately 343,000 km2 of Algeria's territory is suitable for bifacial PV deployment. Overall, about 14.4% of Algeria's national surface emerges as favorable for potential large-scale deployment. These findings highlight the significant contribution of remote sensing-driven spatial analysis in guiding future solar energy development and accelerating the clean energy transition.
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage.
The rapid rollout of industrial photovoltaic (PV) generation demands robust tools for benchmarking performance and guiding investment. Traditional efficiency techniques – non-parametric data envelopment analysis (DEA) and parametric stochastic-frontier analysis (SFA) – offer complementary strengths, yet each has recognized limitations when applied in isolation. This study proposes a hybrid framework that (i) estimates multiple DEA and SFA frontiers; (ii) learns their predictive patterns with Random Forest, Gradient Boosting and XGBoost models; and (iii) integrates the resulting efficiency scores through an optimized weighting scheme derived from a simplex-lattice mixture design of experiments. The approach is demonstrated on a statewide dataset covering 70 immediate geographic regions of Minas Gerais, Brazil, comprising installed capacity, number of PV units, solar-radiation estimates and realized energy output. Optimization yields a parsimonious composite based primarily on the DEA and SFA models, reducing the principal-component loss metric by 18 % relative to equal weighting. Spatial analysis reveals a clear north–south efficiency gradient correlated with solar insolation; nonetheless, centrally located industrial hubs also achieve high scores, underscoring the moderating role of infrastructure. The findings highlight regions where policy support or technology upgrades could deliver the greatest marginal gains and illustrate how mixture-design optimization can reconcile diverse frontier estimates, offering a generalizable decision-analytics toolkit for renewable-energy assessment.
Vision Language Models (VLM) for Remote Sensing (RS) imagery have recently attracted significant interest, with tasks such as image captioning, cross-modal retrieval and visual question answering. However, progress is hampered by the limited size and diversity of existing captioning datasets and by the need for domain-specific models that can handle remote-sensing imagery, which differs substantially from everyday photos. This work introduces a framework that integrates RS imagery with OpenStreetMap (OSM) data to support RS scene imagery description and reasoning. The framework employs a remote-sensing image captioner together with an AI-driven agent to produce aligned im-age-map interpretations, enabling geo-aware scene description and region-specific question answering. The captioner is benchmarked across three encoder-decoder architectures, achieving ($\text{BLEU}-4=0.5132, \text{CIDEr}=2.0164, \text{METEOR}= 0.6828,\ \text{ROUGE-L} =0.7087$). The agent is evaluated using an Evidence-Constrained LLM-as-Judge protocol, obtaining an average Factual Accuracy score of 4.59/5 and Completeness (4.07/5). The results point to a promising direction for integrating VLMs with a geospatial knowledge base to enhance multimodal analysis and text-image representation in environmental and geographic applications.
Timimoun Province in south-central Algeria receives some of the world's highest direct normal irradiance (DNI), yet its CSP potential remains largely unmapped. This work presents a remote-sensing-based framework for generating the first 30m CSP suitability atlas for the region. Multisource Earth-observation layers-including satellitederived DNI, SRTM terrain data, and Landsat land cover-are combined with infrastructural proximity metrics. A fuzzy Analytic Hierarchy Process (FAHP) is applied using six criteria (DNI, slope, aspect, distance to roads, distance to buildings, and distance to transmission lines), with expertvalidated weights (CR<0.05). Three grid-access conditions were explored to quantify infrastructure impacts on siting. Under a strict high-grid access scenario, ≈83% of the province is excluded, and only 3-5% of land qualifies as highly suitable. Allowing access to lower-grid connections reduces the exclusion footprint and increases high-suitability areas to ≈13%. When all distance constraints are removed, suitability expands to ≈39%, revealing that roughly one-fifth of prime land is currently stranded due to limited grid reach.
Algeria combines exceptional solar resources with emerging European demand for low carbon molecules, positioning the country as a strategic candidate for large scale green hydrogen. This study develops a spatial decision support framework that links a national photovoltaic direct current yield (PV DC yield) atlas to a hydrogen production atlas, techno economic evaluation, and screening level environmental co benefits. Two PV configurations, horizontal and optimally tilted are assessed on a land normalized basis and converted to hydrogen production. To reflect deployment realism, the analysis applies multiple infrastructure and accessibility scenarios that differentiate feasible development envelopes. In addition, an export oriented outlook is examined to assess how Algeria could contribute up to 10 million tons of hydrogen per year (10 Mt HQ & sdot;yr & iexcl;1) to European markets through alternative pathways, including pipeline based delivery, desalination supported water supply, and electricity export concepts, with explicit attention to water availability and sourcing. The maps reveal strong geographic gradients in PV productivity that propagate into hydrogen density and cost competitiveness. Optimal tilt systematically increases PV yield and hydrogen production density and reduces levelized cost of hydrogen (LCOH) across the territory, while scenario filters primarily reshape the reachable portfolio of low cost sites rather than the physics based ranking. For the modeled assumptions, LCOH spans approximately 4.6-5.2 & euro;/kg HQ, with tilted PV delivering a consistent cost improvement relative to horizontal PV. Environmental indicators indicate substantial benefits versus grey hydrogen, with avoided COQ and natural gas displacement concentrating in the highest productivity zones. Overall, the atlas results provide an integrated
This study investigates an approach to examine the influence of urbanization-induced land use changes on surface runoff. The research leverages the SCS-CN method, integrating remote sensing and machine learning, to analyze land use and cover (LULC) changes over the years 2000 to 2040. Initial land use classification (2000–2020) was performed using the SVM algorithm, while a novel temporal approach was applied to predict LULC for the years 2025, 2030, and 2040. The accuracy of the LULC prediction model was validated, achieving an overall accuracy of 85.05
The increasing frequency of wildfires in wetlands demands rapid and accurate monitoring tools to mitigate regional ecological damage and assess carbon emission impacts. While Deep Learning (DL) models offer high precision for environmental monitoring, their operational deployment by environmental agencies is often hindered by the prohibitive cost and time required to generate dense, pixel-wise manual annotations. To address this labeling bottleneck, which critically delays disaster response and restoration policies, this work proposes an uncertainty-aware Weakly Supervised Semantic Segmentation (WSSS) framework tailored for high-resolution multispectral imagery. Our approach relies solely on image-level tags (e.g., “burned” or “unburned”) to train segmentation models, significantly democratizing and accelerating the mapping process. Technically, we adapt the Puzzle-CAM architecture to handle RGB-NIR data and integrate an Imbalanced Class and Uncertainty-aware loss (LICU) during the training stage. This loss function is specifically designed to mitigate the label noise inherent in the WSSS, particularly at the fuzzy boundaries between burned areas and spectrally similar wetland features. Experimental results demonstrate that our method produces sharp, boundary-adherent segmentation masks, achieving an IoU of 83.74% for the burned class. Notably, our approach outperforms recent contrast-based WSSS baselines and excels in detecting small, fragmented fire scars (1%–20% coverage). Furthermore, while a performance gap of approximately 11 percentage points remains when benchmarked against cutting-edge fully supervised foundation models, these findings suggest that uncertainty-aware WSSS is a viable, cost-effective alternative for operational fire monitoring, providing robust spatial data for environmental management, carbon accounting, and data-driven ecosystem analysis.
This study evaluates convolutional U-Net and transformer-based SegFormer architectures using very high-resolution (20 cm) airborne imagery for landslide segmentation in the Emilia-Romagna region in northern Italy. In May 2023, the region experienced two extreme rainfall events that triggered a large number of landslides, motivating the need for accurate post-event mapping. Results show that transformer models consistently outperform CNNs, with SegFormer MiT-B2 achieving strong baseline performance $(\mathrm{F} 1=84.57 \%, \text{IoU}=55.11 \%, \text{mIoU}= 75.81 {\%}$). Regarding loss-function, the Tversky configuration with $\alpha=0.6$ and $\beta=0.4$ yields the best overall improvement, increasing performance to $\mathrm{F}1 = 85.43 \%, \text{IoU}=56.11 {\%}$, and $\text{mIoU}=75.97 {\%}$. These results suggest the effectiveness of hierarchical transformer encoders and carefully tuned loss functions for accurate and reliable landslide mapping.
Landslides require rapid, data-driven post-event assessment to support emergency response, yet most existing approaches focus on stand-alone landslide detection or mapping accuracy, with limited integration into operational decision-support workflows. This study presents MAS-LAND (Multi-Agent System for Landslide Detection and Rapid Response), a multi-agent, LLM-enhanced framework structured around three main stages: (1) post-event landslide detection, (2) infrastructure exposure assessment, and (3) automated reporting for civil protection operations. The framework is applied to the May 2023 Emilia-Romagna disaster and leverages very highresolution post-event orthophotos and transformer-based semantic segmentation to identify newly occurred landslides. Among the evaluated models, SegFormer MIT-b2 achieved the strongest performance, detecting 603 of 806 landslides in the test dataset (Recall = 74.81%, F1 = 71.14%) and delivering robust multi-class segmentation results (macro F1 = 71.09%) despite pronounced class imbalance. In the second stage, infrastructure exposure is assessed through automated building footprint extraction using a pretrained RT-DETR model and road-network data derived from OpenStreetMap, enabling deterministic prioritization of intervention needs via a transparent, rulebased decision matrix. In the third stage, a reporting agent powered by Llama-3.3-70B, operated in deterministic mode, synthesizes analytical outputs into standardized operational summaries tailored for civil protection use. A key contribution of this work is the delivery of a fully integrated, post-event pipeline capable of transforming heterogeneous geospatial inputs into actionable intelligence within a few minutes of image acquisition and processing, as this ability is largely absent from previous landslide studies. The framework also produces structured, machine-readable outputs at every stage, forming a scalable knowledge base for systematic storage and reuse of event information. Overall, the results demonstrate the potential of multi-agent, hybrid AI-geospatial architectures to enhance situational awareness, improve reproducibility, and support timely decision-making of Civil Protection authorities during landslide emergencies.
Transformer has shown to be a very effective finding to solve numerous learning tasks for various application fields, such as the image captioning task, which this work will focus on. Its widespread success is owed to two main ingredients: 1) an attention mechanism and 2) positional encoding. This article is interested in the first ingredient, showing that the vanilla attention mechanism may be improved by exploiting not only the context conveyed in a data sequence under analysis, but also in all sequences forming the complete training dataset. In particular, we introduce the concept of attention history to better capture and model contextual information. Two different strategies to compute attention history before its injection in the attention mechanism are described. The proposed solution is validated and discussed on four reference remote sensing image captioning datasets.
Pansharpening is the process of fusing a high-resolution panchromatic image with a low-resolution multispectral image to yield a high-resolution multispectral image with enhanced spacial detail. Efficiently preserving spatial detail in the output image and processing of large volumes of images emerges as the main concerns of image pansharpening. To address these challenges, we leverage the recently introduced state-space models, which have been proven as competitive alternative to convolution neural networks and Transformers. We adapt MambaIR, a recently proposed state-space model for image restoration, to the pansharpening task. To efficiently inject spacial details, we adopt an approach similar to the successful FusionNet pansharpening network, where the nonlinear injection model is extracted through a deep convolution network. Combining these two major contributions yield MambaFuse, our innovative deep learning approach for pansharpening. Extensive experimental evaluation, involving qualitative and quantitative assessment on two urban area satellite datasets, and comparison with recent deep learning approaches demonstrate the soundness of our proposed MambaFuse pansharpening approach. For reproducibility purposes, we provide the source code and the experimental setups of our models at: code https://github.com/faridbi/MambaFuse.
The detection of fire and smoke is critical for early warning systems to avert disasters. This paper presents a deep learning-based segmentation method designed for the precise detection of fire and smoke. We have developed a comprehensive multi-label segmentation dataset and introduced an enhanced deep-learning approach specifically tailored for fire and smoke segmentation. To address the issue of class imbalance, we implemented a loss function overlay strategy to refine the segmentation of underrepresented classes. Experimental results validate the robustness and precision of our proposed method, with fire and smoke classes achieving a Pixel Accuracy (PA) of 99.8% and 98.4%, and a Recall of 97.2% and 92.5%.The model performance surpasses several comparative methods, boasting an overall F1 value of 89.5%.
The high volume of emergency room patients often necessitates head CT examinations to rule out ischemic, hemorrhagic, or other organic pathologies. A system that enhances the diagnostic efficacy of head CT imaging in emergency settings through structured reporting would significantly improve clinical decision making. Currently, no AI solutions address this need. Thus, our research aims to develop an automatic radiology reporting system by directly analyzing brain anomalies in head CT data. We propose a multi-branch CNN-LSTM fusion network-driven system for enhanced radiology reporting in emergency settings. We preprocessed head CT scans by resizing all slices, selecting those with significant variability, and applying PCA to retain 95% of the original data variance, ultimately saving the most representative five slices for each scan. We linked the reports to their respective slice IDs, divided them into individual captions, and preprocessed each. We performed an 80-20 split of the dataset for ten times, with 15% of the training set used for validation. Our model utilizes a pretrained VGG16, processing groups of five slices simultaneously, and features multiple end-to-end LSTM branches, each specialized in predicting one caption, subsequently combined to form the ordered reports after a BERT-based semantic evaluation. Our system demonstrates effectiveness and stability, with the postprocessing stage refining the syntax of the generated descriptions. However, there remains an opportunity to empower the evaluation framework to more accurately assess the clinical relevance of the automatically-written reports. Part of future work will include transitioning to 3D and developing an improved version based on vision-language models.
Normal mixture models are widely used to represent data arising from latent subpopulations. We propose a Design-of-Experiments (DOE) and Response Surface Methodology (RSM) framework to estimate the weights of a bimodal Gaussian mixture when component families are known. The procedure is non-iterative: rather than alternating Expectation Maximization (EM) steps, it performs a double-stage method - fit a quadratic response surface to the sample log-likelihood over the weight simplex and solve one constrained optimization - followed by a final Maximum Likelihood re-estimation of means and variances. This yields predictable runtime (driven by design size) and reduced sensitivity to initialization. The pipeline uses 1) k-medians to obtain preliminary component parameters and 99% confidence intervals (CIs) for component proportions; 2) builds a simplex-lattice mixture design within those CI bounds; 3) fits a quadratic response surface to log-likelihood; and 4) optimizes this surface under sum-to-one constraints. We validate the method in 27 Monte Carlo scenarios (n = 100, 500, 1000; low/medium/high differentiation and three weight settings). In medium/high separation, it attains comparable likelihoods to EM while achieving more favorable BIC in multiple scenarios and indistinguishable AIC in many, whereas EM is preferable under low separation. Two real data sets - Old Faithful (Waiting variable) and Photovoltaic Energy (Production variable) - further confirm applicability, with lower AIC/BIC in Old Faithful and lower BIC in PV; clustering agreement is high (kappa approximate to 0.99 - 1.00). Overall, DOE-RSM offers a simple, interpretable, and often more parsimonious method, and constitutes a non-iterative alternative for mixture-weight estimation.
Understanding tropical cyclones (TCs) impact is critical for implementing effective post–cyclone management methods. While previous researches have assessed the effects of TCs in Oman, this study uniquely examines and quantifies the impact of the Shaheen tropical cyclone (STC) using advanced image resolution and analysis, including artificial intelligence in the form of Multiple Deep learning models, as well as using very high-resolution (41 cm) satellite imagery. This study found significant vegetation loss, with 85.8 ha (51.3 This study used high–resolution satellite imagery to assess the impact of the Shaheen tropical cyclone. The study fused pre–impact and post–impact imagery, which was processed and used to train AI models to detect changes in vegetation, buildings, and water surfaces. The AI models achieved high accuracy (94
Enrico Blanzieri合作论文数and Communication Technology;University of Trento;DIT - Department of Information 5