Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.
This paper investigates a novel concept of time series geolocalization, where the goal is to infer the geographic origin of each raw time series. Successful geolocalization can provide spatial context to time series, enabling downstream location-aware applications. We formalize the problem, adapt core ideas from image geolocalization to establish strong baselines, and propose GeoGNN, a two-tower architecture. During training, GeoGNN's spatial tower learns embeddings of geographic cell candidates by leveraging the geographic adjacency graph, while the temporal tower extracts informative representations from time series. During inference, each temporal representation is matched against candidate geographic embeddings using dot-product similarity, combined with an auxiliary classification head, to predict the time series' associated geographic origin. Experiments on large-scale, countrywide electricity-consumption datasets demonstrate that GeoGNN achieves the best performance across datasets and enhances both fine- and coarse-grained geolocalization accuracy by 27
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.
Advances in artificial intelligence (AI) and multimodal sensing are driving a paradigm shift in the geospatial sciences, moving from task specific GeoAI models toward general purpose Geospatial Foundation Models (GeoFMs). While these models offer unprecedented opportunities for Earth monitoring, geographic knowledge discovery and addressing societal challenges such as natural disaster management, challenges remain regarding multimodal alignment, spatial reasoning, spatial distribution shifts, and generalizability. This article introduces the first section of a special issue dedicated to advancing the state-of-the-art in GeoFMs and their applications, featuring diverse research on neurosymbolic AI, heterogeneous graph learning, self-supervised learning, spatial retrieval-augmented generation, a genealogical review, and additive compositionality in urban representation learning. Together, these contributions offer new insights toward the next generation of geospatial intelligence.
Learning general-purpose representations of geographic locations has become essential to geospatial tasks such as population estimation and environmental monitoring. To obtain such representations, multimodal geo-foundation models often use contrastive learning (CL) to align satellite imagery with geo-coordinates, implicitly assuming that cross-modal (shared) information suffices for downstream tasks. However, not all task-relevant information is shared between modalities, and retaining modality-specific (unique) features can improve task performance. Prior methods retain unique information through extra training objectives or databases, increasing training complexity and computation. Motivated by the conventional wisdom that earlier layers capture general input features while later layers become task-specific, we hypothesize that early layers in CL models consist unique information that is lost toward the final layer. Through a comprehensive layerwise analysis of modality gap, representation similarity, and mutual information, we confirm this trend and find that fusing intermediate (more unique) and final (more shared) representations outperforms state-of-the-art models across diverse geospatial benchmarks. Our findings reveal underutilized information diversity in CL models and show that simple layerwise fusion is an efficient path to richer geo-embeddings.
The rapid advancement of generative models has made the detection of AI-generated images a critical challenge for both research and society. Recent works have shown that most state-of-the-art fake image detection methods overfit to their training data and catastrophically fail when evaluated on curated hard test sets with strong distribution shifts. In this work, we argue that it is more principled to learn a tight decision boundary around the real image distribution and treat the fake category as a sink class. To this end, we propose SimLBR, a simple and efficient framework for fake image detection using Latent Blending Regularization (LBR). Our method significantly improves cross-generator generalization, achieving up to +24.85\% accuracy and +69.62\% recall on the challenging Chameleon benchmark. SimLBR is also highly efficient, training orders of magnitude faster than existing approaches. Furthermore, we emphasize the need for reliability-oriented evaluation in fake image detection, introducing risk-adjusted metrics and worst-case estimates to better assess model robustness. All code and models will be released on HuggingFace and GitHub.
Self-supervised learning (SSL) has emerged as a powerful pretraining strategy to learn transferable representations from unlabeled data. Yet, it remains unclear how long SSL models should be pretrained for such representations to emerge. Contrary to the prevailing heuristic that longer pretraining translates to better downstream performance, we identify a transferability trade-off: across diverse SSL settings, intermediate checkpoints often yield stronger out-of-domain (OOD) generalization, whereas additional pretraining primarily benefits in-domain (ID) accuracy. From this observation, we hypothesize that SSL progresses through learning phases that can be characterized through the lens of critical periods (CP). Prior work on CP has shown that supervised learning models exhibit early phases of high plasticity, followed by a consolidation phase where adaptability declines but task-specific performance keeps increasing. Since traditional CP analysis depends on supervised labels, for SSL we rethink CP in two ways. First, we inject deficits to perturb the pretraining data and measure the quality of learned representations via downstream tasks. Second, to estimate network plasticity during pretraining we compute the Fisher Information matrix on pretext objectives, quantifying the sensitivity of model parameters to the supervisory signal defined by the pretext tasks. We conduct several experiments to demonstrate that SSL models do exhibit their own CP, with CP closure marking a sweet spot where representations are neither underdeveloped nor overfitted to the pretext task. Leveraging these insights, we propose CP-guided checkpoint selection as a mechanism for identifying intermediate checkpoints during SSL that improve OOD transferability. Finally, to balance the transferability trade-off, we propose CP-guided self-distillation, which selectively distills layer representations from the sweet spot (CP closure) checkpoint into their overspecialized counterparts in the final pretrained model.
Object detection in remote sensing demands extensive, high-quality annotations-a process that is both laborintensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing earlystage localization accuracy. Extensive evaluations on multiple remote sensing datasets, along with a real-world user study, demonstrate that our framework not only reduces annotation effort but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface is available under https://github.com/mburgesCVI/ICCV_AL4FM.
In this paper, we propose IRTR-DETR, an Interactive and Real-Time Rotated DEtection TRansformer that extends IRT-DETR to predict rotated bounding boxes. IRTR-DETR maintains the Human-In-The-Loop (HIL) workflow of IRTDETR but introduces rotation-aware heads for improved detection of objects with arbitrary orientations. Similarly to IRTDETR, IRTR-DETR can be trained with a small labeled sample set in an interactive setting, but we show that it can also be pretrained on related but not identical data-such as a building damage dataset-before being applied to tasks like identifying buildings under construction. We demonstrate the efficacy of our approach on the publicly available Tiny-DOTA and xBD dataset, as well as two study-cases on proprietary datasets of greenhouses and houses under construction (“waffle homes”). Detecting greenhouses is highly relevant in the context of damage assessment, while “waffle homes” aid understanding typical floorplans and building codes in different areas, both thereby supporting population modeling, emergency response, and policy planning. Our method outperforms the state of the art in interactive rotated object detection on the Tiny-DOTA dataset by 5.7 percent, and improves upon the non interactive RTDETR by 7.85 to 19.39 percent (depending on the number of provided samples) while maintaining its real-time efficiency.
Unwarned population distributions accounting for routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (LSHD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (approximately 90 m). Although LSHD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LSHD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LSHD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LSHD baseline dataset, we contrast unwarned and residential (WorldPop) population distributions by (1) exploring a practical application of flood risk assessment and (2) evaluating their congruence with outcomes of collective human activities (subnational CO2 emissions). Finally, we discuss plans to address current LSHD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.
Advances in artificial intelligence, hardware accelerators, and data processing architectures, continue to infiltrate the geospatial information sciences, with a transformative impact on many societal challenges. Recent breakthroughs in deep learning have brought forward an automated capability to learn representational features from massive and complex data, including text, images, and videos. In tandem, rapid innovations in sensing technologies enhance the collection of geospatial data in even higher resolution and throughput, supporting the observation, mapping, and analysis of different events and phenomena on the Earth's surface with unprecedented detail. Combined, these developments are offering the potential for breakthroughs in geographic knowledge discovery, impacting decision-making in areas such as humanitarian mapping, intelligent transport systems, urban expansion analysis, health data analysis and epidemiology, the study of climate change, handling natural disasters, the general monitoring of the Earth's surface, and achieving sustainability.
The Annual Meeting of the American Association of Geographers (AAG) in 2023 marked a five-year milestone since the first Geospatial Artificial Intelligence (GeoAI) Symposium was held at AAG in 2018. In the past five years, progress has been made while open questions remain. In this context, we organized an AAG panel and invited five panellists to discuss the advances and limitations in GeoAI research. The panellists commended the successes, such as the development of spatially explicit models, the production of large-scale geographic datasets, and the use of GeoAI to address real-world problems. The panellists also shared their thoughts on limitations in current GeoAI research, which were considered as opportunities to engage theories in geography, enhance model explainability, quantify uncertainty, and improve model generalizability. This article summarizes the presentations from the panellists and also provides after-panel thoughts from the organizers. We hope that this article can make these thoughts more accessible to interested readers and help stimulate new ideas for future breakthroughs.
As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.
In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.
Many machine learning, computer vision, artificial intelligence-inspired approaches have been developed to process and analyze voluminous, heterogeneous, and distributed remote sensing data. The success of these effective imagery analytics has the potential to enable us end-to-end applications. However, determining the most effective data and algorithms remains challenging. Therefore it is crucial to have appropriate benchmarking methods and designs to ensure the effective adaptation of GeoAI systems and to leverage the rich remote sensing data for various applications. This chapter discusses recent developments and challenges (data, metrics, protocol-related) in benchmarking for GeoAI systems. We also highlight important considerations to support and maximize the impact of GeoAI benchmarking.
Fabio Pacifici合作论文数DigitalGlobe, Inc.4