Wildfires are becoming increasingly destructive and costly each year, affecting lives, damaging infrastructure, and degrading ecosystems. To address this growing threat, the fire and land management community need smarter, data-driven tools to understand the landscape and plan their essential prescribed burns that reduce hazardous fuels. With the ultimate goal to minimize the devastation of wildfires by enabling proactive and data-driven fuel management at landscape scale, this paper presents an approach that builds heterogeneous remote sensing data into a temporal-spatial knowledge graph, then queries it with a Large Language Model (LLM) based agent, providing a natural language interface to an extensive system of granular landscape knowledge and metrics. We demonstrate how we build knowledge graphs from LiDAR-derived vegetation metrics using GraphDB enabling precise location and time-based insights. We present a user-facing system intended to respond to queries about the effects of prescribed burns over time. Built around an LLM Agent (e.g., OpenAI, LLaMA) orchestrated with LangChain and LangGraph, the system allows users to interact with complex fire and fuel data through a natural language chat interface. It also includes a web search tool for retrieving external fire-related content to enrich responses. While this work operates as a standalone knowledge system, it was first envisioned as a method of generating conversational narratives while navigating a virtual 3D forest environment in our prototype immersive visualization application called Immersive Forest.
In recent years, digital twins have emerged as an advanced data-driven technology for modeling and monitoring complex physical systems to inform actionable decision-making. One area where digital twins can be especially impactful is urgent computing applications, such as natural disasters and hazards, where timely and reliable decision support is needed. However, it is challenging to streamline dynamic data with modeling and simulation methods for urgent computing applications due to the use of heterogeneous multimodal data sources, Internet of Things (IoT) networks, and high-performance computing (HPC) systems to power edge-to-cloud capabilities. In this work, we describe a reference architecture that uses federation and dynamic composability to power digital twins for complex physical systems. We demonstrate the usefulness of our approach with the Firemap architecture on WIFIRE, an end-to-end cyberinfrastructure for real-time and data-driven simulation, prediction, and visualization of wildfire behavior.
The increasing frequency and severity of wildfires in the Western United States demand improved fire response tools. Initial Attack Fire Response within the first few hours of ignition is critical in preventing fires from escalating. WIFIRE Firemap has been instrumental in supporting early fire suppression efforts through real-time fire behavior modeling. However, wildfires often burn for days or weeks, necessitating longer-term predictive capabilities. To address this challenge, we extended Firemap to forecast fire spread from the first few hours to five days. This advancement integrates two long-term fire behavior models, ELMFIRE and GridFire, enabling real-time, data-driven decision support. The enhanced Firemap platform improves strategic wildfire response planning, allowing firefighters and emergency managers to anticipate fire spread on extended timelines. We present how these extensions were used during the Los Angeles firestorms of 2025, demonstrating their potential to mitigate wildfire risks, protect communities, and improve firefighting strategies, and make recommendations for effective use of extended attack tools for decision support.
Driven by dangerous Santa Ana winds and fueled by dry vegetation, the 2025 Eaton and Palisades wildfires in California caused historic levels of devastation, ultimately becoming the second and third most destructive fires in California history. Burning at the same time and drawing from the same resources, these fires burned a combined total of 16,251 structures. The first several hours of an emerging wildfire are a crucial period for fire officials to assess potential damage and develop a timely and appropriate response. A method to quickly generate accurate estimates of structural damage is essential to providing this crucial rapid response to wildfires. In this paper, we present a machine learning approach for automated assessment of structural damage caused by wildfires. By leveraging multiple data sources in model development (satellite-based building footprints, expert-labeled post-fire damage points, fire perimeters, and aerial thermal imagery) and innovative data processing techniques, the approach can be used to identify various levels of structural damage from just aerial thermal imagery during operational use. The resulting system offers an effective approach for rapid and reliable assessment of burned structures, suitable for operational wildfire damage assessment. Results on the Eaton and Palisades Fires demonstrate the effectiveness of this method and its applicability to real-world scenarios.
This project aims at developing an AI system to provide a reliable assessment of the structural damage caused by wildfires in the first burn period. Our approach uses multimodal data, including multispectral aerial images, historical post-fire damage assessment data, and building footprints, to create an association between damage data and structure footprints. We use these associations to generate features and use machine learning methods to assess the level of damage to structures. The resulting AI-driven system can be used to provide wildfire-induced structural damage assessments in near-real-time using only aerial images for future fires. We provide damage assessment results on several megafires in California to demonstrate the applicability of our approach to real wildfire scenarios.
In recent years, frequent and highly destructive megafires become one of the biggest climate-induced disasters. Fire behavior models using data from many emerging sources can inform decision support tools to respond to and mitigate such megafires. Emerging edge sensing and computing technologies within the fire environment can enhance the speed, reliability, and efficiency of wildland fire management, leading to better prevention, faster response times, and more effective mitigation of fire-related disasters. However, a unified system that streamlines the integration of edge technology advances within fire science and management workflows is needed. This paper presents the design and demonstrated case studies of the WIFIRE Edge Platform that facilitates the integration of sensing and AI capabilities at the edge. The initial attack and prescribed burn concept scenarios are described, highlighting the sensor deployment and utilization at the fire front.
Reliable performance metrics are necessary prerequisites to building large-scale end-to-end integrated workflows for collaborative scientific research, particularly within context of use-inspired decision making platforms with many concurrent users and when computing real-time and urgent results using large data. This work is a building block for the National Data Platform, which leverages multiple use-cases including the WIFIRE Data and Model Commons for wildfire behavior modeling and the EarthScope Consortium for collaborative geophysical research. This paper presents an artificial intelligence and machine learning (AI/ML) approach to performance assessment and optimization of scientific workflows. An associated early AI/ML framework spanning performance data collection, prediction and optimization is applied to wildfire science applications within the WIFIRE BurnPro3D (BP3D) platform for proactive fire management and mitigation.
Wildland fire modeling tools can ingest high resolution 3D vegetation models as inputs. However, data used to build the surface fuels in these models is often at a 30-meter resolution, which does not necessarily provide sufficient detail for accurate modeling of fires. Terrestrial laser scans are increasingly being used to collect detailed vegetation data that could be integrated with new approaches to fuel and fire modeling, but manual segmentation of scans is not scalable beyond a small number of scans. There is a need to automatically segment these high resolution point clouds as they are collected in the field, such that they may be leveraged by fuel and fire models for wildland fire response and mitigation and other applied climate science. This paper summarizes our early work on a labeling, visualization and machine learning pipeline for detailed segmentation of fuels. Specific contributions are: (1) a labeling approach involving 3 dimensional segmentation of point clouds using a point cloud processing engine; (2) a visualization approach using a computer graphics engine; and (3) early results from a deep learning modeling approach for fuel segmentation by category (live and dead) and size class (1, 10, 100 and 1000 hour fuels).
The FAIR and CARE data principles are critical to ensuring widespread and equitable access to open data. They provide guidelines for what should be contained in metadata and how certain types of data should be handled. This study examines how large, collective data hubs such as the WIFIRE Data Commons can implement the FAIR and CARE data principles through evaluating the datasets hosted on the platform. An automation pipeline was developed to check for specified criteria in these principles, allowing fast integration of the principles on a large scale data hub. This pipeline can be expanded to check for all FAIR and CARE criteria, and similar pipelines can be created for a variety of other data hubs. Automating for the FAIR and CARE principles will help simplify organization of open data, allowing for a greater expansion of open science.
Long-standing fire suppression policies, global warming and human influence at the urban-wildland interface are fueling a global wildfire crisis. Prescribed burns are increasingly being recognized as an essential procedure to reduce fuel (biomass) overgrowth and mitigate the size and severity of uncontrolled wildfires. Next-generation three-dimensional (3D) fire behavior simulations, which can help land managers to identify risks and improve planning for successful prescribed burns, depend on accurate 3D fuel structure models. We introduce TrueTrees, a workflow that integrates tree-level observations from airborne lidar surveys into FastFuels 3D fuel models. The workflow is optimized and distributed to allow processing of point cloud data for typical burn units within minutes. TrueTrees is implemented into a prototype of BurnPro3D, a user-friendly science-driven decision support platform for prescribed burn planners.
Research has shown that climate change creates warmer temperatures and drier conditions, leading to longer wildfire seasons and increased wildfire risks in the United States. These factors have, in turn, led to increases in the frequency, extent, and severity of wildfires in recent years. Given the danger posed by wildland fires to people, property, wildlife, and the environment, there is an urgent need to provide tools for effective wildfire management. Early detection of wildfires is essential to minimizing potentially catastrophic destruction. To that end, in this paper, we present our work on integrating multiple data sources into SmokeyNet, a deep learning model using spatiotemporal information to detect smoke from wildland fires. We present Multimodal SmokeyNet and SmokeyNet Ensemble for multimodal wildland fire smoke detection using satellite-based fire detections, weather sensor measurements, and optical camera images. An analysis is provided to compare these multimodal approaches to the baseline SmokeyNet in terms of accuracy metrics, as well as time-to-detect, which is important for the early detection of wildfires. Our results show that incorporating weather data in SmokeyNet improves performance numerically in terms of both F1 and time-to-detect over the baseline with a single data source. With a time-to-detect of only a few minutes, SmokeyNet can be used for automated early notification of wildfires, providing a useful tool in the fight against destructive wildfires.
This paper shows that the prediction capability of wildfire progression can be improved by estimation of a single prevailing wind vector parametrized by a wind speed and a wind direction to drive a wildfire simulation created by FARSITE. Estimations of these wind vectors are achieved in this work by a gradient-free optimization via a grid search that compares wildfire model simulations with measured wildfire perimeters, where noisy observations are modeled as uncertainties on the locations of the vertices of the measured wildfire perimeters. Two optimizations are established to acquire the optimal wind speed and wind direction. To formulate a perimeter optimization, an uncertainty-weighted least-squares error is computed between the vertices of the simulated and measured wildfire perimeter. The challenge in this approach is to match the number of vertices on the simulated and measured wildfire perimeter via interpolation of perimeter points and their uncertainties. For a surface area optimization, an uncertainty-weighted surface area error is introduced to capture the surface of the union minus the intersection of the simulated and measured wildfire perimeter. The challenge in this approach is to formulate a surface area error, weighted by the uncertainties on the locations of the vertices of the measured wildfire perimeter. The optimization in this paper is based on an iterative refinement of a grid of the wind vector and provides robustness to intermittent erroneous results produced by FARSITE, while allowing parallel execution of wildfire model calculations. This paper is an extension of the work in Tan et al., (2021). Results on wind vector estimation are illustrated on two historical wildfire events: the 2019 Maria Fire that burned south of the community of Santa Paula in the area of Somis, CA, and the 2019 Cave Fire that started in the Santa Ynez Mountains of Santa Barbara County.
Timely prediction of debris flow probabilities in areas impacted by wildfires is crucial to mitigate public exposure to this hazard during post-fire rainstorms. This paper presents a machine learning approach to amend an existing dataset of post-fire debris flow events with additional features reflecting existing vegetation type and geology, and train traditional and deep learning methods on a randomly selected subset of the data. The developed methods achieve AUC (area under the receiver operational characteristic curve) values of 0.93 (random forest) and 0.92 (neural network) on the test set, representing a significant improvement over a logistic regression model currently used (AUC 0.79). The paper also overviews a distributed, Kubernetesbased big data processing pipeline to efficiently retrieve features in areas impacted by new fires, and deploy the methods for real-time prediction of debris flow hazards.
Increasing social acceptance of prescribed burns is an important element of ramping up these controlled burns to the scale required to effectively mitigate destructive wildfires through reduction of excessive fire fuel loads. As part of a Design Challenge, students created concept designs for physical or virtual installations that would increase public understanding and acceptance of prescribed burns as an important tool for ending devastating megafires. The proposals defined how the public would interact with the installation and the learning goals for participants. This poster provides an overview of the virtual reality (VR) pipeline created to develop working prototypes of the immersive experiences and VR games that were proposed by the finalists in the design challenge.
This paper shows how a gradient-free optimization method is used to improve the prediction capabilities of wildfire progression by estimating the wind conditions driving a FARSITE wildfire model. To characterize the performance of the prediction of the perimeter as a function of the wind conditions, an uncertainty weighting is applied to each vertex of the measured fire perimeter and a weighted least-squares error is computed between the predicted and measured fire perimeter. In addition, interpolation of the measured fire perimeter and its uncertainty is adopted to match the number of vertices on the predicted and measured fire perimeter. The gradient-free optimization based on iterative refined gridding provides robustness to intermittent erroneous results produced by FARSITE and quickly find optimal wind conditions by paralleling the wildfire model calculations. Results on wind condition estimation are illustrated on two historical wildfire events: the 2019 Maria fire that burned south of the community of Santa Paula in the area of Somis, CA, and the 2019 Cave fire that started in the Santa Ynez Mountains of Santa Barbara County.
This paper shows how the wildfire simulation tool FARSITE is augmented with data assimilation capabilities that exploit the notion of barrier points and a constraint-point ensemble Kalman filtering to update wildfire perimeter predictions. Based on observations of the actual fire perimeter, stationary points on the fire perimeter are identified as barrier points and combined with a recursive update of the initial fire perimeter. It is shown that the combination of barrier point identification and using the barrier points as constraints in the ensemble Kalman filter gives a significant improvement in the forward prediction of the fire perimeter. The results are illustrated on the use case of the 2016 Sandfire that burned in the Angeles National Forest, east of the Santa Clarita Valley in Los Angeles County, California.
High-resolution satellite imagery is a rich source of data applicable to a variety of domains, ranging from demo-graphics and land use to agriculture and hazard assessment. We have developed an end-to-end analysis pipeline that uses deep learning and unsupervised learning to process high-resolution satellite imagery and have applied it to various applications in previous work. As high-resolution satellite imagery is large-volume data, scalability is important to be able to analyze data from large geographical areas. To add scalability to our process, we converted our original pipeline, implemented using the Caffe deep learning library and the Python machine learning library Scikit-Learn, to other platforms that make use of distributed computation. Specifically, to add scalability, we use Keras for deep learning, and evaluate two different distributed platforms, Spark and Dask, for unsupervised learning. We report on results in scaling up our satellite analysis pipeline.
The advances in data, computing and networking over the last two decades led to a shift in many application domains that includes machine learning on big data as a part of the scientific process, requiring new capabilities for integrated and distributed hardware and software infrastructure. This paper contributes a workflow-driven approach for dynamic data-driven application development on top of a new kind of networked Cyberinfrastructure called CHASE-CI. In particular, we present: 1) The architecture for CHASE-CI, a network of distributed fast GPU appliances for machine learning and storage managed through Kubernetes on the high-speed (10-100Gbps) Pacific Research Platform (PRP); 2) A machine learning software containerization approach and libraries required for turning such a network into a distributed computer for big data analysis; 3) An atmospheric science case study that can only be made scalable with an infrastructure like CHASE-CI; 4) Capabilities for virtual cluster management for data communication and analysis in a dynamically scalable fashion, and visualization across the network in specialized visualization facilities in near real-time; and, 5) A step-by-step workflow and performance measurement approach that enables taking advantage of the dynamic architecture of the CHASE-CI network and container management infrastructure.
Data-intensive science communities are progressively adopting FAIR practices that enhance the visibility of scientific breakthroughs and enable reuse. At the core of this movement, research objects contain and describe scientific information and resources in a way compliant with the FAIR principles and sustain the development of key infrastructure and tools. This paper provides an account of the challenges, experiences and solutions involved in the adoption of FAIR around research objects over several Earth Science disciplines. During this journey, our work has been comprehensive, with outcomes including: an extended research object model adapted to the needs of earth scientists; the provisioning of digital object identifiers (DOI) to enable persistent identification and to give due credit to authors; the generation of content-based, semantically rich, research object metadata through natural language processing, enhancing visibility and reuse through recommendation systems and third-party search engines; and various types of checklists that provide a compact representation of research object quality as a key enabler of scientific reuse. All these results have been integrated in ROHub, a platform that provides research object management functionality to a wealth of applications and interfaces across different scientific communities. To monitor and quantify the community uptake of research objects, we have defined indicators and obtained measures via ROHub that are also discussed herein.
Shava Smallen合作论文数San Diego Supercomputer Center2