About 25% of the world’s population live in informal urban settlements containing densely packed buildings (approximately 8,000 houses per square-km) which do not lend themselves favorably to state-of-the-art satellite-based building segmentation methods due to, for example, occlusion, vegetation, shadows and low resolution. To address these challenges, we introduce a novel instance segmentation and counting approach for dense buildings. Our system first extracts a conservative set of tentative building center points using a deep network for jumpstarting a Segment Anything Model 2 (SAM2) module to produce an initial over-segmentation. Second, we use a graph neural network to refine the over-segmented regions into polygons representing accurate building masks. Experiments show that our approach achieves higher accuracy in instance segmentation and counting especially in challenging densely packed building areas in Brazil, Mexico, India, Pakistan, and Kenya, for instance.
Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models.
In this study, we investigate atmospheric new particle formation (NPF) across 65 d in the Bolivian central Andes at two locations: the mountaintop Chacaltaya station (CHC, 5.2 km above sea level) and an urban site in El Alto–La Paz (EAC), 19 km apart and at 1.1 km lower altitude. We classified the days into four categories based on the intensity of NPF, determined by the daily maximum concentration of 4–7 nm particles: (1) high at both sites, (2) medium at both, (3) high at EAC but low at CHC, and (4) low at both. These categories were then named after their emergent and most prominent characteristics: (1) Intense-NPF, (2) Polluted, (3) Volcanic, and (4) Cloudy. This classification was premised on the assumption that similar NPF intensities imply similar atmospheric processes. Our findings show significant differences across the categories in terms of particle size and volume, sulfuric acid concentration, aerosol compositions, pollution levels, meteorological conditions, and air mass origins. Specifically, intense NPF events (1) increased Aitken mode particle concentrations (14–100 nm) significantly on 28 % of the days when air masses passed over the Altiplano. At CHC, larger Aitken mode particle concentrations (40–100 nm) increased from 1.1 × 103 cm−3 (background) to 6.2 × 103 cm−3, and this is very likely linked to the ongoing NPF process. High pollution levels from urban emissions on 24 % of the days (2) were found to interrupt particle growth at CHC and diminish nucleation at EAC. Meanwhile, on 14 % of the days, high concentrations of sulfate and large particle volumes (3) were observed, correlating with significant influences from air masses originating from the actively degassing Sabancaya volcano and a depletion of positive 2–4 nm ions at CHC but not at EAC. During these days, reduced NPF intensity was observed at CHC but not at EAC. Lastly, on 34 % of the days, overcast conditions (4) were associated with low formation rates and air masses originating from the lowlands east of the stations. In all cases, event initiation (∼ 09:00 LT) generally occurred about half an hour earlier at CHC than at EAC and was likely modulated by the daily solar cycle. CHC at dawn is in an air mass representative of the regional residual layer with minimal local surface influence due to the barren landscape. As the day progresses, upslope winds bring in air masses affected by surface emissions from lower altitudes, which may include anthropogenic or biogenic sources. This influence likely develops gradually, eventually creating the right conditions for an NPF event to start. At EAC, the start of NPF was linked to the rapid growth of the boundary layer, which favored the entrainment of air masses from above. The study highlights the role of NPF in modifying atmospheric particles and underscores the varying impacts of urban versus mountain top environments on particle formation processes in the Andean region.
In recent years, generative networks have achieved high quality results in 3D-aware image synthesis. However, most prior approaches focus on outside-in generation of a single object or face, as opposed to full inside-looking-out scenes. Those that do generate scenes typically require depth/pose information, or do not provide camera positioning control. We introduce EpipolarGAN, an omnidirectional Generative Adversarial Network for interior scene synthesis that does not need depth information, yet allows for direct control over the camera viewpoint. Rather than conditioning on an input position, we directly resample the input features to simulate a change of perspective. To reinforce consistency between viewpoints, we introduce an epipolar loss term that employs feature matching along epipolar arcs in the feature-rich intermediate layers of the network. We validate our results with comparisons to recent methods, and we formulate a generative reconstruction metric to evaluate multi-view consistency.
We introduce University of Texas - GLObal Building heights for Urban Studies (UT-GLOBUS), a dataset providing building heights and urban canopy parameters (UCPs) for more than 1200 city or locales worldwide. UT-GLOBUS combines open-source spaceborne altimetry (ICESat-2 and GEDI) and coarse-resolution urban canopy elevation data with a machine-learning model to estimate building-level information. Validation using LiDAR data from six U.S. cities showed UT-GLOBUS-derived building heights had a root mean squared error (RMSE) of 9.1 meters. Validation of mean building heights within 1-km2 grid cells, including data from Hamburg and Sydney, resulted in an RMSE of 7.8 meters. Testing the UCPs in the urban Weather Research and Forecasting (WRF-Urban) model resulted in a significant improvement (55% in RMSE) in intra-urban air temperature representation compared to the existing table-based local climate zone approach in Houston, TX. Additionally, we demonstrated the dataset’s utility for simulating heat mitigation strategies and building energy consumption using WRF-Urban, with test cases in Chicago, IL, and Austin, TX. Street-scale mean radiant temperature simulations using the SOlar and LongWave Environmental Irradiance Geometry (SOLWEIG) model, incorporating UT-GLOBUS and LiDAR-derived building heights, confirmed the dataset’s effectiveness in modeling human thermal comfort in Baltimore, MD (daytime RMSE = 2.85°C). Thus, UT-GLOBUS can be used for modeling urban hazards with significant socioeconomic and biometeorological risks, enabling finer scale urban climate simulations and overcoming previous limitations due to the lack of building information.
Due to their importance in weather and climate assessments, there is significant interest to represent cities in numerical prediction models. However, getting high resolution multi-faceted data about a city has been a challenge. Further, even when the data were available the integration into a model is even more of a challenge due to the parametric needs, and the data volumes. Further, even if this is achieved, the cities themselves continually evolve rendering the data obsolete, thus necessitating a fast and repeatable data capture mechanism. We have shown that by using AI/graphics community advances we can create a seamless opportunity for high resolution models. Instead of assuming every physical and behavioral detail is sensed, a generative and procedural approach seeks to computationally infer a fully detailed 3D fit-for-purpose model of an urban space. We present a perspective building on recent success results of this generative approach applied to urban design and planning at different scales, for different components of the urban landscape, and related applications. The opportunities now possible with such a generative model for urban modeling open a wide range of opportunities as this becomes mainstream.
The generation of large-scale urban layouts has garnered substantial interest across various disciplines. Prior methods have utilized procedural generation requiring manual rule coding or deep learning needing abundant data. However, prior approaches have not considered the context-sensitive nature of urban layout generation. Our approach addresses this gap by leveraging a canonical graph representation for the entire city, which facilitates scalability and captures the multi-layer semantics inherent in urban layouts. We introduce a novel graph-based masked autoencoder (GMAE) for city-scale urban layout generation. The method encodes attributed buildings, city blocks, communities and cities into a unified graph structure, enabling self-supervised masked training for graph autoencoder. Additionally, we employ scheduled iterative sampling for 2.5D layout generation, prioritizing the generation of important city blocks and buildings. Our approach achieves good realism, semantic consistency, and correctness across the heterogeneous urban styles in 330 US cities. Codes and datasets are released at: https://github.com/Arking1995/COHO.
Accurate tree inventories are critical for urban forest management but challenging to obtain, as many urban trees are on private property (backyards, etc.) and are excluded from public inventories. Here, we examined the feasibility of tree species identification in a large heterogenous urban area (>850 km(2)) by using multi-temporal PlanetScope images (3.2 m resolution, multi-spectral) and inventory data from more than 20,000 ground observations within the urban forest of the Greater Chicago area. Our approach achieved an overall classification accuracy of 0.60 and 0.71 for 18 species and ten genera, respectively, but varied from moderate to high for certain species (0.59-0.92) and genera (0.61-0.91). In particular, we identified key host tree species (Fraxinus americana, F. pennsylvanica, and Acer saccharinum) for two damaging invasive insects, emerald ash borer (EAB, Agrilus planipennis) and Asian longhorn beetle (ALB, Anoplophora glabripennis), with over 0.80 accuracies. In addition, we demonstrated that including images from the autumn months (September-November), either for a single-season model or a combined multiple-season model, improved the identification accuracy of temperate deciduous trees. Further, the high classification accuracy of support vector machine (SVM) over random forest (RF) and neural network (NN) approaches suggests that future work might benefit from comparing multiple classification methods to select the approach that maximizes species classification accuracy. Our study demonstrated the potential for applying multi-temporal high-resolution images in urban tree classification, which can be used for urban forest management at a large spatial scale.
Generative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality.
Atmospheric new particle formation (NPF) and associated production of secondary particulate matter dominate aerosol particle number concentrations and submicron particle mass loadings in many environments globally. Our recent investigations show that atmospheric NPF produces a significant amount of particles on days when no clear NPF event has been observed/identified. Furthermore, it has been observed in different environments all around the world that growth rates of nucleation mode particles vary little, usually much less than the measured concentrations of condensable vapors. It has also been observed that the local clustering, which in many cases acts as a starting point of regional new particle formation (NPF), can be described with the formation of intermediate ions at the smallest sizes. These observations, together with a recently developed ranking method, lead us to propose a paradigm shift in atmospheric NPF investigations. In this opinion paper, we will summarize the traditional approach of describing atmospheric NPF and describe an alternative method, covering both particle formation and initial growth. The opportunities and remaining challenges offered by the new approach are discussed.
Text-to-video generation has been dominated by diffusion-based or autoregressive models. These novel models provide plausible versatility, but are criticized for improper physical motion, shading and illumination, camera motion, and temporal consistency. The film industry relies on manually-edited Computer-Generated Imagery (CGI) using 3D modeling software. Human-directed 3D synthetic videos address these shortcomings, but require tight collaboration between movie makers and 3D rendering experts. We introduce an automatic synthetic video generation pipeline based on Vision Large Language Model (VLM) agent collaborations. Given a language description of a video, multiple VLM agents direct various processes of the generation pipeline. They cooperate to create Blender scripts which render a video following the given description. Augmented with Blender-based movie making knowledge, the Director agent decomposes the text-based video description into sub-processes. For each sub-process, the Programmer agent produces Python-based Blender scripts based on function composing and API calling. The Reviewer agent, with knowledge of video reviewing, character motion coordinates, and intermediate screenshots, provides feedback to the Programmer agent. The Programmer agent iteratively improves scripts to yield the best video outcome. Our generated videos show better quality than commercial video generation models in five metrics on video quality and instruction-following performance. Our framework outperforms other approaches in a user study on quality, consistency, and rationality.
Object compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial manual effort from professionals, and is hardly scalable. Thus, with the recent advances in generative models, in this work, we propose a selfsupervised framework for object compositing by leveraging the power of conditional diffusion models. Our framework can hollistically address the object compositing task in a unified model, transforming the viewpoint, geometry, color and shadow of the generated object while requiring no manual labeling. To preserve the input object's characteristics, we introduce a content adaptor that helps to maintain categori-cal semantics and object appearance. A data augmentation method is further adopted to improve the fidelity of the generator. Our method outperforms relevant baselines in both realism and faithfulness of the synthesized result images in a user study on various real-world images.
We present a novel approach to perform instance segmentation and counting for densely packed self-similar trees using a top-view RGB image sequence. We propose a solution that leverages pixel content, shape, and self-occlusion. First, we perform an initial over-segmentation of the image sequence and aggregate structural characteristics into a contour graph with temporal information incorporated. Second, using a graph convolutional network and its inherent local messaging passing abilities, we merge adjacent tree crown patches into a final set of tree crowns. Per various studies and comparisons, our method is superior to all prior methods and results in high-accuracy instance segmentation and counting despite the trees being tightly packed. Finally, we provide various forest image sequence datasets suitable for subsequent benchmarking and evaluation captured at different altitudes and leaf conditions.
Urban and environmental researchers seek to obtain building features (e.g., building shapes, counts, and areas) at large scales. However, blurriness, occlusions, and noise from prevailing satellite images severely hinder the performance of image segmentation, super-resolution, or deep-learning-based translation networks. In this article, we combine globally available satellite images and spatial geometric feature datasets to create a generative modeling framework that enables obtaining significantly improved accuracy in per-building feature estimation and the generation of visually plausible building footprints. Our approach is a novel design that compensates for the degradation present in satellite images by using a novel deep network setup that includes segmentation, generative modeling, and adversarial learning for instance-level building features. Our method has proven its robustness through large-scale prototypical experiments covering heterogeneous scenarios from dense urban to sparse rural. Results show better quality over advanced segmentation networks for urban and environmental planning, and show promise for future continental-scale urban applications.
Modeling and designing urban building layouts is of significant interest in computer vision, computer graphics, and urban applications. A building layout consists of a set of buildings in city blocks defined by a network of roads. We observe that building layouts are discrete structures, consisting of multiple rows of buildings of various shapes, and are amenable to skeletonization for mapping arbitrary city block shapes to a canonical form. Hence, we propose a fully automatic approach to building layout generation using graph attention networks. Our method generates realistic urban layouts given arbitrary road networks, and enables conditional generation based on learned priors. Our results, including user study, demonstrate superior performance as compared to prior layout generation networks, support arbitrary city block and varying building shapes as demonstrated by generating layouts for 28 large cities.
Here we introduce a new method, termed “nanoparticle ranking analysis”, for characterizing new particle formation (NPF) from atmospheric observations. Using daily variations of the particle number concentration at sizes immediately above the continuous mode of molecular clusters, here in practice 2.5–5 nm (i.e. ΔN2.5−5), we can determine the occurrence probability and estimate the strength of atmospheric NPF events. After determining the value of ΔN2.5−5 for all the days during a period under consideration, the next step of the analysis is to rank the days based on this simple metric. The analysis is completed by grouping the days either into a number of percentile intervals based on their ranking or into a few modes in the distribution of log (ΔN2.5−5) values. Using 5 years (2018–2022) of data from the SMEAR II station in Hyytiälä, Finland, we found that the days with higher (lower) ranking values had, on average, both higher (lower) probability of NPF events and higher (lower) particle formation rates. The new method provides probabilistic information about the occurrence and intensity of NPF events and is expected to serve as a valuable tool to define the origin of newly formed particles at many types of environments that are affected by multiple sources of aerosol precursors.
An abundance of impervious surfacesImpervious surfaces like building roofs in densely populated cities make green roofs a suitable solution for urban heat islandUrban heat island (UHI) mitigation. Therefore, we employ random forest (RF) regression to predict the impact of green roofs on the surface UHI (SUHI) in Liege, Belgium. While there have been several studies identifying the impact of green roofs on UHIUrban heat island, fewer studies utilize a remote-sensing-based approach to measure impact on Land Surface TemperaturesLand surface temperatures (LST) that are used to estimate SUHI. Moreover, the RF algorithm, can provide useful insights. In this study, we use LSTLand surface temperatures obtained from Landsat-8 imagery and relate it to 2D and 3D morphological parameters that influence LST and UHI effects. Additionally, we utilise parameters that influence wind (e.g., frontal area index). We simulate the green roofs by assigning suitable values of normalised difference-vegetation index and built-up index to the buildings with flat roofs. Results suggest that green roofs decrease the average LSTLand surface temperatures.
Abstract Taking the examples of Hurricane Florence (2018) over the Carolinas and Hurricane Harvey (2017) over the Texas Gulf Coast, the study attempts to understand the performance of slab, single‐layer Urban Canopy Model (UCM), and Building Environment Parameterization (BEP) in simulating hurricane rainfall using the Weather Research and Forecasting (WRF) model. The WRF model simulations showed that for an intense, large‐scale event such as a hurricane, the model quantitative precipitation forecast over the urban domain was sensitive to the model urban physics. The spatial and temporal verification using the modified Kling‐Gupta efficiency and Method for Object based Diagnostic and Evaluation in Time Domain suggests that UCM performance is superior to the BEP scheme. Additionally, using the BEP urban physics scheme over UCM for landfalling hurricane rainfall simulations has helped simulate heavy rainfall hotspots.
We propose a framework to create projectively-correct and seam-free cube-map images using generative adversarial learning. Deep generation of cube-maps that contain the correct projection of the environment onto its faces is not straightforward as has been recognized in prior work. Our approach extends an existing framework, StyleGAN3, to produce cube-maps instead of planar images. In addition to reshaping the output, we include a cube-specific volumetric initialization component, a projective resampling component, and a modification of augmentation operations to the spherical domain. Our results demonstrate the network's generation capabilities trained on imagery from various 3D environments. Additionally, we show the power and quality of our GAN design in an inversion task, combined with navigation capabilities, to perform novel view synthesis.
Herein, we introduce a novel methodology to generate urban morphometric parameters that takes advantage of deep neural networks and inverse modeling. We take the example of Chicago, USA, where the Urban Canopy Parameters (UCPs) available from the National Urban Database and Access Portal Tool (NUDAPT) are used as input to the Weather Research and Forecasting (WRF) model. Next, the WRF simulations are carried out with Local Climate Zones (LCZs) as part of the World Urban Data Analysis and Portal Tools (WUDAPT) approach. Lastly, a third novel simulation, Digital Synthetic City (DSC), was undertaken where urban morphometry was generated using deep neural networks and inverse modeling, following which UCPs are re-calculated for the LCZs. The three experiments (NUDAPT, WUDAPT, and DSC) were compared against Mesowest observation stations. The results suggest that the introduction of LCZs improves the overall model simulation of urban air temperature. The DSC simulations yielded equal to or better results than the WUDAPT simulation. Furthermore, the change in the UCPs led to a notable difference in the simulated temperature gradients and wind speed within the urban region and the local convergence/divergence zones. These results provide the first successful implementation of the digital urban visualization dataset within an NWP system. This development now can lead the way for a more scalable and widespread ability to perform more accurate urban meteorological modeling and forecasting, especially in developing cities. Additionally, city planners will be able to generate synthetic cities and study their actual impact on the environment.
Manuel Menezes de Oliveira Neto合作论文数Instituto de Informatica, Universidade Federal do Rio Grande do Sul5
Scott Cohen合作论文数Adobe Systems4
Anselmo Lastra合作论文数Department of Computer Science4