Traffic volumes are rising globally, creating a growing need for accurate and scalable data collection to address mobility challenges and enhance transport systems. Yet, traditional methods remain costly and time-consuming despite advances in automated monitoring. This study explores the feasibility of using open webcam data in combination with the state-of-the-art object detection model YOLOv8 out-of-the-box for road user monitoring. Publicly accessible webcam imagery presents challenges such as high variability in image quality, road user occlusion, and environmental factors like poor visibility due to weather conditions. To assess their potential for traffic monitoring, we utilize open webcam data from Germany to evaluate the performance of YOLOv8′s model variants, testing 110 parameter combinations with a manually labeled reference dataset. Among the tested out-of-the-box model variants, YOLOv8x achieved the highest performance, with an F1-score of 0.75. This optimized model was applied to about 500,000 open webcam scenes to monitor the change of road users before and during the COVID-19 pandemic. The analysis revealed a 9.5% overall reduction in road users volume, with motorized road users declining significantly while bicycles increased by 25.2%. This reflects mobility patterns observed during the COVID-19 pandemic, where restrictions led to a significant shift towards cycling as an alternative mode of transport. The results are plausible as they mirror broader trends in active mobility observed in various urban contexts. Our findings demonstrate the potential of leveraging open webcam data and pre-trained object detection models for scalable, cost-effective transport monitoring.
Reliable and up-to-date sub-national data on the gross domestic product (GDP) data remain scarce in many regions across the globe. This limits the ability to analyze spatial economic disparities and to design evidence-based development policies. In this study, we investigate whether high-resolution GDP estimates can be derived from multisource remote sensing and auxiliary inputs using a deep learning fusion framework. We combine optical day-and night-time satellite imagery, and auxiliary geospatial layers with several backbone architectures and fusion strategies to assess the robustness of multimodal learning for economic prediction. Across extensive experiments in Brazil, we evaluate multiple encoders (ResNet-18, EfficientNet-B3, and SwinV2-T), alternative fusion layers (concatenation, attention pooling, graph-based multilayer perceptron, and mixture-of-experts), and diverse input modality combinations. Our best-performing model, an EfficientNet-B3 encoder with concatenation fusion using all available input modalities, achieves an R-2 value of 0.87 for GDP prediction at a spatial resolution of 5 km & times; 5 km, demonstrating that our multimodal approach effectively captures the complex relationships between spatial patterns and economic activity. These findings highlight the potential of multimodal remote sensing to complement traditional statistical sources by providing spatially consistent, high-resolution representations of economic activity.
Few research studies have focused on the nature of the relationship between nighttime lights and electricity consumption at subnational levels in South America, a region with heterogeneous geography and urbanization levels and complex socioeconomic dynamics. This study shows that it is possible to estimate, for Bolivia, a wide range of indicators of electricity consumption at the municipality level and two temporal scales using features derived from nighttime lights and other spatial data sources, in combination with readily available municipality characteristics. The prediction errors for annual electricity consumption range between 13% MAPE for average residential consumption and 59% MAPE for average commercial consumption. Similar accuracies are obtained when predicting monthly values. For both annual and monthly electricity consumption, we highlight the variation in estimation accuracy for various municipality subsets and show that prediction can be significantly improved when selecting municipalities based on population size, energy poverty, or levels of sustainable development.
The concept of the 15-minute city has gained a lot of popularity in recent years. Numerous studies have looked at a wide variety of cities, providing insights into the extent to which the 15-min-city has been achieved, or the specific x-minuteness that has been achieved so far. In a context of competition between cities in terms of quality of life and sustainability, statements about their performance in terms of x-minuteness dominate. However, to date there has been no research on the quantitative measurement and comparison of x-minuteness for different urban structure types. Urban structure types are defined by specific combinations of characteristics of the built environment. In this study, the pedestrian x-minuteness is computed for the fifteen most populated cities in Germany and compared across six characteristic urban structure types. The results underline that pedestrian accessibility varies greatly based on different urban structural types and that it is important to look more closely at the spatial composition of cities at the neighbourhood level. This spatially differentiated information helps to target improvements more efficiently and, thus, to increase quality of life more equitably.
Indian cities have recently emerged as some of the locations with the worst air pollution on Earth. However, the sources of these increases and potential mitigation paths are complex, involving an interplay between urbanization, economic growth, energy systems and environmental management. Here, we analyse an extensive set of satellite and ground-level air quality monitoring data sets to create a detailed examination of geographical and temporal patterns of $$N{O}_{2}$$ pollution across all districts and most large cities in India. For the period 2005–19, we find that urbanization has an inverse relationship with $$N{O}_{2}$$ increase, but at the same time, the most urbanized districts also have the highest $$N{O}_{2}$$ burden. Districts at intermediate levels of urbanization ( $$30-40 \%$$ ) show both high levels of $$N{O}_{2}$$ and the largest increases in $$N{O}_{2}$$ pollution. Our analysis also reveals that the region of the Indo-Gangetic Plain (IGP) is of particular concern both in terms of high absolute levels of $$N{O}_{2}$$ and relative $$N{O}_{2}$$ increase over the period from 2005 to 2019. Studying $$P{M}_{10}$$ levels across 106 cities, we find that they exceed national air quality standards in over 80% of cases, with the problem also being particularly acute in IGP’s urban areas. Scaling analysis reveals that $$N{O}_{2}$$ concentrations increase super-linearly with population in cities in the IGP, in contrast to the sublinear scaling observed for other cities. This implies that larger cities in the IGP are characterized by both higher $$N{O}_{2}$$ concentrations and increased per capita exposure. Moreover, recent mitigation efforts in the largest cities could signal the flow of technologies and strategies down the urban hierarchy and impact the rising levels of air pollution across the nation and over time. For now, the impetus of development and poverty alleviation implies that environmental challenges are likely to worsen, unless immediate and sustained pollution mitigation can address the severe air quality problem in India, especially across the Indo-Gangetic Plain.
Automated delineation of settlements at the level of single buildings is advancing on global scale due to developments in image processing techniques. However, in morphologically complex poverty areas (e.g., slums or informal settlements) automated image classification reaches limits. This gap of systematic, accurate geodata hinders the impetus of ‘better data for better decisions’ by the UN Sustainable Development Goals. To bridge this gap, we present a dataset of >320,000 building footprints derived from satellite imagery in 44 poverty areas across the globe. We use ‘Manual Visual Image Interpretation’ for consistent delineations of rooftops proxying building footprints. The method has been applied to 1. classify different morphological types representing poor living environments based on spatial features; 2. document multitemporal dynamics at different spatial scales; and, 3. assess interpreter-related uncertainties across interpreters. We release these data along with interpretation guidelines and validation. We provide a robust empirical basis for the global systematization and comparative analysis of settlement forms proxying poverty, while offering training and validation data for automated image classification.
This paper explores the potential of public webcams as a source of data for transport research. Eight different open-source object detection models were tested on three publicly accessible webcams located in the city of Brunswick, Germany. Fifteen images at different lighting conditions (bright light, dusk, and night) were selected from each webcam and manually labelled with regard to the following six categories: cars, persons, bicycles, trucks, trams, and buses. The manual counts in these six categories were then compared to the number of counts found by the object detection models. The results show that public webcams constitute a useful source of data for transport research. In bright light conditions, applying out-of-the-box object detection models can yield reliable counts of cars or persons in public squares, streets, and junctions. However, the detection of cars and persons was not reliably accurate at dusk or night. Thus, different object detection models might have to be used to generate accurate counts in different lighting conditions. Furthermore, the object detection models worked less well for identifying trams, buses, bicycles, and trucks. Hence fine-tuning and adapting the models to the specific webcams might be needed to achieve satisfactory results for these four types of traffic participants.
Cities are called upon to become more sustainable. One urban planning concept to improve local accessibility that has been implemented by a number of cities around the world is the x-minute city. Yet, there is no clear consensus on which physical aspects of the urban fabric influence the x-minuteness of an area or city. With this in mind, this study shows that exploring one aspect, urban structure types, has a crucial impact on x-minuteness by comparing different areas in Berlin, Germany. The results show that the x-minuteness of an area is significantly dependent on the urban structural type. Furthermore, we show that different urban structural types not only affect the x-minuteness but are also directly related to the spatial distribution of amenities.
Urban trees have proven an effective adaptation and mitigation strategy to attenuate the negative consequences of ongoing urbanization and climate change. However, area-wide data on trees in cities is still largely missing. Furthermore, ecosystem services of urban trees have only been quantified at the individual tree level for few selected trees within the urban forest to date. This study presents an integrative approach based on remote sensing for detection and characterization of urban trees, which serves as input for the modeling of ecosystem services using the process-based tree growth model "CityTree". In addition, the remote sensing-based parameter setting is compared to in-situ measurements to assess the plausibility of the ecosystem services estimation. The results show that tree dimensions from remote sensing generally agree well with in-situ measurements. Furthermore, the classification of tree genera based on multi-temporal very-high resolution (VHR) satellite data achieved good accuracies around 70%. Finally, this study demonstrates the effectiveness of remote sensing-based individual tree parameters for the estimation of ecosystem services, while the comparison with in-situ parameters confirmed the plausibility of the results.
Slums are densely populated urban areas characterized by substandard housing and squalor. These areas often lack basic infrastructure and services, making them challenging to manage and improve. Mapping slums on a large scale is particularly difficult due to their complex, dynamic, and non-uniform nature and the scarcity of available data. This study employs deep learning and uncertainty estimation to map slum areas in 55 diverse cities. We tackle the issues of scarce labels, noisy labels, and class imbalance to produce reliable slum probability maps. Applied to a large dataset from the Global South, our method approximates probability values to each image, yielding a more granular view of slum morphology within each city. We achieve this by combining Monte Carlo sampling through test-time augmentation, test-time dropout on overlapping image tiles, and an ensemble of four CNNs. The resulting probability maps not only quantify prediction uncertainty but also reveal morphological patterns of slum settlements. Our findings underscore the diversity and complexity of slums, contributing to a more complete understanding of these environments. A key result is our large-scale workflow, which, despite limited data, groups predictions by probability into multiple slum categories defined by their morphological traits. This approach offers a substantial improvement over traditional binary slum classification methods that focus solely on typical slum morphologies.
Pollination is essential for maintaining biodiversity and ensuring food security, and in Europe it is primarily mediated by four insect orders (Coleoptera, Diptera, Hymenoptera, Lepidoptera). However, traditional monitoring methods are costly and time consuming. Although recent automation efforts have focused on butterflies and bees, flies, a diverse and ecologically important group of pollinators, have received comparatively little attention, likely due to the challenges posed by their subtle morphological differences. In this study, we investigate the application of Convolutional Neural Networks (CNNs) for classifying 15 European pollinating fly families and quantifying the associated classification uncertainty. In curating our dataset, we ensured that the images of Diptera captured diverse visual characteristics relevant for classification, including wing morphology and general body habitus. We evaluated the performance of three CNNs, ResNet18, MobileNetV3, and EfficientNetB4 and estimated the prediction confidence using Monte Carlo methods, combining test-time augmentation and dropout to approximate both aleatoric and epistemic uncertainty. We demonstrate the effectiveness of these models in accurately distinguishing fly families. We achieved an overall accuracy of up to 95.61%, with a mean relative increase in accuracy of 5.58% when comparing uncropped to cropped images. Furthermore, cropping images to the Diptera bounding boxes not only improved classification performance across all models but also increased mean prediction confidence by 8.56%, effectively reducing misclassifications among families. This approach represents a significant advance in automated pollinator monitoring and has promising implications for both scientific research and practical applications.
The physical forms of cities emerge from the interplay of diverse processes shaped by various factors, including political, cultural, economic, and geographic influences. As such, this physical aspect, the urban fabric is deeply heterogeneous at multiple levels, from the intra-urban to the global scales. Although peering into these two scales provided the field of urban morphology great insights, the combination of both scales has, to the best of the authors knowledge, never been investigated, mainly because of lack of data. Yet, this combination of such scales could enable the understanding of global and local processes of homogenization or specification of the urban fabric and the way they embed themselves in nowadays urbanization. The recent evolutions in data quality, coverage, comprehensiveness and consistency makes such cross-scaled investigations now possible. Previous work proposed a universal typology of intra-urban patterns relying on a global classification of intra-urban morphology. Based on these results, this study aims to localize the distinct intra-urban patterns across the globe to characterize their geographical distributions. By categorizing these geographical distributions into six main modes, ranging from the most local to the most global, we assess for each type of intra-urban patterns their global spread. This allows to quantify the status quo on the homogeneity or heterogeneity of the global urban fabric. We find that although close to half of the global urban fabric is composed of very widely spread patterns, a non-neglectable number of patterns exist only in very specific regions of the globe. We thus show empirically that in its current status quo, the global urban fabric leans toward a global homogeneity, yet at the same time, local heterogeneities are persistent on a worldwide scale. This informs us about the dissemination of urban planning practices and paradigms and enables us to critically ponder on their driving forces.
Air pollution is a severe threat to urban residents worldwide. Growing cities and an increasing number of urban dwellers make urban air quality an important issue for the majority of the world's population. While from economic and ecologic perspectives, extensive literature suggests that higher concentration of people in cities make urban areas more resource-efficient and environmentally sustainable, there is still unresolved ambiguity in the evaluation whether the same applies to urban air quality. Related case studies come to different findings, assuming larger cities to be either cleaner or more polluted than smaller cities in terms of air pollution; however, usually they consider only single countries which hinders generalizable answers to this question. Reasons for the variety of findings can be identified in the underlying data for air quality and in the varying spatial delineation of urban boundaries. In this study, we present a global analysis of urban air quality for more than 10,000 cities using scaling laws. Scaling laws define a relationship based upon a power law between the population size of a city and a certain characteristic of the city, e.g., in this study its level of air pollution. We rely on a satellite-derived globally homogenous data set for nitrogen dioxide (NO2) which is one of the major air pollutants and a proxy for air quality. Results reveal that globally, NO2 levels for cities scale linearly; however, certain regions and countries show strong superlinear and sublinear scaling behavior, indicating a strong regional dependency. We argue that one reason for this may lie in national policies for cleaner air in cities. Besides a homogenous data set for air quality, especially a harmonized definition for urban boundaries worldwide has proven to be of high importance for globally consistent and comparable results.
Accessibility to public transport is a fundamental component of connecting individuals to urban services. Guided by the UN Habitat Sustainable Development Goal 11.2, which aims to ensure accessible, safe, affordable, and sustainable transport systems for all, our study focuses specifically on accessibility as a key dimension of achieving this goal and its implications for social and spatial equity. In this study, we employ the walking distance indicator proposed by the responsible working group of UN Habitat to calculate accessibility to public transport. Because underlying population data are an essential parameter for the indicator, we compare three distinct population datasets - cadaster-based population data, remote sensing-based population data, and a global dataset - to investigate spatial variations in accessibility across the city of Medellín, Colombia. Furthermore, we examine the impact of both formal public transport and the local semiformal minibus system (paratransit), analyzing differences across formal and informal settlement types of the city, as well as the influence of socio-economic factors. Our findings suggest that remote sensing based population data can serve as a valuable data source, albeit with limitations for global population data. Particularly, our results highlight the significance of the semiformal local minibus system in enhancing accessibility to public transport, despite ongoing expansions of the metro system by responsible authorities, which have led to considerable improvements in accessibility. Notably, we observe that residents with lower socio-economic status and those living in informal settlements experience longer walking distances to public transport stops, highlighting spatial and socio-economic disparities in accessibility. Overall, our study underscores the complex interplay between transport infrastructure, socio-economic factors, and urban development, highlighting the need for targeted interventions to address spatial and socio-economic disparities in public transport accessibility.
In informal settlements, accurate population data, such as the number of people per building, is often lacking, outdated, or incomplete, yet such information is crucial for assessing the population at risk in hazard-prone areas, such as parts of Medellín, Colombia. This study presents a methodology to estimate the population at risk from landslides in informal areas by combining deep learning-based building detection on aerial remote sensing images with a population estimation. A Mask R-CNN instance segmentation model is trained to detect individual buildings in orthophotos of informal settlements from 2019, using official cadastral data from 2018. Population estimates are made based on the size of the detected buildings. When applied to informal settlements within landslide-prone regions, our methodology identifies 10,180 more buildings and 28,575 more people compared to official cadastral data, revealing an underestimation of people at risk in these areas. This approach demonstrates the value of combining instance segmentation with population estimation to enhance risk assessment in data-limited settings.
Pollinating insects provide essential ecosystem services, and using time-lapse photography to automate their observation could improve monitoring efficiency. Computer vision models, trained on clear citizen science photos, can detect insects in similar images with high accuracy, but their performance in images taken using time-lapse photography is unknown. We evaluated the generalisation of three lightweight YOLO detectors (YOLOv5-nano, YOLOv5-small, YOLOv7-tiny), previously trained on citizen science images, for detecting ~ 1,300 flower-visiting arthropod individuals in nearly 24,000 time-lapse images captured with a fixed smartphone setup. These field images featured unseen backgrounds and smaller arthropods than the training data. YOLOv5-small, the model with the highest number of trainable parameters, performed best, localising 91.21% of Hymenoptera and 80.69% of Diptera individuals. However, classification recall was lower (80.45% and 66.90%, respectively), partly due to Syrphidae mimicking Hymenoptera and the challenge of detecting smaller, blurrier flower visitors. This study reveals both the potential and limitations of such models for real-world automated monitoring, suggesting they work well for larger and sharply visible pollinators but need improvement for smaller, less sharp cases.
The physical dimension of cities and its spatial patterns play a crucial role in shaping society and urban dynamics. Understanding the complexity of urban systems requires a detailed assessment of their physical structure. Urban geography has long focused on framing typologies to represent common patterns in the urban fabric using various methodologies. However, only recent advancements in computational methods and global land cover data have enabled to comprehensively identify typologies of urban patterns at the city scale through new unsupervised approaches. Nevertheless, typologies of finer-grained patterns at intra-urban scale have not yet been explored comprehensively at a global level. In this paper, building upon these advances, we explore the intra-urban patterns of more than 1500 cities across the globe. We rely on a Local Climate Zone land cover classification to represent the multidimensional variabilities of intra-urban morphology. Adapting a deep learning based unsupervised clustering approach, we find a typology of 138 intra-urban patterns. Analyzing the results of this data-driven approach, we prove that each pattern identified is unique, i.e. statistically different, in its composition and configuration. With this study summarizing the global diversity of the urban fabric, we reveal that any city of the world can be described as a specific assemblage of a fraction of these 138 universal patterns. These universal patterns reveal a predominance at a global scale of built-up forms of low density in the intra-urban fabric.
Motorized traffic often causes road noise directly in front of our homes and windows. Yet long-term exposure to noise impact life's quality and can potentially cause negative effects on human health. Furthermore, social and behavioral effects have been measured. To protect people's health and well-being from such noise, the European Noise Directive (END, 2002/49/EC) obliges countries to produce strategic noise maps every five years for large agglomerations and along major roads, which are then used for noise action planning. Besides that, the official noise maps are a valuable data source for environmental exposure analyses. However, the END has some limitations. The definition of urban agglomerations is vague, different input parameterizations lead to data inconsistencies across administrative units, undefined post processing methods introduce geometric artifacts, and topological errors incompliant to the common Simple Features Implementation Specification hinder working with the published geodata. The aim of this article is to provide practical insights for end-users and stipulate for concise regulations. Moreover, we highlight that these variations limit the comparability of maps in environmental impact assessments. We compile 84 separate noise assessments in Germany reported according to the END to review shape and structure of the geographic data. Graphical representations are used to show in particular how vertices are connected to polygons in noise contour maps and that these geometric alterations effect the eventual statistics on exposed population shares. We aggregate spatial metrics to assess the reported data's spatial properties in an automatic manner, e.g. when receiving data in future mapping rounds. Along with our quality assessment, a nation-wide dataset on road traffic noise was produced. Depicting the yearly averaged noise level indicator Lden, which integrates exposure at day, evening and night, for 2017, it serves as common ground for environmental health analyses. The examination of different raster to polygon conversion implementations is fundamental to other geodata managers outside the domain of noise mapping, as well.
The European Noise Directive mandates the mapping of noise - high, continuous sound pressure levels considered to be a major health threat. However, the strictest rulesets apply to specific regions only and the majority of residential areas are unmapped. Transfer learning was deployed to close spatial data gaps between the official, strategic road traffic noise maps. The three most suitable hyperparameter configurations achieved weighted Kappa values (a measure of ordinal agreement) ranging between 0.889 and 0.956 during repeated cross-validation. The best model achieved an overall classification accuracy of 90.7 % when tested against held-out samples. 7.8 % of predictions exhibited minor deviations within +/- 5 dB(A). The model was subsequently deployed to predict road traffic noise across Germany at 10 x 10 Meter resolution for 2017. The results suggest a total of 13.1 million people exposed to yearly averaged road traffic noise (Lden) above 55 dB(A) and stress need for improved noise policies.