‘Not in Education, Employment or Training’ (NEET) categorisation is widely used to monitor societal disengagement in young people, but its binary form obscures policy relevant variation (e.g., conflating short disruption with sustained disengagement). We used 26,513 linked administrative records to construct monthly participation sequences over 24 months post-compulsory schooling. Clustering techniques revealed six trajectories: three ‘stable’ (education, employment, training) and three ‘unstable’ trajectories. The unstable trajectories reflected persistent disengagement (‘High Risk’), intermittent disruption (‘Moderate Risk’), and administrative uncertainty (‘Low Contact’). Multinomial models showed that unstable trajectories were associated with vulnerability indicators (e.g., children’s social care involvement, deprivation, exclusions). Low Contact trajectories closely resembled High Risk profiles, indicating that administrative non contact is a meaningful risk signal. School suspensions were associated with a higher likelihood of Steady Employment relative to continued education, suggesting employment led reengagement can support young people with fractured school relationships. Training trajectories were notably less stable than education or employment, challenging assumptions that all participation is equivalent. We conclude: (i) binary NEET metrics mask heterogeneity in an unhelpful manner: (ii) trajectory based measures support prioritising prevention, targeting early support, and identifying persistent non contact as a participation/safeguarding concern rather than a measurement artefact.
Motorbikes in Hanoi cause congestion and have road safety and environmental impacts. A recent directive will progressively ban petrol (gasoline) motorbikes from inner urban districts by 2030. This paper describes statistical and spatial analyses of a travel survey of ~25,700 Hanoi residents to model the relationships between different socio-economic factors and support for such a ban. It applies a standard global binomial logistic regression which is extended to a spatial regression to quantify how these relationships vary spatially. A novel Generalized Additive Model (GAM) smooth framework was used to construct spatially varying coefficient (SVC) models. Critically, the approach undertakes model selection, and evaluates multiple models (~350,000) to determine which variables to include in spatial smooths. The best ranked model by AIC included spatial smooths for all but one of the predictor variables indicating the presence strong spatial dependencies in response to predictor relationships. The global model identified car ownership, car use and private property ownership as the strongest positive predictors of ban support, and distance from public transport as the dominant negative predictor. The SVC GAM provided a substantially better fit than the global model, capturing almost twice as much variation in the outcome. Three groups of response to predictor relationships were identified. The mapped coefficient surfaces indicate how the response to predictor relationships vary over space, highlight a transit rich inner core and where the relationship with ban support varied. These findings have direct policy implications and demonstrate the inferential value of using non-stationary approaches to examine such data.
We examined the variation in 16-18 year olds’ participation trajectories across secondary schools, after accounting for student characteristics. In England, close to one million young people are currently not in education, employment or training (NEET), despite substantial public investment in expanding training and employment opportunities. Research and policy have predominantly focused on identifying individual characteristics associated with disengagement. However, there is evidence to suggest that structural contexts also matter (Evans et al 2026). Moreover, binary measures of NEET obscure meaningful variation in the pathways young people follow, limiting understanding of how trajectories of participation are shaped. Analysis of linked administrative data for 24,635 young people in the North of England showed substantial school differences in the likelihood of the post-16 trajectories followed by the young people they served. School differences were greatest for students experiencing circumstances associated with increased risk of poor outcomes, with predicted probabilities of persistently disengaged or disrupted trajectories ranging from 8% to 71%, depending on the secondary school attended. These findings challenge monocausal student‑deficit accounts of youth disengagement and suggest that schools may represent important risk‑mitigating environments, particularly for young people experiencing greater disadvantage. A critical priority for future research and policy is testing approaches that enable schools to help vulnerable young people make successful post-16 transitions.
Tourism plays a vital role in economic development, but fluctuations in demand, particularly during crises such as the COVID-19 pandemic, highlight the limitations of traditional data sources for timely and detailed monitoring. This study investigates the potential of social media data (Twitter now X) as a complementary data source for estimating domestic and international tourist arrivals. A three-step pipeline was developed to identify tourists from other non-tourist social media users: (1) users were classified using a Large Language Model (LLM) to exclude non-individual accounts; (2) home location was inferred through geocoding of self-reported textual information; and (3) tourists were classified based on the United Nations World Tourism Organization (UNWTO) criteria of spatial origin, trip duration, and travel purpose. Social media derived estimates of tourist arrivals in Bali between 2019 and 2021 were compared with official statistics and showed strong correlations, with coefficients of 0.999 for international tourist arrivals and 0.896 for domestic arrivals. As well as monthly counts, the approach provided fine-grained insights into tourist behaviour, including inferred home origins, destination hotspots, and daily and hourly arrival patterns - dimensions that are typically absent from conventional datasets. While the consistency of social media data declined during the COVID-19 pandemic, the findings demonstrate that social media can be used to generate timely, flexible, and spatially detailed observations that complement official statistics. A number of discussion points and areas of further work are identified to refine classification methods, to enhance scalability and the applicability of the approach in diverse contexts.
Following the emergence of powerful generative large language models (LLMs), there has been a flurry of interest in the use of LLMs to control agents in agent-based models. Proponents argue that using the information about humans and human behaviour contained within an LLM could lead to the creation of agents who exhibit more complex and nuanced behaviour than those whose actions are driven by traditional behavioural frameworks. This paper begins to explore the use of a specific concept that underpins LLMs; that of embeddings. An embedding is a vector-based numerical representation of a piece of text that captures aspects of its meaning and context. We hypothesise that conceptualising agents’ characteristics through embeddings, rather than with discrete state variables, may offer a more nuanced and expressive foundation for representing agent characteristics and behaviours. We demonstrate the potential of this approach by recreating the Schelling residential segregation model using rich text descriptions of household agents and converting these to embeddings as a means of defining agents. The results show how agents can self-organise into more diverse and emergent clusters than is possible when they are defined with a small number of discrete attributes. This offers a path toward more realistic, high-dimensional representations of agent heterogeneity.
The COVID-19 pandemic significantly disrupted global tourism, impacting visitor perceptions and destination reputations. To quantify the impacts for the Indonesian tourism centre Bali, this study examines 3.98 million georeferenced Twitter posts from 2019 to 2021, analysing temporal shifts, spatial distribution, and thematic associations across key tourism topics. Using a RoBERTa-based sentiment classification model, we identify variations between domestic and international tourists, revealing a dominant but fluctuating positive sentiment that declined among international visitors during the pandemic. At key phases of the crisis, tourist sentiment became more aligned, reflecting heightened uncertainty. Spatially, Kuta, Denpasar, and Ubud remained central hubs for both positive and negative sentiment, with persistent concerns regarding waste management and traffic congestion. Topic-based analysis highlighted strong positive sentiment toward ‘Attractions’ and ‘General’ tourism aspects, while ‘Accessibility’ and ‘Amenities’ received more criticism. Additionally, sentiment trends varied by nationality, with UK and US tourists expressing more negative sentiment, while Singaporean and Filipino visitors remained consistently positive. These insights offer valuable guidance for Bali’s post-pandemic recovery, emphasising sustainable tourism management, targeted marketing, and infrastructure improvements.
Agent-based models are flexible tools that allow modellers to capture heterogeneity in agent attributes, characteristics, and behaviours. In this paper, heterogeneity is defined as agent granularity, referring to the level of detail used to describe agent attributes, behaviours, interaction processes, and decision-making rules. However, the increased complexity associated with greater levels of heterogeneity, and hence more parameters, can make the already challenging process of model calibration even more difficult. While modellers recognise the importance of calibration, the issue of uniquely determining model input based on a given output, known as parameter identification, is often overlooked. A central point of this study is that identifiability crucially depends on the outcomes or summary statistics chosen for calibration: even a well-specified model may become empirically uninformative if the selected statistics are not sufficiently sensitive to parameter variation. This paper argues that one significant impact of increasing heterogeneity in an agent-based model is the parameter identification problem, where the effects of model inputs cannot be uniquely distinguished in model outputs. To address this issue, the paper presents a comparative study of homogeneous and heterogeneous scenarios in agent-based models. Using a simple contagion case study model and approximate Bayesian computation for calibration, the study demonstrates that introducing heterogeneity reduces the accuracy of parameter calibration compared to the homogeneous case. This decline in accuracy is attributed to the difficulty in isolating the effects of the additional parameters introduced by heterogeneity. Rather than proposing computational fixes, the paper situates these findings within the broader methodological debate between KISS ("Keep It Simple, Stupid") and KIDS ("Keep It Descriptive, Stupid") strategies, highlighting how the trade-off between descriptive realism and tractability directly shapes the reliability of inference from ABMs.
Digital twins are used across many industries to enable better decision making. However, while policy makers at all levels (including city, national and supranational scales) have expressed a desire to integrate digital twins into their workflows, this adoption has been slow to materialise. In this paper, we discuss the key issues associated with policy digital twins, and the ways in which they differ from, and are similar to, their counterparts in other areas. We describe how multi-level agent based modelling can be used within policy digital twins to include the effects of human behaviours on outcomes; an aspect that is often largely overlooked. We also describe how digital twins can be designed for policy use cases, and present as a case study the design of a policy digital twin incorporating multi-level agent based modelling to aid a UK city council (local authority) in delivering energy transition policy. After describing both the design method used and the resultant digital twin, we discuss the effectiveness of both, as well as how the ways in which different contexts might shape the future architecture of the digital twin.
The impact of student-level risk factors on not in education, employment or training (NEET) rates (e.g. low attainment, absenteeism and socio-economic disadvantage) are well-documented. However, there is limited research on how school characteristics influence NEET rates, despite recognition that inclusive school environments can have a positive effect on education outcomes. In this work, we hypothesized that proxy measures of 'inclusivity' would affect sustained post-16 engagement, and we tested this hypothesis using 3 years of administrative data for secondary schools in England while controlling for known student and school local area risks. Our results indicate that schools with lower suspension rates, higher student progress ('Progress 8') and onsite post-16 provision had lower rates of students becoming NEET. Single-sex and faith schools also exhibited reduced NEET rates. These results suggest school culture and inclusivity play an important role in shaping student trajectories. The proportion of 16-17 year olds in England who become NEET has remained stubbornly high for more than a decade, putting these individuals at risk of long-term adverse outcomes. Our results suggest that policies promoting inclusive school environments, supportive disciplinary practices and clear post-16 pathways may help increase sustained engagement in education and training.
This study provides the first national-level longitudinal analysis directly linking mobility data to police-recorded crime across local authorities in England and Wales, while controlling for stable structural differences between places. Its purpose is to examine how changes in human mobility are associated with changes in crime, taking advantage of the disruption in mobility caused by the COVID-19 pandemic. Negative binomial regression models were applied across 328 local authorities from March 2020 to October 2022, with area fixed effects, to assess how mobility across different activity spaces relates to police-recorded crime while accounting for unobserved, time-invariant local characteristics. The exploratory and modelling analyses showed that the mobility-crime relationship varied considerably by offence type and across local authorities. Increased residential mobility was consistently associated with lower levels of property, violent, and public order crimes, while mobility around retail, grocery, transit, and workplace areas was generally associated with higher crime once stable local differences were controlled. Sensitivity analyses suggested the results are robust to plausible levels of unobserved confounding. The impact of mobility on crime during and after the pandemic in England and Wales was not uniform but was shaped by local conditions and time-invariant characteristics. The findings underline the need for place-specific approaches in crime prevention and suggest that mobility-related interventions should be tailored to the structural and social context of each area.
Social media data offers urban planners insights into human activities in urban green spaces (UGSs). While recent methods like text-based word frequency analysis provide new perspectives on UGS, they are often lack stationary and non-continuous in nature. This limits their ability to capture the complexity and diversity of UGS use. This study conducts a structural topic model (STM) analysis of geo-referenced Tweets posted in London to investigate the dynamics of UGS-related topics before-, during- and after the COVID-19 outbreaks. Additionally, an approach of inverse distance weighting (IDW) was used to investigate the spatial patterns of topics probabilities. The results found that there were seven main topics categories expressed in UGS over study periods. Specifically, the increasing trends in topics proportions were found for the topics Nature engagement and Dog walking, indicating that these activities became increasingly popular during the pandemic. However, the topic Social events showed a decline in topic proportion, which might be the results of restriction measures such as practicing social distance. This study further discussed the potential factors that affecting the dynamics of these topics in spatial and temporal patterns. The results can potentially support future UGS planning and management especially during a time of crisis.
The distribution of wealth is central to economic, social, and environmental dynamics. The release of high-frequency distributional data and the rapid pace of the complex global economy makes 'real-time' predictions about the distribution of wealth and income increasingly relevant. For instance, during the COVID-19 pandemic in spring 2020, the stock markets experienced a crash followed by a surge within a brief period, evidently reshaping the wealth distribution in the US. Yet economic data, when first released, can be uncertain and need to be readjusted - again specifically so during crisis moments like the pandemic when information about household consumption and business returns is patchy and drastically different from "business- as-usual". Our motivation here is to develop one way of overcoming the problem of uncertain 'real-time' data and enable economic simulation methods, such as agent-based models, to accurately predict in 'real-time' when combined with newly released data. Therefore, we tested two distinct, parsimonious agent-based models of wealth distribution, calibrated with US data from 1990 to 2022, in conjunction with data assimilation. Data assimilation is essentially applied control theory - a set of algorithms aiming to improve model predictions by integrating 'real-time' observational data into a simulation. The algorithm we employed is the Ensemble Kalman Filter (EnKF), which performs well in the context of computationally expensive problems. Our findings reveal that while the base models already align well historically, the EnKF enables a superior fit to the data.
Space Syntax comprises a set of techniques that emphasize both material and immaterial characteristics of urban space. However, its inherent lack of a time component and land-use variables limits its effectiveness in investigating crime dynamics at the micro-urban scale. On the other hand, ABMs are simulations of interacting agents capable of perceiving their environment, influencing each other, and making autonomous decisions without central control, whose behaviour depends on time and environmental components. Nonetheless, ABMs face challenges in terms of the required amount of input data and model transparency, as well as regards the matter of validation. Despite being based on different modelling approaches – embodying the top–down versus bottom–up contraposition – both Space Syntax and ABMs can qualitatively and quantitatively analyse crucial components for crime analysis, such as flows, copresence, and visibility, deemed key aspects underlying the crime opportunity concept. This paper establishes a bridge between Space Syntax and ABM through the study of crime in urban settings, scrutinising the fundamental metrics that can significantly contribute to the identification of high-risk locations – specifically, pedestrian flow and visibility. It presents two new models, respectively, using Space Syntax methods and ABM, applied to two separate case studies in the historic city centre of Pisa, Italy. It has three objectives: first, to compare the diverse outcomes the distinct approaches provide in terms of people flow and visibility estimation in the built environment; second, to discuss their potential in forecasting risky areas; and, lastly, to propose an integration between Space Syntax metrics for movement within ABMs as parameters to guide agents’ movement. In essence, this study proposes a methodology oriented towards creating a model capable of simulating the relations between people’s behaviour, urban configuration, environmental conditions, and crime distribution, thereby representing a useful decision-support tool for crime prevention purposes and for broadly exploring and modelling pedestrian behaviour.
Glasgow, a major city in the UK, faces ongoing challenges with traffic congestion and high levels of air pollution. In response, Glasgow City Council introduced Scotland’s first Low Emission Zone (LEZ) in June 2023. This study investigates the potential impacts of the LEZ on traffic flow, emissions and public health outcomes using an agent-based modelling approach. Using the SUMO traffic simulator, the work in-progress model simulates the complexity of urban traffic dynamics and the associated emissions. Preliminary results highlight emissions fluctuations influenced by traffic congestion and vehicle stops, with significant emissions reductions observed under LEZ-compliant scenarios. The study also identifies critical challenges, including traffic collisions due to inadequate infrastructure and limitations of real-time traffic data. These findings highlight the need for calibrated traffic patterns and infrastructure improvements to maximise the effectiveness of the current LEZ simulation.
An agent-based model is a form of complex systems model that is capable of simulating how the micro-level behavior of individual system entities contributes to macro-level system outcomes. Researchers draw on theory and evidence to identify the key elements of a given system and specify behaviors of agents that simulate the individual entities of that system—be they cells, animals, or people. The model is then used to run simulations in which agents interact with one another and the resulting outcomes are observed. These models enable researchers to explore proposed causal explanations of real-world outcomes, experiment with the impacts that potential interventions might have on system behavior, or generate counterfactual scenarios against which real-world events can be compared. In this review, we discuss the application of agent-based modeling within the field of criminology as well as key challenges and future directions for research.
Disasters fundamentally alter human mobility patterns, yet traditional modelling approaches fail to capture the complex temporal dynamics and interdependencies that characterise these disruptions. This paper introduces the Disaster-aware Hawkes process, a novel temporal point process framework specifically designed to model human mobility during and after disasters. Our approach extends standard self-exciting Hawkes processes through five key innovations: a regime-switching baseline intensity function, category-specific excitation parameters, heterogeneous recovery rates across location categories, and post-disaster bounce-back parameters. We apply our model to a comprehensive mobile device dataset from Auckland, New Zealand during Cyclone Gabrielle in February 2023, comprising 5.85 million mobility records from 111,539 devices. The Disaster-aware Hawkes process achieves a 25.68% overall improvement in prediction accuracy compared to standard approaches, with an 80.00% improvement during the disaster period. Beyond enhanced prediction, our model enables novel analyses of cascade effects, revealing how disruptions propagate through mobility networks and identifying critical dependencies invisible to traditional methods. We demonstrate that optimised temporal distribution of recovery interventions can improve system outcomes by 45.8% compared to conventional simultaneous deployment strategies. These findings provide valuable insights for disaster preparedness, response, and recovery planning, while the methodological framework offers a powerful new approach for analysing the complex dynamics of human mobility under disruption.
This paper investigates the optimal locations of health service facilities during the Hajj, a major temporary city event in Mina, Saudi Arabia, attended by millions of pilgrims. Given the logistical challenges and historical accident risks during this dense gathering, effective placement of health facilities is crucial for ensuring pilgrim safety and accessibility to services. The study first employs location-allocation models (LAM) within a network GIS framework to determine optimal facility locations based on both static and dynamic population distributions of pilgrims throughout the day. These models facilitate the strategic placement of services closer to pilgrim activities, potentially enhancing service accessibility and reducing travel times for medical assistance. Additionally, agent-based modelling (ABM) complements the LAM by simulating crowd movements and interactions, helping identify high-risk areas for congestion and accidents. This dynamic approach offers insights into crowd behavior under various scenarios, including different times of day and road closure impacts, thereby supporting more responsive urban planning and crowd management strategies. Recommendations for policy include the use of portable health facilities and the strategic placement of services to accommodate shifting crowd densities.
The paper discusses the relevance of the latest advances in data science and artificial intelligence for urban systems research. It has a particular focus on the importance of recent innovations in the context of ‘wicked’ urban problems which continue to confront decision-makers within practical policy settings. It is argued that the latest advances in AI such as large language models offer the potential for transformative research, but only if properly specified within the unique and distinctive context of geographical space. The idea of a digital twin requires careful articulation to support the management of expectations and appropriate alignment within a social setting. At the end of the day, AI is not a panacea for the problems of cities, nor is it a substitute for imaginative policy design or interventions through consensus and good government. However in a world which is characterised by vast riches of data alongside enormous complexity of process, the investment in new tools and methods is a social and intellectual imperative in driving human understanding to new levels.
As the world rapidly urbanises and cities become larger and more complex, understanding pedestrian dynamics is paramount. New data sources, particularly those that measure pedestrian counts (i.e. ‘footfall’), offer potential as a means of better understanding the fundamental spatio-temporal structures that characterise aggregate pedestrian behaviour. However, footfall data are often complex and influenced by a wide range of social, spatial and temporal factors, which complicates interpretation. This paper applies principal component analysis (PCA) to hourly pedestrian count data from Melbourne, Australia, to extract the key temporal signatures that underpin observed urban footfall patterns. PCA can reduce the dimensionality of noisy pedestrian flow data, revealing dominant activity patterns such as weekday commuting cycles and weekend leisure activities. By subsequently analysing pedestrian volumes through the lens of these components, we start to expose the underlying types of pedestrian activities that characterise different neighbourhoods. In addition, we can distinguish multiple overlapping activity patterns within a single location, identifying changes in urban functionality and detecting shifts in mobility trends. The impacts of external shocks, such as the COVID-19 pandemic, are particularly stark. These findings shed light on the intricacies of urban mobility and suggest that there is value in the use of PCA as a means to better understand urban dynamics.