Residential differentiation reflects the complex patterns by which social groups distribute themselves across urban spaces, fundamentally shaping social, economic, and spatial structures. This paper reviews the methodological development of geodemographic classification, tracing its evolution from early social area analysis and factorial ecology through to contemporary approaches. We critically evaluate this lineage of methods for quantifying residential patterns, and identifying persistent limitations in capturing the non-linear complexities of contemporary urban environments. Building on this review, we explore potential future directions involving learned representations of the social landscape, which may offer alternatives to traditional linear dimensionality reduction techniques. Drawing on recent empirical work applying deep learning to geodemographic classification, we consider how such approaches might address identified limitations while acknowledging that their advantages over established methods remain context-dependent and require further empirical validation. We emphasise that any adoption of these techniques must prioritise transparency and interpretability. The paper concludes by outlining potential directions for future research, including how learned representations might be integrated within existing geodemographic workflows.
This paper describes a new dataset release containing harmonised census tables from the 2021 and 2022 UK Censuses. The release is the first unified dataset covering all four UK nations at the smallest available geographic level: Output Areas in England, Wales, and Scotland, and Data Zones in Northern Ireland. The UK's three census agencies: ONS (England and Wales), NRS (Scotland), and NISRA (Northern Ireland) release their data separately, each with distinct variables, formats, and disclosure controls. Through a process of matching, standardisation, and aggregation, 190 comparable variables are produced. The dataset is made available as a series of topic tables indexed across all 239,023 of the UK's small-area geographies. By providing a standardised dataset, this work enables seamless UK-wide analyses, facilitating cross-national comparisons and supporting research and public policy development.
Financial precarity, the state of economic insecurity characterised by unpredictable employment and declining social protection significantly impacts cognitive functioning, emotional stability and social inclusion. This condition stems from multiple interconnected factors: poor quality and unpredictable work, unmanaged debt, insecure asset wealth and insufficient financial resource. Despite extensive research on financial precarity's individual impacts, its geographical distribution and associated social-spatial inequalities remain poorly understood. This paper addresses this gap by introducing a new geodemographic classification of financial precarity across Great Britain. Our classification system uses small-area measurements encompassing employment patterns, income levels, asset holdings, debt obligations, and lifestyle characteristics at the neighbourhood level. By mapping financial precarity at a fine spatial scale, this research reveals how economic vulnerability varies across different localities, highlighting the uneven geography of financial insecurity between rural and urban areas, city centres and peripheries, coastal and inland communities, and how the classification groups are interwoven to the variegated patterns in and around major urban areas. This small-area approach provides sufficient detail to identify spatial patterns while enabling comparisons between local areas, offering new insights into the geographic dimensions of economic precarity in contemporary Britain.
This paper evaluates the performance of a large language model (LLM) based semantic search tool relative to a traditional keyword-based search for data discovery. Using real-world search behaviour, we compare outputs from a bespoke semantic search system applied to UKRI data services with the Consumer Data Research Centre (CDRC) keyword search. Analysis is based on 131 of the most frequently used search terms extracted from CDRC search logs between December 2023 and October 2024. We assess differences in the volume, overlap, ranking, and relevance of returned datasets using descriptive statistics, qualitative inspection, and quantitative similarity measures, including exact dataset overlap, Jaccard similarity, and cosine similarity derived from BERT embeddings. Results show that the semantic search consistently returns a larger number of results than the keyword search and performs particularly well for place based, misspelled, obscure, or complex queries. While the semantic search does not capture all keyword based results, the datasets returned are overwhelmingly semantically similar, with high cosine similarity scores despite lower exact overlap. Rankings of the most relevant results differ substantially between tools, reflecting contrasting prioritisation strategies. Case studies demonstrate that the LLM based tool is robust to spelling errors, interprets geographic and contextual relevance effectively, and supports natural-language queries that keyword search fails to resolve. Overall, the findings suggest that LLM driven semantic search offers a substantial improvement for data discovery, complementing rather than fully replacing traditional keyword-based approaches.
Retail centres, including town centres, high streets and retail parks, are critical to the economic vitality and social cohesion of communities across the United Kingdom. Yet spatial definitions of these centres remain inconsistent, hampering comparative analysis and evidence-based planning. This paper presents a new open-source methodology for delineating retail centre boundaries at national scale. The approach integrates Foursquare Places point-of-interest data with Open- StreetMap land-use polygons, applying a comprehensive quality-filtering procedure before two parallel detection pathways: H3 hexagonal tessellation with network-based connected- component identification for traditional walkable centres, and brand-matching heuristics for car-oriented retail parks. Two further innovations refine the output: land-use-informed removal of false positives in non-retail employment areas, and the application of Leiden community detection to partition large urban agglomerations into functionally distinct sub-centres. The methodology delineates 9,477 retail centres across the United Kingdom, compared with 6,423 in the previous national dataset, with substantially improved coverage in Scotland and Northern Ireland and explicit, reproducible identification of retail parks. Validation through stakeholder workshops involving academic researchers, local authority planners, and commercial sector representatives confirms that the resulting boundaries align with planning practice and local knowledge. A case study of Liverpool’s night-time economy demonstrates the framework’s practical utility, showing how a consistent spatial unit enables integration of heterogeneous sources, including event listings, live-music gig data, and late-opening commercial units, to characterise urban evening activity. Built on openly available data, the framework is reproducible and supports ongoing monitoring of UK retail landscapes
The UK retail landscape has undergone a profound change in past decades with popular debates largely focusing on decline of the traditional retail spaces. This, predominantly driven by technological advancement and corresponding changes in consumer behaviour, was exacerbated by the Covid-19 pandemic and the subsequent disruption to supply chains and cost of living crisis. This study provides a comprehensive, data-driven descriptive analysis and new evidence on the transformation and economic performance of British retail centres over the five-year pre- and post-pandemic period (2019-2023), which is a crucial period offering a valuable perspective within three different periods: pre-pandemic, the Covid-19 pandemic and the initial post-pandemic 'recovery'. Using longitudinal retailer occupancy data, this study presents a picture of the British retail landscape that is far from uniform, and shows that the decline was predominantly driven by ongoing trends of digitalisation within retailing and services and exacerbated by the temporary closure of 'non-essential' shops during the pandemic. Our findings also provide empirical evidence that Covid-19, when combined with pre-existing trends, prompted further demise of many 'traditional' retailers on high streets, evidenced by increasing vacancies. On the contrary and importantly, we find several trends which are facilitating reorientation and growth in the traditional retail centres and have emerged in the past five years. These changes are conceptualised within existing frameworks of retail resilience to economic shocks in particular retail centres economic cycle and their evolutionary trajectories. The new evidence can be used to substantiate the wider debates on the economic performance of British retail centres and their regeneration in the 'new retail' post Covid-19 era.
Like many countries globally, the private rental sector in England and Wales contains some of the lowest quality and energy inefficient properties, despite being home to some of the most vulnerable households. We present a new data product that classifies small areas based on the energy (in)efficiency characteristics of private rental properties. Newly available Energy Performance Certificate (EPC) data enables us to analyse detailed energy and housing characteristics for 3.9 million private rentals (∼78.8% of total sector), the most comprehensive dataset of its kind, using k-means clustering. Demographic datasets allow us to explore wider socio-spatial inequalities, and uncertainties associated with granular - but at-times incomplete - EPC data. The classification can be used to evidence how inefficiency is spatially concentrated and fragmented, with a diverse range of energy and housing conditions shaping the everyday lives of tenants.
Transportation systems worldwide are facing numerous challenges, including congestion, environmental impacts, and safety concerns. This study used a systematic literature review to investigate how advanced technologies (e.g., IoT, AI, digital twins, and optimization methods) support smart transportation planning. Specifically, this study examines the interrelationships between transportation challenges, proposed solutions, and enabling technologies, providing insights into how these innovations support smart mobility initiatives. A systematic literature review, following PRISMA guidelines, identified 26 peer-reviewed articles published between 2013 and 2024, including studies that examined smart transportation technologies. To quantitatively assess relationships among key concepts, a Sentence BERT-based natural language processing approach was employed to compute alignment scores between transportation challenges, technological solutions, and implementation strategies. The findings highlight the fact that real-time data collection, predictive analytics, and digital twin simulations significantly enhance traffic flow, safety, and operational efficiency while mitigating environmental impacts. The analysis further reveals strong correlations between traffic congestion and public transit optimization, reinforcing the effectiveness of integrated, data-driven strategies. Additionally, IoT-based sensor networks and AI-driven decision-support systems are shown to play a critical role in sustainable urban mobility by enabling proactive congestion management, multimodal transportation planning, and emission reduction strategies. From a policy perspective, this study underscores the need for investment in urban-scale data infrastructures, the integration of digital twin modeling into long-term planning frameworks, and the alignment of optimization tools with public transit improvements to foster equitable and efficient mobility. These findings offer actionable recommendations for policymakers, engineers, and planners, guiding data-driven resource allocation and legislative strategies that support sustainable, adaptive, and technologically advanced transportation ecosystems.
This work provides a thorough Energy Deprivation Segmentation (EDS) for Great Britain, which aims to address the complex and varied aspects of energy poverty in different small regions. By proposing a reproducible analytical framework, we combine many data sources to provide a comprehensive segmentation that encompasses various dimensions such as energy efficiency, accessibility, demand and supply, housing conditions, and financial vulnerability. The results indicate notable disparities in energy deprivation based on social and spatial factors. We observed higher degrees of deprivation in the peripheral areas of major cities and suburbs in the northern regions of England, southern regions of Wales, and central regions of Scotland. The created EDS identifies six top-level Supergroups and 14 finer Groups and was validated internally and externally to confirm its robustness and applicability. This segmentation offers a more comprehensive insights into the characteristics and distribution of energy-deprived neighbourhoods than traditional measures. This research facilitates policymakers to design targeted strategies and resource allocation to combat specific vulnerabilities within communities and foster sustainable and equitable urban growth. Additionally, a practical tool is provided for monitoring and evaluating the effectiveness of policies aimed at reducing energy poverty.
The strategic placement of pedestrian crossings is critical for promoting road safety in urban environments. While previous research has used Geographic Information Systems (GIS) to identify underserved areas, existing approaches often fail to integrate empirical data with stakeholder expertise, or directly link road safety concerns with wider spatial equity considerations. This study develops a novel GIS-Multi Criteria Decision Analysis (GIS-MCDA) framework to evaluate the suitability of sites for new pedestrian crossings, incorporating four key criteria: existing provision of pedestrian infrastructure, road safety risks, proximity to urban amenities, and socio-demographic vulnerability. Applied to Liverpool City Region as a policy case study, the framework integrates advanced spatial analysis techniques with stakeholder informed weights, to assess the suitability of postcodes for new pedestrian crossings. Site suitability scores are generated which reflect both the considerations relevant to key stakeholders, as well as the distribution of controlled pedestrian crossings, road traffic collisions, urban amenities and census-derived vulnerability indicators. The results reveal significant spatial inequalities where, specific locations with high socio-demographic vulnerability also experience the greatest road safety risks and poorest access to pedestrian crossings, despite proximity to essential services. The framework successfully pinpoints priority locations where road safety and spatial inequality concerns compounded negatively, providing evidence-based guidance for data-driven, stakeholder-informed pedestrian infrastructure investment.
Objectives The talk will: (i) Describe the development of a large language model (LLM) powered semantic search tool for UKRI data catalogues, and (ii) Examine the concerns and opportunities of using this tool among researchers for data discovery. Methods A semantic search tool was developed integrating the data catalogues of Administrative Data Research UK, Consumer Data Research Centre, and UK Data Service. We used OpenAI’s vector embedding service to convert these metadata into embeddings, allowing natural language search to be used rather than keywords only. We assessed the acceptability and suitability of this tool using four focus groups. Participants were recruited across academic researchers, PhD researchers, data services staff, and local government / third sector analysts (n=36). Data collected from focus groups were analysed using thematic analysis. Results The key themes identified in focus groups were: (i) Current data discovery techniques are dependent on keyword strategies for searching (including the dominance of using Google). There is need to support training for using any LLM based resources. (ii) There was low trust of LLMs, especially in academic researchers. Participants were concerned that results may be erroneous. Being able to ‘explain’ why a search result was returned was viewed as valuable. (iii) Having a resource that collates all metadata in one place was powerful for helping researchers find data. This could be improved through leveraging the power of LLMs to summarise large quantities of information about datasets to make data discovery more efficient. Our talk will detail steps towards addressing these challenges. Conclusion Although Large Language Model’s can be useful for supporting federated data discovery among researchers, tools need to be developed that are responsible, trustworthy and open if researchers are going to use them.
This paper provides an analysis of working from home patterns in England using data from the 2021 Census to understand (1) how patterns of working from home (WFH) in England have shifted since the COVID‐19 pandemic and (2) whether human mobility indicators, specifically Google Community Mobility Reports, provide a reliable proxy for WFH patterns recorded by the 2021 Census, providing a formal evaluation of the reliability of such datasets, whose applications have grown exponentially over the COVID‐19 pandemic. We find that WFH patterns recorded by the 2021 Census were unique compared with previous UK censuses, reflecting an unprecedented increase likely caused by persistent changes to employment during the COVID‐19 pandemic, with a clear social gradient emerging across the country. We also find that Google mobility in ‘Residential’ and ‘Workplace’ settings provides a reliable measurement of the distribution of WFH populations across Local Authorities, with varying uncertainties for mobility indicators collected in different settings. These findings provide insights into the utility of such datasets to support population research in intercensal periods, where shifts may be occurring, but can be difficult to quantify empirically.
The UK residential sector is energy inefficient and has an overwhelming reliance on natural gas as a heating source. For the UK to meet its 2050 net zero obligations, the sector will need to go through a process of decarbonisation. Previous studies acknowledge the spatial disparities of household energy consumption, but have neglected how consumption varies over time. This paper advances such shortcomings via a sequence and clustering analysis to identify common gas consumption trajectories within neighbourhoods in England and Wales between 2010 and 2020. Four clusters are identified: “Very High to High Consumption”; “High to Medium Consumption”; “Medium to Low Consumption” and “Low to Very Low Consumption”. The clusters were contextualised using spatial datasets representing the socio-economic and built environment. Across all clusters, the proportion of energy inefficient dwellings were high, but there was a trend of high consumption associated with lower proportions of energy efficient dwellings. The results provide useful insight to policy makers and practitioners about where best to target electrification and retrofitting measures to facilitate a cleaner and more equitable residential sector. Policy targeting of areas with continual high gas consumption will accelerate the decarbonisation process, whilst targeting areas who continually under consume will likely enhance household health and well-being.
Observed regional variation in geotagged social media text is often attributed to dialects, where features in language are assumed to exhibit region-specific properties. While dialects are seen as a key component in defining the identity of regions, there are a multitude of other geographic properties that may be captured within natural language text. In our work, we consider locational mentions that are directly embedded within comments on the social media website Reddit, providing a range of associated semantic information, and enabling deeper representations between locations to be captured. Using a large corpus of geoparsed Reddit comments from UK-related local discussion subreddits, we first extract embedded semantic information using a large language model, aggregated into local authority districts, representing the semantic footprint of these regions. These footprints broadly exhibit spatial autocorrelation, with clusters that conform with the national borders of Wales and Scotland. London, Wales, and Scotland also demonstrate notably different semantic footprints compared with the rest of Great Britain.
Spatial inequality is a common urban phenomena in cities around the world, where stark contrasts in a variety of different social and economic outcomes paint a vivid picture of compound inequalities. Tackling these influences from a policy perspective remains challenging, as political economies often span multiple actors and municipal bodies, lacking effective policy instruments to challenge multiple forms of inequality at once. This paper provides a new data-driven perspective, which seeks to improve how policy is developed when trying to mitigate the impacts of compound inequality. Utilising a place-based approach, we present an evidence base which has been co-produced with policymakers, comprising composite spatial indicators and a city dashboard for Liverpool City Region. The assembled evidence base highlights clear patterns of compound inequality across the region, identifying places in greatest need of support. In the paper we discuss how this evidence base is now being used to distribute investment from the City Region Sustainable Transport Settlements, generating positive outcomes for people and places across the region. Finally, we conclude by reflecting on the benefits of building collaborative relationships between academics and policymakers, and the utility of our approach, which uses urban indicators and city dashboards, which we argue can secure a more equitable future for cities globally.
This paper explores cognitive place associations; conceptualised as a place-based mental model that derives subconscious links between geographic locations. Utilising a large corpus of online discussion data from the social media website Reddit, we experiment on the extraction of such geographic knowledge from unstructured text. First we construct a system to identify place names found in Reddit comments, disambiguating each to a set of coordinates where possible. Following this, we build a collective picture of cognitive place associations in the United Kingdom, linking locations that co-occur in user comments and evaluating the effect of distance on the strength of these associations. Exploring these geographies nationally, associations were shown to be typically weaker over greater distances. This distance decay is also highly regional, rural areas typically have greater levels of distance decay, particularly in Wales and Scotland. When comparing major cities across the UK, we observe distinct distance decay patterns, influenced primarily by proximity to other cities. This paper explores cognitive place associations; conceptualised as a place-based mental model that derives subconscious links between geographic locations. Utilising a large corpus of online discussion data from the social media website Reddit, we experiment on the extraction of such geographic knowledge from unstructured text. The strength of cognitive association is typically weaker over greater distances, but this distance decay is spatially heterogeneous.image
This paper outlines the creation of the London Output Area Classification (LOAC) from the 2021 Census, set within the broader context of geodemographic classification systems in the United Kingdom. The LOAC 2021 was developed in collaboration with the Greater London Authority (GLA) and offers an enhanced, statistically robust typology adept at capturing the unique spatial, socio-economic and built characteristics of London's residential neighbourhoods. The paper asserts the critical importance of nuanced, area-specific geodemographic classifications for urban areas with unique geography relative to the national extent.
Inadequate supply of transport infrastructure is often seen as a barrier to a sustainable future for cities globally. Such barriers often perpetuate significant inequalities in who can and who cannot benefit from sustainable transport opportunities, and as a result there is momentum for transformative urban planning to promote sustainable transportation equity. This study introduces a new set of two-dimensional indicators, merging elements of supply and demand, to identify barriers and imbalances in sustainable transport equity. The accessibility indicators, which are generated for bus, rail, and cycle infrastructure, consider the proximity of administrative areas to good quality transport infrastructure, as well as mode-specific demand, to clearly identify areas where the supply of infrastructure is inadequate to support local populations. We present a policy case study for Liverpool City Region, which demonstrates how these indicators can be used in an analytical framework to support transformative urban planning in long-term. In particular, the indicators reveal policy priority areas where demand for sustainable transport is greater than supply, as well as neighbourhoods where multiple transport inequalities are intersecting spatially, highlighting the need for specific types of infrastructure investment to promote sustainable transport equity (e.g. more frequent services, additional cycle paths). Our framework lays the foundations for improved decision-making in urban systems, through development of mode-specific sustainable transport indicators at small area levels, which harmonise elements of supply and demand for the first time.
Higher education is a key global market and considerable literature has focused on investigating the determinants of international student mobility (ISM). However, less is known about the extent to which the relative influence of these factors is moderated by local conditions and vary across origin countries. Drawing on a unique data set of undergraduate applications from the UK Colleges and Admissions Service, we analyse variations in the contextual determinants of ISM flows to the United Kingdom across countries of origin over a 10-year period (2009-2019). We run a suite of negative binomial gravity models to understand the key influences of ISM and uncover the spatial heterogeneity of these influences. Our findings reveal a nonlinear relationship between the level of development of origin countries and ISM flows. Although countries from higher development levels are more likely to send students to the United Kingdom, there appears to be a dip in applications at the mid-levels of development. Given the nonlinearity of this relationship, we seek to understand how countries across different levels of development respond to the typical factors that are seen to influence flows of international students. We also see substantial heterogeneity of the influence of different factors for origin countries, with some countries being influenced by employment opportunities and others by cultural and linguistic ties. However, this variation is not necessarily determined by the countries' level of development. Our findings have implications for policy makers, educators and researchers seeking to navigate and influence global student mobility trends. Our study highlights the need for tailored strategies to attract and retain international students from specific origin countries, recognising the multifaceted nature of ISM determinants.
This paper reviews and assesses the prospects for developing geographically enabled research ready data (RRD) with reference to current UK initiatives. Examples of projects for which such data have been provisioned are given.