Urban environments generate rich temporal patterns of activity across diverse venues such as restaurants, bars, and coffee shops. Popularity dynamics over time provide valuable insight into daily routines, collective habits, and cultural differences in urban life. In this study, we compile a large-scale dataset of venue popularity data from cities in Brazil (BR) and the United States (U.S.) and evaluate a wide range of time series classification methods. Our objective is to assess how different approaches capture temporal signatures of urban activity and to examine their robustness across geographic and cultural contexts. The results show that dictionary-based and shapelet methods (e.g., WEASEL-D, RDST, TDE) consistently outperform deep learning approaches, achieving peak F1-scores above 0.73. Incorporating metadata and automated feature extraction (TSFEL) further improves classification accuracy, confirming the value of enriched feature spaces. However, cross-country evaluation reveals substantial challenges in model transferability, with most classifiers suffering severe performance degradation when applied across national contexts. Only one method (Signature) demonstrates stable generalization in both directions. These findings demonstrate that urban popularity patterns contain identifiable temporal signatures that are both classifiable and culturally specific. The study highlights the strengths of dictionary and shapelet-based methods, the benefits of feature engineering, and the limitations of deep learning for this task. Beyond methodological contributions, the results offer practical guidance for urban analytics, recommendation systems, and smart-city applications, where understanding and predicting population behavior over time is essential for effective service delivery and planning.
User-generated content from location-based platforms provides valuable insights into urban behavior, but the lack of explicit demographic information limits analyses of social representativeness. In particular, understanding gender differences in the use of urban space remains challenging due to the absence of structured user attributes. In this work, we investigate the use of natural language processing techniques to infer binary gender labels from textual reviews in Google Places. We evaluate two transformer-based approaches: a fine-tuned BERT classifier and a BERT-based model augmented with linguistic features for gender classification from review text. Experiments conducted on a large-scale dataset of place reviews show that the augmented BERT model achieves high performance, reaching an average F1-score of 0.95. Beyond predictive performance, we explore how inferred gender proxy labels can support urban representativeness analysis. Using New York City as a case study, we analyze the spatial distribution of gender imbalance across ZIP codes and assess the extent to which these patterns align with an external benchmark (Foursquare). These findings highlight both the potential and the limitations of using inferred demographic proxy attributes to study urban representativeness. Our results demonstrate that contextual language models can support demographic inference in location-based social data, enabling new perspectives on urban behavior while raising important considerations regarding bias, uncertainty, and representativeness.
Large Language Models (LLMs) are increasingly used as proxies for human perception in urban analysis, yet it remains unclear whether persona prompting produces meaningful and reproducible behavioral diversity. We investigate whether distinct personas influence urban sentiment judgments generated by multimodal LLMs. Using a factorial set of personas spanning gender, economic status, political orientation, and personality, we instantiate multiple agents per persona to evaluate urban scene images from the PerceptSent dataset and assess both within-persona consistency and cross-persona variation. Results show strong convergence among agents sharing a persona, indicating stable and reproducible behavior. However, cross-persona differentiation is limited: economic status and personality induce statistically detectable but practically modest variation, while gender shows no measurable effect and political orientation only negligible impact. Agents also exhibit an extremity bias, collapsing intermediate sentiment categories common in human annotations. As a result, performance remains strong on coarse-grained polarity tasks but degrades as sentiment resolution increases, suggesting that simple label-based persona prompting does not capture fine-grained perceptual judgments. To isolate the contribution of persona conditioning, we additionally evaluate the same model without personas. Surprisingly, the no-persona model sometimes matches or exceeds persona-conditioned agreement with human labels across all task variants, suggesting that simple label-based persona prompting may add limited annotation value in this setting.
Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated against empirical mobility patterns. We present CityBehavEx, an interactive LLM-assisted urban simulation platform that scales to city-size populations, exposes agent behavior for inspection, supports empirical validation, and generates mobility patterns that better match real-world spatial, temporal, and semantic distributions. Instead of invoking large language models for every agent action, CityBehavEx combines established human mobility models with fine-tuned cross-encoders that estimate semantic alignment between agent profiles, schedules, and activity transitions. This design enables large-scale simulations, as demonstrated in a case study of 100,000 agents over 75 days in under one hour on a single consumer GPU. The platform allows users to define simulation regions, launch experiments, inspect trajectories and activity traces, debug unrealistic behaviors, and validate generated routines against real-world mobility, time-use, and semantic metrics.
This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further reveals persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.
We present CityHood, an interactive and explainable recommendation system that suggests cities and neighborhoods based on users' areas of interest. The system models user interests leveraging large-scale Google Places reviews enriched with geographic, socio-demographic, political, and cultural indicators. It provides personalized recommendations at city (Core-Based Statistical Areas - CBSAs) and neighborhood (ZIP code) levels, supported by an explainable technique (LIME) and natural-language explanations. Users can explore recommendations based on their stated preferences and inspect the reasoning behind each suggestion through a visual interface. The demo illustrates how spatial similarity, cultural alignment, and interest understanding can be used to make travel recommendations transparent and engaging. This work bridges gaps in location-based recommendation by combining a kind of interest modeling, multi-scale analysis, and explainability in a user-facing system.
Understanding how visual content communicates sentiment is critical in an era where online interaction is increasingly dominated by this kind of media on social platforms. However, this remains a challenging problem, as sentiment perception is closely tied to complex, scene-level semantics. In this paper, we propose an original framework, MLLMsent, to investigate the sentiment reasoning capabilities of Multimodal Large Language Models (MLLMs) through three perspectives: (1) using those MLLMs for direct sentiment classification from images; (2) associating them with pre-trained LLMs for sentiment analysis on automatically generated image descriptions; and (3) fine-tuning the LLMs on sentiment-labeled image descriptions. Experiments on a recent and established benchmark demonstrate that our proposal, particularly the fine-tuned approach, achieves state-of-the-art results outperforming Lexicon-, CNN-, and Transformer-based baselines by up to 30.9
This study investigates methods using a global data source, Google Places, to identify culturally similar urban areas without relying on difficult-to-access data like user preferences shown through checkins. We propose and assess a simple method requiring only information about place types and their frequency in the studied areas, and a more advanced method that enhances venue categories using Scenes Theory it helps us understand the cultural significance of everyday urban life. We tested our methods in 14 cities worldwide and all US states. The results suggest that a straightforward approach based on category frequencies can highlight major cultural differences. However, the Scenes Theory-based method provides a better understanding of cultural nuances, as the ones supported by survey data.
Location-Based Social Networks (LBSNs) provide a rich foundation for modeling urban behavior through iNETs (Interest Networks), which capture how user interests are distributed throughout urban spaces. This study compares iNETs across platforms (Google Places and Foursquare) and spatial granularities, showing that coarser levels reveal more consistent cross-platform patterns, while finer granularities expose subtle, platform-specific behaviors. Our analysis finds that, in general, user interest is primarily shaped by geographic proximity and venue similarity, while socioeconomic and political contexts play a lesser role. Building on these insights, we develop a multi-level, explainable recommendation system that predicts high-interest urban regions for different user types. The model adapts to behavior profiles – such as explorers, who are driven by proximity, and returners, who prefer familiar venues – and provides natural-language explanations using explainable AI (XAI) techniques. To support our approach, we introduce h3-cities, a tool for multi-scale spatial analysis, and release a public demo for interactively exploring personalized urban recommendations. Our findings contribute to urban mobility research by providing scalable, context-aware, and interpretable recommendation systems.
This thesis presents an automatic, generic framework for extracting urban perceptions from Location-Based Social Network (LBSN) data. The framework is organized into five key layers: Data Collection, Preprocessing and Embeddings, Model Training, Knowledge Extraction, and Applications. By leveraging deep learning techniques, including advanced sentence embedding methods, the framework captures both lexical and semantic nuances in textual data, thereby efficiently extracting user perceptions of urban environments. This approach eliminates the need for labor-intensive field surveys and manual data extraction, allowing scalable real-time analysis. We validated the framework by applying it to selected urban areas in Chicago, New York City, and London, demonstrating its effectiveness in uncovering valuable insights about urban perceptions. Furthermore, a comparative evaluation using a public dataset derived from volunteers’ perceptions in a controlled experiment revealed a high level of agreement between the two sets of results. As a proof-of-concept, we introduce Real-Estate Urban Perceptions (REAL-UP), an innovative tool designed to enhance the real estate marketplace. REAL-UP provides interactive 2D maps that integrate traditional real-estate data (e.g., rent prices and property types) with enriched information on neighborhood emotions, sentiments, and brief narrative reviews generated by a Large Language Model (LLM) based on LBSN messages.
Location-Based Social Networks (LBSNs) are valuable for understanding urban behavior and providing useful data on user preferences. Modeling their data into graphs like interest networks (iNETs) offers important insights for urban area recommendations, mobility forecasting, and public policy development. This study uses check-ins and venue reviews to compare the iNETs resulting from two distinct LBSNs, Foursquare and Google Places. Although these two LBSNs differ in nature, with data varying in regularity and purpose, their resulting iNETs reveal similar urban behavior patterns. When analyzing the impact of socioeconomic, political, and geographic factors on iNET edges — each edge representing users' interests in a pair of regions — only geographic factors showed a significant influence. When studying the granularity of area sizes to model iNETs, we highlight important trade-offs between larger and smaller sizes. Additionally, we propose a methodology to identify clusters of geographically neighboring areas where user interest is strongest, which can be advantageous for understanding urban space usage.
Location-Based Social Networks (LBSNs) can help model users’ interests in urban areas in several ways. In the present work, we focus on Interest Networks (iNETs), which result from modeling LBSN data into graphs. The present study provides insights into which areas are frequently visited together by getting data from two distinct LBSNs, Foursquare and Google Places. Although the studied LBSNs differ in nature, with data varying in regularity and purpose, both modeled iNETs revealed similar urban behavior patterns and were likewise impacted by socioeconomic and geographic factors. Also, we discuss the development of a tool to empower urban studies and the by-products of this research.
Finding the best short- or long-term accommodation is troublesome in unknown areas. Current tools provided by the real-estate market offer valuable information regarding the property, such as price, photos, and descriptions of the space; however, this market has little explored other relevant information regarding the surrounding area, such as what is nearby and users' subjective perception of the property's area. To address this gap, we propose REAL-UP, an interactive tool designed to enrich real-estate marketplaces. In addition to information commonly provided by such applications, e.g., rent price, REAL-UP also provides subjective neighborhood information based on Location-Based Social Networks (LBSNs) messages. This novel tool helps to represent complex users' subjective perceptions of urban areas, which could ease the process of finding the best accommodation.
Detecting keywords in texts is a task of paramount importance for many text mining applications. Graph-based techniques have been commonly used to automatically find the key concepts in texts. However, the integration of valuable information provided by embeddings to enrich the graph structure has not been widely used. In this context, this paper aims to address the following question: can the quality of extracted keywords from a co-occurrence network be enhanced by integrating embeddings to enrich the network structure? In the adopted model, texts are represented as co-occurrence networks, where nodes are words and edges are established either by contextual or semantical similarity. Two embedding approaches were used: Word2vec and Bidirectional Encoder Representations from Transformers (BERT). The results indicate that using virtual edges can effectively enhance the discriminative capacity of co-occurrence networks. The best performance was achieved by incorporating a limited proportion of virtual (embedding) edges. A comparison of the structural and dynamical network metrics demonstrated that the degree, PageRank, and accessibility metrics exhibited superior performance in the proposed model.
O conhecimento a respeito das características dos diferentes grupos culturais que existem no mundo e a identificação de similaridades culturais entre suas respectivas áreas de ocupação podem trazer diversos benefícios econômicos e sociais, como a recomendação de locais sob critérios culturais. Pesquisas referentes ao estudo dessas diferentes culturas são realizadas, em grande parte, de maneira tradicional, as quais são caras e não escalam. Dessa forma, este trabalho consiste em obter características relevantes de áreas urbanas utilizando dados geolocalizados de fontes da web, e aplicar uma metodologia que enriquece esses dados obtidos para a geração de uma assinatura cultural de áreas urbanas. Em uma aplicação prática da proposta, o resultado se mostra muito coerente, separando os bairros de Curitiba em clusters com características culturais distintas.
This study aims to propose an approach for spatiotemporal integration of bus transit, which enables users to change bus lines by paying a single fare. This could increase bus transit efficiency and, consequently, help to make this mode of transportation more attractive. Usually, this strategy is allowed for a few hours in a non-restricted area; thus, certain walking distance areas behave like "virtual terminals." For that, two data-driven algorithms are proposed in this work. First, a new algorithm for detecting itineraries based on bus GPS data and the bus stop location. The proposed algorithm's results show that 90 of the database detected valid itineraries by excluding invalid markings and adding times at missing bus stops through temporal interpolation. Second, this study proposes a bus stop clustering algorithm to define suitable areas for these virtual terminals where it would be possible to make bus transfers outside the physical terminals. Using real-world origin-destination trips, the bus network, including clusters, can reduce traveled distances by up to 50 making twice as many connections on average.
This work supports the problematization of the high rate of traffic accidents due to the increased use of automobiles for utility or transport in Brazil. Every year, the lives of approximately 1.3 million people are interrupted due to traffic accidents worldwide. In this regard, this study proposes the use of data mining as an approach to analyze and explore datasets of traffic accidents that occurred on Brazilian highways between the years 2017 and 2022, as provided by the Federal Highway Police. Additionally, vehicle price data were included, allowing for a more comprehensive analysis that also considers the financial value of the vehicles. The goal is to assess the predictive capability of classification models regarding the severity of accidents, focusing on vehicle characteristics and environmental factors. By applying classification algorithms and machine learning explainability techniques, we acquired relevant knowledge regarding the studied data, contributing to understanding and preventing accidents. As a result, the attributes related to vehicle characteristics had a more positive impact on the predictive capability of the models when compared to the attributes describing the environment and other variables.
In this Ph.D. thesis, we proposed an automatic and generic framework composed of 5 major layers (Data Collection, Preprocessing and Embeddings, Models Training, Knowledge Extraction, and Applications) to extract the user’s urban perception from Location-Based Social Network (LBSN) data. The framework employs advanced deep learning algorithms, including sentence embeddings, to capture lexical and semantic relationships in textual data and, thus, effectively extract content related to urban perceptions from LBSNs. Moreover, this framework circumvents the need for labor-intensive field surveys or manual extraction processes, enabling scalable and real-time analysis of urban perception. Studying some urban areas from Chicago, New York City, and London, we demonstrate the framework’s effectiveness in extracting valuable insights related to urban perceptions from LBSN data. Furthermore, we conducted a comparative evaluation using a public dataset derived from volunteers’ perceptions in a controlled experiment, where it was possible to observe that both results yielded a very similar level of agreement. Finally, we introduce a novel tool called Real-Estate Urban Perceptions (REAL-UP), which aims to enhance the real-estate marketplace as a proof-of-concept for our work. REAL-UP provides rich knowledge regarding urban areas in the form of interactive 2D maps, more specifically, the emotion and sentiment perceived, and a short review generated by a Large Language Model (LLM) based on LBSN messages, for every city’s neighborhood, in addition to information commonly provided by such applications, as rent price, property type, and so on.
The synergy of edge computing and Machine Learning (ML) holds immense potential for revolutionizing Internet of Things (IoT) applications, particularly in scenarios characterized by high-speed, continuous data generation. Offline ML algorithms struggle with streaming data as they rely on static datasets for model construction. In contrast, Online Machine Learning (OML) adapts to changing environments by training the model with each new observation in real-time. However, developing OML algorithms introduces complexities such as bias and variance considerations, making the selection of suitable estimators challenging. In this challenging landscape, ensemble learning emerges as a promising approach, offering a strategic framework to navigate the bias-variance tradeoff and enhance prediction accuracy by amalgamating outputs from diverse ML models. This paper introduces a novel ensemble method tailored for edge computing environments, designed to efficiently operate on resource-constrained devices while accommodating various online learning scenarios. The primary objective is to enhance predictive accuracy at the edge, thereby empowering IoT applications with robust decision-making capabilities. Our study addresses the critical challenges of ML in resource-constrained edge computing environments, offering practical insights for enhancing predictive accuracy and scalability in IoT applications. To validate our ensemble’s efficacy, we conducted comprehensive experimental evaluations leveraging both synthetic and real-world datasets. The results indicate that our ensemble surpassed state-of-the-art data stream algorithms and ensemble regressors across a range of regression metrics, underlining its superior predictive prowess. Furthermore, we scrutinized the ensemble’s performance within the realm of auto-scaling for Virtual Network Function (VNF)-based applications situated at the network’s edge, thereby elucidating its applicability and scalability in real-world scenarios.
Urban transportation planning in densely populated areas is a problem in constant need of efficient solutions. Graphs can represent urban street networks and be used to train algorithms, enriching decisions with information learned from structural and topological data of cities. Relational Fusion Networks (RFNs) are Graph Neural Networks specifically tailored for learning on road networks. This work explores the use of RFNs in estimating free-flow travel times and includes experiments on relevant cities from all continents. Results demonstrate the significance of fusion functions and city characteristics in both the learning process of RFNs on regression tasks and the capacity to extrapolate acquired knowledge to different cities.