
Abstract Background: Problem gambling causes harm, but operational identification often relies on heuristic thresholds or sparse manual reviews. Routine logs are heavy-tailed and temporal, complicating monitoring. Methods: We analyzed de-identified records from one online operator across transactions, bets, sessions and payments. Thirty-day windows were leakage-audited and summarized with exceedance-frequency and magnitude features. Window embeddings were learned with a hierarchical conditional variational autoencoder using a responsible-gambling-proxy teacher and a student fine-tuned on sparse analyst labels from the training split. Backlog-aware label inference augmented training labels under missing-not-at-random assessment. Regularized Gaussian hidden Markov models produced dynamic regimes and a three-class operational proxy definition, evaluated on held-out labels and capacity-constrained alerting analyses. Results: Agreement was moderate and stream-dependent. Balanced accuracy ranged from 0.380 for transactions to 0.624 for payments, although payments requires caution because test support was skewed ([514, 22, 1386] low/medium/high) and high-risk argmax recall was 0.025. Bets had the strongest macro-averaged precision-recall harmonic mean (0.543); sessions had the highest high-risk argmax recall (0.757). Label augmentation effects were heterogeneous. In a common-support top-ten-per-week comparison, model detection exceeded responsible-gambling-ranked and random queues in every stream (55-78% versus 43-68% and 31-54%). Median alert-to-observed-assessment intervals were 65-210 days. Conclusions: The framework supports auditable, capacity-constrained monitoring, with cautious interpretation under selective review and class imbalance.
Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult to scale. We investigate whether Multimodal Large Language Models (MLLMs) can assess perceived urban safety from street-view imagery while accounting for the observer-dependent nature of perception. Using Place Pulse 2.0, we evaluate four open and proprietary MLLMs across 56 cities under a Neutral prompt and socio-demographic personas defined by gender, age, and race or ethnicity. We also analyse the keywords generated to justify each classification. All four models display comparable zero-shot capability, with city-macro F1 scores of 65–69
With the world population increasing, new living areas must be created to accommodate the increased number of people. The designs of most existing cities have a series of well-documented limitations. For example, some cities have serious traffic congestion problems, while others are unsafe. There is a lot of research on methods of improving parts of existing cities, and in many cases the improvement of a city is a constant and ongoing process. However, this research is limited by a few key factors: it is applied to the designs of current cities; it typically focuses on specific, localised issues; and it typically aims to fulfill a single objective, rather than examining the trade-off involved when analysing all aspects of a city at once. Recently, cities such as The Line (Neom) project in Saudi Arabia, and Brasilia and Canberra before it, have been designed and built from scratch. These projects often make bold claims that their designs are “optimal” in some regard. We introduce a novel “living metric” to quantify the degree of hierarchical living structure (a property associated with organically developed cities) in a city’s road network, and validate it by applying it to six real cities. As a proof of concept, we also demonstrate that this and two further metrics - population density (space efficiency) and mean circuity (travel efficiency) - can be applied to hypothetical city designs. To do this, we introduce CitySprout, a novel procedural road network generator, and generate cities of four distinct morphologies, or “grammars”. We study current examples of planned cities, and discuss the fundamental limitations of attempting to optimise cities based only on the underlying road network. We then demonstrate that different urban morphologies perform better at different metrics, highlighting that optimising a road network is inherently metric-dependent and therefore cannot converge on the “optimal” urban form. We find that a Grid structure maximises population density (under a set of clearly defined assumptions about population distribution); an Organic structure exhibits the most hierarchical “living structure”; and a Line structure minimises the required detours for cross-city travel. These differences highlight the trade-off between competing urban objectives.
This paper presents the Politus Dataset, a large-scale political public opinion dataset from Turkey that leverages social media data and state-of-the-art artificial intelligence methodologies to advance the field of computational social science. By integrating cutting-edge deep learning, natural language processing, and large language models with robust social scientific conceptualization, we identify key political indicators, such as ideologies, demographic attributes, and political leader job approval rates across both time and location. Our approach addresses critical measurement and representation errors in public opinion research and employs rigorous computational social sciences approaches to mitigate measurement errors and techniques such as Multilevel Regression with Poststratification to enhance representativeness. Additionally, this dataset incorporates strong privacy-preserving measures, including differential privacy, ensuring compliance with legal frameworks like the GDPR while maintaining data utility. Through this innovative and interdisciplinary methodology, the dataset delivers holistic, granular, and ethically sound insights into the dynamics of political opinion in Turkey, thereby contributing substantively to both public opinion research and privacy preservation scholarship.
Sociality borne by language, as is the predominant digital trace on text-based social media platforms, harbours the raw material for exploring a multitude of social phenomena. Distinctively, the messaging service Telegram provides functionalities that allow for socially interactive as well as one-to-many communication. The Telegram dataset presented here contains over 5800 groups and channels discussing conspiracy-related topics with 63 million messages, originating from a data-hoarding initiative named the “Schwurbelarchiv” (from German schwurbeln: speaking nonsense). Uniquely, it includes the transcriptions of 2.5 million audio and video files. Our contribution is a processed, research-ready version of this data hoard: we parse, clean, and validate the raw archive, pseudonymise user data, and transcribe roughly 126,000 hours of audio and video content. In its original form the archive was stored in a format that is difficult to process and largely inaccessible for systematic research. This dataset publication details the structure, scope, and methodological specifics of the Schwurbelarchiv, emphasising its relevance for further research on the German-language conspiracy-related discourse. We validate its predominantly German origin by linguistic and temporal markers and situate it within the context of similar datasets. We describe process and extent of the transcription of multimedia files. Thanks to this effort the dataset uniquely supports analysis of text from originally multimodal sources like voice messages and videos to investigate online social dynamics and content dissemination. Researchers can employ this resource to explore societal dynamics for example related to conspiracy theories, misinformation, political extremism, and social network structures.
Graph neural networks (GNNs) achieve state-of-the-art performance on link prediction, yet the link formation mechanisms they implicitly learn on complex networks remain poorly understood. From a statistical–mechanical and data-driven perspective, we ask which local structural descriptors drive their predictions, how these depend on network type, and to what extent they go beyond classical heuristics. We address these questions for SEAL, a representative subgraph-based GNN, on 27 undirected real-world networks spanning social-like, collaboration, biological and technological domains. For each candidate edge, SEAL operates on its enclosing subgraph. In parallel, we compute a common set of interpretable structural descriptors—neighborhood overlap (common neighbors, Adamic–Adar, resource allocation, Jaccard), shortest-path distance, degree statistics, local clustering and triangle counts—and train global surrogate models (decision trees and explainable boosting machines, EBMs) to approximate SEAL’s outputs and directly predict ground-truth links. This yields quantitative feature-importance profiles characterising both the label-side link formation behaviour and the GNN’s decision function. Across the 27 networks, short distances, neighborhood overlap and triadic closure systematically dominate link formation, while degree-based features play a secondary role. Triangle importance correlates strongly with global clustering, and PCA/t-SNE embeddings of the global importance vectors summarise how structural importance profiles vary across networks. We treat these embeddings as exploratory visualizations of cross-network patterns rather than as evidence for statistically significant domain separation. Degree-preserving randomization suppresses triangle and similarity-related features and leaves distance plus degree as the dominant mechanism, with limited loss in predictive performance. Overall, SEAL largely reuses classical structural heuristics as effective low-order interactions, while exploiting additional higher-order subgraph patterns in specific network domains. Clinical trial number: not applicable.
Large-scale human mobility datasets play increasingly critical roles in many algorithmic systems, business processes, and policy decisions. Unfortunately, there has been little focus on understanding bias and other fundamental shortcomings of these datasets and how they impact downstream analyses and prediction tasks. In this work, we study ‘data production’, quantifying not only whether individuals are represented in big digital datasets, but also how they are represented in terms of how much data they produce. We study one GPS mobility dataset which is collected from anonymized smartphones for ten major US cities and find that data points can be more unequally distributed between users than wealth. We build models to predict the number of data points we can expect to be produced by the composition of demographic groups living in census tracts, and find strong effects of wealth, ethnicity, and education on data production. While we find that bias is a ubiquitous phenomenon, occurring in all ten cities, we further find that each city suffers from its own manifestation of it, and that location-specific models are required to model bias for each city. This work raises serious questions about general approaches to debias human mobility data and urges further research.
Mobility data provides rich insights into human behavior, enabling applications in transportation, retail, and marketing, among others. However, it also poses significant privacy risks, as the unicity of individual mobility patterns may enable user identification within the dataset. Existing measures that aim to quantify the degree of unicity in users’ characteristics, here referred to as user exposure, primarily focus on spatio-temporal unicity (e.g., Uniqueness), while overlooking broader behavioral traits that may also increase identification risks. To address this limitation, we have proposed MoBES (Mobility Behavior-based Exposure Score), a method for quantifying user exposure in mobility datasets. MoBES models user mobility behavior in a multidimensional behavioral space and estimates exposure based on behavioral distance. It also provides a continuous value score, enabling a more refined, nuanced analysis of user exposure. Building on our preliminary study of MoBES, this work considerably broadens and deepens the analysis of its design considerations for user exposure based on mobility behavior. To this end, we evaluate MoBES on two distinct Call Detail Record (CDR) datasets and analyze its behavior in a privacy-preserving setting where location data is perturbed using Differential Privacy (DP). Using two datasets enables a comprehensive comparative analysis, offering deeper insight into how behavioral patterns influence exposure levels. Our findings show that, although MoBES often aligns with traditional spatio-temporal measures, it also reveals some potential behavioral vulnerabilities that such measures fail to capture. Furthermore, we find that privacy protection mechanisms can effectively reduce MoBES exposure scores; however, the extent of this reduction varies across behavioral dimensions. This variation reveals that even under privacy protection, certain metrics could still be explored by potential adversarial methods.
We present a statistical model for the joint analysis of temporal social network data collected as relational events (e.g., digital traces) and as a series of relational states (e.g., repeated surveys). This model effectively combines behavioral and cognitive measures of social ties. It enables hypothesis tests on the coevolution between digital behavioral traces of relationships (e.g., social interactions online) and their cognitive perceptions (e.g., friendships). We introduce a new estimation routine, based on the expectation-maximization algorithm, to address the challenge posed by observing digital and survey data at different frequencies. We apply this framework in an empirical study of an emerging community of first-year undergraduate students at a Swiss university. We combine repeated survey measures of friendship perceptions with time-stamped social media connections collected through the Facebook API. Our results suggest that the online and offline networks coevolve: changes in one network provide valuable information for modeling the respective other. Additionally, indirect connections in the offline network were associated with the creation of direct ties in the online network. Both networks exhibit similarities in transitivity, homophily, and gender effects. However, they differ in preferential attachment: there is evidence for degree popularity online and against it in the offline network. The paper highlights the broad applicability of the model for future dynamic social network studies that combine relational events and relational states.
Financial markets are prototypical nonequilibrium complex systems in which shocks propagate through time-varying, nonlinear channels that linear correlation misses. We develop a unified statistical-mechanical framework to study oil–equity interdependence over 2015–2025 by coupling (i) rolling mutual-information (MI) dependency networks estimated with the KSG k-NN estimator and filtered via PMFG to extract a sparse backbone, (ii) turbulence and nonlinear-dynamics diagnostics (Mahalanobis distance, permutation entropy, largest Lyapunov exponent) that quantify departures from typical states, and (iii) regime detection via a z-scored composite and PELT segmentation, with forecasting through an MI-informed network autoregression. We uncover a persistent, tightly knit equity core anchored by U.S. benchmarks and a structurally peripheral oil block that couples more strongly to energy equities than to broad markets; network density remains ∼0.62–0.65 while clustering stays high (∼0.75–0.80), with episodic rewiring during stress (COVID-19; 2022–2024). Turbulence spikes coincide with entropy dips and weaker Lyapunov exponents, signaling transient complexity loss. Embedding MI network spillovers improves point forecasts relative to a VAR with a lower RMSE/MAE. The framework is robust to MI estimation, filtering, and covariance regularization and yields compact early-warning indicators of fragility for portfolio and systemic-risk surveillance. We also quantify motif over/under-expression (triangles and wedges under a label-shuffle null), showing persistent equity-clique closure and episodic equity–oil mixing aligned with detected regimes.
The rapid expansion of artificial intelligence (AI) is documented in heterogeneous sources that differ in terminology, granularity, and temporal patterns. Existing scientometric approaches often rely on single-source datasets and flat topic labels, limiting their ability to capture multi-level conceptual change or differences between communities. This paper introduces a metric-based framework for analyzing topic trends across heterogeneous sources using hierarchical taxonomies. Rather than proposing a new taxonomy or labelling method, we show how hierarchical topic structures can support aggregation and comparison across levels of abstraction. The framework provides general metrics quantifying topic popularity, granularity, diversity, temporal dynamics, and cross-source temporal uniqueness. We illustrate the approach using three contemporary sources (arXiv, Hugging Face, and the Deep Learning Weekly newsletter) mapped to a taxonomy through an adapted labelling pipeline. The analyses highlight distinct patterns in topical focus, detail depth, and temporal development across sources. The framework is extensible and offers a foundation for scalable, and multi-source monitoring of AI’s evolving landscape. In addition, an interactive tool was developed to assist exploration of the trend analysis results.
The increasing volume of suspicious transaction reports (STRs), covering transactions that may be generated by criminal activity, presents both opportunities and challenges for anti-money laundering (AML) investigations. Objective STRs, automatically triggered based on predefined criteria, are generally de-prioritized by investigators compared to subjective STRs filed on expert judgment. In this study, we analyze the investigative value of objective STRs using five years of data from the Dutch Anti-Money Laundering Centre, part of the Dutch Investigation Service for Financial and Tax Crime (FIOD), one of the national agencies involved in identifying and investigating money laundering. We apply a network science framework to analyze how objective STRs contribute to (i) expanding the number of unique entities linked to investigated cases, (ii) connecting previously disconnected dossiers, and (iii) improving predictive accuracy for identifying entities under investigation. Although objective STRs represent a small fraction of reports, they increase network coverage and connectivity, adding new entities and dossiers, including cases under investigation. However, including objective transaction data when computing centrality-based predictors does not improve the accuracy of identifying entities under active investigation. These findings suggest that, while objective reports help flag potential cases, more descriptive reporting is needed to support investigations.
Abstract Large Language Models (LLMs) are increasingly used to simulate public opinion and societal behavior as a scalable complement to traditional surveys. Yet their deployment in this role raises concerns about fairness, social representation, socio-demographic bias, and cross-cultural validity. In this study, we conduct a fairness audit of LLM-simulated public opinion, examining both cross-national differences (Chile vs. the United States) and disparities across socio-demographic groups within each country using fairness metrics. Using nationally representative survey data from both countries, we evaluate patterns of predictive accuracy and group-level bias. We find substantial disparities. LLMs reproduce U.S. survey responses more faithfully than Chilean ones, consistent with their predominantly U.S.-centric training data. Moreover, the structure of bias varies across contexts: in the United States, disparities are most pronounced along race and political identity, whereas in Chile, gender, education, and religion emerge as more salient axes of inequalities. These findings reveal the uneven social grounding of LLMs and the risk of epistemic injustice when globally trained models are applied to underrepresented regions. Our results serve as a cautionary tale for researchers using LLMs to simulate public opinion, particularly in underrepresented and cross-cultural contexts.
Abstract Walking to school promotes children’s health and sustainable mobility, serving as a key metric for child-friendly cities. However, the share of children walking to school has continued to decline worldwide. The understanding of child-environment interactions is limited by previous studies’ reliance on linear assumptions and single data sources, which obscures their inherent complexity. To address this, we developed a multi-scale analytical framework that integrates street view imagery, perceptual surveys, and geospatial data from a sample of 906 children. We used XGBoost and SHAP to disentangle the complex associations between the built environment, objective street features, subjective perceptions, and walking frequency. The model demonstrated robust performance ( R 2 = 0.60 $R^{2}=0.60$ ). The results show three main findings: (1) building density is the primary predictor; (2) subjective perceptions and objective street features outperform traditional road network morphology indicators in explanatory power; and (3) most environmental factors exhibit nonlinear relationships, manifesting as threshold effects or diminishing marginal returns. This study underscores that children’s walking to school depends not merely on physical access or safety, but is significantly associated with subjective perceptions and objective street features once basic physical thresholds are met. These findings shift the focus of child-friendly planning from traditional infrastructural metrics to perceptual and experiential dimensions.
Multimedia content has become increasingly prominent in online communication, yet its role in shaping information diffusion in non-algorithmic platforms remains insufficiently understood. This study investigates how content format relates to message visibility and cross-channel diffusion within Spanish alternative Telegram channels. Using a longitudinal dataset of over 1.7 million messages (2019–2025), we analyze engagement through within-channel performance metrics and propagation through inter-channel forwarding networks. Building on theories of cognitive fluency, heuristic processing, and network-based diffusion, we examine the role of multimedia content in driving visibility and information spread. Our results show that multimedia messages are systematically more likely to achieve above-average visibility and to enter cross-channel diffusion circuits, even when controlling for channel-level heterogeneity. Moreover, multimedia diffusion occurs within denser and more centralized forwarding networks, suggesting that content format is associated with distinct structural pathways of amplification. These findings indicate that, in the absence of algorithmic curation, content format functions as a primary driver of both attention and diffusion. Rather than representing a purely stylistic feature, multimedia emerges as a structural component of information dynamics within alternative Telegram ecosystems.
Science is a dynamic process involving actors at different scales: from papers to countries going through authors and institutions. Nonetheless, the role of institutions - which is critical for assessing impact of some policies - has been largely understudied. To bridge this gap, we study the scientific portfolio of 26 institutions across the globe in the period 2000-2020. By analyzing 2.4 Mi papers, we find a correlation between portfolio dynamism and diversification, with more dynamic institutions tending to produce research across a broader range of non-related topics. Additionally, we find that not all research topics are equal in terms of impact: High-yielding topics show a significantly positive correlation between relative specialization and scientific impact per paper and are concentrated around a few research areas. Our results reveal robust patterns in university portfolio evolution that open the door to the development of mechanistic theories that explain the effect of university policies in the evolution of their portfolios.
Conventional textbook analysis often faces fragmented frameworks, obscuring the relationship between content structure and cognitive development. To address this gap, we propose an integrative computational modeling framework that combines complex network analysis with cognitive taxonomy. Specifically, textbook knowledge networks (TKNs) are constructed by representing knowledge points as nodes, with links defined by their co-occurrence within paragraphs. Cognitive attributes derived from Bloom’s taxonomy are then incorporated, enabling a joint assessment of structural topology and cognitive demand. Analyzing seven Chinese and U.S. physics textbooks (encompassing 18,209 paragraphs) reveals shared topological features and distinct cultural paradigms in textbook design. Chinese textbooks exhibit dense, low-modularity networks that prioritize systematic coherence, whereas U.S. textbooks adopt modular, long-path structures that promote cross-theme reasoning. Both systems exhibit heterogeneous core–periphery architecture (r<0) that scaffolds progressive knowledge integration. Exponential random graph models (ERGM) and mixed-effects modeling uncover four universal principles of high-quality TKNs: self-organization, thematic homophily, centrality-preferential attachment, and cross-cultural stability in topology–cognition coupling. Crucially, network centrality robustly predicts cognitive strength (β =0.408–0.765, p < 0.001), demonstrating topology is closely associated with cognitive development. Furthermore, Chinese textbooks display a bidirectional reinforcement between structure and cognition, while U.S. textbooks demonstrate an implicit alignment with cognitive progression through a low-intervention approach. These findings not only elucidate the fundamental relationship between textbook design and cognitive development, but also introduce a topology–cognition coupling analysis that provides an empirical foundation for cross-cultural textbook optimization and concept-centered instruction, contributing to global STEM education reform.
This study investigates the stereotypical portrayal of female major characters (FMCs) in fiction books. Building on prior indications that male authors create relatively stereotypical FMCs, this study analyzes public data and reader perceptions of 13,869 novels published between 1485 and 2025, with the majority originating from 1975 onwards, to examine whether the effect of author gender might be explained by genre confounding. FMC stereotypicality was measured on four dimensions: warmth-communion, agency-competence, occupational rank, and appearance-focus with all character annotations being collected by a search-enabled AI agent. Human-in-the-loop-validation of the agent’s search traces indicated an annotation accuracy between 80
Large-scale urban social media data can provide substantial insights into the real-time development of cities around the globe, illuminating phenomena such as gentrification, urban decay, and resilience to major adverse events. This study utilizes a dataset of over 147.8 million georeferenced tweets from multiple cities to demonstrate their potential for analyzing the emotional and temporal impacts of major events, including the U.S. presidential elections and the Covid-19 pandemic. By employing a sentiment indicator and an anxiety indicator, we highlight the importance of establishing robust baselines that are not only city-specific but also long-term, population-based, and user-based. We demonstrate the value of integrating georeferenced data with long-term analysis to uncover spatial and temporal patterns in public emotional responses, offering new perspectives on the dynamics of crises, such as climate change, and societal resilience.
Dictionary-based text analysis, where researchers select keywords to measure constructs such as public sentiment, anxiety, or political attitudes in large text corpora, is widely used in computational social science. However, keyword selection is rarely subjected to the same psychometric scrutiny applied to survey instruments: studies seldom report reliability, evaluate internal structure, or test whether the measurement holds across subpopulations or time points. Moreover, few existing methods enable the construction of measures that reflect theoretical or expected relationships among keywords. This paper proposes a method that brings these capabilities to text analysis by applying Confirmatory Factor Analysis (CFA) to word embeddings. Keywords are treated as observed indicators of a latent construct, and their semantic relationships, operationalized as centered cosine similarities between embedding vectors, serve as the input correlation matrix for CFA estimation. The framework enables researchers to estimate factor loadings and model fit indices (CFI, TLI, RMSEA, SRMR), compute reliability coefficients (Cronbach’s alpha, Omega), and test measurement invariance across groups or time periods using multigroup models with structured means. Moreover, the method allows researchers to compare latent construct intensity across groups or time periods, transforming keyword-based text measures from descriptive indicators into formally comparable latent variables. The method is demonstrated through an empirical application of the discourse of war anxiety during Russia’s 2022 invasion of Ukraine. A Monte Carlo simulation further examines the behavior of fit indices under random keyword selection. The approach complements existing text analysis methods and can be implemented using standard software, such as the lavaan R package.