Online social networks play a central role in shaping public discourse, yet con- ducting controlled experimental research on such platforms remains challenging due to limited access, lack of transparency in ranking algorithms, and restricted intervention capabilities. This paper presents TWON, a scalable and modular social media platform designed to enable controlled, reproducible experimen- tation on user behavior, information diffusion, and algorithmic interventions. The platform further incorporates large language model (LLM) capabilities for content generation, moderation, and agent-based simulation, enabling hybrid experimental designs that combine human participants with automated agents. The system is validated through multiple empirical deployments across diverse research contexts, including (i) a disinformation study analyzing user engagement with manipulated news content (272 participants), (ii) a scientific communication study evaluating pre-bunking and uncertainty interventions (1200 participants), (iii) a toxicity prevention study leveraging real-time AI-assisted comment rewrit- ing (574 participants), and (iv) large-scale agent-based simulations exploring conversational dynamics under different ranking strategies (54 agents). Across these studies, TWON supports flexible experimental configurations and cap- tures fine-grained behavioral data, enabling systematic analysis of engagement patterns and intervention effects.
We present a data-driven multi-agent framework for simulating social media activity using large language models (LLMs). The system combines persona-based agents, a temporal fusion transformer (TFT) for predicting posting times, retrieval-based context grounding, and configurable ranking mechanisms. These components allow agents to post, comment, and interact within a dynamic network. We evaluate the simulation across three domains: Technology, Climate Change, and COVID-19. We compare the generated data with real Reddit data using temporal, linguistic, and network metrics. When agents follow real-world event timing predictions, the simulation produces realistic patterns, participation distributions, and cascade structures. However, qualitative analysis shows that while structural diffusion patterns are present, semantic coherence often weakens in longer interaction chains. Two main interaction types appear: opinion–argument threads and question–answer threads. Homophily is mainly visible at the stylistic and emotional level, while strong topical alignment is less consistent, especially in large-scale simulations. We also analyze how ranking strategies affect network structure. Chronological ranking preserves diversity. Engagement-based ranking increases influence inequality. Hybrid ranking increases clustering and socially close exposure. These structural patterns are often associated with echo chambers. Overall, the framework provides a controlled environment for studying temporal activity, ranking effects, and structural interaction dynamics in simulated social networks.
Agent-based modeling (ABM) provides a powerful framework for exploring how individual behaviors and interactions give rise to collective social dynamics. However, most ABMs rely on handcrafted or parameterized agent rules that are not empirically grounded, thereby limiting their realism and validation against observed data. To address this gap, we constructed a large-scale, empirically grounded dataset from Reddit to support the development and evaluation of agent-based social simulations. The dataset includes 33 technology-focused, 14 climate-focused, and 7 COVID-related aggregated agents, encompassing around one million posts and comments. Using publicly available posts and comments, we define agent categories based on content and interaction patterns, derive inter-agent relationships from temporal commenting behaviors, and build a directed, weighted network that reflects empirically observed user connections. The resulting dataset enables researchers to calibrate and benchmark agent behavior, network structure, and information diffusion processes against real social dynamics. Our quantitative analysis reveals clear topic-dependent differences in how users interact. Climate discussions show dense, highly connected networks with sustained engagement, COVID-related interactions are sparse and mostly one-directional, and technology discussions are organized around a small number of central hubs. Manual qualitative analysis further shows that agent interactions follow realistic patterns of timing, similarity between users, and sentiment change.
The proliferation of fake news across social media, headlines, and news articles poses major challenges for automated detection, particularly in multilingual and cross-media settings affected by data imbalance. We propose a fake news detection framework based on LLM-driven, feature-guided text augmentation. The method generates realistic synthetic samples across languages, media types, and text granularities while preserving meaning and stylistic coherence. Experiments with classical and transformer-based models (Random Forest, Logistic Regression, BERT, XLM-R) across social media, headlines, and multilingual news datasets show consistent improvements in performance. For inherently balanced datasets (e.g., social media), synthetic augmentation yields negligible but stable performance changes. Across imbalanced scenarios, synthetic augmentation substantially improves minority-class recall and F1-score (e.g., fake news recall from 0.57 to 0.86), while preserving majority-class performance, leading to more balanced and reliable classifiers, whereas oversampling significantly degrades results due to overfitting on duplicated language patterns. Overall, a hybrid semantic- and style-based model proves to be the most robust strategy, outperforming oversampling and matching or exceeding baseline performance across datasets.
As the complexity and number of machine learning (ML) models grows, well-documented ML models are essential for developers and companies to use or adapt them to their specific use cases. Model metadata, already present in unstructured format as model cards in online repositories such as Hugging Face, could be more structured and machine readable while also incorporating environmental impact metrics such as energy consumption and carbon footprint. Our work extends the existing State of the Art by defining a structured schema for ML model metadata focusing on machine-readable format and support for integration into a knowledge graph (KG) for better organization and querying, enabling a wider set of use cases. Furthermore, we present an example wireless localization model metadata dataset consisting of 22 models trained on 4 datasets, integrated into a Neo4j-based KG with 113 nodes and 199 relations.
In this paper, we introduce a dataset of multilingual news articles covering the 2021 Tokyo Olympics. A total of 10,940 news articles were gathered from 1,918 different publishers, covering 1,350 sub-events of the 2021 Olympics, and published between July 1, 2021, and August 14, 2021. These articles are written in nine languages from different language families and in different scripts. To create the dataset, the raw news articles were first retrieved via a service that collects and analyzes news articles. Then, the articles were grouped using an online clustering algorithm, with each group containing articles reporting on the same sub-event. Finally, the groups were manually annotated and evaluated. The development of this dataset aims to provide a resource for evaluating the performance of multilingual news clustering algorithms, for which limited datasets are available. It can also be used to analyze the dynamics and events of the 2021 Tokyo Olympics from different perspectives. The dataset is available in CSV format and can be accessed from the CLARIN.SI repository.
This paper presents BAR-Analytics, a web-based, open-source platform designed to analyze news dissemination across geographical, economic, political, and cultural boundaries. Using the Russian–Ukrainian and Israeli–Palestinian conflicts as case studies, the platform integrates four analytical methods: propagation analysis, trend analysis, sentiment analysis, and temporal topic modeling. Over 350,000 articles were collected and analyzed, with a focus on economic disparities and geographical influences using metadata enrichment. We evaluate the case studies using coherence, sentiment polarity, topic frequency, and trend shifts as key metrics. Our results show distinct patterns in news coverage: the Israeli–Palestinian conflict tends to have more negative sentiment with a focus on human rights, while the Russia–Ukraine conflict is more positive, emphasizing election interference. These findings highlight the influence of political, economic, and regional factors in shaping media narratives across different conflicts.
Machine learning models deployed in non-stationary environments degrade silently, since as the input distribution drifts their accuracy decays without an error signal and without labels to reveal it. Sustaining reliable AI therefore requires a concept-drift detector that acts as an external observer of the deployed model, monitoring it using unlabeled operational data alone, so that an MLOps actuator triggers retraining and redeployment only when it is warranted. This paper contributes two concept drift detectors, namely Confidence-Filtered Pseudo-Label Transfer (CFPT) and TabAutoDrift, which combine representation learning with statistical testing to compute an expected utility score that signals whether a deployed model should be retrained, without requiring ground-truth labels after deployment. The detectors are evaluated on two emerging, label-scarce wireless application domains in which post deployment ground truth is effectively unavailable, namely outdoor fingerprinting-based localization and link-anomaly detection. They outperform the classical detectors ADWIN, DDM, and CUSUM, attaining a drift-detection F1-score between 0.88 and 0.94 in the fingerprinting use case and between 0.80 and 1.00 in the link-anomaly use case, up to 0.24 higher than the strongest classical detector. Interpreted as reliability decisions, this precision indicates that the proposed detectors signal retraining more dependably.
This paper introduces an innovative deep learning-based optimization method specifically designed for data derived from stochastic processes. Addressing the prevalent issue of rapid overfitting in real-world scenarios with limited historical data, our approach focuses on denoising optimization. The method effectively balances the simultaneous optimization of latent data representation and target variables, leading to enhanced model performance. We rigorously test our approach using five diverse real-world datasets. Our study is structured into three parts: an ablation study to validate the individual components of our method, a statistical analysis using the Wilcoxon rank-sum test to confirm the superiority of our method against five research hypotheses, and a detailed exploration of parameter visualization and fine-tuning. The comprehensive evaluation demonstrates that our method not only outperforms existing techniques but also significantly contributes to the advancement of deep learning models for stochastic processes. The findings underscore the potential of our method as a robust solution to the challenges in modeling stochastic processes with deep learning, offering new avenues for efficient and accurate predictions.
User engagement on social media platforms is influenced by historical context, time constraints, and reward-driven interactions. This study presents an agent-based simulation approach that models user interactions, considering past conversation history, motivation, and resource constraints. Utilizing German Twitter data on political discourse, we fine-tune AI models to generate posts and replies, incorporating sentiment analysis, irony detection, and offensiveness classification. The simulation employs a myopic best-response model to govern agent behavior, accounting for decision-making based on expected rewards. Our results highlight the impact of historical context on AI-generated responses and demonstrate how engagement evolves under varying constraints.
The dissemination of information worldwide is significantly facilitated by the news media, with many events having global relevance across various regions. However, certain news events receive limited coverage restricted to specific geographic areas, due to the barriers that hinder the spread of information. These barriers can be attributed to political, geographical, economic, cultural, or linguistic factors. In this research, we propose an approach for classifying these barriers by extracting semantic information from news articles using Wikipedia-concepts. Our methodology involves the collection of news articles, each annotated to indicate the specific barrier types, leveraging metadata from news publishers. Subsequently, we employ Wikipedia-concepts, in conjunction with the content of the news articles, as features to determine the barriers to news dissemination. Our approach is then compared with traditional text classification techniques, deep learning methods, and transformer-based models. We have performed experiments on news articles from ten categories of topics including health, sports, business, etc. The findings indicate that 1) Utilizing semantic knowledge yields distinct concepts across the ten categories, thereby enhancing the effectiveness and speed of the classification model. 2) The proposed approach, incorporating Wikipedia-concepts-based semantic knowledge, leads to improved performance in barrier classification when compared to using solely the body text of news articles. Specifically, there is an increase in the average F1-scores for four out of five barriers, with the economic barrier rising from 0.65 to 0.68, the linguistic barrier from 0.71 to 0.72, the political barrier from 0.68 to 0.70, and the geographical barrier from 0.63 to 0.68.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Modern text generation metrics use semantic representations of words to assess the quality of a text generation model without considering the fluency of the generated text. This paper proposes a novel text generation metric that combines adequacy and fluency to measure the quality of the generated text. When computing the final score using optimal transport, the metric considers semantic meaning and word order. We evaluate the metric on text translation data sets consisting of 20 language pairs from various language families and scripts. Using a novel statistic for measuring word order sensitivity, we analyze its adequacy-based performance using Pearson’s r and Kendall’s $\tau $ correlation coefficients and their sensitivity to fluency-related modifications. Results show that the proposed metric is the most sensitive to fluency-related changes among all top-performing embedding-based metrics, which were found to be relatively invariant to variations in word order. The proposed metric’s overall adequacy-based performance is lower than the best embedding-based metric but higher than the n-gram matching metrics. Our code is publicly available on GitHub (https://github.com/eriknovak/metric-OPWScore) under the BSD-2-Clause license.
This paper introduces an innovative methodology for conducting demand analysis within various domains across multiple countries, presenting insights derived from a comprehensive analysis conducted in seven distinct nations. The proposed methodology provides a systematic process for gathering, processing, and analyzing textual data to discern global talent demand within specified domains. Three fundamental steps characterize the methodology: identification of areas of interest, calculation of relative demand, and execution of analysis and forecasting. Areas of interest within the job market are pinpointed through a combination of manual curation and automated techniques, ensuring the inclusion of only pertinent job postings. Relative demand for jobs within each specified domain is then computed for each country, furnishing a standardized metric for cross-country comparisons. Subsequently, analysis and forecasting unveil trends and patterns in job demand, enabling stakeholders to anticipate shifts in the job market landscape. Application of this methodology to scrutinize demand analysis across seven countries - the US, Canada, the UK, France, India, Singapore, and Australia - reveals substantial variations in demand across diverse regions, along with correlations and instances of skill shortages. These findings offer invaluable insights for policymakers, businesses, and researchers, facilitating informed decision-making and fostering growth and innovation within the specified domains.
This paper introduces a conceptual architecture design aimed at enhancing interactions with cognitive digital twins of countries through an Large Language Model (LLM) agent. By leveraging sophisticated data retrieval and summarization techniques, the architecture integrates data from diverse sources, including environmental sensors, web pages, and human inputs, to create a dynamic and comprehensive digital twin. The LLM agent facilitates intuitive conversational interfaces, allowing users to query and interact with the digital twin in a natural manner. Through advanced natural language processing and prompt engineering, the agent can understand complex queries, retrieve relevant data, and provide transparent and explainable insights. Additionally, the system incorporates a feedback loop for continuous improvement based on user interactions. This approach addresses significant challenges in data acquisition and management, offering a scalable solution for creating accurate and real-time representations of countries. The architecture aims to empower decision-makers with precise, actionable insights for policy-making, urban planning, and resource management, demonstrating a significant step towards realizing the potential of digital twins in understanding and managing complex national systems.
Graph processing is increasingly popular given the wide range of phenomena represented as graphs (e.g., social media networks, pharmaceutical drug compounds, or fraud networks, among others). The increasing amount of data available requires new approaches to efficiently ingest and process such data. In this research, we describe a solution at a conceptual level in the context of the Graph-Massivizer architecture. Graph-Inceptor aims to bridge the void among ETL tools enabling data transformations required for graph creation and enrichment and supporting connectors to multiple graph storages at a massive scale. Furthermore, it aims to enhance ETL operations by learning from data content and load and making decisions based on machine-learning-based predictive analytics.
Identifying political bias in news headlines holds significant importance as it influences the dissemination and consumption of news stories. However, employing conventional methods to do so poses a formidable challenge, as the short headline text is often complex and lacks sufficient syntactic and semantic context. Existing approaches fail to acknowledge the potential of commonsense reasoning in facilitating text comprehension, although it has been shown to aid numerous downstream applications. To this end, to facilitate comprehension and compensate for the lack of context, we propose leveraging inferential commonsense knowledge to simplify, interpret, and explain events that are not explicitly stated in headlines. Furthermore, to fully utilise its potential and deal with the unnecessary noise it may introduce, we present a method for emphasising significant inferences. Using this knowledge, we introduce a novel framework, IC-BAIT, short for Inferential Commonsense aware BiAs IdenTifier, which is a flexible neural network framework designed to enhance political bias prediction in news headlines. We also present two bias-annotated datasets: MediaBias and GoodNews. Experiments on both datasets demonstrate that IC-BAIT significantly enhances the performance of the baseline models used in the framework. Experiments on the datasets show that IC-BAIT improves the baseline models in terms of accuracy (2.0-10.0%), macro-averaged F-1 (2.2-22.2%), Jaccard-score (up to 15.1%), and micro-averaged F-1 (up to 18.6%). Our in-depth qualitative analysis reveals the scenarios in which the selected knowledge is beneficial and when it is detrimental. Datasets and scripts are available at https://github.com/Swati17293/IC-BAIT.
Blaz Fortuna合作论文数Text and Web Mining group at Department of Knowledge Technologies11