Causal reasoning is fundamental to human understanding and information organization. People prefer causal explanations because they offer coherence, predictability, and a sense of control. Conversational structures shape how knowledge and perspectives are shared, validated, and amplified in networked publics. Understanding the structural effects of causal language can reveal pathways to fostering deeper, more meaningful interactions online. In this work, we investigate how causal language influences the topology and temporal evolution of discussion threads in online conversations with a dataset of 17 million posts across 200 subreddits in 2023 on Reddit. Our results show that causal language is consistently associated with deeper, more sustained conversations, with effects emerging early in the lifecycle of a thread, as demonstrated through a counterfactual experiment. Importantly, emotional responses do not differ substantially between causal language and non-causal language, suggesting that structural depth arises from framing itself rather than affective escalation. A lightweight qualitative analysis shows that causal framing titles prompt users to elaborate more with reasoning and contribute personal experiences, supporting deeper multi-turn exchanges. These findings suggest that causal language acts not merely as a stylistic device, but as a cognitively grounded and structurally influential signal that shapes the topological structures of online conversations.
How do shared narratives emerge in decentralized online networks? Prior research using simplified group coordination tasks (e.g., face-naming) shows network structure shapes group consensus, but the underlying cognitive mechanisms remain unclear. Here, we examine how network structure influences the emergence and semantic content of shared narrative beliefs in experimental online social networks, using natural language processing measures and agent-based modeling. Media content with complex causal structure attenuates network structure effects by encouraging longer exploration of background knowledge. Yet network structure still shapes the narrative content communicated. An embedding-based narrative alignment measure shows that fully connected groups orient their interactions more toward communicating causes of an event, whereas locally connected networks emphasize the event's effects. A group's network structure also influences representational and language change in personal narratives: participants in fully connected networks showed the largest increase in causal language in personal narratives written after interaction, which also orient more around the narrative's causal events.
In 2022, Hyundai Sonatas became the most frequently stolen car model in Los Angeles, mirroring a nationwide trend. The sudden popularity of this mid-sized sedan among car thieves has, according to some, been linked to the viral spread of videos on TikTok and YouTube that showcased a security vulnerability that made certain Kia and Hyundai models easy to steal. These videos later acquired the moniker ‘Kia Boys’ or ‘Kia Challenge.’ In this work, we show that the surge in Hyundai Sonata thefts in Los Angeles was only weakly tied to the ‘Kia Boys’ trend, having begun much earlier during COVID-19 pandemic lockdowns. Any lasting effects of the viral social media trend were fully absorbed into offline car theft dynamics by the time social media interest peaked in 2022. We use historical data on Honda Civics and Toyota Camrys, car models popular among pre-social media generations of car thieves, to develop a model for the offline temporal and spatial evolution of theft preferences. The model suggests that the effects of the ‘Kia Boys’ vulnerability are likely to persist for decades, irrespective of the role social media played in its inception, and that these effects will be increasingly concentrated in low-income communities over time.
Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metrics. We show that communication patterns can categorize teams into performance predictive behavioral tribes, as early as 10
Language is far more than a communication tool; it encodes a wealth of information about a person's identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21-50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.
Automated methods for the assessment of replicability of scientific claims offer a scalable complement to replication studies and traditional peer review. Drawing on a large dataset of claims, human judgments, and a limited set of replication outcomes, we developed and evaluated three distinct artificial intelligence systems designed to predict human expert assessments of replicability using diverse methodologies—including synthetic prediction markets, interpretable feature-based modeling, knowledge graph reasoning, and semantic parsing with argument structures. While these systems achieved modest calibration to human judgment distributions, they failed to discriminate between replicable and non-replicable claims. Our findings suggest that while machine assessments of research replicability may complement human reasoning, their current performance limitations and opportunities for bias demand careful evaluation before real-world application.
Networked environments shape how information embedded in narratives influences individual and group beliefs and behavior. This raises key questions about how group communication around narrative media impacts belief formation and how such mechanisms contribute to the emergence of consensus or polarization. Language data from generative agents offer insight into how naturalistic forms of narrative interactions (such as hashtag generation) evolve in response to social rewards within networked communication settings. To investigate this, we developed an agent-based modeling and simulation framework composed of networks of interacting Large Language Model (LLM) agents. We benchmarked the simulations of four state-of-the-art LLMs against human group behaviors observed in a prior network experiment (Study 1) and against naturally occurring hashtags from Twitter (Study 2). Quantitative metrics of network coherence (e.g., entropy of a group's responses) reveal that while LLMs can approximate human-like coherence in sanitized domains (Study 1's experimental data), effective integration of background knowledge and social context in more complex or politically sensitive narratives likely requires careful and structured prompting.
For centuries, women have been cast as the source of harm in public narratives, from witch hunts in early modern Europe to contemporary stereotypes about emotional instability. These cultural patterns reflect enduring biases in how people attribute causality and assign blame, often portraying women as agents of disruption and men as figures of rational authority. In this study, we examine how such gendered causal attributions appear in everyday language. Leveraging three complete 24-hour datasets of all English-language posts on Twitter, and using language models, we extract cause-and-effect relationship pairs and identify gendered attribution of causal agents. We then analyze how gender attribution relates to sentiment, the kinds of effects invoked, and the diffusion of posts through the social networks. Our findings reveal that female-attributed causes are more often associated with negative sentiment and emotional or relational outcomes, whereas male-attributed causes are more frequently linked to positive sentiment and abstract, structural effects. Moreover, male-attributed narratives spread more widely across communities. These results suggest that longstanding gender stereotypes continue to appear in how people express and amplify causal narratives in public discourse, in decentralized, high-velocity environments like social media.
In the evolving landscape of text-to-3D technology, Dreamfusion [9] optimizes implicit representations like NeRF using Score Distillation Sampling (SDS) but faces limitations in both fidelity and speed. Specifically, it faces the multi-head Janus issue and exhibits a relatively slow optimization process. We present OrientDream, a camera orientation conditioned framework for efficient, multi-view consistent 3D generation from text prompts. OrientDream achieves this by pre-training a 2D text-to-image diffusion module with camera orientation features and utilizing data from MVImgNet. To shorten training time, we introduced a decoupled back-propagation technique, allowing for multiple updates of implicit parameters per optimization cycle. Our experiments reveal that our method not only produces high-quality NeRF models with consistent multi-view properties but also achieves an optimization speed significantly greater than existing methods, as quantified by comparative metrics.
Many openly non-binary gender individuals participate in social networks. However, the relationship between gender and online interactions is not well understood, which may result in disparate treatment by large language models. We investigate individual identity on Twitter, focusing on gender expression as represented by users chosen pronouns. We find that non-binary groups tend to receive less attention in the form of likes and followers. We also find that nonbinary users send and receive tweets with above-average toxicity. The study highlights the importance of considering gender as a spectrum, rather than a binary, in understanding online interactions and expression.
The recent proliferation of short form video social media sites such as TikTok has been effectively utilized for increased visibility, communication, and community connection amongst trans/nonbinary creators online. However, these same platforms have also been exploited by right-wing actors targeting trans/nonbinary people, enabling such anti-trans actors to efficiently spread hate speech and propaganda. Given these divergent groups, what are the differences in network structure between anti-trans and pro-trans communities on TikTok, and to what extent do they amplify the effects of anti-trans content? In this paper, we collect a sample of TikTok videos containing pro and anti-trans content, and develop a taxonomy of trans related sentiment to enable the classification of content on TikTok, and ultimately analyze the reply network structures of pro-trans and anti-trans communities. In order to accomplish this, we worked with hired expert data annotators from the trans/nonbinary community in order to generate a sample of highly accurately labeled data. From this subset, we utilized a novel classification pipeline leveraging Retrieval-Augmented Generation (RAG) with annotated examples and taxonomy definitions to classify content into pro-trans, anti-trans, or neutral categories. We find that incorporating our taxonomy and its logics into our classification engine results in improved ability to differentiate trans related content, and that Results from network analysis indicate many interactions between posters of pro-trans and anti-trans content exist, further demonstrating targeting of trans individuals, and demonstrating the need for better content moderation tools
The rich and dynamic information environment of social media provides researchers, policy makers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using these data to understand social behavior is difficult due to heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect, and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method by successfully detecting major and smaller events on three different datasets, including (1) a Los Angeles Tweet dataset between Jan. and Aug. 2020, in which we revealed the complex psychological impact of the BlackLivesMatter movement and the COVID-19 pandemic, (2) a dataset related to abortion rights discussions in USA, in which we uncovered the strong emotional reactions to the overturn of Roe v. Wade and state abortion bans, and (3) a dataset about the 2022 French presidential election, in which we discovered the emotional and moral shift from positive before voting to fear and criticism after voting. The capability of our method allows for better sensing and monitoring of population's reactions during crises using online data.
The rise of Large Language Models (LLMs) offers transformative potential for interpreting complex legal frameworks, such as Title 18 Section 175 of the US Code, which governs biological weapons. These systems hold promise for advancing legal analysis and compliance monitoring in sensitive domains. However, this capability comes with a troubling contradiction: while LLMs can analyze and interpret laws, they also demonstrate alarming vulnerabilities in generating unsafe outputs, such as actionable steps for bioweapon creation, despite their safeguards. To address this challenge, we propose a methodology that integrates knowledge graph construction with Retrieval-Augmented Generation (RAG) to systematically evaluate LLMs' understanding of this law, their capacity to assess legal intent (mens rea), and their potential for unsafe applications. Through structured experiments, we assess their accuracy in identifying legal violations, generating prohibited instructions, and detecting unlawful intent in bioweapons-related scenarios. Our findings reveal significant limitations in LLMs' reasoning and safety mechanisms, but they also point the way forward. By combining enhanced safety protocols with more robust legal reasoning frameworks, this research lays the groundwork for developing LLMs that can ethically and securely assist in sensitive legal domains - ensuring they act as protectors of the law rather than inadvertent enablers of its violation.
Esports analytics has gained traction in recent years, leveraging machine learning (ML) to predict in-game events and enhance strategic decision making. This study develops structured datasets from Counter-Strike: Global Offensive (CS: GO) competitive matches to address two predictive tasks: round result prediction and player death prediction. Using data from professional Dust 2 matches in 2022, we extract key playerand team-level features such as health, spatial positioning, and economy-related metrics. Various ML models, including Logistic Regression, Decision Trees, and XGBoost, are evaluated and benchmarked against random guessing and majority-class baselines. Results show that XGBoost consistently outperforms other models, effectively capturing gameplay dynamics and providing accurate predictions. These findings offer valuable insights for esports strategy optimization, coaching, real-time decision support, and applications in live-game analysis and betting. The dataset and methodology also establish a foundation for future esports research and predictive modeling across other game environments.
The present research on team collaboration is typically performed through qualitative interview based studies or social network measurements of connectedness through co-play. In this study, we take the unique approach to build networks from direct messages between players in the massive online game World of Tanks where players self-organize into clans with specific roles assigned from military rankings (from Private to Commander). We explore the relationship between team communication volume and skill level, the impact of communication features on clan rating, and the differences in communication hierarchy between high and low-rated clans. Our findings reveal that higher-rated clans send more pre-battle chat messages, suggesting that effective communication and strategic planning are key to team performance. Evidence shows teams who use voice chat during battle are significantly higher ranked. Finally, we reveal that the highest rated clans have more connected lower-ranked members emphasizing that these teams are "only as strong as their weakest link." This research is guided by the Transactive Memory Systems and Collective Intelligence theories which serve to expand the contribution of this research outside of games to other forms of virtual collaboration.
Automated emotion detection is widely used in applications ranging from well-being monitoring to high-stakes domains like mental health and hiring. However, models often rely on annotations that reflect dominant cultural norms, limiting model ability to recognize emotional expression in dialects often excluded from training data distributions, such as African American Vernacular English (AAVE). This study examines emotion recognition model performance on AAVE compared to General American English (GAE). We analyze 2.7 million tweets geo-tagged within Los Angeles. Texts are scored for strength of AAVE using computational approximations of dialect features. Annotations of emotion presence and intensity are collected on a dataset of 875 tweets with both high and low AAVE densities. To assess model accuracy on a task as subjective as emotion perception, we calculate community-informed "silver" labels where AAVE-dense tweets are labeled by African American, AAVE-fluent (ingroup) annotators. On our labeled sample, GPT and BERT-based models exhibit false positive prediction rates of anger on AAVE more than double than on GAE. SpanEmo, a popular text-based emotion model, increases false positive rates of anger from 25 percent on GAE to 60 percent on AAVE. Additionally, a series of linear regressions reveals that models and non-ingroup annotations are significantly more correlated with profanity-based AAVE features than ingroup annotations. Linking Census tract demographics, we observe that neighborhoods with higher proportions of African American residents are associated with higher predictions of anger (Pearson's correlation r = 0.27) and lower joy (r = -0.10). These results find an emergent safety issue of emotion AI reinforcing racial stereotypes through biased emotion classification. We emphasize the need for culturally and dialect-informed affective computing systems.
We evaluate the relative forecasting performance of three statistical models and a prediction market for several outcomes decided during the November 2024 elections in the United States—the winner of the presidency, the popular vote, fifteen competitive states in the Electoral College, eleven Senate races, and thirteen House races. We argue that conventional measures of predictive accuracy such as the average daily Brier score reward modeling flaws that result in predicable reversals, as long as such movements are in a direction that is aligned with the eventual outcome. Instead, we adopt a test based on the idea that the strength of a model can be measured by the profitability of a trader who believes its forecasts and bets on the market based on this belief. The results of this test depend on the risk preferences with which the trader is endowed, but we show that within a large parameter range this does not lead to ranking reversals. We find that all models failed to beat the market in the headline contract but some did so convincingly in contracts referencing less visible races.