Understanding which posts spark conversation, and how large those conversations grow, is vital for moderation, resource allocation, and anticipating information cascades on Reddit and other social platforms. We study discussion initiation and growth on Reddit by modelling whether a root post receives any comments and how large the resulting thread becomes. Using reconstructed threads from r/politics, r/CryptoCurrency, and r/Conspiracy, we extract compact textual, semantic, temporal, domain, and author features from each post. We train subreddit-specific classifiers with small, transparent feature sets and use SHAP for interpretation. Across communities, the external domain a post links to, and, in news ecosystems, the domain's centrality, consistently emerge as predictors of both the start and scale of discussion. Author activity is also predictive: posts from highly active users are more likely to receive comments. Simple textual cues help too: longer subjects and fewer question marks are associated with a higher likelihood of eliciting replies. Community context moderates these effects: in r/politics, linking familiar mid-tier but well-connected news sources is associated with larger threads, while the r/Conspiracy and r/CryptoCurrency communities prefer novel sources. Predicting whether a discussion will start is notably easier than forecasting its eventual size, as adjacent size classes are often confounded. Still, a concise, interpretable feature set captures a substantial proportion of the predictive signal. Our results suggest practical applications for triage: flagging posts likely to trigger substantial discussion could support targeted, pre-emptive moderation and fact-checking without relying on complex, opaque models.
This study investigates who should bear the responsibility of combating the spread of misinformation in social networks. Should that be the online platforms or their users? Should that be done by debunking the "fake news" already in circulation or by investing in preemptive efforts to prevent their diffusion altogether? We seek to answer such questions in a stylized opinion dynamics framework, where agents in a network aggregate the information they receive from peers and/or from influential external sources, with the aim of learning a ground truth among a set of competing hypotheses. In most cases, we find centralized sources to be more effective at combating misinformation than distributed ones, suggesting that online platforms should play an active role in the fight against fake news. In line with literature on the "backfire effect", we find that debunking in certain circumstances can be a counterproductive strategy, whereas some targeted strategies (akin to "deplatforming") and/or preemptive campaigns turn out to be quite effective. Despite its simplicity, our model provides useful guidelines that could inform the ongoing debate on online disinformation and the best ways to limit its damaging effects.
An abundance of literature has shown that the injection of noise into complex socio-economic systems can improve their resilience. This study aims to understand whether the same applies in the context of information diffusion in social networks. Specifically, we aim to understand whether the injection of noise in a social network of agents seeking to uncover a ground truth among a set of competing hypotheses can build resilience against disinformation. We implement two different stylized policies to inject noise in a social network, i.e., via random bots and via randomized recommendations, and find both to improve the population's overall belief in the ground truth. Notably, we find noise to be as effective as debunking when disinformation is particularly strong. On the other hand, such beneficial effects may lead to a misalignment between the agents' privately held and publicly stated beliefs, a phenomenon which is reminiscent of cognitive dissonance.
We present a stochastic imitation-based model of opinion dynamics in which agents balance social conformity with responsiveness to an external signal. The model captures how populations evolve between two binary opinion states, driven by peer influence and noisy external information. Through both memory-less and memory-based implementations, we identify a critical threshold of social sensitivity that separates an ergodic phase—where agents collectively track the external signal—from a non-ergodic phase characterised by persistent consensus and reduced adaptability to external changes. Analytical results and simulations reveal that memory in decision-making smooths the transition and lowers the critical threshold for ergodicity breaking. Extending the model to various network structures confirms the robustness of the observed phase transition. We further discuss empirical methodologies for estimating the critical threshold and show how the model may be applied to real-world domains. Our findings contribute to understanding how social conformity, memory effects and randomness jointly shape collective behaviour, with implications for predicting social tipping points and influencing large-scale social dynamics.
We examine the disruption of researchers with long-lived careers in Computer Science and Physics. Despite the epistemological differences between such disciplines, we consistently find that a researcher’s most disruptive publication does not occur at random during their career, as it cannot be explained by a null model. Such publication is accompanied by a peak year in which researchers publish other work that exhibits a higher level of disruption than average. Through a series of linear models, we show that the disruption achieved by a researcher during their peak year is higher when it is preceded by a long period of focus and low productivity. These findings are in stark contrast with the dynamics of academic impact. In these dynamics, researchers are incentivized by the prevalent paradigms of scientific evaluation to pursue high productivity and incremental—less disruptive—work, as evidenced by extensive literature.
In Evolutionary game theory the payoffs are typically fixed or shaped by external environmental variables. Here, we introduce an endogenous-feedback model in which the game played coevolves directly with the population state: the payoff matrix is a time-dependent function of the level of cooperation. This allows strategic incentives to be continuously modified by the collective behavior they generate. Even in the simplest case of linear and instantaneous feedback, the model reveals feedback-induced regimes, termed chimera games, in which stable cooperation arises despite being incompatible with the predictions of standard fixed-game dynamics. We further show that delayed feedback can destabilize these equilibria and generate sustained oscillations, while nonlinear feedback reshapes equilibrium structure and introduces path dependence. Our results show how cooperation can be promoted, suppressed, or destabilized by incentives generated endogenously by the very same population's collective behavior. We conclude by outlining how our framework connects to real-world systems shaped by endogenous feedback.
We propose a stylized model of a complex economy to explore the economic tradeoffs imposed by the so called "green transition" -- the shift towards more sustainable production paradigms -- using tools from the Statistical Mechanics of disordered systems. Namely, we promote the parameters of a standard input-output economic model to random variables, in order to characterize the typical statistical properties of its equilibria. A central feature of our work is the explicit inclusion of a waste variable as a proxy for the unwanted byproducts, such as emissions, generated by different production processes. We find that the interplay between economic development and waste gives rise to a double phase transition, separating a region of viable economic activity from two distinct shutdown regimes. Notably, more wasteful ("brown") economies support a larger number of technologies per good but fail to activate most of them, limiting their productive capacity. In contrast, "greener" economies, while having access to fewer technologies, tend to achieve higher economic welfare, up to a critical point where rapid growth under tight waste constraints may destabilize them. Our results highlight the structural tensions between technological development and environmental sustainability, suggesting that the green transition may require careful navigation to avoid abrupt collapses in economic activity. The nature of the phase transitions we observe in the model is linked to the shrinkage of feasible economic configurations under increasing constraints, akin to algorithmic phase transitions studied in Computer Science.
The emergence of the disruption score provides a new perspective that differs from traditional metrics of citations and novelty in research evaluation. Motivated by current studies on the differences among these metrics, we examine the relationship between disruption scores and citation counts. Intuitively, one would expect disruptive scientific work to be rewarded by high volumes of citations and, symmetrically, impactful work to also be disruptive. A number of recent studies have instead shown that such intuition is often at odds with reality. In this paper, we break down the relationship between impact and disruption with a detailed correlation analysis in two large data sets of publications in Computer Science and Physics. We find that highly disruptive papers tend to receive a higher number of citations than average. Contrastingly, the opposite is not true, as we do not find highly cited papers to be particularly disruptive. Notably, these results qualitatively hold even within individual scientific careers, as we find that—on average—an author’s most disruptive work tends to be well cited, whereas their most cited work does not tend to be disruptive. We discuss the implications of our findings in the context of academic evaluation systems, and show how they can contribute to reconcile seemingly contradictory results in the literature.
Participants in socio-economic systems are often ranked based on their performance. Rankings conveniently reduce the complexity of such systems to ordered lists. Yet, it has been shown in many contexts that those who reach the top are not necessarily the most talented, as chance plays a role in shaping rankings. Nevertheless, the role played by chance in determining success, i.e. serendipity, is underestimated, and top performers are often imitated by others under the assumption that adopting their strategies will lead to equivalent results. We investigate the tradeoff between imitation and serendipity in an agent-based model. Agents in the model receive payoffs based on their actions and may switch to different actions by either imitating others or through random selection. When imitation prevails, most agents coordinate on a single action, leading to non-meritocratic outcomes, as a minority of them accumulate the majority of payoffs. Yet, such agents are not necessarily the most skilled ones. When serendipity dominates, instead, we observe more egalitarian outcomes. The two regimes are separated by a sharp transition, which we characterize analytically in a simplified setting. We discuss the implications of our findings in a variety of contexts, ranging from academic research to business.
The Great Gatsby Curve measures the relationship between income inequality and intergenerational income persistence. By utilizing genealogical data of over 245,000 mentor-mentee pairs and their academic publications from 22 different disciplines, this study demonstrates that an academic Great Gatsby Curve exists as well, in the form of a positive correlation between academic impact inequality and the persistence of impact across academic generations. We also provide a detailed breakdown of academic persistence, showing that the correlation between the impact of mentors and that of their mentees has increased over time, indicating an overall decrease in academic intergenerational mobility. We analyze such persistence across a variety of dimensions, including mentorship types, gender, and institutional prestige.
Establishing an independent academic identity is a central yet insufficiently understood challenge for early-career researchers. However, limited resources and mentor-driven research agendas often constrain early efforts toward autonomy. To provide large-scale quantitative evidence on how junior researchers develop independence, we introduce a framework that traces how mentees diverge from their mentors in both research topics and collaboration networks, and how these divergences relate to long-term scientific impact. Analyzing over 500,000 mentee-mentor pairs in Chemistry, Neuroscience, and Physics across six decades, we find that high-impact scientists often initiate work in secondary areas of their mentors' expertise while adaptively establishing distinct research trajectories. This pattern is most pronounced among mentees who eventually surpass their mentors' impact. We identify an inverted U-shaped relationship between topic divergence and mentees' enduring impact, with moderate divergence yielding the highest scientific impact, revealing an independence paradox in scientific careers. This pattern holds whether topic divergence is measured by citation network or semantic thematic distance. We further reveal that excessive direct mentor-mentee collaborations correlate with lower mentee impact, whereas expanding professional networks to include mentors' collaborators is beneficial. These findings not only offer actionable guidance for early-career researchers navigating independence but also inform institutional policies that promote mentorship structures supporting intellectual innovation and recognizing original contributions in promotion evaluations.
We develop a model of opinion dynamics where agents in a social network seek to learn a ground truth among a set of competing hypotheses. Agents in the network form private beliefs about such hypotheses by aggregating their neighbors' publicly stated beliefs, in an iterative fashion. This process allows us to keep track of scenarios where private and public beliefs align, leading to population-wide consensus on the ground truth, as well as scenarios where the two sets of beliefs fail to converge. The latter scenario - which is reminiscent of the phenomenon of cognitive dissonance - is induced by injecting 'conspirators' in the network, i.e., agents who actively spread disinformation by not communicating accurately their private beliefs. We show that the agents' cognitive dissonance non-trivially reaches its peak when conspirators are a relatively small minority of the population, and that such an effect can be mitigated - although not erased - by the presence of 'debunker' agents in the network.
We examine the innovation of researchers with long-lived careers in Computer Science and Physics. Despite the epistemological differences between such disciplines, we consistently find that a researcher's most innovative publication occurs earlier than expected if innovation were distributed at random across the sequence of publications in their career, and is accompanied by a peak year in which researchers publish other work which is more innovative than average. Through a series of linear models, we show that the innovation achieved by a researcher during their peak year is higher when it is preceded by a long period of low productivity. These findings are in stark contrast with the dynamics of academic impact, which researchers are incentivised to pursue through high productivity and incremental - less innovative - work by the currently prevalent paradigms of scientific evaluation.
How difficult is it for an early career academic to climb the ranks of their discipline? We tackle this question with a comprehensive bibliometric analysis of 57 disciplines, examining the publications of more than 5 million authors whose careers started between 1986 and 2008. We calibrate a simple random walk model over historical data of ranking mobility, which we use to 1) identify which strata of academic impact rankings are the most/least mobile and 2) study the temporal evolution of mobility. By focusing our analysis on cohorts of authors starting their careers in the same year, we find that ranking mobility is remarkably low for the top- and bottom-ranked authors and that this excess of stability persists throughout the entire period of our analysis. We further observe that mobility of impact rankings has increased over time, and that such rise has been accompanied by a decline of impact inequality, which is consistent with the negative correlation that we observe between such two quantities. These findings provide clarity on the opportunities of new scholars entering the academic community, with implications for academic policymaking.
It is well known that the probability distribution of high-frequency financial returns is characterized by a leptokurtic, heavy-tailed shape. This behavior undermines the typical assumption of Gaussian log-returns behind the standard approach to risk management and option pricing. Yet, there is no consensus on what class of probability distributions should be adopted to describe financial returns and different models used in the literature have demonstrated, to varying extent, an ability to reproduce empirically observed stylized facts. In order to provide some clarity, in this paper we perform a thorough study of the most popular models of return distributions as obtained in the empirical analyses of high-frequency financial data. We compare the statistical properties and simulate the dynamics of non-Gaussian financial fluctuations by means of Monte Carlo sampling from the different models in terms of realistic tail exponents. Our findings show a noticeable consistency between the considered return distributions in the modeling of the scaling properties of large price changes. We also discuss the convergence rate to the asymptotic distributions of the non-Gaussian stochastic processes and we study, as a first example of possible applications, the impact of our results on option pricing in comparison with the standard Black and Scholes approach.
We examine the tension between academic impact - the volume of citations received by publications - and scientific disruption. Intuitively, one would expect disruptive scientific work to be rewarded by high volumes of citations and, symmetrically, impactful work to also be disruptive. A number of recent studies have instead shown that such intuition is often at odds with reality. In this paper, we break down the relationship between impact and disruption with a detailed correlation analysis in two large data sets of publications in Computer Science and Physics. We find that highly disruptive papers tend to be cited at higher rates than average. Contrastingly, the opposite is not true, as we do not find highly impactful papers to be particularly disruptive. Notably, these results qualitatively hold even within individual scientific careers, as we find that - on average - an author's most disruptive work tends to be well cited, whereas their most cited work does not tend to be disruptive. We discuss the implications of our findings in the context of academic evaluation systems, and show how they can contribute to reconcile seemingly contradictory results in the literature.
Online platforms implement digital reputation systems in order to steer individual user behaviour towards outcomes that are deemed desirable on a collective level. At the same time, most online platforms are highly decentralised environments, leaving their users plenty of room to pursue different strategies and diversify behaviour. We provide a statistical characterisation of the user behaviour emerging from the interplay of such competing forces in Stack Overflow, a long-standing knowledge sharing platform. Over the 11 years covered by our analysis, we represent the interactions between users and topics as bipartite networks. We find such networks to display nested structures akin to those observed in ecological systems, demonstrating that the platform’s user base consistently self-organises into specialists and generalists, i.e., users who focus on narrow and broad sets of topics, respectively. We relate the emergence of these behaviours to the platform’s reputation system with a series of data-driven models, and find specialisation to be statistically associated with a higher ability to post the best answers to a question. We contrast our findings with observations made in top-down environments—such as firms and corporations—where generalist skills are consistently found to be more successful.
Systemic risk analysis has become a very important undertaking in most central banks after the Global Financial Crisis (GFC). This paper describes the Colombian credit system of banks and firms as a bipartite network of lenders and borrowers. To such network, we apply a spectral method to identify the most central actors, and a variant of the DebtRank algorithm to identify the banks and firms that would be the most vulnerable to shocks in the system, and the most impactful in propagating them. We perform our analysis with a multi-layer approach, analysing networks of loans in the Commercial, Housing, and Microcredit domain. Our analyses reveal a rich and heterogeneous systemic risk profile across the Colombian credit system, and highlight the presence of considerable network effects that would contribute to shape the propagation of shocks from the real economy to the banking system.