This study investigates the evolving readability of financial reporting by analyzing Item 7 of the 10-K reports over a 26-year period, utilizing a dataset of nearly 200,000 reports retrieved from SEC EDGAR filings. Our analysis reveals a significant decline in the readability of these reports over time, measured using the Fog Index. Specifically, we find that the number of years of schooling required to comprehend these texts increases by nearly one month each year, indicating a growing inaccessibility of financial reports for a substantial portion of the population. To contextualize these findings, we extend our analysis to include data from diverse financial corpora, encompassing almost 10 million documents. While most financial texts have shown a systematic increase in readability over the past decades, the Wall Street Journal emerges as a notable exception, exhibiting a moderate decline in readability-though at a much slower rate compared to Item 7. This study highlights the widening gap in financial text accessibility and underscores the need for more readable financial reporting.
Online users’ growing awareness of their digital footprint coupled with tightening regulation, has led e-commerce and social media platforms to develop and promote the use of privacy control tools. By and large, these tools directly address the concerns most often raised by users, and focus on offering individual users control over the visibility of their attributes and behaviors. At the same time, platforms, users, and regulators have overlooked the fact that many digital platforms maintain extensive data about their users’ social networks, including network topology and the attributes of the surrounding peers. This data could be used to reveal user behavior and expose details pertaining to users with private profiles, which would violate their privacy expectations. In this work we use data from Steam, an online gaming platform with an active in-house social network, to empirically demonstrate how data exposed by users with public profiles can be used to defeat privacy controls and violate users’ expectations of privacy. Specifically, we show that the topology of a private user’s local network is sufficient to reveal her attributes even if no other information about her or her peers’ attributes is available. Used in concert with data on the individual attributes of her peers, the two sets of data improve inference accuracy. Next, we conduct simulations that test several propagation scenarios for privacy-protecting behaviors. We find that the proliferation of privacy-safe behaviors offers limited protection to private-profile users, raising concerns about the effectiveness of leaving privacy controls in users’ hands. Finally, we offer a solution that could mitigate privacy concerns associated with a user’s social network. If online platforms concealed the social ties involving those users concerned about their privacy, they would thereby protect these users while maintaining the usability of the rest of the social network.
In recent years, researchers have explored fractal dimension, spectral dimension, and multifractal analysis as ways of describing the emergent hierarchical structure of complex networks. However, fractality implies an infinite recursion that is impossible for finite networks describing real-world biological systems. We show that there is a substantial finite size effect on two widely used empirical methods (box-covering and sandbox) of estimating generalized fractal dimensions. As a partial solution to this issue we introduce here a generalized method for calculating network spectral dimension using a memory-biased random walk (MBRW). To observe the impact of network size, we start with an ensemble of networks representing a variety of biological systems, identify their community structures using Infomap, and use a modified stochastic block model to generate networks with similar community structure but varying size. We find that, compared to shortest-path-based generalized fractal dimension methods, the MBRW generalized spectral dimension (Dq) shows a clearer and more consistent ordering of networks by community structure for all orders (q) considered. We also find that, among the measures of multifractality and multispectrality, only MBRW multi-spectrality (range of Dq values) changes in a consistent direction under randomization of each level of community structure. Our results show that network size is an important consideration when comparing the fractal or spectral dimensions of real-world networks and that observing the interaction between network structure and an agent acting in time with memory provides insights into network structure not available through calculations based on purely topological features.
Identity cues appear ubiquitously alongside content in social media today. Some also suggest universal identification, with names and other cues, as a useful deterrent to harmful behaviours online. Unfortunately, we know little about the effects of identity cues on opinions and online behaviours. Here we used a large-scale longitudinal field experiment to estimate the extent to which identity cues affect how people form opinions about and interact with content online. We randomly assigned content produced on a social news aggregation website to ‘identified’ and ‘anonymous’ conditions to estimate the causal effect of identity cues on how viewers vote and reply to content. The effects of identity cues were significant and heterogeneous, accounting for between 28% and 61% of the variation in voting associated with commenters’ production, reputation and reciprocity. Our results also showed that identity cues cause people to vote on content faster (consistent with heuristic processing) and to vote according to content producers’ reputations, production history and reciprocal votes with content viewers. These results provide evidence that rich-get-richer dynamics and inequality in social content evaluation are mediated by identity cues. They also provide insights into the evolution of status in online communities. From a practical perspective, we show via simulation that social platforms may improve content quality by including votes on anonymized content as a ranking signal.
Social animals, including humans, have a broad range of personality traits, which can be used to predict individual behavioral responses and decisions. Current methods to quantify individual personality traits in humans rely on self-report questionnaires, which require time and effort to collect and rely on active cooperation. However, personality differences naturally manifest in social interactions such as online social networks. Here, we demonstrate that the topology of an online social network can be used to characterize the personality traits of its members. We analyzed the directed social graph formed by the users of the LiveJournal (LJ) blogging platform. Individual users personality traits, inferred from their self-reported domains of interest (DOIs), were associated with their network measures. Empirical clustering of DOIs by topological similarity exposed two main self-emergent DOI groups that were in alignment with the personality meta-traits plasticity and stability. Closeness, a global topological measure of network centrality, was significantly higher for bloggers associated with plasticity (vs. stability). A local network motif (a triad of 3 connected bloggers) that correlated with closeness also separated the personality meta-traits. Finally, topology-based classification of DOIs (without analyzing the content of the blogs) attained > 70% accuracy (average AUC of the test-set). These results indicate that personality traits are evident and detectable in network topology. This has serious implications for user privacy. But, if used responsibly, network identification of personality traits could aid in early identification of health-related risks, at the population level.
Wikipedia is a major source of information utilized by internet users around the globe for fact-checking and access to general, encyclopedic information. For researchers, it offers an unprecedented opportunity to measure how societies respond to events and how our collective perception of the world evolves over time and in response to events. Wikipedia use and the reading patterns of its users reflect our collective interests and the way they are expressed in our search for information – whether as part of fleeting, zeitgeist-fed trends or long-term – on most every topic, from personal to business, through political, health-related, academic and scientific. In a very real sense, events are defined by how we interpret them and how they affect our perception of the context in which they occurred, rendering Wikipedia invaluable for understanding events and their context. This paper introduces WikiShark (www.wikishark.com) – an online tool that allows researchers to analyze Wikipedia traffic and trends quickly and effectively, by (1) instantly querying pageview traffic data; (2) comparing traffic across articles; (3) surfacing and analyzing trending topics; and (4) easily leveraging findings for use in their own research.
Empirical studies show that epidemiological models based on an epidemic’s initial spread rate often fail to predict the true scale of that epidemic. Most epidemics with a rapid early rise die out before affecting a significant fraction of the population, whereas the early pace of some pandemics is rather modest. Recent models suggest that this could be due to the heterogeneity of the target population’s susceptibility. We study a computer malware ecosystem exhibiting spread mechanisms resembling those of biological systems while offering details unavailable for human epidemics. Rather than comparing models, we directly estimate reach from a new and vastly more complete data from a parallel domain, that offers superior details and insight as concerns biological outbreaks. We find a highly heterogeneous distribution of computer susceptibilities, with nearly all outbreaks initially over-affecting the tail of the distribution, then collapsing quickly once this tail is depleted. This mechanism restricts the correlation between an epidemic’s initial growth rate and its total reach, thus preventing the majority of epidemics, including initially fast-growing outbreaks, from reaching a macroscopic fraction of the population. The few pervasive malwares distinguish themselves early on via the following key trait: they avoid infecting the tail, while preferentially targeting computers unaffected by typical malware.
Wikipedia is a major source of information utilized by internet users around the globe for fact-checking and access to general, encyclopedic information. For researchers, it offers an unprecedented opportunity to measure how societies respond to events and how our collective perception of the world evolves over time and in response to events. Wikipedia use and the reading patterns of its users reflect our collective interests and the way they are expressed in our search for information – whether as part of fleeting, zeitgeist-fed trends or long-term – on most every topic, from personal to business, through political, health-related, academic and scientific. In a very real sense, events are defined by how we interpret them and how they affect our perception of the context in which they occurred, rendering Wikipedia invaluable for understanding events and their context. This paper introduces WikiShark (www.wikishark.com) – an online tool that allows researchers to analyze Wikipedia traffic and trends quickly and effectively, by (1) instantly querying pageview traffic data; (2) comparing traffic across articles; (3) surfacing and analyzing trending topics; and (4) easily leveraging findings for use in their own research.
Purpose To compare the clinical outcomes after thoracic endovascular aortic repair (TEVAR) with a bare stent to those after TEVAR alone in patients with complicated acute type B aortic dissection (cATBAD). Materials and Methods A prospective, randomized trial was conducted at 2 medical centers in China between 2010 and 2013. Patients with cATBAD were randomly assigned to receive TEVAR with a bare stent (n=42) or TEVAR only (n=42). Patients were scheduled to undergo computed tomography angiography at 3, 6, and 12 months and then annually to 5 years. The primary endpoint was all-cause mortality at 5 years; secondary outcomes were a composite of complications (endoleak, stent-graft–induced new entry, aortic rupture, and secondary intervention) and aortic remodeling at 1 and 5 years. Results All-cause death occurred in 1 (2.4%) patient in the TEVAR with bare stent group (lung cancer) and 5 patients (11.9%) in the TEVAR group (4 aorta-related) during the 5-year follow-up (log-rank p=0.025). The 1- and 5-year rates of complications and secondary interventions did not differ between the groups. Patients in the TEVAR with bare stent group had higher increases in the thoracic true lumen diameter (19.7±3.6 vs 17.0±6.2 mm, p=0.018) and abdominal true lumen diameter (13.7±4.8 vs 7.2±6.1 mm, p<0.001) and a higher incidence of complete false lumen thrombosis (80.9% vs 47.6%, p=0.005) at the 1-year follow-up. However, no between-group differences in the changes of aortic remodeling parameters were observed between the 1- and 5-year follow-up periods. Conclusion The addition of a distal bare stent to a thoracic stent-graft during TEVAR was associated with significantly improved long-term survival in cATBAD patients vs TEVAR only, likely due to the prevention of true lumen collapse and improvement of complete false lumen thrombosis of the dissected aorta.
A general conjecture is that successful products attain their popularity through influence of adopters on their peers and product information disseminating over the social network. Indeed, many studies have confirmed the existence of local peer effects and contagion. But others have shown that peer influence has a marginal, if any, effect on cascades of adoptions. In this work, we study this discrepancy by analyzing video games propagating over the social network of gamers on Steam, the world's largest video game platform. A major identification problem – distinguishing homophily from peer influence – is a challenge in any peer influence study based on observational data. To overcome it, we introduce a novel method, Revealed Preference-based Matching Estimation, that estimates the impact of peer influence on adoption by using unsupervised machine-learning algorithm to match product adopters to users based solely on similarity of their past adoption. This procedure is applied to thousands of products and reveals how peer influence changes over their lifecycle, thus allowing us to draw general conclusions about the entire ecosystem. Results show that most products belong to one of two distinct groups, each exhibiting a characteristic temporal pattern of adoption: products that exhibit substantial peer influence; and products that do not, for which adoption is driven by preferences. Considering the reach of products in each group, surprisingly, we found that local peer effects are stronger in less popular products. Even more surprising is the fact that almost all blockbusters (products adopted by millions of users) did not exhibit substantial peer influence at any stage of their lifecycle. These results shed light on the discrepancy between observed local peer effects and the lack of peer influence in large adoption cascades that are characteristic of successful products.
In computational social science, epidemic-inspired spread models have been widely used to simulate information diffusion. However, recent empirical studies suggest that simple epidemic-like models typically fail to generate the structure of real-world diffusion trees. Such discrepancy calls for a better understanding of how information spreads from person to person in real-world social networks. Here, we analyse comprehensive diffusion records and associated social networks in three distinct online social platforms. We find that the diffusion probability along a social tie follows a power-law relationship with the numbers of disseminator’s followers and receiver’s followees. To develop a more realistic model of information diffusion, we incorporate this finding together with a heterogeneous response time into a cascade model. After adjusting for observational bias, the proposed model reproduces key structural features of real-world diffusion trees across the three platforms. Our finding provides a practical approach to designing more realistic generative models of information diffusion.
Open collaboration platforms have fundamentally changed the way that knowledge is produced, disseminated, and consumed. Although the community governance and open collaboration model of Wikipedia confers many benefits, its decentralized nature can leave questions of information poverty and skewness to the mercy of the system's natural dynamics. In this paper, we leverage a large-scale natural experiment to gain a causal understanding of how exogenous content contributions to Wikipedia articles affect the attention that they attract and how that attention spills over to other articles in the information network. We find a positive feedback loop: content contribution leads to significant and long-lasting increases of attention and future contribution. Unfortunately, this also suggests that impoverished regions of information networks are likely to remain so in the absence of intervention. However, our analysis reveals a potential solution. Articles in impoverished regions of information networks are particularly positioned to benefit from the phenomenon of attention spillovers. Using a simulation that is calibrated with real-world link traffic of the Wikipedia network, we show that an attention contagion policy, which focuses editorial effort coherently on impoverished regions, can lead to as much as a twofold gain in attention relative to unguided contributions.
Regulators, practitioners, and researchers are expressing growing concern over the readability of financial disclosures. Several recent regulatory guidelines are aimed specifically at simplifying the language of financial reporting in order to ensure that the reports can be read and understood by the public at large. However, quantitative scientific evidence of the evolving linguistic complexity of finance is scarce. In this work we introduce various methods for measuring the linguistic complexity of financial texts. Some of these methods rely on advanced Natural Language Processing (NLP) techniques unavailable when earlier studies were conducted. We apply these methods to decades’ worth of texts from various domains. We have found that 10-K reports have grown substantially longer, more complex and less readable. In terms of education necessary to comprehend them, this increased complexity translates to additional 2.2 years of schooling over the course of 18 years. A similar pattern of using highly complex language is also evident in financial news. The latter finding suggests that financial reality is becoming increasingly difficult to describe and explain. In contrast, the language of other corpora, including general news and scientific publications, has become more readable over the same period of time.
Identity cues – pseudonyms or real names, often displayed with profile photos – appear ubiquitously alongside content in social media. In this paper, we seek to understand to what extent these cues affect how people form opinions about content they consume online. We present results from a large scale (N=1.7×10^7), two-year-long field experiment with a novel anonymization condition. We assigned content produced on a social news discussion website to “identified�? and “anonymous�? conditions to estimate the causal effect of identity cues on how viewers interact with content in terms of ratings and reply comments. Our results show that identity cues cause people to rate content faster (consistent with heuristic processing) and to rate according to con- tent producers’ production, reputation, and reciprocal ratings with content viewers. Our results provide insight into the evolution of status in online communities and evidence for rich-get-richer dynamics that are mediated by identity cues. The methods we use can help platform providers detect and correct for an important source of bias in crowdsourced ratings that cause inequality in feedback.
Community structure and its detection in complex networks has been the subject of many studies in the recent years. Towards this goal, we have created a novel approach based on the analysis of the motion of a memory-biased random walker, i.e. an entity that traverses the network with some tendency to follow or avoid pathways it has previously traversed. We found that the walker tends to remain inside communities, that is, subsets of the network nodes which are more connected to each other, rather than to the rest of the network. Based on this trait of the MBRW we developed a method to detect communities and tested its performance on a range of networks with different levels of community structure. In all tested cases, the method proved to be at least as effective as Girvan–Newman or Infomap while outperforming them when communities were less well defined.
People share billions of pieces of content such as news, videos, and photos through social media every day. Marketers are interested in the extent to which such content propagates and, importantly, which factors make widespread propagation more likely. Extant research considers various factors, such as content attributes (e.g., newness), source traits (e.g., expertise), and network structure (e.g., connectivity). This research builds on prior work by introducing a novel behavior-focused transmitter characteristic that is positively associated with content propagation in social media: activity, or how frequently a person transmits content. Evidence for this effect comes from five studies and different paradigms. First, two studies using data from large social media platforms (Twitter and LiveJournal) show that content posted by higher-activity transmitters — whom we refer to as “social pumps” — propagates more than content posted by lower-activity transmitters. Second, three experiments explore the mechanism driving this effect, showing that social media users receiving content from a social pump are more likely to retransmit it (a necessary behavior for achieving aggregate-level propagation) because they infer that content from a social pump is more likely to be current, and therefore more attractive as something to pass along through retransmission.
Most analyses of the social structure of a network implicitly assume that the relationships in the network are relatively stable. We present evidence that this is not the case. The focal network of this study grew in bursts rather than monotonously over time, and the bursts were highly localized. Links were added and deleted in nearby localities and are not randomly dispersed throughout the network. Also changes in structure lead to simultaneous changes in self-stated interests of its members. For SNA marketing applications the findings suggest interesting improvements. Local bursts around a seed can change the structure of the network dramatically and therefore a marketer’s influence and his chances of success. Therefore, network measurements should be carried out more frequently and closer to the actual implementation of a seeding campaign. To detect these abrupt, dramatic local changes marketers also use a finer resolution. Further, recommendation algorithms that simultaneously account for changes in network structure and content should be applied.
Models of network evolution are based on the implicit assumption that network growth is continuous, uniform, and steady. Using the data collected from a large online-blogging platform, we show that the addition and removal of network ties by users do not occur sporadically at isolated nodes spread all over the network, as assumed by the vast majority of stochastic network models, but rather occur in brief bursts of intense local activity.These bursts of network growth and attrition (addition and removal of network ties) are highly localized around focal nodes. Such network changes coincide with nearly instantaneous densification of the ties between the affected nodes, resulting in an increase of local clustering. Furthermore, we find that these network changes are tightly coupled to the dynamics of individual attributes, particularly the increase in homology between neighboring nodes (homophily) within the scope of the burst. Coincidence of the localized network change with the increase in homophily suggests a strong coupling between the selection and influence processes that lead to simultaneous elevation of assortativity and clustering.
The probability distribution of number of ties of an individual in a social network follows a scale-free power-law. However, how this distribution arises has not been conclusively demonstrated in direct analyses of people's actions in social networks. Here, we perform a causal inference analysis and find an underlying cause for this phenomenon. Our analysis indicates that heavy-tailed degree distribution is causally determined by similarly skewed distribution of human activity. Specifically, the degree of an individual is entirely random - following a "maximum entropy attachment" model - except for its mean value which depends deterministically on the volume of the users' activity. This relation cannot be explained by interactive models, like preferential attachment, since the observed actions are not likely to be caused by interactions with other people.