This study strengthens the foundations of multi-venue market modeling by attempting an independent replication of Wah and Wellman's 2016 model of latency arbitrage in a fragmented market. We find that faithful replication is hindered by missing implementation details in the original paper and limited quantitative reporting. We demonstrate that increasing the number of simulation runs beyond the original design allows for the creation of bootstrap confidence intervals to support rigorous tests of quantitative alignment, compensating for lacking distributional information (e.g. variance). We also demonstrate that increased complexity across the modeled scenarios corresponds with increased difficulty aligning to the original results. We draw on a codebase released by the original authors in connection with a later paper to recover additional implementation details; however, we reject quantitative alignment between that codebase and the published results. Combining information from the paper and the released code, we achieve relational equivalence for most metrics but reject quantitative alignment for model settings where latency is non-zero. We show that many of the qualitative takeaways from the original paper on the effects of market fragmentation and latency arbitrage are sensitive to the specifics of a `greedy strategy' extension given to the zero-intelligence (ZI) trader agents. Under an alternative interpretation of this strategy, we find that market fragmentation decreases execution times in all experiments and increases trader welfare in most experiments. Finally, to facilitate future replication, critique, and extension, we provide an ODD (Overview, Design concepts, Details) protocol for our implementations of the model.
In 2001, Rama Cont introduced a now-widely used set of 'stylized facts' to synthesize empirical studies of financial time series, resulting in 11 qualitative properties presumed to be universal to all financial markets. Here, we replicate Cont's analyses for a convenience sample of stocks drawn from the U.S. stock market following a fundamental shift in market regulation. Our study relies on the same authoritative data as that used by the U.S. regulator. We find conclusive evidence in the modern market for eight of Cont's original facts, while we find weak support for one additional fact and no support for the remaining two. Our study represents the first test of the original set of 11 stylized facts against the same stocks, therefore providing insight into how Cont's stylized facts should be viewed in the context of modern stock markets.
We present our Agent-Based Market Microstructure Simulation (ABMMS), an Agent-Based Financial Market (ABFM) that captures much of the complexity present in the US National Market System for equities (NMS). Agent-Based models are a natural choice for understanding financial markets. Financial markets feature a constrained action space that should simplify model creation, produce a wealth of data that should aid model validation, and a successful ABFM could strongly impact system design and policy development processes. Despite these advantages, ABFMs have largely remained an academic novelty. We hypothesize that two factors limit the usefulness of ABFMs. First, many ABFMs fail to capture relevant microstructure mechanisms, leading to differences in the mechanics of trading. Second, the simple agents that commonly populate ABFMs do not display the breadth of behaviors observed in human traders or the trading systems that they create. We investigate these issues through the development of ABMMS, which features a fragmented market structure, communication infrastructure with propagation delays, realistic auction mechanisms, and more. As a baseline, we populate ABMMS with simple trading agents and investigate properties of the generated data. We then compare the baseline with experimental conditions that explore the impacts of market topology or meta-reinforcement learning agents. The combination of detailed market mechanisms and adaptive agents leads to models whose generated data more accurately reproduce stylized facts observed in actual markets. These improvements increase the utility of ABFMs as tools to inform design and policy decisions.
The U.S. stock market is one of the largest and most complex marketplaces in the global financial system. Over the past several decades, this market has evolved at multiple structural and temporal scales. New exchanges became active, and others stopped trading, regulations have been introduced and adapted, and technological innovations have pushed the pace of trading activity to blistering speeds. These developments have supported the growth of a rich machine-trading ecology that leads to qualitative differences in trading behavior at human and machine time scales. We conduct a longitudinal analysis of comprehensive market data to quantify nonstationary dynamics throughout this system. We quantify the relationship between fluctuations in the number of active trading venues and realized opportunity costs experienced by market participants. We find that information asymmetries, in the form of quote dislocations, predict market-wide volatility indicators. Lastly, we uncover multiple micro-to-macro level pathways, including those exhibiting evidence of self-organized criticality.
The U.S. stock market, more precisely known as the National Market System (NMS), is fragmented into various trading venues. The heterogeneity across this set of venues spans many dimensions; to include geographic location, price discovery mechanisms and fee structures. The prevailing models in the scientific community lag behind in replicating the complexity of today's NMS. In this study, we introduce a new generation of market model, with an explicit focus on an initial representation of the complexity and heterogeneity described above. As an extension of previous work we present the motivation and an overview of the literature relevant to the study of dynamics in multi-exchange markets. We also employ the ODD + D protocol to document our model formulation and its evolutionary trajectory from its predecessors. Experiments are described which show the relational, structural, equivalence between this model and real-world markets.
Using the most comprehensive source of commercially available data on the US National Market System, we analyze all quotes and trades associated with Dow 30 stocks in calendar year 2016 from the vantage point of a single and fixed frame of reference. We find that inefficiencies created in part by the fragmentation of the equity marketplace are relatively common and persist for longer than what physical constraints may suggest. Information feeds reported different prices for the same equity more than 120 million times, with almost 64 million dislocation segments featuring meaningfully longer duration and higher magnitude. During this period, roughly 22% of all trades occurred while the SIP and aggregated direct feeds were dislocated. The current market configuration resulted in a realized opportunity cost totaling over $160 million, a conservative estimate that does not take into account intra-day offsetting events.
Using the most comprehensive, commercially-available dataset of trading activity in U.S. equity markets, we catalog and analyze quote dislocations between the SIP National Best Bid and Offer (NBBO) and a synthetic BBO constructed from direct feeds. We observe a total of over 3.1 billion dislocation segments in the Russell 3000 during trading in 2016, roughly 525 per second of trading. However, these dislocations do not occur uniformly throughout the trading day. We identify a characteristic structure that features more dislocations near the open and close. Additionally, around 23% of observed trades executed during dislocations. These trades may have been impacted by stale information, leading to estimated opportunity costs on the order of $ 2 billion USD. A subset of the constituents of the S&P 500 index experience the greatest amount of opportunity cost and appear to drive inefficiencies in other stocks. These results quantify impacts of the physical structure of the U.S. National Market System.
Correction to: Scientific Reports 7: Article number: 44499; published online: 20 March 2017; updated: 19 March 2018 The original HTML version of this Article contained typographical errors in the legend of Figure 2. “In our smart grid models, the initiating failure ① potentially causes overloads ③ which causes an edge failure and #x02464; a loss of power at the “sink” node.
Both the scientific community and the popular press have paid much attention to the speed of the Securities Information Processor—the data feed consolidating all trades and quotes across the US stock market. Rather than the speed of the Securities Information Processor (SIP), we focus here on its accuracy. Relying on Trade and Quote data, we provide various measures of SIP latency relative to high-speed data feeds between exchanges, known as direct feeds. We use first differences to highlight not only the divergence between the direct feeds and the SIP, but also the fundamental inaccuracy of the SIP. We find that as many as 60% or more of trades are reported out of sequence for stocks with high trade volume, therefore skewing simple measures, such as returns. While not yet definitive, this analysis supports our preliminary conclusion that the underlying infrastructure of the SIP is currently unable to keep pace with the trading activity in today’s stock market.
This study addresses a critical regulatory shortfall by developing a platform to extend stress testing from a microprudential approach to a dynamic, macroprudential approach. This paper describes the ensuing agent-based model for analyzing the vulnerability of the financial system to asset- and funding-based fire sales. The model captures the dynamic interactions of agents in the financial system extending from the suppliers of funding through the intermediation and transformation functions of the bank/dealers to the financial institutions that use the funds to trade in the asset markets. The model replicates the key finding that it is the reaction to initial losses, rather than the losses themselves, that determine the extent of a crisis. By building on a detailed mapping of the transformations and dynamics of the financial system, the agent-based model provides an avenue toward risk management that can illuminate the pathways for the propagation of key crisis dynamics such as fire sales and funding runs.
Increased interconnection between critical infrastructure networks, such as electric power and communications systems, has important implications for infrastructure reliability and security. Others have shown that increased coupling between networks that are vulnerable to internetwork cascading failures can increase vulnerability. However, the mechanisms of cascading in these models differ from those in real systems and such models disregard new functions enabled by coupling, such as intelligent control during a cascade. This paper compares the robustness of simple topological network models to models that more accurately reflect the dynamics of cascading in a particular case of coupled infrastructures. First, we compare a topological contagion model to a power grid model. Second, we compare a percolation model of internetwork cascading to three models of interdependent power-communication systems. In both comparisons, the more detailed models suggest substantially different conclusions, relative to the simpler topological models. In all but the most extreme case, our model of a “smart” power network coupled to a communication system suggests that increased power-communication coupling decreases vulnerability, in contrast to the percolation model. Together, these results suggest that robustness can be enhanced by interconnecting networks with complementary capabilities if modes of internetwork failure propagation are constrained.
The emergence and global adoption of social media has rendered possible the real-time estimation of population-scale sentiment, an extraordinary capacity which has profound implications for our understanding of human behavior. Given the growing assortment of sentiment-measuring instruments, it is imperative to understand which aspects of sentiment dictionaries contribute to both their classification accuracy and their ability to provide richer understanding of texts. Here, we perform detailed, quantitative tests and qualitative assessments of 6 dictionary-based methods applied to 4 different corpora, and briefly examine a further 20 methods. We show that while inappropriate for sentences, dictionary-based methods are generally robust in their classification accuracy for longer texts. Most importantly they can aid understanding of texts with reliable and meaningful word shift graphs if (1) the dictionary covers a sufficiently large portion of a given text’s lexicon when weighted by word usage frequency; and (2) words are scored on a continuous scale.
Both the scientific community and the popular press have paid much attention to the speed of the Securities Information Processor - the data feed consolidating all trades and quotes across the US stock market. Rather than the speed of the Securities Information Processor, or SIP, we focus here on its importance to efficient, price discovery. Via extensions to a previous market model, we experiment with four different coupling mechanisms which operate across the US stock market. Of the four, we find that the SIP contributes most to efficient price discovery.
During liquidity shocks such as occur when margin calls force the liquidation of leveraged positions, there is a widening disparity between the reaction speed of the liquidity demanders and the liquidity providers. Those who are forced to sell typically must take action within the span of a day, while those who are providing liquidity do not face similar urgency. Indeed, the flurry of activity and increased volatility of prices during the liquidity shocks might actually reduce the speed with which many liquidity providers come to the market. To analyze these dynamics, we build upon previous agent-based models of financial markets, and specifically the Preis et. al (Europhys Lett 75(3):510–516, 2006) model, to develop an order-book model with heterogeneity in trader decision cycles. The model demonstrates an adherence to important stylized facts such as a leptokurtic distribution of returns, decay of autocorrelations over moderate to long time lags, and clustering volatility. Consistent with empirical analysis of recent market events, we demonstrate the impact of heterogeneous decision cycles on market resilience and the stochastic properties of market prices.
During liquidity shocks such as occur when margin calls force the liquidation of leveraged positions, there is a widening disparity between the reaction speed of the liquidity demanders and the liquidity providers. Those who are forced to sell typically must take action within the span of a day, while those who are providing liquidity do not face similar urgency. Indeed, the flurry of activity and increased volatility of prices during the liquidity shocks might actually reduce the speed with which many liquidity providers come to the market. To analyze these dynamics, we build upon previous agent-based models of financial markets to develop an order-book model with heterogeneity in trader decision cycles. The model demonstrates an adherence to important stylized facts such as a leptokurtic distribution of returns, decay of autocorrelations over moderate to long time lags, and clustering volatility. We show that the heterogeneity in decision cycles can increase the severity of market shocks, and even absent a shock can have notable effects on the stochastic properties of market prices.
The emergence and global adoption of social media has rendered possible the real-time estimation of population-scale sentiment, bearing profound implications for our understanding of human behavior. Given the growing assortment of sentiment measuring instruments, comparisons between them are evidently required. Here, we perform detailed tests of 6 dictionary-based methods applied to 4 different corpora, and briefly examine a further 20 methods. We show that a dictionary-based method will only perform both reliably and meaningfully if (1) the dictionary covers a sufficiently large enough portion of a given text's lexicon when weighted by word usage frequency; and (2) words are scored on a continuous scale.
We demonstrate that the concerns expressed by Garcia et al. are misplaced, due to (1) a misreading of our findings in [1]; (2) a widespread failure to examine and present words in support of asserted summary quantities based on word usage frequencies; and (3) a range of misconceptions about word usage frequency, word rank, and expert-constructed word lists. In particular, we show that the English component of our study compares well statistically with two related surveys, that no survey design influence is apparent, and that estimates of measurement error do not explain the positivity biases reported in our work and that of others. We further demonstrate that for the frequency dependence of positivity---of which we explored the nuances in great detail in [1]---Garcia et al. did not perform a reanalysis of our data---they instead carried out an analysis of a different, statistically improper data set and introduced a nonlinearity before performing linear regression.
Using human evaluation of 100,000 words spread across 24 corpora in 10 languages diverse in origin and culture, we present evidence of a deep imprint of human sociality in language, observing that (1) the words of natural human language possess a universal positivity bias; (2) the estimated emotional content of words is consistent between languages under translation; and (3) this positivity bias is strongly independent of frequency of word usage. Alongside these general regularities, we describe inter-language variations in the emotional spectrum of languages which allow us to rank corpora. We also show how our word evaluations can be used to construct physical-like instruments for both real-time and offline measurement of the emotional content of large-scale texts.
Society's techno-social systems are becoming ever faster and more computer-orientated. However, far from simply generating faster versions of existing behaviour, we show that this speed-up can generate a new behavioural regime as humans lose the ability to intervene in real time. Analyzing millisecond-scale data for the world's largest and most powerful techno-social system, the global financial market, we uncover an abrupt transition to a new all-machine phase characterized by large numbers of subsecond extreme events. The proliferation of these subsecond events shows an intriguing correlation with the onset of the system-wide financial collapse in 2008. Our findings are consistent with an emerging ecology of competitive machines featuring 'crowds' of predatory algorithms, and highlight the need for a new scientific theory of subsecond financial phenomena.
Complex, dynamic networks underlie many systems, and understanding these networks is the concern of a great span of important scientific and engineering problems. Quantitative description is crucial for this understanding yet, due to a range of measurement problems, many real network datasets are incomplete. Here we explore how accidentally missing or deliberately hidden nodes may be detected in networks by the e ect of their absence on predictions of the speed with which information flows through the network. We use Symbolic Regression (SR) to learn models relating information flow to network topology. These models show localized, systematic, and nonrandom discrepancies when applied to test networks with intentionally masked nodes, demonstrating the ability to detect the presence of missing nodes and where in the network those nodes are likely to reside.