Individuals in social systems are embedded in collective decision-making hierarchies, such as households, neighborhoods, communities, organizations, etc. The locus of agency in such systems is dispersed across the system, and can variously be viewed as individual, distributed, and shared agency. Here we propose a general notion of network agency that subsumes these descriptions and also allows for integrating related notions, such as peer influence. In our view, the social system can be seen as a multi-layer network, where each layer corresponds to different aggregations of the underlying units, representing different kinds of perception and decision-making. We illustrate this general framework with an agent-based model of the ongoing forced migration from Ukraine. In our model, individuals perceive hazards (conflict events), but decisions to migrate are taken at the household level, where peer influence from other households in the neighborhood is also taken into account. We present this model in detail to elucidate our concept of network agency. We also calibrate the model with data on daily refugee flows and show that our model is able to estimate the scale of the daily refugee flow from Ukraine for the first two months with a Root Mean Squared Percentage Error (RMSPE) of 0.24, outperforming state-of-the-art, which had an RMSPE of 0.77. Moreover, our model also captures the daily trend of outflow with a Pearson Correlation Coefficient (PCC) of 0.98. We also perform sensitivity analysis of the model and analyze the significant parameters of the model, which in turn tells us how different agencies are significant in different contexts.
We present MacKenzie, a HPC-driven multi-cluster workflow system that was used repeatedly to configure and execute fine-grained US national-scale epidemic simulation models during the COVID-19 pandemic. Mackenzie supported federal and Virginia policymakers, in real-time, for a large number of “what-if” scenarios during the COVID-19 pandemic, and continues to be used to answer related questions as COVID-19 transitions to the endemic stage of the disease. MacKenzie is a novel HPC meta-scheduler that can execute US-scale simulation models and associated workflows that typically present significant big data challenges. The meta-scheduler optimizes the total execution time of simulations in the workflow, and helps improve overall human productivity.As an exemplar of the kind of studies that can be conducted using Mackenzie, we present a modeling study to understand the impact of vaccine-acceptance in controlling the spread of COVID-19 in the US. We use a 288 million node synthetic social contact network (digital twin) spanning all 50 US states plus Washington DC, comprised of 3300 counties, with 12 billion daily interactions. The highly-resolved agent-based model used for the epidemic simulations uses realistic information about disease progression, vaccine uptake, production schedules, acceptance trends, prevalence, and social distancing guidelines. Computational experiments show that, for the simulation workload discussed above, MacKenzie is able to scale up well to 10 K CPU cores.Our modeling results show that, when compared to faster and accelerating vaccinations, slower vaccination rates due to vaccine hesitancy cause averted infections to drop from 6.7M to 4.5M, and averted total deaths to drop from 39.4 K to 28.2 K across the US. This occurs despite the fact that the final vaccine coverage is the same in both scenarios. We also find that if vaccine acceptance could be increased by 10% in all states, averted infections could be increased from 4.5M to 4.7M (a 4.4% improvement) and total averted deaths could be increased from 28.2 K to 29.9 K (a 6% improvement) nationwide.
In this article we develop an information-theoretic framework of multiple sequence alignments (MSAs), based on sub-sampling. The key component of this framework is an information-theoretical potential defined on pairs of sites (links) within the MSA. This potential quantifies the expected drop in variation of information between the two constituent sites. The expectation is taken with respect to all possible sub-alignments, obtained by removing a finite, fixed number of rows. We show that the potential is zero for linked sites representing columns, for which symbols are in bijective correspondence and that it is strictly positive, otherwise. It is furthermore shown that the potential assumes its unique minimum for links at which each symbol pair appears with the same multiplicity. We then show that the established drop of the variation of information exceeds finite-size effects inherent to the construction of the potential. Finally, we provide as a proof of concept an application of our results to a specific MSA composed of the inverse fold solutions of three distinguished secondary structures.
The ongoing Russian aggression against Ukraine has forced over eight million people to migrate out of Ukraine. Understanding the dynamics of forced migration is essential for policy-making and for delivering humanitarian assistance. Existing work is hindered by a reliance on observational data which is only available well after the fact. In this work, we study the efficacy of a data-driven agent-based framework motivated by social and behavioral theory in predicting outflow of migrants as a result of conflict events during the initial phase of the Ukraine war. We discuss policy use cases for the proposed framework by demonstrating how it can leverage refugee demographic details to answer pressing policy questions. We also show how to incorporate conflict forecast scenarios to predict future conflict-induced migration flows. Detailed future migration estimates across various conflict scenarios can both help to reduce policymaker uncertainty and improve allocation and staging of limited humanitarian resources in crisis settings.
A new perspective is introduced regarding the analysis of Multiple Sequence Alignments (MSA), representing aligned data defined over a finite alphabet of symbols. The framework is designed to produce a block decomposition of an MSA, where each block is comprised of sequences exhibiting a certain site-coherence. The key component of this framework is an information theoretical potential defined on pairs of sites (links) within the MSA. This potential quantifies the expected drop in variation of information between the two constituent sites, where the expectation is taken with respect to all possible sub-alignments, obtained by removing a finite, fixed collection of rows. It is proved that the potential is zero for linked sites representing columns, whose symbols are in bijective correspondence and it is strictly positive, otherwise. It is furthermore shown that the potential assumes its unique minimum for links at which each symbol pair appears with the same multiplicity. Finally, an application is presented regarding anomaly detection in an MSA, composed of inverse fold solutions of a fixed tRNA secondary structure, where the anomalies are represented by inverse fold solutions of a different RNA structure.
Large-scale population displacements arising from conflict-induced forced migration generate uncertainty and introduce several policy challenges. Addressing these concerns requires an interdisciplinary approach that integrates knowledge from both computational modeling and social sciences. We propose a generalized computational agent-based modeling framework grounded by Theory of Planned Behavior to model conflict-induced migration outflows within Ukraine during the start of that conflict in 2022. Existing migration modeling frameworks that attempt to address policy implications primarily focus on destination while leaving absent a generalized computational framework grounded by social theory focused on the conflict-induced region. We propose an agent-based framework utilizing a spatiotemporal gravity model and a Bi-threshold model over a Graph Dynamical System to update migration status of agents in conflict-induced regions at fine temporal and spatial granularity. This approach significantly outperforms previous work when examining the case of Russian invasion in Ukraine. Policy implications of the proposed framework are demonstrated by modeling the migration behavior of Ukrainian civilians attempting to flee from regions encircled by Russian forces. We also showcase the generalizability of the model by simulating a past conflict in Burundi, an alternative conflict setting. Results demonstrate the utility of the framework for assessing conflict-induced migration in varied settings as well as identifying vulnerable civilian populations.
We present a novel framework enhancing the prediction of whether novel lineage poses the threat of eventually dominating the viral population. The framework is based purely on genomic sequence data, without requiring prior established biological analysis. Its building blocks are sets of coevolving sites in the alignment (motifs), identified via coevolutionary signals. The collection of such motifs forms a relational structure over the polymorphic sites. Motifs are constructed using distances quantifying the coevolutionary coupling of pairs and manifest as coevolving clusters of sites. We present an approach to genomic surveillance based on this notion of relational structure. Our system will issue an alert regarding a lineage, based on its contribution to drastic changes in the relational structure. We then conduct a comprehensive retrospective analysis of the COVID-19 pandemic based on SARS-CoV-2 genomic sequence data in GISAID from October 2020 to September 2022, across 21 lineages and 27 countries with weekly resolution. We investigate the performance of this surveillance system in terms of its accuracy, timeliness, and robustness. Lastly, we study how well each lineage is classified by such a system.
Disease surveillance systems provide early warnings of disease outbreaks before they become public health emergencies. However, pandemics containment would be challenging due to the complex immunity landscape created by multiple variants. Genomic surveillance is critical for detecting novel variants with diverse characteristics and importation/emergence times. Yet, a systematic study incorporating genomic monitoring, situation assessment, and intervention strategies is lacking in the literature. We formulate an integrated computational modeling framework to study a realistic course of action based on sequencing, analysis, and response. We study the effects of the second variant's importation time, its infectiousness advantage and, its cross-infection on the novel variant's detection time, and the resulting intervention scenarios to contain epidemics driven by two-variants dynamics. Our results illustrate the limitation in the intervention's effectiveness due to the variants' competing dynamics and provide the following insights: i) There is a set of importation times that yields the worst detection time for the second variant, which depends on the first variant's basic reproductive number; ii) When the second variant is imported relatively early with respect to the first variant, the cross-infection level does not impact the detection time of the second variant. We found that depending on the target metric, the best outcomes are attained under different interventions' regimes. Our results emphasize the importance of sustained enforcement of Non-Pharmaceutical Interventions on preventing epidemic resurgence due to importation/emergence of novel variants. We also discuss how our methods can be used to study when a novel variant emerges within a population.
A synthetic population is a simplified microscopic representation of an actual population. Statistically representative at the population level, it provides valuable inputs to simulation models (especially agent-based models) in research areas such as transportation, land use, economics, and epidemiology. This article describes the datasets from the Synthetic Sweden Mobility (SySMo) model using the state-of-art methodology, including machine learning (ML), iterative proportional fitting (IPF), and probabilistic sampling. The model provides a synthetic replica of over 10 million Swedish individuals (i.e., agents), their household characteristics, and activity-travel plans. This paper briefly explains the methodology for the three datasets: Person, Households, and Activity-travel patterns. Each agent contains socio-demographic attributes, such as age, gender, civil status, residential zone, personal income, car ownership, employment, etc. Each agent also has a household and corresponding attributes such as household size, number of children ≤ 6 years old, etc. These characteristics are the basis for the agents’ daily activity-travel schedule, including type of activity, start-end time, duration, sequence, the location of each activity, and the travel mode between activities.
We propose a novel mathematical paradigm for the study of genetic variation in sequence alignments. This framework originates from extending the notion of pairwise relations, upon which current analysis is based on, to k-ary dissimilarity. This dissimilarity naturally leads to a generalization of simplicial complexes by endowing simplices with weights, compatible with the boundary operator. We introduce the notion of k-stances and dissimilarity complex, the former encapsulating arithmetic as well as topological structure expressing these k-ary relations. We study basic mathematical properties of dissimilarity complexes and show how this approach captures watershed moments of viral dynamics in the context of SARS-CoV-2 and H1N1 flu genomic data.
We present a novel framework facilitating the rapid detection of variants of interest (VOI) and concern (VOC) in a viral multiple sequence alignment (MSA). The framework is purely based on the genomic sequence data, without requiring prior established biological analysis. The framework’s building blocks are sets of co-evolving sites (motifs), identified via co-evolutionary signals within the MSA. Motifs form a weighted simplicial complex, whose vertices are sites that satisfy a certain nucleotide diversity. Higher dimensional simplices are constructed using distances quantifying the co-evolutionary coupling of pairs and in the context of our method maximal motifs manifest as clusters. The framework triggers an alert via a cluster with a significant fraction of newly emerging polymorphic sites. We apply our method to SARS-CoV-2, analyzing all alerts issued from November 2020 through August 2021 with weekly resolution for England, USA, India and South America. Within a week at most a handful of alerts, each of which involving on the order of 10 sites are triggered. Cross referencing alerts with a posteriori knowledge of VOI/VOC-designations and lineages, motif-induced alerts detect VOIs/VOCs rapidly, typically weeks earlier than current methods. We show how motifs provide insight into the organization of the characteristic mutations of a VOI/VOC, organizing them as co-evolving blocks. Finally we study the dependency of the motif reconstruction on metric and clustering method and provide the receiver operating characteristic (ROC) of our alert criterion.
Non-pharmaceutical interventions (NPIs) constitute the front-line responses against epidemics. Yet, the interdependence of control measures and individual microeconomics, beliefs, perceptions and health incentives, is not well understood. Epidemics constitute complex adaptive systems where individual behavioral decisions drive and are driven by, among other things, the risk of infection. To study the impact of heterogeneous behavioral responses on the epidemic burden, we formulate a two risk-groups mathematical model that incorporates individual behavioral decisions driven by risk perceptions. Our results show a trade-off between the efforts to avoid infection by the risk-evader population, and the proportion of risk-taker individuals with relaxed infection risk perceptions. We show that, in a structured population, privately computed optimal behavioral responses may lead to an increase in the final size of the epidemic, when compared to the homogeneous behavior scenario. Moreover, we find that uncertain information on the individuals' true health state may lead to worse epidemic outcomes, ultimately depending on the population's risk-group composition. Finally, we find there is a set of specific optimal planning horizons minimizing the final epidemic size, which depend on the population structure.
This paper presents a novel virus surveillance framework, completely independent of phylogeny-based methods. The framework issues timely alerts with an accuracy exceeding 85% that are based on the co-evolutionary relations between sites of the viral multiple sequence array (MSA). This set of relations is formalized via a motif complex, whose dynamics contains key information about the emergence of viral threats without the referencing of strain prevalence. Our notion of threat is centered at the emergence of a certain type of critical cluster consisting of key co-evolving sites. We present three case studies, based on GISAID data from UK, US and New York, where we perform our surveillance. We alert on May 16, 2022, based on GISAID data from New York, to a critical cluster of co-evolving sites mapping to the Pango-designation, BA.5. The alert specifies a cluster of seven genomic sites, one of which exhibits D3N on the M (membrane) protein–the distinguishing mutation of BA.5, three encoding ORF6:D61L and the remaining three exhibiting the synonymous mutations C26858T, C27889T and A27259C. New insight is obtained: when projected onto sequences, this cluster splits into two, mutually exclusive blocks of co-evolving sites (m:D3N,nuc:C27889T) linked to the five reverse mutations (nuc:C26858T,nuc:A27259C,ORF6:D61L). We furthermore provide an in depth analysis of all major signaled threats, during which we discover a specific signature concerning linked reverse mutation in the critical cluster.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis work was partially supported by the VDH Grant PV-BII VDH COVID-19 Modeling Program VDH-21-501-0135.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesAll data produced in the present study are available upon reasonable request to the authors
Responding to a rapidly evolving pandemic like COVID-19 is challenging, and involves anticipating novel variants, vaccine uptake, and behavioral adaptations. Human judgment systems can complement computational models by providing valuable real-time forecasts. We report findings from a study conducted on Metaculus, a community forecasting platform, in partnership with the Virginia Department of Health, involving six rounds of forecasting during the Omicron BA.1 wave in the United States from November 2021 to March 2022. We received 8355 probabilistic predictions from 129 unique users across 60 questions pertaining to cases, hospitalizations, vaccine uptake, and peak/trough activity. We observed that the case forecasts performed on par with national multi-model ensembles and the vaccine uptake forecasts were more robust and accurate compared to baseline models. We also identified qualitative shifts in Omicron BA.1 wave prognosis during the surge phase, demonstrating rapid adaptation of such systems. Finally, we found that community estimates of variant characteristics such as growth rate and timing of dominance were in line with the scientific consensus. The observed accuracy, timeliness, and scope of such systems demonstrates the value of incorporating them into pandemic policymaking workflows.
We study allocation of COVID-19 vaccines to individuals based on the structural properties of their underlying social contact network. Using a realistic representation of a social contact network for the Commonwealth of Virginia, we study how a limited number of vaccine doses can be strategically distributed to individuals to reduce the overall burden of the pandemic. We show that allocation of vaccines based on individuals' degree (number of social contacts) and total social proximity time is significantly more effective than the usually used age-based allocation strategy in reducing the number of infections, hospitalizations and deaths. The overall strategy is robust even: (i) if the social contacts are not estimated correctly; (ii) if the vaccine efficacy is lower than expected or only a single dose is given; (iii) if there is a delay in vaccine production and deployment; and (iv) whether or not non-pharmaceutical interventions continue as vaccines are deployed. For reasons of implementability, we have used degree, which is a simple structural measure and can be easily estimated using several methods, including the digital technology available today. These results are significant, especially for resource-poor countries, where vaccines are less available, have lower efficacy, and are more slowly distributed.
This paper describes an integrated, data-driven operational pipeline based on national agent-based models to support federal and state-level pandemic planning and response. The pipeline consists of (i) an automatic semantic-aware scheduling method that coordinates jobs across two separate high performance computing systems; (ii) a data pipeline to collect, integrate and organize national and county-level disaggregated data for initialization and post-simulation analysis; (iii) a digital twin of national social contact networks made up of 288 Million individuals and 12.6 Billion time-varying interactions covering the US states and DC; (iv) an extension of a parallel agent-based simulation model to study epidemic dynamics and associated interventions. This pipeline can run 400 replicates of national runs in less than 33 h, and reduces the need for human intervention, resulting in faster turnaround times and higher reliability and accuracy of the results. Scientifically, the work has led to significant advances in real-time epidemic sciences.
Human mobility is a primary driver of infectious disease spread. However, existing data is limited in availability, coverage, granularity, and timeliness. Data-driven forecasts of disease dynamics are crucial for decision-making by health officials and private citizens alike. In this work, we focus on a machine-learned anonymized mobility map (hereon referred to as AMM) aggregated over hundreds of millions of smartphones and evaluate its utility in forecasting epidemics. We factor AMM into a metapopulation model to retrospectively forecast influenza in the USA and Australia. We show that the AMM model performs on-par with those based on commuter surveys, which are sparsely available and expensive. We also compare it with gravity and radiation based models of mobility, and find that the radiation model's performance is quite similar to AMM and commuter flows. Additionally, we demonstrate our model's ability to predict disease spread even across state boundaries. Our work contributes towards developing timely infectious disease forecasting at a global scale using human mobility datasets expanding their applications in the area of infectious disease epidemiology.
COVID-19 is an infectious disease caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). The viral genome is considered to be relatively stable and the mutations that have been observed and reported thus far are mainly focused on the coding region. This article provides evidence that macrolevel pandemic dynamics, such as social distancing, modulate the genomic evolution of SARS-CoV-2. This view complements the prevalent paradigm that microlevel observables control macrolevel parameters such as death rates and infection patterns. First, we observe differences in mutational signals for geospatially separated populations such as the prevalence of A23404G in CA versus NY and WA. We show that the feedback between macrolevel dynamics and the viral population can be captured employing a transfer entropy framework. Second, we observe complex interactions within mutational clades. Namely, when C14408T first appeared in the viral population, the frequency of A23404G spiked in the subsequent week. Third, we identify a noncoding mutation, G29540A, within the segment between the coding gene of the N protein and the ORF10 gene, which is largely confined to NY (>95%). These observations indicate that macrolevel sociobehavioral measures have an impact on the viral genomics and may be useful for the dashboard-like tracking of its evolution. Finally, despite the fact that SARS-CoV-2 is a genetically robust organism, our findings suggest that we are dealing with a high degree of adaptability. Owing to its ample spread, mutations of unusual form are observed and a high complexity of mutational interaction is exhibited.
Contact tracing (CT) is an important and effective intervention strategy for controlling an epidemic. Its role becomes critical when pharmaceutical interventions are unavailable. CT is resource intensive, and multiple protocols are possible, therefore the ability to evaluate strategies is important. We describe a high-performance, agent-based simulation model for studying CT during an ongoing pandemic. This work was motivated by the COVID-19 pandemic, however framework and design are generic and can be applied in other settings. This work extends our HPC-oriented ABM framework EpiHiper to efficiently represent contact tracing. The main contributions are: (i) Extension of EpiHiper to represent realistic CT processes. (ii) Realistic case study using the VA network motivated by our collaboration with the Virginia Department of Health.
Samarth Swarup合作论文数Network Dynamics and Simulation Science Lab,
Virginia Bioinformatics Institute,
Virginia Tech19