Bacterial infections are one of the most common causes of sepsis. If not treated quickly, sepsis can result in organ dysfunction, organ failure, or even death. In the presence of bacterial sepsis, it can be challenging to determine whether the infection is caused by Gram-negative or Gram-positive bacteria, which is essential for appropriate and timely treatment. Signaling molecules including cytokines and chemokines can be used as biomarkers of infection and evaluation of the host immune response. Human lung epithelial A549 cells were exposed to various Gram-positive and Gram-negative pathogen associated molecular patterns (PAMPs) including lipopolysaccharide, lipoteichoic acid and Pam3CSK4. A549 cells were treated with these agonists for 24 hours and mRNA expression levels of 84 human cytokines and chemokines were measured to evaluate immune response. Over 30 biological replicates were evaluated for each group and untreated control samples were compared to treated samples to determine upregulation or downregulation of each gene. All three treatment groups stimulated a significant upregulation of five pro-inflammatory cytokines and chemokines: CCL5, CCL20, CCL22, CXCL8, and LTB. The level of expression varied across all three treatment groups, highlighting the diverse cellular responses to three different types of molecules. These immune profiles reveal significant biomarkers of bacterial infection and can potentially provide a technique to quickly diagnose sepsis. This work was supported by the Defense Threat Reduction Agency (DTRA) to H.M. and C.A.M. (Award #CB11015). Cytokines and Chemokines and Their Receptors (CCR)
Mosquitoes are a key virus vector that poses significant health threats globally, affecting 700 million individuals and causing 1 million deaths annually. Accurately predicting mosquito abundance and dispersion remains a challenge. Complex interactions between mosquito dynamics and various environmental factors, notably hydrology, contribute to this challenge. Existing models typically focus on precipitation and temperature and often overlook further impacts of hydrological variables within mosquito modeling. In this study, we developed an artificial intelligence‐based model for mosquito dynamics, explicitly accounting for different hydrological variables, such as precipitation, soil moisture and streamflow. Using Toronto, Canada, as a case study, we identified causal relationships between changes in mosquito populations, hydrological factors, vegetation (e.g., leaf area index), and climate variables (e.g., daylight length, precipitation, and temperature). We embedded these relationships into a Long Short‐Term Memory (LSTM) Neural Network Model capable of accurately detecting mosquito dynamics across annual, seasonal, and monthly time scales. The LSTM is able to explain, on average, approximately 40% of the variance in the observed mosquito abundance data. Using the calibrated model, we predicted that the summer season mosquito abundance would increase by ∼16% and ∼19% under an intermediate greenhouse emission scenario, Shared Socioeconomic Pathway (SSP) 2–4.5, and a high greenhouse emission scenario, SSP5‐8.5, respectively. We expect that this model can serve as a valuable tool and inform science‐based decisions affecting mosquito dynamics and public health. It can also build a foundation for future risk analysis at the regional and larger scales.
Mosquito borne diseases pose a significant health risk for humans. In North America, Culex mosquitoes are a major vector for several diseases including West Nile Virus and St. Louis Encephalitis. In many instances, models used to predict the spread of mosquito borne disease rely on a quantification of mosquito abundance. In our work, we present a novel age-structured partial differential equation model for simulating Culex mosquito abundance. The model is constructed using a system of two dimensional coupled advection reaction equations, in which the first dimension represents the age of the mosquitoes within a growth-stage population and the second dimension is time. We form six mosquito growth-stage populations by subdividing the mosquito life cycle into six stages: egg, larvae/pupae, and four adult gonotrophic cycles. Each growth-stage population is coupled through the boundary conditions on the age of the mosquito, which advances the population through its life cycle. The model also includes a population of diapausing adults represented using an ordinary differential equation. The solution curves for each equation provide the distribution of mosquitoes over time for each growth-stage population. This model provides information on the relative abundance of mosquitoes as well as the abundance of mosquitoes at specific ages. We simulate mosquito abundance for the Greater Ontario Area and compare the simulated adult abundance to mosquito trap count data. The model produces mosquito abundance patterns similar to those observed in trap count data.
The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: 1) type of biochemical signature (transcripts vs. proteins), data curation methods (pre- and post-processing), and 3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly modulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.
Universal and early recognition of pathogens occurs through recognition of evolutionarily conserved pathogen associated molecular patterns (PAMPs) by innate immune receptors and the consequent secretion of cytokines and chemokines. The intrinsic complexity of innate immune signaling and associated signal transduction challenges our ability to obtain physiologically relevant, reproducible and accurate data from experimental systems. One of the reasons for the discrepancy in observed data is the choice of measurement strategy. Immune signaling is regulated by the interplay between pathogen-derived molecules with host cells resulting in cellular expression changes. However, these cellular processes are often studied by the independent assessment of either the transcriptome or the proteome. Correlation between transcription and protein analysis is lacking in a variety of studies. In order to methodically evaluate the correlation between transcription and protein expression profiles associated with innate immune signaling, we measured cytokine and chemokine levels following exposure of human cells to the PAMP lipopolysaccharide (LPS) from the Gram-negative pathogen Pseudomonas aeruginosa. Expression of 84 messenger RNA (mRNA) transcripts and 69 proteins, including 35 overlapping targets, were measured in human lung epithelial cells. We evaluated 50 biological replicates to determine reproducibility of outcomes. Following pairwise normalization, 16 mRNA transcripts and 6 proteins were significantly upregulated following LPS exposure, while only five (CCL2, CSF3, CXCL5, CXCL8/IL8, and IL6) were upregulated in both transcriptomic and proteomic analysis. This lack of correlation between transcription and protein expression data may contribute to the discrepancy in the immune profiles reported in various studies. The use of multiomic assessments to achieve a systems-level understanding of immune signaling processes can result in the identification of host biomarker profiles for a variety of infectious diseases and facilitate countermeasure design and development.
Importance Social and environmental determinants of health (SDOH and EDOH) may contribute significantly to suicide rates among U.S. veterans. Objective To identify key predictive variables for assessing suicide related death rates (SRR), which include suicide deaths, suicide firearm deaths, and suicide nonfirearm deaths and vulnerability areas. Design, Setting, and Participants This case control study utilized Electronic Health Record (EHR) data, which included demographic and mental health information spanning from January 1, 2006, to December 31, 2016. The base cohort considered all veterans from the VHA outpatient database during the above period. Patients from the base cohort who died by suicide were identified through the National Death Index and considered as cases. Given the significantly larger number of alive patients compared to deceased patients, which caused the dataset to be extremely unbalanced and potentially biased, control participants were selected at a ratio of 4 controls to 1 case from those who were still alive. Cases of suicide related death were matched with four controls based on birth year, cohort entry date, sex, and follow up duration. Comprehensive data on social determinants (SDOH), geographic and gun related factors, quality of access to healthcare, environmental determinants (EDOH), and food insecurity were gathered from various sources at the midpoint of the study in 2011. Data analysis was carried out from January 2023 to January 2024. Exposures Suicide related deaths associated with SDOH and EDOH. Main Outcomes and Measures A hierarchical clustering method was employed to downselect the large number of variables, while Cox regression models were used to identify key predictive variables for SRR and areas of vulnerability. Results Out of a total of 9,819,080 veterans, 28,302 were identified as having died by suicide. These cases were matched with 113,208 control participants. The majority of the cohort was male (137,264 [97%]) and White (101,533 [72%]), with a significant portion being Black veterans (18,450 [13.12%]). The average age (SD) was 64.77 (17.56) years. We found that Social Determinants of Health (SDOH) and Environmental Determinants of Health (EDOH) were significantly associated with an increased risk of suicide. By incorporating SDOH and EDOH into the model, the performance (AUC) improved from 0.70 to 0.73. Conclusions and Relevance In this study, veterans who died by suicide using firearms exhibited distinct characteristics based on SDOH and EDOH, particularly in gun related variables, compared to those who died by nonfirearm methods. Our analysis indicated that veterans living in areas with more social issues, higher temperatures, and higher altitudes are at a higher risk of all means suicide. Furthermore, regions such as Montana, Wyoming, West Virgina and Arkansas, characterized by higher gun owernship are predicted to have the highest vulnerability based on veteran suicide firearm rates. Gun ownership and gun laws grades showed as strong predictors rather than rurality. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study was funded by awards # MVP011 (I01-CX001729) and HSRD SDR 21-150 from the U.S. Dept. of Veterans Affairs. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This work was supported by the Department of Veterans Affairs. Work was conducted with the approval of the VA Central IRB under project number MVP011 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
As temperatures change worldwide, the pattern and competency of disease vectors will change, altering the global distribution of both the burden of infectious disease and the risk of the emergence of those diseases into new regions. To evaluate the risk of potential summer dengue outbreaks triggered by infected travelers under various climate scenarios, we develop an SEIR-type model, run numerical simulations, and conduct sensitivity analyses under a range of temperature profiles. Our model extends existing theoretical frameworks for studying dengue dynamics by introducing temperature dependence of two key parameters: the mosquito extrinsic incubation period and the lifespan of mosquitoes, which empirical data suggests are both highly temperature dependent. We find that changing temperature significantly alters dengue risk in an inverted U-shape, with temperatures in the range 27-31°C producing the highest risk. As temperatures increase beyond 31°C, the determinants of dengue risk begin to shift from mosquito biting rate and carrying capacity to the duration of the human infectious period, suggesting that changing temperatures not only alter dengue risk but also the potential efficacy of control measures. To illustrate the role of spatial and temporal temperature heterogeneity, we select five US cities where the primary dengue vector, the mosquito Aedes aegypti, has been observed, and which have had dengue cases in the past: Los Angeles, Houston, Miami, Brownsville, and Phoenix. Our analysis suggests that an increase of 3°C leads to an approximate doubling of the risk of dengue in Los Angeles and Houston, but a reduction of risk in Miami, Brownsville, and Phoenix due to extreme heat.
Climate change is arguably one of the most pressing issues affecting the world today and requires the fusion of disparate data streams to accurately model its impacts. Mosquito populations respond to temperature and precipitation in a nonlinear way, making predicting climate impacts on mosquito-borne diseases an ongoing challenge. Data-driven approaches for accurately modeling mosquito populations are needed for predicting mosquito-borne disease risk under climate change scenarios. Many current models for disease transmission are continuous and autonomous, while mosquito data is discrete and varies both within and between seasons. This study uses an optimization framework to fit a non-autonomous logistic model with periodic net growth rate and carrying capacity parameters for 15 years of daily mosquito time-series data from the Greater Toronto Area of Canada. The resulting parameters accurately capture the inter-annual and intra-seasonal variability of mosquito populations within a single geographic region, and a variance-based sensitivity analysis highlights the influence each parameter has on the peak magnitude and timing of the mosquito season. This method can easily extend to other geographic regions and be integrated into a larger disease transmission model. This method addresses the ongoing challenges of data and model fusion by serving as a link between discrete time-series data and continuous differential equations for mosquito-borne epidemiology models.
The logistic equation has been extensively used to model biological phenomena across a variety of disciplines and has provided valuable insight into how our universe operates. Incorporating time-dependent parameters into the logistic equation allows the modeling of more complex behavior than its autonomous analog, such as a tumor's varying growth rate under treatment, or the expansion of bacterial colonies under varying resource conditions. Some of the most commonly used numerical solvers produce vastly different approximations for a non-autonomous logistic model with a periodically-varying growth rate changing signum. Incorrect, inconsistent, or even unstable approximate solutions for this non-autonomous problem can occur from some of the most frequently used numerical methods, including the lsoda, implicit backwards difference, and Runge-Kutta methods, all of which employ a black-box framework. Meanwhile, a simple, manually-programmed Runge-Kutta method is robust enough to accurately capture the analytical solution for biologically reasonable parameters and consistently produce reliable simulations. Consistency and reliability of numerical methods are fundamental for simulating non-autonomous differential equations and dynamical systems, particularly when applications are physically or biologically informed.
The COVID-19 pandemic has highlighted a need for better understanding of countries’ vulnerability and resilience to not only pandemics but also disasters, climate change, and other systemic shocks. A comprehensive characterization of vulnerability can inform efforts to improve infrastructure and guide disaster response in the future. In this paper, we propose a data-driven framework for studying countries’ vulnerability and resilience to incident disasters across multiple dimensions of society. To illustrate this methodology, we leverage the rich data landscape surrounding the COVID-19 pandemic to characterize observed resilience for several countries (USA, Brazil, India, Sweden, New Zealand, and Israel) as measured by pandemic impacts across a variety of social, economic, and political domains. We also assess how observed responses and outcomes (i.e., resilience) of the COVID-19 pandemic are associated with pre-pandemic characteristics or vulnerabilities, including (1) prior risk for adverse pandemic outcomes due to population density and age and (2) the systems in place prior to the pandemic that may impact the ability to respond to the crisis, including health infrastructure and economic capacity. Our work demonstrates the importance of viewing vulnerability and resilience in a multi-dimensional way, where a country’s resources and outcomes related to vulnerability and resilience can differ dramatically across economic, political, and social domains. This work also highlights key gaps in our current understanding about vulnerability and resilience and a need for data-driven, context-specific assessments of disaster vulnerability in the future.
The dynamics of human infectious diseases are challenging to understand, particularly when a pathogen spreads spatially over a large region. We present a stochastic, spatially-heterogeneous model framework derived from the foundational SEIR compartmental model. These models utilize a graph structure of spatial locations, facilitating mobility via random walks while progressing through disease states, parameterized by the net probability flux between locations. The analysis is bolstered by Approximate Bayesian Computation, by which epidemiological and mobility parameter distributions are estimated, including an empirically adjusted reproductive number, while model structure proposals are compared using Bayes Factors. The utility of this novel class of models is demonstrated through application to the 2014-2016 Ebola outbreak in West Africa. The flexibility of such models, whose complexity may be adjusted as desired, and complementary methods of analysis enable the exploration of various spatial divisions and mobility schema, while maintaining the essential spatiotemporal disease dynamics.
Dengue virus remains a significant public health challenge in Brazil, and seasonal preparation efforts are hindered by variable intra- and interseasonal dynamics. Here, we present a framework for characterizing weekly dengue activity at the Brazilian mesoregion level from 2010–2016 as time series properties that are relevant to forecasting efforts, focusing on outbreak shape, seasonal timing, and pairwise correlations in magnitude and onset. In addition, we use a combination of 18 satellite remote sensing imagery, weather, clinical, mobility, and census data streams and regression methods to identify a parsimonious set of covariates that explain each time series property. The models explained 54% of the variation in outbreak shape, 38% of seasonal onset, 34% of pairwise correlation in outbreak timing, and 11% of pairwise correlation in outbreak magnitude. Regions that have experienced longer periods of drought sensitivity, as captured by the “normalized burn ratio,” experienced less intense outbreaks, while regions with regular fluctuations in relative humidity had less regular seasonal outbreaks. Both the pairwise correlations in outbreak timing and outbreak trend between mesoresgions were best predicted by distance. Our analysis also revealed the presence of distinct geographic clusters where dengue properties tend to be spatially correlated. Forecasting models aimed at predicting the dynamics of dengue activity need to identify the most salient variables capable of contributing to accurate predictions. Our findings show that successful models may need to leverage distinct variables in different locations and be catered to a specific task, such as predicting outbreak magnitude or timing characteristics, to be useful. This advocates in favor of “adaptive models” rather than “one-size-fits-all” models. The results of this study can be applied to improving spatial hierarchical or target-focused forecasting models of dengue activity across Brazil.
54% of the variation in outbreak shape, 38% of seasonal onset, 34% of pairwise correlation in outbreak timing, and 11% of pairwise correlation in outbreak magnitude. Regions that have experienced longer periods of drought sensitivity, as captured by the “normalized burn ratio,” experienced less intense outbreaks, while regions with regular fluctuations in relative humidity had less regular seasonal outbreaks. Both the pairwise correlations in outbreak timing and outbreak trend between mesoresgions were best predicted by distance. Our analysis also revealed the presence of distinct geographic clusters where dengue properties tend to be spatially correlated. Forecasting models aimed at predicting the dynamics of dengue activity need to identify the most salient variables capable of contributing to accurate predictions. Our findings show that successful models may need to leverage distinct variables in different locations and be catered to a specific task, such as predicting outbreak magnitude or timing characteristics, to be useful. This advocates in favor of “adaptive models” rather than “one-size-fits-all” models. The results of this study can be applied to improving spatial hierarchical or target-focused forecasting models of dengue activity across Brazil.
Background Dengue fever is a mosquito-borne infection transmitted by Aedes aegypti and mainly found in tropical and subtropical regions worldwide. Since its re-introduction in 1986, Brazil has become a hotspot for dengue and has experienced yearly epidemics. As a notifiable infectious disease, Brazil uses a passive epidemiological surveillance system to collect and report cases; however, dengue burden is underestimated. Thus, Internet data streams may complement surveillance activities by providing real-time information in the face of reporting lags. Methods We analyzed 19 terms related to dengue using Google Health Trends (GHT), a free-Internet data-source, and compared it with weekly dengue incidence between 2011 to 2016. We correlated GHT data with dengue incidence at the national and state-level for Brazil while using the adjusted R squared statistic as primary outcome measure (0/1). We used survey data on Internet access and variables from the official census of 2010 to identify where GHT could be useful in tracking dengue dynamics. Finally, we used a standardized volatility index on dengue incidence and developed models with different variables with the same objective. Results From the 19 terms explored with GHT, only seven were able to consistently track dengue. From the 27 states, only 12 reported an adjusted R squared higher than 0.8; these states were distributed mainly in the Northeast, Southeast, and South of Brazil. The usefulness of GHT was explained by the logarithm of the number of Internet users in the last 3 months, the total population per state, and the standardized volatility index. Conclusions The potential contribution of GHT in complementing traditional established surveillance strategies should be analyzed in the context of geographical resolutions smaller than countries. For Brazil, GHT implementation should be analyzed in a case-by-case basis. State variables including total population, Internet usage in the last 3 months, and the standardized volatility index could serve as indicators determining when GHT could complement dengue state level surveillance in other countries.
Predicting an infectious disease can help reduce its impact by advising public health interventions and personal preventive measures. Novel data streams, such as Internet and social media data, have recently been reported to benefit infectious disease prediction. As a case study of dengue in Brazil, we have combined multiple traditional and non-traditional, heterogeneous data streams (satellite imagery, Internet, weather, and clinical surveillance data) across its 27 states on a weekly basis over seven years. For each state, we nowcast dengue based on several time series models, which vary in complexity and inclusion of exogenous data. The top-performing model varies by state, motivating our consideration of ensemble approaches to automatically combine these models for better outcomes at the state level. Model comparisons suggest that predictions often improve with the addition of exogenous data, although similar performance can be attained by including only one exogenous data stream (either weather data or the novel satellite data) rather than combining all of them. Our results demonstrate that Brazil can be nowcasted at the state level with high accuracy and confidence, inform the utility of each individual data stream, and reveal potential geographic contributors to predictive performance. Our work can be extended to other spatial levels of Brazil, vector-borne diseases, and countries, so that the spread of infectious disease can be more effectively curbed.
ABSTRACTPredicting an infectious disease can help reduce its impact by advising public health interventions and personal preventive measures. While availability of heterogeneous data streams and sensors such as satellite imagery and the Internet have increased the opportunity to indirectly measure, understand, and predict global dynamics, the data may be prohibitively large and/or require intensive data management while also requiring subject matter experts to properly exploit the data sources (e.g., deriving features from fundamentally different data sets). Few efforts have quantitatively assessed the predictive benefit of novel data streams in comparison to more traditional data sources, especially at fine spatio-temporal resolutions. We have combined multiple traditional and non-traditional data streams (satellite imagery, Internet, weather, census, and clinical surveillance data) and assessed their combined ability to predict dengue in Brazil’s 27 states on a weekly and yearly basis over seven years. For each state, we nowcast dengue based on several time series models, which vary in complexity and inclusion of exogenous data. We also predict yearly cumulative risk by municipality and state. The top-performing model and utility of predictive data varies by state, implying that forecasting and nowcasting efforts in the future may be made more robust by and benefit from the use of multiple data streams and models. One size does not fit all, particularly when considering state-level predictions as opposed to the whole country. Our first-of-its-kind high resolution flexible system for predicting dengue incidence with heterogeneous (and still sometimes sparse) data can be extended to multiple applications and regions.