Opioid overdose related deaths have increased dramatically in recent years. Combating the opioid epidemic requires better understanding of the epidemiology of opioid poisoning (OP). To discover trends and patterns of opioid poisoning and the demographic and regional disparities, we analyzed large scale patient visits data in New York State (NYS). Demographic, spatial, temporal and correlation analyses were performed for all OP patients extracted from the claims data in the New York Statewide Planning and Research Cooperative System (SPARCS) from 2010 to 2016, along with Decennial US Census and American Community Survey zip code level data. 58,481 patients with at least one OP diagnosis and a valid NYS zip code address were included. Main outcome and measures include OP patient counts and rates per 100,000 population, patient level factors (gender, age, race and ethnicity, residential zip code), and zip code level social demographic factors. The results showed that the OP rate increased by 364.6%, and by 741.5% for the age group > 65 years. There were wide disparities among groups by race and ethnicity on rates and age distributions of OP. Heroin and non-heroin based OP rates demonstrated distinct temporal trends as well as major geospatial variation. The findings highlighted strong demographic disparity of OP patients, evolving patterns and substantial geospatial variation.
Introduction: To discover trends and patterns of opioid poisoning and the demographic and regional disparities by analyzing large scale patient visits data in New York State (NYS). Methods: Demographic, spatial, temporal and correlation analyses were performed for all OP patients extracted from the New York Statewide Planning and Research Cooperative System (SPARCS) from 2010 to 2016, along with Decennial US Census and American Community Survey zip code level data. The study is based on claims data. 58,481 patients with at least one OP diagnosis and a valid NYS zip code address were included. OP patient counts and rates per 100,000 population; patient level factors (gender, age, race and ethnicity, residential zip code); zip code level social demographic factors. Analyses were completed between 2017 and 2019. Results: In this study of 58,481 patients with opioid poisoning (OP) in New York State from 2010 to 2016, the OP rate increased by 364.6%, and by 741.5% for the age group > 65 years. There were wide disparities among groups by race and ethnicity on rates and age distributions of OP. Heroin and non-Heroin based OP rates show distinct temporal trends as well as major geospatial variation. Conclusions: The findings highlight strong demographic disparity of OP patients, evolving patterns and substantial geospatial variation.
INTRODUCTION:Not enough is known about the epidemiology of opioid poisoning to tailor interventions to help address the growing opioid crisis in the U.S. The objective of this study is to expand the current understanding of opioid poisoning through the use of data analytics to evaluate geographic, temporal, and sociodemographic differences of opioid poisoning- related hospital visits in a region of New York State with high opioid poisoning rates. METHODS:This retrospective cohort study utilized patient-level New York State all-payer hospital data (2010-2016) combined with Census data to evaluate geographic, patient, and community factors for 9,714 Long Island residents with an opioid poisoning-related inpatient or outpatient hospital facility discharge. Temporal, 7-year opioid poisoning rates and trends were evaluated, and geographic maps were generated. Overall, significance tests and tests for linear trend were based upon logistic regression. Analyses were completed between 2017 and 2018. RESULTS:Since 2010, Long Island and New York State opioid poisoning hospital visit rates have increased 2.5- to 2.7-fold (p<0.001). Opioid poisoning hospital visit rates decreased for men, white patients, and self-payers (p<0.001) and increased for Medicare payers (p<0.001). Communities with high opioid poisoning rates had lower median home values, higher percentages of high school graduates, were younger, and more often white patients (p<0.01). Maps displayed geographic patterns of communities with high opioid poisoning rates overall and by age group. CONCLUSIONS:Findings highlight the changing demographics of the opioid poisoning epidemic and utility of data analytics tools to identify regions and patient populations to focus interventions. These population identification techniques can be applied in other communities and interventions.
Opioid related deaths are increasing dramatically in recent years, and opioid epidemic is worsening in the United States. Combating opioid epidemic becomes a high priority for both the U.S. government and local governments such as New York State. Analyzing patient level opioid related hospital visits provides a data driven approach to discover both spatial and temporal patterns and identity potential causes of opioid related deaths, which provides essential knowledge for governments on decision making. In this paper, we analyzed opioid poisoning related hospital visits using New York State SPARCS data, which provides diagnoses of patients in hospital visits. We identified all patients with primary diagnosis as opioid poisoning from 2010-2014 for our main studies, and from 2003-2014 for temporal trend studies. We performed demographical based studies, and summarized the historical trends of opioid poisoning. We used frequent item mining to find co-occurrences of diagnoses for possible causes of poisoning or effects from poisoning. We provided zip code level spatial analysis to detect local spatial clusters, and studied potential correlations between opioid poisoning and demographic and social-economic factors.
Opioid-abuse epidemic in the United States has escalated to national attention due to the dramatic increase of opioid overdose deaths. Analyzing opioid-related social media has the potential to reveal patterns of opioid abuse at a national scale, understand opinions of the public, and provide insights to support prevention and treatment. Reddit is a community based social media with more reliable content curated by the community through voting. In this study, we collected and analyzed all opioid related discussions from January 2014 to October 2017, which contains 51,537 posts by 16,162 unique users. We analyzed the data to understand the psychological categories of the posts, and performed topic modeling to reveal the major topics of interest. We also characterized the extent of social support received from comments and scores by each post. Last, we analyzed statistically significant difference in the posts between anonymous and non-anonymous users.
Recent years have witnessed an explosion of geospatial data, especially in the form of Volunteered Geographic Information (VGI). As a prominent example, OpenStreetMap (OSM) creates a free editable map of the world from a large number of contributors. On the other hand, social media platforms such as Twitter or Instagram supply dynamic social feeds at population level. As much of such data is geo-tagged, there is a high potential on integrating social media with OSM to enrich OSM with semantic annotations, which will complement existing objective description oriented annotations to provide a broader range of annotations. In this paper, we propose a comprehensive framework on integrating social media data and VGI data to derive knowledge about geographical objects, specifically, top relevant annotations from tweets for objects in OSM. We first integrate geo-tagged tweets with OSM data with scalable spatial queries running on MapReduce. We propose a frequency based method for annotating boundary based geographic objects (a polygon), and a probability based method for annotating point based geographic objects (Latitude and Longitude), with consideration of noise. We evaluate our methods using a large geo-tagged tweets corpus and representative geographic objects from OSM, which demonstrates promising results through ground-truth comparison and case studies. We are able to produce up to 80% correct names for geographical objects and discover implicitly relevant information, such as popular exhibitions of a museum, the nicknames or visitors' impression to a tourism attraction.
Increased accessibility of health data provides unique opportunities to discover spatio-temporal patterns of diseases. For example, New York State SPARCS (Statewide Planning and Research Cooperative System) data collects patient level detail on patient demographics, diagnoses, services, and charges for each hospital inpatient stay and outpatient visit. Such data also provides home addresses for each patient. This paper presents our preliminary work on spatial, temporal, and spatial-temporal analysis of disease patterns for New York State using SPARCS data. We analyzed spatial distribution patterns of typical diseases at ZIP code level. We performed temporal analysis of common diseases based on 12 years' historical data. We then compared the spatial variations for diseases with different levels of clustering tendency, and studied the evolution history of such spatial patterns. Case studies based on asthma demonstrated that the discovered spatial clusters are consistent with prior studies. We visualized our spatial-temporal patterns as animations through videos.
Increased accessibility of health data made available by the government provides unique opportunity for spatial analytics with much higher resolution to discover patterns of diseases, and their correlation with spatial impact indicators. This paper demonstrated our vision of integrative spatial analytics for public health by linking the New York Cancer Mapping Dataset with datasets containing potential spatial impact indicators. We performed spatial based discovery of disease patterns and variations across New York State, and identify potential correlations between diseases and demographic, socio-economic and environmental indicators. Our methods were validated by three correlation studies: the correlation between stomach cancer and Asian race, the correlation between breast cancer and high education population, and the correlation between lung cancer and air toxics. Our work will allow public health researchers, government officials or other practitioners to adequately identify, analyze, and monitor health problems at the community or neighborhood level for New York State.
Recent years have witnessed an explosion of geospatial data, especially in the form of Volunteered Geographic Information (VGI). As a prominent example, OpenStreetMap (OSM) creates a free editable map of the world from a large number of contributors. On the other hand, social media platforms such as Twitter or Instagram supply dynamic social feeds at population level. As much of such data is geo-tagged, there is a high potential on integrating social media with OSM to enrich OSM with semantic annotations, which will complement existing objective description oriented annotations to provide a broader range of annotations. In this paper, we propose a comprehensive framework on integrating social media data and VGI data to derive knowledge about geographical objects, specifically, top relevant annotations from tweets for objects in OSM. We first integrate geo-tagged tweets with OSM data with scalable spatial queries running on MapReduce. We propose a frequency based method for annotating boundary based geographic objects, and a probability based method for annotating point based geographic objects, with consideration of noise. We evaluate our methods using a large geo-tagged tweets corpus and representative geographic objects from OSM, which demonstrates promising results through ground-truth comparison and case studies. We are able to produce up to 80% correct names for geographical objects and discover implicitly relevant information, such as popular exhibitions of a museum, the nicknames or visitors' impression to a tourism attraction.
Social media platforms have become a major gateway to receive and analyze public opinions. Understandingusers can provide invaluable context information of their social media posts and significantly improve traditional opinion analysis models. Demographic attributes,such as ethnicity, gender, age, among others,have been extensively applied to characterize social mediausers. While studies have shown that user groups formed by demographic attributes can have coherent opinions towards political issues, these attributes are often not explicitly coded by users through their profiles.Previous work has demonstrated the effectiveness of different user signals such as users’ posts and names in determining demographic attributes. Yet, these efforts mostly evaluate linguistic signals from users’ postsand train models from artificially balanced datasets. In this paper, we propose a comprehensive list of user signals:self-descriptions and posts aggregated from users’ friends and followers, users’ profile images, and users’ names.We provide a comparative study of these signalsside-by-side in the tasks on inferring three major demographic attributes, namely ethnicity, gender, and age.We utilize a realistic unbalanced datasets that share similar demographic makeups in Twitter for training modelsand evaluation experiments. Our experiments indicate that self-descriptions provide the strongest signal for ethnicity and age inference and clearly improve the overall performance when combined with tweets. Profile images for gender inference have the highest precision score with overall score close to the best result in our setting. This suggests that signals in self descriptions and profile images have potentials to facilitate demographic attribute inferences in Twitter, and are promising for future investigation.
The growth of spatial big data has been explosive thanks to cost-effective and ubiquitous positioning technologies, and the generation of data from multiple sources in multi-forms. Such emerging spatial data has high potential to create new insights and values for our life through spatial analytics. However, spatial data analytics faces two major challenges. First, spatial data is both data-and compute-intensive due to the massive amounts of data and the multi-dimensional nature, which requires high performance spatial computing infrastructure and methods. Second, spatial big data sources are often isolated, for example, OpenStreetMap, census data and Twitter tweets are independent data sources. This leads to incompleteness of information and sometimes limited data accuracy, thus limited values from the data. Integrating spatial big data analytics by consolidating multiple data sources provides significant potential for data quality improvement in terms of completeness and accuracy, and much increased values derived from the data. In this paper, we present our vision of a high performance integrated spatial big data analytics framework. We provide a scalable spatial query based data integration engine with MapReduce, and demonstrate integrated spatial data analytics through a few use cases in our preliminary work. We then present our future plan on integrated spatial big data analytics for improving public health research and applications.
If a drug can treat a specific disease, this drug can probably treat the diseases with similar phenotype. Therefore, large-scale computing similarity of the disease phenotype can help to find the new treatment. We downloaded 3742 diseases phenotype information from OMIM database, and 13721 Mesh words related to anatomy and disease symptoms from Mesh vocabulary thesaurus. Each Mesh word was searched in all of 3742 diseases phenotype annotation text. Finally, we got a Mesh vocabulary list for every disease. Then disease phenotype pairwise similarity matrix was systematically calculated using semantic analysis approach. We found that most of the diseases associated biological pathways include tumor biological pathways, insulin signaling, hypertrophic cardiomyopathy pathways and cell adhesion pathway. The probability of involving same KEGG biological pathway increases with the disease phenotype similarity. This also shows our method is reliable. Disease phenotype similarity can be used as the supplement of disease genomic space similarity, and may have potential application value in drug discovery.
The aim of this study is to develop backpropagation neural networks (BPNN) for better prediction of ventilatory function in children and adolescents. Nine hundred and ninety-nine healthy children and adolescents (500 males and 499 females) aged 10-18 years, all of the Han Nationality, were selected from Inner Mongolia Autonomous Region, and their heights, weights, and ventilatory functions were measured respectively by means of physical examination and spirometric test. Using the approaches of BPNN and stepwise multiple regression, the prediction models and equations for forced vital capacity (FVC), forced expiratory volume in one second (FEV1), peak expiratory flow (PEF), forced expiratory flow at 25% of forced vital capacity (FEF25%), forced expiratory flow at 50% of forced vital capacity (FEF50%), maximal mid-expiratory flow (MMEF) and forced expiratory flow at 75% of forced vital capacity (FEF75%) were established. Through analyzing mean squared difference (MSD) and correlation coefficient (R) of the ventilatory function indexes, the present study compared the results of BPNN, linear regression equation based on this work (LR's equation), prediction equations based on the studies of Ip et al. (Ip's equation) and Zapletal et al. (Zapletal's equation). The results showed, regardless of sex, the BPNN prediction models appeared to have smaller MSD and higher R values, compared with those from the other prediction equations; and the LR's equation also had smaller MSD and higher R values compared with those from Ip's and Zapletal's equations. The coefficients of variance (CV) for FEF50%, MMEF and FEF75% were higher than those of the other ventilatory function parameters, and their increasing percentages of R values (ΔR, relative to R values by LR's equation) derived by BPNN were correspondingly higher than those of the other indexes. In sum, BPNN approach for ventilatory function prediction outperforms the traditional regression methods. When CV of a certain ventilatory function parameter is higher, the superiority of BPNN would be more significant compared with traditional regression methods.