Tuberculosis (TB) is still a major global health challenge, killing over 1.5 million people each year, and hence, there is a need to identify and develop novel treatments for Mycobacterium tuberculosis (M. tuberculosis). The prevalence of infections caused by nontuberculous mycobacteria (NTM) is also increasing and has overtaken TB cases in the United States and much of the developed world. Mycobacterium abscessus (M. abscessus) is one of the most frequently encountered NTM and is difficult to treat. We describe the use of drug-disease association using a semantic knowledge graph approach combined with machine learning models that has enabled the identification of several molecules for testing anti-mycobacterial activity. We established that niclosamide (M. tuberculosis IC90 2.95 μM; M. abscessus IC90 59.1 μM) and tribromsalan (M. tuberculosis IC90 76.92 μM; M. abscessus IC90 147.4 μM) inhibit M. tuberculosis and M. abscessus in vitro. To investigate the mode of action, we determined the transcriptional response of M. tuberculosis and M. abscessus to both compounds in axenic log phase, demonstrating a broad effect on gene expression that differed from known M. tuberculosis inhibitors. Both compounds elicited transcriptional responses indicative of respiratory pathway stress and the dysregulation of fatty acid metabolism.
INTRODUCTION:The advancement of the Army's National Emergency Tele-Critical Care Network (NETCCN) and planned evolution to an Intelligent Medical System rest on a digital transformation characterized by the application of analytic rigor anchored and machine learning.The goal is an enduring capability for telecritical care in support of the Nation's warfighters and, more broadly, for emergency response, crisis management, and mass casualty situations as the number and intensity of disasters increase nationwide. That said, technology alone is unlikely to solve the most pressing issues in operational medicine and combat casualty care.MATERIALS AND METHODS:A total performance system (TPS) creates opportunities to address vulnerabilities and overcome barriers to success. As applied during the NETCCN project, the TPS captures the best performance-centric information and know-how, increasing the potential to save lives, improve readiness, and accomplish missions.RESULTS:The purpose of this project was to apply a performance-based readiness model to aid in the evaluation of Army telehealth technologies. Through various user-facing surveys, polls, and reporting techniques, the project aimed to measure the perceived value of telehealth technologies within a sample of the project team member population. By providing a detailed approach to the collection of lessons learned, researchers were able to determine the importance of information and methods versus a focus on technology alone. The use of an emoji-based feedback assessment indicated that most lessons learned were helpful to the project team.CONCLUSIONS:Through the NETCCN TPS, we have been able to address product-related measures, knowledge of product efficacy, project metrics, and many implementation considerations that can be further investigated by setting and engagement type. Through the Technology in Disaster Environments learning accelerator, it was possible to rapidly acquire, process, organize, and disseminate best practices and learnings in near real time, providing a critical feedback and improvement loop.
In this paper, we introduce the Analysis Platform for Risk, Resilience, and Expenditure in Disasters (APRED)-a disaster-analytic platform developed for crisis practitioners and economic developers across the United States (US). APRED provides practitioners with a centralized platform for exploring disaster resilience and vulnerability profiles of all counties across the US. The platform comprises five sections including: (1) Disaster Resilience Index, (2) Business Vulnerability Index, (3) Disaster Declaration History, (4) County Profile, and (5) Storm History sections. We further describe our end-to-end human-centered design and engineering process that involved contextual inquiry, community-based participatory design, and rapid prototyping with the support of US Economic Development Administration representatives and regional economic developers across the US. Findings from our study revealed that distributed cognition, content heuristic, shareability, and human-centered systems are crucial considerations for developing data-intensive visualization platforms for resilience planning. We discuss the implications of these findings and inform future research on developing sociotechnical visualization platforms to support resilience planning.
The COVID-19 global pandemic has changed every facet of our lives overnight and has resulted in many challenges and opportunities. Utilizing the Lens of Vulnerability we investigate how disparities in technology adoption affect activities of daily living. In this paper, we analyze the existing literature and case studies regarding how the lifestyles of socially vulnerable populations have changed during the pandemic in terms of technology adoption. Socially vulnerable populations, such as racial and ethnic minorities, people with disabilities, older adults, children, and the socially isolated, are specifically addressed because they are groups of people who have been significantly and disproportionately affected by the pandemic. This paper emphasizes that despite seeing changes in and research on technology adoption across healthcare, employment, and education, the impact of COVID-19 in government and social services and activities of daily living is underdeveloped. The study concludes by offering practical and academic recommendations and future research directions. Lessons learned from the current pandemic and an understanding of the differential technology adoption for activities of daily living amid a disaster will help emergency managers, academics, and government officals prepare for and respond to future crises.
Pandemic-tracking apps may form a future infrastructure for public health surveillance. Yet, there has been relatively little exploration of the potential societal implications of such an infrastructure. In semi-structured interviews with 23 participants from India, the Middle East and North Africa (MENA), and the United States, we discussed attitudes and preferences regarding the deployment of apps that support contact tracing to contain the spread of COVID-19. Through interpretive analysis, we examined the relationship between persistent discomfort and vulnerability when using such apps. Such an examination yielded three temporal forms of vulnerability: real, anticipatory, and speculative. By identifying and defining the temporalities of vulnerability through an analysis of people's pandemic-related thoughts and experiences, we develop the overlapping discourses of humanistic infrastructure studies and infrastructural speculation. In doing so, we explore the concept of vulnerability itself and present implications for the study of vulnerability in Human-Computer Interaction (HCI) and for the oversight of app-based public health surveillance.
Pandemic-tracking apps may form a future infrastructure for public health surveillance. Yet, there has been relatively little exploration of the potential societal implications of such an infrastructure. In semi-structured interviews with 23 participants from India, the Middle East and North Africa (MENA), and the United States, we discussed attitudes and preferences regarding the deployment of apps that support contact tracing to contain the spread of COVID-19. Through interpretive analysis, we examined the relationship between persistent discomfort and vulnerability when using such apps. Such an examination yielded three temporal forms of vulnerability: real, anticipatory, and speculative. By identifying and defining the temporalities of vulnerability through an analysis of people's pandemic-related thoughts and experiences, we develop the overlapping discourses of humanistic infrastructure studies and infrastructural speculation. In doing so, we explore the concept of vulnerability itself and present implications for the study of vulnerability in Human-Computer Interaction (HCI) and for the oversight of app-based public health surveillance.
Genome wide association studies (GWAS) can reveal important genotype–phenotype associations, however, data quality and interpretability issues must be addressed. For drug discovery scientists seeking to prioritize targets based on the available evidence, these issues go beyond the single study. Here, we describe rational ranking, filtering and interpretation of inferred gene–trait associations and data aggregation across studies by leveraging existing curation and harmonization efforts. Each gene–trait association is evaluated for confidence, with scores derived solely from aggregated statistics, linking a protein-coding gene and phenotype. We propose a method for assessing confidence in gene–trait associations from evidence aggregated across studies, including a bibliometric assessment of scientific consensus based on the iCite Relative Citation Ratio, and meanRank scores, to aggregate multivariate evidence. This method, intended for drug target hypothesis generation, scoring and ranking, has been implemented as an analytical pipeline, available as open source, with public datasets of results, and a web application designed for usability by drug discovery scientists, at https://unmtid-shinyapps.net/tiga/ .
Background: The COVID-19 pandemic has highlighted the inability of health systems to leverage existing system infrastructure in order to rapidly develop and apply broad analytical tools that could inform state-and national-level policymaking, as well as patient care delivery in hospital settings. The COVID-19 pandemic has also led to highlighted systemic disparities in health outcomes and access to care based on race or ethnicity, gender, income-level, and urban-rural divide. Although the United States seems to be recovering from the COVID-19 pandemic owing to widespread vaccination efforts and increased public awareness, there is an urgent need to address the aforementioned challenges. Objective: This study aims to inform the feasibility of leveraging broad, statewide datasets for population health-driven decision-making by developing robust analytical models that predict COVID-19-related health care resource utilization across patients served by Indiana's statewide Health Information Exchange. Methods: We leveraged comprehensive datasets obtained from the Indiana Network for Patient Care to train decision forest-based models that can predict patient-level need of health care resource utilization. To assess these models for potential biases, we tested model performance against subpopulations stratified by age, race or ethnicity, gender, and residence (urban vs rural). Results: For model development, we identified a cohort of 96,026 patients from across 957 zip codes in Indiana, United States. We trained the decision models that predicted health care resource utilization by using approximately 100 of the most impactful features from a total of 1172 features created. Each model and stratified subpopulation under test reported precision scores >70%, accuracy and area under the receiver operating curve scores >80%, and sensitivity scores approximately >90%. We noted statistically significant variations in model performance across stratified subpopulations identified by age, race or ethnicity, gender, and residence (urban vs rural). Conclusions: This study presents the possibility of developing decision models capable of predicting patient-level health care resource utilization across a broad, statewide region with considerable predictive performance. However, our models present statistically significant variations in performance across stratified subpopulations of interest. Further efforts are necessary to identify root causes of these biases and to rectify them.
BACKGROUND The COVID-19 pandemic has highlighted the inability of health systems to leverage existing system infrastructure to rapidly develop and apply broad analytical tools that could inform state and national-level policymaking as well as patient care delivery at hospital settings. COVID-19 has also led to highlighted systemic disparities in health outcomes and access to care based on race/ethnicity, gender, income-level and urban-rural divide. While the US seems to be recovering from the COVID-19 pandemic due to widespread vaccination efforts and increased public awareness, there is an urgent need to address the aforementioned challenges. OBJECTIVE Inform the feasibility of leveraging broad, statewide datasets for population-health driven decision making by developing robust analytical models that predict COVID-19 related healthcare resource utilization across patients served by Indiana’s statewide Health Information Exchange (HIE). METHODS We leveraged comprehensive datasets obtained from the Indiana Network for Patient Care (INPC) to train decision forest-based models that predicted patient-level need of healthcare resource utilization. To assess models for potential biases, we tested model performance against sub-populations stratified by age, race/ethnicity, gender, and residence (urban vs. rural). RESULTS We identified a cohort of 96,190 patients from 957 zip codes spread across the state of Indiana. We trained decision models that predicted healthcare resource utilization using the most impactful features (~100) out of a total of 1172 features created. Each model and stratified sub-population under test reported precision scores > 70%, accuracy and AUC ROC scores > 80%, and sensitivity scores ~>90%. We noted statistically significant variations in model performance across stratified sub-populations identified by age, race/ethnicity, gender, and residence (urban vs. rural). CONCLUSIONS This study presents the possibility of developing decision models capable of predicting patient-level healthcare resource utilization across a broad statewide region with considerable predictive performance. However, our models present statistically significant variations in performance across stratified sub-populations of interest. Further efforts are necessary to identify root causes of these biases and to rectify them. CLINICALTRIAL NA
Stroke is a common disabling disease that severely affects the daily life of patients. Accumulating evidence indicates that rehabilitation therapy can improve movement function. However, no clear guidelines have specific and effective rehabilitation therapy schemes, and the development of new rehabilitation techniques has been relatively slow. This study used a text mining approach, the ABC model, to identify an existing rehabilitation candidate therapy method that is most likely to be repositioned for stroke. In the model, we built the internal links of stroke (A), assessment scales (B), and rehabilitation therapies (C) in PubMed and the links were related to upper limb function measurements for patients with stroke. In the first step, using E-utility, we retrieved both stroke-related assessment scales and rehabilitation therapy records and then compiled two datasets, which were called Stroke_Scales and Stroke_Therapies, respectively. In the next step, we crawled all rehabilitation therapies co-occurring with the Stroke_Therapies and then named them as All_Therapies. Therapies that were already included in Stroke_Therapies were deleted from All_Therapies; therefore, the remaining therapies were the potential rehabilitation therapies, which could be repositioned for stroke after subsequent filtration by a manual check. We identified the top-ranked repositioning rehabilitation therapy and subsequently examined its clinical validation. Hand-arm bimanual intensive training (HABIT) was ranked the first in our repositioning rehabilitation therapies and had the most interaction links with Stroke_Scales. HABIT significantly improved clinical scores on assessment scales [Fugl-Meyer Assessment (FMA) and action research arm test (ARAT)] in the clinical validation study for acute stroke patients with upper limb dysfunction. Therefore, based on the ABC model and clinical validation, HABIT is a promising repositioned rehabilitation therapy for stroke, and the ABC model is an effective text mining approach for rehabilitation therapy repositioning. The findings in this study would be helpful in clinical knowledge discovery.
Representation learning provides new and powerful graph analytical approaches and tools for the highly valued data science challenge of mining knowledge graphs. Since previous graph analytical methods have mostly focused on homogeneous graphs, an important current challenge is extending this methodology for richly heterogeneous graphs and knowledge domains. The biomedical sciences are such a domain, reflecting the complexity of biology, with entities such as genes, proteins, drugs, diseases, and phenotypes, and relationships such as gene co-expression, biochemical regulation, and biomolecular inhibition or activation. Therefore, the semantics of edges and nodes are critical for representation learning and knowledge discovery in real world biomedical problems. In this paper, we propose the edge2vec model, which represents graphs considering edge semantics. An edge-type transition matrix is trained by an Expectation-Maximization approach, and a stochastic gradient descent model is employed to learn node embedding on a heterogeneous graph via the trained transition matrix. edge2vec is validated on three biomedical domain tasks: biomedical entity classification, compound-gene bioactivity prediction, and biomedical information retrieval. Results show that by considering edge-types into node embedding learning in heterogeneous graphs, edge2vec significantly outperforms state-of-the-art models on all three tasks. We propose this method for its added value relative to existing graph analytical methodology, and in the real world context of biomedical knowledge discovery applicability.
Stroke is a common disabling disease severely affecting the daily life of the patients. There is evidence that rehabilitation therapy can improve the movement function. However, there are no clear guidelines that identify specific, effective rehabilitation therapy schemes, and the development of new rehabilitation techniques has been fairly slow. One informatics translational approach, called ABC model in Literature-based Discovery, was used to mine an existing rehabilitation candidate which is most likely to be repositioned for stroke. As in the classic ABC model originated from Don Swanson, we built the internal links of stroke (A), assessment scales (B), rehabilitation therapies (C) in PubMed relating to upper limb function measurements for stroke patients. In the first step, with E-utility we retrieved both stroke related assessment scales and rehabilitation therapies records, and complied two datasets called Stroke_Scales and Stroke_Therapies, respectively. In the next step, we crawled all rehabilitation therapies co-occurred with the Stroke_Theapies, named as All_Therapies. Therapies that were already included in Stroke_Therapies were deleted from All_Therapies, so that the remaining therapies were the potential rehabilitation therapies, which could be repositioned for stroke after subsequent filtration by manual check. We identified the top ranked repositioning rehabilitation therapy following by subsequent clinical validation. Hand-arm bimanual intensive training (HABIT) ranked the first in our repositioning rehabilitation therapies list, with the most interaction links with Stroke_Scales. HABIT showed a significant improvement in clinical scores on assessment scales of Fugl-Meyer Assessment and Action Research Arm Test in the clinical validation on upper limb function for acute stroke patients. Based on the ABC model and clinical validation of the results, we put forward that HABIT as a promising rehabilitation therapy for stroke, which shows that the ABC model is an effective text mining approach for rehabilitation therapy repositioning. The results seem to be promoted in clinical knowledge discovery. Author Summary In the present study, we proposed a text mining approach to mining terms related to disease, rehabilitation therapy, and assessment scale from literature, with a subsequent ABC inference analysis to identify relationships of these terms across publications. The clinical validation demonstrated that our approach can be used to identify potential repositioning rehabilitation therapy strategies for stroke. Specifically, we identified a promising rehabilitation method called HABIT previously used in pediatric congenital hemiplegia. A subsequent clinical trial confirmed this as a highly promising rehabilitation therapy for stroke.
There are many reasons data science teams should use a well-defined process to manage and coordinate their efforts, such as improved collaboration, efficiency and stakeholder communication. This paper explores the current methodology data science teams use to manage and coordinate their efforts. Unfortunately, based on our survey results, most data science teams currently use an ad hoc project management approach. In fact, 82% of the data scientists surveyed did not follow an explicit process. However, it is encouraging to note that 85% of the respondents thought that adopting an improved process methodology would improve the teams' outcomes. Based on these results, we described six possible process methodologies teams could use. To conclude, we outlined plans to describe best practices for data science team processes and to develop a process evaluation framework.
Background Netpredictor is an R package for prediction of missing links in any given unipartite or bipartite network. The package provides utilities to compute missing links in a bipartite and well as unipartite networks using Random Walk with Restart and Network inference algorithm and a combination of both. The package also allows computation of Bipartite network properties, visualization of communities for two different sets of nodes, and calculation of significant interactions between two sets of nodes using permutation based testing. The application can also be used to search for top-K shortest paths between interactome and use enrichment analysis for disease, pathway and ontology. The R standalone package (including detailed introductory vignettes) and associated R Shiny web application is available under the GPL-2 Open Source license and is freely available to download. Results We compared different algorithms performance in different small datasets and found random walk supersedes rest of the algorithms. The package is developed to perform network based prediction of unipartite and bipartite networks and use the results to understand the functionality of proteins in an interactome using enrichment analysis. Conclusion The rapid application development envrionment like shiny, helps non programmers to develop fast rich visualization apps and we beleieve it would continue to grow in future with further enhancements. We plan to update our algorithms in the package in near future and help scientist to analyse data in a much streamlined fashion.
Tuberculosis (TB) is the world’s leading infectious killer with 1.8 million deaths in 2015 as reported by WHO. It is therefore imperative that alternate routes of identification of novel anti-TB compounds are explored given the time and costs involved in new drug discovery process. Towards this, we have developed RepTB . This is a unique drug repurposing approach for TB that uses molecular function correlations among known drug-target pairs to predict novel drug-target interactions. In this study, we have created a Gene Ontology based network containing 26,404 edges, 6630 drug and 4083 target nodes. The network, enriched with molecular function ontology, was analyzed using Network Based Inference (NBI). The association scores computed from NBI are used to identify novel drug-target interactions. These interactions are further evaluated based on a combined evidence approach for identification of potential drug repurposing candidates. In this approach, targets which have no known variation in clinical isolates, no human homologs, and are essential for Mtb’s survival and or virulence are prioritized. We analyzed predicted DTIs to identify target pairs whose predicted drugs may have synergistic bactericidal effect. From the list of predicted DTIs from RepTB, four TB targets, namely, FolP1 (Dihydropteroate synthase), Tmk (Thymidylate kinase), Dut (Deoxyuridine 5′-triphosphate nucleotidohydrolase) and MenB (1,4-dihydroxy-2-naphthoyl-CoA synthase) may be selected for further validation. In addition, we observed that in some cases there is significant chemical structure similarity between predicted and reported drugs of prioritized targets, lending credence to our approach. We also report new chemical space for prioritized targets that may be tested further. We believe that with increasing drug-target interaction dataset RepTB will be able to offer better predictive value and is amenable for identification of drug-repurposing candidates for other disease indications too.
BACKGROUND:There are a huge variety of data sources relevant to chemical, biological and pharmacological research, but these data sources are highly siloed and cannot be queried together in a straightforward way. Semantic technologies offer the ability to create links and mappings across datasets and manage them as a single, linked network so that searching can be carried out across datasets, independently of the source. We have developed an application called PIBAS FedSPARQL that uses semantic technologies to allow researchers to carry out such searching across a vast array of data sources.RESULTS:PIBAS FedSPARQL is a web-based query builder and result set visualizer of bioinformatics data. As an advanced feature, our system can detect similar data items identified by different Uniform Resource Identifiers (URIs), using a text-mining algorithm based on the processing of named entities to be used in Vector Space Model and Cosine Similarity Measures. According to our knowledge, PIBAS FedSPARQL was unique among the systems that we found in that it allows detecting of similar data items. As a query builder, our system allows researchers to intuitively construct and run Federated SPARQL queries across multiple data sources, including global initiatives, such as Bio2RDF, Chem2Bio2RDF, EMBL-EBI, and one local initiative called CPCTAS, as well as additional user-specified data source. From the input topic, subtopic, template and keyword, a corresponding initial Federated SPARQL query is created and executed. Based on the data obtained, end users have the ability to choose the most appropriate data sources in their area of interest and exploit their Resource Description Framework (RDF) structure, which allows users to select certain properties of data to enhance query results.CONCLUSIONS:The developed system is flexible and allows intuitive creation and execution of queries for an extensive range of bioinformatics topics. Also, the novel "similar data items detection" algorithm can be particularly useful for suggesting new data sources and cost optimization for new experiments. PIBAS FedSPARQL can be expanded with new topics, subtopics and templates on demand, rendering information retrieval more robust.
Highly chemically similar drugs usually possess similar biological activities, but sometimes, small changes in chemistry can result in a large difference in biological effects. Chemically similar drug pairs that show extreme deviations in activity represent distinctive drug interactions having important implications. These associations between chemical and biological similarity are studied as discontinuities in activity landscapes. Particularly, activity cliffs are quantified by the drop in similar activity of chemically similar drugs. In this paper, we construct a landscape using a large drug-target network and consider the rises in similarity and variation in activity along the chemical space. Detailed analysis of structure and activity gives a rigorous quantification of distinctive pairs and the probability of their occurrence.
Geoffrey Fox合作论文数Department of Physics, College of Arts and Sciences, Indiana University;Department of Intelligent Systems Engineering, Indiana University;Community Grid Laboratory, Indiana University;Digital Science Center of Pervasive Technology Institute;School of Engineering and Applied Science, University of Virginia10
Gary Wiggins合作论文数Indiana University3