BackgroundThe use of data standards is low across the health care system, and converting data to a common data model (CDM) is usually required to undertake international research. One such model is the Observational Medical Outcomes Partnership (OMOP) CDM. It has gained substantial traction across researchers and those who have developed data platforms. The Observational Health Care Data Sciences and Informatics (OHDSI) partnership manages OMOP and provides many open-source tools to assist in converting data to the OMOP CDM. The challenge, however, is in the skills, knowledge, know-how, and capacity within teams to convert their data to OMOP. The European Health Care Data Evidence Network provided funds to allow data owners to bring in external resources to do the required conversions. The Carrot software (University of Nottingham) is a new set of open-source tools designed to help address these challenges while not requiring data access by external resources. ObjectiveThe use of data protection rules is increasing, and privacy by design is a core principle under the European and UK legislations related to data protection. Our aims for the Carrot software were to have a standardized mechanism for managing the data curation process, capturing the rules used to convert the data, and creating a platform that can reuse rules across projects to drive standardization of process and improve the speed without compromising on quality. Most importantly, we aimed to deliver this design-by-privacy approach without requiring data access to those creating the rules. MethodsThe software was developed using Agile approaches by both software engineers and data engineers, who would ultimately use the system. Experts in OMOP were used to ensure the approaches were correct. An incremental release program was initiated to ensure we delivered continuous progress. ResultsCarrot has been delivered and used on a project called COVID-Curated and Open Analysis and Research Platform (CO-CONNECT) to assist in the process of allowing datasets to be discovered via a federated platform. It has been used to create over 45,000 rules, and over 5 million patient records have been converted. This has been achieved while maintaining our principle of not allowing access to the underlying data by the team creating the rules. It has also facilitated the reuse of existing rules, with most rules being reused rather than manually curated. ConclusionsCarrot has demonstrated how it can be used alongside existing OHDSI tools with a focus on the mapping stage. The COVID-Curated and Open Analysis and Research Platform project successfully managed to reuse rules across datasets. The approach is valid and brings the benefits expected, with future work continuing to optimize the generation of rules. International Registered Report Identifier (IRRID)RR1-10.2196/60917
The use of data standards is low across the healthcare system and therefore to undertake international research it is usually required to convert data to a common data model. One such model is the Observational Medical Outcomes Partnership (OMOP) Common Data Model. It has gained significant traction across researchers and those who have developed data platforms. The Observational Healthcare Data Sciences and Informatics (OHDSI) partnership manage OMOP and provide many open-source tooling to assist those with data to convert their data to the OMOP CDM. The challenge, however, is in the skills, knowledge, know-how and capacity within teams to convert their data to OMOP. The European Health Data Evidence Network (EHDEN) provided funds to allow data owners to bring in external resource to do the required conversions and therefore creating a once in time conversion of data. The Carrot software is a new set of open-source tools designed to help address these challenges while not requiring data access by external resources. Data protection rules are increasing and privacy by design is a core principle under the European and UK legislations related to data protection. Our aims for the Carrot software were to have a standardised mechanism for managing the data curation process, capturing the rules used to convert the data, and creating a platform that can re-use rules across projects to drive standardisation of process, improve the speed, and without compromising on quality. Most importantly, the privacy by design approach was to deliver this approach without requiring those creating the rules to have access to any of the data. Carrot has been delivered and has been used on a project called CO-CONNECT to assist in the process of allowing datasets to be discovered via a federated platform. It has been used to create over forty five thousand rules and over 5 million of patient records have been converted. This has been achieved while maintaining our principles of ensuring this can be achieved with no access to the underlying data by the team creating the rules. It has also facilitated the re-use of existing rules, with the majority of rules being re-used rather than manually curated. Carrot has demonstrated how it can be utilised alongside existing OHDSI tools with a focus on the mapping stage. In the CO-CONNECT project it successfully managed to re-use rules across datasets. The approach is valid and brought the benefits expected with future work continuing to optimise the generation of rules.
Additional file 1. scRNAseq data ONS76 HD-MB 03 - ALL SIGNIFICANT GENES PER CLUSTER. This is and Excel spreadsheet that contains data relating to all the significantly altered genes with their associated Log2 fold change and p-values.
The most common malignant brain tumour in children, medulloblastoma (MB), is subdivided into four clinically relevant molecular subgroups, although targeted therapy options informed by understanding of different cellular features are lacking. Here, by comparing the most aggressive subgroup (Group 3) with the intermediate (SHH) subgroup, we identify crucial differences in tumour heterogeneity, including unique metabolism-driven subpopulations in Group 3 and matrix-producing subpopulations in SHH. To analyse tumour heterogeneity, we profiled individual tumour nodules at the cellular level in 3D MB hydrogel models, which recapitulate subgroup specific phenotypes, by single cell RNA sequencing (scRNAseq) and 3D OrbiTrap Secondary Ion Mass Spectrometry (3D OrbiSIMS) imaging. In addition to identifying known metabolites characteristic of MB, we observed intra- and internodular heterogeneity and identified subgroup-specific tumour subpopulations. We showed that extracellular matrix factors and adhesion pathways defined unique SHH subpopulations, and made up a distinct shell-like structure of sulphur-containing species, comprising a combination of small leucine-rich proteoglycans (SLRPs) including the collagen organiser lumican. In contrast, the Group 3 tumour model was characterized by multiple subpopulations with greatly enhanced oxidative phosphorylation and tricarboxylic acid (TCA) cycle activity. Extensive TCA cycle metabolite measurements revealed very high levels of succinate and fumarate with malate levels almost undetectable particularly in Group 3 tumour models. In patients, high fumarate levels (NMR spectroscopy) alongside activated stress response pathways and high Nuclear Factor Erythroid 2-Related Factor 2 (NRF2; gene expression analyses) were associated with poorer survival. Based on these findings we predicted and confirmed that NRF2 inhibition increased sensitivity to vincristine in a long-term 3D drug treatment assay of Group 3 MB. Thus, by combining scRNAseq and 3D OrbiSIMS in a relevant model system we were able to define MB subgroup heterogeneity at the single cell level and elucidate new druggable biomarkers for aggressive Group 3 and low-risk SHH MB.
Background COVID-19 data have been generated across the United Kingdom as a by-product of clinical care and public health provision, as well as numerous bespoke and repurposed research endeavors. Analysis of these data has underpinned the United Kingdom’s response to the pandemic, and informed public health policies and clinical guidelines. However, these data are held by different organizations, and this fragmented landscape has presented challenges for public health agencies and researchers as they struggle to find relevant data to access and interrogate the data they need to inform the pandemic response at pace. Objective We aimed to transform UK COVID-19 diagnostic data sets to be findable, accessible, interoperable, and reusable (FAIR). Methods A federated infrastructure model (COVID - Curated and Open Analysis and Research Platform [CO-CONNECT]) was rapidly built to enable the automated and reproducible mapping of health data partners’ pseudonymized data to the Observational Medical Outcomes Partnership Common Data Model without the need for any data to leave the data controllers’ secure environments, and to support federated cohort discovery queries and meta-analysis. Results A total of 56 data sets from 19 organizations are being connected to the federated network. The data include research cohorts and COVID-19 data collected through routine health care provision linked to longitudinal health care records and demographics. The infrastructure is live, supporting aggregate-level querying of data across the United Kingdom. Conclusions CO-CONNECT was developed by a multidisciplinary team. It enables rapid COVID-19 data discovery and instantaneous meta-analysis across data sources, and it is researching streamlined data extraction for use in a Trusted Research Environment for research and public health analysis. CO-CONNECT has the potential to make UK health data more interconnected and better able to answer national-level research questions while maintaining patient confidentiality and local governance procedures.
Staphylococcus aureus is a serious human and animal pathogen threat exhibiting extraordinary capacity for acquiring new antibiotic resistance traits in the pathogen population worldwide. The development of fast, affordable and effective diagnostic solutions capable of discriminating between antibiotic-resistant and susceptible S. aureus strains would be of huge benefit for effective disease detection and treatment. Here we develop a diagnostics solution that uses Matrix-Assisted Laser Desorption/Ionisation–Time of Flight Mass Spectrometry (MALDI-TOF) and machine learning, to identify signature profiles of antibiotic resistance to either multidrug or benzylpenicillin in S. aureus isolates. Using ten different supervised learning techniques, we have analysed a set of 82 S. aureus isolates collected from 67 cows diagnosed with bovine mastitis across 24 farms. For the multidrug phenotyping analysis, LDA, linear SVM, RBF SVM, logistic regression, naïve Bayes, MLP neural network and QDA had Cohen’s kappa values over 85.00%. For the benzylpenicillin phenotyping analysis, RBF SVM, MLP neural network, naïve Bayes, logistic regression, linear SVM, QDA, LDA, and random forests had Cohen’s kappa values over 85.00%. For the benzylpenicillin the diagnostic systems achieved up to (mean result ± standard deviation over 30 runs on the test set): accuracy = 97.54% ± 1.91%, sensitivity = 99.93% ± 0.25%, specificity = 95.04% ± 3.83%, and Cohen’s kappa = 95.04% ± 3.83%. Moreover, the diagnostic platform complemented by a protein-protein network and 3D structural protein information framework allowed the identification of five molecular determinants underlying the susceptible and resistant profiles. Four proteins were able to classify multidrug-resistant and susceptible strains with 96.81% ± 0.43% accuracy. Five proteins, including the previous four, were able to classify benzylpenicillin resistant and susceptible strains with 97.54% ± 1.91% accuracy. Our approach may open up new avenues for the development of a fast, affordable and effective day-to-day diagnostic solution, which would offer new opportunities for targeting resistant bacteria.
Hypertrophic cardiomyopathy (HCM) is characterized by increased left ventricular wall thickness that can lead to devastating conditions such as heart failure and sudden cardiac death. Despite extensive study, the mechanisms mediating many of the associated clinical manifestations remain unknown and human models are required. To address this, human-induced pluripotent stem cell (hiPSC) lines were generated from patients with a HCM-associated mutation (c.ACTC1G301A) and isogenic controls created by correcting the mutation using CRISPR/Cas9 gene editing technology. Cardiomyocytes (hiPSC-CMs) were differentiated from these hiPSCs and analyzed at baseline, and at increased contractile workload (2 Hz electrical stimulation). Released extracellular vesicles (EVs) were isolated and characterized after a 24-h culture period and transcriptomic analysis performed on both hiPSC-CMs and released EVs. Transcriptomic analysis of cellular mRNA showed the HCM mutation caused differential splicing within known HCM pathways, and disrupted metabolic pathways. Analysis at increasing contraction frequency showed further disruption of metabolic gene expression, with an additive effect in the HCM background. Intriguingly, we observed differences in snoRNA cargo within HCM released EVs that specifically altered when HCM hiPSC-CMs were subjected to increased workload. These snoRNAs were predicted to have roles in post-translational modifications and alternative splicing, processes differentially regulated in HCM. As such, the snoRNAs identified in this study may unveil mechanistic insight into unexplained HCM phenotypes and offer potential future use as HCM biomarkers or as targets in future RNA-targeting therapies.
Streptococcus uberis is one of the leading pathogens causing mastitis worldwide. Identification of S. uberis strains that fail to respond to treatment with antibiotics is essential for better decision making and treatment selection. We demonstrate that the combination of supervised machine learning and matrix-assisted laser desorption ionization/time of flight (MALDI-TOF) mass spectrometry can discriminate strains of S. uberis causing clinical mastitis that are likely to be responsive or unresponsive to treatment. Diagnostics prediction systems trained on 90 individuals from 26 different farms achieved up to 86.2% and 71.5% in terms of accuracy and Cohen's kappa. The performance was further increased by adding metadata (parity, somatic cell count of previous lactation and count of positive mastitis cases) to encoded MALDI-TOF spectra, which increased accuracy and Cohen's kappa to 92.2% and 84.1% respectively. A computational framework integrating protein-protein networks and structural protein information to the machine learning results unveiled the molecular determinants underlying the responsive and unresponsive phenotypes.
Influenza A virus is a major global pathogen of humans, and there is an unmet need for effective antivirals. Current antivirals against influenza A virus directly target the virus and are vulnerable to mutational resistance. Harnessing an effective host antiviral response is an attractive alternative. We show that brief exposure to low, non-toxic doses of thapsigargin (TG), an inhibitor of the sarcoplasmic/endoplasmic reticulum (ER) Ca2+ ATPase pump, promptly elicits an extended antiviral state that dramatically blocks influenza A virus production. Crucially, oral administration of TG protected mice against lethal virus infection and reduced virus titres in the lungs of treated mice. TG-induced ER stress unfolded protein response appears as a key driver responsible for activating a spectrum of host antiviral defences that include an enhanced type I/III interferon response. Our findings suggest that TG is potentially a viable host-centric antiviral for the treatment of influenza A virus infection without the inherent problem of drug resistance.