Risk prediction in patients with heart failure (HF) is essential to improve the tailoring of preventive, diagnostic, and therapeutic strategies for the individual patient, and effectively use health care resources. Risk scores derived from controlled clinical studies can be used to calculate the risk of mortality and HF hospitalizations. However, these scores are poorly implemented into routine care, predominantly because their calculation requires considerable efforts in practice and necessary data often are not available in an interoperable format. In this work, we demonstrate the feasibility of a multi-site solution to derive and calculate two exemplary HF scores from clinical routine data (MAGGIC score with six continuous and eight categorical variables; Barcelona Bio-HF score with five continuous and six categorical variables). Within HiGHmed, a German Medical Informatics Initiative consortium, we implemented an interoperable solution, collecting a harmonized HF-phenotypic core data set (CDS) within the openEHR framework. Our approach minimizes the need for manual data entry by automatically retrieving data from primary systems. We show, across five participating medical centers, that the implemented structures to execute dedicated data queries, followed by harmonized data processing and score calculation, work well in practice. In summary, we demonstrated the feasibility of clinical routine data usage across multiple partner sites to compute HF risk scores. This solution can be extended to a large spectrum of applications in clinical care.
COVID-19 has challenged the healthcare systems worldwide. To quickly identify successful diagnostic and therapeutic approaches large data sharing approaches are inevitable. Though organizational clinical data are abundant, many of them are available only in isolated silos and largely inaccessible to external researchers. To overcome and tackle this challenge the university medicine network (comprising all 36 German university hospitals) has been founded in April 2020 to coordinate COVID-19 action plans, diagnostic and therapeutic strategies and collaborative research activities. 13 projects were initiated from which the CODEX project, aiming at the development of a Germany-wide Covid-19 Data Exchange Platform, is presented in this publication. We illustrate the conceptual design, the stepwise development and deployment, first results and the current status.
This article describes some use case studies and self-assessments of FAIR status of de.NBI services to illustrate the challenges and requirements for the definition of the needs of adhering to the FAIR (findable, accessible, interoperable and reusable) data principles in a large distributed bioinformatics infrastructure. We address the challenge of heterogeneity of wet lab technologies, data, metadata, software, computational workflows and the levels of implementation and monitoring of FAIR principles within the different bioinformatics sub-disciplines joint in de.NBI. On the one hand, this broad service landscape and the excellent network of experts are a strong basis for the development of useful research data management plans. On the other hand, the large number of tools and techniques maintained by distributed teams renders FAIR compliance challenging.
ABSTRACT The clinical course of COVID-19 is highly variable, however, underlying host factors and determinants of severe disease are still unknown. Based on single-cell transcriptomes of nasopharyngeal and bronchial samples from clinically well-characterized patients presenting with moderate and critical severities, we reveal the different types and states of airway epithelial cells that are vulnerable for SARS-CoV-2 infection. In COVID-19 patients, we observed a two- to threefold increase of cells expressing the SARS-CoV-2 entry receptor ACE2 within the airway epithelial cell compartment. ACE2 is upregulated in epithelial cells through Interferon signals by immune cells suggesting that the viral defense system may increase the number of potentially susceptible cells in the respiratory epithelium. Infected epithelial cells recruit and activate immune cells by chemokine signaling. Recruited T lymphocytes and inflammatory macrophages were hyperactivated and showed a strong interaction with epithelial cells. In critical patients, increased expression of CCL2, CCL3, CCL5, CXCL9, CXCL10, IL8, IL1B and TNF in macrophages was identified as a likely cause of a hyperinflammatory lung pathology. Moreover, we observed exacerbated epithelial cell death, likely leading to lung injury and respiratory failure in fatal cases. Our study provides novel insights into the pathophysiology of COVID-19 and suggests an immunomodulatory therapy along the CCL2, CCL3/CCR1 axis as promising option to prevent and treat critical course of COVID-19.
AbstractIn COVID-19, hypertension and cardiovascular diseases have emerged as major risk factors for critical disease progression. Concurrently, the impact of the main anti-hypertensive therapies, angiotensin-converting enzyme inhibitors (ACEi) and angiotensin receptor blockers (ARB), on COVID-19 severity is controversially discussed. By combining clinical data, single-cell sequencing data of airway samples andin vitroexperiments, we assessed the cellular and pathophysiological changes in COVID-19 driven by cardiovascular disease and its treatment options. Anti-hypertensive ACEi or ARB therapy, was not associated with an altered expression of SARS-CoV-2 entry receptorACE2in nasopharyngeal epithelial cells and thus presumably does not change susceptibility for SARS-CoV-2 infection. However, we observed a more critical progress in COVID-19 patients with hypertension associated with a distinct inflammatory predisposition of immune cells. While ACEi treatment was associated with dampened COVID-19-related hyperinflammation and intrinsic anti-viral responses, under ARB treatment enhanced epithelial-immune cell interactions were observed. Macrophages and neutrophils of COVID-19 patients with hypertension and cardiovascular comorbidities, in particular under ARB treatment, exhibited higher expression ofCCL3, CCL4, and its receptorCCR1, which associated with critical COVID-19 progression. Overall, these results provide a potential explanation for the adverse COVID-19 course in patients with cardiovascular disease, i.e. an augmented immune response in critical cells for the disease course, and might suggest a beneficial effect of clinical ACEi treatment in hypertensive COVID-19 patients.
In coronavirus disease 2019 (COVID-19), hypertension and cardiovascular diseases are major risk factors for critical disease progression. However, the underlying causes and the effects of the main anti-hypertensive therapies—angiotensin-converting enzyme inhibitors (ACEIs) and angiotensin receptor blockers (ARBs)—remain unclear. Combining clinical data (n = 144) and single-cell sequencing data of airway samples (n = 48) with in vitro experiments, we observed a distinct inflammatory predisposition of immune cells in patients with hypertension that correlated with critical COVID-19 progression. ACEI treatment was associated with dampened COVID-19-related hyperinflammation and with increased cell intrinsic antiviral responses, whereas ARB treatment related to enhanced epithelial–immune cell interactions. Macrophages and neutrophils of patients with hypertension, in particular under ARB treatment, exhibited higher expression of the pro-inflammatory cytokines CCL3 and CCL4 and the chemokine receptor CCR1. Although the limited size of our cohort does not allow us to establish clinical efficacy, our data suggest that the clinical benefits of ACEI treatment in patients with COVID-19 who have hypertension warrant further investigation. Single-cell analysis reveals how anti-hypertensive drugs affect the risk of severe disease in patients with COVID-19 who have hypertension.
To investigate the immune response and mechanisms associated with severe coronavirus disease 2019 (COVID-19), we performed single-cell RNA sequencing on nasopharyngeal and bronchial samples from 19 clinically well-characterized patients with moderate or critical disease and from five healthy controls. We identified airway epithelial cell types and states vulnerable to severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection. In patients with COVID-19, epithelial cells showed an average three-fold increase in expression of the SARS-CoV-2 entry receptor ACE2, which correlated with interferon signals by immune cells. Compared to moderate cases, critical cases exhibited stronger interactions between epithelial and immune cells, as indicated by ligand–receptor expression profiles, and activated immune cells, including inflammatory macrophages expressing CCL2, CCL3, CCL20, CXCL1, CXCL3, CXCL10, IL8, IL1B and TNF. The transcriptional differences in critical cases compared to moderate cases likely contribute to clinical observations of heightened inflammatory tissue damage, lung injury and respiratory failure. Our data suggest that pharmacologic inhibition of the CCR1 and/or CCR5 pathways might suppress immune hyperactivation in critical COVID-19. Single-cell analysis of COVID-19 patient samples identifies activated immune pathways that correlate with severe disease.
In this Article, author Benedikt Brors was erroneously associated with affiliation number '8' (Department of Developmental Neurobiology, St Jude Children's Research Hospital, Memphis, Tennessee, USA); the author's two other affiliations (affiliations '3' and '7', both at the German Cancer Research Center (DKFZ)) were correct. This has been corrected online.
BACKGROUND:Bidirectional promoters (BPs) are prevalent in eukaryotic genomes. However, it is poorly understood how the cell integrates different epigenomic information, such as transcription factor (TF) binding and chromatin marks, to drive gene expression at BPs. Single-cell sequencing technologies are revolutionizing the field of genome biology. Therefore, this study focuses on the integration of single-cell RNA-seq data with bulk ChIP-seq and other epigenetics data, for which single-cell technologies are not yet established, in the context of BPs.RESULTS:We performed integrative analyses of novel human single-cell RNA-seq (scRNA-seq) data with bulk ChIP-seq and other epigenetics data. scRNA-seq data revealed distinct transcription states of BPs that were previously not recognized. We find associations between these transcription states to distinct patterns in structural gene features, DNA accessibility, histone modification, DNA methylation and TF binding profiles.CONCLUSIONS:Our results suggest that a complex interplay of all of these elements is required to achieve BP-specific transcriptional output in this specialized promoter configuration. Further, our study implies that novel statistical methods can be developed to deconvolute masked subpopulations of cells measured with different bulk epigenomic assays using scRNA-seq data.
BioModels serves as a central repository of mathematical models representing biological processes. It offers a platform to make mathematical models easily shareable across the systems modelling community, thereby supporting model reuse. To facilitate hosting a broader range of model formats derived from diverse modelling approaches and tools, a new infrastructure for BioModels has been developed that is available at http://www.ebi.ac.uk/biomodels. This new system allows submitting and sharing of a wide range of models with improved support for formats other than SBML. It also offers a version-control backed environment in which authors and curators can work collaboratively to curate models. This article summarises the features available in the current system and discusses the potential benefit they offer to the users over the previous system. In summary, the new portal broadens the scope of models accepted in BioModels and supports collaborative model curation which is crucial for model reproducibility and sharing.
Pan-cancer analyses that examine commonalities and differences among various cancer types have emerged as a powerful way to obtain novel insights into cancer biology. Here we present a comprehensive analysis of genetic alterations in a pan-cancer cohort including 961 tumours from children, adolescents, and young adults, comprising 24 distinct molecular types of cancer. Using a standardized workflow, we identified marked differences in terms of mutation frequency and significantly mutated genes in comparison to previously analysed adult cancers. Genetic alterations in 149 putative cancer driver genes separate the tumours into two classes: small mutation and structural/copy-number variant (correlating with germline variants). Structural variants, hyperdiploidy, and chromothripsis are linked to TP53 mutation status and mutational signatures. Our data suggest that 7-8% of the children in this cohort carry an unambiguous predisposing germline variant and that nearly 50% of paediatric neoplasms harbour a potentially druggable event, which is highly relevant for the design of future clinical trials.
Incomplete understanding of the metastatic process hinders personalized therapy. Here we report the most comprehensive whole-genome study of colorectal metastases vs. matched primary tumors. 65% of somatic mutations originate from a common progenitor, with 15% being tumor- and 19% metastasis-specific, implicating a higher mutation rate in metastases. Tumor- and metastasis-specific mutations harbor elevated levels of BRCAness. We confirm multistage progression with new components ARHGEF7/ARHGEF33. Recurrently mutated non-coding elements include ncRNAs RP11-594N15.3, AC010091, SNHG14, 3' UTRs of FOXP2, DACH2, TRPM3, XKR4, ANO5, CBL, CBLB, the latter four potentially dual protagonists in metastasis and efferocytosis-/PD-L1 mediated immunosuppression. Actionable metastasis-specific lesions include FAT1, FGF1, BRCA2, KDR, and AKT2-, AKT3-, and PDGFRA-3' UTRs. Metastasis specific mutations are enriched in PI3K-Akt signaling, cell adhesion, ECM and hepatic stellate activation genes, suggesting genetic programs for site-specific colonization. Our results put forward hypotheses on tumor and metastasis evolution, and evidence for metastasis-specific events relevant for personalized therapy.
The One Touch Pipeline (OTP) is an automation platform managing Next-Generation Sequencing (NGS) data and calling bioinformatic pipelines for processing these data. OTP handles the complete digital process from import of raw sequence data via alignment of sequencing reads to identify genomic events in an automated and scalable way. Three major goals are pursued: firstly, reduction of human resources required for data management by introducing automated processes. Secondly, reduction of time until the sequences can be analyzed by bioinformatic experts, by executing all operations more reliably and quickly. Thirdly, storing all information in one system with secure web access and search capabilities. From software architecture perspective, OTP is both information center and workflow management system. As a workflow management system, OTP call several NGS pipelines that can easily be adapted and extended according to new requirements. As an information center, it comprises a database for metadata information as well as a structured file system. Based on complete and consistent information, data management and bioinformatic pipelines within OTP are executed automatically with all steps book-kept in a database.
The International Cancer Genome Consortium (ICGC)’s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.
Knittel, G. and Liedgens, P. and Korovkina, D. and Seeger, J.M. and Al-Baldawi, Y. and Al-Maarri, M. and Fritz, C. and Vlantis, K. and Bezhanova, S. and Scheel, A.H. and Wolz, O.O. and Reimann, M. and Moeller, P. and Lopez, C. and Schlesner, M. and Lohneis, P. and Weber, A.N.R. and Truemper, L. and Staudt, L.M. and Ortmann, M. and Pasparakis, M. and Siebert, R. and Schmitt, C.A. and Klatt, A.R. and Wunderlich, F.T. and Schaefer, S.C. and Persigehl, T. and MontesinosRongen, M. and Odenthal, M. and Buettner, R. and Frenzel, L.P. and Kashkar, H. and Reinhardt, H.C.
Here, we report a novel segmentation-based method, TEPIC, to predict TF binding by combining sets of open-chromatin regions with position weight matrices. TEPIC can be applied to various open-chromatin data, e.g. DNaseI-seq and NOMe-seq, using either peaks or footprints as input. TEPIC computes TF affinities and uses open-chromatin signal intensity as quantitative measures of TF binding strength. Using elastic net regression, we show that incorporating low affinity binding sites improves our ability to explain gene expression. In our application, gene-based scores computed by TEPIC with one open-chromatin assay as input nearly reach the quality of several TF ChIP-seq datasets.
The binding and contribution of transcription factors (TF) to cell specific gene expression is often deduced from open-chromatin measurements to avoid costly TF ChIP-seq assays. Thus, it is important to develop computational methods for accurate TF binding prediction in open-chromatin regions (OCRs). Here, we report a novel segmentation-based method, TEPIC, to predict TF binding by combining sets of OCRs with position weight matrices. TEPIC can be applied to various open-chromatin data, e.g. DNaseI-seq and NOMe-seq. Additionally, Histone-Marks (HMs) can be used to identify candidate TF binding sites. TEPIC computes TF affinities and uses open-chromatin/HM signal intensity as quantitative measures of TF binding strength. Using machine learning, we find low affinity binding sites to improve our ability to explain gene expression variability compared to the standard presence/absence classification of binding sites. Further, we show that both footprints and peaks capture essential TF binding events and lead to a good prediction performance. In our application, gene-based scores computed by TEPIC with one open-chromatin assay nearly reach the quality of several TF ChIP-seq datasets. Finally, these scores correctly predict known transcriptional regulators as illustrated by the application to novel DNaseI-seq and NOMe-seq data for primary human hepatocytes and CD4+ T-cells, respectively.
The impact of epigenetics on the differentiation of memory T (Tmem) cells is poorly defined. We generated deep epigenomes comprising genome-wide profiles of DNA methylation, histone modifications, DNA accessibility, and coding and non-coding RNA expression in naive, central-, effector-, and terminally differentiated CD45RA(+) CD4(+) Tmem cells from blood and CD69(+) Tmem cells from bone marrow (BMTmem). We observed a progressive and proliferation-associated global loss of DNA methylation in heterochromatic parts of the genome during Tmem cell differentiation. Furthermore, distinct gradually changing signatures in the epigenome and the transcriptomesupported a linear model of memory development in circulating T cells, while tissue-resident BM-Tmembranched off with a unique epigenetic profile. Integrative analyses identified candidate master regulators of Tmem cell differentiation, including the transcription factor FOXP1. This study highlights the importance of epigenomic changes for Tmem cell biology and demonstrates the value of epigenetic data for the identification of lineage regulators.
Medulloblastoma is a highly malignant paediatric brain tumour currently treated with a combination of surgery, radiation and chemotherapy, posing a considerable burden of toxicity to the developing child. Genomics has illuminated the extensive intertumoral heterogeneity of medulloblastoma, identifying four distinct molecular subgroups. Group 3 and group 4 subgroup medulloblastomas account for most paediatric cases; yet, oncogenic drivers for these subtypes remain largely unidentified. Here we describe a series of prevalent, highly disparate genomic structural variants, restricted to groups 3 and 4, resulting in specific and mutually exclusive activation of the growth factor independent 1 family proto-oncogenes, GFI1 and GFI1B. Somatic structural variants juxtapose GFI1 or GFI1B coding sequences proximal to active enhancer elements, including super-enhancers, instigating oncogenic activity. Our results, supported by evidence from mouse models, identify GFI1 and GFI1B as prominent medulloblastoma oncogenes and implicate ‘enhancer hijacking’ as an efficient mechanism driving oncogene activation in a childhood cancer.