Molecular biology holds a vast potential for tackling climate change and biodiversity loss. Yet, it is largely absent from the current strategies. We call for a community-wide action to bring molecular biology to the forefront of climate change solutions.
The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.
Technological advances have continuously driven the generation of biomolecular data and the development of bioinformatics infrastructure, which enables data reuse for scientific discovery. Several types of data management resources have arisen, such as data deposition databases, addedvalue databases or knowledgebases, and biology-driven portals. In this review, we provide a unique overview of the gradual evolution of these resources and discuss the goals and features that must be considered in their development. With the increasing application of genomics in the health care context and with 60 to 500 million whole genomes estimated to be sequenced by 2022, biomedical research infrastructure is transforming, too. Systems for federated access, portable tools, provision of reference data, and interpretation tools will enable researchers to derive maximal benefits from these data. Collaboration, coordination, and sustainability of data resources are key to ensure that biomedical knowledge management can scale with technology shifts and growing data volumes.
Drug discovery and development pipelines are long, complex and depend on numerous factors. Machine learning (ML) approaches provide a set of tools that can improve discovery and decision making for well-specified questions with abundant, high-quality data. Opportunities to apply ML occur in all stages of drug discovery. Examples include target validation, identification of prognostic biomarkers and analysis of digital pathology data in clinical trials. Applications have ranged in context and methodology, with some approaches yielding accurate predictions and insights. The challenges of applying ML lie primarily with the lack of interpretability and repeatability of ML-generated results, which may limit their application. In all areas, systematic and comprehensive high-dimensional data still need to be generated. With ongoing efforts to tackle these issues, as well as increasing awareness of the factors needed to validate ML approaches, the application of ML can promote data-driven decision making and has the potential to speed up the process and reduce failure rates in drug discovery and development.
Next-Generation Sequencing (NGS) technologies are expected to play a crucial role in the surveillance of infectious diseases, with their unprecedented capabilities for the characterisation of genetic information underlying the virulence and antimicrobial resistance (AMR) properties of microorganisms. In the implementation of any novel technology for regulatory purposes, important considerations such as harmonisation, validation and quality assurance need to be addressed. NGS technologies pose unique challenges in these regards, in part due to their reliance on bioinformatics for the processing and proper interpretation of the data produced. Well-designed benchmark resources are thus needed to evaluate, validate and ensure continued quality control over the bioinformatics component of the process. This concept was explored as part of a workshop on "Next-generation sequencing technologies and antimicrobial resistance" held October 4-5 2017. Challenges involved in the development of such a benchmark resource, with a specific focus on identifying the molecular determinants of AMR, were identified. For each of the challenges, sets of unsolved questions that will need to be tackled for them to be properly addressed were compiled. These take into consideration the requirement for monitoring of AMR bacteria in humans, animals, food and the environment, which is aligned with the principles of a “One Health” approach.
Summary Objectives: To highlight and provide insights into key developments in translational bioinformatics between 2014 and 2016. Methods: This review describes some of the most influential bioinformatics papers and resources that have been published between 2014 and 2016 as well as the national genome sequencing initiatives that utilize these resources to routinely embed genomic medicine into healthcare. Also discussed are some applications of the secondary use of patient data followed by a comprehensive view of the open challenges and emergent technologies. Results: Although data generation can be performed routinely, analyses and data integration methods still require active research and standardization to improve streamlining of clinical interpretation. The secondary use of patient data has resulted in the development of novel algorithms and has enabled a refined understanding of cellular and phenotypic mechanisms. New data storage and data sharing approaches are required to enable diverse biomedical communities to contribute to genomic discovery. Conclusion: The translation of genomics data into actionable knowledge for use in healthcare is transforming the clinical landscape in an unprecedented way. Exciting and innovative models that bridge the gap between clinical and academic research are set to open up the field of translational bioinformatics for rapid growth in a digital era.
The Global Alliance for Genomics and Health (GA4GH), the standards-setting body in genomics for healthcare, aims to accelerate biomedical advancement globally. We describe the differences between healthcare- and research-driven genomics, discuss the implications of global, population-scale collections of human data for research, and outline mission-critical considerations in ethics, regulation, technology, data protection, and society. We present a crude model for estimating the rate of healthcare-funded genomes worldwide that accounts for the preparedness of each country for genomics, and infers a progression of cancer-related sequencing over time. We estimate that over 60 million patients will have their genome sequenced in a healthcare context by 2025. This represents a large technical challenge for healthcare systems, and a huge opportunity for research. We identify eight major practical, principled arguments to support the position that virtual cohorts of 100 million people or more would have tangible research benefits.
We have designed and developed a data integration and visualization platform that provides evidence about the association of known and potential drug targets with diseases. The platform is designed to support identification and prioritization of biological targets for follow-up. Each drug target is linked to a disease using integrated genome-wide data from a broad range of data sources. The platform provides either a target-centric workflow to identify diseases that may be associated with a specific target, or a disease-centric workflow to identify targets that may be associated with a specific disease. Users can easily transition between these target- and disease-centric workflows. The Open Targets Validation Platform is accessible at https://www.targetvalidation.org.
Researchers in the pharmaceutical industry require information from different sources to make decisions about the association between a potential drug target and a disease. This information is dispersed and not easily accessible to them without the support of specialised data scientists (bioinformaticians). The Open Targets* web platform aims to support researchers in identifying early drug targets faster and with more confidence. The platform integrates biological data from several public databases, presents them in a coherent way and allows researchers to interpret the presented information more easily. Our poster outlines how we applied a range of User Experience (UX) methods (interviews, observations, design workshops, paper prototyping and user testing) to understand the needs of our users. These methods allowed us to design intuitive and reusable visualisations for the Open Targets platform following an iterative Agile development framework. This approach enabled the members of our multidisciplinary team to collaborate with each other and with drug discovery researchers to develop a comprehensive and intuitive target identification platform. *Open Targets was formerly called Centre for Therapeutic Target Validation (CTTV)
ABSTRACT The hepatitis C virus (HCV) NS4B protein is an antiviral therapeutic target for which small-molecule inhibitors have not been shown to exhibit in vivo efficacy. We describe here the in vitro and in vivo antiviral activity of GSK8853, an imidazo[1,2- a ]pyrimidine inhibitor that binds NS4B protein. GSK8853 was active against multiple HCV genotypes and developed in vitro resistance mutations in both genotype 1a and genotype 1b replicons localized to the region of NS4B encoding amino acids 94 to 105. A 20-day in vitro treatment of replicons with GSK8853 resulted in a 2-log drop in replicon RNA levels, with no resistance mutation breakthrough. Chimeric replicons containing NS4B sequences matching known virus isolates showed similar responses to a compound with genotype 1a sequences but altered efficacy with genotype 1b sequences, likely corresponding to the presence of known resistance polymorphs in those isolates. In vivo efficacy was tested in a humanized-mouse model of HCV infection, and the results showed a 3-log drop in viral RNA loads over a 7-day period. Analysis of the virus remaining at the end of in vivo treatment revealed resistance mutations encoding amino acid changes that had not been identified by in vitro studies, including NS4B N56I and N99H. Our findings provide an in vivo proof of concept for HCV inhibitors targeting NS4B and demonstrate both the promise and potential pitfalls of developing NS4B inhibitors.
GSK2485852 (referred to here as GSK5852) is a hepatitis C virus (HCV) NS5B polymerase inhibitor with 50% effective concentrations (EC(50)s) in the low nanomolar range in the genotype 1 and 2 subgenomic replicon system as well as the infectious HCV cell culture system. We have characterized the antiviral activity of GSK5852 using chimeric replicon systems with NS5B genes from additional genotypes as well as NS5B sequences from clinical isolates of patients infected with HCV of genotypes 1a and 1b. The inhibitory activity of GSK5852 remained unchanged in these intergenotypic and intragenotypic replicon systems. GSK5852 furthermore displays an excellent resistance profile and shows a<5-fold potency loss across the clinically important NS5B resistance mutations P495L, M423T, C316Y, and Y448H. Testing of a diverse mutant panel also revealed a lack of cross-resistance against known resistance mutations in other viral proteins. Data from both the newer 454 sequencing method and traditional population sequencing showed a pattern of mutations arising in the NS5B RNA-dependent RNA polymerase in replicon cells exposed to GSK5852. GSK5852 was more potent than HCV-796, an earlier inhibitor in this class, and showed greater reductions in HCV RNA during long-term treatment of replicons. GSK5852 is similar to HCV-796 in its activity against multiple genotypes, but its superior resistance profile suggests that it could be an attractive component of an all-oral regimen for treating HCV.
Improving drug attrition remains a challenge in pharmaceutical discovery and development. A major cause of early attrition is the demonstration of safety signals which can negate any therapeutic index previously established. Safety attrition needs to be put in context of clinical translation (i.e. human relevance) and is negatively impacted by differences between animal models and human. In order to minimize such an impact, an earlier assessment of pharmacological target homology across animal model species will enhance understanding of the context of animal safety signals and aid species selection during later regulatory toxicology studies. Here we sequenced the genomes of the Sus scrofa Göttingen minipig and the Canis familiaris beagle, two widely used animal species in regulatory safety studies. Comparative analyses of these new genomes with other key model organisms, namely mouse, rat, cynomolgus macaque, rhesus macaque, two related breeds (S. scrofa Duroc and C. familiaris boxer) and human reveal considerable variation in gene content. Key genes in toxicology and metabolism studies, such as the UGT2 family, CYP2D6, and SLCO1A2, displayed unique duplication patterns. Comparisons of 317 known human drug targets revealed surprising variation such as species-specific positive selection, duplication and higher occurrences of pseudogenized targets in beagle (41 genes) relative to minipig (19 genes). These data will facilitate the more effective use of animals in biomedical research.
GSK2336805 is an inhibitor of hepatitis C virus (HCV) with picomolar activity on the standard genotype 1a, 1b, and 2a subgenomic replicons and exhibits a modest serum shift. GSK2336805 was not active on 22 RNA and DNA viruses that were profiled. We have identified changes in the N-terminal region of NS5A that cause a decrease in the activity of GSK2336805. These mutations in the genotype 1b replicon showed modest shifts in compound activity (<13-fold), while mutations identified in the genotype 1a replicon had a more dramatic impact on potency. GSK2336805 retained activity on chimeric replicons containing NS5A patient sequences from genotype 1 and patient and consensus sequences for genotypes 4 and 5 and part of genotype 6. Combination and cross-resistance studies demonstrated that GSK2336805 could be used as a component of a multidrug HCV regimen either with the current standard of care or in combination with compounds with different mechanisms of action that are still progressing through clinical development.
SIRT6 is involved in inflammation, aging and metabolism potentially by modulating the functions of both NFκB and HIF1α. Since it is possible to make small molecule activators and inhibitors of Sirtuins we wished to establish biochemical and cellular assays both to assist in drug discovery efforts and to validate whether SIRT6 represents a valid drug target for these indications. We confirmed in cellular assays that SIRT6 can deacetylate acetylated-histone H3 lysine 9 (H3K9Ac), however this deacetylase activity is unusually low in biochemical assays. In an effort to develop alternative assay formats we observed that SIRT6 overexpression had no influence on TNFα induced nuclear translocation of NFκB, nor did it have an effect on nuclear mobility of RelA/p65. In an effort to identify a gene expression profile that could be used to identify a SIRT6 readout we conducted genome-wide expression studies. We observed that overexpression of SIRT6 had little influence on NFκB-dependent genes, but overexpression of the catalytically inactive mutant affected gene expression in developmental pathways.
Abstract The recently updated and complete genome of the C57BL/6J mouse strain provides a model mammalian system for genetics, comparative genomics and evolutionary studies. The extensive freely available resources of mouse genetics and breeds and similarity between mouse and much of human biology makes the mouse genome the primary choice to enable discernment of the biological function of human and other mammalian genes. Of huge importance is that along with human, mouse is currently the only mammalian genome to be sequenced to completeness allowing the investigation of lineage specific biology. In particular the mouse genome sequence provides an unrivalled resource for medical bioscience, in encouraging a deeper understanding of the shared mammalian evolutionary history of potential drug targets. Key Concepts: The availability of high quality genome sequence is the cornerstone of comparative genomics. Although many mammalian genomes have been sequenced with high coverage they are considered ‘drafts’. Only the mouse and human genomes are characterised as ‘complete’ and do not contain gaps in their genome coverage. The comparison of genes and genomes allows the investigation of evolutionary history of species. The availability of multiple animal genomes provides the raw material for the discipline of genome zoology.
SIRT6 is a member of the Sirtuin family of histone deacetylases that has been implicated in inflammatory, aging and metabolic pathways. Some of its actions have been suggested to be via physical interaction with NFκB and HIF1α and transcriptional regulation through its histone deacetylase activity. Our previous studies have investigated the histone deacetylase activity of SIRT6 and explored its ability to regulate the transcriptional responses to an inflammatory stimulus such as TNFα. In order to develop a greater understanding of SIRT6 function we have sought to identify SIRT6 interacting proteins by both yeast-2-hybrid and co-immunoprecipitation studies. We report a number of interacting partners which strengthen previous findings that SIRT6 functions in base excision repair (BER), and novel interactors which suggest a role in nucleosome and chromatin remodeling, the cell cycle and NFκB biology.
ABSTRACT There is a global emergence of multidrug-resistant (MDR) strains of Klebsiella pneumoniae , a Gram-negative enteric bacterium that causes nosocomial and urinary tract infections. While the epidemiology of K. pneumoniae strains and occurrences of specific antibiotic resistance genes, such as plasmid-borne extended-spectrum β-lactamases (ESBLs), have been extensively studied, only four complete genomes of K. pneumoniae are available. To better understand the multidrug resistance factors in K. pneumoniae , we determined by pyrosequencing the nearly complete genome DNA sequences of two strains with disparate antibiotic resistance profiles, broadly drug-susceptible strain JH1 and strain 1162281, which is resistant to multiple clinically used antibiotics, including extended-spectrum β-lactams, fluoroquinolones, aminoglycosides, trimethoprim, and sulfamethoxazoles. Comparative genomic analysis of JH1, 1162281, and other published K. pneumoniae genomes revealed a core set of 3,631 conserved orthologous proteins, which were used for reconstruction of whole-genome phylogenetic trees. The close evolutionary relationship between JH1 and 1162281 relative to other K. pneumoniae strains suggests that a large component of the genetic and phenotypic diversity of clinical isolates is due to horizontal gene transfer. Using curated lists of over 400 antibiotic resistance genes, we identified all of the elements that differentiated the antibiotic profile of MDR strain 1162281 from that of susceptible strain JH1, such as the presence of additional efflux pumps, ESBLs, and multiple mechanisms of fluoroquinolone resistance. Our study adds new and significant DNA sequence data on K. pneumoniae strains and demonstrates the value of whole-genome sequencing in characterizing multidrug resistance in clinical isolates.
Next-generation sequencing (NGS) technologies represent a paradigm shift in sequencing capability. The technology has already been extensively applied to biological research, resulting in significant and remarkable insights into the molecular biology of cells. In this review, we focus on current and potential applications of the technology as applied to the drug discovery and development process. Early applications have focused on the oncology and infectious disease therapeutic areas, with emerging use in biopharmaceutical development and vaccine production in evidence. Although this technology has great potential, significant challenges remain, particularly around the storage, transfer and analysis of the substantial data sets generated.