Summary Beacon v2 is an API specification established by the Global Alliance for Genomics and Health initiative (GA4GH) that defines a standard for federated discovery of genomic and phenotypic data. Here, we present the Beacon v2 Reference Implementation (B2RI), a set of open-source software tools that allow lighting up a local Beacon instance ‘out-of-the-box’. Along with the software, we have created detailed ‘Read the Docs’ documentation that includes information on deployment and installation. Availability and implementation The B2RI is released under GNU General Public License v3.0 and Apache License v2.0. Documentation and source code is available at: https://b2ri-documentation.readthedocs.io. Supplementary information Supplementary data are available at Bioinformatics online.
Since its launch in 2008, the European Genome-Phenome Archive (EGA) has been leading the archiving and distribution of human identifiable genomic data. In this regard, one of the community concerns is the potential usability of the stored data, as of now, data submitters are not mandated to perform any quality control (QC) before uploading their data and associated metadata information. Here, we present a new File QC Portal developed at EGA, along with QC reports performed and created for 1 694 442 files [Fastq, sequence alignment map (SAM)/binary alignment map (BAM)/CRAM and variant call format (VCF)] submitted at EGA. QC reports allow anonymous EGA users to view summary-level information regarding the files within a specific dataset, such as quality of reads, alignment quality, number and type of variants and other features. Researchers benefit from being able to assess the quality of data prior to the data access decision and thereby, increasing the reusability of data (https://ega-archive.org/blog/data-upcycling-powered-by-ega/).
Beacon is a basic data discovery protocol issued by the Global Alliance for Genomics and Health (GA4GH). The main goal addressed by version 1 of the Beacon protocol was to test the feasibility of broadly sharing human genomic data, through providing simple "yes" or "no" responses to queries about the presence of a given variant in datasets hosted by Beacon providers. The popularity of this concept has fostered the design of a version 2, that better serves real-world requirements and addresses the needs of clinical genomics research and healthcare, as assessed by several contributing projects and organizations. Particularly, rare disease genetics and cancer research will benefit from new case level and genomic variant level requests and the enabling of richer phenotype and clinical queries as well as support for fuzzy searches. Beacon is designed as a "lingua franca" to bridge data collections hosted in software solutions with different and rich interfaces. Beacon version 2 works alongside popular standards like Phenopackets, OMOP, or FHIR, allowing implementing consortia to return matches in beacon responses and provide a handover to their preferred data exchange format. The protocol is being explored by other research domains and is being tested in several international projects.
The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.
The European Genome-phenome Archive (EGA - https://ega-archive.org/) is a resource for long term secure archiving of all types of potentially identifiable genetic, phenotypic, and clinical data resulting from biomedical research projects. Its mission is to foster hosted data reuse, enable reproducibility, and accelerate biomedical and translational research in line with the FAIR principles. Launched in 2008, the EGA has grown quickly, currently archiving over 4,500 studies from nearly one thousand institutions. The EGA operates a distributed data access model in which requests are made to the data controller, not to the EGA, therefore, the submitter keeps control on who has access to the data and under which conditions. Given the size and value of data hosted, the EGA is constantly improving its value chain, that is, how the EGA can contribute to enhancing the value of human health data by facilitating its submission, discovery, access, and distribution, as well as leading the design and implementation of standards and methods necessary to deliver the value chain. The EGA has become a key GA4GH Driver Project, leading multiple development efforts and implementing new standards and tools, and has been appointed as an ELIXIR Core Data Resource.
Background: Whole-exome sequencing (WES) has become an efficient diagnostic test for patients with likely monogenic conditions such as rare idiopathic diseases or sudden unexplained death. Yet, many cases remain undiagnosed. Here, we report the added diagnostic yield achieved for 101 WES cases re-analyzed 1 to 7 years after initial analysis. Methods: Of the 101 WES cases, 51 were rare idiopathic disease cases and 50 were postmortem "molecular autopsy" cases of early sudden unexplained death. Variants considered for reporting were prioritized and classified into three groups: (1) diagnostic variants, pathogenic and likely pathogenic variants in genes known to cause the phenotype of interest; (2) possibly diagnostic variants, possibly pathogenic variants in genes known to cause the phenotype of interest or pathogenic variants in genes possibly causing the phenotype of interest; and (3) variants of uncertain diagnostic significance, potentially deleterious variants in genes possibly causing the phenotype of interest. Results: Initial analysis revealed diagnostic variants in 13 rare disease cases (25.4%) and 5 sudden death cases (10%). Re-analysis resulted in the identification of additional diagnostic variants in 3 rare disease cases (5.9%) and 1 sudden unexplained death case (2%), which increased our molecular diagnostic yield to 31.4% and 12%, respectively. Conclusions: The basis of new findings ranged from improvement in variant classification tools, updated genetic databases, and updated clinical phenotypes. Our findings highlight the potential for re-analysis to reveal diagnostic variants in cases that remain undiagnosed after initial WES.
[This corrects the article on p. 72 in vol. 4, PMID: 29181379.].
Purpose: Nail-Patella syndrome is a dominantly inherited genetic disorder characterized by abnormalities of the nails, knees, elbows, and pelvis. Nail abnormalities are the most constant feature of Nail-Patella syndrome. Pathogenic mutations in a single gene, LMX1B , a mesenchymal determinant of dorsal-ventral patterning, explain approximately 95% of Nail-Patella syndrome cases. However, 5% of cases remain unexplained. Methods: Here, we present exome sequencing and analysis of four generations of a family with a dominantly inherited Nail-Patella-like disorder (nail dysplasia with some features of Nail-Patella syndrome) who tested negative for LMX1B mutation. Results: We identify a loss-of-function mutation in WIF1 (NM_007191 p.W15*), which is involved in mesoderm segmentation, as the suspected cause of the Nail-Patella-like disorder observed in this family. Conclusions: Mutation of WIF1 is a potential novel cause of a Nail-Patella-like disorder. Testing of additional patients negative for LMX1B mutation is needed to confirm this finding and further clarify the phenotype. Genet Med advance online publication 06 April 2017
Whole genome and exome sequencing usually include reads containing mitochondrial DNA (mtDNA). Yet, state-of-the-art pipelines and services for human nuclear genome variant calling and annotation do not handle mitochondrial genome data appropriately. As a consequence, any researcher desiring to add mtDNA variant analysis to their investigations is forced to explore the literature for mtDNA pipelines, evaluate them, and implement their own instance of the desired tool. This task is far from trivial, and can be prohibitive for non-bioinformaticians.
Despite its remarkable importance in the arena of drug design, serotonin 1A receptor (5-HT1A) has been elusive to the X-ray crystallography community. This lack of direct structural information not only hampers our knowledge regarding the binding modes of many popular ligands (including the endogenous neurotransmitter-serotonin), but also limits the search for more potent compounds. In this paper we shed new light on the 3D pharmacological properties of the 5-HT1A receptor by using a ligand-guided approach (ALiBERO) grounded in the Internal Coordinate Mechanics (ICM) docking platform. Starting from a homology template and set of known actives, the method introduces receptor flexibility via Normal Mode Analysis and Monte Carlo sampling, to generate a subset of pockets that display enriched discrimination of actives from inactives in retrospective docking. Here, we thoroughly investigated the repercussions of using different protein templates and the effect of compound selection on screening performance. Finally, the best resulting protein models were applied prospectively in a large virtual screening campaign, in which two new active compounds were identified that were chemically distinct from those described in the literature.
To understand cellular processes at the molecular level we need to improve our knowledge of protein-protein interactions, from a structural, mechanistic, and energetic point of view. Current theoretical studies and computational docking simulations show that protein dynamics plays a key role in protein association and support the need for including protein flexibility in modeling protein interactions. Assuming the conformational selection binding mechanism, in which the unbound state can sample bound conformers, one possible strategy to include flexibility in docking predictions would be the use of conformational ensembles originated from unbound protein structures. Here we present an exhaustive computational study about the use of precomputed unbound ensembles in the context of protein docking, performed on a set of 124 cases of the Protein-Protein Docking Benchmark 3.0. Conformational ensembles were generated by conformational optimization and refinement with MODELLER and by short molecular dynamics trajectories with AMBER. We identified those conformers providing optimal binding and investigated the role of protein conformational heterogeneity in protein-protein recognition. Our results show that a restricted conformational refinement can generate conformers with better binding properties and improve docking encounters in medium-flexible cases. For more flexible cases, a more extended conformational sampling based on Normal Mode Analysis was proven helpful. We found that successful conformers provide better energetic complementarity to the docking partners, which is compatible with recent views of binding association. In addition to the mechanistic considerations, these findings could be exploited for practical docking predictions of improved efficiency.
DNA replication initiation is a vital and tightly regulated step in all replicons and requires an initiator factor that specifically recognizes the DNA replication origin and starts replication. RepB from the promiscuous streptococcal plasmid pMV158 is a hexameric ring protein evolutionary related to viral initiators. Here we explore the conformational plasticity of the RepB hexamer by i) SAXS, ii) sedimentation experiments, iii) molecular simulations and iv) X-ray crystallography. Combining these techniques, we derive an estimate of the conformational ensemble in solution showing that the C-terminal oligomerisation domains of the protein form a rigid cylindrical scaffold to which the N-terminal DNA-binding/catalytic domains are attached as highly flexible appendages, featuring multiple orientations. In addition, we show that the hinge region connecting both domains plays a pivotal role in the observed plasticity. Sequence comparisons and a literature survey show that this hinge region could exists in other initiators, suggesting that it is a common, crucial structural element for DNA binding and manipulation.
Studies of long-lived individuals have revealed few genetic mechanisms for protection against age-associated disease. Therefore, we pursued genome sequencing of a related phenotype-healthy aging-to understand the genetics of disease-free aging without medical intervention. In contrast with studies of exceptional longevity, usually focused on centenarians, healthy aging is not associated with known longevity variants, but is associated with reduced genetic susceptibility to Alzheimer and coronary artery disease. Additionally, healthy aging is not associated with a decreased rate of rare pathogenic variants, potentially indicating the presence of disease-resistance factors. In keeping with this possibility, we identify suggestive common and rare variant genetic associations implying that protection against cognitive decline is a genetic component of healthy aging. These findings, based on a relatively small cohort, require independent replication. Overall, our results suggest healthy aging is an overlapping but distinct phenotype from exceptional longevity that may be enriched with disease-protective genetic factors. VIDEO ABSTRACT.
During the last decade we witnessed how computational docking methods became a crucial tool in the search for new drug candidates. The ‘central dogma’ of small molecule docking is that compounds that dock correctly into the receptor are more likely to display biological activity than those that do not dock. This ‘dogma’, however, possesses multiple twists and turns that may not be obvious to novice dockers. The first premise is that the compounds must dock; this implies: (i) availability of data, (ii) realistic representation of the chemical entities in a form that can be understood by the computer and the software, and, (iii) exhaustive sampling of the protein-ligand conformational space. The second premise is that, after the sampling, all docking solutions must be ranked correctly with a score representing the physico-chemical foundations of binding. The third premise is that ‘correctness’ must be defined unambiguously, usually by comparison with ‘static’ experimental data (or lack thereof). Each of these premises involves some degree of simplification of reality, and overall loss in the accuracy of the docking predictions.In this chapter we will revise our latest experiences in receptor-based docking when dealing with all three above-mentioned issues. First, we will explain the theoretical foundation of ICM docking, along with a brief explanation on how we measure performance. Second, we will contextualize ICM by showing its performance in single and multiple receptor conformation schemes with the Directory of Useful Decoys (DUD) and the Pocketome. Third, we will describe which strategies we are using to represent protein plasticity, like using multiple crystallographic structures or Monte Carlo (MC) and Normal Mode Analysis (NMA) sampling methods, emphasizing how to overcome the associated pitfalls (e.g., increased number of false positives). In the last section, we will describe ALiBERO, a new tool that is helping us to improve the discriminative power of X-ray structures and homology models in screening campaigns.
BACKGROUND:Most of the proteins in the Protein Data Bank (PDB) are oligomeric complexes consisting of two or more subunits that associate by rotational or helical symmetries. Despite the myriad of superimposition tools in the literature, we could not find any able to account for rotational symmetry and display the graphical results in the web browser.RESULTS:BioSuper is a free web server that superimposes and calculates the root mean square deviation (RMSD) of protein complexes displaying rotational symmetry. To the best of our knowledge, BioSuper is the first tool of its kind that provides immediate interactive visualization of the graphical results in the browser, biomolecule generator capabilities, different levels of atom selection, sequence-dependent and structure-based superimposition types, and is the only web tool that takes into account the equivalence of atoms in side chains displaying symmetry ambiguity. BioSuper uses ICM program functionality as a core for the superimpositions and displays the results as text, HTML tables and 3D interactive molecular objects that can be visualized in the browser or in Android and iOS platforms with a free plugin.CONCLUSIONS:BioSuper is a fast and functional tool that allows for pairwise superimposition of proteins and assemblies displaying rotational symmetry. The web server was created after our own frustration when attempting to superimpose flexible oligomers. We strongly believe that its user-friendly and functional design will be of great interest for structural and computational biologists who need to superimpose oligomeric proteins (or any protein). BioSuper web server is freely available to all users at http://ablab.ucsd.edu/BioSuper.
After decades of using urea as denaturant, the kinetic role of this molecule in the unfolding process is still undefined: does urea actively induce protein unfolding or passively stabilize the unfolded state? By analyzing a set of 30 proteins (representative of all native folds) through extensive molecular dynamics simulations in denaturant (using a range of force-fields), we derived robust rules for urea unfolding that are valid at the proteome level. Irrespective of the protein fold, presence or absence of disulphide bridges, and secondary structure composition, urea concentrates in the first solvation shell of quasi-native proteins, but with a density lower than that of the fully unfolded state. The presence of urea does not alter the spontaneous vibration pattern of proteins. In fact, it reduces the magnitude of such vibrations, leading to a counterintuitive slow down of the atomic-motions that opposes unfolding. Urea stickiness and slow diffusion is, however, crucial for unfolding. Long residence urea molecules placed around the hydrophobic core are crucial to stabilize partially open structures generated by thermal fluctuations. Our simulations indicate that although urea does not favor the formation of partially open microstates, it is not a mere spectator of unfolding that simply displaces to the right of the folded ←→ unfolded equilibrium. On the contrary, urea actively favors unfolding: it selects and stabilizes partially unfolded microstates, slowly driving the protein conformational ensemble far from the native one and also from the conformations sampled during thermal unfolding.
Receptor models generated by homology or even obtained by crystallography often have their binding pockets suboptimal for ligand docking and virtual screening applications due to insufficient accuracy or induced fit bias. Knowledge of previously discovered receptor ligands provides key information that can be used for improving docking and screening performance of the receptor. Here, we present a comprehensive ligand-guided receptor optimization (LiBERO) algorithm that exploits ligand information for selecting the best performing protein models from an ensemble. The energetically feasible protein conformers are generated through normal mode analysis and Monte Carlo conformational sampling. The algorithm allows iteration of the conformer generation and selection steps until convergence of a specially developed fitness function which quantifies the conformer's ability to select known ligands from decoys in a small-scale virtual screening test. Because of the requirement for a large number of computationally intensive docking calculations, the automated algorithm has been implemented to use Linux clusters allowing easy parallel scaling. Here, we will discuss the setup of LiBERO calculations, selection of parameters, and a range of possible uses of the algorithm which has already proven itself in several practical applications to binding pocket optimization and prospective virtual ligand screening.
Docking and virtual screening (VS) reach maximum potential when the receptor displays the structural changes needed for accurate ligand binding. Unfortunately, these conformational changes are often poorly represented in experimental structures or homology models, debilitating their docking performance. Recently, we have shown that receptors optimized with our LiBERO method (Ligand-guided Backbone Ensemble Receptor Optimization) were able to better discriminate active ligands from inactives in flexible-ligand VS docking experiments. The LiBERO method relies on the use of ligand information for selecting the best performing individual pockets from ensembles derived from normal-mode analysis or Monte Carlo. Here we present ALiBERO, a new computational tool that has expanded the pocket selection from single to multiple, allowing for automatic iteration of the sampling-selection procedure. The selection of pockets is performed by a dual method that uses exhaustive combinatorial search plus individual addition of pockets, selecting only those that maximize the discrimination of known actives compounds from decoys. The resulting optimized pockets showed increased VS performance when later used in much larger unrelated test sets consisting of biologically active and inactive ligands. In this paper we will describe the design and implementation of the algorithm, using as a reference the human estrogen receptor alpha.