
Life science research now heavily relies on all sorts of databases for genome sequences, transcription, protein three-dimensional (3D) structures, protein–protein interactions, phenotypes and so forth. The knowledge accumulated by all the omics research is so vast that a computer-aided search of data is now a prerequisite for starting a new study. In addition, a combinatory search throughout these databases has a chance to extract new ideas and new hypotheses that can be examined by wet-lab experiments. By virtually integrating the related databases on the Internet, we have built a new web application that facilitates life science researchers for retrieving experts’ knowledge stored in the databases and for building a new hypothesis of the research target. This web application, named VaProS, puts stress on the interconnection between the functional information of genome sequences and protein 3D structures, such as structural effect of the gene mutation. In this manuscript, we present the notion of VaProS, the databases and tools that can be accessed without any knowledge of database locations and data formats, and the power of search exemplified in quest of the molecular mechanisms of lysosomal storage disease. VaProS can be freely accessed at http://p4d-info.nig.ac.jp/vapros/ .
structural life science big data.The newly developed methods are applied to the promotion of structural life science in Japan.The efforts to improve the methods for information retrieval from genome and protein data in this project have achieved some helpful results, not only for the project but also for the entire scientific community in this field.Hence we organized this special issue of Big Data Analysis in Structural and Functional Genomics.This special issue includes six papers concerning methods for big data analysis and applications of the methods to life science research.K. Yura et al. reported a new type of computer application, named VaProS, which searches numerous life science databases on the Internet for relevant data and visualizes the integrated results.T. Kawabata introduced a web tool, called HOMCOS, which systematically connects both genome information and protein structure information.K. Kinoshita et al. described the Natural Ligand Database (NLDB), which connects the modified ligands in the PDB, the protein structural database, and the natural ligands in organisms.K. Nagata et al. applied new information retrieval methods to the data analysis of G protein-coupled receptors (GPCRs) and analyzed the relationship between orphan class-A GPCR proteins and diseases.T. Shirai et al. reported an improvement in the method to compare ligands in the PDB.K. Tomii et al. demonstrated the importance of improving the substitution matrix for highly sensitive homology searches against the current huge sequence databases.Using big data in structural and functional genomics through sophisticated retrieval methods should provide new perspectives on problems to be solved.We hope that the databases and tools described here are the ones that pave the way for building novel hypotheses and conducting new experiments in this field.
Nowadays, in scientific fields such as Structural Biology or Vaccinology, there is an increasing need of fast, effective and reproducible gene cloning and expression processes. Consequently, the implementation of robotic platforms enabling the automation of protocols is becoming a pressing demand. The main goal of our study was to set up a robotic platform devoted to the high-throughput automation of the polymerase incomplete primer extension cloning method, and to evaluate its efficiency compared to that achieved manually, by selecting a set of bacterial genes that were processed either in the automated platform (330) or manually (94). Here we show that we successfully set up a platform able to complete, with high efficiency, a wide range of molecular biology and biochemical steps. 329 gene targets (99 %) were effectively amplified using the automated procedure and 286 (87 %) of these PCR products were successfully cloned in expression vectors, with cloning success rates being higher for the automated protocols respect to the manual procedure (93.6 and 74.5 %, respectively).
The period 2000–2015 brought the advent of high-throughput approaches to protein structure determination. With the overall funding on the order of $2 billion (in 2010 dollars), the structural genomics (SG) consortia established worldwide have developed pipelines for target selection, protein production, sample preparation, crystallization, and structure determination by X-ray crystallography and NMR. These efforts resulted in the determination of over 13,500 protein structures, mostly from unique protein families, and increased the structural coverage of the expanding protein universe. SG programs contributed over 4400 publications to the scientific literature. The NIH-funded Protein Structure Initiatives alone have produced over 2000 scientific publications, which to date have attracted more than 93,000 citations. Software and database developments that were necessary to handle high-throughput structure determination workflows have led to structures of better quality and improved integrity of the associated data. Organized and accessible data have a positive impact on the reproducibility of scientific experiments. Most of the experimental data generated by the SG centers are freely available to the community and has been utilized by scientists in various fields of research. SG projects have created, improved, streamlined, and validated many protocols for protein production and crystallization, data collection, and functional analysis, significantly benefiting biological and biomedical research.
Protein database search for public databases is a fundamental step in the target selection of proteins in structural and functional genomics and also for inferring protein structure, function, and evolution. Most database search methods employ amino acid substitution matrices to score amino acid pairs. The choice of substitution matrix strongly affects homology detection performance. We earlier proposed a substitution matrix named MIQS that was optimized for distant protein homology search. Herein we further evaluate MIQS in combination with LAST, a heuristic and fast database search tool with a tunable sensitivity parameter m, where larger m denotes higher sensitivity. Results show that MIQS substantially improves the homology detection and alignment quality performance of LAST across diverse m parameters. Against a protein database consisting of approximately 15 million sequences, LAST with m = 105 achieves better homology detection performance than BLASTP, and completes the search 20 times faster. Compared to the most sensitive existing methods being used today, CS-BLAST and SSEARCH, LAST with MIQS and m = 106 shows comparable homology detection performance at 2.0 and 3.9 times greater speed, respectively. Results demonstrate that MIQS-powered LAST is a time-efficient method for sensitive and accurate homology search.
More than 800 G protein-coupled receptor (GPCR) genes have been discovered in the human genome. Towards the next step in GPCR research, we performed a knowledge-driven analysis of orphan class-A GPCRs that may serve as novel targets in drug discovery. We examined the relationship between 61 orphan class-A GPCR genes and diseases using the Online Mendelian Inheritance in Man (OMIM) database and the DDSS tool. The OMIM database contains data on disease-related variants of the genes. Particularly, the variants of GPR101, GPR161, and GPR88 are related to the genetic diseases: growth hormone-secreting pituitary adenoma 2, pituitary stalk interruption syndrome (not confirmed), and childhood-onset chorea with psychomotor retardation, respectively. On the other hand, the Drug Discovery and Diagnostic Support System (DDSS) tool suggests that 48 out of the 61 orphan receptor genes are related to diseases, judging from their co-occurrences in abstracts of biomedical literature. Notably, GPR50 and GPR3 are related to as many as 25 and 24 disease-associated keywords, respectively. GPR50 is related to 17 keywords of psychiatric disorders, whereas GPR3 is related to 11 keywords of neurological disorders. The aforementioned five orphan GPCRs were characterized genetically, structurally and functionally using the structural life science data cloud VaProS, so as to evaluate their potential as next targets in drug discovery.
We present a new method for predicting protein–ligand-binding sites based on protein three-dimensional structure and amino acid conservation. This method involves calculation of the van der Waals interaction energy between a protein and many probes placed on the protein surface and subsequent clustering of the probes with low interaction energies to identify the most energetically favorable locus. In addition, it uses amino acid conservation among homologous proteins. Ligand-binding sites were predicted by combining the interaction energy and the amino acid conservation score. The performance of our prediction method was evaluated using a non-redundant dataset of 348 ligand-bound and ligand-unbound protein structure pairs, constructed by filtering entries in a ligand-binding site structure database, LigASite. Ligand-bound structure prediction (bound prediction) indicated that 74.0 % of predicted ligand-binding sites overlapped with real ligand-binding sites by over 25 % of their volume. Ligand-unbound structure prediction (unbound prediction) indicated that 73.9 % of predicted ligand-binding residues overlapped with real ligand-binding residues. The amino acid conservation score improved the average prediction accuracy by 17.0 and 17.6 points for the bound and unbound predictions, respectively. These results demonstrate the effectiveness of the combined use of the interaction energy and amino acid conservation in the ligand-binding site prediction.
Tight control of protein synthesis is necessary for cells to respond and adapt to environmental changes rapidly. Eukaryotic translation initiation factor (eIF) 2B, the guanine nucleotide exchange factor for eIF2, is a key target of translation control at the initiation step. The nucleotide exchange activity of eIF2B is inhibited by the stress-induced phosphorylation of eIF2. As a result, the level of active GTP-bound eIF2 is lowered, and protein synthesis is attenuated. eIF2B is a large multi-subunit complex composed of five different subunits, and all five of the subunits are the gene products responsible for the neurodegenerative disease, leukoencephalopathy with vanishing white matter. However, the overall structure of eIF2B has remained unresolved, due to the difficulty in preparing a sufficient amount of the eIF2B complex. To overcome this problem, we established the recombinant expression and purification method for eIF2B from the fission yeast Schizosaccharomyces pombe. All five of the eIF2B subunits were co-expressed and reconstructed into the complex in Escherichia coli cells. The complex was successfully purified with a high yield. This recombinant eIF2B complex contains each subunit in an equimolar ratio, and the size exclusion chromatography analysis suggests it forms a heterodecamer, consistent with recent reports. This eIF2B increased protein synthesis in the reconstituted in vitro human translation system. In addition, disease-linked mutations led to subunit dissociation. Furthermore, we crystallized this functional recombinant eIF2B, and the crystals diffracted to 3.0 Å resolution.
Mutations in Plasmodium falciparum gene kelch13 (pfkelch13) are strongly and causally associated with resistance to anti-malarial drug artemisinin, but their effects on PfKelch13 structure and function remain unclear. Utilizing the publicly available three-dimensional structure of PfKech13 (PDB ID: 4yy8), we find that most of the mutations in its propeller domain occur in two spatial clusters. Of these, one cluster is enriched in surface exposed residues which may drive PfKelch13-centered protein interactions, and the second cluster mostly contains residues which are buried and whose mutations may destabilize PfKelch13 structure. The most prevalent resistant mutations C580Y and Y493H are distal from the above two clusters. The C580Y mutation creates sterically unfavourable contacts while Y493H possibly alters the hydrophobic core of the propeller domain. These analyses will facilitate further experimental studies aimed at understanding how mutations in pfkelch13 lead to artemisinin resistance.
Premeltons are examples of emergent-structures (i.e., structural-solitons) that arise spontaneously in DNA due to the presence of nonlinear-excitations in its structure. They are of two kinds: B–B (or A–A) premeltons form at specific DNA-regions to nucleate site-specific DNA melting. These are stationary and, being globally-nontopological, undergo breather-motions that allow drugs and dyes to intercalate into DNA. B–A (or A–B) premeltons, on the other hand, are mobile, and being globally-topological, act as phase-boundaries transforming B- into A-DNA during the structural phase-transition. They are not expected to undergo breather motions. A key feature of both types of premeltons is the presence of an intermediate structural-form in their central regions (proposed as being a transition-state intermediate in DNA-melting and in the B- to A-transition), which differs from either A- or B-DNA. Called beta-DNA, this is both metastable and hyperflexible—and contains an alternating sugar-puckering pattern along the polymer backbone combined with the partial unstacking (in its lower energy-forms) of every-other base-pair. Beta-DNA is connected to either B- or to A-DNA on either side by boundaries possessing a gradation of nonlinear structural-change, these being called the kink and the antikink regions. The presence of premeltons in DNA leads to a unifying theory to understand much of DNA physical chemistry and molecular biology. In particular, premeltons are predicted to define the 5′ and 3′ ends of genes in naked-DNA and DNA in active-chromatin, this having important implications for understanding physical aspects of the initiation, elongation and termination of RNA-synthesis during transcription. For these and other reasons, the model will be of broader interest to the general-audience working in these areas. The model explains a wide variety of data, and carries with it a number of experimental predictions—all readily testable—as will be described in this review.
The fast heuristic graph match algorithm for small molecules, COMPLIG, was improved by adding a structural superposition process to verify the atom–atom matching. The modified method was used to classify the small molecule ligands in the Protein Data Bank (PDB) by their three-dimensional structures, and 16,660 types of ligands in the PDB were classified into 7561 clusters. In contrast, a classification by a previous method (without structure superposition) generated 3371 clusters from the same ligand set. The characteristic feature in the current classification system is the increased number of singleton clusters, which contained only one ligand molecule in a cluster. Inspections of the singletons in the current classification system but not in the previous one implied that the major factors for the isolation were differences in chirality, cyclic conformations, separation of substructures, and bond length. Comparisons between current and previous classification systems revealed that the superposition-based classification was effective in clustering functionally related ligands, such as drugs targeted to specific biological processes, owing to the strictness of the atom–atom matching.
NLDB (Natural Ligand DataBase; URL: http://nldb.hgc.jp) is a database of automatically collected and predicted 3D protein–ligand interactions for the enzymatic reactions of metabolic pathways registered in KEGG. Structural information about these reactions is important for studying the molecular functions of enzymes, however a large number of the 3D interactions are still unknown. Therefore, in order to complement such missing information, we predicted protein–ligand complex structures, and constructed a database of the 3D interactions in reactions. NLDB provides three different types of data resources; the natural complexes are experimentally determined protein–ligand complex structures in PDB, the analog complexes are predicted based on known protein structures in a complex with a similar ligand, and the ab initio complexes are predicted by docking simulations. In addition, NLDB shows the known polymorphisms found in human genome on protein structures. The database has a flexible search function based on various types of keywords, and an enrichment analysis function based on a set of KEGG compound IDs. NLDB will be a valuable resource for experimental biologists studying protein–ligand interactions in specific reactions, and for theoretical researchers wishing to undertake more precise simulations of interactions.
The HOMCOS server (http://homcos.pdbj.org) was updated for both searching and modeling the 3D complexes for all molecules in the PDB. As compared to the previous HOMCOS server, the current server targets all of the molecules in the PDB including proteins, nucleic acids, small compounds and metal ions. Their binding relationships are stored in the database. Five services are available for users. For the services “Modeling a Homo Protein Multimer” and “Modeling a Hetero Protein Multimer”, a user can input one or two proteins as the queries, while for the service “Protein-Compound Complex”, a user can input one chemical compound and one protein. The server searches similar molecules by BLAST and KCOMBU. Based on each similar complex found, a simple sequence-replaced model is quickly generated by replacing the residue names and numbers with those of the query protein. A target compound is flexibly superimposed onto the template compound using the program fkcombu. If monomeric 3D structures are input as the query, then template-based docking can be performed. For the service “Searching Contact Molecules for a Query Protein”, a user inputs one protein sequence as the query, and then the server searches for its homologous proteins in PDB and summarizes their contacting molecules as the predicted contacting molecules. The results are summarized in “Summary Bars” or “Site Table”display. The latter shows the results as a one-site-one-row table, which is useful for annotating the effects of mutations. The service “Searching Contact Molecules for a Query Compound” is also available.
ZFAT is a transcriptional regulator, containing eighteen C2H2-type zinc-fingers and one AT-hook, involved in autoimmune thyroid disease, apoptosis, and immune-related cell survival. We determined the solution structures of the thirteen individual ZFAT zinc-fingers (ZF) and the tandemly arrayed zinc-fingers in the regions from ZF2 to ZF5, by NMR spectroscopy. ZFAT has eight uncommon bulged-out helix-containing zinc-fingers, and six of their structures (ZF4, ZF5, ZF6, ZF10, ZF11, and ZF13) were determined. The distribution patterns of the putative DNA-binding surface residues are different among the ZFAT zinc-fingers, suggesting the distinct DNA sequence preferences of the N-terminal and C-terminal zinc-fingers. Since ZFAT has three to five consecutive tandem zinc-fingers, which may cooperatively function as a unit, we also determined two tandemly arrayed zinc-finger structures, between ZF2 to ZF4 and ZF3 to ZF5. Our NMR spectroscopic analysis detected the interaction between ZF4 and ZF5, which are connected by an uncommon linker sequence, KKIK. The ZF4–ZF5 linker restrained the relative structural space between the two zinc-fingers in solution, unlike the other linker regions with determined structures, suggesting the involvement of the ZF4–ZF5 interfinger linker in the regulation of ZFAT function.
Multiprotein complexes play essential roles in all cells and X-ray crystallography can provide unparalleled insight into their structure and function. Many of these complexes are believed to be sufficiently stable for structural biology studies, but the production of protein–protein complexes using recombinant technologies is still labor-intensive. We have explored several strategies for the identification and cloning of heterodimers and heterotrimers that are compatible with the high-throughput (HTP) structural biology pipeline developed for single proteins. Two approaches are presented and compared which resulted in co-expression of paired genes from a single expression vector. Native operons encoding predicted interacting proteins were selected from a repertoire of genomes, and cloned directly to expression vector. In an alternative approach, Helicobacter pylori proteins predicted to interact strongly were cloned, each associated with translational control elements, then linked into an artificial operon. Proteins were then expressed and purified by standard HTP protocols, resulting to date in the structure determination of two H. pylori complexes.
Vectors designed for protein production in Escherichia coli and by wheat germ cell-free translation were tested using 21 well-characterized eukaryotic proteins chosen to serve as controls within the context of a structural genomics pipeline. The controls were carried through cloning, small-scale expression trials, large-scale growth or synthesis, and purification. Successfully purified proteins were also subjected to either crystallization trials or 1 H– 15 N HSQC NMR analyses. Experiments evaluated: (1) the relative efficacy of restriction/ligation and recombinational cloning systems; (2) the value of maltose-binding protein (MBP) as a solubility enhancement tag; (3) the consequences of in vivo proteolysis of the MBP fusion as an alternative to post-purification proteolysis; (4) the effect of the level of LacI repressor on the yields of protein obtained from E. coli using autoinduction; (5) the consequences of removing the His tag from proteins produced by the cell-free system; and (6) the comparative performance of E. coli cells or wheat germ cell-free translation. Optimal promoter/repressor and fusion tag configurations for each expression system are discussed.
The methylmalonyl Co-A mutase-associated GTPase MeaB from Methylobacterium extorquens is involved in glyoxylate regulation and required for growth. In humans, mutations in the homolog methylmalonic aciduria associated protein (MMAA) cause methylmalonic aciduria, which is often fatal. The central role of MeaB from bacteria to humans suggests that MeaB is also important in other, pathogenic bacteria such as Mycobacterium tuberculosis . However, the identity of the mycobacterial MeaB homolog is presently unclear. Here, we identify the M. tuberculosis protein Rv1496 and its homologs in M. smegmatis and M. thermoresistibile as MeaB. The crystal structures of all three homologs are highly similar to MeaB and MMAA structures and reveal a characteristic three-domain homodimer with GDP bound in the G domain active site. A structure of Rv1496 obtained from a crystal grown in the presence of GTP exhibited electron density for GDP, suggesting GTPase activity. These structures identify the mycobacterial MeaB and provide a structural framework for therapeutic targeting of M. tuberculosis MeaB.
The MazG family proteins, which are highly conserved in bacteria, are nucleoside triphosphate pyrophosphohydrolases that hydrolyze all canonical nucleoside triphosphates, and are also involved in removing noncanonical nucleoside triphosphates to prevent their incorporation into DNA or RNA. The primary structure of TM0360 from Thermotoga maritima MSB8 suggested that TM0360 is a MazG-related nucleoside triphosphate pyrophosphohydrolase. The crystal structure of the TM0360 protein was determined by the MAD technique at 2.0 Å resolution. The asymmetric unit contains an intact dimer molecule. The overall structure of TM0360 is similar to the known structures of the dimeric MazG protein and dUTPases. The putative NTP binding pocket in TM0360, identified by considering the probable NTP-interacting residues and structural features, suggested that TM0360 resembles the C-terminal domain of Escherichia coli MazG, although TM0360 may be a truncated paralog of the N-terminal domain of T. maritima MazG (TM0913), according to its primary structure. The putative function of TM0360 is discussed, based on structural homology.
Working with a combination of ProMOL (a plugin for PyMOL that searches a library of enzymatic motifs for local structural homologs), BLAST and Pfam (servers that identify global sequence homologs), and Dali (a server that identifies global structural homologs), we have begun the process of assigning functional annotations to the approximately 3,500 structures in the Protein Data Bank that are currently classified as having "unknown function". Using a limited template library of 388 motifs, over 500 promising in silico matches have been identified by ProMOL, among which 65 exceptionally good matches have been identified. The characteristics of the exceptionally good matches are discussed.
The adiponectin receptors (AdipoR1 and AdipoR2) are membrane proteins with seven transmembrane helices. These receptors regulate glucose and fatty acid metabolism, thereby ameliorating type 2 diabetes. The full-length human AdipoR1 and a series of N-terminally truncated mutants of human AdipoR1 and AdipoR2 were expressed in insect cells. In small-scale size exclusion chromatography, the truncated mutants AdipoR1Δ88 (residues 89–375) and AdipoR2Δ99 (residues 100–386) eluted mostly in the intact monodisperse state, while the others eluted primarily as aggregates. However, gel filtration chromatography of the large-scale preparation of the tag-affinity-purified AdipoR1Δ88 revealed the presence of an excessive amount of the aggregated state over the intact state. Since aggregation due to contaminating nucleic acids may have occurred during the sample concentration step, anion-exchange column chromatography was performed immediately after affinity chromatography, to separate the intact AdipoR1Δ88 from the aggregating species. The separated intact AdipoR1Δ88 did not undergo further aggregation, and was successfully purified to homogeneity by gel filtration chromatography. The purified AdipoR1Δ88 and AdipoR2Δ99 proteins were characterized by thermostability assays with 7-diethylamino-3-(4-maleimidophenyl)-4-methyl coumarin, thin layer chromatography of bound lipids, and surface plasmon resonance analysis of ligand binding, demonstrating their structural integrities. The AdipoR1Δ88 and AdipoR2Δ99 proteins were crystallized with the anti-AdipoR1 monoclonal antibody Fv fragment, by the lipidic mesophase method. X-ray diffraction data sets were obtained at resolutions of 2.8 and 2.4 Å, respectively.