Testing the effect of rare variants on phenotypic variation is difficult due to the need for extremely large cohorts to identify associated variants given expected effect sizes. An alternative approach is to investigate the effect of rare genetic variants on low-level genomic traits, such as gene expression or DNA methylation (DNAm), as effect sizes are expected to be larger for low-level compared to higher-order complex traits. Here, we investigate DNAm in healthy ageing populations - the Lothian Birth cohorts of 1921 and 1936 and identify both transient and stable outlying DNAm levels across the genome. We find an enrichment of rare genetic variants within 1kb of DNAm sites in individuals with stable outlying DNAm, implying genetic control of this extreme variation. Using a family-based cohort, the Brisbane Systems Genetics Study, we observed increased sharing of DNAm outliers among more closely related individuals, consistent with these outliers being driven by rare genetic variation. We demonstrated that outlying DNAm levels have a functional consequence on gene expression levels, with extreme levels of DNAm being associated with gene expression levels towards the tails of the population distribution. Overall, this study demonstrates the role of rare variants in the phenotypic variation of low-level genomic traits, and the effect of extreme levels of DNAm on gene expression.
The extent to which the impact of regulatory genetic variants may depend on other factors, such as the expression levels of upstream transcription factors, remains poorly understood. Here we report a framework in which regulatory variants are first aggregated into sets, and using these as estimates of the total cis-genetic effects on a gene we model their non-additive interactions with the expression of other genes in the genome. Using 1220 lymphoblastoid cell lines across platforms and independent datasets we identify 74 genes where the impact of their regulatory variant-set is linked to the expression levels of networks of distal genes. We show that these networks are predominantly associated with tumourigenesis pathways, through which immortalised cells are able to rapidly proliferate. We consequently present an approach to define gene interaction networks underlying important cellular pathways such as cell immortalisation.
This study highlights dangers in over-interpreting fine-mapping results. Chundru et al. show that genotype imputation accuracy has a large impact on fine-mapping accuracy. They used DNA methylation at CpG-sites with a variant... Genetic variants disrupting DNA methylation at CpG dinucleotides (CpG-SNP) provide a set of known causal variants to serve as models to test fine-mapping methodology. We use 1716 CpG-SNPs to test three fine-mapping approaches (Bayesian imputation-based association mapping, Bayesian sparse linear mixed model, and the J-test), assessing the impact of imputation errors and the choice of reference panel by using both whole-genome sequence (WGS), and genotype array data on the same individuals (n = 1166). The choice of imputation reference panel had a strong effect on imputation accuracy, with the 1000 Genomes Project Phase 3 (1000G) reference panel (n = 2504 from 26 populations) giving a mean nonreference discordance rate between imputed and sequenced genotypes of 3.2% compared to 1.6% when using the Haplotype Reference Consortium (HRC) reference panel (n = 32,470 Europeans). These imputation errors had an impact on whether the CpG-SNP was included in the 95% credible set, with a difference of ∼23% and ∼7% between the WGS and the 1000G and HRC imputed datasets, respectively. All of the fine-mapping methods failed to reach the expected 95% coverage of the CpG-SNP. This is attributed to secondary cis genetic effects that are unable to be statistically separated from the CpG-SNP, and through a masking mechanism where the effect of the methylation disrupting allele at the CpG-SNP is hidden by the effect of a nearby SNP that has strong linkage disequilibrium with the CpG-SNP. The reduced accuracy in fine-mapping a known causal variant in a low-level biological trait with imputed genetic data has implications for the study of higher-order complex traits and disease.
Despite the fundamental importance of single nucleotide polymorphisms (SNPs) to human evolution there are still large gaps in our understanding of the forces that shape their distribution across the genome. SNPs have been shown to not be distributed evenly, with directly adjacent SNPs found unusually frequently. Why this is the case is unclear. We illustrate how neighbouring SNPs that can’t be explained by a single mutation event (that we term here sequential dinucleotide mutations, SDMs) are driven by distinct mutational processes and selective pressures to SNPs and multinucleotide polymorphisms (MNPs). By studying variation across multiple populations, including a novel cohort of 1,358 Scottish genomes, we show that, SDMs are over twice as common as MNPs and like SNPs, display distinct mutational spectra across populations. These biases are though not only different to those observed among SNPs and MNPs, but also more divergent between human population groups. We show that the changes that make up SDMs are not independent, and identify a distinct mutational profile, CA → CG → TG, that is observed an order of magnitude more often than other SDMs, including others that involve the gain and subsequent deamination of CpG sites. This suggests these specific changes are driven by a distinct process. In coding regions particular SDMs are favoured, and especially those that lead to the creation of single codon amino acids. Intriguingly selection has favoured particular pathways through the amino acid code, with epistatic selection appearing to have disfavoured sequential non-synonymous changes.
Genome-wide association studies have to date identified multiple coronary artery disease (CAD)-associated loci; however, for most of these loci the mechanism by which they affect CAD risk is unclear. The CAD-associated locus 7q32.2 is unusual in that the lead variant, rs11556924, is not in strong linkage disequilibrium with any other variant and introduces a coding change in ZC3HC1, which encodes NIPA. In this study, we show that rs11556924 polymorphism is associated with lower regulatory phosphorylation of NIPA in the risk variant, resulting in NIPA with higher activity. Using a genome-editing approach we show that this causes an effective decrease in cyclin-B1 stability in the nucleus, thereby slowing its nuclear accumulation. By perturbing the rate of nuclear cyclin-B1 accumulation, rs11556924 alters the regulation of mitotic progression resulting in an extended mitosis. This study shows that the CAD-associated coding polymorphism in ZC3HC1 alters the dynamics of cell-cycle regulation by NIPA.
Complement is an important part of the immune system. It is initiated through three different pathways known as the classical, lectin, and alternative pathway. The multimolecular C1 complex of the classical pathway consists of a subcomponent, C1q, which binds to a tetramer comprising two C1r and two C1s proteases. A detailed description of the structure of the C1 complex is essential to fully understand how the complex acts on pathogens. A variety of different models have been proposed, which differ mainly in the way the proteases interact with C1q. In this study, we have used a combination of homology‐based structure prediction and molecular dynamics to predict a partial structure of the C1s/C1r/C1r/C1s tetramer. For computational expediency the study was restricted to the CUB 1 ‐EGF‐CUB 2 domains which are directly involved in the formation of the tetramer and its interaction with C1q; the catalytic fragments (CCP 1 ‐CCP 2 ‐SP), which mediate C1 activation and subsequent cleavage of substrates, were omitted. A systematic molecular dynamics (MD) study of several possible dimeric combinations suggest that the tetramer is formed when a pair of C1r/C1s dimers form a “doughnut” via a C1s/C1s head‐to‐tail interaction, which is stabilized by several putative salt bridges at the dimer interface. This result is consistent with biochemical data which have shown that self assembly requires the formation of C1r‐C1s contacts and that electrostatic interactions play a key role. Furthermore, it identifies a number of putative binding residues that can be tested using site‐directed mutagenesis. Proteins 2012. © 2012 Wiley Periodicals, Inc
Antibodies are routinely used as research tools, in diagnostic assays and increasingly as therapeutics. Ideally, these applications require antibodies with high sensitivity and specificity; however, many commercially available antibodies are limited in their use as they cross-react with non-related proteins. Here we describe a novel method to characterize antibody specificity. Six commercially available monoclonal and polyclonal antibodies were screened on high-density protein arrays comprising of ~10,000 recombinant human proteins (Imagenes). Two of the six antibodies examined; anti-pICln and anti-GAPDH, bound exclusively to their target antigen and showed no cross-reactivity with non-related proteins. However, four of the antibodies, anti-HSP90, anti-HSA, anti-bFGF and anti-Ro52, showed strong cross-reactivity with other proteins on the array. Antibody-antigen interactions were readily confirmed using Western immunoblotting. In addition, the redundant nature of the protein array used, enabled us to define the epitopic region within HSP90 of the anti-HSP90 antibody, and identify possible shared epitopes in cross-reacting proteins. In conclusion, high-density protein array technology is a fast and effective means for determining the specificity of antibodies and can be used to further improve the accuracy of antibody applications.
We have studied the initial stages of hydrolysis in the monomeric aspartic proteinases by performing ab initio calculations (Hartree–Fock, MP2, DFT) on the active site of endothiapepsin. The crystallographic coordinates of a difluorostatone inhibitor complexed with endothiapepsin were used to generate a template for the initial enzyme/substrate complex. The active site was modelled as a formic acid–formate anion moiety (representing the catalytic aspartates, Asp-32 and Asp-215) and a bound water molecule. Acetamide was used to represent the substrate, and residues Gly-34, Ser-35, Gly-217, and Thr-218, which all interact with the active site, were modelled using formamide and methanol. The calculations predict that T218 plays a direct role in catalysis by forming a strong hydrogen bond with the water oxygen during the initial stages of nucleophilic attack. S35 plays an indirect role by lowering the pKa of Asp-32 and facilitating the formation of the gem–diol intermediate. The predicted reaction pathway is unusual because it involves two local minima, with very short Owater...Csubstrate distances, which are stabilised by T218.
The Internet consists of a vast inhomogeneous reservoir of data. Developing software that can integrate a wide variety of different data sources is a major challenge that must be addressed for the realisation of the full potential of the Internet as a scientific research tool. This article presents a semi-automated object-oriented programming system for integrating web-based resources. We demonstrate that the current Internet standards (HTML, CGI [common gateway interface], Java™, etc.) can be exploited to develop a data retrieval system that scans existing web interfaces and then uses a set of rules to generate new Java™ code that can automatically retrieve data from the Web. The validity of the software has been demonstrated by testing it on several biological databases. We also examine the current limitations of the Internet and discuss the need for the development of universal standards for web-based data.
Mitogen-activated protein kinase (MAPK) cascades are universal and highly conserved signal transduction modules in eucaryotes, including plants. These protein phosphorylation cascades link extracellular stimuli to a wide range of cellular responses. However, the underlying mechanisms are so far unknown as information about phosphorylation substrates of plant MAPKs is lacking. In this study we addressed the challenging task of identifying potential substrates for Arabidopsis thaliana mitogen-activated protein kinases MPK3 and MPK6, which are activated by many environmental stress factors. For this purpose, we developed a novel protein microarray-based proteomic method allowing high throughput study of protein phosphorylation. We generated protein microarrays including 1,690 Arabidopsis proteins, which were obtained from the expression of an almost nonredundant uniclone set derived from an inflorescence meristem cDNA expression library. Microarrays were incubated with MAPKs in the presence of radioactive ATP. Using a threshold-based quantification method to evaluate the microarray results, we were able to identify 48 potential substrates of MPK3 and 39 of MPK6. 26 of them are common for both kinases. One of the identified MPK6 substrates, 1-aminocyclopropane-1-carboxylic acid synthase-6, was just recently shown as the first plant MAPK substrate in vivo, demonstrating the potential of our method to identify substrates with physiological relevance. Furthermore we revealed transcription factors, transcription regulators, splicing factors, receptors, histones, and others as candidate substrates indicating that regulation in response to MAPK signaling is very complex and not restricted to the transcriptional level. Nearly all of the 48 potential MPK3 substrates were confirmed by other in vitro methods. As a whole, our approach makes it possible to shortlist candidate substrates of mitogen-activated protein kinases as well as those of other protein kinases for further analysis. Follow-up in vivo experiments are essential to evaluate their physiological relevance.
There is burgeoning interest in protein microarrays, but a source of thousands of nonredundant, purified proteins was not previously available. Here we show a glass chip containing 2413 nonredundant purified human fusion proteins on a polymer surface, where densities up to 1600 proteins/cm(2) on a microscope slide can be realized. In addition, the polymer coating of the glass slide enables screening of protein interactions under nondenaturing conditions. Such screenings require only 200-microl sample volumes, illustrating their potential for high-throughput applications. Here we demonstrate two applications: the characterization of antibody binding, specificity, and cross-reactivity; and profiling the antibody repertoire in body fluids, such as serum from patients with autoimmune diseases. For the first application, we have incubated these protein chips with anti-RGSHis(6), anti-GAPDH, and anti-HSP90beta antibodies. In an initial proof of principle study for the second application, we have screened serum from alopecia and arthritis patients. With analysis of large sample numbers, identification of disease-associated proteins to generate novel diagnostic markers may be possible.
We have studied the initial stages of hydrolysis in the aspartic proteinases by performing ab initio Hartree–Fock calculations on a small portion of the active site of endothiapepsin. The crystallographic coordinates of a difluorostatone inhibitor complexed with endothiapepsin were used to generate a template for the initial enzyme–substrate complex. The active site was modelled as a formic acid–formate anion moiety (representing the catalytic aspartates, Asp-32 and -215) and a bound water molecule. Residues Gly-34, Ser-35, Gly-217, and Thr-218, which all form hydrogen bonds to the active site, were modelled using two acetamide and two methanol molecules. The calculations predict that nucleophilic attack by the active site water results in the formation of a tetrahedral oxyanion intermediate which subsequently abstracts the acidic proton from Asp-32 to form a gem-diol intermediate. Thr-218 appears to play an important role in catalysis by forming an additional hydrogen bond with the water oxygen during the early stages of nucleophilic attack. Ser-35 lowers the pKa of Asp-32 and this facilitates the formation of the gem-diol intermediate.
The serine and cysteine proteinases represent two important classes of enzymes that use a catalytic triad to hydrolyze peptides and esters. The active site of the serine proteinases consists of three key residues, Asp...His...Ser. The hydroxyl group of serine functions as a nucleophile and the imidazole ring of histidine functions as a general acid/general base during catalysis. Similarly, the active site of the cysteine proteinases also involves three key residues: Asn, His, and Cys. The active site of the cysteine proteinases is generally believed to exist as a zwitterion (Asn...His(+)...Cys(-)) with the thiolate anion of the cysteine functioning as a nucleophile during the initial stages of catalysis. Curiously, the mutant serine proteinases, thiol subtilisin and thiol trypsin, which have the hybrid Asp...His...Cys triad, are almost catalytically inert. In this study, ab initio Hartree-Fock calculations have been performed on the active sites of papain and the mutant serine proteinase S195C rat trypsin. These calculations predict that the active site of papain exists predominately as a zwitterion (Cys(-)...His(+)...Asn). However, similar calculations on S195C rat trypsin demonstrate that the thiol mutant is unable to form a reactive thiolate anion prior to catalysis. Furthermore, structural comparisons between native papain and S195C rat trypsin have demonstrated that the spatial juxtapositions of the triad residues have been inverted in the serine and cysteine proteinases and, on this basis, I argue that it is impossible to convert a serine proteinase to a cysteine proteinase by site-directed mutagenesis.
Theoretical one-electron reduction potentials, E1, have been determined for a set of eight nitroarene hypoxic cell radiosensitisers using a combination of classical statistical mechanics and quantum mechanical methods. Gas-phase electron affinities were calculated using ab initio Hartree–Fock calculations and relative hydration energies were computed using the free energy perturbation (FEP) method. The results were used to estimate the relative one-electron reduction potentials for these molecules in solution. In general, the computed results are in good agreement with experiment although further work is required to determine the limitations of the method. Nevertheless, the method shows sufficient promise to be of value in the rational design of improved oxidative agents for use as hypoxia-selective radiosensitisers and bioreductivity activated cytotoxins.
We have performed ab initio Hartree-Fock self-consistent field calculations on the active site of endothiapepsin. The active site was modeled as a formic acid/formate anion moiety (representing the catalytic aspartates, Asp-32 and -215) and a bound water molecule. Residues Gly-34, Ser-35, Gly-217, and Thr-218, which all form hydrogen bonds to the active site, were modeled using formamide and methanol molecules. The water molecule, which is generally believed to function as the attacking nucleophile in catalysis, was allowed to bind to the active site in four distinct configurations. The geometry of each configuration was optimized using two basis sets (4-31G and 4-31G*). The results indicate that in the native enzyme the nucleophilic water is bound in a catalytically inert configuration. However, by rotating the carboxyl group of Asp-32 by about 90 degrees the water molecule can be reorientated to attack the scissile bond of the substrate. A model of the bound enzyme-substrate complex was constructed from the crystal structure of a difluorostatone inhibitor complexed with endothiapepsin. This model suggests that the substrate itself initiates the reorientation of the nucleophilic water immediately prior to catalysis by forcing the carboxyl group of Asp-32 to rotate. The theoretical results predict that the active site of endothiapepsin undergoes a large distortion during substrate binding and this observation has been used to explain some of the kinetics results which have been reported for mutant aspartic proteinases.
Dienelactone hydrolase (DLH), an enzyme from the beta-ketoadipate pathway, catalyses the hydrolysis of dienelactone to maleylacetate. DLH is unusual because it is the only known naturally occurring enzyme which contains the catalytic triad Cys ... His ... Asp. This triad has previously been created artificially in the mutant serine proteases, thiol subtilisin and thiol trypsin. In both cases the mutant enzymes exhibited activities several orders of magnitude lower than the wild type enzymes; the low reactivity has generally been attributed to the inability of these enzymes to form a catalytically active thiolate anion (Cys- ... His+ ... Asp-). The crystal structure of DLH suggests that the native enzyme exists predominantly in a catalytically inert configuration; the triad cysteine is neutral and points away from the active site binding cleft. However, a crystallographic analysis of C123S DLH complexed with an isostructural inhibitor (dienelactam) indicates that substrate binding induces a prototropic rearrangement of the active site prior to catalysis which results in the formation of a highly nucleophilic thiolate anion. We have performed ab initio SCF/MP2 calculations on a relatively small portion of the active site of DLH to examine the details of this activation process. Our calculations provide supporting evidence that the conformational changes observed in the crystal structure due to inhibitor (or substrate) binding facilitate the formation of a reactive thiolate anion. In particular, substrate binding alters the position of Glu36; the carboxylate side chain of Glu36 is pushed towards C123 enabling it to abstract the thiol proton thus creating a catalytically active thiolate anion.(ABSTRACT TRUNCATED AT 250 WORDS)
The aspartic proteinases are a family of proteolytic enzymes known to be involved in several medical disorders (such as hypertension, AIDS and cancer) and are therefore an important therapeutic target for the rational design of inhibitors. The crystal structures of several enzyme/inhibitor complexes have already been reported. Many of these inhibitors contain a hydroxyl group which forms hydrogen bonds to the carboxyl groups of the active site aspartates, D32 and D215, (pepsin numbering). In many cases it is not possible to identify unambiguously the hydrogen bonds formed between the inhibitor and the catalytic aspartates from the crystal structures, and therefore one cannot determine which of these inhibitors are good transition state analogues. We have performed ab initio Hartree-Fock calculations on a small portion of the active site of endothiapepsin complexed with three potent inhibitors: pepstatin A, and two peptide analogues which contain difluorostatone and phosphostatine to identify the hydrogen bonds formed between the inhibitor and the active site aspartates. Our calculations suggest that inhibitors which contain a gem-diol group (e.g. difluorostatone-containing inhibitors) are more likely to be good transition state analogues.
We have performed ab initio Hartree-Fock calculations on the active site of endothiapepsin and two “mutants” S35A and T218A. The active site which carries a formal negative charge to effect hydrolysis was modelled as a formic acid/formate anion moiety and a water molecule. The four nearest hydrogen bonding residues (Gly34, Ser35, Gly217 and Thr218) were modelled as formamide and methanol molecules. The most stable active site configuration for wild-type endothiapepsin and the two mutants has the water molecule bound across the shortest OD32 · OD215 diagonal. The inner oxygen of D32 is protonated and the lone pairs of the water oxygen point away from the active site cleft. In this configuration the water molecule would be expected to be catalytically inert suggesting that a prototropic rearrangement of the active site must occur during substrate binding to enable the water molecule to attack the scissile bond.
We have performed ab initio Hartree-Fock calculations on the active site of endothiapepsin and two "mutants" S35A and T218A. The active site which carries a formal negative charge to effect hydrolysis was modelled as a formic acid/formate anion moiety and a water molecule. The four nearest hydrogen bonding residues (Gly34, Ser35, Gly217 and Thr218) were modelled as formamide and methanol molecules. The most stable active site configuration for wild-type endothiapepsin and the two mutants has the water molecule bound across the shortest O(D32)...O(D215) diagonal. The inner oxygen of D32 is protonated and the lone pairs of the water oxygen point away from the active site cleft. In this configuration the water molecule would be expected to be catalytically inert suggesting that a prototropic rearrangement of the active site must occur during substrate binding to enable the water molecule to attack the scissile bond.
We have performed ab initio self-consistent field (SCF) and configuration interaction (CI) calculations on the active site of the aspartic proteinases pepsin and endothiapepsin. The active site, which carries a formal negative charge to effect hydrolysis, was modeled as a formic acid/formate anion moiety and a water molecule, and the nearest hydrogen bonding residues (Gly34, Ser35, Gly217, and Thr218, with respect to the residue numbering in endothiapepsin) were modeled as formamide and methanol molecules. Four possible binding modes for the active-site water molecule were considered. In contrast to previous theoretical studies, we predict that the most stable form has the water molecule forming a bifurcated hydrogen bond to the inner oxygens of Asp32 and -215, with Asp32 being ionized. The calculations suggest that the water molecule prefers to bind across the shortest OD32 ... OD215 diagonal of the active-site carboxyl groups and therefore the binding mode of the water molecule for all the native aspartic proteinases can be readily predicted by measuring these distances.