Bacillus thuringiensis strains are well known for the production of insecticidal proteins upon sporulation and these proteins are deposited in parasporal crystalline inclusions. The majority of these insect-specific toxins exhibit three domains in the mature toxin sequence. However, other Cry toxins are structurally and evolutionarily unrelated to this three-domain family and little is known of their three dimensional structures, limiting our understanding of their mechanisms of action and our ability to engineer the proteins to enhance their function. Among the non-three domain Cry toxins, the Cry34Ab1 and Cry35Ab1 proteins from B. thuringiensis strain PS149B1 are required to act together to produce toxicity to the western corn rootworm (WCR) Diabrotica virgifera virgifera Le Conte via a pore forming mechanism of action. Cry34Ab1 is a protein of ∼14 kDa with features of the aegerolysin family (Pfam06355) of proteins that have known membrane disrupting activity, while Cry35Ab1 is a ∼44 kDa member of the toxin_10 family (Pfam05431) that includes other insecticidal proteins such as the binary toxin BinA/BinB. The Cry34Ab1/Cry35Ab1 proteins represent an important seed trait technology having been developed as insect resistance traits in commercialized corn hybrids for control of WCR. The structures of Cry34Ab1 and Cry35Ab1 have been elucidated to 2.15 Å and 1.80 Å resolution, respectively. The solution structures of the toxins were further studied by small angle X-ray scattering and native electrospray ion mobility mass spectrometry. We present here the first published structure from the aegerolysin protein domain family and the structural comparisons of Cry34Ab1 and Cry35Ab1 with other pore forming toxins.
The duplicated and the highly repetitive nature of the maize genome has historically impeded the development of true single nucleotide polymorphism (SNP) markers in this crop. Recent advances in genome complexity reduction methods coupled with sequencing-by-synthesis technologies permit the implementation of efficient genome-wide SNP discovery in maize. In this study, we have applied Complexity Reduction of Polymorphic Sequences technology (Keygene N.V., Wageningen, The Netherlands) for the identification of informative SNPs between two genetically distinct maize inbred lines of North and South American origins. This approach resulted in the discovery of 1,123 putative SNPs representing low and single copy loci. In silico and experimental (Illumina GoldenGate (GG) assay) validation of putative SNPs resulted in mapping of 604 markers, out of which 188 SNPs represented 43 haplotype blocks distributed across all ten chromosomes. We have determined and clearly stated a specific combination of stringent criteria (> 0.3 minor allele frequency, > 0.8 GenTrainScore and > 0.5 Chi_test100 score) necessary for the identification of highly polymorphic and genetically stable SNP markers. Due to these criteria, we identified a subset of 120 high-quality SNP markers to leverage in GG assay-based marker-assisted selection projects. A total of 32 high-quality SNPs represented 21 haplotypes out of 43 identified in this study. The information on the selection criteria of highly polymorphic SNPs in a complex genome such as maize and the public availability of these SNP assays will be of great value for the maize molecular genetics and breeding community.
Computational analyses of protein structure-function relationships have traditionally been based on sequence homology, fold family analysis and 3D motifs/templates. Previous structurebased approaches characterize and compare active sites based on global shape and electrostatic properties. But, these methodologies are unable to capture similarities between diverse active sites that span multiple fold families despite catalyzing the same reaction (convergent evolution). In this work, we extend previous feature-based analyses of active sites by defining a system of localized geometric and electrostatic descriptors that identify localized patterns of protein-ligand interactions. Singular Value Decomposition is used to identify linear combinations of features with maximum information content which are then used to compute the class conditional probability density distribution of active sites using kernel density estimation. We successfully tested our algorithm on a database that contained examples of adenine, citrate, nicotinamide, phosphate, pyridoxal and ribose binding proteins with over 75% accuracy.
The enrichment and recall of known inhibitors in a virtual screen are correlated with the probability of finding effective inhibitors through this process. In practice, a large number of false positives are ranked higher than known inhibitors in many virtual screen results. In this paper, we use the interaction of known inhibitors across a range of decoy active sites in order to formulate a modified ranking score, Rscore. This ranking scheme seeks to normalize the DOCK score of a compound based on its interaction with decoy active sites, and uses a linear programming formulation to optimize Rscore for inhibitors versus non-inhibitors. We show an increase in recall of known inhibitors by greater than 20% in most of the test cases considered.
TEXTAL is a computer program that automatically-interprets electron density maps to determine the atomic structures of proteins through X-ray crystallography. Electron density maps are traditionally interpreted by visually fitting atoms into density Patterns. This manual process can be time-consuming and error prone, even for expert crystallographers, Noise in the data and limited resolution make map interpretation challenging. To automate the process, TEXTAL employs a variety of AI and Pattern-recognition techniques that emulate the decision-making processes of domain experts. In this article, we discuss the various ways AI technology is used in TEXTAL, including neural networks, case-based reasoning, nearest neighbor learning and linear discriminant analysis. The AI and pattern-recognition approaches have proven to be effective for building protein models even with medium resolution data. TEXTAL is a successfully deployed application; it is being used in more than 100 crystallography labs from 20 countries.
Non-crystallographic symmetry (NCS) averaging is a well known method for improving the quality of an electron-density map and thus aiding structure determination. Prior methods of NCS-operator determination based on estimated heavy-atom positions are prone to errors arising from inaccuracies in these coordinates or differences in the relative orientations of domains between molecules. In this paper, two real-space methods to determine NCS relationships from initial electron-density maps are presented. A brute-force method identifies matching regions in a map by local density correlation. A feature-based algorithm uses rotation-invariant features to reduce the computational time taken by the brute-force algorithm by filtering out regions that are likely to have dissimilar density patterns. This makes the feature-based algorithm faster and as accurate as the brute-force approach. Neither method requires the positions of heavy atoms or any information regarding the protein sequence. Both methods have been tested on a diverse range of experimentally phased maps and the correct NCS relationships were accurately identified for almost all of the test cases. The NCS operators obtained by the feature-based algorithm were used to perform NCS averaging and an improvement in map correlation was observed for some cases.
UNLABELLED X-ray crystallography is the most widely used method to determine the 3D structure of protein molecules. One of the most difficult steps in protein crystallography is model-building, which consists of constructing a backbone and then amino acid side chains into an electron density map. Interpretation of electron density maps represents a major bottleneck in protein structure determination pipelines, and thus, automated techniques to interpret maps can greatly improve the throughput. We have developed WebTex, a simple and yet powerful web interface to TEXTAL, a program that automates this process of fitting atoms into electron density maps. TEXTAL can also be downloaded for local installation. AVAILABILITY Web interface, downloadable binaries and documentation at http://textal.tamu.edu
TEXTAL is a successfully deployed system for automated model-building in protein X-ray crystallography. It represents a novel solution to an important, complex real-world, problem using various AI and pattern recognition algorithms. TEXTAL takes a model-building approach based on real-space density pattern recognition, similar to how a human crystallographer would work. TEXTAL first tries to predict the coordinates of the alpha-carbon (C/spl alpha/) atoms in the protein's connected backbone using a neural network. It then analyzes the density patterns around each C/spl alpha/ atom and searches a database of previously solved structures for regions with similar patterns. TEXTAL determines the best match, retrieves the coordinates for that region, and fits them to the unknown density. TEXTAL concatenates these local models into a global model and subjects them to various subsequent refinements to produce a complete protein model automatically.
This paper reports on TEXTAL™, a deployed application that uses a variety of AI techniques to automate the process of determining the 3D structure of proteins by x-ray crystallography. The TEXTAL™ project was initiated in 1998, and the application is currently deployed in three ways: (1) a web-based interface called WebTex, operational since June 2002; (2) as the automated model-building component of an integrated crystallography software called PHENIX, first released in July 2003; (3) binary distributions, available since September 2004. TEXTAL™ and its sub-components are currently being used by crystallographers around the world, both in the industry and in academia. TEXTAL™ saves up to weeks of effort typically required to determine the structure of one protein; the system has proven to be particularly helpful when the quality of the data is poor, which is very often the case. Automated protein modeling systems like TEXTAL™ are critical to the structural genomics initiative, a worldwide effort to determine the 3D structure of all proteins in a high-throughput mode, thereby keeping up with the rapid growth of genomic sequence databases.
Thiol protease Cathepsins play a significant role in proteolysis especially in certain desease states in humans.Mode of inhibition by different protein inhibitors in not clear yet.Keeping this in mind modeling studies were initiated with previously known and rcently identified inhibitors of stefin family with intact and truncated nterminal inhibitor.We find that role of N-terminal wedge as well as the second hairpin loop is very important.Inhibition mechanism also differs for endo and exo peptidase.
X-ray crystallography is the most widely used method for determining the three-dimensional structures of proteins and other macromolecules. One of the most difficult steps in crystallography is interpreting the 3D image of the electron density cloud surrounding the protein. This is often done manually by crystallographers and is very time-consuming and error-prone. The difficulties stem from the fact that the domain knowledge required for interpreting electron density data is uncertain. Thus crystallographers often have to resort to intuitions and heuristics for decision-making. The problem is compounded by the fact that in most cases, data available is noisy and blurred. TEXTAL ™ is a system designed to automate this challenging process of inferring the atomic structure of proteins from electron density data. It uses a variety of AI and pattern recognition techniques to try to capture and mimic the intuitive decision-making processes of experts in solving protein structures. The system has been quite successful in determining various protein structures, even with average quality data. The initial structure built by TEXTAL ™ can be used for subsequent manual refinement by a crystallographer, and combined with post-processing routines to generate a more complete model.