Automated model-building software aims at the objective interpretation of crystallographic diffraction data by means of the construction or completion of macromolecular models. Automated methods have rapidly gained in popularity as they are easy to use and generate reproducible and consistent results. However, the process of model building has become increasingly hidden and the user is often left to decide on how to proceed further with little feedback on what has preceded the output of the built model. Here, ArpNavigator, a molecular viewer tightly integrated into the ARP/wARP automated model-building package, is presented that directly controls model building and displays the evolving output in real time in order to make the procedure transparent to the user.
The identification and modelling of ligands into macromolecular models is important for understanding molecule's function and for designing inhibitors to modulate its activities. We describe new algorithms for the automated building of ligands into electron density maps in crystal structure determination. Location of the ligand-binding site is achieved by matching numerical shape features describing the ligand to those of density clusters using a "fragmentation-tree" density representation. The ligand molecule is built using two distinct algorithms exploiting free atoms with inter-atomic connectivity and Metropolis-based optimisation of the conformational state of the ligand, producing an ensemble of structures from which the final model is derived. The method was validated on several thousand entries from the Protein Data Bank. In the majority of cases, the ligand-binding site could be correctly located and the ligand model built with a coordinate accuracy of better than 1 Å. We anticipate that the method will be of routine use to anyone modelling ligands, lead compounds or even compound fragments as part of protein functional analyses or drug design efforts.
Page s312 s312characteristic is shown by a point at its own ruler, and the rulers are plotted together as a set of lines with the same origin, forming a hub and spokes.The points for a given model marked on these lines are connected to form a polygon.A polygon strongly compressed or dilated along some axes reveals unusually low or high values of corresponding characteristics.Different parts of the rulers are colored differently to reflect the frequency (red color for a low frequency, blue for a high frequency) with which the corresponding values are observed in a reference set of structures determined previously.Polygon vertices in 'red zones' indicate parameters which lie outside typical values.The reference set of structures can be selected by the resolution, by their size of structures or by other characteristics.The list of model characteristics to be shown in the polygon is also variable.In particular, in addition to (or instead of) the average values of distortion of stereochemical parameters it may include their maximal values to indicate local problems if they exist.Both the stand-alone Tcl/tk version of the program for macromolecules and the python version incorporated into PHENIX [2] are available.As an extra control tool and independently of the POLYGON, the typical values for the R-and R free -factors and for their difference at a given resolution can be obtained as linear functions of the logarithm of the resolution [3].
The interpretation of a 20 Å resolution electron-density map using segmentation and pattern-recognition-based identification of domain shapes is described.
Traditionally crystallographic protein modelling and particularly model completion involves graphical display software for manual user interference.During the last decade, a number of software packages -such as ARP/wARP [1]emerged that do a large part of the crystallographic map interpretation and model building automatically without much manual work.This has tremendously shortened the time in which a macromolecular model could be built, from months and sometimes years to days or even a few hours.Having done a computer intensive map interpretation step there is generally no elaborate graphical communication or reporting framework, which would indicate how a model has progressed from a visually hardly interpretable density to an almost complete set of polypeptide fragments.As a result, the modelling needs to be viewed, judged and completed where needed elsewhere.At the same time, the recent past has shown that programs that allow semi-automatic modelling coupled with instantaneous graphical display -such as Coot [2] -are becoming very popular.To bridge the gap between the automation that the ARP/ wARP package offers and intuitive communication via graphical components, we have developed a graphical front end to ARP/wARP.It allows to run less time consuming tasks from the ARP/wARP suite, e.g.fitting ligands [3] or secondary structure [4].The user can choose the input from the graphical items and get results displayed immediately.The approach behind such a graphical tool is to not only provide in total more information than logfiles do on their own, but also at the same time to instantaneously allow the user to choose from minimal intervention to very informed guidance of the software.
The 'Buccaneer' software for automated protein model building provides an effective tool for building an initial protein model into an experimentally phased electron density map.While it works at higher resolutions, it has been particularly useful in the lower resolution range, e.g. from 2.5A to 3.5A, when good phases are available [1].The effectiveness of the procedure across a range of resolutions depends on the use of a problem-specific search function for locating likely alpha-carbon positions, where the search function is determined from a standard library structure and tailored to the resolution and data quality of the problem at hand.Some of the same methods have been implemented in the 'Coot' graphics software to allow interactive location of secondary structure features.More recently, the focus of this work has shifted to the rebuilding of molecular replacement models.Initial results of this work will be presented.The integration of 'Buccaneer' into automation pipelines has also involved the development of a new 'classical' density modification program, 'Parrot'.
ARP/wARP is a software suite to build macromolecular models in X-ray crystallography electron density maps. Structural genomics initiatives and the study of complex macromolecular assemblies and membrane proteins all rely on advanced methods for 3D structure determination. ARP/wARP meets these needs by providing the tools to obtain a macromolecular model automatically, with a reproducible computational procedure. ARP/wARP 7.0 tackles several tasks: iterative protein model building including a high-level decision-making control module; fast construction of the secondary structure of a protein; building flexible loops in alternate conformations; fully automated placement of ligands, including a choice of the best-fitting ligand from a 'cocktail'; and finding ordered water molecules. All protocols are easy to handle by a nonexpert user through a graphical user interface or a command line. The time required is typically a few minutes although iterative model building may take a few hours.
A new space group determination algorithm is described that follows the structure solution step rather than preceeding it.It is based on an analysis of the average phase differences between symmetry equivalent reflections following the solution of the structure in space group P1.It is shown that this method performs very well for cases that are troublesome for space group determination algorithms based on an analysis of systematic extinct reflections.Examples illustrating the method include the analysis of weak data sets, faulty data sets, centrosymmetric/non-centrosymmetric ambiguity, ambiguous cases without systematic extinct reflections, data sets missing critical reflections, powder data and data from incommensurate crystal structures.The algorithm can be used as well for the analysis of missed higher symmetry.
The efficiency of the ligand-building module of ARP/wARP version 6.1 has been assessed through extensive tests on a large variety of protein-ligand complexes from the PDB, as available from the Uppsala Electron Density Server. Ligand building in ARP/wARP involves two main steps: automatic identification of the location of the ligand and the actual construction of its atomic model. The first step is most successful for large ligands. The second step, ligand construction, is more powerful with X-ray data at high resolution and ligands of small to medium size. Both steps are successful for ligands with low to moderate atomic displacement parameters. The results highlight the strengths and weaknesses of both the method of ligand building and the large-scale validation procedure and help to identify means of further improvement.
A crucial point in enabling automated protein model building at low resolution is to automatically obtain a good starting model together with its associated restraints for refinement. Having such a model, additional parts of the molecule can then be constructed iteratively. As a starting model the recurring motifs of helices and strands are a good choice due to their pronounced features even at low resolution. For the treatment of crystallographic electron density maps at low resolution ranging down to 4 Å the version 6.1 of the ARP/wARP software suite [1] contains a dedicated module that will trace the respective portions of the protein main thus reducing the need for manual intervention. The key elements of the method for accomplishing this task are successive filtering steps based on discriminant analysis applied to geometric features of helical fragments of increasing size. The helix building software has been applied to more than a thousand test cases from the PDB and, on average, more than 60% of the helices could be correctly identified, more than 80% of which were predicted with the correct chain direction. Furthermore the results are nearly independent of the resolution of the data. A similar approach can be used for the location of strands and results of new developments in terms of a more complete capture of secondary structure will be presented. Finally we will show, how the models built with the module can be made use of in successive steps of protein model building.
A server to run remote ARP/wARP jobs for automated model building in macromolecular crystallography was set up at EMBL-Hamburg, based on a 16 processor Linux Cluster.The project aimed to provide a user-friendly service for macromolecular crystallographers, with the following intentions: 1) to allow users to go far beyond routine data acqusition.2) to obtain data for software development and project tracking 3) to provide users with the latest executables 4). to provide users with CPU power.The server was opened to the community in July 2004.This paper will describe the way remote job submission is carried out.An example will be given to lead the reader through the process.Statistics on the number of jobs run at the cluster, type and frequency of software and hardware errors, and a variety of other lessons learnt will be presented.Plans for scaling up and improvement of the services will also be described.
Group leader: Victor Lamzin Staff scientist: Andrea Schmidt Postdoctoral fellows: Olga Kirillova, Gerrit Langer*, Tilo Strutz* PhD student: Petrus Zwart* Technicians: Venkataraman Parthasarathy*, Babu Pothineni*, Katja Schirwitz Visitors: Serge Cohen, Zbigniew Dauter, Francisco Fernandez Perez*, Christian Jelsch, Mattheos Kakaris, Joergen Koepke, Olga V. Koroleva, Richard J. Morris, Peter Østergaard, Tatiana V. Pegasova, Anastassis Perrakis, Denis V. Rebrikov, Wojciech Rypniewski, Jozef Sevcic , Elena V. Stepanova, Clemens Vonrhein