In previous articles in this series, a novel probabilistic method was described which is capable of estimating triplet invariants using the Patterson map as prior information. The first experimental tests demonstrated the superiority of the new method compared with the traditional Cochran estimate. The advantages were so significant that the ab initio solution of macromolecular structures was considered to be feasible even when the data resolution is worse than 2 Å. However, several questions remained unanswered. For example: (i) which and how many Patterson peaks should be used to optimize a direct-methods phasing procedure applied to experimental data up to 2.2 Å resolution?, (ii) is the presence of heavy atoms a necessary ingredient for the validity of the method?, (iii) which and how many reflections must be used in the triplet search?, (iv) is the information contained in the Patterson map able to identify negative cosine triplets? and (v) may a computer program be made that routinary solves macromolecular structures with data resolution up to 2.2 Å? This article recalls these five unresolved questions and answers them. In particular, criteria have been defined to determine both the number of Patterson peaks to be actively used for triplet estimation and the number of reflections to be used in the triplet search. It has also been shown that the presence of heavy atoms is a necessary ingredient for success of the theory. In particular, the theory is unable to accurately identify triplet invariants with a negative cosine, but rather can identify enantiomorph-sensitive triplets. A paradox of the theory is discussed and resolved. Finally, a computer program is presented that is capable of automatically, with a few directives, solving some of the test structures at non-atomic resolution (proteins and nucleic acids) with data resolution up to 2.2 Å, but not in a straightforward way. The limitations of the computer program and its prospects are discussed.
In the previous articles in this series, it was shown how the use of the Patterson map as a priori information significantly improves the estimation of triplet invariants and makes possible the ab initio solution of macromolecular structures with data resolutions up to 2.2 Å, provided that heavy atoms are present in the unit cell. For the above estimation, the position and intensity of the most representative Patterson peaks were used. No effort towards Patterson deconvolution is necessary. Triplet invariants, however, depend not only on individual Patterson peaks, but also on pairs of interatomic vectors that start or converge on the same atom. Such pairs presuppose the presence of a third interatomic vector, and therefore in this article we legitimately speak of Patterson peak triples. The new mathematical formulation allows their algebra to be described: the informative contribution coming from them has been incorporated into the technique of joint probability distribution functions, and the result is a concentration factor that significantly improves the estimation of triplet invariants. Experimental tests on proteins and nucleic acids show that structural complexity and non-atomic resolution are no longer an unsurpassable obstacle for crystal structure solution.
A new theory for the probabilistic estimation of first-rank one-phase semi-invariants is presented. In this approach, atomic positions are treated as primitive random variables but are constrained by the a priori knowledge of interatomic vectors. This information is always available, thus allowing the new technique to be considered an ab initio probabilistic method conditioned by the knowledge of the Patterson map. The theoretical foundation for the estimation of triplet invariants was outlined in the first paper of this series [Giacovazzo (2019). Acta Cryst. A 75 , 142–157]. Subsequent experimental tests, shown in the second paper of this series [Burla et al. (2024). J. Appl. Cryst. 57 , 1011–1022], have demonstrated the significant superiority of this new approach over existing methods. The improvements were so notable that it has been suggested this technique could be valuable for the ab initio solution of macromolecular structures. This work expands the probabilistic approach to include the estimation of first-rank one-phase semi-invariants, The hope is that they can contribute to the ab initio solution of macromolecular structures. Only in this way can one-phase semi-invariants go from being a historical curiosity to an effective tool for solving macromolecular structures.
Quartet invariants play a minor role in modern direct methods. In practice, only the quartets whose cosine is estimated to be negative are used, as they have no correlation with the triplet invariants. However, their role remains marginal: in fact, the quartet relations are of order 1/N while the triplet relations are of order 1/√N. The reliability of the quartets is therefore relatively low, in particular for the quartets estimated to be negative. Two papers have recently appeared (Papers I and II of this series) that describe procedures able to exploit the information contained in the Patterson map to estimate the triplet phases. The improvements in estimates are notable, apparently capable of resolving macromolecular structures even at non-atomic resolution. It therefore seems useful to develop a theory of quartet invariants that is able to exploit the Patterson information. This is the main purpose of this article. The method of joint probability distribution functions is used to obtain a von Mises-type distribution which associates a probability with each quartet phase. It is expected that the Patterson map, used as a priori information, can significantly increase the reliabilities of quartet invariants, particularly those whose cosine is estimated to be negative. The quartets may thus be able to play a more prominent role in future.
Direct methods have practically solved the phase problem for small–medium-size molecules but have substantially failed in macromolecular crystallography. They have two main limitations: a strong dependence on structural complexity and the need to work with atomic-resolution data. Many attempts have been made to broaden their field of applicability, for example the use of some a priori information to make the estimate of the triplet invariant phases more effective. Unfortunately none of these new approaches allowed the successful application of direct methods to proteins and nucleic acids. Direct methods are still a niche tool in macromolecular crystallography. In a recent publication [Giacovazzo (2019). Acta Cryst. A 75 , 142–157] the method of joint probability distributions has been modified to take into account new sources of prior information, one of which is relevant to this article: the Patterson map. In practice, it has been shown that with prior knowledge of the interatomic vectors one is able to modify the classic Cochran reliability parameter for estimating the triplet invariant phases. The article was essentially theoretical in nature, and no attempt was described to test the practical usefulness of the new probabilistic formulas. This work is therefore the first application of the new method. It is shown that the use of the Patterson map as prior information substantially improves the Cochran estimate of triplet phases; the phase error distribution for the new estimates, even if it is related to macromolecular structures, becomes similar to that obtained for medium-size structures. In some ways, it is as if the use of the Patterson information reduces the structural complexity, thus allowing a more general use of direct methods in macromolecular crystallography. Atomic resolution no longer seems to be a necessary ingredient for the applicability of direct methods; tests show that the apparent reduction in structural complexity also occurs in macromolecular structures with experimental data having a resolution of 2.3 Å. A number of test structures have been used to show the potential of the new technique.
A description of REMO22, a new molecular replacement program for proteins and nucleic acids, is provided. This program, as with REMO09, can use various types of prior information through appropriate conditional distribution functions. Its efficacy in model searching has been validated through several test cases involving proteins and nucleic acids. Although REMO22 can be configured with different protocols according to user directives, it has been developed primarily as an automated tool for determining the crystal structures of macromolecules. To evaluate REMO22’s utility in the current crystallographic environment, its experimental results must be compared favorably with those of the most widely used Molecular Replacement (MR) programs. To accomplish this, we chose two leading tools in the field, PHASER and MOLREP. REMO22, along with MOLREP and PHASER, were included in pipelines that contain two additional steps: phase refinement (SYNERGY) and automated model building (CAB). To evaluate the effectiveness of REMO22, SYNERGY and CAB, we conducted experimental tests on numerous macromolecular structures. The results indicate that REMO22, along with its pipeline REMO22 + SYNERGY + CAB, presents a viable alternative to currently used phasing tools.
Patterson superposition techniques are a historical method for solving the structures of small molecules ab initio, provided they contain heavy atoms in the unit cell. In the 1990s, they were combined with effective EDM procedures and succeeded in the crystal structure solution of macromolecular structures with resolution data up to 1.6–1.9 Å. In this paper we enlarge the concept of Patterson superposition by replacing it with the vector superposition concept. We show, indeed, that besides Patterson other Fourier syntheses may also be used for the superposition of the interatomic vectors. Five Fourier syntheses are described and used in the practical applications. We show that even macromolecular structures with 2.2 Å data resolution may be solved via the new approach.
CAB, a recently described automated model-building (AMB) program, has been modified to work effectively with nucleic acids. To this end, several new algorithms have been introduced and the libraries have been updated. To reduce the input average phase error, ligand heavy atoms are now located before starting the CAB interpretation of the electron-density maps. Furthermore, alternative approaches are used depending on whether the ligands belong to the target or to the model chain used in the molecular-replacement step. Robust criteria are then applied to decide whether the AMB model is acceptable or whether it must be modified to fit prior information on the target structure. In the latter case, the model chains are rearranged to fit prior information on the target chains. Here, the performance of the new AMB program CAB applied to various nucleic acid structures is discussed. Other well documented programs such as Nautilus, ARP/wARP and phenix.autobuild were also applied and the experimental results are described.
In this study, the properties of observed, difference, and hybrid syntheses (hybrid indicates a combination of observed and difference syntheses) are investigated from two points of view. The first has a statistical nature and aims to estimate the amplitudes of peaks corresponding to the model atoms, belonging or not belonging to the target structure; the amplitudes of peaks related to the target atoms, missed or shared with the model; and finally, the quality of the background. The latter point deals with the practical features of Fourier syntheses, the special role of weighted syntheses, and their usefulness in practical applications. It is shown how the properties of the various syntheses may vary according to the available structural model and, in particular, how weighted hybrid syntheses may act like an observed and difference or a full hybrid synthesis. The theoretical results obtained in this paper suggest new Fourier syntheses using novel Fourier coefficients: their main features are first discussed from a mathematical point of view. Extended experimental applications show that they meet the basic mission of the Fourier syntheses, enhancing peaks corresponding to the missed target atoms, depleting peaks corresponding to the model atoms not belonging to the target, and significantly reducing the background. A comparison with the results obtained via the most popular modern Fourier syntheses is made, suggesting a role for the new syntheses in modern procedures for phase extension and refinement. The most promising new Fourier synthesis has been implemented in the current version of SIR2014.
Obtaining high-quality models for nucleic acid structures by automated model building programs (AMB) is still a challenge. The main reasons are the rather low resolution of the diffraction data and the large number of rotatable bonds in the main chains. The application of the most popular and documented AMB programs (e.g., PHENIX.AUTOBUILD, NAUTILUS and ARP/wARP) may provide a good assessment of the state of the art. Quite recently, a cyclic automated model building (CAB) package was described; it is a new AMB approach that makes the use of BUCCANEER for protein model building cyclic without modifying its basic algorithms. The applications showed that CAB improves the efficiency of BUCCANEER. The success suggested an extension of CAB to nucleic acids-in particular, to check if cyclically including NAUTILUS in CAB may improve its effectiveness. To accomplish this task, CAB algorithms designed for protein model building were modified to adapt them to the nucleic acid crystallochemistry. CAB was tested using 29 nucleic acids (DNA and RNA fragments). The phase estimates obtained via molecular replacement (MR) techniques were automatically submitted to phase refinement and then used as input for CAB. The experimental results from CAB were compared with those obtained by NAUTILUS, ARP/wARP and PHENIX.AUTOBUILD.
Although the success of molecular-replacement techniques requires the solution of a six-dimensional problem, this is often subdivided into two three-dimensional problems. REMO09 is one of the programs which have adopted this approach. It has been revisited in the light of a new probabilistic approach which is able to directly derive conditional distribution functions without passing through a previous calculation of the joint probability distributions. The conditional distributions take into account various types of prior information: in the rotation step the prior information may concern a non-oriented model molecule alone or together with one or more located model molecules. The formulae thus obtained are used to derive figures of merit for recognizing the correct orientation in the rotation step and the correct location in the translation step. The phases obtained by this new version of REMO09 are used as a starting point for a pipeline which in its first step extends and refines the molecular-replacement phases, and in its second step creates the final electron-density map which is automatically interpreted by CAB, an automatic model-building program for proteins and DNA/RNA structures.
The standard method of joint probability distribution functions, so crucial for the development of direct methods, has been revisited and updated. It consists of three steps: identification of the reflections which may contribute to the estimation of a given structure invariant or seminvariant, calculation of the corresponding joint probability distribution, and derivation of the conditional distribution of the invariant or seminvariant phase given the values of some diffracted amplitudes. In this article the conditional distributions are derived directly without passing through the second step. A good feature of direct methods is that they may work in the absence of any prior information: that is also their weakness. Different types of prior information have been taken into consideration: interatomic distances, interatomic vectors, Patterson peaks, structural model. The method of directly deriving the conditional distributions has been applied to those cases. Some new formulas have been obtained estimating two-, three- and four-phase invariants. Special attention has been dedicated to the practical aspects of the new formulas, in order to simplify their possible use in direct phasing procedures.
The program Buccaneer, a well known fast and efficient automatic model-building program, is also a tool for phase refinement: indeed, input phases are used to calculate electron-density maps that are interpreted in terms of a molecular model, from which new phase estimates may be obtained. This specific property is shared by all other automatic model-building programs and allows their cyclic use, as is usually performed in other phase-refinement methods (for example electron-density modification techniques). Buccaneer has been included in a cyclic procedure, called CAB, aimed at increasing the rate of success of Buccaneer and the quality of the molecular models provided. CAB has been tested on 81 protein structures that were solved via molecular-replacement, anomalous dispersion and ab initio methods. The corresponding phases were submitted to a phase-refinement process that synergically combines current phase-refinement techniques and out-of-mainstream refinement methods [Burla et al. (2017), Acta Cryst. D73, 877-888]. The phases thus obtained were used as input for CAB. The experimental results were compared with those obtained by the sole use of Buccaneer: it is shown that CAB improves the Buccaneer results, both in completeness and in accuracy.
Crystallographic least-squares techniques, the main tool for crystal structure refinement of small and medium-size molecules, are for the first time used for ab initio phasing. It is shown that the chief obstacle to such use, the least-squares severe convergence limits, may be overcome by a multi-solution procedure able to progressively recognize and discard model atoms in false positions and to include in the current model new atoms sufficiently close to correct positions. The applications show that the least-squares procedure is able to solve many small structures without the use of important ancillary tools: e.g. no electron-density map is calculated as a support for the least-squares procedure.
The method of the joint probability distribution function was applied in order to estimate the normal structure factor amplitudes of the anomalous scatterer substructure in a FEL experiment. The two-wavelength case was examined. In this, the prior knowledge of the moduli vertical bar F-1(+)vertical bar, vertical bar F-1(-)vertical bar, vertical bar F-2(+)vertical bar, vertical bar F-2(-)vertical bar was used to predict the value of vertical bar F-oa vertical bar, which is the structure factor amplitude arising from the normal scattering of the heavy atom anomalous scatterers. The mathematical treatment provides a solid theoretical basis for the RIP (Radiation-damage Induced Phasing) method, which was originally proposed in order to take the radiation damage induced by synchrotron radiation sources into account. This was further adapted to exploit FEL data, where the crystal damage is usually more massive.
The method of the joint probability distribution function was applied in order to estimate the normal structure factor amplitudes of the anomalous scatterer substructure in a FEL experiment. The two-wavelength case was examined. In this, the prior knowledge of the moduli | F 1 + | , | F 1 − | , | F 2 + | , | F 2 − | was used to predict the value of | F 0 a | , which is the structure factor amplitude arising from the normal scattering of the heavy atom anomalous scatterers. The mathematical treatment provides a solid theoretical basis for the RIP (Radiation-damage Induced Phasing) method, which was originally proposed in order to take the radiation damage induced by synchrotron radiation sources into account. This was further adapted to exploit FEL data, where the crystal damage is usually more massive.
Difference electron densities do not play a central role in modern phase refinement approaches, essentially because of the explosive success of the EDM (electron-density modification) techniques, mainly based on observed electron-density syntheses. Difference densities however have been recently rediscovered in connection with the VLD (Vive la Difference) approach, because they are a strong support for strengthening EDM approaches and for ab initio crystal structure solution. In this paper the properties of the most documented difference electron densities, here denoted as F - Fp, mF - Fp and mF - DFp syntheses, are studied. In addition, a fourth new difference synthesis, here denoted as {\overline F_q} synthesis, is proposed. It comes from the study of the same joint probability distribution function from which the VLD approach arose. The properties of the {\overline F_q} syntheses are studied and compared with those of the other three syntheses. The results suggest that the {\overline F_q} difference may be a useful tool for making modern phase refinement procedures more efficient.
This study clarifies why, in the phantom derivative (PhD) approach, randomly created structures can help in refining phases obtained by other methods. For this purpose the joint probability distribution of target, model, ancil and phantom derivative structure factors and its conditional distributions have been studied. Since PhD may use n phantom derivatives, with n ≥ 1, a more general distribution taking into account all the ancil and derivative structure factors has been considered, from which the conditional distribution of the target phase has been derived. The corresponding conclusive formula contains two components. The first is the classical Srinivasan & Ramachandran term, relating the phases of the target structure with the model phases. The second arises from the combination of two correlations: that between model and derivative (the first is a component of the second) and that between derivative and target. The second component mathematically codifies the information on the target phase arising from model and derivative electron-density maps. The result is new, and explains why a random structure, uncorrelated with the target structure, adds useful information on the target phases, provided a model structure is known. Some experimental tests aimed at checking if the second component really provides information on ϕ (the target phase) were performed; the favourable results confirm the correctness of the theoretical calculations and of the corresponding analysis.
Ab initio and non-ab initio phasing methods are often unable to provide phases of sufficient quality to allow the molecular interpretation of the resulting electron-density maps. Phase extension and refinement is therefore a necessary step: its success or failure can make the difference between solution and nonsolution of the crystal structure. Today phase refinement is trusted to electron-density modification (EDM) techniques, and in practice to dual-space methods which try, via suitable constraints in direct and in reciprocal space, to generate higher quality electron-density maps. The most popular EDM approaches, denoted here as mainstream methods, are usually part of packages which assist crystallographers in all of the structure-solution steps from initial phasing to the point where the molecular model perfectly fits the known features of protein chemistry. Other phase-refinement approaches that are based on different sources of information, denoted here as out-of-mainstream methods, are not frequently employed. This paper aims to show that mainstream and out-of-mainstream methods may be combined and may lead to dramatic advances in the present state of the art. The statement is confirmed by experimental tests using molecular-replacement, SAD-MAD and ab initio techniques.
The efficient multipurpose figure of merit MPF has been defined and characterized. It may be very helpful in phasing procedures. Indeed, it might be used for establishing the centric or acentric nature of an unknown structure, for identifying the presence of some pseudotranslational symmetry, for recognizing the correct solution in multisolution approaches and for estimating the quality of structure models as they become available during the phasing process. Thus, phase improvement or deterioration may be monitored and useless models may be discarded to save computing time. It is also shown that MPF may be applied in different phasing approaches, no matter if ab initio or non ab initio.