Automated cluster analysis is used to examine the conformation and configuration of pyranose sugars. Previous findings on this issue are confirmed, importantly from an analysis that requires no prior knowledge of the significant factors determining the conformational classification. The findings on the conformations adopted in the crystalline solid state are found to be different to existing quantum chemical calculations performed for D-glucose in the gas phase, but consistent with empirically determined conformations in the solution state. The use of this clustering analysis in studying chirality in the determined structures is discussed, as is the ability of this type of method to examine higher dimensions within the metric multi-dimensional scaling formalism.
The use of clustering to classify powder X-ray diffraction patterns is presented. The associated visualization techniques are also presented. The methods provide an easy way to handle large volumes of data.
A method of validating organic and organometallic crystal structures solved using powder diffraction data is presented. It uses searches and comparisons with structures in the Cambridge Structural Database coupled with multivariate analysis and Clustering techniques. An example using sulphonamides is presented. The associated software (dSNAP) is available as a free download.
The η’-precipitate is the main strengthening agent in age-hardening Al-Mg-Zn alloys. A new structure model of the phase has been derived by a combination of electron microscopy techniques and synchrotron X-ray diffraction. Samples for intensity measurements were prepared from an alloy casting with large matrix grains alloyed with 88.1 Al, 10.2 Zn, and 1.68 Mg. wt per cent, after aging by a standard procedure. Electron microscope images reveal a distribution of precipitates 3-10 nm wide, with varying degree of stacking disorder, in four different lattice orientations, embedded in the aluminiummatrix. Electron diffraction patterns confirmed the hexagonal lattice, a = 0.496 nm, c = 1.405 nm reported earlier. Intensities were measured by the precession technique [1], which suppressed multiple scattering viamatrix reflections, but incurred a high background of diffuse scattering. Systematic absences are consistent with three hexagonal, P63mc (186), P62c (190) and P63/mmc (194) or two trigonal space groups, P31c (159) and P31c (163). Due to the extensive overlap and high background level, we could not distinguish between amodified version in (190) of the earlier model [2] or a trigonal structure in (163) from electron diffraction data. Three-dimensional synchrotron data were collected at the Swiss Norwegian Beamlines at ESRF from two single aluminium grains. More than 2000 intensities were extracted from each of the four orientations of these precipitate. The preliminary analysis points towards an average structure best described in the centrosymmetric space group P31c. A trigonal model adopted for an average η’-structure can be described as a faulted stacking of units of the Laves phase stable MgZn2, in a way that retains the trigonal stacking as in the parent FCC aluminium lattice.
The ''Next Generation of small molecule Crystallographic Software'' is a joint project between the University of Oxford and Durham University, with an external management advisory group and scientific advisory committee of experienced senior crystallographers from many Universities.The consortium was set up to react to the challenges facing current crystallographic software and the knowledge contained within them.To date, almost all widely used small molecule programs are written by single individuals or small groups to software standards dating from the 1970's.The algorithm details are not documented and this software is neither extendable nor supportable.There is a risk of loss of technical knowledge as authors retire or leave the field, hence the pressing need to bridge the gap between previous and future generations of crystallographers.The aim of the project is to develop Open Source software solutions for Crystallography, which we consider to be a natural way for a project to evolve, since we adhere to software culture that encourages code-sharing.This will ensure the emergence of new science and will complement existing macromolecular crystallographic developments within the domain of small molecule crystallography.To this end, we will implement a pilot design and develop a new program to modern software standards, unhindered by legacy code, supported by detailed documentation, and which will guarantee maintenance and sustainability.It will provide a software kernel on which other researches can build new applications.The initial phase of the project will include all the fields covered by mainstream crystallographic software, such as CRYSTALS, GSAS, JANA, PLATON, SHELX, TOPAS, XP, and include data quality analysis, model building, electron density maps and related analysis, refinement, structure analyses and validation.The presentation will explain why it is important to develop such a project for crystallography now, and what the key challenges will be in completing it.The development team recognises that while current crystallographers are generally content with existing software, the next generation will requite something altogether more integrated and sophisticated, and their aim is to anticipate these needs.After introducing our conceptual view of the project, we will show how we intend to develop methods and algorithms within the new software architecture and at the same time benefit from existing programs.Other parties will be strongly encouraged to contribute their knowledge and ideas to the main development code base.
Extended abstract of a paper presented at Microscopy and Microanalysis 2005 in Honolulu, Hawaii, USA, July 31--August 4, 2005
A computer program that automatically classifies and clusters structural fragments extracted from mining the Cambridge Structural Database is described. The methodology is based on cluster analysis and multivariate data processing of distance matrix information describing the extracted fragments. Coupled with the calculations is a set of visualization tools that enable the user to view and verify the proposed classification scheme, and further explore it in varying levels of detail. Two examples are presented: the first is based on a simple difluoroalkene fragment and the second, more complex, on a chiral vicinal dialcohol, R-1(OH) CHCH( OH) R-2.
In high-throughput crystallography, it is possible to accumulate over 1000 powder diffraction patterns on a series of related compounds, often polymorphs. A method is presented that can analyse such data, automatically sort the patterns into related clusters or classes, characterize each cluster and identify any unusual samples containing, for example, unknown or unexpected polymorphs. Mixtures may be analysed quantitatively if a database of pure phases is available. A key component of the method is a set of visualization tools based on dendrograms, cluster analysis, pie charts, principal-component-based score plots and metric multidimensional scaling. Applications to pharmaceutical data and inorganic compounds are presented. The procedures have been incorporated into the PolySNAP commercial computer software.
In high-throughput crystallography experiments, it is possible to measure over 1000 powder diffraction patterns on a series of related compounds, often polymorphs or salts, in less than one week. The analysis of these patterns poses a difficult statistical problem. A computer program is presented that can analyse such data, automatically sort the patterns into related clusters or classes, characterize each cluster and identify any unusual samples containing, for example, unknown or unexpected polymorphs. Mixtures may be analysed quantitatively if a database of pure phases is available. A key component of the method is a set of visualization tools based on dendrograms and pie charts, as well as principal-component analysis and metric multidimensional scaling as a source of three-dimensional score plots. The procedures have been incorporated into the computer program PolySNAP, which is available commercially from Bruker-AXS.
A new integrated approach to full powder diffraction pattern analysis is described. This new approach incorporates wavelet-based data pre-processing, non-parametric statistical tests for full-pattern matching, and singular value decomposition to extract quantitative phase information from mixtures. Every measured data point is used in both qualitative and quantitative analyses. The success of this new integrated approach is demonstrated through examples using several test data sets. The methods are incorporated within the commercial software program SNAP-1D, and can be extended to high-throughput powder diffraction experiments.
In two previous papers [Gilmore, Barr & Paisley (2004). J. Appl. Cryst. 37, 231–242; Barr, Dong & Gilmore (2004). J. Appl. Cryst. 37, 243–252], it was demonstrated how to generate a correlation matrix by comparing full powder diffraction patterns, and then partition the diffractograms into groups using multivariate statistics and associated classification procedures. For clustering the patterns into related sets, dendrograms, metric multidimensional scaling and three-dimensional principal-components analysis score plots are employed. However, sometimes cluster membership for certain patterns is not always very clear or other ambiguities may arise; this paper describes cluster validation techniques using silhouettes and fuzzy clustering. The two methods operate in a complementary way: in some cases silhouettes are the most useful, and in others fuzzy clustering is more applicable. These procedures are available as options in the commercial computer program PolySNAP.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Investigating the structure of a membrane protein using electron diffraction can be very difficult, particularly because of the potential lack of accuracy and the limited amount of data collected because of limitations due to tilting.When these issues become a problem a method which can still be used for solving or enhancing the potential density map is the technique of maximum entropy, an implementation of which is included in the program MICE [1].The advantage of this method over other similar ones is its ability to incorporate prior information, such as reflection phases from electron microscopy, the molecular envelope, any known partial structure and to employ entropy optimization techniques.Incorporation of a molecular envelope into the maximum entropy calculation can be vital to impose values on the unit cell where it is most likely to find new places of density.The envelope can easily be produced using the CCP4 program suite [2], however only if a small fragment of the structure is known.To give a starting point for maximum entropy a basis set of known reflection phases is required.These are normally gained through the permutation of the observed reflections, but here they can be found either directly from electron microscopy or the most accurate phases from the Fourier Transform of the partial structure can be used.Using this information maximum entropy can be shown to be a very powerful tool even at relatively low resolutions, enhancing partially known structures and even calculating unmeasured reflections from the missing cone.
The tools of modern direct methods are examined and their limitations for solving protein structures discussed. Direct methods need atomic resolution data (1.1-1.2 A) for structures of around 1000 atoms if no heavy atom is present. For low-resolution data, alternative approaches are necessary and these include maximum entropy, symbolic addition, Sayre's equation, group scattering factors and electron microscopy.
Least-squares refinement is unusual in the context of electron crystallography because of the sparsity of the measured intensity data set and the problems of systematic errors due to multiple dynamical scattering. With 120 unique hkl electron diffraction intensities measured from polymorphic form III of isotactic poly(1-butene), conditions for improving an existing structural model derived from initial direct structure analysis have been evaluated. The polymer crystallizes in space group P2(1)2(1)2(1) with a = 12.38, b = 8.88, c = 7.56 A and there are 8 unique atoms in the asymmetric unit. Starting with atomic positions resulting from Fourier refinement, four cycles of least-squares refinement, where the positional shifts of atomic positions were constrained, produced better bonding parameters than found before while lowering the conventional crystallographic residual, based on absolute value(F), from an overall value of R = 0.26 to R = 0.185 for the 58 most intense reflections where magnitude of absolute value(Fh(obs)) > or = 4sigma (Fh(obs)) or 0.216 for the complete data set of 120 reflections. The weighted residuals based on magnitude of absolute value(F)2 fell from 0.50 to 0.41 for the complete data set. This refinement was not improved however when attempts were made to fill in very weak intensities by default values. Also, effects of multiple-scattering perturbations were found in the irregularity of the final isotropic thermal parameters.