The Macromolecular Structure Database (MSD) (http://www.ebi.ac.uk/msd/) [H. Boutselakis, D. Dimitropoulos, J. Fillon, A. Golovin, K. Henrick, A. Hussain, J. Ionides, M. John, P. A. Keller, E. Krissinel et al. (2003) E-MSD: the European Bioinformatics Institute Macromolecular Structure Database. Nucleic Acids Res., 31, 458-462.] group is one of the three partners in the worldwide Protein DataBank (wwPDB), the consortium entrusted with the collation, maintenance and distribution of the global repository of macromolecular structure data [H. Berman, K. Henrick and H. Nakamura (2003) Announcing the worldwide Protein Data Bank. Nature Struct. Biol., 10, 980.]. Since its inception, the MSD group has worked with partners around the world to improve the quality of PDB data, through a clean up programme that addresses inconsistencies and inaccuracies in the legacy archive. The improvements in data quality in the legacy archive have been achieved largely through the creation of a unified data archive, in the form of a relational database that stores all of the data in the wwPDB. The three partners are working towards improving the tools and methods for the deposition of new data by the community at large. The implementation of the MSD database, together with the parallel development of improved tools and methodologies for data harvesting, validation and archival, has lead to significant improvements in the quality of data that enters the archive. Through this and related projects in the NMR and EM realms the MSD continues to improve the quality of publicly available structural data.
The aMAZE database (http://www.amaze.ulb.ac.be) manages information on the molecular functions of genes and proteins, their interactions and the biochemical processes in which they participate. Its data model embodies general rules for associating molecules and interactions into large complex networks that can be analysed using graph theory methods. The processes represented include metabolic pathways, protein-protein interactions, gene regulation, transport and signal transduction. These processes are mapped into their spatial localisation. A distinct feature of aMAZE is its Object-Oriented, modular and open user interface. Queries are invoked through dedicated modules, data can be linked to external sources, interactively browsed and transferred between modules, and new modules can be readily added. Available modules also include, a custom-built Diagram Editor for the automatic layout, display, and interactive modification of pathway diagrams, and procedures for analysing network graphs. http://www.beilstein-institut.de/bozen2002/proceedings/Wodak/Wodak.pdf
This paper describes how biological function can be represented in terms of molecular activities and processes. It presents several key features of a data model that is based on a conceptual description of the network of interactions between molecular entities within the cell and between cells. This model is implemented in the aMAZE database that presently deals with information on metabolic pathways, gene regulation, sub- or supracellular locations, and transport. It is shown that this model constitutes a useful generalisation of data representations currently implemented in metabolic pathway databases, and that it can furthermore include multiple schemes for categorising and classifying molecular entities, activities, processes and localisations. In particular, we highlight the flexibility offered by our system in representing multiple molecular activities and their control, in viewing biological function at different levels of resolution and in updating this view as our knowledge evolves.
Determining the biological function of a myriad of genes, and understanding how they interact to yield a living cell, is the major challenge of the post genomesequencing era. The complexity of biological systems is such that this cannot be envisaged without the help of powerful computer systems capable of representing and analysing the intricate networks of physical and functional interactions between the different cellular components. In this review we try to provide the reader with an appreciation of where we stand in this regard. We discuss some of the inherent problems in describing the different facets of biological function, give an overview of how information on function is currently represented in the major biological databases, and describe different systems for organising and categorising the functions of gene products. In a second part, we present a new general data model, currently under development, which describes information on molecular function and cellular processes in a rigourous manner. The model is capable of representing a large variety of biochemical processes, including metabolic pathways, regulation of gene expression and signal transduction. It also incorporates taxonomies for categorising molecular entities, interactions and processes, and it offers means of viewing the information at different levels of resolution, and dealing with incomplete knowledge. The data model has been implemented in the database on protein function and cellular processes ‘aMAZE’ (http://www.ebi.ac.uk/research/pfbp/), which presently covers metabolic pathways and their regulation. Several tools for querying, displaying, and performing analyses on such pathways are briefly described in order to illustrate the practical applications enabled by the model.
We have constructed a morphologically divided redshift distribution of faint field galaxies using a statistically unbiased sample of 196 galaxies brighter than I = 21.5 for which detailed morphological information (from the Hubble Space Telescope) as well as ground-based spectroscopic redshifts are available. Galaxies are classified into 3 rough morphological types according to their visual appearance (E/S0s, Spirals, Sdm/dE/Irr/Pec's), and redshift distributions are constructed for each type. The most striking feature is the abundance of low to moderate redshift Sdm/dE/Irr/Pec's at I < 19.5. This confirms that the faint end slope of the luminosity function (LF) is steep (alpha < -1.4) for these objects. We also find that Sdm/dE/Irr/Pec's are fairly abundant at moderate redshifts, and this can be explained by strong luminosity evolution. However, the normalization factor (or the number density) of the LF of Sdm/dE/Irr/Pec's is not much higher than that of the local LF of Sdm/dE/Irr/Pec's. Furthermore, as we go to fainter magnitudes, the abundance of moderate to high redshift Irr/Pec's increases considerably. This cannot be explained by strong luminosity evolution of the dwarf galaxy populations alone: these Irr/Pec's are probably the progenitors of present day ellipticals and spiral galaxies which are undergoing rapid star formation or merging with their neighbors. On the other hand, the redshift distributions of E/S0s and spirals are fairly consistent those expected from passive luminosity evolution, and are only in slight disagreement with the non-evolving model.
The problem of automated separation of stars and galaxies on photographic plates is revisited with two goals in mind : First, to separate galaxies from everything else (as opposed to most previous work, in which galaxies were lumped together with all other non-stellar images). And second, to search optically for galaxies at low Galactic latitudes (an area that has been largely avoided in the past). This paper demonstrates how an artificial neural network can be trained to achieve both goals on Schmidt plates of the Digitised Sky Survey. Here I present the method while its application to large numbers of plates is deferred to a later paper. Analysis is also provided of the way in which the network operates and the results are used to counter claims that it is a complicated and incomprehensible tool.
The applicability of the highly idealized secondary infall model to ''realistic'' initial conditions is investigated. The collapse of protohalos seeded by 3 sigma density perturbations to an Einstein-de Sitter universe is studied here for a variety of scale-free power spectra with spectral indices ranging from n = 1 to -2. Initial conditions are set by the constrained realization algorithm and the dynamical evolution is calculated both analytically and numerically. The analytic calculation is based on the simple secondary infall model where spherical symmetry is assumed. A full numerical simulation is performed by a Tree N-body code where no symmetry is assumed. A hybrid calculation has been performed by using a monopole term code, where no symmetry is imposed on the particles but the force is approximated by the monopole term only. The main purpose of using such code is to suppress off-center mergers. In all cases studied here the rotation curves calculated by the two numerical codes are in agreement over most of the mass of the halos, excluding the very inner region, and these are compared with the analytically calculated ones.The main result obtained here, which reinforces the findings of many N-body experiments, is that the collapse proceeds ''gently'' and not via violent relaxation. There is a strong correlation of the final energy of individual particles with the initial one. In particular, we find a preservation of the ranking of particles according to their binding energy. In cases where the analytic model predicts nonincreasing rotation curves its predictions are confirmed by the simulations. Otherwise, sensitive dependence on initial conditions is found, and the analytic model fails completely. In the cosmological context power spectra with n greater than or equal to - 1 yields (in the mean) nonincreasing rotation curves, and in such cases the secondary infall model is expected to be a useful tool in calculating the final virialized structure of collapsing halos.