We present the parallel module of a program called LASSAP which is intended to raise numerous limitations of current sequence comparison programs and which is able to handle homologies searches at a large scale level. This module is MIMD oriented, taking advantages of the intrinsic parallelism of most problems. It can run on various architectures (shared memory and message passing) and achieves near perfect speed-up on highly parallel machines (more than 100 processors). Work in progress focus on the way to deal with a network of parallel super-computers, taking advantage of the specificities of each of them.
Results from current homology-based search tools, designed to locate biologically relevant sequences, present a level of uncertainty from an intellectual property standpoint.
The Z-value is an attempt to estimate the statistical significance of a Smith–Waterman dynamic alignment score (SW-score) through the use of a Monte–Carlo process. It partly reduces the bias induced by the composition and length of the sequences.This paper is not a theoretical study on the distribution of SW-scores and Z-values. Rather, it presents a statistical analysis of Z-values on large datasets of protein sequences, leading to a law of probability that the experimental Z-values follow.First, we determine the relationships between the computed Z-value, an estimation of its variance and the number of randomizations in the Monte–Carlo process. Then, we illustrate that Z-values are less correlated to sequence lengths than SW-scores.Then we show that pairwise alignments, performed on ‘quasi-real’ sequences (i.e., randomly shuffled sequences of the same length and amino acid composition as the real ones) lead to Z-value distributions that statistically fit the extreme value distribution, more precisely the Gumbel distribution (global EVD, Extreme Value Distribution). However, for real protein sequences, we observe an over-representation of high Z-values.We determine first a cutoff value which separates these overestimated Z-values from those which follow the global EVD. We then show that the interesting part of the tail of distribution of Z-values can be approximated by another EVD (i.e., an EVD which differs from the global EVD) or by a Pareto law.This has been confirmed for all proteins analysed so far, whether extracted from individual genomes, or from the ensemble of five complete microbial genomes comprising altogether 16956 protein sequences.
Bacillus subtilis is the best-characterized member of the Gram-positive bacteria. Its genome of 4,214,810 base pairs comprises 4,100 protein-coding genes. Of these protein-coding genes, 53% are represented once, while a quarter of the genome corresponds to several gene families that have been greatly expanded by gene duplication, the largest family containing 77 putative ATP-binding transport proteins. In addition, a large proportion of the genetic capacity is devoted to the utilization of a variety of carbon sources, including many plant-derived molecules. The identification of five signal peptidase genes, as well as several genes for components of the secretion apparatus, is important given the capacity of Bacillus strains to secrete large amounts of industrially important enzymes. Many of the genes are involved in the synthesis of secondary metabolites, including antibiotics, that are more typically associated with Streptomyces species. The genome contains at least ten prophages or remnants of prophages, indicating that bacteriophage infection has played an important evolutionary role in horizontal gene transfer, in particular in the propagation of bacterial pathogenesis.
MOTIVATION:This paper presents LASSAP, a new software package for sequence comparison. LASSAP is a programmable, high-performance system designed to raise current limitations of sequence comparison programs in order to fit the needs of large-scale analysis. LASSAP provides an API (Application Programming Interface) allowing the integration of any generic pairwise-based algorithm.RESULTS:Whatever pairwise algorithm is used in LASSAP, it shares with all other algorithms numerous enhancements such as: (i) intra- and inter-databank comparisons; (ii) computational requests (selections and computations are achieved on the fly); (iii) frame translations on queries and databanks; (iv) structured results allowing easy and powerful post-analysis; (v) performance improvements by parallelization and the driving of specialized hardware. LASSAP currently implements all major sequence comparison algorithms (Fasta, Blast, Smith/Waterman), and other string matching and pattern matching algorithms. LASSAP is both an integrated software for end-users and a framework allowing the integration and the combination of new algorithms. LASSAP is used in different projects such as the building of PRODOM, the exhaustive comparison of yeast sequences, and the subfragments matching problem of TREMBL.
Cet article presente plusieurs machines paralleles specialisees pour la comparaison de sequences biologiques. Ce sont des machines principalement bâties autour d'un reseau lineaire de processeurs. Leurs performances depassent de plusieurs ordres de grandeur celles des machines programmables, permettant ainsi de faire face a l'accroissement extremement rapide des banques de sequences. This article presents several machines dedicated to biological sequence comparison. These machines are parallel machines based primarily on linear arrays. Their performance is several orders-of-magnitude better than that ofprogrammable machines, allowing them to face the challenge of the extremely fast growth of biological-sequence databases.
This paper describes the mathematical and computational techniques used at Genethon to obtain within one year a YAC contig map covering more than 50% of the whole human genome. The fingerprinting approach used has already yielded more than 1,000 contigs totaling more than 3,600 clones, after 16,896 clones from the CEPH YAC library have been analyzed. The resulting map will be a powerful tool for the identification of unknown genes, particularly those responsible for genetic diseases.