Summary: MMseqs2 taxonomy is a new tool to assign taxonomic labels to metagenomic contigs. It extracts all possible protein fragments from each contig, quickly retains those that can contribute to taxonomic annotation, assigns them with robust labels and determines the contig’s taxonomic identity by weighted voting. Its fragment extraction step is suitable for the analysis of all domains of life. MMseqs2 taxonomy is 2–18 faster than state-of-the-art tools and also contains new modules for creating and manipulating taxonomic reference databases as well as reporting and visualizing taxonomic assignments. Availability and implementation: MMseqs2 taxonomy is part of the MMseqs2 free open-source software package available for Linux, macOS and Windows at https://mmseqs.com. Contact: soeding@mpibpc.mpg.de or eli.levy.karin@gmail.com Supplementary information: Supplementary data are available at Bioinformatics online.
Generation of context profile library N = 1 million training profiles of length l = 2d+1 were generated as described in the main text and Figure 2. Each training profile is represented by a count profile cn(j, x), which specifies the counts of amino acid x ∈ {1, . . . , 20} at position j ∈ {−d, . . . , d}. These counts are obtained by multiplying the sequence profile tn(j, x) by the effective number of sequences Nn(j) at position j in the alignment from which training profile tn(j, x) was calculated: cn(j, x) = Nn(j)tn(j, x) (see next section for details). Here, we describe how these N profiles are clustered in order to obtain a set of K context profiles which recur frequently among the training profiles and which together can describe all training profiles. More precisely, we seek to determine context profiles p = (p1, . . . , pK) and their prior probabilities α = (α1, . . . , αK) that maximize the likelihoodP (c|p, α) that the training profile counts c = (c1, . . . , cN ) were generated by the context profiles. We model the distribution of counts cn(j, x) in each column j by a multinomial distribution. Since cn(j, x) can be real-valued, however, we replace the factorials in the multinomial distribution by Gamma functions (n! = Γ(n + 1)). The probability for context profile pk to have emitted counts cn(j, x) (j ∈ {−d, . . . , d}, x ∈ {1, . . . , 20}) is
With over 47,000 citations, the sequence search method BLAST has been an essential tool in biological research since its development in 1990. By accounting for the influence of sequence context on the mutation probabilities of amino acids, context-specific BLAST (CS-BLAST) achieves two-fold higher sensitivity for distantly related protein sequences, at the same speed and error rate.
Motivation: An estimated 25% of all eukaryotic proteins contain repeats, which underlines the importance of duplication for evolving new protein functions. Internal repeats often correspond to structural or functional units in proteins. Methods capable of identifying repeated segments or domains at the sequence level can therefore assist in predicting domain structures, inferring hypotheses about function and mechanism, and investigating the evolution of proteins from smaller
MOTIVATION:Protein homology detection and sequence alignment are at the basis of protein structure prediction, function prediction and evolution.RESULTS:We have generalized the alignment of protein sequences with a profile hidden Markov model (HMM) to the case of pairwise alignment of profile HMMs. We present a method for detecting distant homologous relationships between proteins based on this approach. The method (HHsearch) is benchmarked together with BLAST, PSI-BLAST, HMMER and the profile-profile comparison tools PROF_SIM and COMPASS, in an all-against-all comparison of a database of 3691 protein domains from SCOP 1.63 with pairwise sequence identities below 20%.Sensitivity: When the predicted secondary structure is included in the HMMs, HHsearch is able to detect between 2.7 and 4.2 times more homologs than PSI-BLAST or HMMER and between 1.44 and 1.9 times more than COMPASS or PROF_SIM for a rate of false positives of 10%. Approximately half of the improvement over the profile-profile comparison methods is attributable to the use of profile HMMs in place of simple profiles. Alignment quality: Higher sensitivity is mirrored by an increased alignment quality. HHsearch produced 1.2, 1.7 and 3.3 times more good alignments ('balanced' score >0.3) than the next best method (COMPASS), and 1.6, 2.9 and 9.4 times more than PSI-BLAST, at the family, superfamily and fold level, respectively.Speed: HHsearch scans a query of 200 residues against 3691 domains in 33 s on an AMD64 2GHz PC. This is 10 times faster than PROF_SIM and 17 times faster than COMPASS.
AbrB is a key transition-state regulator of Bacillus subtilis. Based on the conservation of a βαβ structural unit, we proposed a β barrel fold for its DNA binding domain, similar to, but topologically distinct from, double-psi β barrels. However, the NMR structure revealed a novel fold, the “looped-hinge helix.” To understand this discrepancy, we undertook a bioinformatics study of AbrB and its homologs; these form a large superfamily, which includes SpoVT, PrlF, MraZ, addiction module antidotes (PemI, MazE), plasmid maintenance proteins (VagC, VapB), and archaeal PhoU homologs. MazE and MraZ form swapped-hairpin β barrels. We therefore reexamined the fold of AbrB by NMR spectroscopy and found that it also forms a swapped-hairpin barrel. The conservation of the core βαβ element supports a common evolutionary origin for swapped-hairpin and double-psi barrels, which we group into a higher-order class, the cradle-loop barrels, based on the peculiar shape of their ligand binding site.
REPPER (REPeats and their PERiodicities) is an integrated server that detects and analyzes regions with short gapless repeats in protein sequences or alignments. It finds periodicities by Fourier Transform (FTwin) and internal similarity analysis (REPwin). FTwin assigns numerical values to amino acids that reflect certain properties, for instance hydrophobicity, and gives information on corresponding periodicities. REPwin uses self-alignments and displays repeats that reveal significant internal similarities. Both programs use a sliding window to ensure that different periodic regions within the same protein are detected independently. FTwin and REPwin are complemented by secondary structure prediction (PSIPRED) and coiled coil prediction (COILS), making the server a versatile analysis tool for sequences of fibrous proteins. REPPER is available at http://protevo.eb.tuebingen.mpg.de/repper.
HHpred is a fast server for remote protein homology detection and structure prediction and is the first to implement pairwise comparison of profile hidden Markov models (HMMs). It allows to search a wide choice of databases, such as the PDB, SCOP, Pfam, SMART, COGs and CDD. It accepts a single query sequence or a multiple alignment as input. Within only a few minutes it returns the search results in a user-friendly format similar to that of PSI-BLAST. Search options include local or global alignment and scoring secondary structure similarity. HHpred can produce pairwise query-template alignments, multiple alignments of the query with a set of templates selected from the search results, as well as 3D structural models that are calculated by the MODELLER software from these alignments. A detailed help facility is available. As a demonstration, we analyze the sequence of SpoVT, a transcriptional regulator from Bacillus subtilis. HHpred can be accessed at http://protevo.eb.tuebingen.mpg.de/hhpred.
We have measured the three-body decay of a Bose–Einstein condensate of rubidium (87Rb) atoms prepared in the doubly polarized ground state F=m F =2. Our data are taken for a peak atomic density in the condensate varying between 2×1014 cm-3 at initial time and 7×1013 cm-3, 16 s later. Taking into account the influence of the uncondensed atoms on the decay of the condensate, we deduce a rate constant for condensed atoms L=1.8 (±0.5) ×10-29 cm6 s-1. For these densities we did not find a significant contribution of two-body processes such as spin dipole relaxation.
We have measured the rate of inelastic collisions in a cloud of doubly polarized ground-state cesium atoms (F = m(F) = 4) confined in a magnetic trap for temperatures T between 8 and 70 mu K. We find a two-body rate coefficient varying as T-0.63. At 8 mu K it reaches 4 X 10(-12) cm(3) s(-1) which is 3 orders of magnitude larger than predicted, ruling out a Bose-Einstein condensation of Cs in this internal state.
Buse-Einstein condensation of an atomic Rubidium gas is achieved using a novel magncric trapping achemc. The iitomi itre first rnipped in a magnetic quadrupole geometry which is then convened inlo B lolie yometry The lriip providea strong confinemenr of the ~ o m i and allows efficient evaporative coohng. I, "WE only three cur~enl carrying coil, (i.r. no iiddifioniil bias cotli) and merely 600 Walls of electric p o w r are dissipaled. Two olthe coils produce the 09.30 QTuA4
We investigate the dipole relaxation of a cesium atomic gas prepared in its lowest hyperfine level (F = 3) and confined in a magnetic trap. We measure a rate ∼ 4 × 10−13, for a field of 0.1 mT and a temperature 1 μK. This value is 2 orders of magnitude larger than for lighter alkalis. It puts strong constraints on the trapping mechanism which could lead to the Bose-Einstein condensation of cesium.
Using forced radio-frequency evaporation, we have cooled cesium atoms prepared in the sublevel F = -m(F) = 3 and confined in a magnetic trap. At the end of the evaporation ramp, the sample contains ~ 7000 atoms at 80 nK, corresponding to a phase space density 3 x 10(-2). A molecular dynamics approach, including the effect of gravity, gives a good account for the experimental data, assuming a scattering length larger than 300 Angstrom.
We have decelerated a cesium atomic beam from thermal velocities down to several tens of m/s within only a 10 cm slowing distance. A bichromatic standing light wave was used to generate a stimulated force exceeding the spontaneous force limit by a factor of similar to 10 and extending over a large, saturation-broadened velocity range. Because of the short slowing distance this method allows production of very intense, continuous beams of slow atoms.
We propose a trap based on an evanescent light wave which is formed on the surface of a conical or a pyramidal cavity in a glass substrate. The far blue-detuned evanescent wave acts as a mirror for atoms, bouncing on it in the gravitational field. In addition, fast and deep Sisyphus cooling in the evanescent wave can be organized for alkali atoms. This leads to the cooling of the atoms down to the recoil-limited temperature and their collection in the field- free region near the bottom of the trap, to form an atomic gas of extremely high density.
We propose a trap based on an evanescent light wave formed on the surface of a pyramidal hollow in a glass substrate. Alkali atoms repeatedly reflected at the evanescent wave and kept in their lower hyperfine ground state by a weak repumping laser beam can be cooled to recoil-limited temperatures by a Sisyphus and a geometric cooling mechanism, which are connected with spontaneous transitions to the upper hyperfine ground state during the reflections. Numerical simulations for 39K, 85Rb and 133Cs predict equilibrium 3D rms momenta of ∼3.5 h̵k. We estimate the collisional loss rates and show that the trap can produce extremely cold and dense samples of atoms and may even reach the point of Bose-Einstein condensation.