The delivery of newly synthesized soluble lysosomal hydrolases to the endosomal system is essential for lysosome function and cell homeostasis. This process relies on the proper trafficking of the mannose 6-phosphate receptors (MPRs) between the trans-Golgi network (TGN), endosomes and the plasma membrane. Many transmembrane proteins regulating diverse biological processes ranging from virus production to the development of multicellular organisms also use these pathways. To explore how cell signaling modulates MPR trafficking, we used high-throughput RNA interference (RNAi) to target the human kinome and phosphatome. Using high-content image analysis, we identified 127 kinases and phosphatases belonging to different signaling networks that regulate MPR trafficking and/or the dynamic states of the subcellular compartments encountered by the MPRs. Our analysis maps the MPR trafficking pathways based on enzymes regulating phosphatidylinositol phosphate metabolism. Furthermore, it reveals how cell signaling controls the biogenesis of post-Golgi tubular carriers destined to enter the endosomal system through a SRC-dependent pathway regulating ARF1 and RAC1 signaling and myosin II activity.
BACKGROUND:The study of gene families is pivotal for the understanding of gene evolution across different organisms and such phylogenetic background is often used to infer biochemical functions of genes. Modern high-throughput experiments offer the possibility to analyze the entire transcriptome of an organism; however, it is often difficult to deduct functional information from that data.RESULTS:To improve functional interpretation of gene expression we introduce Ortho2ExpressMatrix, a novel tool that integrates complex gene family information, computed from sequence similarity, with comparative gene expression profiles of two pre-selected biological objects: gene families are displayed with two-dimensional matrices. Parameters of the tool are object type (two organisms, two individuals, two tissues, etc.), type of computational gene family inference, experimental meta-data, microarray platform, gene annotation level and genome build. Family information in Ortho2ExpressMatrix bases on computationally different protein family approaches such as EnsemblCompara, InParanoid, SYSTERS and Ensembl Family. Currently, respective all-against-all associations are available for five species: human, mouse, worm, fruit fly and yeast. Additionally, microRNA expression can be examined with respect to miRBase or TargetScan families. The visualization, which is typical for Ortho2ExpressMatrix, is performed as matrix view that displays functional traits of genes (differential expression) as well as sequence similarity of protein family members (BLAST e-values) in colour codes. Such translations are intended to facilitate the user's perception of the research object.CONCLUSIONS:Ortho2ExpressMatrix integrates gene family information with genome-wide expression data in order to enhance functional interpretation of high-throughput analyses on diseases, environmental factors, or genetic modification or compound treatment experiments. The tool explores differential gene expression in the light of orthology, paralogy and structure of gene families up to the point of ambiguity analyses. Results can be used for filtering and prioritization in functional genomic, biomedical and systems biology applications. The web server is freely accessible at http://bioinf-data.charite.de/o2em/cgi-bin/o2em.pl.
Insertions and deletions occur during evolution of biological sequences resulting in gaps in sequence alignments. The quality of an alignment depends on the placement of the gaps. Reliable pairwise as well as multiple sequence alignments are useful in inferring protein protein interacton sites through residue conservation[23], [24]. It has been reported that the Zipfian distribution best approximates the observed gap-lengths in the sequence alignments. The probability of a gap of length N decreases, inversely related to length, as a function of N−c for some suitable c. We have analysed four different gap scoring models: affine, log, power and the new Zipf that is based on Zipfian distribution. When tested on pairwise alignments from the BAliBASE benchmark suite, the widely used affine gaps were outperformed by the three other models. Log, Power and Zipf gap models performed comparably well.
Studying genetic variations in the human genome is important for understanding phenotypes and complex traits, including rare personal variations and their associations with disease. The interpretation of polymorphisms requires reliable methods to isolate natural genetic variations, including combinations of variations, in a format suitable for downstream analysis. Here, we describe a strategy for targeted isolation of large regions (∼35 kb) from human genomes that is also applicable to any genome of interest. The method relies on recombineering to fish out target fosmid clones from pools and thereby circumvents the laborious need to plate and screen thousands of individual clones. To optimize the method, a new highly recombineering-efficient bacterial host, including inducible TrfA for fosmid copy number amplification, was developed. Various regions were isolated from human embryonic stem cell lines and a personal genome, including highly repetitive and duplicated ones. The maternal and paternal alleles at the MECP2/IRAK 1 loci were distinguished based on identification of novel allele-specific single-nucleotide polymorphisms in regulatory regions. Additionally, we applied further recombineering to construct isogenic targeting vectors for patient-specific applications. These methods will facilitate work to understand the linkage between personal variations and disease propensity, as well as possibilities for personal genome surgery.
Motivation: We noted that the sumoylation site in C/EBP homologues is conserved beyond the canonical consensus sequence for sumoylation. Therefore, we investigated whether this pattern might define a more general protein motif. Results: We undertook a survey of the human proteome using a regular expression based on the C/EBP motif. This revealed significant enrichment of the motif using different Gene Ontology terms (e.g. ‘transcription’) that pertain to the nucleus. When considering requirements for the motif to be functional (evolutionary conservation, structural accessibility of the motif and proper cell localization of the protein), more than 130 human proteins were retrieved from the UniProt/Swiss-Prot database. These candidates were particularly enriched in transcription factors, including FOS, JUN, Hif-1α, MLL2 and members of the KLF, MAF and NFATC families; chromatin modifiers like CHD-8, HDAC4 and DNA Top1; and the transcriptional regulatory kinases HIPK1 and HIPK2. The KEPEmotif appears to be restricted to the metazoan lineage and has three length variants—short, medium and long—which do not appear to interchange. Contact: toby.gibson@embl.de Supplementary information: Supplementary data are available at Bioinformatics online.
MOTIVATION KEN-box-mediated target selection is one of the mechanisms used in the proteasomal destruction of mitotic cell cycle proteins via the APC/C complex. While annotating the Eukaryotic Linear Motif resource (ELM, http://elm.eu.org/), we found that KEN motifs were significantly enriched in human protein entries with cell cycle keywords in the UniProt/Swiss-Prot database-implying that KEN-boxes might be more common than reported. RESULTS Matches to short linear motifs in protein database searches are not, per se, significant. KEN-box enrichment with cell cycle Gene Ontology terms suggests that collectively these motifs are functional but does not prove that any given instance is so. Candidates were surveyed for native disorder prediction using GlobPlot and IUPred and for motif conservation in homologues. Among >25 strong new candidates, the most notable are human HIPK2, CHFR, CDC27, Dab2, Upf2, kinesin Eg5, DNA Topoisomerase 1 and yeast Cdc5 and Swi5. A similar number of weaker candidates were present. These proteins have yet to be tested for APC/C targeted destruction, providing potential new avenues of research.
SIRW (http://sirw.embl.de/) is a World Wide Web interface to the Simple Indexing and Retrieval System (SIR) that is capable of parsing and indexing various flat file databases. In addition it provides a framework for doing sequence analysis (e.g. motif pattern searches) for selected biological sequences through keyword search. SIRW is an ideal tool for the bioinformatics community for searching as well as analyzing biological sequences of interest.
Multidomain proteins predominate in eukaryotic proteomes. Individual functions assigned to different sequence segments combine to create a complex function for the whole protein. While on-line resources are available for revealing globular domains in sequences, there has hitherto been no comprehensive collection of small functional sites/motifs comparable to the globular domain resources, yet these are as important for the function of multidomain proteins. Short linear peptide motifs are used for cell compartment targeting, protein-protein interaction, regulation by phosphorylation, acetylation, glycosylation and a host of other post-translational modifications. ELM, the Eukaryotic Linear Motif server at http://elm.eu.org/, is a new bioinformatics resource for investigating candidate short non-globular functional motifs in eukaryotic proteins, aiming to fill the void in bioinformatics tools. Sequence comparisons with short motifs are difficult to evaluate because the usual significance assessments are inappropriate. Therefore the server is implemented with several logical filters to eliminate false positives. Current filters are for cell compartment, globular domain clash and taxonomic range. In favourable cases, the filters can reduce the number of retained matches by an order of magnitude or more.
SUMMARY:SIR is a Simple Indexing and Retrieval tool for indexing and searching biological flat file databases. SIR is a cross-platform solution entirely written in Python. Since the package is very small and installation is trivial, this would be an ideal solution for database providers to provide a custom retrieval tool to access them.AVAILABILITY:The modules will be made available at http://www.EMBLHeidelberg.de/~chenna/PySAT/sir.html
Expressed sequence tags (ESTs) are randomly sequenced cDNA clones. Currently, nearly 3 million human and 2 million mouse ESTs provide valuable resources that enable researchers to investigate the products of gene expression. The EST databases have proven to be useful tools for detecting homologous genes, for exon mapping, revealing differential splicing, etc. With the increasing availability of large amounts of poorly characterised eukaryotic (notably human) genomic sequence, ESTs have now become a vital tool for gene identification, sometimes yielding the only unambiguous evidence for the existence of a gene expression product. However, BLAST-based Web servers available to the general user have not kept pace with these developments and do not provide appropriate tools for querying EST databases with large highly spliced genes, often spanning 50 000-100 000 bases or more. Here we describe Gene2EST (http://woody.embl-heidelberg.de/gene2est/), a server that brings together a set of tools enabling efficient retrieval of ESTs matching large DNA queries and their subsequent analysis. RepeatMasker is used to mask dispersed repetitive sequences (such as Alu elements) in the query, BLAST2 for searching EST databases and Artemis for graphical display of the findings. Gene2EST combines these components into a Web resource targeted at the researcher who wishes to study one or a few genes to a high level of detail.
Always look on the bright side of life and at a method for debugging CGI programs on the command line.
MOTIVATION:While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way.RESULTS:We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.
Mutations in the recently cloned AIRE-1 gene[1, 2]cause autoimmune polyendocrinopathy-candidiasis-ectodermal dystrophy (APECED)—a recessive systemic disease that is also known as APS 1 or PGA I (see OMIM Number 240300 at http://www.ncbi.nlm. nih.gov/omim/). A rather variable set of symptoms characterize APECED; frequently, these include autoantibodies to the adrenal gland and chronic candidial infection. Unusually, for an autoimmune-associated gene, AIRE-1 is found outside the major histocompatibility complex (on chromosome 6) at chromosomal location 21q22.3. APECED is quite uncommon, except in certain populations, such as Finns (1:25 000), Iranian Jews (1:9000) and Sardinians, where there are higher carrier frequencies. Partly because of the variability of the symptoms, APECED is a rather recently characterized disease; modern chromosomal mapping techniques have been essential in verifying genetic homogeneity[3]and should continue to be useful in diagnosing APECED in populations where the disease is less common and, hence, more difficult to diagnose.
A specialized, interdisciplinary database on various types of related information on methanogenic bacteria is described. Derived from other sequence databases etc., this database collects information from many sources, including unpublished work from research laboratories working in this field, and makes them accessible from a single source, to interested scientists, free of cost. It is presently held in eight 48 T.P.I. floppy disks and can be run on any IBM PC under DOS 3.0 or above, making this database of particular interest to researchers with limited resources and on-line search/access facilities.