Characterisation of gender differences throughout peer-review publication process as revealed by thorough analysis of Royal Society of Chemistry submissions, publications and citation data.
In this paper, we explore the application of artificial neural network ('deep learning') methods to the problem of detecting chemical-protein interactions in PubMed abstracts. We present here a system using multiple Long Short Term Memory layers to analyse candidate interactions, to determine whether there is a relation and which type. A particular feature of our system is the use of unlabelled data, both to pre-train word embeddings and also pre-train LSTM layers in the neural network. On the BioCreative VI CHEMPROT test corpus, our system achieves an F score of 61.51% (56.10% precision, 67.84% recall).
Mass spectrometry information has long offered the potential of discovering biomarkers that would enable clinicians to diagnose disease, and treat it with targeted therapies. PeptideAtlas currently provides access to large-scale spectra data and identification information. This data, and the generation of targeted peptide information, represents the first step in the process of locating disease biomarkers. Reaching the goal of clinical proteomics requires that this data be integrated with additional information from disease literature and genomic studies. Here we describe PeptideAtlas and associated methods for mining the data, as well as the software tools necessary to support large-scale integration and mining.
BACKGROUND:The advances in high-throughput sequencing technologies and growth in data sizes has highlighted the need for scalable tools to perform quality assurance testing. These tests are necessary to ensure that data is of a minimum necessary standard for use in downstream analysis. In this paper we present the SAMQA tool to rapidly and robustly identify errors in population-scale sequence data.RESULTS:SAMQA has been used on samples from three separate sets of cancer genome data from The Cancer Genome Atlas (TCGA) project. Using technical standards provided by the SAM specification and biological standards defined by researchers, we have classified errors in these sequence data sets relative to individual reads within a sample. Due to an observed linearithmic speedup through the use of a high-performance computing (HPC) framework for the majority of tasks, poor quality data was identified prior to secondary analysis in significantly less time on the HPC framework than the same data run using alternative parallelization strategies on a single server.CONCLUSIONS:The SAMQA toolset validates a minimum set of data quality standards across whole-genome and exome sequences. It is tuned to run on a high-performance computational framework, enabling QA across hundreds gigabytes of samples regardless of coverage or sample type.
The advent of new high-throughput sequencing technologies has led to a flood of genomic data which overwhelms the capabilities of single processor machines. We present a MapReduce pipeline called Howdah that supports the analysis of genomic sequence data allowing multiple tests to be plugged in to a single MapReduce job. The pipeline is used to detect chromosomal abnormalities such as insertions, deletions and translocations as well as single nucleotide polymorphisms (SNPs).
In the life sciences, the need to balance the costs and benefits of introducing software processes into a research environment presents a distinct set of challenges due to the cultural disconnect between life sciences research and software engineering. The Institute for Systems Biology's research informatics team has studied these challenges and developed a software process to address them.
A robust probabilistic classification technique, using expectation maximization of finite mixture models, is used to analyze multi-frequency fisheries acoustic data. The number of clusters is chosen using the Bayesian Information Criterion. Probabilities of membership to clusters are used to classify each sample. The utility of the technique is demonstrated using two examples: the Gulf of Alaska representing a low-diversity, well-known system; and the Mid-Atlantic Ridge, a species-rich, relatively unknown system.
Posttranslational histone modifications participate in modulating the structure and function of chromatin. Promoters of transcribed genes are enriched with K4 trimethylation and hyperacetylation on the N-terminal tail of histone H3. Recently, PHD finger proteins, like Yng1 in the NuA3 HAT complex, were shown to interact with H3K4me3, indicating a biochemical link between K4 methylation and hyperacetylation. By using a combination of mass spectrometry, biochemistry, and NMR, we detail the Yng1 PHD-H3K4me3 interaction and the importance of NuA3-dependent acetylation at K14. Furthermore, genome-wide ChIP-Chip analysis demonstrates colocalization of Yng1 and H3K4me3 in vivo. Disrupting the K4me3 binding of Yng1 altered K14ac and transcription at certain genes, thereby demonstrating direct in vivo evidence of sequential trimethyl binding, acetyltransferase activity, and gene regulation by NuA3. Our data support a general mechanism of transcriptional control through which histone acetylation upstream of gene activation is promoted partially through availability of H3K4me3, "read" by binding modules in select subunits.
One of the strangest paradoxes of the silicon era is the dichotomy between ’enjoyable’ recreational computer activities and ’mundane’ work-based computer operations. How can an activity as pointless as a computer game have so much appeal? The answer to this lies in the user interface, and not the functionality, of the program. Computer games rely heavily on an interface which is natural and enjoyable to use. We believe that an interface should appeal to the user, and to do so must capture the user's interest and imagination. To this end, we have been using high performance graphics to generate meaningful three dimensional representations for our graphical user interface. We propose new metaphors for both query construction and result representation.
UNLABELLEDSeqExpress, a gene-expression analysis suite, has been extended to offer a number of cluster generation, refinement and visualization techniques. The cluster generation methods have been specialized to deal with aspects of the sparseness and extreme values that occur within microarray data. The results of such cluster analysis can then be refined using either: a functional enrichment based procedure, which examines each cluster to see if it possesses an unusually high or low concentration of ontology terms; or by using Expectation-Maximization to find a mixture of model based distributions within the datasets. Visualizations are provided both to explore and compare the results of the cluster generation algorithms. In addition, a tool has been developed which integrates SeqExpress with the Gene-Expression Omnibus repository. The tool provides seamless access to the large number of experimental results in the repository, so that they can be visualized and analysed locally using SeqExpress.AVAILABILITYSeqExpress is available as a 6 MB download from http://www.seqexpress.com and runs under Windows. A server-based version is available and is required for the GEO integration. SeqExpress is not affiliated with any academic institution, funding body or commercial organization and is free to use by all.
SUMMARY SeqExpress is a stand-alone desktop application for the identification of relevant genes within collections of microarray or SAGE experiments. A number of analysis, filtering and visualization tools are provided to aid in the selection of groups of genes. If R is installed then the application can use this to provide further analysis. AVAILABILITY SeqExpress is available at: http://www.seqexpress.com
This paper discusses the designs behind Amaze, a graphical user interface to an object oriented database system. Amaze uses three dimensional graphics to visualise both query construction and result representation. The designs behind Amaze are introduced, and the visualisations of a prototype implementation are shown. It is felt that just as word processors have changed how people interact with document processing, so to could database visualisation alter how people interact with database systems.
The yeast TRP3 gene encodes a bifunctional protein with anthranilate synthase II and indoleglycerol-phosphate synthase activities. Replacing ten consecutive non-preferred codons in the indoleglycerol-phosphate synthase region of the TRP3 gene with synonymous preferred codons (to create the TRP3pr gene; translational pause replaced) causes a 1.5-fold reduction in relative indoleglycerol-phosphate synthase activity [Crombie, T., Swaffield, J.C. & Brown, A.J.P. (1992) J. Mol. Biol. 228, 7-12]. Here, we report that both the anthranilate synthase II and indoleglycerol-phosphate synthase domains are affected to similar extents when the translational pause is removed. Also, structural modelling of the yeast indoleglycerol-phosphate synthase domain against the X-ray crystal structure of indoleglycerol-phosphate synthase from Escherichia coli indicates that the translational pause lies in a region of structural divergence between similar structures. To probe the role of cytoplasmic heat-shock protein 70 (Hsp 70) chaperones in Trp3 protein folding, anthranilate synthase and indoleglycerol-phosphate synthase activities were measured in ssa and ssb mutants. Neither indoleglycerol-phosphate synthase nor anthranilate synthase were affected significantly in the ssb mutant. However, depletion of Hsp70 proteins encoded by the SSA genes led to decreased anthranilate synthase and indoleglycerol-phosphate synthase activities from the TRP3 gene, suggesting that both domains depend to some extent upon the SSA chaperone family. The data are consistent with roles for both the translational pause and Ssa chaperones in Trp3 protein folding in vivo.
P. M. D. Gray合作论文数University of Aberdeen;Department of Computing Science2