Current approaches to protein identification rely heavily on database matching of fragmentation spectra or precursor peptide ions. We have developed a method for MALDI TOF-TOF instrumentation that uses peptide masses and their measurement errors to confirm protein identifications from a first pass MS/MS database search. The method uses MS1-level spectral data that have heretofore been ignored by most search engines. This approach uses the distribution of mass errors of peptide matches in the MS1 spectrum to develop a probability model that is independent of the MS/MS database search identifications. Peptide mass matches can come from both precursor ions that have been fragmented as well as those that are tentatively identified by accurate mass alone. This additional corroboration enables us to confirm protein identifications to MS/MS-based scores that are otherwise considered to be only of moderate quality. Straightforward and easily applicable to current proteomic analyses, this tool termed "Protein-Processor" provides a robust and invaluable addition to current protein identification tools.
Alterations in DNA methylation have been reported to occur during development and aging; however, much remains to be learned regarding post-natal and age-associated epigenome dynamics, and few if any investigations have compared human methylome patterns on a whole genome basis in cells from newborns and adults. The aim of this study was to reveal genomic regions with distinct structure and sequence characteristics that render them subject to dynamic post-natal developmental remodeling or age-related dysregulation of epigenome structure. DNA samples derived from peripheral blood monocytes and in vitro differentiated dendritic cells were analyzed by methylated DNA Immunoprecipitation (MeDIP) or, for selected loci, bisulfite modification, followed by next generation sequencing. Regions of interest that emerged from the analysis included tandem or interspersed-tandem gene sequence repeats (PCDHG, FAM90A, HRNR, ECEL1P2), and genes with strong homology to other family members elsewhere in the genome (FZD1, FZD7 and FGF17). Our results raise the possibility that selected gene sequences with highly homologous copies may serve to facilitate, perhaps even provide a clock-like function for, developmental and age-related epigenome remodeling. If so, this would represent a fundamental feature of genome architecture in higher eukaryotic organisms.
We present the first application of the quality threshold (QT) clustering algorithm to mass spectrometry (MS) data. The unique abilities of QT clustering to yield precision nodes that are commensurate with the mass measurement precision of the instrument are exploited to generate a consensus spectrum out of multiple replicate spectra. The spectral dot product and confidence intervals are used as a tool for evaluating the similarity and reproducibility between the consensus and replicates. The method is equally applicable to high and low resolution measurements. This paper demonstrates applications to linear spectra from a matrix assisted laser desorption ionization (MALDI) time of flight (TOF) instrument as well as peptide fragmentation data obtained from a TOF/TOF after unimolecular decomposition. The advantages of clustering to mitigate the inherent precision the shortcomings of MALDI data are discussed.
The zebrafish pineal gland (epiphysis) is a site of melatonin production, contains photoreceptor cells, and functions as a circadian clock pace maker. Here, we have used microarray technology to study the zebrafish pineal transcriptome. Analysis of gene expression at three larval and two adult stages revealed a highly dynamic transcriptional profile, revealing many genes that are highly expressed in the zebrafish pineal gland. Statistical analysis of the data based on Gene Ontology annotation indicates that many transcription factors are highly expressed during larval stages, whereas genes dedicated to phototransduction are preferentially expressed in the adult. Furthermore, several genes were identified that exhibit day/night differences in expression. Among the multiple candidate genes suggested by these data, we note the identification of a tissue-specific form of the unc119 gene with a possible role in pineal development.
Publisher Summary De novo sequencing involves deducing a peptide's sequence purely from mass spectral fragmentation data. There are roughly two ways to view the relationship between de novo sequencing and database search algorithms. First is complementary, and the second is alternative. Approaches that emphasize the complementary relationship between de novo sequences and database searches treat de novo sequences as a means to enhance the quality and reliability of database searches. Regimes originating from this approach utilize the spectra to derive one or more highly reliable sequence tags. These tags guide the database matching process, and since they ostensibly represent the most prominent and reliable features of the spectra, the tags also impart a higher degree of confidence to the database results. De novo sequencing approaches also differ according to the way in which they determine the fragment type (y or b) and score the sequence information. These differences in approach tend to generate great differences between de novo sequences obtained even from the same spectra. There are two principal approaches to the implementation of partial de novo sequencing called the global and local paradigms.
The effects of laser fluence on ion formation in MALDI were studied using a tandem TOF mass spectrometer with a Nd-YAG laser and alpha-cyano hydrocinnamic acid matrix. Leucine enkephalin ionization and fragmentation were followed as a function of laser fluence ranging from the threshold of ion formation to the maximum available, that is, about 280-930 mJ/mm2. The most notable finding was the appearance of immonium ions at fluence values close to threshold, increasing rapidly and then tapering in intensity with the appearance of typical backbone fragment ions. The data suggest the presence of two distinct environments for ion formation. One is associated with molecular desorption at low values of laser fluence that leads to extensive immonium ion formation. The second becomes dominant at higher fluences, is associated initially with backbone type fragments, but, at the highest values of fluence, progresses to immonium fragments. This second environment is suggestive of ion desorption from large pieces of material ablated from the surface. Arrhenius rate law considerations were used to estimate temperatures associated with the onset of these two processes.
We introduce the use of a peptide composition lookup table indexed by residual mass and number of amino acids for de novo sequencing of polypeptides. Polypeptides of 1600 Daltons (Da) or more can be sequenced effectively through exhaustive compositional analysis of MS/MS spectra obtained by unimolecular decomposition (without CID) in a MALDI TOF/TOF despite a fragment mass accuracy of 50 mDa. Peaks are referenced against the lookup table to obtain a complete profile of amino acid combinations, and combinations are assembled into series of increasing length. Concatenating the differences between successive entries in compositional series yields peptide sequences that can be scored and ranked according to signal intensity. While the current work involves measurements acquired on MALDI TOF-TOF, such general treatment of the data anticipates extension to other types of mass analyzers.
We describe a web-based program called 'DBParser' for rapidly culling, merging, and comparing sequence search engine results from multiple LC-MS/MS peptide analyses. DBParser employs the principle of parsimony to consolidate redundant protein assignments and derive the most concise set of proteins consistent with all of the assigned peptide sequences observed in an experiment or series of experiments. The resulting reports summarize peptide and protein identifications from multidimensional experiments that may contain a single data set or combine data from a group of data sets, all related to a single analytical sample. Additionally, the results of multiple experiments, each of which may contain several data sets, can be compared in reports that identify features that are common or different. DBParser actively links to the primary mass spectral data and to public online databases such as NCBI, GO, and Swiss-Prot in order to structure contextually specific reports for biologists and biochemists.
This book is designed specifically for candidates preparing for the MRCS Viva examination. The format of the exam has been used in the book's structure, with over 1000 questions to illustrate the key points of over 200 topics. The book has been divided into 6 main chapters corresponding to a viva 'station' in the exam. Each chapter begins with a check-list of the main topics and themes covered by the exam, and information is then subdivided logically within the chapter itself. Key 'pass-or-fail' concepts are covered, and in some cases, topics are covered in more detail than will be asked in the exam, so the candidate can be confident that they will be fully prepared.