BACKGROUND:Interactions between the epigenome and structural genomic variation are potentially bi-directional. In one direction, structural variants may cause epigenomic changes in cis. In the other direction, specific local epigenomic states such as DNA hypomethylation associate with local genomic instability.METHODS:To study these interactions, we have developed several tools and exposed them to the scientific community using the Software-as-a-Service model via the Genboree Workbench. One key tool is Breakout, an algorithm for fast and accurate detection of structural variants from mate pair sequencing data.RESULTS:By applying Breakout and other Genboree Workbench tools we map breakpoints in breast and prostate cancer cell lines and tumors, discriminate between polymorphic breakpoints of germline origin and those of somatic origin, and analyze both types of breakpoints in the context of the Human Epigenome Atlas, ENCODE databases, and other sources of epigenomic profiles. We confirm previous findings that genomic instability in human germline associates with hypomethylation of DNA, binding sites of Suz12, a key member of the PRC2 Polycomb complex, and with PRC2-associated histone marks H3K27me3 and H3K9me3. Breakpoints in germline and in breast cancer associate with distal regulatory of active gene transcription. Breast cancer cell lines and tumors show distinct patterns of structural mutability depending on their ER, PR, or HER2 status.CONCLUSIONS:The patterns of association that we detected suggest that cell-type specific epigenomes may determine cell-type specific patterns of selective structural mutability of the genome.
Pregnancy is associated with significant change in leukocyte composition/function in the establishment/adaptation to pregnancy and parturition. We characterized methylation profiles longitudinally in a group of healthy pregnant women. The methylation profiles were characterized in maternal buffy coat DNA longitudinally in pregnancy: early (<16 gestational weeks GW), mid (24-28 GW), delivery and 6 weeks postpartum in 14 nulligravid, non-smoking, normotensive women. Bisulfite modified genomic DNA was run on the Illumina Methylation Assay 27K platform. Mean methylation levels at each CpG site were compared across all time points using a paired t-test and adjusted for multiple comparisons. Global trends between time points were described as % hypo-methylated. Pathway analysis was performed on differentially methylated genes. There were 2692 unique CpG sites differentially methylated at one or more time points in pregnancy compared to baseline postpartum levels. Most CpGs become differentially methylated early in pregnancy and remain so throughout (870, 32.31%), change during mid-pregnancy and persist through delivery (690, 25.63%) or are altered only at delivery (574, 21.32%). The vast majority of differentially methylated sites become hypo-methylated compared with baseline (early 91%, mid 84%, and delivery 81%). Pathway analysis of subfractions identified expected processes such as innate immune response and leukocyte migration but also processes unique to different time points in pregnancy, such as smooth muscle contraction and mammary gland development at the time of delivery for example (all p<10 -7). The differential methylation of maternal leukocyte DNA that starts in the first trimester continues throughout gestation and trends toward hypomethylation. Pathway analysis demonstrates that epigenetic regulation plays a key role in physiologic processes associated not only with immune function but other normal maternal adaptations in healthy pregnancy.
Background Microbial metagenomic analyses rely on an increasing number of publicly available tools. Installation, integration, and maintenance of the tools poses significant burden on many researchers and creates a barrier to adoption of microbiome analysis, particularly in translational settings. Methods To address this need we have integrated a rich collection of microbiome analysis tools into the Genboree Microbiome Toolset and exposed them to the scientific community using the Software-as-a-Service model via the Genboree Workbench. The Genboree Microbiome Toolset provides an interactive environment for users at all bioinformatic experience levels in which to conduct microbiome analysis. The Toolset drives hypothesis generation by providing a wide range of analyses including alpha diversity and beta diversity, phylogenetic profiling, supervised machine learning, and feature selection. Results We validate the Toolset in two studies of the gut microbiota, one involving obese and lean twins, and the other involving children suffering from the irritable bowel syndrome. Conclusions By lowering the barrier to performing a comprehensive set of microbiome analyses, the Toolset empowers investigators to translate high-volume sequencing data into valuable biomedical discoveries.
Background: There is an increasing usage of ion mobility-mass spectrometry (IMMS) in proteomics. IMMS combines the features of ion mobility spectrometry (IMS) and mass spectrometry ( MS). It separates and detects peptide ions on a millisecond time-scale. IMS separates peptide ions based on drift time that is determined by the collision cross-section of each peptide ion in a given experiment condition. A peptide ion's collision cross-section is related to the ion size and shape resulted from the peptide amino acid sequence and their modifications. This inherent relation between the drift time of peptide ion and peptide sequence indicates that the drift time of peptide ions can be used to infer peptide sequence and therefore, for peptide identification.Results: This paper describes an artificial neural networks (ANNs) regression model for the prediction of peptide ion drift time in IMMS. Each peptide in this work was represented using three descriptors (i.e., molecular weight, sequence length and a two-dimensional sequence index). An ANN predictor consisting of four input nodes, three hidden nodes and one output node was constructed for peptide ion drift time prediction. For the model training and testing, a 10-fold cross-validation strategy was employed for three datasets each containing different charge states. Dataset one contains 212 singly-charged peptide ions, dataset two has 306 doubly-charged peptide ions, and dataset three has 77 triply-charged peptide ions. Our proposed method achieved 94.4%, 93.6% and 74.2% prediction accuracy for singly-, doubly- and triply-charged peptide ions, respectively.Conclusions: An ANN-based method has been developed for predicting the drift time of peptide ions in IMMS. The results achieved here demonstrate the effectiveness and efficiency of the prediction model. This work can enhance the confidence of protein identification by combining with current database search approaches for protein identification.
Background Understanding the proteome, the structure and function of each protein, and the interactions among proteins will give clues to search useful targets and biomarkers for pharmaceutical design. Peptide drift time prediction in IMMS will improve the confidence of peptide identification by limiting the peptide search space during MS/MS database searching and therefore reducing false discovery rate (FDR) of protein identification. A peptide drift time prediction method was proposed here using an artificial neural networks (ANN) regression model. We test our proposed model on three peptide datasets with different charge state assignment (see Table 1). The results can be found in Figure 1, where a higher prediction performance was achieved, over 0.9 for CI and C2, as well as 0.75 for C3.