Plant architecture is a major determinant of planting density, which enhances productivity potential for crops per unit area. Genomic prediction is well positioned to expedite genetic gain of plant architectural traits since they are typically highly heritable. Additionally, the adaptation of genomic prediction models to query predictive abilities of markers tagging certain genomic regions could shed light on the genetic architecture of these traits. Here, we leveraged transcriptional networks from a prior study that contextually described developmental progression during tassel and leaf organogenesis in maize (Zea mays) to inform genomic prediction models for architectural traits. Since these developmental processes underlie tassel branching and leaf angle, 2 important agronomic architectural traits, we tested whether genes prioritized from these networks quantitatively contribute to the genetic architecture of these traits. We used genomic prediction models to evaluate the ability of markers in the vicinity of prioritized network genes to predict breeding values of tassel branching and leaf angle traits for 2 diversity panels in maize and diversity panels from sorghum (Sorghum bicolor) and rice (Oryza sativa). Predictive abilities of markers near these prioritized network genes were similar to those using whole-genome marker sets. Notably, markers near highly connected transcription factors from core network motifs in maize yielded predictive abilities that were significantly greater than expected by chance in not only maize but also closely related sorghum. We expect that these highly connected regulators are key drivers of architectural variation that are conserved across closely related cereal crop species.
Understanding the impact of regulatory variants on complex phenotypes is a significant challenge because the genes and pathways that are targeted by such variants and the cell type context in which regulatory variants operate are typically unknown. Cell-type-specific long-range regulatory interactions that occur between a distal regulatory sequence and a gene offer a powerful framework for examining the impact of regulatory variants on complex phenotypes. However, high-resolution maps of such long-range interactions are available only for a handful of cell types. Furthermore, identifying specific gene subnetworks or pathways that are targeted by a set of variants is a significant challenge. We have developed L-HiC-Reg, a Random Forests regression method to predict high-resolution contact counts in new cell types, and a network-based framework to identify candidate cell-type-specific gene networks targeted by a set of variants from a genome-wide association study (GWAS). We applied our approach to predict interactions in 55 Roadmap Epigenomics Mapping Consortium cell types, which we used to interpret regulatory single nucleotide polymorphisms (SNPs) in the NHGRI-EBI GWAS catalogue. Using our approach, we performed an in-depth characterization of fifteen different phenotypes including schizophrenia, coronary artery disease (CAD) and Crohn's disease. We found differentially wired subnetworks consisting of known as well as novel gene targets of regulatory SNPs. Taken together, our compendium of interactions and the associated network-based analysis pipeline leverages long-range regulatory interactions to examine the context-specific impact of regulatory variation in complex phenotypes.
Differential chromatin interaction analysis between MCF-7 and T-47D based on Hi-C data.
Some of inherited human genetic variation can contribute to important phenotypic diversity, such as the varying degrees of individual susceptibility to developing certain health conditions and individual response to therapeutic interventions. To date, over 490,000 genotypephenotype associations have been discovered through large-scale genome-wide association studies (GWAS) [1]; however, molecular functions of most of these discovered GWAS variants remain unknown. There are several technical challenges hindering our understanding: (1) the effect size of a typical genetic variant, as measured in terms of the odds ratio of genotype occurrence in case versus control populations, is very small, suggesting that macroscopic systems-level phenotypic differences modulated by each variant may also be small and difficult to detect; (2) most reported variants reside in non-proteincoding regions of the human genome, indicating that they are likely affecting the regulation of some unknown target genes’ expression; and, (3) the discovered variants may not be functional themselves, but be merely in genetic linkage disequilibrium with other functional variants. A promising approach to address these challenges is to integrate genomic, epigenomic, transcriptomic and machine learning methods to identify functional genetic variants and characterize their mode of action in regulating target genes. One particular mode of regulatory function amenable to this integrative analysis is altering the binding affinity of transcription factors (TF) to DNA recognition sequences [2]. That is, assuming that a causative variant perturbs the binding activity of a TF, one can focus on the variants that are genetically linked to a given GWAS variant and located in transcriptionally active open chromatin regions annotated via epigenomic profiling – e.g., DNase-seq, ATAC-seq, and histone modification signatures of enhancers and promoters, often available in public databases such as the Encyclopedia of DNA Elements (ENCODE), Roadmap Epigenomics Mapping Consortium (REMC) and Gene Expression Omnibus (GEO) [3–5]. The ability of these epigenomically filtered candidate variants to perturb the binding activity of a specific TF can then be assessed computationally by training machine learning algorithms on TF ChIP-seq and HT-SELEX-seq data to learn the salient features of preferred DNA recognition sequences and to predict how the variants in the context of surrounding nucleotides alter the strength of TF-DNA interaction [2, 6–12]. Allelespecific binding preferences of predicted TFs can be verified by searching for skewed allele frequencies of the candidate variants in raw ChIP-seq reads, appropriately taking into account potential mapping biases. Target genes that are differentially expressed between case and control populations as a result of the predicted perturbation of TF binding activity may then be identified via expression quantitative trait loci and allele-specific expression analyses using processed and raw RNA-seq data from The Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) projects [13–15]; further support can
Background. Large-scale genome-wide association studies (GWAS) have implicated thousands of germline genetic variants in modulating individuals' risk to various diseases, including cancer. At least 25 risk loci have been identified for low-grade gliomas (LGGs), but their molecular functions remain largely unknown. Methods. We hypothesized that GWAS loci contain causal single nucleotide polymorphisms (SNPs) that reside in accessible open chromatin regions and modulate the expression of target genes by perturbing the binding affinity of transcription factors (TFs). We performed an integrative analysis of genomic and epigenomic data from The Cancer Genome Atlas and other public repositories to identify candidate causal SNPs within linkage disequilibrium blocks of LGG GWAS loci. We assessed their potential regulatory role via in silico TF binding sequence perturbations, convolutional neural network trained onTF binding data, and simulated annealing-based interpretation methods. Results. We built an interactive website (http://education.knoweng.org/alg3/) summarizing the functional footprinting of 280 variants in 25 LGG GWAS regions, providing rich information for further computational and experimental scrutiny. We identified as case studies PHLDB1 and SLC25A26 as candidate target genes of rs12803321 and rs11706832, respectively, and predicted the GWAS variant rs648044 to be the causal SNP modulating ZBTB16, a known tumor suppressor in multiple cancers. We showed that rs648044 likely perturbed the binding affinity of the TF MAFF, as supported by RNA interference and in vitro MAFF binding experiments. Conclusions. The identified candidate (causal SNP, target gene, TF) triplets and the accompanying resource will help accelerate our understanding of the molecular mechanisms underlying genetic risk factors for gliomas.
Over the past decade, hundreds of genome-wide association studies (GWAS) have implicated genetic variants in various diseases, including cancer. However, only a few of these variants have been functionally characterized to date, mainly because the majority of the variants reside in non-coding regions of the human genome with unknown function. A comprehensive functional annotation of the candidate variants is thus necessary to fill the gap between the correlative findings of GWAS and the development of therapeutic strategies. By integrating large-scale multi-omics datasets such as the Cancer Genome Atlas (TCGA) and the Encyclopedia of DNA Elements (ENCODE), we performed multivariate linear regression analysis of expression quantitative trait loci, sequence permutation test of transcription factor binding perturbation, and modeling of three-dimensional chromatin interactions to analyze the potential molecular functions of 2,813 single nucleotide variants in 93 genomic loci associated with estrogen receptor-positive breast cancer. To facilitate rapid progress in functional genomics of breast cancer, we have created “Analysis of Breast Cancer GWAS” (ABC-GWAS), an interactive database of functional annotation of estrogen receptor-positive breast cancer GWAS variants. Our resource includes expression quantitative trait loci, long-range chromatin interaction predictions, and transcription factor binding motif analyses to prioritize putative target genes, causal variants, and transcription factors. An embedded genome browser also facilitates convenient visualization of the GWAS loci in genomic and epigenomic context. ABC-GWAS provides an interactive visual summary of comprehensive functional characterization of estrogen receptor-positive breast cancer variants. The web resource will be useful to both computational and experimental biologists who wish to generate and test their hypotheses regarding the genetic susceptibility, etiology, and carcinogenesis of breast cancer. ABC-GWAS can also be used as a user-friendly educational resource for teaching functional genomics. ABC-GWAS is available at http://education.knoweng.org/abc-gwas/.
Genome-wide association studies (GWAS) have hitherto identified several germline variants associated with cancer susceptibility, but the molecular functions of these risk modulators remain largely uncharacterized. Recent studies have begun to uncover the regulatory potential of noncoding GWAS SNPs using epigenetic information in corresponding cancer cell types and matched normal tissues. However, this approach does not explore the potential effect of risk germline variants on other important cell types that constitute the microenvironment of tumor or its precursor. This paper presents evidence that the breast-cancer-associated variant rs3903072 may regulate the expression of CTSW in tumor-infiltrating lymphocytes. CTSW is a candidate tumor-suppressor gene, with expression highly specific to immune cells and also positively correlated with breast cancer patient survival. Integrative analyses suggest a putative causative variant in a GWAS-linked enhancer in lymphocytes that loops to the 3’ end of CTSW through three-dimensional chromatin interaction. Our work thus poses the possibility that a cancer-associated genetic variant could regulate a gene not only in the cell of cancer origin but also in immune cells in the microenvironment, thereby modulating the immune surveillance by T lymphocytes and natural killer cells and affecting the clearing of early cancer initiating cells.
SUMMARY:Next-generation sequencing (NGS) techniques are revolutionizing biomedical research by providing powerful methods for generating genomic and epigenomic profiles. The rapid progress is posing an acute challenge to students and researchers to stay acquainted with the numerous available methods. We have developed an interactive online educational resource called Sequencing Techniques Engine for Genomics (SequencEnG) to provide a tree-structured knowledge base of 66 different sequencing techniques and step-by-step NGS data analysis pipelines comparing popular tools. SequencEnG is designed to facilitate barrier-free learning of current NGS techniques and provides a user-friendly interface for searching through experimental and analysis methods. AVAILABILITY AND IMPLEMENTATION:SequencEnG is part of the project Knowledge Engine for Genomics (KnowEnG) and is freely available at http://education.knoweng.org/sequenceng/.
Abstract Previous genome-wide association studies (GWAS) have identified several common genetic variants that may significantly modulate cancer susceptibility. However, the precise molecular mechanisms behind these associations remain largely unknown; it is often not clear whether discovered variants are themselves functional or merely genetically linked to other functional variants. Here, we provide an integrated method for identifying functional regulatory variants associated with cancer and their target genes by combining analyses of expression quantitative trait loci, a modified version of allele-specific expression that systematically utilizes haplotype information, transcription factor (TF)–binding preference, and epigenetic information. Application of our method to a breast cancer susceptibility region in 5p12 demonstrates that the risk allele rs4415084-T correlates with higher expression levels of the protein-coding gene mitochondrial ribosomal protein S30 (MRPS30) and lncRNA RP11-53O19.1. We propose an intergenic SNP rs4321755, in linkage disequilibrium (LD) with the GWAS SNP rs4415084 (r2 = 0.988), to be the predicted functional SNP. The risk allele rs4321755-T, in phase with the GWAS rs4415084-T, created a GATA3-binding motif within an enhancer, resulting in differential GATA3 binding and chromatin accessibility, thereby promoting transcription of MRPS30 and RP11-53O19.1. MRPS30 encodes a member of the mitochondrial ribosomal proteins, implicating the role of risk SNP in modulating mitochondrial activities in breast cancer. Our computational framework provides an effective means to integrate GWAS results with high-throughput genomic and epigenomic data and can be extended to facilitate rapid functional characterization of other genetic variants modulating cancer susceptibility. Significance: Unification of GWAS results with information from high-throughput genomic and epigenomic profiles provides a direct link between common genetic variants and measurable molecular perturbations. Cancer Res; 78(7); 1579–91. ©2018 AACR.
Abstract Genome-wide association studies (GWAS) have identified genetic variants that may significantly modulate breast cancer susceptibility. However, the precise molecular mechanisms behind these associations remain largely unknown; often, it is not even clear whether the GWAS variants are functional themselves or just genetically linked to other functional variants. We here provide an integrated method for identifying functional regulatory variants associated with breast cancer and their target genes by combining the analyses of expression quantitative trait loci (eQTL), a modified version of allele-specific expression (ASE) systematically utilizing haplotype information, transcription factor (TF) binding preference, and epigenetic information. Application of our method to the breast cancer susceptibility region in 5p12 demonstrates that the GWAS risk allele rs4415084-T is correlated with higher expression levels of the protein-coding gene MRPS30 and lncRNA RP11-53O19.1. We propose that an intergenic SNP, in linkage disequilibrium (LD) with the GWAS SNP rs4415084, is the predicted functional SNP. We provide multiple levels of evidence that the risk allele of the predicted functional SNP, in phase with the GWAS risk allele rs4415084-T, creates a GATA3 binding motif within a regulatory element, resulting in differential GATA3 binding and chromatin accessibility, which thereby promote the transcription of MRPS30 and RP11-53O19.1. MRPS30 encodes a member of the mitochondrial ribosomal proteins, implicating the risk SNP's role in modulating mitochondrial activities in breast cancer. Our computational framework can be extended to facilitate the rapid functional characterization of other genetic variants modulating cancer susceptibility and provides an effective way of integrating GWAS results with high-throughput genomic and epigenomic data. Citation Format: Yi Zhang, Mohith Manjunath, Shilu Zhang, Deborah Chasman, Sushmita Roy, Jun S. Song. Integrative genomic analysis discovers the causative regulatory mechanisms of a breast cancer-associated genetic variant [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1220.
Genome-wide association studies (GWAS) have hitherto identified several genetic variants associated with cancer susceptibility, but the molecular functions of these risk modulators remain largely uncharacterized. Recent studies have begun to uncover the regulatory potential of non-coding GWAS SNPs by using epigenetic information in corresponding cancer cell types and matched normal tissues. However, this approach does not explore the potential effect of risk germline variants on other important cell types that constitute the microenvironment of tumor or its precursor. This paper presents evidence that the breast cancer-associated variant rs3903072 may regulate the expression of CTSW in tumor infiltrating lymphocytes. CTSW is a candidate tumor-suppressor gene, with expression highly specific to immune cells and also positively correlated with breast cancer patient survival. Integrative analyses suggest a putative causative variant in a GWAS-linked enhancer in lymphocytes that loops to the 3’ end of CTSW through three-dimensional chromatin interaction. Our work thus poses the possibility that a cancer-associated genetic variant might regulate a gene not only in the cell of cancer origin, but also in immune cells in the microenvironment, thereby modulating the immune surveillance by T lymphocytes and natural killer cells and affecting the clearing of early cancer initiating cells.
Abbreviations: GWAS genome-wide association study; eQTL expression quantitative trait loci; ASE allele-specific expression; TF transcription factor; LD linkage disequilibrium; FPKM fragments per kilobase of transcript per million mapped reads; LCASE local chromosome allele-specific expression; DHS DNase I hypersensitive sites; PWM position weight matrix; TCGA The Cancer Genome Atlas; ER+ estrogen receptor positive; TAD topologically associated domain; MAF minor allele frequency; RPKM reads per kilobase of transcript per million mapped reads; SNP single nucleotide polymorphism; MAPQ mapping quality; ChIP-seq chromatin immunoprecipitation sequencing; ASB allele-specific binding; ncRNA non-coding RNA; TSS transcription start site; DNase-seq DNase I
Background Clustering is one of the most common techniques in data analysis and seeks to group together data points that are similar in some measure. Although there are many computer programs available for performing clustering, a single web resource that provides several state-of-the-art clustering methods, interactive visualizations and evaluation of clustering results is lacking. Methods ClusterEnG (acronym for Clustering Engine for Genomics) provides a web interface for clustering data and interactive visualizations including 3D views, data selection and zoom features. Eighteen clustering validation measures are also presented to aid the user in selecting a suitable algorithm for their dataset. ClusterEnG also aims at educating the user about the similarities and differences between various clustering algorithms and provides tutorials that demonstrate potential pitfalls of each algorithm. Conclusions The web resource will be particularly useful to scientists who are not conversant with computing but want to understand the structure of their data in an intuitive manner. The validation measures facilitate the process of choosing a suitable clustering algorithm among the available options. ClusterEnG is part of a bigger project called KnowEnG (Knowledge Engine for Genomics) and is available at http://education.knoweng.org/clustereng.
Summary Clustering is one of the most common techniques used in data analysis to discover hidden structures by grouping together data points that are similar in some measure into clusters. Although there are many programs available for performing clustering, a single web resource that provides both state-of-the-art clustering methods and interactive visualizations is lacking. ClusterEnG (acronym for Clustering Engine for Genomics) provides an interface for clustering big data and interactive visualizations including 3D views, cluster selection and zoom features. ClusterEnG also aims at educating the user about the similarities and differences between various clustering algorithms and provides clustering tutorials that demonstrate potential pitfalls of each algorithm. The web resource will be particularly useful to scientists who are not conversant with computing but want to understand the structure of their data in an intuitive manner. Availability ClusterEnG is part of a bigger project called KnowEnG (Knowledge Engine for Genomics) and is available at http://education.knoweng.org/clustereng . Contact songi@illinois.edu
Plane wave propagation in periodic ordered granular media comprising of elastic spherical particles is investigated. The spheres are under zero precompression and are assumed to interact via the Hertzian contact potential. Various two- and three-dimensional granular structures such as hexagonal packing (2D and 3D), face-centered cubic and body-centered cubic packings are considered in the present study, with the plane impact either normal or oblique to the granular system. For the normal impact case, 1D chains equivalent to the 2D and 3D structures are obtained. A universal relation between the wavefront speed and the force amplitude is derived, valid for all the granular structures studied. In the angular impact case, the shear component of the amplitude of the particle velocity is found to initially decay exponentially and further in a series of linear regimes. By employing simpler models, semi-analytical predictions are obtained for the decay of shearing effect.
Uniform planar impact on a two-dimensional square packing of spheres with intruders at interstitial locations is investigated. An equivalent one-dimensional granular chain model is proposed with appropriate scaling and is verified numerically. Numerical observations demonstrate the existence of a new family of plane solitary waves with different profiles at unique combinations of material properties. In particular, a special case of a solitary wave whose profile is similar to that of the homogeneous chains is also reported. Material combinations that cause solitary waves are systematically extracted for a wide range of material properties. For the solitary wave similar to that of a homogeneous chain, a quasicontinuum approximation is employed to predict the shape and width of the solitary wave, showing good agreement with the numerical results. Finally, an asymptotic analysis is conducted to predict the solitary wave solutions.
Elastic solitary waves resulting from Hertzian contact in one-dimensional (1-D) granular chains have demonstrated promising properties for wave tailoring such as amplitude-dependent wave speed and acoustic band gap zones. However, as load increases, plasticity or other material nonlinearities significantly affect the contact behavior between particles and hence alter the elastic solitary wave formation. This restricts the possible exploitation of solitary wave properties to relatively low load levels (up to a few hundred Newtons). In this work, a method, which we term preconditioning, based on contact pre-yielding is implemented to increase the contact force elastic limit of metallic beads in contact and consequently enhance the ability of 1-D granular chains to sustain high-amplitude elastic solitary waves. Theoretical analyses of single particle deformation and of wave propagation in a 1-D chain under different preconditioning levels are presented, while a complementary experimental setup was developed to demonstrate such behavior in practice. The experimental results show that 1-D granular chains with preconditioned beads can sustain high amplitude (up to several kN peak force) solitary waves. The solitary wave speed is affected by both the wave amplitude and the preconditioning level, while the wave spatial wavelength is still close to 5 times the preconditioned bead size. Comparison between the theoretical and experimental results shows that the current theory can capture the effect of preconditioning level on the solitary wave speed.
Numerical simulations are performed to study wave propagation in an infinite domain of square packed spheres with intruders at interstitial locations. Randomness is uniformly distributed in the densities of the spheres and the packing contains no precompression. The simulations reveal the two regimes of force amplitude decay: exponential and power law, similar to those previously reported in 1D random granular chains. The source of transition between the regimes is identified as a function of the amplitude of the leading pulse and the backscatter. An analysis of the contours of force amplitude demonstrates the anisotropy of wave propagation with respect to the distance traveled by the wave and the level of randomness. An investigation of the evolution of ensemble kinetic energy shows a gradual increase with increasing randomness, and the results are compared with the corresponding 1D random chains. The study of angular decay provides a way to identify suitable sphere properties to design optimized granular structures.
The influence of randomness on wave propagation in one-dimensional chains of spherical granular media is investigated. The interaction between the elastic spheres is modeled using the classical Hertzian contact law. Randomness is introduced in the discrete model using random distributions of particle mass, Young's modulus, or radius. Of particular interest in this study is the quantification of the attenuation in the amplitude of the impulse associated with various levels of randomness: two distinct regimes of decay are observed, characterized by an exponential or a power law, respectively. The responses are normalized to represent a vast array of material parameters and impact conditions. The virial theorem is applied to investigate the transfer from potential to kinetic energy components in the system for different levels of randomness. The level of attenuation in the two decay regimes is compared for the three different sources of randomness and it is found that randomness in radius leads to the maximum rate of decay in the exponential regime of wave propagation.