Meiosis is a fundamental process responsible for sexual reproduction and generating genetic diversity in the progeny. Its successful completion requires fine-tuning of expression programs of many genes: promoting expression of genes involved in meiotic processes and suppressing genes whose expression may interfere with meiosis. Molecular mechanisms involved in meiotic transcriptome regulation and controlling meiosis progression vary between plants, animals, and fungi and remain elusive. We found that the Plural abnormalities of meiosis1 (Pam1) gene in maize controls meiosis progression by tethering transcriptome processing to the meiosis-specific chromosome axis. Pam1 encodes an RNA binding protein that becomes associated with chromosomes during early meiotic prophase I, binds transcripts of a large number of meiosis-related genes, and affects their splicing by interacting with the CCR4-NOT RNA processing protein complex. Disrupting Pam1 function results in a wide array of severe meiosis defects affecting chromosome condensation and dynamics, nuclear envelope and cytoskeleton organization, as well as the overall meiosis progression. Pam1 controls only a subset of meiotic genes and processes, indicating that several programs directing transcriptome architecture collectively regulate meiosis. RNA-binding proteins have been found to control meiosis progression in fungi and animals, and it is now shown to be also the case in plants. Interestingly, these proteins all exhibit distinct modes of action and evolutionary origins, presenting a remarkable case of convergent evolution. Uncovering mechanisms controlling meiosis progression should enable engineering meiosis to benefit crop improvement efforts.
Polyploidy is ubiquitous across North American prairies, which provide essential ecosystem services and rich soil for agriculture. Yet the mechanism driving polyploid abundance is unclear. Multiple hypotheses have been proposed including polyploid abundance is proportional to the opportunity for whole genome duplication (WGD), and WGD alters phenotypes that may increase fitness. We tested these two hypotheses together in the mixed-ploidy species Andropogon gerardi , a dominant grass species in endangered North American tallgrass prairies. Leveraging a novel, phased allopolyploid reference genome, we found the A. gerardi hexaploid arose after the C 4 grassland expansion in the early Pleistocene, when glacial cycles likely increased secondary contact between the diploid progenitors. We sequenced A. gerardi from 25 popula-tions and examined cytotype performance and morphology in a controlled environment to investigate the consequences of the contemporary mixed-ploidy populations. We found the 9 x A. gerardi cytotype is a neopolyploid and a result of recurrent WGD events. Further, we demonstrate the 9 x neopolyploids have greater growth and a decreased stomatal pore index, which is adaptive in xeric climates where the 9 x cy-totype is most common. Together, our results support both hypotheses for polyploid abundance in North America: WGD is a product of opportunity and can have immediate fitness consequences. Although the changes to fitness may provide an advantage to 9 x A. gerardi , the establishment of 9 x may lower overall population fitness due to the lower reproductive viability of 9 x individuals. Polyploid species are abundant in North American prairies and make up many of the dominant species in the ecosystem. This prominence could be a result of whole genome duplication conferring an advantage that increases the frequency of polyploids or could simply indicate that the opportunity for whole genome duplication is higher in this ecosystem, or both. Through examining three polyploidization events in A. gerardi , the dominant species in endangered tallgrass North American prairies, we found whole genome duplication is both surprisingly common and confers traits that are beneficial in some environments.
Meiotic recombination is an important evolutionary process because it can increase the amount of genetic variation within populations through the breakage of unfavorable linkages and creation of novel allelic combinations. Despite the plethora of knowledge about population-level benefits of recombination and numerous theoretical studies examining how recombination rates can evolve over time, there is a lack of empirical evidence for any hypotheses that have been put forward. To alleviate this gap in knowledge, we characterized the evolution of the recombination landscape in Zea mays ssp. mays (maize) during its domestication from Zea mays ssp. parviglumis (teosinte), explored hypotheses that permitted the evolution of the maize recombination landscape and tied these alterations to changes in the genetic basis of recombination. Using experimental populations and the population genomics approach of ancestral recombination graph (ARG) inference, our data demonstrated that maize had a 12% increase in its genome-wide recombination rate during domestication. Although the maize and teosinte recombination landscapes are highly correlated, r = 0.85 at 1Mb resolution, maize has evolved to have higher recombining regions in interstitial chromosome regions, compared to teosinte which only harbors high recombining regions sub-telomerically. Our data show that the re patterning of COs towards interstitial chromosome regions came from reduced CO interference levels within maize. Supporting the idea that CO interference is reduced within maize, we found evidence for selection acting on trans acting recombination-modifiers that participate in the class I CO pathway or CO interference directly. Lastly, we showed that the re-patterning of COs was beneficial to maize evolution because regions that significantly increased in recombination were targeted to gene-rich regions harboring domestication related loci. Because we found regions with significant increases in recombination and a lower deleterious mutation load, compared to regions with decreases in recombination, we concluded that the domestication-related variation in these regions, in which selection acted upon during domestication, was shielded from the Hill-Robertson effect. In conclusion, the re patterning of CO events during domestication allowed maize to adapt and evolve at a faster rate than previously understood. ### Competing Interest Statement The authors have declared no competing interest.
KEY MESSAGE:Two read depth methods were jointly used in next-generation sequencing data to identify deletions in maize population. GWAS by deletions were analyzed for gene expression pattern and classical traits, respectively. Many studies have confirmed that structural variation (SV) is pervasive throughout the maize genome. Deletion is one type of SV that may impact gene expression and cause phenotypic changes in quantitative traits. In this study, two read count approaches were used to analyze the deletions in the whole-genome sequencing data of 270 maize inbred lines. A total of 19,754 deletion windows overlapped 12,751 genes, which were unevenly distributed across the genome. The deletions explained population structure well and correlated with genomic features. The deletion proportion of genes was determined to be negatively correlated with its expression. The detection of gene expression quantitative trait loci (eQTL) indicated that local eQTL were fewer but had larger effects than distant ones. The common associated genes were related to basic metabolic processes, whereas unique associated genes with eQTL played a role in the stress or stimulus responses in multiple tissues. Compared with the eQTL detected by SNPs derived from the same sequencing data, 89.4% of the associated genes could be detected by both markers. The effect of top eQTL detected by SNPs was usually larger than that detected by deletions for the same gene. A genome-wide association study (GWAS) on flowering time and plant height illustrated that only a few loci could be consistently captured by SNPs, suggesting that combining deletion and SNP for GWAS was an excellent strategy to dissect trait architecture. Our findings will provide insights into characteristic and biological function of genome-wide deletions in maize.
Ribosomal repeats occupy 5% of a plant genome, yet there has been little study of their diversity in the modern age of genomics. Ribosomal copy number and expression variation present an opportunity to tap a novel source of diversity. In the present study, we estimated the ribosomal DNA (rDNA) copy number and ribosomal RNA (rRNA) expression for a population of maize inbred lines and investigated the potential role of rDNA and rRNA dosage in regulating global gene expression. Extensive variation was found in both ribosomal DNA copy number and ribosomal RNA expression among maize inbred lines. However, rRNA abundance was not consistent with the copy number of the rDNA. We have not found that the rDNA gene dosage has a regulatory role in gene expression; however, thousands of genes are identified to be coregulated with rRNA expression, including genes participating in ribosome biogenesis and other functionally relevant pathways. We further investigated the potential roles of copy number and the expression level of rDNA on agronomic traits and found that both correlated with flowering time but through different regulatory mechanisms. This comprehensive analysis suggested that rRNA expression variation is a valuable source of functional diversity that affects gene expression variation and field-based phenotypic changes.
The maize W22 inbred has served as a platform for maize genetics since the mid twentieth century. To streamline maize genome analyses, we have sequenced and de novo assembled a W22 reference genome using short-read sequencing technologies. We show that significant structural heterogeneity exists in comparison to the B73 reference genome at multiple scales, from transposon composition and copy number variation to single-nucleotide polymorphisms. The generation of this reference genome enables accurate placement of thousands of Mutator (Mu) and Dissociation (Ds) transposable element insertions for reverse and forward genetics studies. Annotation of the genome has been achieved using RNA-seq analysis, differential nuclease sensitivity profiling and bisulfite sequencing to map open reading frames, open chromatin sites and DNA methylation profiles, respectively. Collectively, the resources developed here integrate W22 as a community reference genome for functional genomics and provide a foundation for the maize pan-genome.
Background Characterization of genetic variations in maize has been challenging, mainly due to deterioration of collinearity between individual genomes in the species. An international consortium of maize research groups combined resources to develop the maize haplotype version 3 (HapMap 3), built from whole genome sequencing data from 1,218 maize lines, covering pre-domestication and domesticated Zea mays varieties across the world. Results A new computational pipeline was set up to process over 12 trillion bp of sequencing data, and a set of population genetics filters were applied to identify over 83 million variant sites. Conclusions We identified polymorphisms in regions where collinearity is largely preserved in the maize species. However, the fact that the B73 genome used as the reference only represents a fraction of all haplotypes is still an important limiting factor.
By 4000 years ago, people had introduced maize to the southwestern United States; full agriculture was established quickly in the lowland deserts but delayed in the temperate highlands for 2000 years. We test if the earliest upland maize was adapted for early flowering, a characteristic of modern temperate maize. We sequenced fifteen 1900-year-old maize cobs from Turkey Pen Shelter in the temperate Southwest. Indirectly validated genomic models predicted that Turkey Pen maize was marginally adapted with respect to flowering, as well as short, tillering, and segregating for yellow kernel color. Temperate adaptation drove modern population differentiation and was selected in situ from ancient standing variation. Validated prediction of polygenic traits improves our understanding of ancient phenotypes and the dynamics of environmental adaptation.
A potential energy surface for the water dimer with explicit dependence on monomer coordinates is presented. The surface was fitted to a set of previously published interaction energies computed on a grid of over a quarter million points in the 12-dimensional configurational space using symmetry-adapted perturbation theory and coupled-cluster methods. The present fit removes small errors in published fits, and its accuracy is critically evaluated. The minimum and saddle-point structures of the potential surface were found to be very close to predictions from direct ab initio optimizations. The computed second virial coefficients agreed well with experimental values. At low temperatures, the effects of monomer flexibility in the virial coefficients were found to be much smaller than the quantum effects.
Background The composition of bacteria in and on the human body varies widely across human individuals, and has been associated with multiple health conditions. While microbial communities are influenced by environmental factors, some degree of genetic influence of the host on the microbiome is also expected. This study is part of an expanding effort to comprehensively profile the interactions between human genetic variation and the composition of this microbial ecosystem on a genome- and microbiome-wide scale. Results Here, we jointly analyze the composition of the human microbiome and host genetic variation. By mining the shotgun metagenomic data from the Human Microbiome Project for host DNA reads, we gathered information on host genetic variation for 93 individuals for whom bacterial abundance data are also available. Using this dataset, we identify significant associations between host genetic variation and microbiome composition in 10 of the 15 body sites tested. These associations are driven by host genetic variation in immunity-related pathways, and are especially enriched in host genes that have been previously associated with microbiome-related complex diseases, such as inflammatory bowel disease and obesity-related disorders. Lastly, we show that host genomic regions associated with the microbiome have high levels of genetic differentiation among human populations, possibly indicating host genomic adaptation to environment-specific microbiomes. Conclusions Our results highlight the role of host genetic variation in shaping the composition of the human microbiome, and provide a starting point toward understanding the complex interaction between human genetics and the microbiome in the context of human evolution and disease.
Gene expression differences between divergent lineages caused by modification of cis regulatory elements are thought to be important in evolution. We assayed genome-wide cis and trans regulatory differences between maize and its wild progenitor, teosinte, using deep RNA sequencing in F1 hybrid and parent inbred lines for three tissue types (ear, leaf and stem). Pervasive regulatory variation was observed with approximately 70% of ∼17,000 genes showing evidence of regulatory divergence between maize and teosinte. However, many fewer genes (1,079 genes) show consistent cis differences with all sampled maize and teosinte lines. For ∼70% of these 1,079 genes, the cis differences are specific to a single tissue. The number of genes with cis regulatory differences is greatest for ear tissue, which underwent a drastic transformation in form during domestication. As expected from the domestication bottleneck, maize possesses less cis regulatory variation than teosinte with this deficit greatest for genes showing maize-teosinte cis regulatory divergence, suggesting selection on cis regulatory differences during domestication. Consistent with selection on cis regulatory elements, genes with cis effects correlated strongly with genes under positive selection during maize domestication and improvement, while genes with trans regulatory effects did not. We observed a directional bias such that genes with cis differences showed higher expression of the maize allele more often than the teosinte allele, suggesting domestication favored up-regulation of gene expression. Finally, this work documents the cis and trans regulatory changes between maize and teosinte in over 17,000 genes for three tissues.
The prolamin-box binding factor1 (pbf1) gene encodes a transcription factor that controls the expression of seed storage protein (zein) genes in maize. Prior studies show that pbf1 underwent selection during maize domestication although how it affected trait change during domestication is unknown. To assay how pbf1 affects phenotypic differences between maize and teosinte, we compared nearly isogenic lines (NILs) that differ for a maize versus teosinte allele of pbf1. Kernel weight for the teosinte NIL (162mg) is slightly but significantly greater than that for the maize NIL (156mg). RNAseq data for developing kernels show that the teosinte allele of pbf1 is expressed at about twice the level of the maize allele. However, RNA and protein assays showed no difference in zein profile between the two NILs. The lower expression for the maize pbf1 allele suggests that selection may have favored this change; however, how reduced pbf1 expression alters phenotype remains unknown. One possibility is that pbf1 regulates genes other than zeins and thereby is a domestication trait. The observed drop in seed weight associated with the maize allele of pbf1 is counterintuitive but could represent a negative pleiotropic effect of selection on some other aspect of kernel composition.
In flowering plants, mitochondrial and chloroplast mRNAs are edited by C-to-U base modification. In plant organelles, RNA editing appears to be generally a correcting mechanism that restores the proper function of the encoded product. Members of the Arabidopsis RNA editing-Interacting Protein (RIP) family have been recently shown to be essential components of the plant editing machinery. We report the use of a strand- and transcript-specific RNA-seq method (STS-PCRseq) to explore the effect of mutation or silencing of every RIP gene on plant organelle editing. We confirm RIP1 to be a major editing factor that controls the editing extent of 75% of the mitochondrial sites and 20% of the plastid C targets of editing. The quantitative nature of RNA sequencing allows the precise determination of overlapping effects of RIP factors on RNA editing. Over 85% of the sites under the influence of RIP3 and RIP8, two moderately important mitochondrial factors, are also controlled by RIP1. Previously uncharacterized RIP family members were found to have only a slight effect on RNA editing. The preferential location of editing sites controlled by RIP7 on some transcripts suggests an RNA metabolism function for this factor other than editing. In addition to a complete characterization of the RIP factors for their effect on RNA editing, our study highlights the potential of RNA-seq for studying plant organelle editing. Unlike previous attempts to use RNA-seq to analyze RNA editing extent, our methodology focuses on sequencing of organelle cDNAs corresponding to known transcripts. As a result, the depth of coverage of each editing site reaches unprecedented values, assuring a reliable measurement of editing extent and the detection of numerous new sites. This strategy can be applied to the study of RNA editing in any organism.
This study compared the relative public health impact in deli meats at retail contaminated with Listeria monocytogenes by either (i) other products or (ii) the retail environment. Modeling was performed using the risk of listeriosis-associated deaths as a public health outcome of interest and using two deli meat products (i.e., ham and turkey, both formulated without growth inhibitors) as model systems. Based on reported data, deli meats coming to retail were assumed to be contaminated at a frequency of 0.4%. Three contamination scenarios were investigated: (i) a baseline scenario, in which no additional cross-contamination occurred at retail, (ii) a scenario in which an additional 2.3% of products were cross-contaminated at retail due to transfer of L. monocytogenes cells from already contaminated ready-to-eat deli meats, and (iii) a scenario in which an additional 2.3% of products were contaminated as a result of cross-contamination from a contaminated retail environment. By using a previously reported L. monocytogenes risk assessment model that uses product-specific growth kinetic parameters, cross-contamination of deli ham and turkey was estimated to increase the relative risk of listeriosis-associated deaths by 5.9- and 6.1-fold, respectively, for contamination from other products and by 4.9- and 5.8-fold, respectively, for contamination from the retail environment. Sensitivity and scenario analyses indicated that the frequency of cross-contamination at retail from any source (other food products or environment) was the most important factor affecting the relative risk of listeriosis-associated deaths. Overall, our data indicate that retail-level cross-contamination of ready-to-eat deli meats with L. monocytogenes has the potential to considerably increase the risk of human listeriosis cases and deaths, and thus precise estimates of cross-contamination frequency are critical for accurate risk assessments.
New DNA sequencing technologies present an exceptional opportunity for novel and creative applications with the potential for breakthrough discoveries. To support such research efforts, the Cornell University Life Sciences Core Laboratories Center has implemented the Illumina HiSeq 2000 and the Roche 454 GS FLX platforms as academic core facility shared research resources. We have established sample handling methods, LIMS tools and BioHPC informatics analysis pipelines in support of these new technologies. Our genomics core laboratory, in collaboration with our epigenomics core and bioinformatics core, provides sample preparation and data generation services and both project consultation and analysis support for a wide range of possible applications, including de novo or reference based genome assembly, detection of genetic variation, transcriptome sequencing, small RNA profiling, and genome-wide epigenomic measurements of methylation and protein-nucleic acid interactions. Implementation of next generation sequencing platforms as shared resources with multidisciplinary core facility support enables cost effective access and broad based use of these technologies.
RP-10 One of the challenges of High Performance Computing (HPC) in biology is resource accessibility.Computational Biology Applications Suite for High Performance Computing (BioHPC) is addressing this challenge at the user level as well as at the research group or administrator/developer level. BioHPC has been developed at the Computational Biology Service Unit (CBSU) of the Cornell University Life Sciences Core Laboratories Center (CLC) under the Microsoft HPC Institute program. This is an application suite that: (a) allows researchers from biological laboratories to submit their jobs to a parallel cluster through an easy-to-use web interface; (b) facilitates analysis of large data sets on a parallel platform; (c) allows biology research groups to set up Microsoft CCS and/or HPC 2008 based clusters for computational biology applications and allows efficient management via the BioHPC interface. The suite documentation and code are accessible at http://BioHPC.net. The suite can be downloaded and installed locally on any Microsoft based cluster.
CF-21 New DNA sequencing technologies presents an exceptional opportunity for novel and creative applications with the potential for breakthrough discoveries. To support such research efforts, the Cornell University Life Sciences Core Laboratories Center has implemented the Illumina Solexa Genome Analyzer IIx and the Roche 454 Genome Sequencer FLX platforms as academic core facility shared research resources. We have established sample handling methods, wikiLIMS tools and informatics analysis pipelines in support of these new technologies. Our DNA sequencing and genotyping core laboratory provides sample preparation and data generation services and in collaboration with the microarrays and informatics core facilities, provides both project consultation and analysis support for a wide range of possible applications, including de novo or reference based genome assembly, detection of genetic variation, transcriptome sequencing, small RNA profiling, and genome-wide epigenomic measurements of protein-nucleic interactions. Implementation of next generation sequencing platforms as shared resources with multi-disciplinary core facility support enables cost effective access and broad based use of these technologies.
CF-3 The Computational Biology Service Unit (CBSU) of the Cornell University Life Sciences Core Laboratories Center (CLC) provides bioinformatics and computational biology resources and services to the university community and to outside investigators. The CBSU is Microsoft Biology Initiative partner and charter member of Microsoft High Performance Computing Institute program. Facility resources include a set of servers and compute clusters with over 1200 cores. The goal of the facility is to meet the increasing need of Cornell investigators for broad support in computational biology and bioinformatics.
The objective of this study was to estimate the relative risk of listeriosis-associated deaths attributable to Listeria monocytogenes contamination in ham and turkey formulated without and with growth inhibitors (GIs). Two contamination scenarios were investigated: (i) prepackaged deli meats with contamination originating solely from manufacture at a frequency of 0.4% (based on reported data) and (ii) retail-sliced deli meats with contamination originating solely from retail at a frequency of 2.3% (based on reported data). Using a manufacture-to-consumption risk assessment with product-specific growth kinetic parameters (i.e., lag phase and exponential growth rate), reformulation with GIs was estimated to reduce human listeriosis deaths linked to ham and turkey by 2.8- and 9-fold, respectively, when contamination originated at manufacture and by 1.9- and 2.8-fold, respectively, for products contaminated at retail. Contamination originating at retail was estimated to account for 76 and 63% of listeriosis deaths caused by ham and turkey, respectively, when all products were formulated without GIs and for 83 and 84% of listeriosis deaths caused by ham and turkey, respectively, when all products were formulated with GIs. Sensitivity analyses indicated that storage temperature was the most important factor affecting the estimation of per annum relative risk.,Scenario analyses suggested that reducing storage temperature in home refrigerators to consistently below 7 degrees C would greatly reduce the risk of human listeriosis deaths, whereas reducing storage time appeared to be less effective. Overall, our data indicate a critical need for further development and implementation of effective control strategies to reduce L. monocytogenes contamination at the retail level.