Grape (Vitis spp.) is an economically and culturally significant crop grown in a wide array of climates, including cooler areas that regularly experience freezing temperatures. To better adapt grapes for cultivation in cooler climates, wild grape relatives and hybrids have been and continue to be used in breeding efforts. The US Department of Agriculture, Agricultural Research Service maintains a collection of cultivated and wild cold-hardy grapes in Geneva, NY, USA. This collection contains more than one dozen species, mostly of North American origin, as well as an extensive set of hybrid breeding lines and cultivars. We demonstrate the genetic variation present in the collection using newly developed rhAmpSeq markers to explore phylogenetic relationships. Our findings match those of previous analyses that showed Eurasian species nested within the North American species, suggesting a North American origin of the Vitis genus. In addition, an analysis of ancestry and genetic distance suggested taxonomic identities of 18 previously unidentified accessions and 36 putatively misidentified accessions. The data presented here advance the understanding of the Vitis clade and provide support for ongoing research, conservation, and breeding efforts.
The genus Vitis is composed of two subgenera, Vitis (2n = 38) and Muscadinia (2n = 40), which are both cultivated for fresh-market, juice, and wine industries. Stenospermocarpic seedless and perfect-flowered vines are highly desired in both Vitis vinifera and muscadine (Muscadinia rotundifolia) breeding programs. Stenospermocarpy has recently been introgressed from V. vinifera to muscadines through conventional breeding despite their differing chromosome number, but no molecular markers for this trait have been validated in muscadine germplasm. Wild vines in both subgenera are dioecious, but perfect-flowered forms were independently selected during domestication to enhance reproductive efficiency and fruit production. This study reports the development and validation of two Kompetitive allele-specific PCR (KASP) markers targeting causal polymorphisms within candidate genes for male sterility (VviINP1) and stenospermocarpy (VviAGL11). Sequence alignments with published M. rotundifolia genomes suggested that the seedless_Arg197Leu_site56 and female_INP_indel_site56 KASP markers might be broadly effective across diverse species within Vitis and Muscadinia. Marker performance was evaluated using a validation panel including 918 Vitis × Muscadinia hybrid seedlings from the University of Arkansas Fruit Breeding Program and a diverse set of cultivated and wild accessions (209 accessions evaluated with the seedless marker and 315 accessions with the flower sex marker). After excluding incomplete phenotype and genotype data, the stenospermocarpic marker (seedless_Arg197Leu_site56) accurately predicted seedlessness in 921 of 924 (99.7%) entries. Additionally, 148 of 203 seedlings that failed to produce fruit across both growing seasons were predicted to be stenospermocarpic with the seedless KASP marker, suggesting that seedless Vitis × Muscadinia hybrids may have partial sterility or lack cold hardiness. The flower sex marker (female_INP_indel_site56) correctly predicted flower sex in all 1138 (100%) entries. Together, these KASP markers provide highly accurate and cost-effective tools for early selection of seedless and perfect-flowered genotypes across Vitis, Muscadinia, and hybrid breeding programs.
Staphylococcus pseudintermedius is a common representative of the normal skin microbiota of dogs and cats but is also a causative agent of a variety of infections. Although primarily a canine/feline bacterium, recent studies suggest an expanded host range including humans. This paper details population genomic analyses of the largest yet assembled and sequenced collection of S. pseudintermedius isolates from across the USA and Canada and assesses these isolates within a larger global population genetic context. We then employ a pan-genome-wide association study analysis of over 1,700 S. pseudintermedius isolates from sick dogs and cats, covering the period 2017-2020, correlating loci at a genome-wide level, with in vitro susceptibility data for 23 different antibiotics. We find no evidence from either core genome phylogenies or accessory genome content for separate lineages colonizing cats or dogs. Some core genome geographic clustering was evident on a global scale, and accessory gene content was noticeably different between various regions, some of which could be linked to known antimicrobial resistance (AMR) loci for certain classes of antibiotics (e.g., aminoglycosides). Analysis of genes correlated with AMR was divided into different categories, depending on whether they were known resistance mechanisms, on a plasmid, or a putatively novel resistance mechanism on the chromosome. We discuss several novel chromosomal candidates for follow-up laboratory experimentation, including, for example, a bacteriocin (subtilosin), for which the same protein from Bacillus subtilis has been shown to be active against Staphylococcus aureus infections, and for which the operon, present in closely related Staphylococcus species, is absent in S. aureus.IMPORTANCEStaphylococcus pseudintermedius is an important causative agent of a variety of canine and feline infections, with recent studies suggesting an expanded host range, including humans. This paper presents global population genomic data and analysis of the largest set yet sequenced for this organism, covering the USA and Canada as well as more globally. It also presents analysis of in vitro antibiotic susceptibility testing results for the North American (NA) isolates, as well as genetic analysis for the global set. We conduct a pan-genome-wide association study analysis of over 1,700 S. pseudintermedius isolates from sick dogs and cats from NA to correlate loci at a genome-wide level with the in vitro susceptibility data for 23 different antibiotics. We discuss several chromosomal loci arising from this analysis for follow-up laboratory experimentation. This study should provide insight regarding the development of novel molecular treatments for an organism of both veterinary and, increasingly, human medical concern.
ABSTRACT Background Fever of unknown origin (FUO) without a respiratory component is a frequent clinical presentation in horses. Multiple pathogens, both tick‐borne and enteric, can be involved as etiologic agents. An additional potential mechanism is intestinal barrier dysfunction. Objectives This case–control study aimed to detect and associate microbial taxa in blood with disease state. Study Design Areas known for a high prevalence of tick‐borne diseases in humans were chosen to survey horses with FUO, which was defined as fever of 101.5°F or higher with no signs of respiratory illness or other recognisable diseases. Blood samples and clinical parameters were obtained from 52 FUO cases and also from matched controls from the same farms. An additional 23 febrile horses without matched controls were included. Methods Broadly targeted polymerase chain reaction (PCR) amplification directed at conserved sequence regions of bacterial 16S rRNA, parasite 18S rRNA, coronavirus RdRp and parvovirus NS1 was performed, followed by deep sequencing. To control for contamination and identify taxa unique to the cases, metagenomic sequences from the controls were subtracted from those of the cases, and additional targeted molecular testing was performed. Sera were also tested for antibodies to equine coronavirus. Results Over 60% of cases had intestinal microbial DNA circulating in the blood. Nineteen percent of cases were attributed to infection with Anaplasma phagocytophilum, of which two were subtyped as human‐associated strains. A novel Erythroparvovirus was detected in two cases and two controls. Serum titres for equine coronavirus were elevated in some cases but not statistically different overall between the cases and controls. Main Limitations Not all pathogens are expected to circulate in blood, which was the sole focus of this study. Conclusions The presence of commensal gut microbes in blood of equine FUO cases is consistent with a compromised intestinal barrier, which is highlighted as a direction for future study.
Equine Erythroparvovirus 1 is a parvovirus that was identified in the blood of four horses in the United States. Here, we report one genome from a horse in New York State. This genome may represent a new species within the genus Erythroparvovirus.
Wild Malus species harbor untapped genetic diversity to advance apple breeding, particularly for disease resistance and stress tolerance. However, existing marker panels, developed mainly using Malus domestica accessions, introduce ascertainment bias and limit detecting rare variants in wild species. We developed and validated a medium-density and cost-effective pan-generic 3 K apple DArTag panel optimized to capture genome-wide variation across the Malus genus. The panel was constructed using conserved, syntenic, and collinear genomic blocks identified within the core genome of 13 Malus accessions for cross-species transferability. The panel was validated across three bi-parental mapping populations totaling 593 progeny. Across these populations, 2461–3234 SNP markers were polymorphic and 1482–2620 were informative. Each population contained over 900 multiallelic micro-haplotype loci, with several hundred loci exhibiting three or four distinct haplotypes. Markers were uniformly distributed across all 17 chromosomes, each containing between 60 and 230 informative SNPs. The panel was further evaluated on 174 diverse germplasm accessions from 20 Malus species. It exhibited strong cross-species transferability, exceptionally low rates of missing data (< 0.5
Grapevine downy mildew, caused by the oomycete pathogen Plasmopara viticola, can lead to economically significant losses in humid climates. An ever-growing catalog of loci for resistance to P. viticola is available to breeders, including Rpv3 on the lower arm of chromosome 18. Widely used in French-American cultivars, Rpv3 is a complex TIR-NBS-LRR locus for which associated SSR markers have provided evidence of multiple alleles with varying degrees of resistance. However, SSRs lack the resolution to detect nuances between alleles and fully characterize the locus. PacBio long-read sequencing enables phased assembly of highly repetitive gene cluster regions, allowing for high resolution comparison among predicted alleles. Rpv3 haplotypes of eight Vitis genomes (‘Catawba’, ‘Chambourcin’, ‘Concord’, ‘Horizon’, MN1264, ‘Norton’, NY84.0101.03, and PN40024) were compared to identify differences in gene structure among Rpv3.1, Rpv3.2, Rpv3.3, co-located Rpv27, and susceptible SSR haplotypes. This region was extracted from each haplotype as delimited by their flanking SSR markers and ranged in length from 0.8 to 1.7 Mb. While there was strong consistency in gene structure within Rpv3.1, both Rpv3.2 and Rpv3.3 showed evidence of divergence between haplotypes. In addition to local alignments, the candidate genes identified in Rpv3.1 were tested for copy number and functional variation across haplotypes. The candidate gene region ranged in length from 85.3–220 kb from SSR UDV737. While the structure of this region and the TNL gene sequences were largely consistent within Rpv3.1 and Rpv3.2, the structure was variable among Rpv3.3 individuals. In spite of this structural variation among the Rpv3.3 SSR haplotype, their TNL gene sequences had strong similarity to Rpv27. The variation in gene structure shown in this study underscores the need for the refinement of allele naming, high-quality genome assemblies, as well as more informative, higher resolution marker systems for marker-assisted selection to improve resistance to grapevine downy mildew.
Infections caused by antimicrobial-resistant Escherichia coli are the leading cause of death attributed to antimicrobial resistance (AMR) worldwide, and the known AMR mechanisms involve a range of functional proteins. Here, we employed a pan-genome wide association study (GWAS) approach on over 1,000 E. coli isolates from sick dogs collected across the US and Canada and identified a strong statistical association (empirical P < 0.01) of AMR, involving a range of antibiotics to a group 1 capsular (CPS) gene cluster. This cluster included genes under relaxed selection pressure, had several loci missing, and had pseudogenes for other key loci. Furthermore, this cluster is widespread in E. coli and Klebsiella clinical isolates across multiple host species. Earlier studies demonstrated that the octameric CPS polysaccharide export protein Wza can transmit macrolide antibiotics into the E. coli periplasm. We suggest that the CPS in question, and its highly divergent Wza, functions as an antibiotic trap, preventing antimicrobial penetration. We also highlight the high diversity of lineages circulating in dogs across all regions studied, the overlap with human lineages, and regional prevalence of resistance to multiple antimicrobial classes.
Arabidopsis (Arabidopsis thaliana) ecotype Col-0 has plastid and mitochondrial genomes encoding over 100 proteins. Public databases (e.g. Araport11) have redundancy and discrepancies in gene identifiers for these organelle-encoded proteins. RNA editing results in changes to specific amino acid residues or creation of start and stop codons for many of these proteins, but the impact of RNA editing at the protein level is largely unexplored due to the complexities of detection. Here, we assembled the nonredundant set of identifiers, their correct protein sequences, and 452 predicted nonsynonymous editing sites of which 56 are edited at lower frequency. We then determined accumulation of edited and/or unedited proteoforms by searching ∼259 million raw tandem MS spectra from ProteomeXchange, which is part of PeptideAtlas (www.peptideatlas.org/builds/arabidopsis/). We identified all mitochondrial proteins and all except 3 plastid-encoded proteins (NdhG/Ndh6, PsbM, and Rps16), but no proteins predicted from the 4 ORFs were identified. We suggest that Rps16 and 3 of the ORFs are pseudogenes. Detection frequencies for each edit site and type of edit (e.g. S to L/F) were determined at the protein level, cross-referenced against the metadata (e.g. tissue), and evaluated for technical detection challenges. We detected 167 predicted edit sites at the proteome level. Minor frequency sites were edited at low frequency at the protein level except for cytochrome C biogenesis 382 at residue 124 (Ccb382-124). Major frequency sites (>50% editing of RNA) only accumulated in edited form (>98% to 100% edited) at the protein level, with the exception of Rpl5-22. We conclude that RNA editing for major editing sites is required for stable protein accumulation.
Grape (Vitis) production and fruit quality traits such as cluster size, berry shape, and timing of fruit development are key aspects when selecting cultivars for commercial production. Molecular markers for some, but not all, of these traits have been identified using biparental or association mapping populations. Previously identified markers were tested for transferability using a small (24 individual) test panel of commercially available grape cultivars. Markers had little to no ability to differentiate grape phenotypes based on the expected characteristics, except the marker for seedlessness. Using a biparental interspecific cross, 43 quantitative trait loci (QTLs) (previously identified and new genomic regions) associated with berry shape, number, size, cluster weight, cluster length, time to flower, veraison, and full color were detected. Kompetitive allele-specific polymerase chain reaction markers designed on newly identified QTLs were tested for transferability using the same panel. Transferability was low when use types were combined, but they were varied when use types were evaluated separately. A comparison of a 4-Mb region at the end of chromosome 18 revealed structural differences among grape species and use types. Table grape cultivars had the highest similarity in structure for this region (>75%) compared with other grape species and commodity types.
This study presents the Maize PeptideAtlas resource (www.peptideatlas.org/builds/maize) to help solve questions about the maize proteome. Publicly available raw tandem mass spectrometry (MS/MS) data for maize collected from ProteomeXchange were reanalyzed through a uniform processing and metadata annotation pipeline. These data are from a wide range of genetic backgrounds and many sample types and experimental conditions. The protein search space included different maize genome annotations for the B73 inbred line from MaizeGDB, UniProtKB, NCBI RefSeq, and for the W22 inbred line. 445 million MS/MS spectra were searched, of which 120 million were matched to 0.37 million distinct peptides. Peptides were matched to 66.2% of proteins in the most recent B73 nuclear genome annotation. Furthermore, most conserved plastid- and mitochondrial-encoded proteins (NCBI RefSeq annotations) were identified. Peptides and proteins identified in the other B73 genome annotations will improve maize genome annotation. We also illustrate the high-confidence detection of unique W22 proteins. N-terminal acetylation, phosphorylation, ubiquitination, and three lysine acylations (K-acetyl, K-malonyl, and K-hydroxyisobutyryl) were identified and can be inspected through a PTM viewer in PeptideAtlas. All matched MS/MS-derived peptide data are linked to spectral, technical, and biological metadata. This new PeptideAtlas is integrated in MaizeGDB with a peptide track in JBrowse.
Meiotic recombination is an important evolutionary process because it can increase the amount of genetic variation within populations through the breakage of unfavorable linkages and creation of novel allelic combinations. Despite the plethora of knowledge about population-level benefits of recombination and numerous theoretical studies examining how recombination rates can evolve over time, there is a lack of empirical evidence for any hypotheses that have been put forward. To alleviate this gap in knowledge, we characterized the evolution of the recombination landscape in Zea mays ssp. mays (maize) during its domestication from Zea mays ssp. parviglumis (teosinte), explored hypotheses that permitted the evolution of the maize recombination landscape and tied these alterations to changes in the genetic basis of recombination. Using experimental populations and the population genomics approach of ancestral recombination graph (ARG) inference, our data demonstrated that maize had a 12% increase in its genome-wide recombination rate during domestication. Although the maize and teosinte recombination landscapes are highly correlated, r = 0.85 at 1Mb resolution, maize has evolved to have higher recombining regions in interstitial chromosome regions, compared to teosinte which only harbors high recombining regions sub-telomerically. Our data show that the re patterning of COs towards interstitial chromosome regions came from reduced CO interference levels within maize. Supporting the idea that CO interference is reduced within maize, we found evidence for selection acting on trans acting recombination-modifiers that participate in the class I CO pathway or CO interference directly. Lastly, we showed that the re-patterning of COs was beneficial to maize evolution because regions that significantly increased in recombination were targeted to gene-rich regions harboring domestication related loci. Because we found regions with significant increases in recombination and a lower deleterious mutation load, compared to regions with decreases in recombination, we concluded that the domestication-related variation in these regions, in which selection acted upon during domestication, was shielded from the Hill-Robertson effect. In conclusion, the re patterning of CO events during domestication allowed maize to adapt and evolve at a faster rate than previously understood. ### Competing Interest Statement The authors have declared no competing interest.
Despite increasing threats of extinction to Elasmobranchii (sharks and rays), whole genome-based conservation insights are lacking. Here, we present chromosome-level genome assemblies for the Critically Endangered great hammerhead (Sphyrna mokarran) and the Endangered shortfin mako (Isurus oxyrinchus) sharks, with genetic diversity and historical demographic comparisons to other shark species. The great hammerhead exhibited low genetic variation, with 8.7% of the 2.77 Gbp genome in runs of homozygosity (ROH) > 1 Mbp and 74.4% in ROH >100 kbp. The 4.98 Gbp shortfin mako genome had considerably greater diversity and <1% in ROH > 1 Mbp. Both these sharks experienced precipitous declines in effective population size (Ne) over the last 250 thousand years. While shortfin mako exhibited a large historical Ne that may have enabled the retention of higher genetic variation, the genomic data suggest a possibly more concerning picture for the great hammerhead, and a need for evaluation with additional individuals.
Fine mapping of quantitative trait loci (QTL) to dissect the genetic basis of traits of interest is essential to modern breeding practice. Here, we employed a multitiered haplotypic marker system to increase fine mapping accuracy by constructing a chromosome-level, haplotype-resolved parental genome, accurate detection of recombination sites, and allele-specific characterization of the transcriptome. In the first tier of this system, we applied the preexisting panel of 2,000 rhAmpSeq core genome markers that is transferable across the entire Vitis genus and provides a genomic resolution of 200 kb to 1 Mb. The second tier consisted of high-density haplotypic markers generated from Illumina skim sequencing data for samples enriched for relevant recombinations, increasing the potential resolution to hundreds of base pairs. We used this approach to dissect a novel Resistance to Plasmopara viticola-33 (RPV33) locus conferring resistance to grapevine downy mildew, narrowing the candidate region to only 0.46 Mb. In the third tier, we used allele-specific RNA-seq analysis to identify a cluster of 3 putative disease resistance RPP13-like protein 2 genes located tandemly in a nonsyntenic insertion as candidates for the disease resistance trait. In addition, combining the rhAmpSeq core genome haplotype markers and skim sequencing-derived high-density haplotype markers enabled chromosomal-level scaffolding and phasing of the grape Vitis × doaniana 'PI 588149' assembly, initially built solely from Pacific Biosciences (PacBio) high-fidelity (HiFi) reads, leading to the correction of 16 large-scale phasing errors. Our mapping strategy integrates high-density, phased genetic information with individual reference genomes to pinpoint the genetic basis of QTLs and will likely be widely adopted in highly heterozygous species.
This study describes a new release of the Arabidopsis thaliana PeptideAtlas proteomics resource providing protein sequence coverage, matched mass spectrometry (MS) spectra, selected PTMs, and metadata. 70 million MS/MS spectra were matched to the Araport11 annotation, identifying ∼0.6 million unique peptides and 18267 proteins at the highest confidence level and 3396 lower confidence proteins, together representing 78.6% of the predicted proteome. Additional identified proteins not predicted in Araport11 should be considered for building the next Arabidopsis genome annotation. This release identified 5198 phosphorylated proteins, 668 ubiquitinated proteins, 3050 N-terminally acetylated proteins and 864 lysine-acetylated proteins and mapped their PTM sites. MS support was lacking for 21.4% (5896 proteins) of the predicted Araport11 proteome – the ‘dark’ proteome. This dark proteome is highly enriched for certain (e.g. CLE, CEP, IDA, PSY) but not other (e.g. THIONIN, CAP,) signaling peptides families, E3 ligases, TFs, and other proteins with unfavorable physicochemical properties. A machine learning model trained on RNA expression data and protein properties predicts the probability for proteins to be detected. The model aids in discovery of proteins with short-half life (e.g. SIG1,3 and ERF-VII TFs) and completing the proteome. PeptideAtlas is linked to TAIR, JBrowse, PPDB, SUBA, UniProtKB and Plant PTM Viewer.
Powdery mildew resistance genes restrict infection attempts at different stages of pathogenesis. Here, a strong and rapid powdery mildew resistance phenotype was discovered from Vitis amurensis 'PI 588631' that rapidly stopped over 97% of Erysiphe necator conidia, before or immediately after emergence of a secondary hypha from appressoria. This resistance was effective across multiple years of vineyard evaluation on leaves, stems, rachises, and fruit and against a diverse array of E. necator laboratory isolates. Using core genome rhAmpSeq markers, resistance mapped to a single dominant locus (here named REN12) on chromosome 13 near 22.8-27.0 Mb, irrespective of tissue type, explaining up to 86.9% of the phenotypic variation observed on leaves. Shotgun sequencing of recombinant vines using skim-seq technology enabled the locus to be further resolved to a 780 kb region, from 25.15 to 25.93 Mb. RNASeq analysis indicated the allele-specific expression of four resistance genes (NLRs) from the resistant parent. REN12 is one of the strongest powdery mildew resistance loci in grapevine yet documented, and the rhAmpSeq sequences presented here can be directly used for marker-assisted selection or converted to other genotyping platforms. While no virulent isolates were identified among the genetically diverse isolates and wild populations of E. necator tested here, NLR loci like REN12 are often race-specific. Thus, stacking of multiple resistance genes and minimal use of fungicides should enhance the durability of resistance and could enable a 90% reduction in fungicides in low-rainfall climates where few other pathogens attack the foliage or fruit.
ABSTRACTWe developed the Maize PeptideAtlas resource (www.peptideatlas.org/builds/maize) to help solve questions about the maize proteome. Publicly available raw tandem mass spectrometry (MS/MS) data for maize were collected from ProteomeXchange and reanalyzed through a uniform processing and metadata annotation pipeline. These data are from a wide range of genetic backgrounds, including the inbred lines B73 and W22, many hybrids and their respective parents. Samples were collected from field trials, controlled environmental conditions, a range of (a)biotic conditions and different tissues, cell types and subcellular fractions. The protein search space included different maize genome annotations for the B73 inbred line from MaizeGDB, UniProtKB, NCBI RefSeq and for the W22 inbred line. 445 million MS/MS spectra were searched, of which 120 million were matched to 0.37 million distinct peptides. Peptides were matched to 66.2% of the proteins (one isoform per protein coding gene) in the most recent B73 nuclear genome annotation (v5). Furthermore, most conserved plastid- and mitochondrial-encoded proteins (NCBI RefSeq annotations) were identified. Peptides and proteins identified in the other searched B73 genome annotations will aid to improve maize genome annotation. We also illustrate high confidence detection of unique W22 proteins. N-terminal acetylation, phosphorylation, ubiquitination, and three lysine acylations (K-acetyl, K-malonyl, K-hydroxyisobutyryl) were identified and can be inspected through a PTM viewer in PeptideAtlas. All matched MS/MS-derived peptide data are linked to spectral, technical and biological metadata. This new PeptideAtlas is integrated with community resources including MaizeGDB athttps://www.maizegdb.org/and a peptide track in JBrowse.