Hatchery supplementation is vital for conserving dwindling fish populations. Effective augmentation requires distinguishing hatchery-origin from wild individuals and accurately identifying species, particularly in systems where closely related species coexist. Genetic monitoring is key to quantifying genetic differences, but conventional markers do not distinguish hybrids, especially backcrosses. Misidentifying hybrids in hatchery programs compromises wild gene pools because hatchery broodstock contributes to numerous offspring being released into the wild. Here, we present a workflow for developing and evaluating the Genotyping-in-Thousands by sequencing (GT-seq) single nucleotide polymorphism (SNP) panel for North American river sturgeons (Scaphirhynchus spp.). This panel is designed to detect complex hybrid classes and to determine parent-offspring relationships. Our species identification panel (S-loci) contains 155 SNPs selected for high genetic differentiation (FST) between Pallid Sturgeon (S. albus) and Shovelnose Sturgeon (S. platorynchus), and the parentage assignment panel (P-loci) includes 112 SNPs with high heterozygosity within Pallid Sturgeon. Simulation analyses demonstrated that our GT-seq S-loci panel reliably classifies pure species, F1, F2 and backcross hybrids, even with up to 70% missing data. The P-loci panel achieves high-confidence parentage assignment with ≥ 80% typed loci, with performance influenced by the proportion of sampled parents. Overall, the novel Scaphirhynchus GT-seq panel developed in this study represents a robust and efficient tool for detecting hybridisation, assigning parentage and providing critical information for management decisions in ongoing Pallid Sturgeon conservation.
ABSTRACT The concept of assisted evolution suggests that hurdles to species restoration in modified environments might be overcome by selective breeding of individuals better‐suited to the contemporary evolutionary landscape. Here, we introduce an experimental program aimed at developing an Atlantic salmon broodstock resilient to dietary thiamine deficiency complex (TDC). Dietary TDC is caused by overconsumption of lipid‐rich prey items containing thiaminase, leading to insufficient maternal allocation of thiamine to eggs and high offspring mortality. Within Lake Champlain, USA, where TDC is widespread due to consumption of invasive Alewife, we conducted offspring survival trials of 114 wild‐reared salmon pairs. We documented among‐family patterns in offspring survival to yolk‐sac absorption. Some females allocated sufficient free thiamine to eggs to avoid severe offspring mortality from TDC, while most did not and suffered complete or severe reproductive failure. Offspring from “resilient” families with high survival were used as founders of a putatively “TDC‐Tolerant” broodstock but at the expense of low effective population size (Ne = 48). We also developed a control broodstock that included additional families that only survived after supplemental thiamine treatment (MAX‐Gen, no selection, Ne = 152) to represent a status quo scenario. From 2016 to 2025 we (1) developed a Genotyping‐in‐Thousands by Sequencing (GT‐seq) marker panel (2) implemented parentage‐based tagging (3) curated two experimental broodstocks and released ~1.6 million of their offspring into Lake Champlain and (4) implemented a sampling program to monitor demographic and evolutionary outcomes. In Lake Champlain, TDC continues to be a major restoration bottleneck afflicting 42%–76% of salmon evaluated annually since 2007 and is an increasing conservation issue globally. Continuation of this work will determine whether assisted evolution improved TDC‐resistance in the wild and also evaluate the genetic and demographic consequences of contrasting broodstock strategies.
Biologists use long-term monitoring of wildlife populations to investigate how populations change over time or in relation to extrinsic pressures. Recently, amplicon-based sequencing approaches with high-quality samples have been used to study or monitor wildlife populations in lieu of traditional microsatellite methodologies. However, the application of these techniques with lower-quality DNA sources has had mixed results and risks the loss of interoperability between microsatellites and genomic data in long-term datasets. We sought to optimize a microhaplotype panel for use with non-invasively collected wolf fecal samples, maintaining the ability to identify and match individuals across two datasets and draw familial inferences. We conducted 16 experiments to investigate pre-treatment of sample extracts, PCR1 cleaning methods, combining PCR2 with normalization using "Nate's Plates", and incubating normalization plates overnight. Additionally, we developed a quantitative PCR (qPCR) assay to quantify wolf nuclear DNA in fecal samples. We quantified 2283 fecal sample extractions on qPCR assays and ultimately genotyped 790 unique fecal samples. We retained 472 unique samples within our final dataset with an average capture rate per individual of 1.63. Finally, we successfully reconnected the two datasets by matching 66 individuals identified through scat sampling with harvest tissue genotypes. We found that noninvasive samples should have a minimum concentration of 0.2 ng/uL accompanied by efforts to reduce primer artifacts. As next generation sequencing technologies become increasingly applied in wildlife populations including using low-quality samples as we have demonstrated, special care will be required to maintain interoperability across datasets.
Abstract GT-seq (Genotyping-in-Thousands by Sequencing) is widely used for high-throughput amplicon genotyping, but most analytical pipelines focus on single SNPs or rely on alignment-based variant calling. Here we present an alignment-free approach for microhaplotype genotyping that leverages the high read depth and low error rates typical of paired-end Illumina and Element sequencing. The pipeline first identifies primer-bounded reads and resolves paired-end sequences into complete amplicon sequences. Within each sample and locus, unique sequences are ranked by read abundance and the top one or two sequences are retained as candidate diploid alleles. These alleles are aggregated across samples to construct a catalog of unique haplotypes for each locus. In a second pass, reads are assigned to catalog haplotypes by exact sequence matching to produce diploid genotypes. Finally, catalog haplotype sequences are positionally compared to identify phased SNP and collapsed indel variation, generating compact microhaplotype representations suitable for population genetic analysis. This approach enables robust, alignment-free microhaplotype inference directly from high-depth amplicon sequencing data.
Population size estimation is critical for determining sustainable harvest levels of hunted species. Spatial capture-recapture via non-invasive genetic sampling is one method employed to estimate population size and is especially useful in habitats where aerial surveys are not possible. Developments in DNA sequencing technologies have led to the application of amplicon-based genotyping-in-thousands by sequencing (GTseq) for genotyping tissue and non-invasive genetic samples. Here we developed and assessed the power of a GTseq single nucleotide polymorphism (SNP) panel for individual identification. We performed population structure and diversity analysis of Sitka black-tailed deer (Odocoileus hemionus sitkensis) tissue samples from native and transplanted populations in Alaska and British Columbia. We then demonstrated that we could perform individual matching on fecal pellet samples from central Southeast Alaska (CSEAK). Our panel contains 248 SNPs and has enough power for individual identification with as few as 30 loci across all sampled populations. We identified population structure aligned with island geography and translocation histories. To select panel SNPs, we sampled our initial genomic data from one island in CSEAK, resulting in ascertainment bias when applying the panel outside of CSEAK. We highlight that this ascertainment bias led to lower locus variability in populations other than CSEAK. For individual identification outside of CSEAK, we recommend subsetting the panel to only use loci that are variable for the population being studied. If intending to estimate patterns of diversity and differentiation across populations outside of CSEAK, we recommend re-designing the panel with additional range wide genomic data.
Minimally invasive samples are often the best option for collecting genetic material from species of conservation concern, but they perform poorly in many genomic sequencing methods due to their tendency to yield low DNA quality and quantity. Genotyping-in-thousands by sequencing (GT-seq) is a powerful amplicon sequencing method that can genotype large numbers of variable-quality samples at a standardized set of single nucleotide polymorphism (SNP) loci. Here, we develop, optimize, and validate a GT-seq panel for the federally threatened northern Idaho ground squirrel (Urocitellus brunneus) to provide a standardized approach for future genetic monitoring and assessment of recovery goals using minimally invasive samples. The optimized panel consists of 224 neutral and 81 putatively adaptive SNPs. DNA collected from buccal swabs from 2016 to 2020 had 73% genotyping success, while samples collected from hair from 2002 to 2006 had little to no DNA remaining and did not genotype successfully. We evaluated our GT-seq panel by measuring genotype discordance rates compared to RADseq and whole-genome sequencing. GT-seq and other sequencing methods had similar population diversity and F-ST estimates, but GT-seq consistently called more heterozygotes than expected, resulting in negative F-IS values at the population level. Genetic ancestry assignment was consistent when estimated with different sequencing methods and numbers of loci. Our GT-seq panel is an effective and efficient genotyping tool that will aid in the monitoring and recovery of this threatened species, and our results provide insights for applying GT-seq for minimally invasive DNA sampling techniques in other rare animals.
Chromosomal inversion polymorphisms have special importance in theAnopheles gambiaecomplex of malaria vector mosquitoes, due to their role in local adaptation and range expansion. The study of inversions in natural populations is reliant on polytene chromosome analysis by expert cytogeneticists, a process that is limited by the rarity of trained specialists, low throughput, and restrictive sampling requirements. To overcome this barrier, we ascertained tag single nucleotide polymorphisms (SNPs) that are highly correlated with inversion status (inverted or standard orientation). We compared the performance of the tag SNPs using two alternative high throughput molecular genotyping approachesvs.traditional cytogenetic karyotyping of the same 960 individualAn. gambiaeandAn. coluzziimosquitoes sampled from Burkina Faso, West Africa. We show that both molecular approaches yield comparable results, and that either one performs as well or better than cytogenetics in terms of genotyping accuracy. Given the ability of molecular genotyping approaches to be conducted at scale and at relatively low cost without restriction on mosquito sex or developmental stage, molecular genotyping via tag SNPs has the potential to revitalize research into the role of chromosomal inversions in the behavior and ongoing adaptation ofAn. gambiaeandAn. coluzziito environmental heterogeneities.
Polymorphic chromosomal inversions have been implicated in local adaptation. In anopheline mosquitoes, inversions also contribute to epidemiologically relevant phenotypes such as resting behavior. Progress in understanding these phenotypes and their mechanistic basis has been hindered because the only available method for inversion genotyping relies on traditional cytogenetic karyotyping, a rate-limiting and technically difficult approach that is possible only for the fraction of the adult female population at the correct gonotrophic stage. Here, we focus on an understudied malaria vector of major importance in sub-Saharan Africa, Anopheles funestus. We ascertain and validate tag single nucleotide polymorphisms (SNPs) using high throughput molecular assays that allow rapid inversion genotyping of the three most common An. funestus inversions at scale, overcoming the cytogenetic karyotyping barrier. These same inversions are the only available markers for distinguishing two An. funestus ecotypes that differ in indoor resting behavior, Folonzo and Kiribina. Our new inversion genotyping tools will facilitate studies of ecotypic differentiation in An. funestus and provide a means to improve our understanding of the roles of Folonzo and Kiribina in malaria transmission.
Minimally invasive sampling (MIS) is widespread in wildlife studies; however, its utility for massively parallel DNA sequencing (MPS) is limited. Poor sample quality and contamination by exogenous DNA can make MIS challenging to use with modern genotyping-by-sequencing approaches, which have been traditionally developed for high-quality DNA sources. Given that MIS is often more appropriate in many contexts, there is a need to make such samples practical for harnessing MPS. Here, we test the ability for Genotyping-in-Thousands by sequencing (GT-seq), a multiplex amplicon sequencing approach, to effectively genotype minimally invasive cloacal DNA samples collected from the Western Rattlesnake (Crotalus oreganus), a threatened species in British Columbia, Canada. As there was no previous genetic information for this species, an optimized panel of 362 SNPs was selected for use with GT-seq from a de novo restriction site-associated DNA sequencing (RADseq) assembly. Comparisons of genotypes generated within and among RADseq and GT-seq for the same individuals found low rates of genotyping error (GT-seq: 0.50%; RADseq: 0.80%) and discordance (2.57%), the latter likely due to the different genotype calling models employed. GT-seq mean genotype discordance between blood and cloacal swab samples collected from the same individuals was also minimal (1.37%). Estimates of population diversity parameters were similar across GT-seq and RADseq data sets, as were inferred patterns of population structure. Overall, GT-seq can be effectively applied to low-quality DNA samples, minimizing the inefficiencies presented by exogenous DNA typically found in minimally invasive samples and continuing the expansion of molecular ecology and conservation genetics in the genomics era.
Abstract Coho salmon were extirpated in the mid‐20th century from the interior reaches of the Columbia River but were reintroduced with relatively abundant source stocks from the lower Columbia River near the Pacific coast. Reintroduction of Coho salmon to the interior Columbia River (Wenatchee River) using lower river stocks placed selective pressures on the new colonizers due to substantial differences with their original habitat such as migration distance and navigation of six additional hydropower dams. We used restriction site‐associated DNA sequencing (RAD‐seq) to genotype 5,392 SNPs in reintroduced Coho salmon in the Wenatchee River over four generations to test for signals of temporal structure and adaptive variation. Temporal genetic structure among the three broodlines of reintroduced fish was evident among the initial return years (2000, 2001, and 2002) and their descendants, which indicated levels of reproductive isolation among broodlines. Signals of adaptive variation were detected from multiple outlier tests and identified candidate genes for further study. This study illustrated that genetic variation and structure of reintroduced populations are likely to reflect source stocks for multiple generations but may shift over time once established in nature.
GT-seq is a genotyping method that leverages large read numbers from Illumina sequencers to genotype hundreds of single nucleotide polymorphisms within pools of multiplex PCR amplicons generated from thousands of individual samples (Campbell et al., 2014). This method produces genotypes that are 99.9% concordant to those produced using TaqMan assays at approximately one-fourth the cost. Since its development, GT-seq panels have been created for several species (Chinook salmon, Coho salmon, sockeye salmon, rainbow trout, and pacific lamprey) and have become the preferred SNP genotyping method in our laboratory. New genotyping software allows genotypes and summary figures to be produced from a lane of raw sequencing data in less than an hour, using a desktop Linux computer.
Next-generation sequencing data can be mined for highly informative single nucleotide polymorphisms (SNPs) to develop high-throughput genomic assays for nonmodel organisms. However, choosing a set of SNPs to address a variety of objectives can be difficult because SNPs are often not equally informative. We developed an optimal combination of 96 high-throughput SNP assays from a total of 4439 SNPs identified in a previous study of Pacific lamprey (Entosphenus tridentatus) and used them to address four disparate objectives: parentage analysis, species identification and characterization of neutral and adaptive variation. Nine of these SNPs are F-ST outliers, and five of these outliers are localized within genes and significantly associated with geography, run-timing and dwarf life history. Two of the 96 SNPs were diagnostic for two other lamprey species that were morphologically indistinguishable at early larval stages and were sympatric in the Pacific Northwest. The majority (85) of SNPs in the panel were highly informative for parentage analysis, that is, putatively neutral with high minor allele frequency across the species' range. Results from three case studies are presented to demonstrate the broad utility of this panel of SNP markers in this species. As Pacific lamprey populations are undergoing rapid decline, these SNPs provide an important resource to address critical uncertainties associated with the conservation and recovery of this imperiled species.
Twelve eulachon (Thaleichthys pacificus, Osmeridae) populations ranging from Cook Inlet, Alaska and along the west coast of North America to the Columbia River were examined by restriction-site-associated DNA (RAD) sequencing to elucidate patterns of neutral and adaptive variation in this high geneflow species. A total of 4104 single-nucleotide polymorphisms (SNPs) were discovered across the genome, with 193 putatively adaptive SNPs as determined by F(ST) outlier tests. Estimates of population structure in eulachon with the putatively adaptive SNPs were similar, but provided greater resolution of stocks compared with a putatively neutral panel of 3911 SNPs or previous estimates with 14 microsatellites. A cline of increasing measures of genetic diversity from south to north was found in the adaptive panel, but not in the neutral markers (SNPs or microsatellites). This may indicate divergent selective pressures in differing freshwater and marine environments between regional eulachon populations and that these adaptive diversity patterns not seen with neutral markers could be a consideration when determining genetic boundaries for conservation purposes. Estimates of effective population size (N(e)) were similar with the neutral SNP panel and microsatellites and may be utilized to monitor population status for eulachon where census sizes are difficult to obtain. Greater differentiation with the panel of putatively adaptive SNPs provided higher individual assignment accuracy compared to the neutral panel or microsatellites for stock identification purposes. This study presents the first SNPs that have been developed for eulachon, and analyses with these markers highlighted the importance of integrating genome-wide neutral and adaptive genetic variation for the applications of conservation and management.
As ectothermic organisms have evolved to differing aquatic climates, the molecular basis of thermal adaptation is a key area of research. In this study, we tested for differential transcriptional response of ecologically divergent populations of redband trout (Oncorhynchus mykiss gairdneri) that have evolved in desert and montane climates. Each pure strain and their F1 cross were reared in a common garden environment and exposed over four weeks to diel water temperatures that were similar to those experienced in desert climates within the species’ range. Gill tissues were collected from the three strains of fish (desert, montane, F1 crosses) at the peak of heat stress and tested for mRNA expression differences across the transcriptome with RNA-seq.
Recent advances in genotyping-by-sequencing have enabled genome-wide association studies in nonmodel species including those in aquaculture programs. As with other aquaculture species, rainbow trout and steelhead (Oncorhynchus mykiss) are susceptible to disease and outbreaks can lead to significant losses. Fish culturists have therefore been pursuing strategies to prevent losses to common pathogens such as Flavobacterium psychrophilum (the etiological agent for bacterial cold water disease [CWD]) and infectious hematopoietic necrosis virus (IHNV) by adjusting feed formulations, vaccine development, and selective breeding. However, discovery of genetic markers linked to disease resistance offers the potential to use marker-assisted selection to increase resistance and reduce outbreaks. For this study we sampled juvenile fish from 40 families from 2-yr classes that either survived or died after controlled exposure to either CWD or IHNV. Restriction site−associated DNA sequencing produced 4661 polymorphic single-nucleotide polymorphism loci after strict filtering. Genotypes from individual survivors and mortalities were then used to test for association between disease resistance and genotype at each locus using the program TASSEL. After we accounted for kinship and stratification of the samples, tests revealed 12 single-nucleotide polymorphism markers that were highly associated with resistance to CWD and 19 markers associated with resistance to IHNV. These markers are candidates for further investigation and are expected to be useful for marker assisted selection in future broodstock selection for various aquaculture programs.
Genotyping‐in‐Thousands by sequencing (GT‐seq) is a method that uses next‐generation sequencing of multiplexed PCR products to generate genotypes from relatively small panels (50–500) of targeted single‐nucleotide polymorphisms (SNPs) for thousands of individuals in a single Illumina HiSeq lane. This method uses only unlabelled oligos and PCR master mix in two thermal cycling steps for amplification of targeted SNP loci. During this process, sequencing adapters and dual barcode sequence tags are incorporated into the amplicons enabling thousands of individuals to be pooled into a single sequencing library. Post sequencing, reads from individual samples are split into individual files using their unique combination of barcode sequences. Genotyping is performed with a simple perl script which counts amplicon‐specific sequences for each allele, and allele ratios are used to determine the genotypes. We demonstrate this technique by genotyping 2068 individual steelhead trout (Oncorhynchus mykiss) samples with a set of 192 SNP markers in a single library sequenced in a single Illumina HiSeq lane. Genotype data were 99.9% concordant to previously collected TaqMan™ genotypes at the same 192 loci, but call rates were slightly lower with GT‐seq (96.4%) relative to Taqman (99.0%). Of the 192 SNPs, 187 were genotyped in ≥90% of the individual samples and only 3 SNPs were genotyped in <70% of samples. This study demonstrates amplicon sequencing with GT‐seq greatly reduces the cost of genotyping hundreds of targeted SNPs relative to existing methods by utilizing a simple library preparation method and massive efficiency of scale.
To elucidate the mechanisms of thermal adaptation and acclimation in ectothermic aquatic organisms from differing climates, we used a common‐garden experiment for thermal stress to investigate the heat shock response of redband trout ( O ncorhynchus mykiss gairdneri ) from desert and montane populations. Evidence for adaptation was observed as expression of heat shock genes in fish from the desert population was more similar to control (unstressed) fish and significantly different ( P ≤ 0.05) from those from the montane population, while F1 crosses were intermediate. High induction of heat shock proteins (Hsps) in the montane strain appeared to improve short‐term survival during first exposure to high water temperatures, but high physiological costs of Hsp production may have led to lower long‐term survival. In contrast, the desert strain had significantly lower heat shock response than the montane fish and F1 crosses, suggesting that these desert fish have evolved alternative mechanisms to deal with thermal stress that provide better balance of physiological costs. Genomewide tests of greater than 10 000 SNPs found multiple SNP s that were significantly associated with survival under thermal stress, including Hsp47 which consistently appeared as a strong candidate gene for adaption to desert climates. Candidate SNP s identified in this study are prime targets to screen more broadly across this species' range to predict the potential for adaptation under scenarios of climate change. These results demonstrate that aquatic species can evolve adaptive responses to thermal stress and provide insight for understanding how climate change may impact ectotherms.
Little is known of the genetic basis of migration despite the ecological benefits migratory species provide to their communities and their rapid global decline due to anthropogenic disturbances in recent years. Using next-generation sequencing of restriction-site-associated DNA (RAD) tags, we genotyped thousands of single nucleotide polymorphisms (SNPs) in two wild populations of migratory steelhead and resident rainbow trout (Oncorhynchus mykiss) from the Pacific Northwest of the United States. One population maintains a connection to the sea, whereas the other population has been sequestered from its access to the ocean for more than 50 years by a hydropower dam. Here we performed a genome-wide association study to identify 504 RAD SNP markers from several genetic regions that were associated with the propensity to migrate both within and between the populations. Our results corroborate those in previous quantitative trait loci studies and provide evidence for additional loci associated with this complex migratory life history. Our results suggest a complex multi-genic basis with several loci of small effect distributed throughout the genome contributing to migration in this species. We also determined that despite being sequestered for decades, the landlocked population continues to harbour genetic variation associated with a migratory life history and ATPase activity. Furthermore, we demonstrate the utility of genotyping-by-sequencing and how RAD-tag SNP data can be readily compared between studies to investigate migration within this species.
Parentage-based tagging (PBT) is a promising alternative to traditional coded-wire tag (CWT) methodologies for monitoring and evaluating hatchery stocks. This approach involves the genotyping of hatchery broodstock and uses parentage assignments to identify the origin and brood year of their progeny. In this study we empirically confirmed that fewer than 100 single nucleotide polymorphisms (SNPs) were needed to accurately conduct PBT, we demonstrated that our selected panel of SNPs was comparable in accuracy to a panel of microsatellites, and we verified that stock assignments made with this panel matched those made using CWTs. We also demonstrated that when sampling of spawners was incomplete, an estimated PBT rate for the offspring could also be predicted with fewer than 100 SNPs. This study in the Snake River basin is one of the first large-scale implementations of PBT in salmonids and lays the foundation for adopting this technology more broadly in the region, thereby allowing the unprecedented ability to mark millions of smolts and an opportunity to address a variety of parentage-based research and management questions.
Unlike most anadromous fishes that have evolved strict homing behaviour, Pacific lamprey (Entosphenus tridentatus) seem to lack philopatry as evidenced by minimal population structure across the species range. Yet unexplained findings of within‐region population genetic heterogeneity coupled with the morphological and behavioural diversity described for the species suggest that adaptive genetic variation underlying fitness traits may be responsible. We employed restriction site–associated DNA sequencing to genotype 4439 quality filtered single nucleotide polymorphism (SNP) loci for 518 individuals collected across a broad geographical area including British Columbia, Washington, Oregon and California. A subset of putatively neutral markers (N = 4068) identified a significant amount of variation among three broad populations: northern British Columbia, Columbia River/southern coast and ‘dwarf’ adults (FCT = 0.02, P ≪ 0.001). Additionally, 162 SNPs were identified as adaptive through outlier tests, and inclusion of these markers revealed a signal of adaptive variation related to geography and life history. The majority of the 162 adaptive SNPs were not independent and formed four groups of linked loci. Analyses with matsam software found that 42 of these outlier SNPs were significantly associated with geography, run timing and dwarf life history, and 27 of these 42 SNPs aligned with known genes or highly conserved genomic regions using the genome browser available for sea lamprey. This study provides both neutral and adaptive context for observed genetic divergence among collections and thus reconciles previous findings of population genetic heterogeneity within a species that displays extensive gene flow.