
Whole genome amplification (WGA), and in particular multiple displacement amplification (MDA), has become a key technique for genomic sequencing of microscopic organisms, yet it introduces artefacts such as palindromic (inverted chimeric) reads that may compromise downstream analyses. We assessed how pervasive palindromic reads generated by MDA impact the assembly of tardigrade (Acutuncus giovanniniae and A. mecnuffi) mitogenomes sequenced with Oxford Nanopore technology. We show that the MDA produces a high proportion of palindromic reads, often exceeding one-third of mitochondrial reads and frequently exhibiting complex multi-inversion structures. These artefacts severely impair long-read assembly, leading to low success rates and inconsistent genome reconstruction. To solve this issue, a strategy based on in silico fragmentation of long reads into short, high-quality fragments, followed by short-read assembly, consistently produced complete and accurate circularised mitochondrial genomes. Our results demonstrate that palindromic read formation can be, in some cases, a limitation of MDA coupled with long-read sequencing, but this issue can be mitigated through read fragmentation. This approach provides a simple, robust and scalable solution for mitogenome assembly from data heavily affected by amplification artefacts, particularly in microscopic taxa where whole genome amplification is often unavoidable.
The choice of optimality criterion is a key consideration in phylogenetic studies. Recent work challenges the notion that more computationally demanding optimisation objectives result in better phylogenetic trees. This finding underscores the importance of comparing trees across optimisation objectives in addition to different models of evolution and data partitions. It is currently cumbersome to optimise trees for alternative objectives because multiple programmes must be used, each with its own optimisation framework. Here, I introduce Treeline for optimising balanced minimum evolution, maximum likelihood and maximum parsimony trees. Treeline explores the optimisation landscape using a new strategy based on perturbing the patristic distance matrix used to initialize candidate trees, which is shown to be particularly effective for the balanced minimum evolution objective. Tests suggest Treeline can be more accurate, memory efficient or faster than existing programmes designed for a single optimisation objective. Treeline's unified nature facilitates comparison of phylogenetic trees across distinct optimisation objectives and models of sequence evolution. Consistent with prior studies, the balanced minimum evolution objective resulted in gene trees that were more consistent with species trees than maximum likelihood or maximum parsimony objectives. With a case study of species from the family Hominidae, including Homo sapiens, I show how Treeline simplifies the process of constructing trees under alternative optimisation objectives and models of evolution. Treeline is part of the DECIPHER package for R and is available from Bioconductor and online (https://DECIPHER.codes/).
Given the high levels of endemism, diversity, and increasing anthropogenic impacts in tropical regions, studies on species interactions and ecological networks are urgently needed to understand community responses to environmental change. Increasingly, molecular methods are used to monitor biodiversity and identify species interactions. In this study, we developed an optimised DNA metabarcoding protocol to construct quantitative dung beetle-vertebrate trophic networks in a heterogenous tropical forest. We show that while the highest probability of detection is within 3 h of feeding, vertebrate DNA can still be detected in most beetles after 24 h, including around 50% of beetles with visually empty guts. Using the vertebrate taxa detected from DNA in the gut contents, we constructed dung beetle-vertebrate networks across 50 forest sites. Our networks displayed high trophic generalism and nestedness. By using group-specific primers, we documented interactions between dung beetles and multiple vertebrate species, including amphibians, birds, and rare mammals, such as the Sunda slow loris and Sunda pangolin. Our optimised molecular techniques contribute towards the development of new methods for comprehensive biodiversity assessments and ecological network studies that are urgently needed to monitor biodiversity in the hyper-diverse tropics.
Metabarcoding sequence data from environmental DNA (eDNA) is rapidly expanding as a powerful method for biodiversity surveys. In order to interpret these data, tools are needed that account for the uncertainty associated with eDNA sampling, sequencing and analysis. The data resulting from eDNA marker gene analysis differ from many traditional methods of biodiversity surveys because they are highly complex, sparse and compositional. Methodological biases produce uncertainty at every step of the sampling and sequencing process. Thus, it is critical that users have a way of interpreting eDNA results that accounts for their compositional nature and models the uncertainty resulting from factors like patchy sampling, PCR amplification biases and variable sequencing depth. Here, we introduce MAMBO: Metabarcoding Analysis using Modeled Bayesian Occurrences. MAMBO simulates in silico replication and models the uncertainty surrounding the sequencing and analysis process. Further, it uses these modelled sequence count data to correlate two sets of marker genes with a Bayesian regression, facilitating the linkage of different groups targeted by these assays. Compared with correlational network analyses, MAMBO overcomes many of the limitations to robust statistical analyses of eDNA marker gene data and provides an opportunity for new insight into ecological patterns over space and time.
Analysing large population genomic datasets requires an interdisciplinary skillset. Beyond a knowledge base in genetics and population biology, population genomic analyses involve computer science and statistics, representing a barrier for researchers without experience in those fields. PopGenHelpR seeks to lower this barrier by enabling researchers to perform population genomic analyses and generate near-publication quality figures in a streamlined and informed fashion. PopGenHelpR allows users to estimate genetic diversity within populations as well as differentiation among populations and individuals from single nucleotide polymorphism data. PopGenHelpR includes commonly used measures such as observed heterozygosity and FST. PopGenHelpR also provides five previously unavailable measures of heterozygosity, including the proportion of heterozygous loci and homozygosity by locus. Additionally, PopGenHelpR integrates widely used visualization tools in population genomics that normally require additional software packages, such as ancestry bar charts, piechart maps, and genetic differentiation heatmaps. Moreover, PopGenHelpR is minimally dependent on other R packages, reducing its chance of being removed from public networks. The PopGenHelpR website also provides tutorials and educational resources. Altogether, PopGenHelpR provides resources for informed analyses and effective visualization, making PopGenHelpR a valuable tool for many researchers. PopGenHelpR is available on the Comprehensive R Archive Network and GitHub (https://github.com/kfarleigh/PopGenHelpR).
Just over two decades ago, DNA-based dietary analysis promised to advance the resolution, sensitivity, and speed with which we could detect and identify trophic interactions. Since then, these approaches have generated a paradigm shift in our understanding across a wide range of natural systems. Although decreasing sequencing costs and increased access to sequencing technologies have significantly broadened adoption in recent years, advances in the methods used for dietary analysis have arguably slowed. We stand now, however, at the brink of the next advance, as traditionally DNA-based ecological studies increasingly apply RNA-based methods. To date, this has most commonly taken place in the context of environmental monitoring, and the application to dietary analysis is still underrepresented despite immense potential and relatively straightforward implementation. Given the reduced stability of some RNA types, the detection of consumed resources via RNA alongside DNA can mitigate many longstanding methodological pitfalls of DNA-based dietary analyses alone by (i) differentiating between living and dead resources, (ii) identifying potential false positives, and (iii) providing temporal context to detections, facilitating the construction of weighted or multilayer trophic networks. Detection of functional RNA may also present an opportunity to ascribe functional contexts to both consumers and resources, and contextualise interactions or their impact on wider trophic networks. With the increasing accessibility of RNA-based methods, their application to community and network ecology may significantly advance our ability to analyse and understand trophic interactions in complex natural systems. Here, we summarise RNA-based dietary analysis methods with reference to recent literature in trophic ecology and identify areas that would benefit from additional research. By summarising the immediately identifiable advances and constraints that dietary RNA presents, we hope to stimulate widespread adoption of these approaches and advance integration of RNA-based analyses into trophic ecological research.
DNA metabarcoding is becoming an increasingly common approach in ecological monitoring of marine and freshwater planktonic communities, yet methodological choices along the metabarcoding workflow and data post-processing approaches remain highly inconsistent across studies, limiting the ability to track biodiversity trends, detect range shifts, or integrate datasets across monitoring programs. This study addresses this methodological bottleneck by combining controlled experimental comparisons with a comprehensive literature synthesis to identify how protocol decisions-from sample preservation and DNA extraction to sequencing platforms and taxonomic assignments-affect the results of COI metabarcoding and its interpretation. Overall biodiversity and community patterns were recovered by all combinations of tested methods, supporting the notion that patterns identified through DNA metabarcoding are robust and comparable across studies. We identify TES (Tris-EDTA-SDS) buffer, optionally paired with at-sea homogenisation, as a practical alternative to ethanol preservation for large-scale monitoring surveys. We show that integrating several classification methods and reference databases for taxonomic assignment improves diversity estimates and confidence in the assignments, and advocate for increased use of tools like BOLDigger that facilitate manual curation of ambiguous/erroneous references. Finally, we demonstrate that introducing stricter filtering thresholds reduces the effect of false positives, pseudogenes and lab-specific contamination, and make comparisons of data generated by different laboratories and methodological configurations more robust, although potentially at the expense of excluding rare taxa. While we intentionally refrain from recommending a universal best practices protocol, this study aims to provide a practical roadmap to help enhance the reliability and reproducibility of marine zooplankton monitoring via DNA metabarcoding.
Advancements in historical genomics increasingly leverage museum collections to study past ecosystems, species interactions and biodiversity. Formalin-fixed, ethanol-preserved specimens, once thought inaccessible to molecular analyses due to DNA degradation, are emerging as valuable genomic resources. If recoverable and reliably attributable, DNA within preservation media could provide a non-destructive alternative to conventional tissue sampling, with the potential to expand molecular access to valuable or irreplaceable specimens. We tested whether preservation media contains recoverable DNA suitable for taxonomic inference. We coupled passive adsorption and active filtration of specimen media with hot alkaline lysis DNA extraction followed by metabarcoding and shotgun metagenomics. DNA was recoverable across samples, including 41 of 61 (~67%) targets in a composite sample. However, detections were dominated by non-target taxa, indicating that preservation media retain a layered mixture of specimen-derived DNA and broader collection-level background. Detection success tracked with preservation chemistry (near-neutral pH and low residual formaldehyde) rather than specimen age. Method choice influenced detections: active filtration increased target detections but admitted more background; passive capture was sparser but more selective; shotgun sequencing retrieved broader vertebrate signals, including reptiles, but was heavily enriched for non-targets. Because both target and non-target taxa were often abundant, read-abundance cut-offs were unreliable for attribution. Spirit-media DNA is therefore best interpreted as a collection-level signal and a screening tool to identify jars with molecular potential (e.g., taxa of conservation or biosecurity interest), rather than as a definitive non-destructive proxy for specimen identity. Prioritising chemically favourable jars and implementing rigorous contamination controls should improve signal interpretability and help unlock the value of preservation media for historical genomics.
In-line barcoding offers a streamlined and scalable alternative to two-step PCR library preparation for 16S rRNA gene amplicon sequencing, enabling cost-effective, high-throughput profiling of microbial communities. Here, we tested 136 and 156 in-line barcoded primer pairs for bacterial and archaeal communities for their performance across environmental samples and a mock standard community. The primers were designed by combining widely used universal 16S rRNA gene primers with existing barcode sets from Illumina kits. The designed primer pairs produced efficient and consistent amplification with minimal dropout and no systematic taxonomic bias. Through clustering and performance-based filtering, we selected final sets of 96 pairs for both bacterial and archaeal communities that work efficiently and well together for direct further use. This in-line tagging strategy is easy to adopt with fewer processing steps and PCR-associated artefacts, allows straightforward sample tracking, and supports reliable large-scale microbiome studies. We also present a framework for evaluating barcode- or primer-induced biases. More broadly, the proposed in-line barcoding strategy can be adapted to any amplicon-sequencing application, as well as targeted sequencing, highlighting its relevance beyond 16S rRNA gene surveys. All validation datasets, open-source processing scripts, and barcode design resources are provided to promote reproducibility and community-wide adoption.
Effective monitoring of hybrid zones is essential for understanding evolutionary dynamics and mitigating species loss caused by human-mediated hybridisation. Conventional methods rely on sampling numerous individuals, which is costly, time-consuming and often impractical for rare or elusive species. Environmental DNA (eDNA) offers a promising alternative for locating hybrid zones but requires the detection of nuclear eDNA, which is typically scarce in natural ecosystems. While a few recent studies have successfully recovered sufficient nuclear eDNA to assess intraspecific variation, its application in hybridisation studies remains untested. This study provides the first empirical validation that nuclear eDNA can screen for hybrid populations. We present an eDNA-based toolkit that employs Kompetitive Allele-Specific PCR (KASP) to genotype a panel of species-diagnostic unlinked nuclear SNPs without sequencing, mapping end-point fluorescence to a hybrid index that reflects ancestry levels at the population scale. In mesocosms housing different combinations of individuals from two crested newt species (Triturus ivanbureschi and T. macedonicus) and their captive-bred F1 hybrids, we compared eDNA-derived ancestry estimates with genotypes obtained from skin swabs of the same individuals placed in the mesocosms. eDNA-based ancestry estimates showed strong concordance with individual genotypes across two eDNA sampling concentrations. This approach represents a promising non-invasive, fast and cost-efficient screening tool, qualities that make it well suited to locate and track putative hybrid zones and a scalable complement to conventional sampling for biodiversity monitoring and conservation.
Metabarcoding of faecal samples is a powerful, non-invasive approach for investigating the feeding ecology of carnivores, revealing prey diversity and unexpected dietary components with greater resolution than traditional methods. However, the approach remains technically demanding, as challenges and potential biases arise at every stage, from scat collection and DNA extraction to primer selection, sequencing, and data interpretation. Methodological details for these steps are often scattered across studies, limiting reproducibility and accessibility for ecologists. Here, we present a comprehensive field-to-sequencer workflow for dietary metabarcoding of terrestrial carnivores using Oxford Nanopore Technologies (ONT), covering all stages from sample collection to ecological interpretation. Drawing on field-collected scats of brown (Parahyaena brunnea) and spotted hyenas (Crocuta crocuta) across arid and semi-arid savannas in Botswana, we illustrate practical decisions, technical considerations, and common pitfalls encountered throughout the process. By integrating field, laboratory, and bioinformatic components into a single, accessible framework, this paper provides a pragmatic reference for ecologists aiming to design robust, transparent, and comparable studies of carnivore diet composition.
Differential gene expression (DGE) analysis enables researchers to investigate the link between gene expression and the phenotypic responses observed in organisms across time, experimental, or field conditions. Accurate quantification of gene expression is essential when performing DGE experiments, with a range of methods having been developed to enable the study of gene expression within a species. Quantifying differences in expression not just within but across multiple species can also be used to reveal the genetic mechanisms underlying phenotypic differences observed between species. Accurate quantification of gene expression across multiple species requires a suitable reference; it should include each species' own expressed transcripts to mitigate reference bias, with the orthology relationships of transcripts being used to facilitate comparison of expression at the gene level. Production of such a reference remains a challenge, despite its necessity for minimising bias during multispecies DGE analysis. Our software BINge specifically aims to address this need through use of a novel approach to modelling orthology which results in multispecies transcript clusters that accurately reflect their locus orthology. Evaluation experiments demonstrate the effectiveness of this approach over existing clustering methods which have not been designed for producing a reference suitable for multispecies DGE analysis. Source code and documentation for BINge are available from the GitHub repository at https://github.com/zkstewart/BINge.
Spiders are renowned for their ecological versatility and silk-based innovations in materials science, yet marine environments remain virtually uncolonized by this predominantly terrestrial lineage. A striking exception is the obligate intertidal spider genus Desis, whose members have evolved extraordinary physiological and behavioural adaptations to persist in wave-swept, saline habitats that oscillate between land and sea. However, the molecular basis of these adaptations has remained largely unexplored. Here, we present a high-quality, chromosome-scale genome of the intertidal spider Desis jiaxiangi, together with a reference genome of the water spider Argyroneta aquatica, integrated with transcriptomic and proteomic data. This multi-omics framework reveals the genomic architecture underlying adaptation to life at the ocean's edge. We uncover expansions of gene families linked to hormone biosynthesis and DNA repair, alongside signatures of adaptive evolution in genes involved osmoregulation, the rate-limiting step of glycolysis, mitochondrial regulation, epithelial tube morphogenesis and circadian rhythm. Notably, we characterize a novel silk spidroin enriched with a unique GVGAKV motif, which may enhance silk hydrophobicity, and detect the duplication burst of hemocyanin genes likely supporting oxygen transport during submersion. Together, these findings reveal convergent molecular strategies for coping with extreme and fluctuating environments, and demonstrate how genomic innovation enables terrestrial lineages to invade marine-influenced ecosystems. Our study establishes Desis as powerful model for understanding adaptation at terrestrial-marine interface.
Wild relatives of domestic animals are crucial reservoirs of genetic diversity, yet pervasive hybridization with domestic animals poses significant conservation challenges. Here, we developed a deep learning-based pipeline, consisting of a multi-layer perceptron for SNP panel selection and a Deep & Cross Network for model training, to discern wild relatives from their closely related domestic animals using genomic SNP data. Leveraging the 1960 genomes from 164 red jungle fowl (RJF; Gallus gallus) and 1796 domestic chicken samples, we applied this pipeline to yield the RJF identification model based on a 285-SNP panel. We employed this model to characterize domestic chickens, RJF, and hybrids in the independent genomic datasets from contemporary samples and historical specimens, respectively. The accuracy was 97.8% for historical samples with missing genotypes. The benchmarking multiple hybrid detection tools indicated that the RJF identification model was effective and practical. The further application to the genomic data from wild boar (Sus scrofa), domestic pigs, and their hybrids validated the pipeline. Our method has potential in not only monitoring genetic diversity in wild relatives of domestic animals but also supporting animal genetic resource conservation and management.
DNA barcode reference libraries provide useful tools for specimen identification, highlighting potential new species and detecting introduced ones. Here, we present a comprehensive DNA barcode library for European ants and, in order to tackle the Linnean, Wallacean and Darwinian shortfalls of this group, we provide an updated checklist, distribution data, mitochondrial genetic diversity maps and mitochondrial gene trees. The European ant fauna is here established to include 55 genera and 650 species (587 of which are native), including one species newly recorded for Europe and novel citations for 26 species from 11 countries. Our genetic dataset includes 6530 georeferenced COI sequences (62.1% d e novo) for 506 species (77.8%) across all genera. On average, 12.9 sequences were obtained per species, and 209 species were sequenced for the first time. We generated intra- and interspecific genetic distance estimates, 52 genus-level trees, mitochondrial genetic diversity and specimen maps for 384 species, as well as haplotype networks for 289 species, available in the Atlas V1.0 'The Mitochondrial Genetic Diversity Maps of European Ants'. We estimate that 56.3% of European ants are monophyletic with respect to the COI gene and can be unambiguously identified by DNA barcoding, though performance varies widely among genera. We observed moderate levels of barcode sharing (19.3%) and of barcode gap presence (47.6%), as well as high levels of intraspecific divergences (up to 17.9%). These findings likely reflect both biological and operational factors and highlight the existence of potential cryptic taxa and the need for taxonomic revisions. The framework presented here aims to facilitate future research, species discovery and conservation of European ants.
Genetic reference databases underpin a wide range of molecular approaches used to study cetacean biodiversity, including environmental DNA (eDNA), yet their reliability depends critically on data completeness, taxonomic accuracy, and metadata quality. Here, we present the first global assessment of mitochondrial sequence availability for cetaceans, evaluating taxonomic coverage, geographic representation, metadata completeness, and the distribution of five commonly targeted mitochondrial markers (12S rRNA, 16S rRNA, D-loop, cytochrome oxidase I, and cytochrome b). We retrieved 17,569 cetacean accessions from the NCBI Nucleotide database and an additional 259 COI-only records from BOLD Systems. Sequence availability was strongly biased toward Delphinidae and Balaenopteridae, whereas several families, notably Ziphiidae, were markedly underrepresented. We also identified discrepancies between database records and currently accepted cetacean taxonomy (e.g., outdated genera, non-accepted species, and collapsed higher-level taxa). Among markers, the D-loop dominated database representation, largely as standalone sequences (12,437 records), reflecting historical and current sequencing priorities and underscoring its continued relevance for population-level studies and eDNA marker development. Only 38% of accessions included geographic metadata, with georeferenced records concentrated primarily in the Americas and the Northwest Pacific, while large regions, including much of Africa, remained poorly represented. Although broad geographic patterns mirrored known family distributions, pronounced regional and taxonomic gaps persist. Our results highlight critical deficiencies in mitochondrial reference databases for cetaceans and emphasise the need for improved metadata standards, targeted sequencing of underrepresented taxa and regions, and open data sharing to enhance the effectiveness and global applicability of eDNA-based cetacean monitoring.
Sex chromosomes in amphibians exhibit substantial variability and often remain largely homomorphic, providing a powerful system for studying sex determination. The Sangzhi horned toad (Boulenophrys sangzhiensis) is an ideal model, but limited genomic resources have hindered insights into sex determination mechanisms. Here, we assembled a 2.8 Gb chromosome-level genome of B. sangzhiensis generated using PacBio HiFi and Hi-C data, achieving a Contig N50 of 30 Mb and identifying 21,775 protein-coding genes. Genome-wide analyses of coverage, FST, SNP density, and linkage disequilibrium identified two putative sex-linked regions on chromosome 2 and 6, showing sex-specific coverage and strong genetic differentiation. Within these sex-linked regions, we identified a Y-specific sequence and several candidate sex-determining genes, including Hsd11b2, Nlrp14, and zinc finger genes. Our findings suggest that sex determination in B. sangzhiensis involves a complex, possibly polymorphic mechanism, with multiple Y haplotypes segregating within populations. Additionally, the sex-linked regions exhibit accumulation of repetitive sequences, multi-copy genes, and chromosomal rearrangements, which may contribute to the evolution of sex chromosomes. This chromosome-level genome provides a valuable resource for understanding the dynamic evolution of sex-determining mechanisms in amphibians.
Subterranean ecosystems host a diverse range of ancient fauna, but studying these ecosystems is challenging due to significant sampling difficulties. Environmental DNA (eDNA) metabarcoding offers a promising approach for monitoring subterranean biodiversity, yet issues such as primer bias and non-target amplification can complicate its effectiveness. Thus, thorough validation of metabarcoding primers is crucial for accurate and comprehensive assessments of subterranean faunal diversity. This study aimed to address the need for robust primer validation through in silico, in vitro and in situ analyses, shedding light on primer performance across various subterranean taxa. The primary objective was to evaluate the effectiveness of COI metabarcoding primers for assessing subterranean faunal diversity. In silico analyses involved curating COI sequences from the Barcode of Life Database (BOLD) and selecting 14 primer combinations for in vitro testing using mock communities. Results revealed varying primer performance in terms of PCR efficiency and detection limits across different taxa. One primer combination (BF1/jgHCO2198) detected 82% of taxa in the mock community, but only at high DNA concentrations of the target taxa. The highest proportion of subterranean taxa detected in a diluted mock community was 68% using the fwhF2/fwhR2n primer combination. For in situ field validation, this same primer set detected 13 out of 16 subterranean taxa identified in haul net samples, along with an additional four taxa not identified by haul net. These findings highlight the potential of COI metabarcoding and the critical importance of primer selection for eDNA studies aimed at conserving subterranean biodiversity.
Hatchery supplementation is vital for conserving dwindling fish populations. Effective augmentation requires distinguishing hatchery-origin from wild individuals and accurately identifying species, particularly in systems where closely related species coexist. Genetic monitoring is key to quantifying genetic differences, but conventional markers do not distinguish hybrids, especially backcrosses. Misidentifying hybrids in hatchery programs compromises wild gene pools because hatchery broodstock contributes to numerous offspring being released into the wild. Here, we present a workflow for developing and evaluating the Genotyping-in-Thousands by sequencing (GT-seq) single nucleotide polymorphism (SNP) panel for North American river sturgeons (Scaphirhynchus spp.). This panel is designed to detect complex hybrid classes and to determine parent-offspring relationships. Our species identification panel (S-loci) contains 155 SNPs selected for high genetic differentiation (FST) between Pallid Sturgeon (S. albus) and Shovelnose Sturgeon (S. platorynchus), and the parentage assignment panel (P-loci) includes 112 SNPs with high heterozygosity within Pallid Sturgeon. Simulation analyses demonstrated that our GT-seq S-loci panel reliably classifies pure species, F1, F2 and backcross hybrids, even with up to 70% missing data. The P-loci panel achieves high-confidence parentage assignment with ≥ 80% typed loci, with performance influenced by the proportion of sampled parents. Overall, the novel Scaphirhynchus GT-seq panel developed in this study represents a robust and efficient tool for detecting hybridisation, assigning parentage and providing critical information for management decisions in ongoing Pallid Sturgeon conservation.
Environmental DNA (eDNA) metabarcoding can rapidly characterise biodiversity, yet its accuracy and effectiveness are limited by incomplete DNA barcode reference databases. We evaluated how comprehensive reference databases that include sequence variation within genomes (intragenomic) and across individuals and species (intergenomic) improve eDNA-based biodiversity assessments. We collected coral tissue and water samples at deep sites offshore Puerto Rico for reference barcoding and eDNA metabarcoding. Genome skimming coral specimens yielded 28S barcodes for 314 of 346 samples (90.8%) and revealed divergent intragenomic 28S lineages in multiple octocoral families. Incorporating local reference barcodes substantially changed ASV taxonomic classifications: 22 ASVs (8.9%) gained genus-level resolution, 19 ASVs (7.7%) were reassigned to different genera, and 14 ASVs (5.7%) lost incorrect genus-level classifications. Thus, incomplete reference databases produce not only unclassified ASVs but also false positive detections and ecologically meaningful misclassifications. When intragenomic 28S lineages were excluded from the reference database, 18 ASVs (7.4%) could not be classified to family or genus, demonstrating that unrecognised intragenomic variation can be mistaken for unsampled taxa. Integrating reference genome skimming and eDNA metabarcoding expanded known coral family richness by 36% at depths shallower than 1000 m and by 181% at depths greater than 1000 m. eDNA also detected two coral families previously unknown off Puerto Rico and nearby islands, underscoring its potential for biodiversity discovery.