
Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.
Replication is central to most experimental and sampling designs, increasing inferential power and capturing fine-scale data heterogeneity. However, its importance remains poorly evaluated in some ecological and evolutionary settings. This is the case of metabarcoding studies using DNA recovered from sedimentary archives, in which biological signals integrate ecological information through depositional and burial processes, yet are commonly inferred from a single sediment core per site. Here, we evaluated the effect of different types of replication using sedimentary DNA metabarcoding data from two genetic markers (mitochondrial COI and nuclear 18S) using a nested sampling design. The design included three intertidal sites, three spatially separated sediment cores per site (biological replicates), two sediment horizons per core, and eight PCR (technical) replicates per sediment sample. Variance partitioning showed that site identity and sediment age group together explained > 70% of the variation in beta diversity, indicating that among-site spatial and stratigraphic differences were the dominant drivers of community composition. PERMANOVA likewise identified non-significant effects of biological replication. Among PCR replicates from the same sediment sample, richness varied substantially, whereas Shannon diversity was more consistent. Despite this variability, differences in community composition among technical replicates remained smaller than those associated with biological replication or site identity, indicating a limited influence on broader ecological patterns. Community composition was highly similar among replicate cores within sites, consistent with stratigraphic coherence. These results indicate limited within-site heterogeneity and suggest that, under stratigraphically coherent conditions, increasing biological replication may provide little additional information, whereas enhancing technical replication and stratigraphic resolution can improve ecological inference from sedimentary DNA metabarcoding datasets.
Target capture is widely used to enrich endogenous DNA from calcium phosphate skeletal material in vertebrates, but its performance on calcium carbonate hard parts widely produced by invertebrates remains poorly understood. Here, we compared DNA recovery from four fresh and 12 ancient (eight radiocarbon-dated to 1671-1135 years old before present) deep-sea vesicomyid clam shells, including species Archivesica marissinica, A. nanshaensis and A. okutanii, using whole-genome sequencing (WGS) or target capture of ultraconserved elements (UCEs). WGS achieved 16.65% on-target read recovery of UCEs from fresh soft tissue, but < 1% from shell specimens. By contrast, UCE capture in the same specimen increased on-target reads by up to 155-fold, reaching 29.84% in fresh shells and up to 72-fold, reaching 19.89% in ancient shells. Target capture of UCEs recovered 142-1001 loci per sample compared to 0-230 with WGS alone. Ancient shells of A. marissinica and A. okutanii, based on reads mapped with bwa-mem2 and bbmap, exhibited characteristic post-mortem DNA damage signals, with average 5'-end C-to-T misincorporation rates of 3.46% and 15.97%, respectively, exceeding the levels observed in fresh A. marissinica shells (maximum 1.24%). UCE-based phylogenetic reconstructions incorporating shell ancient DNA recovered two major clades within Pliocardiinae, consistent with published phylogenomic trees. Together, these findings demonstrate that target-capture enrichment enables effective recovery of highly degraded DNA from ancient mollusc shells and supports robust phylogenetic inference at the intrageneric scale, expanding the utility of shells-one of the most abundant invertebrate remains-for evolutionary, biogeographic and conservation studies.
The sustainability of Southeast Asian fisheries hinges on high-throughput tools for monitoring early life-stage fish biodiversity. However, applying DNA metabarcoding to hyper-diverse tropical ichthyoplankton requires rigorous calibration to ensure quantitative reliability. We systematically evaluated the metabarcoding workflow using controlled mock communities, revealing that taxonomic recovery is governed by a stochastic limit of detection at a normalised proxy biomass threshold of ≤ 0.05. Quantitative analysis confirmed a significant linear relationship between specimen size and read abundance (R2 up to 0.817), demonstrating that biomass-driven template competition induces frequent false negatives for low-biomass taxa, a phenomenon exacerbated by increasing community complexity (ANOVA: p < 0.001). To mitigate these systemic biases, we applied a size-stratified specimen-balancing strategy intended to increase the representation of small-bodied components in natural bulk samples. Applying this optimised workflow to field samples from Khanh Hoa, Vietnam, we identified 139 species and unmasked a North-South biogeographic dichotomy (PERMANOVA: R2 = 53%, p = 0.001) driven by transect-scale environmental gradients and local hydrography. Notably, we identified diversity hotspots requiring > 200,000 reads for saturation, suggesting these sites act as critical larval retention zones. The contrast between functional management zones was highly significant (p = 0.002), with the conservation area (Zone B) exhibiting higher alpha richness and a nine-fold increase in unique indicator species compared to high-activity areas (18 vs. 2). Our work demonstrates that comprehensive validation is vital for accurate metabarcoding, offering a robust framework to understand how ecological gradients and localised human pressures shape Vietnam's critical marine spawning sites and nursery grounds.
The proliferation of high-dimensional data in ecology and evolutionary biology raises the promise of statistical and machine learning models that are highly predictive and interpretable. However, high-dimensional data are commonly burdened with an inherent trade-off: in-sample prediction of outcomes will improve as additional variables are included in the model, but this may come at the cost of poor predictive accuracy and limited generalizability for future or unsampled observations (out-of-sample prediction). To confront this problem of overfitting, sparse models can focus on key variables by correctly placing low weight on unimportant variables. We compared nine methods to quantify their performance in variable selection and prediction using simulated data with different sample sizes, numbers of variables, and strengths of effects. Overfitting was typical for many methods and simulation scenarios. Despite this, in-sample and out-of-sample prediction converged on the true predictive target for simulations with more observations, larger causal effects, and fewer variables. Accurate variable selection to support process-based understanding will be unattainable for many realistic sampling schemes in ecology and evolution. We use our analyses to characterize data attributes for which statistical learning is possible, and illustrate how some sparse methods can achieve predictive accuracy while mitigating and learning the extent of overfitting.
Traditional classification of flagellates (i.e., flagellated protists) relies on morphological traits or a single molecular marker, which suffer from subjectivity and limited data sources. This study proposes a multimodal deep learning model, Residual Multi-Feature Attention-50 (ResMFA50), that integrates photomicrographs and small subunit ribosomal RNA (SSU rRNA) gene sequences of flagellates. The dual-branch architecture (DNA sequence and image branches) extracts local and global features, while the Multi-Feature Attention (MFA) mechanism dynamically fuses heterogeneous data. Experiments were conducted on a dataset comprising 296 SSU rRNA gene sequences and 308 standardized photomicrographs, evaluated using 10-fold cross-validation. The results demonstrate that ResMFA50 achieves an accuracy of 92.5% in classifying flagellates at a batch size of eight, which is significantly higher than the accuracies achieved by SVM (82.4%), Random Forest (84.2%), EfficientNet (83.6%), ResNet50 (87.4%), and MMNet (91.3%). Moreover, ablation experiments comparing early, intermediate, and late fusion strategies demonstrate that the proposed late fusion scheme consistently outperforms other fusion timings, achieving improvements of 3.2%-3.8% over early fusion across different batch sizes. This study establishes a methodological foundation for modelling multimodal biological data in complex systems, advancing deep learning applications in integrative taxonomy. This advantage is attributed to the dual-channel global pooling mechanism (Global Average/Max Pooling fusion), which balances the variance-bias trade-off through complementary strategies of spatial statistical smoothing and local salient feature detection, enhancing robustness to data scale expansion.
Quantitative monitoring of fish using environmental DNA (eDNA) typically relies on quantitative real-time PCR (qPCR) and droplet digital PCR (ddPCR), yet both methods are constrained in throughput, making it difficult to meet the demands of basin-scale surveys. High-throughput qPCR (HT-qPCR) has emerged as a platform capable of simultaneously processing large numbers of targets and samples, but its performance for fish eDNA quantification and its comparability with ddPCR remain poorly characterized. Here, we selected nine widely distributed freshwater fishes and designed 19 species-specific primer pairs (2-3 per species). All assays performed well on a high-throughput chip-based qPCR (HT-qPCR) platform, with amplification efficiencies ranging from 91.07% to 105.83% and 78.9% of assays falling within the optimal 95%-105% interval. Applying HT-qPCR to quantify fish eDNA across six aquatic regions, Cyprinus carpio (C. carpio) achieved a 100% detection rate. For the same fish species, the two primer sets exhibited measurable differences in detection rate and quantitative estimates. HT-qPCR and ddPCR estimates were strongly correlated (r = 0.858), and Deming regression indicated near-equivalence (HT-qPCR = 0.52 + 0.99 × ddPCR). Bland-Altman analysis revealed a consistent ~3-fold overestimation by HT-qPCR relative to ddPCR, which is generally considered the more accurate method, suggesting that a correction factor could harmonize data across platforms. These results validate the utility of HT-qPCR for catchment-scale quantification of fish eDNA concentrations and provide methodological innovation and technical support for aquatic ecosystem monitoring and assessment.
Estimating malaria parasite migration is crucial for informing elimination strategies, particularly by identifying regions with higher parasite migration that may serve as transmission sources and intervention targets. Methods that infer spatial variations in gene flow from georeferenced genetic data, including Estimated Effective Migration Surfaces (EEMS) and related approaches such as estimating Migration And Population-size Surfaces (MAPS), have become widely used to visualize barriers and corridors of organism movement. However, when sampling locations are sparse or unevenly distributed, spatial gene-flow maps can contain artefacts that are difficult to interpret, and methods such as EEMS provide posterior probability summaries that indicate whether inferred migration rates are statistically supported by the genomic data. However, these summaries do not directly capture whether a spatially inferred feature, such as a migration barrier or corridor, lies in a region with sufficient sampling locations to be geographically reliable. A high posterior probability in a sparsely sampled region reflects model confidence given available data, not spatial sampling adequacy. In this study, we developed a sample location-aware (SLA) filtering workflow that applies topological skeletons and kernel density estimation to assess sample density for each contour and filter out areas with lower sample density. This workflow can be applied to any gene-flow inference approach that produces spatially explicit migration features (e.g., contours, surfaces, or high/low-migration regions) from genotype data and sample coordinates. Using simulated genomic data and different barrier configurations under systematic site subsampling scenarios, we show that posterior-only filtering can retain spurious barrier features, whereas SLA filtering consistently increases the precision of inferred low-migration regions across sample sites. In an empirical case study of Plasmodium falciparum from Cambodia, SLA filtering produced more stable estimates of parasite migration patterns under location subsampling than posterior probability filtering alone. These results demonstrate that incorporating an explicit measure of geographic sampling density around each inferred migration feature improves the reliability and interpretability of spatial gene-flow migration patterns for malaria parasites and provides a general diagnostic and filtering approach for a broad class of georeferenced population-genetic migration mapping methods.
Metabarcoding is a powerful tool for biodiversity comparisons, where standard-size DNA barcodes (> 500 bases) offer better taxonomic resolution than shorter ones. Still, the choice of sequencing platforms and bioinformatics pipelines may strongly affect inferred diversity due to various technical biases. We assessed the relative performance of Illumina MiSeq i100 (2 × 500 paired-end), PacBio Revio and Oxford Nanopore MinION sequencing and bioinformatics pipelines, using full-length ITS amplicon sequencing datasets from a 103-species mock community and 45 composite soil samples. Despite numerous low-quality reads, PacBio yielded the lowest overall error rate and highest number of taxa. Illumina revealed the highest proportion of chimeric and index-switched reads, along with a strong bias towards shorter amplicons. MinION data analysed using PRONAME and Minovar-a bioinformatics pipeline presented here-had the largest proportion of low-quality data, and rare taxa were lost during data filtering and read polishing steps. Although Minovar enabled amplicon sequence variant (ASV) level precision for common taxa, we recommend clustering ASVs into OTUs. For PacBio, standard filtering approaches outperformed the ASV approach because they retained rare taxa. For Illumina, a stringent ASV approach or removal of rare OTUs would limit artefacts. Across all platforms, excess PCR cycles promoted chimeric and low-quality reads and lost quantitativity in biodiversity assessments. With moderate differences in effect sizes, all analytical approaches supported the conclusion that sampling design determines how we see soil biodiversity responses to land use. For biodiversity surveys based on the full-length ITS metabarcoding, we recommend using PacBio sequencing with standard, non-ASV pipelines.
Spotted sea bass (Lateolabrax maculatus) is an ecologically and economically important species that supports a large-scale aquaculture industry in China. However, most farmed populations have not undergone systematic breeding programs and are characterised by undocumented empirical selection. Understanding genetic diversity dynamics and gene evolution patterns for farmed populations is critical for designing appropriate breeding strategies. Here, we assembled a chromosome-level reference genome using HiFi and Hi-C sequencing technologies, resulting in a highly contiguous assembly with a total size of 636.65 Mb and a scaffold N50 of 27.27 Mb. Furthermore, we performed population genomic analyses using 1107 representative wild and farmed individuals. Significant genetic differentiation and reduced genetic diversity were observed in farmed populations relative to wild populations. Pan-genome of 1107 accessions further captured 86.48 Mb of non-reference novel sequences (NRNS) and 206 novel genes absent from the reference genome. Gene presence/absence variation (PAV) analysis revealed that farmed populations harbour fewer genes than wild counterparts, implying gene loss and negative selection during the domestication process. Lost genes were enriched in functions related to immune response and nervous system development, highlighting wild populations as important genetic resources for future breeding programs. Notably, selection signatures and comparative PAV analyses across independent farmed populations revealed parallel domestication effects at both SNP and gene PAV levels, suggesting similar selective pressures despite varied aquaculture practices and domestication histories. Our study provides the first pan-genome of spotted sea bass that deciphers genetic diversity dynamics and gene evolution patterns during domestication selection and establishes valuable genomic resources for sustainable genetic improvement of this species.
DNA metabarcoding using relative read abundance (RRA) is commonly applied to estimate herbivore diet composition, yet its quantitative accuracy remains uncertain. We assessed taxonomic resolution and quantitative performance of RRA from DNA metabarcoding compared to metagenomic sequencing and hybridization capture, using deer scats from feeding trials and recreated diet samples using plant tissues. All methods recovered plant composition in recreated diets (R2 = 0.59-0.82), indicating accurate scaling with biomass in the absence of digestion, with only minor bias from amplicon length in DNA metabarcoding. In contrast, RRA from scat samples performed poorly (R2 < 0.01) across all methods largely due to differential plant digestibility. Correcting for digestibility, measured with acid detergent lignin and acid-insoluble ash, was strongly supported in mixed-effects models and improved prediction of dietary composition, although species-level variation remained. For metagenomic sequencing and hybridization capture, we also evaluated Relative Genome Coverage (RGC), a novel relative abundance metric quantifying the proportion of each plant's chloroplast genome covered by mapped reads, normalized for genome length. RGC further improved correlations in recreated diets (R2 = 0.82-0.84) and, with hybridization capture, largely overcame digestibility-related biases in scat samples (R2 = 0.57) without correction. When such corrections are infeasible, hybridization capture with uncorrected RGC may achieve higher quantitative accuracy in scat samples. Our results provide practical guidance for improving molecular herbivore diet analysis and highlight the importance of accounting for digestion-related biases.
Quantifying abundances of unicellular eukaryotes (protists) remains a central challenge in microbial ecology, as methodological differences can strongly influence abundance estimates and ecological interpretation. Although molecular tools have thus far greatly improved our understanding of protists, high rRNA gene copy numbers limit quantitative inferences. Digital PCR (dPCR) has emerged as a promising tool for absolute quantification, yet its application for unicellular eukaryotes and its comparability to established cell-based methods remain insufficiently explored. Here, we develop species-specific dPCR assays for two important freshwater ciliates (Urotricha castalia and Urotricha pseudofurcata) and establish gene copy number correction factors to enable highly accurate quantitative abundance estimates. We assess assay performance using controlled laboratory experiments and apply the approach to environmental samples, directly benchmarking dPCR against catalyzed reporter deposition-FISH (CARD-FISH). Under controlled conditions, dPCR and CARD-FISH yielded comparable accuracy, with dPCR showing superior precision. In field applications, method-dependent differences emerged, reflecting both methodological constraints and biological variability. Notably, dPCR provided an overall higher sensitivity, enabling robust detection of low-abundance taxa. Our results highlight dPCR as a scalable and sensitive approach that, when combined with appropriate correction strategies, represents a significant step towards more reliable molecular quantification of protists. At the same time, differences between methods underscore the value of integrating molecular and microscopy-based approaches. We propose that combining dPCR with tools such as CARD-FISH can offer complementary insights into protist population dynamics. Such integrative frameworks provide a powerful path forward for improving abundance estimates and advancing quantitative microbial ecology.
Accurate identification of centromeres in telomere-to-telomere (T2T) genomes remains difficult due to the rapid evolution of centromeric repeats and their lack of conserved sequence features. In this study, we present EasyCen, a lightweight sequence-based framework for centromere identification and repeat-architecture profiling across various eukaryotes. Rather than relying on repeat annotation or homology, EasyCen recognises centromeres based on recurrent positional features of repetitive DNA. Besides centromere localisation, EasyCen incorporates a repeat-pair profiling module for exploratory characterisation of internal repeat organisation. Benchmarking on Arabidopsis thaliana and Mus musculus showed high accuracy (typically > 85% coordinate overlap with published annotations) and substantially reduced runtime. Cross-species analysis revealed a 'beads-on-a-string'-like repeat pattern in both mouse and human centromeres, associated with GC- and CpG-rich subdomains. EasyCen performs effectively without pre-existing repeat libraries, making it particularly useful for large, repeat-rich, or non-model genomes. Our analyses further suggest that certain organisational features of centromeric repeats may recur across diverse eukaryotic lineages despite rapid sequence turnover.
Genotyping-in-thousands by sequencing (GT-seq) panels are powerful tools in ecological, evolutionary and conservation genomics, yet the optimization process critical for robust and reproducible genotyping remains poorly formalized. Here, we present an iterative workflow for GT-seq panel optimization that emphasizes systematic refinement, quality control and structured decision-making to improve panel performance across diverse populations and study contexts. We illustrate this framework through the development and optimization of a GT-seq panel for white-tailed deer, a widely distributed and ecologically important North American species. From an initial set of 1200 candidate SNPs selected from a commercial microarray (OVSNP60, containing 72,728 SNPs) and prioritized for high heterozygosity, primers were designed for 646 loci. The final optimized panel contains 508 high-performing markers retained after iterative removal of overamplifying primer pairs, adjustment of primer concentrations, PCR conditions and bioinformatic filtering. The overall proportion of SNPs with more than 70% genotype rate increased from 25.5% in the first optimization round to 87.8% in the final round. Consequently, the overall genotype rate increased from 39.4% to 84%. We also identify key quality-control checkpoints and practical criteria to guide panel refinement and ensure consistent performance. By prioritizing optimization as an integral component of GT-seq panel development, this work provides a reproducible framework for generating robust, high-throughput genotyping tools in non-model species and underscores the importance of iterative refinement to maximize data quality and utility.
Noninvasive genetic sampling is widely used in ecology and conservation to identify predators and their diets but recovering individual-level information from consumed prey remains largely unexplored. We evaluated whether individual prey can be reliably genotyped from carnivore scats and assessed limitations associated with degraded and mixed DNA sources. We developed a 31-locus SNP panel optimized for genotyping degraded elk (Cervus canadensis) DNA using amplicon sequencing. We validated prey genotyping by matching elk genotypes recovered from carnivore scats ('carcass scats') collected at cougar (Puma concolor) kill sites to genotypes from corresponding elk carcasses. We also genotyped scats collected throughout the study area, identified as containing elk using DNA metabarcoding ('survey scats'). To avoid misidentifying multiple individuals in a scat as a unique genotype, we evaluated artificial mixtures of prey DNA to assess the ability to detect and filter samples containing mixed DNA. Elk genotypes recovered from carcass scats matched associated carcass genotypes, confirming accurate recovery of individual prey DNA from carnivore scats. Genotyping success was 88% in fresh carcass scats and 74% in survey scats of unknown age. Heterozygosity excess filtering removed most mixed samples, although one of 24 mixtures with equal DNA contributions from two individuals produced a false unique genotype. Our results demonstrate that carnivore scats can serve as a reliable source of individual-level prey DNA under appropriate conditions. This method provides a new data stream on individual prey mortality, predation rates, and scavenging dynamics, processes that have previously been difficult to quantify without invasive capture and collaring techniques.
Metabarcoding of environmental and ancient environmental DNA (eDNA and sedaDNA) is a powerful approach for studying and monitoring marine communities. However, its effectiveness is limited by the availability of comprehensive and well-curated reference databases, particularly for protists. Here, we introduce pr2-wormifier, a bioinformatics pipeline designed to create customized and improved reference databases for 18S rRNA-based metabarcoding. This pipeline integrates sequences from PR2 and NCBI with taxonomic information from the World Register of Marine Species (WoRMS) and AlgaeBase, allowing for refined taxonomic assignments at the genus and species levels. pr2-wormifier enables users to tailor reference databases to specific taxonomic groups or geographic regions, enhancing the resolution and accuracy of biodiversity assessments. We benchmarked the pipeline using a sedimentary ancient DNA dataset from the Baltic Sea, focusing on marine protists, especially ciliates and dinoflagellates. The customized database generated by pr2-wormifier identified more sequences at the genus and species levels than PR2 alone, while maintaining taxonomic consistency and quality. Our results demonstrate that pr2-wormifier addresses common limitations of existing databases, such as low taxonomic resolution and missing taxa and facilitates more reliable classification in metabarcoding studies. By enabling the creation of locally relevant, taxonomically curated databases, pr2-wormifier offers a flexible and scalable solution for improving the identification of protists in environmental and paleoenvironmental research.
Advances in non-invasive genetic sampling and long-term genetic monitoring programmes have enabled collection of large individual genotype datasets for many wildlife populations, often accompanied by rich field metadata that place the genotyped individuals in time and space. These datasets allow reconstruction of multigenerational pedigrees and have the potential to provide valuable insights into population demography, reproduction, dispersal, social structure and genetic processes. But while the tools for construction of pedigrees keep improving, their interpretation remains challenging. Integrating multigenerational pedigree data with field metadata creates significant complexity, yet specialized tools to facilitate the interpretation of such datasets remain scarce. Here we introduce wild pedigree exploreR (wpeR), an R package designed to simplify exploration, organization and interpretation of complex pedigrees. The package enables users to link reconstructed pedigrees with genetic sample metadata, enabling evaluation of biological plausibility of inferred relationships, but also allowing exploration of other characteristics of individuals and populations in spatial and temporal contexts. wpeR implements a linear workflow through which the pedigree data is imported, formatted, organized into families and integrated with field metadata. The resulting dataset can be visualized through temporal plots that track individuals and families over time, as well as with spatial outputs representing parent-offspring relationships and individual movement patterns as geographic features that can be either directly visualized on maps within R, or exported to be further explored with common GIS tools. wpeR allows exploration of lineage relationships within their ecological context, bridging the gap between statistically reconstructed pedigrees and their biological interpretation. It provides a scalable and flexible framework for analyzing these complex data, providing a practical tool for researchers and managers working with genetic monitoring datasets.
Persistent biodiversity data shortfalls undermine our capacity to detect species, map their distributions and characterize their spatial genetic structure, limiting robust biogeographic analyses and the development of effective conservation strategies. This particularly affects hyperdiverse invertebrate groups where hidden diversity remains largely undocumented. This study develops and demonstrates the potential of an integrated high-throughput sequencing (HTS) framework to improve the representation of hidden diversity in regional species inventories and to help close critical gaps in our understanding of species distributions and genetic diversity from a conservation biogeography perspective. Focusing on the Canary Islands (Spain), the workflow combines megabarcoding of more than 4000 mesofauna specimens to generate a curated species-level molecular reference library with community DNA metabarcoding of 168 soil samples. This approach enables consistent taxonomic assignment across insular landscapes and increases the spatial and genetic resolution of occurrence data. We identified 145 species of mites and springtails, including 49 species newly recorded for the archipelago and numerous genetically distinct lineages likely representing undescribed taxa, highlighting all the biodiversity that remains to be described. Integration of the barcode library with metabarcoding data produced 1440 species occurrences, revealing extensive distributional gaps, multiple range expansions and strong within-island phylogeographic structuring, indicating prevalent diversification at fine spatial scales. These results highlight a deep, taxonomically broad underestimation of soil biodiversity and demonstrate that this integrative approach provides a transferable model for advancing the biogeography, evolutionary understanding and conservation of dark and cryptic taxa across broad taxonomic and conservation-relevant contexts.
Whole genome amplification (WGA), and in particular multiple displacement amplification (MDA), has become a key technique for genomic sequencing of microscopic organisms, yet it introduces artefacts such as palindromic (inverted chimeric) reads that may compromise downstream analyses. We assessed how pervasive palindromic reads generated by MDA impact the assembly of tardigrade (Acutuncus giovanniniae and A. mecnuffi) mitogenomes sequenced with Oxford Nanopore technology. We show that the MDA produces a high proportion of palindromic reads, often exceeding one-third of mitochondrial reads and frequently exhibiting complex multi-inversion structures. These artefacts severely impair long-read assembly, leading to low success rates and inconsistent genome reconstruction. To solve this issue, a strategy based on in silico fragmentation of long reads into short, high-quality fragments, followed by short-read assembly, consistently produced complete and accurate circularised mitochondrial genomes. Our results demonstrate that palindromic read formation can be, in some cases, a limitation of MDA coupled with long-read sequencing, but this issue can be mitigated through read fragmentation. This approach provides a simple, robust and scalable solution for mitogenome assembly from data heavily affected by amplification artefacts, particularly in microscopic taxa where whole genome amplification is often unavoidable.
The choice of optimality criterion is a key consideration in phylogenetic studies. Recent work challenges the notion that more computationally demanding optimisation objectives result in better phylogenetic trees. This finding underscores the importance of comparing trees across optimisation objectives in addition to different models of evolution and data partitions. It is currently cumbersome to optimise trees for alternative objectives because multiple programmes must be used, each with its own optimisation framework. Here, I introduce Treeline for optimising balanced minimum evolution, maximum likelihood and maximum parsimony trees. Treeline explores the optimisation landscape using a new strategy based on perturbing the patristic distance matrix used to initialize candidate trees, which is shown to be particularly effective for the balanced minimum evolution objective. Tests suggest Treeline can be more accurate, memory efficient or faster than existing programmes designed for a single optimisation objective. Treeline's unified nature facilitates comparison of phylogenetic trees across distinct optimisation objectives and models of sequence evolution. Consistent with prior studies, the balanced minimum evolution objective resulted in gene trees that were more consistent with species trees than maximum likelihood or maximum parsimony objectives. With a case study of species from the family Hominidae, including Homo sapiens, I show how Treeline simplifies the process of constructing trees under alternative optimisation objectives and models of evolution. Treeline is part of the DECIPHER package for R and is available from Bioconductor and online (https://DECIPHER.codes/).