With the explosion of chromosome-scale genome assemblies being generated in recent years, there is vast potential for comparative genomics analyses through detecting multi-genome synteny. While existing tools can detect synteny blocks between multiple genomes, their text-based outputs make it challenging to intuitively explore large-scale synteny patterns. Interpretable, information-rich and easy-to-use synteny visualization tools are imperative to enable important biological insights from the synteny block data output by the aforementioned utilities. Here, we present ntSynt-viz, a command-line tool for automated sorting, normalization and plotting of multi-genome synteny blocks. We show how ntSynt-viz provides clearer and more easily interpretable chromosome painting ribbon plots compared to the state-of-the-art tools NGenomeSyn and plotsr when evaluating synteny between 14 human genomes, and compared to NGenomeSyn when comparing 9 hoverfly genomes. As plotsr is limited to comparing genomes with equal chromosome numbers, it was not applicable to the hoverfly dataset. Furthermore, we demonstrate how ntSynt-viz can also be applied to visualize syntenic patterns encoded in pangenome graphs, using a Minigraph-Cactus graph built from 16 Drosophila genomes. We expect that ntSynt-viz will provide crucial insights into large-scale synteny patterns between divergent genomes, thereby advancing research into key evolutionary questions.
Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10 250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score ± SD: 0.012 ± 0.052), Black (0.013 ± 0.042), White (0.055 ± 0.022), and South Asian (0.057 ± 0.098) participants, whereas higher scores among Hispanic (0.083 ± 0.060) and Middle Eastern or Central Asian (0.122 ± 0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.
Advanced long-read sequencing technologies, such as those from Oxford Nanopore Technologies and Pacific Biosciences, are finding a wide use in de novo genome sequencing projects. However, long reads typically have higher error rates relative to short reads. If left unaddressed, subsequent genome assemblies may exhibit high base error rates that compromise the reliability of downstream analysis. Several specialized error correction tools for genome assemblies have since emerged, employing a range of algorithms and strategies to improve base quality. However, despite these efforts, many genome assembly workflows still produce regions with elevated error rates, such as gaps filled with unpolished or ambiguous bases. To address this, we introduce GoldPolish-Target, a modular targeted sequence polishing pipeline. Coupled with GoldPolish, a linear-time genome assembly algorithm, GoldPolish-Target isolates and polishes user-specified assembly loci, offering a resource-efficient means for polishing targeted regions of draft genomes. Experiments using Drosophila melanogaster and Homo sapiens datasets demonstrate that GoldPolish-Target can reduce insertion/deletion (indel) and mismatch errors by up to 49.2
Long-read sequencing platforms such as the Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) platforms now offer sufficient read lengths, throughput, and accuracy at competitive costs to analyze polymorphic regions of the human genome, including the highly complex human leukocyte antigen (HLA) gene cluster-a cornerstone of human immunity. Here, we present a streamlined protocol for predicting HLA signatures from whole-genome shotgun (WGS) long-read sequencing data by directly streaming sequence alignments into HLAminer. This method is as simple as running minimap2, scales efficiently with the number of sequences, and works with any read aligner compatible with the SAM file format-eliminating the need to store bulky alignment files on disk. We provide a step-by-step guide for predicting HLA class I and class II alleles from third-generation long-read sequencing data and demonstrate the robustness of predictions even with older, less accurate WGS nanopore datasets and relatively low (10×) sequencing coverage. Code availability: HLAminer is available under the BC Cancer software license agreement (academic use) at https://github.com/bcgsc/HLAminer. © 2025 The Author(s). Current Protocols published by Wiley Periodicals LLC. Basic Protocol: HLA prediction from streamed ONT or PacBio long-read alignments.
The ever-growing global health threat of antibiotic resistance is compelling researchers to explore alternatives to conventional antibiotics. Antimicrobial peptides (AMPs) are emerging as a promising solution to fill this need. Naturally occurring AMPs are produced by all forms of life as part of the innate immune system. High-throughput bioinformatics tools have enabled fast and large-scale discovery of AMPs from genomic, transcriptomic, and proteomic resources of selected organisms. Public protein sequence databases, comprising over 200 million records and growing, serve as comprehensive compendia of sequences from a broad range of source organisms. Yet, large-scale in silico probing of those databases for novel AMP discovery using modern deep learning techniques has rarely been reported. In the present study, we propose an AMP mining workflow to predict novel AMPs from the UniProtKB/Swiss-Prot database using the AMP prediction tool, AMPlify, as its discovery engine. Using this workflow, we identified 8008 novel putative AMPs from all eukaryotic sequences in the database. Focusing on the practical use of AMPs as suitable antimicrobial agents with applications in the poultry industry, we prioritized 40 of those AMPs based on their similarities to known chicken AMPs in predicted structures. In our tests, 13 out of the 38 successfully synthesized peptides showed antimicrobial activity against Escherichia coli and/or Staphylococcus aureus. AMPlify and the companion scripts supporting the AMP mining workflow presented herein are publicly available at https://github.com/bcgsc/AMPlify.
Abstract High‐quality, accessible mitogenome sequences are crucial for comparative genomics, phylogenetic constructions and environmental DNA (eDNA) applications. With advancements in sequencing technologies, genomic sequencing reads are available for a wide variety of species. Although these data sets contain mitochondrial sequences, this information remains largely untapped due to the limited specialized tools for assembling mitogenomes from short‐read libraries. The present study introduces the Mitochondrial Genome Reference‐grade Assembly and Standardization Pipeline (mtGrasp), a streamlined and memory‐efficient pipeline for assembling complete and standardized mitogenomes from short‐read metazoan DNA libraries (https://github.com/bcgsc/mtGrasp). We assembled 23 short‐read libraries and found that mtGrasp ran 2–7 times faster and used 3–4 times less memory, on average, compared to the state‐of‐the‐art mitogenome assembly pipelines GetOrganelle and MitoZ. Additionally, mtGrasp successfully assembled mitogenomes of varying completeness across different animal taxa. We illustrate that standardizing existing mitogenome annotations with mtGrasp would improve the accuracy of downstream phylogenetic analysis and demonstrate the utility of mtGrasp in designing robust targeted quantitative polymerase chain reaction‐based eDNA assays using mitogenomes assembled from museum voucher specimens. We expect the reference‐grade mitogenomes generated with freely available mtGrasp to advance the development of robust, high‐quality eDNA tools and aid in a variety of comparative genomic analyses.
Antimicrobial resistance is a rapidly escalating global health concern, largely driven by the overuse and misuse of antibiotics in agriculture and clinical settings. There is an urgent and currently unmet need for effective therapeutics against multi-drug resistant (MDR) pathogens. Antimicrobial peptides (AMPs) are a diverse class of cationic peptides that have the potential to overcome extant resistance mechanisms. We evaluated the antimicrobial efficacy of 13 structurally distinct antimicrobial peptides against a panel of drug-resistant Escherichia coli strains. Toxicity assays, including hemolysis and cell viability tests, were performed to determine therapeutic indices and assess the clinical potential of those select peptides. Our findings indicate that, relative to susceptible reference strains, the tested AMPs retain robust bioactivity against known MDR isolates, exhibiting only marginal or no decrease in antimicrobial efficacy. Among them, TeRu4 emerged as the lead candidate, with a minimum inhibitory concentration of 0.5 μg/mL and a therapeutic index exceeding 256. This study underscores the potential of AMPs to act as powerful alternatives to traditional antibiotics, offering new possibilities to address public health concerns surrounding drug-resistant bacterial infections.
Globally, coastal waters experience degradation from pollution associated with multiple discharges, including industrial and agricultural runoff, and municipal wastewater. Certain benthic infaunal taxa are tolerant of high nutrient input and anoxic conditions, while others are sensitive to these conditions. Using these indicator taxa as proxies for assessing organic enrichment is well established to characterize subsequent pollution impacts. Conventional assessment of macroinfauna involves the detailed analysis of each individual specimen within a sample by taxonomic experts, a resource intensive process. As an alternative, we developed sensitive quantitative polymerase chain reaction (qPCR) assays to detect these indicator taxa in a scalable and reliable way. Using whole genome shotgun sequencing, we generated full mitogenome sequences of selected indicator macroinfaunal polychaetes routinely used for monitoring programs in Pacific Northwest marine environments. These sequences were used to design five new, rigorously validated environmental DNA (eDNA) assays capable of detecting low levels of DNA that can be isolated from environmental samples. For nine sites at a wastewater treatment plant outfall in Vancouver, British Columbia, we tested three eDNA sample collection types: active filtration, a passive dip filter from water containing collected macroinfauna, and active filtration from water collected near the sea floor. Generalized linear models indicated that eDNA signal strength correlated with organism count particularly with passive dip sample collection type. eDNA occupancy modelling techniques estimated detection probabilities corresponding with organism count. The present study emphasizes the value of integrating eDNA into marine outfall monitoring efforts to enhance the assessment of environmental effects.
Timely and accurate assessment of the presence of at-risk or invasive species is critical for effective responses to climate change and human impacts. For example, at-risk species are often difficult to find, while invasive species are often well established before their infiltration is detected using conventional surveying methods. However, all organisms release genetic material such as DNA into their surroundings, leaving traces of themselves that can be detected using environmental DNA (eDNA) methods. These approaches are powerful tools in the conservation toolbox, as they are transforming how risk assessments and the evaluation of mitigation and remediation effectiveness are done. Despite this, poorly performing tools hinder broad adoption of eDNA-based detection methods, due in part to their associated high false negatives and false positives that can impair effective management decision-making. iTrackDNA is a multi-year, large-scale applied research project that is addressing these concerns with researchers and end users from various sectors across North America. It is building end-user capacity through innovative, accessible, socially responsible genomics-based analytical eDNA tools for effective decision-making by publishing 125 quantitative real-time polymerase chain reaction (qPCR) primer/probe sets designed to detect key invertebrates, fish, amphibians, birds, reptiles, and mammals in coastal and inland ecosystems important to North America, with an emphasis on Canada. These 125 assays were designed to meet or exceed the new Canadian Standards Association (CSA) consensus-based and multi-stakeholder national standards for eDNA (CSA W214:21 and CSA W219:23). Herein, we describe how we applied eDNA assay design and validation approaches across a wide range of animal taxa to achieve compliance.
Identification of conserved genomic sequences and their utilisation as anchor points for clade detection and/or characterisation is a mainstay in ecological studies. For environmental DNA (eDNA) assays, effective processing of large genomic datasets is crucial for reliable species detection in biodiversity monitoring. While considerable focus has been on developing robust species-targeted assays, eDNA assays with broader taxonomic coverage (e.g., detecting any species within a taxonomic group such as fish), can significantly streamline environmental monitoring, especially when detecting individual species' DNA proves challenging. Designing such assays requires identifying conserved regions representing the target taxonomic group, a chiefly manual task that is often labor-intensive and error-prone, particularly when working with large sequence datasets. To address these challenges, we present unikseq2, an enhanced, alignment-free, k-mer-based tool for identifying unique and conserved sequences. It introduces a new functionality to identify sequence conservation among target species, enabling more informed marker selection for applications such as universal primer design. This automates sequence selection in large-scale mitochondrial genome datasets eliminating the need for manual inspection of computationally costly multiple sequence alignments. Herein, we demonstrate unikseq2's capabilities by developing and validating eDNA assays for various taxa, including Osteichthyes (bony fishes), the Salmonidae family (salmon and trout), Myotis bats and Cervus deer. Unikseq2-based eDNA assays allow for accurate detection across multiple taxonomic levels, from genus to class, enhancing the flexibility, scalability and reliability of eDNA tools in environmental monitoring. By leveraging genomic data from public repositories, unikseq2 supports efficient, reproducible assay design, making it an invaluable tool for a wide range of ecological and biodiversity research applications.
Accurately characterizing human diversity is foundational to equitable genomics. In this study, we present a large-scale analysis comparing self-declared ancestry with genetically inferred ancestry in 10,250 participants from the pan-Canadian HostSeq cohort. Using the ntRoot algorithm on whole genome sequencing data, we inferred both global and local continental-level ancestry and assessed concordance with self-identified sociocultural categories. High agreement was observed among individuals self-identifying as White (concordance rate=98.8%), Black (97.2%), East Asian (96.1%), and South Asian (89.9%), while substantial discordance was found in those self-identifying as Hispanic (concordance rate=74.6%), Middle Eastern / Central Asian (67.9%) or Indigenous (40.7%). We quantified agreement using Cohen’s kappa (κ = -0.01 unweighted; 0.35 weighted) and assessed admixture complexity with Shannon entropy, revealing a strong relationship between discordance and ancestry heterogeneity. Principal component analysis further revealed that tightly clustered genetic profiles often corresponded with lower admixture complexity, whereas broader, overlapping distributions were observed in groups with more heterogeneous ancestry and complex sociocultural histories. These findings underscore the complex interplay between sociocultural identity and genomic data, with discordance patterns reflecting the historical and cultural complexity of human populations. By quantifying this relationship systematically with ntRoot, our approach provides a framework for moving rigid categorical labels toward more nuanced genome-derived ancestry characterization that can improve both scientific rigor and representational equity in genomics. Author Summary Study concept: RLW. Software implementation: RLW. Data analysis: RLW. Manuscript development: RLW. Manuscript editing: RLW, IB. Funding acquisition: IB. ### Competing Interest Statement The authors have declared no competing interest. Canadian Institutes of Health Research, https://ror.org/01gavpb45, PJT-183608
Antimicrobial resistance is a critical public health concern, necessitating the exploration of alternative treatments. While antimicrobial peptides (AMPs) show promise, assessing their toxicity using traditional wet lab methods is both time-consuming and costly. We introduce tAMPer, a novel multi-modal deep learning model designed to predict peptide toxicity by integrating the underlying amino acid sequence composition and the three-dimensional structure of peptides. tAMPer adopts a graph-based representation for peptides, encoding ColabFold-predicted structures, where nodes represent amino acids and edges represent spatial interactions. Structural features are extracted using graph neural networks, and recurrent neural networks capture sequential dependencies. tAMPer's performance was assessed on a publicly available protein toxicity benchmark and an AMP hemolysis data we generated. On the latter, tAMPer achieves an F1-score of 68.7%, outperforming the second-best method by 23.4%. On the protein benchmark, tAMPer exhibited an improvement of over 3.0% in the F1-score compared to current state-of-the-art methods. We anticipate tAMPer to accelerate AMP discovery and development by reducing the reliance on laborious toxicity screening experiments.
Using an alignment-free single nucleotide variant prediction framework that leverages integrated variant call sets from the 1000 Genomes Project, we demonstrate accurate ancestry inference predictions on over 600 human genome sequencing datasets, including complete genomes, draft assemblies, and >280 independently-generated datasets. The method presented, ntRoot, infers super-population ancestry along an input human genome in 1h15m or less on 30X sequencing data, and will be an enabling technology for cohort studies.
We present ntRoot, a computationally-lightweight and alignment-free methods for inferring human super-population-level global and local ancestry from whole genome assemblies or raw sequencing data types, and demonstrate its utility on over 300 datasets.
The rocky reefs of British Columbia’s (BC) coast are a productive ecosystem, home to 38 rockfish species (Genus: Sebastes) that are culturally and economically important. Quantitatively assessing rockfish populations is vital to support conservation and stock assessment needs. Self-contained underwater breathing apparatus (SCUBA) diving surveys are a commonly used monitoring method in BC. However, this resource-intensive approach is challenging, particularly for cryptic or deeper species. Herein, we compared environmental DNA (eDNA) detection methods with SCUBA diving surveys to capture overall rockfish biodiversity. We employed two eDNA methods: 1) a targeted quantitative real-time polymerase chain reaction (qPCR) approach to monitor species of particular importance to First Nations collaborators and decision makers, and 2) a metabarcoding approach for assessing community composition using the previously published MiSebastes assay. Both approaches are confounded by the little DNA sequence divergence among species and high sequence variation within species. Overcoming these challenges using a whole mitochondrial approach with the mtGrasp and unikseq pipelines, we generated highly useful eDNA tools. We found that eDNA methods were highly comparable to dive surveys, as both methods indicated a similar ecological reality, including species detections and distributions. Though there are certain species that cannot be distinguished by the MiSebastes assay, eDNA metabarcoding still detected more rockfish species overall. Both eDNA methods show potential for use alongside conventional methods for scalable incorporation into community-based monitoring programs.
Antibiotic resistance is recognized as an imminent and growing global health threat. New antimicrobial drugs are urgently needed due to the decreasing effectiveness of conventional small-molecule antibiotics. Antimicrobial peptides (AMPs), a class of host defense peptides, are emerging as promising candidates to address this need. The potential sequence space of amino acids is combinatorially vast, making it possible to extend the current arsenal of antimicrobial agents with a practically infinite number of new peptide-based candidates. However, mining naturally occurring AMPs, whether directly by wet lab screening methods or aided by bioinformatics prediction tools, has its theoretical limit regarding the number of samples or genomic/transcriptomic resources researchers have access to. Further, manually designing novel synthetic AMPs requires prior field knowledge, restricting its throughput. In silico sequence generation methods are gaining interest as a high-throughput solution to the problem. Here, we introduce AMPd-Up, a recurrent neural network based tool for de novo AMP design, and demonstrate its utility over existing methods. Validation of candidates designed by AMPd-Up through antimicrobial susceptibility testing revealed that 40 of the 58 generated sequences possessed antimicrobial activity against Escherichia coli and/or Staphylococcus aureus. These results illustrate that AMPd-Up can be used to design novel synthetic AMPs with potent activities.
In recent years, the landscape of reference-grade genome assemblies has seen substantial diversification. With such rich data, there is pressing demand for robust tools for scalable, multi-species comparative genomics analyses, including detecting genome synteny, which informs on the sequence conservation between genomes and contributes crucial insights into species evolution. Here, we introduce ntSynt, a scalable utility for computing large-scale multi-genome synteny blocks using a minimizer graph-based approach. Through extensive testing utilizing multiple ∼3 Gbp genomes, we demonstrate how ntSynt produces synteny blocks with coverages between 79–100% in at most 2h using 34 GB of memory, even for genomes with appreciable (>15%) sequence divergence. Compared to existing state-of-the-art methodologies, ntSynt offers enhanced flexibility to diverse input genome sequences and synteny block granularity. We expect the macrosyntenic genome analyses facilitated by ntSynt will have broad utility in generating critical evolutionary insights within and between species across the tree of life. ### Competing Interest Statement The authors have declared no competing interest.
Objectives Antibiotic resistance is a rising global threat to human health and is prompting researchers to seek effective alternatives to conventional antibiotics, which include antimicrobial peptides (AMPs). Recently, we have reported AMPlify, an attentive deep learning model for predicting AMPs in databases of peptide sequences. In our tests, AMPlify outperformed the state-of-the-art. We have illustrated its use on data describing the American bullfrog ( Rana [Lithobates] catesbeiana ) genome. Here we present the model files and training/test data sets we used in that study. The original model (the balanced model) was trained on a balanced set of AMP and non-AMP sequences curated from public databases. In this data note, we additionally provide a model trained on an imbalanced set, in which non-AMP sequences far outnumber AMP sequences. We note that the balanced and imbalanced models would serve different use cases, and both would serve the research community, facilitating the discovery and development of novel AMPs. Data description This data note provides two sets of models, as well as two AMP and four non-AMP sequence sets for training and testing the balanced and imbalanced models. Each model set includes five single sub-models that form an ensemble model. The first model set corresponds to the original model trained on a balanced training set that has been described in the original AMPlify manuscript, while the second model set was trained on an imbalanced training set.
Abstract Motivation K-mer hashing is a common operation in many foundational bioinformatics problems. However, generic string hashing algorithms are not optimized for this application. Strings in bioinformatics use specific alphabets, a trait leveraged for nucleic acid sequences in earlier work. We note that amino acid sequences, with complexities and context that cannot be captured by generic hashing algorithms, can also benefit from a domain-specific hashing algorithm. Such a hashing algorithm can accelerate and improve the sensitivity of bioinformatics applications developed for protein sequences. Results Here, we present aaHash, a recursive hashing algorithm tailored for amino acid sequences. This algorithm utilizes multiple hash levels to represent biochemical similarities between amino acids. aaHash performs ∼10× faster than generic string hashing algorithms in hashing adjacent k-mers. Availability and implementation aaHash is available online at https://github.com/bcgsc/btllib and is free for academic use.