Abstract Background Oxford Nanopore Technologies (ONT) sequencing is increasingly used for whole-genome sequencing (WGS) across a wide range of applications. However, the platform has evolved rapidly through updates to flow cell chemistry and basecalling algorithms, altering the characteristics of the resulting sequencing data. Read simulators provide synthetic datasets with known ground truth, enabling controlled development and evaluation of methods. However, many existing simulators were developed for earlier versions of ONT sequencing or use generic long-read assumptions, and their realism for contemporary ONT data is unclear. Results We benchmarked six ONT-compatible read simulators (Badread, LongISLND, lrsim, NanoSim, PBSIM3 and SimLoRD) using a microbial genome reference and ONT R10.4.1 reads as the empirical standard. Each tool was configured to maximise realism, including training on empirical reads when supported. We compared simulated and real datasets with respect to read length, read accuracy, FASTQ quality scores and sequence error profiles. No simulator reproduced all metrics of the real data well. PBSIM3 most closely reproduced read length, read accuracy and FASTQ quality scores, making it a strong simulator for broad read-level realism. However, it did not capture important features of the real error profile, including context-dependent substitution rates and homopolymer-length errors. Badread and LongISLND better reproduced some aspects of the error profile, but showed other departures from the real data. Conclusion PBSIM3 is a good general-purpose choice for many ONT WGS simulation tasks because it reproduced several key read-level properties well. However, Badread or LongISLND may be preferable for applications where error structure is more important. No evaluated tool was realistic across all tested metrics, highlighting a gap for improved long-read simulators.
ABSTRACT Enteric fever is endemic to many low- and middle-income countries (LMICs), particularly those in sub-Saharan Africa, South and South-East Asia. The causative agents are typhoidal serovars of Salmonella enterica, including Typhi ( S. Typhi) and Paratyphi A (SPA). There are no vaccines currently licensed for SPA, leaving antimicrobials as the only therapeutic option. Multi-drug resistance (MDR) S. Typhi is increasingly prevalent, but to date has not been detected in SPA. In Australia, cases of SPA are notifiable. Here we report on the genomic epidemiology of 208 cases of SPA in returned travellers to Australia, and their close contacts, from 2018 to 2025. A total of 15 unique genotypes were detected, and these were correlated with geographical regions of reported travel. There was a low incidence of antimicrobial resistance with only a single isolate carrying acquired resistance genes. Mutations in quinolone resistance determining regions were common across the genotypes, detected in 95.7% of isolates. A single isolate in a traveller returning from India was resistant to several first line antibiotics including: ampicillin, amoxicillin plus clavulanic acid, ceftriaxone, azithromycin and ciprofloxacin. The isolate carried a plasmid encoding an extended spectrum beta-lactamase ( bla CTX-M-231 ), two macrolide resistance genes ( mphA and ermB ) and a quinolone resistance gene ( qnrS1 ). Elements of the pangenome were explored, with stable maintenance of small plasmids encoding hypothetical proteins detected in four genotypes. Copy number variation in genes encoding surface antigen biosynthesis genes were detected in six genotypes. These biosynthesis genes are targets for one of the two SPA vaccines in development, and the potential variation in surface antigens could have implications for vaccine efficacy. Linking epidemiological data with genomic studies of SPA provides an opportunity to improve understanding of the emergence, spread and risk of drug-resistant SPA infections, and to better inform empirical treatment guidelines in returned travellers. AUTHOR SUMMARY Enteric fever is a disease characterised by a prolonged fever, fatigue and diarrhoea. The burden of disease is disproportionately experienced in children <5years in low-and middle-income countries. One of the causative agents is Salmonella enterica serovar Paratyphi A (SPA). There is no licensed vaccine for SPA, leaving antibiotics as the only treatment option. Increasing antimicrobial resistance (AMR) in SPA is of concern. Enteric fever is not endemic in Australia but can be acquired by individuals travelling to high-risk regions, such as Sub-Saharan Africa, South and South-East Asia. Using genomics on SPA isolates collected from returned travellers allows for informal surveillance of disease across a range of geographical regions. Rates of AMR were low, however there was a single case with a multi-drug-resistant profile. Concerningly this resistance could be easily spread between pathogens, highlighting the need to continued surveillance of this disease. Further, differences in genes involved in surface antigens was detected, which may have implications for the development of vaccines.
Escherichia coli resistant to third-generation cephalosporins (3GCs) is a WHO priority pathogen due to its antimicrobial resistance (AMR). In high-income countries, such as Australia, 3GC-resistant E. coli are a common cause of extra-intestinal infections in both healthcare and community settings. Long-term targeted surveillance efforts of AMR in E. coli routinely identify E. coli as a leading pathogen in bacteraemic infections. To date, there has been limited detailed genomic analysis of the drug-resistant E. coli circulating in Australian clinical settings. Here, we sought to explore the genomic diversity of 3GC-resistant isolates (mediated by extended-spectrum beta-lactamase or AmpC), collected from four hospital networks in Melbourne, Australia. We establish the population structure, identifying ten main lineages in addition to multiple other sequence types, demonstrating 3GC resistance has emerged in multiple genetic backgrounds. We show diversity of accessory genome features, including surface antigens, AMR and plasmid profiles. A total of 117 serotypes and 47 capsular loci were detected, with diversity observed within and between main lineages. We identified 17 unique 3GC resistance mechanisms disseminated across the E. coli population, which co-occurred in different combinations of AMR genes and plasmid replicons. We explored the use of genomic clustering as an approach to detect different population dynamics, identifying 99 clusters of which only 15 had more than 5 isolates. This study provides a comprehensive snapshot of these drug-resistant E. coli in Australia over this time period and will serve as a baseline for future studies of clinical and community drug-resistant isolates in Australia.
Abstract Shigella flexneri is the leading causative agent of shigellosis globally. The public health threat posed by S. flexneri is compounded by its emergence as a sexually transmissible infection, importance of international travel in driving dissemination, and the increasing prevalence of antimicrobial resistance (AMR). A rapid and robust computational method is needed to enhance genomic surveillance and systematically explore features of the population structure of this WHO priority pathogen, which is scalable and readily implementable across jurisdictions, particularly as vaccine development efforts are underway. Here, we present Flex-It, a genomic framework and genotyping scheme implemented in Mykrobe for S. flexneri serotypes 1-5, X & Y, compatible with previous approaches used to describe S. flexneri’s population structure. To develop Flex-It, we curated a retrospective dataset of 5,819 publicly available S. flexneri genomes. We characterised the global population structure for S. flexneri , exploring geographical and temporal traits, and showed the granular diversity of AMR and serotype profiles. We applied Flex-It to >13,000 genomes routinely generated by public health laboratories from Australia, the UK and the USA across a ten-year period. We found significant genotype diversity in all three locations, with the emergence of genotypes with converged resistance to all major drugs currently used for treatment. Flex-It provides an open-source, novel genotyping method that rapidly characterises S. flexneri and its ciprofloxacin resistance determinants in <1 minute from both short and long whole-genome sequencing reads. Flex-It provides the community with a standardised nomenclature to monitor the emergence and spread of S. flexneri lineages.
Abstract Background In Australia, the burden of shigellosis is predominantly in returning travellers or in men who have sex with men (MSM). Here, we combine genomic data with comprehensive epidemiological data on sexual exposure and international travel to explore population dynamics of Shigella sonnei and the expansion of multi-drug resistant (MDR) and extensively drug-resistant (XDR) sub-lineages. Methods A population-level study of all cultured Shigella sonnei isolates in the state of Victoria, Australia, was undertaken between January 2002 and December 2024. Antimicrobial susceptibility testing, whole-genome sequencing, and bioinformatic analyses of 1,305 Shigella sonnei isolates were performed at the Microbiological Diagnostic Unit Public Health Laboratory. Enhanced metadata on source attribution including travel and sexual exposure were collected through surveillance forms or by interviews. Results This study highlights significant shifts in Shigella sonnei cases in Victoria from sensitive strains to MDR and then XDR, particularly in the MSM-associated groups but also associated with a large point source outbreak. We describe an historical pattern of shifting genotype prevalence, and replacement to more varied and higher proportions of antimicrobial resistance over the last decade, resulting in the establishment of two distinct but highly concerning XDR sub-lineages within Victoria. Conclusions Our genomic-epidemiological analyses highlight that drug-resistant Shigella sonnei remains an ongoing public health threat, and the importance of ongoing surveillance. We determined local evolutionary trajectories and identified expanding sub-lineages that informed shifts in clinical management and antimicrobial recommendations over time, including the use of azithromycin and carbapenems. Placing these local dynamics within the broader global epidemiology, we link how regional evolution interconnects with international dissemination, proving valuable context for guiding local, national and global strategies for prevent outbreaks and antimicrobial resistance.
ColV/ColBM and ColIa/senB F virulence plasmids feature prominently in Escherichia coli associated with urinary tract and bloodstream infections globally. Australian-sourced E. coli that carry these plasmids were examined among 5,471 isolates (3,316 sequenced by the APG and AusGEM programmes; 2,155 from public databases) spanning years 1986-2020 from humans (n=2,996/5,471; 54.8%), wild animals (n=870/5,471; 15.9%), livestock (n=649/5,471; 11.9%), companion animals (n=375/5,471; 6.9%), environmental sources (n=292/5,471; 5.3%) and food (n=289/5,471; 5.3%). Putative plasmid reconstruction, assisted by a plasmid database comprising 23,700 complete plasmid sequences, identified 22,534 putative plasmids of which 21,814 (96.8%) represented 547 known plasmid clusters. E. coli harbouring ColV-associated putative plasmids, particularly plasmid cluster AA176 [F replicon sequence type (RST): F18:A-:B1 (repFII-18:repFIA-null:repFIB-1)] was identified among phylogenetically diverse strains from humans, livestock, particularly poultry, and food. Closely related isolates, defined as ≤10 core-genome multilocus sequence type allelic distance, that carried either ColV or ColIa/senB-associated putative (F) plasmids were identified across multiple sources and diverse phylogenetic backgrounds. ColIa/senB-associated putative (F) plasmid clusters AA337 (RST: F29:A-:B10) and AA171 (RST: F2:A1:B20) were associated with phylogenetically closely related isolates from humans, wild animals and companion animals, but their absence in E. coli sourced from food and livestock was notable. E. coli carrying ColV and ColIa/senB plasmids frequently exhibit genotypic multidrug resistance, many with critically important antimicrobial resistance genes, highlighting their role in the evolution of clinically problematic lineages. Our study has important epidemiological considerations for understanding the spread of extraintestinal pathogenic and hybrid E. coli lineages across the One Health spectrum.
Diarrhoeal pathogens impose a substantial global health burden, disproportionately affecting low- and middle-income countries (LMICs). However, in these settings, health-seeking behaviours, suboptimal microbiological capacity, and challenges in establishing genomics capacity constrain effective surveillance, including surveillance of antimicrobial resistance (AMR). In contrast, high-income countries routinely generate and share large volumes of diarrhoeal pathogen genomes through established systems, with a significant proportion originating from travellers returning from LMICs. These data reveal strong geographical structuring of lineages and clinically relevant AMR patterns, demonstrating untapped potential to support improvements in geographically granulated surveillance to support antimicrobial treatment recommendations. In this opinion article, we outline the potential to integrate traveller-derived microbial genomic data into LMIC public health decision-making and highlight the scientific, ethical, practical, and governance considerations for implementation.
Mycoplasma genitalium is a sexually transmitted infection where antimicrobial resistance poses marked challenges to patient care and public health. Despite this, genetic studies of M. genitalium have remained limited due to the fastidious culture requirements. To address this, we developed a targeted and culture-independent sequencing approach to enable the genomic study. A global data set of 220 M. genitalium genomes was compiled, comprising contemporary genomes from Victoria, Australia, and previously publicly available genomes. Major phylogenetic lineages were identified of which multiple show evidence of transmitted antimicrobial resistance. Macrolide resistance was pervasive and has reached concerning rates both in Australia and globally (160/220; 72.7%), while known fluoroquinolone resistance–associated mutations correlated with treatment failure (aOR 22.67; 95% CI 4.08-168.4). Of genomes with metadata, most were derived from gay bisexual and other men who have sex with men (55/113; 48.7%), while one M. genitalium lineage showed a higher percentage of women (17/63; 27%) and heterosexual men (21/43; 48.8%). Genome-based typing schemes are shown to offer higher resolution and robustness when compared with single-locus approaches. Together these findings clarify M. genitalium population structure and drivers of resistance, with genomics adding value in identifying effective treatments and guiding public health responses.
Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a species has been used to identify associations between the microbiome and disease. However, this approach overlooks intra-species genetic variation and is susceptible to spurious correlations arising from the compositional nature of abundance data and microbial load. Fast, k-mer-based algorithms can now accurately estimate strain-level Average Nucleotide Identity (ANI) in metagenomes. Despite its value as an orthogonal metric for strain-level analysis, methods for conducting ANI-based association studies remain limited. To address this, we developed StrainSpy, a statistical algorithm that identifies associations between containment ANI and variables of interest across a wide range of study designs, including longitudinal and multi-cohort designs. Re-analysis of a study examining gut microbiota recovery in 12 healthy adults following antibiotic exposure revealed novel strain-level associations, including a reduction in strain-level diversity despite species persistence. Applying StrainSpy to a multi-cohort analysis of 3,414 colorectal cancer metagenomes identified novel strain-level associations with colorectal cancer. However, in a separate collection of microbiome-immunotherapy studies, no individual strain was consistently associated across cohorts. Importantly, across both datasets, StrainSpy informed containment ANI-based machine learning models achieved comparable accuracy to traditional abundance-based methods. StrainSpy is publicly available as an R package github.com/gtonkinhill/strainspy .
Phylodynamic analyses infer epidemiological parameters from pathogen genome sequences for enhanced genomic surveillance in public health. Pathogen genome sequences and their associated sampling dates are the essential data in every analysis. However, sampling dates are usually associated with hospitalisation or testing and can sometimes be used to identify individual patients, posing a threat to patient confidentiality. To lower this risk, sampling dates are often given with reduced date-resolution to the month or year, which can potentially bias inference. Here, we introduce a practical guideline on when date-rounding biases the inference of epidemiologically important parameters across a diverse range of empirical and simulated datasets. We show that the direction of bias varies for different parameters, datasets, and tree priors, while compounding with lower date-resolution and higher substitution rates. We also find that bias decreases for datasets with longer sampling intervals, implying that our guideline is most applicable to emerging datasets. We conclude by discussing future solutions that prioritise patient confidentiality and propose a method for safer sharing of sampling dates that translates them them uniformly by a random number.
The 'silent pandemic' of antimicrobial resistance (AMR) represents a significant global public health threat. AMR genes in bacteria are often carried on mobile elements, such as plasmids. The horizontal movement of plasmids allows AMR genes and resistance to key therapeutics to disseminate in a population. However, the quantification of the movement of plasmids remains challenging with existing computational approaches. Here, we introduce a novel method that allows us to reconstruct and quantify the movement of plasmids in bacterial populations over time. To do so, we model chromosomal and plasmid DNA co-evolution using a joint coalescent and plasmid transfer process in a Bayesian phylogenetic network approach. This approach reconstructs differences in the evolutionary history of plasmids and chromosomes to reconstruct instances where plasmids likely move between bacterial lineages while accounting for parameter uncertainty. We apply this new approach to a five-year dataset of Shigella, exploring the plasmid transfer rates of five different plasmids with different AMR and virulence profiles. In doing so, we reconstruct the co-evolution of the large Shigella virulence plasmid with the chromosome DNA. We quantify higher plasmid transfer rates of three small plasmids that move between lineages of Shigella sonnei. Finally, we determine the recent dissemination of a multidrug-resistant plasmid between S. sonnei and S. flexneri lineages in multiple independent events and through steady growth in prevalence since 2010. This approach has a strong potential to improve our understanding of the evolutionary dynamics of AMR-carrying plasmids as they are introduced, circulate, and are maintained in bacterial populations.
Serovars of Salmonella are significant bacterial pathogens and are leading contributors to the global burden of diarrhoeal disease. Salmonella pathogenicity islands (SPIs) are essential for the survival and success of this genus, enabling colonisation, invasion, and survival in hostile environments. While genomics has transformed efforts to understand the evolution, dissemination, and antimicrobial resistance of members, its use to explore virulence determinants that contribute to the pathogenicity of specific organisms and severity of infection remains varied. Here, we discuss the importance of SPIs to the evolution of Salmonella, the implications in the shift of identification of SPIs from molecular microbiology to genomic-based approaches, and examine current efforts to explore the distribution and prevalence of SPIs in large-scale datasets of Salmonella genomes.
Typhoid fever results from systemic infection with Salmonella enterica serovar Typhi (Typhi) and causes 10 million illnesses annually. Disease control relies on prevention (water, sanitation, and hygiene interventions or vaccination) and effective antimicrobial treatment. Antimicrobial-resistant (AMR) Typhi lineages have emerged and become established in many parts of the world. Knowledge of local pathogen populations informed by genomic surveillance, including of lineages (defined by the GenoTyphi scheme) and AMR determinants, is increasingly used to inform local treatment guidelines and to inform vaccination strategy. Current tools for genotyping Typhi require multiple read alignment or assembly steps and have not been validated for analysis of data generated with Oxford Nanopore Technologies (ONT) long-read sequencing devices. Here, we introduce Typhi Mykrobe, a command line software tool for rapid genotyping of Typhi lineages, AMR determinants, and plasmid replicons direct from sequencing reads. We validated Typhi Mykrobe lineage genotyping by comparison with the current standard read mapping-based approach and demonstrated 99.8 https://github.com/typhoidgenomics/genotyphi . Typhi Mykrobe provides rapid and sensitive genotyping of Typhi genomes direct from Illumina and ONT reads, although lower accuracy was observed for R9 ONT data. It demonstrated accurate assignment of GenoTyphi lineage, detection of AMR determinants and prediction of corresponding AMR phenotypes, and identification of plasmid replicons.
The Coxiellaceae bacterial family, within the order Legionellales, is defined by a collection of poorly characterized obligate intracellular bacteria. The zoonotic pathogen and causative agent of human Q fever, Coxiella burnetii, represents the best-characterized member of this family. Coxiellaceae establish replicative niches within diverse host cells and rely on their host for survival, making them challenging to isolate and cultivate within a laboratory setting. Here, we describe a new genus within the Coxiellaceae family that has been previously shown to infect economically significant freshwater crayfish. Using culture-independent long-read metagenomics, we reconstructed the complete genome of this novel organism and demonstrate that the species previously referred to as Candidatus Coxiella cheraxi represents a novel genus within this family, herein denoted Candidatus Paracoxiella cheracis. Interestingly, we demonstrate that Candidatus P. cheracis encodes a complete, putatively functional Dot/Icm type 4 secretion system that likely mediates the intracellular success of this pathogen. In silico analysis defined a unique repertoire of Dot/Icm effector proteins and highlighted homologs of several important C. burnetii effectors, including a homolog of CpeB that was demonstrated to be a Dot/Icm substrate in C. burnetii.IMPORTANCEUsing long-read sequencing technology, we have uncovered the full genome sequence of Candidatus Paracoxiella cheracis, a pathogen of economic importance in aquaculture. Analysis of this sequence has revealed new insights into this novel member of the Coxiellaceae family, demonstrating that it represents a new genus within this poorly characterized family of intracellular organisms. Importantly, the genome sequence reveals invaluable information that will support diagnostics and potentially both preventative and treatment strategies within crayfish breeding facilities. Candidatus P. cheracis also represents a new member of Dot/Icm pathogens that rely on this system to establish an intracellular niche. Candidatus P. cheracis possesses a unique cohort of putative Dot/Icm substrates that constitute a collection of new eukaryotic cell biology-manipulating effector proteins.
The critical role of plasmids, particularly in the dissemination of AMR and virulence in nosocomial pathogens like Klebsiella pneumoniae , underpins the need for robust and scalable tools for plasmid identification and reconstruction from large-scale datasets of short-read sequence data that are routinely generated in clinical and public health settings. Here, we sought to evaluate all available tools to determine which are best suited for reconstructing the plasmidome of K. pneumoniae and related species from the species complex (KpSC). We used a publicly available dataset of 568 diverse KpSC isolates that had high-quality short-read Illumina data and matched complete genomes generated from hybrid assemblies incorporating additional long-read sequence data. This allowed us to investigate which tool perform best at recovering plasmid sequences when only short-read data are available. None of the tools tested offered a comprehensive or consistently reliable solution to assembling plasmids from short-read sequence data of KpSC. Among the six tools that were benchmarked, performance varied across total runtime, RAM usage, prediction accuracy and sensitivity. Future tools developed in this space should offer meaningful advancements over existing tools and be rigorously evaluated using large, standardised bacterial datasets that reflect the diversity and complexity of plasmid content to ensure comparability across benchmarking studies. ### Competing Interest Statement The authors have declared no competing interest. Australian Research Council, DE250100677 National Health and Medical Research Council, https://ror.org/011kf5r70, GNT1195210, GNT2009163
BACKGROUND:Non-typhoidal Salmonella is a globally important bacterial pathogen, typically associated with foodborne gastrointestinal infection. Some non-typhoidal Salmonella serovars can also colonise typically sterile sites in people to cause invasive non-typhoidal Salmonella disease. Salmonella enterica serovar Panama is responsible for a substantial number of cases of human bloodstream infection, but despite its global dissemination, numerous outbreaks, and a reported association with invasive non-typhoidal Salmonella disease, S enterica serovar Panama (S Panama) is understudied. We aimed to describe the genomic epidemiology and evolutionary history of S Panama to provide a vital baseline of understanding for this globally important serovar. METHODS:In this genomic epidemiology study, we analysed S Panama genomes derived from historical collections, national surveillance datasets, and publicly available epidemiological and whole-genome sequencing data which span the years 1931-2019. Maximum likelihood and Bayesian phylodynamic approaches were used to investigate population structure and evolutionary history and to infer geotemporal dissemination. A combination of different bioinformatic approaches with short-read and long-read data were used to characterise geographical and clade-specific trends in antimicrobial resistance (AMR) and genetic markers for invasiveness. FINDINGS:We analysed 836 S Panama genomes, of which 559 (67%) were sequenced as part of this study. The collection represents all inhabited continents and includes isolates collected between 1931 and 2019. We identified the presence of four geographically linked S Panama clades (C1 [ie, the Latin America and the Caribbean clade; n=338], C2 [ie, the European clade; n=124], C3 [ie, the Martinique clade; n=131], and C4 [ie, the Asia and Oceania clade; n=104]) and regional trends in AMR profiles. Most isolates (715 [86%] of 836) were pan-susceptible to antibiotics and belonged to clades circulating in Latin America and the Caribbean (64%, n=458). Most antibiotic-resistant isolates in our collection (113 [93%] of 121) fell within clades C4 (ie, the Asia and Oceania clade) and C2 (ie, the European clade), the latter of which had the highest invasiveness index values based on the conservation of 196 extraintestinal predictor genes. INTERPRETATION:This first large-scale phylogenetic analysis of S Panama has revealed important information about the population structure, AMR, global ecology, and genetic markers of invasiveness of the identified genomic subtypes. Our findings provide an important baseline for understanding S Panama infection. The presence of multidrug-resistant clades with elevated invasiveness index values should be monitored through ongoing surveillance, as such clades could pose an increased public health risk. FUNDING:UK Research and Innovation Global Challenges Research Fund and Biotechnology and Biological Sciences Research Council, UK Medical Research Council, Wellcome Trust, John Lennon Memorial Scholarship, Institut Pasteur, Santé publique France, Fondation Le Roch-Les Mousquetaires, Investissement d'Avenir Programme, and Australian National Health and Medical Research Council.
Shigellosis is a leading cause of diarrheal mortality worldwide. Shigella boydii is one of four Shigella species that contributes to this burden, however studies on S. boydii are limited. Here we combined epidemiological and genomic data to better understand S. boydii circulating both in Australia and globally. Between 1991 and 2019, there were 294 cases of S. boydii infections notified to the National Notifiable Diseases Surveillance System by Australian states and territories, with an increasing trend in notifications observed from 2013. Of cases whose place of acquisition was known, 54% (111/206) were acquired overseas, mainly from South-East Asia (57%; 63/111). Our genomic analysis included 250 S. boydii isolates: 44 from Victoria, Australia spanning 22 years (2001-2022) and 206 international isolates spanning 91 years (1930-2020). Phylogenomic analyses identified five major S. boydii phylogenetic lineages circulating globally. The Australian isolates were distributed across all five lineages, but the highest proportion was in Lineage 3. Antimicrobial resistance was common in both international and Australian isolates with > 60% of isolates classified as multi-drug-resistant. Resistance to the main clinically relevant antimicrobials was rare in S. boydii. Ciprofloxacin resistance was detected in seven S. boydii, however reduced susceptibility to ciprofloxacin was detected in 56 isolates and found in both Australian and international data. Importantly, resistance mechanisms to third-generation cephalosporins and macrolides were also detected. This study is the largest genomic analysis of S. boydii to date, providing insights into the population structure, epidemiology and emerging AMR threats in this neglected Shigella species.
Phylogenetic analyses are crucial for understanding microbial evolution and infectious disease transmission. Bacterial phylogenies are often inferred from SNP alignments, with SNPs as the fundamental signal within these data. SNP alignments can be reduced to a ‘strict core’ by removing those sites that do not have data present in every sample. However, as sample size and genome diversity increase, a strict core can shrink markedly, discarding potentially informative data. Here, we propose and provide evidence to support the use of a ‘soft core’ that tolerates some missing data, preserving more information for phylogenetic analysis. Using large datasets of Neisseria gonorrhoeae and Salmonella enterica serovar Typhi, we assess different core thresholds. Our results show that strict cores can drastically reduce informative sites compared to soft cores. In a 10 000-genome alignment of Salmonella enterica serovar Typhi, a 95% soft core yielded ten times more informative sites than a 100% strict core. Similar patterns were observed in N. gonorrhoeae. We further evaluated the accuracy of phylogenies built from strict- and soft-core alignments using datasets with strong temporal signals. Soft-core alignments generally outperformed strict cores in producing trees displaying clock-like behaviour; for instance, the N. gonorrhoeae 95% soft-core phylogeny had a root-to-tip regression R 2 of 0.50 compared to 0.21 for the strict-core phylogeny. This study suggests that soft-core strategies are preferable for large, diverse microbial datasets. To facilitate this, we developed Core-SNP-filter (https://github.com/rrwick/Core-SNP-filter), an open-source software tool for generating soft-core alignments from whole-genome alignments based on user-defined thresholds.