The Knosp grading system classifies the extent of cavernous sinus invasion (CSI) in pituitary adenomas (PAs). Knosp grade is predictive of surgical remission rates and is used in decision-making for the management of these adenomas. This study evaluates the rate and accuracy of Knosp grade reporting of PAs by radiologists. This is a retrospective observational study of 100 consecutive patients with pituitary macroadenomas who underwent pituitary MR imaging between March 2023 and March 2025. The rate of CSI reporting in radiologist reports of the scans was determined, and the reported grade was compared with the Knosp grade calculated by the study panel. Radiologist reports contained a specific comment regarding CSI, or its absence, in 73/100 (73
The chemotherapeutic agent 5-fluorouracil (5-FU), widely used in the treatment of head and neck cancer (HNC), also exhibits broad antimicrobial activity, yet fluoropyrimidine resistance within HNC-associated microbiota remains poorly characterised. We assessed 5-FU susceptibility and resistance-associated genomic features in 101 Streptococcus isolates obtained from tumour tissue and oral swabs of 31 HNC patients using minimum inhibitory concentration assays integrated with whole-genome sequencing and pangenome analysis. Resistance to 5-FU was prevalent across multiple Streptococcus species and was primarily associated with species identity rather than resistance phenotype or anatomical niche. Resistant isolates showed functional convergence in pathways related to multidrug efflux, stress response, DNA repair, cell-envelope biosynthesis, and virulence, whereas sensitive isolates were enriched for genes involved in core metabolism, nutrient acquisition, and colonisation. Species-resolved analyses revealed heterogeneous, polygenic resistance architectures rather than conserved resistance determinants. Together, these findings suggest that 5-FU exposure may act as an ecological selective pressure shaping microbial functional potential within tumour- and oral communities in HNC.
Bacteriophage (phage) genome annotation is essential for understanding their functional potential and suitability for use as therapeutic agents. Here, we introduce Phold, an annotation framework utilizing protein structural information that combines the ProstT5 protein language model and structural alignment tool Foldseek. Phold assigns annotations using a database of over 1.36 million predicted phage protein structures with high-quality functional labels. Benchmarking reveals that Phold outperforms existing sequence-based homology approaches in functional annotation sensitivity whilst maintaining speed, consistency, and scalability. Applying Phold to diverse cultured and metagenomic phage genomes shows it consistently annotates over 50% of genes on an average phage and 40% on an average archaeal virus. Comparisons of phage protein structures to other protein structures across the tree of life reveal that phage proteins commonly have structural homology to proteins shared across the tree of life, particularly those that have nucleic acid metabolism and enzymatic functions. Phold is available as free and open-source software at https://github.com/gbouras13/phold.
BACKGROUND:Chronic rhinosinusitis (CRS) is frequently associated with polymicrobial biofilms involving Staphylococcus aureus and Pseudomonas aeruginosa. Interactions between these organisms are thought to influence disease severity, but the epithelial effects of exoproteins derived from patient-matched cocultures remain poorly defined. METHODS:Clinical isolates of S. aureus and P. aeruginosa (n = 3 each) co-isolated from three CRS patients were cultured in a Transwell system as same-patient or cross-patient pairs. Cell-free exoproteins were applied to primary human nasal epithelial cells. Epithelial repair was assessed using a scratch assay, while cytotoxicity, oxidative stress, and inflammatory responses were evaluated by lactate dehydrogenase release, intracellular reactive oxygen species measurement, and interleukin-6 secretion. Exoprotein profiles were further characterized using data-independent acquisition proteomics. RESULTS:Exoproteins derived from same-patient cocultures consistently impaired epithelial wound closure compared with P. aeruginosa monocultures, although the timing of inhibition varied among patients. These exoproteins also induced greater epithelial cytotoxicity, elevated intracellular reactive oxygen species levels, and increased interleukin-6 secretion compared with monocultures or cross-patient cocultures. In contrast, cross-patient pairings produced epithelial responses similar to monoculture conditions. Proteomic analysis indicated that patient origin was a major determinant of exoproteome organization, with same-patient cocultures showing increased abundance of proteins associated with redox balance and metabolic regulation. CONCLUSIONS:Patient-matched S. aureus-P. aeruginosa interactions were associated with increased epithelial stress and inflammatory responses, together with differences in the exoproteomic profile. These findings suggest that host-specific bacterial interactions may contribute to virulence and possibly recalcitrance in CRS.
Achromobacter species are emerging multidrug-resistant (MDR) pathogens in people with cystic fibrosis. Their increasing resistance has grown an interest in phage therapy as an alternative treatment strategy. However, the factors governing phage susceptibility remain poorly understood, thereby limiting the rational selection of phage candidates. Using 15 strictly lytic Achromobacter phages and 7 clinical cystic fibrosis isolates representing Achromobacter insolitus and Achromobacter xylosoxidans, we demonstrate substantial variation in infection efficiency across all 105 phage-host combinations, variation that could not be discerned from qualitative plaque assays alone. We integrated complete bacterial and phage genomes with quantitative efficiency-of-plating (EOP) assays and lineage-aware Bayesian mixed-effects modelling to show that phage infectivity in Achromobacter is governed predominantly by bacterial lineage and strain identity, accounting for 90% of total variance in log-normalised EOP, with individual strains varying substantially in permissiveness irrespective of species membership. After accounting for this lineage structure, no individual defence system, antimicrobial resistance gene class, or phage tail cluster retained a statistically significant independent or interaction association with infectivity. Together, these findings demonstrate that bacterial strain identity is the primary driver of infection outcome. Host defence systems and phage tail-associated genes remain biologically plausible contributors; their independent effect could not be resolved after accounting for lineage structure, indicating that infection outcomes are largely strain-dependent. This work shifts the question from which individual traits predict infection to how strain lineage and specific host-phage combinations jointly determine infectivity and argues that quantitative phenotyping of individual phage-host pairs is essential for guiding phage candidate selection and supporting rational cocktail design against multidrug-resistant Achromobacter infections in cystic fibrosis.
Abstract Chronic Rhinosinusitis (CRS) is a common chronic inflammation of the paranasal sinus mucosa. Staphylococcus aureus contributes to its severity through biofilm formation. In this study, we isolated eight sequential methicillin-resistant S. aureus (MRSA) isolates from a patient with severe CRS over a period of 672 days (T1-T8). The isolates were phenotypically and genomically characterised, and the extracellular biofilm proteome analysed. We identified an accumulation of mutations that included the acquisition of an IS21 family insertion sequence inactivating the icaR gene and nucleotide variants in various genes including the transcription repair coupling factor ( mfd) . The genomic changes were associated with a switch to a mucoid phenotype from T3 onwards (Day 178), with a significant increase in biofilm-forming capacity and the secretion of multiple enterotoxins. Targeted mutagenesis confirmed mfd is a regulator of strain mucoidy with enhanced biofilm and enterotoxin production. These findings support mfd as a target for novel anti-virulence therapies.
Abstract Viruses are abundant, ancestral and potentially fast-evolving biological entities. As a result, their encoded proteins are diverse and identifying homologous relationships between sequences is as important for phylogeny and functional annotation as it is challenging. Traditional methods group viral proteins by sequence similarity, build HMM profiles for each protein family, and cluster further via profile comparisons. Here, we present an improved framework where HMM sensitivity is boosted by enriching reference virus HMM profiles with tens of millions of metagenomic sequences. This increases diversity within most protein families, raising the diversity index from less than 2 for 92.7% of clusters to a median value of 6. This enrichment of the profiles more than triples the number of homologies detected compared to the raw profiles. First-step clusters are then grouped more effectively using these relationships and further unified via structural predictions and comparisons. The sequence-enrichment strategy excels at linking small proteins, while structures better connect highly structured ones like tail and head proteins. Applied to 1.42 million proteins, our method yields 56,560 families—far fewer than 200,018 (sequence-based) or 135,048 (raw HMM)—revealing that prior approaches vastly overestimated viral protein diversity. The strategy of enriching the diversity of sequences of interest with external sequences, combined with the complementary use of structural information, highlights deep evolutionary links, offering a more accurate picture of viral protein evolution.
Motivation Assembly graphs are a fundamental data structure used by genome and metagenome assemblers to represent sequences and their overlap information, facilitating the assembler in constructing longer genomic fragments. Apart from their core use in assemblers, assembly graphs have become increasingly important in a range of downstream applications such as metagenomic binning, plasmid detection, viral genome resolution, and haplotype phasing. However, there is a need for a comprehensive tool that allows programmatic access to manipulate assembly graphs (e.g. parse, convert, filter, and analyze) across different assembly graph formats.Results Here we present agtools, an open-source Python framework to manipulate assembly graphs produced by commonly used assemblers. agtools provides a command-line interface for tasks such as assembly graph format conversion, segment filtering, and component extraction. It also exposes a Python package interface to load, query, and analyze assembly graphs from popular genome and metagenome assemblers. This enables streamlined assembly-graph-based analyses that can be integrated into other bioinformatics software and workflows.Availability and implementation The source code of agtools is hosted on GitHub at https://github.com/Vini2/agtools and the documentation is available at https://agtools.readthedocs.io/. agtools can also be installed from Bioconda (https://anaconda.org/bioconda/agtools) and PyPI (https://pypi.org/project/agtools/).
Bacteriophages (phages) play essential roles in microbial systems, yet most phage proteins remain poorly characterised. Protein tertiary and quaternary structure information contributes valuable information about protein function. As many phage proteins function as homooligomers, complexes that consist of multiple identical subunits, there is great interest in computationally predicting their configurations. Here we present a computational framework, the Phage Homomer Level Estimate and Generation Method (PHLEGM) for inferring homooligomeric states directly from the protein sequence by combining AlphaFold-Multimer modelling with inter-subunit interface quality assessment. We proceeded to experimentally validate two out of nine predicted homooligomers using size exclusion chromatography and complementary hydrodynamic techniques. These efforts confirmed our predictions for a dimer and a trimer, highlighting the value of experimentally benchmarked computational predictions and showing the challenges of heterologous phage protein production. Applied to >22,000 phage protein sequences in the PHROGs database, our approach revealed extensive diversity in phage homooligomeric protein complexes. Benchmarking against protein language model-based predictors on a curated reference set of known phage homooligomers demonstrated superior accuracy of our structure-based method, achieving robust performance in classifying protein homooligomeric states, with the highest accuracy observed for trimers and higher-order complexes. These results highlight the value of computational predictions to decipher the complexities of the vast viral sequence space. All predicted complex structures and functional inferences are made publicly available to support structural and functional studies of phage proteins.
Viral metagenomics is an increasingly powerful tool for understanding the function and structure of viruses across the diverse environments of our planet. However, decoding the functional potential of prokaryotic viral metagenomes is extremely challenging. Pharokka, Phold, and Phynteny are complementary open-source prokaryotic viral genome annotation tools that utilize a variety of bioinformatics approaches to maximally annotate viral metagenomes. This article describes a protocol for installing and running these tools on a viral metagenomic dataset, followed by visualization of annotations using our client-side Phold Plot web assembly application. © 2026 The Author(s). Current Protocols published by Wiley Periodicals LLC. Basic Protocol 1: Prokaryotic viral metagenome annotation with Pharokka Basic Protocol 2: Enhanced prokaryotic viral metagenome protein annotation using protein structures with Phold Basic Protocol 3: Further prokaryotic viral metagenome protein annotation using genome synteny and protein language models with Phynteny Basic Protocol 4: Visualization of prokaryotic viral metagenome annotations with Phold Plot web assembly application.
Disruption of the oral microbiome is increasingly implicated in head and neck cancer (HNC), yet the genomic adaptations that commensal bacteria acquire in tumour-associated environments remain unclear. We performed genome-resolved analyses of 101 complete Streptococcus genomes from the tumours and oral cavities of 31 HNC patients. Phylogenomic analysis identified 35 species, including ten novel species belonging to the Mitis group. The Streptococcus genus shared 29 core genes, with analysis of accessory genomes (1.7-2.5 Mbp) showing extensive horizontal gene transfer (HGT), supported by 245 ICE clusters, 82 prophages and 4 plasmid groups. Comparison with 391 published genomes from the oral cavities of healthy individuals showed that tumour-associated isolates exhibited niche-specific expansions of carbohydrate-active enzymes and enrichment of genes involved in sugar transport, thiamine biosynthesis and antimicrobial resistance. Together, these findings reveal distinct HGT-driven genomic remodelling in tumour-associated Streptococcus and provide the first comprehensive genomic resource for examining microbiome adaptation in HNC. Graphical Abstract
Abstract The functional annotation of protein sequences has undergone tremendous progress over recent years, but still too-many protein sequences remain as so-called hypothetical proteins after applying state-of-the-art genome annotation software pipelines. Here, we introduce Baktfold, a new command line software tool for the ultra-sensitive but taxon-independent fast annotation of protein sequences across the microbial tree of life. Baktfold conducts sequential protein structure-based searches against four complementary structure databases. Protein sequences are transformed into Foldseek 3Di tokens via the ProstT5 protein language model and subsequently searched against structure databases via Foldseek. All results are exported in GFF3 and INSDC-compliant flat files as well as comprehensive JSON files facilitating automated downstream analysis 100% interoperable with the popular bacterial annotation tool Bakta. We compared Baktfold’s performance in terms of wallclock runtime and functional annotation of hypothetical proteins from various sources including bacterial and archaeal isolates, plasmids, metagenomic-assembled genomes and micro-eukaryotes. When benchmarked on over three hundred thousand species representatives across the prokaryotic tree of life, Baktfold’s median overall bacterial genome annotation rate is 87.8% compared to 72.9% with Bakta, while Baktfold’s median bacterial annotation rate of remaining hypothetical proteins is 50.1% (n=290258). For archaea, Baktfold’s overall median annotation rate is 71.5% compared to Prokka’s 35.8%, with a median archaeal annotation rate of hypothetical proteins of 68.0% (n=14058), making Baktfold the most sensitive automated archaeal annotation method by far. Baktfold is implemented in Python 3 and runs on MacOS and Linux systems. It is freely available under a MIT license at https://github.com/gbouras13/baktfold . Data Summary Baktfold was developed in Python as a command line application for Linux and MacOS The complete source code and documentation are available on GitHub under an MIT license: https://github.com/gbouras13/baktfold The Baktfold database is hosted at Zenodo ( https://zenodo.org/records/17347516 ) mirrored on HuggingFace ( https://huggingface.co/datasets/gbouras13/baktfold-db ) Baktfold is available via bioconda ( https://anaconda.org/bioconda/baktfold ) and PyPI ( https://pypi.org/project/baktfold/ ) Baktfold can also be run without local installation using Google Colab at https://colab.research.google.com/github/gbouras13/baktfold/blob/main/run_baktfold . ipynb All supplementary code, data and files required to reproduce the results of this manuscript are available at https://github.com/gbouras13/baktfold-analysis (code and small data) and https://zenodo.org/records/19333697 (large data)
The rapid rate of virus discovery renders manual curation by taxonomy experts increasingly impractical, creating a need for reliable software that can reproducibly assign viral contigs to taxa at all fifteen ranks of the virus taxonomy. We led an open community challenge for the computational taxonomic classification of viruses and assembled a dataset of virus sequences combining expert-curated and metagenomic sequences. Seventeen teams contributed a total of thirty-four automated, fully reproducible classification pipelines. Most tools correctly assigned viruses belonging to established species, genera, or families, but viruses that are unclassified at those lower ranks remain challenging. This study provides datasets, open-source software, novel approaches, and recommendations to benchmark computational taxonomic classification of viruses, and support organizing the many viruses discovered in big omics data.
Antimicrobial resistance is a growing global health threat, necessitating alternatives to conventional antibiotics. Bacteriophages, viruses that specifically target bacteria, represent a promising option, and phage-loaded electrospun fibers have recently gained attention as wound dressings for localized phage therapy. However, the influence of phage morphology and scaffold design has been largely overlooked. This study investigates how phage morphology and structure, in conjunction with scaffold design and processing conditions, may influence the biological performance of electrospun scaffolds. A bilayer scaffold was developed comprising a supportive polycaprolactone (PCL)/gelatin (70:30) layer and a polyvinyl alcohol (PVA) top layer loaded with bacteriophages. Two phage types, short-tailed podovirus APTC-SL.1 and long-tailed myovirus APTC-Efa.20, were incorporated into PVA fibers to evaluate their antibacterial activity against Staphylococcus lugdunensis and Enterococcus faecalis, respectively. The fibers were characterized using XRD, FTIR, TGA, optical microscopy, SEM, TEM, wettability analysis, and in vitro degradation tests. Biological assessments included antimicrobial testing, phage viability, and phage release. The bilayer scaffold containing short-tailed phages preserved phage viability and produced clear zones of lysis against S. lugdunensis, with ≈8.15% viability retained after electrospinning and relatively controlled release, whereas long-tailed phages showed no antibacterial activity. These results suggest that phage structure and morphology, together with electrospinning conditions and scaffold architecture, may play an important role in maintaining phage functionality in wound dressing applications, while acknowledging that host–phage interactions may also contribute to the observed differences.
Chronic infections in cystic fibrosis (CF) emerge from gradual ecological transitions in the airway microbiome, yet early predictive markers remain poorly defined. We developed a new autoencoder-based framework that outperforms read-based or metagenome-assembled genome-based analyses at capturing the continuum from health-associated commensals to pathogen-dominated, antibiotic-tolerant communities. This improvement is achieved by integrating taxonomic and functional data from 127 sputum and bronchoalveolar lavage metagenomes from 64 people with CF into latent "Clusters of Phylogeny and Functions" (COPFs). Coupled with gradient-boosted random forests, COPFs predicted Pseudomonas aeruginosa colonisation, multidrug resistance, and impending infection up to a year before clinical detection. The multidrug-resistant P. aeruginosa signature showed the same resistance-mechanism evolution as found in laboratory experiments. The inclusion of eukaryotic markers revealed persistent Aspergillus fumigatus signatures even during culture-negative intervals. Applying our South Australian-trained model to over 1,000 global metagenomes from 22 independent CF datasets, we achieved 94% accuracy in predicting P. aeruginosa status across platforms and geographies, validating the model's universal utility. Our results demonstrate that combining datasets with deep learning reveals conserved ecological and metabolic mechanisms in disease progression, transforming metagenomics into a predictive framework for managing chronic infections.
Public microbial genomes encode an immense record of biological diversity, evolution and molecular function, but much of this information remains difficult to reuse because raw sequencing data are not uniformly assembled, quality controlled, annotated or searchable at scale. Here we present AllTheBacteria, an open, community-built resource that transforms public bacterial short-read whole-genome sequencing reads into a uniformly processed discovery platform. The current analysed release contains 2,440,377 high-quality bacterial and archaeal genomes from 11,273 species, together with standardized taxonomic assignments, genome annotations, antimicrobial resistance calls, antiphage-defence annotations, protein structure predictions and AI-ready sequence tables. We show that this infrastructure enables applications that would otherwise be impractical, from global sequence search and outbreak contextualization to pangenome method development, antimicrobial resistance reservoir mapping and antiphage-defence ecology. As a stringent experimental demonstration, we mined 3,919,096 encrypted peptide fragments from AllTheBacteria proteomes using our deep learning model APEX 1.1, identifying 1,867 candidates with predicted antimicrobial activity. We synthesized 24 representative peptides and tested them against 20 clinically relevant bacterial strains, including antibiotic-resistant pathogens. Multiple peptides showed low-micromolar activity, membrane-responsive conformational transitions and selective envelope perturbation. A lead molecule, ATB20, reduced Acinetobacter baumannii burden in a murine skin abscess model with efficacy comparable to polymyxin B and no overt toxicity. Together, these results establish AllTheBacteria as both a foundational community resource for microbiology and a renewable engine for AI-guided antimicrobial discovery.
BACKGROUND:Staphylococcus aureus biofilms play a crucial role in chronic rhinosinusitis (CRS), leading to the persistence of symptoms. Severe CRS patients are frequently infected with S. aureus strains that exhibit higher biofilm properties (e.g., biomass, exoprotein production) compared to S. aureus from controls. S. aureus biofilms resist antibiotic treatment; however, the relationship between bacterial biofilm properties, antibiotic susceptibility, and CRS severity has not yet been defined and is the subject of this study. METHODS:S. aureus clinical isolates and reference strains, and matched clinical datasets were collected from CRS patients and controls (n = 35). Antimicrobial susceptibility of the isolates to clindamycin, mupirocin, clarithromycin, doxycycline, and amoxicillin-clavulanic acid was determined by minimum inhibitory concentration (MIC) and minimum biofilm eradication concentration (MBEC). RESULTS:S. aureus MBEC values (n = 35) were significantly higher (up to 11 times) than the MIC for all five antibiotics (p < 0.001). Among the various antibiotics tested, mupirocin had the strongest antibiofilm activity and amoxicillin-clavulanic acid the weakest: at antibiotic concentrations that are deemed to indicate susceptibility or intermediate resistance when testing in planktonic form, 80% biofilm eradication was reached for all isolates using mupirocin and only 5/35 isolates using amoxicillin-clavulanic acid. The biofilm metabolic activity, biomass, colony-forming units, and exoprotein production were positively correlated with the MBEC values for amoxicillin-clavulanic acid and clarithromycin, but not for the other antibiotics. Lund-Mackay and Lund-Kennedy disease severity scores showed positive correlations with the MBEC values of clarithromycin. CONCLUSION:These findings show that whilst severe CRS patients are frequently infected with S. aureus strains that exhibit higher biofilm-mediated virulence, these biofilms are also more difficult to control with standard of care antibiotics. Better personalized therapies are required to manage biofilm-mediated infections in severe CRS patients.
Motivation:Phage therapy offers a viable alternative for bacterial infections amid rising antimicrobial resistance. Its success relies on selecting safe and effective phage candidates that require comprehensive genomic screening to identify potential risks. However, this process is often labor intensive and time-consuming, hindering rapid clinical deployment. Results:We developed Sphae, an automated bioinformatics pipeline designed to streamline the therapeutic potential of a phage in under 10 minutes. Using Snakemake workflow manager, Sphae integrates tools for quality control, assembly, genome assessment, and annotation tailored specifically for phage biology. Sphae automates the detection of key genomic markers, including virulence factors, antimicrobial resistance genes, and lysogeny indicators such as integrase, recombinase, and transposase, which could preclude therapeutic use. Among the 65 phage sequences analyzed, 28 showed therapeutic potential, 8 failed due to low sequencing depth, 22 contained prophage or virulent markers, and 23 had multiple phage genomes. This workflow produces a report to assess phage safety and therapy suitability quickly. Sphae is scalable and portable, facilitating efficient deployment across most high-performance computing and cloud platforms, accelerating the genomic evaluation process. Availability and implementation:Sphae source code is freely available at https://github.com/linsalrob/sphae, with installation supported on Conda, PyPi, Docker containers.
KEY POINTS:Long-read-metagenomic sequencing is the best method for analyzing the sinus microbiome. 16s-rRNA-sequencing (both long and short read) results in PCR amplification bias that significantly distorts the sinus microbiome.
Phages, the viruses of bacteria, harbor an incredibly diverse repertoire of proteins capable of manipulating their bacterial hosts, inspiring many medical and biotechnological applications. However, to date, only a limited subset of that repertoire can be exploited, due to the difficulties in functionally elucidating these proteins. In this study, we investigated several structure-informed approaches to annotate hypothetical proteins from Pseudomonas infecting phages. We curated a representative dataset of over 10,000 proteins derived from NCBI, for which we predicted protein structures with ColabFold and assessed structural similarity via FoldSeek against the PDB, AlphaFold, and Phold databases. We evaluated multiple annotation strategies, including sequence-based (Pharokka), and structure-based (FoldSeek, Phold) methods. Our results show that up to 43 % of truly unannotated proteins can be functionally annotated when combining structure-informed approaches with UniProt-derived annotations. We highlight the complementarity of different databases and the importance of annotation quality filtering. This work provides a valuable resource of predicted structures and annotations, and offers insights into optimizing structure-based annotation pipelines for viral proteins, paving the way for deeper exploration of phage biology and its applications.