Bioinformatics tools are increasingly important for diagnostics in clinical care and precision medicine, but despite a very active bioinformatics research community, implementation and adaptation is slow. Drawing on multidisciplinary expertise, we have identified key systemic barriers on the journey from research to implementation, using the Danish healthcare ecosystem as the example. We find the main obstacles to be regulatory uncertainty, fragmented data access, and limited infrastructure for implementation. Cultural resistance to commercialization and workforce gaps further impedes progress. We believe that these challenges reflect broader international trends and could be generally applicable. Consensus recommendations include centralized data and regulatory resources, cross-sector collaboration models, and pilot initiatives to support scalable implementation. These findings offer a roadmap for translating bioinformatics innovation into clinical practice.
Abstract Gene expression-based prognostic models have shown promise for predicting recurrence in colorectal cancer (CRC), but their clinical implementation remains limited. The NanoString nCounter platform provides a practical alternative to RNA sequencing and microarrays through standardized, cost-effective gene expression profiling that is compatible with routine clinical samples. In this study, we evaluated whether NanoString nCounter gene expression data improve prediction of recurrence following curative CRC surgery. Gene expression profiles from the NanoString PanCancer IO 360™ panel were analyzed in two independent CRC cohorts (cohort A, n = 189; cohort B, n = 131). Differential gene expression analyses and Cox proportional hazards models were used to assess the prognostic value of gene expression alone and in combination with established clinical risk factors. Model performance was evaluated by five-fold cross-validation and external validation between cohorts using the concordance index (C-index) and Kaplan-Meier risk stratification. The two cohorts differed significantly in recurrence-free survival, and differential expression analysis demonstrated marked cohort-specific transcriptional patterns. Ninety-one recurrence-associated genes were identified in cohort A, whereas no significant genes were detected in cohort B, with poor agreement in gene-level differential expression between cohorts (Pearson r = 0.128). Across all prediction models, external performance was modest, and inclusion of gene expression data did not improve prediction beyond clinical variables. The clinical baseline model, incorporating age, UICC stage, and tumor site, consistently achieved the highest cross-cohort performance, with UICC stage emerging as the strongest predictor of recurrence. Although overall discrimination was moderate, the baseline model successfully stratified patients into significantly different high- and low-risk groups across cohorts. These findings indicate that prognostic gene expression signatures derived from NanoString data showed limited reproducibility across independent cohorts and provided little additional predictive value beyond established clinical factors. The results highlight the importance of external validation and suggest that robust clinical variables remain the most reliable predictors of recurrence risk in this setting.
Metagenomic sequencing has provided great advantages in the characterisation of microbiomes, but currently available analysis tools lack the ability to combine subspecies-level taxonomic resolution and accurate abundance estimation with functional profiling of assembled genomes. To define the microbiome and its associations with human health, improved tools are needed to enable comprehensive understanding of the microbial composition and elucidation of the phylogenetic and functional relationships between the microbes. Here, we present MAGinator, a freely available tool, tailored for profiling of shotgun metagenomics datasets. MAGinator provides de novo identification of subspecies-level microbes and accurate abundance estimates of metagenome-assembled genomes (MAGs). MAGinator utilises the information from both gene- and contig-based methods yielding insight into both taxonomic profiles and the origin of genes and genetic content, used for inference of functional content of each sample by host organism. Additionally, MAGinator facilitates the reconstruction of phylogenetic relationships between the MAGs, providing a framework to identify clade-level differences.
Polygenic risk scores (PRSs) are expected to play a critical role in precision medicine. Currently, PRS predictors are generally based on linear models using summary statistics, and more recently individual-level data. However, these predictors mainly capture additive relationships and are limited in data modalities they can use. We developed a deep learning framework (EIR) for PRS prediction which includes a model, genome-local-net (GLN), specifically designed for large-scale genomics data. The framework supports multi-task learning, automatic integration of other clinical and biochemical data, and model explainability. When applied to individual-level data from the UK Biobank, the GLN model demonstrated a competitive performance compared to established neural network architectures, particularly for certain traits, showcasing its potential in modeling complex genetic relationships. Furthermore, the GLN model outperformed linear PRS methods for Type 1 Diabetes, likely due to modeling non-additive genetic effects and epistasis. This was supported by our identification of widespread non-additive genetic effects and epistasis in the context of T1D. Finally, we constructed PRS models that integrated genotype, blood, urine, and anthropometric data and found that this improved performance for 93% of the 290 diseases and disorders considered. EIR is available at https://github.com/arnor-sigurdsson/EIR.
AbstractMotivationMetagenomic sequencing has provided great advantages in the characterization of microbiomes, but currently available analysis tools lack the ability to combine strain-level taxonomic resolution and abundance estimation with functional profiling of assembled genomes. In order to define the microbiome and its associations with human health, improved tools are needed to enable comprehensive understanding of the microbial composition and elucidation of the phylogenetic and functional relationships between the microbes.ResultsHere, we present MAGinator, a freely available tool, tailored for the profiling of shotgun metagenomics datasets. MAGinator providesde novoidentification of subspecies-level microbes and accurate abundance estimates of metagenome-assembled genomes (MAGs). MAGinator utilises the information from both gene- and contig-based methods yielding insight into both taxonomic profiles and the origin of genes as well as genetic content, used for inference of functional content of each sample by host organism. Additionally, MAGinator facilitates the reconstruction of phylogenetic relationships between the MAGs, providing a framework to identify clade-level differences within subspecies MAGs.Availability and implementationMAGinator is available as a Python module athttps://github.com/Russel88/MAGinatorContactTrine Zachariasen,trine_zachariasen@hotmail.com
Crops with deeper rooting is an emerging tool for better exploitation of soil resources. However, there is a need for more in-depth understanding on how the increased rooting depth may be achieved. In this study a novel approach for obtaining deeper rooting has been proposed. Crops with assumed similar capacity for subsoil exploration: sugar beet (Beta vulgaris) and chicory (Cichorium intybus var. foliosum) were intercropped. Repeated measurements of biomass, deep root growth, and nutrient uptake were conducted to monitor plant competitive dynamics in the intercrop and sole crops. It was found that the intercrop positively affected biomass production with Land Equivalent Ratio close to or greater than 1 (0.99 - 1.14). Similarly, the strongest root growth over time was observed for the intercrop (from 98 +/- 48 to 304 +/- 28 cm depth). Moreover, the effect from the interspecific interactions in the intercrop varied over time. In the first half of the season yield advantage and the observed enhanced contribution to the uptake of N, Mg, Mn, Zn, and Na in the intercrop were driven by the sugar beet. Later in the growing season, yield advantage, deep root growth, and contribution to the uptake of S, Fe, Cu, and Al in the intercrop were driven by the chicory. This has also been confirmed by the root quantification analysis, which showed that in the end of the season intercrop consisted of 84 % and 98 % roots from the chicory at 1 and 2.5 m depth, respectively. This study concluded that intercropping two crops with similar root characteristics, sugar beet and chicory, can still lead to complementary interactions showing potential for efficient deep soil exploration by roots and yield advantage in comparison with the sole crops.
Motivation Metagenomic sequencing has provided great advantages in the characterization of microbiomes, but currently available analysis tools lack the ability to combine strain-level taxonomic resolution and abundance estimation with functional profiling of assembled genomes. In order to define the microbiome and its associations with human health, improved tools are needed to enable comprehensive understanding of the microbial composition and elucidation of the phylogenetic and functional relationships between the microbes. Results Here, we present MAGinator, a freely available tool, tailored for the profiling of shotgun metagenomics datasets. MAGinator provides de novo identification of subspecies-level microbes and accurate abundance estimates of metagenome-assembled genomes (MAGs). MAGinator utilises the information from both gene- and contig-based methods yielding insight into both taxonomic profiles and the origin of genes as well as genetic content, used for inference of functional content of each sample by host organism. Additionally, MAGinator facilitates the reconstruction of phylogenetic relationships between the MAGs, providing a framework to identify clade-level differences within subspecies MAGs. Availability and implementation MAGinator is available as a Python module at Contact Trine Zachariasen, trine_zachariasen{at}hotmail.com ### Competing Interest Statement The authors have declared no competing interest.
High-throughput genome sequencing technologies enable the investigation of complex genetic interactions, including the horizontal gene transfer of plasmids and bacteriophages. However, identifying these elements from assembled reads remains challenging due to genome sequence plasticity and the difficulty in assembling complete sequences. In this study, we developed a classifier, using random forest, to identify whether sequences originated from bacterial chromosomes, plasmids, or bacteriophages. The classifier was trained on a diverse collection of 23,211 chromosomal, plasmid, and bacteriophage sequences from hundreds of bacterial species. In order to adapt the classifier to incomplete sequences, each complete sequence was subsampled into 5,000 nucleotide fragments and further subdivided into k-mers. This three-class classifier succeeded in identifying chromosomes, plasmids, and bacteriophages using k-mer distributions of complete and partial genome sequences, including simulated metagenomic scaffolds with minimum performance of 0.939 area under the receiver operating characteristic curve (AUC). This classifier, implemented as SourceFinder, has been made available as an online web service to help the community with predicting the chromosomal, plasmid, and bacteriophage sources of assembled bacterial sequence data (https://cge.food.dtu.dk/ services/SourceFinder/). IMPORTANCE Extra-chromosomal genes encoding antimicrobial resistance, metal resistance, and virulence provide selective advantages for bacterial survival under stress conditions and pose serious threats to human and animal health. These accessory genes can impact the composition of microbiomes by providing selective advantages to their hosts. Accurately identifying extra-chromosomal elements in genome sequence data are critical for understanding gene dissemination trajectories and taking preventative measures. Therefore, in this study, we developed a random forest classifier for identifying the source of bacterial chromosomal, plasmid, and bacteriophage sequences.
ABSTRACT Blood and urine biomarkers are an essential part of modern medicine, not only for diagnosis, but also for their direct influence on disease. Many biomarkers have a genetic component, and they have been studied extensively with genome-wide association studies (GWAS) and methods that compute polygenic scores (PGSs). However, these methods generally assume both an additive allelic model and an additive genetic architecture for the target outcome, and thereby risk not capturing non-linear allelic effects nor epistatic interactions. Here, we trained and evaluated deep-learning (DL) models for PGS prediction of 34 blood and urine biomarkers in the UK Biobank cohort, and compared them to linear methods. For lipid traits, the DL models greatly outperformed the linear methods, which we found to be consistent across diverse populations. Furthermore, the DL models captured non-linear effects in covariates, non-additive genotype (allelic) effects, and epistatic interactions between SNPs. Finally, when using only genome-wide significant SNPs from GWAS, the DL models performed equally well or better for all 34 traits tested. Our findings suggest that DL can serve as a valuable addition to existing methods for genotype-phenotype modelling in the era of increasing data availability.
Background Semi-quantitative bacterial culture is the reference standard to diagnose urinary tract infection, but culture is time-consuming and can be unreliable if patients are receiving antibiotics. Metagenomics could increase diagnostic accuracy and speed by sequencing the microbiota and resistome directly from urine. We aimed to compare metagenomics to culture for semi-quantitative pathogen and resistome detection from urine. Methods In this proof-of-concept study, we prospectively included consecutive urine samples from a clinical diagnostic laboratory in Amsterdam. Urine samples were screened by DNA concentration, followed by PCR-free metagenomic sequencing of randomly selected samples with a high concentration of DNA (culture positive and negative). A diagnostic index was calculated as the product of DNA concentration and fraction of pathogen reads. We compared results with semi-quantitative culture using area under the receiver operating characteristic curve (AUROC) analyses. We used ResFinder and PointFinder for resistance gene detection and compared results to phenotypic antimicrobial susceptibility testing for six antibiotics commonly used for urinary tract infection treatment: nitrofurantoin, ciprofloxacin, fosfomycin, cotrimoxazole, ceftazidime, and ceftriaxone. Findings We screened 529 urine samples of which 86 were sequenced (43 culture positive and 43 culture negative). The AUROC of the DNA concentration-based screening was 0.85 (95% CI 0.81-0.89). At a cutoff value of 6.0 ng/mL, culture positivity was ruled out with a negative predictive value of 91% (95% CI 87-93; 26 of 297 samples), reducing the number of samples requiring sequencing by 56% (297 of 529 samples). The AUROC of the diagnostic index was 0.87 (95% CI 0.79-0.95). A diagnostic index cutoff value of 17.2 yielded a positive predictive value of 93% (95% CI 85-97) and a negative predictive value of 69% (55-80), correcting for a culture-positive prevalence of 66%. Gram-positive pathogens explained eight (89%) of the nine false-negative metagenomic test results. Agreement of phenotypic and genotypic antimicrobial susceptibility testing varied between 71% (22 of 31 samples) and 100% (six of six samples), depending on the antibiotic tested. Interpretation This study provides proof-of-concept of metagenomic semi-quantitative pathogen and resistome detection for the diagnosis of urinary tract infection. The findings warrant prospective clinical validation of the value of this approach in informing patient management and care. Copyright (c) 2022 The Author(s). Published by Elsevier Ltd.
Plasmids play a major role facilitating the spread of antimicrobial resistance between bacteria. Understanding the host range and dissemination trajectories of plasmids is critical for surveillance and prevention of antimicrobial resistance. Identification of plasmid host ranges could be improved using automated pattern detection methods compared to homology-based methods due to the diversity and genetic plasticity of plasmids. In this study, we developed a method for predicting the host range of plasmids using machine learning-specifically, random forests. We trained the models with 8,519 plasmids from 359 different bacterial species per taxonomic level; the models achieved Matthews correlation coefficients of 0.662 and 0.867 at the species and order levels, respectively. Our results suggest that despite the diverse nature and genetic plasticity of plasmids, our random forest model can accurately distinguish between plasmid hosts. This tool is available online through the Center for Genomic Epidemiology (https://cge.cbs.dtu.dk/services/PlasmidHostFinder/). IMPORTANCE Antimicrobial resistance is a global health threat to humans and animals, causing high mortality and morbidity while effectively ending decades of success in fighting against bacterial infections. Plasmids confer extra genetic capabilities to the host organisms through accessory genes that can encode antimicrobial resistance and virulence. In addition to lateral inheritance, plasmids can be transferred horizontally between bacterial taxa. Therefore, detection of the host range of plasmids is crucial for understanding and predicting the dissemination trajectories of extrachromosomal genes and bacterial evolution as well as taking effective countermeasures against antimicrobial resistance.
Lactococcus lactis strains are important components in industrial starter cultures for cheese manufacturing. They have many strain-dependent properties, which affect the final product. Here, we explored the use of machine learning to create systematic, high-throughput screening methods for these properties. Fast acidification of milk is such a strain-dependent property. To predict the maximum hourly acidification rate (V max ), we trained Random Forest (RF) models on four different genomic representations: Presence/absence of gene families, counts of Pfam domains, the 8 nucleotide long subsequences of their DNA (8-mers), and the 9 nucleotide long subsequences of their DNA (9-mers). V max was measured at different temperatures, volumes, and in the presence or absence of yeast extract. These conditions were added as features in each RF model. The four models were trained on 257 strains, and the correlation between the measured V max and the predicted V max was evaluated with Pearson Correlation Coefficients (PC) on a separate dataset of 85 strains. The models all had high PC scores: 0.83 (gene presence/absence model), 0.84 (Pfam domain model), 0.76 (8-mer model), and 0.85 (9-mer model). The models all based their predictions on relevant genetic features and showed consensus on systems for lactose metabolism, degradation of casein, and pH stress response. Each model also predicted a set of features not found by the other models.
The emergence of artemisinin-resistant Plasmodium falciparum parasites in Southeast Asia threatens malaria control and elimination. The interconnectedness of parasite populations may be essential to monitor the spread of resistance. Combining a published barcoding system of geographically restricted single-nucleotide polymorphisms (SNPs), mainly mitochondria of P. falciparum with SNPs in the K13 artemisinin resistance marker, could elucidate the parasite population structure and provide insight regarding the spread of drug resistance. We explored the diversity of mitochondrial SNPs (bp position 611-2825) and identified K13 SNPs from malaria patients in the districts of India (Ranchi), Tanzania (Korogwe), and Senegal (Podor, Richard Toll, Kaolack, and Ndoffane). DNA was amplified using a nested PCR and Sanger-sequenced. Overall, 199 K13 sequences (India: N = 92; Tanzania: N = 48; Senegal: N = 59) and 237 mitochondrial sequences (India: N = 93; Tanzania: N = 48; Senegal: N = 96) were generated. SNPs were identified by comparisons with reference genomes. We detected previously reported geographically restricted mitochondrial SNPs (T2175C and G1367A) as markers for parasites originating from the Indian subcontinent and several geographically unrestricted mitochondrial SNPs. Combining haplotypes with published P. falciparum mitochondrial genome data suggested possible regional differences within India. All three countries had G1692A, but Tanzanian and Senegalese SNPs were well-differentiated. Some mitochondrial SNPs are reported here for the first time. Four nonsynonymous K13 SNPs were detected: K189T (India, Tanzania, Senegal); A175T (Tanzania); and A174V and R255K (Senegal). This study supports the use of mitochondrial SNPs to determine the origin of the parasite and suggests that the P. falciparum populations studied were susceptible to artemisinin during sampling because all K13 SNPs observed were outside the propeller domain for artemisinin resistance.
Abstract Summary Here, we present an automated pipeline for Download Of NCBI Entries (DONE) and continuous updating of a local sequence database based on user-specified queries. The database can be created with either protein or nucleotide sequences containing all entries or complete genomes only. The pipeline can automatically clean the database by removing entries with matches to a database of user-specified sequence contaminants. The default contamination entries include sequences from the UniVec database of plasmids, marker genes and sequencing adapters from NCBI, an E.coli genome, rRNA sequences, vectors and satellite sequences. Furthermore, duplicates are removed and the database is automatically screened for sequences from green fluorescent protein, luciferase and antibiotic resistance genes that might be present in some GenBank viral entries, and could lead to false positives in virus identification. For utilizing the database, we present a useful opportunity for dealing with possible human contamination. We show the applicability of DONE by downloading a virus database comprising 37 virus families. We observed an average increase of 16 776 new entries downloaded per month for the 37 families. In addition, we demonstrate the utility of a custom database compared to a standard reference database for classifying both simulated and real sequence data. Availabilityand implementation The DONE pipeline for downloading and cleaning is deposited in a publicly available repository (https://bitbucket.org/genomicepidemiology/done/src/master/). Supplementary information Supplementary data are available at Bioinformatics online.
Background Quinoa ( Chenopodium quinoa Willd.) is an ancient grain crop that is tolerant to abiotic stress and has favorable nutritional properties. Downy mildew is the main disease of quinoa and is caused by infections of the biotrophic oomycete Peronospora variabilis Gaüm. Since the disease causes major yield losses, identifying sources of downy mildew tolerance in genetic resources and understanding its genetic basis are important goals in quinoa breeding. Results We infected 132 South American genotypes, three Danish cultivars and the weedy relative C. album with a single isolate of P. variabilis under greenhouse conditions and observed a large variation in disease traits like severity of infection, which ranged from 5 to 83%. Linear mixed models revealed a significant effect of genotypes on disease traits with high heritabilities (0.72 to 0.81). Factors like altitude at site of origin or seed saponin content did not correlate with mildew tolerance, but stomatal width was weakly correlated with severity of infection. Despite the strong genotypic effects on mildew tolerance, genome-wide association mapping with 88 genotypes failed to identify significant marker-trait associations indicating a polygenic architecture of mildew tolerance. Conclusions The strong genetic effects on mildew tolerance allow to identify genetic resources, which are valuable sources of resistance in future quinoa breeding.
Cassava is Africa’s most important food security crop and sustains about 700 million people globally. Survey interviews of 320 farmers in three regions of Tanzania to identify their production characteristics, and interviews with 20 international whitefly/virus experts were conductedto identify adaptation strategies to lessen the impacts of cassava whiteflies and viruses due to climate change in Tanzania. Structured and pre-tested interview schedules were conducted using a multistage sampling technique. Most of the farmers (66.8%) produced cassava primarily for food, and relied mainly on their friends (43.8%) and their farms (41.9%) for cassava planting materials. Farmers significantly differed in their socio-economic and production characteristics except for gender and access to extension support (P < 0.01). A significant association was found between extension support, sources of planting materials, and reasons for growing cassava with both the control of cassava viruses and the control of whiteflies by the farmers. A significantly higher number of farmers controlled cassava viruses (38.1%) than cassava whiteflies (19.7%). The adaptation strategies most recommended by experts were: integrating pest and disease management programs, phytosanitation, and applying novel vector management techniques.The experts also recommended capacity building through the training of stakeholders, establishing monitoring networks to get updates on cassava pests and disease statuses, incorporating pest and disease adaptation planning into the general agricultural management plans, and developing climate change-pest/disease models for accessing the local and national level impacts that can facilitate more specific adaptation planning in order to enhance the farmers’ adaptive capacities.
Antimicrobial resistance causes thousands of deaths annually worldwide. Understanding the regions of the genome that are involved in antimicrobial resistance is important for developing mitigation strategies and preventing transmission.
Traditional genotyping methods for infection control of antimicrobial-resistant bacteria in healthcare settings have been supplemented by whole-genome sequencing (WGS), often relying on a gene-based approach, e.g., core genome multilocus sequence typing (cgMLST), to cluster-related samples. In this study, we compared clusters of methicillin-resistant Staphylococcus aureus (MRSA) and Enterococcus faecium analyzed with the commercial cgMLST software Ridom SeqSphere+ and with an open-source single-nucleotide polymorphism (SNP)-based phylogenetic analysis pipeline (PAPABAC). A total of 5,655 MRSA and 2,572 E. faecium patient isolates, collected between 2013 and 2018, were processed. Clusters of 1,844 MRSA and 1,355 E. faecium isolates were compared to cgMLST results, and epidemiological data were included when available. The phylogenies inferred by the two different technologies were highly concordant, and the MRSA SNP tree re-captured known hospital-related outbreaks and epidemiologically linked samples. PAPABAC has the advantage over Ridom SeqSphere+ to generate stable, referable clusters without the need for sequence assembly, and it is a free-of-charge, open-source alternative to the commercial software.
For detection of clonal outbreaks in clinical settings, we present a complete pipeline that generates a single-nucleotide polymorphisms-distance matrix from a set of sequencing reads. Importantly, the program is able to handle a separate mix of both short reads from the Illumina sequencing platforms and long reads from Oxford Nanopore Technologies' (ONT) platforms as input. MINTyper performs automated reference identification, alignment, alignment trimming, optional methylation masking, and pairwise distance calculations. With this approach, we could rapidly and accurately cluster a set of DNA sequenced isolates, with a known epidemiological relationship to confirm the clustering. Functions were built to allow for both high-accuracy methylation-aware base-called MinION reads (hac_m Q10) and fast generated lower-quality reads (fast Q8) to be used, also in combination with Illumina data. With fast Q8 reads a higher number of base pairs were excluded from the calculated distance matrix, compared with the high-accuracy methylation-aware Q10 base-calling of ONT data. Nonetheless, when using different qualities of ONT data with corresponding input parameters, the clustering of isolates were nearly identical.
Additional file 4: Table S2. Precision, recall and F1 scores obtained with the CCMetagen analysis with assembled sequence reads.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark51