Motivation The Adverse Outcome Pathways (AOP)-Wiki, a knowledge database for AOPs, requires an efficient way to present an overview of its content for the reconstruction of networks by experts in a given domain. We have developed the AOP-networkFinder, a user-friendly tool that retrieves AOPs of interest, allows network generation and cleaning, and finally visualizes networks built around the retrieved AOPs. Our tool constructs AOP networks by connecting AOPs that use the same Key Events (KEs) in a versatile but controlled manner. Genes related to these KEs are also displayed. The constructed networks can then be exported as images or to Cytoscape for further fine-tuning and statistical analysis.Results The AOP-networkFinder allows users to comprehensively identify relationships between KEs and visualize the overall structure of an AOP both quickly and easily. This is immensely beneficial to researchers who need to understand the complex interplay between different KEs and the overall pathway they are studying and helps them to build further networks of interest while logging relevant information about changes within the network. These efforts are in line with the Findable, Accessible, Interoperable, and Reusable principles, which are crucial attributes for any developed databases and tools for optimizing (re)use in a dynamically changing landscape of AOP-Wiki.Availability and implementation The AOP-networkFinder is an open-source application and is available online at aop-networkfinder.no, in the 'Computational Toxicology at Norwegian Institute of Public Health' Zenodo community at DOI 10.5281/zenodo.11068434, in the GitHub repository at github.com/folkehelseinstituttet/AOPnetworkFinder_v1, as well as in a Docker image at hub.docker.com/r/nurre123/aop_network_finder. The software is available under the GNU Affero General Public License (AGPL), v3.0. The tool uses the AOP-Wiki SPARQL endpoint to retrieve AOP data.
The ability of epithelial monolayers to self-organize into a dynamic polarized state, where cells migrate in a uniform direction, is essential for tissue regeneration, development, and tumor progression. However, the mechanisms governing long-range polar ordering of motility direction in biological tissues remain unclear. Here, we investigate the self-organizing behavior of quiescent epithelial monolayers that transit to a dynamic state with long-range polar order upon growth factor exposure. We demonstrate that the heightened self-propelled activity of monolayer cells leads to formation of vortex-antivortex pairs that undergo sequential annihilation, ultimately driving the spread of long-range polar order throughout the system. A computational model, which treats the monolayer as an active elastic solid, accurately replicates this behavior, and weakening of cell-to-cell interactions impedes vortex-antivortex annihilation and polar ordering. Our findings uncover a mechanism in epithelia, where elastic solid material characteristics, activated self-propulsion, and topology-mediated guidance converge to fuel a highly efficient polar self-ordering activity.
Purpose A large proportion of Common variable immunodeficiency (CVID) patients has duodenal inflammation with increased intraepithelial lymphocytes (IEL) of unknown aetiology. The histologic similarities to celiac disease, lead to confusion regarding treatment (gluten-free diet) of these patients. We aimed to elucidate the role of epigenetic DNA methylation in the aetiology of duodenal inflammation in CVID and differentiate it from true celiac disease.Methods DNA was isolated from snap-frozen pieces of duodenal biopsies and analysed for differences in genome-wide epigenetic DNA methylation between CVID patients with increased IEL (CVID_IEL; n = 5) without IEL (CVID_N; n = 3), celiac disease (n = 3) and healthy controls (n = 3).Results The DNA methylation data of 5-methylcytosine in CpG sites separated CVID and celiac diseases from healthy controls. Differential methylation in promoters of genes were identified as potential novel mediators in CVID and celiac disease. There was limited overlap of methylation associated genes between CVID_IEL and Celiac disease. High frequency of differentially methylated CpG sites was detected in over 100 genes nearby transcription start site (TSS) in both CVID_IEL and celiac disease, compared to healthy controls. Differential methylation of genes involved in regulation of TNF/cytokine production were enriched in CVID_IEL, compared to healthy controls.Conclusion This is the first study to reveal a role of epigenetic DNA methylation in the etiology of duodenal inflammation of CVID patients, distinguishing CVID_IEL from celiac disease. We identified potential biomarkers and therapeutic targets within gene promotors and in high-frequency differentially methylated CpG regions proximal to TSS in both CVID_IEL and celiac disease.
Stool samples for fecal immunochemical tests (FIT) are collected in large numbers worldwide as part of colorectal cancer screening programs. Employing FIT samples from 1034 CRCbiome participants, recruited from a Norwegian colorectal cancer screening study, we identify, annotate and characterize more than 18000 DNA viruses, using shotgun metagenome sequencing. Only six percent of them are assigned to a known taxonomic family, with Microviridae being the most prevalent viral family. Linking individual profiles to comprehensive lifestyle and demographic data shows 17/25 of the variables to be associated with the gut virome. Physical activity, smoking, and dietary fiber consumption exhibit strong and consistent associations with both diversity and relative abundance of individual viruses, as well as with enrichment for auxiliary metabolic genes. We demonstrate the suitability of FIT samples for virome analysis, opening an opportunity for large-scale studies of this enigmatic part of the gut microbiome. The diverse viral populations and their connections to the individual lifestyle uncovered herein paves the way for further exploration of the role of the gut virome in health and disease.
Background Shotgun metagenome sequencing data obtained from a host environment will usually be contaminated with sequences from the host organism. Host sequences should be removed before further analysis to avoid biases, reduce downstream computational load, or ensure privacy in the case of a human host. The tools that we identified, as designed specifically to perform host contamination sequence removal, were either outdated, not maintained, or complicated to use. Consequently, we have developed HoCoRT, a fast and user-friendly tool that implements several methods for optimised host sequence removal. We have evaluated the speed and accuracy of these methods. Results HoCoRT is an open-source command-line tool for host contamination removal. It is designed to be easy to install and use, offering a one-step option for genome indexing. HoCoRT employs a variety of well-known mapping, classification, and alignment methods to classify reads. The user can select the underlying classification method and its parameters, allowing adaptation to different scenarios. Based on our investigation of various methods and parameters using synthetic human gut and oral microbiomes, and on assessment of publicly available data, we provide recommendations for typical datasets with short and long reads. Conclusions To decontaminate a human gut microbiome with short reads using HoCoRT, we found the optimal combination of speed and accuracy with BioBloom, Bowtie2 in end-to-end mode, and HISAT2. Kraken2 consistently demonstrated the highest speed, albeit with a trade-off in accuracy. The same applies to an oral microbiome, but here Bowtie2 was notably slower than the other tools. For long reads, the detection of human host reads is more difficult. In this case, a combination of Kraken2 and Minimap2 achieved the highest accuracy and detected 59% of human reads. In comparison to the dedicated DeconSeq tool, HoCoRT using Bowtie2 in end-to-end mode proved considerably faster and slightly more accurate. HoCoRT is available as a Bioconda package, and the source code can be accessed at https://github.com/ignasrum/hocort along with the documentation. It is released under the MIT licence and is compatible with Linux and macOS (except for the BioBloom module).
Many high-throughput sequencing datasets can be represented as objects with coordinates along a reference genome. Currently, biological investigations often involve a large number of such datasets, for example representing different cell types or epigenetic factors. Drawing overall conclusions from a large collection of results for individual datasets may be challenging and time-consuming. Meaningful interpretation often requires the results to be aggregated according to metadata that represents biological characteristics of interest. In this light, we here propose the hierarchical Genomic Suite HyperBrowser (hGSuite), an open-source extension to the GSuite HyperBrowser platform, which aims to provide a means for extracting key results from an aggregated collection of high-throughput DNA sequencing data. The hGSuite utilizes a metadata-informed data cube to calculate various statistics across the multiple dimensions of the datasets. With this work, we show that the hGSuite and its associated data cube methodology offers a quick and accessible way for exploratory analysis of large genomic datasets. The web-based toolkit named hGsuite Hyperbrowser is available at https://hyperbrowser.uio.no/hgsuite under a GPLv3 license.
Abstract The ability of epithelial monolayers to self-organize into a dynamic polarized state is essential for tissue regeneration, development, and tumor progression. However, the mechanisms governing long-range polar ordering in biological tissues remain unclear. Here we investigate the self-organizing behavior of quiescent epithelial monolayers that transit from a disordered, static state to a dynamic state with long-range polar order upon growth factor exposure. We demonstrate that the heightened self-propelled activity of monolayer cells leads to formation of ±1 topological defect pairs that subsequently undergo sequential annihilation, ultimately driving the spread of long-range polar order throughout the system. A computational model, which treats the monolayer as an active elastic solid, accurately replicates this behavior, and weakening of cell-to-cell interactions impedes defect annihilation and polar ordering. Furthermore, we have identified a set of fundamental rules that describe how topological defects can function as local organizers to spread order through the system. Lastly, we show that transient proliferation and re-annihilation of defect pairs facilitate a 180 degrees reorientation of the collective cell flow. Our study reveals a crucial role of topological defects in generating collective motion with long-range order, and provides insights into the underlying principles of active matter physics in biological systems.
SummaryAdaptive immune receptor (AIR) repertoires (AIRRs) record past immune encounters with exquisite specificity. Therefore, identifying identical or similar AIR sequences across individuals is a key step in AIRR analysis for revealing convergent immune response patterns that may be exploited for diagnostics and therapy. Existing methods for quantifying AIRR overlap do not scale with increasing dataset numbers and sizes. To address this limitation, we developed CompAIRR, which enables ultra-fast computation of AIRR overlap, based on either exact or approximate sequence matching. CompAIRR improves computational speed 1000-fold relative to the state of the art and uses only one-third of the memory: on the same machine, the exact pairwise AIRR overlap of 104AIRRs with 105sequences is found in ∼17 minutes, while the fastest alternative tool requires 10 days. CompAIRR has been integrated with the machine learning ecosystem immuneML to speed up various commonly used AIRR-based machine learning applications.Availability and implementationCompAIRR code and documentation are available athttps://github.com/uio-bmi/compairr. Docker images are available athttps://hub.docker.com/r/torognes/compairr. The scripts used for benchmarking and creating figures, and all raw data, may be found athttps://github.com/uio-bmi/compairr-benchmarking.
MOTIVATION:Previously we presented swarm, an open-source amplicon clustering programme that produces fine-scale molecular operational taxonomic units (OTUs) that are free of arbitrary global clustering thresholds. Here, we present swarm v3 to address issues of contemporary datasets that are growing towards tera-byte sizes. RESULTS:When compared with previous swarm versions, swarm v3 has modernized C++ source code, reduced memory footprint by up to 50%, optimized CPU-usage and multithreading (more than 7 times faster with default parameters), and it has been extensively tested for its robustness and logic. AVAILABILITY AND IMPLEMENTATION:Source code and binaries are available at https://github.com/torognes/swarm. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
DNA loop extrusion emerges as a key process establishing genome structure and function. We introduce MoDLE, a computational tool for fast, stochastic modeling of molecular contacts from DNA loop extrusion capable of simulating realistic contact patterns genome wide in a few minutes. MoDLE accurately simulates contact maps in concordance with existing molecular dynamics approaches and with Micro-C data and does so orders of magnitude faster than existing approaches. MoDLE runs efficiently on machines ranging from laptops to high performance computing clusters and opens up for exploratory and predictive modeling of 3D genome structure in a wide range of settings.
[This corrects the article DOI: 10.1093/narcan/zcaa019.].
BACKGROUND:Studies of shifts in microbial community composition has many applications. For studies at species or subspecies levels, the 16S amplicon sequencing lacks resolution and is often replaced by full shotgun sequencing. Due to higher costs, this restricts the number of samples sequenced. As an alternative to a full shotgun sequencing we have investigated the use of Reduced Metagenome Sequencing (RMS) to estimate the composition of a microbial community. This involves the use of double-digested restriction-associated DNA sequencing, which means only a smaller fraction of the genomes are sequenced. The read sets obtained by this approach have properties different from both amplicon and shotgun data, and analysis pipelines for both can either not be used at all or not explore the full potential of RMS data.RESULTS:We suggest a procedure for analyzing such data, based on fragment clustering and the use of a constrained ordinary least square de-convolution for estimating the relative abundance of all community members. Mock community datasets show the potential to clearly separate strains even when the 16S is 100% identical, and genome-wide differences is < 0.02, indicating RMS has a very high resolution. From a simulation study, we compare RMS to shotgun sequencing and show that we get improved abundance estimates when the community has many very closely related genomes. From a real dataset of infant guts, we show that RMS is capable of detecting a strain diversity gradient for Escherichia coli across time.CONCLUSION:We find that RMS is a good alternative to either metabarcoding or shotgun sequencing when it comes to resolving microbial communities at the strain level. Like shotgun metagenomics, it requires a good database of reference genomes and is well suited for studies of the human gut or other communities where many reference genomes exist. A data analysis pipeline is offered, as an R package at https://github.com/larssnip/microRMS . Video abstract.
C-type lectin-like domain family 16 member A (CLEC16A) is associated with autoimmune disorders, including multiple sclerosis (MS), but its functional relevance is not completely understood. CLEC16A is expressed in several immune cells, where it affects autophagic processes and receptor expression. Recently, we reported that the risk genotype of an MS-associated single nucleotide polymorphism in CLEC16A intron 19 is associated with higher expression of CLEC16A in CD4(+) T cells. Here, we show that CLEC16A expression is induced in CD4(+) T cells upon T cell activation. By the use of imaging flow cytometry and confocal microscopy, we demonstrate that CLEC16A is located in Rab4a-positive recycling endosomes in Jurkat TAg T cells. CLEC16A knock-down in Jurkat cells resulted in lower cell surface expression of the T cell receptor, however, this did not have a major impact on T cell activation response in vitro in Jurkat nor in human, primary CD4(+) T cells.
DNA methylation (5mC) and hydroxymethylation (5hmC) are chemical modifications of cytosine bases which play a crucial role in epigenetic gene regulation. However, cost, data complexity and unavailability of comprehensive analytical tools is one of the major challenges in exploring these epigenetic marks. Hydroxymethylation-and Methylation-Sensitive Tag sequencing (HMST-seq) is one of the most cost-effective techniques that enables simultaneous detection of 5mC and 5hmC at single base pair resolution. We present HMST-Seq-Analyzer as a comprehensive and robust method for performing simultaneous differential methylation analysis on 5mC and 5hmC data sets. HMST-Seq-Analyzer can detect Differentially Methylated Regions (DMRs), annotate them, give a visual overview of methylation status and also perform preliminary quality check on the data. In addition to HMST-Seq, our tool can be used on whole-genome bisulfite sequencing (WGBS) and reduced representation bisulfite sequencing (RRBS) data sets as well. The tool is written in Python with capacity to process data in parallel and is available at (https://hmst-seq.github.io/hmst/).
We aimed to investigate awareness of colorectal cancer (CRC) lifestyle risk factors, willingness to participate in CRC screening, and preferences concerning channels for information on CRC prevention in the general population, including the target age of the upcoming Norwegian national CRC screening program. The present study was a cross-sectional online survey of adults aged 39 to 55 years registered as Kantar Web Panel respondents in Norway. The survey included demographic characteristics, multiple choice knowledge questions of lifestyle risk factors for CRC, attitudes towards CRC screening, and preferred channels for receiving information on CRC prevention. Of 4375 participants invited, 2007 (46%) answered the survey. The average number of correctly identified lifestyle risk factors for CRC was 7.3 of ten. Women were significantly more likely than men, and those with university or college education more likely than those with lower education to correctly identify at least eight risk factors (odds ratio, OR = 1.53, 95% confidence interval, CI 1.25-1.87, and OR = 1.51, 95% CI 1.23-1.86, respectively). The number of correctly identified risk factors was positively associated with willingness to participate in CRC screening (P for trend < 0.001). The national public work force and the Norwegian Cancer Society were selected by 76% and 69% of the participants, respectively, to be trustworthy sources of information on CRC prevention. Awareness of CRC risk factors was associated with willingness to participate in CRC screening. The national public work force and Cancer Society can be generally accepted sources of CRC preventive information.
BACKGROUND:Advances in whole genome sequencing strategies have provided the opportunity for genomic and comparative genomic analysis of a vast variety of organisms. The analysis results are highly dependent on the quality of the genome assemblies used. Assessment of the assembly accuracy may significantly increase the reliability of the analysis results and is therefore of great importance.RESULTS:Here, we present a new tool called NucBreak aimed at localizing structural errors in assemblies, including insertions, deletions, duplications, inversions, and different inter- and intra-chromosomal rearrangements. The approach taken by existing alternative tools is based on analysing reads that do not map properly to the assembly, for instance discordantly mapped reads, soft-clipped reads and singletons. NucBreak uses an entirely different and unique method to localise the errors. It is based on analysing the alignments of reads that are properly mapped to an assembly and exploit information about the alternative read alignments. It does not annotate detected errors. We have compared NucBreak with other existing assembly accuracy assessment tools, namely Pilon, REAPR, and FRCbam as well as with several structural variant detection tools, including BreakDancer, Lumpy, and Wham, by using both simulated and real datasets.CONCLUSIONS:The benchmarking results have shown that NucBreak in general predicts assembly errors of different types and sizes with relatively high sensitivity and with lower false discovery rate than the other tools. Such a balance between sensitivity and false discovery rate makes NucBreak a good alternative to the existing assembly accuracy assessment tools and SV detection tools. NucBreak is freely available at https://github.com/uio-bmi/NucBreak under the MPL license.
Background In spite of the major breakthroughs in the second-generation sequencing technologies and the developments of a plethora of assemblers over the last ten years, the resulting genome assemblies may still be fragmented and contain errors. It is typical in genome projects with second-generation reads involved to run multiple assemblers with different parameters and choose the best assembly. However, such an approach is always a trade-off between the strengths and weaknesses of the assemblies. To exploit the advantages of different assemblers, an alternative approach that combines the best parts of several assemblies into one may be applied. The existing tools based on such an approach assist in elongation of assembly fragments and/or improvement of assembly accuracy. Though there has been progress with such a strategy, there is still room for improvement of the existing tools. Results We present NucMerge, a tool for improving genome assembly accuracy by incorporating information derived from an alternative assembly and paired-end Illumina reads from the same genome. The tool corrects insertion, deletion, substitution, and inversion errors and locates different inter- and intra-chromosomal rearrangement errors. NucMerge was compared to two existing alternatives, namely Metassembler and GAM-NGS. Conclusions The benchmarking results show that NucMerge has generally better performance than the other tools tested, providing accuracy improvement of more assemblies. NucMerge is freely available at https://github.com/uio-bmi/NucMerge under the MPL license.
Both a DNA lesion and an intermediate for antibody maturation, uracil is primarily processed by base excision repair (BER), either initiated by uracil-DNA glycosylase (UNG) or by single-strand selective monofunctional uracil DNA glycosylase (SMUG1). The relative in vivo contributions of each glycosylase remain elusive. To assess the impact of SMUG1 deficiency, we measured uracil and 5-hydroxymethyluracil, another SMUG1 substrate, in Smug1 -/- mice. We found that 5-hydroxymethyluracil accumulated in Smug1 -/- tissues and correlated with 5-hydroxymethylcytosine levels. The highest increase was found in brain, which contained about 26-fold higher genomic 5-hydroxymethyluracil levels than the wild type. Smug1 -/- mice did not accumulate uracil in their genome and Ung -/- mice showed slightly elevated uracil levels. Contrastingly, Ung -/- Smug1 -/- mice showed a synergistic increase in uracil levels with up to 25-fold higher uracil levels than wild type. Whole genome sequencing of UNG/SMUG1-deficient tumours revealed that combined UNG and SMUG1 deficiency leads to the accumulation of mutations, primarily C to T transitions within CpG sequences. This unexpected sequence bias suggests that CpG dinucleotides are intrinsically more mutation prone. In conclusion, we showed that SMUG1 efficiently prevent genomic uracil accumulation, even in the presence of UNG, and identified mutational signatures associated with combined UNG and SMUG1 deficiency.