CO$_2$ concentrations above air level (0.04%) are beneficial for the growth of cyanobacteria. However, very high CO$_2$ levels inhibit growth, limiting the usability of cyanobacteria for carbon capture and biotechnological applications. The transcriptomic changes in the cyanobacterium Synechococcus sp. PCC 7002 were investigated at varying growth rates governed by different gas-phase CO$_2$ concentrations, ranging from limiting (0.04%) to optimal (4% and 8%) to inhibitory high (30%). Compared to optimal CO$_2$ concentrations, large differences in the transcriptome were observed in limiting and inhibiting CO$_2$. At 30% CO$_2$, genes encoding CO$_2$ uptake mechanisms, photosynthetic electron transfer, and light-harvesting antennae proteins were down-regulated compared to lower CO$_2$. Genes involved in the ribosomal machinery and biosynthetic pathways were down-regulated at 30% and 0.04% CO$_2$, consistent with the observed reduced growth. The genes most strongly up-regulated at 30% CO$_2$ were primarily associated with stress responses, but did not closely resemble the transcriptomic changes at other stress conditions previously described. The small RNA PsrR1 was strongly up-regulated at 30% CO$_2$ and is likely to be involved in the regulation of growth. These transcriptomic insights are essential for the engineering of fast-growing cyanobacteria at high CO$_2$, which can be applied in carbon capture from industrial point sources.
Cultivation of photosynthetic microbes under elevated CO2 is widely used in biotechnology to enhance growth. However, very high CO2 levels inhibit growth for reasons that are not always clear. We established a method that allowed cultivation under different CO2 concentrations and controlled pH to study the cyanobacterium Synechococcus sp. PCC 7002. Growth rates, cell composition, and transcriptomic profiles were investigated under limiting (0.04%), optimal (1-4%), and inhibitory high (15-30%) gas phase CO2 contents, while maintaining pH similar to 7.5. At intermediate light (200 mu mol m(-2) s(-1)), the growth rate at 30% CO2 was similar to 45% of that under optimal conditions. Under low light (<100 mu mol m(-2) s(-1)), growth was less inhibited by high CO2 than at higher light intensities. Pigmentation and photosynthesis likewise decreased under high CO2, except that carotenoids indicative of high-light stress increased. Transcriptome sequencing confirmed that the photosynthesis machinery was downregulated and stress responses were upregulated under high CO2. Genetic knockdown of CO2 uptake (ndhF3 and ndhF4) or overexpression of a heterologous Na+/H+ antiporter (nhaS3) improved tolerance to high CO2, while knockdown of the CcmR transcriptional regulator reduced it. We conclude that the inhibiting species under very high inorganic carbon conditions was CO2, and not HCO3-, and that reduced light intensity or genetically suppressing active CO2 uptake improved the tolerance of Synechococcus to very high CO2 conditions. This work provides a basis for cultivation under a very wide range of CO2 concentrations, thereby broadening the scope for metabolic manipulation and biotechnological applications of cyanobacteria.
Cyanobacteria are one of the oldest and most abundant groups of prokaryotes and are crucial for research in climate, ecology, medicine, and agriculture. Despite intensive efforts in metabolic engineering of cyanobacteria, the mechanisms of gene regulation, particularly through regulatory RNA structures, are often ignored. We computationally searched 202 cyanobacterial genomes for putative conserved RNA structures (CRSs) in the upstream and downstream intergenic regions of 931 orthologous gene groups with the comparative genomics tool CMfinder. The predicted structures were scored according to their local phylogeny and filtered for a maximal false discovery rate of 10%. The screen identified 402 CRSs that match known RNA families (Rfam and Rho-independent bacterial terminators) and 409 novel CRSs. The structures are not limited to either low or high nucleotide conservation, and about half have a high level of significant covariation. The majority of novel CRSs are supported by transcription in at least one species in public RNA-seq data. The regulatory associations of CRSs are discussed in different metabolic pathways, such as photosynthesis, nitrogen fixation, and CO$_2$ metabolism. This resource will support future research on the regulatory mechanisms of RNA in cyanobacteria.
Abstract One strategy for CO2 mitigation is using photosynthetic microorganisms to sequester CO2 under high concentrations, such as in flue gases. While elevated CO2 levels generally promote growth, excessively high levels inhibit growth through uncertain mechanisms. This study investigated the physiology of the cyanobacterium Synechocystis sp. PCC 6803 under very high CO2 concentrations and yet stable pH around 7.5. The growth rate of the wild type (WT) at 200 µmol photons m−2 s−1 and a gas phase containing 30% CO2 was 2.7-fold lower compared to 4% CO2. Using a CRISPR interference mutant library, we identified genes that, when repressed, either enhanced or impaired growth under 30% or 4% CO2. Repression of genes involved in light harvesting (cpc and apc), photochemical electron transfer (cytM, psbJ, and petE), and several genes with little or unknown functions promoted growth under 30% CO2, while repression of key regulators of photosynthesis (pmgA) and CO2 capture and fixation (ccmR, cp12, and yfr1) increased growth inhibition under 30% CO2. Experiments confirmed that WT cells were more susceptible to light inhibition under 30% than under 4% CO2 and that a light-harvesting-impaired ΔcpcG mutant showed improved growth under 30% CO2 compared to the WT. These findings suggest that enhanced fitness under very high CO2 involves modifications in light harvesting, electron transfer, and carbon metabolism, and that the native regulatory machinery is insufficient, and in some cases obstructive, for optimal growth under 30% CO2. This genetic profiling provides potential targets for engineering cyanobacteria with improved photosynthetic efficiency and stress resilience for biotechnological applications. Key points • Synechocystis growth was inhibited under very high CO 2 . • Inhibition of growth under very high CO 2 was light dependent. • Repression of photosynthesis genes improved growth under very high CO 2 . Graphical Abstract
Long intergenic non-coding RNAs (lincRNAs) and single nucleotide polymorphisms (SNPs) have been associated with cancers for years, yet the molecular mechanism is mostly unclear. The secondary structure of lincRNAs is often crucial for their biological function but can be disrupted by SNPs. However, only a small number of studies investigated how lincRNAs function through secondary structure. Given that the vast majority of cancer-related SNPs are located mainly in non-coding regions, there is a large potential for associating SNPs to disrupted structure in lincRNAs. To estimate the structural impacts of cancer-associated SNPs on lincRNAs, we predicted local secondary structures for lincRNAs and computed structural distances between structural ensembles of the wild-type and mutant sequences. Manual literature curation was performed to study the function of lincRNAs that are structurally disrupted by cancer-associated SNPs. By integrating with RBP binding sites annotation, we estimated the impacts of structural changes from cancer-associated SNPs on the protein binding of lincRNAs. We predict 559 SNPs to cause significant structural disruption in 231 lincRNAs using RNAsnp (P-value < 0.1). In addition, we find that these disrupted regions have the potential to alter the binding ability of lincRNAs with RBPs. An example is the structural change in the lincRNA small nucleolar RNA host gene 25 (SNHG25), which overlaps the binding site of protein insulin like growth factor 2 mRNA binding protein 2 (IGF2BP2) in glioblastoma multiforme. The results show the importance of the lincRNA secondary structure in understanding their biological function, especially the structural changes from SNPs in cancer. The predicted structural change in lincRNA SNHG25 holds a potential insight into the mechanism of protein IGF2BP2 in recognizing RNA methylation signals in glioblastoma multiforme.
Patents are essential for transferring scientific discoveries to meaningful products that benefit societies. While the academic community focuses on the number of citations to rank scholarly works according to their “scientific merit,” the number of citations is unrelated to the relevance for patentable innovation. To explore associations between patents and scholarly works in publicly available patent data, we propose to utilize statistical methods that are commonly used in biology to determine gene-disease associations. We illustrate their usage on patents related to biotechnological trends of high relevance for food safety and ecology, namely the CRISPR-based gene editing technology (>60,000 patents) and cyanobacterial biotechnology (>33,000 patents). Innovation trends are found through their unexpected large changes of patent numbers in a time-series analysis. From the total set of scholarly works referenced by all investigated patents (~254,000 publications), we identified ~1,000 scholarly works that are statistical significantly over-represented in the references of patents from changing innovation trends that concern immunology, agricultural plant genomics, and biotechnological engineering methods. The detected associations are consistent with the technical requirements of the respective innovations. In summary, the presented data-driven analysis workflow can identify scholarly works that were required for changes in innovation trends, and, therefore, is of interest for researches that would like to evaluate the relevance of publications beyond the number of citations.
IntroductionChimeric antigen receptor-expressing T cells (CAR T cells) have revolutionized cancer treatment, particularly in B cell malignancies. However, the use of autologous T cells for CAR T therapy presents several limitations, including high costs, variable efficacy, and adverse effects linked to cell phenotype.MethodsTo overcome these challenges, we developed a strategy to generate universal and safe anti-CD19 CAR T cells with a defined memory phenotype. Our approach utilizes CRISPR/Cas9 technology to target and eliminate the B2M and TRAC genes, reducing graft-versus-host and host-versus-graft responses. Additionally, we selected less differentiated T cells to improve the stability and persistence of the universal CAR T cells. The safety of this method was assessed using our CRISPRroots transcriptome analysis pipeline, which ensures successful gene knockout and the absence of unintended off-target effects on gene expression or transcriptome sequence.ResultsIn vitro experiments demonstrated the successful generation of functional universal CAR T cells. These cells exhibited potent lytic activity against tumor cells and a reduced cytokine secretion profile. The CRISPRroots analysis confirmed effective gene knockout and no unintended off-target effects, validating it as a pioneering tool for on/off-target and transcriptome analysis in genome editing experiments.DiscussionOur findings establish a robust pipeline for manufacturing safe, universal CAR T cells with a favorable memory phenotype. This approach has the potential to address the current limitations of autologous CAR T cell therapy, offering a more stable and persistent treatment option with reduced adverse effects. The use of CRISPRroots enhances the reliability and safety of gene editing in the development of CAR T cell therapies.ConclusionWe have developed a potent and reliable method for producing universal CAR T cells with a defined memory phenotype, demonstrating both efficacy and safety in vitro. This innovative approach could significantly improve the therapeutic landscape for patients with B cell malignancies.
RNA secondary structures play essential roles in the formation of the tertiary structure and function of a transcript. Recent genome-wide studies highlight significant potential for RNA structures in the mammalian genome. However, a major challenge is assigning functional roles to these structured RNAs. In this study, we conduct a guilt-by-association analysis of clusters of computationally predicted conserved RNA structure (CRSs) in human untranslated regions (UTRs) to associate them with gene functions. We filtered a broad pool of similar to 500 000 human CRSs for UTR overlap, resulting in 4734 and 24 754 CRSs from the 5 ' and 3 ' UTR of protein-coding genes, respectively. We separately clustered these CRSs for both sets using RNAscClust, obtaining 793 and 2403 clusters, each containing an average of five CRSs per cluster. We identified overrepresented binding sites for 60 and 43 RNA-binding proteins co-localizing with the clustered CRSs. Furthermore, 104 and 441 clusters from the 5 ' and 3 ' UTRs, respectively, showed enrichment for various Gene Ontologies, including biological processes such as 'signal transduction', 'nervous system development', molecular functions like 'transferase activity' and the cellular components such as 'synapse' among others. Our study shows that significant functional insights can be gained by clustering RNA structures based on their structural characteristics.
The European Cooperation in Science and Technology (COST) is an intergovernmental organization dedicated to funding and coordinating scientific and technological research in Europe, fostering collaboration among researchers and institutions across countries. Recently, COST Action funded the ''Genome Editing to treat Human Diseases'' (GenE-HumDi) network, uniting various stakeholders such as pharmaceutical companies, academic institutions, regulatory agencies, biotech firms, and patient advocacy groups. GenE-HumDi’s primary objective is to expedite the application of genome editing for therapeutic purposes in treating human diseases. To achieve this goal, GenE-HumDi is organized in several working groups, each focusing on specific aspects. These groups aim to enhance genome editing technologies, assess delivery systems, address safety concerns, promote clinical translation, and develop regulatory guidelines. The network seeks to establish standard procedures and guidelines for these areas to standardize scientific practices and facilitate knowledge sharing. Furthermore, GenE-HumDi aims to communicate its findings to the public in accessible yet rigorous language, emphasizing genome editing’s potential to revolutionize the treatment of many human diseases. The inaugural GenE-HumDi meeting, held in Granada, Spain, in March 2023, featured presentations from experts in the field, discussing recent breakthroughs in delivery methods, safety measures, clinical translation, and regulatory aspects related to gene editing.
AbstractYield improvements in cell factories can potentially be obtained by fine-tuning the regulatory mechanisms for gene candidates. In pursuit of such candidates, we performed RNA-sequencing of two α-amylase producing Bacillus strains and predict hundreds of putative novel non-coding transcribed regions. Surprisingly, we found among hundreds of non-coding and structured RNA candidates that non-coding genomic regions are proportionally undergoing the highest changes in expression during fermentation. Since these classes of RNA are also understudied, we targeted the corresponding genomic regions with CRIPSRi knockdown to test for any potential impact on the yield. From differentially expression analysis, we selected 53 non-coding candidates. Although CRISPRi knockdowns target both the sense and the antisense strand, the CRISPRi experiment cannot link causes for yield changes to the sense or antisense disruption. Nevertheless, we observed on several instances with strong changes in enzyme yield. The knockdown targeting the genomic region for a putative antisense RNA of the 3′ UTR of the skfA-skfH operon led to a 21% increase in yield. In contrast, the knockdown targeting the genomic regions of putative antisense RNAs of the cytochrome c oxidase subunit 1 (ctaD), the sigma factor sigH, and the uncharacterized gene yhfT decreased yields by 31 to 43%.
Frontotemporal dementia (FTD) is a common cause of early-onset dementia, with no current treatment options. FTD linked to chromosome 3 (FTD3) is a rare sub-form of the disease, caused by a point mutation in the Charged Multivesicular Body Protein 2B (CHMP2B). This mutation causes neuronal phenotypes, such as mitochondrial deficiencies, accompanied by metabolic changes and interrupted endosomal-lysosomal fusion. However, the contribution of glial cells to FTD3 pathogenesis has, until recently, been largely unexplored. Glial cells play an important role in most neurodegenerative disorders as drivers and facilitators of neuroinflammation. Microglia are at the center of current investigations as potential pro-inflammatory drivers. While gliosis has been observed in FTD3 patient brains, it has not yet been systematically analyzed. In the light of this, we investigated the role of microglia in FTD3 by implementing human induced pluripotent stem cells (hiPSC) with either a heterozygous or homozygous CHMP2B mutation, introduced into a healthy control hiPSC line via CRISPR-Cas9 precision gene editing. These hiPSC were differentiated into microglia to evaluate the pro-inflammatory profile and metabolic state. Moreover, hiPSC-derived neurons were cultured with conditioned microglia media to investigate disease specific interactions between the two cell populations. Interestingly, we identified two divergent inflammatory microglial phenotypes resulting from the underlying mutations: a severe pro-inflammatory profile in CHMP2B homozygous FTD3 microglia, and an "unresponsive" CHMP2B heterozygous FTD3 microglial state. These findings correlate with our observations of increased phagocytic activity in CHMP2B homozygous, and impaired protein degradation in CHMP2B heterozygous FTD3 microglia. Metabolic mapping confirmed these differences, revealing a metabolic reprogramming of the CHMP2B FTD3 microglia, displayed as a compensatory up-regulation of glutamine metabolism in the CHMP2B homozygous FTD3 microglia. Intriguingly, conditioned CHMP2B homozygous FTD3 microglia media caused neurotoxic effects, which was not evident for the heterozygous microglia. Strikingly, IFN-γ treatment initiated an immune boost of the CHMP2B heterozygous FTD3 microglia, and conditioned microglia media exposure promoted neural outgrowth. Our findings indicate that the microglial profile, activity, and behavior is highly dependent on the status of the CHMP2B mutation. Our results suggest that the heterozygous state of the mutation in FTD3 patients could potentially be exploited in form of immune-boosting intervention strategies to counteract neurodegeneration.
Alzheimer's disease (AD) is a progressive and irreversible brain disorder, which can occur either sporadically, due to a complex combination of environmental, genetic, and epigenetic factors, or because of rare genetic variants in specific genes (familial AD, or fAD). A key hallmark of AD is the accumulation of amyloid beta (Aβ) and Tau hyperphosphorylated tangles in the brain, but the underlying pathomechanisms and interdependencies remain poorly understood. Here, we identify and characterise gene expression changes related to two fAD mutations (A79V and L150P) in the Presenilin-1 (PSEN1) gene. We do this by comparing the transcriptomes of glutamatergic forebrain neurons derived from fAD-mutant human induced pluripotent stem cells (hiPSCs) and their individual isogenic controls generated via precision CRISPR/Cas9 genome editing. Our analysis of Poly(A) RNA-seq data detects 1111 differentially expressed coding and non-coding genes significantly altered in fAD. Functional characterisation and pathway analysis of these genes reveal profound expression changes in constituents of the extracellular matrix, important to maintain the morphology, structural integrity, and plasticity of neurons, and in genes involved in calcium homeostasis and mitochondrial oxidative stress. Furthermore, by analysing total RNA-seq data we reveal that 30 out of 31 differentially expressed circular RNA genes are significantly upregulated in the fAD lines, and that these may contribute to the observed protein-coding gene expression changes. The results presented in this study contribute to a better understanding of the cellular mechanisms impacted in AD neurons, ultimately leading to neuronal damage and death.
Non-coding RNAs are key regulatory players in bacteria. Many computationally predicted non-coding RNAs, however, lack functional associations. An example is the Bacillaceae-1 RNA motif, whose Rfam model consists of two hairpin loops. We find the motif conserved in nine of 13 non-pathogenic strains of the genus Bacillus but only in one pathogenic strain. To elucidate functional characteristics, we studied 118 hits of the Rfam model in 11 Bacillus spp. and found two distinct classes based on the ensemble diversity of their RNA secondary structure and the genomic context concerning the ribosomal RNA (rRNA) cluster. Forty hits are associated with the rRNA cluster, of which all 19 hits upstream flanking of 16S rRNA have a reverse complementary structure of low structural diversity. Fifty-two hits have large ensemble diversity, of which 38 are located between two coding genes. For eight hits in Bacillus subtilis, we investigated public expression data under various conditions and observed either the forward or the reverse complementary motif expressed. Five hits are associated with the rRNA cluster. Four of them are located upstream of the 16S rRNA and are not transcriptionally active, but instead, their reverse complements with low structural diversity are expressed together with the rRNA cluster. The three other hits are located between two coding genes in non-conserved genomic loci. Two of them are independently expressed from their surrounding genes and are structurally diverse. In summary, we found that Bacillaceae-1 RNA motifs upstream flanking of ribosomal RNA clusters tend to have one stable structure with the reverse complementary motif expressed in B. subtilis. In contrast, a subgroup of intergenic motifs has the thermodynamic potential for structural switches.
Abstract The CRISPR-Cas9 genome editing tool is used to study genomic variants and gene knockouts, and can be combined with transcriptomic analyses to measure the effects of such alterations on gene expression. But how can one be sure that differential gene expression is due to a successful intended edit and not to an off-target event, without performing an often resource-demanding genome-wide sequencing of the edited cell or strain? To address this question we developed CRISPRroots: CRISPR–Cas9-mediated edits with accompanying RNA-seq data assessed for on-target and off-target sites. Our method combines Cas9 and guide RNA binding properties, gene expression changes, and sequence variants between edited and non-edited cells to discover potential off-targets. Applied on seven public datasets, CRISPRroots identified critical off-target candidates that were overlooked in all of the corresponding previous studies. CRISPRroots is available via https://rth.dk/resources/crispr.
Accelerated evolution of any portion of the genome is of significant interest, potentially signaling positive selection of phenotypic traits and adaptation. Accelerated evolution remains understudied for structured RNAs, despite the fact that an RNA's structure is often key to its function. RNA structures are typically characterized by compensatory (structure-preserving) basepair changes that are unexpected given the underlying sequence variation, i.e., they have evolved through negative selection on structure. We address the question of how fast the primary sequence of an RNA can change through evolution while conserving its structure. Specifically, we consider predicted and known structures in vertebrate genomes. After careful control of false discovery rates, we obtain 13 de novo structures (and three known Rfam structures) that we predict to have rapidly evolving sequences-defined as structures where the primary sequences of human and mouse have diverged at least twice as fast (1.5 times for Rfam) as nearby neutrally evolving sequences. Two of the three known structures function in translation inhibition related to infection and immune response. We conclude that rapid sequence divergence does not preclude RNA structure conservation in vertebrates, although these events are relatively rare.
Abstract Background Bacillus subtilis is a Gram-positive bacterium used as a cell factory for protein production. Over the last decades, the continued optimization of production strains has increased yields of enzymes, such as amylases, and made commercial applications feasible. However, current yields are still significantly lower than the theoretically possible yield based on the available carbon sources. In its natural environment, B. subtilis can respond to unfavorable growth conditions by differentiating into motile cells that use flagella to swim towards available nutrients. Results In this study, we analyze existing transcriptome data from a B. subtilis α-amylase production strain at different time points during a 5-day fermentation. We observe that genes of the fla/che operon, essential for flagella assembly and motility, are differentially expressed over time. To investigate whether expression of the flagella operon affects yield, we performed CRISPR-dCas9 based knockdown of the fla/che operon with sgRNA target against the genes flgE, fliR, and flhG, respectively. The knockdown resulted in inhibition of mobility and a striking 2–threefold increase in α-amylase production yield. Moreover, replacing flgE (required for flagella hook assembly) with an erythromycin resistance gene followed by a transcription terminator increased α-amylase yield by about 30%. Transcript levels of the α-amylase were unaltered in the CRISPR-dCas9 knockdowns as well as the flgE deletion strain, but all manipulations disrupted the ability of cells to swim on agar. Conclusions We demonstrate that the disruption of flagella in a B. subtilis α-amylase production strain, either by CRISPR-dCas9-based knockdown of the operon or by replacing flgE with an erythromycin resistance gene followed by a transcription terminator, increases the production of α-amylase in small-scale fermentation.
The production of the alpha-amylase (AMY) enzyme in Bacillus subtilis at a high rate leads to the accumulation of unfolded AMY, which causes secretion stress. The over-expression of the PrsA chaperone aids enzyme folding and reduces stress. To identify affected pathways and potential mechanisms involved in the reduced growth, we analyzed the transcriptomic differences during fed-batch fermentation between a PrsA over-expressing strain and control in a time-series RNA-seq experiment. We observe transcription in 542 unannotated regions, of which 234 had significant changes in expression levels between the samples. Moreover, 1,791 protein-coding sequences, 80 non-coding genes, and 20 riboswitches overlapping UTR regions of coding genes had significant changes in expression. We identified putatively regulated biological processes via gene-set over-representation analysis of the differentially expressed genes; overall, the analysis suggests that the PrsA over-expression affects ATP biosynthesis activity, amino acid metabolism, and cell wall stability. The investigation of the protein interaction network points to a potential impact on cell motility signaling. We discuss the impact of these highlighted mechanisms for reducing secretion stress or detrimental aspects of PrsA over-expression during AMY production.
Stellate cells are principal neurons in the entorhinal cortex that contribute to spatial processing. They also play a role in the context of Alzheimer’s disease as they accumulate Amyloid beta early in the disease. Producing human stellate cells from pluripotent stem cells would allow researchers to study early mechanisms of Alzheimer’s disease, however, no protocols currently exist for producing such cells. In order to develop novel stem cell protocols, we characterize at high resolution the development of the porcine medial entorhinal cortex by tracing neuronal and glial subtypes from mid-gestation to the adult brain to identify the transcriptomic profile of progenitor and adult stellate cells. Importantly, we could confirm the robustness of our data by extracting developmental factors from the identified intermediate stellate cell cluster and implemented these factors to generate putative intermediate stellate cells from human induced pluripotent stem cells. Six transcription factors identified from the stellate cell cluster including RUNX1T1, SOX5, FOXP1, MEF2C, TCF4, EYA2 were overexpressed using a forward programming approach to produce neurons expressing a unique combination of RELN, SATB2, LEF1 and BCL11B observed in stellate cells. Further analyses of the individual transcription factors led to the discovery that FOXP1 is critical in the reprogramming process and omission of RUNX1T1 and EYA2 enhances neuron conversion. Our findings contribute not only to the profiling of cell types within the developing and adult brain’s medial entorhinal cortex but also provides proof-of-concept for using scRNAseq data to produce entorhinal intermediate stellate cells from human pluripotent stem cells in-vitro.
A large part of our current understanding of gene regulation in Gram-positive bacteria is based on Bacillus subtilis , as it is one of the most well studied bacterial model systems. The rapid growth in data concerning its molecular and genomic biology is distributed across multiple annotation resources. Consequently, the interpretation of data from further B. subtilis experiments becomes increasingly challenging in both low- and large-scale analyses. Additionally, B. subtilis annotation of structured RNA and non-coding RNA (ncRNA), as well as the operon structure, is still lagging behind the annotation of the coding sequences. To address these challenges, we created the B. subtilis genome atlas, BSGatlas, which integrates and unifies multiple existing annotation resources. Compared to any of the individual resources, the BSGatlas contains twice as many ncRNAs, while improving the positional annotation for 70 % of the ncRNAs. Furthermore, we combined known transcription start and termination sites with lists of known co-transcribed gene sets to create a comprehensive transcript map. The combination with transcription start/termination site annotations resulted in 717 new sets of co-transcribed genes and 5335 untranslated regions (UTRs). In comparison to existing resources, the number of 5′ and 3′ UTRs increased nearly fivefold, and the number of internal UTRs doubled. The transcript map is organized in 2266 operons, which provides transcriptional annotation for 92 % of all genes in the genome compared to the at most 82 % by previous resources. We predicted an off-target-aware genome-wide library of CRISPR–Cas9 guide RNAs, which we also linked to polycistronic operons. We provide the BSGatlas in multiple forms: as a website (https://rth.dk/resources/bsgatlas/), an annotation hub for display in the UCSC genome browser, supplementary tables and standardized GFF3 format, which can be used in large scale -omics studies. By complementing existing resources, the BSGatlas supports analyses of the B. subtilis genome and its molecular biology with respect to not only non-coding genes but also genome-wide transcriptional relationships of all genes.