Yinggehai Basin is an important high-temperature and high-pressure gas producing basin in the South China Sea, and Ledong 8-1 diapir is located in the southeast of the basin, which is the focus of natural gas migration and accumulation. However, the study of fluid inclusions in this gas-bearing structure is not enough, and the accumulation period and accumulation process are not clear, which restricts the progress of exploration in this area. Taking Ledong 8-1 as the research object, this paper uses microscopic lithofacies observation, microscopic temperature measurement of inclusions, laser Raman composition identification, in situ quantitative spectroscopy based on inclusion data to recover paleo-pressure and indirect projection dating of uniform temperature-burial history to identify gas inclusions of different composition types and recover the paleo-pressure evolution history of the Huangliu Formation. The accumulation period and process of natural gas were determined. The results show that there are four types of gas inclusions in Ledong 8-1 diapir, corresponding to three phases of gas charging. The charging time of the first phase of natural gas is 2.2-1.6Ma, the gas composition is mainly CH4, and the pressure coefficient of Huangliu group is 1.13-1.42, indicating weak overpressure. The second phase of natural gas charging occurred at 1.6-0.6Ma, and the gas composition was mainly CO2. Since 0.6Ma, the third stage of natural gas was charged into the formation, and the pressure coefficient of Huangliu formation was as high as 1.8, indicating strong overpressure. It is speculated that the diapiric was opened at this stage, and the natural gas formed in the early stage of Huangliu formation and the newly generated natural gas from deep source rocks migrated upward and adjusted. Studying the fluid inclusions in Ledong 8-1 area of Yinggehai Basin and clarifying the process of gas accumulation in this area is of great significance to the oil and gas exploration in the whole basin, and it is also the key to promote the increase of oil and gas storage and production in the whole basin.
The Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.
People often value the sensual, celebratory, and health aspects of food, but behind this experience exists many other value-laden agricultural production, distribution, manufacturing, and physiological processes that support or undermine a healthy population and a sustainable future. The complexity of such processes is evident in both every-day food preparation of recipes and in industrial food manufacturing, packaging and storage, each of which depends critically on human or machine agents, chemical or organismal ingredient references, and the explicit instructions and implicit procedures held in formulations or recipes. An integrated ontology landscape does not yet exist to cover all the entities at work in this farm to fork journey. It seems necessary to construct such a vision by reusing expert-curated fit-to-purpose ontology subdomains and their relationship, material, and more abstract organization and role entities. The challenge is to make this merger be, by analogy, one language, rather than nouns and verbs from a dozen or more dialects which cannot be used directly in statements about some aspect of the farm to fork journey without expensive translation or substantial dialect education in order to understand a particular text or domain of knowledge. This work focuses on the ontology components – object and data properties and annotations – needed to model food processes or more general process modelling within the context of the Open Biological and Biomedical Ontology Foundry and congruent ontologies. Ideally these components can be brought together in a general process ontology that can be specialized not only for the food domain but for carrying out other protocols as well. Many operations involved in food identification, preparation, transportation and storage – shaking, boiling, mixing, freezing, labeling, shipping – are actually common to activities from manufacturing and laboratory work to local or home food preparation.
Since its creation in 2016, the FoodOn food ontology has become an interconnected partner in various academic and government projects that span agricultural and public health domains. This paper examines recent data interoperability capabilities arising from food-related ontologies belonging to, or compatible with, the encyclopedic Open Biological and Biomedical Ontology Foundry (OBO) ontology platform, and how research organizations and industry might utilize them for their own projects or for data exchange. Projects are seeking standardized vocabulary across many food supply activities ranging from agricultural production, harvesting, preparation, food processing, marketing, distribution and consumption, as well as more indirect health, economic, food security and sustainability analysis and reporting tools. To satisfy this demand for controlled vocabulary requires establishing domain specific ontologies whose curators coordinate closely to produce recommended patterns for food system vocabulary.
Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce "RNA-Bloom2 [ https://github.com/bcgsc/RNA-Bloom ]", a reference-free assembly method for long-read transcriptome sequencing data. Using simulated datasets and spike-in control data, we show that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods. Furthermore, we find that RNA-Bloom2 requires 27.0 to 80.6% of the peak memory and 3.6 to 10.8% of the total wall-clock runtime of a competing reference-free method. Finally, we showcase RNA-Bloom2 in assembling a transcriptome sample of Picea sitchensis (Sitka spruce). Since our method does not rely on a reference, it further sets the groundwork for large-scale comparative transcriptomics where high-quality draft genome assemblies are not readily available.
Background: Nanopore sequencing is crucial to metagenomic studies as its kilobase-long reads can contribute to resolving genomic structural differences among microbes. However, sequencing platform-specific challenges, including high base-call error rate, nonuniform read lengths, and the presence of chimeric artifacts, necessitate specifically designed analytical algorithms. The use of simulated datasets with characteristics that are true to the sequencing platform under evaluation is a cost-effective way to assess the performance of bioinformatics tools with the ground truth in a controlled environment. Results: Here, we present Meta-NanoSim, a fast and versatile utility that characterizes and simulates the unique properties of nanopore metagenomic reads. It improves upon state-of-the-art methods on microbial abundance estimation through a base-level quantification algorithm. Meta-NanoSim can simulate complex microbial communities composed of both linear and circular genomes and can stream reference genomes from online servers directly. Simulated datasets showed high congruence with experimental data in terms of read length, error profiles, and abundance levels. We demonstrate that Meta-NanoSim simulated data can facilitate the development of metagenomic algorithms and guide experimental design through a metagenome assembly benchmarking task. Conclusions: The Meta-NanoSim characterization module investigates read features, including chimeric information and abundance levels, while the simulation module simulates large and complex multisample microbial communities with different abundance profiles. All trained models and the software are freely accessible at GitHub: https://github.com/bcgsc/NanoSim.
BACKGROUND Associations between increased dietary fat and decreased carbohydrate intake with circulating HDL and non-HDL cholesterol have not been conclusively determined. OBJECTIVE We assessed these relations in 8 European observational human studies participating in the European Nutritional Phenotype Assessment and Data Sharing Initiative (ENPADASI) using harmonized data. METHODS Dietary macronutrient intake was recorded using study-specific dietary assessment tools. Main outcome measures were lipoprotein cholesterol concentrations: HDL cholesterol (mg/dL) and non-HDL cholesterol (mg/dL). A cross-sectional analysis on 5919 participants (54% female) aged 13-80 y was undertaken using the statistical platform DataSHIELD that allows remote/federated nondisclosive analysis of individual-level data. Generalized linear models (GLM) were fitted to assess associations between replacing 5% of energy from carbohydrates with equivalent energy from total fats, SFAs, MUFAs, or PUFAs with circulating HDL cholesterol and non-HDL cholesterol. GLM were adjusted for study source, age, sex, smoking status, alcohol intake and BMI. RESULTS The replacement of 5% of energy from carbohydrates with total fats or MUFAs was statistically significantly associated with 0.67 mg/dL (95% CI: 0.40, 0.94) or 0.99 mg/dL (95% CI: 0.37, 1.60) higher HDL cholesterol, respectively, but not with non-HDL cholesterol concentrations. The replacement of 5% of energy from carbohydrates with SFAs or PUFAs was not associated with HDL cholesterol, but SFAs were statistically significantly associated with 1.94 mg/dL (95% CI: 0.08, 3.79) higher non-HDL cholesterol, and PUFAs with -3.91 mg/dL (95% CI: -6.98, -0.84) lower non-HDL cholesterol concentrations. A statistically significant interaction by sex for the association of replacing carbohydrates with MUFAs and non-HDL cholesterol was observed, showing a statistically significant inverse association in males and no statistically significant association in females. We observed no statistically significant interaction by age. CONCLUSIONS The replacement of dietary carbohydrates with fats had favorable effects on lipoprotein cholesterol concentrations in European adolescents and adults when fats were consumed as MUFAs or PUFAs but not as SFAs.
Recent advances in single-cell RNA sequencing technologies have made detection of transcripts in single cells possible. The level of resolution provided by these technologies can be used to study changes in transcript usage across cell populations and help investigate new biology. Here, we introduce RNA-Scoop, an interactive cell cluster and transcriptome visualization tool to analyze transcript usage across cell categories and clusters. The tool allows users to examine differential transcript expression across clusters and investigate how usage of specific transcript expression mechanisms varies across cell groups.
Robust recommendations for healthy diets and nutrition require careful synthesis of available evidence. Given the increasing volume of research articles generated, the retrieval and synthesis of evidence are increasingly becoming laborious and time-consuming. Information technology could help to reduce workload for humans. To guide supervised learning however, human identification of key study characteristics is necessary. Reporting guidelines recommend that authors include essential content in articles and could generate manually labeled training data for automated evidence retrieval and synthesis. Here, we present a semiautomated approach to annotate, link, and track the content of nutrition research manuscripts. We used the STROBE extension for nutritional epidemiology (STROBE-nut) reporting guidelines to manually annotate a sample of 15 articles and converted the semantic information into linked data in a Neo4j graph database through an automated process. Six summary statistics were computed to estimate the reporting completeness of the articles. The content structure, presence of essential study characteristics as well as the reporting completeness of the articles are visualized automatically from the graph database. The archived linked data are interoperable through their annotations and relations. A graph database with linked data on essential study characteristics can enable Natural Language Processing in nutrition.
BACKGROUND:Compared with second-generation sequencing technologies, third-generation single-molecule RNA sequencing has unprecedented advantages; the long reads it generates facilitate isoform-level transcript characterization. In particular, the Oxford Nanopore Technology sequencing platforms have become more popular in recent years owing to their relatively high affordability and portability compared with other third-generation sequencing technologies. To aid the development of analytical tools that leverage the power of this technology, simulated data provide a cost-effective solution with ground truth. However, a nanopore sequence simulator targeting transcriptomic data is not available yet.FINDINGS:We introduce Trans-NanoSim, a tool that simulates reads with technical and transcriptome-specific features learnt from nanopore RNA-sequncing data. We comprehensively benchmarked Trans-NanoSim on direct RNA and complementary DNA datasets describing human and mouse transcriptomes. Through comparison against other nanopore read simulators, we show the unique advantage and robustness of Trans-NanoSim in capturing the characteristics of nanopore complementary DNA and direct RNA reads.CONCLUSIONS:As a cost-effective alternative to sequencing real transcriptomes, Trans-NanoSim will facilitate the rapid development of analytical tools for nanopore RNA-sequencing data. Trans-NanoSim and its pre-trained models are freely accessible at https://github.com/bcgsc/NanoSim.
Isoform detection and discovery in single cells is central to improving our understanding of tumor heterogeneity. However, since current interactive RNA-seq visualization tools are designed for bulk RNA-Seq data, they have limited utility in displaying and analyzing single-cell RNA-seq data. Here, we introduce RNA-Scoop, a tool which visualizes isoform expression in single cells through an interactive t-SNE plot. The input of RNA-Scoop consists of a single JSON file, which specifies the paths to a GTF file containing the isoforms of interest, a matrix file containing their expression levels in each cell, and files containing labels for the matrix rows and columns. Users can select genes for the isoform view, where all isoforms of selected genes are displayed. The t-SNE plot allows users to zoom in and out of different areas and select cells via lasso selection. Upon selection, displayed isoforms are colored according to their average level of expression in the selected cells. Expression per cluster is visualized through a dot plot. Additionally, isoforms are selectable, enabling users to highlight the cells in which isoforms of interest are expressed. Through these easy-to-use features, RNA-Scoop simplifies the interrogation of isoforms and cell types in thousands of cells.
As a long-read sequencing technique, Oxford Nanopore Technology (ONT) has shown unprecedented potential in metagenomic studies. However, the challenges associated with ONT reads, such as high error rate and non-uniform error distributions, necessitate analytical tools designed specifically for long reads. To facilitate the development and benchmarking, simulated datasets with known ground truth are desirable. Here, we present Meta-NanoSim, a fast and lightweight ONT read simulator that characterizes and simulates the unique properties of ONT metagenomes, including abundance levels, chimeric reads, and reads that span both ends of a circular genome. Provided with the empirical profiles and abundance profile learnt from experimental dataset, multi-sample multi-replicate metagenome datasets are generated to simulate microbial communities with both circular and linear genomes. To demonstrate its performance, we train Meta-NanoSim with two mock microbial community standards and compare the simulation results against state-of-the-art tools. Further, we showcase the application of Meta-NanoSim through benchmarking ONT metagenome assemblers on our simulated datasets. Gold standards provided by Meta-NanoSim will facilitate the development of algorithms and pipelines in metagenomics, including functional gene prediction, species detection, comparative metagenomics, and clinical diagnosis. As such, we expect Meta-NanoSim to have an enabling role in the field.
Despite the rapid advance in single-cell RNA sequencing (scRNA-seq) technologies within the last decade, single-cell transcriptome analysis workflows have primarily used gene expression data while isoform sequence analysis at the single-cell level still remains fairly limited. Detection and discovery of isoforms in single cells is difficult because of the inherent technical shortcomings of scRNA-seq data, and existing transcriptome assembly methods are mainly designed for bulk RNA samples. To address this challenge, we developed RNA-Bloom, an assembly algorithm that leverages the rich information content aggregated from multiple single-cell transcriptomes to reconstruct cell-specific isoforms. Assembly with RNA-Bloom can be either reference-guided or reference-free, thus enabling unbiased discovery of novel isoforms or foreign transcripts. We compared both assembly strategies of RNA-Bloom against five state-of-the-art reference-free and reference-based transcriptome assembly methods. In our benchmarks on a simulated 384-cell data set, reference-free RNA-Bloom reconstructed 37.9%–38.3% more isoforms than the best reference-free assembler, whereas reference-guided RNA-Bloom reconstructed 4.1%–11.6% more isoforms than reference-based assemblers. When applied to a real 3840-cell data set consisting of more than 4 billion reads, RNA-Bloom reconstructed 9.7%–25.0% more isoforms than the best competing reference-based and reference-free approaches evaluated. We expect RNA-Bloom to boost the utility of scRNA-seq data beyond gene expression analysis, expanding what is informatically accessible now.
We discuss efforts in improving the value of nutrition research. We organised the paper in five research stages: Stage 1: research priority setting; Stage 2: research design, conduct and analysis; Stage 3: research regulation and management; Stage 4: research accessibility and Stage 5: research reporting and publishing. Along the stages of the research cycle, varied initiatives exist to improve the quality and added value of nutrition research. However, efforts are focused on single stages of the research cycle without vision of the research system as a whole. Although research on nutrition research has been limited, it has potential to improve the quality of nutrition research and develop new tools and instruments for this purpose. A comprehensive assessment of the magnitude of research waste in nutrition and consensus on priority actions is needed. The nutrition research community at large needs to have open discussions on the usefulness of these tools and lead suitable efforts to enhance nutrition research across the stages of the research cycle. Capacity building is essential and considerations of nutrition research quality are vital to be integrated in training efforts of nutrition researchers.
As a widespread RNA processing machinery, alternative polyadenylation plays a crucial role in gene regulation. To help decipher its underlying mechanism and understand its impact, it is desirable to comprehensively profile 3’-untranslated region cleavage and associated polyadenylation sites. State-of-the-art polyadenylation site detection tools are influenced either by library preparation or manually selected features. Here we present Termin(A)ntor, a deep neural network-based profiling pipeline to predict polyadenylation sites from RNA-seq data. We show how Termin(A)ntor outperforms competing tools in sensitivity and precision on experimental transcriptome sequence data. We also demonstrate applications of Termin(A)ntor with both short-read and long-read sequencing technologies.
There is an increased interest in the application of information technology to advance nutritional research. In nutrition science, a graph database enables the creation of multilateral logic relationships throughout the database, which can be used to electronically store, visualize, and scale the outputs of nutritional research. It provides a knowledge structure to standardize nutritional research outputs, which is both human- and machine-readable in a Resource Description Framework format. However, the development of various specific graph databases may cause difficulties for data integration and decrease human-readability. In this article, we propose an approach to develop a graph database according to the Data, Information, Knowledge, and Wisdom or “DIKW” pyramid for nutritional epidemiologic data. Then, authoritative ontologies are suggested to construct the nodes and edges of the graph database to facilitate data integration. Finally, the findability and re-usability of the knowledge in the graph database are showcased using the SPARQL and SQWRL query languages.
We present RNA-Bloom, a de novo RNA-seq assembly algorithm that leverages the rich information content in single-cell transcriptome sequencing (scRNA-seq) data to reconstruct cell-specific isoforms. We benchmark RNA-Bloom’s performance against leading bulk RNA-seq assembly approaches, and illustrate its utility in detecting cell-specific gene fusion events using sequencing data from HiSeq-4000 and BGISEQ-500 platforms. We expect RNA-Bloom to boost the utility of scRNA-seq data, expanding what is informatically accessible now.
Background: The use of linked data in the Semantic Web is a promising approach to add value to nutrition research. An ontology, which defines the logical relationships between well-defined taxonomic terms, enables linking and harmonizing research output. To enable the description of domain-specific output in nutritional epidemiology, we propose the Ontology for Nutritional Epidemiology (ONE) according to authoritative guidance for nutritional epidemiology. Methods: Firstly, a scoping review was conducted to identify existing ontology terms for reuse in ONE. Secondly, existing data standards and reporting guidelines for nutritional epidemiology were converted into an ontology. The terms used in the standards were summarized and listed separately in a taxonomic hierarchy. Thirdly, the ontologies of the nutritional epidemiologic standards, reporting guidelines, and the core concepts were gathered in ONE. Three case studies were included to illustrate potential applications: (i) annotation of existing manuscripts and data, (ii) ontology-based inference, and (iii) estimation of reporting completeness in a sample of nine manuscripts. Results: Ontologies for “food and nutrition” (n = 37), “disease and specific population” (n = 100), “data description” (n = 21), “research description” (n = 35), and “supplementary (meta) data description” (n = 44) were reviewed and listed. ONE consists of 339 classes: 79 new classes to describe data and 24 new classes to describe the content of manuscripts. Conclusion: ONE is a resource to automate data integration, searching, and browsing, and can be used to assess reporting completeness in nutritional epidemiology.
BACKGROUND:Alternative polyadenylation (APA) results in messenger RNA molecules with different 3' untranslated regions (3' UTRs), affecting the molecules' stability, localization, and translation. APA is pervasive and implicated in cancer. Earlier reports on APA focused on 3' UTR length modifications and commonly characterized APA events as 3' UTR shortening or lengthening. However, such characterization oversimplifies the processing of 3' ends of transcripts and fails to adequately describe the various scenarios we observe.RESULTS:We built a cloud-based targeted de novo transcript assembly and analysis pipeline that incorporates our previously developed cleavage site prediction tool, KLEAT. We applied this pipeline to elucidate the APA profiles of 114 genes in 9939 tumor and 729 tissue normal samples from The Cancer Genome Atlas (TCGA). The full set of 10,668 RNA-Seq samples from 33 cancer types has not been utilized by previous APA studies. By comparing the frequencies of predicted cleavage sites between normal and tumor sample groups, we identified 77 events (i.e. gene-cancer type pairs) of tumor-specific APA regulation in 13 cancer types; for 15 genes, such regulation is recurrent across multiple cancers. Our results also support a previous report showing the 3' UTR shortening of FGF2 in multiple cancers. However, over half of the events we identified display complex changes to 3' UTR length that resist simple classification like shortening or lengthening.CONCLUSIONS:Recurrent tumor-specific regulation of APA is widespread in cancer. However, the regulation pattern that we observed in TCGA RNA-seq data cannot be described as straightforward 3' UTR shortening or lengthening. Continued investigation into this complex, nuanced regulatory landscape will provide further insight into its role in tumor formation and development.
Background Joint data analysis from multiple nutrition studies may improve the ability to answer complex questions regarding the role of nutritional status and diet in health and disease. Objective The objective was to identify nutritional observational studies from partners participating in the European Nutritional Phenotype Assessment and Data Sharing Initiative (ENPADASI) Consortium, as well as minimal requirements for joint data analysis. Methods A predefined template containing information on study design, exposure measurements (dietary intake, alcohol and tobacco consumption, physical activity, sedentary behavior, anthropometric measures, and sociodemographic and health status), main health-related outcomes, and laboratory measurements (traditional and omics biomarkers) was developed and circulated to those European research groups participating in the ENPADASI under the strategic research area of "diet-related chronic diseases." Information about raw data disposition and metadata sharing was requested. A set of minimal requirements was abstracted from the gathered information. Results Studies (12 cohort, 12 cross-sectional, and 2 case-control) were identified. Two studies recruited children only and the rest recruited adults. All studies included dietary intake data. Twenty studies collected blood samples. Data on traditional biomarkers were available for 20 studies, of which 17 measured lipoproteins, glucose, and insulin and 13 measured inflammatory biomarkers. Metabolomics, proteomics, and genomics or transcriptomics data were available in 5, 3, and 12 studies, respectively. Although the study authors were willing to share metadata, most refused, were hesitant, or had legal or ethical issues related to sharing raw data. Forty-one descriptors of minimal requirements for the study data were identified to facilitate data integration. Conclusions Combining study data sets will enable sufficiently powered, refined investigations to increase the knowledge and understanding of the relation between food, nutrition, and human health. Furthermore, the minimal requirements for study data may encourage more efficient secondary usage of existing data and provide sufficient information for researchers to draft future multicenter research proposals in nutrition.