Abstract Human endogenous retroviruses (HERVs), as remnants of ancient exogenous retrovirus infected and integrated into germ cells, comprise ∼8% of the human genome. These HERVs have been implicated in numerous diseases, and extensive research has been conducted to uncover their specific roles. Despite these efforts, a comprehensive source of HERV-disease association still needs to be added. To address this gap, we introduce the HervD Atlas (https://ngdc.cncb.ac.cn/hervd/), an integrated knowledgebase of HERV-disease associations manually curated from all related published literature. In the current version, HervD Atlas collects 60 726 HERV-disease associations from 254 publications (out of 4692 screened literature), covering 21 790 HERVs (21 049 HERV-Terms and 741 HERV-Elements) belonging to six types, 149 diseases and 610 related/affected genes. Notably, an interactive knowledge graph that systematically integrates all the HERV-disease associations and corresponding affected genes into a comprehensive network provides a powerful tool to uncover and deduce the complex interplay between HERVs and diseases. The HervD Atlas also features a user-friendly web interface that allows efficient browsing, searching, and downloading of all association information, research metadata, and annotation information. Overall, the HervD Atlas is an essential resource for comprehensive, up-to-date knowledge on HERV-disease research, potentially facilitating the development of novel HERV-associated diagnostic and therapeutic strategies.
How N6-methyladenosine (m6A), the most abundant mRNA modification, contributes to primate tissue homeostasis and physiological aging remains elusive. Here, we characterize the m6A epitranscriptome across the liver, heart and skeletal muscle in young and old nonhuman primates. Our data reveal a positive correlation between m6A modifications and gene expression homeostasis across tissues as well as tissue-type-specific aging-associated m6A dynamics. Among these tissues, skeletal muscle is the most susceptible to m6A loss in aging and shows a reduction in the m6A methyltransferase METTL3. We further show that METTL3 deficiency in human pluripotent stem cell-derived myotubes leads to senescence and apoptosis, and identify NPNT as a key element downstream of METTL3 involved in myotube homeostasis, whose expression and m6A levels are both decreased in senescent myotubes. Our study provides a resource for elucidating m6A-mediated mechanisms of tissue aging and reveals a METTL3-m6A-NPNT axis counteracting aging-associated skeletal muscle degeneration.
Genome-wide association study has identified fruitful variants impacting heritable traits. Nevertheless, identifying critical genes underlying those significant variants has been a great task. Transcriptome-wide association study (TWAS) is an instrumental post-analysis to detect significant gene-trait associations focusing on modeling transcription-level regulations, which has made numerous progresses in recent years. Leveraging from expression quantitative loci (eQTL) regulation information, TWAS has advantages in detecting functioning genes regulated by disease-associated variants, thus providing insight into mechanisms of diseases and other phenotypes. Considering its vast potential, this review article comprehensively summarizes TWAS, including the methodology, applications and available resources.
Transcriptome-wide association studies (TWASs), as a practical and prevalent approach for detecting the associations between genetically regulated genes and traits, are now leading to a better understanding of the complex mechanisms of genetic variants in regulating various diseases and traits. Despite the ever-increasing TWAS outputs, there is still a lack of databases curating massive public TWAS information and knowledge. To fill this gap, here we present TWAS Atlas (https://ngdc.cncb.ac.cn/twas/), an integrated knowledgebase of TWAS findings manually curated from extensive literature. In the current implementation, TWAS Atlas collects 401,266 high-quality human gene-trait associations from 200 publications, covering 22,247 genes and 257 traits across 135 tissue types. In particular, an interactive knowledge graph of the collected gene-trait associations is constructed together with single nucleotide polymorphism (SNP)-gene associations to build up comprehensive regulatory networks at multi-omics levels. In addition, TWAS Atlas, as a user-friendly web interface, efficiently enables users to browse, search and download all association information, relevant research metadata and annotation information of interest. Taken together, TWAS Atlas is of great value for promoting the utility and availability of TWAS results in explaining the complex genetic basis as well as providing new insights for human health and disease research.
Abstract The National Genomics Data Center (NGDC), part of the China National Center for Bioinformation (CNCB), provides a family of database resources to support global research in both academia and industry. With the explosively accumulated multi-omics data at ever-faster rates, CNCB-NGDC is constantly scaling up and updating its core database resources through big data archive, curation, integration and analysis. In the past year, efforts have been made to synthesize the growing data and knowledge, particularly in single-cell omics and precision medicine research, and a series of resources have been newly developed, updated and enhanced. Moreover, CNCB-NGDC has continued to daily update SARS-CoV-2 genome sequences, variants, haplotypes and literature. Particularly, OpenLB, an open library of bioscience, has been established by providing easy and open access to a substantial number of abstract texts from PubMed, bioRxiv and medRxiv. In addition, Database Commons is significantly updated by cataloguing a full list of global databases, and BLAST tools are newly deployed to provide online sequence search services. All these resources along with their services are publicly accessible at https://ngdc.cncb.ac.cn.
Background Long non-coding RNA (lncRNA) exhibits a crucial role in multiple human malignancies. The expression of lncRNA LINC00511, reportedly, is aberrantly up-regulated in several types of tumors. Our research was aimed at deciphering the role and mechanism of LINC00511 in the progression of cervical cancer (CC). Method Quantitative real-time polymerase chain reaction (qRT-PCR) was performed to quantify the expression levels of LINC00511, miR-497-5p and MAPK1 mRNA in CC tissues and cell lines. Cell counting kit-8 (CCK-8), 5-bromo-2’-deoxyuridine (BrdU) and Transwell assays were conducted for detecting the proliferation, migration and invasion of CC cells. Dual-luciferase reporter gene experiments were performed to verify the targeting relationships amongst LINC00511, miR-497-5p and MAPK1. Besides, MAPK1 expression in CC cells was detected via Western blot after LINC00511 and miR-497-5p were selectively regulated. Results Up-regulation of LINC00511 expression in CC tissues and cell lines was observed, which was in association with tumor size, clinical stage and lymph node metastasis of the patients. LINC00511 overexpression facilitated the proliferation, migration and invasion of CC cells, while opposite effects were observed after knockdown of LINC00511. Mechanistically, LINC00511 was capable of targeting miR-497-5p and up-regulating MAPK1 expression. Conclusion LINC00511/miR-497-5p/MAPK1 axis regulates CC progression.
Somatic variants act as critical players during cancer occurrence and development. Thus, an accurate and robust method to identify them is the foundation of cutting-edge cancer genome research. However, due to low accessibility and high individual-/sample-specificity of the somatic variants in tumor samples, the detection is, to date, still crammed with challenges, particularly when lacking paired normal samples as control. To solve this burning issue, we developed a tumor-only somatic and germline variant identification method (TSomVar) using the random forest algorithm established on sample-specific variant datasets derived from genotype imputation, reads-mapping level annotation and functional annotation. We trained TSomVar by using genomic variant datasets of three major cancer types: colorectal cancer, hepatocellular carcinoma and skin cutaneous melanoma. Compared with existing tumor-only somatic variant identification tools, TSomVar shows excellent performances in somatic variant detection with higher accuracy and better capability of recalling for test datasets from colorectal cancer and skin cutaneous melanoma. In addition, TSomVar is equipped with the competence of accurately identifying germline variants in tumor samples. Taken together, TSomVar will undoubtedly facilitate and revolutionize somatic variant explorations in cancer research.
With the proliferating studies of human cancers by single-cell RNA sequencing technique (scRNA-seq), cellular heterogeneity, immune landscape and pathogenesis within diverse cancers have been uncovered successively. The exponential explosion of massive cancer scRNA-seq datasets in the past decade are calling for a burning demand to be integrated and processed for essential investigations in tumor microenvironment of various cancer types. To fill this gap, we developed a database of Cancer Single-cell Expression Map (CancerSCEM, https://ngdc.cncb.ac.cn/cancerscem), particularly focusing on a variety of human cancers. To date, CancerSCE version 1.0 consists of 208 cancer samples across 28 studies and 20 human cancer types. A series of uniformly and multiscale analyses for each sample were performed, including accurate cell type annotation, functional gene expressions, cell interaction network, survival analysis and etc. Plus, we visualized CancerSCEM as a user-friendly web interface for users to browse, search, online analyze and download all the metadata as well as analytical results. More importantly and unprecedentedly, the newly-constructed comprehensive online analyzing platform in CancerSCEM integrates seven analyze functions, where investigators can interactively perform cancer scRNA-seq analyses. In all, CancerSCEM paves an informative and practical way to facilitate human cancer studies, and also provides insights into clinical therapy assessments.
N6-Methyladenosine (m6A) messenger RNA methylation is a well-known epitranscriptional regulatory mechanism affecting central biological processes, but its function in human cellular senescence remains uninvestigated. Here, we found that levels of both m6A RNA methylation and the methyltransferase METTL3 were reduced in prematurely senescent human mesenchymal stem cell (hMSC) models of progeroid syndromes. Transcriptional profiling of m6A modifications further identified MIS12, for which m6A modifications were reduced in both prematurely senescent hMSCs and METTL3-deficient hMSCs. Knockout of METTL3 accelerated hMSC senescence whereas overexpression of METTL3 rescued the senescent phenotypes. Mechanistically, loss of m6A modifications accelerated the turnover and decreased the expression of MIS12 mRNA while knockout of MIS12 accelerated cellular senescence. Furthermore, m6A reader IGF2BP2 was identified as a key player in recognizing and stabilizing m6A-modified MIS12 mRNA. Taken together, we discovered that METTL3 alleviates hMSC senescence through m6A modification-dependent stabilization of the MIS12 transcript, representing a novel epitranscriptional mechanism in premature stem cell senescence.
The BIG Data Center at Beijing Institute of Genomics (BIG) of the Chinese Academy of Sciences provides freely open access to a suite of database resources in support of worldwide research activities in both academia and industry. With the vast amounts of omics data generated at ever-greater scales and rates, the BIG Data Center is continually expanding, updating and enriching its core database resources through big-data integration and value-added curation, including BioCode (a repository archiving bioinformatics tool codes), BioProject (a biological project library), BioSample (a biological sample library), Genome Sequence Archive (GSA, a data repository for archiving raw sequence reads), Genome Warehouse (GWH, a centralized resource housing genome-scale data), Genome Variation Map (GVM, a public repository of genome variations), Gene Expression Nebulas (GEN, a database of gene expression profiles based on RNA-Seq data), Methylation Bank (MethBank, an integrated databank of DNA methylomes), and Science Wikis (a series of biological knowledge wikis for community annotations). In addition, three featured web services are provided, viz., BIG Search (search as a service; a scalable inter-domain text search engine), BIG SSO (single sign-on as a service; a user access control system to gain access to multiple independent systems with a single ID and password) and Gsub (submission as a service; a unified submission service for all relevant resources). All of these resources are publicly accessible through the home page of the BIG Data Center at http://bigd.big.ac.cn.