MOTIVATION:Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS:scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION:scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.
Glycans influence protein structure and function, modulate tissue development, regulate organ function, drive pathogen interactions, control inflammation and immunity, and impact most disease processes. However, glycan and glycosylation data are frequently difficult to access, existing across heterogeneous resources, with inconsistent representation and limited linkage to the genes, proteins, and protein sites that define their biological context. This fragmentation limits integration with genomics, proteomics, and other omics data, and constrains the discovery of glycan-mediated relationships. GlyGen is a knowledgebase designed to harmonize and integrate glycan, protein, and glycosylation data within a unified, glycosylation-centric data model. The model currently captures 58,841 glycans, 243,702 proteins, and 183,573 glycosylation sites, and extends to motifs, diseases, biomarkers, as well as germline and somatic sequence variations. Data are collected from public resources and literature, standardized using identifiers such as UniProt and GlyTouCan accessions, and integrated through ontology-driven workflows with evidence tracking and attribution. GlyGen supports free access via a web portal, APIs, downloads, and semantic web technologies. The current release (www.glygen.org) integrates diverse datasets across 16 organisms with extensive cross-references and provides an interoperable framework for advancing the integration of glycoscience into the broader biomedical data ecosystem.
While GlyTouCan provides stable identifiers for referencing glycan structures, they are not organized semantically. GNOme, a glycan naming and subsumption ontology and a member of the OBOFoundry, organizes GlyTouCan accessions for automated reasoning and interactive browsing of glycan structures by subsumption. GNOme makes it quick and easy to discover glycans with a specific degree of characterization; provides a text-based table of common synonyms for specific structures and compositions; enumerates glycan subsumption relationships for automated reasoning; and assigns each glycan to well-defined categories based on their degree of characterization. As an OBOFoundry ontology, GNOme can be readily integrated with other OBOFoundry ontologies and standards initiatives that need to refer to glycans with various degrees of characterization. GNOme is integrated with GlyGen, a glycoinformatics knowledge base, providing navigation to "related glycans," and expanding the utility of species and glycan classification annotations. GNOme is available at https://gnome.glyomics.org/ and via GlyGen, the OBO Foundry, and GitHub.
Over 50% of human proteins are estimated to be glycosylated, making glycosylation one of the most common post-translational modifications (PTMs) of proteins. A glycoinformatics resource such as the GlyGen knowledgebase, consisting of experimentally verified sequence-specific glycosylation sites, is critical for advancing research in glycobiology. Unfortunately, most experimental studies report glycosylation sites in free text format in scientific literature, mentioning gene names and amino acid positions without providing protein sequence identifiers, making it difficult to mine reported sites that can be mapped onto specific protein sequences. We have developed GlycoSiteMiner, which is an automated literature mining-based pipeline that extracts experimentally verified protein sequence-specific glycosylation sites from PubMed abstracts. The pipeline employs ML/AI algorithms to filter out incorrectly identified sites and has been applied to 33 million PubMed abstracts, identifying 1118 new sequence-specific glycosylation sites that were not previously present in the GlyGen resource.
Dynamic changes in protein glycosylation impact human health and disease progression. However, current resources that capture disease and phenotype information focus primarily on the macromolecules within the central dogma of molecular biology (DNA, RNA, proteins). To gain a better understanding of organisms, there is a need to capture the functional impact of glycans and glycosylation on biological processes. A workshop titled “Functional impact of glycans and their curation” was held in conjunction with the 16th Annual International Biocuration Conference to discuss ongoing worldwide activities related to glycan function curation. This workshop brought together subject matter experts, tool developers, and biocurators from over 20 projects and bioinformatics resources. Participants discussed four key topics for each of their resources: (i) how they curate glycan function-related data from publications and other sources, (ii) what type of data they would like to acquire, (iii) what data they currently have, and (iv) what standards they use. Their answers contributed input that provided a comprehensive overview of state-of-the-art glycan function curation and annotations. This report summarizes the outcome of discussions, including potential solutions and areas where curators, data wranglers, and text mining experts can collaborate to address current gaps in glycan and glycosylation annotations, leveraging each other’s work to improve their respective resources and encourage impactful data sharing among resources. Database URL: https://wiki.glygen.org/Glycan_Function_Workshop_2023
Cross referencing to genomic and imaging resources for individual cases. Example of cross-referencing (A) on the Clinical tab of PDC’s Explore page; B, on the Clinical tab and External References section of PDC study summary pages.
Motivation:Understanding genetic variation at the single-cell level is crucial for insights into cellular heterogeneity, clonal evolution, and gene expression regulation, but there is a scarcity of tools for visualizing and analyzing cell-level genetic variants. Results:We introduce scSNViz, a comprehensive R-based toolset for visualization and analysis of cell-specific expressed Single Nucleotide Variants (sceSNVs) within cell-barcoded single-cell RNA-sequencing (scRNA-seq) data. ScSNViz offers 3D sceSNV visualization capabilities for dimensionally reduced scRNA-seq gene expression data, compatibility with popular scRNA-seq processing tools like Seurat, cell-type classification tools such as SingleR and scType, and trajectory inference computation using Slingshot. Furthermore, scSNViz conducts estimation, summary, and graphical representation of statistical metrics pertaining to sceSNVs distribution and expression across individual cells. It also provides support for the analysis of individual sceSNVs as well as sets comprising multiple expressed sceSNVs of interest. Availability:ScSNViz is implemented as user-friendly R-scripts, freely available on https://horvathlab.github.io/NGS/scSNViz , supported by help utilities, and requiring no specialized bioinformatics skills for use.
Abstract Proteomics has emerged as a powerful tool for studying cancer biology, developing diagnostics, and therapies. With the continuous improvement and widespread availability of high-throughput proteomic technologies, the generation of large-scale proteomic data has become more common in cancer research, and there is a growing need for resources that support the sharing and integration of multi-omics datasets. Such datasets require extensive metadata including clinical, biospecimen, and experimental and workflow annotations that are crucial for data interpretation and reanalysis. The need to integrate, analyze, and share these data has led to the development of NCI’s Proteomic Data Commons (PDC), accessible at https://pdc.cancer.gov. As a specialized repository within the NCI Cancer Research Data Commons (CRDC), PDC enables researchers to locate and analyze proteomic data from various cancer types and connect with genomic and imaging data available for the same samples in other CRDC nodes. Presently, PDC houses annotated data from more than 160 datasets across 19 cancer types, generated by several large-scale cancer research programs with cohort sizes exceeding 100 samples (tumor and associated normal when available). In this article, we review the current state of PDC in cancer research, discuss the opportunities and challenges associated with data sharing in proteomics, and propose future directions for the resource. Significance: The Proteomic Data Commons (PDC) plays a crucial role in advancing cancer research by providing a centralized repository of high-quality cancer proteomic data, enriched with extensive clinical annotations. By integrating and cross-referencing with complementary genomic and imaging data, the PDC facilitates multi-omics analyses, driving comprehensive insights, and accelerating discoveries across various cancer types.
Recent technological advances in glycobiology have resulted in a large influx of data and the publication of many papers describing discoveries in glycoscience. However, the terms used in describing glycan structural features are not standardized, making it difficult to harmonize data across biomolecular databases, hampering the harvesting of information across studies and hindering text mining and curation efforts. To address this shortcoming, the Glycan Structure Dictionary has been developed as a reference dictionary to provide a standardized list of widely used glycan terms that can help in the curation and mapping of glycan structures described in publications. Currently, the dictionary has 190 glycan structure terms with 297 synonyms linked to 3,332 publications. For a term to be included in the dictionary, it must be present in at least 2 peer-reviewed publications. Synonyms, annotations, and cross-references to GlyTouCan, GlycoMotif, and other relevant databases and resources are also provided when available. The purpose of this effort is to facilitate biocuration, assist in the development of text mining tools, improve the harmonization of search, and browse capabilities in glycoinformatics resources and help to map glycan structures to function and disease. It is also expected that authors will use these terms to describe glycan structures in their manuscripts over time. A mechanism is also provided for researchers to submit terms for potential incorporation. The dictionary is available at https://wiki.glygen.org/Glycan_structure_dictionary.
Abstract Motivation In single-cell RNA-sequencing (scRNA-seq) data, stratification of sequencing reads by cellular barcode is necessary to study cell-specific features. However, apart from gene expression, the analyses of cell-specific features are not sufficiently supported by available tools designed for high-throughput sequencing data. Results We introduce SCExecute, which executes a user-provided command on barcode-stratified, extracted on-the-fly, single-cell binary alignment map (scBAM) files. SCExecute extracts the alignments with each cell barcode from aligned, pooled single-cell sequencing data. Simple commands, monolithic programs, multi-command shell scripts or complex shell-based pipelines are then executed on each scBAM file. scBAM files can be restricted to specific barcodes and/or genomic regions of interest. We demonstrate SCExecute with two popular variant callers—GATK and Strelka2—executed in shell-scripts together with commands for BAM file manipulation and variant filtering, to detect single-cell-specific expressed single nucleotide variants from droplet scRNA-seq data (10X Genomics Chromium System). In conclusion, SCExecute facilitates custom cell-level analyses on barcoded scRNA-seq data using currently available tools and provides an effective solution for studying low (cellular) frequency transcriptome features. Availability and implementation SCExecute is implemented in Python3 using the Pysam package and distributed for Linux, MacOS and Python environments from https://horvathlab.github.io/NGS/SCExecute. Supplementary information Supplementary data are available at Bioinformatics online.
Pan-cancer analysis of TCGA and CPTAC (proteomics) data shows that SULF1 and SULF2 are oncogenic in a number of human malignancies and associated with poor survival outcomes. Our studies document a consistent upregulation of SULF1 and SULF2 in HNSC which is associated with poor survival outcomes. These heparan sulfate editing enzymes were considered largely functional redundant but single-cell RNAseq (scRNAseq) shows that SULF1 is secreted by cancer-associated fibroblasts in contrast to the SULF2 derived from tumor cells. Our RNAScope and patient-derived xenograft (PDX) analysis of the HNSC tissues fully confirm the stromal source of SULF1 and explain the uniform impact of this enzyme on the biology of multiple malignancies. In summary, SULF2 expression increases in multiple malignancies but less consistently than SULF1, which uniformly increases in the tumor tissues and negatively impacts survival in several types of cancer even though its expression in cancer cells is low. This paradigm is common to multiple malignancies and suggests a potential for diagnostic and therapeutic targeting of the heparan sulfatases in cancer diseases.
ABSTRACT SULF1 and SULF2 are oncogenic in a number of human malignancies, including head and neck squamous cell carcinoma (HNSC). The function of these two heparan sulfate editing enzymes was previously considered largely redundant but the biology of cancer suggests differences that we explore in our RNAseq and RNAScope studies of HNSC and in a pan cancer analysis using the TCGA and CPTAC (proteomics) data. Our studies document a consistent upregulation of SULF1 and SULF2 in HNSC which is associated with poor survival outcomes. SULF2 expression increases in multiple malignancies but less consistently than SULF1, which uniformly increases in the tumor tissues and negatively impacts survival in several types of cancer. Meanwhile, SULF1 showed low expression in cancer cell lines and a scRNAseq study of HNSC shows that SULF1 is not supplied by epithelial tumor cells, like SULF2, but is secreted by cancer associated fibroblasts. Our RNAScope and PDX analysis of the HNSC tissues fully confirm the stromal source of SULF1 and explain the uniform impact of this enzyme on the biology of multiple malignancies. In summary, the SULF1 enzyme, supplied by a subset of cancer associated fibroblasts, is upregulated and negatively impacts HNSC survival at an early stage of the disease progression while the SULF2 enzyme, supplied by tumor cells, impacts survival at later stages of HNSC. This paradigm is common to multiple malignancies and suggests a potential for diagnostic and therapeutic targeting of the heparan sulfatases in cancer diseases.
We demonstrate a novel variant calling strategy using barcode-stratified alignments on 25 tumor and normal 10XGenomics scRNA-seq datasets (>200,000 cells). Our approach identified 24,528 exonic non-dbSNP single cell expressed (sce)SNVs, a third of which are shared across multiple samples. The novel sceSNVs include unreported somatic and germline variants, as well as RNA-originating variants; some are expressed in up to 17% of the cells, and many are found in known cancer genes. Our findings suggest that there is an unacknowledged repertoire of expressed genetic variants, possibly recurrent and common across samples, in the normal and cancer transcriptome.
Asymmetric allele content in the transcriptome can be indicative of functional and selective features of the underlying genetic variants. Yet, imbalanced alleles, especially from diploid genome regions, are poorly explored in cancer. Here we systematically quantify and integrate the variant allele fraction from corresponding RNA and DNA sequence data from patients with breast cancer acquired through The Cancer Genome Atlas (TCGA). We test for correlation between allele prevalence and functionality in known cancer-implicated genes from the Cancer Gene Census (CGC). We document significant allele-preferential expression of functional variants in CGC genes and across the entire dataset. Notably, we find frequent allele-specific overexpression of variants in tumor-suppressor genes. We also report a list of over-expressed variants from non-CGC genes. Overall, our analysis presents an integrated set of features of somatic allele expression and points to the vast information content of the asymmetric alleles in the cancer transcriptome. The cancer phenotype is largely driven by somatic mutations, whose carcinogenic effects are ultimately inter- vened by the transcription process 1–3 . As a mediator between genotype and phenotype, the tumor transcriptome reflects both advantage- selective pressure, and direct effects of the mutations on the transcription process. Hence, the tumor transcriptome is highly informative about the somatic functionality, especially through allele-specific approaches that can confine expressed structures to particular mutant alleles 1–4 . Several studies have explored the allele-specific transcriptional landscape of cancer 1, 5–10 . Preferentially expressed alleles are reported to play a role in epithelial ovarian cancer 7 , as well as in microRNA-implicated carcinogenesis, an example of which is miR-31 dysregulation in lung cancer 8 . Imbalanced allele expression can be caused by both large chromosomal alterations, such as copy number alterations (CNAs), and single nucleotide somatic mutations 1 . Nucleotide somatic mutations can affect the transcriptome through alteration of regulatory, splicing, or expression-rate modifying sites. Such effects commonly manifest in cis-fashion and directly impact the transcript abundance of the mutation bearing allele 1, 11, 12 . Mutations can also indirectly imbalance the allele content through changing the protein functions to either advance or impair the tumor growth. Functional mutations that provide selective advantage are referred to as drivers, and they are commonly targeted by either positive or negative selection forces to retain or deplete the growth-affecting allele 13–16 . Accordingly, somatic allele imbalance, including the extremes of loss or over-expression, can indicate tumorigenic functionality.