Abstract Glycans coat the surface of all cells, and every glycan is recognised by specific glycan-binding proteins (GBPs). There are no general tools that can accurately estimate the binding strength between glycan and GBP from the amino acid sequence of the GBP and the molecular structure of the glycan, represented as SMILES string. We describe models for predicting such binding strengths developed as a part of a Capstone Course at the University of Alberta. The models are trained on a dataset that combines BindingDB, a published database of small-molecule protein interactions, and data from glycan arrays measured by Consortium of Functional Glycomics (CFG). In this hybrid dataset of protein-ligand interactions the ligands are both glycans from CFG and small molecules from BindingDB; similarly, proteins include GBP and proteins from BindingDB. Three models are presented (i) ProMax which fuses ESM-2, MolFormer, and MolCLR features; (ii) APEX which constrains learning to a predetermined form, a physical model of binding; (iii) UltraMax adds inter-atomic distances for the ligands. To address the dataset’s severe long-tail distribution, the models employ tail-aware losses for rare high-binding instances. Trained and evaluated on approximately one million protein–ligand pairs using hold-out splits for unseen molecules, the three models provide a unified framework for quantitative glycan–protein binding prediction. We observed that learning glycan-protein binding is harder than the similar task of learning small-molecule-protein interactions. Simple mirror-inversion tests led us to postulate that insufficient use of chiral features is an important source of difficulty in learning these interactions.
Glycans regulate multiple physiological processes, including immune recognition and cancer progression. In disease, altered glycan landscapes are interpreted by human lectins. Functional glycan-lectin interactions are difficult to profile because glycans are not genome-encoded and their changes are poorly captured by existing multimodal methods. We present two platforms, single-cell outlining and transcriptome sequencing (scGOAT-seq) and GlycoScope, which use human lectins to enable functional glycan accessibility into single-cell and spatial multiomic measurements. ScGOAT-seq quantifies lectin-accessible glycan states with gene expression, while GlycoScope enables multiplexed in situ co-detection of glycans and proteins in tissues. Applying these approaches to immune cells, we identify stimulus-specific glycan remodeling and show that distinct Siglec-ligand-defined programs stratify immune activation states not captured by traditional methods; in follicular lymphoma, GlycoScope, resolves spatial glycan programs associated with malignant B cells and localized immune microenvironments. The presented methods provide a general framework for integrating functional glycan accessibility into single-cell and spatial multiomics.
We describe a machine-learned (ML) model, MCNet, which predicts interactions between proteins and glycans. MCNet predicted quantitative interactions between glycan-binding proteins (GBPs) and enantiomers of common glycans, which were not part of the original training datasets. l-glycans are rare in nature but are important in consideration of safety of putative mirror-image life-forms. Current ML models that predict properties of glycans from their monosaccharide composition cannot extrapolate properties of mirror glycans. Instead, MCNet uses an atom-level description of the glycan to output an estimate of binding to GBPs. MCNet is trained using data from glycan microarrays and affinity measurements unified using a "fraction bound" parameter. Trained MCNet predicted unexpected binding of l-glucose to some fucose-binding GBPs. Both glycan and lectin arrays conformed these predictions. ML models akin to MCNet reach beyond traditional glycobiology and make it possible to anticipate interaction between biomolecules in mirror-life forms and present-day life-forms.
Cross-chiral recognition in glycobiology is the interactions between biologically conventional proteins and the enantiomers of biological glycans (e.g., L-proteins binding with L-hexoses) from organisms across all kingdoms of life. By symmetry, it also describes the interactions of chirally mirrored proteins with normal D-glycans. Knowledge of cross-chiral recognition is critical to understanding the potential interactions of existing life forms with artificial mirror-life forms, but currently known rules of protein-glycan interaction are insufficient. To build a methodology for learning such interactions, we constructed machine learning models that predict binding strength between proteins and glycans represented as graphs of atoms, rather than monosaccharides. Atomic q-gram and Morgan fingerprint (MF) based representation of glycans made it possible to train ML models that predict lectin binding properties of glycans, glycomimetic compounds, and enantiomers of all natural glycans. Critical to this training was merging disparate data—some with relative fluorescence units (RFU) from glycan microarrays and others with Kd values from ITC—using a universal "fraction bound" parameter f at a specific lectin concentration. A fully-connected neural network architecture, MCNet takes a MF and concentration (C) as inputs and returns f for 147 lectins. Performance of MCNet is comparable to the GlyNet models, and by proxy to other state-of-the art models that predict strength of protein-glycan interactions. MCNet effectively predicts binding of glycomimetic compounds to Galectins 1, 3, and 7. Breaking from a monosaccharide-based description makes it possible for MCNet to predict cross-chiral recognition. We employed a Liquid Glycan Array to validate some predictions, such as the lack of interactions of L-mannose with D-mannose binding lectins, purified ConA, and DC-SIGN displayed on cells, and weak binding of L-Man to galactose-binding lectins. MCNet's atom-level input makes it possible to agglomerate protein-glycan data from diverse glycans across all kingdoms of life and non-glycan structures (e.g. glycomimetic compounds). The universal fraction bound parameter makes it possible to unify disparate quantitative observations (Kd/IC50, RFU, chromatographic retention times, etc.). We believe that such an approach will facilitate a merger of knowledge from diverse glycobiology datasets and predict protein interactions with uncommon/unnatural glycans not attainable from current ML models. ### Competing Interest Statement The authors have declared no competing interest.
Glycans constitute a significant fraction of biomolecular diversity on the surface of cells across all the species in all kingdoms of life. As the structure of glycans is not encoded by the DNA of the host organisms, it is impossible to use cutting-edge DNA technology to study the role of cellular glycosylation or to understand how cell-surface glycome is recognized by glycan-binding proteins (GBPs). To address this gap, we recently described a genetically-encoded liquid glycan array (LiGA) platform that allows profiling of glycan:GBP interactions on the surface of live cells in vitro and in vivo using next-generation sequencing (NGS). LiGA is a library of DNA-barcoded bacteriophages coated with 5-1500 copies of a glycan; the DNA barcode inside each bacteriophage encodes the structure and density of the displayed glycans. Deep sequencing of the glycophages associated with live cells yields a glycan-binding profile of GBPs displayed on the surface of such cells. This protocol provides detailed instructions of using LiGA to probe cell surface receptors and includes information on the preparation of glycophages, analysis by MALDI-TOF MS, the assembly of a LiGA library, and its deep-sequencing. Using the protocol detailed in this report, we measure a glycan-binding profile of the immunomodulatory SiglecLJ1, -2, -6, -7, and -9 expressed on the surface of different cell types and uncover previously unknown environment-dependent recognition of glycans by Siglec-receptors on the surface of live cells. Protocols similar to the one described in this report will make it possible to measure the precise glycan-binding profile of any GPBs displayed on the surface of any cell types.
Selective detection of disease-associated changes in the glycocalyx is an emerging field in modern targeted therapies. Detecting minor glycan changes on the cell surface is a challenge exacerbated by the lack of correspondence between cellular DNA/RNA and glycan structures. We demonstrate that multivalent displays of lectins on DNA-barcoded phages—liquid lectin array (LiLA)—detect subtle differences in density of glycans on cells. LiLA constructs displaying 73 copies of diCBM40 (CBM) lectin per virion (φ-CBM73) exhibit non-linear ON/OFF-like recognition of sialoglycans on the surface of normal and cancer cells. A high-valency φ-CBM290 display, or soluble CBM protein, cannot amplify the subtle differences detected by φ-CBM73. Similarly, multivalent displays of CBM and Siglec-7 detect differences in the glycocalyx between stem-like and non-stem populations in cancer. Multivalent display of lectins offer in situ detection of minor differences in glycocalyx in cells both in vitro and in vivo not feasible to currently available technologies.
The M13 phage platform is a stable and monodisperse nanoscale carrier, which can be modified with different molecules by chemical conjugation strategies. Here, we describe M13 phage acylated on pVIII protein with a dibenzocyclooctyne reacting with azido glycan to yield 30-1500 copy numbers of glycan per phage and monitored by MALDI-TOF spectrometry to generate multivalent glycoconjugates that contain desired densities of glycans. We prepared the liquid glycan arrays (LiGA) such that both the structure and density of glycans were encoded in the DNA of the bacteriophage. The LiGA can be used to validate the binding properties of glycans to purified lectins and explore the effect of glycan density on such binding. From a mixture of multivalent glycan probes, LiGAs can also identify the glycoconjugates with optimal avidity necessary for binding to lectins on living cells in vitro and live animals in vivo.
Selective detection of disease-associated changes in the cellular glycocalyx is a foundation of modern targeted therapies. Detecting minor changes in the density and identity of glycans on the cell surface is a technological challenge exacerbated by lack of 1:1 correspondence between cellular DNA/RNA and glycan structures on cell surface. We demonstrate that multivalent displays of up to 300 lectins on DNA-barcoded M13 phage on a liquid lectin array (LiLA), detects subtle differences in composition and density of glycans on cells ex vivo and in immune cells or organs in animals. For example, constructs displaying 73 copies of diCBM40 lectin per 700×5 nm virion (φ-CBM73) exhibit non-linear ON/OFF-like recognition of sialoglycans on the surface of normal and cancer cells. In contrast, a high-valency φ-CBM290 display, or soluble diCBM40, exhibit canonical progressive scaling in binding with increased epitope density; these constructs cannot amplify the subtle differences detected by φ-CBM73. Similarly, multivalent displays of diCBM40 and Siglec-7 detect differences in the glycocalyx between stem-like and non-stem populations in cancer cells that are not detected with soluble lectins. Multivalent display of lectins on M13 scaffold with protected DNA inside the phage offer non-destructive detection of minor differences in glycocalyx in cells in vitro and in vivo not feasible to currently available technologies.### Competing Interest StatementR.D. is shareholder of the start-up company 48Hour Discovery Inc. that licensed the patent application (WO2018141058A1) describing LiGA technology.
Cellular glycosylation is characterized by chemical complexity and heterogeneity, which is challenging to reproduce synthetically. Here we show chemoenzymatic synthesis on phage to produce a genetically-encoded liquid glycan array (LiGA) of complex type N -glycans. Implementing the approach involved by ligating an azide-containing sialylglycosyl-asparagine to phage functionalized with 50–1000 copies of dibenzocyclooctyne. The resulting intermediate can be trimmed by glycosidases and extended by glycosyltransferases yielding a phage library with different N -glycans. Post-reaction analysis by MALDI-TOF MS allows rigorous characterization of N -glycan structure and mean density, which are both encoded in the phage DNA. Use of this LiGA with fifteen glycan-binding proteins, including CD22 or DC-SIGN on cells, reveals optimal structure/density combinations for recognition. Injection of the LiGA into mice identifies glycoconjugates with structures and avidity necessary for enrichment in specific organs. This work provides a quantitative evaluation of the interaction of complex N -glycans with GBPs in vitro and in vivo.
Advances in diagnostics, therapeutics, vaccines, transfusion, and organ transplantation build on a fundamental understanding of glycan-protein interactions. To aid this, we developed GlyNet, a model that accurately predicts interactions (relative binding strengths) between mammalian glycans and 352 glycan-binding proteins, many at multiple concentrations. For each glycan input, our model produces 1257 outputs, each representing the relative interaction strength between the input glycan and a particular protein sample. GlyNet learns these continuous values using relative fluorescence units (RFUs) measured on 599 glycans in the Consortium for Functional Glycomics glycan arrays and extrapolates these to RFUs from additional, untested glycans. GlyNet's output of continuous values provides more detailed results than the standard binary classification models. After incorporating a simple threshold to transform such continuous outputs the resulting GlyNet classifier outperforms those standard classifiers. GlyNet is the first multi-output regression model for predicting protein-glycan interactions and serves as an important benchmark, facilitating development of quantitative computational glycobiology.
Shotgun metagenomics studies have improved our understanding of microbial population dynamics and have revealed significant contributions of microbes to gut homeostasis. They also allow in silico inference of the metagenome. While they link the microbiome with metabolic abnormalities associated with disease phenotypes, they do not capture microbial gene expression patterns that occur in response to the multitude of stimuli that constantly ambush the gut environment. Metatranscriptomics closes that gap, but its implementation is more expensive and tedious. We assessed the metabolic perturbations associated with gut inflammation using shotgun metagenomics and metatranscriptomics. Shotgun metagenomics detected changes in abundance of bacterial taxa known to be SCFA producers, which favors gut homeostasis. Bacteria in the phylum Firmicutes were found at decreased abundance, while those in phyla Bacteroidetes and Proteobacteria were found at increased abundance. Surprisingly, inferring the coding capacity of the microbiome from shotgun metagenomics data did not result in any statistically significant difference, suggesting functional redundancy in the microbiome or poor resolution of shotgun metagenomics data to profile bacterial pathways, especially when sequencing is not very deep. Obviously, the ability of metatranscriptomics libraries to detect transcripts expressed at basal (or simply low) levels is also dependent on sequencing depth. Nevertheless, metatranscriptomics informed about contrasting roles of bacteria during inflammation. Functions involved in nutrient transport, immune suppression and regulation of tissue damage were dramatically upregulated, perhaps contributed by homeostasis-promoting bacteria. Functions ostensibly increasing bacteria pathogenesis were also found upregulated, perhaps as a consequence of increased abundance of Proteobacteria. Bacterial protein synthesis appeared downregulated. In summary, shotgun metagenomics was useful to profile bacterial population composition and taxa relative abundance, but did not inform about differential gene content associated with inflammation. Metatranscriptomics was more robust for capturing bacterial metabolism in real time. Although both approaches are complementary, it is often not possible to apply them in parallel. We hope our data will help researchers to decide which approach is more appropriate for the study of different aspects of the microbiome.
Cation and anion channelrhodopsins (CCRs and ACRs, respectively) primarily from two algal species, Chlamydomonas reinhardtii and Guillardia theta, have become widely used as optogenetic tools to control cell membrane potential with light. We mined algal and other protist polynucleotide sequencing projects and metagenomic samples to identify 75 channelrhodopsin homologs from four channelrhodopsin families, including one revealed in dinoflagellates in this study. We carried out electrophysiological analysis of 33 natural channelrhodopsin variants from different phylogenetic lineages and 10 metagenomic homologs in search of sequence determinants of ion selectivity, photocurrent desensitization, and spectral tuning in channelrhodopsins. Our results show that association of a reduced number of glutamates near the conductance path with anion selectivity depends on a wider protein context, because prasinophyte homologs with a glutamate pattern identical to that in cryptophyte ACRs are cation selective. Desensitization is also broadly context dependent, as in one branch of stramenopile ACRs and their metagenomic homologs, its extent roughly correlates with phylogenetic relationship of their sequences. Regarding spectral tuning, we identified two prasinophyte CCRs with red-shifted spectra to 585 nm. They exhibit a third residue pattern in their retinalbinding pockets distinctly different from those of the only two types of red-shifted channelrhodopsins known (i.e., the CCR Chrimson and RubyACRs). In cryptophyte ACRs we identified three specific residue positions in the retinal-binding pocket that define the wavelength of their spectral maxima. Lastly, we found that dinoflagellate rhodopsins with a TCP motif in the third transmembrane helix and a metagenomic homolog exhibit channel activity. IMPORTANCE Channelrhodopsins are widely used in neuroscience and cardiology as research tools and are considered prospective therapeutics, but their natural diversity and mechanisms remain poorly characterized. Genomic and metagenomic sequencing projects are producing an ever-increasing wealth of data, whereas biophysical characterization of the encoded proteins lags behind. In this study, we used manual and automated patch clamp recording of representative members of four channelrhodopsin families, including a family in dinoflagellates that we report in this study. Our results contribute to a better understanding of molecular determinants of ionic selectivity, photocurrent desensitization, and spectral tuning in channelrhodopsins.
Abnormal cell surface glycosylation plays a major role in disease processes such as immune evasion. However, the underlying role of glycans is yet to be fully understood. Binding information obtained from glycan arrays can provide critical starting points for downstream applications such as the development of carbohydrate-based inhibitors, vaccines, and other therapeutics. However, it is challenging to use powerful techniques like DNA deep sequencing to analyze glycan recognition due to the lack of 1:1 correspondence between DNA and glycan structures. Therefore, we have developed Liquid Glycan Array (LiGA), a technology that allows for genetic encoding of glycans. LiGA provides a 1:1 correspondence between the glycan displayed in multiple copies on a bacteriophage carrier and the phage genetic material. LiGA is generated by acylation of phage pVIII protein with a dibenzocyclooctyne, followed by ligation of azido-modified glycans. The display of glycans on each phage virion can be controlled from 30-1500 copies to probe the critical variables in glycan recognition: valency and density. A simple pulldown of the LiGA along with lectins followed by deep sequencing of the DNA in the bound phage decodes the recognized glycans. LiGA is target agnostic and measures binding profile of lectins expressed on intact cells, such as hCD22 (Siglec-2) and DC-SIGN (Dendritic Cell-Specific Intercellular adhesion molecule-3-Grabbing Non-integrin), and in live mice (Nat. Chem. Bio. 17, 806-816, 2021). From a mixture of 50-100 multivalent glycan probes, LiGA identifies the glycan-phage conjugates with optimal valency and density for binding to antibodies and lectins on cells in vitro and in vivo. Sialic acid-binding immunoglobulin-type lectins (Siglecs) expressed on the surface of immune cells are exploited by cancer to evade immune response. We applied LiGA to study the binding specificity of Siglec-7, a cell surface receptor that cancer cells use to evade immune response from natural killer (NK) cells. Additionally, we explored the roles of valency and density in ganglioside interaction with Siglec-1 using a cell-based assay. Building on these successes, we plan to use LiGA to identify the valency and affinity required by trans- glycan to overcome the cis- masking on the surface of immune cells.
The Central Dogma of Biology does not allow for the study of glycans using DNA sequencing. We report a “Liquid Glycan Array” (LiGA) platform comprising a library of DNA ‘barcoded’ M13 virions that display 30-1500 copies of glycans per phage. A LiGA is synthesized by acylation of phage pVIII protein with a dibenzocyclooctyne, followed by ligation of azido-modified glycans. Pulldown of the LiGA with lectins followed by deep sequencing of the barcodes in the bound phage decodes the optimal structure and density of the recognized glycans. The LiGA is target agnostic and can measure the glycan-binding profile of lectins such as CD22 on cells in vitro and immune cells in a live mouse. From a mixture of multivalent glycan probes, LiGAs identifies the glycoconjugates with optimal avidity necessary for binding to lectins on living cells in vitro and in vivo ; measurements that cannot be performed with canonical glass slide-based glycan arrays. Dedication The paper is dedicated to Laura L. Kiessling on the occasion of her 60th birthday.
This protocol is part of a collection of eighteen protocols used to isolate total RNA from plant tissue. (RNA Isolation from Plant Tissue Collection: https://www.protocols.io/view/rna-isolation-from-plant-tissue-439gyr6) and was originally published as part of Appendix S1 of "Evaluating Methods for Isolating Total RNA and Predicting the Success of Sequencing Phylogenetically Diverse Plant Transcriptomes" Marc T. J. Johnson et al. PLOS ONE, November 21, 2012. https://doi.org/10.1371/journal.pone.0050226
RNA-Seq data is inherently nonuniform for different transcripts because of differences in gene expression. This makes it challenging to decide how much data should be generated from each sample. How much should one spend to recover the less expressed transcripts? The sequencing technology used is another consideration, as there are inevitably always biases against certain sequences. To investigate these effects, we first looked at high-depth libraries from a set of well-annotated organisms to ascertain the impact of sequencing depth on de novo assembly. We then looked at libraries sequenced from the Universal Human Reference RNA (UHRR) to compare the performance of Illumina HiSeq and MGI DNBseq™ technologies. On the issue of sequencing depth, the amount of exomic sequence assembled plateaued using data sets of approximately 2 to 8 Gbp. However, the amount of genomic sequence assembled did not plateau for many of the analyzed organisms. Most of the unannotated genomic sequences are single-exon transcripts whose biological significance will be questionable for some users. On the issue of sequencing technology, both of the analyzed platforms recovered a similar number of full-length transcripts. The missing “gap” regions in the HiSeq assemblies were often attributed to higher GC contents, but this may be an artefact of library preparation and not of sequencing technology. Increasing sequencing depth beyond modest data sets of less than 10 Gbp recovers a plethora of single-exon transcripts undocumented in genome annotations. DNBseq™ is a viable alternative to HiSeq for de novo RNA-Seq assembly.
BACKGROUND:The 1000 Plant transcriptomes initiative (1KP) explored genetic diversity by sequencing RNA from 1,342 samples representing 1,173 species of green plants (Viridiplantae). FINDINGS:This data release accompanies the initiative's final/capstone publication on a set of 3 analyses inferring species trees, whole genome duplications, and gene family expansions. These and previous analyses are based on de novo transcriptome assemblies and related gene predictions. Here, we assess their data and assembly qualities and explain how we detected potential contaminations. CONCLUSIONS:These data will be useful to plant and/or evolutionary scientists with interests in particular gene families, either across the green plant tree of life or in more focused lineages.
The carbohydrate-rich cell walls of land plants and algae have been the focus of much interest given the value of cell wall-based products to our current and future economies. Hydroxyproline-rich glycoproteins (HRGPs), a major group of wall glycoproteins, play important roles in plant growth and development, yet little is known about how they have evolved in parallel with the polysaccharide components of walls. We investigate the origins and evolution of the HRGP superfamily, which is commonly divided into three major multigene families: the arabinogalactan proteins (AGPs), extensins (EXTs), and proline-rich proteins. Using motif and amino acid bias, a newly developed bioinformatics pipeline, we identified HRGPs in sequences from the 1000 Plants transcriptome project (www.onekp.com). Our analyses provide new insights into the evolution of HRGPs across major evolutionary milestones, including the transition to land and the early radiation of angiosperms. Significantly, data mining reveals the origin of glycosylphosphatidylinositol (GPI)-anchored AGPs in green algae and a 3- to 4-fold increase in GPI-AGPs in liverworts and mosses. The first detection of cross-linking (CL)-EXTs is observed in bryophytes, which suggests that CL-EXTs arose though the juxtaposition of preexisting SPn EXT glycomotifs with refined Y-based motifs. We also detected the loss of CL-EXT in a few lineages, including the grass family (Poaceae), that have a cell wall composition distinct from other monocots and eudicots. A key challenge in HRGP research is tracking individual HRGPs throughout evolution. Using the 1000 Plants output, we were able to find putative orthologs of Arabidopsis pollen-specific GPI-AGPs in basal eudicots.
Australian Research Council Centre of Excellence in Plant Cell Walls, School of BioSciences, University of Melbourne, Parkville, Victoria 3010, Australia (K.L.J., A.M.C., A.L., A.B., M.S.D.); Departments of Biological Sciences and Medicine, University of Alberta, Edmonton, Alberta, Canada, and BGI-Shenzhen, Bei Shan Industrial Zone, Yantian District, Shenzhen, China (G.K.-S.W., E.J.C.); Florida Museum of Natural History, Department of Biology, University of Florida, Gainsville, Florida 32611 (D.E.S., N.W.M.); Botanical Institute, Cologne Biocenter, University of Cologne, D50674 Cologne, Germany (M.M., B.M.); Department of Biology, University of British Columbia, Kelowna, British Columbia V1V 1V7, Canada (M.K.D.) Department of Plant Biology, University of Georgia, Athens, Georgia 3062 (J.L.-M.); University Herbarium and Department of Integrative Biology, University of California, Berkeley, California 94720 (C.J.R.); New York Botanical Garden, Bronx, New York 10458 (D.W.S.); Department of Botany, University of British Columbia, Vancouver, British Columbia V6T 1Z4, Canada (S.W.G.); Key Laboratory of Genome Science and Information, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing 100101, China (X.W., S.W.); Division of Biological Sciences and Bond Life Sciences Center, University of Missouri, Columbia, Missouri 65211 (J.C.P.); Department of Horticulture, Michigan State University, East Lansing, Michigan 48823 (P.P.E.); and School of Agriculture, Food, and Wine, University of Adelaide, Waite Research Institute, Glen Osmond, South Australia 5064, Australia (C.J.S.)