Journal Article Accepted manuscript Meeting report of the GlySpace alliance and GaLSIC symposium Get access Kiyoko F Aoki-Kinoshita, Kiyoko F Aoki-Kinoshita Glycan and Life Systems Integration Center (GaLSIC), Soka University, Tokyo 192-8677, Japan Search for other works by this author on: Oxford Academic PubMed Google Scholar Frederique Lisacek, Frederique Lisacek SIB Swiss Institute of Bioinformatics & Computer Science Department, University of Geneva, Geneva, Switzerland Search for other works by this author on: Oxford Academic PubMed Google Scholar Raja Mazumder, Raja Mazumder The Department of Biochemistry & Molecular Medicine, The George Washington University Medical Center, Washington, DC 20037, United States Search for other works by this author on: Oxford Academic PubMed Google Scholar Rene Ranzinger, Rene Ranzinger Complex Carbohydrate Research Center, University of Georgia, 315 Riverbend Road, Athens, GA, 30602-4712, United States Search for other works by this author on: Oxford Academic PubMed Google Scholar Michael Tiemeyer, Michael Tiemeyer Complex Carbohydrate Research Center, University of Georgia, 315 Riverbend Road, Athens, GA, 30602-4712, United States Search for other works by this author on: Oxford Academic PubMed Google Scholar Issaku Yamada, Issaku Yamada The Noguchi Institute, Tokyo 173-0003, Japan Search for other works by this author on: Oxford Academic PubMed Google Scholar Nicolle H Packer Nicolle H Packer School of Natural Sciences, Macquarie University, North Ryde, Sydney, NSW 2109, Australia Corresponding author. Nicolle H. Packer, School of Natural Sciences, Macquarie University, North Ryde, Sydney, NSW 2109, Australia, [email protected] Search for other works by this author on: Oxford Academic PubMed Google Scholar Glycobiology, cwaf019, https://doi.org/10.1093/glycob/cwaf019 Published: 28 March 2025 Article history Received: 14 March 2025 Accepted: 16 March 2025 Published: 28 March 2025
The motifs exposed on surface glycans of mucins serve as receptors to a range of proteins expressed by microorganisms. The specificity of microbial glycan-binding proteins, i.e., lectins, towards human mucin epitopes is the result of co-evolution. The binding of microbes to mucins is described in several databases hosted on the UniLectin portal. In particular, UniLectin3D provides structural details on the recognition process and LectomeXplore allows for the identification of putative lectins in the genomes of mucin-associated microbiome species. The usage of these resources is illustrated in this chapter.
Lectins are ubiquitous proteins that interact with glycans in a variety of molecular processes and as such, also play a role in diseases, whether infectious, chronic or cancer-related. The systematic study of lectins is therefore essential, in particular for understanding cell-cell communication. Accumulated protein three-dimensional structural data in the past decades boosted advance in AI-based prediction and opened up new options to characterise lectins that are known to often be multimeric and multivalent. This article reviews the methods to obtain structures of lectins, the current data available for lectin 3D structures and their interactions, how this knowledge is used to classify these proteins and shows that the combination of an array of bioinformatics tools should make the prediction of binding specificity possible in a near future.
The MIRAGE (Minimum Information Required for A Glycomics Experiment) guidelines for mass spectrometry (MS) data were initially developed to standardize the reporting of instrumentation, data acquisition and analytical details of the MS-based identification of released glycans. However, the growing interest in the study of intact glycoproteins and recent advances in MS-based glycoproteomics now necessitate a revision and expansion of these guidelines. This update includes an enhanced section focused on glycan structure analysis (glycomics) and introduces a new component tailored to the specific requirements of glycoproteomics. It addresses both shared and unique aspects of each approach and highlights glycoinformatics resources designed to facilitate data submission in compliance with the updated standards.
MatrixDB, a member of the International Molecular Exchange consortium (IMEx), is a curated interaction database focused on interactions established by extracellular matrix (ECM) constituents including proteins, proteoglycans, glycosaminoglycans and ECM bioactive fragments. The architecture of MatrixDB was upgraded to ease interaction data export, allow versioning and programmatic access and ensure sustainability. The new version of the database includes more than twice the number of manually curated and experimentally-supported interactions. High-confidence predicted interactions were imported from the Integrated Interactions Database to increase the coverage of the ECM interactome. ECM and ECM-associated proteins of five species (human, murine, bovine, avian and zebrafish) were annotated with matrisome divisions and categories, which are used for computational analyses of ECM -omic datasets. Biological pathways from the Reactome Pathway Knowledgebase were also added to the biomolecule description. New transcriptomic and expanded proteomic datasets were imported in MatrixDB to generate cell- and tissue-specific ECM networks using the newly developed in-house Network Explorer integrated in the database. MatrixDB is freely available at https://matrixdb.univ-lyon1.fr. [GRAPHICS] .
The UniLectin portal (https://unilectin.unige.ch/) was designed in 2019 with the goal of centralising curated and predicted data on carbohydrate-binding proteins known as lectins. UniLectin is also intended as a support for the study of lectomes (full lectin set) of organisms or tissues. The present update describes the inclusion of several new modules and details the latest (https://unilectin.unige.ch/humanLectome/), covering our knowledge of the human lectome and comprising 215 unevenly characterised lectins, particularly in terms of structural information. Each HumanLectome entry is protein-centric and compiles evidence of carbohydrate recognition domain(s), specificity, 3D-structure, tissue-based expression and related genomic data. Other recent improvements regarding interoperability and accessibility are outlined.
Glycosylation is a unique posttranslational modification that dynamically shapes the surface of cells. Glycans attached to proteins or lipids in a cell or tissue are studied as a whole and collectively designated as a glycome. UniCarb-DB is a glycomic spectral library of tandem mass spectrometry (MS/MS) fragment data. The current version of the database consists of over 1500 entries and over 1000 unique structures. Each entry contains parent ion information with associated MS/MS spectra, metadata about the original publication, experimental conditions, and biological origin. Each structure is also associated with the GlyTouCan glycan structure repository allowing easy access to other glycomic resources. The database can be directly utilized by mass spectrometry (MS) experimentalists through the conversion of data generated by MS into structural information. Flexible online search tools along with a downloadable version of the database are easily incorporated in either commercial or open-access MS software. This chapter highlights UniCarb-DB online search tool to browse differences of isomeric structures between spectra, a peak matching search between user-generated MS/MS spectra and spectra stored in UniCarb-DB and more advanced MS tools for combined quantitative and qualitative glycomics.
Dynamic changes in protein glycosylation impact human health and disease progression. However, current resources that capture disease and phenotype information focus primarily on the macromolecules within the central dogma of molecular biology (DNA, RNA, proteins). To gain a better understanding of organisms, there is a need to capture the functional impact of glycans and glycosylation on biological processes. A workshop titled “Functional impact of glycans and their curation” was held in conjunction with the 16th Annual International Biocuration Conference to discuss ongoing worldwide activities related to glycan function curation. This workshop brought together subject matter experts, tool developers, and biocurators from over 20 projects and bioinformatics resources. Participants discussed four key topics for each of their resources: (i) how they curate glycan function-related data from publications and other sources, (ii) what type of data they would like to acquire, (iii) what data they currently have, and (iv) what standards they use. Their answers contributed input that provided a comprehensive overview of state-of-the-art glycan function curation and annotations. This report summarizes the outcome of discussions, including potential solutions and areas where curators, data wranglers, and text mining experts can collaborate to address current gaps in glycan and glycosylation annotations, leveraging each other’s work to improve their respective resources and encourage impactful data sharing among resources. Database URL: https://wiki.glygen.org/Glycan_Function_Workshop_2023
The development of a stable human gut microbiota occurs within the first year of life. Many open questions remain about how microfloral species are influenced by the composition of milk, in particular its content of human milk oligosaccharides (HMOs). The objective is to investigate the effect of the human HMO glycome on bacterial symbiosis and competition, based on the glycoside hydrolase (GH) enzyme activities known to be present in microbial species. We extracted from UniProt a list of all bacterial species catalysing glycoside hydrolase activities (EC 3.2.1.-), cross-referencing with the BRENDA database, and obtained a set of taxonomic lineages and CAZy family data. A set of 13 documented enzyme activities was selected and modelled within an enzyme simulator according to a method described previously in the context of biosynthesis. A diverse population of experimentally observed HMOs was fed to the simulator, and the enzymes matching specific bacterial species were recorded, based on their appearance of individual enzymes in the UniProt dataset. Pairs of bacterial species were identified that possessed complementary enzyme profiles enabling the digestion of the HMO glycome, from which potential symbioses could be inferred. Conversely, bacterial species having similar GH enzyme profiles were considered likely to be in competition for the same set of dietary HMOs within the gut of the newborn. We generated a set of putative biodegradative networks from the simulator output, which provides a visualisation of the ability of organisms to digest HMO and mucin-type O-glycans. B. bifidum, B. longum and C. perfringens species were predicted to have the most diverse GH activity and therefore to excel in their ability to digest these substrates. The expected cooperative role of Bifidobacteriales contrasts with the surprising capacities of the pathogen. These findings indicate that potential pathogens may associate in human gut based on their shared glycoside hydrolase digestive apparatus, and which, in the event of colonisation, might result in dysbiosis. The methods described can readily be adapted to other enzyme categories and species as well as being easily fine-tuneable if new degrading enzymes are identified and require inclusion in the model.
DNA, RNA, and proteins are synthesized using template molecules, but glycosylation is not believed to be constrained by a template. However, if cellular environment is the only determinant of glycosylation, all sites should receive the same glycans on average. This template-free assertion is inconsistent with observations of microheterogeneity-wherein each site receives distinct and reproducible glycan structures. Here, we test the assumption of template-free glycan biosynthesis. Through structural analysis of site-specific glycosylation data, we find protein-sequence and structural features that predict specific glycan features. To quantify these relationships, we present a new amino acid substitution matrix that describes glycoimpact-how glycosylation varies with protein structure. High-glycoimpact amino acids co-evolve with glycosites, and glycoimpact is high when estimates of amino acid conservation and variant pathogenicity diverge. We report hundreds of disease variants near glycosites with high-glycoimpact, including several with known links to aberrant glycosylation (e.g., Oculocutaneous Albinism, Jakob-Creutzfeldt disease, Gerstmann-Straussler-Scheinker, and Gaucher's Disease). Finally, we validate glycoimpact quantification by studying oligomannose-complex glycan ratios on HIV ENV, differential sialylation on IgG3 Fc, differential glycosylation on SARS-CoV-2 Spike, and fucose-modulated function of a tuberculosis monoclonal antibody. In all, we show glycan biosynthesis is accurately guided by specific, genetically-encoded rules, and this presents a plausible refutation to the assumption of template-free glycosylation. ### Competing Interest Statement This work is associated with a provisional patent filed by the authors, and Augment Biologics, founded by BK and NEL.
Glycosylation is described as a non-templated biosynthesis. Yet, the template-free premise is antithetical to the observation that different N-glycans are consistently placed at specific sites. It has been proposed that glycosite-proximal protein structures could constrain glycosylation and explain the observed microheterogeneity. Using site-specific glycosylation data, we trained a hybrid neural network to parse glycosites (recurrent neural network) and match them to feasible N-glycosylation events (graph neural network). From glycosite-flanking sequences, the algorithm predicts most human N-glycosylation events documented in the GlyConnect database and proposed structures corresponding to observed monosaccharide composition of the glycans at these sites. The algorithm also recapitulated glycosylation in Enhanced Aromatic Sequons, SARS-CoV-2 spike, and IgG3 variants, thus demonstrating the ability of the algorithm to predict both glycan structure and abundance. Thus, protein structure constrains glycosylation, and the neural network enables predictive in silico glycosylation of uncharacterized or novel protein sequences and genetic variants.
For decades, lectins have been used as probes in glycobiology and this usage has gradually spread to other domains of Life Science. Nowadays, researchers investigate glycan recognition with lectins in diverse biotechnology and clinical applications, addressing key questions regarding binding specificity. The latter is documented in scattered and heterogeneous sources, and this situation calls for a centralized and easy-access reference. To address this need, an on-line solution called BiotechLec (https://www.unilectin.eu/biotechlec) is proposed in a new section of UniLectin, a platform dedicated to lectin molecular knowledge.
Recent technological advances in glycobiology have resulted in a large influx of data and the publication of many papers describing discoveries in glycoscience. However, the terms used in describing glycan structural features are not standardized, making it difficult to harmonize data across biomolecular databases, hampering the harvesting of information across studies and hindering text mining and curation efforts. To address this shortcoming, the Glycan Structure Dictionary has been developed as a reference dictionary to provide a standardized list of widely used glycan terms that can help in the curation and mapping of glycan structures described in publications. Currently, the dictionary has 190 glycan structure terms with 297 synonyms linked to 3,332 publications. For a term to be included in the dictionary, it must be present in at least 2 peer-reviewed publications. Synonyms, annotations, and cross-references to GlyTouCan, GlycoMotif, and other relevant databases and resources are also provided when available. The purpose of this effort is to facilitate biocuration, assist in the development of text mining tools, improve the harmonization of search, and browse capabilities in glycoinformatics resources and help to map glycan structures to function and disease. It is also expected that authors will use these terms to describe glycan structures in their manuscripts over time. A mechanism is also provided for researchers to submit terms for potential incorporation. The dictionary is available at https://wiki.glygen.org/Glycan_structure_dictionary.
Glycosylation is a common post-translational modification of brain proteins including cell surface adhesion molecules, synaptic proteins, receptors and channels, as well as intracellular proteins, with implications in brain development and functions. Using advanced state-of-the-art glycomics and glycoproteomics technologies in conjunction with glycoinformatics resources, characteristic glycosylation profiles in brain tissues are increasingly reported in the literature and growing evidence shows deregulation of glycosylation in central nervous system disorders, including aging associated neurodegenerative diseases. Glycan signatures characteristic of brain tissue are also frequently described in cerebrospinal fluid due to its enrichment in brain-derived molecules. A detailed structural analysis of brain and cerebrospinal fluid glycans collected in publications in healthy and neurodegenerative conditions was undertaken and data was compiled to create a browsable dedicated set in the GlyConnect database of glycoproteins (https://glyconnect.expasy.org/brain). The shared molecular composition of cerebrospinal fluid with brain enhances the likelihood of novel glycobiomarker discovery for neurodegeneration, which may aid in unveiling disease mechanisms, therefore, providing with novel therapeutic targets as well as diagnostic and progression monitoring tools.
Glycosaminoglycans (GAGs) are complex polysaccharides exhibiting a vast structural diversity and fulfilling various functions mediated by thousands of interactions in the extracellular matrix, at the cell surface, and within the cells where they have been detected in the nucleus. It is known that the chemical groups attached to GAGs and GAG conformations comprise "glycocodes" that are not yet fully deciphered. The molecular context also matters for GAG structures and functions, and the influence of the structure and functions of the proteoglycan core proteins on sulfated GAGs and vice versa warrants further investigation. The lack of dedicated bioinformatic tools for mining GAG data sets contributes to a partial characterization of the structural and functional landscape and interactions of GAGs. These pending issues will benefit from the development of new approaches reviewed here, namely (i) the synthesis of GAG oligosaccharides to build large and diverse GAG libraries, (ii) GAG analysis and sequencing by mass spectrometry (e.g., ion mobility-mass spectrometry), gas-phase infrared spectroscopy, recognition tunnelling nanopores, and molecular modeling to identify bioactive GAG sequences, biophysical methods to investigate binding interfaces, and to expand our knowledge and understanding of glycocodes governing GAG molecular recognition, and (iii) artificial intelligence for in-depth investigation of GAGomic data sets and their integration with proteomics.
Introduction: One of the main challenges in bioinformatics has been and still is, the comparison of entities through the development of algorithms for similarity scoring and data clustering according to biologically relevant aspects. Glycoinformatics also faces this challenge, in particular regarding the automated comparison of protein and/or tissue glycomes, that remains a relatively uncharted territory. Methods: Low and high throughput experimental glycomic and glycoproteomic results were collected, revealing a bias toward N-linked glycomes. Then, N-glycomes were considered and represented as networks of related glycan compositions as opposed to lists of glycans. They were processed and compared through a java application generating graphs and another producing a similarity matrix based on graph content. Several scoring schemes (e.g., Jaccard index or cosine) were tested and evaluated using the Matthews Correlation Coefficient, in order to capture a meaningful protein and tissue N-glycome similarity. Results: Assuming that a glycome corresponds to a well-connected graph of glycan compositions, graph comparison has revealed gaps that can be interpreted as inconsistencies. The outcome of systematic graph comparison is both formal and practical. In principle, it is shown that the idiosyncrasy of current glycome data limits the definition of appropriate estimates for systematically comparing N-glycomes. Yet, several potentially interesting criteria could be identified in a series of use cases detailed in the study. Discussion: Differentially expressed glycomes are usually compared manually, but the resulting work tends to remain in publications due to the lack of dedicated tools. Even manually, cross-comparison is challenging mostly because different sets of features are used from one study to the other. The work presented here enables laying down guidelines for developing a software tool comparing glycomes based on appropriate definitions of similarity and suitable methods for its evaluation and implementation.