Adaptive immune receptor repertoire sequencing (AIRR-seq) enables large-scale profiling of B- and T-cell receptor diversity and has become a cornerstone of modern computational immunology. However, AIRR-seq provides only a partial and lossy molecular snapshot of immune dynamics, lacking explicit ground truth for clonal ancestry, lineage trajectories, antigen specificity, and longitudinal immune evolution. This limitation complicates benchmarking, method validation, and mechanistic interpretation of repertoire analysis pipelines. Here, we introduce UnivAIRRse, a unified hierarchical framework that organizes AIRR simulators within a shared conceptual coordinate system spanning five operational levels, from observed sequence data to the theoretical generative potential of the adaptive immune system. By explicitly distinguishing sequence-, clonal-, specificity-, repertoire-, and generative-level representations, UnivAIRRse enables systematic comparison of simulator assumptions, biological scope, abstraction level, and application focus. To our knowledge, this is the first review to formalize such a unified structure across biological, computational, and functional layers of AIRR simulation. Using this framework, we review how simulation supports benchmarking, strengthens computational inference, and enables multi-scale investigation of immune repertoire formation and evolution. We identify persistent limitations in existing simulators, including incomplete biological context, limited modularity, restricted interoperability, and overreliance on AIRR-seq as a molecular proxy for complex spatiotemporal immune processes. To operationalize this framework, we provide an interactive web-based AIRR Simulation Landscape Explorer (publicly available at https://www.imgt.org/AIRR-Simulator/ ) that enables dynamic filtering and comparison of simulators across biological scope, abstraction level, output fidelity, and application focus. Finally, we outline emerging directions toward digital-twin–ready immune simulation, emphasizing modular architectures, longitudinal multi-omic integration, uncertainty quantification, and dynamic model updating. By providing a coherent conceptual and operational coordinate system, UnivAIRRse establishes a foundation for reproducible, interpretable, and clinically actionable modeling of adaptive immune repertoires, bridging current simulation practices with the next generation of predictive and personalized immunological modeling.
Over thousands of years, selective breeding of domestic dogs (Canis lupus familiaris) has led to extensive genetic diversity, underscoring the need for refined genomic and immunogenetic investigations. This study presents a comprehensive analysis and biocuration of the immunoglobulin kappa light chain locus (IGK) across multiple dog breeds to identify breed-specific variations and assess their relevance to canine immunology and veterinary diagnostics. We examined nine canine genome assemblies to characterize structural variations, polymorphisms, and gene diversity, with the goal of enriching the IMGT® reference database and expanding its representativeness across breeds. Through in-depth annotation of seven breeds—Bernese Mountain Dog, Boxer, Cairn Terrier, Labrador Retriever, Great Dane, Basenji, and German Shepherd—we identified 40 genes and 97 alleles, highlighting both conserved regions and unique breed-specific variants. Variants were validated in silico against Sanger sequencing data. Importantly, discrepancies were observed in the CanFam3.1 Boxer reference genome, indicating possible sequencing or assembly artifacts, challenges in gene and allele nomenclature standardization, and a low-density genomic segment within the IGK locus. These findings refine current knowledge of IGK locus diversity and enhance IMGT® database accuracy, supporting future studies on immunogenetic variability, somatic hypermutation, and immune response mechanisms in canine health and disease.
A bstract High-throughput B-cell receptor sequencing has transformed the analysis of adaptive immunity, but benchmarking clonal grouping and lineage reconstruction methods remains limited by the absence of datasets with known evolutionary histories. Here we present Ancestra, a lineage-explicit simulator of B-cell receptor heavy-chain affinity maturation. Ancestra models stochastic V(D)J recombination, context-dependent somatic hypermutation, affinity-based selection and clonal expansion while recording complete parent-child relationships and mutation events. The framework generates BCR heavy-chain sequence datasets together with their corresponding ground-truth lineage trees, enabling direct benchmarking of lineage-aware analytical methods. Across simulations, Ancestra recapitulates key properties of human repertoires, including complementarity-determining region 3 length distributions, amino-acid usage patterns, junctional mutation patterns consistent with IMGT criteria and heterogeneous branching topologies. Simulated lineages also reveal multi-label lineage trees, in which identical nucleotide sequences can arise independently along distinct evolutionary paths. Ancestra provides a practical foundation for rigorous benchmarking of lineage-aware immune repertoire analysis.
Cancer remains one of the leading causes of mortality worldwide, accounting for approximately 9.7 million deaths in 2022. Faced with this significant public health challenge, therapeutic monoclonal antibodies (mAbs) have emerged as promising alternatives that may minimize the side effects associated with conventional treatments such as radiotherapy and chemotherapy. To support mAb research and development, IMGT®, the international ImMunoGeneTics information system, has established two standardized data sources namely IMGT/mAb-DB, a comprehensive database for mAbs, and, more recently, IMGT/mAb-KG, a dedicated knowledge graph for mAbs. Despite these advances, the development of therapeutic mAbs remains both time-consuming and financially burdensome—costs can reach up to 2.8 billion. To address this challenge and accelerate cancer treatment, mAb repurposing represents a promising alternative. Compared to existing general drug repurposing frameworks, in this study, we leveraged a specialized subset of an expert-curated mAb-specific data model dedicated to the oncology domain IMGT/mAbOnco-KG, to develop a scientific hypothesis generation application for mAb repurposing. We conducted an exhaustive benchmarking of 28 transductive knowledge graph embedding models, identifying BoxE as the superior architecture for capturing asymmetric relations inherent in mAb-target-disease interactions. A user-friendly web interface provides access to the tool, incorporating a novel trackback support allowing researchers to visualize the biological subgraphs supporting each prediction. Our application demonstrates the potential of transductive knowledge graph embedding techniques in the oncology domain by enabling the repurposing of existing mAbs for new therapeutic uses. Using this tool, we have identified two novel mAbs, loncastuximab tesirine and glofitamab, both currently undergoing clinical trials for the treatment of chronic lymphocytic leukemia. This decision-support tool thus facilitates the discovery of new therapeutic opportunities by effectively repositioning existing mAbs for oncological indications, potentially accelerating the development of cancer therapies and addressing critical public health needs. To handle new clinical indication entities, our future research direction is to build a foundation model in immunogenetics. In addition, due to the probabilistic nature of the generation process, experimental validation by specialists in mAbs engineering field is an important next step. Not applicable.
IMGT®, the international ImMunoGeneTics information system®, has advanced its comprehensive platform for the analysis of immunoglobulin (IG) and T-cell receptor (TR) genes through the development of new automated and scalable tools. This article presents major updates aligned with IMGT's three axes of research. Axis I introduces dynamic resources such as IMGT/GeneTables, IMGT/AssemblyComparison, IMGT/CDRLengths, IMGT/MultiGenomeViewer, enabling real-time access to annotated genomic data, and IMGT/StatAssembly to assess quality of IG/TR loci in assemblies. Axis II enhances repertoire analysis with a redesigned IMGT/GeneFrequency tool and new customization features in IMGT/V-QUEST, supporting flexible exploration of IG and TR gene expression. Axis III improves the accurate prediction of peptide-MHC thanks to IMGT/RobustpMHC. Additionally, the IMGT knowledge graph and its therapeutic extension, IMGT/mAb-KG, provide semantically structured access to >100 million immunogenetic triplets, integrating IMGT databases and linking IMGT content to external biomedical resources. These developments promote standardization, interoperability, and integrative analysis across immunogenetics and clinical applications, reinforcing IMGT's (accessible at https://www.imgt.org) role as a core reference in the era of FAIR data and personalized medicine.
The immensity of the T cell receptor (TCR) repertoire, shaped by combinatorial diversity, imprecise rearrangements, and clonal selection, makes it difficult to comprehend. The immgenT Project generated single-cell RNA and TCRseq to map paired αβTCR repertoires across 734 mouse T cell samples from diverse tissues, lineages and challenge conditions. Compositional analysis uncovered some extreme junctional architectures. Previously unreported recurrent recombinations suggested non-randomness in VDJ joining, broadening the precedent of quasi-invariant iNKT and MAIT TCRs. These, with public clonotypes linked to self or environmental factors, contribute to a highly skewed distribution of clonotype frequencies. Tissue analyses reveal compartmentalized and tissue-specific clonal expansions. Unproductive rearrangements of one V gene appeared to suppress the rearrangements of the same V gene on the second chromosome, suggesting that unproductive TCR transcripts may act as regulatory lncRNAs. This organism-wide look into the TCR repertoire offers novel insights on the evolutionary and immunological pressures on repertoire selection. ### Competing Interest Statement BV is an employee and stockholder of Repertoire, CB is an advisor and stockholder of Repertoire. Other authors declare no competing interests. National Institute of Allergy and Infectious Diseases, https://ror.org/043z4tv69, AI072073 Scientific Research National Center Université de Montpellier, https://ror.org/051escj72 Institut Universitaire de France, https://ror.org/055khg266 Institut du Développement et des Resources en Informatique Scientifique
Domestic ferrets (Mustela putorius furo) are important for modeling human respiratory diseases. However, ferret B and T cell receptors have not been completely identified or annotated, limiting immune repertoire studies. Here we performed long read transcriptome sequencing of ferret splenocyte and lymph node samples to obtain over 120,000 high-quality full-length immunoglobin (Ig) and T cell receptor (TCR) transcripts. We constructed a complete reference set of the constant regions of ferret Ig and TCR isotypes and chain types. We also systematically annotated germline Ig and TCR variable (V), diversity (D), joining (J), and constant (C) genes on a recent ferret reference genome assembly. We designed new ferret-specific immune repertoire profiling assays by targeting positions in constant regions without allelic diversity across 11 ferret genome assemblies, and experimentally validated them using a commercially compatible single-cell-based platform. These improved resources and assays will enable future studies to fully capture ferret immune repertoire diversity.
Over millennia, the selective breeding of dogs ( Canis lupus familiaris ) has generated remarkable genetic diversity among breeds, highlighting the need for comprehensive genomic and immunogenetic studies. This research provides detailed immunoglobulin kappa light chain locus (IGK) analysis across multiple dog breeds. It aims to uncover breed-specific genetic variations and their implications for immunology and veterinary medicine. The primary objectives were to do the biocuration of the IGK locus in nine canine genome assemblies, investigate structural variations, polymorphisms, and gene diversity, and to enrich the IMGT® database with comprehensive IGK data from diverse breeds, creating a more inclusive genetic resource. Our extensive annotation of breeds, including the Bernese Mountain Dog, Boxer, Cairn Terrier, Labrador Retriever, Great Dane, Basenji, and German Shepherd, identified 40 genes and 97 alleles, revealing both conserved genes and unique variants across these breeds, with in silico validation through Sanger sequencing. Notably, we analyzed discrepancies in the first reference assembly from the Boxer breed (Canfam3.1), highlighting potential errors in assembly, challenges in gene and allele nomenclature, and a low-density region within the canine IGK locus. This study not only refines the understanding of IGK locus diversity but also contributes to the IMGT® databases, advancing future research on immunogenetic variability, somatic mutations, and immune response dynamics in canine health and disease. ### Competing Interest Statement The authors have declared no competing interest. IMGT® is granted access to the High-Performance Computing (HPC) resources of Meso{at}LR and of the Centre Informatique National de lEnseignement Superieur (CINES), to Tres Grand Centre de Calcul (TGCC) of the Commissariat a lEnergie Atomique et aux Energies Alternatives (CEA) and Institut du developpement et des ressources en informatique scientifique (IDRIS) [036029 (2010-2024)] made by GENCI (Grand Equipement National de Calcul Intensif). IMGT® is currently supported by the Centre National de la Recherche Scientifique (CNRS) and the University of Montpellier. TGKs PhD thesis and internship project at IMGT® were funded by the São Paulo Research Foundation (FAPESP) in Brazil (Grants number 23/05951-5 and 21/09982-7). We acknowledge the support of Immun4Cure IHU Institute for innovative immunotherapies in autoimmune diseases (France 2030/ANR-23-IHUA-0009) as well as the support of the Institut Universitaire de France
Antibodies, or immunoglobulins (IG), are central to the vertebrate adaptive immune system, yet the genomic architecture of IG loci remains poorly characterized in many nonhuman primates. In this study we present the first comprehensive genomic analysis of the immunoglobulin (IG) loci (IGH, IGL, and IGK) in two critically endangered orangutan species; Pongo abelii (Sumatran orangutan) and Pongo pygmaeus (Bornean orangutan) across multiple genome assemblies. Using IMGT-standardized biocuration framework combined with read-level structural validation, we identified previously undocumented haplotype-specific variation, including multigene duplications, asymmetric gene absence, and species-specific expansions of variable gene families. Recombination signal sequence (RSS) and switch region analyses revealed conserved regulatory motifs with potential implications for V(D)J recombination and class-switch recombination. These findings underscore the complexity and evolutionary adaptability of IG loci in great apes and highlight the value of orangutans as key references for understanding immune system evolution in the Hominidae lineage.
Unraveling the genetic complexity of the human immunoglobulin heavy (IGH) chain locus provides valuable insights into the mechanisms underlying the efficacy and specificity of the adaptive immune response. Despite its crucial role, the IGH locus remains insufficiently characterized, with its allelic diversity and polymorphisms inadequately investigated. In this study, we present an analysis of the human IGH locus, incorporating 15 human genome assemblies from diverse ancestries, including African, European, Asian, Saudi, and mixed backgrounds. Through our examination of both maternal and paternal assemblies, we uncover novel IGH alleles, copy number variations (CNV), and polymorphisms, particularly within the variable (IGHV) region. Our findings reveal extensive and previously uncharacterized genetic variability in the constant (IGHC) region and distinct IMGT CNV forms across individuals. This research contributes to a significant enrichment of the IMGT® IGH reference directory, databases, tools and web resources, and lays the groundwork for an IMGT® haplotype database which can be progressively enriched as additional datasets become available. Such a resource promises to propel personalized immunogenomics forward, with exciting applications in cancer immunotherapy, COVID-19, and other immune-related diseases.
T cells play a pivotal role in the immune system, relying on their somatically rearranged T cell receptor (TCR) to recognize peptide-MHC complexes. A comprehensive and extensively used set of monoclonal antibodies (mAbs) against TCR variable regions was generated in the previous century. The separate identification of mAb-specific TCR-V proteins and TRV genes has resulted in multiple nomenclatures, making their relationships unclear. To formally re-establish this link and determine patterns of reactivity within TRV subfamilies, we sorted T cells from C57BL/6 mice positive for any one of a panel of 22 anti-V mAbs and determined their TRV genes by single-cell TCRseq. RNAseq data revealed consistently higher expression of repeated elements from the ERV1-family LTR RLTR6Mm (mapping to Gm20400) in cells utilizing TRBV segments encoded within a 66 kb genomic region between TRBV23 and TRBV30. Our findings provide a comprehensive resource for anti-mouse TCR mAb specificity and insight into V-gene usage biases and T cell function.
The human immunoglobulin light chain loci, kappa (IGK) and lambda (IGL), are structurally complex genomic regions with germline gene content that is not yet fully characterized. These loci are marked by extensive gene duplication, allelic diversity, and segmental duplications, features that contribute critically to the adaptive immune response. In this study, we present a comprehensive IMGT annotation of IGK and IGL using two high-quality human reference assemblies (GRCh38 and T2T-CHM13) along with 142 and 125 additional chromosomal-level haploid assemblies, respectively for each locus, from individuals representing all major human superpopulations. Detailed gene and allele annotation of the reference assemblies led to the identification of 5 novel IGKV genes and 8 new IGKV alleles, 16 new IGLV genes, and 22 novel IGLV alleles. These were confirmed through assembly read validation, presence in whole genome sequencing datasets, and recurrence in multiple assemblies. Gene-level identification across the broader dataset enabled assessment of structural variation (SV) at both loci. IGL displayed high conservation, with recurrent absence observed for only one gene. In contrast, IGK exhibited greater variability, including complete loss of the distal region in certain assemblies. This structural diversity was analyzed across superpopulations, allowing us to map potential patterns of gene presence and absence across different ancestral groups. All newly identified genes were consistently observed across individuals and genomic backgrounds. This work enhances the structural resolution of the IGK and IGL loci and expands the IMGT reference directory with newly described germline genes and alleles. The results provide a more complete view of light chain genomic diversity and serve as a valuable resource for studies of antibody gene repertoires, immunogenetic variation, monoclonal antibody development and population-level diversity. ### Competing Interest Statement The authors have declared no competing interest.
Monoclonal antibodies (mAbs) and fusion proteins for immune applications (FPIA) play a crucial role in treating autoimmune diseases and cancers by targeting cell-surface proteins and triggering multiple immune mechanisms. These functions are mediated by the crystallizable fragment (Fc) region of mAbs and fusion proteins, whose interaction with Fc gamma receptors (FcγRs) can be modulated through Fc amino acid (AA) engineering. To aid research in this area, we developed the IMGT/FcVariantsExplorer tool (https://www.imgt.org/fcvariantsexplorer/) to identify engineered AA changes or variants within the Fc region in mAb and fusion proteins sequences from IMGT/2Dstructure-DB, the AA sequence database of IMGT®, the international ImMunoGeneTics information system®. We used the IMGT® nomenclature of engineered Fc variants involved in antibody effector properties and formats, applying a standardized classification in five categories: ‘Effector,’ ‘Half-life,’ ‘Physicochemical properties,’ ‘Structure,’ and ‘Hybrid.’ We analyzed sequences from 1,107 mAbs and fusion proteins, identifying 483 entries with Fc AA changes, resulting in 211 unique Fc variants in the dataset. We also used web scraping to retrieve associated biological data from literature. All data have been integrated into IMGT/mAb-DB, with links to sequences in IMGT/2Dstructure-DB, enabling users to query Fc variants by their ‘Category’ or ‘Effect.’ This curated dataset reveals key trends in antibody engineering.
AbstractThe naked mole-rat (Heterocephalus glaber) is a long-lived rodent species showing resistance to the development of cancer. Although naked mole-rats have been reported to lack natural killer (NK) cells, γδ T cell-based immunity has been suggested in this species, which could represent an important arm of the immune system for antitumor responses. Here, we investigate the biology of these unconventional T cells in peripheral tissues (blood, spleen) and thymus of the naked mole-rat at different ages by TCR repertoire profiling and single-cell gene expression analysis. Using our own TCR annotation in the naked mole-rat genome, we report that the γδ TCR repertoire is dominated by a public invariant Vγ4-2/Vδ1-4 TCR, containing the complementary-determining-region-3 (CDR3)γ CTYWDSNYAKKLF / CDR3δ CALWELRTGGITAQLVF that are likely generated by short-homology-repeat-driven DNA rearrangements. This invariant TCR is specifically found in γδ T cells expressing genes associated with NK cytotoxicity and is generated in both the thoracic and cervical thymus of the naked mole-rat until adult life. Our results indicate that invariant Vγ4-2/Vδ1-4 NK-like effector T cells in the naked mole-rat can contribute to tumor immunosurveillance by γδ TCR-mediated recognition of a common molecular signal.
Dynamic changes in protein glycosylation impact human health and disease progression. However, current resources that capture disease and phenotype information focus primarily on the macromolecules within the central dogma of molecular biology (DNA, RNA, proteins). To gain a better understanding of organisms, there is a need to capture the functional impact of glycans and glycosylation on biological processes. A workshop titled “Functional impact of glycans and their curation” was held in conjunction with the 16th Annual International Biocuration Conference to discuss ongoing worldwide activities related to glycan function curation. This workshop brought together subject matter experts, tool developers, and biocurators from over 20 projects and bioinformatics resources. Participants discussed four key topics for each of their resources: (i) how they curate glycan function-related data from publications and other sources, (ii) what type of data they would like to acquire, (iii) what data they currently have, and (iv) what standards they use. Their answers contributed input that provided a comprehensive overview of state-of-the-art glycan function curation and annotations. This report summarizes the outcome of discussions, including potential solutions and areas where curators, data wranglers, and text mining experts can collaborate to address current gaps in glycan and glycosylation annotations, leveraging each other’s work to improve their respective resources and encourage impactful data sharing among resources. Database URL: https://wiki.glygen.org/Glycan_Function_Workshop_2023
The accurate prediction of peptide-major histocompatibility complex (MHC) class I binding probabilities is a critical endeavor in immunoinformatics, with broad implications for vaccine development and immunotherapies. While recent deep neural network based approaches have showcased promise in peptide-MHC (pMHC) prediction, they have two shortcomings: (i) they rely on hand-crafted pseudo-sequence extraction, (ii) they do not generalize well to different datasets, which limits the practicality of these approaches. While existing methods rely on a 34 amino acid pseudo-sequence, our findings uncover the involvement of 147 positions in direct interactions between MHC and peptide. We further show that neural architectures can learn the intricacies of pMHC binding using even full sequences. To this end, we present PerceiverpMHC that is able to learn accurate representations on full-sequences by leveraging efficient transformer based architectures. Additionally, we propose IMGT/RobustpMHC that harnesses the potential of unlabeled data in improving the robustness of pMHC binding predictions through a self-supervised learning strategy. We extensively evaluate RobustpMHC on eight different datasets and showcase an overall improvement of over 6% in binding prediction accuracy compared to state-of-the-art approaches. We compile CrystalIMGT, a crystallography-verified dataset presenting a challenge to existing approaches due to significantly different pMHC distributions. Finally, to mitigate this distribution gap, we further develop a transfer learning pipeline.
The evolutionary conserved Notch signaling pathway functions as a mediator of direct cell–cell communication between neighboring cells during development. Notch plays a crucial role in various fundamental biological processes in a wide range of tissues. Accordingly, the aberrant signaling of this pathway underlies multiple genetic pathologies such as developmental syndromes, congenital disorders, neurodegenerative diseases, and cancer. Over the last two decades, significant data have shown that the Notch signaling pathway displays a significant function in the mature brains of vertebrates and invertebrates beyond neuronal development and specification during embryonic development. Neuronal connection, synaptic plasticity, learning, and memory appear to be regulated by this pathway. Specific mutations in human Notch family proteins have been linked to several neurodegenerative diseases including Alzheimer’s disease, CADASIL, and ischemic injury. Neurodegenerative diseases are incurable disorders of the central nervous system that cause the progressive degeneration and/or death of brain nerve cells, affecting both mental function and movement (ataxia). There is currently a lot of study being conducted to better understand the molecular mechanisms by which Notch plays an essential role in the mature brain. In this study, an in silico analysis of polymorphisms and mutations in human Notch family members that lead to neurodegenerative diseases was performed in order to investigate the correlations among Notch family proteins and neurodegenerative diseases. Particular emphasis was placed on the study of mutations in the Notch3 protein and the structure analysis of the mutant Notch3 protein that leads to the manifestation of the CADASIL syndrome in order to spot possible conserved mutations and interpret the effect of these mutations in the Notch3 protein structure. Conserved mutations of cysteine residues may be candidate pharmacological targets for the potential therapy of CADASIL syndrome.
Through the analysis of immunoglobulin genes at the IGH, IGK, and IGL loci from four Gorilla gorilla gorilla genome assemblies, IMGT® provides an in-depth overview of these loci and their individual variations in a species closely related to humans. The similarity between gorilla and human IG gene organization allowed the assignment of gorilla IG gene names based on their human counterparts. This study revealed significant findings, including variability in the IGH locus, the presence of known and new copy number variations (CNVs), and the accurate estimation of IGHG genes. The IGK locus displayed remarkable homogeneity and lacked the gene duplication seen in humans, while the IGL locus showed a previously unconfirmed CNV in the J-C cluster. The curated data from these analyses, available on the IMGT website, enhance our understanding of gorilla immunogenetics and provide valuable insights into primate evolution.
Breast milk, often referred to as "liquid gold," is a complex biofluid that provides essential nutrients, immune factors, and developmental cues for newborns. Recent advancements in the field of exosome research have shed light on the critical role of exosomes in breast milk. Exosomes are nanosized vesicles that carry bioactive molecules, including proteins, lipids, nucleic acids, and miRNAs. These tiny messengers play a vital role in intercellular communication and are now being recognized as key players in infant health and development. This paper explores the emerging field of milk exosomics, emphasizing the potential of exosome fingerprinting to uncover valuable insights into the composition and function of breast milk. By deciphering the exosomal cargo, we can gain a deeper understanding of how breast milk influences neonatal health and may even pave the way for personalized nutrition strategies.