The marginal zone (MZ) of the spleen harbors B cells that play an indispensable role in immune defense, orchestrating rapid responses to bloodborne pathogens. Under steady-state conditions, MZ B cell numbers are maintained through the balance of de novo generation from precursors, proliferative self-renewal, and loss. The mechanisms governing this homeostatic control remain elusive. Further, the developmental pathways underlying the establishment and continued supplementation of the MZ B cell compartment are not fullyelucidated. To address these gaps, we combined multiple fate-mapping tools and mathematical models to study MZ B cell dynamics in mice across the life course. Our analyses find evidence of quorum sensing mechanisms that regulate both the accumulation of mature MZ B cells during early life and their maintenance throughout adulthood. Specifically, we demonstrate that they derive predominantly from transitional B cell precursors with an efficiency that increases with age, reaching stable levels only in adulthood. MZ B cells compensate for this early developmental inefficiency through cell density-dependent proliferation, ensuring the timely establishment of a stable pool. Collectively, these findings unveil critical roles of quorum sensing and immune system maturation in the maintenance of this vital B cell subset.
Infections and vaccinations elicit coordinated humoral and cellular adaptive immune responses that together provide protection. In addition to antibodies, pathogen-specific memory T and B cells persist in blood and tissues, but it remains unclear how their composition and spatial distribution relate to serum antibody titers, the most common correlate of vaccine-induced protection. Understanding these relationships is essential for predicting vaccine efficacy and optimizing immunization strategies. We analyzed tissues from 58 adult human organ donors vaccinated against SARS-CoV-2, including individuals with and without prior infection. Using multivariate imputation, dimensionality reduction, and correlation, regression, and causal analyses, we identified immune signatures linking memory B cell, CD4 T cell, and CD8 T cell subsets in spleen, lung, and lung-draining lymph nodes with antibody titers and neutralizing activity. Our analyses indicate that humoral immunity is driven primarily by virus-specific B cells and CD4 T cells in lymphoid tissues rather than blood, whereas tissue-localized CD8 T cell responses, although correlated with antibody levels, develop independently. These findings demonstrate that cross-sectional immune profiling across multiple tissues recapitulates established immunological principles and reveal that serum antibody responses emerge from coordinated cellular immune responses distributed throughout the body.
Multiple sclerosis (MS) is a complex immune-mediated disorder with polygenic and multicellular underpinnings, necessitating cell-type-specific molecular studies to delineate dysregulated pathways. Here, we profile 1,075 transcriptomes from 167 patients with MS and 42 healthy participants across six peripheral immune cell-type-states. MS-associated transcriptional differences are more pronounced in primary (unstimulated) immune cells than in in vitro-stimulated counterparts. We identify shared and cell-type-specific transcriptional alterations at the level of genes, pathways, and co-expressed gene modules, prioritizing regulators, such as ZBTB16, across T cells and monocytes, and replicating six MS-associated modules in independent datasets. The top T cell module is enriched for MS susceptibility genes and affects proliferation. The top monocyte module implicates dysregulated TNF-α/NF-κB signaling, for which an in silico drug screen and in vitro validation nominate alvespimycin as a candidate modulator. Together, these findings define stable peripheral immune dysregulation signatures in MS that may serve as diagnostic or prognostic biomarkers in at-risk individuals.
Three ingredients create the strikingly variable immune trajectories observed across human populations: (i) genetic variation, (ii) individual-scale generation of additional genetic diversity solely in lymphocytes, and (iii) collision of these processes with a changing environment through life, triggering plastic responses that can have lasting effects. We explore what generates or reduces immune variation, integrating recent empirical advances with potential overarching evolutionary principles. We posit that the multiple components of the immune system each present trade-offs (such as self-defense-self-harm, and investments in memory-retaining flexibility for future responses) that vary through life, and that variation in how immune components balance these trade-offs increases homeostasis to external and internal threats. Disentangling the mechanisms that shape immune diversity could inform effective strategies to manage microbial interactions and mitigate the burden of immune-mediated chronic diseases.
Tissue-resident memory T cells (T RM ) protect from repeat infections within organs and barrier sites. The breadth and duration of such protection are defined at minimum by three quantities: the rate at which new T RM are generated from precursors, their rate of self-renewal, and their rate of loss through death, egress, or differentiation. Quantifying these processes individually is challenging. Here we combine genetic fate mapping tools and mathematical models to untangle these basic homeostatic properties of CD4 + T RM in the skin and gut lamina propria (LP) of healthy adult mice. We show that CD69 + CD4 + T RM in skin reside for ∼24 days and self-renew more slowly, such that clones halve in size approximately every 5 weeks, and approximately 2% of cells are replaced daily from precursors. CD69 + CD4 + T RM in LP have shorter residencies (∼14 days) and are maintained largely by immigration (4–6% per day). We also find evidence that the continuous replacement of CD69 + CD4 + T RM at both sites derives from circulating effector-memory CD4 + T cells, in skin possibly via a local CD9 − intermediate. Our approach maps the ontogeny of CD4 + T RM in skin and LP and exposes their dynamic and distinct behaviours, with continuous seeding and erosion potentially impacting the duration of immunity at these sites.
Ensembl (www.ensembl.org) is an open platform integrating publicly available genomics data across the tree of life with a focus on eukaryotic species related to human health, agriculture and biodiversity. This year has seen a continued expansion in the number of species represented, with >4800 eukaryotic and >31 300 prokaryotic genomes available. The new Ensembl site, currently in beta, has continued to develop, currently holding >2700 eukaryotic genome assemblies. The new site provides genome, gene, transcript, homology and variation views, and will replace the current Rapid Release site; this represents a key step towards provision of a single integrated Ensembl site. Additional activities have included developing improved regulatory annotation for human, mouse and agricultural species, and expanding the Ensembl Variant Effect Predictor tool. To learn more about Ensembl, help and documentation are available along with an extensive training program that can be accessed via our training pages.
Regulatory T cells (T reg cells) are critical regulators of adaptive immunity and the pathophysiology of antitumoral immunity. T reg cells are both generated during thymic development and induced from peripheral conventional T cells. How these distinct pathways contribute to the homeostasis of circulating T reg cells in health and disease remains unclear. We addressed this question using multiple fate-mapping mouse systems and modeling. Naive and effector/memory (EM) T reg cells exhibit distinct dynamics but are both continuously replenished by de novo generation throughout life. The predominant precursors of circulating EM T reg cells are naive thymic T reg cells and not conventional T cells, a process driven by self rather than foreign antigen recognition. Using the same fate reporters and three tumor models, we demonstrate that infiltrating T reg cells specifically derive from preexisting EM T reg cells. In summary, we define a linear ontogeny of T reg cells from the thymus to EM, driven by self-antigen recognition, that then gives rise to tumor-infiltrating T reg cells.
Tissue-resident memory T cells (TRM) provide regionalized immunity through potent effector responses and optimal localization in tissues. While mouse models have offered significant insights into TRM against single or heterosubtypic infection, they fail to recapitulate the complexity of human tissue exposed to a lifetime of antigens. Here, we developed a mouse model to track Ag-specific TRM each independently responding to sequentially administered respiratory virus, enabling examination of how pre-existing TRM impact newly-recruited populations, and vice-versa. Although prior exposure to IAV (X31) infection did not hinder the effector response to IBV, establishment of memory was impaired. Notably, lung IBV-specific CD4 T cells shifted from a CD69+CXCR6+ TRM phenotype after single infection, to CXCR6-FR4+ Tfh-like cells in previously infected mice, localizing away from airway epithelia toward B cell clusters. However, depleting the original IAV-specific T cells restored the IBV-specific TRM response in double infected mice to that of a single IBV infection. Finally, IBV infection followed by X31 priming and PR8 rechallenge elicited greater morbidity than heterosubtypically rechallenged mice without a prior infection, suggesting competition alters TRM phenotype, localization, and protective capacity. Our data highlight the impact of our immunological past on future T cell responses, and the importance of our formative exposures in establishing an optimally protective niche. NIH AI150680 Mucosal and Regional Immunology (MUC)
Memory T cells are maintained in tissues as circulating effector-memory (TEM) and tissue-resident (TRM) populations for protective immunity, though the role of site and subset in memory persistence remains undefined. Here, we investigated age-associated dynamics of human T cells in lymphoid organs, mucosal sites, and blood over 10 decades of life using retrospective radiocarbon (14C) birth dating, along with cellular, transcriptome, and epigenetic profiling. Memory T cells across peripheral sites exhibited continuous turnover with mean lifespans of 1-2 years, while the spleen contained longer-lived T cells. Over age, TEM cells expressed senescent markers and a GZMK transcriptional signature, while TRM cells maintained site-specific resident phenotypes without exhibiting features of senescence. Both TEM and TRM cells showed age-associated DNA hypomethylation, though TRM cells exhibited more epigenetically regulated genes. Together, our findings reveal asynchronous aging of human memory T cells by subset and site, as well as persistence of TRM cells without immunosenescence.
Mechanistic models of dynamic, interacting cell populations have yielded many insights into the growth and resolution of immune responses. Historically these models have described the behavior of pre-defined cell types based on small numbers of phenotypic markers. The ubiquity of deep phenotyping therefore presents a new challenge; how do we confront tractable and interpretable mathematical models with high-dimensional data? To tackle this problem, we studied the development and persistence of lung-resident memory CD4 and CD8 T cells (TRM) in mice infected with influenza virus. We developed an approach in which dynamical model parameters and the population structure are inferred simultaneously. This method uses deep learning and stochastic variational inference and is trained on the single-cell flow-cytometry data directly, rather than on the kinetics of pre-identified clusters. We show that during the resolution phase of the immune response, memory CD4 and CD8 T cells within the lung are phenotypically diverse, with subsets exhibiting highly distinct and time-dependent dynamics. TRM heterogeneity is maintained long-term by ongoing differentiation of relatively persistent Bcl-2hi CD4 and CD8 TRM subsets which resolve into distinct functional populations. Our approach yields new insights into the dynamics of tissue-localized immune memory, and is a novel basis for interpreting time series of high-dimensional data, broadly applicable to diverse biological systems.
Conversational search faces incomplete and informal follow-up questions. Prior works address these by contextualizing user utterances with cues derived from the previous turns of the conversation. This approach works well when the conversation centers on prominent entities, for which knowledge bases (KBs) or language models (LMs) can provide rich background. This work addresses the unexplored direction where user questions are about tail entities, not featured in KBs and sparsely covered by LMs. We devise a new method, called CONSENT, for selectively contextualizing a user utterance with turns, KB-linkable entities, and mentions of tail and out-of-KB (OKB) entities. CONSENT derives relatedness weights from Sentence-BERT similarities and employs an integer linear program (ILP) for judiciously selecting the best context cues for a given set of candidate answers. This method couples the contextualization and answer-ranking stages, and jointly infers the best choices for both.
Next basket recommendation (NBR) is a special type of sequential recommendation that is increasingly receiving attention. So far, most NBR studies have focused on optimizing the accuracy of the recommendation, whereas optimizing for beyond-accuracy metrics, e.g., item fairness and diversity remains largely unexplored. Recent studies into NBR have found a substantial performance difference between recommending repeat items and explore items. Repeat items contribute most of the users' perceived accuracy compared with explore items. Informed by these findings, we identify a potential "short-cut" to optimize for beyond-accuracy metrics while maintaining high accuracy. To leverage and verify the existence of such short-cuts, we propose a plug-and-play two-step repetition-exploration (TREx) framework that treats repeat items and explores items separately, where we design a simple yet highly effective repetition module to ensure high accuracy, while two exploration modules target optimizing only beyond-accuracy metrics. Experiments are performed on two widely-used datasets w.r.t. a range of beyond-accuracy metrics, viz. five fairness metrics and three diversity metrics. Our experimental results show that: (i) we can achieve state-of-the-art performance w.r.t. accuracy via the designed repetition module in TREx; and (ii) the simple TREx framework achieves "better" beyond-accuracy performance than existing sophisticated methods. Prima facie, this appears to be good news: we can achieve high accuracy and improved beyond-accuracy metrics at the same time. However, we argue that the real-world value of our algorithmic solution, TREx, is likely to be limited and reflect on the reasonableness of the evaluation setup. We end up challenging existing evaluation paradigms, particularly in the context of beyond-accuracy metrics, and provide insights for researchers to navigate potential pitfalls and determine reasonable metrics to consider when optimizing for accuracy and beyond-accuracy metrics.
Click logs collect user interaction with information retrieval systems (e.g., search engines). Clicks therefore become implicit feedback for such systems, and are further used to train click models , which in turn improve the quality of search and recommendations results. Click models based on expectation maximization (EM) are known to be effective and robust against various biases. Training EM-based models is challenging due to the size of click logs, and can take many hours when using sequential tools like PyClick. Alternatives, such as ParClick, employ parallelism and show significant speed-up. However, ParClick only works on single-node multi-core systems. To further scale up and out, in this work we introduce MassiveClicks, the first massively parallel, distributed, multi-GPU framework for EM-based click-models training. MassiveClicks relies on efficient GPU kernels, balanced data-partitioning policies, and distributed computing to improve the performance of EM-based model training, outperforming ParClick by orders of magnitude when using GPUs and/or multiple nodes. Additionally, the framework supports heterogeneous GPU architectures, variable numbers of GPUs per node, allows for multi-node multi-core CPU-based training when no GPUs are available.
Generative information retrieval (Gen-IR) is a fast-growing interdisciplinary research area that investigates how to leverage advances in generative Artificial Intelligence (AI) to improve information retrieval systems. Gen-IR has attracted interest from the information retrieval, natural language processing, and machine learning communities, among others. Since the dawn of Gen-IR last year, there has been an explosion of Gen-IR systems that have launched and are now widely used. Interest in this area across academia and industry is only expected to continue to grow as new research challenges and application opportunities arise. The goal of this proposed workshop, The Second Workshop on Generative Information Retrieval (Gen-IR @ SIGIR 2024) is to provide an interactive venue for exploring a broad range of foundational and applied Gen-IR research. The workshop will focus on tasks such as generative document retrieval, grounded answer generation, generative recommendation, and generative knowledge graphs, all through the lens of model training, model behavior, and broader issues. The workshop will be highly interactive, favoring panel discussions, poster sessions, and roundtable discussions over one-sided keynotes and paper talks.
Quantifying the kinetics with which memory T cell populations are generated and maintained is essential for identifying the determinants of the duration of immunity. The quality and persistence of circulating CD4+ effector memory (TEM) and central memory (TCM) T cells in mice appear to shift with age, but it is unclear whether these changes are driven by the aging host environment, by cell age effects, or both. Here we address these issues by combining DNA labelling methods, established fate-mapping systems, a novel reporter mouse strain, and mathematical models. Together, these allow us to quantify the dynamics of both young and established circulating memory CD4+ T cell subsets, within both young and old mice. We show that that these cells and their descendents become more persistent the longer they reside within the TCM and TEM pools. This behaviour may limit memory CD4 T cell diversity by skewing TCR repertoires towards clones generated early in life, but may also compensate for functional defects in new memory cells generated in old age.
Sustained Notch2 signals induce trans-differentiation of Follicular B (FoB) cells into Marginal Zone B (MZB) cells in mice, but the physiology underlying this differentiation pathway is still elusive. Here, we demonstrate that most B cells receive a basal Notch signal, which is intensified in pre-MZB and MZB cells. Ablation or constitutive activation of Notch2 upon T-cell-dependent immunization reveals an interplay between antigen-induced activation and Notch2 signaling, in which FoB cells that turn off Notch2 signaling enter germinal centers (GC), while high Notch2 signaling leads to generation of MZB cells or to initiation of plasmablast differentiation. Notch2 signaling is dispensable for GC dynamics but appears to be re-induced in some centrocytes to govern expansion of IgG1 + GCB cells. Mathematical modelling suggests that antigen-activated FoB cells make a Notch2 dependent binary fate-decision to differentiate into either GCB or MZB cells. This bifurcation might serve as a mechanism to archive antigen-specific clones into functionally and spatially diverse B cell states to generate robust antibody and memory responses.
Foxp3 + Regulatory T cells (Treg) are a subset of CD4 + T cells that play critical functions in maintaining tolerance to self antigens and suppressing autoimmunity, regulating immune responses to pathogens and have a role in the pathophysiology of anti-tumoural immunity. Treg ontogeny is complex since they are generated following recognition of self antigens in the thymus during normal T cell development (thymic Treg), but are also induced from mature conventional T cells when activated by foreign antigen with appropriate additional cues (inducible Treg). How these distinct ontogenic pathways contribute to the maintenance and function of the mature Treg compartment in health and disease remains unclear. Here, we use a combination of fate mapping approaches in mice to map the ontogeny of Treg subsets throughout life and estimate rates of production, loss and self-renewal. We find that naive and effector/memory (EM) Treg subsets exhibit distinct dynamics but are both continuously replenished by de novo generation throughout life. Using an inducible Foxp3-dependent Cre fate reporter system, we show that naive Treg and not conventional T cells, are the predominant precursors of EM Treg in adults. Tonic development of new EM Treg is not influenced by foreign antigens from commensals, rather suggesting a role for self recognition. To investigate the ontogeny of Treg development in malignant disease, we used the same fate reporter systems to characterise the Treg infiltrate of three different model tumours. In all three cases, we found that Treg derived from pre-existing, EM Treg. Together, these results reveal a predominantly linear pathway of Treg development from thymic origin to EM Treg associated with pathophysiology of malignant disease, that is driven by self antigen recognition throughout.
Most conversational passage retrieval systems try to resolve conversational dependencies by using an intermediate query resolution step. To do so, they synthesize conversational data or assume the availability of large-scale question rewriting datasets. To relax those conditions, we propose a zero-shot unified resolution–retrieval approach, that (i) contextualizes and (ii) expands query embeddings using the conversation history and without fine-tuning on conversational data. Contextualization biases the last user question embeddings towards the conversation. Query expansion is used in two ways: (i) abstractive expansion generates embeddings based on the current question and previous history, whereas (ii) extractive expansion tries to identify history term embeddings based on attention weights from the retriever. Our experiments demonstrate the effectiveness of both contextualization and unified expansion in improving conversational retrieval. Contextualization does so mostly by resolving anaphoras to the conversation and bringing their embeddings closer to the important resolution terms that were omitted. By adding embeddings to the query, expansion targets phenomena of ellipsis more explicitly, with our analysis verifying its effectiveness on identifying and adding important resolutions to the query. By combining contextualization and expansion, we find that our zero-shot unified resolution–retrieval methods are competitive and can even outperform supervised methods.
TableQA is the task of answering questions over tables of structured information, returning individual cells or tables as output. TableQA research has focused primarily on high-resource languages, leaving medium- and low-resource languages with little progress due to scarcity of annotated data and neural models. We address this gap by introducing a fully automatic large-scale tableQA data generation process for low-resource languages with limited budget. We incorporate our data generation method on two Indic languages, Bengali and Hindi, which have no tableQA datasets or models. TableQA models trained on our large-scale datasets outperform state-of-the-art LLMs. We further study the trained models on different aspects, including mathematical reasoning capabilities and zero-shot cross-lingual transfer. Our work is the first on low-resource tableQA focusing on scalable data generation and evaluation procedures. Our proposed data generation method can be applied to any low-resource language with a web presence. We release datasets, models, and code (https://github.com/kolk/Low-Resource-TableQA-Indic-languages).
Ensembl (https://www.ensembl.org) is a freely available genomic resource that has produced high-quality annotations, tools, and services for vertebrates and model organisms for more than two decades. In recent years, there has been a dramatic shift in the genomic landscape, with a large increase in the number and phylogenetic breadth of high-quality reference genomes, alongside major advances in the pan-genome representations of higher species. In order to support these efforts and accelerate downstream research, Ensembl continues to focus on scaling for the rapid annotation of new genome assemblies, developing new methods for comparative analysis, and expanding the depth and quality of our genome annotations. This year we have continued our expansion to support global biodiversity research, doubling the number of annotated genomes we support on our Rapid Release site to over 1700, driven by our close collaboration with biodiversity projects such as Darwin Tree of Life. We have also strengthened support for key agricultural species, including the first regulatory builds for farmed animals, and have updated key tools and resources that support the global scientific community, notably the Ensembl Variant Effect Predictor. Ensembl data, software, and tools are freely available.