Affinity maturation is the Darwinian process by which antibodies improve antigen binding through somatic hypermutation and selection. The adaptive landscape, which defines the set of antibody-specific mutations that improve functional characteristics like antigen binding, has been explored in only a handful of antibodies. Identifying the sites of adaptive mutations in a given antibody sequence, and how these sites vary across the antibody repertoire, can inform the design of therapeutic antibodies. We develop a parameter-free population genetic framework that leverages the statistics of convergent affinity maturation in B cell lineages sharing similar naive sequences, called public clonotypes, to identify beneficial mutations. Applying this framework to more than 10,000 public clonotypes represented by multiple lineages across 20 healthy individuals, we identify widespread signatures of clonotype-dependent selection of individual mutations. We estimate the prevalence and typical fitness effects of mutations across the V gene at the single-site level, uncovering a general tradeoff between prevalence and fitness effect. These inferred landscapes broadly reproduce the statistics of convergent mutation in antibodies specific to SARS-CoV-2 and influenza. Finally, we use our framework to benchmark predictions from existing antibody language models, and show that while these models are dominated by non-selective signatures, a simple renormalization procedure can expose signatures of clonotype-dependent positive selection consistent with our predictions.
Cellular diversification in processes from development to cancer progression and affinity maturation is often linked to the appearance of new mutations, generating genetic heterogeneity. Describing the underlying coupled genetic and growth processes that result in the observed diversity in cell populations is informative about the timing, drivers and outcomes of cell fates. Current approaches based on phylogenetic methods do not cover the entire range of evolutionary rates, often making artificial assumptions about the timing of events. We introduce CBA, a probabilistic method that infers the division, degradation and mutation rates from the observed genetic diversity in a population of cells. It uses a summarized backbone tree, intermediary between the true cell tree and the allelic tree representing the ancestral relationships between types, called a monogram, which allows for efficient sampling of possible phylogenies consistent with the observed mutational signatures. We demonstrate the accuracy of our method on simulated data and compare its performance to standard phylogenetic approaches.
The specific region of an antibody responsible for binding to an antigen, known as the paratope, is essential for immune recognition. Accurate identification of this small yet critical region can accelerate the development of therapeutic antibodies. Determining paratope locations typically relies on modeling the antibody structure, which is computationally intensive and difficult to scale across large antibody repertoires. We introduce Paraplume, a sequence-based paratope prediction method that leverages embeddings from protein language models (PLMs), without the need for structural input and achieves superior performance across multiple benchmarks compared to current methods. In addition, reweighting PLM embeddings using Paraplume predictions yields more informative sequence representations, improving downstream tasks such as binder classification and epitope binning. Applied to large antibody repertoires, Paraplume reveals that antigen-specific somatic hypermutations are associated with larger paratopes, suggesting a potential mechanism for affinity enhancement. Our findings position PLM-based paratope prediction as a powerful, scalable alternative to structure-dependent approaches, opening new avenues for understanding antibody evolution.
Our adaptive immune system relies on the persistence over long times of a diverse set of antigen-experienced B cells to encode our memories of past infections and to protect us against future ones. While longitudinal repertoire sequencing promises to track the long-term dynamics of many B cell clones simultaneously, sampling and experimental noise make it hard to draw reliable quantitative conclusions. Leveraging statistical inference, we infer the dynamics of memory B cell clonal dynamics and conversion to plasmablasts, which includes clone creation, degradation, abundance fluctuations, and differentiation. We find that memory B cell clones degrade slowly, with a halflife of 10 y. Based on the inferred parameters, we predict that it takes about 50 y to renew 50% of the repertoire, with most observed clones surviving for a lifetime. We infer that, on average, 1 out of 100 memory B cells differentiates into a plasmablast each year, more than expected from purely antigen-stimulated differentiation, and that plasmablast clones degrade with a half-life of about one year in the absence of memory imports. Our method is general and could be applied to other longitudinal repertoire sequencing B cell subsets.
Cells integrate signals and make decisions about their future state in short amounts of time. A lot of theoretical effort has gone into asking how to best design gene regulatory circuits that fulfill a given function, yet little is known about the constraints that performing that function in a small amount of time imposes on circuit architectures. Using an optimization framework, we explore the properties of a class of promoter architectures that distinguish small differences in transcription factor concentrations under time constraints. We show that the full temporal trajectory of gene activity allows for faster decisions than its integrated activity represented by the total number of transcribed mRNA. The topology of promoter architectures that allow for rapidly distinguishing low transcription factor concentrations result in a low, shallow, and non cooperative response, while at high concentrations, the response is high and cooperative. In the presence of non-cognate ligands, networks with fast and accurate decision times need not be optimally selective, especially if discrimination is difficult. While optimal networks are generically out of equilibrium, the energy associated with that irreversibility is only modest, and negligible at small concentrations. Instead, our results highlight the crucial role of rate-limiting steps imposed by biophysical constraints.
Combinatorial pairing of independently recombined T cell receptor (TCR) α- and β-chains is central to diversifying the TCR repertoire. Although sequence motifs in one chain correlate with epitope recognition, the extent to which a single chain dictates specificity remains unclear. Here, we systematically tested TCR chain coupling constraints by enforcing the pairing of individual chains with hundreds of thousands of partners. Although most chains paired stably, the preservation of epitope specificity was rare and highly variable, with the frequency of compatible partners ranging from ~10 to <0.1%. This approach identified >70,000 epitope-specific TCRs across 10 epitopes. Our work illuminates the distinct contributions of TCR chains, highlights the limitations of single-chain data, and provides an experimental and analytical framework for refining TCR-peptide-major histocompatibility complex specificity inference.
Biomolecular condensates form on timescales of seconds in cells upon environmental or compositional changes. Condensate formation is thus argued to act as a mechanism for sensing such changes and quickly initiating downstream processes, such as forming stress granules in response to heat stress and amplifying cyclic GMP-AMP synthase enzymatic activity upon detection of cytosolic DNA. Here, we study a dynamical model of droplet nucleation and growth to demonstrate how phase separation allows cells to discriminate small concentration differences on finite, biologically relevant timescales. We propose optimal sensing protocols, which use the sharp onset of phase separation. We show how, given experimentally measured rates, cells can achieve rapid and robust sensing of concentration differences of 1% on a timescale of minutes, offering an alternative to classical biochemical mechanisms.
The composition of a polyclonal antibody response is hard to measure experimentally but contains vital information about the robustness of immunity. Here, we argue that the statistics of neutralization titers alone can be used to make quantitative predictions about the composition of the response, circumventing challenges arising through sequencing and monoclonal antibody expression. We show that the response against influenza within a cohort can be either driven by a collective phenomenon where many antibodies contribute to neutralization, or dominated by just a few strong binders, leading to a broad distribution of titers across individuals described by a Gumbel distribution from extreme value theory. Comparing titers across cohorts, we find that Gumbel statistics accurately describe individuals prior to an immune challenge. We propose an equilibrium binding model that quantitatively captures titer data and illustrates the structure of the polyclonal response. Our approach extends generically to immune responses to other pathogens.
T cells activate and expand upon interaction with cognate antigen, derived from pathogens or mutated proteins. T cell clones can be identified by their T cell receptor (TCR) which can act as a unique barcode to track their expansion. Longitudinal TCR sequencing can be used to track T cell responses to a large array of stimuli. However, experimental identification of T cell clones of interest is challenging, especially when information about the driving antigen is lacking. Computational identification based on clonal dynamics is an antigen-agnostic alternative. However, it is subject to sequencing noise and biological variability, and relies on the choice of particular time points that are compared to find expanding and contracting clones. We present CloneSearch, a method to identify expanding and contracting T cell clones from longitudinal TCR sequencing which is agnostic to the time of stimulus and can account for the noise these clones are subject to. We show that CloneSeach can recapitulate previously identified responses from published data, and expand the analysis to show identification of previously undetected responses from these same datasets. We make CloneSearch available at https://github.com/mm523/CloneSearch .
Generative models derived from large protein sequence alignments define complex fitness landscapes, but their utility for accurately modeling non-equilibrium evolutionary dynamics remains unclear. In this work, we perform a rigorous comparative analysis of three simulation schemes, designed to mimic evolution in silico by local sampling of the probability distribution defined by a generative model. We compare standard independent Markov Chain Monte Carlo, Monte Carlo on a phylogenetic tree, and a population genetics dynamics, benchmarking their outputs against deep sequencing data from four distinct in vitro evolution experiments. We find that standard Monte Carlo fails to reproduce the correct phylogenetic structure and generates unrealistic, gradual mutational sweeps. Performing Monte Carlo on a tree inferred from data improves phylogenetic fidelity and historical accuracy. The population genetics scheme successfully captures phylogenetic correlations, mutational abundances, and selective sweeps as emergent properties, without the need to infer additional information from data. However, the latter choice come at the price of not sampling the proper generative model distribution at long times. Our findings highlight the crucial role of phylogenetic correlations and finite-population effects in shaping evolutionary trajectories on fitness landscapes. These models therefore provide powerful tools for predicting complex adaptive paths and for reliably extrapolating evolutionary dynamics beyond current experimental limitations.
Administration of HIV-1 neutralizing antibodies can suppress viremia and prevent infection in vivo. However, clinical use is challenged by broad envelope sequence diversity and rapid emergence of viral escape1-9. Here, we performed single B cell profiling of 32 top HIV-1 elite neutralizers to identify broadly neutralizing antibodies (bNAbs) with highest potency and breadth for clinical application. From 831 expressed monoclonal antibodies, we identified 04_A06, a new VH1-2-encoded CD4 binding site bNAb with remarkable breadth and potency against extended multiclade pseudovirus panels (GeoMean IC50 = 0.059 μg/ml, breadth = 98.5%, 332 virus strains). Moreover, 04_A06 was not susceptible to classic viral CD4bs escape variants and maintained full viral suppression in HIV-1-infected humanized mice. Structural analyses revealed that antiviral activity is mediated by an unusually long 11-amino acid heavy chain insertion. This insertion facilitates inter-protomer contacts and interactions with highly conserved residues on the adjacent gp120 protomer. Finally, 04_A06 demonstrated high activity against contemporaneously circulating viruses from the Antibody Mediated Prevention (AMP) trials (GeoMean IC50 = 0.082 μg/ml, breadth = 98.4%, 191 virus strains) and in silico modeling for 04_A06LS predicted HIV-1 prevention efficacy of >93%. Thus, 04_A06 will provide unique opportunities for effective treatment and prevention strategies of HIV-1 infection.
Neural correlations play a critical role in sensory information coding. They are of two kinds: signal correlations, when neurons have overlapping sensitivities, and noise correlations from network effects and shared noise. In experiments from early sensory systems and cortex, many pairs of neurons typically show both types of correlations to be positive and large, especially between nearby neurons with similar stimulus sensitivity. However, theoretical arguments have suggested that stimulus and noise correlations should have opposite signs to improve coding, at odds with experimental observations. We analyze retinal recording in response to a large variety of stimuli, and show that, contrary to common belief, large noise correlations are beneficial for coding, even if aligned with signal correlations. To understand this result, we develop a theory of visual information coding by correlated neurons, which resolves that paradox. We show that noise correlations are always beneficial if they are strong enough, unless neurons are perfectly correlated by the stimulus. Finally, using neuronal recordings and modeling, we show that for high dimensional stimuli noise correlation benefits the encoding of fine-grained details of visual stimuli, at the expense of large-scale features, which are already well encoded.
Biomolecular condensates form on timescales of seconds in cells upon environmental or compositional changes. Condensate formation is thus argued to act as a mechanism for sensing such changes and quickly initiating downstream processes, such as forming stress granules in response to heat stress and amplifying cGAS enzymatic activity upon detection of cytosolic DNA. Here, we show that phase separation allows cells to discriminate small concentration differences on finite, biologically relevant timescales. We propose optimal sensing protocols, which use the sharp onset of phase separation. We show how, given experimentally measured rates, cells can achieve rapid and robust sensing of concentration differences of 1% on a timescale of minutes, offering an alternative to classical biochemical mechanisms.
T cells recognize a wide range of pathogens using surface receptors that interact directly with pep-tides presented on major histocompatibility complexes (MHC) encoded by the HLA loci in humans. Understanding the association between T cell receptors (TCR) and HLA alleles is an important step towards predicting TCR-antigen specificity from sequences. Here we analyze the TCR alpha and beta repertoires of large cohorts of HLA-typed donors to systematically infer such associations, by looking for overrepresentation of TCRs in individuals with a common allele.TCRs, associated with a specific HLA allele, exhibit sequence similarities that suggest prior antigen exposure. Immune repertoire sequencing has produced large numbers of datasets, however the HLA type of the corresponding donors is rarely available. Using our TCR-HLA associations, we trained a computational model to predict the HLA type of individuals from their TCR repertoire alone. We propose an iterative procedure to refine this model by using data from large cohorts of untyped individuals, by recursively typing them using the model itself. The resulting model shows good predictive performance, even for relatively rare HLA alleles.
Administration of HIV-1 neutralizing antibodies can suppress viremia and prevent infection in vivo. However, clinical use is challenged by envelope diversity and rapid viral escape. Here, we performed single B cell profiling of 32 top HIV-1 elite neutralizers to identify broadly neutralizing antibodies with highest antiviral activity. From 831 expressed monoclonal antibodies, we identified 04_A06, a VH1-2-encoded broadly neutralizing antibody to the CD4 binding site with remarkable breadth and potency against multiclade pseudovirus panels (geometric mean half-maximal inhibitory concentration = 0.059 µg ml−1, breadth = 98.5
The competition for resources is a defining feature of microbial communities. In many contexts, from soils to host-associated communities, highly diverse microbes are organized into metabolic groups or guilds with similar resource preferences. The resource preferences of individual taxa that give rise to these guilds are critical for understanding fluxes of resources through the community and the structure of diversity in the system. However, inferring the metabolic capabilities of individual taxa, and their competition with other taxa, within a community is challenging and unresolved. Here we address this gap in knowledge by leveraging dynamic measurements of abundances in communities. We show that simple correlations are often misleading in predicting resource competition. We show that spectral methods such as the cross-power spectral density (CPSD) and coherence that account for time-delayed effects are superior metrics for inferring the structure of resource competition in communities. We first demonstrate this fact on synthetic data generated from consumer-resource models with time-dependent resource availability, where taxa are organized into groups or guilds with similar resource preferences. By applying spectral methods to oceanic plankton time-series data, we demonstrate that these methods detect interaction structures among species with similar genomic sequences. Our results indicate that analyzing temporal data across multiple timescales can reveal the underlying structure of resource competition within communities.
Cells use signalling pathways as windows into the environment to gather information, transduce it into their interior, and use it to drive behaviours. MAPK (ERK) is a highly conserved signalling pathway in eukaryotes, directing multiple fundamental cellular behaviours such as proliferation, migration, and differentiation, making it of few central hubs in the signalling circuitry of cells. Despite this versatility of behaviors, population-level measurements have reported low information content (1 bit) relayed through the ERK pathway, rendering the population barely able to distinguish the presence or absence of stimuli. Here, we contrast the information transmitted by a single cell and a population of cells. Using a combination of optogenetic experiments, data analysis based on information theory framework, and numerical simulations we quantify the amount of information transduced from the receptor to ERK, from responses to singular, brief and sparse input pulses. We show that single cells are indeed able to resolve between graded stimuli, yielding over 2 bit of information, however showing a large population heterogeneity.
The T cell’s ability to recognize antigens relies on the diversity of the T cell receptor (TCR) arising from V(D)J recombination of the two TCR chains, complemented by their random pairing. While it is generally recognized that both TCR chains are crucial for epitope recognition, their relative contribution remains unclear. To estimate the extent to which a single TCR chain determines epitope specificity, we introduced an exogenous TCR chain with known specificity into PBMC samples. This enabled it to pair with diverse endogenous TCR chains, determining the frequency and features of paired repertoires for MAIT, NKT, and conventional HLA class I and II epitopes. For each tested chain, we identified hundreds to thousands of novel, specific receptors showing antigen selection biases. Comparisons with existing data demonstrated that these engineered paired TCRs accurately replicate the natural repertoire. We found suitable partner chain frequencies varied greatly depending on the specific chain tested. The resulting data, including over 70,000 unique epitope-specific paired chain TCRs, broadens our understanding of each chain’s role in epitope recognition and quantifies a previously underestimated amount of single chain sharing across epitope specificities. Utilizing such comprehensive epitope-specific positive and negative data to train TCR-epitope prediction algorithms can enhance the accuracy of TCR specificity predictions. The work was funded by AI136514 and AI150747. Immune Response Regulation: Molecular Mechanisms (IRM)