The T-cell receptor (TCR) repertoire records an individual's immunological history, but most unique CDR3 sequences in any one person are private and uninformative about anyone else. A small subset, however, recurs predictably across unrelated donors. These public TCRβ sequences are independently generated by multiple mechanisms: enrichment by thymic positive selection on a largely shared self-peptide-MHC ligandome, further amplified by selection on common foreign antigens and convergent recombination, together, they are what makes a personal repertoire computationally legible: they provide the shared coordinate system on which otherwise incommensurable repertoires can be aligned and compared. This review takes public TCR sequences as its protagonist. We trace the biology of TCR publicity through new measurements on a 1.5-billion-sequence meta-repertoire (7943 samples, 41 studies) that quantify five interrelated properties: the universe is finite, with Chao2-bounded ceilings (i.e., a lower estimate) of ≈1.97 × 109 amino acid and ≈8.25 × 109 nucleotide CDR3β sequences of which ≈22.5% and ≈12% have already been observed; the rarefaction curve already bends within [0, 7943] by a factor of ×3.3 (amino acid) and ×1.6 (nucleotide) relative to a linear-at-initial-rate extrapolation; publicness is a cohort-scale statistic, with the heavy tail of the publicness distribution growing predictably from N = 200 to N = 7943; super-public sequences are ≈0.014% of unique amino acid CDR3s but carry 8.45% of the total observation mass; and the recapture rate of a CDR3 already observed in another donor reaches 73.3% (amino acid) and 36.2% (nucleotide) at N ≈ 7900. We then trace the family of computational methods that exploit public sequences as features: frequency vectors, sequence-similarity networks, self-supervised transformer embeddings (CVC, Beaker, TCR-BERT, SCEPTR), and our own anchor-based graph neural networks (GraffiTee). We summarize their performance across cancer detection, autoimmunity, infectious disease, immunotherapy monitoring, and immunological aging, alongside the structural confounders (HLA, age, sex, sequencing platform) that bound generalizability. We connect repertoire-level classification to TCR-pMHC binding prediction and antigen identity, and ask what stands between current capabilities and a universal TCR-based diagnostic.
The distinctive characteristics of an individual’s T cell receptor repertoire are crucial in recognizing and responding to a diverse array of antigens, contributing to immune specificity and adaptability. The repertoire, famously vast due to a series of cellular mechanisms, can be quantified using repertoire sequencing. In this study, we sampled the repertoire of 85 women: ovarian cancer patients (OC) and healthy donors (HD), generating a dataset of T cell clones and their abundance. For the alpha chain we obtained 6.4·106 reads, with an average of 75936 clones per sample, and an average of 30607 clonotypes per sample. For the beta chain we obtained 13.6·106 reads, with an average of 160400 clones per sample, and an average of 70071 clonotypes per sample. The changes in dynamics of the repertoire can be observed in response to disease, with specific clones undergoing clonal expansion and contraction. The data provided here offers a unique view of immune system behavior in health and disease and can be used to stratify OC and HD.
The T-cell receptor (TCR) repertoire, unique to each individual, forms a critical component of the adaptive immune system and serves as a defense mechanism against diverse pathogens. This repertoire encompasses a vast diversity of sequences, which can be profiled through high-throughput sequencing technologies. In this study, we present and provide a dataset of TCR sequencing data derived from 215 blood samples collected from newly diagnosed colorectal cancer (CRC) patients prior to any surgical or systemic treatment. Patients were classified according to the TNM staging system, enabling the inclusion of stage and risk labels specific to CRC. Dynamic changes in the TCR repertoire, such as the expansion or contraction of specific clonotypes, are known to reflect immune responses to disease and therapeutic interventions. This dataset provides a comprehensive snapshot of the pre-treatment TCR repertoire in CRC, offering valuable insights into immune system behavior in the context of cancer. The data can serve as a resource for further research into CRC immunology, biomarker discovery, and risk assessment studies.
The immune system’s defense abilities rely on the diversity of T and B lymphocytes. T Cell Receptors (TCRs) are generated through V(D)J recombination, where distinct genetic elements combine and undergo modifications, creating extensive variability. In breast cancer, the most frequently diagnosed cancer in women, early detection sometimes helps with highly effective and potentially curative treatment. The TCR repertoire may provide information about tumor status. To test this, we investigated the peripheral blood TCR repertoire and its association with tumor status. We collected blood samples from 98 women, including patients and healthy donors. Following TCR profiling, machine learning of these data was able to show an association between TCR profiles and breast cancer presence or absence with high accuracy (average AUC of 0.96). Our findings imply the immune system retains tumor-relevant, TCR-related, signals detectable in blood. This information could potentially benefit future derivatives from this knowledge, either in the field of detection or treatment.
The purpose of this study is to evaluate the feasibility of using RNA sequencing data as substrate for the computational extraction of T cell receptor sequences. Data from hundreds of thousands of samples is available as RNA sequencing. However, the use of these data for repertoires has not been contrasted against a gold standard. We conducted a benchmarking analysis, comparing T cell receptor data extracted from RNA sequencing to those obtained from T cell receptor sequencing (as gold standard) of the same tissue samples. The focus was on the extraction of Complementarity-Determining Region 3 (CDR3) sequences. To evaluate the influence of sequencing read lengths, samples were analyzed using both 75 base pair single-end and 150 base pair paired-end sequencing methods. In addition we calculated T cell abundance in these samples to test for any correlation between reads and abundance. The findings reveal a significant, perhaps too great, discrepancy between the ability to extract Complementarity-Determining Region 3 sequences from RNA sequencing data and the results obtained from TCR sequencing. The lack of significant improvement with longer read lengths, combined with the absence of correlation to T cell abundance, emphasize the necessity of using T cell receptor sequencing methodologies.
This study investigates the T-cell receptor (TCR) repertoires in colorectal cancer (CRC) patients by analyzing three distinct datasets: one bulk sequencing dataset of 205 patients with various tumor stages, all newly diagnosed at Sheba Medical Center between 2017 and 2022, with minimal recruitment in 2014 and 2016, and two (public) single-cell sequencing datasets of 10 and 12 patients. Despite the significant variability in the TCR repertoire and the low likelihood of sequence overlap, our analysis reveals an interesting set of TCR sequences across these data. Notably, we observe elevated presence of mucosal-associated invariant T (MAIT) cells in both metastatic and non-metastatic patients. Furthermore, we identify nine identical TCR alpha and TCR beta pairs that appear in both single-cell datasets, with 13 out of 18 sequences from these sequences also appearing in the bulk data. Clinical risk analysis over the bulk dataset, using a subset of these unique sequences, demonstrates a correlation between TCR repertoire disease stage and risk. These findings enhance our understanding of the TCR landscape in CRC and underscore the potential of TCR sequences as biomarkers for disease outcome.
Introduction: We present YADA, a cellular content deconvolution algorithm for estimating cell type proportions in heterogeneous cell mixtures based on gene expression data. YADA utilizes curated gene signatures of cell type-specific marker genes, either obtained intrinsically from pure cell type expression matrices or provided by the user. Methods: YADA implements an accessible and extensible deconvolution framework uniquely capable of handling marker genes alone as inputs. Adoption barriers are lowered significantly by relying solely on literature-supported cell type-specific signatures rather than full transcriptomic profiles from purified isolates. However, flexible inputs do not necessitate sacrificing rigor - predictions match metrics of current methodologies through an integrated optimization scheme balancing multiple inference algorithms. Efficiency optimizations via compiled runtimes enable rapid execution. Packaging as an importable Python toolkit promotes community enhancement while retaining codebase extensibility. Results: Validation studies demonstrate that YADA matches or exceeds the performance of current deconvolution methods on benchmark datasets. To demonstrate the utility and enable immediate usage, we provide an online Jupyter Notebook implementation coupled with tutorials. Conclusion: YADA provides an accurate, efficient, and extensible Python-based toolkit for cellular deconvolution analysis of heterogeneous gene expression data.
Pathway analysis is a powerful approach for elucidating insights from gene expression data and associating such changes with cellular phenotypes. The overarching objective of pathway research is to identify critical molecular drivers within a cellular context and uncover novel signaling networks from groups of relevant biomolecules. In this work, we present PathSingle, a Python-based pathway analysis tool tailored for single-cell data analysis. PathSingle employs a unique graph-based algorithm to enable the classification of diverse cellular states, such as T cell subtypes. Designed to be open-source, extensible, and computationally efficient, PathSingle is available at https://github.com/zurkin1/PathSingle under the MIT license. This tool provides researchers with a versatile framework for uncovering biologically meaningful insights from high-dimensional single-cell transcriptomics data, facilitating a deeper understanding of cellular regulation and function.
The T cell receptor (TCR) repertoire is an extraordinarily diverse collection of TCRs essential for maintaining the body’s homeostasis and response to threats. In this study, we compiled an extensive dataset of more than 4200 bulk TCR repertoire samples, encompassing 221,176,713 sequences, alongside 6,159,652 single-cell TCR sequences from over 400 samples. From this dataset, we then selected a representative subset of 5 million bulk sequences and 4.2 million single-cell sequences to train two specialized Transformer-based language models for bulk (CVC) and single-cell (scCVC) TCR repertoires, respectively. We show that these models successfully capture TCR core qualities, such as sharing, gene composition, and single-cell properties. These qualities are emergent in the encoded TCR latent space and enable classification into TCR-based qualities such as public sequences. These models demonstrate the potential of Transformer-based language models in TCR downstream applications.
Abstract Head and neck squamous cell carcinoma (HNSCC) is a global health problem. Successful immunotherapy in HNSCC is currently based on the ability to block the interaction of PD1 and PDL1 and abolish the inhibition of CD8+ T cells, thus enhancing the antitumor activity. T cells recognize peptide antigens in the context of MHC molecules through the clonally distributed T Cell receptor (TCR), which offers the means to detect and track specific T cells. We characterized TCR signatures in unique HNSCC paired samples (including L = Lymphocyte, S = Saliva, T = Tumor) to profile TCR repertoire diversity and clonality in different compartments. A total of 48 HNSCC samples corresponding to 14 patients (2 assayed 2X) were obtained from Johns Hopkins University. DNA samples were submitted for TCR sequencing (TCR-seq) via enrichment of the human TCRB locus using Human-TCRB-PD4bx. CDR3 sequences were downloaded from Adaptive ImmuneAnalyzer. For each sequence, V, D and J regions were extracted and translation to amino acid done on the CDR3 region. 8 paired (T, L, S) samples were submitted for shallow sequencing survey depth and 8 for deep assay depth. In order to better understand the TCR repertoire, we assessed the number of unique immune receptor clonotypes, their relative abundance, repertoire overlap, clonotype tracking and sequence length distribution. As expected, the deep assay revealed more clones than the survey. Using the TCR-seq deep assay, 9 out of 10 S samples harbored at least 10% of the number of clonotypes found in their paired T. As expected, L harbored a greater number and more diverse clones than T and S samples and showed partial overlap with T and S samples. Comparing the overlap between T, L and S based on a shared CDR3.aa repertoire: by shallow survey, 3/5 S samples had at least 1 overlap with the primary T. Using the deep assay, 9 out of 10 patients (90%) showed overlap between the S and paired T. We found top public clonotypes occupying the larger part in HPV negative samples, and surprisingly no known CDR3.aa sequences annotated to HPV in HPV positive samples but annotated to other common viruses (CMV, EBV). Our study demonstrates a robust TCR repertoire in S that corresponds at least partially to the clones observed in T and circulating L. The presence of HPV did not yet yield CDR3.aa specific sequences. Future single-cell transcriptomes may identify signatures of CD8+ and CD4+ neoantigen-reactive tumor-infiltrating L in the paired S samples in addition to the viral or tumor-associated antigens present in bulk assays. Our data need to be further validated in larger well characterized cohorts that include S from cancer-free subjects. These early results support the use of saliva as a potential surrogate for diagnostic immune-profiling of tumors, for early detection and follow-up. We further envision creating a personalized TCR repertoire for individual patients from an initial tumor sample biopsy to monitor their tumor immune dynamics using saliva. Citation Format: Dieila Giomo De Lima, Mariana Brait, Laura Palmieri, Fernando T. Zamuner, Esther Broner, Kellie N. Smith, Timothy Westlake, Or Malca, Sol Efroni, Ido Sloma, David Sidransky. Saliva TCR repertoire as a tool for head and neck cancer immunophenotype monitoring [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 2121.
Background: In the global effort to discover biomarkers for cancer prognosis, prediction tools have become essential resources. TCR (T cell receptor) repertoires contain important features that differentiate healthy controls from cancer patients or differentiate outcomes for patients being treated with different drugs. Considering, tools that can easily and quickly generate and identify important features out of TCR repertoire data and build accurate classifiers to predict future outcomes are essential. Results: This paper introduces GENTLE (GENerator of T cell receptor repertoire features for machine LEarning): an open-source, user-friendly web-application tool that allows TCR repertoire researchers to discover important features; to create classifier models and evaluate them with metrics; and to quickly generate visualizations for data interpretations. We performed a case study with repertoires of TRegs (regulatory T cells) and TConvs (conventional T cells) from healthy controls versus patients with breast cancer. We showed that diversity features were able to distinguish between the groups. Moreover, the classifiers built with these features could correctly classify samples ('Healthy' or 'Breast Cancer')from the TRegs repertoire when trained with the TConvs repertoire, and from the TConvs repertoire when trained with the TRegs repertoire. Conclusion: The paper walks through installing and using GENTLE and presents a case study and results to demonstrate the application's utility. GENTLE is geared towards any researcher working with TCR repertoire data and aims to discover predictive features from these data and build accurate classifiers. GENTLE is available on https://github.com/dhiego22/gentle and https://share.streamlit.io/dhiego22/gentle/main/gentle.py.
PDF file - 125K, Correlated expression of E2F1 and RBM38 is associated with increased survival in breast cancer patients
Immunotherapy is now an essential tool for cancer treatment, and the unique features of an individual's T cell receptor repertoire are known to play a key role in its effectiveness. The repertoire, famously vast due to a cascade of cellular mechanisms, can be quantified using repertoire sequencing. In this study, we sampled the repertoire over several time points following treatment with anti-CTLA-4, in a syngeniec mouse model for colorectal cancer, generating a longitudinal dataset of T cell clones and their abundance. The dynamics of the repertoire can be observed in response to treatment over a period of four weeks, as clonal expansion of specific clones ascends and descends. The data made available here can be used to determine treatment and predict its effect, while also providing a unique look at the behavior of the immune system over time.
Upon the development of a therapeutic, a successful response to a global pandemic relies on efficient worldwide distribution, a process constrained by our global shipping network. Most existing strategies seek to maximize the outflow of the therapeutics, hence optimizing for rapid dissemination. Here we find that this intuitive approach is, in fact, counterproductive. The reason is that by focusing strictly on the quantity of disseminated therapeutics, these strategies disregard the way in which this quantity distributes across destinations. Most crucially—they overlook the interplay of the therapeutic spreading patterns with those of the pathogens. This results in a discrepancy between supply and demand, that prohibits efficient mitigation even under optimal conditions of superfluous flow. To solve this, we design a dissemination strategy that naturally follows the predicted spreading patterns of the pathogens, optimizing not just for supply volume, but also for its congruency with the anticipated demand. Specifically, we show that epidemics spread relatively uniformly across all destinations, prompting us to introduce an equality constraint into our dissemination that prioritizes supply homogeneity. This strategy may, at times, slow down the supply rate in certain locations, however, thanks to its egalitarian nature, which mimics the flow of the pathogens, it provides a dramatic leap in overall mitigation efficiency, potentially saving more lives with orders of magnitude less resources.
Biochemical pathways analysis is an effective tool for understanding changes in gene expression data and associating such changes with cellular phenotypes. Pathway research aims to identify associated proteins within a cell using pathways and at building new pathways from a group of molecules of interest. Using pathway-based methods we gain insight into different functions of relevant molecules and find direct and indirect relations between them. We present PathWeigh, a Python-based tool for pathway analysis and graph presentation. The tool is open-sourced, extendable and runtime efficient. PathWeigh is available at https://github.org/zurkin1/Pathweigh and is released under MIT license. A sample Python notebook is provided with examples of running the tool.
T cells are at the core of human health. Their unique ability to produce an overwhelming repertoire of receptors which they use to interact with threats and to assist with healthy functions has made them an interesting subject for Data Science. One such function, involved with multiple conditions including cancer, response to viral threats, autoimmune disease and more, is the ability to produce what are termed Public clones. Those Public Clones are T cells that are shared between individuals. Some of those clones are even shared between a high percentage of all observed samples. Yet, the reason for this sharing, as well as the DNA sequence that might characterize them is still unknown. Here, using a BERT-based language model, we show that a latent space built by self supervised learning provides distinct areas for Public and Private sequences. We continue to show that these embeddings could be successfully used for binary classification to tell apart Public and Private sequences.
The partial success of tumor immunotherapy induced by checkpoint blockade, which is not antigen-specific, suggests that the immune system of some patients contain antigen receptors able to specifically identify tumor cells. Here we focused on T-cell receptor (TCR) repertoires associated with spontaneous breast cancer. We studied the alpha and beta chain CDR3 domains of TCR repertoires of CD4 T cells using deep sequencing of cell populations in mice and applied the results to published TCR sequence data obtained from human patients. We screened peripheral blood T cells obtained monthly from individual mice spontaneously developing breast tumors by 5 months. We then looked at identical TCR sequences in published human studies; we used TCGA data from tumors and healthy tissues of 1,256 breast cancer resections and from 4 focused studies including sequences from tumors, lymph nodes, blood and healthy tissues, and from single cell dataset of 3 breast cancer subjects. We now report that mice spontaneously developing breast cancer manifest shared, Public CDR3 regions in both their alpha and beta and that a significant number of women with early breast cancer manifest identical CDR3 sequences. These findings suggest that the development of breast cancer is associated, across species, with biomarker, exclusive TCR repertoires.
Restoration of T cell repertoire diversity after allogeneic bone marrow transplantation (allo-BMT) is crucial for immune recovery. T cell diversity is produced by rearrangements of germline gene segments (V (D) and J) of the T cell receptor (TCR) α and β chains, and selection induced by binding of TCRs to MHC-peptide complexes. Multiple measures were proposed for this diversity. We here focus on the V-gene usage and the CDR3 sequences of the beta chain. We compared multiple T cell repertoires to follow T cell repertoire changes post-allo-BMT in HLA-matched related donor and recipient pairs. Our analyses of the differences between donor and recipient complementarity determining region 3 (CDR3) beta composition and V-gene profile show that the CDR3 sequence composition does not change during restoration, implying its dependence on the HLA typing. In contrast, V-gene usage followed a time-dependent pattern, initially following the donor profile and then shifting back to the recipients' profile. The final long-term repertoire was more similar to that of the recipient's original one than the donor’s; some recipients converged within months, while others took multiple years. Based on the results of our analyses, we propose that donor-recipient V-gene distribution differences may serve as clinical biomarkers for monitoring immune recovery.
Biology of the response to anti-CTLA-4 involves the dynamics of specific T cell clones. Reasons for clinical success and failure of this treatment are still largely unknown. Here, we quantified the dynamics of the T cell receptor (TCR) repertoire, throughout 4 weeks involving treatment with anti-CTLA-4, in a syngeneic mouse model for colorectal cancer. These dynamics show an initial increase in clonality in tandem with a decrease in diversity, effects which gradually subside. Furthermore, response to treatment is tightly connected to the shared and public parts of the T cell repertoire. We were able to recognize time-dependent behaviors of specific TCR sequences and cell types and to show the response is dominated by specific motifs. We see that a single, specific time point might be useful to inform a physician of the true response to treatmentThe research further highlights the importance of temporal analyses of the immune response.