DNA methyltransferase accessibility protocol for individual templates (MAPit) enables localization of DNA-protein interactions by probing the native chromatin architecture in nuclei or permeabilized cells using exogenously applied DNA methyltransferases. By preserving the integrity of DNA in chromatin, MAPit enables simultaneous mapping of nucleosome positioning, transcription factor binding, and endogenous CpG methylation in a single assay-while maintaining linkage between regulatory modules along chromatin fibers. Although MAPit and other sequencing-based approaches can be applied genome-wide, there is a growing demand for orthogonal approaches that validate target discovery and offer deeper molecular resolution at defined loci. Here, we describe a detailed protocol for flap-enabled next-generation capture (FENGC), a cost-effective method for targeted, multiplexed enrichment of DNA sequences for epigenetic and genetic analysis. FENGC precisely excises and enriches tens to hundreds of user-defined targets, generating sequencing-ready libraries with superior on-target rates compared to existing methods, all at a fraction of the cost. When paired with MAPit, the MAPit-FENGC platform enables single-molecule resolution of CpG methylation patterns in the context of chromatin architecture and transcription factor occupancy.
Genome-wide studies have identified significant allelic associations between genetic variants in or near the IKZF1 gene and multiple autoimmune disorders. IKZF1, encoding the transcription factor IKAROS, produces at least 10 distinct transcripts. To explore the impact of alternative splicing of IKZF1 on the function of mature T cells and the risk of autoimmunity, we generated a panel of human T-cell clones with truncating mutations in IKZF1 exons 4, 6, or both. Differences in gene expression, chromatin accessibility, and protein abundance among clones were assessed by RNA-seq, ATAC-seq, and immunoblotting. Clones with single targeting events clustered separately from double-targeted clones on multiple parameters, but overall, clone responses were highly heterogeneous. Perturbation of IKZF1 splicing resulted in significant differences in expression and chromatin accessibility of other autoimmunity-associated genes and elicited compensatory expression changes in other IKAROS family members. Our results suggest that even modest alterations of IKZF1 splicing can have significant effects on gene expression and function in mature T cells, potentially contributing to autoimmunity in susceptible individuals.
How cell type, sex and disease interact and affect gene expression and splicing is an important, but complicated question. Visualizing and testing specific hypotheses around these complex interactions is an important first step to identifying molecular components underpinning complex disease. Using a meta-analytical framework, we develop an analytical path for identifying testable molecular hypotheses of complex interactions between splicing, sex, disease and cell type. We focus on type 1 diabetes (T1D) but the approach is generalizable to any complex disease with defined candidate loci. Previous studies report T1D-associated splicing in candidate genes, differences in disease effects across immune cell types, sex effects on splicing and cell-type-specific splicing. However, identifying and interpreting complex interactions between sex, splicing and disease are challenging. Here we demonstrate how a gene expression study of T1D, designed to evaluate these interactions can be analyzed in a straightforward manner. We find that sex-dependent T1D-associated splicing is markedly more prevalent in CD4⁺ T cells than in CD8⁺ T cells, affecting 72% of T1D candidate genes in CD4⁺ cells compared to 30% in CD8⁺ cells. We pinpoint exons whose rate of inclusion is affected by the interaction of sex and disease. We use long-read RNAseq to identify novel intron retention events and splice sites which are quantified with short-reads leading to a richer description of the regulatory impact of T1D on alternative splicing. We identify a set of candidate isoforms for follow-up molecular studies in BACH2 , a transcription factor known to be relevant in disease prevalence.
PURPOSE:Women treated with radiation therapy (RT) for breast cancer have an increased risk of developing radiation-associated contralateral breast cancer (CBC). Predicting CBC events is challenging because of the complex interplay of genomic, treatment, personal, and clinical factors. This study investigated computational methods that integrate genome-wide single-nucleotide polymorphisms and nongenomic data to develop a risk stratification model for developing CBC in women treated with RT for their first primary breast cancer. METHODS AND MATERIALS:This study used a subset of the population-based Women's Environmental Cancer and Radiation Epidemiology study that included 633 CBC cases and 1253 individually matched unilateral breast cancer controls who were treated with RT and had single-nucleotide polymorphism data available from a genome-wide association study. The study population was split into training, validation, and test sets for rigorous modeling and validation. Three data integration methods were compared in terms of their ability to stratify CBC risk: (1) naive integration; (2) sequential integration; and (3) sequential iterative integration. A biological analysis of the final model was performed using gene set enrichment analysis and protein-protein interaction analysis with gene annotation information informed by the model. RESULTS:The best-performing integration method was the sequential iterative integration equipped with the mixed-effect random forest algorithm. This approach achieved an area under the curve of 0.64 to stratify CBC risk in the test set, representing moderate predictive power. Calibration analysis showed good agreement between the lowest and highest risk bins stratified using sorted predicted values in the test set, resulting in an odds ratio of 3.27 for both predicted and observed CBC occurrence. Gene set enrichment analysis and protein-protein interaction analysis revealed that genes with high importance scores were associated with pathways relevant to lipid and fatty acid metabolism as well as breast cancer sensitivity to tamoxifen. CONCLUSIONS:The mixed-effect random forest approach demonstrated the potential for integrating high-dimensional genomic and low-dimensional nongenomic data to stratify CBC risk.
Type 1 diabetes (T1D) results from the autoimmune destruction of the insulin-producing β cells. Genetic factors account for approximately 50% of the risk for T1D but, by the late 1990s, the genetic basis was limited. The Type 1 Diabetes Genetics Consortium (T1DGC) was formed in 2002 to accelerate discovery of genes contributing to T1D risk through a grant from the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) to assemble existing data and samples from affected sib-pair families and to establish new collections. In recognition of the 75th anniversary of the NIDDK, this manuscript highlights the contributions made by the T1DGC to understanding the genetic basis of T1D using both family (for linkage) and case-control (for genome-wide association) designs. The T1DGC conducted large-scale genetic research and used fine mapping to define risk regions. The T1DGC data, results, and samples have been made available to the scientific community, leading to the discovery of more than 100 loci associated with T1D risk, many with small effects and relevant to autoimmune pathways. The T1DGC not only expanded the list of genes contributing to disease risk but also identified noncoding genetic variation in disease-relevant cell types that contribute to the etiology of T1D. The success of the T1DGC and the NIDDK investment in the global consortium is highlighted in its continuing effect on mapping genetic variants to their function and identifying pathways that provide new targets for the prediction, prevention, and treatment of T1D.
Biological datasets often consist of thousands or millions of variables, e.g. genetic variants or biomarkers, and when sample sizes are large it is common to find many associated with an outcome of interest, for example, disease risk in a GWAS, at high levels of statistical significance, but with very small effects. The False Discovery Rate (FDR) is used to identify effects of interest based on ranking variables according to their statistical significance. Here, we develop a complementary measure to the FDR, the priorityFDR, that ranks variables by a combination of effect size and significance, allowing further prioritisation among a set of variables that pass a significance or FDR threshold. Applying to the largest GWAS of type 1 diabetes to date (15,573 cases and 158,408 controls), we identified 26 independent genetic associations, including two newly-reported loci, with qualitatively lower priorityFDRs than the remaining 175 signals. We detected putatively causal type 1 diabetes risk genes using Mendelian Randomisation, and found that these were located disproportionately close to low priorityFDR signals (p = 0.005), as were genes in the IL-2 pathway (p = 0.003). Selecting variables on both effect size and significance can lead to improved prioritisation for mechanistic follow-up studies from genetic and other large biological datasets.
Genome-wide association studies have identified SH2B3 as an important non-MHC gene for islet autoimmunity and type 1 diabetes (T1D). In this study, we found a single SH2B3 haplotype significantly associated with increased risk for human T1D. Fine mapping has demonstrated the most credible causative variant is the single nucleotide rs3184504*T polymorphism in SH2B3. To better characterize the role of SH2B3 in T1D, we used mouse modeling and found a T cellintrinsic role for SH2B3 regulating peripheral tolerance. SH2B3 deficiency had minimal effect on TCR signaling or proliferation across antigen doses, yet enhanced cell survival and cytokine signaling including common gamma chain-dependent and interferon-gamma receptor signaling. SH2B3 deficient naïve CD8+ T cells showed augmented STAT5-MYC and effector-related gene expression partially reversed with blocking autocrine IL-2 in culture. Using the RIP-mOVA model, we found CD8+ T cells lacking SH2B3 promoted early islet destruction and diabetes without requiring CD4+ T cell help. SH2B3-deficient cells demonstrated increased survival and reduced activation-induced cell death. Lastly, we created a spontaneous NOD.Sh2b3-/- mouse model and found markedly increased incidence and accelerated T1D across sexes. Collectively, these studies identify SH2B3 as a critical mediator of peripheral T cell tolerance limiting the T cell response to self-antigens.
Abstract Breast cancer includes several subtypes with distinct characteristic biological, pathologic, and clinical features. Elucidating subtype-specific genetic etiology could provide insights into the heterogeneity of breast cancer to facilitate the development of improved prevention and treatment approaches. In this study, we conducted pairwise case–case comparisons among five breast cancer subtypes by applying a case–case genome-wide association study (CC-GWAS) approach to summary statistics data of the Breast Cancer Association Consortium. The approach identified 13 statistically significant loci and eight suggestive loci, the majority of which were identified from comparisons between triple-negative breast cancer (TNBC) and luminal A breast cancer. Associations of lead variants in 12 loci remained statistically significant after accounting for previously reported breast cancer susceptibility variants, among which, two were genome-wide significant. Fine mapping implicated putative functional/causal variants and risk genes at several loci, e.g., 3q26.31/TNFSF10, 8q22.3/NACAP1/GRHL2, and 8q23.3/LINC00536/TRPS1, for TNBC as compared with luminal cancer. Functional investigation further identified rs16867605 at 8q22.3 as a SNP that modulates the enhancer activity of GRHL2. Subtype-informative polygenic risk scores (PRS) were derived, and patients with a high subtype-informative PRS had an up to two-fold increased risk of being diagnosed with TNBC instead of luminal cancers. The CC-GWAS PRS remained statistically significant after adjusting for TNBC PRS derived from traditional case–control GWAS in The Cancer Genome Atlas and the African Ancestry Breast Cancer Genetic Consortium. The CC-GWAS PRS was also associated with overall survival and disease-specific survival among patients with breast cancer. Overall, these findings have advanced our understanding of the genetic etiology of breast cancer subtypes, particularly for TNBC. Significance: The discovery of subtype-informative genetic risk variants for breast cancer advances our understanding of the etiologic heterogeneity of breast cancer, which could accelerate the identification of targets and personalized strategies for prevention and treatment.
The T cell Ubiquitin Ligand (TULA) protein family contains two members, UBASH3A and UBASH3B, that display similarities in protein sequence and domain structure. Both TULA proteins act to repress T cell activation via a combination of overlapping and nonredundant functions. UBASH3B acts mainly as a phosphatase that suppresses proximal T cell receptor (TCR) signaling. In contrast, UBASH3A acts primarily as an adaptor protein, interacting with other proteins (including UBASH3B) in T cells upon TCR stimulation and resulting in downregulation of TCR signaling and NF-κB signaling. Human genetic and functional studies have revealed another notable distinction between UBASH3A and UBASH3B: numerous genome-wide association studies have identified statistically significant associations between genetic variants in and around the UBASH3A gene and at least seven different autoimmune diseases, suggesting a key role of UBASH3A in autoimmunity. However, the evidence for an independent role of UBASH3B in autoimmune disease is limited. This review summarizes key findings regarding the roles of TULA proteins in T cell biology and autoimmunity, highlights the commonalities and differences between UBASH3A and UBASH3B, and speculates on the individual and joint effects of TULA proteins on T cell signaling.
ImportanceHeterogeneity in development of estrogen receptor (ER)-specific first primary breast cancer exists due to deleterious germline variants in moderate- to high-penetrance breast cancer susceptibility genes, but it is unknown if these associations occur in ER-specific CBC.ObjectiveTo determine the association of deleterious germline variants in breast cancer susceptibility genes with ER-specific CBC development and whether ER status of the first primary breast cancer modifies these associations.Design, Setting, and ParticipantsThis case-control study included CBC cases and matched unilateral breast cancer controls from The Women’s Environment, Cancer, and Radiation Epidemiology (WECARE) Study, a population-based case-control study. Eligible women were diagnosed between 1985 and 2000 with data and biospecimens collected from 2001 to 2004. Eligible participants were women younger than 55 years at first invasive breast cancer diagnosis. Participants were matched on age, diagnosis year, cancer registry region, and race and ethnicity, and countermatched on radiation treatment. For cases, CBC occurred 1 year or more following first breast cancer diagnosis. Analyses were performed from May to October 2024.ExposuresCHEK2 1100delC and deleterious variants in ATM, BRCA1, and BRCA2.Main Outcome and MeasureDevelopment of CBC, measured as a rate ratio (RR).ResultsA total of 1290 women were included in analysis (median [IQR] age at first diagnosis, 47 [42-51] years). The ER-positive CBC rate for women with deleterious ATM variants was 4 times higher than for women without deleterious ATM variants (RR, 4.84; 95% CI, 1.11-21.08; P = .04); no women with ER-negative CBC carried deleterious ATM variants. The ER-positive CBC rates for women with deleterious variants in BRCA2 or CHEK2 1100delC were 5 to 6 times higher than for women without deleterious variants in BRCA2 or CHEK2 1100delC, respectively (BRCA2: RR, 5.88; 95% CI, 2.61-13.26, P < .001; CHEK2 1100delC: RR, 6.06; 95% CI, 1.26-29.04; P = .02). The ER-negative CBC rate for women with deleterious BRCA1 variants was 26 times higher than for women without deleterious BRCA1 variants (RR, 26.16; 95% CI, 8.01-85.44; P < .001). First primary breast cancer ER status did not modify associations between deleterious variants and ER-specific CBC development.Conclusions and RelevanceIn this case-control study of CBC, deleterious variants in breast cancer susceptibility genes were differentially associated with ER-specific CBC development. Germline variation profile may inform estimates of outcomes for ER-specific CBC subtypes.
Background Contralateral breast cancer (CBC) is the most common second primary cancer diagnosed in breast cancer survivors, yet the understanding of the genetic susceptibility of CBC, particularly with respect to common variants, remains incomplete. This study aimed to investigate the genetic basis of CBC to better understand this malignancy. Findings We performed a genome-wide association analysis in the Women’s Environmental Cancer and Radiation Epidemiology (WECARE) Study of women with first breast cancer diagnosed at age < 55 years including 1161 with CBC who served as cases and 1668 with unilateral breast cancer (UBC) who served as controls. We observed two loci (rs59657211, 9q32, SLC31A2 / FAM225A and rs3815096, 6p22.1, TRIM31 ) with suggestive genome-wide significant associations ( P < 1 × 10 –6 ). We also found an increased risk of CBC associated with a breast cancer-specific polygenic risk score (PRS) comprised of 239 known breast cancer susceptibility single nucleotide polymorphisms (SNPs) (rate ratio per 1-SD change: 1.25; 95% confidence interval 1.14–1.36, P < 0.0001). The protective effect of chemotherapy on CBC risk was statistically significant only among patients with an elevated PRS ( P heterogeneity = 0.04). The AUC that included the PRS and known breast cancer risk factors was significantly elevated. Conclusions The present GWAS identified two previously unreported loci with suggestive genome-wide significance. We also confirm that an elevated risk of CBC is associated with a comprehensive breast cancer susceptibility PRS that is independent of known breast cancer risk factors. These findings advance our understanding of genetic risk factors involved in CBC etiology.
Background: Contralateral breast cancer (CBC) is associated with younger age at first diagnosis, family history and pathogenic germline variants (PGVs) in genes such as BRCA1, BRCA2 and PALB2. However, data regarding genetic factors predisposing to CBC among younger women who are BRCA1/2/PALB2-negative remain limited. Methods: In this nested case-control study, participants negative for BRCA1/2/PALB2 PGVs were selected from the WECARE Study. The burden of PGVs in established breast cancer risk genes was compared in 357 cases with CBC and 366 matched controls with unilateral breast cancer (UBC). The samples were sequenced in two phases. Whole exome sequencing was used in Group 1, 162 CBC and 172 UBC (mean age at diagnosis: 42 years). A targeted panel of genes was used in Group 2, 195 CBC and 194 UBC (mean age at diagnosis: 50 years). Comparisons of PGVs burdens between CBC and UBC were made in these groups, and additional stratified sub-analysis was performed within each group according to the age at diagnosis and the time from first breast cancer (BC). Results: The PGVs burden in Group 1 was significantly higher in CBC than in UBC (p = 0.002, OR = 2.5, 95CI: 1.2–5.6), driven mainly by variants in CHEK2 and ATM. The proportions of PGVs carriers in CBC and UBC in this group were 14.8% and 5.8%, respectively. There was no significant difference in PGVs burden between CBC and UBC in Group 2 (p = 0.4, OR = 1.4, 95CI: 0.7–2.8), with proportions of carriers being 8.7% and 8.2%, respectively. There was a significant association of PGVs in CBC with younger age. Metanalysis combining both groups confirmed the significant association between the burden of PGVs and the risk of CBC (p = 0.006) with the significance driven by the younger cases (Group 1). Conclusion: In younger BRCA1/BRCA2/PALB2-negative women, the aggregated burden of PGVs in breast cancer risk genes was associated with the increased risk of CBC and was inversely proportional to the age at onset.
There is increasing interest in the effects of low-dose ionizing radiation (IR) on plants as might occur during spaceflight, or as a consequence of human activities, such as nuclear power generation, that may result in the release of radioactive materials into the environment. High IR doses have long been used for the induction of mutations in plants with the goal of generating desirable traits for agribusiness. Less is known about the responses of plants to acute low doses of IR exposure. Here, we take a multi-omics approach to characterize the response to low dose IR in Arabidopsis thaliana . We adapt the Methyltransferase Accessibility Protocol for individual templates (MAPit) technique for use in plants allowing us to assay the epigenetic response to acute low-dose IR (10 cGy and 100 cGy) 72 hr after exposure, and, in parallel, use RNA sequencing to profile the transcription response at 1, 3, 24 and 72 hr after exposure. IR exposures as low as 10 cGy elicit robust genetic responses in A. thaliana detectable as early as 1 hr after exposure. Further examination revealed dose-dependent changes in gene expression, chromatin accessibility and DNA methylation that implicate the ethylene signalling pathway and abiotic stress response as underlying the transcriptional and epigenetic changes associated with IR. These changes are observable up to 72 hr post-exposure, suggesting that they are maintained well after the initial acute exposure. Our findings indicate that A. thaliana executes a multi-modal response to low-dose IR through induction and regulation of the ethylene response pathway.
Genome-wide association studies have identified numerous loci with allelic associations to Type 1 Diabetes (T1D) risk. Most disease-associated variants are enriched in regulatory sequences active in lymphoid cell types, suggesting that lymphocyte gene expression is altered in T1D. Here we assay gene expression between T1D cases and healthy controls in two autoimmunity-relevant lymphocyte cell types, memory CD4 + /CD25 + regulatory T cells (Treg) and memory CD4 + /CD25 - T cells, using a splicing event-based approach to characterize tissue-specific transcriptomes. Limited differences in isoform usage between T1D cases and controls are observed in memory CD4 + /CD25 - T-cells. In Tregs, 402 genes demonstrate differences in isoform usage between cases and controls, particularly RNA recognition and splicing factor genes. Many of these genes are regulated by the variable inclusion of exons that can trigger nonsense mediated decay. Our results suggest that dysregulation of gene expression, through shifts in alternative splicing in Tregs, contributes to T1D pathophysiology.
The proportions and phenotypes of immune cell subsets in peripheral blood undergo continual and dramatic remodeling throughout the human life span, which complicates efforts to identify disease-associated immune signatures in type 1 diabetes (T1D). We conducted cross-sectional flow cytometric immune profiling on peripheral blood from 826 individuals (stage 3 T1D, their first-degree relatives, those with ≥2 islet autoantibodies, and autoantibody-negative unaffected controls). We constructed an immune age predictive model in unaffected participants and observed accelerated immune aging in T1D. We used generalized additive models for location, shape, and scale to obtain age-corrected data for flow cytometry and complete blood count readouts, which can be visualized in our interactive portal (ImmScape); 46 parameters were significantly associated with age only, 25 with T1D only, and 23 with both age and T1D. Phenotypes associated with accelerated immunological aging in T1D included increased CXCR3+ and programmed cell death 1–positive (PD-1+) frequencies in naive and memory T cell subsets, despite reduced PD-1 expression levels on memory T cells. Phenotypes associated with T1D after age correction were predictive of T1D status. Our findings demonstrate advanced immune aging in T1D and highlight disease-associated phenotypes for biomarker monitoring and therapeutic interventions.
UBASH3A is a negative regulator of T cell activation and IL-2 production and plays key roles in autoimmunity. Although previous studies revealed the individual effects of UBASH3A on risk for type 1 diabetes (T1D; a common autoimmune disease), the relationship of UBASH3A with other T1D risk factors remains largely unknown. Given that another well-known T1D risk factor, PTPN22, also inhibits T cell activation and IL-2 production, we investigated the relationship between UBASH3A and PTPN22. We found that UBASH3A, via its Src homology 3 (SH3) domain, physically interacts with PTPN22 in T cells, and that this interaction is not altered by the T1D risk coding variant rs2476601 in PTPN22. Furthermore, our analysis of RNA-seq data from T1D cases showed that the amounts of UBASH3A and PTPN22 transcripts exert a cooperative effect on IL2 expression in human primary CD8+ T cells. Finally, our genetic association analyses revealed that two independent T1D risk variants, rs11203203 in UBASH3A and rs2476601 in PTPN22, interact statistically, jointly affecting risk for T1D. In summary, our study reveals novel interactions, both biochemical and statistical, between two independent T1D risk loci, and suggests how these interactions may affect T cell function and increase risk for T1D.
The Network for Pancreatic Organ donors with Diabetes (nPOD) is the largest biorepository of human pancreata and associated immune organs from donors with type 1 diabetes (T1D), maturity-onset diabetes of the young (MODY), cystic fibrosis-related diabetes (CFRD), type 2 diabetes (T2D), gestational diabetes, islet autoantibody positivity (AAb+), and without diabetes. nPOD recovers, processes, analyzes, and distributes high-quality biospecimens, collected using optimized standard operating procedures, and associated de-identified data/metadata to researchers around the world. Herein describes the release of high-parameter genotyping data from this collection. 372 donors were genotyped using a custom precision medicine single nucleotide polymorphism (SNP) microarray. Data were technically validated using published algorithms to evaluate donor relatedness, ancestry, imputed HLA, and T1D genetic risk score. Additionally, 207 donors were assessed for rare known and novel coding region variants via whole exome sequencing (WES). These data are publicly-available to enable genotype-specific sample requests and the study of novel genotype:phenotype associations, aiding in the mission of nPOD to enhance understanding of diabetes pathogenesis to promote the development of novel therapies.
For polygenic traits, associations with genetic variants can be detected over many chromosome regions, owing to the availability oflarge sample sizes. Most variants, however, have small effects on disease risk and, therefore, unravelling the causal variants, target genes, and biology of these variants is challenging. Here, we define the Bigger or False Discovery Rate (BFDR) as the probability that either a variant is a false-positive or a randomly drawn, true-positive association exceeds it in effect size. Using the BFDR, we identified 302 previously unreported signals with larger effect associations with type 1 diabetes and autoimmune thyroid disease. Out of 239 genome-wide significant signals in both diseases, only 66 (28%) show evidence for having a large effect using the BFDR, further demonstrating how using a combination of effect size and significance, rather than significance alone, is important in identifying SNPs and candidate genes for further investigation.