The basic-helix-loop-helix Per-Arnt-Sim (PAS) homology domain (bHLH-PAS) transcription factor (TF) family comprises critical sensors or actuators of physiological (hypoxia, tryptophan metabolites, neuronal activity, and appetite) and environmental (diet-derived metabolites and pollutants) stimuli regulating genes involved in signal adaptation and homeostasis. Despite the importance of this TF family, the mechanisms underlying specificity of DNA binding and target gene regulation by the bHLH-PAS subfamily remain unresolved. We systematically analysed cognate DNA binding hierarchies of prototypical bHLH-PAS family members (ARNT, ARNT2, HIF1α, HIF2α, AhR, NPAS4, SIM1), revealing large DNA binding footprints (12-15 bp) and unique mechanisms of DNA binding specificity involving preferential DNA sequences flanking the core motif. Flank-encoded DNA binding specificity discerns otherwise identical core sequence binding by SIM1 and the HIFs, mediated through N-terminal HIFα-DNA interactions. We also reveal an intimate relationship between DNA shape and core and flank TF binding that allows motif sequence flexibility and underpins multimodal mechanisms for achieving TF binding specificity. Furthermore, novel downstream SIM1 PAS-loop/DNA interactions are associated with AT-rich sequences contributing to DNA binding and transcriptional activity; these interactions are critical for TF biological function underpinning a monogenic cause of human hyperphagic obesity in a recapitulated SIM1.R171H knock-in mouse model.
Seminal fluid elicits an immune response in the uterine mucosa after mating that impacts embryo implantation and pregnancy, but the underlying molecular and cellular events are unclear. In this study, we report RNA sequencing to analyze the uterine response to seminal fluid after mating. Females exposed to seminal fluid of intact males exhibited gene expression changes on D3.5 post-coitum (pc) just prior to embryo implantation, compared to females mated with males surgically rendered seminal plasma deficient. Functional enrichment analysis revealed genes related to T cell activation amongst those with the largest fold-changes. Using flow cytometry we then showed profound changes in uterine T cell abundance and phenotype regulated by seminal fluid contact. While CD4+ and CD8+ T cells were elevated by seminal fluid, the most conspicuous change was in CD4-CD8- T cells expressing γδ T cell receptors (TCR). Mating with intact males caused a 8.3-fold increase in γδ T cell abundance compared to estrous virgin females, and a 22.4-fold increase in the proportion of γδ T cells expressing proliferation marker Ki67. Vγ6+ cells were the most abundant subpopulation in the uterus, followed by Vγ4+ and Vγ1+ T cells, and all three were similarly expanded after mating. Seminal plasma was critical for γδ T cell accumulation and activation in the endometrium, and similar changes occurred in uterine-draining lymph nodes but not spleen. These findings identify γδ T cells as prominent in the immune response to seminal fluid and imply key roles in uterine immune regulation and reproductive success.
Acute lymphoblastic leukaemia is a highly heterogeneous malignancy characterised by various genomic alterations that influence disease progression and therapeutic outcomes. Gene fusions involving the immunoglobulin heavy chain gene represent a complex and diverse category. These fusions often result in enhancer hijacking, upregulation of partner proto-oncogenes and contribute to leukemogenesis. This review highlights the mechanisms underlying IGH gene fusions, the critical role they play in ALL pathogenesis, and current detection technologies.
Preeclampsia is a hypertensive disorder of pregnancy with major maternal and fetal consequences. While the molecular basis of early-onset preeclampsia is well studied, the mechanisms underlying late-onset disease—and how they differ by fetal sex—remain poorly understood. Placental transcriptomic profiling at term can reveal persistent molecular alterations reflecting cumulative disease processes. We conducted a cross-sectional observational analysis of placental gene expression using RNA sequencing in a subset of 58 term placentas (21 male-bearing and 37 female-bearing pregnancies) drawn from two large prospective birth cohorts. Pregnancies were classified based on a clinical diagnosis of late-onset preeclampsia (diagnosed ≥ 20 weeks’ gestation according to ISSHP criteria) or as uncomplicated pregnancies. We then assessed for differential gene expression. Cell type proportions were estimated using CIBERSORTx from a placenta-specific reference single-cell dataset. Weighted gene co-expression network analysis identified modules of co-expressed genes associated with late-onset preeclampsia and fetal sex. Differential gene expression analysis identified 150 genes with altered expression in male-bearing placentas from pregnancies with late-onset preeclampsia compared to those from uncomplicated pregnancies. No differentially expressed genes were identified in female-bearing placentas. Cell type deconvolution revealed increased abundance of CD14 + monocytes and CD8 + activated T cells (log odds of 1.42 and 1.44 respectively) and reduced fetal GZMK natural killer cells (log odds of 0.60) in male-bearing placentas from affected pregnancies. In female-bearing placentas, late-onset preeclampsia was associated with increased fetal nucleated red blood cells and maternal plasma cells (log odds of 1.33 and 1.40 respectively). Male-specific co-expression analysis identified gene modules enriched for biological processes including RNA processing, immune regulation, and metabolism. Placental transcription and cellular responses to late-onset preeclampsia differ by fetal sex. Evidence of altered immune cell composition and gene co-expression in male-bearing placentas suggests a sex-specific vulnerability. These findings highlight the importance of considering fetal sex in molecular investigation and clinical management of preeclampsia. Preeclampsia is a common pregnancy complication marked by high blood pressure, but how it affects the placenta, especially in later pregnancy and depending on the baby’s sex, is not well understood. In this study, we analysed placental tissue from pregnancies with and without late-onset preeclampsia using RNA sequencing. By separating the data based on whether the neonate was male or female, we found striking differences in gene expression. Only placentas from male-bearing pregnancies showed significant changes in gene expression linked to preeclampsia. These changes involved genes related to immune response, metabolism and vascular function. We also used computational tools to estimate what types of cells were present in each placental sample. In male-bearing pregnancies affected by late-onset preeclampsia, there was a notable increase in certain immune cells, suggesting an altered immune response and increased inflammation. In contrast, female-bearing pregnancies affected by late-onset preeclampsia showed an increase in cell composition for two blood related cell types, but no significant gene expression differences. By grouping genes that worked together into networks, we identified several groups, especially in placentas from male-bearing pregnancies, that were strongly associated with biological processes known to be disrupted in preeclampsia, such as blood vessel formation, extracellular matrix remodelling, and hormone regulation. These findings emphasise the importance of considering fetal sex in pregnancy research and could help guide future sex-specific diagnostic or treatment strategies. Sex-specific response: Significant placental gene expression changes between preeclamptic and uncomplicated pregnancies were found only in male-bearing pregnancies. Immune involvement: Increased abundance of CD14 + monocytes and CD8 + T cells in placenta from male-bearing pregnancies suggests a heightened immune response. ECM and vascular remodelling: Key gene modules in male-bearing placentas were enriched for extracellular matrix organisation and vascular function. Transcriptomic network disruption: Co-expression analysis revealed distinct gene modules associated with preeclampsia in male-bearing pregnancies, reflecting changes in metabolism and gene signaling. Female resilience hypothesis: Despite altered cell composition, female-bearing placentas showed no significant gene expression changes, suggesting possible adaptive resilience.
Understanding of molecular mechanisms contributing to the pathophysiology of endometriosis, and upstream drivers of lesion formation, remains limited. Using a C57Bl/6 mouse model in which decidualized endometrial tissue is injected subcutaneously in the abdomen of recipient mice, we generated a comprehensive profile of gene expression in decidualized endometrial tissue (n=4), and in endometriosis-like lesions at Day 7 (n=4) and Day 14 (n=4) of formation. High-throughput mRNA sequencing allowed identification of genes and pathways involved in the initiation and progression of endometriosis-like lesions. We observed distinct patterns of gene expression with substantial differences between the lesions and the decidualized endometrium that remained stable across the two lesion timepoints, and showed similarity to transcriptional changes implicated in human endometriosis lesion formation. Pathway enrichment analysis revealed several immune and inflammatory response-associated canonical pathways, multiple potential upstream regulators, and involvement of genes not previously implicated in endometriosis pathogenesis, including IRF2BP2 and ZBTB10, suggesting novel roles in disease progression. Collectively, the provided data will be a useful resource to inform research on the molecular mechanisms contributing to endometriosis-like lesion development in this mouse model.
Stem cell-based therapy is a potential alternative strategy for brain repair, with neural stem cells (NSC) presenting as the most promising candidates. Obtaining sufficient quantities of NSC for clinical applications is challenging, therefore alternative cell types, such as neural crest-derived dental pulp stem cells (DPSC), may be considered. Human DPSC possess neurogenic potential, exerting positive effects in the damaged brain through paracrine effects. However, a method for conversion of DPSC into NSC has yet to be developed. Here, overexpression of octamer-binding transcription factor 4 (OCT4) in combination with neural inductive conditions was used to reprogram human DPSC along the neural lineage. The reprogrammed DPSC demonstrated a neuronal-like phenotype, with increased expression levels of neural markers, limited capacity for sphere formation, and enhanced neuronal but not glial differentiation. Transcriptomic analysis further highlighted the expression of genes associated with neural and neuronal functions. In vivo analysis using a developmental avian model showed that implanted DPSC survived in the developing central nervous system and respond to endogenous signals, displaying neuronal phenotypes. Therefore, OCT4 enhances the neural potential of DPSC, which exhibited characteristics aligning with neuronal progenitors. This method can be used to standardise DPSC neural induction and provide an alternative source of neural cell types.
Genomic information is increasingly used to inform medical treatments and manage future disease risks. However, any personal and societal gains must be carefully balanced against the risk to individuals contributing their genomic data. Expanding our understanding of actionable genomic insights requires researchers to access large global datasets to capture the complexity of genomic contribution to diseases. Similarly, clinicians need efficient access to a patient's genome as well as population-representative historical records for evidence-based decisions. Both researchers and clinicians hence rely on participants to consent to the use of their genomic data, which in turn requires trust in the professional and ethical handling of this information. Here, we review existing and emerging solutions for secure and effective genomic information management, including storage, encryption, consent, and authorization that are needed to build participant trust. We discuss recent innovations in cloud computing, quantum-computing-proof encryption, and self-sovereign identity. These innovations can augment key developments from within the genomics community, notably GA4GH Passports and the Crypt4GH file container standard. We also explore how decentralized storage as well as the digital consenting process can offer culturally acceptable processes to encourage data contributions from ethnic minorities. We conclude that the individual and their right for self-determination needs to be put at the center of any genomics framework, because only on an individual level can the received benefits be accurately balanced against the risk of exposing private information.
Cells undergo a major epigenome reconfiguration when reprogrammed to human induced pluripotent stem cells (hiPS cells). However, the epigenomes of hiPS cells and human embryonic stem (hES) cells differ significantly, which affects hiPS cell function1-8. These differences include epigenetic memory and aberrations that emerge during reprogramming, for which the mechanisms remain unknown. Here we characterized the persistence and emergence of these epigenetic differences by performing genome-wide DNA methylation profiling throughout primed and naive reprogramming of human somatic cells to hiPS cells. We found that reprogramming-induced epigenetic aberrations emerge midway through primed reprogramming, whereas DNA demethylation begins early in naive reprogramming. Using this knowledge, we developed a transient-naive-treatment (TNT) reprogramming strategy that emulates the embryonic epigenetic reset. We show that the epigenetic memory in hiPS cells is concentrated in cell of origin-dependent repressive chromatin marked by H3K9me3, lamin-B1 and aberrant CpH methylation. TNT reprogramming reconfigures these domains to a hES cell-like state and does not disrupt genomic imprinting. Using an isogenic system, we demonstrate that TNT reprogramming can correct the transposable element overexpression and differential gene expression seen in conventional hiPS cells, and that TNT-reprogrammed hiPS and hES cells show similar differentiation efficiencies. Moreover, TNT reprogramming enhances the differentiation of hiPS cells derived from multiple cell types. Thus, TNT reprogramming corrects epigenetic memory and aberrations, producing hiPS cells that are molecularly and functionally more similar to hES cells than conventional hiPS cells. We foresee TNT reprogramming becoming a new standard for biomedical and therapeutic applications and providing a novel system for studying epigenetic memory.
Background Receptivity of the uterus is essential for embryo implantation and progression of mammalian pregnancy. Acquisition of receptivity involves major molecular and cellular changes in the endometrial lining of the uterus from a non-receptive state at ovulation, to a receptive state several days later. The precise molecular mechanisms underlying this transition and their upstream regulators remain to be fully characterized. Here, we aimed to generate a comprehensive profile of the endometrial transcriptome in the peri-ovulatory and peri-implantation states, to define the genes and gene pathways that are different between these states, and to identify new candidate upstream regulators of this transition, in the mouse. Results High throughput RNA-sequencing was utilized to identify genes and pathways expressed in the endometrium of female C57Bl/6 mice at estrus and on day 3.5 post-coitum (pc) after mating with BALB/c males (n = 3–4 biological replicates). Compared to the endometrium at estrus, 388 genes were considered differentially expressed in the endometrium on day 3.5 post-coitum. The transcriptional changes indicated substantial modulation of uterine immune and vascular systems during the pre-implantation phase, with the functional terms Angiogenesis, Chemotaxis, and Lymphangiogenesis predominating. Ingenuity Pathway Analysis software predicted the activation of several upstream regulators previously shown to be involved in the transition to receptivity including various cytokines, ovarian steroid hormones, prostaglandin E2, and vascular endothelial growth factor A. Our analysis also revealed four candidate upstream regulators that have not previously been implicated in the acquisition of uterine receptivity, with growth differentiation factor 2, lysine acetyltransferase 6 A, and N-6 adenine-specific DNA methyltransferase 1 predicted to be activated, and peptidylprolyl isomerase F predicted to be inhibited. Conclusions This study confirms that the transcriptome of a receptive uterus is vastly different to the non-receptive uterus and identifies several genes, regulatory pathways, and upstream drivers not previously associated with implantation. The findings will inform further research to investigate the molecular mechanisms of uterine receptivity.
B-cell acute lymphoblastic leukaemia (B-ALL) is characterised by diverse genomic alterations, the most frequent being gene fusions detected via transcriptomic analysis (mRNA-seq). Due to its hypervariable nature, gene fusions involving the Immunoglobulin Heavy Chain (IGH) locus can be difficult to detect with standard gene fusion calling algorithms and significant computational resources and analysis times are required. We aimed to optimize a gene fusion calling workflow to achieve best-case sensitivity for IGH gene fusion detection. Using Nextflow, we developed a simplified workflow containing the algorithms FusionCatcher, Arriba, and STAR-Fusion. We analysed samples from 35 patients harbouring IGH fusions (IGH::CRLF2 n = 17, IGH::DUX4 n = 15, IGH::EPOR n = 3) and assessed the detection rates for each caller, before optimizing the parameters to enhance sensitivity for IGH fusions. Initial results showed that FusionCatcher and Arriba outperformed STAR-Fusion (85-89% vs. 29% of IGH fusions reported). We found that extensive filtering in STAR-Fusion hindered IGH reporting. By adjusting specific filtering steps (e.g., read support, fusion fragments per million total reads), we achieved a 94% reporting rate for IGH fusions with STAR-Fusion. This analysis highlights the importance of filtering optimization for IGH gene fusion events, offering alternative workflows for difficult-to-detect high-risk B-ALL subtypes.
Horns, a form of headgear carried by Bovidae, have ethical and economic implications for ruminant production species such as cattle and goats. Hornless (polled) individuals are preferred. In cattle, four genetic variants (Celtic, Friesian, Mongolian and Guarani) are associated with the polled phenotype, which are clustered in a 300-kb region on chromosome 1. As the variants are intergenic, the functional effect is unknown. The aim of this study was to determine if the POLLED variants affect chromatin structure or disrupt enhancers using publicly available data. Topologically associating domains (TADs) were analyzed using Angus- and Brahman-specific Hi-C reads from lung tissue of an Angus (Celtic allele) cross Brahman (horned) fetus. Predicted bovine enhancers and chromatin immunoprecipitation sequencing peaks for histone modifications associated with enhancers (H3K27ac and H3K4me1) were mapped to the POLLED region. TADs analyzed from Angus- and Brahman-specific Hi-C reads were the same, therefore, the Celtic variant does not appear to affect this level of chromatin structure. The Celtic variant is located in a different TAD from the Friesian, Mongolian, and Guarani variants. Predicted enhancers and histone modifications overlapped with the Guarani and Friesian variants but not the Celtic or Mongolian variants. This study provides insight into the mechanisms of the POLLED variants for disrupting horn development. These results should be validated using data produced from the horn bud region of horned and polled bovine fetuses.
Background Sea snakes underwent a complete transition from land to sea within the last ~ 15 million years, yet they remain a conspicuous gap in molecular studies of marine adaptation in vertebrates. Results Here, we generate four new annotated sea snake genomes, three of these at chromosome-scale ( Hydrophis major , H . ornatus and H. curtus ), and perform detailed comparative genomic analyses of sea snakes and their closest terrestrial relatives. Phylogenomic analyses highlight the possibility of near-simultaneous speciation at the root of Hydrophis , and synteny maps show intra-chromosomal variations that will be important targets for future adaptation and speciation genomic studies of this system. We then used a strict screen for positive selection in sea snakes (against a background of seven terrestrial snake genomes) to identify genes over-represented in hypoxia adaptation, sensory perception, immune response and morphological development. Conclusions We provide the best reference genomes currently available for the prolific and medically important elapid snake radiation. Our analyses highlight the phylogenetic complexity and conserved genome structure within Hydrophis . Positively selected marine-associated genes provide promising candidates for future, functional studies linking genetic signatures to the marine phenotypes of sea snakes and other vertebrates.
The search for novel microRNA (miRNA) biomarkers in plasma is hampered by haemolysis, the lysis and subsequent release of red blood cell contents, including miRNAs, into surrounding fluid. The biomarker potential of miRNAs comes in part from their multicompartment origin and the long-lived nature of miRNA transcripts in plasma, giving researchers a functional window for tissues that are otherwise difficult or disadvantageous to sample. The inclusion of red-blood-cell-derived miRNA transcripts in downstream analysis introduces a source of error that is difficult to identify posthoc and may lead to spurious results. Where access to a physical specimen is not possible, our tool will provide an in silico approach to haemolysis prediction. We present DraculR, an interactive Shiny/R application that enables a user to upload miRNA expression data from a short-read sequencing of human plasma as a raw read counts table and interactively calculate a metric that indicates the degree of haemolysis contamination. The code, DraculR web tool and its tutorial are freely available as detailed herein.
AimsTo investigate epigenomic indices of diabetic kidney disease (DKD) susceptibility among high-risk populations with type 2 diabetes mellitus.MethodsKDIGO (Kidney Disease: Improving Global Outcomes) clinical guidelines were used to classify people living with or without DKD. Differential gene methylation of DKD was then assessed in a discovery Aboriginal Diabetes Study cohort (PROPHECY, 89 people) and an external independent study from Thailand (THEPTARIN, 128 people). Corresponding mRNA levels were also measured and linked to levels of albuminuria and eGFR.ResultsIncreased DKD risk was associated with reduced methylation and elevated gene expression in the PROPHECY discovery cohort of Aboriginal Australians and these findings were externally validated in the THEPTARIN diabetes registry of Thai people living with type 2 diabetes mellitus.ConclusionsNovel epigenomic scores can improve diagnostic performance over clinical modelling using albuminuria and GFR alone and can distinguish DKD susceptibility.
Progesterone receptor (PGR) plays diverse roles in reproductive tissues and thus coordinates mammalian fertility. In the ovary, rapid acute induction of PGR is the key determinant of ovulation through transcriptional control of a unique set of genes that culminates in follicle rupture. However, the molecular mechanisms for this specialized PGR function in ovulation is poorly understood. We have assembled a detailed genomic profile of PGR action through combined ATAC-seq, RNA-seq and ChIP-seq analysis in wildtype and isoform-specific PGR null mice. We demonstrate that stimulating ovulation rapidly reprograms chromatin accessibility in two-thirds of sites, correlating with altered gene expression. An ovary-specific PGR action involving interaction with RUNX transcription factors was observed with 70% of PGR-bound regions also bound by RUNX1. These transcriptional complexes direct PGR binding to proximal promoter regions. Additionally, direct PGR binding to the canonical NR3C motif enable chromatin accessibility. Together these PGR actions mediate induction of essential ovulatory genes. Our findings highlight a novel PGR transcriptional mechanism specific to ovulation, providing new targets for infertility treatments or new contraceptives that block ovulation.
Multiple myeloma (MM) is an incurable malignancy characterised by uncontrolled proliferation of plasma cells (PCs) in the bone marrow (BM). MM is a genetically heterogeneous disease with each patient's PCs harbouring unique genetic mutations, however the development of MM tumours is not only dependent on the underlying genetics but also on the selective pressures applied by the BM microenvironment. Hence, we hypothesise that identifying the dependencies which promote MM cell outgrowth in the BM microenvironment will allow for the identification of new drug targets. This project aims to use an in vivo functional screen, combining CRISPR-Cas9 gene editing, using a single guide RNA (sgRNA) library targeting thousands of genes, with a murine model of MM, to identify novel genes involved in MM tumour development. The Bassik human apoptosis and cancer CRISPR knockout library (Addgene #101926) was used to transduce MM cells with 31,324 unique sgRNAs targeting 3,015 genes and 1,500 control regions. The human MM cell line OPM2 was transduced with the Cas9 transgene (FuCas9GFP), followed by the sgRNA expression vector (pMCB320) generating OPM2-Cas9-sgRNA cells. These cells were then injected (5x10 5 cells/mouse, n=8 mice + 1 tumour-naive control) into the tibiae of immunodeficient NOD.Cg-Prkdc scid Il2rg tm1Wjl/SzJ (NSG) mice. Four weeks post-injection, the primary tumour within the injected tibia was isolated and sgRNA frequencies were assessed via next generation sequencing and compared with that of the initial library and the injected cells using the MAGeCK algorithm. To identify genes that play roles in MM tumour formation in patients, we assessed the expression of our identified genes in CD138-selected BM PCs from newly diagnosed MM patients (n=155) when compared with normal PCs from healthy donors (n=5) in microarray dataset E-MTAB-363 (ArrayExpress). Furthermore, to identify potential pathway involvement of our identified genes we performed gene set enrichment analysis (GSEA) using the University of California San Diego and Broad Institute developed GSEA software, comparing to Hallmark Gene Sets. Comparison of the in vitro cultured cells with the initial library identified 34 genes with undetectable (p<0.05) sgRNAs post in vitro culture, suggesting these genes play key roles in the survival of OPM2 cells. This list contained a number of genes known to control MM cell survival, such as IRF4 and CCND2. Additionally, sgRNAs targeting a further 115 genes were significantly depleted (p<0.05) in vitro and were completely undetectable in vivo, suggesting that these are implicated in cell proliferation and tumour formation in vivo. This gene list contained some already known MM drivers such as MYC, KRAS and DIS3. Furthermore, we identified 28 genes where CRISPR knockout had no significant effect in vitro but were critical dependencies for tumour development, with sgRNAs being undetectable (p<0.05) in vivo. This list contained genes known to be important for MM and BM stromal cell interaction such as ALCAM and SPP1. Notably, 25% of the genes identified as critical dependencies for OPM2 cells in vitro and/or in vivo are significantly upregulated >2-fold (p<0.05; limma) in MM patient PCs when compared with normal PCs (E-MTAB-363), suggesting these may play a role in tumour development in patients. Gene set enrichment analysis revealed that the genes that were essential in vitro were predominantly involved in cell survival related pathways such as cell cycle regulation (G2M checkpoint) and DNA replication/repair; however, the uniquely in vivo genes are predominantly involved in cell metabolism pathways such as MTORC1 signalling, oxidative phosphorylation and glycolysis. Overall, this study has successfully used a functional genetic screen to identify novel gene dependencies required for in vitro MM cell growth and in vivo MM tumour development. A secondary genetic screen will be performed in vivo to validate and refine this list of candidate genes. These validated genes will be further investigated using patient derived data sets to determine if they represent common vulnerabilities in patients and in multiple MM cell lines to investigate their specific mechanistic and pathway involvement. Identification and understanding the driving dependencies which promote MM cell outgrowth in the BM microenvironment will allow for the identification and development of new drug targets to improve patient outcomes.
Chromosomal rearrangements involving the KMT2A gene occur frequently in acute lymphoblastic leukaemia (ALL). KMT2A-rearranged ALL (KMT2Ar ALL) has poor long-term survival rates and is the most common ALL subtype in infants less than 1 year of age. KMT2Ar ALL frequently occurs with additional chromosomal abnormalities including disruption of the IKZF1 gene, usually by exon deletion. Typically, KMT2Ar ALL in infants is accompanied by a limited number of cooperative le-sions. Here we report a case of aggressive infant KMT2Ar ALL harbouring additional rare IKZF1 gene fusions. Comprehensive genomic and transcriptomic analyses were performed on sequential samples. This report highlights the genomic complexity of this particular disease and describes the novel gene fusions IKZF1::TUT1 and KDM2A::IKZF1.
The basic-Helix-Loop-Helix Per-Arnt-Sim (PAS) homology domain (bHLH-PAS) transcription factor (TF) family comprises critical biological sensors of physiological (hypoxia, tryptophan metabolites, neuronal activity, and appetite) and environmental (diet derived metabolites and environmental pollutants) stimuli to regulate genes involved in signal adaptation and homeostasis 1 . bHLH TFs bind DNA as homo or heterodimers via E-box (CANNTG) response elements, however the DNA binding specificity of the PAS domain-containing bHLH subfamily remains unresolved 1 . We systematically analysed cognate DNA binding hierarchies of prototypical bHLH-PAS family members (ARNT, ARNT2, HIF1α, HIF2α, AhR, NPAS4, SIM1) and demonstrate distinct core (NNCGTG) specificities for different heterodimer classes. The results also show that bHLH-PAS TFs bind over a large footprint 12-15bp and recognise preferential DNA sequences flanking the core. For example, specificity beyond otherwise identical core binding by SIM1 and the HIFs is mediated through N-terminal HIFα-DNA interactions. We also reveal an intimate relationship between DNA shape and both core and flanking TF binding allowing motif sequence flexibility and underpinning TF binding specificity. Furthermore, DNA-shape affinity relationships revealed that novel downstream PAS-A-loop DNA interactions are associated with AT-rich sequences that lead to high-affinity binding, and that loss of this function underpins a monogenic cause of human hyperphagic obesity in a recapitulated SIM1.R171H knock-in mouse model. Importantly, models of protein-DNA binding accurately predict in vivo occupancy, while response element methylation blocks DNA binding and predicts cell type specific chromatin occupancy. These data provide a definitive and accurate map of bHLH-PAS TF specificity and target selectivity through novel flanking protein-DNA interactions that are crucial for in vivo biological function.
Plant DNA preserved in ancient specimens has recently gained importance as a tool in comparative genomics, allowing the investigation of evolutionary processes in plant genomes through time. However, recovering the genomic information contained in such specimens is challenging owing to the presence of secondary substances that limit DNA retrieval. In this chapter, we provide a DNA extraction protocol optimized for the recovery of DNA from degraded plant materials. The protocol is based on a commercially available DNA extraction kit that does not require handling of hazardous reagents.
Background Genome-wide association studies (GWAS) have enabled the discovery of single nucleotide polymorphisms (SNPs) that are significantly associated with many autoimmune diseases including type 1 diabetes (T1D). However, many of the identified variants lie in non-coding regions, limiting the identification of mechanisms that contribute to autoimmune disease progression. To address this problem, we developed a variant filtering workflow called 3DFAACTS-SNP to link genetic variants to target genes in a cell-specific manner. Here, we use 3DFAACTS-SNP to identify candidate SNPs and target genes associated with the loss of immune tolerance in regulatory T cells (Treg) in T1D. Results Using 3DFAACTS-SNP, we identified from a list of 1228 previously fine-mapped variants, 36 SNPs with plausible Treg-specific mechanisms of action. The integration of cell type-specific chromosome conformation capture data in 3DFAACTS-SNP identified 266 regulatory regions and 47 candidate target genes that interact with these variant-containing regions in Treg cells. We further demonstrated the utility of the workflow by applying it to three other SNP autoimmune datasets, identifying 16 Treg-centric candidate variants and 60 interacting genes. Finally, we demonstrate the broad utility of 3DFAACTS-SNP for functional annotation of all known common (> 10% allele frequency) variants from the Genome Aggregation Database (gnomAD). We identified 9376 candidate variants and 4968 candidate target genes, generating a list of potential sites for future T1D or other autoimmune disease research. Conclusions We demonstrate that it is possible to further prioritise variants that contribute to T1D based on regulatory function, and illustrate the power of using cell type-specific multi-omics datasets to determine disease mechanisms. Our workflow can be customised to any cell type for which the individual datasets for functional annotation have been generated, giving broad applicability and utility.