Early diagnosis of lung cancer is critical for timely intervention and reducing mortality. The immune system and cancer are intricately linked, which provides a unique opportunity to monitor changes in the immune system as a biomarker of cancer development. We collected bulk blood transcriptome from 432 lung cancer cases, 8154 healthy controls, and 14,187 samples with other diseases from 241 datasets for discovery. We also obtained a prospectively enrolled cohort of 371 subjects (172 with lung cancer) and 454 subjects from the Framingham Heart Study (42 with lung cancer) for validation. Furthermore, we integrated single-cell RNA sequencing profiles of 1,022,063 cells from 260 blood, lymph node, or lung tissue samples from lung cancer patients and other samples across 15 datasets. We performed a multi-cohort blood transcriptome meta-analysis and identified 6 genes consistently differentially expressed between lung cancer and other samples. Using the 6-gene signature, we defined a lung cancer score that was primarily derived from myeloid cells and was consistently higher in tumor-associated macrophages and fibroblasts than their normal counterparts. In the prospectively enrolled cohort, the diagnostic classifier using the lung cancer score had an AUROC of 0.822 (95% CI: 0.78-0.864) for distinguishing patients with lung cancer from control or benign samples. The classifiers could also potentially reduce the need for additional testing in 37% of patients with benign lung conditions at 90% sensitivity. Importantly, the lung cancer score was also significantly associated with an elevated risk of future lung cancer diagnosis in the Framingham cohort study. Together, we identified a robust blood-based immune gene signature for early detection of lung cancer, which has potential for further clinical development to aid early cancer detection and diagnosis.
Understanding breast cancer genetic risk relies on identifying causal variants and candidate target genes in risk loci identified by genome-wide association studies (GWAS), which remains challenging. Since most loci fall in active gene regulatory regions, we developed a novel approach facilitated by pinpointing the variants with greater regulatory potential in the disease’s tissue of origin. Through genome-wide differential allelic expression (DAE) analysis, using microarray data from 64 normal breast tissue samples, we mapped the variants associated with DAE (daeQTLs). Then, we intersected these with GWAS data to reveal candidate risk regulatory variants and analysed their cis-acting regulatory potential. Finally, we validated our approach by extensive functional analysis of the 5q14.1 breast cancer risk locus. We observed widespread gene expression regulation by cis-acting variants in breast tissue, with 65% of coding and noncoding expressed genes displaying DAE (daeGenes). We identified over 54 K daeQTLs for 6761 (26%) daeGenes, including 385 daeGenes harbouring variants previously associated with BC risk. We found 1431 daeQTLs mapped to 93 different loci in strong linkage disequilibrium with risk-associated variants (risk-daeQTLs), suggesting a link between risk-causing variants and cis-regulation. There were 122 risk-daeQTL with stronger cis-acting potential in active regulatory regions with protein binding evidence. These variants mapped to 41 risk loci, of which 29 had no previous report of target genes and were candidates for regulating the expression levels of 65 genes. As validation, we identified and functionally characterised five candidate causal variants at the 5q14.1 risk locus targeting the ATG10 and ATP6AP1L genes, likely acting via modulation of alternative transcription and transcription factor binding. Our study demonstrates the power of DAE analysis and daeQTL mapping to identify causal regulatory variants and target genes at breast cancer risk loci, including those with complex regulatory landscapes. It additionally provides a genome-wide resource of variants associated with DAE for future functional studies.
BACKGROUND:Lung cancer is the leading cause of cancer-related death in the world. In contrast to many other cancers, a direct connection to modifiable lifestyle risk in the form of tobacco smoke has long been established. More than 50% of all smoking-related lung cancers occur in former smokers, 40% of which occur more than 15 years after smoking cessation. Despite extensive research, the molecular processes for persistent lung cancer risk remain unclear. We thus set out to examine whether risk stratification in the clinic and in the general population can be improved upon by the addition of genetic data and to explore the mechanisms of the persisting risk in former smokers. METHODS:We analysed transcriptomic data from accessible airway tissues of 487 subjects, including healthy volunteers and clinic patients of different smoking statuses. We developed a computational model to assess smoking-associated gene expression changes and their reversibility after smoking is stopped, comparing healthy subjects to clinic patients with and without lung cancer. RESULTS:We find persistent smoking-associated immune alterations to be a hallmark of the clinic patients. Integrating previous GWAS data using a transcriptional network approach, we demonstrate that the same immune- and interferon-related pathways are strongly enriched for genes linked to known genetic risk factors, demonstrating a causal relationship between immune alteration and lung cancer risk. Finally, we used accessible airway transcriptomic data to derive a non-invasive lung cancer risk classifier. CONCLUSIONS:Our results provide initial evidence for germline-mediated personalized smoke injury response and risk in the general population, with potential implications for managing long-term lung cancer incidence and mortality.
PDF file - 293K, Figure S1. SETD8 is overexpressed in bladder cancer. Figure S2. SETD8 is overexpressed in CML, HCC and pancreatic cancer. Figure S3. SETD8 expression in various types of cell lines. Figure S4. PIP box in SETD8 and amino acid sequence alignment of PCNA. Figure S5. Chromatogram of amino acids obtained after acid hydrolysis of PCNA. Figure S6. An in vivo methylation of PCNA by SETD8. Figure S7. Validation of the anti-mono-methylated K248 PCNA antibody. Figure S8. Validation of methylation status of endogenous PCNA using a specific antibody. Figure S9. SW780 cells were pre-treated with siEGFP and siSETD8 for 24 hours and then, cells were treated with 100 g/ml of cycloheximide (CHX) for 0, 4 and 8 hours. Figure S10. Methylated PCNA enhances the interaction with FEN1. Figure S11 Effects of PCNA methylation on the sensitivity to H2O2 stress. Figure S12. A knockdown effect of endogenous PCNA in HeLa cells was confirmed by quantitative real-time PCR. Figure S13. Expression levels of SETD8 in 78 normal tissues.
Searchable abstracts of presentations at key conferences in endocrinology ISSN 1470-3947 (print) | ISSN 1479-6848 (online)
Supplementary Tables 1-4 from Tagging Single Nucleotide Polymorphisms in Cell Cycle Control Genes and Susceptibility to Invasive Epithelial Ovarian Cancer
Supplementary Tables 1-2 from Telomere Length in Prospective and Retrospective Cancer Case-Control Studies
PDF file - 69K, Table S1. Primer sequences for quantitative RT-PCR. Table S2. siRNA sequences. Table S3. Sequences of oligonucleotides for Okazaki fragment maturation assay. Table S4. Statistical analysis of SETD8 expression levels in clinical bladder tissues. Table S5. Gene expression profile of SETD8 in cancer tissues analyzed by cDNA microarray. Table S6. Correlation of SETD8 and PCNA expressions at the protein level.
Supplementary Figures 1-9, Tables 1-2, Methods from Demethylation of RB Regulator MYPT1 by Histone Demethylase LSD1 Promotes Cell Cycle Progression in Cancer Cells
Supplementary Table S1 from Tagging Single-Nucleotide Polymorphisms in Antioxidant Defense Enzymes and Susceptibility to Breast Cancer
Supplementary Information and Tables 1-5 from Association Study of 69 Genes in the Ret Pathway Identifies Low-penetrance Loci in Sporadic Medullary Thyroid Carcinoma
Supplementary Table 1 from Association Study of Prostate Cancer Susceptibility Variants with Risks of Invasive Ovarian, Breast, and Colorectal Cancer
<p>Supplementary Material: Figure S1: Nuclear localisation and fluorescent properties of FRET constructs Figure S2: Quantification of FRET experiment Figure S3: Relative mRNA expression of ESR1 target genes Figure S4: Characterisation of MCF-7 clones overexpressing NFIB and YBX1 Figure S5: Separated growth curves shown in Figure 4 Figure S6: Kaplan-Meier survival curves for breast cancer cases stratified by YBX1 expression Figure S7: Response of breast cancer PDX models to tamoxifen treatment Table S1: Primers used in RT-PCR Table S2: Antibodies used in Western blots</p>
Motivation: Dendrogram is a classical diagram for visualizing binary trees. Although efficient to represent hierarchical relations, it provides limited space for displaying information on the leaf elements, especially for large trees. Results: Here, we present TreeAndLeaf, an R/Bioconductor package that implements a hybrid layout strategy to represent tree diagrams with focus on the leaves. The TreeAndLeaf package combines force-directed graph and tree layout algorithms using a single visualization system, allowing projection of multiple layers of information onto a graph-tree diagram. The Supplementary Information Information provides two case studies that use breast cancer data from epidemiological and experimental studies.
Background Breast cancer (BC) genome-wide association studies (GWAS) have identified hundreds of risk-loci that require novel approaches to reveal the causal variants and target genes within them. As causal variants are most likely regulators of gene expression, we hypothesize that their identification is facilitated by pinpointing the variants with greater regulatory potential within risk-loci. Methods We performed genome-wide differential allelic expression (DAE) analysis using microarrays data from 64 normal breast tissue samples. Then, we mapped the variants associated with DAE (daeQTLs) and intersected these with GWAS data to reveal candidate risk regulatory variants. Finally, we validated our approach by functionally analysing the 5q14.1 breast cancer risk-locus. Results We found widespread gene expression regulation by cis-acting variants in breast tissue, with 80% of coding and non-coding expressed genes displaying DAE (daeGenes). We identified over 23K daeQTLs for 2753 (16%) daeGenes, including at 154 known BC risk-loci. And in 31 of these risk-loci, we found risk-associated variant(s) and daeQTLs in strong linkage disequilibrium suggesting that the risk-causing variants are cis-regulatory, and in 27 risk-loci we propose 37 candidate target genes. As validation, we identified five candidate causal variants at the 5q14.1 risk-locus targeting the ATG10, RPS23, and ATP6AP1L genes, likely via modulation of miRNA binding, alternative transcription, and transcription factor binding. Conclusion Our study shows the power of DAE analysis and daeQTL mapping to identify causal regulatory variants and target genes at BC risk loci, including those with complex regulatory landscapes, and provides a genome-wide resource of variants associated with DAE for future functional studies.
Background: Improving lung cancer risk assessment is required because current early-detection screening criteria miss most cases. We therefore examined the utility for lung cancer risk assessment of a DNA Repair score obtained from OGG1, MPG, and APE1 blood tests. In addition, we examined the relationship between the level of DNA repair and global gene expression. Methods: We conducted a blinded case-control study with 150 non-small cell lung cancer case patients and 143 control individuals. DNA Repair activity was measured in peripheral blood mononuclear cells, and the transcriptome of nasal and bronchial cells was determined by RNA sequencing. A combined DNA Repair score was formed using logistic regression, and its correlation with disease was assessed using cross-validation; correlation of expression to DNA Repair was analyzed using Gene Ontology enrichment. Results: DNA Repair score was lower in case patients than in control individuals, regardless of the case's disease stage. Individuals at the lowest tertile of DNA Repair score had an increased risk of lung cancer compared to individuals at the highest tertile, with an odds ratio (OR) of 7.2 (95% confidence interval [CI] = 3.0 to 17.5; P < .001), and independent of smoking. Receiver operating characteristic analysis yielded an area under the curve of 0.89 (95% CI = 0.82 to 0.93). Remarkably, low DNA Repair score correlated with a broad upregulation of gene expression of immune pathways in patients but not in control individuals. Conclusions: The DNA Repair score, previously shown to be a lung cancer risk factor in the Israeli population, was validated in this independent study as a mechanism-based cancer risk biomarker and can substantially improve current lung cancer risk prediction, assisting prevention and early detection by computed tomography scanning.
The recent outbreak of the severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2), which causes coronavirus disease 2019 (COVID-19), has led to a worldwide pandemic. One week after initial symptoms develop, a subset of patients progresses to severe disease, with high mortality and limited treatment options. To design novel interventions aimed at preventing spread of the virus and reducing progression to severe disease, detailed knowledge of the cell types and regulating factors driving cellular entry is urgently needed. Here we assess the expression patterns in genes required for COVID-19 entry into cells and replication, and their regulation by genetic, epigenetic and environmental factors, throughout the respiratory tract using samples collected from the upper (nasal) and lower airways (bronchi). Matched samples from the upper and lower airways show a clear increased expression of these genes in the nose compared to the bronchi and parenchyma. Cellular deconvolution indicates a clear association of these genes with the proportion of secretory epithelial cells. Smoking status was found to increase the majority of COVID-19 related genes including ACE2 and TMPRSS2 but only in the lower airways, which was associated with a significant increase in the predicted proportion of goblet cells in bronchial samples of current smokers. Both acute and second hand smoke were found to increase ACE2 expression in the bronchus. Inhaled corticosteroids decrease ACE2 expression in the lower airways. No significant effect of genetics on ACE2 expression was observed, but a strong association of DNA- methylation with ACE2 and TMPRSS2- mRNA expression was identified in the bronchus.