BACKGROUND:Tumour-associated autoantibodies (TAAbs) are promising biomarkers for cancer detection, but their induction and clinical relevance in lung cancer remain unclear. METHODS:Serum samples from 695 individuals were analysed for TAAb profiling by protein-array screening and two-stage ELISA validation. Diagnostic models were constructed with identified TAAbs and compared with conventional tumour markers. Potential mechanisms, clinical and prognostic features of TAAb seropositivity were analysed and its presence in prediagnostic sera was evaluated to assess the potential for early detection. RESULTS:Six TAAbs for small cell lung cancer (SCLC) and four for lung adenocarcinoma (LUAD) were identified, demonstrating excellent diagnostic performance (AUC > 0.8) and outperforming ProGRP and CEA. TAAb induction correlated with antigen overexpression, somatic mutations and HLA class II amino acid polymorphisms. TAAb panel seropositivity was associated with older age and advanced stage in both subtypes, and predicted poor survival in SCLC but a favourable outcome in advanced LUAD. In prediagnostic sera, the TAAb concentration increased progressively, with detectability up to 2 years before clinical diagnosis. CONCLUSIONS:Distinct TAAb panels were identified for SCLC and LUAD, serving as accurate diagnostic markers that enable early detection and as indicators of prognosis in different clinical contexts.
Autoimmune hypothyroidism (Hashimoto's thyroiditis) is common and has a strong genetic component. Here we performed multi-ancestry genome-wide association meta-analyses encompassing 48,694 Hashimoto's thyroiditis cases, using a precise case definition, and 1,044,134 controls. We identified 155 significant (P < 5 × 10-8) independent genetic associations, of which 45 variants and 19 loci were not previously associated with hypothyroidism. Six loci were specific for individuals of European ancestry reference populations. Functional enrichment analyses of Hashimoto's thyroiditis-associated genes highlighted immune cells and the spleen, underpinning the importance of T cells in Hashimoto's thyroiditis development. This observation was further supported by 161 significant colocalizations with expression quantitative trait loci in immune cells and 40 in thyroid tissue (for example, TG, VAV3, IRF5), highlighting the interplay between the immune system and the thyroid. Mendelian randomization indicated causal effects of Hashimoto's thyroiditis on cardiovascular traits and expected associations with thyroid hormone levels.
Oculopharyngodistal myopathy (OPDM) is a hereditary muscle disease caused by CGG/CCG repeat expansions in six genes. Although the clinical features are often similar, such as ptosis, dysphagia, and distal muscle weakness, the age at onset vary widely, and the mechanisms underlying this variation remain unclear. In particular, the contributions of repeat size, flanking sequence variation, and DNA methylation to phenotype have not been systematically explored using single-molecule resolution. We applied CRISPR/Cas9-targeted nanopore sequencing (nCATS) to genomic DNA from 91 individuals carrying expanded CGG repeats in three OPDM-related genes (LRP12, GIPC1, and NOTCH2NLC). This approach enabled the simultaneous analysis of CGG repeat length, flanking sequence architecture, single nucleotide variant haplotypes, structural variation, and CpG methylation profiles. Genotype–phenotype correlations were evaluated by integrating molecular and clinical data. Expanded LRP12 and GIPC1 alleles in the patients showed respective single nucleotide variant patterns around repeat regions, suggesting founder haplotypes. Repeat regions essentially comprised pure CGG expansions, but exhibited size variability, even within patients. Additionally, LRP12-expanded repeats lacked flanking nucleotide sequences present in non-expanded repeats, whereas GIPC1 expanded repeats contained specific discontinued CGG patterns in their 5'-regions. Structural variations were also identified in some patients. A significant inverse correlation was observed between repeat length and age at onset in patients with GIPC1 or NOTCH2NLC expansions, while this was disturbed by higher methylation of upstream regions in patients with LRP12 expansions, leading to delayed onset. This study highlights gene-specific differences in CGG repeat architecture and epigenetic regulation in OPDM. Founder haplotypes, expanded allele-specific flanking sequences, and the combined effects of repeat size and methylation contribute to patient regional frequency, repeat stability, and clinical variability, respectively, offering insight into disease pathomechanism and potential therapeutic targets.
The approximately 40-week gestational period is central to human reproduction, yet the genetic architecture of diverse gestational phenotypes and their links to maternal late-life health remain unclear. In 111 phenotypes from up to 121,579 Chinese pregnancies (median n = 78,535 per phenotype), we identified 4,688 independent genome-wide significant signals, including 1,703 new associations. Gestation-specific effects were observed for 7.8% of variants across 30 phenotypes; 18.7% of signals for 24 longitudinal hematological traits exhibited genotype-by-gestational-timing interactions across five antenatal and postpartum periods. Dynamic genetic effects were enriched in growth-regulatory and hormone-regulatory pathways, reflecting maternal-fetal interactions. Genetic correlation and Mendelian randomization analyses with 80 diseases and medication traits in BioBank Japan females revealed shared genetic overlaps and potential causal links between gestational phenotypes and maternal mid-life and late-life health. These results establish a dynamic genetic atlas of human gestation, providing a framework for precision maternal health.
Biological age estimates are increasingly used to study aging, disease risk, and mortality, yet their predictive uncertainty is rarely quantified. Consequently, conventional age-gap measures can treat deviations as equally informative even when the underlying biological age predictions differ substantially in reliability. We developed a framework for uncertainty-aware biological aging that generates calibrated prediction intervals and individualized probabilities of accelerated or decelerated aging alongside point estimates. We applied this framework to the UK Biobank Pharma Proteomics Project, evaluating three composite and eleven organ-specific biological age clocks. Predictive uncertainty varied substantially both within and across clocks, revealing that apparently extreme age gaps can differ markedly in the strength of evidence supporting accelerated or decelerated aging. In particular, low-accuracy clocks, including many organ-specific clocks, provided little evidence for confidently accelerated or decelerated aging. Beyond biological age gaps, prediction-interval width was independently associated with disease risk and mortality, particularly for composite, brain, and immune clocks, suggesting that predictive uncertainty captures an additional dimension of biological aging that may reflect increased molecular heterogeneity and dysregulation associated with aging and disease. We replicated these findings in Biobank Japan and an independent clinical cohort from Stanford. By incorporating individual-specific predictive uncertainty, our framework provides a more informative characterization of biological aging and enables improved individual-level risk stratification for disease prevention and longitudinal monitoring.
BACKGROUND:Moyamoya disease (MMD) has a strong genetic basis, with the rare RNF213 variant (rs112735431) representing a major risk factor, while the broader genetic architecture and disease-relevant vascular cell types remain incompletely understood. METHODS:We conducted a genome-wide association study in Japanese individuals (n=47 656; 401 MMD cases and 47 255 controls). Population-level features at MMD risk loci were examined by regional allele frequency and haplotype analyses. We performed single-nucleus RNA-seq of superficial temporal arteries from patients with MMD (n=3). Cell type-specific enrichment of genome-wide association study signals was assessed using the Single-Cell Disease Relevance Score. Endothelial signatures were validated by integration with publicly available single-cell data sets from controls (n=5) and immunohistochemistry for candidate markers (n=1). RESULTS:Beyond rs112735431, we identified a genome-wide significant signal in the HDAC9-TWIST1 region (P=3.3×10-14; odds ratio, 1.77). Conditional analysis on rs112735431 revealed a protective RNF213 missense variant, p.Asn1331Gly (rs8074015; P=3.7×10-9; odds ratio, 0.53), whose minor allele was mutually exclusive with rs112735431-A on haplotypes. Population analysis revealed geographic variation and extended haplotype structure of the rs112735431-A allele in Japan. Single-nucleus RNA-seq identified a mesenchymal-like endothelial cell (MEC) population with selective FN1 expression. Genome-wide association study-prioritized disease genes were strongly enriched in MECs. MECs showed mesenchymal pathway activation with a regulatory program distinct from canonical endothelial states. The proportion of MECs was markedly increased in MMD (72% versus 28% in controls), and FN1 expression in endothelial regions was confirmed by immunohistochemistry. CONCLUSIONS:Our findings identify a protective RNF213 variant that is mutually exclusive with the known rs112735431-A allele. Genetic risk converges on an MEC state markedly expanded in MMD.
Rare bi-allelic variation is a major contributor to human disease risk, yet its effects are difficult to study at scale in population cohorts owing to the limited number of individuals with putatively deleterious bi-allelic genotypes and the challenges of accurately phasing low-frequency variants. Here, we present recessive, gene-based analyses of rare and low-frequency variants in up to 948,690 exome- or whole-genome-sequenced individuals across six biobanks with linked electronic health records. Through statistical phasing, we inferred putatively damaging compound-heterozygous genotypes, increasing the number of bi-allelic damaging genotypes by 19%. Restricting to predicted loss-of-function (pLoF) variants, we identified 5,563 genes harboring bi-allelic genotypes, a 19.8% increase in putative knockouts. We then considered all low-frequency variants (minor allele frequency [MAF] <5%) and performed gene-based recessive association testing using putatively damaging bi-allelic genotypes, identifying 58 significant associations (false discovery rate [FDR] ≤1% or prec≤7.5 × 10-7) after meta-analysis and Cauchy combination of nonsynonymous annotations. Comparing recessive and additive models, we found 17 instances where recessive effects were more pronounced, including several previously unreported associations, such as HBB with heart failure (prec = 2.6 × 10-14; padd = 0.98), LECT2 with height (prec = 3.7 × 10-14; padd = 4.1 × 10-10), and ENSG00000267561 with height (prec = 2.9 × 10-9; padd = 0.37). This study demonstrates the potential of federated approaches to study the effects of rare bi-allelic variation.
Background: Sex hormone alterations, such as estrogen deficiency or testosterone excess, substantially increase cardiovascular disease (CVD) risk in females. Dietary fibre and its microbial by-products, short-chain fatty acids (SCFAs), have cardioprotective effects, but it remains unclear whether these benefits extend to females with an altered sex hormone profile. In this study, we aim to investigate whether dietary fibre intake, measured via plasma acetate-the most abundant SCFA-is associated with improved cardiovascular outcomes in females with altered sex hormone profiles. Methods: This cohort study included 116,235 female participants from the UK Biobank and Biobank Japan with up to 10 years of follow-up. We analysed early menopause (as a surrogate for estrogen insufficiency) and plasma free testosterone (in a subset). The primary outcome was major adverse cardiovascular events (MACE). Secondary outcomes were blood pressure. Proteomics analyses explored potential mechanisms. Results: Acetate levels were associated with lower 10-year MACE incidence (-0.618/1000 woman-year, HR=0.900, p=0.002) and systolic blood pressure (-0.231 mmHg per 1 SD, p<0.001) in the UK Biobank. High acetate levels attenuated the increased MACE risk associated with early menopause (HR=1.158, p=0.057) compared with low acetate (HR=1.425, p<0.001), with similar patterns replicated in Biobank Japan (high: HR=1.322, p=0.090; low: HR=1.385, p=0.042). Proteomics analyses suggested a mechanism involving pro-inflammatory proteins. Moreover, high acetate levels attenuated the increased MACE associated with elevated free testosterone in the UK Biobank (high: HR=1.238, p=0.024; low: HR=1.056, p=0.666). A significant interaction between acetate and free testosterone on systolic blood pressure indicated that the effect of rising testosterone on blunting acetate's effect (0.167, 95% CI: [5.212x10-2-2.818x10-1], p=0.004) was partially mediated by central obesity (waist-to-hip ratio). Conclusions: Higher plasma acetate levels were associated with lower cardiovascular risk, particularly in females with early menopause or elevated free testosterone, potentially via inflammatory pathways. These findings underscore the importance of hormonal context in shaping cardiometabolic resilience and support personalised CVD prevention strategies for females with altered sex hormone profiles, including increasing dietary fibre intake. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement F.Z.M. is supported by a Senior Medical Research Fellowship from the Sylvia and Charles Viertel Charitable Foundation, a National Heart Foundation Future Leader Fellowship (105663), and National Health & Medical Research Council (NHMRC) Emerging Leader Fellowship (GNT2017382). S.N. was supported by AMED (JP24tm0424228, JP24tm0524009, JP25kk0305032, and JP256f0137004), Takeda Science Foundation, and Japan Foundation for Applied Enzymology. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This research has been conducted using the UK Biobank Resource under Application Number 86879. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Deidentified clinical data of UK Biobank and Biobank Japan is available upon request. The R codes for free testosterone conversion and plotting are available at https://github.com/ChrYang/SexHormone.
Prioritizing causal variants in a regulatory region of the genome remains challenging. Here we introduce the Expression Modifier Score (EMS) v2, allowing prioritization of regulatory variants with high precision. EMSv2 achieves higher prediction performance compared to alternative methods, especially in tissues with low sample size such as brains. We show that the power gain is attributed to implementation of (1) features accounting for long-range DNA sequence interaction, (2) customization of loss-function in training, and (3) multi-task learning framework. We then apply EMSv2 to an independent eQTL data from a Japanese population to demonstrate that EMSv2 outperforms alternative methods in regulatory variant prioritization and can be utilized for functionally-informed fine-mapping in a distinct population. We also show that EMSv2 can be utilized in combination with the gene-level polygenic prioritization score (PoPS) to prioritize complex trait-causal regulatory variants. Our work accelerates regulatory variant prioritization in the human genome. Expression Modifier Score v2 integrates Enformer-derived features and multi-task learning to prioritize regulatory variants and support interpretation of noncoding genetic associations.
Genetic effects on gene expression are often cell type-specific and obscured in bulk analyses. To resolve this context-dependent regulation, we performed a federated cis-eQTL meta-analysis across 12 PBMC datasets (2,032 individuals, 2.5 million cells). Across six immune cell types, we identified cis-eQTLs for 6,592 genes and fine-mapped 14,985 independent loci. Notably, the 42% of eQTLs that were undetected in a bulk eQTL study on 43,301 whole blood samples also showed stronger enrichment for disease GWAS loci. We further identified three genome-wide significant and 65 suggestive loci affecting the abundance of (rare) immune cell types and validated these using previously reported hematological GWAS and bulk-derived trans-eQTLs. Integrating single-cell cis-eQTLs with bulk trans-eQTLs enabled us to anchor 6,382 trans-eGenes (37.2% novel) to upstream regulators and reconstruct directed gene regulatory relationships. For example, a hemorrhoidal disease-associated variant showed a CD4+ T cell-specific cis-eQTL on BACH1 that colocalized with 45 immune and metabolic trans-eGenes. These results demonstrate the power of single-cell QTL meta-analysis in interpreting complex trait genetics.
Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system characterized by demyelination disseminated in space and time. Here we performed a genome-wide association study (GWAS) using 688 MS cases and 205,199 controls from the Japanese population and identified significant associations in the major histocompatibility complex region and a population-specific risk variant in 11q24. Through cross-population GWAS meta-analyses using a total of 29,374 cases and 1,843,563 controls from 4 ancestral populations, we identified 22 novel susceptibility loci. Integration of GWAS and single-cell and single-nucleus RNA sequencing of peripheral blood mononuclear cells and subcortical lesions from patients with MS revealed enrichment of genetic risk factors for MS in CD4+ T helper cell lineage and regulatory T cells, as well as in endothelial cells. Furthermore, spatial transcriptomics of subcortical lesions demonstrated spatial and temporal heterogeneity in associations with MS genetic risk. Our study demonstrates the value of investigation of spatiocellular features of disease genetics across diverse populations and omics modalities.
Thyroid diseases are common and highly heritable. We performed a meta-analysis of genome-wide association studies from 19 biobanks for five thyroid diseases: thyroid cancer (ThC), benign nodular goiter, Graves’ disease, lymphocytic thyroiditis and primary hypothyroidism. We analyzed genetic association data from ~2.9 million genomes and identified 313 known and 570 new independent loci linked to thyroid diseases. We discovered genetic correlations between ThC, benign nodular goiter and autoimmune thyroid diseases ( rg = 0.16–0.97). Telomere maintenance genes contributed to benign and malignant thyroid nodular disease risk, whereas cell cycle, DNA repair and damage response genes were associated with ThC. We propose a paradigm that explains genetic predisposition to benign and malignant thyroid nodules. We found polygenic risk score associations with ThC risk of structural disease recurrence, tumor size, multifocality, lymph node metastases and extranodal extension. Polygenic risk scores identified individuals with aggressive ThC in a biobank, creating an opportunity for genetically informed population screening.
Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics1-5. Most of these variants individually have weak effects6 and lie in non-coding gene-regulatory elements7-10, for which we lack a complete understanding of how single-nucleotide alterations modulate transcriptional processes to affect human phenotypes5,11-15. To address this problem, we measured the activity of 221,412 fine-mapped trait-associated variants using a massively parallel reporter assay16-20 in 5 diverse cell types. We show that this assay effectively discriminates between likely causal variants and controls, and identified 13,121 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% of them can plausibly be explained by the disruption of a known transcription factor binding motif. We investigated the mechanisms of 136 variants using saturation mutagenesis and assigned affected transcription factors for 91% of variants without a clear canonical mechanism. Finally, we detected regulatory epistasis at 11% of tested regulatory variants in close proximity and identified multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants that underlie complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.
Many non-coding variants influence complex traits and diseases through gene regulation, yet the mechanisms linking these variants to downstream biology remain poorly understood. Here, we present eQTLGen Phase 2, a comprehensive genome-wide analysis of gene expression quantitative trait loci (eQTLs) in 43,301 blood samples from 52 datasets. Beyond local ciseffects, this sample size enabled the first systematic mapping of trans-eQTLs at scale. We identify cis-eQTLs for nearly all expressed genes (94.7%) and trans-eQTLs for over half (56.2%). Second, by colocalizing cis-eQTLs with trans-eQTLs, we infer a directed gene regulatory network comprising 47,554 directed gene regulatory relationships. These networks reveal how genetic perturbations in upstream regulators produce dose-dependent downstream effects, supported by Perturb-seq and ChIP-seq data. Third, integrating this network with 87 genome-wide association studies allows us to systematically prioritize trait-relevant pathways and candidate genes. Variants exerting both cis- and trans-effects are markedly more likely to colocalize with trait associations than cis-only variants, delineating a subset of functionally active cis-eQTLs from a large group with limited downstream impact. This distinction provides a conceptual framework for identifying regulatory variants that truly mediate complex trait biology. Together, these results provide a publicly available resource of cis- and trans-eQTLs and an in vivo scaffold for human gene-regulatory networks, elucidating how propagation of cis-effects modulates complex disease.
Environmental differences in genetic effect sizes, namely, gene-environment interactions, may uncover the genetic encoding of phenotypic plasticity1-3. We provide a cross-population atlas of gene-environment interactions comprising 440,210 individuals from European and Japanese populations, with replication in 539,794 individuals from diverse populations. By decomposing the contributions from age, sex and lifestyles, we delineate the aetiology of these gene-environment interactions, including a reverse-causality from a disease-related dietary change. Genome-wide analyses uncovered missing heritability and trait-trait relationships connected by the synergistic effects of genome and environments, which systematically affected polygenic prediction accuracy and cross-population portability. Single-cell projection revealed aging shift of pathways and cell types responsible for genetic regulation. Omics-level gene-environment analyses identified multiple sex-discordant genetic effects in lipid metabolism, informing clinical trial failures for genetically supported drug development. Our comprehensive gene-environment study decodes the dynamics of genetic associations, offering insights into complex trait biology, personalized medicine and drug development.
The gut microbiome has emerged as an important environmental factor in the pathogenesis of autoimmune diseases. Advances in high-throughput sequencing technologies have enabled comprehensive characterization of the gut microbiome, providing detailed insights into its composition and functional potential. These approaches have been widely applied in autoimmune disease research, revealing disease-associated alterations in the gut microbiome of patients with conditions such as rheumatoid arthritis and systemic lupus erythematosus. In addition, microbiome sequencing data can be leveraged to investigate the gut virome, including viruses residing in the intestinal ecosystem. This review summarizes current evidence linking autoimmune diseases and the gut microbiome, with a particular focus on studies employing microbiome sequencing-based analyses.