Urinary tract infections (UTIs) are traditionally viewed as environmentally driven, yet their inherited susceptibility remains largely unexplored. We conducted a cross-biobank genome-wide association study of recurrent UTIs in 1,860,836 individuals (213,869 cases and 1,646,967 controls). We identified 36 genetic susceptibility loci and performed tissue-based multi-omic mapping to prioritize candidate causal genes. UTI risk alleles preferentially modulated epithelial gene expression in kidney and bladder, converging on urinary epithelia structure and function. PSCA, encoding a secreted epithelial surface protein, emerged as the strongest candidate under genetic control; the gene product is constitutively secreted into the urine from kidney papilla and bladder epithelia, binds uropathogenic E. coli, and inhibits bacterial growth in vitro. Our findings define the polygenic architecture of UTIs and highlight the critical role of uroepithelial surface defenses, providing a new framework for host-directed, non-antibiotic interventions.
Biological age estimates are increasingly used to study aging, disease risk, and mortality, yet their predictive uncertainty is rarely quantified. Consequently, conventional age-gap measures can treat deviations as equally informative even when the underlying biological age predictions differ substantially in reliability. We developed a framework for uncertainty-aware biological aging that generates calibrated prediction intervals and individualized probabilities of accelerated or decelerated aging alongside point estimates. We applied this framework to the UK Biobank Pharma Proteomics Project, evaluating three composite and eleven organ-specific biological age clocks. Predictive uncertainty varied substantially both within and across clocks, revealing that apparently extreme age gaps can differ markedly in the strength of evidence supporting accelerated or decelerated aging. In particular, low-accuracy clocks, including many organ-specific clocks, provided little evidence for confidently accelerated or decelerated aging. Beyond biological age gaps, prediction-interval width was independently associated with disease risk and mortality, particularly for composite, brain, and immune clocks, suggesting that predictive uncertainty captures an additional dimension of biological aging that may reflect increased molecular heterogeneity and dysregulation associated with aging and disease. We replicated these findings in Biobank Japan and an independent clinical cohort from Stanford. By incorporating individual-specific predictive uncertainty, our framework provides a more informative characterization of biological aging and enables improved individual-level risk stratification for disease prevention and longitudinal monitoring.
IgA nephropathy (IgAN), IgA vasculitis (IgAV), focal segmental glomerulosclerosis (FSGS), membranous nephropathy (MN), and minimal change disease (MCD) account for the majority of idiopathic glomerulo-nephropathies (GN). These disorders involve immune system dysregulation and have a complex genetic architecture. Currently, there are no adequately powered blood transcriptomic datasets coupled to genetic data from patients with GN that can delineate disease-context specific genetic effects on the blood immune cell transcriptome. We performed whole genome sequencing coupled with bulk blood transcriptome sequencing on 1,822 participants from the CureGN study, a prospective cohort of participants with a kidney biopsy diagnosis of primary GN. We generated disease-context specific transcriptome-wide maps of gene expression QTL (eQTL), splicing QTL (sQTL), and double strand RNA-editing QTL (edQTL) for FSGS (N=447), IgAN (N=403), IgAV (N=123), MCD (N=408), and MN (N=441), as well as cross-disease maps for all 1,822 participants. Our QTL mapping identified 16,068 eGenes, 4,644 sGenes and 4,611 edQTLs with an FDR<0.05 in at least one GN type. Approximately 5-10% of the QTL signals were unique to a specific GN type, while ~90% were shared between at least two conditions. Colocalization analysis demonstrated that ~80% of shared eGenes between traits also shared the same causal variants, whereas ~2% had distinct causal variants, suggesting context-specific regulatory effects. Cross-phenotype QTL mapping uncovered 6,466 eGenes, 2,705 sGenes and 5,321 edQTLs not previously detected in GTEx. Age, eGFR, and proteinuria-interaction QTL analyses identified hundreds of loci modified by age or disease severity. Lastly, integrative analyses with GWAS nominated new candidate genes for each of the five GN types under study. In summary, we generated comprehensive maps of GN-context-specific genetic effects on the blood transcriptome, providing a powerful new resource for integrative gene discovery studies of primary GN.
Production livestock provide a natural system for studying gene regulation under physiologically demanding conditions shaped by rapid growth, environmental exposure, and immune challenges. Using farm pigs from the PigGTEx resource, we apply quantile regression to reveal context-dependent genetic effects on gene expression across tissues. We identify quantile-specific eQTLs missed by standard linear models that preferentially localize to distal regulatory elements and three-dimensional genome architecture, in contrast to the promoter-proximal bias of canonical eQTLs. Genes with quantile-dependent eQTLs are more intolerant to loss-of-function variants and exhibit distinct patterns of enrichment in GO functional categories, indicating their likely functional significance. Cross-species comparisons reveal substantial overlap between pig and human eGenes across tissues, indicating conservation of regulatory architecture. Notably, many quantile-specific eQTLs influence tail expression states and involve genes relevant to human disease. For example, we identify a cis-eQTL affecting the conserved transcriptional regulator BCL6B in pig blood that modulates enhancer activity and reduces expression at lower quantiles. In contrast, BCL6B is minimally expressed in resting human blood and lacks detectable cis-regulatory variation under baseline conditions, consistent with its reported induction during immune activation. These findings demonstrate that pig eQTL maps can reveal context-dependent regulatory variation at loci that remain silent or weakly variable in human cohorts.
OBJECTIVE:Anti-β2-glycoprotein I (anti-β2GPI) antibodies are central to antiphospholipid syndrome. The APOH locus has been associated with anti-β2GPI, but the causal variant and thrombotic risk implications remain unclear. METHODS:We performed a multi-ancestry genome-wide association study (GWAS) of quantitative total anti-β2GPI antibody levels in 5,969 participants from the Multi-Ethnic Study of Atherosclerosis (MESA). The relationship between genetically determined anti-β2GPI levels and venous thromboembolism (VTE) risk was evaluated using two-sample Mendelian randomization (MR) and Bayesian colocalization. Candidate causal variants were prioritized by fine-mapping and integrative functional genomic analyses. Molecular dynamics simulations were used to investigate the structural effects. RESULTS:We identified a genome-wide significant association at the APOH locus (lead SNP rs1801690-G, β = 0.21, p = 1.08 × 10-10). Paradoxically, genetic variation at the APOH locus associated with higher anti-β2GPI levels was associated with reduced VTE risk (β = -0.25, p = 3.95 × 10-6; PP4 = 0.97), likely reflecting horizontal pleiotropy rather than a true protective antibody effect. Fine-mapping and functional genomics prioritized the missense variant rs1801690 (W335S) as the most likely causal variant. Molecular dynamics simulations support a dual-effect mechanism whereby W335S impairs phospholipid binding in Domain V while increasing epitope exposure in Domains I-II. CONCLUSIONS:The APOH genetic signal associated with higher anti-β2GPI levels paradoxically reduced VTE risk, likely mediated by W335S via epitope accessibility and disrupted phospholipid binding. These findings provide human genetic evidence that phospholipid binding by DV is essential for anti-β2GPI-mediated thrombosis. The clinical utility of APOH genotype warrants further investigation.
Anti-β2-glycoprotein I (anti-β2GPI) antibodies are central to the pathogenesis of antiphospholipid syndrome (APS), an autoimmune disease characterized by a strong predisposition to venous thromboembolism (VTE). In this study, we conducted a multi-ancestry genome-wide association study (GWAS) of quantitative total anti-β2GPI levels in 5,969 participants enrolled in the Multi-Ethnic Study of Atherosclerosis (MESA) and identified a genome-wide significant association at the APOH locus. Paradoxically, genetically determined increases in anti-β2GPI levels at this locus were associated with lower VTE risk. Fine-mapping and functional genomics prioritized the missense variant rs1801690 (W335S) in β2GPI (apolipoprotein H, [APOH]) as the most likely causal variant. This variant has an allele frequency of 5-6% in European and East Asian ancestries but only 1% in African ancestries. Integrating prior experimental studies, molecular dynamics simulations and structure-based epitope prediction, we propose a dual-effect mechanism whereby W335S reduces thrombotic risk by disrupting phospholipid binding in Domain V, yet increases autoantibody production through conformational changes that enhance epitope exposure in Domains I and II. These findings mechanistically uncouple autoantibody formation from thrombotic risk in carriers of the W335S variant, and suggest that APOH genotype may represent a clinically relevant genetic biomarker with potential utility for thrombotic risk stratification in anti-β2GPI-positive individuals.
Background:Idiopathic pulmonary fibrosis (IPF) and telomere length (TL) are both strongly linked to rare and common genetic variation. Shortened TL itself may be causal for IPF. Whether rare and common variants compete or cooperate to confer genetic risk of IPF uniformly is unknown. Methods:We used whole genome sequencing (WGS) data from a discovery case-control cohort sequenced at Columbia (777 IPF, 2905 controls) and validated findings using WGS data from Trans-Omics for Precision Medicine (TOPMed, 1148 IPF, 5202 controls) and the UK Biobank (UKBB, 2739 IPF, 395331 controls). In all cohorts, we identified rare damaging variants in disease-associated genes and computed control-normalized polygenic risk scores for IPF (IPF-PRS) and telomere length (TL-PRS). Telomere length of blood leukocytes was measured using a qPCR assay for two cohorts. We determined the association of the MUC5B rs35705950 polymorphism, an IPF-PRS excluding MUC5B (IPF-PRS-noMUC5B), and a TL-PRS with IPF risk in the overall cohort and in subgroups stratified by genetic endotypes (rare variant carriers, non-carriers stratified by TL cutoffs). We calculated cross-validated area under the receiver operator curve (AUC) and compared the liability of IPF explained by genetic variables. Findings:We identified independent associations between IPF risk and rare variants, the MUC5B SNP, and both polygenic scores in the discovery cohort and replicated these findings in the TOPMed and UKBB cohorts. The adjusted effect size of the TL-PRS, which includes >180 SNPs not previously associated with IPF, was comparable to the IPF-PRS-noMUC5B in the discovery (ORTL-PRS 1.63 [95% CI 1.47, 1.81] vs. ORIPF-PRS 1.60 [1.44, 1.77]) and replication cohorts (TOPMed ORTL-PRS 1.47 [1.36, 1.59] vs. ORIPF-PRS 1.37 [1.25, 1.50]; UKBB ORTL-PRS 1.24 [1.19, 1.29] vs. ORIPF-PRS 1.25 [1.21, 1.30]). The TL-PRS incrementally improved disease prediction beyond known IPF common and rare genetic predictors and clinical variables in discovery (combined AUC: 0.89, pDelong = 0.006), TOPMed (combined AUC: 0.89, pDelong = 0.01), and UKBB cohorts (combined AUC: 0.77, pDelong = 0.03). Rare and common variants jointly contributed to genetic liability of IPF. The TL-PRS increased liability of IPF explained by 13% in the discovery cohort and 8% and 13% in the TOPMed and UKBB cohorts, respectively. In IPF subjects with damaging rare variants, the TL-PRS was consistently associated with disease risk whereas the IPF-PRS-noMUC5B was not. The TL-PRS also conferred nominally greater odds of disease risk than the IPF-PRS-noMUC5B in patients with shorter TL, in the discovery and UKBB cohorts. Together, 23-43% of IPF cases have damaging rare variants or telomeres <10th percentile, where the TL-PRS represents a major unrecognized genetic risk factor. Interpretation:Common and rare genetic variation confer context-specific genetic risk in IPF competitively and cooperatively. In contrast to known IPF common risk variants, the TL-PRS, which includes >180 genetic loci not previously associated with IPF, increases the risk of disease specifically in certain IPF endotypes. Polygenic risk from telomere-associated common variants is a key feature of IPF genetic heterogeneity. Funding:National Institutes of Health (NIH), Medical Research Council (MRC), National Institute for Health and Care Research (NIHR).
Genome-wide association studies (GWASs) in ancestrally diverse populations are rapidly expanding, opening up unique opportunities for novel gene discoveries and increased utility of genetic findings in non-European individuals. A popular technique to identify putative causal variants at GWAS loci is via statistical fine-mapping. Despite tremendous efforts, fine-mapping remains a very challenging task, even in the relatively simple scenario of studies with a single, homogeneous population. For studies with admixed individuals, such as within Latin America and the Caribbean, methods for gene discovery are still limited. Here, we propose a Bayesian model for fine-mapping in admixed populations, CARMA-X, that addresses some of the unique challenges of admixed individuals. The proposed method includes an estimation method for the linkage disequilibrium (LD) matrix that accounts for small reference panels for admixed individuals, heterogeneity across populations and cross-ancestry LD, and a Bayesian hypothesis test that leads to robust fine-mapping when relying on external reference panels of modest size for LD estimation. Using simulations, we compare performance with recently proposed fine-mapping methods for multi-ancestry studies and show that the proposed model provides higher power while controlling false discoveries, especially when using an out-of-sample LD matrix. We further illustrate our approach through applications to two Latin American genetic studies, the Estudio Familiar de Influencia Genética en Alzheimer (EFIGA) study in the Dominican Republic and the Mexican Biobank, where we show the benefit of modeling ancestry-specific effects by prioritizing putative causal variants and genes, including several findings driven by ancestry-specific effects in the African and Native American ancestries.
Polygenic risk scores (PRS) are widely used in post-GWAS analyses to predict complex traits across humans, animals, and plants. While significant progress has been made in developing new PRS methods, much less attention has been given to quantifying the uncertainty associated with these predictions. In this work, we propose a method for individualized uncertainty quantification based on quantile regression. When paired with conformal prediction, this approach enables the construction of prediction intervals with guaranteed coverage, offering lower and upper bounds within which the phenotype is likely to fall with high probability. We apply this framework to data from the UK Biobank and the ProgeNIA/SardiNIA studies, showing that the resulting prediction intervals: (1) maintain valid coverage under minimal model assumptions, (2) provide more realistic individualized estimates of uncertainty by allowing for asymmetry and individual-specific interval lengths, and (3) exhibit reduced uncertainty compared to existing methods. Overall, we present a novel framework for individualized uncertainty quantification in PRS analyses and highlight the importance of incorporating uncertainty into predictive modeling. ### Competing Interest Statement The authors have declared no competing interest.
LiDAR and photogrammetry are active and passive remote sensing techniques for point cloud acquisition, respectively, offering complementary advantages and heterogeneous. Due to the fundamental differences in sensing mechanisms, spatial distributions and coordinate systems, their point clouds exhibit significant discrepancies in density, precision, noise, and overlap. Coupled with the lack of ground truth for large-scale scenes, integrating the heterogeneous point clouds is a highly challenging task. This paper proposes a self-supervised registration network based on a masked autoencoder, focusing on heterogeneous LiDAR and photogrammetric point clouds. At its core, the method introduces a multi-scale masked training strategy to extract robust features from heterogeneous point clouds under self-supervision. To further enhance registration performance, a rotation-translation embedding module is designed to effectively capture the key features essential for accurate rigid transformations. Building upon the robust representations, a transformer-based architecture seamlessly integrates local and global features, fostering precise alignment across diverse point cloud datasets. The proposed method demonstrates strong feature extraction capabilities for both LiDAR and photogrammetric point clouds, addressing the challenges of acquiring ground truth at the scene level. Experiments conducted on two real-world datasets validate the effectiveness of the proposed method in solving heterogeneous point cloud registration problems.
Electronic health records have been increasingly adopted as useful resources for genomic research. However, case-control labeling of clinical data from electronic health records is challenging and most studies utilize phenotype codes to define case/control labels, resulting in suboptimal downstream analyses. Here we describe the liability threshold phenotypic integration, a method combining genetic relatedness with phenotypic data, including binary and continuous traits such as diagnosis codes, family disease history, laboratory measurements and biomarkers, to derive new continuous phenotypes for target diseases. The model utilizes an automatic trait selection algorithm that increases performance in disease risk prediction and provides insights into nontarget traits associated with the target disease. Our simulations and applications to the eMERGE network and the UK Biobank data demonstrate consistent performance gains in disease risk prediction and genome-wide association study power compared to conventional phenotype codes, models that solely incorporate family history and the phenotype imputation method SoftImpute, with similar false-positive rate control.
Key PointsFamily history of kidney disease was not associated with kidney disease progression in the context of established CKD.Family history of diabetes was a risk factor of CKD progression independently of diabetes status, polygenic risk, and traditional risk factors.BackgroundA family history of health conditions may reflect shared genetic and/or environmental risk. It is not well known to what extent family history affects outcomes among patients with CKD. In this study, we investigated the associations of family history of CKD, diabetes, and other conditions with common comorbidities and kidney disease progression among patients with CKD.MethodsWe performed an observational study of two prospective CKD cohorts, 2573 adults and children from the Cure Glomerulopathy Network and 3939 Chronic Renal Insufficiency Cohort adult participants. Self-reported first-degree family history of CKD, diabetes, and other common diseases was tested for associations with the risk of comorbidities and CKD progression using multivariable models.ResultsFamily history of common comorbid conditions was associated with higher risk of these conditions in the context of CKD, including approximately by over three-fold for diabetes (adjusted odds ratio [OR], 3.37; 95% confidence interval [CI], 2.73 to 4.15), 48% for cancer (adjusted OR, 1.48; 95% CI, 1.05 to 2.09), and 69% for cardiovascular disease (adjusted OR, 1.69; 95% CI, 1.36 to 2.10 in combined cohorts). While polygenic risk score (PRS) for CKD was associated with kidney disease progression (adjusted hazards ratio, 1.11; 95% CI, 1.06 to 1.16 in combined cohorts), family history of kidney disease was not an independent risk factor of disease progression in the context of existing CKD. By contrast, family history of diabetes was significantly associated with a higher risk of CKD progression independently of diabetes occurrence or PRS for diabetes (adjusted hazards ratio, 1.19; 95% CI, 1.05 to 1.35 in combined cohorts).ConclusionsBroad collection of family history in the context of CKD improved clinical risk stratification. Family history of diabetes was consistently associated with a higher risk of CKD progression independently of diabetes status or PRS for diabetes in both cohorts.
Background: Colorectal cancer (CRC) is a leading cause of cancer death, and the incidence and mortality rates among young adults are rising. Although a subset of CRC cases presents with a family history, suggesting a hereditary component, the specific genetic underpinnings remain incompletely understood, particularly in early-onset CRC (EOCRC). This study aimed to discover novel risk genes for EOCRC using exome sequencing and gene-based rare variant burden testing. Methods: Our cohort consisted of 212 European-ancestry cases (174 diagnosed with CRC and 38 with significant polyps) from the South Australian Young Onset Colorectal Polyp and Cancer Study (SAYO) and 31,699 unaffected controls from the Simons Foundation Powering Autism Research for Knowledge (SPARK) cohort. After filtering for ancestry, relatedness, variant quality, and population allele frequency, we performed gene-set and individual-gene burden tests using predicted deleterious missense and loss-of-function variants. Statistical significance was assessed using permutation-corrected binomial testing. An independent validation was conducted in the UK Biobank. Results: Loss-of-function variants in known CRC tumor suppressor genes were significantly enriched in SAYO cases. Gene-level analyses identified MEIKIN as a novel EOCRC susceptibility candidate (p value = 1.0 × 10−7), with supporting enrichment of deleterious missense and loss-of-function variants in distal colon cancer cases from the UK Biobank. Additional genes (STK25, PGBD4, DIRAS3, ATG3, RPS6KA4, and DDX42) demonstrated borderline significance, implicating pathways related to kinetochore assembly, autophagy regulation, and immune signaling. Both predicted gain-of-function and loss-of-function variants contributed to the EOCRC risk, supporting heterogeneous mechanisms of CRC pathogenesis. Conclusions: This study identified novel candidate risk genes for EOCRC, underscoring the role of rare variants and expanding our understanding of the genetic architecture of CRC. Future studies should include functional validation and replication studies on other ancestries to confirm and extend these results.
BACKGROUND:Understanding the genetic basis of human diseases has become integral to drug development and precision medicine. Recent advancements have enabled the identification of molecular pathways driving diseases, leading to targeted treatment strategies. The increasing investment in rare diseases by the biotech industry underscores the importance of genetic evidence in drug discovery and approval processes. Here we studied a monogenic Mendelian kidney disease, TRPC6-associated podocytopathy (TRPC6-AP), to present its natural history, genetic spectrum, and clinicopathological associations in a large cohort of patients with causal variants in TRPC6, in order to help define the specific features of disease and further facilitate drug development and clinical trials design. METHODS:the study involved 64 individuals from 39 families with TRPC6 causal missense variants. Clinical data, including age of onset, laboratory results, response to treatment, kidney biopsy findings, and genetic information, were collected from multiple centers nationally and internationally. Exome or targeted sequencing was performed and variant classification was based on strict criteria. Structural and functional analyses of TRPC6 variants were conducted to understand their impact on protein function. In depth re-analysis of light and electron microscopy specimens for 9 available kidney biopsies was conducted to identify pathological features and correlates of TRPC6-AP. RESULTS:Large-scale sequencing data did not support causality for TRPC6 protein-truncating variants. We identified 21 unique TRPC6 missense variants, clustering in three distinct regions of the protein, and with different effects on TRPC6 3D protein structure. Kidney biopsy analysis revealed FSGS patterns of injury in most cases, along with distinctive podocyte features including diffuse foot process effacement and swollen cell bodies. The majority of patients presented in adolescence or early adulthood but with ample variation (average 22, SD ± 14 years), with frequent progression to kidney failure but with variability in time between presentation and ESKD. CONCLUSIONS:This study provides insights into the genetic spectrum, clinicopathological associations, and natural history of TRPC6-AP.
Traditional statistical anomaly detection methods struggle with the high-dimensional and highly dynamic characteristics. With the advancement in computing and monitoring capabilities, machine learning-based anomaly detection methods for dynamic system equipment have become a major research focus. However, due to the infinite number of dynamic operating states inherent in dynamic equipment, it is difficult to enumerate them all. Therefore, this paper introduces a dynamic feature extraction method based on a deep temporal model with an encoder-decoder structure. This method accumulates long-term historical dynamic features and constructs a rule engine for online anomaly detection based on inference engine-driven rule extraction, aiming to cover the system's dynamic operating space as comprehensively as possible. Taking a typical dynamic thermal system, gas turbine, as an example, the paper performs feature extraction, feature library construction, and anomaly rule extraction on the training dataset. The proposed anomaly detection method based on feature distance are verified. With the accumulation of time and personnel calibration, rule extraction can be further quantified, and there is potential to cluster and classify dynamic operating conditions within the feature space.