It is essential for understanding neural network decisions to interpret the functionality (also known as concepts) of neurons. Existing approaches describe neuron concepts by generating natural language descriptions, thereby advancing the understanding of the neural network's decision-making mechanism. However, these approaches assume that each neuron has well-defined functions and provides discriminative features for neural network decision-making. In fact, some neurons may be redundant or may offer misleading concepts. Thus, the descriptions for such neurons may cause misinterpretations of the factors driving the neural network’s decisions. To address the issue, we introduce a verification of neuron functions, which checks whether the generated concept highly activates the corresponding neuron. Furthermore, we propose a Select–Hypothesize–Verify framework for interpreting neuron functionality. This framework consists of: 1) selecting activation samples that best capture a neuron’s well-defined functional behavior through activation-distribution analysis; 2) forming hypotheses about concepts for the selected neurons; and 3) verifying whether the generated concepts accurately reflect the functionality of the neuron. Extensive experiments show that our method produces more accurate neuron concepts. Our generated concepts activate the corresponding neurons with a probability approximately 1.5 times that of the current state-of-the-art method.
Spread through air spaces (STAS) is a characteristic invasive pattern of lung adenocarcinoma (LUAD), which is associated with a high recurrence rate and poor prognosis. This research introduced the automatic deep learning multimodal detection network (DAFNet) for predicting STAS based on preoperative whole-lung CT scans. In contrast to conventional approaches necessitating manual tumor delineation, DAFNet performs comprehensive end-to-end analysis of pulmonary imaging data through integrated multimodal data fusion and multiscale feature extraction methodologies. A retrospective analysis was performed on 1164 patients with LUAD (511 STAS-positive and 653 negative) from two centers, with a training-to-test split ratio of 70:30 (814 in training set and 350 in test set). In the test set, DAFNet demonstrated an area under the receiver operating characteristic curve (AUROC) of 0.90 (95% confidence interval 0.86-0.94), significantly outperforming predictive models utilizing clinical examination numerical data alone (AUROC 0.72) and radiomics features independently (AUROC 0.65). The implementation of an adaptive gate fusion mechanism combined with a DINOv3-based pre-trained architecture substantially improved predictive accuracy for STAS status determination. These findings establish DAFNet as a promising fully automated, non-invasive diagnostic tool for preoperative STAS prediction in lung adenocarcinoma, thereby advancing personalized surgical planning and promoting AI-driven oncological applications through enhanced clinical translation potential.
Plasma proteomics can provide a dynamic molecular readout of human health, but models that learn generalizable protein-expression patterns in population cohorts remain limited. Here we show that ProLM, a BERT-based plasma proteomics model pretrained on 15,499 relatively healthy UK Biobank participants, captures baseline protein-expression relationships and supports prediction of 16 common chronic diseases. After disease-specific fine-tuning, the ProLM-derived proteomic risk score outperformed the Age+Sex model for all 16 diseases, the cardiovascular disease (ASCVD) risk equation for 14 diseases and a 35-variable clinical PANEL score for 11 diseases. Model interpretation highlighted proteins including GDF15 whose expression changed more than 15 years before clinical diagnosis, and key findings were externally evaluated in the China Kadoorie Biobank. These results support plasma proteomics pretrained models as tools for early chronic-disease risk stratification, while prospective validation is needed before clinical implementation.
Aging reshapes global disease burdens, yet the regulatory roles of long non-coding RNAs (lncRNAs) in age-related disorders remain incompletely characterized. We developed iLDA-SGCN, a graph-based computational framework that integrates singular value decomposition (SVD) with dual graph convolutional networks (GCNs) to predict lncRNA-disease associations. SVD first derives compact low-dimensional representations from the lncRNA-disease association matrix. Two complementary GCN modules then learn topology-aware embeddings: a correlation-map GCN operating on the bipartite lncRNA-disease network, and a similarity-map GCN operating on fused homogeneous graphs of lncRNAs and diseases constructed from MeSH semantic similarity and Gaussian association-profile kernels. Finally, association scores are estimated with an inner-product decoder optimized with a class-imbalance-aware loss function. Across five-fold cross-validation on LncRNADisease and MNDR datasets, iLDA-SGCN outperformed five competitive methods (SDLDA, LDNFSGB, IPCARF, LDASR, and LDA-VGHB) in terms of AUC (area under the ROC curve) and AUPR (area under the precision-recall curve). The model achieved AUC/AUPR of 0.960/0.968 on MNDR and 0.896/0.901 on LncRNADisease, with only a marginal precision shortfall versus LDA-VGHB on LncRNADisease. Ablation studies showed both GCN modules improved over a fully connected backbone, with the similarity-map GCN contributing the largest gains; the full model performed best overall. In case studies across eight prototypical age-related diseases, iLDA-SGCN identified HOTAIR, MALAT1, PVT1, MEG3, H19, LSINCT5, UCA1, and other candidates, yielding 33 candidates potentially involved in age-related disease mechanisms that require further experimental validation. Collectively, iLDA-SGCN integrates semantic and topological information to prioritize candidate lncRNA-disease associations related to aging, providing testable hypotheses for downstream mechanistic studies.
Background:Lower educational attainment is associated with obesity and type 2 diabetes (T2D), but prospective evidence, external validation, mediator patterns, and genetic triangulation have rarely been integrated. Methods:UK Biobank analyses included 501,932 participants at baseline and 474,659 participants without baseline diabetes prospectively. Logistic and Cox models estimated associations of higher educational attainment with prevalent and incident diabetes. T2D pathway analyses evaluated adiposity, health behaviors, and cardiometabolic biomarkers. NHANES 2011-2018 provided weighted cross-sectional validation. Public GWAS summary statistics were used for genetic correlation, Mendelian randomization (MR), tissue enrichment, candidate-gene analyses, shared-locus analyses, and S-PrediXcan transcriptome-wide association analyses (TWAS). Results:In fully adjusted UK Biobank models, higher educational attainment was associated with lower odds of prevalent T2D (OR 0.76, 95% CI 0.72-0.79) and lower risk of incident T2D (HR 0.70, 95% CI 0.68-0.72). Detailed qualification categories showed lower prevalent T2D odds for college/university degree versus no qualifications (OR 0.69, 95% CI 0.65-0.72). In NHANES, college graduation or above was associated with lower prevalent T2D odds (OR 0.69, 95% CI 0.61-0.78; n=20,502). Adiposity, smoking, alcohol use, lipids, blood pressure, and C-reactive protein attenuated the education-T2D association. Genetically predicted educational attainment was inversely associated with BMI and T2D, and BMI/T2D brain TWAS identified 19 shared multi-SNP gene signals, including NPC1, HSD17B12, MAP2K5, DHX36, BHMT, and LEPROT. Conclusions:Higher educational attainment was consistently associated with lower T2D risk across UK Biobank, NHANES, and genetic analyses. The results highlight modifiable metabolic and behavioral pathways relevant to T2D prevention.
Enclosure is a widely applied strategy for grassland degradation mitigation and ecological restoration. However, the effects of long-term enclosure on nutrient distribution and microbial community assembly within soil aggregate micro-habitats of alpine grasslands remain poorly understood. This study compared and analyzed the nutrient characteristics, microbial community composition, co-occurrence networks, and assembly processes of soil aggregates of different particle sizes inside and outside the enclosure based on a 37-year enclosure experiment. The results showed that long-term enclosure altered the distribution of nutrients among aggregates, with small macroaggregates (SMA) being identified as the most responsive fraction. Enclosure increased the relative contribution of stochastic processes to community assembly; fungal communities were consistently dominated by drift, whereas bacterial dispersal limitation was observed only in large macroaggregates (LMA). Enclosure decreased network complexity and modified the composition of the major microbial taxa. The dominant bacterial groups shifted from Acidobacteria to Actinobacteria, while Ascomycota dominated the fungal communities. The relative abundance of these key taxa was positively correlated with soil organic carbon and available phosphorus and negatively correlated with the β-nearest taxon index (βNTI). A partial least squares path model indicated that nutrients exerted indirect effects on community assembly processes through these keystone taxa. Owing to its pore architecture and alkaline pH, SMA serves as a central microhabitat that facilitates interactions among nutrients, keystone taxa, and assembly dynamics. This study clarifies the nutrient-driven mechanisms governing microbial community assembly within soil aggregate micro-habitats under prolonged enclosure, offering insights into the co-evolution of soil structural organization and ecosystem functioning in grassland restoration.
SARS-CoV-2 coronavirus emerged in 2019, leading to the Coronavirus disease 2019 (COVID-19). Expression of viral entry factors such as ACE2 and TMPRSS2 is higher in the testis, particularly in Sertoli cells, Leydig cells, and spermatogonia. To understand COVID-19's impact on testicular cell populations and gene expression, we analyzed testicular tissue samples from 28 COVID-19 patients and compared them with 23 non-diseased controls. COVID-19 samples showed increased immune cell infiltration, thrombosis, and reduced numbers of testicular cells. There was a significant decrease in Sertoli cells and spermatogonial stem cells (SSC) among COVID-19 patients, associated with high levels of DNA damage and apoptosis in these cell types. To explore the pathways through which the virus affects testicular function, we profiled 112,657 single-nucleus transcriptomes from the testes of 4 COVID-19 patients and 4 controls. We found that COVID-19 infection alters multiple transcriptome clusters and induces a new COVID-19-specific cluster. To confirm that these transcriptome changes are COVID-19-specific, we compared our results with those from the brains of patients with COVID-19 and influenza. We observed an average of 144 dysregulated genes unique to COVID-19, regardless of tissue type. Lastly, we examined whether the SSC phenotype seen in fatal COVID-19 cases was also present in recovered patients. Indeed, recovered patients also exhibited high DNA damage and a reduced SSC population at 3, 6, and 12 months post-infection. Additionally, embryos derived from recovered patients’ sperm showed lower fertilization rates compared to control-derived embryos and fewer live births. Overall, our findings demonstrate that SARS-CoV-2 disrupts spermatogenesis, alters the transcriptional landscape, and affects human testicular architecture and function well after the acute phase of infection, with potential long-term consequences for male fertility.
BackgroundAlzheimer’s disease (AD), Parkinson’s disease (PD) and Lewy body dementia (LBD) overlap clinically, pathologically and genetically, complicating interpretation of cross-disorder genome-wide association study (GWAS) signals.MethodsWe analysed 322,963 UK Biobank participants with bidirectional time-varying Cox models, one-year and two-year lag analyses, and competing-risk sensitivity models to quantify AD-PD clinical co-occurrence. We then analysed European-ancestry AD, PD and LBD GWAS summary statistics using linkage disequilibrium score regression (LDSC), GCTA-mtCOJO/GSMR, MAGMA, stratified LDSC, brain eQTL/mQTL SMR with HEIDI filtering, and Bayesian colocalization for selected methylation probes. Conditional loci were compared with original GWAS loci to separate shared liability from retained disorder-predominant associations.ResultsPD was associated with subsequent AD (fully adjusted HR 2.27, 95% CI 1.94–2.65; P = 6.40E-25), and AD was associated with subsequent PD (HR 3.14, 95% CI 2.56–3.85; P = 2.10E-28). Lag and competing-risk sensitivity analyses remained concordant. LDSC estimated positive genetic correlations for AD-PD (rg = 0.20; P = 0.0086) and PD-LBD (rg = 0.61; P = 0.0005). Conditioning reduced genome-wide significant loci from 14 to 9 for AD, from 24 to 21 for PD and from 5 to 2 for LBD. Retained loci included AD signals near CR1, BIN1, CLU, SPI1, MS4A, PICALM, ABCA7 and APOE; PD signals near GBA, NUCKS1, TMEM163, STK39, GAK/TMEM175, BST1, SNCA, LRRK2, MAPT and RIT2; and LBD signals near SNCA/MMRN1 and APOE. MAGMA and S-LDSC highlighted amyloid, lipid, immune, synaptic-vesicle and brain-tissue enrichment patterns. Brain QTL analyses prioritized retained eQTL and mQTL signals, and colocalization supported shared PD-GWAS/mQTL signals at HLA-DRB5, ARHGAP27, CRHR1, MAPT and KANSL1.ConclusionAD and PD show bidirectional clinical co-occurrence, whereas conditional genetic analyses retain a smaller set of disease-predominant loci and regulatory signals across AD, PD and LBD. These findings refine cross-disorder interpretation and nominate loci for independent genetic and functional validation.
Cognitive performance has been found to be associated with the complex structure of human cerebral cortex. However, due to the limitations of previous cortical parcellation atlases, the cortical genetic patterns determining cognitive performance remain unknown. Here, we utilized the latest Human Connectome Project Multi-Modal Parcellation (HCP-MMP) atlas to divide the cerebral cortex into 180 regions per hemisphere. We investigated the shared genetic architecture between four types of magnetic resonance imaging (MRI)-derived cortical phenotypes and cognitive performance using large-scale genome-wide association studies (N for cortical phenotypes = 36,843; N for cognitive performance = 257,828). We observed extensive genetic overlap between cortical surface area, volume, and local gyrification index (LGI) with cognitive performance, particularly the subregions in the insula, cingulate cortex, and ventromedial prefrontal cortex, many of which were novel findings. However, the thickness of some prefrontal regions was negatively correlated with cognitive performance. We identified 18 and 312 shared genetic loci for global and regional cortical phenotypes with cognitive performance, respectively. These genetic loci were involved in a substantial number of biological processes related to neuronal development, cell growth, and neuronal death or apoptosis. The cortical patterns defined by these shared loci were established entirely along the sensorimotor-association (S-A) axis. These findings provide new insights into the genetic relationship between cognitive performance and the human cerebral cortex under a more refined multimodal cortical parcellation scheme.
Enzyme activities are jointly regulated by microbial traits and soil physicochemical properties, but their relative importance remains unclear in alpine grasslands, where cold climate and nutrient limited. To address this, we combined metagenomic analyses of microbial functional genes with assays of potential C- and N-cycling enzyme activities across five alpine grassland types on the Tibetan Plateau. We tested whether enzyme activities are driven more by biotic factors (microbial community composition, functional gene abundance) or abiotic filters (soil properties). Seven enzymes were measured, including cellulase, beta-glucosidase, and alpha-N-acetylglucosaminidase, and their associations with microbial taxa, CAZy/KO gene abundances, and soil properties were evaluated. The results showed weak or non-significant linear correlations between enzyme activities and the corresponding gene abundances (all P > 0.05), even after accounting for multiple environmental factors. Piecewise structural equation modeling revealed that total phosphorus (TP) strongly promoted cellulase gene (K01179) abundance (beta = 0.592, P < 0.01), while soil organic matter (SOM) significantly enhanced cellulase activity (beta = 0.587, P < 0.05). In contrast, microbial taxonomic composition contributed little once abiotic factors (e.g. pH, TP, SOM, C/N) were considered. Collectively, TP and SOM act as dominant abiotic filters that decoupled functional gene abundance from enzyme activities. These findings highlight in alpine grasslands, soil properties, rather than community or gene abundance, primarily regulate enzyme activities. This provides new mechanistic insight into soil enzyme dynamics under nutrient limitation and extreme environments, emphasizing that environmental filters override microbial traits in shaping ecosystem function.
Studies have indicated that COVID-19 infection may accelerate the aging process in organisms. However, it remains unknown whether contracting COVID-19 affects life expectancy. Furthermore, the underlying biological mechanisms behind these findings are still unclear. We conducted a prospective cohort study on 56,504 participants of European ancestry from the UK Biobank who reported the time and number of COVID-19 infection between January 2020 and September 2023. The parental average longevity was used as a proxy for their own longevity. Linear regression was used to assess the relationship between COVID-19 infection and longevity. Furthermore, we investigated the shared genetic basis between COVID-19 and longevity using large-scale genome-wide association studies (GWAS) for COVID-19 (122,616 cases and 2,475,240 controls) and longevity (3,484 cases and 25,483 controls). Mendelian randomization (MR) and mediation analysis were utilized to assess causal relationships and potential mediators between COVID-19 susceptibility and longevity. Shared genetic loci between the two phenotypes were identified using conjunctional false discovery rate (conjFDR) statistical frameworks. After controlling for relevant covariates, COVID-19 infection might not be significantly correlated with longevity. In all MR methods, generalized summary-data-based Mendelian randomization (GSMR) analysis revealed a significant decrease in longevity due to severe COVID-19 infection (OR = 0.91, 95
Early identification of individuals at high risk for chronic diseases is crucial for prevention and intervention, yet current risk assessment tools are disease-specific, require extensive clinical data collection, and cannot provide multidisease risk profiles from a single measurement. Several protein large language models have been developed for tasks such as protein structure prediction, function prediction, and sequence design. However, none of these models can be directly applied in clinical settings to predict an individual's future disease risk. Here, we present a multimodal proteomics Transformer (Proformer) model that integrates protein expression, sequence, and function information for multidisease risk assessment. We trained Proformer using real proteomics data from 47 124 individuals from the UK Biobank to evaluate its performance in discriminating the risk of 20 common chronic diseases. Proformer achieved state-of-the-art (SOTA) performance in all 20 diseases compared with five common machine learning and deep learning models. Compared to three common clinical predictors, Proformer's 10-year discriminative performance outperforms Age + Sex model for 19 diseases, outperforms the ASCVD risk score for 16 diseases, and outperforms the panel composed of 35 clinical variables for 11 diseases. These results were replicated in the Scotland and Wales cohort from UK Biobank. In conclusion, Proformer enabled users to directly obtain a 10-year risk report for common chronic diseases by inputting their individual proteomics data.
The shared genetic signals between human cerebral cortex and substance use disorders (SUDs) remain largely unknown. Here, we utilized the Human Connectome Project Multi-Modal Parcellation (HCP-MMP) to divide each hemisphere into 180 regions and investigated the genetic overlap between cortical surface area/thickness of these novel regions and four types of SUDs (N > 1 million). We identified 17 and 282 shared genetic loci between global and regional cortical phenotypes and SUDs. The anatomical patterns of genetic overlap were similar for problematic alcohol use and nicotine use, with substantial overlap in the TGd, insula, primary motor cortex, and posterior cingulate cortex. The cortical patterns of SUDs were established along the anatomical and functional hierarchies in the sensorimotor-association (S-A) cortical axis, but were independent of evolutionary hierarchies. Mendelian randomization analyses indicated that genetically predicted reduced surface area of the ventromedial prefrontal cortex (area 25) and frontal opercular area 3 (FOP3), posterior dorsal superior temporal sulcus (STSdp), and posterior insular area 2 (PoI2) were associated with increased risk of cannabis use disorder and opioid use, respectively. Reduced thickness of retrosplenial complex (RSC) was associated with increased risk of problematic alcohol use. However, reduced thickness of fusiform face complex (FFC) was associated with decreased risk of nicotine use. In summary, we provided novel insights into the shared genetic etiology between cortical phenotypes and SUDs under a more refined multimodal cortical parcellation scheme.
Introduction: Lenvatinib is the first-line therapy of hepatocellular carcinoma (HCC) and the high frequency of lenvatinib resistance hinders the improvement of HCC treatment. Since NADPH plays vital roles in antioxidant defense and reductive biosynthesis, cancer cells exert NADPH metabolic adaptation to support their malignant activities, including drug resistance. However, the underlying mechanisms need to be further studied. Objectives: This study aims to delineate the latent mechanism by which HCC cells modulate NADPH metabolic adaptation and lenvatinib resistance. Methods: Using high-throughput screening, we screened LINC01532 as a critical regulator in NADPH metabolic adaptation. The function of LINC01532 in drug resistance of HCC cells was analyzed by in vitro and in vivo model. NADPH assay, malondialdehyde (MDA) assay, and glutathione (GSH) detection assay were carried out to explore the role of LINC01532 in NADPH metabolism. Furthermore, RNA-binding protein immunoprecipitation, RNA pull-down assay, co-immunoprecipitation, and chromatin immunoprecipitation experiments were utilized to uncover the underlying mechanisms. Results: High expression of LINC01532 predicted poorer prognosis in HCC patients. LINC01532 stimulated NADPH production and blunted lenvatinib-induced cell death, leading to drug resistance. Mechanistically, LINC01532 bound to hnRNPK and promoted CDK2-mediated phosphorylation of hnRNPK, which facilitated G6PD pre-mRNA splicing, resulting in high expression of G6PD and upregulated NADPH synthesis. The elevated NADPH cleared reactive oxygen species (ROS), supported biomass synthesis, and epigenetically modulated gene expression. Inhibition of LINC01532 significantly enhanced lenvatinib sensitivity of HCC cells. The m6A modification induced by mTORC1 promoted the expression of LINC01532 in HCC cells. Conclusion: Collectively, our findings demonstrate that LINC01532 confers lenvatinib resistance of HCC cells by modulating NADPH metabolic adaptation. LINC01532 might be a prognostic or therapeutic target for HCC.
Objectives: This bibliometric analysis investigates recent research trends in biologics and small molecules for treating inflammatory bowel disease (IBD) based on literature from the past decade. Methods: This cross-sectional study involved analyzing data retrieved from the Web of Science Core Collection (WoSCC) database to examine the evolution and thematic trends of biological agents and small-molecular drugs for IBD conducted between 1 January 2014, and 20 September 2024. VOSviewer software was utilized to assess co-authorship, co-occurrence, co-citation, and network visualization, followed by a further discussion on significant sub-themes. Results: From 2014 to 20 September 2024, the annual number of global publications increased by 23%, reflecting an acceleration in research activity. The journal “Inflammatory Bowel Diseases” published the highest number of manuscripts (579 publications) and garnered the most citations (13,632 citations), followed by the “Journal of Crohn’s & Colitis” (480 publications) and “Alimentary Pharmacology & Therapeutics” (250 publications). The United States led in productivity with 1943 publications and 66,320 citations, with UC San Diego (291) and authors Sandborn and Vermeire (180) topping the list. The co-occurrence cluster analysis of the top 100 keywords resulted in the formation of six distinct clusters: Disease Mechanisms, Drug Development, Surgical Interventions, Therapeutic Drug Monitoring (TDM), Immunological Targets, and Emerging Therapies. Burst terms (TNF-α inhibitors, JAK inhibitors, and trough-level optimization) highlight trends toward personalized biologics and small-molecule regimens. Conclusions: The bibliometric analysis indicates that IBD therapeutic research and clinical applications focus on biologics and small molecules, with research trends leaning toward precise therapy conversion or the combination in non-responders. Future work will assess monotherapy, the combination, and conversion therapies and investigate new drugs targeting inflammatory pathways.
Soybean (Glycine max) provides vegetable oils and proteins for human consumption. Its production depends on seeds and other production-related agronomic traits. How the seed traits are regulated in soybean remains largely unclear. In this study, we identified a miR172a-ERF416/413 module for the regulation of seed traits. The miR172a can cleave the targets ERF416 and ERF413 to affect the downstream gene expression for the reduction of soybean seed size and weight. Both the MIR172a-overexpressing transgenic soybean plants and the erf416/413 mutants produced smaller seeds than the control. Consistently, the ERF416-overexpressing transgenic soybean plants generated larger seeds. ERF416 and ERF413 were directly targeted to the promoter of GmKIX8-1 and GmSWEET10a to regulate their gene expression for seed size/weight control. Interestingly, the erf416/413 mutants showed higher seed yield per plant and higher total seed fatty acid (FA) content, whereas the MIR172a-transgenic soybean had lower total seed FA content compared with the control cultivar, suggesting that miR172a and ERF416/413 may function in FA accumulation through different pathways. Haplotypes of the ERF416 promoter region were further analyzed and Hap1 was correlated with higher gene expression and higher seed weight, while Hap3 was correlated with higher total seed lipid content. Our study revealed a new module for seed trait control. Manipulation of such alleles should facilitate breeding for high-oil and high-yield soybean cultivars.
Missing values in nuclear magnetic resonance metabolomics data compromise downstream clinical interpretation. Here, we present MetImputBERT, an imputation method based on a pretrained BERT framework. MetImputBERT uses the masks in the masked language model to simulate missing values and leverages predictions and reconstructions to these positions to simulate the imputation process. The learning of MetImputBERT is driven by minimizing the reconstruction error. MetImputBERT was pretrained on the largest metabolomics dataset to date, comprising data from over 230 000 individuals in the UK Biobank. When new datasets with missing values were encountered, MetImputBERT loaded the pretrained parameters and directly imputed the missing values by inferring their reconstructed estimates. MetImputBERT outperformed commonly used methods-K-nearest neighbors, multiple imputation by chained equations, and singular value decomposition-in imputation performance on two independent test sets. We provide an open-source Python tool that allows users to quickly impute missing values in their own NMR metabolomics data without any additional training.
Soybean is one of the most important oilseed crops, and its seed oil content directly determines the economic value and industrial applicability worldwide. However, how soybean seed oil accumulation is regulated remains less understood. Here, through RNA-seq analysis and screening for the interacting proteins of a positive oil regulator GmNFYA, we identified an AP2/ERF-type transcription factor GmERFA, which acts as a negative regulator of oil accumulation. Knocking out GmERFA and its homologue by genome editing increased seed total fatty acid content, while overexpression of GmERFA leads to a reduced fatty acid level in transgenic soybean. GmERFA interacts with GmNFYA to inhibit its transcriptional activation of GmbZIP123 and GmZF392, both of which promotes seed oil accumulation. The GmERFA also directly binds to the promoter regions of GmbZIP123 and GmZF392 and represses their gene expression. Through further analysis of more than 300 soybean accessions, an elite allele of ERFA with Hap3 promoter is identified to correlate with lower promoter activity, lower gene expression but higher seed oil content. The Hap3 ERFA may be selected and fixed during soybean domestication. Together, our study discovers a brake gene for oil accumulation and may function in a novel molecular network GmERFA-GmNFYA-GmbZIP123/GmZF392 at the later stage of soybean seed development. Manipulation of the gene, its elite allele, and the whole pathway should benefit breeding for high oil cultivars in soybean.
Early prediction of chronic diseases from routine blood tests has potential to transform public health prevention strategies. Here, we developed MetaboLM, a transformer-based language model pre-trained on plasma metabolomics data from 83,744 relatively healthy UK Biobank participants. After fine-tuning with metabolomics data from individuals diagnosed with 16 common chronic diseases, MetaboLM demonstrated excellent performance in disease prediction and stratification, and generated a metabolomic risk score (MetaboRS) capable of predicting disease onset more than 10 years in advance. MetaboRS outperformed established demographic predictors in 16 diseases, and outperformed atherosclerotic cardiovascular disease (ASCVD) risk equations in 13 diseases. Furthermore, interpretability analysis of the attention mechanism identified key metabolites related to disease prediction. These findings underscore the potential of metabolomic language models and derived risk scores for predicting the risk of multiple diseases and for other potential downstream applications.
The causative mechanisms underlying the genetic relationships of neurodegenerative diseases with epigenetic aging and human longevity remain obscure. We aimed to detect causal associations and shared genetic etiology of neurodegenerative diseases with epigenetic aging and human longevity. We obtained large-scale genome-wide association study summary statistics data for four measures of epigenetic age (GrimAge, PhenoAge, IEAA, and HannumAge) (N = 34,710), multivariate longevity (healthspan, lifespan, and exceptional longevity) (N = 1,349,462), and for multiple neurodegenerative diseases (N = 6618-482,730), including Lewy body dementia, Alzheimer's disease (AD), Parkinson's disease, amyotrophic lateral sclerosis, and multiple sclerosis. Main analyses were conducted using multiplicative random effects inverse-variance weighted Mendelian randomization (MR), and conditional/conjunctional false discovery rate (cond/conjFDR) approach. Shared genomic loci were functionally characterized to gain biological understanding. Evidence showed that AD patients had 0.309 year less in exceptional longevity (IVW beta = -0.309, 95% CI: -0.38 to -0.24, p = 1.51E-19). We also observed suggestively significant causal evidence between AD and GrimAge age acceleration (IVW beta = -0.10, 95% CI: -0.188 to -0.013, p = 0.02). Following the discovery of polygenic overlap, we identified rs78143120 as shared genomic locus between AD and GrimAge age acceleration, and rs12691088 between AD and exceptional longevity. Among these loci, rs78143120 was novel for AD. In conclusion, we observed that only AD had causal effects on epigenetic aging and human longevity, while other neurodegenerative diseases did not. The genetic overlap between them, with mixed effect directions, suggested complex shared genetic etiology and molecular mechanisms.