Genes whose expression is affected, in a consistent manner, by GWAS-identified risk variants and the disease process, constitute preferred drug targets. We herein combine cis-eQTL analysis in 27 sorted blood cell populations and 43 intestinal cell types identified by single cell RNA-Seq in the ileum, colon and rectum, with information on gene expression in patients, to search for putative drug targets for inflammatory bowel disease. We detect >95 K cis-eQTL that affect >13 K e-genes and cluster in >24 K regulatory modules. We uncover matching regulatory modules for 140 risk loci, implicating >300 e-genes not previously connected with inflammatory bowel disease, and find 152 matching e-genes whose expression is perturbed in the blood or gut of patients. We identify entrectinib, a small molecule inhibiting the NRLP3 inflammasome by binding NEK7, as a possible repurposing candidate. Inflammatory bowel disease risk variants can identify genes and pathways involved in disease mechanisms. Here, the authors integrate blood and intestinal eQTL maps with patient gene expression data to prioritise candidate drug targets and identify entrectinib as a potential repurposing candidate.
The production of multiple transcripts per gene is a process regulated by inherited genetic variants and epitranscriptomic modifications, and plays a prominent role in modulating complex traits and diseases. To simultaneously characterize the effect of genetic variants on transcript abundance and N6-methyladenosine (m⁶A) modifications, we produced long-read native poly(A) RNA-seq data for 60 genetically different lymphoblastoid cell lines (LCLs) from the 1000 Genomes/Geuvadis project. We identified a high diversity of both annotated (40.2%) and unannotated (59.8%) transcripts, with only a small proportion expressed across individuals (35% and 18.4%, respectively). In a genome-wide genetic analysis on transcripts, we identified 87 trQTLs (transcript quantitative trait loci), of which 55 were not detected as eQTLs using a larger published short-read RNAseq dataset (317 samples). A population wide characterization of m⁶A methylation DRACH motifs (canonical m⁶A consensus sequence) identified an average of 4.04 m⁶A modifications on 4,325 genes. Genetic association analysis of highly variable modifications from 1,272 genes identified m⁶A modification quantitative trait loci (m⁶A-QTLs) for 5 transcripts. Even with a limited number of genetic associations, colocalization analysis of trQTLs and m⁶A-QTLs identified 8 candidate transcripts mediating GWAS traits, with approximately 75% of the colocalized trQTLs implicating novel risk transcripts, and an enrichment of GWAS lead variants involved in trQTLs. Overall, the simultaneous characterization of transcripts and post-transcriptional modifications identified genetic effects on transcription often missed when using other sequencing technologies.
OBJECTIVE:To delineate organ-specific and systemic drivers of metabolic dysfunction-associated steatotic liver disease (MASLD), we applied integrative causal inference across clinical, imaging, and proteomic domains in individuals with and without type 2 diabetes (T2D). METHODS:Bayesian network analyses and complementary two-sample Mendelian randomization were used to quantify causal pathways linking adipose distribution, glycemia, and insulin dynamics with liver fat in the IMI-DIRECT prospective cohort study. Data included frequently sampled metabolic challenge tests, MRI-derived abdominal and hepatic fat content, serological biomarkers, and Olink plasma proteomics from 331 adults with new-onset T2D and 964 adults without diabetes, with harmonized protocols enabling replication. RESULTS:High basal insulin secretion rate (BasalISR), estimated via C-peptide deconvolution, emerged as the primary potential causal driver of liver fat accumulation in both cohorts. BasalISR, a clearance-independent measure of β-cell insulin output distinct from peripheral insulin levels, was independently linked to hepatic steatosis. Visceral adipose tissue exhibited bidirectional associations with liver fat, suggesting a self-reinforcing metabolic loop. Of 446 analyzed proteins, 34 mapped to these metabolic networks (27 in the non-diabetes network, 18 in the T2D network, and 11 shared). Key proteins directly associated with liver fat included GUSB, ALDH1A1, LPL, IGFBP1/2, CTSD, HMOX1, FGF21, AGRP, and ACE2. Sex-stratified analyses identified GUSB in females and LEP in males as the strongest protein predictors of liver fat. CONCLUSIONS:BasalISR may better capture early β-cell-driven disturbances contributing to MASLD. These findings outline a multifactorial, sex- and disease stage-specific proteo-metabolic architecture of hepatic steatosis and identify potential biomarkers or therapeutic targets.
Multiomic profiling provides a comprehensive physiological overview at the molecular level, but understanding of its spatiotemporal dynamics remains limited in human populations. We profiled longitudinal whole-blood gene expression and metabolite levels in 335 females over 8 years. Levels of 5061 genes and 181 metabolites changed over time, with individual trajectories often diverging from population-level trends. Longitudinally variable genes showed cell type specificity and enrichment for aging-relevant pathways, including cardiometabolic and neurodegenerative disorders. Longitudinal trajectories were further shaped by genetics, circadian rhythm, seasonality, and environmental pollutant exposures. Integrative analyses revealed extensive static and time-variable cross-omic connectivity. Longitudinal profiling offers insight into the temporal evolution of age-related conditions at the molecular level, and understanding individual variation within these longitudinal patterns will be essential for future precision medicine approaches.
ABSTRACT Metabolic diseases such as type 2 diabetes (T2D) arise through complex interactions between physiological, molecular, and environmental processes. Clinical traits including age, sex, adiposity, and glycaemic status are strongly associated with disease risk and progression, yet most molecular studies examine these factors independently and assume relatively static molecular regulation. Consequently, how physiological state dynamically reshapes molecular organisation across omics layers remains poorly understood. Here, we integrated transcriptomic, proteomic, metabolomic, and genetic data from 3,027 individuals in the IMI DIRECT cohort to characterise the joint molecular effects of age, sex, body mass index (BMI), and glycated haemoglobin (HbA1c). We identified widespread associations between these traits and molecular phenotypes. However, interaction analyses revealed a more complex context-dependent regulation, showing that the molecular effect of one trait frequently depends on the state of another, with sex-specific effects of age being more prominent. We also investigated relationships between different types of molecular phenotypes and how these relationships are modulated by metabolic disease relevant traits, demonstrating that cross-omic molecular coordination is itself dynamically remodelled by physiological and metabolic state. Probabilistic causal inference identified a directionally structured network of age-associated molecules, revealing pathways through which age effects propagate across omics layers, showcased in the example of the mTOR signalling pathway. Integration of this directed network with genetic colocalisation analyses also identified a sub-network relevant for T2D. Collectively, our findings demonstrate that metabolic disease relevant traits not only independently influence molecular phenotype abundance but also jointly reshape the directional organisation of cross-omic molecular networks. These results support a model in which metabolic disease susceptibility emerges through dynamic rewiring of interconnected molecular systems and provide a framework for context-dependent biomarker discovery, disease stratification, and precision metabolic medicine.
Abstract Transcriptomic profiling of peripheral blood offers a promising, non-invasive approach for disease diagnosis and monitoring. However, its clinical translation is hindered by limited knowledge of the natural temporal variation. Here, we present a comprehensive reference map of longitudinal transcriptomic variability, based on RNA-sequencing of 333 healthy individuals sampled at three time points over six months. We find that 85% of genes and 99% of transcripts exhibit greater intra-individual than inter-individual variation, primarily driven by dynamic regulation of housekeeping pathways. In contrast, immune-related transcripts –particularly those linked to T and B cell activity– are strikingly stable over time. Gene expression levels drive inter-individual differences, while splicing variation contributes more to intra-individual fluctuation. In an independent twin cohort (148 monozygotic, 166 dizygotic), genes with high inter-individual variability show greater heritability, suggesting genetic control of steady-state expression. By integrating extensive clinical and environmental data, we trace temporal expression changes to genetic, compositional, and external factors, and identify robust seasonal and sex-specific signatures. These findings were validated in a third, cross-sectional cohort of 3,480 individuals. The observed temporal variation patterns have important implications for cohort-based transcriptomic analyses, as they may limit discovery and reproducibility of expression quantitative trait loci and increase the risk of spurious associations in cross-sectional studies. This resource provides a critical baseline for distinguishing disease-associated transcriptomic changes from normal physiological variation, advancing the reliability of blood-based biomarkers in clinical practice.
Zinc is essential for many physiological processes and its deficiency is highly prevalent worldwide. Its complex homeostasis involves membrane transporters from the SLC39/ZIP and SLC30/ZnT protein families. We conducted a genome-wide association study (GWAS) meta-analysis of urinary zinc levels in three European-ancestry cohorts (N = 10,113), followed by in silico and in vivo studies to elucidate their underlying public health and physiological relevance. We identified eleven genome-wide significant signals with six mapping to SLC39/ZIP and SLC30/ZnT gene regions. The lead signal (rs3008217C>G, p = 2.42E-110) in the SLC30A2 gene region which explained 6.1% of urinary zinc variation strongly colocalized with its expression in kidney tubules. Low phenotypic and genetic correlations between plasma and urinary zinc levels indicated distinct genetic regulation. High urinary zinc correlated with an unfavorable cardiometabolic profile, and Mendelian randomization analyses suggested causal roles for diabetes increasing urinary zinc levels, and elevated urinary zinc increasing stroke risk. Analyzing country-level allele frequencies and zinc deficiency prevalences revealed a 3-fold higher genetic zinc excretion risk in sub-Saharan Africa compared to Europe, significantly correlating with nutritional zinc deficiency prevalence. Although mutations in SLC30A2 are linked to insufficient zinc in human milk, we found no association with common variants using data generated from 387 mothers. Mice experiments showed that dietary zinc deficiency decreased urinary but not plasma zinc levels, and upregulated kidney Slc30a2 expression. This first GWAS on urinary zinc highlights the involvement of zinc transporters in its genetic regulation, as well as its role as a non-invasive biomarker for cardiometabolic diseases.
Here we report the results from exploratory analysis using a Bayesian network approach of data originally derived from a large North European study of type 2 diabetes (T2D) conducted by the IMI DIRECT consortium. 3029 individuals (795 with T2D and 2234 without) within 7 different study centres provided data comprising genotypes, proteins, metabolites, gene expression measurements and many different clinical variables. The main aim of the current study was to demonstrate the utility of our previously developed method to fit Bayesian networks by performing exploratory analysis of this dataset to identify possible causal relationships between these variables. The data was analysed using the BayesNetty software package, which can handle mixed discrete/continuous data with missing values. The original dataset consisted of over 16,000 variables, which were filtered down to 260 variables for analysis. Even with this reduction, no individual had complete data for all variables, making it impossible to analyse using standard Bayesian network methodology. However, using the recently proposed novel imputation method implemented in BayesNetty we computed a large average Bayesian network from which we could infer possible associations and causal relationships between variables of interest. Our results confirmed many previous findings in connection with T2D, including possible mediating proteins and genes, some of which have not been widely reported. We also confirmed potential causal relationships with liver fat that were identified in an earlier study that used the IMI DIRECT dataset but was limited to a smaller subset of individuals and variables (namely individuals with complete data at pre-defined variables of interest). In addition to providing valuable confirmation, our analyses thus demonstrate a proof-of-principle of the utility of the method implemented within BayesNetty. The full final average Bayesian network generated from our analysis is freely available and can be easily interrogated further to address specific focussed scientific questions of interest.
Previous case–control studies have reported aberrations of the gut microbiota in individuals with prediabetes. The primary objective of the present study was to explore the dynamics of the gut microbiota of individuals with prediabetes over 4 years with a secondary aim of relating microbiota dynamics to temporal changes of metabolic phenotypes. The study included 486 European patients with prediabetes. Gut microbiota profiling was conducted using shotgun metagenomic sequencing and the same bioinformatics pipelines at study baseline and after 4 years. The same phenotyping protocols and core laboratory analyses were applied at the two timepoints. Phenotyping included anthropometrics and measurement of fasting plasma glucose and insulin levels, mean plasma glucose and insulin under an oral glucose tolerance test (OGTT), 2-h plasma glucose after an OGTT, oral glucose insulin sensitivity index, Matsuda insulin sensitivity index, body mass index, waist circumference, and systolic and diastolic blood pressure. Measures of the dynamics of bacterial microbiota were related to concomitant changes in markers of host metabolism. Over 4 years, significant declines in richness were observed in gut bacterial and viral species and microbial pathways accompanied by significant changes in the relative abundance and the genetic composition of multiple bacterial species. Additionally, bacterial-viral interactions diminished over time. Despite the overall reduction in bacterial richness and microbial pathway richness, 80 dominant core bacterial species and 78 core microbial pathways were identified at both timepoints in 99
While it is generally known that metabolic disorders and circadian dysfunction are intertwined, how the two systems affect each other is not well understood, nor are the genetic factors that might exacerbate this pathological interaction. Blood chemistry is profoundly changed in metabolic disorders, and we have previously shown that serum factors change cellular clock properties. To investigate if circulating factors altered in metabolic disorders have circadian modifying effects, and whether these effects are of genetic origin, we measured circadian rhythms in U2OS cell in the presence of serum collected from diabetic, obese, or control subjects. We observed that circadian period lengthening in U2OS cells was associated with serum chemistry that is characteristic of insulin resistance. Characterizing the genetic variants that altered circadian period length by genome-wide association analysis, we found that one of the top variants mapped to the E3 ubiquitin ligase MARCH1 involved in insulin sensitivity. Confirming our data, the serum circadian modifying variants were also enriched in type 2 diabetes and chronotype variants identified in the UK Biobank cohort. Finally, to identify serum factors that might be involved in period lengthening, we performed detailed metabolomics, and found that the circadian modifying variants are particularly associated with branched chain amino acids, whose levels are known to correlate with diabetes and insulin resistance. Overall, our multi-omics data showed comprehensively that systemic factors serve as a path through which metabolic disorders influence circadian system, and these can be examined in human populations directly by simple cellular assays in common cultured cells.
Genes whose expression is affected in a consistent manner by GWAS-identified risk variants and the disease process, constitute preferred drug targets. We herein combine integrated cis-eQTL analysis in 27 blood cell populations and 43 intestinal cell types of the ileum, colon and rectum, and information on gene expression in patients, to search for putative drug targets for inflammatory bowel disease (IBD). We detect >95K cis-eQTL that affect >13K e-genes and cluster in >24K regulatory modules (RM). We uncover matching RM for 140 risk loci, implicating >300 e-genes not previously connected to IBD, and find 152 IBD-matching e-genes whose expression is perturbed in the blood or gut of patients. We identify entrectinib, a small molecule inhibiting the NRLP3 inflammasome by binding NEK7, as a promising repurposing candidate for IBD. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This project was conducted with funding from the SYSCID H2020 grant (ref. 733100, the MyQuant (ref. 30770923) and BRIDGE (O.0006.22 RG3124) projects from the Excellence of Science (EOS) program (FNRS, Federation Wallonie-Bruxelles and FWO, Flemish Community), the CLIMAX (WELBIO CR 2022 A) project from WELBIO (Walloon Region), the RHEAQT (T.0096.19) and IBD GI Seq (T.0190.19) projects from the FNRS (Federation Wallonie Bruxelles), the ARC RHEACT WITH HSPC project from the University of Liege. Computational resources have been provided by the Consortium des Equipements de Calcul Intensif (CeCI), funded by the Fonds de la Recherche Scientifique de Belgique (F.R.S. FNRS) under Grant No. 2.5020.11 and by the Walloon Region. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Written informed consent was obtained prior to donation in agreement with the recommendations of the declaration of Helsinki for experiments involving human subjects. The experimental protocol was approved by the Ethics committee of the CHU Liege (reference number: 2017/214). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
Transposable elements (TEs) are prevalent repeats in the human genome, play a significant role in the regulome, and their disruption can contribute to tumorigenesis. However, TE influence on gene expression in cancer remains unclear. Here, we analyze 275 normal colon and 276 colorectal cancer samples from the SYSCOL cohort, discovering 10,231 and 5,199 TE-expression quantitative trait loci (eQTLs) in normal and tumor tissues, respectively, of which 376 are colorectal cancer specific eQTLs, likely due to methylation changes. Tumor-specific TE-eQTLs show greater enrichment of transcription factors, compared to shared TE-eQTLs suggesting specific regulation of their expression in tumor. Bayesian networks reveal 1,766 TEs as mediators of genetic effects, altering the expression of 1,558 genes, including 55 known cancer driver genes and show that tumor-specific TE-eQTLs trigger the driver capability of TEs. These insights expand our knowledge of cancer drivers, deepening our understanding of tumorigenesis and presenting potential avenues for therapeutic interventions.
Files that were used for mito-nuclear eQTL analyses.
BACKGROUND:Expression quantitative trait loci (eQTL) studies provide insights into regulatory mechanisms underlying disease risk. Expanding studies of gene regulation to underexplored populations and to medically relevant tissues offers potential to reveal yet unknown regulatory variants and to better understand disease mechanisms. Here, we performed eQTL mapping in subcutaneous (S) and visceral (V) adipose tissue from 106 Greek individuals (Greek Metabolic study, GM) and compared our findings to those from the Genotype-Tissue Expression (GTEx) resource.RESULTS:We identified 1,930 and 1,515 eGenes in S and V respectively, over 13% of which are not observed in GTEx adipose tissue, and that do not arise due to different ancestry. We report additional context-specific regulatory effects in genes of clinical interest (e.g. oncogene ST7) and in genes regulating responses to environmental stimuli (e.g. MIR21, SNX33). We suggest that a fraction of the reported differences across populations is due to environmental effects on gene expression, driving context-specific eQTLs, and suggest that environmental effects can determine the penetrance of disease variants thus shaping disease risk. We report that over half of GM eQTLs colocalize with GWAS SNPs and of these colocalizations 41% are not detected in GTEx. We also highlight the clinical relevance of S adipose tissue by revealing that inflammatory processes are upregulated in individuals with obesity, not only in V, but also in S tissue.CONCLUSIONS:By focusing on an understudied population, our results provide further candidate genes for investigation regarding their role in adipose tissue biology and their contribution to disease risk and pathogenesis.
Venous thromboembolism (VTE) is a common disease with high heritability. However, only a small portion of the genetic variance of VTE can be explained by known genetic risk factors. Neutrophil extracellular traps (NETs) have been associated with prothrombotic activity. Therefore, the genetic basis of NETs could reveal novel risk factors for VTE. A recent genome-wide association study of plasma cell-free DNA (cfDNA) levels in the Genetic Analysis of Idiopathic Thrombophilia 2 (GAIT-2) Project showed a significant associated locus near ORM1. We aimed to further explore this candidate region by next-generation sequencing, copy number variation (CNV) quantification, and expression analysis using an extreme phenotype sampling design involving 80 individuals from the GAIT-2 Project. The RETROVE study with 400 VTE cases and 400 controls was used to replicate the results. A total of 105 genetic variants and a multiallelic CNV (mCNV) spanning ORM1 were identified in GAIT-2. Of these, 17 independent common variants, a region of 22 rare variants, and the mCNV were significantly associated with cfDNA levels. In addition, eight of these common variants and the mCNV influenced ORM1 expression. The association of the mCNV and cfDNA levels was replicated in RETROVE (p-value = 1.19 × 10-6). Additional associations between the mCNV and thrombin generation parameters were identified. Our results reveal that increased mCNV dosages in ORM1 decreased gene expression and upregulated cfDNA levels. Therefore, the mCNV in ORM1 appears to be a novel marker for cfDNA levels, which could contribute to VTE risk.
BACKGROUND:Human plasma contains a wide variety of circulating proteins. These proteins can be important clinical biomarkers in disease and also possible drug targets. Large scale genomics studies of circulating proteins can identify genetic variants that lead to relative protein abundance.METHODS:We conducted a meta-analysis on genome-wide association studies of autosomal chromosomes in 22,997 individuals of primarily European ancestry across 12 cohorts to identify protein quantitative trait loci (pQTL) for 92 cardiometabolic associated plasma proteins.RESULTS:We identified 503 (337 cis and 166 trans) conditionally independent pQTLs, including several novel variants not reported in the literature. We conducted a sex-stratified analysis and found that 118 (23.5%) of pQTLs demonstrated heterogeneity between sexes. The direction of effect was preserved but there were differences in effect size and significance. Additionally, we annotate trans-pQTLs with nearest genes and report plausible biological relationships. Using Mendelian randomization, we identified causal associations for 18 proteins across 19 phenotypes, of which 10 have additional genetic colocalization evidence. We highlight proteins associated with a constellation of cardiometabolic traits including angiopoietin-related protein 7 (ANGPTL7) and Semaphorin 3F (SEMA3F).CONCLUSION:Through large-scale analysis of protein quantitative trait loci, we provide a comprehensive overview of common variants associated with plasma proteins. We highlight possible biological relationships which may serve as a basis for further investigation into possible causal roles in cardiometabolic diseases.
IntroductionMigraine is a complex disorder with genetic and environmental inputs. Cumulative evidence implicates oxidative stress (OS) in migraine pathophysiology while genetic variability may influence an individuals' oxidative/antioxidant capacity. Aim of the current study was to investigate the impact of eight common OS-related genetic variants [rs4880 (SOD2), rs1001179 (CAT), rs1050450 (GPX1), rs1695 (GSTP1), rs1138272 (GSTP1), rs1799983 (NOS3), rs6721961 (NFE2L2), rs660339 (UCP2)] in migraine susceptibility and clinical features in a South-eastern European Caucasian population. MethodsGenomic DNA samples from 221 unrelated migraineurs and 265 headache-free controls were genotyped for the selected genetic variants using real-time PCR (melting curve analysis). ResultsAlthough allelic and genotypic frequency distribution analysis did not support an association between migraine susceptibility and the examined variants in the overall population, subgroup analysis indicated significant correlation between NOS3 rs1799983 and migraine susceptibility in males. Furthermore, significant associations of CAT rs1001179 and GPX1 rs1050450 with disease age-at-onset and migraine attack duration, respectively, were revealed. Lastly, variability in the CAT, GSTP1 and UCP2 genes were associated with sleep/weather changes, alcohol consumption and physical exercise, respectively, as migraine triggers. DiscussionHence, the current findings possibly indicate an association of OS-related genetic variants with migraine susceptibility and clinical features, further supporting the involvement of OS and genetic susceptibility in migraine.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark21