Meiotic recombination is a fundamental process that generates genetic diversity by creating new combinations of existing alleles1. Whereas crossovers in humans are well characterized2, the more frequent non-crossovers that lead to gene conversion remain challenging to study. Here we show that single high-fidelity long sequencing reads from sperm can capture both crossovers and non-crossovers, which enables effectively arbitrary sample sizes for analysis from a single male. We analysed 2,382 candidate non-crossovers in 15 sperm samples from 13 donors, and identified a consistent component with properties distinct from PRDM9-induced recombination. This phenomenon was not associated with meiotic double-strand break sites identified by DMC1 binding, the crossover recombination map or GC-biased gene conversion, but was associated with genomic fragile sites. This component is also seen in paternal non-crossover gene conversions in pedigree data3. Applying the same analysis to 12 blood samples4, we observed non-crossover gene conversions with similar properties, but very few crossover events. Further, we demonstrate variation between donors for the different types of recombination, even when they share the same PRDM9 genotype. We suggest that a substantial fraction of the non-crossover gene conversion events seen in sperm arise prior to meiosis.
Fatty acids are important as structural components, energy sources, and signaling mediators. While studies have extensively explored genetic regulation of fatty acids in serum and other bodily fluids, their regulation within adipose tissue, a crucial regulator of cardiovascular and metabolic health, remains unclear. Here, we investigated the genetic regulation of 18 fatty acids in subcutaneous adipose tissue from 569 female twins from TwinsUK. Using twin models, the heritability of fatty acids ranged from 5% to 59%, indicating a substantial genetic regulation of fatty acid levels within adipose tissue, which was also tissue specific in many cases. Genome-wide association studies identified 10 significant loci, in SCD, ADAMTSL1, ZBTB41, SNTB1, EXOC6B, ACSL3, LINC02055, MKRN2/TSEN2, FADS1, and HAPLN across 13 fatty acids or fatty acid product-to-precursor ratios. Using adipose gene expression and methylation, which were concurrently measured in these samples, we detected five fatty acid-associated signals that colocalized with expression quantitative trait locus (eQTL) and methylation quantitative trait locus (meQTL) signals, highlighting fatty acids that are regulated by molecular processes within adipose tissue. We explored links between polygenic scores of common metabolic traits and adipose fatty acid levels and identified associations between polygenic scores of BMI, body-fat distribution, and triglycerides and several fatty acids, indicating these risk scores impact local adipose tissue content. Overall, our results identified local genetic regulation of fatty acids within adipose tissue and highlighted their links with renal and cardio-metabolic health.
Abstract Background X chromosome inactivation (XCI) is the mechanism which randomly silences one X chromosome to equalise gene expression between 46, XX females and 46, XY males. Though XCI is expected to result in a random pattern of mosaicism across tissues, some females display a significantly unbalanced ratio in immune cells, termed XCI-skew, in which ≥75% of cells have the same X inactivated. XCI-skew is associated with adverse health outcomes and its prevalence increases with age – particularly after midlife - yet the specific risk factors have yet to be identified. The menopausal transition, which is driven by profound shifts in sex hormone levels, has significant impact on chronic disease risk yet the molecular and cellular effects are incompletely understood. We hypothesised that the menopausal transition may impact XCI-skew. Methods Using XCI data measured in blood-derived DNA from 1,395 females from the TwinsUK population cohort, along with questionnaires, genetic data, and sex hormone measures, we carried out a cross-sectional study to assess the impact of the menopausal transition and sex hormones on XCI-skew. Results We demonstrate that early menopause (<45yrs) is associated with increased risk of XCI-skew. In subset analyses across those who had a surgically induced or natural menopause, we find the association restricted to those who underwent a surgical menopause. We next identify a low polygenic score (PGS) for testosterone levels is significantly associated with XCI-skew, which we replicate in an independent dataset (n=149), while a PGS for age at natural menopause is not associated. Finally, using longitudinal measures across two time points spanning ∼18 years we show XCI-skew is a stable cellular phenotype that typically increases over time. Discussion These data represent the first environmental and genetic risk factors of XCI-skew, both of which implicate endogenous sex hormone levels, particularly testosterone. We propose XCI-skew may have clinical relevance in postmenopausal females.
Multiomic profiling provides a comprehensive physiological overview at the molecular level, but understanding of its spatiotemporal dynamics remains limited in human populations. We profiled longitudinal whole-blood gene expression and metabolite levels in 335 females over 8 years. Levels of 5061 genes and 181 metabolites changed over time, with individual trajectories often diverging from population-level trends. Longitudinally variable genes showed cell type specificity and enrichment for aging-relevant pathways, including cardiometabolic and neurodegenerative disorders. Longitudinal trajectories were further shaped by genetics, circadian rhythm, seasonality, and environmental pollutant exposures. Integrative analyses revealed extensive static and time-variable cross-omic connectivity. Longitudinal profiling offers insight into the temporal evolution of age-related conditions at the molecular level, and understanding individual variation within these longitudinal patterns will be essential for future precision medicine approaches.
Abstract Background Subcutaneous adipose tissue, the largest fat depot buffering fatty acids, correlates with cardio-metabolic diseases. Diet plays a major role in fatty acid metabolism and correlates with cardio-metabolic diseases. Studies have revealed clear links between plasma fatty acids with diet and cardio-metabolic diseases, but limited research focuses on fatty acids in subcutaneous adipose tissue. Methods In TwinsUK, 18 subcutaneous adipose tissue fatty acid proportions were measured ( N = 569) alongside dietary intake and clinical phenotypes. We investigated the association of adipose fatty acid levels and: (1) proxies of circadian factors; (2) eight dietary patterns; (3) cardio-metabolic risk phenotypes; and (4) the mediation effects of adipose fatty acids on the associations between diet and cardio-metabolic diseases. Results (1) Proxies of circadian factors: We observed no association between adipose fatty acids and fasting hours or time of visit. (2) Dietary pattern: We identified associations between eight dietary patterns and adipose fatty acids, where the Healthy Diet Indicator showed the most associations. Most dietary scores positively correlated with polyunsaturated fatty acid (Beta = 0.10 [95% CI: 0.005, 0.19] − 0.23 [0.14, 0.32], P FDR < 0.05), but Dietary Approaches to Stop Hypertension negatively correlated with two polyunsaturated fatty acids, arachidonic acid (Beta = −0.09 [−0.17,−0.006], P FDR = 0.04) and adrenic acid (Beta = −0.09 [−0.18, −0.004], P FDR = 0.04). Follow-up analysis showed that the intake of fat-rich food and fat-related nutrients were highly associated with fatty acid proportions. (3) Cardio-metabolic risk phenotypes including adiposity, lipids, glycemic and cardiovascular factors were associated with adipose fatty acids, where saturated and unsaturated fatty acids showed opposite directions. (4) Mediation effect: Adipose tissue proportion of arachidonic acid putatively mediated 24.7% of the total effect of Dietary Approaches to Stop Hypertension on visceral to gluteofemoral fat ratio. Conclusion Adipose tissue fatty acid proportions represent dietary intakes and are associated with cardio-metabolic phenotypes. We provide a resource summarising the association between dietary measurements (dietary scores, nutrients and foods) and adipose tissue fatty acid proportions.
Abstract Transcriptomic profiling of peripheral blood offers a promising, non-invasive approach for disease diagnosis and monitoring. However, its clinical translation is hindered by limited knowledge of the natural temporal variation. Here, we present a comprehensive reference map of longitudinal transcriptomic variability, based on RNA-sequencing of 333 healthy individuals sampled at three time points over six months. We find that 85% of genes and 99% of transcripts exhibit greater intra-individual than inter-individual variation, primarily driven by dynamic regulation of housekeeping pathways. In contrast, immune-related transcripts –particularly those linked to T and B cell activity– are strikingly stable over time. Gene expression levels drive inter-individual differences, while splicing variation contributes more to intra-individual fluctuation. In an independent twin cohort (148 monozygotic, 166 dizygotic), genes with high inter-individual variability show greater heritability, suggesting genetic control of steady-state expression. By integrating extensive clinical and environmental data, we trace temporal expression changes to genetic, compositional, and external factors, and identify robust seasonal and sex-specific signatures. These findings were validated in a third, cross-sectional cohort of 3,480 individuals. The observed temporal variation patterns have important implications for cohort-based transcriptomic analyses, as they may limit discovery and reproducibility of expression quantitative trait loci and increase the risk of spurious associations in cross-sectional studies. This resource provides a critical baseline for distinguishing disease-associated transcriptomic changes from normal physiological variation, advancing the reliability of blood-based biomarkers in clinical practice.
Abstract Cysteine oxidation analyses require the preservation of the redox state present at harvest and quantitative scaling to relate oxidation to protein copy numbers, rather than only providing fractional oxidation data. Here, we present ‘ReCap’, a Redox Capture workflow combining Oxi-DIA, an enrichment-free isotope-encoded DIA workflow, with Oxi-Stop, a simple oxygen-exclusion strategy for cryopreserved tissue. In mouse brains, Oxi-DIA quantified 17,809 cysteine sites belonging to 6,085 protein groups in every sample, enabling matched measurements of residue-resolved oxidation and protein abundance. Atmospheric oxygen exposure during 14 days of cryopreservation distorted the measured cysteine redox state. The resultant increase of an estimated 5.3176 × 10 11 µg −1 oxidised cysteine molecules was mitigated by Oxi-Stop, which minimised exogenous oxidation during cryopreservation. Copy-number scaling altered the interpretation of cysteine oxidation values. Although cysteine oxidation was detected across 2,371 sites and 1,439 proteins, 20 sites on abundant proteins accounted for 44% of the oxidised signal. ReCap advances redox proteomics from providing a site catalogue into a biologically weighted map of redox information, revealing cysteine oxidation as a sparse, ordered and quantitatively concentrated signal.
Cardiovascular disease progression is characterised by the dysregulation of lipid metabolism and pro-atherogenic effects of adipose tissue signalling. Recent findings from the analysis of transcriptomic data in bulk tissue has enabled these insights and revealed important changes in gene expression. However, few studies have explored these molecular mechanisms before the onset of cardiovascular disease. We explore associations between future lipid-regulating drug use and cardiometabolic traits (n = 103), including DXA scans of body composition at baseline and follow-up 5–10 years later, in a cohort of British twins (n up to 6963). Utilising transcriptomic profiles from a subset of twins (n = 766), we explore the associations between baseline adipose tissue gene expression, clinical traits, and future lipid-regulating drug usage. We then test the joint predictive capacity of clinical traits plus gene expression compared to traditional risk scores using an automated machine learning approach. We find 44 traits are associated with lipid-regulating drug usage including measurements of abdominal fat tissue, cardiovascular health, and lipid metabolism (FDR 5%). Then, we present that adipose tissue gene expression levels at baseline are associated cross-sectionally with 19 of these 44 traits (FDR 5%). By comparing adipose gene expression levels between individuals prescribed lipid-regulating drugs in the future and controls, we discover that genes associated with 16 of these 19 traits produced greater log2-fold changes, suggesting shared mechanisms. We reveal 15 differentially expressed genes comparing future lipid-regulating drug users and controls at baseline (FDR 10%), including some implicated in angiogenesis: ESM1, RCAN2, and SOCS3. Functional enrichment with 1212 significantly differentially expressed genes (p < 0.05) included molecular mechanisms related to abnormal cardiovascular system electrophysiology (p = 1.89 × 10−3), arrhythmia (p = 4.02 × 10−3), and mitochondrial pathways (p = 1.12 × 10−3). Finally, we confirm inclusion of gene expression levels as features in machine learning models achieves a better AUC (0.919) compared to traditional risk predictors. These findings highlight the potential of bulk transcriptomic data to improve risk stratification for lipid-regulating drug use, offering new insights into the RNA biology of adipose tissue and advancing approaches for cardiovascular disease prevention.
As we age, many tissues become colonized by microscopic clones carrying somatic driver mutations1-7. Some of these clones represent a first step towards cancer whereas others may contribute to ageing and other diseases. However, our understanding of this phenomenon remains limited due to the challenge of detecting mutations in small clones. Here we introduce a new version of nanorate sequencing (NanoSeq)8, a duplex sequencing method with an error rate lower than five errors per billion base pairs, which is compatible with whole-exome and targeted capture. Deep sequencing of polyclonal samples with single-molecule sensitivity simultaneously profiles large numbers of clones, providing accurate mutation rates, signatures and driver frequencies in any tissue. Applying targeted NanoSeq to 1,042 non-invasive samples of oral epithelium and 371 blood samples from a twin cohort, we report an extremely rich selection landscape, with 46 genes under positive selection in oral epithelium, more than 62,000 driver mutations and evidence of negative selection in essential genes. High-resolution maps of selection across coding and non-coding sites are obtained for many genes: a form of in vivo saturation mutagenesis. Multivariate regression models enable mutational epidemiology studies on how exposures and cancer risk factors, such as age, tobacco or alcohol, alter the acquisition or selection of somatic mutations. Accurate single-molecule sequencing provides a powerful tool to study early carcinogenesis, cancer prevention and the role of somatic mutations in ageing and disease.
Perfluorooctanoic acid (PFOA) and Perfluorooctanesulfonic Acid (PFOS) are synthetic substances with long half-lives. Their presence is widespread and pervasive, and they are noted for their environmental persistence. Research has shown these chemicals to be associated with dyslipidaemia, although few studies have considered the long-term associations in the general population. The aim of this study was to consider the longitudinal and cross-sectional associations with lipid phenotypes. We investigated the association of these chemicals with total cholesterol (TC), low-density lipoprotein (LDL), high-density lipoprotein (HDL), triglycerides (TG), and the total cholesterol: high-density lipoprotein ratio (TC:HDL), in a healthy unselected British population of twins (n = 2069), measured at three timepoints between 1996 and 2014. Serum levels of PFOA and PFOS decreased over time during this period. We demonstrate longitudinal associations across serum levels of both PFOA and PFOS, finding positive associations with TC (PFOA:β = 0.51, p = 1.9e−07; PFOS:β = 0.24, p = 3.8e−05) and LDL (PFOA:β = 0.61, p = 1.7e−11; PFOS:β = 0.42, p = 1.6e−14), and consistent negative associations with HDL and PFOA (β = −0.12, p = 0.003) and PFOS (β = −0.25, p = <2e−16). We also observe cross-sectional associations of PFAS with lipids across all three timepoints.
Previous studies have reported inverse associations between total (poly)phenol intake or specific subclasses of (poly)phenols, estimated from dietary questionnaires, and cardiovascular disease (CVD) risk. However, no studies have examined (poly)phenol-rich dietary patterns and their corresponding urinary metabolite profiles in relation to CVD risk. This study investigated the associations between a (poly)phenol-rich dietary score (PPS-D), its urinary metabolic signature (PPS-M), and longitudinal CVD risk in the TwinsUK cohort. We included 3110 participants (followed up for 11.20 ± 7.03 years) from TwinsUK who completed the EPIC-Norfolk Food Frequency Questionnaire with longitudinal data. A subset of 200 participants provided spot urine samples, in which 114 (poly)phenol metabolites were quantified using ultra-high-performance liquid chromatography-mass spectrometry (UHPLC-MS) to objectively measure (poly)phenol exposure. Associations between the PPS-D or PPS-M and CVD risk scores (ASCVD risk score and HeartScore), and biomarkers of CVD risk (blood pressure and lipid profile) were assessed using linear mixed models, adjusting for covariates and multiple testing (FDR < 0.05). PPS-D was negatively associated with ASCVD risk score (stdBeta: − 0.05 (− 0.07, − 0.04)) and Heartscore (stdBeta: − 0.03 (− 0.04, − 0.01)) (FDR-adjusted p < 0.01) in the overall population (n = 3,110). In the subgroup with urinary metabolites (n = 200), such significant associations were partially replicated through metabolites of flavonoids, phenolic acids, and tyrosols that significantly negatively associated with the ASCVD risk score, HeartScore, diastolic blood pressure (DBP), and positively with high-density lipoprotein cholesterol (HDL-C). In addition, a higher PPS-M was correlated with elevated HDL-C and lower blood pressure, ASCVD risk score, and HeartScore. Higher adherence to a (poly)phenol-rich diet is associated with lower CVD risk, with consistent associations observed through urinary metabolite profiles, highlighting the long-term cardiovascular benefits of (poly)phenol consumption, particularly flavonoids and phenolic acids.
Mutations that occur in the cell lineages of sperm or eggs can be transmitted to offspring. In humans, positive selection of driver mutations during spermatogenesis can increase the birth prevalence of certain developmental disorders1-3. Until recently, characterizing the extent of this selection in sperm has been limited by the error rates of sequencing technologies. Here we used the duplex sequencing method NanoSeq4 to sequence 81 bulk sperm samples from individuals aged 24-75 years. Our findings revealed a linear accumulation of 1.67 (95% confidence interval of 1.41-1.92) mutations per year per haploid genome driven by two mutational signatures associated with human ageing. Deep targeted and exome NanoSeq5 of sperm samples identified more than 35,000 germline coding mutations. We detected 40 genes (31 newly identified) under significant positive selection in the male germline that have activating or loss-of-function mechanisms and are involved in diverse cellular pathways. Most of the positively selected genes are associated with developmental or cancer predisposition disorders in children, whereas four of the genes exhibited increased frequencies of protein-truncating variants in healthy populations. We show that positive selection during spermatogenesis drives a 2-3-fold increased risk of known disease-causing mutations, which results in 3-5% of sperm from middle-aged to older individuals with a pathogenic mutation across the exome. These findings shed light on germline selection dynamics and highlight a broader increased disease risk for children born to fathers of advanced age than previously appreciated.
Heavy metals in our direct environment have profound effects on human health and while some are essential for life, others can be toxic. In vivo studies often focus on clinical features caused by overexposure to, or by deprivation of a heavy metal. However, to understand the cellular impact of heavy metals on health, studies in healthy volunteers before symptom onset are needed. Here, we explored the impact of mercury, lead and selenium in over 800 British female twins on multi-tissue gene expression levels as an intermediate phenotype. Total mercury, lead and selenium concentrations were determined in plasma as a proxy for heavy metal exposure. We identified significant associations between total mercury levels measured in plasma, that fall within normal ranges, and expression of 873 genes within skin tissue, including PUSL1, SAMD10, ERCC1, MRPL17, NDUFB8, SELENOH, SEC31A, and KAT7P1. Functional analysis of genes associated with total mercury levels in plasma show a strong link to the mitochondrial oxidative phosphorylation pathway (p-value = 3.02 × 10−10). Analysis of mitochondrial-specific gene expression supported involvement of genes of oxidative phosphorylation complexes (MT-ND4L, and MT-ND5), which are encoded in mitochondrial DNA. These results suggest that mercury is likely detrimental to the energy metabolism of mitochondria. We also tested for associations between total mercury levels in plasma and gene expression in adipose and whole blood samples, but did not identify significant associations in these tissues, nor with lead or selenium in any tissue. Our results demonstrate that subtoxic mercury exposure leaves a clear molecular signature. It also underscores the necessity of conducting tissue-specific association studies to accurately capture the molecular impact of environmental exposures, as only relevant tissues will manifest a response to environmental exposures.
Complete characterization of the genetic effects on gene expression is needed to elucidate tissue biology and the etiology of complex traits. Here, we analyzed 2,344 subcutaneous adipose tissue samples and identified 34K conditionally distinct expression quantitative trait locus (eQTL) signals in 18K genes. Over half of eQTL genes exhibited at least two eQTL signals. Compared to primary signals, non-primary signals had lower effect sizes, lower minor allele frequencies, and less promoter enrichment; they corresponded to genes with higher heritability and higher tolerance for loss of function. Colocalization of eQTL with conditionally distinct genome-wide association study signals for 28 cardiometabolic traits identified 3,605 eQTL signals for 1,861 genes. Inclusion of non-primary eQTL signals increased colocalized signals by 46%. Among 30 genes with ≥2 pairs of colocalized signals, 21 showed a mediating gene dosage effect on the trait. Thus, expanded eQTL identification reveals more mechanisms underlying complex traits and improves understanding of the complexity of gene expression regulation.
As we age, many tissues become colonised by microscopic clones carrying somatic driver mutations. Some of these clones represent a first step towards cancer whereas others may contribute to ageing and other diseases. However, our understanding of the clonal landscapes of human tissues, and their impact on cancer risk, ageing and disease, remains limited due to the challenge of detecting somatic mutations present in small numbers of cells. Here, we introduce a new version of nanorate sequencing (NanoSeq), a duplex sequencing method with error rates <5 errors per billion base pairs, which is compatible with whole-exome and targeted gene sequencing. Deep sequencing of polyclonal samples with single-molecule sensitivity enables the simultaneous detection of mutations in large numbers of clones, yielding accurate somatic mutation rates, mutational signatures and driver mutation frequencies in any tissue. Applying targeted NanoSeq to 1,042 non-invasive samples of oral epithelium and 371 samples of blood from a twin cohort, we found an unprecedentedly rich landscape of selection, with 49 genes under positive selection driving clonal expansions in the oral epithelium, over 62,000 driver mutations, and evidence of negative selection in some genes. The high number of positively selected mutations in multiple genes provides high-resolution maps of selection across coding and non-coding sites, a form of in vivo saturation mutagenesis. Multivariate regression models enable mutational epidemiology studies on how carcinogenic exposures and cancer risk factors, such as age, tobacco or alcohol, alter the acquisition and selection of somatic mutations. Accurate single-molecule sequencing has the potential to unveil the polyclonal landscape of any tissue, providing a powerful tool to study early carcinogenesis, cancer prevention and the role of somatic mutations in ageing and disease. ### Competing Interest Statement I.M., M.R.S., and P.J.C are co-founders, shareholders and consultants for Quotient Therapeutics Ltd. ### Funding Statement I.M. is funded by Cancer Research UK (C57387/A21777), the Dr Josef Steiner Cancer Research Foundation and the Wellcome Trust. TwinsUK is funded by the Wellcome Trust, Medical Research Council, Versus Arthritis, European Union Horizon 2020, Chronic Disease Research Foundation (CDRF), Zoe Ltd, the National Institute for Health and Care Research (NIHR) Clinical Research Network (CRN) and Biomedical Research Centre based at Guy's and St Thomas' NHS Foundation Trust in partnership with King's College London. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was carried out under TwinsUK BioBank ethics, approved by North West Liverpool Central Research Ethics Committee (REC reference 19/NW/0187), IRAS ID 258513 and earlier approvals granted to TwinsUK by the St Thomas' Hospital Research Ethics Committee, later London Westminster Research Ethics Committee (REC reference EC04/015). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study will be available through managed access via the TwinsUK Resource Executive Committee (TREC).
Whole-skin DNA methylation variation has been implicated in several diseases, including melanoma, but its genetic basis has not yet been fully characterized. Using bulk skin tissue samples from 414 healthy female UK twins, we performed twin-based heritability and methylation quantitative trait loci (meQTL) analyses for >400,000 DNA methylation sites. We find that the human skin DNA methylome is on average less heritable than previously estimated in blood and other tissues (mean heritability: 10.02%). meQTL analysis identified local genetic effects influencing DNA methylation at 18.8% (76,442) of tested CpG sites, as well as 1,775 CpG sites associated with at least one distal genetic variant. As a functional follow-up, we performed skin expression QTL (eQTL) analyses in a partially overlapping sample of 604 female twins. Colocalization analysis identified over 3,500 shared genetic effects affecting thousands of CpG sites (10,067) and genes (4,475). Mediation analysis of putative colocalized gene-CpG pairs identified 114 genes with evidence for eQTL effects being mediated by DNA methylation in skin, including in genes implicating skin disease such as ALOX12 and CSPG4. We further explored the relevance of skin meQTLs to skin disease and found that skin meQTLs and CpGs under genetic influence were enriched for multiple skin-related genome-wide and epigenome-wide association signals, including for melanoma and psoriasis. Our findings give insights into the regulatory landscape of epigenomic variation in skin.
The objective assessment of habitual (poly)phenol-rich diets in nutritional epidemiology studies remains challenging. This study developed and evaluated the metabolic signature of a (poly)phenol-rich dietary score (PPS) using a targeted metabolomics method comprising 105 representative (poly)phenol metabolites, analyzed in 24 h of urine samples collected from healthy volunteers. The metabolites that were significantly associated with PPS after adjusting for energy intake were selected to establish a metabolic signature using a combination of linear regression followed by ridge regression to estimate penalized weights for each metabolite. A metabolic signature comprising 51 metabolites was significantly associated with adherence to PPS in 24 h urine samples, as well as with (poly)phenol intake estimated from food frequency questionnaires and diaries. Internal and external data sets were used for validation, and plasma, spot urine, and 24 h urine samples were compared. The metabolic signature proposed here has the potential to accurately reflect adherence to (poly)phenol-rich diets, and may be used as an objective tool for the assessment of (poly)phenol intake.
Objectives Systemic lupus erythematosus (SLE) shows a marked female bias in prevalence. X chromosome inactivation (XCI) is the mechanism which randomly silences one X chromosome to equalise gene expression between 46, XX females and 46, XY males. Though XCI is expected to result in a random pattern of mosaicism across tissues, some females display a significantly skewed ratio in immune cells, termed XCI-skew. We tested whether XCI was abnormal in females with SLE and hence contributes to sexual dimorphism. Methods We assayed XCI in whole blood DNA in 181 female SLE cases, 796 female healthy controls and 10 twin pairs discordant for SLE. Using regression modelling and intra-twin comparisons, we assessed the effect of SLE on XCI and combined clinical, cellular and genetic data via a polygenic score to explore underlying mechanisms. Results Accommodating the powerful confounder of age, XCI-skew was reduced in females with SLE compared with controls (p=1.3x10(-5)), with the greatest effect seen in those with more severe disease. Applying an XCI threshold of >80%, we observed XCI-skew in 6.6% of SLE cases compared with 22% of controls. This difference was not explained by differential white cell counts, medication or genetic susceptibility to SLE. Instead, XCI-skew correlated with a biomarker for type I interferon-regulated gene expression. Conclusions These results refute current views on XCI-skew in autoimmunity and suggest, in lupus, XCI patterns of immune cells reflect the impact of disease state, specifically interferon signalling, on the haematopoietic stem cells from which they derive.