Mass spectrometry (MS) is the technology standard for expression proteomics, but statistical analysis of the resulting data is complicated by the occurrence of missing values. Missing values remain ubiquitous in MS-based proteomics data and are especially frequent in the emerging fields of single-cell and spatial proteomics. The limpa package implements new methods for quantification and differential expression analysis of MS proteomics data, including probabilistic information recovery from missing values. limpa summarises peptide-level data to estimate an expression value for every protein in every sample. Expression values that are supported by fewer detected peptides or involve more missing values are treated as less precise and are downweighted in the differential expression analysis, maximising statistical power while avoiding false discoveries. limpa produces a linear model object suitable for downstream analysis with the limma package, allowing complex experimental designs and other downstream tasks such as the gene ontology or pathway analysis. ### Competing Interest Statement The authors have declared no competing interest. Chan Zuckerberg Initiative (United States)Chan Zuckerberg Initiative (United States), https://ror.org/02qenvm24, 2021-237445 NHMRCNHMRC, , 1058892, 2025645 CSL (Australia)CSL (Australia), https://ror.org/044tc0x05, Translation Data Science Scholarship
BACKGROUND:Imaging-based spatial transcriptomics technologies allow us to explore spatial gene expression profiles at the cellular level. Cell type annotation of imaging-based spatial data is challenging due to the small gene panel, but it is a crucial step for downstream analyses. Many good reference-based cell type annotation tools have been developed for single-cell RNA sequencing and sequencing-based spatial transcriptomics data. However, the performance of the reference-based cell type annotation tools on imaging-based spatial transcriptomics data has not been well studied yet. RESULTS:We compared performance of five reference-based methods (SingleR, Azimuth, RCTD, scPred and scmapCell) with the marker-gene-based manual annotation method on an imaging-based Xenium data of human breast cancer. A practical workflow has been demonstrated for preparing a high-quality single-cell RNA reference, evaluating the accuracy, and estimating the running time for reference-based cell type annotation tools. CONCLUSIONS:SingleR was the best performing reference-based cell type annotation tool for the Xenium platform, being fast, accurate and easy to use, with results closely matching those of manual annotation.
edgeR is an R/Bioconductor software package for differential analyses of sequencing data in the form of read counts for genes or genomic features. Over the past 15 years, edgeR has been a popular choice for statistical analysis of data from sequencing technologies such as RNA-seq or ChIP-seq. edgeR pioneered the use of the negative binomial distribution to model read count data with replicates and the use of generalized linear models to analyze complex experimental designs. edgeR implements empirical Bayes moderation methods to allow reliable inference when the number of replicates is small. This article announces edgeR version 4, which includes new developments across a range of application areas. Infrastructure improvements include support for fractional counts, implementation of model fitting in C and a new statistical treatment of the quasi-likelihood pipeline that improves accuracy for small counts. The revised package has new functionality for differential methylation analysis, differential transcript expression, differential transcript and exon usage, testing relative to a fold-change threshold and pathway analysis. This article reviews the statistical framework and computational implementation of edgeR, briefly summarizing all the existing features and functionalities but with special attention to new features and those that have not been described previously.
Hair follicles cycle through expansion, regression and quiescence. To investigate the role of MCL‑1, a BCL‑2 family protein with anti‑apoptotic and apoptosis‑unrelated functions, we delete Mcl‑1 within the skin epithelium using constitutive and inducible systems. Constitutive Mcl‑1 deletion does not impair hair follicle organogenesis but leads to gradual hair loss and elimination of hair follicle stem cells. Acute Mcl‑1 deletion rapidly depletes activated hair follicle stem cells and completely blocks depilation‑induced hair regeneration in adult mice, while quiescent hair follicle stem cells remain unaffected. Single‑cell RNA‑seq profiling reveals the engagement of P53 and DNA mismatch repair signaling in hair follicle stem cells upon depilation‑induced activation. Trp53 deletion rescues hair regeneration defects caused by acute Mcl‑1 deletion, highlighting a critical interplay between P53 and MCL‑1 in balancing proliferation and death. The ERBB pathway plays a central role in sustaining the survival of adult activated hair follicle stem cells by promoting MCL‑1 protein expression. Remarkably, the loss of a single Bak allele, a pro‑apoptotic Bcl‑2 effector gene, rescues Mcl‑1 deletion‑induced defects in both hair follicles and mammary glands. These findings demonstrate the pivotal role of MCL‑1 in inhibiting proliferation stress‑induced apoptosis when quiescent stem cells activate to fuel tissue regeneration.
Hepatocytes are organized into distinct zonal subsets across the liver lobule, yet their contributions to liver homeostasis and regeneration remain controversial. Here, we developed multiple genetic lineage-tracing mouse models to systematically address this. We found that the liver lobule can be divided into two major zonal and molecular hepatocyte populations marked by Cyp2e1 or Gls2. Pericentral Cyp2e1+ and periportal Gls2+ hepatocytes maintain their own lineage during adult homeostasis, while Cyp2e1+ hepatocytes fuel neonatal liver growth. The Gls2+ and Cyp2e1+ populations can rapidly regenerate one another when one of the populations is severely damaged. Midlobular Ccnd1+ hepatocytes are enriched in the Cyp2e1+ zone in adult liver but have limited contributions to regeneration upon partial hepatectomy and severe pericentral injury. Remarkably, Lgr5+ hepatocytes, a unique Cyp2e1+ subset, contribute significantly to liver replenishment upon periportal injuries. Our findings unravel that zonal hepatocytes mainly self-maintain during homeostasis but exhibit complex plasticity in repair upon injury.
Closely related genes typically display common essential functions but also functional diversification, ensuring retention of both genes throughout evolution. The histone lysine acetyltransferases KAT6A (MOZ) and KAT6B (QKF/MORF), sharing identical protein domain structure, are mutually exclusive catalytic subunits of a multiprotein complex. Mutations in either KAT6A or KAT6B result in congenital intellectual disability disorders in human patients. In mice, loss of function of either gene results in distinct, severe phenotypic consequences. Here we show that, surprisingly, 4-fold overexpression of Kat6b rescues all previously described developmental defects in Kat6a mutant mice, including rescuing the absence of hematopoietic stem cells. Kat6b restores acetylation at histone H3 lysines 9 and 23 and reverses critical gene expression anomalies in Kat6a mutant mice. Our data suggest that the target gene specificity of KAT6A can be substituted by the related paralogue KAT6B, despite differences in amino acid sequence, if KAT6B is expressed at sufficiently high levels.
Loss of the gene encoding the histone acetyltransferase KAT6B (MYST4/MORF/QKF) causes developmental brain abnormalities as well as behavioral and cognitive defects in mice. In humans, heterozygous variants in the KAT6B gene cause two cognitive disorders, Say-Barber-Biesecker-Young-Simpson syndrome (SBBYSS; OMIM:603736) and genitopatellar syndrome (GTPTS; OMIM:606170). Although the effects of KAT6B homozygous and heterozygous mutations have been documented in humans and mice, KAT6B gain-of-function effects have not been reported. Here, we show that overexpression of the Kat6b gene in mice caused aggression, anxiety, and spontaneous epilepsy. Kat6b overexpression led to an increase in histone H3 lysine 9 acetylation and upregulation of genes driving nervous system development and neuronal differentiation. Kat6b overexpression additionally promoted neural stem cell proliferation and favored neuronal over astrocyte differentiation in vivo and in vitro. Our results suggest that, in addition to loss-of-function alleles, gain-of-function KAT6B alleles may be detrimental for brain development.
Fibroblasts form a major component of the stroma in normal mammary tissue and breast tumors. Here, we have applied longitudinal single-cell transcriptome profiling of >45,000 fibroblasts in the mouse mammary gland across five different developmental stages and during oncogenesis. In the normal gland, diverse stromal populations were resolved, including lobular-like fibroblasts, committed preadipocytes and adipogenesis-regulatory, as well as cycling fibroblasts in puberty and pregnancy. These specialized cell types appear to emerge from CD34high mesenchymal progenitor cells, accompanied by elevated Hedgehog signaling. During late tumorigenesis, heterogeneous cancer-associated fibroblasts (CAFs) were identified in mouse models of breast cancer, including a population of CD34- myofibroblastic CAFs (myCAFs) that were transcriptionally and phenotypically similar to senescent CAFs. Moreover, Wnt9a was demonstrated to be a regulator of senescence in CD34- myCAFs. These findings reflect a diverse and hierarchically organized stromal compartment in the normal mammary gland that provides a framework to better understand fibroblasts in normal and cancerous states.
The MYST family histone acetyltransferase gene, KAT6B (MYST4, MORF, QKF) is mutated in two distinct human congenital disorders characterised by intellectual disability, facial dysmorphogenesis and skeletal abnormalities; the Say-Barber-Biesecker-Young-Simpson variant of Ohdo syndrome and Genitopatellar syndrome. Despite its requirement in normal skeletal development, the cellular and transcriptional effects of KAT6B in skeletogenesis have not been thoroughly studied. Here, we show that germline deletion of the Kat6b gene in mice causes premature ossification in vivo, resulting in shortened craniofacial elements and increased bone density, as well as shortened tibias with an expanded pre-hypertrophic layer, as compared to wild type controls. Mechanistically, we show that the loss of KAT6B in mesenchymal progenitor cells promotes transition towards an osteoblast-progenitor state with upregulation of gene targets of RUNX2, a master regulator of osteoblast development and concomitant downregulation of SOX9, a critical gene in chondrocyte development. Moreover, we find that compound heterozygosity at Kat6b and Runx2 loci partially rescues the reduction in ossification of Runx2 heterozygous, but not homozygous mice, suggesting that KAT6B may limit the action of RUNX2, possibly through a role in maintaining progenitors in an undifferentiated state. Moreover, our results show that KAT6B has essential roles in regulating the expression of a large number of genes involved in skeletogenesis and bone development.
Differential transcript usage (DTU) refers to changes in the relative abundance of transcript isoforms of the same gene between experimental conditions, even when the total expression of the gene does not change. DTU analysis requires the quantification of individual isoforms from RNA-seq data, which has a high level of uncertainty due to transcript overlap and read-to-transcript ambiguity (RTA). Popular DTU analysis methods do not directly account for the RTA overdispersion within their statistical frameworks, leading to reduced statistical power or poor error rate control, particularly in scenarios with small sample sizes. This article presents limma and edgeR analysis pipelines that account for RTA during DTU assessment. Leveraging recent advancements in the limma and edgeR Bioconductor packages, we propose DTU analysis pipelines optimized for small and large datasets with a unified interface via the diffSplice function. The pipelines make use of divided counts to remove RTA-induced dispersion from transcript isoform counts and account for the sparsity in transcript-level counts. Simulations and analyses of real data from mouse mammary epithelial cells demonstrate that the diffSplice pipelines provide greater power, improved efficiency, and improved false discovery rate control compared to existing specialized DTU methods.
Antibody production by B cells is essential for protective immunity. The clonal selection theory posits that each mature B cell has a unique immunoglobulin receptor generated through random gene recombination and, when stimulated to differentiate into an antibody-secreting cell, has the capacity to produce only a single antibody specificity. It follows from this ‘one-cell-one-antibody’ dogma that single-cell RNA-seq profiling of antibody-secreting cells should find that each cell expresses only a single form of each of the immunoglobulin heavy and light chains. However, when using GRCh38 as the genome reference, we found that many antibody-secreting cells appeared to express multiple immunoglobulin isotypes. When the newly published T2T-CHM13 genome was used instead as the genome reference, every antibody-secreting cell was found to express a unique isotype, and read mapping quality was also improved. We show that the superior performance of T2T-CHM13 was due to its European origin matching the genetic background of the query samples. On the other hand, T2T-CHM13 failed to appropriately fit the ‘one-cell-one-antibody’ dogma when applied to data derived from East Asia. Our results show that read assignment to human immunoglobulin isotype genes is very sensitive to the ancestral origin of the genome reference.
Heterozygous mutations in the histone lysine acetyltransferase gene KAT6B (MYST4/MORF/QKF) underlie neurodevelopmental disorders, but the mechanistic roles of KAT6B remain poorly understood. Here, we show that loss of KAT6B in embryonic neural stem and progenitor cells (NSPCs) impaired cell proliferation, neuronal differentiation, and neurite outgrowth. Mechanistically, loss of KAT6B resulted in reduced acetylation at histone H3 lysine 9 and reduced expression of key nervous system development genes in NSPCs and the developing cortex, including the SOX gene family, in particular Sox2, which is a key driver of neural progenitor proliferation, multipotency and brain development. In the fetal cortex, KAT6B occupied the Sox2 locus. Loss of KAT6B caused a reduction in Sox2 promoter activity in NSPCs. Sox2 overexpression partially rescued the proliferative defect of Kat6b-/- NSPCs. Collectively, these results elucidate molecular requirements for KAT6B in brain development and identify key KAT6B targets in neural precursor cells and the developing brain.
Inheritance of a BRCA2 pathogenic variant conveys a substantial life-time risk of breast cancer. Identification of the cell(s)-of-origin of BRCA2-mutant breast cancer and targetable perturbations that contribute to transformation remains an unmet need for these individuals who frequently undergo prophylactic mastectomy. Using preneoplastic specimens from age-matched, premenopausal females, here we show broad dysregulation across the luminal compartment in BRCA2mut/+ tissue, including expansion of aberrant ERBB3lo luminal progenitor and mature cells, and the presence of atypical oestrogen receptor (ER)-positive lesions. Transcriptional profiling and functional assays revealed perturbed proteostasis and translation in ERBB3lo progenitors in BRCA2mut/+ breast tissue, independent of ageing. Similar molecular perturbations marked tumours bearing BRCA2-truncating mutations. ERBB3lo progenitors could generate both ER+ and ER- cells, potentially serving as cells-of-origin for ER-positive or triple-negative cancers. Short-term treatment with an mTORC1 inhibitor substantially curtailed tumorigenesis in a preclinical model of BRCA2-deficient breast cancer, thus uncovering a potential prevention strategy for BRCA2 mutation carriers.
The histone lysine acetyltransferase KAT6B (MYST4, MORF, QKF) is the target of recurrent chromosomal translocations causing hematological malignancies with poor prognosis. Using Kat6b germline deletion and overexpression in mice, we determined the role of KAT6B in the hematopoietic system. We found that KAT6B sustained the fetal hematopoietic stem cell pool but did not affect viability or differentiation. KAT6B was essential for normal levels of histone H3 lysine 9 (H3K9) acetylation but not for a previously proposed target, H3K23. Compound heterozygosity of Kat6b and the closely related gene, Kat6a, abolished hematopoietic reconstitution after transplantation. KAT6B and KAT6A cooperatively promoted transcription of genes regulating hematopoiesis, including the Hoxa cluster, Pbx1, Meis1, Gata family, Erg, and Flt3. In conclusion, we identified the hematopoietic processes requiring Kat6b and showed that KAT6B and KAT6A synergistically promoted HSC development, function, and transcription. Our findings are pertinent to current clinical trials testing KAT6A/B inhibitors as cancer therapeutics.
Variants in the poorly characterised oncoprotein, MORC2, a chromatin remodelling ATPase, lead to defects in epigenetic regulation and DNA damage response. The C-terminal domain (CTD) of MORC2, frequently phosphorylated in DNA damage, promotes cancer progression, but its role in chromatin remodelling remains unclear. Here, we report a molecular characterisation of full-length, phosphorylated MORC2, demonstrating its preference for binding open chromatin and functioning as a DNA sliding clamp. We identified a phosphate interacting motif within the CTD that dictates ATP hydrolysis rate and cooperative DNA binding. The DNA binding impacts several structural domains within the ATPase region. We provide the first visual proof that MORC2 induces chromatin remodelling through ATP hydrolysis-dependent DNA compaction, regulated by its phosphorylation state. These findings highlight phosphorylation of MORC2 CTD as a key modulator of chromatin remodelling, presenting it as a potential therapeutic target.
Background The left and right ventricles of the human heart are functionally and developmentally distinct such that genetic or acquired insults can cause dysfunction in one or both ventricles resulting in heart failure. The left ventricle is most clinically relevant in research as its dysfunction is the most dominant cause of heart failure whereby right ventricular involvement can exacerbate the condition. However, the molecular composition of the left ventricular adult human myocardium relative to the right ventricle in health and in heart failure has yet to be thoroughly explored. Methods We performed unbiased quantitative mass spectrometry analyses on the myocardium of pre-mortem cryopreserved non-diseased human hearts to compare the proteome (n = 27) and metabolome (n = 25) between the normal left and right ventricles. We then characterised the proteome and metabolome of the left and right ventricles within end-stage dilated cardiomyopathy (n = 14 and 13) and ischaemic cardiomyopathy (n = 19-17), respectively. All analyses featured a mix of paired and unpaired samples. Intra-condition comparative analyses were performed to identify differences of molecular abundance between the ventricles, and intra-ventricular analyses were performed between sexes of non-diseased hearts. Novel and innovative techniques were used to merge datasets, increasing the sample size and statistical power. KEGG and Gene Ontology databases were used to perform enrichment analyses and inform metabolic trends. Results Constituents of gluconeogenesis, glycolysis, lipogenesis, lipolysis, fatty acid catabolism, the citrate cycle and oxidative phosphorylation were down-regulated in the non-diseased left ventricle, while glycogenesis, pyruvate and ketone metabolism were up-regulated. Inter-ventricular significance of these metabolic pathways was then found to be diminished within end-stage dilated cardiomyopathy and ischaemic cardiomyopathy, while heart failure-associated pathways were increased in the left ventricle relative to the right within ischaemic cardiomyopathy, such as fluid sheer-stress, increased glutamine to glutamate ratio, and down-regulation of contractile proteins, indicating a left ventricular pathological bias. Conclusions The inter-ventricular molecular analyses within this study aides to fill a critical gap in our understanding of the metabolic differences between the human left and right ventricular myocardium and may be used to inform future therapeutic targets for heart failure processes in one or both the ventricles. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Trial Not applicable ### Funding Statement This work was supported by the National Health and Medical Research Council (NHMRC) of Australia and the National Heart Foundation (NHF) of Australia. The contents of the published material are solely the responsibility of the individual authors and do not reflect the view of NHMRC or the NHF. This work was also supported by philanthropic donations to the University of Sydney and by a grant from the R.T. Hall Trust. Metabolomics Workbench raw data availability is supported by NIH grant U2C-DK119886 and OT2-OD030544 grants. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The methods of harvesting, storage and use of donated human myocardium were approved by the Human Research Ethics Committee at The University of Sydney (USYD 2021/122). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Raw data and experimental conditions are available via online public repositories. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium (<http://proteomecentral.proteomexchange.org>) via the PRIDE[126][1] partner repository with the dataset identifiers PXD014826 for the 2018 dataset and PXD042155 for the 2020 dataset. The 2018 mass spectrometry proteomic raw data files can be linked to this study via Supplementary Table 31. The mass spectrometry metabolomics data have been deposited to the Metabolomics Workbench[127][2] under the project ID PR001684 (<http://dx.doi.org/10.21228/M81H71>) with study IDs ST002716 (2018 dataset) and ST002717 (2020 dataset). [1]: #ref-126 [2]: #ref-127
edgeR is an R/Bioconductor software package for differential analyses of sequencing data in the form of read counts for genes or genomic features. Over the past 15 years, edgeR has been a popular choice for statistical analysis of data from sequencing technologies such as RNA-seq or ChIP-seq. edgeR pioneered the use of the negative binomial distribution to model read count data with replicates and the use of generalized linear models to analyse complex experimental designs. edgeR implements empirical Bayes moderation methods to allow reliable inference when the number of replicates is small. This article announces edgeR version 4, which includes new developments across a range of application areas. Infrastructure improvements include support for fractional counts, implementation of model fitting in C++, and a new statistical treatment of the quasi-likelihood pipeline that improves accuracy for small counts. The revised package has new functionality for differential methylation analysis, differential transcript expression, differential transcript and exon usage, testing relative to a fold-change threshold and pathway analysis. This article reviews the statistical framework and computational implementation of edgeR, briefly summarizing all the existing features and functionalities but with special attention to new features and those that have not been described previously.### Competing Interest StatementThe authors have declared no competing interest.
Differential transcript expression analysis of RNA-seq data is becoming an increasingly popular tool to assess changes in expression of individual transcripts between biological conditions. Software designed for transcript-level differential expression analyses account for the uncertainty of transcript quantification, the read-to-transcript ambiguity (RTA), in statistical analyses via resampling methods. Bootstrap sampling is a popular resampling method that is implemented in the RNA-seq quantification tools kallisto and Salmon. However, bootstrapping is computationally intensive and provides replicate counts with low resolution when the number of sequence reads originating from a gene is low. For lowly expressed genes, bootstrap sampling results in noisy replicate counts for the associated transcripts, which in turn leads to non reproducible and unrealistically high RTA overdispersion for those transcripts. Gibbs sampling is a more efficient and high resolution algorithm implemented in Salmon. Here we leverage the latest developments of edgeR 4.0 to present an improved differential transcript expression analysis pipeline with Salmon’s Gibbs sampling algorithm. The new bias-corrected quasi-likelihood method with adjusted deviances for small counts from edgeR, combined with the efficient Gibbs sampling algorithm from Salmon, provides faster and more accurate DTE analyses of RNA-seq data. Comprehensive simulations and test data show that the presented analysis pipeline is more powerful and efficient than previous differential transcript expression pipelines while providing correct control of the false discovery rate. ### Competing Interest Statement The authors have declared no competing interest.
Börjeson-Forssman-Lehmann syndrome (BFLS) is an X-linked intellectual disability and endocrine disorder caused by pathogenic variants of plant homeodomain finger gene 6 (PHF6). An understanding of the role of PHF6 in vivo in the development of the mammalian nervous system is required to advance our knowledge of how PHF6 mutations cause BFLS. Here, we show that PHF6 protein levels are greatly reduced in cells derived from a subset of patients with BFLS. We report the phenotypic, anatomical, cellular and molecular characterization of the brain in males and females in two mouse models of BFLS, namely loss of Phf6 in the germline and nervous system-specific deletion of Phf6. We show that loss of PHF6 resulted in spontaneous seizures occurring via a neural intrinsic mechanism. Histological and morphological analysis revealed a significant enlargement of the lateral ventricles in adult Phf6-deficient mice, while other brain structures and cortical lamination were normal. Phf6 deficient neural precursor cells showed a reduced capacity for self-renewal and increased differentiation into neurons. Phf6 deficient cortical neurons commenced spontaneous neuronal activity prematurely suggesting precocious neuronal maturation. We show that loss of PHF6 in the foetal cortex and isolated cortical neurons predominantly caused upregulation of genes, including Reln, Nr4a2, Slc12a5, Phip and ZIC family transcription factor genes, involved in neural development and function, providing insight into the molecular effects of loss of PHF6 in the developing brain.