Next-generation sequencing requires accuracy, reproducibility, and standardized reference materials. The Sequencing Quality Control (SEQC-2) multicenter studies on paired breast cancer and B cell lines generated extensive genomic datasets, but integrated epigenomic and proteomic references remain limited. Here, we performed Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq), Methyl-seq, RNA sequencing (RNA-seq), and proteomic profiling to establish comprehensive multi-omics reference materials. We identified >7,700 protein groups, with 95% of genes encoding a single peptide isoform. Protein expression from CpG island (CGI)-overlapping transcripts was higher than non-CGI transcripts in both cell lines. Certain SNVs were incorporated into mutated peptides. Chromatin accessibility was regulated by CG density: CG-rich regions showed lower methylation, greater accessibility, and higher gene/protein expression, whereas CG-poor regions exhibited higher methylation, reduced accessibility, and cell line-specific expression patterns. These datasets provide well-defined genomic, epigenomic, transcriptomic, and proteomic characterizations that can serve as benchmarks for validating omics assays and bioinformatics methods, offering a valuable community resource.
Prenatal e-cigarette exposure (PeCE) is increasingly prevalent and has been associated with adverse neurodevelopmental outcomes, yet how maternal vaping perturbs early brain development remains poorly understood. We integrated spatial transcriptomics, snRNA-seq and lipidomics to define neonatal rat brain responses to PeCE in rats at regional and cellular resolution. PeCE induced pronounced spatial heterogeneity in vulnerability, with the striatum exhibiting the strongest developmental transcriptional disruption. PeCE disrupts lipid metabolic homeostasis and induced molecular signatures suggestive of enhance Ca 2+ signaling, dopaminergic responsiveness and region-specific synaptic stress programs. In the striatum, PeCE induced transcriptional programs consistent with altered lipid utilization, silent synapse-like molecular features, suppressed dendritic spine development and delayed D1-medium spiny neuron maturation. PeCE-sensitive genes were enriched in human autism spectrum disorders and neurodevelopment risk loci. Together, these findings identify disrupted metabolic and developmental reprogramming as a central feature of neonatal brain vulnerability to maternal vaping and provide a mechanistic framework linking PeCE to neurodevelopmental disease risk.
Background:Single-cell RNA-sequencing (scRNA-seq) has emerged as a powerful tool for cancer research, enabling in-depth characterization of tumor heterogeneity at the single-cell level. Recently, several scRNA-seq copy number variation (scCNV) inference methods have been developed, expanding the application of scRNA-seq to study genetic heterogeneity in cancer using transcriptomic data. However, the fidelity of these methods has not been investigated systematically. Methods:We benchmarked five commonly used scCNV inference methods: HoneyBADGER, CopyKAT, CaSpER, inferCNV, and sciCNV. We evaluated their performance across four different scRNA-seq platforms using data from our previous multicenter study. We evaluated scCNV performance further using scRNA-seq datasets derived from mixed samples consisting of five human lung adenocarcinoma cell lines and also sequenced tissues from a small cell lung cancer patient and used the data to validate our findings with a clinical scRNA-seq dataset. Results:We found that the sensitivity and specificity of the five scCNV inference methods varied, depending on the selection of reference data, sequencing depth, and read length. CopyKAT and CaSpER outperformed other methods overall, while inferCNV, sciCNV, and CopyKAT performed better than other methods in subclone identification. We found that batch effects significantly affected the performance of subclone identification in mixed datasets in most methods we tested. Conclusion:Our benchmarking study revealed the strengths and weaknesses of each of these scCNV inference methods and provided guidance for selecting the optimal CNV inference method using scRNA-seq data.
A variety of newly developed next-generation sequencing technologies are making their way rapidly into the research and clinical applications, for which accuracy and cross-lab reproducibility are critical, and reference standards are much needed. Our previous multicenter studies under the SEQC-2 umbrella using a breast cancer cell line with paired B-cell line have produced a large amount of different genomic data including whole genome sequencing (Illumina, PacBio, Nanopore), HiC, and scRNA-seq with detailed analyses on somatic mutations, single-nucleotide variations (SNVs), and structural variations (SVs). However, there is still a lack of well-characterized reference materials which include epigenomic and proteomic data. Here we further performed ATAC-seq, Methyl-seq, RNA-seq, and proteomic analyses and provided a comprehensive catalog of the epigenomic landscape, which overlapped with the transcriptomes and proteomes for the two cell lines. We identified >7,700 peptide isoforms, where the majority (95%) of the genes had a single peptide isoform. Protein expression of the transcripts overlapping CGIs were much higher than the protein expression of the non-CGI transcripts in both cell lines. We further demonstrated the evidence that certain SNVs were incorporated into mutated peptides. We observed that open chromatin regions had low methylation which were largely regulated by CG density, where CG-rich regions had more accessible chromatin, low methylation, and higher gene and protein expression. The CG-poor regions had higher repressive epigenetic regulations (higher DNA methylation) and less open chromatin, resulting in a cell line specific methylation and gene expression patterns. Our studies provide well-defined reference materials consisting of two cell lines with genomic, epigenomic, transcriptomic, scRNA-seq and proteomic characterizations which can serve as standards for validating and benchmarking not only on various omics assays, but also on bioinformatics methods. It will be a valuable resource for both research and clinical communities.
With recent developments in the field of tissue engineering, innovative products are seeking access to the market to be applied in patients. Therefore, a framework must be defined in which tissue-engineered products can prove their safety and efficacy. For these products, the challenge to live up to the standards of drug development arises not only with respect to practical implementation but also with respect to legislation and ethics. Within this chapter, an overview is given on the development of the standards for this new class of products and on the status of directives and regulation in this rapidly changing field.
Extracellular vesicles (EVs) carry diverse bioactive components including nucleic acids, proteins, lipids and metabolites that play versatile roles in intercellular and interorgan communication. The capability to modulate their stability, tissue-specific targeting and cargo render EVs as promising nanotherapeutics for treating heart, lung, blood and sleep (HLBS) diseases. However, current limitations in large-scale manufacturing of therapeutic-grade EVs, and knowledge gaps in EV biogenesis and heterogeneity pose significant challenges in their clinical application as diagnostics or therapeutics for HLBS diseases. To address these challenges, a strategic workshop with multidisciplinary experts in EV biology and U.S. Food and Drug Administration (USFDA) officials was convened by the National Heart, Lung and Blood Institute. The presentations and discussions were focused on summarizing the current state of science and technology for engineering therapeutic EVs for HLBS diseases, identifying critical knowledge gaps and regulatory challenges and suggesting potential solutions to promulgate translation of therapeutic EVs to the clinic. Benchmarks to meet the critical quality attributes set by the USFDA for other cell-based therapeutics were discussed. Development of novel strategies and approaches for scaling-up EV production and the quality control/quality analysis (QC/QA) of EV-based therapeutics were recognized as the necessary milestones for future investigations.
Background Accurate detection of somatic mutations is challenging but critical in understanding cancer formation, progression, and treatment. We recently proposed NeuSomatic, the first deep convolutional neural network-based somatic mutation detection approach, and demonstrated performance advantages on in silico data. Results In this study, we use the first comprehensive and well-characterized somatic reference data sets from the SEQC2 consortium to investigate best practices for using a deep learning framework in cancer mutation detection. Using the high-confidence somatic mutations established for a cancer cell line by the consortium, we identify the best strategy for building robust models on multiple data sets derived from samples representing real scenarios, for example, a model trained on a combination of real and spike-in mutations had the highest average performance. Conclusions The strategy identified in our study achieved high robustness across multiple sequencing technologies for fresh and FFPE DNA input, varying tumor/normal purities, and different coverages, with significant superiority over conventional detection approaches in general, as well as in challenging situations such as low coverage, low variant allele frequency, DNA damage, and difficult genomic regions
Although emerging evidence reveals that vaping alters the function of the central nervous system, the effects of maternal vaping on offspring brain development remain elusive. Using a well-established in utero exposure model, we performed single-nucleus ATAC-seq (snATAC-seq) and RNA sequencing (snRNA-seq) on prenatally e-cigarette-exposed rat brains. We found that maternal vaping distorted neuronal lineage differentiation in the neonatal brain by promoting excitatory neurons and inhibiting lateral ganglionic eminence-derived inhibitory neuronal differentiation. Moreover, maternal vaping disrupted calcium homeostasis, induced microglia cell death, and elevated susceptibility to cerebral ischemic injury in the developing brain of offspring. Our results suggest that the aberrant calcium signaling, diminished microglial population, and impaired microglia-neuron interaction may all contribute to the underlying mechanisms by which prenatal e-cigarette exposure impairs neonatal rat brain development. Our findings raise the concern that maternal vaping may cause adverse long-term brain damage to the offspring.
Gene counts matrices extracted from scRNA-seq datasets processed with different pipelines.
Clinical applications of precision oncology require accurate tests that can distinguish true cancer-specific mutations from errors introduced at each step of next-generation sequencing (NGS). To date, no bulk sequencing study has addressed the effects of cross-site reproducibility, nor the biological, technical and computational factors that influence variant identification. Here we report a systematic interrogation of somatic mutations in paired tumor–normal cell lines to identify factors affecting detection reproducibility and accuracy at six different centers. Using whole-genome sequencing (WGS) and whole-exome sequencing (WES), we evaluated the reproducibility of different sample types with varying input amount and tumor purity, and multiple library construction protocols, followed by processing with nine bioinformatics pipelines. We found that read coverage and callers affected both WGS and WES reproducibility, but WES performance was influenced by insert fragment size, genomic copy content and the global imbalance score (GIV; G > T/C > A). Finally, taking into account library preparation protocol, tumor content, read coverage and bioinformatics processes concomitantly, we recommend actionable practices to improve the reproducibility and accuracy of NGS experiments for cancer mutation detection. Recommendations are given on optimal read coverage and selection of calling algorithm to maximize the reproducibility of cancer mutation detection in whole-genome or whole-exome sequencing.
Single-cell RNA sequencing (scRNA-seq) is developing rapidly, and investigators seeking to use this technology are left with a variety of options for both experimental platform and bioinformatics methods. There is an urgent need for scRNA-seq reference datasets for benchmarking of different scRNA-seq platforms and bioinformatics methods. To be broadly applicable, these should be generated from renewable, well characterized reference samples and processed in multiple centers across different platforms. Here we present a benchmark scRNA-seq dataset that includes 20 scRNA-seq datasets acquired either as mixtures or as individual samples from two biologically distinct cell lines for which a large amount of multi-platform whole genome sequencing data are also available. These scRNA-seq datasets were generated from multiple popular platforms across four sequencing centers. We believe the datasets we describe here will provide a resource that meets this need by allowing evaluation of various bioinformatics methods for scRNA-seq analyses, including but not limited to data preprocessing, imputation, normalization, clustering, batch correction, and differential analysis.
The lack of samples for generating standardized DNA datasets for setting up a sequencing pipeline or benchmarking the performance of different algorithms limits the implementation and uptake of cancer genomics. Here, we describe reference call sets obtained from paired tumor–normal genomic DNA (gDNA) samples derived from a breast cancer cell line—which is highly heterogeneous, with an aneuploid genome, and enriched in somatic alterations—and a matched lymphoblastoid cell line. We partially validated both somatic mutations and germline variants in these call sets via whole-exome sequencing (WES) with different sequencing platforms and targeted sequencing with >2,000-fold coverage, spanning 82% of genomic regions with high confidence. Although the gDNA reference samples are not representative of primary cancer cells from a clinical sample, when setting up a sequencing pipeline, they not only minimize potential biases from technologies, assays and informatics but also provide a unique resource for benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses. Tumor–normal paired DNA samples from a breast cancer cell line and a matched lymphoblastoid cell line enable calibration of clinical sequencing pipelines and benchmarking ‘tumor-only’ or ‘matched tumor–normal’ analyses.
Comparing diverse single-cell RNA sequencing (scRNA-seq) datasets generated by different technologies and in different laboratories remains a major challenge. Here we address the need for guidance in choosing algorithms leading to accurate biological interpretations of varied data types acquired with different platforms. Using two well-characterized cellular reference samples (breast cancer cells and B cells), captured either separately or in mixtures, we compared different scRNA-seq platforms and several preprocessing, normalization and batch-effect correction methods at multiple centers. Although preprocessing and normalization contributed to variability in gene detection and cell classification, batch-effect correction was by far the most important factor in correctly classifying the cells. Moreover, scRNA-seq dataset characteristics (for example, sample and cellular heterogeneity and platform used) were critical in determining the optimal bioinformatic method. However, reproducibility across centers and platforms was high when appropriate bioinformatic methods were applied. Our findings offer practical guidance for optimizing platform and software selection when designing an scRNA-seq study.
We characterized two reference samples for NGS technologies: a human triple-negative breast cancer cell line and a matched normal cell line. Leveraging several whole-genome sequencing (WGS) platforms, multiple sequencing replicates, and orthogonal mutation detection bioinformatics pipelines, we minimized the potential biases from sequencing technologies, assays, and informatics. Thus, our “truth sets” were defined using evidence from 21 repeats of WGS runs with coverages ranging from 50X to 100X (a total of 140 billion reads). These “truth sets” present many relevant variants/mutations including 193 COSMIC mutations and 9,016 germline variants from the ClinVar database, nonsense mutations in BRCA1/2 and missense mutations in TP53 and FGFR1. Independent validation in three orthogonal experiments demonstrated a successful stress test of the truth set. We expect these reference materials and “truth sets” to facilitate assay development, qualification, validation, and proficiency testing. In addition, our methods can be extended to establish new fully characterized reference samples for the community.
Background aims. Connective tissue progenitors (CTPs) embody the heterogeneous stem and progenitor cell populations present in native tissue. CTPs are essential to the formation and remodeling of connective tissue and represent key targets for tissue-engineering and cell-based therapies. To better understand and characterize CTPs, we aimed to compare the (i) concentration and prevalence, (ii) early in vitro biological behavior and (iii) expression of surface-markers and transcription factors among cells derived from marrowspace (MS), trabecular surface (TS), and adipose tissues (AT). Methods. Cancellousbone and subcutaneous-adipose tissues were collected from 8 patients. Cells were isolated and cultured. Colony formation was assayed using Colonyze software based on ASTM standards. Cell concentration ([Cell]), CTP concentration ([CTP]) and CTP prevalence (PCTP) were determined. Attributes of culture-expanded cells were compared based on (i) effective proliferation rate and (ii) expression of surface-markers CD73, CD90, CD105, SSEA-4, SSEA-3, SSEA-1/CD15, Cripto-1, E-Cadherin/CD324, Ep-CAM/CD326, CD146, hyaluronan and transcription factors Oct3/4, Sox-2 and Nanog using flow cytometry. Results. Mean [Cell], [CTP] and P-CTP were significantly different between MS and TS samples (P = 0.03, P = 0.008 and P = 0.0003), respectively. AT-derived cells generated the highest mean total cell yield at day 6 of culture-4-fold greater than TS and more than 40-fold greater than MS per million cells plated. TS colonies grew with higher mean density than MS colonies (290 +/- 11 versus 150 +/- 11 cell per mm(2); P = 0.0002). Expression of classical-mesenchymal stromal cell (MSC) markers was consistently recorded (>95%) from all tissue sources, whereas all the other markers were highly variable. Conclusions. The prevalence and biological potential of CTPs are different between patients and tissue sources and lack variation in classical MSC markers. Other markers are more likely to discriminate differences between cell populations in biological performance. Understanding the underlying reasons for variation in the concentration, prevalence, marker expression and biological potential of CTPs between patients and source tissues and determining the means of managing this variation will contribute to the rational development of cell-based clinical diagnostics and targeted cell-based therapies.
Mimicking developmental events has been proposed as a strategy to engineer tissue constructs for regenerative medicine. However, this approach has not yet been investigated for skeletal tissues. Here, it is demonstrated that ectopic implantation of day-14.5 mouse embryonic long bone anlagen, dissociated into single cells and randomly incorporated in a bioengineered construct, gives rise to epiphyseal growth plate-like structures, bone and marrow, which share many morphological and molecular similarities to epiphyseal units that form after transplanting intact long bone anlage, demonstrating substantial robustness and autonomy of complex tissue self-assembly and the overall organogenesis process. In vitro studies confirm the self-aggregation and patterning capacity of anlage cells and demonstrate that the model can be used to evaluate the effects of large and small molecules on biological behaviour. These results reveal the preservation of self-organizing and self-patterning capacity of anlage cells even when disconnected from their developmental niche and subjected to system perturbations such as cellular dissociation. These inherent features make long bone anlage cells attractive as a model system for tissue engineering technologies aimed at creating constructs that have the potential to self-assemble and self-pattern complex architectural structures.
The matricellular protein SMOC (Secreted Modular Calcium binding protein) is conserved phylogenetically from vertebrates to arthropods. We showed previously that SMOC inhibits bone morphogenetic protein (BMP) signaling downstream of its receptor via activation of mitogen-activated protein kinase (MAPK) signaling. In contrast, the most prominent effect of the Drosophila orthologue, pentagone (pent), is expanding the range of BMP signaling during wing patterning. Using SMOC deletion constructs we found that SMOC-∆EC, lacking the extracellular calcium binding (EC) domain, inhibited BMP2 signaling, whereas SMOC-EC (EC domain only) enhanced BMP2 signaling. The SMOC-EC domain bound HSPGs with a similar affinity to BMP2 and could expand the range of BMP signaling in an in vitro assay by competition for HSPG-binding. Together with data from studies in vivo we propose a model to explain how these two activities contribute to the function of Pent in Drosophila wing development and SMOC in mammalian joint formation.
In an attempt to identify the cell-associated protein(s) through which SMOC (Secreted Modular Calcium binding protein) induces mitogen-activated protein kinase (MAPK) signaling, the epidermal growth factor receptor (EGFR) became a candidate. However, although in 32D/EGFR cells, the EGFR was phosphorylated in the presence of a commercially available human SMOC-1 (hSMOC-1), only minimal phosphorylation was observed in the presence of Xenopus SMOC-1 (XSMOC-1) or human SMOC-2. Analysis of the commercial hSMOC-1 product demonstrated the presence of pro-EGF as an impurity. When the pro-EGF was removed, only minimal EGFR activation was observed, indicating that SMOC does not signal primarily through EGFR and its receptor remains unidentified. Investigation of SMOC/pro-EGF binding affinity revealed a strong interaction that does not require the C-terminal extracellular calcium-binding (EC) domain of SMOC or the EGF domain of pro-EGF. SMOC does not appear to potentiate or inhibit MAPK signaling in response to pro-EGF, but the interaction could provide a mechanism for retaining soluble pro-EGF at the cell surface.
The U.S. Food and Drug Administration applies regulatory flexibility to balance benefits and risks to subjects in cell-therapy clinical trials.
Background Subtilisin-like Proprotein Convertase 7 (SPC7) is a member of the subtilisin/kexin family of pro-protein convertases. It cleaves many pro-proteins to release their active proteins, including members of the bone morphogenetic protein (BMP) family of signaling molecules. Other SPCs are known to be required during embryonic development but corresponding data regarding SPC7 have not been reported previously. Methodology/Principal Findings We demonstrated that Xenopus SPC7 (SPC7) was expressed predominantly in the developing brain and eye, throughout the neural plate initially, then more specifically in the lens and retina primordia as development progressed. Since no prior functional information has been reported for SPC7, we used gain- and loss-of-function experiments to investigate the possibility that it may also convey patterning or tissue specification information similarly to Furin, SPC4, and SPC6. Overexpression of SPC7 was without effect. In contrast, injection of SPC7 antisense morpholino oligonucleotides (MO) into a single blastomere at the 2- or 4-cell stage produced marked disruption of head structures; anophthalmia was salient. Bilateral injections suppressed head and eye formation completely. In parallel with suppression of eye and brain development by SPC7 knockdown, expression of early anterior neural markers (Sox2, Otx2, Rx2, and Pax6) and late eye-specific markers (β-Crystallin and Opsin), and of BMP target genes such as Tbx2 and Tbx3, was reduced or eliminated. Taken together, these findings suggest a critical role for SPC7–perhaps, at least in part, due to activation of one or more BMPs–in early patterning of the anterior neural plate and its derivatives. Conclusion/Significance SPC7 is required for normal development of the eye and brain, possibly through processing BMPs, though other potential substrates cannot be excluded.