Background: In complex biological processes, there exists a tipping point (pre-disease state) when the system undergoes a sudden and dramatic shift to a contrasting state. Accurate detection of the pre-disease state is critical for preventive medicine. However, precise detection of the pre-disease state proves challenging due to the clinical single-sample problem. Methods: To address this challenge, in this study, we introduce a novel single-sample pre-disease state detection method based on the change in local network enrichment level. Results: We validated the proposed method on five independent real datasets, including one influenza virus infection time-course dataset and four tumor datasets. Experimental results confirmed that the proposed method can accurately identify the pre-disease state prior to overt disease onset. Further analysis verified key genes identified by the proposed method in pre-disease state are associated with viral infection and immune dysregulation for the influenza dataset, and tumor metastasis for the tumor datasets. Conclusions: These results demonstrate that this method is a robust and biologically interpretable tool for single-sample pre-disease state detection, with great potential for clinical translation in individualized preventive medicine.
Inflammatory bowel disease (IBD) is characterized by immune dysregulation and oxidative stress. The persistence activated proinflammatory myeloid cells are associated with the nonresponse to anti-TNF therapy in IBD patients. Reactive oxygen species (ROS) act as a pivotal pathogenic hub, bridging TNF-driven NF-kappa B activation to NLRP3-inflammasome signaling and fueling a chronic inflammatory microenvironment (IME). Thus, we developed a dual-functional infliximab antibody-Prussian blue nanoparticle conjugate (PB@IFX) that concurrently scavenges ROS and neutralizes TNF. Using a TNBS-induced chronic colitis model integrated with single-cell transcriptomics, we demonstrate that PB@IFX effect reduces inflammatory myeloid cells (Spp1* macrophages, Ly6c* monocytes, neutrophils, and pDCs) and activated fibroblasts, while restoring epithelial integrity and expanding regulatory T cells. Consequently, PB@IFX suppresses the expression of genes encoding proinflammatory mediators and nicotinamide adenine dinucleotide phosphate (NADPH) oxidase genes, and suppresses neutrophil recruitment and neutrophil extracellular trap (NET)-associated transcriptional programs. Furthermore, PB@IFX reprograms macrophage polarization toward a pro-resolving phenotype. The reshaped IME of the intestine supports recovery of microbiome balance, increasing beneficial bacteria and reducing pathogenic bacteria. Importantly, long-term biodistribution and biosafety studies confirmed favorable clearance and biocompatibility. Our findings illustrate a synergistic nanotherapeutic strategy that simultaneously resolves immune dysregulation and oxidative damage while providing mechanistic insights into successful IBD therapy and identifying potential targets for enhanced treatment efficacy.
N-6-methyladenosine (m(6)A) is the most common internal modification in eukaryotic mRNA, essential for post-transcriptional regulation. MeRIP-seq is widely used for m(6)A profiling, but RNA degradation challenges accurate analysis. While the effects of sample handling on RNA are known, the impact of tissue sample storage durations on m(6)A remains unclear. We investigated how sample storage durations (0, 2, 12, 24, and 48 h at room temperature) affect RNA integrity, m(6)A peaks, and transcriptomes in mouse liver. RNA integrity declined with time, reducing m(6)A peak number and reproducibility. Prolonged storage diverged m(6)A profiles, increased unreported peaks, shifted RRACH motifs, and redistributed peaks to intergenic/intronic regions. Integrated data showed opposing changes in m(6)A and expression for some genes, suggesting storage degradation confounds epitranscriptomic interpretation. Sample storage durations are critical for m(6)A accuracy, emphasizing the need for standardized handling in MeRIP-seq studies.
Objectives: To investigate the clinical diagnostic and prognostic value of extrachromosomal circular DNA (eccDNA) in breast cancer, eccDNA profiles were constructed for 81 breast cancer tumor tissues and 33 adjacent non-tumor tissues. Methods: The distribution characteristics of eccDNA across functional genomic elements and repetitive sequences were systematically analyzed. Furthermore, a diagnostic model for differentiating malignant and normal breast tissues, as well as a prognostic prediction model, was developed using a random forest algorithm. Results: EccDNA in breast cancer tissues harbor a higher proportion of functional elements and repetitive sequences, with their annotated genes significantly enriched in tumor- and immune-related pathways. However, no significant differences in eccDNA features were observed across breast cancer subtypes or pathological stages. In the validation cohort, the eccDNA-based diagnostic model achieved an AUC of 0.83, with repetitive elements and enhancer-associated features contributing the most to diagnostic performance. The prognostic model achieved an AUC of 0.78, with repetitive element annotations also showing strong prognostic relevance. Conclusions: These findings highlight the promising potential of eccDNA in the development of precision diagnostics and prognostic systems for breast cancer.
Extrachromosomal circular DNA (eccDNA) has emerged as a potential biomarker for disease due to its stable closed circular structure. However, the diagnostic utility of eccDNA remains underexplored. In this study, we demonstrate that the characteristics of eccDNA associated with genomic repetitive elements change in breast cancer patient tissues and plasma. These changes can serve as signatures for accurate cancer classification. We profiled eccDNA annotated to repeat elements across the genome in tissues and plasma, aggregating each repeat element to the superfamily and subfamily level. Our findings indicate that eccDNA associated with repetitive elements in cancer exhibits regular patterns of enrichment or depletion in specific elements, particularly at the family level. Additionally, these repeat element changes are present in different subtypes of breast cancer, correlated with varying hormone receptor expression. Although there are differences in the landscapes of eccDNA on repetitive elements between cancer tissues and paired plasma, the unique characteristics of eccDNA associated with repetitive sequences in the plasma of cancer patients facilitate better differentiation from normal individuals. These analyses reveal that changes in eccDNA associated with repeat sequences in human cancers can be used as diagnostic biomarkers for cancer patients.
Diffuse large B-cell lymphoma (DLBCL) is a common but difficult to treat type of non Hodgkin lymphoma. At present, platinum nanoparticles with enzyme mimicking activity - platinum nanoenzymes - have been well applied in the treatment of DLBCL disease, but the molecular mechanism of their treatment is still unclear and there is a certain degree of drug resistance. Extrachromosomal circular DNA (eccDNA) elements have been shown to be widely present in eukaryotic genomes and associated with drug resistance. However, the characteristics and biological functions of eccDNA in DLBCL are still unclear. In this study, we investigated the transcriptome changes, eccDNA feature changes, and potential biological functions of eccDNA before and after treatment of A20 cells with platinum nanoenzymes. We identified a total of 7248 eccDNAs, most of which are distributed on each chromosome of mice within 1000bp in size. Bioinformatics analysis showed the detection of 5 important genes, among which Adgre1 was associated with multiple eccDNA differentially expressed genes in subsequent analysis, and may promote tumor progression through multiple mechanisms in drug therapy for cancer. We found that overexpression of Adgre1 may be an important factor in reducing the effect of platinum nanoenzyme treatment on diffuse large B-cell lymphoma cell line A20.
Introduction:Extrachromosomal circular DNA (eccDNA) represents a class of circular DNA molecules derived from chromosomes with diverse roles in disease. Long eccDNAs (typically 1-5 kb) pose detection challenges due to their large size, hindering functional studies. We propose HyenaCircle, a novel deep learning model leveraging large language model and third-generation sequencing data to predict long eccDNA formation. Methods:Full-length eccDNAs within 1-5 kb were identified by FLED algorithm for Nanopore sequencing data, extended by 100-bp flanking sequences, and paired with 20,000 length-matched negative controls from eccDNA-depleted genomic regions. HyenaCircle was built by adapting the pretrained HyenaDNA model with a designed classifier head. The strategies of data augmentation, regularization and class imbalance weighting were applied to increase model robustness. Results:HyenaCircle achieved comparable performance with a validation AUROC of 0.715 and recall of 0.776. It surpassed DNABERT by 5.9% in AUROC and demonstrated stable convergence. Hyperparameter optimization confirmed batch size 16 and learning rate 5 × 10-5 as optimal. The ablation studies revealed flanking sequences are important, as their removal reduced model stability. The model also showed superior stability over the baseline HyenaDNA architecture. Conclusion:HyenaCircle integrated third-generation sequencing data and large language model for long eccDNA prediction, which outperformed the existing model. Our work demonstrates that the HyenaDNA architecture enables effective long-sequence genomic modeling and provides a new insight for eccDNA prediction and identification.
Environmental DNA (eDNA) derived from aquatic vertebrates has recently been employed to estimate species presence. However, the accuracy of these estimations is contingent upon the degradation rate of eDNA. In this study, we introduced the eDNA Integrity Index (eDI) to adjust eDNA concentration for the purpose of estimating carp biomass. The adjusted eDNA concentration is designated as the Biomass Index (BI). We investigated the degradation rate of eDNA through a series of simulation experiments, followed by experiments conducted in tanks and ponds. In all experimental setups, eDNA concentration exhibited a gradual decline, while eDI demonstrated rapid fluctuations following the removal of fish species. Notably, the eDI decreased to nearly zero within two days, whereas eDNA remained detectable for over a month. In experiments involving multiple species raised in conjunction, we observed no significant mutual interference among different species concerning eDI. Furthermore, temperature was determined to have a minimal impact on eDI. Although both eDNA concentration and BI were positively correlated with carp biomass across all experiments, BI exhibited a stronger correlation (R2 > 0.95), was more sensitive to variations in biomass, and provided a more accurate estimate of carp biomass. We successfully applied this methodology to estimate the biomass of carp in a fishpond, demonstrating that precise biomass data can reflect the potential distribution of common carp in natural environments. We present a non-invasive, straightforward, rapid, and accurate method for biomass estimation. SYNOPSIS: We developed an environmental DNA integrity-based method which can estimate species biomass sensitively for species distribution and resource investigations.
Crohn's Disease (CD) is a type of inflammatory bowel disease (IBD) with an unknown etiology, characterized by gastrointestinal immune-related inflammation. Currently, there is no specific drug for treating IBD, and existing treatments only alleviate symptoms with significant side effects that vary among individuals. Numerous studies have shown a close relationship between gut microbiota and Crohn's disease. However, due to the unclear pathogenesis of Crohn's disease, microbial-based therapies have not been widely adopted. In this study, we used a TNBS (2, 4, 6-trinitrobenzene sulfonic acid)-induced rat model of Crohn's disease to treat the model group with Prussian Blue Nanozymes (PB). Based on metagenomic analysis, we examined the diversity, community changes, and species differences in gut microbiota before and after treatment, identifying 20 disease-related species, 16 PB treatment-related species, and significant community changes. This provides insights for further understanding the pathogenesis of Crohn's disease and the therapeutic mechanisms of Prussian Blue Nanozymes.
Considerable interest has focused on cell-free DNA (cfDNA) methylation for noninvasive prenatal diagnosis. We investigated maternal plasma DNA methylation using a simple protocol. Eighty plasma samples from pregnant women were separated for cfDNA methylation determination, including 13 from patients with gestational diabetes mellitus (GDM). Four CpG sites of the LHX3 gene were detected. The results show that methylation-sensitive high-resolution melting (MS-HRM) may be used for semi-quantitative determination of cfDNA methylation. There are extensive cfDNA methylation modifications at different gestational ages and different patterns were observed at different CpG sites. Hypermethylation was observed in the second trimester of pregnancy and GDM. Significant differences were observed in the two CpG sites in LHX3 between those with GDM and in the second trimester. It is suggested that cfDNA methylation and its level may be biomarkers for noninvasive prenatal diagnosis.
Extrachromosomal circular DNAs (eccDNAs) are a unique class of chromosome-originating circular DNA molecules, which are closely linked to oncogene amplification. Due to recent technological advances, particularly in high-throughput sequencing technology, bioinformatics methods based on sequencing data have become primary approaches for eccDNA identification and functional analysis. Currently, eccDNA-relevant databases incorporate previously identified eccDNA and provide thorough functional annotations and predictions, thereby serving as a valuable resource for eccDNA research. In this review, we collected around 20 available eccDNA-associated bioinformatics tools, including identification tools and annotation databases, and summarized their properties and capabilities. We evaluated some of the eccDNA detection methods in simulated data to offer recommendations for future eccDNA detection. We also discussed the current limitations and prospects of bioinformatics methodologies in eccDNA research.
The biogenesis and functions of extrachromosomal circular DNA(eccDNA)have been studied for several decades.However,the heterogeneity of eccDNA is largely ignored.In this study,we purified and sequenced eccDNA and RNA using a method that simultaneously extracts DNA and RNA from cultured cells treated with iron nanoparticles.
ObjectivesSpinal muscular atrophy (SMA) is an autosomal recessive disease that is one of the most common in childhood neuromuscular disorders. Our screenings are more meaningful programs in preventing birth defects, providing a significant resource for healthcare professionals, genetic counselors, and policymakers involved in designing strategies to prevent and manage SMA.MethodWe screened 39,647 participants from 2020 to the present by quantitative real-time PCR, including 7,231 pre-pregnancy participants and 32,416 pregnancy participants, to detect the presence of SMN1 gene EX7 and EX8 deletion in the DNA samples provided by the subjects. To validate the accuracy of our findings, we also utilized the Multiplex Ligation-dependent Probe Amplification (MLPA) to confirm the reliability of screening results obtained by quantitative real-time PCR.ResultAmong the 39,647 participants who were screened, 726 participants were the carriers of SMN1. The overall carrier rate was calculated to be 1.83% (95% confidence interval: 0.86–2.8%). After undergoing screening, a total of 592 pregnancy carriers were provided with genetic counseling and only 503 of their spouses (84.97, 95% confidence interval: 82.09–87.85%) voluntarily underwent SMA screening.ConclusionThis study provides crucial insights into the prevalence and distribution of SMA carriers among the female population. The identification of 726 asymptomatic carriers highlights the necessity of comprehensive screening programs to identify at-risk individuals and ensure appropriate interventions are in place to minimize the impact of SMA-related conditions.
Considering the importance of accurate information of full-length (FL) transcripts in functional analysis, researchers prefer to develop new sequencing methods base on third-generation sequencing (TGS) rather than short-read sequencing. Several...
Background: The cell-free RNA (cf-RNA) of spent embryo medium (SEM) has aroused a concern of academic and clinical researchers for its potential use in non-invasive embryo screening. However, comprehensive characterization of cf-RNA from SEM still presents significant technical challenges, primarily due to the limited volume of SEM. Hence, there is urgently need to a small input liquid volume and ultralow amount of cf-RNA library preparation method to unbiased cf-RNA sequencing from SEM. (75) Result: Here, we report a high sensitivity agarose amplification-based cf-RNA sequencing method (SEM-Acf) for human preimplantation SEM cf-RNA analysis. It is a cf-RNA sequencing library preparation method by adding agarose amplification. The agarose amplification sensitivity (0.005 pg) and efficiency (105.35 %) were increased than that of without agarose addition (0.45 pg and 96.06 %) by - 90 fold and 9.29 %, respectively. Compared with SMART sequencing (SMART-seq), the correlation of gene expression was stronger in different SEM samples by using SEM-Acf. The cf-RNA number of detected and coverage uniformity of 3 ' end were significantly increased. The proportion of 5(' )end adenine, alternative splicing events and short fragments (<400 bp) were increased. It is also found that 4-mer end motifs of cf-RNA fragments was significantly differences between different embryonic stage by day3 spent cleavage medium and day5/6 spent blastocyst medium. (141) Significance: This study established an efficient SEM amplification and library preparation method. Additionally, we successfully described the characterizations of SEM cf-RNA in preimplantation embryo using SEM-Acf, including expression features and fragment lengths. SEM-Acf facilitates the exploration of cf-RNA as a noninvasive embryo screening biomarker, and opens up potential clinical utilities of small input liquid volume and ultralow amount cf-RNA sequencing. (59).
In gene quantification and expression analysis, issues with sample selection and processing can be serious, as they can easily introduce irrelevant variables and lead to ambiguous results. This study aims to investigate the extent and mechanism of the impact of sample selection and processing on ribonucleic acid (RNA) sequencing. RNA from PBMCs and blood samples was investigated in this study. The integrity of this RNA was measured under different storage times. All the samples underwent high-throughput sequencing for comprehensive evaluation. The differentially expressed genes and their potential functions were analyzed after the samples were placed at room temperature for 0h, 4h and 8h, and different feature changes in these samples were also revealed. The sequencing results showed that the differences in gene expression were higher with an increased storage time, while the total number of genes detected did not change significantly. There were five genes showing gradient patterns over different storage times, all of which were protein-coding genes that had not been mentioned in previous studies. The effect of different storage times on seemingly the same samples was analyzed in this present study. This research, therefore, provides a theoretical basis for the long-term consideration of whether sample processing should be adequately addressed.
The hippocampus is an important part of the limbic system in the human brain that has essential roles in spatial navigation and cognitive functions. It is still unknown how gene expression changes in single-cell in different spatial locations of the hippocampus of Parkinson's disease. The purpose of this study was to analyze the gene expression features of single cells in different spatial locations of mouse hippocampus, and to explore the effects of gene expression regulation on learning and memory mechanisms. Here, we obtained 74 single-cell samples from different spatial locations in a mouse hippocampus through microdissection technology, and used single-cell RNA-sequencing and spatial transcriptome sequencing to visualize and quantify the single-cell transcriptome features of tissue sections. The results of differential expression analysis showed that the expression of Sv2b, Neurod6, Grp and Stk32b genes in a hippocampus single cell at different locations was significantly different, and the marker genes of CA1, CA3 and DG subregions were identified. The results of gene function enrichment analysis showed that the up-regulated differentially expressed genes Tubb2a, Eno1, Atp2b1, Plk2, Map4, Pex5l, Fibcd1 and Pdzd2 were mainly involved in neuron to neuron synapse, vesicle-mediated transport in synapse, calcium signaling pathway and neurodegenerative disease pathways, thus affecting learning and memory function. It revealed the transcriptome profile and heterogeneity of spatially located cells in the hippocampus of PD for the first time, and demonstrated that the impaired learning and memory ability of PD was affected by the synergistic effect of CA1 and CA3 subregions neuron genes. These results are crucial for understanding the pathological mechanism of the Parkinson's disease and making precise treatment plans.
Background: To investigate the relationship between dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) radiomic features and the expression activity of hallmark pathways and to develop prediction models of pathway-level heterogeneity for breast cancer (BC) patients. Methods: Two radiogenomic cohorts were analyzed (n = 246). Tumor regions were segmented semiautomatically, and 174 imaging features were extracted. Gene set enrichment analysis (GSEA) and gene set variation analysis (GSVA) were performed to identify significant imaging-pathway associations. Random forest regression was used to predict pathway enrichment scores. Five-fold cross-validation and grid search were used to determine the optimal preprocessing operation and hyperparameters. Results: We identified 43 pathways, and 101 radiomic features were significantly related in the discovery cohort (p-value < 0.05). The imaging features of the tumor shape and mid-to-late post-contrast stages showed more transcriptional connections. Ten pathways relevant to functions such as cell cycle showed a high correlation with imaging in both cohorts. The prediction model for the mTORC1 signaling pathway achieved the best performance with the mean absolute errors (MAEs) of 27.29 and 28.61% in internal and external test sets, respectively. Conclusions: The DCE-MRI features were associated with hallmark activities and may improve individualized medicine for BC by noninvasively predicting pathway-level heterogeneity.
Highly transcribed noncoding elements (HTNEs) are critical noncoding elements with high levels of transcriptional capacity in particular cohorts involved in multiple cellular biological processes. Investigation of HTNEs with persistent aberrant expression in abnormal tissues could be of benefit in exploring their roles in disease occurrence and progression. Breast cancer is a highly heterogeneous disease for which early screening and prognosis are exceedingly crucial. In this study, we developed a HTNE identification framework to systematically investigate HTNE landscapes in breast cancer patients and identified over ten thousand HTNEs. The robustness and rationality of our framework were demonstrated via public datasets. We revealed that HTNEs had significant chromatin characteristics of enhancers and long noncoding RNAs (lncRNAs) and were significantly enriched with RNA-binding proteins as well as targeted by miRNAs. Further, HTNE-associated genes were significantly overexpressed and exhibited strong correlations with breast cancer. Ultimately, we explored the subtype-specific transcriptional processes associated with HTNEs and uncovered the HTNE signatures that could classify breast cancer subtypes based on the properties of hormone receptors. Our results highlight that the identified HTNEs as well as their associated genes play crucial roles in breast cancer progression and correlate with subtype-specific transcriptional processes of breast cancer.
Cell-free DNA molecules are released into the plasma via apoptotic or necrotic events and active release mechanisms, which carry the genetic and epigenetic information of its origin tissues. However, cfDNA is the mixture of various cell fragments, and the efficient enrichment of cfDNA fragments with diagnostic value remains a great challenge for application in the clinical setting. Evidence from recent years shows that cfDNA fragmentomics’ characteristics differ in normal and diseased individuals without the need to distinguish the source of the cfDNA fragments, which makes it a promising novel biomarker. Moreover, cfDNA fragmentomics can identify tissue origins by inferring epigenetic information. Thus, further insights into the fragmentomics of plasma cfDNA shed light on the origin and fragmentation mechanisms of cfDNA during physiological and pathological processes in diseases and enhance our ability to take the advantage of plasma cfDNA as a molecular diagnostic tool. In this review, we focus on the cfDNA fragment characteristics and its potential application, such as fragment length, end motifs, jagged ends, preferred end coordinates, as well as nucleosome footprints, open chromatin region, and gene expression inferred by the cfDNA fragmentation pattern across the genome. Furthermore, we summarize the methods for deducing the tissue of origin by cfDNA fragmentomics.