The limited spatial resolution of mainstream spatial transcriptomic technologies captures transcriptomic mixtures from multiple cells per spot, obscuring crucial single-cell information. While numerous methods leverage single-cell RNA sequencing references to infer cellular composition from ST data, they primarily rely on fixed cell type labels, overlooking the intrinsic hierarchical heterogeneity (subtypes within broad types) of cellular populations and its association with spatial organization. To address this limitation, HIDF, a Hierarchical Iterative Deconvolution Framework is proposed. HIDF progressively resolves cellular heterogeneity from coarse to fine granularity, it employs a hierarchical iterative optimization mechanism guided by the cluster-tree to recover single-cell spatial distributions. This process is further stabilized and enhanced by incorporating dual regularization constraints (spatial neighborhood and cross-level regularization). Comprehensive benchmarking demonstrates that HIDF outperforms existing methods on simulated and real tissue datasets. In addition, HIDF not only reveals cell type distributions consistent with known tissue functions but also uncovers spatially heterogeneous patterns of cell subtypes undetectable by conventional methods.
Spatial transcriptomics technologies enable the generation of gene expression profiles while preserving spatial context, providing the potential for in-depth understanding of spatial-specific tissue heterogeneity. Leveraging gene and spatial data effectively is fundamental to accurately identifying spatial domains in spatial transcriptomics analysis. However, many existing methods have not yet fully exploited the local neighborhood details within spatial information. To address this issue, we introduce SpaGIC, a novel graph-based deep learning framework integrating graph convolutional networks and self-supervised contrastive learning techniques. SpaGIC learns meaningful latent embeddings of spots by maximizing both edge-wise and local neighborhood-wise mutual information of graph structures, as well as minimizing the embedding distance between spatially adjacent spots. We evaluated SpaGIC on seven spatial transcriptomics datasets across various technology platforms. The experimental results demonstrated that SpaGIC consistently outperformed existing state-of-the-art methods in several tasks, such as spatial domain identification, data denoising, visualization, and trajectory inference. Additionally, SpaGIC is capable of performing joint analyses of multiple slices, further underscoring its versatility and effectiveness in spatial transcriptomics research.
With the development of intelligent electric power inspection technology, electric power companies use UAVs, cameras, robots and other equipment for power transmission line inspections, and take a large number of inspection images. In recent years, the image multi-target detection model has been maturely applied in the intelligent inspection of electric power. The current image intelligent detection model mainly relies on a large number of image sample labels. Power companies have a small number of labeled image samples, and a large number of unlabeled. But manual data annotation is very expensive, and cost lots of time and money. Moreover, there will be certain errors in manual annotation, so it is necessary to repeat the annotation and verification to reduce the error. But this will also increase the cost. This article proposes an active learning model production method for electric power inspection image multi-target detection. We select samples for annotation according to some strategies, and select labeled samples for verification. This method can reduce the cost of sample annotation, and improve the model. productivity. Compared with the traditional random labeling strategy, our method can achieve the same accuracy with only 20% of the annotation cost. With 50% of the annotation cost, the mAP index will increase by 12%.
To improve the performance of systolic array processors, this paper designs and implements a novel architecture which supports dynamic dataflows. First, we design three typical systolic array processors, including the output stationary, weight stationary, and input stationary systolic array. Second, we detailedly evaluate the performance of these processors, and find that none of them always perform best in all environments. Finally, based on the characteristics of three different dataflows, this paper designs a novel systolic array processor with dynamic dataflows. The experimental results show that the proposed systolic array processor achieves the best performance for a variety of computing environments.
Efficient peptide and protein identifications from data-independent acquisition mass spectrometric (DIA-MS) data typically rely on a project-specific spectral library with a suitable size. Here, we describe subLib, a computational strategy for optimizing the spectral library for a specific DIA data set based on a comprehensive spectral library, requiring the preliminary analysis of the DIA data set. Compared with the pan-human library strategy, subLib achieved a 41.2% increase in peptide precursor identifications and a 35.6% increase in protein group identifications in a test data set of six colorectal tumor samples. We also applied this strategy to 389 carcinoma samples from 15 tumor data sets: up to a 39.2% increase in peptide precursor identifications and a 19.0% increase in protein group identifications were observed. Our strategy for spectral library size optimization thus successfully proved to deepen the proteome coverages of DIA-MS data.
To address the increasing need for detecting and validating protein biomarkers in clinical specimens, mass spectrometry (MS)-based targeted proteomic techniques, including the selected reaction monitoring (SRM), parallel reaction monitoring (PRM), and massively parallel data-independent acquisition (DIA), have been developed. For optimal performance, they require the fragment ion spectra of targeted peptides as prior knowledge. In this report, we describe a MS pipeline and spectral resource to support targeted proteomics studies for human tissue samples. To build the spectral resource, we integrated common open-source MS computational tools to assemble a freely accessible computational workflow based on Docker. We then applied the workflow to generate DPHL, a comprehensive DIA pan-human library, from 1096 data-dependent acquisition (DDA) MS raw files for 16 types of cancer samples. This extensive spectral resource was then applied to a proteomic study of 17 prostate cancer (PCa) patients. Thereafter, PRM validation was applied to a larger study of 57 PCa patients and the differential expression of three proteins in prostate tumor was validated. As a second application, the DPHL spectral resource was applied to a study consisting of plasma samples from 19 diffuse large B cell lymphoma (DLBCL) patients and 18 healthy control subjects. Differentially expressed proteins between DLBCL patients and healthy control subjects were detected by DIA-MS and confirmed by PRM. These data demonstrate that the DPHL supports DIA and PRM MS pipelines for robust protein biomarker discovery. DPHL is freely accessible at https://www.iprox.org/page/project.html?id=IPX0001400000.
Efficient peptide and protein identification from data-independent acquisition mass spectrometric (DIA-MS) data typically rely on an experiment-specific spectral library with a suitable size. Here, we report a computational strategy for optimizing the spectral library for a specific DIA dataset based on a comprehensive spectral library, which is accomplished by a priori analysis of the DIA dataset. This strategy achieved up to 44.7% increase in peptide identification and 38.1% increase in protein identification in the test dataset of six colorectal tumor samples compared with the comprehensive pan-human library strategy. We further applied this strategy to 389 carcinoma samples from 15 tumor datasets and observed up to 39.2% increase in peptide identification and 19.0% increase in protein identification. In summary, we present a computational strategy for spectral library size optimization to achieve deeper proteome coverage of DIA-MS data.
To answer the increasing need for detecting and validating protein biomarkers in clinical specimens, proteomic techniques are required that support the fast, reproducible and quantitative analysis of large clinical sample cohorts. Targeted mass spectrometry techniques, specifically SRM, PRM and the massively parallel SWATH/DIA technique have emerged as a powerful method for biomarker research. For optimal performance, they require prior knowledge about the fragment ion spectra of targeted peptides. In this report, we describe a mass spectrometric (MS) pipeline and spectral resource to support data-independent acquisition (DIA) and parallel reaction monitoring (PRM) based biomarker studies. To build the spectral resource we integrated common open-source MS computational tools to assemble an open source computational workflow based on Docker. It was then applied to generate a comprehensive DIA pan-human library (DPHL) from 1,096 data dependent acquisition (DDA) MS raw files, and it comprises 242,476 unique peptide sequences from 14,782 protein groups and 10,943 SwissProt-annotated proteins expressed in 16 types of cancer samples. In particular, tissue specimens from patients with prostate cancer, cervical cancer, colorectal cancer, hepatocellular carcinoma, gastric cancer, lung adenocarcinoma, squamous cell lung carcinoma, diseased thyroid, glioblastoma multiforme, sarcoma and diffuse large B-cell lymphoma (DLBCL), as well as plasma samples from a range of hematologic malignancies were collected from multiple clinics in China, the Netherlands and Singapore and included in the resource. This extensive …
Objective: To screen the serum protein molecular markers of postmenopausal osteoporosis by the proteomics analysis using Tandem Mass Tag (TMT) coupled with liquid chromatography-tandem mass spectrometry (LC-MS/MS). Methods: Serum protein samples were recruited from 10 cases of postmenopausal patients with osteoporosis and 10 cases of postmenopausal women without osteoporosis and the high abundance ratios protein was removed, differentiation protein was extracted and labeled with TMT reagent. Then, mass spectrometric detection, data analysis of differentially expressed proteins, and analysis of biological information were carried out. Results: 87 significantly differentially expressed proteins were screened from the differentiated protein expression profile by LC-ESI-MS/MS combined with TMT labeling, including 50 proteins up-regulated and 37 proteins down-regulated. Differentially expressed proteins were analyzed by GO annotation, these proteins are mainly involved in 15 kinds of biological processes, seven kinds of cellular component and six kinds of molecular function. RAB7A, TSP1, GAS6, SPP24 were screened as candidate proteins which were related to the mechanism of bone remodeling of osteoporosis. By STRING10.0 protein interaction network analysis tools, RAB7A, TSP1, GAS6 were located in the center of the interaction network. SPP24 was located at edge of the network, but it is directly related to the protein BMP2 of bone remodeling. Conclusion: These results provide that the proteomics analysis by using TMT combined with LC-ESI-MS/MS was a feasible method for screening the molecular biomarkers. It suggests that RAB7A, TSP1, GAS6 and SPP24 may be a useful biomarker which can be used in diagnosis and treatment of postmenopausal osteoporosis.