The Sepsis-3 criteria operationalized organ dysfunction using the original Sequential Organ Failure Assessment (SOFA-1) score, which was updated to SOFA-2 in October 2025 to align with modern intensive care unit (ICU) practices. However, the impact of adopting SOFA-2 for sepsis detection under the Sepsis-3 criteria has not yet been evaluated. We conducted a retrospective multicenter cohort study using three large-scale ICU databases from the United States and the Netherlands. Adult patients with suspected infection within 72 h of ICU admission were included. Sepsis was independently identified according to Sepsis-3 criteria, utilizing either the SOFA-1 or SOFA-2 score. We systematically compared diagnostic concordance, the timeliness of sepsis detection, clinical outcomes and predictive performance of prognostic models between the two scoring systems. The primary outcome was ICU mortality, while secondary outcomes included hospital mortality and 28-day survival. The study cohort comprised 74,615 adult patients with suspected infection. The diagnostic concordance of sepsis between SOFA-1 and SOFA-2 reached 89.62
BACKGROUND:Septic shock (SS) is a syndrome with high mortality. Early forewarning and diagnosis of SS, which are critical in reducing mortality, are still challenging in clinical management. OBJECTIVE:We propose a simple and fast risk-stratified forewarning model for SS to help physicians recognize patients in time. Moreover, further insights can be gained from the application of the model to improve our understanding of SS. METHODS:A total of 5125 patients with sepsis from the Medical Information Mart for Intensive Care-IV (MIMIC-IV) database were divided into training, validation, and test sets. In addition, 2180 patients with sepsis from the eICU Collaborative Research Database (eICU) were used as an external validation set. We developed a simplified risk-stratified early forewarning model for SS based on the weight of evidence and logistic regression, which was compared with multi-feature complex models, and clinical characteristics among risk groups were evaluated. RESULTS:Using only vital signs and rapid arterial blood gas test features according to feature importance, we constructed the Septic Shock Risk Predictor (SORP), with an area under the curve (AUC) of 0.9458 in the test set, which is only slightly lower than that of the optimal multi-feature complex model (0.9651). A median forewarning time of 13 hours was calculated for SS patients. 4 distinct risk groups (high, medium, low, and ultralow) were identified by the SORP 6 hours before onset, and the incidence rates of SS in the 4 risk groups in the postonset interval were 88.6% (433/489), 34.5% (262/760), 2.5% (67/2707), and 0.3% (4/1301), respectively. The severity increased significantly with increasing risk in both clinical features and survival. Clustering analysis demonstrated a high similarity of pathophysiological characteristics between the high-risk patients without SS diagnosis (NS_HR) and the SS patients, while a significantly worse overall survival was shown in NS_HR patients. On further exploring the characteristics of the treatment and comorbidities of the NS_HR group, these patients demonstrated a significantly higher incidence of mean blood pressure <65 mmHg, significantly lower vasopressor use and infused volume, and more severe renal dysfunction. The above findings were further validated by multicenter eICU data. CONCLUSIONS:The SORP demonstrated accurate forewarning and a reliable risk stratification capability. Among patients forewarned as high risk, similar pathophysiological phenotypes and high mortality were observed in both those subsequently diagnosed as having SS and those without such a diagnosis. NS_HR patients, overlooked by the Sepsis-3 definition, may provide further insights into the pathophysiological processes of SS onset and help to complement its diagnosis and precise management. The importance of precise fluid resuscitation management in SS patients with renal dysfunction is further highlighted. For convenience, an online service for the SORP has been provided.
Molecular profiling of exosomes is a promising non-invasive approach for cancer diagnosis. However, the rapid, robust, and sensitive detection of exosomal proteins remains challenging. In this study, we develop a dual-modal aptamer sensor (aptasensor) using reversible exosomes isolation and 2-sulfo-acridone (2-SA) mediated signal amplification for the colorimetric and fluorescent detection of exosomal proteins. We use aptamers to specifically recognize and capture exosomal transmembrane proteins, and the captured exosomes load a large amount of 2SA. Then the captured exosomes are released efficiently by disulfide bond-modified complementary DNA strands. The exosomes are lysed, and loaded 2-SA is released. The released 2-SA oxidizes 3,3 ',5,5 '-tetramethylbenzidine (TMB) under 365 nm LED light irradiation at physiological pH condition, changing the solution from colorless to blue, which enables the rapid screening and identification of samples as negative or positive, especially in pointof-care testing (POCT) settings. Furthermore, 2-SA is able to emit a bright fluorescence signal, which can enable the precise quantitative detection of exosomes with a limit of detection of 1.3x105 particles mL-1. Finally, the disulfide bond was reduced by glutathione, resulting in the dissociation between aptamers and their complementary DNA strands and enabling the next cycle of exosomes capture and release. We evaluate the ability of the aptasensor to accurately classify cancers using machine learning algorithms by analyzing different exosomal transmembrane proteins from blood samples. We believe this aptasensor will be a promising tool in the field of cancer diagnosis because it provides a dual-modal readout signal to achieve rapid qualitative screening and quantitative detection.
Background Septic shock (SS) is a highly fatal and heterogeneous syndrome. Identifying distinct clinical phenotypes provides valuable insights into the underlying pathophysiological mechanisms and may help to propose precise clinical management strategies. Methods Latent profile analysis (LPA), a model-based unsupervised method, was used for phenotyping in the MIMIC cohort, and the model was externally independently validated in the eICU and AUMC cohorts. Results Three phenotypes, labeled phenotype I, II, and III, were derived. These phenotypes varied in demographics, clinical features, comorbidities, patterns of organ dysfunction, organ support, and prognosis. Phenotype I, characterized by the most severe organ dysfunction (especially liver), the youngest age, and the highest BMI, had the highest mortality (p < 0.001). Phenotype II, with moderate mortality, was characterized by severe renal injury. In contrast, phenotype III, associated with the oldest age and the fewest comorbidities, exhibited significantly lower mortality. Phenotype I patients had the steepest survival curves and demonstrated an ultra-high risk of death, particularly within the first few days after SS onset. Conclusions The individualized identification of phenotypes is well suited to clinical practice. The three SS phenotypes differed significantly in pathophysiological and clinical outcomes, which are crucial for informing management decisions and prognosis.
Abstract Background Common methods of identification of differentially methylated genes (DMGs) mainly detect differences between case and control groups, which cannot tell whether a gene is differentially methylated in a specific disease sample (first scenario), and are not applicable for the study with no normal control (one-phenotype, second scenario). Also, these methods have low detection capacity at the control-limited (third) scenario. Results we developed a method, termed RankDMG, to analyze DNA methylation data in the three special scenarios. For the individualized DMG analysis, RankDMG showed remarkable performances in simulated and real data, independent of measured platforms. Using DMGs detected by common methods as ‘gold standard’, the DMGs identified by RankDMG using only one-phenotype data were comparable to those detected by common methods using case-control samples. Moreover, even when the number of disease samples reduced to five, RankDMG could also identify disease-related DMGs for control-limited data. Conclusion RankDMG provides a novel tool to dissect the inter-individual heterogeneity of tumor at epigenetic level, and it could analyze the one-phenotype and control-limited methylation data. RankDMG is provided as an open source tool via https://github.com/FunMoy/RankDMG.
Establishing an RNA-associated interaction repository facilitates the system-level understanding of RNA functions. However, as these interactions are distributed throughout various resources, an essential prerequisite for effectively applying these data requires that they are deposited together and annotated with confidence scores. Hence, we have updated the RNA-associated interaction database RNAInter (RNA Interactome Database) to version 4.0, which is freely accessible at http://www.rnainter.org or http://www.rna-society.org/rnainter/. Compared with previous versions, the current RNAInter not only contains an enlarged data set, but also an updated confidence scoring system. The merits of this 4.0 version can be summarized in the following points: (i) a redefined confidence scoring system as achieved by integrating the trust of experimental evidence, the trust of the scientific community and the types of tissues/cells, (ii) a redesigned fully functional database that enables for a more rapid retrieval and browsing of interactions via an upgraded user-friendly interface and (iii) an update of entries to >47 million by manually mining the literature and integrating six database resources with evidence from experimental and computational sources. Overall, RNAInter will provide a more comprehensive and readily accessible RNA interactome platform to investigate the regulatory landscape of cellular RNAs.
Liquid-liquid phase separation (LLPS) partitions cellular contents, underlies the formation of membraneless organelles and plays essential biological roles. To date, most of the research on LLPS has focused on proteins, especially RNA-binding proteins. However, accumulating evidence has demonstrated that RNAs can also function as 'scaffolds' and play essential roles in seeding or nucleating the formation of granules. To better utilize the knowledge dispersed in published literature, we here introduce RNAPhaSep (http://www.rnaphasep.cn), a manually curated database of RNAs undergoing LLPS. It contains 1113 entries with experimentally validated RNA self-assembly or RNA and protein co-involved phase separation events. RNAPhaSep contains various types of information, including RNA information, protein information, phase separation experiment information and integrated annotation from multiple databases. RNAPhaSep provides a valuable resource for exploring the relationship between RNA properties and phase behaviour, and may further enhance our comprehensive understanding of LLPS in cellular functions and human diseases.
Owing to the remarkable heterogeneity of gastric cancer (GC), population-level differentially expressed genes (DEGs) identified using case-control comparison cannot indicate the dysregulated frequency of each DEG in GC. In this work, first, the individual-level DEGs were identified for 1,090 GC tissues without paired normal tissues using the RankComp method. Second, we directly compared the gene expression in a cancer tissue to that in paired normal tissue to identify individual-level DEGs among 448 paired cancer-normal gastric tissues. We found 25 DEGs to be dysregulated in more than 90% of 1,090 GC tissues and also in more than 90% of 448 GC tissues with paired normal tissues. The 25 genes were defined as universal DEGs for GC. Then, we measured 24 paired cancer-normal gastric tissues by RNA-seq to validate them further. Among the universal DEGs, 4 upregulated genes (BGN, E2F3, PLAU, and SPP1) and 1 downregulated gene (UBL3) were found to be cancer genes already documented in the COSMIC or F-Census databases. By analyzing protein-protein interaction networks, we found 12 universally upregulated genes, and we found that their 284 direct neighbor genes were significantly enriched with cancer genes and key biological pathways related to cancer, such as the MAPK signaling pathway, cell cycle, and focal adhesion. The 13 universally downregulated genes and 16 direct neighbor genes were also significantly enriched with cancer genes and pathways related to gastric acid secretion. These universal DEGs may be of special importance to GC diagnosis and treatment targets, and they may make it easier to study the molecular mechanisms underlying GC.
Currently, due to the low quality of RNA caused by degradation or low abundance, the accuracy of gene expression measurements by transcriptome sequencing (RNA-seq) is very challenging for non-research-oriented clinical samples, majority of which are preserved in hospitals or tissue banks worldwide with complete pathological information and follow-up data. Molecular signatures consisting of several genes are rarely applied to such samples. To utilize these resources effectively, 45 stage II non-research-oriented samples which were formalin-fixed paraffin-embedded (FFPE) colorectal carcinoma samples (CRC) using RNA-seq have been analysed. Our results showed that although gene expression measurements were significantly affected, most cancer features, based on the relative expression orderings (REOs) of gene pairs, were well preserved. We then developed two REO-based signatures, which consisted of 136 gene pairs for early diagnosis of CRC, and 4500 gene pairs for predicting post-surgery relapse risk of stage II and III CRC. The performance of our signatures, which included hundreds or thousands of gene pairs, was more robust for non-research-oriented clinical samples, compared to that of two published concise REO-based signatures. In conclusion, REO-based signatures with relatively more gene pairs could be robustly applied to non-research-oriented CRC samples.
It is meaningful to assess the risk of cancer incidence among patients with precancerous colorectal lesions. Comparing the within-sample relative expression orderings (REOs) of colorectal cancer patients measured by multiple platforms with that of normal colorectal tissues, a qualitative transcriptional signature consisting of 1,840 gene pairs was identified in the training data. Within an evaluation dataset of 16 active and 18 inactive (remissive) ulcerative colitis subjects, the median incidence risk score of colorectal carcinoma was 0.6402 in active ulcerative colitis subjects, significantly higher than that in remissive subjects (0.3114). Evaluation of two other independent datasets yielded similar results. Moreover, we found that the score significantly positively correlated with the degree of dysplasia in the case of colorectal adenomas. In the merged dataset, the median incidence risk score was 0.9027 among high-grade adenoma samples, significantly higher than that among low-grade adenomas (0.8565). In summary, the developed incidence risk score could well predict the incidence risk of precancerous colorectal lesions and has value in clinical application.
MOTIVATION:For some specific tissues, such as the heart and brain, normal controls are difficult to obtain. Thus, studies with only a particular type of disease samples (one phenotype) cannot be analyzed using common methods, such as significance analysis of microarrays, edgeR and limma. The RankComp algorithm, which was mainly developed to identify individual-level differentially expressed genes (DEGs), can be applied to identify population-level DEGs for the one-phenotype data but cannot identify the dysregulation directions of DEGs.RESULTS:Here, we optimized the RankComp algorithm, termed PhenoComp. Compared with RankComp, PhenoComp provided the dysregulation directions of DEGs and had more robust detection power in both simulated and real one-phenotype data. Moreover, using the DEGs detected by common methods as the 'gold standard', the results showed that the DEGs detected by PhenoComp using only one-phenotype data were comparable to those identified by common methods using case-control samples, independent of the measurement platform. PhenoComp also exhibited good performance for weakly differential expression signal data.AVAILABILITY AND IMPLEMENTATION:The PhenoComp algorithm is available on the web at https://github.com/XJJ-student/PhenoComp.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
The non-cancerous components in tumor tissues, e.g., infiltrating stromal cells and immune cells, dilute tumor purity and might confound genomic mutation profile analyses and the identification of pathological biomarkers. It is necessary to systematically evaluate the influence of tumor purity. Here, using public gastric cancer samples from The Cancer Genome Atlas (TCGA), we firstly showed that numbers of mutation, separately called by four algorithms, were significant positively correlated with tumor purities (all p < 0.05, Spearman rank correlation). Similar results were also observed in other nine cancers from TCGA. Notably, the result was further confirmed by six in-house samples from two gastric cancer patients and five in-house samples from two colorectal cancer patients with different tumor purities. Furthermore, the metastasis mechanism of gastric cancer may be incorrectly characterized as numbers of mutation and tumor purities of 248 lymph node metastatic (N + M0) samples were both significantly lower than those of 121 non-metastatic (N0M0) samples (p < 0.05, Wilcoxon rank-sum test). Similar phenomena were also observed that tumor purities could confound the analysis of histological subtypes of cancer and the identification of microsatellite instability status (MSI) in both gastric and colon cancer. Finally, we suggested that the higher tumor purity, such as above 70%, rather than 60%, could be better to meet the requirement of mutation calling. In conclusion, the influence of tumor purity on the genomic mutation profile and pathological analyses should be fully considered in the further study.
BACKGROUND AND PURPOSE:Currently, 5-fluorouracil (5-FU)-based adjuvant chemoradiotherapy (ACRT) is a preferred regimen for post-surgery gastric cancer (GC). However, the survival outcome of 5-FU-based ACRT varies greatly among different GC patients. Thus, it is necessary to classify which patients may benefit from 5-FU-based ACRT.MATERIALS AND METHODS:We collected 577 GC and 84 adjacent normal samples for training and 675 GC samples for validation. Based on the within-sample relative expression orderings (REOs) of gene expression levels, reversal gene pairs were selected, and the pairs correlating with overall survival (OS) of GC patients receiving 5-FU-based ACRT were identified as candidates. Finally, an optimized set of candidate gene pairs was selected as a classification signature in training data and validated in validation data.RESULTS:A signature consisting of 34 gene pairs was identified in training data and validated in three independent datasets. The classified low-risk group had better OS than the classified high-risk group. We also analyzed the recurrent free survival or disease free survival (RFS/DFS) of the validation datasets, and the similar results were shown. Furthermore, although the signature was identified based on the OS of GC patients receiving ACRT, it was not a prognostic signature for patients treated with surgery alone, but may be a potential signature for 5-FU-based chemotherapy alone.CONCLUSIONS:The signature can accurately classify GC patients who may benefit from 5-FU-based ACRT, which could aid clinicians in tailoring more effective GC treatments.
For estrogen receptor (ER)-negative breast cancer patients, paclitaxel (P), doxorubicin (A) and cyclophosphamide (C) neoadjuvant chemotherapy (NAC) is the standard therapeutic regimen. Pathologic complete response (pCR) and residual disease (RD) are common surrogate measures of chemosensitivity. After NAC, most patients still have RD; of these, some partially respond to NAC, whereas others show extreme resistance and cannot benefit from NAC but only suffer complications resulting from drug toxicity. Here we developed a qualitative transcriptional signature, based on the within-sample relative expression ordering (REO) of gene pairs, to identify extremely resistant samples to PAC NAC. Using gene expression data for ER-negative breast cancer patients including 113 pCR samples and 137 RD samples from four datasets, we selected 61 gene pairs with reversal REO patterns between the two groups as the resistance signature, denoted as NR61. Samples with more than 37 signature gene pairs that had the same REO patterns within the extremely resistant group were defined as having extreme resistance; otherwise, they were considered responders. In the GSE25055 and GSE25065 dataset, the NR61 signature could correctly identify 44 (97.8%) of the 45 pCR samples and 22 (95.7%) of the 23 pCR samples as responder samples, respectively; it also identified 13 (16.9%) of 77 RD samples and 8 (21.1%) of 38 RD samples as extremely resistant samples, respectively. Survival analysis showed that the distant relapse-free survival (DRFS) time of the 14 extremely resistant cases was significantly shorter than that of the 108 responders (P < 0.01; HR = 3.84; 95% CI = 1.91-7.70) in GSE25055. Similar results were obtained in GSE25065. Moreover, in the integrated data of the two datasets with 94 responders and 21 extremely resistant samples identified from RD patients, the former had significantly longer DRFS than the latter (P < 0.01; HR = 2.22; 95% CI = 1.26-3.90). In summary, our signature could effectively identify patients who completely respond to PAC NAC, as well as cases of extreme resistance, which can assist decision-making on the clinical therapy for these patients.
The heterogeneity of cancer is a big obstacle for cancer diagnosis and treatment. Prioritizing combinations of driver genes that mutate in most patients of a specific cancer or a subtype of this cancer is a promising way to tackle this problem. Here, we developed an empirical algorithm, named PathMG, to identify common and subtype-specific mutated sub-pathways for a cancer. By analyzing mutation data of 408 samples (Lung-data1) for lung cancer, three sub-pathways each covering at least 90% of samples were identified as the common sub-pathways of lung cancer. These sub-pathways were enriched with mutated cancer genes and drug targets and were validated in two independent datasets (Lung-data2 and Lung-data3). Especially, applying PathMG to analyze two major subtypes of lung cancer, lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LSCC), we identified 13 subtype-specific sub-pathways with at least 0.25 mutation frequency difference between LUAD and LSCC samples in Lung-data1, and 12 of the 13 sub-pathways were reproducible in Lung-data2 and Lung-data3. Similar analyses were done for colorectal cancer. Together, PathMG provides us a novel tool to identify potential common and subtype-specific sub-pathways for a cancer, which can provide candidates for cancer diagnoses and sub-pathway targeted treatments.
Background The amount of RNA per cell, namely the transcriptome size, may vary under many biological conditions including tumor. If the transcriptome size of two cells is different, direct comparison of the expression measurements on the same amount of total RNA for two samples can only identify genes with changes in the relative mRNA abundances, i.e., cellular mRNA concentration, rather than genes with changes in the absolute mRNA abundances. Results Our recently proposed RankCompV2 algorithm identify differentially expressed genes (DEGs) through comparing the relative expression orderings (REOs) of disease samples with that of normal samples. We reasoned that both the mRNA concentration and the absolute abundances of these DEGs must have changes in disease samples. In simulation experiments, this method showed excellent performance for identifying DEGs between normal and disease samples with different transcriptome sizes. Through analyzing data for ten cancer types, we found that a significantly higher proportion of the DEGs with absolute mRNA abundance changes overlapped or directly interacted with known cancer driver genes and anti-cancer drug targets than that of the DEGs only with mRNA concentration changes alone identified by the traditional methods. The DEGs with increased absolute mRNA abundances were enriched in DNA damage-related pathways, while DEGs with decreased absolute mRNA abundances were enriched in immune and metabolism associated pathways. Conclusions Both the mRNA concentration and the absolute abundances of the DEGs identified through REOs comparison change in disease samples in comparison with normal samples. In cancers these genes might play more important upstream roles in carcinogenesis.
Background: Previously reported transcriptional signatures for predicting the prognosis of stage I-III bladder cancer (BLCA) patients after surgical resection are commonly based on risk scores summarized from quantitative measurements of gene expression levels, which are highly sensitive to the measurement variation and sample quality and thus hardly applicable under clinical settings. It is necessary to develop a signature which can robustly predict recurrence risk of BLCA patients after surgical resection. Methods: The signature is developed based on the within-sample relative expression orderings (REOs) of genes, which are qualitative transcriptional characteristics of the samples. Results: A signature consisting of 12 gene pairs (12-GPS) was identified in training data with 158 samples. In the first validation dataset with 114 samples, the low-risk group of 54 patients had a significantly better overall survival than the high-risk group of 60 patients (HR = 3.59, 95% CI: 1.34~9.62, p = 6.61 × 10-03). The signature was also validated in the second validation dataset with 57 samples (HR = 2.75 × 1008, 95% CI: 0~Inf, p = 0.05). Comparison analysis showed that the transcriptional differences between the low- and high-risk groups were highly reproducible and significantly concordant with DNA methylation differences between the two groups. Conclusions: The 12-GPS signature can robustly predict the recurrence risk of stage I-III BLCA patients after surgical resection. It can also aid the identification of reproducible transcriptional and epigenomic features characterizing BLCA metastasis.
BACKGROUND:Pluripotent stem cell-derived cardiomyocytes (PSC-CMs) are widely used models for regenerative medicine and disease research. However, PSC-CMs are usually immature in morphology and functionality and the maturity of PSC-CMs could not be determined accurately. In order to reasonably interpret the experimental results obtained by PSC-CMs, it is necessary to evaluate the maturity of PSC-CMs and find the key genes related to maturation.METHODS:Using the gene expression profiles of normal adult cardiac tissue and embryonic stem cell (ESC) samples, we identified gene pairs with identically relative expression orderings (REOs) within adult cardiac tissue but reversely identical in ESCs. Then, for a PSC-CM model, we calculated the maturity score as the percentage of these gene pairs that exhibit the same REOs in adult cardiac tissue. Lastly, the CellComp method was used to identify the maturation-related genes.RESULTS:The maturity score increased gradually from 0.8401 for 18-week fetal cardiac tissue to 0.9997 for adult cardiac tissue. For four human PSC-CM models, the mature scores increased with prolonged culture time but were all below 0.8. The genes involved in energy metabolism, angiogenesis, immunity, and proliferation were dysregulated in the 1-year PSC-CMs compared with adult cardiac tissue.CONCLUSION:We proposed a qualitative transcriptional signature to score the maturity degree of PSC-CMs. This score can reasonably track the maturity of PSC-CMs and be used to compare different PSC-CM culture methods.
FOLFOX (5-fluorouracil, leucovorin and oxaliplatin) is one of the main chemotherapy regimens for colorectal cancer (CRC), but only half of CRC patients respond to this regimen. Using gene expression profiles of 96 metastatic CRC patients treated with FOLFOX, we first selected gene pairs whose within-sample relative expression orderings (REO) were significantly associated with the response to FOLFOX using the exact binomial test. Then, from these gene pairs, we applied an optimization procedure to obtain a subset that achieved the largest F-score in predicting pathological response of CRC to FOLFOX. The REO-based qualitative transcriptional signature, consisting of five gene pairs, was developed in the training dataset consisting of 96 samples with an F-score of 0.90. In an independent test dataset consisting of 25 samples with the response information, an F-score of 0.82 was obtained. In three other independent survival datasets, the predicted responders showed significantly better progression-free survival than the predicted non-responders. In addition, the signature showed a better predictive performance than two published FOLFOX signatures across different datasets and is more suitable for CRC patients treated with FOLFOX than 5-fluorouracil-based signatures. In conclusion, the REO-based qualitative transcriptional signature can accurately identify metastatic CRC patients who may benefit from the FOLFOX regimen.
Currently, using biopsy specimens for the early diagnosis of colorectal cancer (CRC) is not entirely reliable due to insufficient sampling amount and inaccurate sampling location. Thus, it is necessary to develop a signature that can accurately identify patients with CRC under these clinical scenarios. Based on the relative expression orderings of genes within individual samples, we developed a qualitative transcriptional signature to discriminate CRC tissues, including CRC adjacent normal tissues from non-CRC individuals. The signature was validated using multiple microarray and RNA sequencing data from different sources. In the training data, a signature consisting of 7 gene pairs was identified. It was well validated in both biopsy and surgical resection specimens from multiple datasets measured by different platforms. For biopsy specimens, 97.6% of 42 CRC tissues and 94.5% of 163 non-CRC (normal or inflammatory bowel disease) tissues were correctly classified. For surgically resected specimens, 99.5% of 854 CRC tissues and 96.3% of 81 CRC adjacent normal tissues were correctly identified as CRC. Notably, we additionally measured 33 CRC biopsy specimens by the Affymetrix platform and 13 CRC surgical resection specimens, with different proportions of tumor epithelial cells, ranging from 40% to 100%, by the RNA sequencing platform, and all these samples were correctly identified as CRC. The signature can be used for the early diagnosis of CRC, which is also suitable for minimum biopsy specimens and inaccurately sampled specimens, and thus has potential value for clinical application.