PURPOSE:Comprehensive tumor biomarker testing is a fundamental step in the selection of highly effective molecularly driven therapies for a variety of solid tumors. The primary objective of this study was to examine racial differences in biomarker testing and clinical trial participation in the United States using a real-world database.METHODS:Patients in a real-world deidentified database diagnosed with advanced/metastatic non-small-cell lung cancer (NSCLC), metastatic colorectal cancer (CRC), or metastatic breast cancer were eligible. Biomarker testing and clinical trial participation was compared between Black and White racial groups using chi-squared test and stepwise logistic regression controlling for baseline covariates.RESULTS:A total of 23,488 patients met eligibility criteria. Next-generation sequencing (NGS) testing rates differed significantly between White versus Black race before first-line therapy (36.6% v 29.7%, P < .0001) and at any given time (54.7% v 43.8%, P < .0001) in the nonsquamous NSCLC cohort. Similar disparities in NGS testing rates at any time during the study were observed among patients with CRC (White 51.6%; Black 41.8%, P < .0001). No differences were observed in the breast cancer cohort. Patients of Black race were less likely to be treated in a clinical trial in the overall NSCLC cohort when compared with White counterparts (3.9% v 2.1%, P = .0002). A statistically significant relationship between biomarker/NGS testing and clinical trial enrollment was observed in all cohorts (P < .003) after adjusting for covariates.CONCLUSION:In a real-world database, significant disparities in NGS-based testing rates were observed between Black and White races in NSCLC and CRC. NGS and any biomarker testing were both associated with trial enrollment in all cohorts. There is a need for interventions to promote access to comprehensive testing for patients with advanced/metastatic tumors.
9005 Background: Cancer racial disparities may exist at many levels in the health care system, from screening to timely diagnosis and treatments received, as well as clinical trial enrollment. This study investigated differences in black versus white race among patients with NSCLC undergoing biomarker testing and clinical trial enrollment in the US. Methods: This retrospective observational study utilized the Flatiron Health database, which includes longitudinal data of patients with advanced/metastatic NSCLC. Patients were eligible if they had evidence of systemic therapy in the database from 1/1/2017 through 10/30/2020. Descriptive analyses summarized differences by race in biomarker testing and trial enrollment. Multivariable regression examined the relationship between these factors. Results: A total of 14,768 patients were eligible: 9,793 (66.3%) were white and 1,288 (8.7%) were black. 76.4% of white patients and 73.6% of black patients underwent at least one single molecular test or comprehensive genomic analysis (p = 0.03). Next-generation sequencing (NGS) was performed among 50.1% of white patients and 39.8% of black patients (p < 0.0001. Trial participation was observed among 3.9% of white and 1.9% of black patients (p = 0.0002). There was a statistically significant association between race (white vs black) and both biomarker testing (ever vs never) and trial participation (yes vs no) (both p < 0.001, unadjusted chi square). Differences in NGS testing, baseline biomarker testing, and race were retained as statistically significant (p < 0.01) in adjusted regression analyses. The receipt of first-line targeted therapy was comparable between white and black patients (10.2% and 9.2%, respectively, p = 0.24); however, this summary did not consider biomarker test results. First line use of pembrolizumab+carboplatin+pemetrexed was observed among 19.8% of white and 22.6% of black patients; carboplatin+paclitaxel was observed among 16.5% and 18.6%, and single-agent pembrolizumab was observed among 14.8% and 11.5%, respectively. Conclusions: The use of NGS-based testing, which is recommended by the National Comprehensive Cancer Network Clinical Guidelines in Oncology for patients with advanced/metastatic NSCLC, is the most notable disparity among black patients, with more than a 10 percentage-point difference in receipt of this testing versus white counterparts. This may in part contribute to the more than double the rate of participation in clinical trials observed among white patients, as many second line and beyond trials utilize molecular targets as inclusion criteria. While multiple factors are known to impact health care disparities, access to and receipt of appropriate biomarker testing may be an attenable goal in order to ensure equal access to quality care.
125 Background: Racial disparities may exist at many levels in the health care system; in oncology, yet little is known about racial disparities in biomarker testing and clinical trial enrollment among patients with mCRC. This study was designed to explore racial differences in comprehensive biomarker testing and clinical trial enrollment in the US using a large real-world database. Methods: This retrospective observational study utilized the Flatiron Health electronic health records database, which includes longitudinal data from patients diagnosed with mCRC. Patients with mCRC were eligible for this study if they had evidence of systemic therapy from 1/1/2017 through 10/30/2020 and were alive for at least 120 days after metastatic diagnosis. Unadjusted analyses summarized differences in biomarker testing and clinical trial enrollment between White and Black race, adjusted regression analyses were conducted using all baseline variables as covariates. These data are de-identified and are not considered human subjects research in accordance with the US Code of Federal Regulations (45 CFR Part 46). Results: A total of 7,879 patients were eligible: 4,803 (61.0%) were White and 838 (10.6%) were Black. Comprehensive testing by next-generation sequencing (NGS) was received by 51.6% and 41.8% of patients who were White and Black, respectively (p < 0.0001). There was no significant difference in clinical trial participation across all lines of therapy (2.9%, White and 2.9% Black). There was a statistically significant relationship between NGS-based testing and clinical trial enrollment (p < 0.0001), however, race was not identified a moderating factor in this relationship in adjusted regression analyses. The receipt of molecularly-targeted therapy was comparable between both races (11.9% and 9.7% for White and Black, respectively; p = 0.06). Patients received FOLFOX+bevacizumab most commonly in the first line (34.3% White; 40.5% Black), all other regimens were within 2 percentage points between racial groups. Targeted agents were each used by less than 7.4% of the study population. Conclusions: The use of NGS-based testing is significantly different by race in this database. The significant relationship between NGS testing and clinical trial enrollment at any time in the database did not appear to be moderated by race; however, descriptive analyses suggest that the ongoing analyses by line of therapy and considering timing of testing may better quantify these relationships. These data may not be generalizable to the entire US population as they are obtained from a single database that is limited to practices using this EHR system.
e18516 Background: Lack of diverse representation in clinical trials negatively impacts the cancer survival of patients and populations unaccounted for in clinical research. Efforts such as the 1993 NIH Revitalization Act have focused on improving the diversity of trial participants in the US. This retrospective study evaluated the racial distribution of oncology clinical trial participants using data published in clinicaltrials.gov from Jan 2010 through Dec 2020. Methods: I2E of Linguamatics (IQVIA, Inc), a natural language processing software, was used to identify participant race in oncology trials. Data extracted included trial identifier, year of completion, sponsor, cancer type, and race. Studies were limited to academic, cooperative group and government studies headquartered in the US. Clinical trial results were compared to the racial distribution of SEER 2010 data using z-test. Results: Data from 35,686 patients (14,220 enrolled to 236 phase 2 and 21,471 enrolled to 47 phase 3 trials) were available for analysis. A summary by race is provided in the Table, excluding unknown, which represented 8.5% of phase 2 and 3.5% of phase 3 trials. The proportions of white/black patients enrolled to phase 2 and phase 3 trials beginning in 2010-12 were 84.4%/11% and 83.1%/9.9%, respectively (total enrollment 84.9%/9.6%). For trials beginning in 2015-17, white/black enrollment represented 88.5%/8.1% of patients enrolled to phase 2 and 86.4%/10.1% of patients in phase 3 trials. Black patients represented 9.6% of all trial participants, in contrast with the SEER data where 12% of all patients were black (p < 0.001). For lung cancer trials, black participants represented only 7.9% of all trial participants whereas in breast cancer trials, 10.2% of participants were black, versus the SEER data specific to these tumor types (black patients represent 10.9%/11.5% of lung/breast cancer diagnoses between 2013 to 2017, both p < 0.01). Conclusions: This study suggests that over the past decade most races (other than white) have been significantly underrepresented in US oncology clinical trials, and is even more pronounced for black patients with lung cancer. Based on this analysis, there is no evidence that trial enrollment distribution, particularly of white versus black participants, has changed since 2010. Data are limited to the relative lack of studies reporting results that began enrollment after 2017. These findings suggest that the development of new strategies to improve the recruitment of racial minorities to oncology clinical trials are warranted.[Table: see text]
The protocol below describes an in silico method for drug repositioning (drug repurposing). The data source is ClinicalTrials.gov , which contains about a quarter of a million clinical studies. Mining such rich and clean clinical summary data could be helpful to many health-related researches. Described here is a method that utilizes serious adverse event data to identify potential new uses of drugs and dietary supplements (repositioning).
Drug repositioning (i.e., drug repurposing) is the process of discovering new uses for marketed drugs. Historically, such discoveries were serendipitous. However, the rapid growth in electronic clinical data and text mining tools makes it feasible to systematically identify drugs with the potential to be repurposed. Described here is a novel method of drug repositioning by mining ClinicalTrials.gov. The text mining tools I2E (Linguamatics) and PolyAnalyst (Megaputer) were utilized. An I2E query extracts “Serious Adverse Events” (SAE) data from randomized trials in ClinicalTrials.gov. Through a statistical algorithm, a PolyAnalyst workflow ranks the drugs where the treatment arm has fewer predefined SAEs than the control arm, indicating that potentially the drug is reducing the level of SAE. Hypotheses could then be generated for the new use of these drugs based on the predefined SAE that is indicative of disease (for example, cancer).
OBJECTIVE To identify the Hyperglycemia and Its Effect After Acute Myocardial Infarction on Cardiovascular Outcomes in Patients With Type 2 Diabetes Mellitus (HEART2D) trial subgroups with treatment difference. RESEARCH DESIGN AND METHODS In 1,115 type 2 diabetic patients who had suffered from an acute myocardial infarction (AMI), the HEART2D trial compared two insulin strategies targeting postprandial or fasting/premeal glycemia on time until first cardiovascular event (cardiovascular death, nonfatal MI, nonfatal stroke, coronary revascularization, or hospitalization for acute coronary syndrome). The HEART2D trial ended prematurely for futility. We used the classification and regression tree (CART) to identify baseline subgroups with potential treatment differences. RESULTS CART estimated the age of >65.7 years to best predict the difference in time to first event. In the subgroup aged >65.7 years (prandial, n = 189; basal, n = 210), prandial patients had a significantly longer time to first event and a lower proportion experienced a first event (n = 56 [29.6%] vs. n = 85 [40.5%]; hazard ratio 0.69 [95% CI 0.49–0.96]; P = 0.029), despite similar A1C levels. CONCLUSIONS Older type 2 diabetic AMI survivors may have a lower risk for a subsequent cardiovascular event with insulin targeting postprandial versus fasting/premeal glycemia.
Previous research demonstrated the use of evolutionary computation for the discovery of transcription factor binding sites (TFBS) in promoter regions upstream of coexpressed genes. However, it remained unclear whether or not composite TFBS elements, commonly found in higher organisms where two or more TFBSs form functional complexes, could also be identified by using this approach. Here, we present an important refinement of our previous algorithm and test the identification of composite elements using NFAT/AP-1 as an example. We demonstrate that by using appropriate existing parameters such as window size, novel-scoring methods such as central bonusing and methods of self-adaptation to automatically adjust the variation operators during the evolutionary search, TFBSs of different sizes and complexity can be identified as top solutions. Some of these solutions have known experimental relationships with NFAT/AP-1. We also indicate that even after properly tuning the model parameters, the choice of the appropriate window size has a significant effect on algorithm performance. We believe that this improved algorithm will greatly augment TFBS discovery.
Studies of gene expression in primary human disease tissue often span several years in order to achieve reasonably large sample sizes and to collect patient clinical information making this data particularly valuable. Due to the lack of a central repository, this data has only been available through disparate and non-publicly accessible sources following publication. We developed disease-to-gene expression mapper (D-GEM) as a publically accessible database and data mining toolbox for microarray data of human primary disease tissue. A statistical pipeline has also been implemented to identify genes over-expressed in disease tissue samples in comparison with normal control samples, or genes whose expression values are associated with clinical parameters such as patient survival rate. One potential application of this data is the identification of pathway specific cancer prognosis markers. By applying a novel, gene signatures for cancer prognosis in the context of known biological pathways in cancer development were identified and confirmed.
Gene expression patterns can reflect gene regulations in human tissues under normal or pathologic conditions. Gene expression profiling data from studies of primary human disease samples are particularly valuable since these studies often span many years in order to collect patient clinical information and achieve a large sample size. Disease-to-Gene Expression Mapper (DGEM) provides a beneficial community resource to access and analyze these data; it currently includes Affymetrix oligonucleotide array datasets for more than 40 human diseases and 1400 samples. The data are normalized to the same scale and stored in a relational database. A statistical-analysis pipeline was implemented to identify genes abnormally expressed in disease tissues or genes whose expressions are associated with clinical parameters such as cancer patient survival. Data-mining results can be queried through a web-based interface at http://dgem.dhcp.iupui.edu/. The query tool enables dynamic generation of graphs and tables that are further linked to major gene and pathway resources that connect the data to relevant biology, including Entrez Gene and Kyoto Encyclopedia of Genes and Genomes (KEGG). In summary, DGEM provides scientists and physicians a valuable tool to study disease mechanisms, to discover potential disease biomarkers for diagnosis and prognosis, and to identify novel gene targets for drug discovery. The source code is freely available for non-profit use, on request to the authors.
Proprotein convertase subtilisin/kexin type 9 (PCSK9) is the most recently identified member of the proprotein convertase family. Genetic and cell biology studies have suggested a critical role of PCSK9 in regulating low-density lipoprotein receptor (LDLR) protein levels and thus modulating plasma LDL cholesterol. Recent data on the molecular basis for PCSK9 action support the model in which PCSK9 is self-cleaved, secreted, and tightly bound to the EGF-A repeat of LDLR extracellular domain. PCSK9 binding to LDLR is essential for the ensuing receptor-mediated endocytosis and is speculated to lock LDLR in a specific conformation that favors degradation in lysosomal compartment instead of recycling back to plasma membrane. We report here a novel human PCSK9 splicing variant, which we named PCSK9sv. PCSK9sv had an in-frame deletion of the eighth exon of 58 amino acids and was expressed in multiple tissues, including liver, small intestine, prostate, uterus, brain, and adipose tissue. Unlike wild-type PCSK9, which is secreted, PCSK9sv expressed in human embryonic kidney HEK293 cells failed to process the prosegment intracellularly and thus was not secreted into the medium. Examination of potential functions revealed that PCSK9sv did not change the LDLR protein levels. Two mutations that have been reported in humans with the associated changes in plasma LDL cholesterol were within exon 8, and thus the expression and function of the two mutants were studied. Both N425S and A443T mutants were processed normally, secreted, and reduced LDLR levels. However, the physiological function of this novel splicing variant of PCSK9 has yet to be determined.
To determine cancer pathway activities in nine types of primary tumors and NCI60 cell lines, we applied an in silica approach by examining gene signatures reflective of consequent pathway activation using gene expression data. Supervised learning approaches predicted that the Ras pathway is active in approximately 70% of lung adenocarcinomas but inactive in most squamous cell carcinomas, pulmonary carcinoids, and small cell lung carcinomas. In contrast, the TGF-beta, TNF-alpha, Src, Myc, E2F3, and beta-catenin pathways are inactive in lung adenocarcinomas. We predicted an active Ras, Myc, Src, and/or E2F3 pathway in significant percentages of breast cancer, colorectal carcinoma, and gliomas. Our results also suggest that Ras may be the most prevailing oncogenic pathway. Additionally, many NCI60 cell lines exhibited a gene signature indicative of an active Ras, Myc, and/or Src, but not E2F3, beta-catenin, TNF-alpha, or TGF-beta pathway. To our knowledge, this is the first comprehensive survey of cancer pathway activities in nine major tumor types and the most widely used NCI60 cell lines. The "gene expression pathway signatures" we have defined could facilitate the understanding of molecular mechanisms in cancer development and provide guidance to the selection of appropriate cell lines for cancer research and pharmaceutical compound screening.
Mining the “meaningful” clues from vast amount of expression profiling data remains to be challenge for biologists. After all the statistical tests, biologists often struggle deciding how to do next with a large list of genes without any obvious theme of mechanism, partly because most statistical analyses do not incorporate understanding of biological systems before hand. Here, we developed a novel method of “gene –pair difference within a sample” to identify phenotype-defining gene signatures, based on the hypothesis that a biological state is governed by the relative difference among different biological processes. For gene expression, it is relative difference among the genes within a sample (an individual, cell, etc), the highest frequency of occurrences a gene contributing to the within sample difference underline the contributions of genes in defining the biological states. We tested the method on three datasets, and identified the most important gene-pairs to drive the phenotypic differences.
Background NCI60 cell lines are derived from cancers of 9 tissue origins and have been invaluable in vitro models for cancer research and anti-cancer drug screen. Although extensive studies have been carried out to assess the molecular features of NCI60 cell lines related to cancer and their sensitivities to more than 100,000 chemical compounds, it remains unclear if and how well these cell lines represent or model their tumor tissues of origin. Identification and confirmation of correct origins of NCI60 cell lines are critical to their usage as model systems and to translate in vitro studies into clinical potentials. Here we report a direct comparison between NCI60 cell lines and primary tumors by analyzing global gene expression profiles. Results Comparative analysis suggested that 51 of 59 cell lines we analyzed represent their presumed tumors of origin. Taking advantage of available clinical information of primary tumor samples used to generate gene expression profiling data, we further classified those cell lines with the correct origins into different subtypes of cancer or different stages in cancer development. For example, 6 of 7 non-small cell lung cancer cell lines were classified as lung adenocarcinomas and all of them were classified into late stages in tumor progression. Conclusion Taken together, we developed and applied a novel approach for systematic comparative analysis and integrative classification of NCI60 cell lines and primary tumors. Our results could provide guidance to the selection of appropriate cell lines for cancer research and pharmaceutical compound screenings. Moreover, this gene expression profile based approach can be generally applied to evaluate experimental model systems such as cell lines and animal models for human diseases.
Stromal Cell-derived factor 1 (SDF-1) is a CXC chemokine that binds to the CXCR4 receptor. Recent publication indicates that the SDF-1/CXCR4 signaling pathway plays a pivotal role during development and in many patho-physiological conditions including hematopoiesis, blood vessel formation, cancer metastasis, angiogenesis and HIV infection. Two human SDF-1 isoforms, SDF-1 alpha and SDF-1 beta, have been reported to date. Here we report the identification of four additional human SDF-1 isoforms derived from alternative splicing events, SDF-1 gamma, SDF-1 delta, SDF-1 epsilon and SDF-1 phi. These SDF-1 splice variants all share the same first three exons but contain different fourth exons. The human SDF-1gene spans over 88 kilobase-pairs on chromosome 10. Using the semi-quantitative RT-PCR method, we determined the tissue distribution of these SDF-1 isoforms. SDF-1 alpha and SDF-1 beta share similar expression patterns and the highest expression were detected in liver, pancreas and spleen. SDF-1 gamma seems to be the human orthologue of recently isolated rat SDF-1 gamma, and its expression was only detected in the heart. SDF-1 delta expression can be detected in several adult tissues but the highest expression was detected in fetal liver. When transfected into HEK293 cells, all the SDF-1 isoforms can be detected as secreted proteins in the cell culture media. The conditioned media from transfected cells can stimulate cell migration in a CXCR4-dependent manner. These data suggest that the novel SDF-1splice variants encode functional proteins. (c) 2006 Elsevier B.V. All rights reserved.
Background The tissue expression pattern of a gene often provides an important clue to its potential role in a biological process. A vast amount of gene expression data have been and are being accumulated in public repository through different technology platforms. However, exploitations of these rich data sources remain limited in part due to issues of technology standardization. Our objective is to test the data comparability between SAGE and microarray technologies, through examining the expression pattern of genes under normal physiological states across variety of tissues. Results There are 42–54% of genes showing significant correlations in tissue expression patterns between SAGE and GeneChip, with 30–40% of genes whose expression patterns are positively correlated and 10–15% of genes whose expression patterns are negatively correlated at a statistically significant level (p = 0.05). Our analysis suggests that the discrepancy on the expression patterns derived from technology platforms is not likely from the heterogeneity of tissues used in these technologies, or other spurious correlations resulting from microarray probe design, abundance of genes, or gene function. The discrepancy can be partially explained by errors in the original assignment of SAGE tags to genes due to the evolution of sequence databases. In addition, sequence analysis has indicated that many SAGE tags and Affymetrix array probe sets are mapped to different splice variants or different sequence regions although they represent the same gene, which also contributes to the observed discrepancies between SAGE and array expression data. Conclusion To our knowledge, this is the first report attempting to mine gene expression patterns across tissues using public data from different technology platforms. Unlike previous similar studies that only demonstrated the discrepancies between the two gene expression platforms, we carried out in-depth analysis to further investigate the cause for such discrepancies. Our study shows that the exploitation of rich public expression resource requires extensive knowledge about the technologies, and experiment. Informatic methodologies for better interoperability among platforms still remain a gap. One of the areas that can be improved practically is the accurate sequence mapping of SAGE tags and array probes to full-length genes. Reviewers This article was reviewed by Dr. I. King Jordan, Dr. Joel Bader, and Dr. Arcady Mushegian.
Biological data have accumulated at an unprecedented pace as a result of improvements in molecular technologies. However, the translation of data into information, and subsequently into knowledge, requires the intricate interplay of data access, visualisation and interpretation. Biological data are complex and are organised either hierarchically or non-hierarchically. For non-hierarchically organised data, it is difficult to view relationships among biological facts. In addition, it is difficult to make changes in underlying data storage without affecting the visualisation interface. Here, we demonstrate a platform where non-hierarchically organised data can be visualised through the application of a customised hierarchy incorporating medical subject headings (MeSH) classifications. This platform gives users flexibility in updating and manipulation. It can also facilitate fresh scientific insight by highlighting biological impacts across different hierarchical branches. An example of the integration of biomarker information from the curated Proteome® database (http://www.biobase-International.com/) using MeSH and the StarTree® visualisation tool is presented.
Abstract Background CC-family chemokine receptor 2 (CCR2) is implicated in the trafficking of blood-borne monocytes to sites of inflammation and is implicated in the pathogenesis of several inflammatory diseases such as rheumatoid arthritis, multiple sclerosis and atherosclerosis. The major challenge in the development of small molecule chemokine receptor antagonists is the lack of cross-species activity to the receptor in the preclinical species. Rabbit models have been widely used to study the role of various inflammatory molecules in the development of inflammatory processes. Therefore, in this study, we report the cloning and characterization of rabbit CCR2. Data regarding the activity of the CCR2 antagonist will provide valuable tools to perform toxicology and efficacy studies in the rabbit model. Results Sequence alignment indicated that rabbit CCR2 shares 80 % identity to human CCR2b. Tissue distribution indicated that rabbit CCR2 is abundantly expressed in spleen and lung. Recombinant rabbit CCR2 expressed as stable transfectants in U-937 cells binds radiolabeled 125I-mouse JE (murine MCP-1) with a calculated Kd of 0.1 nM. In competition binding assays, binding of radiolabeled mouse JE to rabbit CCR2 is differentially competed by human MCP-1, -2, -3 and -4, but not by RANTES, MIP-1α or MIP-1β. U-937/rabbit CCR2 stable transfectants undergo chemotaxis in response to both human MCP-1 and mouse JE with potencies comparable to those reported for human CCR2b. Finally, TAK-779, a dual CCR2/CCR5 antagonist effectively inhibits the binding of 125I-mouse JE (IC50 = 2.3 nM) to rabbit CCR2 and effectively blocks CCR2-mediated chemotaxis. Conclusion In this study, we report the cloning of rabbit CCR2 and demonstrate that this receptor is a functional chemotactic receptor for MCP-1.
Transcription factors are key regulatory elements that control gene expression. The TRANSFAC database represents the largest repository for experimentally derived transcription factor binding sites (TFBS). Understanding TFBS, which are typically conserved during evolution, helps us identify genomic regions related to human health and disease, and regions that might be predictive of patient outcomes. Here we present a statistical analysis of all TFBS in the TRANSFAC database. Our analysis suggests that current definition of TFBS core regions in TRANSFAC should be re-examined so as to capture a more precise notion of "cores." We offer insight into more appropriate definitions of TFBS consensus sequences and core regions. These revised definitions provide a better understanding of the nature of transcription factor-DNA binding and assist with developing algorithms for de novo TFBS discovery as well as finding novel variants of known TFBS.