Objectives:Electronic Health Record (EHR) data are increasingly used in cancer research, yet the fidelity of this data when exchanged between systems remains poorly quantified. This study investigated the agreement in essential biomarker data after they are passed from the EHR into the cancer registry and Fast Healthcare Interoperability Resources (FHIR) extracts. Materials and Methods:This single-institution retrospective study compared demographics and 6 biomarkers from 30 lung cancer patients seen between July 2020 and July 2022. Manual review from the EHR served as the gold standard, with concordance tested between the source EHR, Institutional Cancer Registry, and FHIR exports. Results:Demographics showed high concordance across databases. In contrast, biomarker data present in the source EHR were missing in 80%-100% of FHIR extracts. The demographic registry variables were highly concordant. Discussion:This study reports a significant loss in biomarker data availability across real-world data (RWD) sources. Results underscore critical gaps in RWD extraction or exchange methods and highlight risks of relying on RWD without validation.
e13674 Background: Clinical trials are essential for improving cancer treatment; yet 80% of trials fail to meet enrollment timelines, often due to inefficient patient-trial matching processes. To address this, a large language model (LLM)-based clinical trial matching platform was designed to prioritize and evaluate trial matches and facilitate human review. It is systematically validated herein, and real-world utility is explored. Methods: The platform consists of three modules: 1.) trial relevance filters trials to the therapeutic area that best aligns with a patient’s overall profile; 2.) for relevant trials, criterion-level eligibility evaluates individual inclusion and exclusion criteria; and 3.) for each criterion, source evidence retrieval extracts supporting evidence from patient electronic health records (EHRs). Each component was tested independently, and the complete platform is being evaluated in a proof-of-concept study with Atrium Health (analyses 1-4). Datasets per analysis: 1.) Trial relevance was analyzed for 75,000 pairs of patient and clinical trial descriptions. This encompassed 185 patients (18 cancer, and to test generalizability, 167 non-cancer). 2.) Criterion-level eligibility : a.) 2,136 criterion-level eligibility decisions were labeled in triplicate, for 359 criteria across 6 cancer patient EHRs from HealthVerity (HV dataset). b.) Generalizability testing leveraged the n2c2 2018 cohort selection dataset. This included adjudication for 13 criteria across 288 diabetes patients. 3.) Source evidence retrieval : on the HV dataset, experts assessed the accuracy of 2,500 pieces of EHR evidence retrieved by the system. 4.) Proof-of-concept : trial relevance and eligibility are being validated against real-world enrollment of 58 trials and 1053 cancer patients from Atrium Health. Results: Relevant trials were returned with sensitivities of 92% and 90% for cancer and non-cancer patients, respectively, and a specificity of 95% (analysis 1). For analysis 2a, the platform maintained criterion-level accuracy within the range of three experts (93%; experts: 92-97%), while selectively seeking human input in 12% of evaluations (experts: 1-3%). Criterion-level accuracy on the n2c2 dataset was 91%, and the system sought human input in 0.1% of evaluations (analysis 2b). For each criterion in the HV dataset, it retrieved source evidence with a mean accuracy of 94% (analysis 3). Proof-of-concept data is expected to be available in time for presentation. Conclusions: A clinical trial matching platform efficiently identified relevant trials and assessed criterion-level patient eligibility with expert accuracy. It reliably surfaced source evidence for improved provenance and interpretability, and these results will soon be validated on real-world cancer trial enrollment data. Such a system could alleviate staff burden, enhance trial efficiency, and democratize trial access for patients.
INTRODUCTION:Stroke is a leading cause of morbidity and mortality, particularly in older adults. Identifying lifestyle factors, such as physical activity (PA), that mitigate stroke risk is critical for stroke prevention, especially in postmenopausal women. We sought to determine the association between levels and types of recreational PA and risk of total, ischemic, and hemorrhagic stroke in postmenopausal women. METHODS:We performed a prospective cohort study conducted within the Women's Health Initiative from 1993 to 1998 with a mean follow-up of 8.5 years. We studied a total of 139,871 postmenopausal women aged 50-79 years without prior cardiovascular disease or stroke at enrollment. Cox regression was used to estimate hazard ratios (HRs) and 95% confidence intervals (CIs). Recreational PA was assessed via questionnaire, including total, light, moderate, and vigorous activities and walking. Incident total, ischemic, and hemorrhagic strokes were recored. HRs and 95% CIs were adjusted for sociodemographic, lifestyle, and clinical factors. RESULTS:During follow-up, 4,642 stroke occurred (3,496 ischemic and 728 hemorrhagic). Higher levels of total PA (per 1 SD MET-hr/wk: HR = 0.90, 95% CI: 0.87-0.93), walking (HR = 0.93, 95% CI: 0.90-0.96), and moderate PA (HR = 0.91, 95% CI: 0.88-0.94) were associated with reduced total stroke risk. Similar inverse associations were found for ischemic stroke. Vigorous PA demonstrated a J-shaped association with ischemic stroke, while light PA was not significantly associated with stroke risk. Total (HR = 0.90, 95% CI: 0.83-0.97) and vigorous PA (HR = 0.88, 95% CI: 0.81-0.96) were inversely associated with hemorrhagic stroke. Associations were consistent across subgroups defined by age, race/ethnicity, blood pressure, hormone therapy use, BMI, and dietary intake. CONCLUSION:Increased recreational PA, particularly moderate, with cautious interpretation of vigorous activity due to its J-shaped association and potential risks, is associated with reduced risks of total and ischemic stroke in postmenopausal women. Our findings support promoting PA as a key strategy for stroke prevention in this population.
Deep learning (DL) has gained prominence in healthcare for its ability to facilitate early diagnosis, treatment identification with associated prognosis, and varying patient outcome predictions. However, because of highly variable medical practices and unsystematic data collection approaches, DL can unfortunately exacerbate biases and distort estimates. For example, the presence of sampling bias poses a significant challenge to the efficacy and generalizability of any statistical model. Even with DL approaches, selection bias can lead to inconsistent, suboptimal, or inaccurate model results, especially for underrepresented populations. Therefore, without addressing bias, wider implementation of DL approaches can potentially cause unintended harm. In this paper, we studied a novel method for bias reduction that leverages the frequency domain transformation via the Gerchberg-Saxton and corresponding impact on the outcome from a racio-ethnic bias perspective.
Real-world studies have suggested decreased trastuzumab emtansine (T-DM1) effectiveness in patients with metastatic breast cancer (mBC) who received prior trastuzumab plus pertuzumab (H + P). However, these studies may have been biased toward pertuzumab-experienced patients with more aggressive disease. Using an electronic health record-derived database, patients diagnosed with mBC on/after 1 January 2011 who initiated T-DM1 in any treatment line (primary cohort) or who initiated second-line T-DM1 following first-line H ± P (secondary cohort) from 22 February 2013 to 31 December 2019 were included. The primary outcome was time from index date to next treatment or death (TTNT). In the primary cohort (n = 757), the percentage of patients with prior P increased from 37% to 73% across the study period, while population characteristics and treatment effectiveness measures were generally stable. Among P-experienced patients from the secondary cohort (n = 246), median time from mBC diagnosis to T-DM1 initiation increased from 10 to 14 months (2013–2019), and median TTNT increased from 4.4 to 10.2 months (2013–2018). Over time, prior H + P prevalence significantly increased with no observable impact on T-DM1 effectiveness. Drug approval timing should be considered when assessing treatment effectiveness within a sequence.
Purpose: Recent clinical trials support de-escalation of adjuvant radiation therapy following lumpectomy in some older women with low-risk HR+ breast cancers planning to take endocrine therapy. The adoption of these find-ings into clinical practice, and the effectiveness of de-escalated therapy in real-world populations, remain under investigation. Materials and methods: We evaluated use of adjuvant radiation therapy and/or endocrine therapy among older women with T1-2 node-negative, HR+ breast cancer in the United States between 2007 and 2011. The study included patients from the Surveillance, Epidemiology and End Results-Medicare linked database and the North Carolina Cancer Information and Population Health Resource database. Results: Radiation therapy was received by 65.5% of patients, with no decrease over time. Older women and those with T2 (compared to T1) tumors were less likely to receive radiation therapy. In propensity-adjusted analyses, both radiation therapy alone (HR 0.75, 95% CI 0.67-0.84) and radiation + endocrine therapy (HR 0.62, 95% CI 0.54-0.69) were associated with significantly lower recurrence risk compared to endocrine therapy alone. Non-adherence to endocrine therapy was common (37%) and similar across groups. With a median follow-up of 48 months (range 13-84), we were not able to detect an association of non-adherence with recurrence risk in endocrine therapy-containing treatment arms. Conclusion: Most older women with stage I HR+ breast cancers continue to receive radiation, at higher rates than patients with node-negative stage II tumors. These findings suggest that while multiple evidence-based treatment options exist in these patients, improvements are needed to ensure that radiation therapy is applied equitably and rationally. (c) 2021 Elsevier Ltd. All rights reserved.
Generating evidence on the use, effectiveness, and safety of new cancer therapies is a priority for researchers, health care providers, payers, and regulators given the rapid pace of change in cancer diagnosis and treatments. The use of real-world data (RWD) is integral to understanding the utilization patterns and outcomes of these new treatments among patients with cancer who are treated in clinical practice and community settings. An initial step in the use of RWD is careful study design to assess the suitability of an RWD source. This pivotal process can be guided by using a conceptual model that encourages predesign conceptualization. The primary types of RWD included are electronic health records, administrative claims data, cancer registries, and specialty data providers and networks. Careful consideration of each data type is necessary because they are collected for a specific purpose, capturing a set of data elements within a certain population for that purpose, and they vary by population coverage and longitudinality. In this review, the authors provide a high-level assessment of the strengths and limitations of each data category to inform data source selection appropriate to the study question. Overall, the development and accessibility of RWD sources for cancer research are rapidly increasing, and the use of these data requires careful consideration of composition and utility to assess important questions in understanding the use and effectiveness of new therapies.
6524 Background: Disparities in health outcomes can be affected by biological factors associated with GA and social determinants of health. These factors can be teased apart using GA data from comprehensive genomic profiling (CGP) in pts with cancer. CGDBs that link EHR data with CGP enable the selection of pts with similar GA. Holding GA constant provides an opportunity to directly study the effects of reported race in health disparities. This study evaluated a published racial disparity (BRCA testing rates in African American [AA] vs White pts with BC) in a population with fixed, similar GA. Methods: The nationwide (US-based) deidentified Flatiron Health and Foundation Medicine (FMI) BC CGDB (Q3 2020) was used. For each pt, GA fractions from 5 geographic ancestry groups (African [AFR]; Admixed American; East Asian; European [EUR]; South Asian) were derived by FMI using an admixture analysis workflow using genes captured in the CGP assay. To focus on BRCA testing in AA vs White pts and find a sufficient population with similar GA but AA or White race, pts with admixture of both EUR and AFR ancestry were selected. The chosen fractions were: Cohort 1=35%-65% AFR and EUR each; Cohort 2=25%-60% AFR and EUR each; Cohort 3=30%-60% AFR. Cohorts overlap but were chosen to increase sample size. In each cohort, documented BRCA testing prevalence, time from diagnosis to BRCA test date, age at BRCA test and overall survival (OS) were compared between races. Other race (OR) and missing race (MR) were also reported. Results: Most pts (4130/6903) in the BC CGDB had ≥75% EUR ancestry; 129 pts had AFR ancestry fractions ≥25% with EUR ancestry >0%. AA pts had the lowest BRCA testing rates (39%, 43%, 44% for Cohorts 1-3, respectively), which were 18%, 10% and 17% lower compared with White pts, respectively (Table). In Cohorts 1-3, AA pts experienced a longer median time between diagnosis and testing (399, 668, 900 days) compared with White pts (93, 667, 106 days). The median age at BRCA test was 16, 9 and 8 years younger in AA pts (49, 47 and 50 years) compared with White pts. Although pts with MR data had the lowest OS compared with the other races within each cohort, the sample size of each arm for all cohorts was too small to make conclusions. Conclusions: This study demonstrated that when holding GA constant, racial disparities persist in BRCA testing patterns and outcome in pts with BC from a CGDB. With increasing availability of linked clinical and genomic data, further exploration of disparities in genetically similar cohorts can provide deeper insight for cancer outcomes and health disparities research.[Table: see text]
IntroductionIncreasingly in pharmacoepidemiology, linking is required to enrich analytic data to more accurately define study populations, enable adjustment for confounding, and improve capture of health outcomes. When creating such novel linked datasets, researchers should consider their suitability to meet research objectives, assess source data completeness and population coverage, and ensure well-defined data governance standards and protections exist. Additionally, while the RECORD-PE guidelines assist in the reporting of studies using observational health data specific to pharmacoepidemiology, they do not address the unique requirements for transparent evaluation and reporting of the data linkage process. Objectives and ApproachWe aimed to 1) provide guidance on data linkage appropriateness and feasibility to plan purposeful and sustainable new linkages that advance pharmacoepidemiological research and 2) generate a checklist with specific recommendations to assist researchers in providing clear and transparent assessment of the linkage process. To develop these guidelines, a working group comprised of members of the International Society of harmacoepidemiology was formed. Recommendations were open for comment by Society members and endorsed by the Society. ResultsGuidance for feasibility assessment was categorized into five domains: (1) research objectives and justification; (2) data quality and completeness; (3) the linkage process; (4) data ownership and governance; and (5) overall value added by linkage. A checklist for evaluation and reporting of data-linkage processes covered five domains including; (1) data sources; (2) linkage variables; (3) linkage methods; (4) linkage results; and (5) linkage evaluation, including validation and verification of the resulting linked data. Conclusion/ImplicationsOur guidelines for data linkage feasibility assessment and reporting can be used to inform the design of sustainable linked data resources and for transparent communication of linkage processes. Together, these guidelines will help various stakeholders to critically assess the potential for bias in research based on linked data and help generate actionable evidence.
519 Background: Outcomes in pts with hepatocellular carcinoma (HCC) vary by epidemiology, degree of hepatic dysfunction and tx. We analyzed the relationship between tx patterns and outcomes to help characterize emerging clinical data in the context of contemporary disease management. Methods: Retrospective observational study of the Flatiron Health de-identified electronic health record–derived database to analyze the relationship between first recorded tx (1tx) and overall survival (OS) in pts diagnosed with HCC (any stage) Jan 2011 to Nov 2018. Tx categories included transplant, resection/SBRT/RFA, TACE/TARE/TAE, tyrosine kinase inhibitor (TKI), cancer immunotherapy (CIT), and others. Descriptive statistics were used to summarize tx distribution and pt characteristics; Kaplan-Meier method was used to estimate OS by tx category. Results: A total of 2134 pts with HCC were categorized by 1tx: transplant (n = 35), resection/SBRT/RFA (n = 408), TACE/TARE/TAE (n = 830), TKI (n = 751), and CIT (n = 20). Pt demographics were generally similar across txs (Table). Overall, pts with HCC had a median OS of 16.6 mo; varying from 71.5 mo in pts receiving transplant to 5.0 mo in pts treated with TKI. Conclusions: Pts receiving systemic tx for HCC have poor prognoses in clinical practice. Despite the limitations of data availability, this study showed a substantial unmet need for more effective HCC tx options. [Table: see text]
The United States has become an epicenter for the coronavirus disease 2019 (COVID-19) pandemic. However, communities have been unequally affected and evidence is growing that social determinants of health may be exacerbating the pandemic. Furthermore, the impact and timing of social distancing at the community level have yet to be fully explored. We investigated the relative associations between COVID-19 mortality and social distancing, sociodemographic makeup, economic vulnerabilities, and comorbidities in 24 counties surrounding 7 major metropolitan areas in the US using a flexible and robust time series modeling approach. We found that counties with poorer health and less wealth were associated with higher daily mortality rates compared to counties with fewer economic vulnerabilities and fewer pre-existing health conditions. Declines in mobility were associated with up to 15% lower mortality rates relative to pre-social distancing levels of mobility, but effects were lagged between 25-30 days. While we cannot estimate causal impact, this study provides insight into the association of social distancing on community mortality while accounting for key community factors. For full transparency and reproducibility, we provide all data and code used in this study.
PURPOSE:The purpose of this paper is to provide guidance on the evaluation of data linkage quality through the development of a checklist for reporting key elements of the linkage process.METHODS:Responding to a call for manuscripts from the International Society for Pharmacoepidemiology (ISPE), a working group including international representation from the academic, industry, and contract research, and regulatory sectors was formed to develop a checklist for evaluation of data linkage performance and reporting data linkage specifically for pharmacoepidemiologic research. This checklist expands on the reporting of studies conducted using observational routinely collected health data specific to pharmacoepidemiology (RECORD-PE) guidelines.RESULTS:A key aspect of data linkage evaluation for pharmacoepidemiology is to articulate how a linkage process was performed and its accuracy in terms of validation and verification of the resulting linked data. This study generates a checklist, which covers domains including data sources, linkage variables, linkage methods, linkage results, and linkage evaluation. For each domain, specific recommendations provide a clear and transparent assessment of the linkage process.CONCLUSIONS:Linking data sources can help to enrich analytic databases to more accurately define study populations, enable adjustment for confounding, and improve the capture of health outcomes. Clear and transparent reporting of data linkage processes will help to increase confidence in the evidence generated from these data by allowing researchers and end users to critically assess the potential for bias owing to the data linkage process.