ObjectivesBisphosphonates (BPs), a class of drugs used to prevent bone loss, exert multifaceted immuno-modulatory activities. The objective of this study was to assess the association between prior BP-use and COVID-19-related outcomes.MethodsClosed-claims data from Komodo Health were used to identify patients with continuous medical and prescription insurance-enrollment during the study period 1/1/2019-6/30/2020 and no missing demographic information. BP-users, identified as patients with ≥1 BP-claim during the pre-observation period 1/1/2019-2/29/2020, were propensity-score-matched to BP-nonusers based on demographic and clinical characteristics. The association between BP-use and the odds ratio (OR) for SARS-CoV-2 testing, COVID-19 diagnosis, and COVID-19-related hospitalization during the observation period 3/1/2020-6/30/2020 were assessed using multivariable-logistic regression controlling for demographic/clinical characteristics. Sensitivity analyses performed included: restricting to females aged >50 diagnosed with osteoporosis and matching an active-control cohort of non-BP bone-medication-users to BP-users within state/insurance-type; assessing the relationship between use of other preventive-medications (statins, antihypertensives, antidiabetics, antidepressants), including impact of BP-use within users/nonusers of other preventive-medications, on COVID-19-related outcomes; and evaluating the effect of BP-use on positive-control outcomes (acute bronchitis, pneumonia) assessed 7/1/2019-12/31/2019 among BP-users 1/1/2019-6/30/2019 and their matched BP-nonuser-pair.Results7,906,603 patients met core eligibility criteria, of which 450,366 BP-users were identified and matched to 450,366 BP-nonusers. Compared to BP-nonusers, BP-users displayed a lower odds for SARS-CoV-2 testing (OR=0.22;95%CI:0.21-0.23;p<0.001), COVID-19 diagnosis (OR=0.23;95%CI:0.22-0.24;p<0.001), and COVID-19-related hospitalization (OR=0.26;95%CI:0.24-0.29;p<0.001). Consistent results were found comparing BP-users to non-BP bone-medication-users [(SARS-CoV-2 testing (OR=0.28;95%CI:0.23-0.35;p<0.001), COVID-19 diagnosis (OR=0.40;95%CI:0.32-0.49;p<0.001), COVID-19-related hospitalization (OR=0.45;95%CI:0.26-0.75;p=0.003)]. Use of other preventive-medications did not display similarly-consistent effects on COVID-19-related outcomes, though the impact of BP-use was maintained within users/nonusers of other preventive-medications. BP-use was also associated with a decreased odds of medical-services for acute bronchitis (OR=0.23;95%CI:0.22-0.23;p<0.001) or pneumonia (OR=0.32;95%CI:0.31-0.34;p<0.001) in 2019.ConclusionsPrior BP-use was associated with a reduced odds of COVID-19-related outcomes during the initial wave of the pandemic in 2020. ObjectivesBisphosphonates (BPs), a class of drugs used to prevent bone loss, exert multifaceted immuno-modulatory activities. The objective of this study was to assess the association between prior BP-use and COVID-19-related outcomes. Bisphosphonates (BPs), a class of drugs used to prevent bone loss, exert multifaceted immuno-modulatory activities. The objective of this study was to assess the association between prior BP-use and COVID-19-related outcomes. MethodsClosed-claims data from Komodo Health were used to identify patients with continuous medical and prescription insurance-enrollment during the study period 1/1/2019-6/30/2020 and no missing demographic information. BP-users, identified as patients with ≥1 BP-claim during the pre-observation period 1/1/2019-2/29/2020, were propensity-score-matched to BP-nonusers based on demographic and clinical characteristics. The association between BP-use and the odds ratio (OR) for SARS-CoV-2 testing, COVID-19 diagnosis, and COVID-19-related hospitalization during the observation period 3/1/2020-6/30/2020 were assessed using multivariable-logistic regression controlling for demographic/clinical characteristics. Sensitivity analyses performed included: restricting to females aged >50 diagnosed with osteoporosis and matching an active-control cohort of non-BP bone-medication-users to BP-users within state/insurance-type; assessing the relationship between use of other preventive-medications (statins, antihypertensives, antidiabetics, antidepressants), including impact of BP-use within users/nonusers of other preventive-medications, on COVID-19-related outcomes; and evaluating the effect of BP-use on positive-control outcomes (acute bronchitis, pneumonia) assessed 7/1/2019-12/31/2019 among BP-users 1/1/2019-6/30/2019 and their matched BP-nonuser-pair. Closed-claims data from Komodo Health were used to identify patients with continuous medical and prescription insurance-enrollment during the study period 1/1/2019-6/30/2020 and no missing demographic information. BP-users, identified as patients with ≥1 BP-claim during the pre-observation period 1/1/2019-2/29/2020, were propensity-score-matched to BP-nonusers based on demographic and clinical characteristics. The association between BP-use and the odds ratio (OR) for SARS-CoV-2 testing, COVID-19 diagnosis, and COVID-19-related hospitalization during the observation period 3/1/2020-6/30/2020 were assessed using multivariable-logistic regression controlling for demographic/clinical characteristics. Sensitivity analyses performed included: restricting to females aged >50 diagnosed with osteoporosis and matching an active-control cohort of non-BP bone-medication-users to BP-users within state/insurance-type; assessing the relationship between use of other preventive-medications (statins, antihypertensives, antidiabetics, antidepressants), including impact of BP-use within users/nonusers of other preventive-medications, on COVID-19-related outcomes; and evaluating the effect of BP-use on positive-control outcomes (acute bronchitis, pneumonia) assessed 7/1/2019-12/31/2019 among BP-users 1/1/2019-6/30/2019 and their matched BP-nonuser-pair. Results7,906,603 patients met core eligibility criteria, of which 450,366 BP-users were identified and matched to 450,366 BP-nonusers. Compared to BP-nonusers, BP-users displayed a lower odds for SARS-CoV-2 testing (OR=0.22;95%CI:0.21-0.23;p<0.001), COVID-19 diagnosis (OR=0.23;95%CI:0.22-0.24;p<0.001), and COVID-19-related hospitalization (OR=0.26;95%CI:0.24-0.29;p<0.001). Consistent results were found comparing BP-users to non-BP bone-medication-users [(SARS-CoV-2 testing (OR=0.28;95%CI:0.23-0.35;p<0.001), COVID-19 diagnosis (OR=0.40;95%CI:0.32-0.49;p<0.001), COVID-19-related hospitalization (OR=0.45;95%CI:0.26-0.75;p=0.003)]. Use of other preventive-medications did not display similarly-consistent effects on COVID-19-related outcomes, though the impact of BP-use was maintained within users/nonusers of other preventive-medications. BP-use was also associated with a decreased odds of medical-services for acute bronchitis (OR=0.23;95%CI:0.22-0.23;p<0.001) or pneumonia (OR=0.32;95%CI:0.31-0.34;p<0.001) in 2019. 7,906,603 patients met core eligibility criteria, of which 450,366 BP-users were identified and matched to 450,366 BP-nonusers. Compared to BP-nonusers, BP-users displayed a lower odds for SARS-CoV-2 testing (OR=0.22;95%CI:0.21-0.23;p<0.001), COVID-19 diagnosis (OR=0.23;95%CI:0.22-0.24;p<0.001), and COVID-19-related hospitalization (OR=0.26;95%CI:0.24-0.29;p<0.001). Consistent results were found comparing BP-users to non-BP bone-medication-users [(SARS-CoV-2 testing (OR=0.28;95%CI:0.23-0.35;p<0.001), COVID-19 diagnosis (OR=0.40;95%CI:0.32-0.49;p<0.001), COVID-19-related hospitalization (OR=0.45;95%CI:0.26-0.75;p=0.003)]. Use of other preventive-medications did not display similarly-consistent effects on COVID-19-related outcomes, though the impact of BP-use was maintained within users/nonusers of other preventive-medications. BP-use was also associated with a decreased odds of medical-services for acute bronchitis (OR=0.23;95%CI:0.22-0.23;p<0.001) or pneumonia (OR=0.32;95%CI:0.31-0.34;p<0.001) in 2019. ConclusionsPrior BP-use was associated with a reduced odds of COVID-19-related outcomes during the initial wave of the pandemic in 2020. Prior BP-use was associated with a reduced odds of COVID-19-related outcomes during the initial wave of the pandemic in 2020.
To estimate expected collision rate in a large US mortality dataset and specifically examine the relationship between collision rate and sample size. We propose the hypothesize that expected collision rate scales linearly with sample size. The hypothesized relationship between collision rate and sample size was validated on Datavant’s Mortality dataset, which contains over 100 million unique Datavant Tokens (Social Security Number + First Name) on a patient-level basis and is based on data from US government sources. A series of random samples were drawn from the dataset and the number of unique individuals (N), unique Tokens (K), and collisions (C) and the collision rate (C/N) were computed at each sample size. Expected vs. observed collision rate were compared. This same analysis was repeated to examine the collision of combinations of Token 1 (Last Name + First Initial of First Name + Gender + Date of Birth) and Token 2 (Last Name (soundex) + First Name (soundex) + Gender + Date of Birth). The sample size threshold to have at least 1 expected collision was found to be between N = 40,000 and N = 70,000. We observed that as sample size increased, the number of collisions and collision rate correspondingly increased. When Token 1 and Token 2 were used together, the resulting distinct combinations of PII were higher compared with when using Token 2 individually (5.65 billion vs. 1.845 billion using 100% of the dataset sample), and the collision rate was substantially lower. The results indicate that collision rate scales about linearly with sample size. This validation helps to further inform how the false positive rate of token-based matching algorithms may change with sample size.
The objective of this study was to assess the relationship between social determinants of health (SDoH) characteristics, healthcare resource utilization (HCRU), adherence, and persistence among patients with hypertension. Data from the 2019 National Health and Wellness Survey were linked to medical and prescription claims from Komodo Health. Participants age≥18 who self-reported a physician-diagnoses of hypertension and had continuous medical and prescription eligibility 1/1/2019-12/31/2019 were included. HCRU outcomes include physician-office visits, outpatient visits, emergency room (ER) visits, hospitalizations, and antihypertensive utilization including beta-blockers (BB), calcium channel-blockers (CCB), and renin-angiotensin system antagonists (RASA). Adherence was calculated by class using pharmacy quality alliance proportion of days covered (PDC), with adherent defined as PDC≥80%. Nonpersistence was defined as a gap in therapy ≥30 days or discontinuation ≥30 days prior to observation period end, and was calculated by class among participants with ≥1 prescription claim. Statistical testing was used to assess differences in outcomes when stratified by SDoH (e.g., gender, race/ethnicity, income). 1763 eligible patients with hypertension were analyzed. Statistically-significant differences in HCRU were found for all SDoH characteristics, including: higher proportion of Hispanic patients with a physician-office visit (85.2%;p=0.038) and a higher proportion of non-Hispanic/black patients with an all-cause ER visit (31.8%;p<0.0001) or hospitalization (11.8%;p=0.037). 60.1% (N=1059) of all patients had ≥1 antihypertensive claim, with higher proportions seen in non-Hispanic/Black (65.9%;p=0.032) or patients with an income $25000-$49000 (67.6%;p<0.001). A larger proportion of Hispanic patients were nonpersistent (72.2%p<0.01) or nonadherent (70.6%;p=0.001) to CCBs, and nonadherent to RASAs (60.5%;p<0.05). Among patents with hypertension significant associations exist between SDoH characteristics, HCRU, adherence, and persistence. Further research is needed to understand why Hispanics are less persistent to their antihypertensives while having more frequent contacts with their physicians. A better understanding of the role SDoH plays in the management of hypertension could improve hypertension control and reduce healthcare costs.
To examine the burden of caregiving in terms of comorbidities, health-related quality of life (HRQoL), and work productivity in multiple countries Data were obtained from the 2017 Japan (N = 30,001), China (N = 19,994), and US (N = 75,004) National Health and Wellness Survey (NHWS), a nationally-representative internet survey of adults. The US NHWS was additionally linked to a US claims database using a HIPAA-compliant matching algorithm. Outcomes included comorbidities, depression (Patient Health Questionnaire 9-item (PHQ-9)) and anxiety (Generalized Anxiety Disorder 7-item (GAD-7) severity, HRQoL (SF-12v2/SF-36v2 mental (MCS) and physical (PCS) component summary scores), and Work Productivity and Activity Impairment Questionnaire. Chi-squared and t-tests or Mann–Whitney tests were used to compare outcomes between caregivers and non-caregivers. Among total respondents of the Japan and China NHWS, 2,821 and 2,900 caregivers were identified, respectively. Compared with non-caregivers, caregivers in Japan and China were more likely to self-report depression and anxiety diagnoses and experience moderate-to-severe anxiety (13.4% vs. 7.4% and 16.2% vs. 3.9%) and depression (17% vs. 10.5% and 28% vs. 8.6%) (all p < .001). Caregivers vs. non-caregivers had significantly lower MCS (45.7 vs. 48.5 and 45.2 vs. 49.1) and PCS (50.5 vs. 52 and 48 vs. 51.2), and higher work (27.8% vs. 20.2% and 35.1% vs. 25%) and activity (27.6% vs. 20.6% and 29.8% vs. 19%) impairment (all p < .001). Similar results were observed among 4,015 caregivers in the US linked sample: caregivers had lower HRQoL and higher work and activity impairment than non-caregivers (all p < .001). Additionally, caregivers vs. non-caregivers were more likely to be diagnosed with anxiety (13.4% vs. 8.2%) and depression (15% vs. 9.7%) according to claims data (all p < .001). Consistent across countries and different data sources (self-reported, claims), caregivers experienced significantly greater anxiety and depression, poorer HRQoL, and higher work impairment.
Linking secondary clinical data with patient-reported data at the patient-level brings together a comprehensive view of the patient but sample sizes can be a challenge. This study demonstrates the fusion of Patient Reported Outcomes (PROs) in surveys with clinical data in claims enabling the study of associations between quality of life and disease-treatment interactions at scale especially for rare diseases. The PROs SF-36v2 PCS, MCS, SF6D, and EQ5D were available in the National Health and Wellness Survey (N=345K). Clinical information from Komodo Health, a large U.S. database of health insurance claims (N=200M), were obtained using ICD, CPT/HCPCS, and NDC codes. 104K patients were linkable in the two data sets. The fusion process was accomplished using an artificial neural network-based predictive model followed by predictive mean matching. The linked data was used to train, validate and test the fusion methodology. The method allows for the simultaneous imputation of the 4 PROs. Results were assessed for the general patient population (GP), type-2 diabetes (T2D), and Myasthenia Gravis (MG), a rare disease. Results were also assessed after stratifying by age and gender. The triplet of numbers corresponds to the 3 cohorts (GP,T2D,MG). The number of patients in the test set was N:(5207,898,100). The difference between the observed and imputed means were: PCS:(0.23,-0.23,-0.22), MCS:(-0.009,0.14,1.5), EQ5D:(0.002,-0.005,-0.01) and SF6D:(0.002,-0.002,-0.004). We failed to reject hypothesis of no difference in all cases. All differences were less than the respective minimal clinical important difference. Similar results were observed when stratified by age and gender. The correlations between the imputed PROs mimic the observed correlations (absolute difference < 0.05). This study shows the suitability of data fusion as a substitute for linkage where overlap between data sources is small to study the effects of clinical variables on PROs.
Quality of life (QOL) measurements are extremely important in outcome research as they reflect the effect of illness and treatment as perceived by the patients. The availability of QOL data, however, becomes a major barrier for including the patient perspective in big data healthcare studies. This research is trying to identify key predictors that can estimate QOL values within reasonable accuracy so they can be used in studies in conjunction with claims or EHR data. National Health and Wellness Survey is self-administered online survey of adults 18 years and over. In China, the survey was administered to the urban population. Type 2 Diabetes (T2D) patient data from 2013 to 2017 were extracted for this analysis. 795 patients were included in the final working data set with 869 variables including QOL measure using EQ5D instruments. An iterative process of data transformation, recoding and LASSO regression was used to identify a subset of 14 variables that descriptively covers the clinical, social and economic aspects of T2D patients. A two-layer feed-forward artificial neural network (ANN) model was fitted to the data set with five-fold cross-validation. The ANN model demonstrated a good fit for the training data set with R-squared value of 0.81. The model also performed very well in the cross-validation with R-squared value of 0.86. Residual analysis supported that a good fit was achieved. QOL outcomes can be reasonably accurately approximated with a small set of predictive variables that are commonly available in large real-world databases. The combination of optimal predictive variables and robust models may enable QOL outcomes to be included in more real-world data studies.
To use nationally-representative patient-reported outcome survey data linked with clinical data from electronic health records (EHR) to evaluate differences in serum uric acid (sUA) levels and health-related quality of life (HRQoL) in patients with gout only and patients with gout and hypertension (GwH). A HIPAA-compliant linking methodology was performed to link Patient-Centered-Research (PaCeR) US data (2015-2018; N=282,368), consisting of patient-reported survey data, with a large US ambulatory EHR database (2012-2018; N=50 million+). Linked PaCeR-EHR respondents were included in the study if they had at least one sUA test result in the EHR taken within 18 months of survey completion. Gout and hypertension patients were identified through self-reported physician diagnoses in PaCeR. Wilcoxon rank-sum tests were used to compare sUA and HRQoL between groups. HRQoL measures included SF-36v2 metrics. A total of 231 PaCeR-EHR respondents had valid sUA tests result in the EHR, of which 5.2% had gout only, 16.5% had GwH, 27.7% had hypertension and 50.6% had neither. Gout patients were more likely to be male, older, and unemployed compared with non-gout patients. Patients with gout (6.6 mg/dL), hypertension (5.8 mg/dL), or GwH (6.1 mg/dL) had significantly higher sUA levels compared to those with neither condition (5.2 mg/dL) (p’s < 0.001). Patients with hypertension (44.4) or GwH (44.6) had lower physical component summary scores compared with those with neither condition (50.6) (p’s < 0.05). Additionally, physical functioning and role-physical domain scores were also lower among patients with hypertension or GwH (p’s <0.001). Mental component summary scores were similar across groups. Linked PaCeR-EHR data enabled examination of a both clinical characteristic (sUA) and patient-reported HRQoL among gout patients with and without comorbid hypertension. Results reveal the humanistic burden of gout and its association with sUA levels. sUA levels were consistently higher with lower physical HRQoL among patients with GwH.
Frequency of HbA1c testing is associated with type 2 diabetes (T2D) management and control, but limited data exists on the relationship between HbA1c testing and patient-reported measures such as HRQoL. This study used patient-reported survey data linked with laboratory tests from electronic health records (EHR) to assess differences in HRQoL associated with HbA1c testing frequency. Patient-Centered-Research (PaCeR) US data (2015-2018; N=345,184), consisting of nationally-representative patient-reported survey data, were linked with a large US ambulatory EHR database (2015-2018; N=50 million+) in a HIPAA-compliant linking methodology. The study population consisted of linked respondents who were diagnosed with T2D and had at least one HbA1c test result in the EHR within six months of survey participation. Study participants were grouped by HbA1c testing frequency (1-2 tests/year, 3-4 tests/year, 5+ tests/year) and compared on sociodemographic and HRQoL measures using two-sample t-tests and chi-square tests. HRQoL measures included SF-36v2 (mental and physical component summary scores [MCS and PCS] and SF-6D health utility index). Of 492 total T2D patients who met the study criteria [mean age = 61.3 years, 54.7% female, 85.6% White, 38.8% uncontrolled HbA1c (≥ 7%)], 58.1% had 1-2 HbA1c tests/year, 26.6% had 3-4 HbA1c tests/year, and 15.2% had 5+ HbA1c tests/year. T2D patients with 1-2 HbA1c tests/year had significantly lower MCS compared to those with 3-4 HbA1c tests/year (mean = 48.8 vs. 51.8, p = 0.01). Vitality (mean = 47.6 vs. 50.0, p = 0.04) and mental health (mean = 48.6 vs. 52.1, p < 0.005) domain scores were also lower among patients with 1-2 HbA1c tests/year compared to patients with 3-4 HbA1c tests/year. There were no differences in HRQoL between the 1-2 tests/year and 5+ tests/year groups. Linking patient-reported outcomes with EHR data facilitated analysis of HbA1c testing with HRQoL. Results showed that less frequent HbA1c testing was associated with lower MCS.
Body weight is a determinant of health-related quality of life (HRQoL). Less is known, however, on how rapid changes in weight affect HrQoL. This study assessed differences in demographics and HRQoL associated with rapid weight change. The National Health and Wellness Survey (NHWS), a nationally-representative survey of adults (≥18 years), was linked to a large US ambulatory electronic health records (EHR) database using a HIPAA-compliant matching algorithm. The study population consisted of NHWS respondents between 2015-2018 with ≥ two weight measurements taken ≥ 30 days apart in the EHR six months before NHWS participation. Comparison groups included those who gained, lost (≥10, ≥20 lbs for both), or maintained their weight over 30 days. Analysis of variance and Chi-squared tests were used to compare demographic characteristics (e.g., age, gender, race, education, income) and HRQoL (SF-36v2 physical component summary (PCS) and mental component summary (MCS)) between groups. Of 1,434 NHWS respondents linked to EHR data (mean age = 57.4 years, 64.2% females), using a threshold of ≥10 lbs, 7.5% gained, 9.3% lost, and 83.2% maintained their weight. Using a threshold of ≥20 lbs, these values were 1.9%, 2.4%, and 95.7%, respectively. There were no demographic differences between weight change categories by threshold. Weight gain, but not weight loss, was associated with HRQoL. Compared with patients who maintained their weight, patients who gained ≥10 lbs had lower PCS scores (mean = 43.0 vs. 46.5, p < 0.01) and MCS scores (mean = 46.1 vs. 48.9, p = 0.02). Differences in HRQoL increased when using a threshold of ≥20 lbs (PCS: mean = 41.1 vs. 46.3, p = 0.01; MCS: 43.9 vs. 48.8, p = 0.03). Rapid weight gain, assessed using weight measurements from EHR, was associated with lower HRQoL. Linking data facilitated a longitudinal analysis combining an objective clinical measure with patient-reported outcomes.
To examine depression and anxiety prevalence and health-related quality of life (HRQoL) in cancer caregivers. The National Health and Wellness Survey (NHWS), a nationally-representative Internet survey of adults (≥18 years), was linked to a large US ambulatory electronic health records (EHR) database using a HIPAA-compliant matching algorithm. The study population included NHWS respondents between 2015-2018 who reported being a caregiver of cancer patients and those who reported not being a caregiver. Prevalence of depression and anxiety among caregivers were estimated using self-reported physician diagnoses and ICD-10 codes in the EHR. Analysis of variance and Chi-squared tests were used to compare depression (Patient Health Questionnaire 9-item (PHQ-9)) and anxiety (Generalized Anxiety Disorder 7-item (GAD-7)) severity and HRQoL (SF-36v2 physical component summary (PCS) and mental component summary (MCS)) between caregivers and non-caregivers. Of 13,928 NHWS respondents linked to EHR data (mean age = 51.6 years, 64.3% females, 78.4% White), 298 (2.1%) reported being caregivers to cancer patients. The prevalence of depression and anxiety were 38.9% and 36.2% using self-reported diagnoses and 18.8% and 17.6% based on EHR diagnoses. Compared with non-caregivers, cancer caregivers were more likely have generalized anxiety disorder (33.8% vs 14.0%, p < 0.001), based on the GAD-7. Cancer caregivers also reported greater severity of depressive symptoms based on the PHQ-9 and were more likely to have mild (24.2% vs. 18.8%), moderate (15.5% vs 8.1%), and severe (18.0 % vs. 7.0%) depression (all p < 0.001), compared with non-caregivers. Compared with non-caregivers, cancer caregivers reported significantly lower MCS (43.5 vs. 48.1, p < 0.001) and PCS (46.3. vs 49.0, p < 0.001). Cancer caregivers reported poorer HRQoL and were more likely to report experiencing anxiety and depressive symptoms. Linking data enabled assessment of both self-reported and clinical diagnoses to characterize burden among cancer caregivers.
Assess the relationship between health-related quality of life (HRQoL) and body mass index (BMI) using nationally representative patient-reported data linked with electronic health records (EHR), and characterize patients with substantial differences in their clinical versus self-reported BMI values. Patient-Centered-Research (PaCeR) US data (2015-2017; N=270,207) consisting of patient-reported information were included in a HIPAA-compliant linking methodology with a large US ambulatory EHR database (2012-2017; N=50 million+). Linking was performed by comparing Protected Health Information (PHI) from EHR and Personal Identifiable Information from PaCeR. Height and weight values were collected from EHR and self-reported data to calculate BMI and categorize patients as underweight (BMI<18.5 kg/m2), normal weight (18.530 kg/m2). Descriptive statistics were calculated as means with standard deviation (SD) for continuous variables and frequencies for categorical variables. Pearson correlation coefficients were calculated between BMI and HRQoL measures including SF-36v2 (mental and physical component summary scores [MCS and PCS] and SF-6D health utility index). Patient characteristics were compared for those with self-reported BMI and those with either discrepant (differing by more than 2 SDs from linked clinical values) or missing self-reported BMI. A total of 6,445 PaCeR-EHR patients had valid clinical BMI values, with 6,235 (96.7%) having self-reported BMI. Among this group, EHR and self-reported BMI were similar: 1.8% vs. 1.7% were underweight, 24.5% vs. 28.1% were normal, 28.8% vs. 32.1% were overweight, and 44.9% vs. 38.1% were obese. EHR and self-reported BMI were also strongly correlated (r=0.89, p<0.001). BMI was weakly associated with lower PCS scores for both EHR (r=-0.27, p<0.001) and self-reported (r=-0.30, p<0.001) BMI. Associations observed between HRQoL indicators and BMI were similar for both EHR and self-reported BMI values. The comparison of EHR and PaCeR data served as a validation check of patient-reported information.
To assess the feasibility of linking a large nationally representative patient-reported database with an electronic health records (EHR) database to enhanced patient data. Patient-Centered-Research (PaCeR) datasets comprising 3 years (2015-2017; total N=270207) of patient-reported data were included in a HIPAA-compliant linking methodology involving 50 million+ patients from an EHR database. Linking was performed by comparing Protected Health Information from EHR and Personal Identifiable Information from PaCeR. Data used in the linking included first and last name, address, zip code, gender, date of birth, email address, and phone number. Once data was linked, the prevalence of diagnosed type 2 diabetes (T2D), rheumatoid arthritis (RA), psoriasis, inflammatory bowel disease (IBD), depression, and migraine was examined for linked, non-linked, and all PaCeR respondents. Post linking, 7266 PaCeR respondents were identified as having linked records in the EHR database. Of these, 941 self-reported a physician's diagnosis for T2D, 308 for RA, 271 for psoriasis, 149 for IBD, 1902 for depression, and 1028 for migraines. Prevalence estimates were highest for the linked respondent subsample, followed by the full PaCeR sample, and lowest for the non-linked subsample. This relationship held for the prevalence of T2D (13.98% vs. 8.91% vs. 8.75%), RA (4.47% vs. 2.92% vs. 2.87%), psoriasis (3.68% vs. 2.72% vs. 2.69%), IBD (1.94% vs. 1.24% vs. 1.22%), depression (25.69% vs. 19.63% vs. 19.44%), and migraine (13.34% vs. 9.81% vs. 9.70%). Linking of PaCeR and EHR databases using HIPAA-compliant methods was successful, giving a sub-sample of linked patients for which both patient-reported data and clinical data can be used to address research questions. Prevalence estimates for linked, non-linked, and the full PaCeR samples were as expected, with the highest prevalence being among those seeking care (linked), and the lowest among those who may or may not be seeking care (non-linked).