We leveraged real-world data from the OneFlorida+ Data Trust to examine sociodemographic biases associated with computable phenotypes (CPs) for type 2 diabetes (T2D) that incorporate prescription and laboratory information. Medical record review was used as a gold-standard to evaluate the discriminative ability of each CP. CPs identified patients ≥ 18 years from OneFlorida+ electronic health records (9/1/2018 – 2/29/2020). To clarify the relationships among the partially overlapping CP definitions, we defined four mutually exclusive (ME) groups of individuals: 1) T2D diagnosis only (ME_Claims); 2) T2D diagnosis and antidiabetic prescription (ME_Meds); 3) T2D diagnosis and HbA1c ≥ 6.5
GLIMMPSE Version 3 is a free, web-based, open-source software tool, which calculates power and sample size for general linear mixed models with Gaussian errors. The software permits power calculations for clinical trials, randomized experiments, and observational studies with clustering, repeated measures, and both, and almost any testable hypothesis. The software has been supported by five United States National Institutes of Health (NIH) grants, is used for more than 14,000 power or sample size calculations per year, has been cited in almost 500 peer-reviewed manuscripts, and used to design more than 200 million dollars in NIH-funded studies. This release provides several new features. The back end has been refactored in Python. The interface has been simplified, requiring user decisions about only one topic per screen. A new menu improves specification of both between-participant and within-participant hypotheses. A recursive algorithm permits computing covariances for up to ten levels of clustering. An updated Monte Carlo simulation using five new examples with clustering, longitudinality, or both, shows accuracy of the power approximations to within 0.01. Five new examples demonstrate power or sample size calculations for 1) a cluster-randomized trial, 2) a longitudinal study with repeated measures, 3) a multilevel study with a multivariate outcome, 4) a multilevel and longitudinal study, and 5) a complex study with a subgroup factor, repeated measures, and intervention-by-location interaction.
We give examples of three features in the design of randomized controlled clinical trials which can increase power and thus decrease sample size and costs. We consider an example multilevel trial with several levels of clustering. For a fixed number of independent sampling units, we show that power can vary widely with the choice of the level of randomization. We demonstrate that power and interpretability can improve by testing a multivariate outcome rather than an unweighted composite outcome. Finally, we show that using a pooled analytic approach, which analyzes data for all subgroups in a single model, improves power for testing the intervention effect compared to a stratified analysis, which analyzes data for each subgroup in a separate model. The power results are computed for a proposed prevention research study. The trial plans to randomize adults to either telehealth (intervention) or in-person treatment (control) to reduce cardiovascular risk factors. The trial outcomes will be measures of the Essential Eight, a set of scores for cardiovascular health developed by the American Heart Association which can be combined into a single composite score. The proposed trial is a multilevel study, with outcomes measured on participants, participants treated by the same provider, providers nested within clinics, and clinics nested within hospitals. Investigators suspect that the intervention effect will be greater in rural participants, who live farther from clinics than urban participants. The results use published, exact analytic methods for power calculations with continuous outcomes. We provide example code for power analyses using validated software.
Tracking trajectories of body size in children provides insight into chronic disease risk. One measure of pediatric body size is body mass index (BMI), a function of height and weight. Errors in measuring height or weight may lead to incorrect assessment of BMI. Yet childhood measures of height and weight extracted from electronic medical records often include values which seem biologically implausible in the context of a growth trajectory. Removing biologically implausible values reduces noise in the data, and thus increases the ease of modeling associations between exposures and childhood BMI trajectories, or between childhood BMI trajectories and subsequent health conditions. We developed open-source algorithms (available on github) for detecting and removing biologically implausible values in pediatric trajectories of height and weight. A Monte Carlo simulation experiment compared the sensitivity, specificity and speed of our algorithms to three published algorithms. The comparator algorithms were selected because they used trajectory information, had open-source code, and had published verification studies. Simulation inputs were derived from longitudinal epidemiological cohorts. Our algorithms had higher specificity, with similar sensitivity and speed, when compared to the three published algorithms. The results suggest that our algorithms should be adopted for cleaning longitudinal pediatric growth data.
Researchers often aim to assess whether repeated measures of an exposure are associated with repeated measures of an outcome. A question of particular interest is how associations between exposures and outcomes may differ over time. In other words, researchers may seek the best form of a temporal model. While several models are possible, researchers often consider a few key models. For example, researchers may hypothesize that an exposure measured during a sensitive period may be associated with repeated measures of the outcome over time. Alternatively, they may hypothesize that the exposure measured immediately before the current time period may be most strongly associated with the outcome at the current time. Finally, they may hypothesize that all prior exposures are important. Many analytic methods cannot compare and evaluate these alternative temporal models, perhaps because they make the restrictive assumption that the associations between exposures and outcomes remains constant over time. Instead, we provide a tutorial describing four temporal models that allow the associations between repeated measures of exposures and outcomes to vary, and showing how to test which temporal model is best supported by the data. By finding the best temporal model, developmental psychopathology researchers can find optimal windows for intervention.
Although superficially similar to data from clinical research, data extracted from electronic health records may require fundamentally different approaches for model building and analysis. Because electronic health record data is designed for clinical, rather than scientific use, researchers must first provide clear definitions of outcome and predictor variables. Yet an iterative process of defining outcomes and predictors, assessing association, and then repeating the process may increase Type I error rates, and thus decrease the chance of replicability, defined by the National Academy of Sciences as the chance of "obtaining consistent results across studies aimed at answering the same scientific question, each of which has obtained its own data."[1] In addition, failure to account for subgroups may mask heterogeneous associations between predictor and outcome by subgroups, and decrease the generalizability of the findings. To increase chances of replicability and generalizability, we recommend using a stratified split sample approach for studies using electronic health records. A split sample approach divides the data randomly into an exploratory set for iterative variable definition, iterative analyses of association, and consideration of subgroups. The confirmatory set is used only to replicate results found in the first set. The addition of the word 'stratified' indicates that rare subgroups are oversampled randomly by including them in the exploratory sample at higher rates than appear in the population. The stratified sampling provides a sufficient sample size for assessing heterogeneity of association by testing for effect modification by group membership. An electronic health record study of the associations between socio-demographic factors and uptake of hepatic cancer screening, and potential heterogeneity of association in subgroups defined by gender, self-identified race and ethnicity, census-tract level poverty and insurance type illustrates the recommended approach.
Background When evaluating the impact of environmental exposures on human health, study designs often include a series of repeated measurements. The goal is to determine whether populations have different trajectories of the environmental exposure over time. Power analyses for longitudinal mixed models require multiple inputs, including clinically significant differences, standard deviations, and correlations of measurements. Further, methods for power analyses of longitudinal mixed models are complex and often challenging for the non-statistician. We discuss methods for extracting clinically relevant inputs from literature, and explain how to conduct a power analysis that appropriately accounts for longitudinal repeated measures. Finally, we provide careful recommendations for describing complex power analyses in a concise and clear manner. Methods For longitudinal studies of health outcomes from environmental exposures, we show how to [1] conduct a power analysis that aligns with the planned mixed model data analysis, [2] gather the inputs required for the power analysis, and [3] conduct repeated measures power analysis with a highly-cited, validated, free, point-and-click, web-based, open source software platform which was developed specifically for scientists. Results As an example, we describe the power analysis for a proposed study of repeated measures of per- and polyfluoroalkyl substances (PFAS) in human blood. We show how to align data analysis and power analysis plan to account for within-participant correlation across repeated measures. We illustrate how to perform a literature review to find inputs for the power analysis. We emphasize the need to examine the sensitivity of the power values by considering standard deviations and differences in means that are smaller and larger than the speculated, literature-based values. Finally, we provide an example power calculation and a summary checklist for describing power and sample size analysis. Conclusions This paper provides a detailed roadmap for conducting and describing power analyses for longitudinal studies of environmental exposures. It provides a template and checklist for those seeking to write power analyses for grant applications.
When designing repeated measures studies, both the amount and the pattern of missing outcome data can affect power. The chance that an observation is missing may vary across measurements, and missingness may be correlated across measurements. For example, in a physiotherapy study of patients with Parkinson’s disease, increasing intermittent dropout over time yielded missing measurements of physical function. In this example, we assume data are missing completely at random, since the chance that a data point was missing appears to be unrelated to either outcomes or covariates. For data missing completely at random, we propose noncentral F power approximations for the Wald test for balanced linear mixed models with Gaussian responses. The power approximations are based on moments of missing data summary statistics. The moments were derived assuming a conditional linear missingness process. The approach provides approximate power for both complete-case analyses, which include independent sampling units where all measurements are present, and observed-case analyses, which include all independent sampling units with at least one measurement. Monte Carlo simulations demonstrate the accuracy of the method in small samples. We illustrate the utility of the method by computing power for proposed replications of the Parkinson’s study.
Although superficially similar to data from clinical research, data extracted from electronic health records (EHRs) may require fundamentally different approaches to analysis and model building. Some outcome and predictor variables may not be well-defined at the start of the study. Selecting specific definitions requires exploratory data analysis. Specifying the rules for computing a new variable inevitably leads to exploratory analyses. Achieving replicability, i.e., a high probability that a similar future study will reach the same conclusions, requires special approaches. We recommend a study design strategy based on stratified sample splitting for studies using EHRs. The split-sample design ensures meeting the goal of replicability. Stratified sampling of EHRs increases generalizability by allowing heterogeneity between subgroups to be tested appropriately with good statistical power. Building a model from EHR data to predict uptake of hepatic cancer screening illustrates the recommended approach.
ObesityVolume 30, Issue 3 p. 565-570 EDITORIAL A practical decision tree to support editorial adjudication of submitted parallel cluster randomized controlled trials Yasaman Jamshidi-Naeini, Yasaman Jamshidi-Naeini orcid.org/0000-0003-4769-2764 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorAndrew W. Brown, Andrew W. Brown orcid.org/0000-0002-1758-8205 Department of Applied Health Science, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorTapan Mehta, Tapan Mehta orcid.org/0000-0002-2016-2344 Department of Health Services Administration, University of Alabama at Birmingham, Birmingham, Alabama, USASearch for more papers by this authorDeborah H. Glueck, Deborah H. Glueck Department of Pediatrics, University of Colorado School of Medicine, University of Colorado Denver, Aurora, Colorado, USASearch for more papers by this authorLilian Golzarri-Arroyo, Lilian Golzarri-Arroyo orcid.org/0000-0002-1221-6701 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorKeith E. Muller, Keith E. Muller Department of Health Outcomes and Biomedical Informatics, College of Medicine, University of Florida, Gainesville, Florida, USASearch for more papers by this authorCarmen D. Tekwe, Carmen D. Tekwe orcid.org/0000-0002-1857-2416 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorDavid B. Allison, Corresponding Author David B. Allison allison@iu.edu orcid.org/0000-0003-3566-9399 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USA Correspondence David B. Allison, Indiana University School of Public Health-Bloomington, 1025 E. 7th St., PH 111, Bloomington, IN, USA 47405. Email: allison@iu.eduSearch for more papers by this author Yasaman Jamshidi-Naeini, Yasaman Jamshidi-Naeini orcid.org/0000-0003-4769-2764 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorAndrew W. Brown, Andrew W. Brown orcid.org/0000-0002-1758-8205 Department of Applied Health Science, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorTapan Mehta, Tapan Mehta orcid.org/0000-0002-2016-2344 Department of Health Services Administration, University of Alabama at Birmingham, Birmingham, Alabama, USASearch for more papers by this authorDeborah H. Glueck, Deborah H. Glueck Department of Pediatrics, University of Colorado School of Medicine, University of Colorado Denver, Aurora, Colorado, USASearch for more papers by this authorLilian Golzarri-Arroyo, Lilian Golzarri-Arroyo orcid.org/0000-0002-1221-6701 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorKeith E. Muller, Keith E. Muller Department of Health Outcomes and Biomedical Informatics, College of Medicine, University of Florida, Gainesville, Florida, USASearch for more papers by this authorCarmen D. Tekwe, Carmen D. Tekwe orcid.org/0000-0002-1857-2416 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USASearch for more papers by this authorDavid B. Allison, Corresponding Author David B. Allison allison@iu.edu orcid.org/0000-0003-3566-9399 Department of Epidemiology and Biostatistics, Indiana University School of Public Health-Bloomington, Bloomington, Indiana, USA Correspondence David B. Allison, Indiana University School of Public Health-Bloomington, 1025 E. 7th St., PH 111, Bloomington, IN, USA 47405. Email: allison@iu.eduSearch for more papers by this author First published: 23 February 2022 https://doi.org/10.1002/oby.23373Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Volume30, Issue3March 2022Pages 565-570 RelatedInformation
How can people use the results?Researchers and doctors can use the methods to help make sure patients with high healthcare needs get the care they need.
We derive a noncentral F power approximation for the Kenward and Roger test. We use a method of moments approach to form an approximate distribution for the Kenward and Roger scaled Wald statistic, under the alternative. The result depends on the approximate moments of the unscaled Wald statistic. Via Monte Carlo simulation, we demonstrate that the new power approximation is accurate for cluster randomized trials and longitudinal study designs. The method retains accuracy for small sample sizes, even in the presence of missing data. We illustrate the method with a power calculation for an unbalanced group-randomized trial in oral cancer prevention.
ObjectiveWe sought to examine the extent to which body mass index (BMI) was available in electronic health records for Florida Medicaid recipients aged 5 to 18 years taking Second-Generation Antipsychotics (SGAP). We also sought to illustrate how clinical data can be used to identify children most at-risk for SGAP-induced weight gain, which cannot be done using process-focused measures.MethodsElectronic health record (EHR) data and Medicaid claims were linked from 2013 to 2019. We quantified sociodemographic differences between children with and without pre- and post-BMI values. We developed a linear regression model of post-BMI to examine pre-post changes in BMI among 4 groups: 1) BH/SGAP+ children had behavioral health conditions and were taking SGAP; 2) BH/SGAP- children had behavioral health conditions without taking SGAP; 3) children with asthma; and 4) healthy children.ResultsOf 363,360 EHR-Medicaid linked children, 18,726 were BH/SGAP+. Roughly 4% of linked children and 8% of BH/SGAP+ children had both pre and post values of BMI required to assess quality of SGAP monitoring. The percentage varied with gender and race-ethnicity. The R2 for the regression model with all predictors was 0.865. Pre-post change in BMI differed significantly (P < .0001) among the groups, with more BMI gain among those taking SGAP, particularly those with higher baseline BMI.ConclusionMeeting the 2030 Centers for Medicare and Medicaid Services goal of digital monitoring of quality of care will require continuing expansion of clinical encounter data capture to provide the data needed for digital quality monitoring. Using linked EHR and claims data allows identifying children at higher risk for SGAP-induced weight gain.
Abstract Hepatitis C virus (HCV) infection is a leading risk factor for hepatocellular carcinoma. We employed a retrospective cohort study design and analyzed 2012–2018 Medicaid claims linked with electronic health records data from the OneFlorida Data Trust, a statewide data repository containing electronic health records data for 15.07 million Floridians from 11 health care systems. Only adult patients at high-risk for HCV (n = 30,113), defined by diagnosis of: HIV/AIDS (20%), substance use disorder (64%), or sexually transmitted infections (22%) were included. Logistic regression examined factors associated with meeting the recommended sequence of HCV testing. Overall, 44.1% received an HCV test. The odds of receiving an initial test were significantly higher for pregnant females (odds ratio [OR]1.99; 95% confidence interval [CI] 1.86–2.12; P < .001) and increased with age (OR 1.01; 95% CI 1.00–1.01; P < .001).Among patients with low Charlson comorbidity index (CCI = 1), non-Hispanic (NH) black patients (OR 0.86; 95% CI 0.81–0.9; P < .001) had lower odds of getting an HCV test; however, NH black patients with CCI = 10 had higher odds (OR 1.41; 95% CI 1.21–1.66; P < .001) of receiving a test. Of those who tested negative during initial testing, 17% received a second recommended test after 6 to 24 months. Medicaid-Medicare dual eligible patients, those with high CCI (OR 1.14; 95% CI 1.11–1.17; P < .001), NH blacks (OR 1.93; 95% CI 1.61–2.32; P < .001), and Hispanics (OR 1.49; 95% CI 1.08–2.06; P = .02) were significantly more likely to have received a second HCV test, while pregnant females (OR 0.71; 95% CI 0.57–0.89; P = .003), had lower odds of receiving it. The majority of patients who tested positive during the initial test (97%) received subsequent testing. We observed suboptimal adherence to the recommended HCV testing among high-risk patients underscoring the need for tailored interventions aimed at successfully navigating high-risk individuals through the HCV screening process. Future interventional studies targeting multilevel factors, including patients, clinicians and health systems are needed to increase HCV screening rates for high-risk populations.
BACKGROUND AND OBJECTIVE:First-line, nonpharmacological therapy is recommended for many pediatric mental health (MH) conditions prior to initiating antipsychotic prescription therapies. Many children do not receive these recommended services, despite the known association between antipsychotic medications and metabolic dysfunction. The main objective of this study was to quantify the association among children's MH diagnosis categories, sociodemographic characteristics and receipt of first-line psychosocial care among children in Florida Medicaid METHODS: Florida Medicaid enrollment, healthcare and pharmacy claims were used for this multivariate analysis. Children were assigned to condition clusters wherein related diagnoses were grouped into clinically relevant categories. A total of 7704 children were included in the final analysis.RESULTS:Twenty-four percent of children in Florida Medicaid do not receive first-line, nonpharmacological psychosocial care. Age was significantly associated with not receiving psychosocial services, with older children less likely to receive. Non-Hispanic White children as well as those living in rural areas had lower odds of receiving behavioral intervention prior to initiating antipsychotics. Children with mood-disorders, behavior problems, anxiety and stress related disorders were more likely to receive first-line psychosocial care.CONCLUSIONS:This study provides an important understanding of the variability in receipt of first-line psychosocial care before antipsychotic medication initiation among children in Medicaid based on sociodemographic and MH health characteristics. These analyses can be used to develop quality improvement initiatives targeted toward children that are most vulnerable for not receiving recommended care.
Objectives: Promoting patient involvement in managing co-occurring physical and mental health conditions is increasingly recognized as critical to improving outcomes and controlling costs in this growing chronically ill population. The main objective of this study was to conduct an economic evaluation of the Wellness Incentives and Navigation (WIN) intervention as part of a longitudinal randomized pragmatic clinical trial for chronically ill Texas Medicaid enrollees with co-occurring physical and mental health conditions. Methods: The WIN intervention used a personal navigator, motivational interviewing, and a flexible wellness expense account to increase patient activation, that is, the patient's knowledge, skills, and confidence in managing their self-care and co-occurring physical and mental health conditions. Regression models were fit to both participant-level quality-adjusted life years (QALYs) and total costs of care (including the intervention) controlling for demographics, health status, poverty, Medicaid managed care plan, intervention group, and baseline health utility and costs. Incremental costs and QALYs were calculated based on the difference in predicted costs and QALYs under intervention versus usual care and were used to calculate the incremental cost-effectiveness ratios (ICERs). Confidence intervals were calculated using Fieller's method, and sensitivity analyses were performed. Results: The mean ICER for the intervention compared with usual care was $12 511 (95% CI $8971-$16 842), with a sizable majority of participants (70%) having ICERs below $40 000. The WIN intervention also produced higher QALY increases for participants who were sicker at baseline compared to those who were healthier at baseline. Conclusion: The WIN intervention shows considerable promise as a cost-effective intervention in this challenging chronically ill population.
AbstractObjectiveTo examine whether the Wellness Incentive and Navigation (WIN) intervention can improve health‐related quality of life (HRQOL) among Medicaid enrollees with co‐occurring physical and behavioral health conditions.Data SourcesAnnual telephone survey data from 2013 to 2016, linked with claims data.Study DesignWe recruited 1259 participants from the Texas STAR + PLUS managed care program and randomized them into an intervention group that received flexible wellness accounts and navigator services or a control group that received standard care. We conducted 4 waves of telephone surveys to collect data on HRQOL, patient activation, and other participant demographic and clinical characteristics.Data Collection/Extraction MethodsThe 3M Clinical Risk Grouping Software was used to extract variables from claims data and group participants based on disease severity.Principal FindingsOur results showed that the WIN intervention was effective in increasing patient activation and HRQOL among Medicaid enrollees with co‐occurring physical and behavioral health conditions. Furthermore, we found that this intervention effect on HRQOL was partially mediated by patient activation.ConclusionsProviding navigator support with wellness account is effective in improving HRQOL among Medicaid enrollees. The pragmatic nature of the trial maximizes the chance of successfully implementing it in state Medicaid programs.
Patient activation, the perceived capacity to manage one’s health, is positively associated with better health outcomes and lower costs. Underlying characteristics influencing patient activation are not completely understood leading to gaps in intervention strategies designed to improve patient activation. We suggest that variability in executive functioning influences patient activation and ultimately has an impact on health outcomes. To examine this hypothesis, 440 chronically ill Medicaid enrollees completed measures of executive functioning, patient activation, and health-related quality of life. Mediation analyses revealed that executive functioning: (a) directly affected patient activation and mental health-related quality of life, (b) indirectly affected mental health-related quality of life through patient activation, and (c) was unrelated to physical health-related quality of life. These data indicate that further study of the relationships among neurocognitive processes, patient activation, and health-related quality of life is needed and reinforces previous work demonstrating the association between patient activation and self-reported outcomes.
The purpose of this design and development case is to share our experiences in the transformation of a face-to-face workshop into a Massive Open Online Course (MOOC) for a prominent MOOC platform. The goal of the workshop and MOOC is to teach learners how to conduct appropriate power and sample size analysis for multilevel and longitudinal studies in social and behavioral health research. Learners include people from across the biomedical research spectrum, from students to full professors. We first describe the design and development frameworks and processes used to create the three-day, face-to-face workshop. Then, we detail the design and development approach to transform this face-to-face workshop into a MOOC. At a macro-design level, we employed backward design (Wiggins & McTighe 1998) as an instructional design framework. At a micro-design level, we used a combination of the first principles of instruction, the cognitive theory of multimedia learning, the nine events of instruction, and design recommendations for MOOCs found in the literature. We report the results of a formative evaluation of the MOOC. Finally, we provide closing remarks, lessons learned, and the next steps for the instructional program.