Behavioral interventions are widely used in clinical trials supported by the National Institutes of Health (NIH). When behavioral interventions involve group-formatted components and/or shared interventionists, they require special design and analytic methods not needed in trials that do not involve these features. The NIH Office of Disease Prevention (ODP) and the NIH Office of Behavioral and Social Sciences Research (OBSSR) offer resources to make it easier for investigators to use appropriate methods to evaluate these interventions. This commentary draws attention to these issues and highlights the ODP and OBSSR resources available to investigators. We urge investigators to take advantage of these resources to learn about and adopt appropriate sample size and analytic methods for trials to evaluate behavioral interventions so that their results will be reliable and reproducible. That is the best way to advance the science of behavioral interventions to improve health.
Background For the public health community, monitoring recently published articles is crucial for staying informed about the latest research developments. However, identifying publications about studies with specific research designs from the extensive body of public health publications is a challenge with the currently available methods. Objective Our objective is to develop a fine-tuned pretrained language model that can accurately identify publications from clinical trials that use a group- or cluster-randomized trial (GRT), individually randomized group-treatment trial (IRGT), or stepped wedge group- or cluster-randomized trial (SWGRT) design within the biomedical literature. Methods We fine-tuned the BioMedBERT language model using a dataset of biomedical literature from the Office of Disease Prevention at the National Institute of Health. The model was trained to classify publications into three categories of clinical trials that use nested designs. The model performance was evaluated on unseen data and demonstrated high sensitivity and specificity for each class. Results When our proposed model was tested for generalizability with unseen data, it delivered high sensitivity and specificity for each class as follows: negatives (0.95 and 0.93), GRTs (0.94 and 0.90), IRGTs (0.81 and 0.97), and SWGRTs (0.96 and 0.99), respectively. Conclusions Our work demonstrates the potential of fine-tuned, domain-specific language models to accurately identify publications reporting on complex and specialized study designs, addressing a critical need in the public health research community. This model offers a valuable tool for the public health community to directly identify publications from clinical trials that use one of the three classes of nested designs.
In June 2022, the NIH Office of Disease Prevention (ODP) issued a Call for Papers for a Supplemental Issue to Prevention Science on Design and Analytic Methods to Evaluate Multilevel Interventions to Reduce Health Disparities. ODP sought to bring together current thinking and new ideas about design and analytic methods for studies aimed at reducing health disparities, including strategies for balancing methodological rigor with design feasibility, acceptability, and ethical considerations. ODP was particularly interested in papers on design and analytic methods for parallel group- or cluster-randomized trials (GRTs), stepped-wedge GRTs, group-level regression discontinuity trials, and other methods appropriate for evaluating multilevel interventions. In this issue, we include 12 papers that report new methods, provide examples of strong applications of existing methods, or provide guidance on developing multilevel interventions to reduce health disparities. These papers provide examples showing that rigorous methods are available for the design and analysis of multilevel interventions to reduce health disparities.
Many individually randomized group treatment (IRGT) trials randomly assign individuals to study arms but deliver treatments via shared agents, such as therapists, surgeons, or trainers. Post-randomization interactions induce correlations in outcome measures between participants sharing the same agent. Agents can be nested in or crossed with trial arm, and participants may interact with a single agent or with multiple agents. These complications have led to ambiguity in choice of models but there have been no systematic efforts to identify appropriate analytic models for these study designs. To address this gap, we undertook a simulation study to examine the performance of candidate analytic models in the presence of complex clustering arising from multiple membership, single membership, and single agent settings, in both nested and crossed designs and for a continuous outcome. With nested designs, substantial type I error rate inflation was observed when analytic models did not account for multiple membership and when analytic model weights characterizing the association with multiple agents did not match the data generating mechanism. Conversely, analytic models for crossed designs generally maintained nominal type I error rates unless there was notable imbalance in the number of participants that interact with each agent.
NHLBI funded seven projects as part of the Disparities Elimination through Coordinated Interventions to Prevent and Control Heart and Lung Disease Risk (DECIPHeR) Initiative. They were expected to collaborate with community partners to (1) employ validated theoretical or conceptual implementation research frameworks, (2) include implementation research study designs, (3) include implementation measures as primary outcomes, and (4) inform our understanding of mediators and mechanisms of action of the implementation strategy. Several projects focused on late-stage implementation strategies that optimally and sustainably delivered two or more evidence-based multilevel interventions to reduce or eliminate cardiovascular and/or pulmonary health disparities and to improve population health in high-burden communities. Projects that were successful in the three-year planning phase transitioned to a 4-year execution phase. NHLBI formed a Technical Assistance Workgroup during the planning phase to help awardees refine study aims, strengthen research designs, detail analytic plans, and to use valid sample size methods. This paper highlights methodological and study design challenges encountered during this process. Important lessons learned included (1) the need for greater emphasis on implementation outcomes, (2) the need to clearly distinguish between intervention and implementation strategies in the protocol, (3) the need to address clustering due to randomization of groups or clusters, (4) the need to address the cross-classification that results when intervention agents work across multiple units of randomization in the same arm, (5) the need to accommodate time-varying intervention effects in stepped-wedge designs, and (6) the need for data-based estimates of the parameters required for sample size estimation.
Despite several ambitious national health initiatives to eliminate health disparities, spanning more than 4 decades, health disparities remain pervasive in the United States. In an attempt to bend the curve in disparities elimination, the National Heart, Lung, and Blood Institute (NHLBI) issued a funding opportunity on Disparities Elimination through Coordinated Interventions to Prevent and Control Heart and Lung Disease Risk (DECIPHeR) in March 2019. Seven implementation research centers and 1 research coordinating center were funded in September 2020 to plan, develop, and test effective implementation strategies for eliminating disparities in heart and lung disease risk. In the 16 articles presented in this issue of Ethnicity & Disease, the DECIPHeR Alliance investigators and their NHLBI program staff address the work accomplished in the first phase of this biphasic research endeavor. Included in the collection are an article on important lessons learned during technical assistance sessions designed to ensure scientific rigor in clinical study designs, and 2 examples of clinical study process articles. Several articles show the diversity of clinical and public health settings addressed including schools, faith-based settings, federally qualified health centers, and other safety net clinics. All strategies for eliminating disparities tackle a cardiovascular or pulmonary disease and related risk factors. In an additional article, NHLBI program staff address expectations in phase 2 of the DECIPHeR program, strategies to ensure feasibility of scaling and spreading promising strategies identified, and opportunities for translating the DECIPHeR research model to other chronic diseases for the elimination of related health disparities.
A Community Genetics carrier screening program for the Jewish community has operated on-site in high schools in Sydney (Australia) for 25 years. During 2020, in response to the COVID-19 pandemic, government-mandated social-distancing, ‘lock-down’ public health orders, and laboratory supply-chain shortages prevented the usual operation and delivery of the annual testing program. We describe development of three responses to overcome these challenges: (1) pivoting to online education sufficient to ensure informed consent for both genetic and genomic testing; (2) development of contactless telehealth with remote training and supervision for collecting genetic samples using buccal swabs; and (3) a novel patient and specimen identification ‘GeneTrustee’ protocol enabling fully identified clinical-grade specimens to be collected and DNA extracted by a research laboratory while maintaining full participant confidentiality and privacy. These telehealth strategies for education, consent, specimen collection and sample processing enabled uninterrupted delivery and operation of complex genetic testing and screening programs even amid pandemic restrictions. These tools remain available for future operation and can be adapted to other programs.
Multiple-period parallel group randomized trials (GRTs) analyzed with linear mixed models can represent time in mean models as continuous or categorical. If time is continuous, random effects are traditionally group- and member-level deviations from condition-specific slopes and intercepts and are referred to as random coefficients (RC) analytic models. If time is categorical, random effects are traditionally group- and member-level deviations from time-specific condition means and are referred to as repeated measures ANOVA (RM-ANOVA) analytic models. Longstanding guidance recommends the use of RC over RM-ANOVA for parallel GRTs with more than two periods because RC exhibited nominal type I error rates for both time parameterizations while RM-ANOVA exhibited inflated type I error rates when applied to data generated using the RC model. However, this recommendation was developed assuming a variance components covariance matrix for the RM-ANOVA, using only cross-sectional data, and explicitly modeling time × group variation. Left unanswered were how well RM-ANOVA with an unstructured covariance would perform on data generated according to the RC mechanism, if similar patterns would be observed in cohort data, and the impact of not modeling time × group variation if such variation was present in the data-generating model. Continuous outcomes for cohort and cross-sectional parallel GRT data were simulated according to RM-ANOVA and RC mechanisms at five total time periods. All simulations assumed time × group variation. We varied the number of groups, group size, and intra-cluster correlation. Analytic models using RC, RM-ANOVA, RM-ANOVA with unstructured covariance, and a Saturated random effects structure were applied to the data. All analytic models specified time × group random effects. The analytic models were then reapplied without specifying random effects for time × group. Results indicated the RC and saturated analytic models maintained the nominal type I error rate in all data sets, RM-ANOVA with an unstructured covariance did not avoid type I error rate inflation when applied to cohort RC data, and analytic models omitting time-varying group random effects when such variation exists in the data were prone to substantial type I error inflation unless the residual error variance is high relative to the time × group variance. The time × group RC and saturated analytic models are recommended as the default for multiple period parallel GRTs.
In cluster randomized trials (CRTs), the hierarchical nesting of participants (level 1) within clusters (level 2) leads to two conceptual populations: clusters and participants. When cluster sizes vary and the goal is to generalize to a hypothetical population of clusters, the unit average treatment effect (UATE), which averages equally at the cluster level rather than equally at the participant level, is a common estimand of interest. From an analytic perspective, when a generalized estimating equations (GEE) framework is used to obtain averaged treatment effect estimates for CRTs with variable cluster sizes, it is natural to specify an inverse cluster size weighted analysis so that each cluster contributes equally and to adopt an exchangeable working correlation matrix to account for within-cluster correlation. However, such an approach essentially uses two distinct weights in the analysis (i.e. both cluster size weights and covariance weights) and, in this article, we caution that it will lead to biased and/or inefficient treatment effect estimates for the UATE estimand. That is, two weights "make a wrong" or lead to poor estimation characteristics. These findings are based on theoretical derivations, corroborated via a simulation study, and illustrated using data from a CRT of a colorectal cancer screening program. We show that, an analysis with both an independence working correlation matrix and weighting by inverse cluster size is the only approach that always provides valid results for estimation of the UATE in CRTs with variable cluster sizes.
We can learn a great deal about the research questions being addressed in a field by examining the study designs used in that field. This manuscript examines the research questions being addressed in prevention research by characterizing the distribution and trends of study designs included in primary and secondary prevention research supported by the National Institutes of Health through grants and cooperative agreements, together with the types of prevention research, populations, rationales, exposures, and outcomes associated with each type of design. The Office of Disease Prevention developed a taxonomy to classify new extramural NIH-funded research projects and created a database with a representative sample of 14,523 research projects for fiscal years 2012–2019. The data were weighted to represent the entirety of the extramural research portfolio. Leveraging this dataset, the Office of Disease Prevention characterized the study designs proposed in NIH-funded primary and secondary prevention research applications. The most common study designs proposed in new NIH-supported prevention research applications during FY12-19 were observational designs (63.3%, 95% CI 61.5%–65.0%), analysis of existing data (44.5%, 95% CI: 42.7–46.3), methods research (23.9%, 95% CI: 22.3–25.6), and randomized interventions (17.2%, 95% CI: 16.1%–18.4%). Observational study designs dominated primary prevention research, while intervention designs were more common in secondary prevention research. Observational designs were more common for exposures that would be difficult to manipulate (e.g., genetics, chemical toxin, and infectious disease (not pneumonia/influenza or HIV/AIDS)), while intervention designs were more common for exposures that would be easier to manipulate (e.g., education/counseling, medication/device, diet/nutrition, and healthcare delivery). Intervention designs were not common for outcomes that are rare or have a long latency (e.g., cancer, neurological disease, Alzheimer’s disease) and more common for outcomes that are more common or where effects would be expected earlier (e.g., healthcare delivery, health related quality of life, substance use, and medication/device). Observational designs and analyses of existing data dominated, suggesting that much of the prevention research funded by NIH continues to focus on questions of association and on questions of identification of risk and protective factors. Randomized and non-randomized intervention designs were included far less often, suggesting that a much smaller fraction of the NIH prevention research portfolio is focused on questions of whether interventions can be used to modify risk or protective factors or to change some other health-related biomedical or behavioral outcome. The much heavier focus on observational studies is surprising given how much we know already about the leading risk factors for death and disability in the USA, because those risk factors account for 74% of the county-level mortality in the USA, and because they play such a vital role in the development of clinical and public health guidelines, whose developers often weigh results from randomized trials much more heavily than results from observational studies. Improvements in death and disability nationwide are more likely to derive from guidelines based on intervention research to address the leading risk factors than from additional observational studies.
Background. This article identifies the most influential methods reports for group-randomized trials and related designs published through 2020. Many interventions are delivered to participants in real or virtual groups or in groups defined by a shared interventionist so that there is an expectation for positive correlation among observations taken on participants in the same group. These interventions are typically evaluated using a group- or cluster-randomized trial, an individually randomized group treatment trial, or a stepped wedge group- or cluster-randomized trial. These trials face methodological issues beyond those encountered in the more familiar individually randomized controlled trial. Methods. PubMed was searched to identify candidate methods reports; that search was supplemented by reports known to the author. Candidate reports were reviewed by the author to include only those focused on the designs of interest. Citation counts and the relative citation ratio, a new bibliometric tool developed at the National Institutes of Health, were used to identify influential reports. The relative citation ratio measures influence at the article level by comparing the citation rate of the reference article to the citation rates of the articles cited by other articles that also cite the reference article. Results. In total, 1043 reports were identified that were published through 2020. However, 55 were deemed to be the most influential based on their relative citation ratio or their citation count using criteria specific to each of the three designs, with 32 group-randomized trial reports, 7 individually randomized group treatment trial reports, and 16 stepped wedge group-randomized trial reports. Many of the influential reports were early publications that drew attention to the issues that distinguish these designs from the more familiar individually randomized controlled trial. Others were textbooks that covered a wide range of issues for these designs. Others were “first reports” on analytic methods appropriate for a specific type of data (e.g. binary data, ordinal data), for features commonly encountered in these studies (e.g. unequal cluster size, attrition), or for important variations in study design (e.g. repeated measures, cohort versus cross-section). Many presented methods for sample size calculations. Others described how these designs could be applied to a new area (e.g. dissemination and implementation research). Among the reports with the highest relative citation ratios were the CONSORT statements for each design. Conclusions. Collectively, the influential reports address topics of great interest to investigators who might consider using one of these designs and need guidance on selecting the most appropriate design for their research question and on the best methods for design, analysis, and sample size.
IntroductionThis manuscript characterizes primary and secondary prevention research in humans and related methods research funded by NIH in 2012‒2019.MethodsThe NIH Office of Disease Prevention updated its prevention research taxonomy in 2019‒2020 and applied it to a sample of 14,523 new extramural projects awarded in 2012–2019. All projects were coded manually for rationale, exposures, outcomes, population focus, study design, and type of prevention research. All results are based on that manual coding.ResultsTaxonomy updates resulted in a slight increase, from an average of 16.7% to 17.6%, in the proportion of prevention research awards for 2012‒2017; there was a further increase to 20.7% in 2019. Most of the leading risk factors for death and disability in the U.S. were observed as an exposure or outcome in <5% of prevention research projects in 2019 (e.g., diet, 3.7%; tobacco, 3.9%; blood pressure, 2.8%; obesity, 4.4%). Analysis of existing data became more common (from 36% to 46.5%), whereas randomized interventions became less common (from 20.5% to 12.3%). Randomized interventions addressing a leading risk factor in a minority health or health disparities population were uncommon.ConclusionsThe number of new NIH awards classified as prevention research increased to 20.7% in 2019. New projects continued to focus on observational studies and secondary data analysis in 2018 and 2019. Additional research is needed to develop and test new interventions or develop methods for the dissemination of existing interventions, which address the leading risk factors, particularly in minority health and health disparities populations.
* Abbreviation: NIH — : National Institutes of Health The National Institutes of Health (NIH) has a special interest group on childhood screening that is staffed jointly by the Eunice Kennedy Shriver National Institute of Child Health and Human Development and the Office of Disease Prevention in the Office of the NIH Director. After a series of internal meetings, the special interest group decided that the most compelling issue to tackle was the lack of an evidence base to support screening recommendations for exposures, behaviors, and conditions that are part of current routine pediatric care. To stimulate methodologic innovation to address this problem, the NIH convened an in-person workshop in Bethesda, Maryland, on May 9–10, 2019, entitled “Methods for Assessing the Impact of Screening in Childhood on Health Outcomes.” Workshop objectives included (1) conceptualizing the universe of child health outcomes pertinent to screening; (2) identifying methodologic challenges in assessing child health outcomes; (3) discussing novel and rigorous approaches to assessing these outcomes; … Address correspondence to Diana W. Bianchi, MD, Eunice Kennedy Shriver National Institute of Child Health and Human Development, National Institutes of Health, Building 31, Room 2A03, 31 Center Dr, Bethesda, MD 20892. E-mail: diana.bianchi{at}nih.gov
Objectives: Multiple myeloma (MM) is a malignant plasma cell neoplasm, requiring the integration of clinical examination, laboratory and radiological investigations for diagnosis. Detection and isotypic identification of the monoclonal protein(s) and measurement of other relevant biomarkers in serum and urine are pivotal analyses. However, occasionally this approach fails to characterize complex protein signatures. Here we describe the development and application of next generation mass spectrometry (MS) techniques, and a novel adaptation of immunofixation, to interrogate non-canonical monoclonal immunoproteins. Methods: Immunoprecipitation immunofixation (IP-IFE) was performed on a Sebia Hydrasys Scan2. Middle-down de novo sequencing and native MS were performed with multiple instruments (21T FT-ICR, Q Exactive HF, Orbitrap Fusion Lumos, and Orbitrap Eclipse). Post-acquisition data analysis was performed using Xcalibur Qual Browser, ProSight Lite, and TDValidator. Results: We adapted a novel variation of immunofixation electrophoresis (IFE) with an antibody-specific immunosubtraction step, providing insight into the clonal signature of gamma-zone monoclonal immunoglobulin (M-protein) species. We developed and applied advanced mass spectrometric techniques such as middle-down de novo sequencing to attain in-depth characterization of the primary sequence of an M-protein. Quaternary structures of M-proteins were elucidated by native MS, revealing a previously unprecedented non-covalently associated heterotetrameric immunoglobulin. Conclusions: Next generation proteomic solutions offer great potential for characterizing complex protein structures and may eventually replace current electrophoretic approaches for the identification and quantification of M-proteins. They can also contribute to greater understanding of MM pathogenesis, enabling classification of patients into new subtypes, improved risk stratification and the potential to inform decisions on future personalized treatment modalities.
In patients with immunoglobulin light-chain (AL) amyloidosis, depth of hematologic response correlates with both organ response and overall survival. Our group has demonstrated that screening with a matrix-assisted laser desorption/ionization-time-of-flight (TOF) mass spectrometry (MS) is a quick, sensitive, and accurate means to diagnose and monitor the serum of patients with plasma cell disorders. Microflow liquid chromatography coupled with electrospray ionization and quadrupole TOF MS adds further sensitivity. We identified 33 patients with AL amyloidosis who achieved amyloid complete hematologic response, who also had negative bone marrow by six-color flow cytometry, and who had paired serum samples to test by MS. These samples were subjected to blood MS. Four patients (12%) were found to have residual disease by these techniques. The presence of residual disease by MS was associated with a poorer time to progression (at 50 months 75% versus 13%, p = 0.003). MS of the blood out-performed serum and urine immunofixation, the serum immunoglobulin free light chain, and six-color flow cytometry of the bone marrow in detecting residual disease. Additional studies that include urine MS and next-generation techniques to detect clonal plasma cells in the bone marrow will further elucidate the full potential of this technique.
Declining levels of physical activity probably contribute to the increasing prevalence of overweight in US youth. In this study, the authors examined cross-sectional and longitudinal associations between physical activity and body composition in sixthand eighth-grade girls. In 2003, girls were recruited from six US states as part of the Trial of Activity for Adolescent Girls. Physical activity was measured using 6 days of accelerometry, and percentage of body fat was calculated using an ageand ethnicity-specific prediction equation. Sixth-grade girls with an average of 12.8 minutes of moderate-to-vigorous physical activity (MVPA) per day (15th percentile) were 2.3 times (95% confidence interval: 1.52, 3.44) more likely to be overweight than girls with 34.7 minutes of MVPA per day (85th percentile), and their percent body fat was 2.64 percentage points greater (95% confidence interval: 1.79, 3.50). Longitudinal analyses showed that percent body fat increased 0.28 percentage points less in girls with a 6.2-minute increase in MVPA than in girls with a 4.5-minute decrease (85th and 15th percentiles of change). Associations between MVPA in sixth grade and incidence of overweight in eighth grade were not detected. More population-based research using Correspondence to Dr. June Stevens, CB 7400, Department of Nutrition, School of Public Health, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599 (e-mail: June_stevens@unc.edu).. NIH Public Access Author Manuscript Am J Epidemiol. Author manuscript; available in PMC 2008 June 1. Published in final edited form as: Am J Epidemiol. 2007 December 1; 166(11): 1298–1305. N IH PA Athor M anscript N IH PA Athor M anscript N IH PA Athor M anscript objective physical activity and body composition measurements is needed to make evidence-based physical activity recommendations for US youth.
High-sensitivity mass spectrometry assays are available to detect monoclonal immunoglobulins. To better assess the prevalence of monoclonal gammopathy of undetermined significance (MGUS), we identified 300 patients diagnosed with MGUS or related gammopathy who had a prior negative work-up for monoclonal proteins as part of the Olmsted County MGUS screening study. Two mass spectrometry-based detection methods (matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) and monoclonal immunoglobulin rapid accurate mass measurements (miRAMM) along with traditional immunofixation were performed on the Olmsted baseline and MGUS diagnostics serum samples. Among the 226 patients considered negative for MGUS based on protein electrophoresis and serum-free light-chain assay, a monoclonal protein could be detected at baseline in 24 patients (10.6%) by immunofixation, 113 patients (50%) by MADLI-TOF mass spectrometry, and 149 patients (65.9%) by miRAMM mass spectrometry. In addition, using miRAMM, some patients demonstrated an oligoclonal to monoclonal transition giving insight into the origin of MGUS. Using the sensitive miRAMM, MGUS is present in 887 of 17,367 persons from the Olmsted County cohort, translating into a prevalence of 5.1% among persons 50 years of age and older. This represents the most accurate prevalence estimate of MGUS thus far.