Abstract Objectives Defects in human cognition commonly result in clinical reasoning failures that can lead to diagnostic errors. Case presentation A 43-year-old female was brought to the emergency department with 4–5 days of confusion, disequilibrium resulting in several falls, and hallucinations. Further investigation revealed tachycardia, diaphoresis, mydriatic pupils, incomprehensible speech and she was seen picking at the air. Given multiple recent medication changes, there was initial concern for serotonin syndrome vs. an anticholinergic toxidrome. She then developed a fever, marked leukocytosis, and worsening encephalopathy. She underwent lumbar puncture and aspiration of an identified left ankle effusion. Methicillin sensitive staph aureus (MSSA) grew from blood, joint, and cerebrospinal fluid cultures within 18 h. She improved with antibiotics and incision, drainage, and washout of her ankle by orthopedic surgery. Conclusions Through integrated commentary on the diagnostic reasoning process from clinical reasoning experts, this case underscores how multiple cognitive biases can cascade sequentially, skewing clinical reasoning toward erroneous conclusions and driving potentially inappropriate testing and treatment. A fishbone diagram is provided to visually demonstrate the major factors that contributed to the diagnostic error. A case discussant describes the importance of structured reflection, a tool to promote metacognitive analysis, and the application of knowledge organization tools such as illness scripts to navigate these cognitive biases.
Clinical reasoning encompasses the process of data collection, synthesis, and interpretation to generate a working diagnosis and make management decisions. Situated cognition theory suggests that knowledge is relative to contextual factors, and clinical reasoning in urgent situations is framed by pressure of consequential, time-sensitive decision-making for diagnosis and management. These unique aspects of urgent clinical care may limit the effectiveness of traditional tools to assess, teach, and remediate clinical reasoning. Using two validated frameworks, a multidisciplinary group of clinicians trained to remediate clinical reasoning and with experience in urgent clinical care encounters designed the novel Rapid Evaluation Assessment of Clinical Reasoning Tool (REACT). REACT is a behaviorally anchored assessment tool scoring five domains used to provide formative feedback to learners evaluating patients during urgent clinical situations. A pilot study was performed to assess fourth-year medical students during simulated urgent clinical scenarios. Learners were scored using REACT by a separate, multidisciplinary group of clinician educators with no additional training in the clinical reasoning process. REACT scores were analyzed for internal consistency across raters and observations. Overall internal consistency for the 41 patient simulations as measured by Cronbach’s alpha was 0.86. A weighted kappa statistic was used to assess the overall score inter-rater reliability. Moderate reliability was observed at 0.56. To our knowledge, REACT is the first tool designed specifically for formative assessment of a learner’s clinical reasoning performance during simulated urgent clinical situations. With evidence of reliability and content validity, this tool guides feedback to learners during high-risk urgent clinical scenarios, with the goal of reducing diagnostic and management errors to limit patient harm.
Introduction The US Department of Defense (DoD) has adopted a model concept of the warrior athlete. Identifying latent disease that could compromise the military operator is critical to the warrior athlete concept. Cardiovascular complaints are the important problem recognized in service members evacuated from combat zones, and the incidence of sudden cardiac death in U.S. military recruits is comparable to or greater than that among National Collegiate Athletic Association Athletes. Nevertheless, the mandatory electrocardiogram (ECG) was removed from official U.S. military accession screening policy in 2002. Inclusion of ECG screening in high risk athletics is increasingly recognized as appropriate by professional organizations such as the American Heart Association and American Medical Society for Sports Medicine, though neither recommends ECG for generalized screening in large, low-risk populations. Materials and Methods The appropriate DoD instructions were reviewed in the context of recent literature regarding the sensitivity and specificity of ECG screening for prevention of sudden cardiac arrest or debilitating arrhythmias. Results Challenges to implementation of ECG as a screening modality in U.S. military accessions include clinician interpretation validity and reliability. Modern interpretation criteria and new interpretation technology each serve to mitigate these recognized limitations. Outside experience with implementation of modern ECG suggest potential benefits are significant in the highest risk military groups. Conclusion Prospective study of ECG screening is needed to determine the impact on cardiovascular outcomes in U.S. military populations.
Many athletes use anabolic-androgenic steroids (AAS) for physical enhancement but the magnitude of these gains and associated adverse effects has not been rigorously quantified. MEDLINE, EMBASE, Cochrane, SPORTDiscus, and PsycINFO were searched to identify randomized placebo-controlled trials of AAS in healthy exercising adults that reported one of the following outcomes: muscular strength, body composition, cardiovascular endurance, or power. Two authors appraised abstracts to identify studies for full-text retrieval; these were reviewed in duplicate to identify included studies. Study quality was assessed using the Cochrane method. Data were extracted in duplicate and pooled using the DerSimonian and Laird random effects model and to calculate the ratio of mean outcome improvement where possible. Pooled standardized mean difference (SMD) in muscle strength between AAS and placebo was 0.27 (95% confidence interval, 0.07-0.47; I2 = 12.7%; 21 studies). Change in strength was 52% greater in the AAS group compared to placebo. The SMD for change in lean mass between AAS and placebo was 0.62 (95% confidence interval, 0.35-0.89; I2 = 26%; 14 studies). Due to missing data, fat mass, cardiovascular endurance, power, and adverse effects were summarized qualitatively. Only 13 of 25 studies reported adverse effects including increased low density lipoprotein (LDL), decreased high density lipoprotein (HDL), irritability, and acne. In healthy exercising adults, AAS use is associated with a small absolute increase in muscle strength and moderate increase in lean mass. However, the transparency and completeness of adverse effect reporting varied, most studies were of short duration, and doses studied may not reflect actual use by athletes.
BACKGROUNDThree full doses of RTS,S/AS01 malaria vaccine provides partial protection against controlled human malaria parasite infection (CHMI) and natural exposure. Immunization regimens, including a delayed fractional third dose, were assessed for potential increased protection against malaria and immunologic responses.METHODSIn a phase 2a, controlled, open-label, study of healthy malaria-naive adults, 16 subjects vaccinated with a 0-, 1-, and 2-month full-dose regimen (012M) and 30 subjects who received a 0-, 1-, and 7-month regimen, including a fractional third dose (Fx017M), underwent CHMI 3 weeks after the last dose. Plasmablast heavy and light chain immunoglobulin messenger RNA sequencing and antibody avidity were evaluated. Protection against repeat CHMI was evaluated after 8 months.RESULTSA total of 26 of 30 subjects in the Fx017M group (vaccine efficacy [VE], 86.7% [95% confidence interval [CI], 66.8%-94.6%]; P < .0001) and 10 of 16 in the 012M group (VE, 62.5% [95% CI, 29.4%-80.1%]; P = .0009) were protected against infection, and protection differed between schedules (P = .040, by the log rank test). The fractional dose boosting increased antibody somatic hypermutation and avidity and sustained high protection upon rechallenge.DISCUSSIONSA delayed third fractional vaccine dose improved immunogenicity and protection against infection. Optimization of the RTS,S/AS01 immunization regimen may lead to improved approaches against malaria.CLINICAL TRIALS REGISTRATIONNCT01857869.
INTRODUCTION:Performance-enhancing drugs (PEDs) are commonly consumed in the United States with high prevalence of use in athlete populations and increased use by deployed service members. Many PEDs may contain anabolic-androgenic steroids (AAS), which are legally restricted and prohibited by many agencies due to their health risk.CASE DESCRIPTION:A unique case of acute pancreatitis associated with the use of the PED "Guerilla Warfare," a labeled AAS-containing supplement, is presented. The patient is a healthy 20-year-old male Marine who presented with multiple episodes of abdominal cramps each day for a month with decreased appetite and nonbilious vomiting. He reported a 6-week history of "Guerilla Warfare" PED use and review of systems identified fatigue and 12 lb reported weight loss. He presented with normal vital signs, tenderness in upper abdominal quadrants, elevated lipase (909 units/L), lactate dehydrogenase (193 units/L), and an enlarged pancreas with surrounding inflammation on computed tomography.SUMMARY:This constitutes the first report of acute pancreatitis with the use of "Guerilla Warfare," and the second reported case with the use of any AAS-containing PED. Increased awareness of significant PED-associated adverse effects by both the civilian and military communities is needed to better characterize these risks moving forward.
BACKGROUND:Supplement adulteration with anabolic-androgenic steroids (AAS) has been reported and AAS-associated drug-induced liver injury is clinically variable.OBJECTIVES:We present two cases of AAS-associated drug-induced liver injury in deployed service members, including the first report of clinical hepatotoxicity with desoxymethyltestosterone. We highlight variable hepatotoxicity patterns of AAS, raise concern with inaccurate supplement labeling and identify educational resources.METHODS:The first case presents with cholestatic jaundice following 10 weeks of prohormone use. Hepatobiliary imaging was unrevealing. Viral, autoimmune, and metabolic etiologies were excluded. Bilirubin normalized by 8 weeks after stopping the supplement. The second case presents with asymptomatic hepatocellular toxicity and marked dyslipidemia identified on service-related physical following 21 days of prohormone use. Aspartate aminotransferase and alanine aminotransferase normalized 4 weeks after supplement cessation; high-density lipoprotein and low-density lipoprotein returned to baseline at 8 weeks. Each supplement was volunteered for analytic testing.RESULTS:Supplement label contents did not match gas chromatography/mass spectrometry analysis; 3 of 4 supplements contained federally regulated AAS.CONCLUSIONS:AAS hepatotoxicity is clinically variable and dyslipidemia may be an important clinical indicator. False labeling introduces clinical risk and threatens mission readiness. Educational resources are available to facilitate information sharing. Supplement analysis informs of clinical risk of specific supplements and facilitates shared clinical decision-making.
BACKGROUND:Cardiac complications are a major cause of postoperative morbidity. The purpose of this study was to determine the rates, risk factors, and time of occurrence for cardiac complications within thirty days after primary unilateral total knee arthroplasty and total hip arthroplasty. METHODS:The American College of Surgeons National Surgical Quality Improvement Program data set from 2006 to 2011 was used to identify all total knee arthroplasties and total hip arthroplasties. Cardiac complications occurring within thirty days after surgery were the primary outcome measure. Patients were designated as having a history of cardiac disease if they had a new diagnosis or exacerbation of chronic congestive heart failure or a history of angina within thirty days before surgery, a history of myocardial infarction within six months, and/or any percutaneous cardiac intervention or other major cardiac surgery at any time. An analysis of the occurrence of all major cardiac complications and deaths within the thirty-day postoperative time frame was performed. RESULTS:For the 46,322 patients managed with total knee arthroplasty or total hip arthroplasty, the cardiac complication rate was 0.33% (n = 153) at thirty days postoperatively. In both the total knee arthroplasty and total hip arthroplasty groups, an age of eighty years or more (odds ratios [ORs] = 27.95 and 3.72), hypertension requiring medication (ORs = 4.74 and 2.59), and a history of cardiac disease (ORs = 4.46 and 2.80) were the three most significant predictors for the development of postoperative cardiac complications. Of the patients with a cardiac complication, the time of occurrence was within seven days after surgery for 79% (129 of the 164 patients for whom the time of occurrence could be determined). CONCLUSIONS:An age of eighty years or more, a history of cardiac disease, and hypertension requiring medication are significant risk factors for developing postoperative cardiac complications following primary unilateral total knee arthroplasty and total hip arthroplasty. Consideration should be given to a preoperative cardiology evaluation and co-management in the perioperative period for individuals with these risk factors.
Background and objective The clinical note documents the clinician's information collection, problem assessment, clinical management, and its used for administrative purposes. Electronic health records (EHRs) are being implemented in clinical practices throughout the USA yet it is not known whether they improve the quality of clinical notes. The goal in this study was to determine if EHRs improve the quality of outpatient clinical notes.Materials and methods A five and a half year longitudinal retrospective multicenter quantitative study comparing the quality of handwritten and electronic outpatient clinical visit notes for 100 patients with type 2 diabetes at three time points: 6 months prior to the introduction of the EHR (before-EHR), 6 months after the introduction of the EHR (after-EHR), and 5 years after the introduction of the EHR (5-year-EHR). QNOTE, a validated quantitative instrument, was used to assess the quality of outpatient clinical notes. Its scores can range from a low of 0 to a high of 100. Sixteen primary care physicians with active practices used QNOTE to determine the quality of the 300 patient notes.Results The before-EHR, after-EHR, and 5-year-EHR grand mean scores (SD) were 52.0 (18.4), 61.2 (16.3), and 80.4 (8.9), respectively, and the change in scores for before-EHR to after-EHR and before-EHR to 5-year-EHR were 18% (p<0.0001) and 55% (p<0.0001), respectively. All the element and grand mean quality scores significantly improved over the 5-year time interval.Conclusions The EHR significantly improved the overall quality of the outpatient clinical note and the quality of all its elements, including the core and non-core elements. To our knowledge, this is the first study to demonstrate that the EHR significantly improves the quality of clinical notes.
Background . The FDA recently approved tenofovir/emtricitabine as pre-exposure prophylaxis (PrEP) to prevent acquisition of HIV among adults. The CDC established guidance on prescribing PrEP. However, there is a paucity of data on how providers should implement PrEP into clinical practice, provider knowledge related to PrEP, and its cost-effectiveness. Methods . A voluntary, anonymous survey was conducted to evaluate the current knowledge, attitudes, and perceptions of PrEP among two groups of primarily infectious disease providers. The link to the 34-question survey was emailed to both the GreaterWashington Infectious Disease Society (GWIDS) and the Armed Forces Infectious Disease Society (AFIDS). This survey assessed provider demographics and the volume of HIV-infected patients in their practice in addition to their knowledge, prescribing patterns, and opinions regarding PrEP. Results . There were 105 responses – 20 (19%) were members of GWIDS, 58 (55%) were members of AFIDS, and 27 (25.7%) were part of both groups. All were physicians, and 94% were adult infectious disease specialists. The majority (60%) of knowledge questions were answered incorrectly. Of the respondents, 36 (34.3%) spent >25% of their time in HIV care. Those who spent >25% of their time in HIV care had a signi fi cantly higher percentage of correct answers in the knowledge component of the survey. Sixty-two (67%) respondents felt that the current literature supports the use of PrEP, 12 (13%) did not think the literature supports its use, and 18 (19.5%) were undecided. When asked whether the cost of PrEP was considered justi fi able, only 23 (25%) said yes, while 35 (38%) said no and 34 (37%) were undecided. Conclusion . There is a signi fi cant amount of uncertainty that remains regard- ing the use of PrEP. This survey demonstrates that knowledge related to the use of PrEP is lacking and suggests training is warranted to ensure providers become fa- miliar with CDC guidance. Given the lack of knowledge amongst providers who spent <25% of their time caring for HIV patients, organizations should consider restricting its use to providers who spend >25% of their time caring for HIV patients. Only a quarter of surveyed providers feel the cost is justi fi ed. Further re- search in this area is necessary to explore options for more cost effective methods of HIV prevention. Disclosures . All authors: No reported disclosures.
PurposeTo study medical students' letters of recommendation (LORs) from their applications to medical school to determine whether these predicted medical school performance, because many researchers have questioned LORs' predictive validity.MethodA retrospective cohort study of three consecutive graduating classes (2007-2009) at the Uniformed Services University of the Health Sciences was performed. In each class, the 27 students who had been elected into the Alpha Omega Alpha (AOA) Honor Medical Society were defined as top graduates, and the 27 students with the lowest cumulative grade point average (GPA) were designated as "bottom of the class" graduates. For each student, the first three LORs (if available) in the application packet were independently coded by two blinded investigators using a comprehensive list of 76 characteristics. Each characteristic was compared with graduation status (top or bottom of the class), and those with statistical significance related to graduation status were inserted into a logistic regression model, with undergraduate GPA and Medical College Admission Test score included as control variables.ResultsFour hundred thirty-seven LORs were included. Of 76 LOR characteristics, 7 were associated with graduation status (P = .05), and 3 remained significant in the regression model. Being rated as "the best" among peers and having an employer or supervisor as the LOR author were associated with induction into AOA, whereas having nonpositive comments was associated with bottom of the class students.ConclusionsLORs have limited value to admission committees, as very few LOR characteristics predict how students perform during medical school.
Background and objective The outpatient clinical note documents the clinician's information collection, problem assessment, and patient management, yet there is currently no validated instrument to measure the quality of the electronic clinical note. This study evaluated the validity of the QNOTE instrument, which assesses 12 elements in the clinical note, for measuring the quality of clinical notes. It also compared its performance with a global instrument that assesses the clinical note as a whole. Materials and methods Retrospective multicenter blinded study of the clinical notes of 100 outpatients with type 2 diabetes mellitus who had been seen in clinic on at least three occasions. The 300 notes were rated by eight general internal medicine and eight family medicine practicing physicians. The QNOTE instrument scored the quality of the note as the sum of a set of 12 note element scores, and its inter-rater agreement was measured by the intraclass correlation coefficient. The Global instrument scored the note in its entirety, and its inter-rater agreement was measured by the Fleiss κ. Results The overall QNOTE inter-rater agreement was 0.82 (CI 0.80 to 0.84), and its note quality score was 65 (CI 64 to 66). The Global inter-rater agreement was 0.24 (CI 0.19 to 0.29), and its note quality score was 52 (CI 49 to 55). The QNOTE quality scores were consistent, and the overall QNOTE score was significantly higher than the overall Global score (p=0.04). Conclusions We found the QNOTE to be a valid instrument for evaluating the quality of electronic clinical notes, and its performance was superior to that of the Global instrument.
Background: Electrocardiogram (ECG) with preparticipation evaluation (PPE) for athletes remains controversial in the United States and diagnostic accuracy of clinician ECG interpretation is unclear.This study aimed to assess reliability and validity of clinician ECG interpretation using expert-validated ECGs according to the 2010 European Society of Cardiology (ESC) interpretation criteria.Methods: This is a blinded, prospective study of diagnostic accuracy of clinician ECG interpretation.Anonymized ECGs were validated for normal and abnormal patterns by blinded expert interpreters according to the ESC interpretation criteria from October 2011 through March 2012.Six pairs of clinician interpreters were recruited from relevant clinical specialties in an academic medical center in March 2012.Each clinician interpreted 85 ECGs according to the ESC interpretation guidelines.Cohen and Fleiss' kappa, sensitivity, and specificity were calculated within specialties and across primary care and cardiology specialty groups.Results: Experts interpreted 189 ECGs yielding a kappa of 0.63, demonstrating "substantial" interrater agreement.A total of 85 validated ECGs, including 26 abnormals, were selected for clinician interpretation.The kappa across cardiology specialists was "substantial" and "moderate" across primary care (0.69 vs 0.52, respectively, P < 0.001).Sensitivity and specificity to detect abnormal patterns were similar between cardiology and primary care groups (sensitivity 93.3% vs 81.3%, respectively, P = 0.31; specificity 88.8% vs 89.8%, respectively, P = 0.91).Conclusions: Clinician ECG interpretation according to the ESC interpretation criteria appears to demonstrate limited reliability and validity.Before widespread adoption of ECG for PPE of U.S. athletes, further research of training focused on improved reliability and validity of clinician ECG interpretation is warranted.
US medical students have been placing increased importance on lifestyle when choosing their specialty. A 2003 study showed that lifestyle explained 55% of the changing trends in specialty choice of US allopathic medical students from 1992–2002.1 It seems intuitive that medical students’ description of which specialties have a favorable lifestyle would be well known. However, this has only been described once using scientific methods.2 That study, conducted by Newton et al, included more than 1,000 students from two medical schools who were to rate the importance of lifestyle in their specialty choice.2 Findings indicated that lifestyle played a significant role in medical students’ decisions to specialize in fields such as radiology, physical medicine/rehabilitation, emergency medicine, ophthalmology, anesthesia, urology, and dermatology, the specialties rated the most lifestyle friendly. Those students who valued lifestyle highly
Surveys are frequently used to collect data in graduate medical education (GME) settings.1 However, if a GME survey is not rigorously designed, the quality of the results is likely to be lower than desirable. In a recent editorial we introduced a framework for developing survey instruments.1 This systematic approach is intended to improve the quality of GME surveys and increase the likelihood of collecting survey data with evidence of reliability and validity. In this article we illustrate how researchers in medical education may operationalize this framework with examples from a survey we developed during the recent integration of 2 independent internal medicine (IM) residency programs.In 2010, the Department of Defense mandated the integration of the Walter Reed Army Medical Center in Washington, DC, with the National Naval Medical Center in Bethesda, Maryland. Prior to this integration, each hospital maintained independently accredited GME programs, including separate IM residency programs. During the merger these IM programs were asked to integrate seamlessly into a unified program.Despite many similarities, the 2 IM programs had important differences that might inhibit successful integration. For example, residents at Walter Reed were accustomed to an overnight, 24-hour call structure that was thought to bring a strong experiential learning element to the program. Yet, this call system risked violating work hour restrictions. Residents at the National Naval Medical Center worked under a night-float system that eliminated the risk of duty hour violations but increased the number of handoffs. Given these programmatic differences, we were interested in understanding how individuals in both programs thought the integration would affect the quality of the IM residency.Evidence-based design processes allow GME researchers to develop a set of survey items that every respondent is likely to interpret the same way, is able to respond to accurately, and is willing and motivated to answer. Six questions, introduced in our first editorial, can guide researchers through this systematic survey design framework. In the sections that follow, we illustrate each step of the process with examples from our own survey design project. Readers interested in the rationale behind the framework or in further details about each step are encouraged to consult our previous article.1Before creating a survey, it is important to consider the research question(s) of interest and the variables (or constructs) the researcher intends to measure. If, for example, the research question relates to the beliefs, opinions, or attitudes of the intended audience, a survey makes sense. On the other hand, if a researcher is more interested in assessing a directly observable behavior, such as residents' skill level for a particular clinical procedure, an observational tool may be a better choice.In the context of the residency merger, we wanted to understand how the integration effort would have an impact on key GME quality elements and program requirements as specified by the Accreditation Council for Graduate Medical Education (ACGME).2 We believed that understanding these factors from the residents' perspective would enable leadership to identify threats to successful integration as well as potential opportunities for process improvement. Further, we felt a survey was the appropriate tool because it would allow us to collect real-time feedback from participants rather than waiting for more objective outcomes, such as in-service exam scores and board scores, which, although valuable, would occur much later. The residents of each program were identified as the target population for our survey.A thorough review of the literature should be the next step in the GME survey design process. This step provides information about how the construct of interest has been defined in previous research. It also helps one identify existing survey scales that could be employed or adapted.Our review revealed that very little has been published on integration of GME programs. However, we found a number of examples of organizational change and restructuring in the business literature. Some of the most widely published, well-studied examples of organizational change were developed by William Bridges.3 We reviewed several of Bridges' survey instruments to identify common themes and items that might be applicable to our survey. Ultimately, we did not use any of these items verbatim. Instead, we adapted several items and used Bridges' work to better define our constructs of 3 separate but related ideas: current satisfaction, perceptions of the impact of the integration on the training experience, and beliefs about the readiness of the training programs to make the transition.GME researchers who find and wish to use or modify relevant survey scales can usually contact authors and request such use. It is worth noting, however, that “previously validated” survey scales require the collection of additional reliability and validity evidence in the specific research context, particularly if the scales are modified in any way or used in populations different from the initial survey audience. For publication, this additional evidence should be reported in the “Methods” and “Results” sections.The goal of this step is to create survey items that adequately represent the construct of interest in a language that respondents can easily understand. One important design consideration is the number of items needed to adequately assess the construct. There is no easy answer to this question. The ideal number of items depends on a number of factors, including the complexity of the construct and the level at which one intends to assess the construct (sometimes referred to as the “grain size” or level of abstraction at which the construct will be measured).4 In general, it is a good idea to develop more items than will ultimately be needed in the final scale because some items will undoubtedly be deleted or revised later in the design process.4The next challenge is to write a set of clear and unambiguous items. Writing good items is as much an art as it is a science. Nonetheless, there is a plethora of item-writing guidance—evidence-based, best practices—that should be used to guide the item-writing process.1,4–8 Reviewing these best practices is beyond the scope of this editorial; however, we have provided a summary of several evidence-based recommendations in table 1 to assist readers.To guide our item-writing process, we selected elements that the ACGME requires in an accredited IM residency.2 In particular, we chose elements that we felt were likely to be affected by reorganization and were also visible to the residents. Our initial draft had questions about every potentially relevant element; it quickly became clear that we had too many items. As such, we refocused our efforts on those issues most likely to be relevant to the intended respondents. In doing so, we were able to cut down our survey from 150 items to a more manageable 45 items.To illustrate other decisions we made during the survey development process, we focus the remainder of this editorial on our didactic quality scale. For this scale, we wanted to know the extent to which participants believed the merger would have an impact on the quality of their IM residency experience. For example, one item asked “How do you think the internal medicine integration will impact the educational quality of morning report?” Because we believed, based on our literature review and discussion with experts, that respondents might think the impact could either be positive or negative, we chose a bipolar scale with a midpoint of “neither positive nor negative impact” and endpoints of “extreme negative impact” on the low side of the response scale and “extreme positive impact” on the high side of the response scale (table 2).To assess survey content, GME researchers should ask a group of experts to review the items. This process, called content or expert validation, involves asking experts to review the draft survey items for clarity, relevance to the construct, and cognitive difficulty.9–11 Experts can also assist in identifying important aspects of the construct that may have been omitted during item development. “Experts” might include those more experienced in survey design, national content experts, or local colleagues knowledgeable about the specific construct of interest. The number of experts needed to conduct a content validation is typically small; 6 to 12 experts will often suffice.10,11Our 9 experts included staff from each IM program, as well as select faculty from our affiliated university who had expertise in survey design. Each expert received an invitation to participate in the content validation, along with the draft survey items and a document outlining our purpose and the specific aspects of the survey on which we wanted him or her to focus. Through this process our experts identified 6 items that were poorly focused; we eliminated these items from the survey. Our experts did not identify any content omissions in the scale items, and they agreed with our use of ACGME quality elements and program requirements as universal attributes of a high-quality GME training program. Finally, our experts identified several items that they felt were difficult to interpret, and these items were revised.After the draft survey items have undergone expert review, it is important to assess how the target population will interpret the items and response options. One way to do this is through a process known as cognitive interviewing or cognitive pretesting.12 Cognitive interviewing typically involves a face-to-face interview during which a respondent reads each item aloud and explains his or her thought process in selecting a particular response. This process allows one to verify that each respondent interprets the items as the researcher intended, performs the expected cognitive steps to generate an accurate response, and responds using the appropriate response anchors. Cognitive interviewing is a qualitative method that should be conducted with a handful of participants who are representative of the target population (typically 4–10 participants). This is a critical step to identify problems with question or response wording that may result in misinterpretation or bias. Cognitive interviewing should be conducted using a standardized methodology, and there are several systematic approaches that can be applied.12Given our small target population (approximately 68 residents in all), we opted to perform cognitive interviews with the chief residents and the program and associate program directors of each program, who we felt represented the closest available analogues to our target population. We performed cognitive interviews using both the think-aloud and retrospective verbal probing techniques.12 In the think-aloud method, each participant is provided with a copy of the draft survey, which he or she reviews while an interviewer reads from a standardized script. The interviewer reads each item, after which the participant is invited to think aloud while processing the question and selecting a response. Although time consuming, this method of interviewing is helpful in identifying items that fail to evoke the desired cognitive response. Retrospective verbal probing is an alternative method that consists of scripted questions administered just after the participant completes the entire survey. This approach conserves time and allows for a more authentic survey experience; however, retrospective verbal probing can introduce bias related to the participant's memory of each question. Through cognitive interviewing, we identified several small but important issues with our item wording, visual design, and survey layout, all of which were revised in our next iteration.Despite the best efforts of GME researchers during the aforementioned survey design process, some survey items may still be problematic.4 Thus, to gain additional validity evidence, pilot testing of the survey instrument should be performed. It is important to pilot test the survey using conditions identical or very similar to those planned for the full-scale survey. Descriptive data from pilot testing can then be used to evaluate the response distributions for individual items and scale composite scores. In addition, these data can be used to analyze item and composite score correlations, all of which are evidence of the internal structure of the survey and its relations to other variables. It is also worth noting that other, more advanced statistical techniques, such as factor analysis, can be used to ascertain the internal structure of a survey.13We pilot tested our survey on 14 residents from Walter Reed's IM program and 20 residents from the National Naval Medical Center's IM program; this represented 50% of the target population. table 2 presents the results of our pilot test for the didactic quality scale, which was designed to assess the perceived impact of the IM integration on specific didactic components within each training program. The scale included 8 questions and used a 5-point, Likert-type response scale ranging from “extreme negative impact” to “extreme positive impact.”After reviewing the item-level statistics using SPSS 20.0 (IBM Corp., New York), we calculated a Cronbach alpha coefficient to assess internal consistency reliability of the 8 items in our didactic quality scale. A Cronbach alpha coefficient can range from 0 to 1 and provides an assessment of the extent to which the scale items are related to one another. As we explained in our first editorial, a group of survey items designed to measure a given construct, such as our 8 items designed to measure didactic quality, should all exhibit moderate to strong positive correlations with one another. If they are not positively correlated, this suggests a potential problem with one or more of the items. It should be noted, however, that Cronbach alpha is sensitive to scale length; all other things being equal, a longer scale will generally have a higher Cronbach alpha. As such, a fairly easy way to increase a scale's internal consistency reliability is to add items. However, this increase in Cronbach alpha must be balanced with the potential for more response error due to an overly long survey that is exhausting for respondents.Although there is no set threshold value for internal consistency reliability, an alpha ≥0.75 is generally considered to be acceptable.14 For our didactic quality scale, the Cronbach alpha was .89, which indicated that our 8 items were highly correlated with one another, as expected. We then calculated a composite score (ie, an unweighted mean score of the 8 items) to create our didactic quality variable, and we inspected the descriptive statistics and histogram of the composite scores (table 2 and figure, respectively). The histogram was normally distributed, which suggested that our respondents were using almost all of the points along our response scale.After conducting the item- and scale-level analyses, as described above, it is reasonable to advance to the full-scale survey project. Of course, if the pilot results indicate poor reliability and/or surprising relationships (ie, one or more items in a given scale do not correlate with the other items as expected), researchers should consider revising existing items, removing poorly performing items, or drafting new items. If significant modifications are made to the survey, a follow-up pilot test of the revised survey may be in order. If only minor modifications are made, such as removing a handful of poorly performing items, it is reasonable to proceed directly to full-scale survey implementation.Developing a high-quality survey takes time. Nevertheless, the benefits of following a rigorous, systematic approach to survey design far outweigh the drawbacks. The example outlined here demonstrates our recommended survey design process. This approach can improve the quality of GME surveys and the likelihood of collecting survey data with evidence of reliability and validity in a given context, with a particular sample, and for a specific purpose. GME researchers are strongly encouraged to follow this or another systematic process when designing surveys and to report validity evidence (ie, the steps of their survey design process) so that readers can critically evaluate the quality of the survey instrument.