The three-step United States Medical Licensing Examination (USMLE) was developed by the National Board of Medical Examiners and the Federation of State Medical Boards to provide medical licensing authorities a uniform evaluation system on which to base licensure. The test results appear to be a good measure of content knowledge and a reasonable predictor of performance on subsequent in-training and certification exams. Nonetheless, it is disconcerting that the test preoccupies so much of students' attention with attendant substantial costs (in time and money) and mental and emotional anguish. There is an increasingly pervasive practice of using the USMLE score, especially the Step 1 component, to screen applicants for residency. This is despite the fact that the test was not designed to be a primary determinant of the likelihood of success in residency. Further, relying on Step 1 scores to filter large numbers of applications has unintended consequences for students and undergraduate medical education curricula. There are many other factors likely to be equally or more predictable of performance during residency. The authors strongly recommend a move away from using test scores alone in the applicant screening process and toward a more holistic evaluation of the skills, attributes, and behaviors sought in future health care providers. They urge more rigorous study of the characteristics of students that predict success in residency, better assessment tools for competencies beyond those assessed by Step 1 that are relevant to success, and nationally comparable measures from those assessments that are easy to interpret and apply.
To the Editor: Prober and colleagues1 make a strong, rational argument for not using United States Medical Licensing Examination (USMLE) Step 1 scores to screen graduate medical education (GME) applicants. The USMLE was designed to certify minimum competency for medical licensure. Unfortunately, residency program directors are inappropriately using USMLE scores to screen applicants. The authors recommend development of systems to collect competency-based evidence that predicts success in GME. We would like to suggest a solution to institute such a change. First, the National Board of Medical Examiners should stop reporting three-digit scores for the USMLE Steps and provide only pass/fail scores and a list of areas of strength and areas for improvement based on performance. To promote habits of lifelong learning, we should require students to address these recommended areas for improvement and reflect on the efforts they make to improve. Program directors would use this evidence of efforts towards lifelong learning instead of USMLE scores when reviewing applications. The mission of medical education is to improve the health of the nation’s population. Failure to achieve this mission can jeopardize GME funding and the entire medical education structure. It is accepted that assessment drives learning and should inform curriculum design. Thus our second suggestion is that major consideration be given by GME programs to evidence of successful participation in a population health initiative during medical school. This will require evidence of working in interprofessional teams, using data for quality improvement and knowledge of population health. This requirement will drive medical schools to help students participate in innovative projects for population health and use systems for 360-degree feedback to assess students in a broad range of competencies. This process will generate the type of data needed by program directors and will complement the evidence for lifelong learning described above. These goals are actually quite feasible. At the Cleveland Clinic Lerner College of Medicine, we have had a “No Tests, No Grades” philosophy since 2004. Students learn in collaborative small groups and receive formative feedback from peers and faculty, which they use to strive for constant improvement and write a reflective essay on their progress each year.2 A system such as this could be modified to implement the proposals described above. Neil B. Mehta, MBBS, MSAssistant dean of education informatics and technology, Cleveland Clinic Lerner College of Medicine at Case Western Reserve University, Cleveland, Ohio; [email protected] Alan Hull, MD, PhDAssociate dean of curricular affairs, Cleveland Clinic Lerner College of Medicine at Case Western Reserve University, Cleveland, Ohio. James Young, MDExecutive dean, Cleveland Clinic Lerner College of Medicine at Case Western Reserve University, Cleveland, Ohio.
We appreciate the letters in response to our recent Commentary.1 We hope that our Commentary and this exchange continue to stimulate others to consider strategies that enhance medical education while clarifying the role of standardized testing in the holistic review of applicants for residency. A br
With assessment systems that are adequate, robust, comprehensive, as well as responsive to local and regional needs, should the location of the medical education institution be irrelevant? Adequate assessment is determined by local needs, along with accepted minimum global standards of practice. If an assessment system is robust, it should be able to predict future behavior and performance to some degree. A comprehensive system would include assessment of all relevant competencies. In order to achieve comprehensiveness, new approaches are needed to demonstrate mastery of competencies that is now inferred from medical school and graduate medical education participation. These are likely to require a novel approach to assessment – gathering natural, real world data longitudinally rather than only through point-in-time tests. Increasingly the world of assessment may be able to provide tools and data that offer individualized assurances of competence.
The introduction of the Step 2 Clinical Skills program to the United States Medical Licensing Examination was a vital step in assuring that physicians seeking a license to practice medicine in this country demonstrate the patient-centered skills that are essential to practice.
To the Editor: The timely importance of the article by McGaghie et al1 cannot be overstated. As a former residency program director and now as senior associate dean for medical education, I have long maintained that the common practice by certain subspecialty residency programs of using United States Medical Licensing Examination (USMLE) scores to screen and/or censor applicants is wholly unfair and seemingly without validity. With their article, McGaghie and colleagues have finally proved the lack of validity for this practice. For residency programs to persist in perpetuating the notion that high board scores = a better resident, when no correlation between USMLE scores and objective measures of trainees' clinical skills has ever been demonstrated, is wholly inconsistent with their teaching residents to adhere to practices that are evidence based. I propose that we, as an academic medical community, take the implications of their work one step further and urge the National Board of Medical Examiners (NBME) to stop the release of students' numerical examination scores. The Step examinations were designed to contribute to medical licensure decisions. While the NBME acknowledges that there are concomitant secondary uses of the scores by third parties—such as in postgraduate residency selection—these uses are not validated. Shouldn't they be stopped? It is a travesty that student affairs deans are annually forced to explain to perfectly capable, sometimes truly outstanding, medical students that their career dreams of being in “X” specialty are categorically eliminated simply because their USLME Step 1 scores were insufficiently high.2 This change will be difficult for some residency programs. It may force them to identify those traits, skill sets, and attitudes that best predict excellence in their particular specialties rather than simply focusing on a number. It might free up medical schools to provide innovative educational opportunities for students rather than devoting excessive curricular time to board preparation. It could result in some previously “unqualified by board scores but otherwise excellent” students to achieve their dreams. It would be a fairer system all around. Jeffrey G. Wong, MD Senior associate dean for medical education and professor of internal medicine, Medical University of South Carolina, Charleston, South Carolina; [email protected].
To the Editor: In their March article, Kerfoot et al1incorrectly report that [T]he National Board of Medical Examiners, the Educational Commission for Foreign Medical Graduates, and the Federation of State Medical Boards have recently endorsed a plan to replace the current three-step licensure examination system with two gateway examination sequences. These will be administered near the end of medical school and at the end of internship. We wish to correct this misperception. In 2006 the United States Medical Licensing Examination (USMLE) program convened the Committee to Evaluate the USMLE Program (CEUP) as part of a self-study process that yielded recommendations for changes to the USMLE examination sequence. Among its findings,2 the CEUP recommended examinations to support state medical board decisions at two points: as physicians prepare to enter supervised practice, that is, residency (decision point 1), and as physicians prepare for full licensure and unsupervised independent practice (decision point 2). This recommendation has sometimes been misconstrued, as Kerfoot and colleagues did, to mandate two examinations or two sequences of examinations to be completed immediately before the two licensing decisions described above. In fact, no plan has been announced or is under development to aggregate all decision-point-1 assessments or move them to near the end of medical school. State medical boards currently determine readiness for supervised practice during postgraduate training based on knowledge of fundamental science, clinical medicine, and clinical skill. In the current USMLE program, separate examinations provide the basis for these decisions. There will also be multiple components in the new USMLE program. Candidates may elect to take components in closer proximity, given the higher degree of integration of fundamental science and clinical medicine that will be present. The USMLE program provides periodic updates as plans proceed to implement recommendations arising from CEUP's review of the USMLE program. Interested readers are encouraged to view links at www.usmle.org/cru for implementation plans and progress reports. Peter J. Katsufrakis, MD, MBA Vice president for assessment programs, National Board of Medical Examiners, Philadelphia, Pennsylvania; [email protected]. Peter V. Scoles, MD Senior vice president for assessment programs, National Board of Medical Examiners, Philadelphia, Pennsylvania. Donald E. Melnick, MD President, National Board of Medical Examiners, Philadelphia, Pennsylvania.
In 2008, Congress amended the Americans with Disabilities Act (ADA) to relax court-imposed limitations on evidence required to warrant protection under the ADA. Since passage of the ADA in 1990, medicine has focused not on evaluating the types of accommodations that would best balance the interests of individuals with disabilities, institutions, and patients but, rather, on the question of whether individuals seeking protection under the law qualify for disability accommodations at all. The medical profession should refocus on the nature of accommodations provided to those with disabilities. In doing so, the intent to support disabled persons seeking careers in medicine must be balanced with ethical obligations to protect patient welfare. Medical schools, graduate medical education programs, licensing and certifying authorities, and assessment organizations should work together to establish evidence-based minimum criteria for the physical and cognitive capabilities required of every physician.
The United States and Canada both have long-standing, highly developed national systems of assessment for medical-licensure based outside the institutions of medical education. This commentary reviews those programs and explores some of the reasons for their implementation and retention for nearly a century. The North American experience may be relevant to dialog about national or European assessments for medical practice.
The author outlines the intertwining roles of the Educational Commission for Foreign Medical Graduates (ECFMG), which is celebrating its 50th anniversary in 2006, and the National Board of Medical Examiners (NBME) in meeting needs for assessment of international medical graduates. Both organizations had early histories focused on a protective role: ensuring that only the most qualified foreign-trained doctors could train or practice in the United States. The two organizations have interacted throughout the ECFMG's 50-year history to improve the assessment of internationally trained doctors. As both the ECFMG and the NBME have matured, their missions have expanded to include improvement of medical education and assessment around the world. Much of the success of each organization in fulfilling its mission can be attributed to their close collaboration through the past 50 years.
This chapter contains section titled: Overview of Professional Regulation of Doctors in the Us Computer-Based Testing in the Regulation of Us Health Professionals Examination for Medical Licensure in the Us—Usmle Computer-Based Usmle CBT Allows Innovative Test Formats-Primumtm™ CLINICAL CASE SIMULATIONS Lessons Learned in Implementing CBT Operational Challenges of Computer-Based Testing Future Potential from CBT References
This article has three key points. The first proposes and illustrates a model for planning effective continuing medical education (CME) and continuing professional development (CPD) and how assessment might fit into it. The second reviews major trends in assessment, particularly with regard to regulation and CME. The third addresses challenges for CME and CPD.
Medical training is undergoing extensive revision in France. A nationwide comprehensive clinical competency examination will be administered for the first time in 2004, relying exclusively on essay-questions. Unfortunately, these questions have psychometric shortcomings, particularly their typically low reliability. High score reliability is mandatory in a high-stakes context. The National Board of Medical Examiners-designed multiple choice-questions (MCQ) are well adapted to assess clinical competency with a high reliability score. The purpose of this study was to test the hypothesis that French medical students could take an American-designed and French-adapted comprehensive clinical knowledge examination with this MCQ format. Two hundred and eighty five French students, from four Medical Schools across France, took an examination composed of 200 MCQs under standardized conditions. Their scores were compared with those of American students. This examination was found assess French students' clinical knowledge with a high level of reliability. French students' scores were slightly lower than those of American students, mostly due to a lack of familiarity with this particular item format, and a lower motivational level. Another study is being designed, with a larger group, to address some of the shortcomings of the initial study. If these preliminary results are replicated, the MCQ format might be a more defendable and sensible alternative to the proposed essay questions.
PURPOSE:The French government, as part of medical education reforms, has affirmed that an examination program for national residency selection will be implemented by 2004. The purpose of this study was to develop a French multiple-choice (MC) examination using the National Board of Medical Examiners' (NBME) expertise and materials.METHOD:The Evaluation Standardisée du Second Cycle (ESSC), a four-hour clinical sciences examination, was administered in January 2002 to 285 medical students at four university test sites in France. The ESSC had 200 translated and adapted MC items selected from the Comprehensive Clinical Sciences Examination (CCSE), an NBME subject test.RESULTS:Less than 10% of the ESSC items were rejected as inappropriate to French practice. Also, the distributions of ESSC item characteristics were similar to those reported with the CCSE. The ESSC also appeared to be very well targeted to examinees' proficiencies and yielded a reliability coefficient of.91. However, because of a higher word count, the ESSC did show evidence of speededness. Regarding overall performance, the mean proficiency estimate for French examinees was about 0.4 SD below that of a CCSE population.CONCLUSIONS:This study provides strong evidence for the usefulness of the model adopted in this first collaborative effort between the NBME and a consortium of French medical schools. Overall, the performance of French students was comparable to that of CCSE students, which was encouraging given the differences in motivation and the speeded nature of the French test. A second phase with the participation of larger numbers of French medical schools and students is being planned.
Practice inevitably narrows over time. Therefore, testing of established doctors requires that their assessment be tailored to a far narrower practice than is appropriate for testing of new doctors who have not yet differentiated. In this paper, we address the conceptual challenges of tailoring physician assessment to individual practice. Testing of established doctors needs to reflect that physicians specialise, often in idiosyncratic ways; otherwise, the testing will not be credible among established doctors and will not reflect the realities of their practice. Despite the importance of these goals, the conceptual and methodological challenges of creating tailored assessments remain daunting.