BACKGROUND:Less than 75% of people prescribed antihypertensive medication are still using treatment after 6 months. Physicians determine treatment, educate patients, manage side effects, and influence patient knowledge and motivation. Although physician communication ability likely influences persistence, little is known about the importance of medical management skills, even though these abilities can be enhanced through educational and practice interventions. The purpose of this study was to determine whether a physician's medical management and communication ability influence persistence with antihypertensive treatment.METHODS:This was a population-based study of 13,205 hypertensive patients who started antihypertensive medication prescribed by a cohort of 645 physicians entering practice in Quebec, Canada, between 1993 and 2007. Medical Council of Canada licensing examination scores were used to assess medical management and communication ability. Population-based prescription and medical services databases were used to assess starting therapy, treatment changes, comorbidity, and persistence with antihypertensive treatment in the first 6 months.RESULTS:Within 6 months after starting treatment, 2926 patients (22.2%) had discontinued all antihypertensive medication. The risk of nonpersistence was reduced for patients who were treated by physicians with better medical management (odds ratio per 2-SD increase in score, 0.74; 95% confidence interval, 0.63-0.87) and communication (0.88; 0.78-1.00) ability and with early therapy changes (odds ratio, 0.45; 95% confidence interval, 0.37-0.54), more follow-up visits, and nondiuretics as the initial choice of therapy. Medical management ability was responsible for preventing 15.8% (95% confidence interval, 7.5%-23.3%) of nonpersistence.CONCLUSION:Better clinical decision-making and data collection skills and early modifications in therapy improve persistence with antihypertensive therapy.
BACKGROUND:A physician's personal and professional characteristics constitute only one, and not necessarily the most important, determining factor of clinical performance. Our study assessed how physician, organizational and systemic factors affect family physicians' performance.METHOD:Our study examined 532 family practitioners who were randomly selected for peer assessment by the College of Physicians and Surgeons of Ontario. A series of multivariate regression analyses examined the impact of physician factors (e.g., demographics, certification) on performance scores in five clinical areas: acute care, chronic conditions, continuity of care and referrals, well care and records. A second series of regressions examined the simultaneous effects of physician, organizational (e.g., practice volume, hours worked, solo practice) and systemic factors (e.g., northern practice location, community size, physician-to-population ratio).RESULTS:OUR STUDY HAD THREE KEY FINDINGS: (a) physician factors significantly influence performance but do not appear to be nearly as important as previously thought; (b) organizational and systemic factors have significant effects on performance after the effects of physician factors are controlled; and (c) physician, organizational and systemic factors have varying effects across different dimensions of clinical performance.CONCLUSIONS:We discuss the implications of our results for performance improvement and physician governance insofar as both need to consider the broader environmental context of medical practice.
Objectives This study aimed to determine if national licensing examinations that measure medical knowledge (QE1) and clinical skills (QE2) predict the quality of care delivered by doctors in future practice.
CONTEXT Problem-based learning (PBL) is an educational strategy designed to enhance self-assessment, self-directed learning and lifelong learning. The present study examines a peer review programme to determine whether the impact of PBL on continuing competence can be detected in practice.OBJECTIVES This study aimed to establish whether McMaster graduates who graduated between 1972 and 1991 were any less likely to be identified as having issues of competence by a systematic peer review programme than graduates of other Ontario medical schools.METHODS We identified a total of 1166 doctors who had graduated after 1972 and had completed a mandated peer review programme. Of these, 108 had graduated from McMaster and 857 from other Canadian schools. School of graduation was cross-tabulated against peer rating. A secondary analysis examined predictors of ratings using multiple regression.RESULTS We found that 4% of McMaster graduates and 5% of other graduates were deemed to demonstrate cause for concern or serious concern, and that 24% of McMaster doctors and 28% of other doctors were rated as excellent. These differences were not significant. Multiple regression indicated that certification by family medicine or a specialty, female gender and younger age were all predictors of practice outcomes, but school of graduation was not.CONCLUSIONS There is no evidence from this study that PBL graduates are better able to maintain competence than graduates of conventional schools. The study highlights potential problems in attempting to link undergraduate educational interventions to doctor performance outcomes.
CONTEXT Poor patient-physician communication increases the risk of patient complaints and malpractice claims. To address this problem, licensure assessment has been reformed in Canada and the United States, including a national standardized assessment of patient-physician communication and clinical history taking and examination skills. OBJECTIVE To assess whether patient-physician communication examination scores in the clinical skills examination predicted future complaints in medical practice. DESIGN, SETTING, AND PARTICIPANTS Cohort study of all 3424 physicians taking the Medical Council of Canada clinical skills examination between 1993 and 1996 who were licensed to practice in Ontario and/or Quebec. Participants were followed up until 2005, including the first 2 to 12 years of practice. MAIN OUTCOME MEASURE Patient complaints against study physicians that were filed with medical regulatory authorities in Ontario or Quebec and retained after investigation. Multivariate Poisson regression was used to estimate the relationship between complaint rate and scores on the clinical skills examination and traditional written examination. Scores are based on a standardized mean (SD) of 500 (100). RESULTS Overall, 1116 complaints were filed for 3424 physicians, and 696 complaints were retained after investigation. Of the physicians, 17.1% had at least 1 retained complaint, of which 81.9% were for communication or quality-of-care problems. Patient-physician communication scores for study physicians ranged from 31 to 723 (mean [SD], 510.9 [91.1]). A 2-SD decrease in communication score was associated with 1.17 more retained complaints per 100 physicians per year (relative risk [RR], 1.38; 95% confidence interval [CI], 1.18-1.61) and 1.20 more communication complaints per 100 practice-years (RR, 1.43; 95% CI, 1.15-1.77). After adjusting for the predictive ability of the clinical decision-making score in the traditional written examination, the patient-physician communication score in the clinical skills examination remained significantly predictive of retained complaints (likelihood ratio test, P < .001), with scores in the bottom quartile explaining an additional 9.2% (95% CI, 4.7%-13.1%) of complaints. CONCLUSION Scores achieved in patient-physician communication and clinical decision making on a national licensing examination predicted complaints to medical regulatory authorities.
It is the obligation of a profession to articulate the special meaning of competence in its field and to foster the good performance of its practitioners through education and discipline. External societal demands for increased accountability, and internal pressures for greater use of measurements of the processes and outcomes of clinical performance, are forcing the medical profession to reevaluate its view of competence and to change the way the profession "manages" the competence of its members. Traditionally, and predicated on the notion of "once in, good for life," medical education has focused on assuring the competence of trainees as they first enter independent professional life. In parallel, professional regulatory authorities have concentrated on apprehending the "false-positives" of the educational system. But viewed from a performance orientation, competence reflects situational relationships among doctors, their patients, and the systems in which they perform and, thus, is only partly dependent on the attributes of individual actors. This shift in thinking has major implications for the practice of medicine, particularly for the process of maintaining and improving performance. In jurisdictions throughout the world, recognition of the need for systematic and accountable ongoing education for practicing doctors is growing. This educational need should not be seen as a mark of weakness or failure but, rather, as the natural consequence of engagement in challenging practice. "Ars longa, vita breva." The profession must address the complex issues of education-in-practice with the same determination and creativity that it previously applied to education at entry to practice.
A fair amount of scrutiny has been given recently to the assessment of medical students' competence before they enter practice. In this issue of the Journal, Epstein provides a timely summary of advances in this arena.1 In contrast, little attention has been paid to the assessment of doctors who are already in practice. As Epstein points out, far from being a fixed attribute or trait, competence comprises multidimensional sets of behaviors that are dependent on both environmental and individual factors.2–4 As a result, the assessment of competence must go beyond the identification of who practitioners are, on the basis . . .
Introduction: The College of Physicians and Surgeons of Ontario developed an enhanced peer assessment (EPA), the goal of which was to provide participating physicians educational value by helping them identify specific learning needs and aligning the assessment process with the principles of continuing education and professional development. In this article, we examine the educational value of the EPA and whether physicians will change their practice as a result of the recommendations received during the assessment. Methods: A group of 41 randomly selected physicians (23 general or family practitioners, 7 obstetrician-gynecologists, and 11 general surgeons) agreed to participate in the EPA pilot. Nine experienced peer assessors were trained in the principles of knowledge translation and the use of practice resources (tool kits) and clinical practice guidelines. The EPA was evaluated through the use of a postassessment questionnaire and focus groups. Results: The physicians felt that the EPA was fair and educationally valuable. Most focus group participants indicated that they implemented recommendations made by the assessor and made changes to some aspect of their practice. The physicians' suggestions for improvement included expanding the assessment beyond the current medical record review and interview format (eg, to include multisource feedback), having assessments occur at regular intervals (eg, every 5 to 10 years), and improving the administrative process by which physicians apply for educational credit for EPA activities. Conclusions: The EPA pilot study has demonstrated that providing detailed individualized feedback and optimizing the one-to-one interaction between assessors and physicians is a promising method for changing physician behavior. The college has started the process of aligning all its peer assessments with the principles of continuing professional development outlined in the EPA model.
INTRODUCTION:The College of Physicians and Surgeons of Ontario, the regulatory authority for physicians in Ontario, Canada, conducts peer assessments of physicians' practices as part of a broad quality assurance program. Outcomes are summarized as a single score and there is no differentiation between performance in various aspects of care. In this study we test the hypothesis that physician performance is multidimensional and that dimensions can be defined in terms of physician-patient encounters.METHODS:Peer assessment data from 532 randomly selected family practitioners were analyzed using factor analysis to assess the dimensional structure of performance. Content validity was confirmed through consultation sessions with 130 physicians. Multiple-item measures were constructed for each dimension and reliability calculated. Analysis of variance determined the extent to which multiple-item measure scores would vary across peer assessment outcomes.RESULTS:Six performance dimensions were confirmed: acute care, chronic conditions, continuity of care and referrals, well care and health maintenance, psychosocial care, and patient records.DISCUSSION:Physician performance is multidimensional, including types of physician-patient encounters and variation across dimensions, as demonstrated by individual practice. A conceptual framework for multidimensional performance may inform the design of meaningful evaluation and educational recommendations to meet the individual performance of practicing physicians.
Background: Clinical skills examinations using standardized patients (SPs) are important in documenting the proficiency of trainees. "Standardized examinees" (SEs) are individuals trained to a specific level of performance; they can be used as internal controls in a high-stakes, clinical skills examination. Purpose: The purpose of this study was to determine whether SEs can be trained to portray a specified level of confidence and whether SPs' checklist scoring is affected by the personal manner of the examinee. Methods: Eight SEs were trained as "students" and trained to achieve a failing score on six cases in an National Board of Medical Examiners (NBME) Prototype Clinical Skills Examination. Four SEs were coached to be confident in manner, and 4 were coached to be insecure. Checklist scores were compared. Seven lay reviewers scored the SEs as confident or insecure on a behavioral assessment form. Results: SEs were not detected as simulations. There was no difference between the checklist scores of confident versus insecure SEs, but their manner was rated as significantly different on all scales in the behavioral assessment. Conclusions: SEs can be trained to a specified performance level and a desired level of confidence. In this small study, personal manner did not affect SPs' checklist scoring. The use of the SEs provides a mechanism to screen for bias in high-stakes SP examinations.
Information and knowledge are important clinical tools. But it9s judgment that really matters, says Daniel Klass
Practice inevitably narrows over time. Therefore, testing of established doctors requires that their assessment be tailored to a far narrower practice than is appropriate for testing of new doctors who have not yet differentiated. In this paper, we address the conceptual challenges of tailoring physician assessment to individual practice. Testing of established doctors needs to reflect that physicians specialise, often in idiosyncratic ways; otherwise, the testing will not be credible among established doctors and will not reflect the realities of their practice. Despite the importance of these goals, the conceptual and methodological challenges of creating tailored assessments remain daunting.
BACKGROUND:If continuing professional development is to work and be sensible, an understanding of clinical practice is needed, based on the daily experiences of doctors within the multiple factors that determine the nature and quality of practice. Moreover, there must be a way to link performance and assessment to ensure that ongoing learning and continuing competence are, in reality, connected. Current understanding of learning no longer holds that a doctor enters practice thoroughly trained with a lifetime's storehouse of knowledge. Rather a doctor's ongoing learning is a 'journey' across a practice lifetime, which involves the doctor as a person, interacting with their patients, other health professionals and the larger societal and community issues.OBJECTIVES:In this paper, we describe a model of learning and practice that proposes how change occurs, and how assessment links practice performance and learning. We describe how doctors define desired performance, compare actual with desired performance, define educational need and initiate educational action.METHOD:To illustrate the model, we describe how doctor performance varies over time for any one condition, and across conditions. We discuss how doctors perceive and respond to these variations in their performance. The model is also used to illustrate different formative and summative approaches to assessment, and to highlight the aspects of performance these can assess.CONCLUSIONS:We conclude by exploring the implications of this model for integrated medical services, highlighting the actions and directions that would be required of doctors, medical and professional organisations, universities and other continuing education providers, credentialling bodies and governments.
Clinical skills assessments have traditionally been scored via experts' ratings of examinee performance. However, this approach to scoring may be impractical in a large-scale context due to logistical and cost considerations as well as the increased probability of rater error. The purpose of this investigation was therefore to identify, using discriminant analysis, weighted score-based models that maximize the accuracy with which mastery level can be estimated for examinees taking a nationally administered standardized patient test. Additionally, the accuracy with which the resulting classification functions can be applied to predict mastery level for a cross-validation sample of examinees was also examined. Results suggest that it might be feasible to implement an automated scoring procedure in a cost-effective manner while still retaining the important facets of the decision-making process of expert raters. Cost-benefit, test development and psychometric implications of these results are important and discussed in the full paper.
The large-scale standardized patient (SP) test in this study assessed the clinical skills of fourth-year medical students in a series of clinical encounters targeting history taking, physical examination, communication, and interpersonal skills. Yearly large-scale field tests have been undertaken over the past seven years in preparation for national administration. The study reported here was conducted in 1998. Students are oriented to the test prior to completing up to 12 15-minute SP encounters (cases). Following each encounter, the SP records history elicited, counseling provided, or physical examination performed using an objective checklist developed by expert clinicians. The checklists may be thought of as a process measure, serving as a reflection of actual behaviors demonstrated by the candidate. Interpersonal skills are assessed using the Patient Perception Questionnaire (PPQ), a six-item instrument with a fivepoint Likert rating scale (uniform for every case). Following each encounter, students are given seven minutes to write a free-response Post-Encounter Note (PEN) (either a list of significant positive and negative history and physical findings or a written chart note documenting findings and counseling). The PEN is specifically tailored to reflect each case. There is no limit to the number of findings students may write. Patient management (diagnosis or therapeutic plans) and interpretation of diagnostic tests are not assessed in these PENs. The PENs potentially reflect a candidate’s ability to determine the most significant findings elicited from the encounter and to accurately record them. While numerous studies have examined the use of checklists with respect to fairness, security, and accuracy, there is limited research investigating the psychometric properties of PENs. Previous studies have examined appropriate methods for scoring the PEN. Soliciting global judgments from experts seems appealing because scores are derived from the expertise of practicing physicians, but global ratings can be unreliable unless the scoring task is highly structured and extensive standardized training is provided. From a national testing perspective, recruiting physicians to score the PENs for thousands of candidates may not be feasible. As a result, many researchers have favored the use of analytic keys to score PENs. A significant advantage of using such scoring keys is the fact that non-physicians can be trained to score the PENs with an accuracy level comparable to that of physicians. Research examining the usefulness of PENs with an SP test has suggested that these scores contribute valuable information to the assessment of clinical skills by providing unique information different from that derived from checklist scores. However, other research indicates that the chart audit scores should not replace the checklist entirely, since the information written by candidates in a simulated medical record may not provide a complete picture of events during an SP encounter. The inclusion of the PEN in an SP test is appealing. First, it is thought that PENs are relatively immune to within-site and crosssite effects. Also, they do not depend on the accurate recording of checklists by SPs. Additionally, threats to security are minimized because the PEN is a free-response instrument and does not reveal checklist content or other exam material. However, before the PEN can be used in large-scale testing, it is important to determine whether the PEN is a reflection of the checklist or whether the PEN contributes unique information about a student’s ability to synthesize and record medical information. The purpose of this study was, therefore, to investigate the relationship between entries recorded in the PEN and actions captured on the checklist. It is hoped that the results of this study will help determine how to best incorporate PEN information into a composite score.
Score validity is of central concern to any organization or school involved in high-stakes testing.1 Validation research entails clearly identifying the purpose for which test scores are to be used so that appropriate empirical evidence can be gathered to substantiate the intended score-based inferences.2 The validity of these score-based interpretations can be weakened by several test-related phenomena, including breaches to the security of the environment. The impacts of various forms of test security breaches need to be clearly addressed to determine the extent to which a priori knowledge of materials might provide an undue advantage to subgroups of examinees. This evidence also ensures that misinterpretation of scores is minimized on the part of the user. This task is especially crucial with performance-based tests such as standardized patient (SP) examinations, given the typically limited nature of case banks, the long exposure of items/cases, and the high costs associated with developing these types of assessments.3 Impact of Security Breaches on Test Performance The literature devoted to assessing the impacts of various forms of security breaches on the performances of students completing SP tests has reported mixed findings. Most investigations undertaken in this area have been aimed at determining whether mean scores on SP tests vary significantly when cases are administered throughout an extended interval, ranging from as little as several weeks4 to as much as an academic year.5 The authors of these studies have reported that mean station or case scores generally remain stable and that the reuse of identical cases, consequently, appears to have only a minimal impact on the scores of students taking the examination at different periods of time throughout the administration cycle.4,6,7,8 However, other research suggests that the reuse of identical cases can yield an increase in overall mean score, prompting a suggestion that the number of common cases be kept at a minimum across forms.5,9,10 Swartz, Colliver, Cohen, and Barrows11,12 examined whether collusion among students did affect overall SP test scores in a more systematized fashion by encouraging students who took the examination in the early stages of administration to share as much information as possible about the cases with students scheduled to be tested at a later date. The authors found little evidence that information-sharing among students affected performance. It is important to underscore that those studies restricted their view of a test security breach to various degrees of (presumed) information-sharing among examinees. It can be argued that complicity among students, although a common form of a test-security breach, is probably one of its most benign manifestations. This is especially likely with low- to moderate-stakes SP examinations, where students' motivation to engage in information sharing is low. In a high-stakes context (e.g., in licensure and certification testing), dishonest coaching organizations and examinees might employ a host of illicit means to obtain and disseminate actual test materials. A study undertaken by De Champlain et al.13 did model the impact of additional, more severe forms of test-security breaches on examinees' performances such as those that would result from students' having access to formal materials prior to taking the examination. The authors reported that disclosing test materials, whether it be directly to a subgroup of examinees or via a dishonest coaching course, led to significant checklist performance gains for a sample of United States medical graduates (USMGs). However, the impact of disclosure on interpersonal skills (IPS) scores was nil. Although informative, it is important to point out that these findings were based on a small and homogeneous sample with respect to examinees' medical education and clinical skill levels. As such, there is a need for this type of research to be replicated with a more varied sample of examinees, to obtain an estimate of disclosure effects that might generalize to a more heterogeneous population of medical students. The purpose of the present study was to model the impact of disclosing test materials on SP examination scores with a sample of international medical graduates. Furthermore, it is hoped that ensuing findings will provide a practical estimate of expected effect size within the context of this type of security breach and with this population. Method Examination. In this investigation, the SP test assessed the clinical (history taking, physical examination, communication) skills and IPS of physicians about to enter supervised practice. SPs are laypeople trained to portray one of a variety of clinical scenarios. Test candidates rotate through these scenarios (or cases) and encounter patients in a setting intended to reflect an ambulatory care clinic. Case-specific checklists are used to assess examinees' clinical skills. These checklists are composed of dichotomously scored items, each of which represents a single action that is expected to be done by the student. A percent-correct score, corresponding to the number of actions done by the student out of the total number of behaviors listed in a given checklist, is computed for all encounters. IPS are assessed with the Patient Perception Questionnaire (PPQ), a case-independent inventory that is composed of six five-point Likert scale items. A percent-correct PPQ score is also computed and reported to each student for all encounters. Both measurement instruments are completed by the SP following each 15-minute encounter with the student. The same ten cases (chosen from the available pool) were administered to all examinees. The cases were selected to reflect the majority of cells contained in the test blueprint with regard to both skill and content domains. Scoring Procedure. In this examination, two SPs were trained to portray each case. For any given case, the performing SP portrayed the actual clinical scenario with the examinee, whereas the monitoring SP observed the encounter as it proceeded on a video screen in a separate room. Each student's final percent-correct checklist score reflected the consensus reached by the performing and monitoring SPs as to what constituted the appropriate response to each item. Videotape review was instituted to arrive at a consensus if two or more discrepancies per checklist were noted in any given encounter. Of the 9,625 checklist item responses recorded (77 students × 125 checklist items across the ten cases), videotape review was necessary for 202 (2.10%). The PPQ percent-correct score was derived from the performing SP. Examinees. Seventy-seven international medical graduates (IMGs), recruited from the Los Angeles metropolitan area, participated in this study and were blinded to its purpose. All examinees were certified by the Educational Commission for Foreign Medical Graduates, i.e., they had successfully passed the following examinations: Step 1 and Step 2 of the United States Medical Licensing Examination and a test of English-language proficiency. The examinees were paid for their participation and randomly assigned to one of two testing conditions: control or security breach (SB). The testing environment for examinees assigned to the control condition (n = 32) was representative of a “normal” assessment situation (i.e., participants received routine prior information about the test but no materials from the examination). In the SB condition, we attempted to model a situation in which actual case materials were disclosed. Examinees in the SB condition (n = 45) were directly provided with the checklists for five of the ten cases to be seen (referred to as the exposed cases) as well as the PPQ, and were given one to two hours to review these materials prior to completing the test. Information pertaining to the five non-exposed cases was not disclosed to any of the examinees participating in this study. Cases included in the exposed and non-exposed sets were matched with respect to the main areas of this SP test's blueprint. Analyses. Two separate analyses of covariance (ANCOVAs) were undertaken to compare the performances of the two groups on the five exposed cases. For both models, the condition factor (control or SB) was treated as the independent variable. The mean percent-correct checklist score on the five non-exposed cases was treated as the covariate in the first ANCOVA, while the mean percent-correct checklist score on the five exposed cases was deemed to be the dependent variable (DV). In the second analysis, the mean percent-correct PPQ score on the five non-exposed was deemed to be the covariate, whereas the mean percent-correct PPQ score on the five exposed cases was treated as the DV. Results Mean scores and standard errors on the five exposed cases for examinees assigned to each of the two conditions, adjusted for initial differences in ability between groups, were as follows: For examinees assigned to the control condition, the adjusted mean percent-correct checklist score was 54.53 (SE = 1.48), and the adjusted mean percent-correct Patient Perception Questionnaire score was 60.87 (SE = 1.18). For examinees assigned to the security breach condition, the adjusted mean percent-correct checklist score was 59.95 (SE = 1.24), and the adjusted mean percent-correct Patient Perception Questionnaire score was 67.03 (SE = 0.99). A significant group main effect was obtained in the first ANCOVA, F(1,74) = 7.66, p =.0071. For the exposed cases, the SB group (adjusted M = 59.95%) significantly outperformed the control group (adjusted M = 54.53) on the checklist. Similarly, the mean PPQ score for examinees assigned to the SB condition (adjusted M = 67.03%) was significantly higher than the mean estimated for the control group (adjusted M = 60.87%), F(1,74) = 15.84, p =.0002. Conclusions Results obtained in the present study with a sample of international medical graduates mirror those reported in previous research with USMGs.13 Disclosing checklist items led to significant performance gains for the examinees assigned to the SB condition. The gain noted in this investigation (5.4%), was, however, slightly lower than that obtained with a sample of USMGs. This is probably attributable to the larger number of cases administered in the test form (ten as opposed to six in the past USMG study). Therefore, the challenge posed to the IMGs was slightly more daunting, as they had to sift through ten cases to identify the clinical scenarios for which they possessed disclosed materials and apply this information accordingly. Nonetheless, the gain noted would concretely translate itself into a 4.4-checklist-item disadvantage over five cases (slightly less than one item per case). This advantage might be inconsequential for most USMGs, who typically perform well above the cut-score on this type of examination.14 However, it could significantly affect decision consistency for IMGs, whose scores tend to cluster in the vicinity of the pass/fail standard in a larger proportion. The control and SB groups also did differ significantly with respect to their mean PPQ scores, a result that was not found with USMGs.13 Interestingly, the difference between the two groups (6.2%) was actually larger than the one resulting from disclosing checklist items. This could reflect a difference in interaction styles that is culturally based. Disclosing simple indicators of IPS (such as the Likert-scale items found on the PPQ) to SB group examinees yielded a mean score that was similar to that typically encountered with U.S. medical students. It is also worth noting that the type of case that was most susceptible to the effects of disclosure appears to be population-dependent. For U.S. medical students, prior research suggested that cases involving largely mechanical physical examination maneuvers were the easiest to memorize and consequently reflected the highest performance gains for those examinees with prior knowledge of materials. Divulging materials for cases that primarily require communication and IPS in the interaction with the patient proved to be the most beneficial for our sample of IMGs. Again, these findings appear be indicative of differences in the way our sample of IMGs interacted with the SPs. These results suggest that providing a clear description of the examination and its goal to all examinees prior to the administration (in some form of information bulletin, for example) is necessary to ensure a common understanding of expected behavior on the part of students. In summary, the results presented in this study provide further evidence that the secure handling of test materials is essential for all examinations, whether they be traditional in format or performance-based. Although the security breach modeled in this investigation was severe (half of the test materials were directly exposed to students), steps can nonetheless be undertaken to minimize the likelihood of materials being disclosed. This, in turn, might lessen the impact of a security breach should checklists or other pertinent information fall into the hands of dishonest individuals. One obvious strategy that should be adopted with all SP tests is to clearly lay out the flow of materials and restrict access solely to concerned staff so that these individuals can be held accountable for receipt and safekeeping of this information. Delivering the measurement instruments via a computer network also seems advisable, given the greater control that the latter medium can afford and the virtual elimination of a “paper trail.” The results of our study also point out the need to increase test development efforts to minimize the likelihood of a security breach. Increasing the pool of available cases enables a more frequent rotation of forms within and across test sites, thus limiting the exposure rate for any given set. Finally, the use of modeled or cloned cases also seems desirable to increase the size of the case pool and thwart those individuals who may have mechanically memorized cases and accompanying materials. Modeled cases are defined as those presenting a similar opening scenario but requiring a different work-up on the part of the student. Cloned cases, on the other hand, call for a similar set of actions on the part of the student but present different contexts. Although informative, our results need to be interpreted in light of several limitations. First, the sample size examined was small, and generalizations should be made with caution. Our sample was also composed of IMGs who were perhaps atypical of the corresponding population, given that they had successfully fulfilled several U.S. medical licensing requirements (passed the USMLE Step 1 and Step 2 and a test of English-language proficiency). Consequently, the effect sizes reported in this study should probably be viewed as lower-bound estimates of what to expect in an operational testing context. Replication of this research with different groups of both IMGs and USMGs seems advisable. This research might also permit us to test the hypothesis that lower-ability students might benefit more from gaining access to materials than would those who are more proficient. From a test-development perspective, pursuing research that focuses on the identification of characteristics that make a case more vulnerable to memorization would also be helpful. Finally, the findings reported in this study underscore the need to develop methods to detect breaches to the security of the testing environment. Research aimed at assessing the usefulness of “tagged” checklist items and other means should be pursued.15 Testing organizations and medical schools should always be vigilant in guarding themselves against dishonest examinees and organizations that may wish to compromise the secure nature of the testing environment. This investigation confirms past findings in that the psychometric properties of the SP examination described appear to be vulnerable to blatant disclosure of testing materials. It is hoped that the results presented in this article will foster future relevant research that will ultimately lead to the implementation of secure SP tests for licensure and other purposes.
Log in or Register Subscribe to journalSubscribe Get new issue alertsGet alerts Enter your Email address: Wolters Kluwer Health may email you for journal alerts and information, but is committed to maintaining your privacy and will not share your personal information without your express consent. For more information, please refer to our Privacy Policy. Subscribe to eTOC Secondary Logo Journal Logo All Articles Images Videos Podcasts Blogs Advanced Search Toggle navigation Subscribe Register Login Articles & Issues Current IssuePrevious IssuesPublished Ahead-of-Print Collections Editorials of Laura Weiss Roberts, MD, MAAM Last PageCOVID-19 and Medical EducationAddressing Race and Racism in Medical EducationeBooksView All For Authors Submit a ManuscriptInformation for AuthorsLanguage Editing ServicesAuthor Permissions Journal Info About the JournalAbout the AAMCJournal MastheadSubmit a ManuscriptAdvertising InformationSubscription ServicesReprints and Back IssuesClassified AdsRights and PermissionsFor ReviewersFor MediaFor Trainees All Articles Images Videos Podcasts Blogs Advanced Search
Log in or Register Subscribe to journalSubscribe Get new issue alertsGet alerts Enter your Email address: Wolters Kluwer Health may email you for journal alerts and information, but is committed to maintaining your privacy and will not share your personal information without your express consent. For more information, please refer to our Privacy Policy. Subscribe to eTOC Secondary Logo Journal Logo All Articles Images Videos Podcasts Blogs Advanced Search Toggle navigation Subscribe Register Login Articles & Issues Current IssuePrevious IssuesPublished Ahead-of-Print Collections Editorials of Laura Weiss Roberts, MD, MAAM Last PageCOVID-19 and Medical EducationAddressing Race and Racism in Medical EducationeBooksView All For Authors Submit a ManuscriptInformation for AuthorsLanguage Editing ServicesAuthor Permissions Journal Info About the JournalAbout the AAMCJournal MastheadSubmit a ManuscriptAdvertising InformationSubscription ServicesReprints and Back IssuesClassified AdsRights and PermissionsFor ReviewersFor MediaFor Trainees All Articles Images Videos Podcasts Blogs Advanced Search
Objectives The purpose of the study was to explore foreign medical graduates' (FMGs) performance on a clinical skills (SPX) examination. The National Board of Medical Examiners (NBME) is in the process of developing an SPX for potential use in the United States Medical Licensing Examination (USMLE). The Educational Commission for Foreign Medical Graduates (ECFMG) is developing the Clinical Skills Assessment (CSA) as an additional requirement for FMGs who wish to be certified by ECFMG.Design Thirty-three FMGs and 151 United States medical students (USMSs) took the SPX during the winter of 1996 as part of the ongoing pilot studies conducted by the NBME. Four clinical skill areas were assessed: history-taking, physical examination, communication and interpersonal skills. The examination used in this research consisted of 12 cases. The examination utilizes standardized patients (SPs) who are trained to document examinee behaviours and evaluate the communication component of the test. The SPs were also trained to evaluate the English proficiency of the candidates. Candidates were also administered the Test of Spoken English developed by the Educational Testing Services (ETS).Setting The examination was conducted in one medical school which served as an SPX centre for NBME pilot studies.Subjects Thirty-three foreign medical students and 151 US medical students.Results The indications were that the majority of candidates in both groups felt the examination was moderately fair but 78% of FMGs felt moderately pressed for time, vs. 80% of the USMSs who did not feel pressed for time. Reliabilities obtained for the various SPX components were somewhat higher for the FMGs reflecting the heterogeneity of this group.Conclusions The NBME-ECFMG collaborative study yielded important information regarding the NBME SPX prototype as a performance measure for FMGs.