Background: The Ministry of Health of the Republic of Panama is currently developing a national examination system that will be used to license graduates to practice medicine in that country, as well as to undertake postgraduate medical training. As part of these efforts, a preliminary project was undertaken between the National Board of Medical Examiners (NBME) and the Faculty of Medicine of the University of Panama to develop a Residency Selection Process Examination (RSPE).Purpose: The purpose of this study was to assess the reliability and validity of RSPE scores for a sample of candidates who wished to obtain a residency slot in Panama.Methods: The RSPE, composed of 200 basic and clinical sciences multiple-choice items, was administered to 261 residency applicants at the University of Panama.Results: The reliability estimate computed was comparable with that reported with other high-stakes examinations (Cronbach's alpha = 0.89). Also, a Rasch examinee proficiency item difficulty plot showed that the RSPE was well targeted to the proficiency levels of candidates. Finally, a moderate correlation was noted between local grade point averages and RSPE scores for University of Panama students (r = 0.38).Conclusions: Findings suggest that it is possible to translate and adapt test materials for use in other contexts. Copyright (C) 2005 by Lawrence Erlbaum Associates, Inc.
PURPOSE:The purpose of this study was to assess whether the interaction of examinee and standardized patient (SP) ethnicity has an impact on data gathering and written communication scores in a large-scale clinical skills assessment used for certification purposes.METHOD:The sample that was the focus of the present investigation was selected from the population of 9,551 international medical graduates (IMGs) who completed the Educational Commission for Foreign Medical Graduates' Clinical Skills Assessment between May 1, 2002, and May 31, 2003. Analyses of covariance were undertaken separately for four cases, adjusting for initial mean differences between candidate groups and controlling for stringency levels of SPs. Over 2,800 SP-IMG encounters were analyzed, ranging from 597 (Case 2) to 915 (Case 3).RESULTS:None of the SP ethnicity/examinee ethnicity interactions were statistically significant.CONCLUSIONS:Findings suggest that there is little advantage to be gained by encountering a SP of similar ethnic makeup. These results are discussed in light of past research undertaken to assess fairness issues with clinical skills examinations.
Medical training is undergoing extensive revision in France. A nationwide comprehensive clinical competency examination will be administered for the first time in 2004, relying exclusively on essay-questions. Unfortunately, these questions have psychometric shortcomings, particularly their typically low reliability. High score reliability is mandatory in a high-stakes context. The National Board of Medical Examiners-designed multiple choice-questions (MCQ) are well adapted to assess clinical competency with a high reliability score. The purpose of this study was to test the hypothesis that French medical students could take an American-designed and French-adapted comprehensive clinical knowledge examination with this MCQ format. Two hundred and eighty five French students, from four Medical Schools across France, took an examination composed of 200 MCQs under standardized conditions. Their scores were compared with those of American students. This examination was found assess French students' clinical knowledge with a high level of reliability. French students' scores were slightly lower than those of American students, mostly due to a lack of familiarity with this particular item format, and a lower motivational level. Another study is being designed, with a larger group, to address some of the shortcomings of the initial study. If these preliminary results are replicated, the MCQ format might be a more defendable and sensible alternative to the proposed essay questions.
PURPOSE:The French government, as part of medical education reforms, has affirmed that an examination program for national residency selection will be implemented by 2004. The purpose of this study was to develop a French multiple-choice (MC) examination using the National Board of Medical Examiners' (NBME) expertise and materials.METHOD:The Evaluation Standardisée du Second Cycle (ESSC), a four-hour clinical sciences examination, was administered in January 2002 to 285 medical students at four university test sites in France. The ESSC had 200 translated and adapted MC items selected from the Comprehensive Clinical Sciences Examination (CCSE), an NBME subject test.RESULTS:Less than 10% of the ESSC items were rejected as inappropriate to French practice. Also, the distributions of ESSC item characteristics were similar to those reported with the CCSE. The ESSC also appeared to be very well targeted to examinees' proficiencies and yielded a reliability coefficient of.91. However, because of a higher word count, the ESSC did show evidence of speededness. Regarding overall performance, the mean proficiency estimate for French examinees was about 0.4 SD below that of a CCSE population.CONCLUSIONS:This study provides strong evidence for the usefulness of the model adopted in this first collaborative effort between the NBME and a consortium of French medical schools. Overall, the performance of French students was comparable to that of CCSE students, which was encouraging given the differences in motivation and the speeded nature of the French test. A second phase with the participation of larger numbers of French medical schools and students is being planned.
Recently, standardized patient assessments and objective structured clinical examinations have been used for high-stakes certification and licensure decisions. In these testing situations, it is important that the assessments are standardized, the scores are accurate and reliable, and the resulting decisions regarding competence are equitable and defensible. For the decisions to be valid, justifiable standards, or cut-scores, must be set. Unfortunately, unlike the body of research specifically dedicated to multiple-choice examinations, relatively little research has been conducted on standard-setting methods appropriate for use with performance-based assessments. The purpose of this article is to provide the reader with some guidance on how to set defensible standards on performance assessments, especially those that utilize standardized patients in simulated medical encounters. Various methods are discussed and contrasted, highlighting the relevant strengths and weaknesses. In addition, based on the prevailing literature and research, ideas for future studies and potential augmentations to current performance-based standard setting protocols are advanced.
Standardized patient (SP) examinations are widely used by medical schools and testing and certification organizations to evaluate clinical and interpersonal skills not readily measurable with written multiple-choice examinations. Albeit valuable, SP examinations bear limitations, mainly decreased reliability of examinee scores attributable to the limited number of cases seen by the student and variations in SP recording, rating, and portrayal accuracy. Regardless of whether SP exams are being used by medical schools for teaching and learning purposes or by medical testing organizations for licensure or certification, it is critical that scores accurately reflect the appropriate clinical skill levels of the examinees. Threats to reliability may increase when exams are administered on a large scale and it becomes necessary to train multiple SPs to portray the same case across multiple testing sites. Much research has focused on quantifying sources of variability in SP exams because any type of unwanted variation could have a deleterious impact on pass/fail decisions. The conclusions of these studies are not easily discerned. Background The majority of initial studies indicated that use of multiple SPs did not cause large discrepancies in total test scores when examinees were randomly assigned to SPs.1 Swanson and Norcini2 found that raters nested within a case explained only 1% to 2% of the observed score variance, and De Champlain et al.3 found that multiple SPs can similarly assess examinees' performances, leading to identical mastery-level decisions for nearly all students tested. However, research has indicated that using multiple SPs may introduce enough error to be consequential at the case level. Raters' marks for the same examinee on a given case have differed, with the level of agreement seemingly influenced by the nature of the case.2 Discrepancies among raters have been large enough to produce statistically significant differences in pass/fail rates of individual cases.4 More recent studies have shown differences in intra-rater reliability,5 and the effects of rater discrepancies on pass/fail decisions are large enough to warrant concern.6 Although testing organizations that accommodate large volumes of examinees are concerned with the effects of administering forms across multiple sites, fewer studies have examined these effects. While some have found little or no differences in scores of candidates taking the same test administered at different sites,5,7 others have indicated that candidates' scores can be influenced by the site at which they take an exam, especially when training is minimal.8 It has been suggested that raters' characteristics may be a more important determinant of rater agreement than are differences in test site or trainers.8 Unfortunately, few studies have examined the relationship of rater characteristics to score variability. It is clear that more research is needed to assess the sources of variation present when multiple SPs portray and score identical cases across multiple testing sites. The most widely used approach to quantifying sources of score variation in SP examinations has been generalizability analysis. However, in large-scale testing when SPs are nested in cases and in test sites, this approach precludes estimation of many potential sources of variation. Furthermore, it does not provide insight into which raters' characteristics are responsible for score fluctuations. It seems desirable to use an approach that could quantify sources of score variation while concurrently examining rater characteristics. The purpose of the current investigation was to use multilevel modeling to quantify the amounts of variation in individual cases that are attributable to SP and site differences and to examine how SPs' gender and experience may explain such variation. Multilevel modeling allows for score variation to be partitioned across several levels (e.g., student, rater, and site) while simultaneously determining the relative weights and significance of predictors at each level (e.g., individual SP and/or site characteristics). This comprehensive approach to understanding sources of score variability in SP encounters will allow us to evaluate training procedures and measurement instruments and to make more informed decisions regarding scoring, calibrating, and equating procedures that will ultimately enhance decision consistency rates. Method SP examination and instrumentation. In the present study, the SP test assessed the clinical skills (history taking, physical examination, and communication) of physicians about to enter supervised practice. Examinees proceeded through (six to ten) cases and encountered patients in a setting intended to reflect an ambulatory clinic. Subsequent to each 15-minute encounter, the SP performing the case recorded the student's performance using a checklist and the patient-perception questionnaire. The checklist had been tailored specifically to the SP's complaint by a group of subject-matter experts and contained ten to 25 dichotomously scored items targeting behaviors deemed critical for success in the encounter. Unlike the checklist, the case-invariant patient-perception questionnaire was composed of seven five-point Likert-type items that allowed the SP to rate the student's interpersonal skills (IPS). The checklist and IPS scores are reported on a percentage-correct scale. The reliability (Cronbach's alpha) of these scores is typically lower than that of traditional multiple-choice examinations due to the limited number of items. Sample and cases. SP cases had been administered with variable frequency across 21 testing sites in the United States in 2000 as part of a large-scale research project. For the purposes of this investigation, four cases were selected from those administered: one biomedical, one grave illness, and two routine counseling. Two routine counseling cases were selected because the second was performed by both men and women SPs. All other cases in the bank were portrayed by only men or only women. The numbers of students who saw the individual cases ranged from 357 to 565. Students who had not already taken Step 2 of the U.S. Medical Licensing Examination (USMLE) were excluded from the sample when the checklist scores were modeled. The final sample for each case reflected students testing at one of eight or nine test sites. Hierarchical linear models (HLMs). The multilevel modeling software package9 was used to: (1) estimate the proportion of score variation between SPs portraying the same case at a given site, (2) estimate the proportion of score variation between training sites, (3) assess the relationship between the skill scores and SPs' characteristics, and (4) quantify the proportion of SP variation explained when SPs' characteristics were added into the model. Because research has suggested that IPS scores are more prone to SP variation, checklist and IPS scores were modeled separately for each of the four cases. A total of seven models were examined (three checklist and four IPS). Checklist scores for the routine counseling case portrayed by men and women were not modeled due to lack of USMLE Step 2 data. The checklist and IPS scores were modeled as a function of the number of encounters performed by the SP over the course of the testing period and gender (solely for the case portrayed by both men and women). The number of encounters served as a proxy for test-administration experience. USMLE Step 2 served as a covariate for checklist models because it was moderately correlated with SP checklist scores. However, because its relationship with IPS scores was weak, these scores were run as intercept-only models with no covariate. Analyses. Prior to entering predictors into the three-level model, a one-way analysis of variance (ANOVA) with random effects was run. This is typically the first step in multilevel modeling and allows for score variation to be partitioned and quantified across various levels (student, site, or SP). It also serves as a base with which to estimate the amount of variation explained as predictors are entered into the model. Subsequent to the ANOVA, predictors were entered in a blockentry fashion. First, USMLE Step 2 was entered at the student level (level 1) to adjust for ability (for checklist scores only). Next, the number of SP encounters was entered at the SP level (level 2) to help explain variation between SPs at a given site. Gender was also added as an SP-level predictor for the routine counseling case performed by both men and women SPs (women SPs are coded 1). Random effects that were not statistically significant were removed from the model to reduce the number of parameters. Scores with little variation among sites and SPs were not modeled with predictors. Results Results from the seven ANOVAs indicated that checklist scores were more robust to variation among SPs within a given site (2–3%), whereas IPS scores varied substantially across SPs in a given site (26–49%). Irrespective of the nature of the case, less than 12% of the variation checklist scores was due to differences among testing sites. There was no variation in IPS scores across testing sites with the exception of one routine counseling case where 16% of the score variation was attributable to site differences. Table 1 provides the ANOVA variance proportions, the regression coefficients from the final model containing predictors, and the proportions of variation explained as SP predictors were entered into the model.TABLE 1: Variance Proportions Estimated in the One-way ANOVA with Random-effects Model; Effect Sizes and Variation Explained for the Final Model Including SP CharacteristicsAs shown in Table 1, the USMLE Step 2 score was significantly related to the checklist scores and therefore appeared to be an appropriate choice as a covariate. Conversely, the SPs' characteristics were not significantly related to skill score. Adding test experience into the model for the first routine counseling case explained about half of the variation among SPs (56%), even though this predictor was not statistically significantly related to IPS score. Adding gender and test experience into the model for the second routine counseling case explained only 8% of the variation among SPs. In the other models, the inclusion of SPs' characteristics did not explain any variation from the ANOVA. For all models, the likelihood ratio chi-square statistic for model fit (not shown) indicated that including SP characteristics was not justified, that is, it did not make a substantial contribution to model fit. Discussion This preliminary investigation has provided important information for the consortia that use multiple SPs to portray a given case across multiple testing sites and within a given site. Results of previous generalizability analyses have suggested that most of the variation in overall examination scores is due to case specificity. However, this study has demonstrated that differences in SPs or training sites can be pronounced in individual cases. Rather than assuming that these effects will "wash out," it appears wise to take a proactive approach to detect and remedy problems with individual cases prior to their live administration in high-stakes examinations. In addition to partitioning score variation into various levels, multilevel models inform us as to the causes of variation inherent to clinical skills encounters. This information is critical to large-scale testing organizations that are considering selection of equating, calibrating, and scoring rubrics. If the sources of score variation are known, then steps can be taken a priori to reduce this variation, and scoring and calibrating routines may be developed to compensate for variability in the measures. Although a major advantage of using multilevel modeling is to explain variation at various levels, the benefit of adding predictors into these models diminishes if there is not a substantial amount of variation to model (generally about 10%). Therefore, based on the negligible amount of SP variability found in checklist scores examined in this study, it is doubtful that test administration experience or any predictor, for that matter, could have explained this variation. The results suggest that it might have been preferable to exclude SP characteristics when modeling the checklist scores and add test site characteristics to the model. The finding that IPS scores were more influenced by variations among SPs than checklist scores was not surprising, given the results of previous research. It was surprising, however, that test experience had little impact in explaining variation in IPS scores. Perhaps the training SPs receive prior to testing with students is extensive enough to make additional experience inconsequential. The aggregated nature of this variable may have also contributed to its lack of statistical significance. Regardless of the cause of the variation, the large proportion of IPS score variation associated with SP differences compels us to consider how IPS instruments are developed and to do so in a way that limits rater subjectivity. For example, asking the rater to record how he or she "felt" may not be adequate. Rather, videotaped encounters that establish base-lines of interpersonal skill levels may be more effective. One finding of the current study not anticipated based on past research was the systematic stringency or leniency across training sites for checklist scores of certain cases. After adjusting for ability, several checklist scores varied by as much as 10% as a function of the testing site. The inter-site variability in checklist scores could have resulted from differences in training across sites or possibly because guidelines to scoring checklists were not clear (i.e., what a student must do to receive credit for completing a given behavior). These findings might also be due to the various levels of adherence to protocols on the part of different trainers or the inadequacy of Step 2 as a covariate. This preliminary study has several limitations that must be addressed. The results from IPS models must be interpreted cautiously because the model did not contain a covariate. It is likely that these proportions would have been much smaller had the IPS scores been adjusted for ability. However, in the three-level HLM, the proportion of variation among SPs is associated with those performing at a given site. Under the assumption that students at a given school (test site) are of the same relative ability and have been trained in a similar manner, we would not anticipate the proportion of variation to be so large. A second limitation was that the data available were very limited. An extensive data set containing student-, SP-, and site-related information would have been desirable in order to explain SP variation in the IPS scores and site variation in the checklist scores. It is critical to note that running HLM for individual cases allows for examination of rater and site characteristics but precludes estimation of SP or site variation at the exam level (as estimated using generalizability analysis). Although HLM is very flexible, the models described in this study are more desirable for pre-testing cases than estimating variance inherent in the entire examination. The most comprehensive approach would be to use generalizability analysis to quantify score variation across the entire examination first and then use HLM to troubleshoot potential causes for variation across SPs or sites. Alternatively, if rater characteristics are not of interest, models can be analyzed across multiple cases at the examination level. This study should be viewed as an important first step in using multilevel modeling to explore and explain the variability of SP examinations. Some results from this investigation confirmed those found in previous research, whereas others offered different perspectives. Given the high stakes involved in a licensing examination, it appears unwise to adopt the attitude that stringency and leniency of SPs wash out over the course of the multi-case examination. Given the opportunity, it is advisable to take actions prior to live administration to reduce the variation in individual cases. Careful consideration of SP and site characteristics should be captured and analyzed statistically so steps can be taken to develop and implement fair and reliable examinations.
Diagnostic assessment of practicing physicians is modest in scale, and normative data for practicing physicians are virtually nonexistent in the U.S. It is the responsibility of state licensing boards to ensure the competency of their licensed physicians. Increasingly, state licensing boards are responding to the concerns of patients, communities, and hospitals regarding the continued clinical competence of physicians. 1 Precipitating events may involve deficiencies in prescription of pharmacotherapy, patient management, and interpersonal skills. The Assessment Center Program (ACP) is a joint activity of the National Board of Medical Examiners (NBME) and the Federation of State Medical Boards (FSMB). It is a program of standardized tests intended to assist state medical boards (SMBs) in their diagnostic assessment of physicians whose clinical competence has been questioned. This type of personalized assessment of physicians has increased since the 1980s, with a sizeable proportion of the work being done in Canada. 2,3 The target candidate of the ACP is a physician referred by a state medical board because of concerns about his or her clinical competence. In a study of doctors in California 4 it was found that disciplined doctors were less likely to be board certified than were a matched group of non-disciplined physicians. Age has also been reported to be a predictor of physicians' being identified for deficiencies in clinical competence. 5 Under a special arrangement with Colorado Personalized Education for Physicians (CPEP), referred doctors are tested at the CPEP facility, located near Denver, Colorado. The full CPEP assessment for generalist doctors requires three days of multi-method testing. For the last three and a half years, ACP has been administering computer-based case simulations (CCSs) to referred doctors at CPEP and separately to a comparison group of physicians. 6–8 The CCSs are a Windows-based simulation of the health care environment designed to assess a physician's skills in patient management. A presentation screen displayed at the start of a case describes the patient, the initial history and vitals, and the location of the clinical encounter (emergency room or outpatient office). From this point, the case is prompt-free and therefore examinee-driven. As one of its goals, the ACP has been exploring the utility of CCSs in the diagnostic assessment of referred doctors. It is anticipated that the case means for the comparison group will be higher than will be those of the referred group. In addition, we have been examining the relationships between selected demographic variables and CCS performances of referred physicians. Method Referred group. At the time of writing, 42 referred group physicians had taken the Windows-based CCS exam since August 1999 as part of the CPEP assessment program. Physicians are referred for a number of reasons, but most referrals involve some concern regarding clinical competence. The mean age of this referred group was 53, and the range was 30 to 76 years. There were eight women in the referred group. All referred physicians except two reported internal medicine, family practice, or general practice as their specialties. Two physicians were residents and therefore not board certified. Of the 40 remaining referred licensed physicians, 23 (58%) were not board certified. Comparison group. In the fall of 1998, a convenience sample of 32 Philadelphia-area, board-certified physicians completed the same eight-case CCS exam. All testing took place at the NBME offices, and the physicians were compensated. Score reports were not provided. Additionally, between July 1999 and October 2000, 49 staff physicians and residents from the emergency medicine, internal medicine, and general surgery departments at the Naval Medical Center in Portsmouth, Virginia (NMCP), volunteered to take the same eight-case CCS exam. Compensation was not provided, and individual score reports were not provided. The total number of physicians in the comparison group was 81. The mean age of the doctors in the comparison group was 38 years, and the range was 27 to 63 years. The comparison group was approximately 15 years younger, on average, than was the referred group. Specialties consisted of internal medicine (48), family practice (20), emergency medicine (8), and general surgery (5). All physicians in the comparison group were board certified with the exception of 14 residents from NMCP. The assessment. In CCSs, the physician is responsible for addressing the patient's concern, diagnosing the problem, and providing treatment. The physician is instructed to assume complete responsibility as the simulated patient's primary care physician. The physician types all orders on a simulated order sheet. The patient's condition evolves as a function of the physician's management choices. At the end of each case the physician is prompted to enter the primary diagnosis. Each action is recorded in a text file called a transaction list, which becomes the source for scoring. All actions are classified as one of three levels of benefit (i.e., most, more, least) or as one of three levels of risk. Raw scores are computed and then weighted to reflect expert ratings. 9,10 Scores are reported to CPEP on a scale of 1 to 9, with higher scores indicating better patient management skills. CCSs have recently been incorporated within the United States Medical Licensing Examination (USMLE) Step 3. Both referred and comparison-group-physicians took the same eight-case CCS exam, which was designed to reflect generalist practice. Specifically, the exam contained six medicine cases, one pediatric case, and one ob—gyn case. We are unable to describe further the content of these cases because of security concerns. All physicians took three orientation cases prior to completing the scored cases. A proctor remained in the testing room to answer technical questions regarding navigating the simulation. Answers to questions related to medical content were not provided. For referred doctors who requested assistance (because of unfamiliarity with computers), the CPEP proctor served as a scribe at the keyboard. Cases were presented in the same order for all examinees. Referred physicians completed a post—CCS interview with a medical director trained in the use of CCS. The medical director conducted a clinical interview to evaluate the decision-making process the referred physician used to manage the simulated patients. Results Case scores. Table 1 displays the descriptive statistics for referred and comparison groups. The case means of the referred group were lower than those of the comparison group for all eight cases. The comparison group's case means ranged from 4.40 for the pediatric case to 6.54 for a medicine case (case 4). The referred group's case means ranged from 3.41, also for the pediatric case, to 5.20 for a medicine case (case 3). To test whether case means differed across the groups, t-tests were computed for independent groups. Mean differences were statistically significant at the Bonferroni-adjusted alpha value of .006 (.05/8 cases) for five cases: cases 1, 2, 4, 6, and 8. All are “medicine-focused” cases except case 2, the one pediatric case. The largest mean difference (2.3) was seen for case 4. The mean difference for the grand mean (across all eight test cases) was also significant (t (121) = −5.970, p = .000).TABLE 1: CCS Performances of Physicians in the Referred and Comparison Groups, Cases 1–8Timing. Except for the pediatric case, the referred group took more time (real time spent on the case), on average, to complete each test case than did the comparison group. Independent t-tests were computed to test whether the case mean times were significantly different from those of the comparison group. Six mean differences were significant at a Bonferroni-adjusted alpha value of .006, specifically those for cases 1, 3, 4, 5, 7, and 8. It is possible that the referred group spent more time on the cases because of the context of their performance; that is, the referred physician was aware that his or her CCS performance was important to the overall evaluation by CPEP. The Pearson correlation between the mean time in minutes and the CCS grand mean for the referred group was significant at −.33 (p = .032). The Pearson correlation between the mean time and the CCS grand mean for the comparison group was .11. CCS scores and demographics of referred doctors. It has been reported that doctors who are older and not certified in a specialty are at increased risk of being identified as having deficiencies in clinical competence. 4,5 We examined this relationship by comparing the performance of the board-certified referred doctors (n = 17) on the CCSs with the scores of the non-certified referred doctors (n = 25). The mean age of the certified referred doctors was 48.3 years and that of the non-certified referred doctors was 55.8 years. The Pearson correlation between age and mean CCS score was significant at −.40 (p = .011). Older referred physicians performed less well on CCSs than did younger referred physicians. This is in line with the fact that many of the older referred physicians are not board certified. In fact, the correlation between age and certification (yes/no) was .37. The CCS mean for board-certified referred physicians was 5.02, and the CCS mean for non—board-certified referred physicians was 4.17. An independent-samples t-test was computed to test whether the mean difference in CCS mean scores (over all eight cases) between board-certified and non—board-certified physicians was significant. The mean difference was significant (t (40) = −2.96, p = .005), with board-certified physicians outperforming their counter-parts. Because age and certification are correlated, we then performed a univariate analysis of variance with certification as the fixed factor, mean CCS score as the dependent variable, and age as the covariate. Certification status was significant in this model. Discussion We have reported on the performances of referred and comparison-group physicians on an eight-case computer-based case simulation (CCS) examination of patient management skills. The purpose of this study was to examine the usefulness of CCSs in assessing patient management skills of referred physicians. Descriptive statistics of CCS performance show that the referred physicians performed less well than did the comparison-group physicians on all eight cases. This was the anticipated outcome, given that the referred physicians' competence had been called into question. The referred physicians scored significantly lower than did the comparison-group physicians on five cases. There was a significant difference in mean CCS scores (over all cases) between board-certified and non—board-certified referred physicians. Other researchers 4 have observed this relationship. Overall, physicians in the referred group took more time than comparison-group physicians took to complete each case. This may be related to the referred group's being older than the comparison group and less comfortable with Windows-based computer programs. In fact, CPEP reports that many referred physicians stated they had no computer experience and requested the proctor to enter orders during the examination. The results of this study provide evidence that CCS is useful in the assessment of the patient management skills of referred physicians. The results show a pattern of comparison-group physicians' performing better, on average, then referred group physicians do. Board-certified physicians perform better on CCSs than do non—board-certified physicians, and there was a significant, negative correlation between mean CCS score and age of the referred physician. It is the goal of this program to continue to collect comparison-group data on CCSs. Currently, the comparison group consists of two combined convenience samples. We recognize that this convenience sample of 81 comparison physicians is not adequate for generalized statements of performance. However, for those organizations involved in assessing physicians' performances, this study provides formative data regarding the patient-management performances of competent, practicing physicians as assessed by CCSs. A database of such information will provide a standard with which to compare referred physicians.
The large-scale standardized patient (SP) test in this study assessed the clinical skills of fourth-year medical students in a series of clinical encounters targeting history taking, physical examination, communication, and interpersonal skills. Yearly large-scale field tests have been undertaken over the past seven years in preparation for national administration. The study reported here was conducted in 1998. Students are oriented to the test prior to completing up to 12 15-minute SP encounters (cases). Following each encounter, the SP records history elicited, counseling provided, or physical examination performed using an objective checklist developed by expert clinicians. The checklists may be thought of as a process measure, serving as a reflection of actual behaviors demonstrated by the candidate. Interpersonal skills are assessed using the Patient Perception Questionnaire (PPQ), a six-item instrument with a fivepoint Likert rating scale (uniform for every case). Following each encounter, students are given seven minutes to write a free-response Post-Encounter Note (PEN) (either a list of significant positive and negative history and physical findings or a written chart note documenting findings and counseling). The PEN is specifically tailored to reflect each case. There is no limit to the number of findings students may write. Patient management (diagnosis or therapeutic plans) and interpretation of diagnostic tests are not assessed in these PENs. The PENs potentially reflect a candidate’s ability to determine the most significant findings elicited from the encounter and to accurately record them. While numerous studies have examined the use of checklists with respect to fairness, security, and accuracy, there is limited research investigating the psychometric properties of PENs. Previous studies have examined appropriate methods for scoring the PEN. Soliciting global judgments from experts seems appealing because scores are derived from the expertise of practicing physicians, but global ratings can be unreliable unless the scoring task is highly structured and extensive standardized training is provided. From a national testing perspective, recruiting physicians to score the PENs for thousands of candidates may not be feasible. As a result, many researchers have favored the use of analytic keys to score PENs. A significant advantage of using such scoring keys is the fact that non-physicians can be trained to score the PENs with an accuracy level comparable to that of physicians. Research examining the usefulness of PENs with an SP test has suggested that these scores contribute valuable information to the assessment of clinical skills by providing unique information different from that derived from checklist scores. However, other research indicates that the chart audit scores should not replace the checklist entirely, since the information written by candidates in a simulated medical record may not provide a complete picture of events during an SP encounter. The inclusion of the PEN in an SP test is appealing. First, it is thought that PENs are relatively immune to within-site and crosssite effects. Also, they do not depend on the accurate recording of checklists by SPs. Additionally, threats to security are minimized because the PEN is a free-response instrument and does not reveal checklist content or other exam material. However, before the PEN can be used in large-scale testing, it is important to determine whether the PEN is a reflection of the checklist or whether the PEN contributes unique information about a student’s ability to synthesize and record medical information. The purpose of this study was, therefore, to investigate the relationship between entries recorded in the PEN and actions captured on the checklist. It is hoped that the results of this study will help determine how to best incorporate PEN information into a composite score.
Score validity is of central concern to any organization or school involved in high-stakes testing.1 Validation research entails clearly identifying the purpose for which test scores are to be used so that appropriate empirical evidence can be gathered to substantiate the intended score-based inferences.2 The validity of these score-based interpretations can be weakened by several test-related phenomena, including breaches to the security of the environment. The impacts of various forms of test security breaches need to be clearly addressed to determine the extent to which a priori knowledge of materials might provide an undue advantage to subgroups of examinees. This evidence also ensures that misinterpretation of scores is minimized on the part of the user. This task is especially crucial with performance-based tests such as standardized patient (SP) examinations, given the typically limited nature of case banks, the long exposure of items/cases, and the high costs associated with developing these types of assessments.3 Impact of Security Breaches on Test Performance The literature devoted to assessing the impacts of various forms of security breaches on the performances of students completing SP tests has reported mixed findings. Most investigations undertaken in this area have been aimed at determining whether mean scores on SP tests vary significantly when cases are administered throughout an extended interval, ranging from as little as several weeks4 to as much as an academic year.5 The authors of these studies have reported that mean station or case scores generally remain stable and that the reuse of identical cases, consequently, appears to have only a minimal impact on the scores of students taking the examination at different periods of time throughout the administration cycle.4,6,7,8 However, other research suggests that the reuse of identical cases can yield an increase in overall mean score, prompting a suggestion that the number of common cases be kept at a minimum across forms.5,9,10 Swartz, Colliver, Cohen, and Barrows11,12 examined whether collusion among students did affect overall SP test scores in a more systematized fashion by encouraging students who took the examination in the early stages of administration to share as much information as possible about the cases with students scheduled to be tested at a later date. The authors found little evidence that information-sharing among students affected performance. It is important to underscore that those studies restricted their view of a test security breach to various degrees of (presumed) information-sharing among examinees. It can be argued that complicity among students, although a common form of a test-security breach, is probably one of its most benign manifestations. This is especially likely with low- to moderate-stakes SP examinations, where students' motivation to engage in information sharing is low. In a high-stakes context (e.g., in licensure and certification testing), dishonest coaching organizations and examinees might employ a host of illicit means to obtain and disseminate actual test materials. A study undertaken by De Champlain et al.13 did model the impact of additional, more severe forms of test-security breaches on examinees' performances such as those that would result from students' having access to formal materials prior to taking the examination. The authors reported that disclosing test materials, whether it be directly to a subgroup of examinees or via a dishonest coaching course, led to significant checklist performance gains for a sample of United States medical graduates (USMGs). However, the impact of disclosure on interpersonal skills (IPS) scores was nil. Although informative, it is important to point out that these findings were based on a small and homogeneous sample with respect to examinees' medical education and clinical skill levels. As such, there is a need for this type of research to be replicated with a more varied sample of examinees, to obtain an estimate of disclosure effects that might generalize to a more heterogeneous population of medical students. The purpose of the present study was to model the impact of disclosing test materials on SP examination scores with a sample of international medical graduates. Furthermore, it is hoped that ensuing findings will provide a practical estimate of expected effect size within the context of this type of security breach and with this population. Method Examination. In this investigation, the SP test assessed the clinical (history taking, physical examination, communication) skills and IPS of physicians about to enter supervised practice. SPs are laypeople trained to portray one of a variety of clinical scenarios. Test candidates rotate through these scenarios (or cases) and encounter patients in a setting intended to reflect an ambulatory care clinic. Case-specific checklists are used to assess examinees' clinical skills. These checklists are composed of dichotomously scored items, each of which represents a single action that is expected to be done by the student. A percent-correct score, corresponding to the number of actions done by the student out of the total number of behaviors listed in a given checklist, is computed for all encounters. IPS are assessed with the Patient Perception Questionnaire (PPQ), a case-independent inventory that is composed of six five-point Likert scale items. A percent-correct PPQ score is also computed and reported to each student for all encounters. Both measurement instruments are completed by the SP following each 15-minute encounter with the student. The same ten cases (chosen from the available pool) were administered to all examinees. The cases were selected to reflect the majority of cells contained in the test blueprint with regard to both skill and content domains. Scoring Procedure. In this examination, two SPs were trained to portray each case. For any given case, the performing SP portrayed the actual clinical scenario with the examinee, whereas the monitoring SP observed the encounter as it proceeded on a video screen in a separate room. Each student's final percent-correct checklist score reflected the consensus reached by the performing and monitoring SPs as to what constituted the appropriate response to each item. Videotape review was instituted to arrive at a consensus if two or more discrepancies per checklist were noted in any given encounter. Of the 9,625 checklist item responses recorded (77 students × 125 checklist items across the ten cases), videotape review was necessary for 202 (2.10%). The PPQ percent-correct score was derived from the performing SP. Examinees. Seventy-seven international medical graduates (IMGs), recruited from the Los Angeles metropolitan area, participated in this study and were blinded to its purpose. All examinees were certified by the Educational Commission for Foreign Medical Graduates, i.e., they had successfully passed the following examinations: Step 1 and Step 2 of the United States Medical Licensing Examination and a test of English-language proficiency. The examinees were paid for their participation and randomly assigned to one of two testing conditions: control or security breach (SB). The testing environment for examinees assigned to the control condition (n = 32) was representative of a “normal” assessment situation (i.e., participants received routine prior information about the test but no materials from the examination). In the SB condition, we attempted to model a situation in which actual case materials were disclosed. Examinees in the SB condition (n = 45) were directly provided with the checklists for five of the ten cases to be seen (referred to as the exposed cases) as well as the PPQ, and were given one to two hours to review these materials prior to completing the test. Information pertaining to the five non-exposed cases was not disclosed to any of the examinees participating in this study. Cases included in the exposed and non-exposed sets were matched with respect to the main areas of this SP test's blueprint. Analyses. Two separate analyses of covariance (ANCOVAs) were undertaken to compare the performances of the two groups on the five exposed cases. For both models, the condition factor (control or SB) was treated as the independent variable. The mean percent-correct checklist score on the five non-exposed cases was treated as the covariate in the first ANCOVA, while the mean percent-correct checklist score on the five exposed cases was deemed to be the dependent variable (DV). In the second analysis, the mean percent-correct PPQ score on the five non-exposed was deemed to be the covariate, whereas the mean percent-correct PPQ score on the five exposed cases was treated as the DV. Results Mean scores and standard errors on the five exposed cases for examinees assigned to each of the two conditions, adjusted for initial differences in ability between groups, were as follows: For examinees assigned to the control condition, the adjusted mean percent-correct checklist score was 54.53 (SE = 1.48), and the adjusted mean percent-correct Patient Perception Questionnaire score was 60.87 (SE = 1.18). For examinees assigned to the security breach condition, the adjusted mean percent-correct checklist score was 59.95 (SE = 1.24), and the adjusted mean percent-correct Patient Perception Questionnaire score was 67.03 (SE = 0.99). A significant group main effect was obtained in the first ANCOVA, F(1,74) = 7.66, p =.0071. For the exposed cases, the SB group (adjusted M = 59.95%) significantly outperformed the control group (adjusted M = 54.53) on the checklist. Similarly, the mean PPQ score for examinees assigned to the SB condition (adjusted M = 67.03%) was significantly higher than the mean estimated for the control group (adjusted M = 60.87%), F(1,74) = 15.84, p =.0002. Conclusions Results obtained in the present study with a sample of international medical graduates mirror those reported in previous research with USMGs.13 Disclosing checklist items led to significant performance gains for the examinees assigned to the SB condition. The gain noted in this investigation (5.4%), was, however, slightly lower than that obtained with a sample of USMGs. This is probably attributable to the larger number of cases administered in the test form (ten as opposed to six in the past USMG study). Therefore, the challenge posed to the IMGs was slightly more daunting, as they had to sift through ten cases to identify the clinical scenarios for which they possessed disclosed materials and apply this information accordingly. Nonetheless, the gain noted would concretely translate itself into a 4.4-checklist-item disadvantage over five cases (slightly less than one item per case). This advantage might be inconsequential for most USMGs, who typically perform well above the cut-score on this type of examination.14 However, it could significantly affect decision consistency for IMGs, whose scores tend to cluster in the vicinity of the pass/fail standard in a larger proportion. The control and SB groups also did differ significantly with respect to their mean PPQ scores, a result that was not found with USMGs.13 Interestingly, the difference between the two groups (6.2%) was actually larger than the one resulting from disclosing checklist items. This could reflect a difference in interaction styles that is culturally based. Disclosing simple indicators of IPS (such as the Likert-scale items found on the PPQ) to SB group examinees yielded a mean score that was similar to that typically encountered with U.S. medical students. It is also worth noting that the type of case that was most susceptible to the effects of disclosure appears to be population-dependent. For U.S. medical students, prior research suggested that cases involving largely mechanical physical examination maneuvers were the easiest to memorize and consequently reflected the highest performance gains for those examinees with prior knowledge of materials. Divulging materials for cases that primarily require communication and IPS in the interaction with the patient proved to be the most beneficial for our sample of IMGs. Again, these findings appear be indicative of differences in the way our sample of IMGs interacted with the SPs. These results suggest that providing a clear description of the examination and its goal to all examinees prior to the administration (in some form of information bulletin, for example) is necessary to ensure a common understanding of expected behavior on the part of students. In summary, the results presented in this study provide further evidence that the secure handling of test materials is essential for all examinations, whether they be traditional in format or performance-based. Although the security breach modeled in this investigation was severe (half of the test materials were directly exposed to students), steps can nonetheless be undertaken to minimize the likelihood of materials being disclosed. This, in turn, might lessen the impact of a security breach should checklists or other pertinent information fall into the hands of dishonest individuals. One obvious strategy that should be adopted with all SP tests is to clearly lay out the flow of materials and restrict access solely to concerned staff so that these individuals can be held accountable for receipt and safekeeping of this information. Delivering the measurement instruments via a computer network also seems advisable, given the greater control that the latter medium can afford and the virtual elimination of a “paper trail.” The results of our study also point out the need to increase test development efforts to minimize the likelihood of a security breach. Increasing the pool of available cases enables a more frequent rotation of forms within and across test sites, thus limiting the exposure rate for any given set. Finally, the use of modeled or cloned cases also seems desirable to increase the size of the case pool and thwart those individuals who may have mechanically memorized cases and accompanying materials. Modeled cases are defined as those presenting a similar opening scenario but requiring a different work-up on the part of the student. Cloned cases, on the other hand, call for a similar set of actions on the part of the student but present different contexts. Although informative, our results need to be interpreted in light of several limitations. First, the sample size examined was small, and generalizations should be made with caution. Our sample was also composed of IMGs who were perhaps atypical of the corresponding population, given that they had successfully fulfilled several U.S. medical licensing requirements (passed the USMLE Step 1 and Step 2 and a test of English-language proficiency). Consequently, the effect sizes reported in this study should probably be viewed as lower-bound estimates of what to expect in an operational testing context. Replication of this research with different groups of both IMGs and USMGs seems advisable. This research might also permit us to test the hypothesis that lower-ability students might benefit more from gaining access to materials than would those who are more proficient. From a test-development perspective, pursuing research that focuses on the identification of characteristics that make a case more vulnerable to memorization would also be helpful. Finally, the findings reported in this study underscore the need to develop methods to detect breaches to the security of the testing environment. Research aimed at assessing the usefulness of “tagged” checklist items and other means should be pursued.15 Testing organizations and medical schools should always be vigilant in guarding themselves against dishonest examinees and organizations that may wish to compromise the secure nature of the testing environment. This investigation confirms past findings in that the psychometric properties of the SP examination described appear to be vulnerable to blatant disclosure of testing materials. It is hoped that the results presented in this article will foster future relevant research that will ultimately lead to the implementation of secure SP tests for licensure and other purposes.
Log in or Register Subscribe to journalSubscribe Get new issue alertsGet alerts Enter your Email address: Wolters Kluwer Health may email you for journal alerts and information, but is committed to maintaining your privacy and will not share your personal information without your express consent. For more information, please refer to our Privacy Policy. Subscribe to eTOC Secondary Logo Journal Logo All Articles Images Videos Podcasts Blogs Advanced Search Toggle navigation Subscribe Register Login Articles & Issues Current IssuePrevious IssuesPublished Ahead-of-Print Collections Editorials of Laura Weiss Roberts, MD, MAAM Last PageCOVID-19 and Medical EducationAddressing Race and Racism in Medical EducationeBooksView All For Authors Submit a ManuscriptInformation for AuthorsLanguage Editing ServicesAuthor Permissions Journal Info About the JournalAbout the AAMCJournal MastheadSubmit a ManuscriptAdvertising InformationSubscription ServicesReprints and Back IssuesClassified AdsRights and PermissionsFor ReviewersFor MediaFor Trainees All Articles Images Videos Podcasts Blogs Advanced Search
Log in or Register Subscribe to journalSubscribe Get new issue alertsGet alerts Enter your Email address: Wolters Kluwer Health may email you for journal alerts and information, but is committed to maintaining your privacy and will not share your personal information without your express consent. For more information, please refer to our Privacy Policy. Subscribe to eTOC Secondary Logo Journal Logo All Articles Images Videos Podcasts Blogs Advanced Search Toggle navigation Subscribe Register Login Articles & Issues Current IssuePrevious IssuesPublished Ahead-of-Print Collections Editorials of Laura Weiss Roberts, MD, MAAM Last PageCOVID-19 and Medical EducationAddressing Race and Racism in Medical EducationeBooksView All For Authors Submit a ManuscriptInformation for AuthorsLanguage Editing ServicesAuthor Permissions Journal Info About the JournalAbout the AAMCJournal MastheadSubmit a ManuscriptAdvertising InformationSubscription ServicesReprints and Back IssuesClassified AdsRights and PermissionsFor ReviewersFor MediaFor Trainees All Articles Images Videos Podcasts Blogs Advanced Search
COLLIVER, JERRY A.; SWARTZ, MARK H.; ROBBS, RANDALL S.; LOFQUIST, MARIANNE; COHEN, DEVRA; VERHULST, STEVEN J.
Log in or Register Subscribe to journalSubscribe Get new issue alertsGet alerts Enter your Email address: Wolters Kluwer Health may email you for journal alerts and information, but is committed to maintaining your privacy and will not share your personal information without your express consent. For more information, please refer to our Privacy Policy. Subscribe to eTOC Secondary Logo Journal Logo All Articles Images Videos Podcasts Blogs Advanced Search Toggle navigation Subscribe Register Login Articles & Issues Current IssuePrevious IssuesPublished Ahead-of-Print Collections Editorials of Laura Weiss Roberts, MD, MAAM Last PageCOVID-19 and Medical EducationAddressing Race and Racism in Medical EducationeBooksView All For Authors Submit a ManuscriptInformation for AuthorsLanguage Editing ServicesAuthor Permissions Journal Info About the JournalAbout the AAMCJournal MastheadSubmit a ManuscriptAdvertising InformationSubscription ServicesReprints and Back IssuesClassified AdsRights and PermissionsFor ReviewersFor MediaFor Trainees All Articles Images Videos Podcasts Blogs Advanced Search