Automatic Item Generation (AIG) refers to the process of using cognitive models to generate test items using computer modules. It is a new but rapidly evolving research area where cognitive and psychometric theory are combined into digital framework. However, assessment of the item quality, usability and validity of AIG relative to traditional item development methods lacks clarification. This paper takes a top-down strong theory approach to evaluate AIG in medical education. Two studies were conducted: Study I—participants with different levels of clinical knowledge and item writing experience developed medical test items both manually and through AIG. Both item types were compared in terms of quality and usability (efficiency and learnability) ; Study II—Automatically generated items were included in a summative exam in the content area of surgery. A psychometric analysis based on Item Response Theory inspected the validity and quality of the AIG-items. Items generated by AIG presented quality, evidences of validity and were adequate for testing student’s knowledge. The time spent developing the contents for item generation (cognitive models) and the number of items generated did not vary considering the participants' item writing experience or clinical knowledge. AIG produces numerous high-quality items in a fast, economical and easy to learn process, even for inexperienced and without clinical training item writers. Medical schools may benefit from a substantial improvement in cost-efficiency in developing test items by using AIG. Item writing flaws can be significantly reduced thanks to the application of AIG's models, thus generating test items capable of accurately gauging students' knowledge.
5012 Background: PROpel (NCT03732820) met its primary endpoint and showed significantly prolonged investigator-assessed rPFS with abi + ola vs abi + pbo at primary analysis (data cut-off [DCO]: 7/30/21; median 24.8 vs 16.6 months (m); hazard ratio [HR] 0.66, 95% confidence interval [CI] 0.54–0.81; P<0.001). HRQoL (based on The Functional Assessment of Cancer Therapy-Prostate [FACT-P] total score) was not different when ola was combined with standard-of-care abi. We present data at the final prespecified overall survival (OS) DCO (10/12/22). Methods: PROpel is a randomised, double-blind trial in 1L mCRPC. Time to pain progression (TTPP) was based on the Brief Pain Inventory-Short Form (BPI-SF) Item 3 ‘worst pain in 24 hours’ and opiate analgesic use (analgesic quantification algorithm) score. Time to first symptomatic skeletal related event (SSRE) was time to use of therapy to prevent/relieve skeletal symptoms, new bone fractures, spinal compression or surgery on bone metastases. HRQoL was assessed by change from baseline (BL) in FACT-P total and subscale scores, BPI-SF pain severity, pain interference and worst pain score between arms using a mixed model for repeated measures. Results: At median follow-up of 33.6 m with abi + ola and 32.1 m with abi + pbo, 17.0% pts (68/399) in the abi + ola arm and 15.1% pts (60/397) in the abi + pbo arm had pain progression (PP) events. The % of pts who had not experienced PP with abi + ola vs abi + pbo was 76.9% vs 77.2% at 24 m and 70.7% vs 71.0% at 36 m. No meaningful difference in TTPP was observed (16% maturity, HR 1.06, 95% CI 0.75–1.50, P=0.75 [nominal], median not reached [NR] either arm). The % of pts who had not had a SSRE with abi + ola vs abi + pbo was 86.1% vs 82.2% at 24 m and 80.8 vs 78.5% at 36 m. No meaningful difference in time to SSRE was observed (12% maturity, HR 0.82, 95% CI 0.55–1.22, P=0.32 [nominal], median NR vs NR). Least-squares mean changes from BL between arms in BPI-SF pain severity (difference, −0.06; 95% CI −0.23–0.12), pain interference (difference, −0.12; 95% CI −0.31–0.06), worst pain score (difference, −0.12; 95% CI −0.35–0.11) and FACT-P total score (difference, −0.54; 95% CI –3.00–1.92) suggest no clinically meaningful difference in HRQoL with abi + ola vs abi + pbo. Least-squares mean change from BL values for FACT-P subscale scores were consistent with FACT-P total score result. Conclusions: PROpel demonstrated a significant delay in rPFS for pts receiving abi + ola vs abi + pbo. Most pts in the trial did not experience a PP event. Abi + ola showed no difference in HRQoL (assessed by FACT-P total and subscale scores, BPI-SF domain and worst pain scores) vs abi + pbo, suggesting pts can derive clinical benefit from abi + ola while maintaining a similar HRQoL compared with a current standard-of-care treatment. Clinical trial information: NCT03732820 .
With the re-emergence of competency-based frameworks in professional education, multisource feedback (MSF) has become a common method for assessing various competencies, including communication, professionalism, and aspects of team-based performance. A wide variety of publications over the past 50 years or more in the business and health literature would seem to support the use of MSF at least for quality improvement (QI) purposes. However, our own experience with using MSF in physicians has been quite mixed, with some physicians embracing the experience and making improvements and others the complete opposite. We decided to review the existing literature on MSF to try to identify key aspects of successful MSF programs. This paper presents a structured critique of the literature on the use of MSF in physician populations. The findings were surprising as key assumptions around the validity and reliability of MSF were not consistently met, key lessons from earlier research were not carried over to present day programs and perhaps most concerning was a lack of evidence for MSF producing meaningful sustained behavior change. From these findings we suggest some key areas of potential improvement in MSF programs.
The purpose of this study was to compare the quality of multiple choice questions (MCQs) developed using automated item generation (AIG) versus traditional methods, as judged by a panel of experts. The quality of MCQs developed using two methods (i.e., AIG or traditional) was evaluated by a panel of content experts in a blinded study. Participants rated a total of 102 MCQs using six quality metrics and made a judgment regarding whether or not each item tested recall or application of knowledge. A Wilcoxon two-sample test evaluated differences in each of the six quality metrics rating scales as well as an overall cognitive domain judgment. No significant differences were found in terms of item quality or cognitive domain assessed when comparing the two item development methods. The vast majority of items (> 90%) developed using both methods were deemed to be assessing higher-order skills. When compared to traditionally developed items, MCQs developed using AIG demonstrated comparable quality. Both modalities can produce items that assess higher-order cognitive skills.
ABSTRACT The purpose of this longitudinal study was to gather extrapolation evidence of validity by assessing whether performance on a national medical licensing exam, in addition to practice and socio-demographic variables, is predictive of future physician performance in practice. The study focused on a cohort of 3,404 physicians who were registered with the College of Physicians and Surgeons of Alberta (CPSA) and who completed the Medical Council of Canada Qualifying Examination (MCCQE) Parts I and II between 1992–2017. Separate multivariate quasi-Poisson regression models were run to assess the degree of relationship between first-time pass/fail status on the MCCQE I and II, and several CPSA socio-demographic variables and several CPSA socio-demographic variables, in addition to complaints/physician and various prescribing flags. Candidates who failed the MCCQE I on their first attempt had 27% more complaints lodged against them, compared to those who passed. Physicians who failed the MCCQE II on their first attempt prescribed 2+ benzodiazepines and 2+ opioids to 30% more patients than those who passed. Conclusions: Performance on the MCCQE Part I and II is an important predictor of physician performance. Combined with other critical variables, these measures provide important evidence to aid in risk modeling efforts and to guide educational interventions for physicians at an early stage of their careers.
Despite the increased emphasis on the use of workplace-based assessment in competency-based education models, there is still an important role for the use of multiple choice questions (MCQs) in the assessment of health professionals. The challenge, however, is to ensure that MCQs are developed in a way to allow educators to derive meaningful information about examinees' abilities. As educators' needs for high-quality test items have evolved so has our approach to developing MCQs. This evolution has been reflected in a number of ways including: the use of different stimulus formats; the creation of novel response formats; the development of new approaches to problem conceptualization; and the incorporation of technology. The purpose of this narrative review is to provide the reader with an overview of how our understanding of the use of MCQs in the assessment of health professionals has evolved to better measure clinical reasoning and to improve both efficiency and item quality.
There exists an assumption that improving medical education will improve patient care. While seemingly logical, this premise has rarely been investigated. In this Invited Commentary, the authors propose the use of big data to test this assumption. The authors present a few example research studies linking education and patient care outcomes and argue that using big data may more easily facilitate the process needed to investigate this assumption. The authors also propose that collaboration is needed to link educational and health care data. They then introduce a grassroots initiative, inclusive of universities in one Canadian province and national licensing organizations that are working together to collect, organize, link, and analyze big data to study the relationship between pedagogical approaches to medical training and patient care outcomes. While the authors acknowledge the possible challenges and issues associated with harnessing big data, they believe that the benefits supersede these. There is a need for medical education research to go beyond the outcomes of training to study practice and clinical outcomes as well. Without a coordinated effort to harness big data, policy makers, regulators, medical educators, and researchers are left with sometimes costly guesses and assumptions about what works and what does not. As the social, time, and financial investments in medical education continue to increase, it is imperative to understand the relationship between education and health outcomes.
Introduction Medical education and regulatory bodies do not often share performance data due to privacy concerns. Innovative approaches are needed to facilitate research while preserving security and privacy. To this end, a privacy preserving protocol was employed linking medical examination and regulatory data to examine future physician competence across the career. Objectives and Approach This study extends previous work linking de-identified Canadian medical licensing examination data with medical regulatory outcomes to answer the following question: is there a predictive relationship between licensing examination scores and post-licensure practice outcomes? A privacy preserving protocol using a third party organization was employed to link data between two disparate organizations - a medical licensing examination organization (MLE) and a medical regulatory authority (MRA). Multiple years of licensing examinations were linked to thirteen years of regulatory assessment outcomes (2004 – 2016) without identifiable data being shared to either party. Results Medical Identification Number for Canada (MINC) was used as a common identifying variable between the two organizations. First, the analytic cohort was created by linking identifying variables of the physicians of interest from both parties, thereby creating a common cohort. The third-party organization then created an encryption key using the common cohort and the MLE examination data. The key was given to the MRA and the encrypted, de-identified examination data was given back to the MLE. Lastly, the MRA data was de-identified, encrypted and transferred to the MLE for analysis. This ensured neither party had access to each other’s encrypted data and the key simultaneously. Conclusion/Implications Privacy preserving protocols enhance opportunities for novel research questions and data linkages within and across sectors; here, results from this analysis may enhance the utility of medical licensing exams by providing evidence for secondary uses. Furthermore, it will offer other physician organizations evidence to support physicians across their career trajectory.
Construct: Valid score interpretation is important for constructs in performance assessments such as objective structured clinical examinations (OSCEs). An OSCE is a type of performance assessment in which a series of standardized patients interact with the student or candidate who is scored by either the standardized patient or a physician examiner.BACKGROUND:In high-stakes examinations, test security is an important issue. Students accessing unauthorized test materials can create an unfair advantage and lead to examination scores that do not reflect students' true ability level. The purpose of this study was to assess the impact of various simulated security breaches on OSCE scores.APPROACH:Seventy-six 3rd-year medical students participated in an 8-station OSCE and were randomized to either a control group or to 1 of 2 experimental conditions simulating test security breaches: station topic (i.e., providing a list of station topics prior to the examination) or egregious security breach (i.e., providing detailed content information prior to the examination). Overall total scores were compared for the 3 groups using both a one-way between-subjects analysis of variance and a repeated measure analysis of variance to compare the checklist, rating scales, and oral question subscores across the three conditions.RESULTS:Overall total scores were highest for the egregious security breach condition (81.8%), followed by the station topic condition (73.6%), and they were lowest for the control group (67.4%). This trend was also found with checklist subscores only (79.1%, 64.9%, and 60.3%, respectively for the security breach, station topic, and control conditions). Rating scale subscores were higher for both the station topic and egregious security breach conditions compared to the control group (82.6%, 83.1%, and 77.6%, respectively). Oral question subscores were significantly higher for the egregious security breach condition (88.8%) followed by the station topic condition (64.3%), and they were the lowest for the control group (48.6%).CONCLUSIONS:This simulation of different OSCE security breaches demonstrated that student performance is greatly advantaged by having prior access to test materials. This has important implications for medical educators as they develop policies and procedures regarding the safeguarding and reuse of test content.
Purpose: The aim of this research was to compare different methods of calibrating multiple choice question (MCQ) and clinical decision making (CDM) components for the Medical Council of Canada’s Qualifying Examination Part I (MCCQEI) based on item response theory. Methods: Our data consisted of test results from 8,213 first time applicants to MCCQEI in spring and fall 2010 and 2011 test administrations. The data set contained several thousand multiple choice items and several hundred CDM cases. Four dichotomous calibrations were run using BILOG-MG 3.0. All 3 mixed item format (dichotomous MCQ responses and polytomous CDM case scores) calibrations were conducted using PARSCALE 4. Results: The 2-PL model had identical numbers of items with chi-square values at or below a Type I error rate of 0.01 (83/3,499 or 0.02). In all 3 polytomous models, whether the MCQs were either anchored or concurrently run with the CDM cases, results suggest very poor fit. All IRT abilities estimated from dichotomous calibration designs correlated very highly with each other. IRT-based pass-fail rates were extremely similar, not only across calibration designs and methods, but also with regard to the actual reported decision to candidates. The largest difference noted in pass rates was 4.78%, which occurred between the mixed format concurrent 2-PL graded response model (pass rate= 80.43%) and the dichotomous anchored 1-PL calibrations (pass rate= 85.21%). Conclusion: Simpler calibration designs with dichotomized items should be implemented. The dichotomous calibrations provided better fit of the item response matrix than more complex, polytomous calibrations.
Construct: Automatic item generation (AIG) is an alternative method for producing large numbers of test items that integrate cognitive modeling with computer technology to systematically generate multiple-choice questions (MCQs). The purpose of our study is to describe and validate a method of generating plausible but incorrect distractors. Initial applications of AIG demonstrated its effectiveness in producing test items. However, expert review of the initial items identified a key limitation where the generation of implausible incorrect options, or distractors, might limit the applicability of items in real testing situations. Background: Medical educators require development of test items in large quantities to facilitate the continual assessment of student knowledge. Traditional item development processes are time-consuming and resource intensive. Studies have validated the quality of generated items through content expert review. However, no study has yet documented how generated items perform in a test administration. Moreover, no study has yet to validate AIG through student responses to generated test items. Approach: To validate our refined AIG method in generating plausible distractors, we collected psychometric evidence from a field test of the generated test items. A three-step process was used to generate test items in the area of jaundice. At least 455 Canadian and international medical graduates responded to each of the 13 generated items embedded in a high-stake exam administration. Item difficulty, discrimination, and index of discrimination estimates were calculated for the correct option as well as each distractor. Results: Item analysis results for the correct options suggest that the generated items measured candidate performances across a range of ability levels while providing a consistent level of discrimination for each item. Results for the distractors reveal that the generated items differentiated the low- from the high-performing candidates. Conclusions: Previous research on AIG highlighted how this item development method can be used to produce high-quality stems and correct options for MCQ exams. The purpose of the current study was to describe, illustrate, and evaluate a method for modeling plausible but incorrect options. Evidence provided in this study demonstrates that AIG can produce psychometrically sound test items. More important, by adapting the distractors to match the unique features presented in the stem and correct option, the generation of MCQs using automated procedure has the potential to produce plausible distractors and yield large numbers of high-quality items for medical education.
With the recent interest in competency-based education, educators are being challenged to develop more assessment opportunities. As such, there is increased demand for exam content development, which can be a very labor-intense process. An innovative solution to this challenge has been the use of automatic item generation (AIG) to develop multiple-choice questions (MCQs). In AIG, computer technology is used to generate test items from cognitive models (i.e. representations of the knowledge and skills that are required to solve a problem). The main advantage yielded by AIG is the efficiency in generating items. Although technology for AIG relies on a linear programming approach, the same principles can also be used to improve traditional committee-based processes used in the development of MCQs. Using this approach, content experts deconstruct their clinical reasoning process to develop a cognitive model which, in turn, is used to create MCQs. This approach is appealing because it: (1) is efficient; (2) has been shown to produce items with psychometric properties comparable to those generated using a traditional approach; and (3) can be used to assess higher order skills (i.e. application of knowledge). The purpose of this article is to provide a novel framework for the development of high-quality MCQs using cognitive models.
The incorporation of learner assessments has become part and parcel of the accreditation process over the past few decades as a means of evaluating program or instructional effectiveness.1 Given the high stakes associated with assessments not only for individual candidate-based decisions but also programs as a whole, it is critical to ensure that scores based on any tools meet certain psychometric standards. At its most elemental level, any test score is intended to reflect the competency domain(s) presumed to underlie an assessment. For example, if a candidate obtains a score of 90% on a direct observation tool, this might be interpreted as reflecting “strong” patient care, even though the latter is, in all likelihood, established on a small number of encounters. Given that high-stakes decisions may be based on such observational tools, it is critical that the sample of performance be reflective of the candidate's true ability in that competency. Reliability refers to the extent to which performance on any assessment (ie, in a restricted number of encounters) is indicative of the candidate's true competency level (ie, in an infinite number of encounters).2 An “unreliable” assessment (ie, one that does not reflect the candidate's true competency level) could have dire consequences not only for the physician's medical education but also for the accreditation of the postgraduate program.Due to restricted testing time, any assessment encompasses a limited sample of encounters that theoretically represents the domain of interest. The selection of 10 patients for inclusion into a direct observation assessment, for example, might be predicated on 3 hours of testing time. However, one could conceive of different sets of 10 patients that could have been selected. The program director who is reviewing a candidate's score of 90% with these 10 patients is not interested in restricting his or her interpretation of that “strong” performance to these 10 specific encounters, but rather generalizes this statement to the (theoretically infinite) pool of encounters from which the sample of 10 was selected.Yet, several sources of measurement error can detract from the accuracy or precision with which the performance on a restricted sample of encounters generalizes to the broader domain. With performance assessments (in addition to the restricted sample of encounters), the examiners, the setting, and other factors can impede a candidate's score. Reliability allows us to estimate how well a score on any assessment (ie, a sample of performance) generalizes to the broader domain(s) of interest. With the previous example, how accurately does a score of 90%, in 10 patient encounters, scored by 10 examiners, generalize to all possible patient encounters and physician examiners? This generalization is quantified with a reliability coefficient.Note that patients and examiners are sources of measurement error, given that any candidate's true score or ability level should not depend on the sample of patients nor the examiners encountered. A candidate's true ability level should be invariant across all these sources of measurement error or facets. In reality, all of these sources will detract from reliability due to the lack of representativeness of the patient encounters selected for an examination and the poor training of examiners.Commonly, Cronbach's α coefficient is computed as the reliability estimate largely because it is readily available in most statistical software packages.3 However, the use of Cronbach's α with examinations that are affected by several sources of measurement error, such as performance-based assessments, is ill-advised. Specifically, this coefficient does not partition all sources of measurement error in the computation of the reliability coefficient; rather, it is restricted to only 1 facet (ie, “patient encounters”) in the previous example. Cronbach's α can thus yield a very misleading (spurious) reliability estimate because of its inability to incorporate (and partition out) all sources of measurement error.Generalizability Theory (G Theory) is a reliability framework that allows us to properly quantify the impact of these error sources in regard to the extent to which we can generalize performance in a restricted sample of conditions (patient encounters, examiners) to broader domains. G Theory is an extension of Cronbach's α that allows the user to prespecify and estimate the impact of all potential sources of measurement error.4G Theory uses analysis of variance modeling to estimate the amount of variability in scores due to sources of measurement error as well as their impact on the reliability coefficient, referred to as a generalizability coefficient (G coefficient). To use G Theory most efficiently and in a helpful manner, careful consideration must be given to all aspects of examination development (eg, the number of raters, how they are to be assigned to candidates, the number of stations, etc). For example, to estimate how much variance is due to raters, the raters need to score some common elements of the assessment (ie, either common patients or candidates).To illustrate the application of G Theory, imagine a 9-patient encounter assessment that is completed by 90 candidates as a requirement in a given postgraduate program. The assessment targets the “patient care” Accreditation Council for Graduate Medical Education competency. Three examiners are assigned to rate different candidates: (1) examiner 1 rates candidates 1 to 30; (2) examiner 2 rates candidates 31 to 60; and (3) examiner 3 rates candidates 61 to 90. In G Theory parlance, this is a p:r × pe design, where p, r, and pe respectively correspond to persons (candidates), raters (examiners), and patient encounters. Persons are nested within raters (p:r), since not all candidates are rated by the same examiner. Furthermore, persons nested within raters are crossed with patient encounter (p:r × pe) because it is assumed in this example that all candidates encounter the same 9 patients in their assessment.The Cronbach's α value for this dataset was 0.86, which users may infer as “highly reliable” (ie, scores generalize well to domains targeted by the examination and allow us to accurately rank order candidates from low to high). However, the reliability estimate is spuriously inflated, as supported by an analysis using G Theory conducted on the same dataset using a p:r × pe design (table).The variance component associated with p:r is akin to true score variance, as it provides an estimate of the amount of score variability due to true differences in ability among candidates. Specifically, 7% of total score variance is due to true difference in ability among candidates, suggesting some modest spread and consequently some capability to differentiate candidates (rank order). The variance component due to patient encounter reflects difficult differences. The small percentage of variance accounted for by this source suggests that encounters were highly comparable in terms of difficulty. The r × pe variance component, which is virtually nil, indicates that the stringency level of the examiners did not differ as a function of the patient encounter. Finally, the p:r × pe, e component is a residual term, which reflects the amount of error in generalizing due to all other sources not specified in the design.Of particular interest in this example is the large amount of variance due to raters (39.5%). Nearly 40% of the variance in the assessment scores is due to differences in stringency between the 3 raters. Note that this effect is completely independent from the abilities of the 90 candidates. A G coefficient of 0.57 was computed—a significantly lower value than Cronbach's α.What might account for the large difference in reliability estimates obtained with the assessment scores? In calculating Cronbach's α, the large differences in candidate scores due to the high variability among examiners (a source of measurement error) gets “confounded,” as true score variance which artificially inflates the reliability coefficient. The highly divergent examiners are “injecting” a high level of score variance due to their own variability as examiners rather than being reflective of differences in candidate ability levels. Since there is no mechanism in the calculation of Cronbach's α to account for examiner variability, this gets incorrectly partitioned as true score variance or true differences among candidate abilities. In G Theory, the error variance due to examiners is correctly partitioned out of true score variance and treated as a source of measurement error, which appropriately lowers the G coefficient value.This example illustrates the pitfalls that can result from the sole use of Cronbach's α coefficient in estimating the reliability of scores with highly complex assessments, such as those commonly used in postgraduate medical education. It is important to point out, however, that Cronbach's α is appropriate in instances where a single rater is involved in the assessment. Also, for those assessments, such as simulations, which may involve clear scoring keys with little to no rater input, reliability can be confidently estimated with Cronbach's α given that there is only 1 source of measurement error (scenario).However, in the example used to illustrate the concept, the high Cronbach's α value (0.86) could lead the medical educator to commit erroneous high-stakes decisions (promotion, graduation, etc), given the limitations of the reliability coefficient and its inability to properly account for the large amount of variability due to examiners. For any assessment that involves several facets (multi-source feedback, direct observation–based rating scales, etc), it is highly recommended that the practitioner complete a generalizability analysis not only to properly estimate reliability, but also to garner information that might be beneficial in improving the assessment for future uses.
Item development is a time-and resource-intensive process. Automatic item generation integrates cognitive modeling with computer technology to systematically generate test items. To date, however, items generated using cognitive modeling procedures have received limited use in operational testing situations. As a result, the psychometric characteristics of generated multiple-choice test items are largely unknown and undocumented. We present item analysis results from one of the first empirical studies designed to evaluate the psychometric properties of generated multiple-choice items using the results from a high stakes national medical licensure examination. The item analysis results for the correct option revealed that the generated items measured examinees' performance across a broad range of ability levels while, at the same time, providing a consistently strong level of discrimination for each item. Results for the incorrect options revealed that the generated items consistently differentiated the low from the high performing examinees.
Examiner effects and content specificity are two well known sources of construct irrelevant variance that present great challenges in performance-based assessments. National medical organizations that are responsible for large-scale performance based assessments experience an additional challenge as they are responsible for administering qualification examinations to physician candidates at several locations and institutions. This study explores the impact of site location as a source of score variation in a large-scale national assessment used to measure the readiness of internationally educated physician candidates for residency programs. Data from the Medical Council of Canada's National Assessment Collaboration were analyzed using Hierarchical Linear Modeling and Rasch Analyses. Consistent with previous research, problematic variance due to examiner effects and content specificity was found. Additionally, site location was also identified as a potential source of construct irrelevant variance in examination scores.
Purpose: This study aims to assess the fit of a number of exploratory and confirmatory factor analysis models to the 2010 Medical Council of Canada Qualifying Examination Part I (MCCQE1) clinical decision-making (CDM) cases. The outcomes of this study have important implications for a range of domains, including scoring and test development. Methods: The examinees included all first-time Canadian medical graduates and international medical graduates who took the MCCQE1 in spring or fall 2010. The fit of one- to five-factor exploratory models was assessed for the item response matrix of the 2010 CDM cases. Five confirmatory factor analytic models were also examined with the same CDM response matrix. The structural equation modeling software program Mplus was used for all analyses. Results: Out of the five exploratory factor analytic models that were evaluated, a three-factor model provided the best fit. Factor 1 loaded on three medicine cases, two obstetrics and gynecology cases, and two orthopedic surgery cases. Factor 2 corresponded to pediatrics, and the third factor loaded on psychiatry cases. Among the five confirmatory factor analysis models examined in this study, three- and four-factor lifespan period models and the five-factor discipline models provided the best fit. Conclusion: The results suggest that knowledge of broad disciplinary domains best account for performance on CDM cases. In test development, particular effort should be placed on developing CDM cases according to broad discipline and patient age domains; CDM testlets should be assembled largely using the criteria of discipline and age.
We present a framework for technology-enhanced scoring of bilingual clinical decision-making (CDM) questions using an open-source scoring technology and evaluate the strength of the proposed framework using operational data from the Medical Council of Canada Qualifying Examination. Candidates' responses from six write-in CDM questions were used to develop a three-stage-automated scoring framework. In Stage 1, the linguistic features from CDM responses were extracted. In Stage 2, supervised machine learning techniques were employed for developing the scoring models. In Stage 3, responses to six English and French CDM questions were scored using the scoring models from Stage 2. Of the 8,007 English and French CDM responses, 7,643 were accurately scored with an agreement rate of 95.4% between human and computer scoring. This result serves as an improvement of 5.4% when compared with the human inter-rater reliability. Our framework yielded scores similar to those of expert physician markers and could be used for clinical competency assessment.
First‐year residents begin clinical practice in settings in which attending staff and senior residents are available to supervise their work. There is an expectation that, while being supervised and as they become more experienced, residents will gradually take on more responsibilities and function independently.