Entrustable professional activities (EPAs) have been proposed as a holistic approach to competency-based assessment. The 13 Core EPAs for Entering Residency (CEPAER) are essential tasks that a medical student should be trusted to perform with indirect supervision upon entering residency, based on demonstrated competence. To study the validity and reliability of workplace-based assessments of the 13 core EPAs as measurements of medical student performance and growth over the Internal Medicine (IM) clerkship. Correlational-based population study. A total of 398 third-year medical students at the University of Minnesota Medical School participated. Students were enrolled in a required 8-week IM clerkship during the 2023–2024 and 2024–2025 academic years. A total of 825 assessors provided EPA ratings with a mean number of 12 per assessor; SD = 15.08. There were 10,034 EPA-based assessments collected (mean per student = 25; SD = 6.1). The most frequently assessed EPAs were EPA 6 (Provide an oral presentation of a clinical encounter; n = 1866; mean per student = 4.69), EPA 5 (Document a clinical encounter in the patient record; n = 1662; mean per student = 4.18), and EPA 2 (Recommend and interpret common diagnostic and screening tests; n = 1421; mean per student = 3.57). Regression analyses indicated statistically significant growth in entrustment scores for EPAs 1, 2, 3, 5, 6, 8, 10, and 12. Generalizability analysis showed that to achieve adequate reliability (Ep2 ≥ 0.80), at least 5 assessments were required to be conducted by 5 raters. EPAs represent a valid and reliable measure for medical student growth during the IM clerkship, particularly for EPAs 1, 2, 3, 5, 6, 8, and 10.
BackgroundA common practice in assessment development, fundamental for fairness and consequently the validity of test score interpretations and uses, is to ascertain whether test items function equally across test-taker groups. Accordingly, we conducted differential item functioning (DIF) analysis, a psychometric procedure for detecting potential item bias, for three preclinical medical school foundational courses based on students' sex and race.MethodsThe sample included 520, 519, and 344 medical students for anatomy, histology, and physiology, respectively, collected from 2018 to 2020. To conduct DIF analysis, we used the Wald test based on the two-parameter logistic model as utilized in the IRTPRO software.ResultsThe three assessments had as many as one-fifth of the items that functioned statistically differentially across one or more of the variables sex and race: 10 out of 49 items (20%), six out of 40 items (15%), 5 out of 45 items (11%) showed statistically significant DIF for Anatomy, Histology, and Physiology courses, respectively. Measurement specialists and subject matter experts independently reviewed the items to identify construct-irrelevant factors as potential sources for DIF as demonstrated in Appendix A. Most identified items were generally poorly written or had unclear images.ConclusionsThe validity of score-based inferences, particularly for group comparisons, requires test items to function equally across test-taker groups. In the present study, we found DIF of some items for sex and race in three content areas. The present approach should be utilized in other medical schools to address the generalizability of the present findings. Item level DIF should also be routinely conducted as part of psychometric analyses for basic sciences courses and other assessments.Clinical trial numberNot applicable.
ABSTRACTPurposeFaculty wellbeing impacts student learning and is a priority among medical schools, especially as a counterbalance to growing burnout. Previous researchers found differences in burnout by sex and race among clinicians, but not for faculty with disabilities. Accordingly, the purpose was to test the association between faculty's wellbeing, burnout, and control over workload and investigate differences in wellbeing attributed to department type and ability status.MethodThe authors developed and administered a comprehensive wellbeing survey to University of Minnesota Medical School faculty, of whom 703 provided complete responses. The authors conducted two‐way ANOVA followed by a post hoc analysis to test for differences in faculty wellbeing domains due to department type (basic sciences, nonsurgical, surgical, and two large departments of Medicine and Pediatrics) and disability status (yes, no). The authors also fitted a two‐way ordinal model since burnout frequency and control over workload were assessed by one ordinal item each.ResultsWellbeing domains were positively correlated with control over workload but negatively associated with burnout. Faculty with disabilities reported less support from their work environment and meeting of their basic needs. Department type had a statistically significant impact on faculty's sense of basic needs, respect, and contribution. Multiple comparisons revealed faculty in basic sciences departments had higher scores within basic needs compared to the departments of Medicine, Pediatrics, and surgical departments, who reported lower levels of respect as well. Results revealed department type and disability status affected the frequency of burnout, as faculty in basic sciences departments reported lower levels of burnout compared to other departments.ConclusionsResults support disaggregating wellbeing by department and ability status for targeted interventions due to differences‐ notably among faculty with disabilities and surgical departments‐ in their assessment of basic needs, work environment, respect, and contribution. Results suggest revisiting interventions in these domains to account for lower reported wellbeing.
Purpose: Limited research has been conducted on the differential predictive validity of the Medical College Admission Test (MCAT) section scores (e.g., Biological and Biochemical Foundations of Living Systems [BBLS], Chemical and Physical Foundations of Biological Systems [CPBS]) for basic science courses (anatomy and histology). Accordingly, the purpose was to assess the predictive validity and differential prediction of these section scores to predict anatomy and histology performance across sex (men vs. women) and race (white vs. nonwhite). Methods: The authors analyzed data from 520 undergraduate medical students (sex: 292 women [56.15%], 228 men [43.85%]) in anatomy and histology courses. The authors utilized multiple linear regression and t-tests to test for statistically significant differences in slopes associated with each section score across sex and race groups. Results: BBLS and CPBS section scores explained more variance in histology compared to anatomy, particularly for the nonwhite and men groups (25% and 29% vs. 24%). For the differential prediction, t-tests were not statistically significant for most analyses, which provided evidence for the comparable predictive validity across the groups or lack of differential prediction, a desired psychometric property for fair assessments. The t-test associated with the BBLS section score, however, was statistically significant between the women and men groups only in anatomy. Conclusions: Current results contribute to the broader evidence supporting the validity of MCAT section scores. The general absence of differential prediction (i.e., comparable predictions across groups) supports the fairness of MCAT-based inferences, bolstering confidence in its use for medical school admissions and promoting equity.
Creating original, integrated multiple-choice questions (MCQs) is time-consuming and onerous for basic science and clinical faculty. We demonstrate that medical students are co-experts to overcome assessment challenges of the faculty. We recruited, trained, and motivated medical students to write 10,000 high-quality MCQs for use in the foundational courses of medical education. These students were ideal because they possessed integrated knowledge (basic sciences and clinical experience). We taught them how to write high-quality MCQs using a writing template and continuous monitoring and support by an item bank curator. The students themselves also benefitted personally and pedagogically from the experience.
Background: There has been recently a common practice in assessment development to ascertain that test items function equally across test-takers’ subgroups, which is fundamental for fairness, and consequently validity of test score interpretations and uses. Accordingly, we conducted differential item functioning (DIF) analysis for three preclinical medical school foundational courses based on students’ sex and race. Methods: The sample included 520, 519, and 344 medical students for anatomy, histology, and physiology, respectively, collected from 2018-2020. To conduct DIF analysis, we used the IRTPRO software based on the item response theory two-parameter logistic model. Results : The three assessments had as many as one-fifth of the items that functioned differentially across one or more of the variables sex and race: 10 (20%) out of 49 items, six (15%) out of 40 items, 5 (11%) out of 45 items showed statistically significant DIF for Anatomy , Histology , and Physiology courses, respectively. Measurement specialists and subject matter experts independently reviewed the items to identify construct-irrelevant factors as potential sources for DIF. Most identified items were generally poorly written or had unclear images. Conclusions: The validity of score-based inferences, particularly for subgroup comparisons, requires test items to function equally across examinee subgroups. In the present study, we found DIF of some items for sex and race in three content areas. The present approach should be explored in other medical schools to address the generalizability of the present findings. Item level DIF should also be routinely conducted as part of psychometric analyses for basic sciences courses and other assessments.
PURPOSE:This study examines the feasibility and psychometric results of an assessment of entrustable professional activities (EPAs) as a core component of the clinical program of assessment in undergraduate medical education, assesses the learning curves for each EPA, explores the time to entrustment, and investigates the dependability of the EPA data based on generalizability theory (G theory) analysis. METHOD:Third-year medical students from the University of Minnesota Medical School in 7 required clerkships from May 2022 through April 2023 were assessed. Students were required to obtain at least 4 EPA assessments per week on average from clinical faculty, residents supervising the students, or assessment and coaching experts. Student ratings were depicted as curves describing their performance over time; regression models were used to fit the curves. RESULTS:The complete class of 240 (138 women [58.0%] and 102 men [42.0%]) third-year medical students at the University of Minnesota Medical School (mean [SD] age at matriculation, 24.2 [2.7] years) participated. There were 32,614 EPA-based assessments (mean [SD], 136 [29.6] assessments per student). Reliability analysis using G theory found that an overall score dependability of 0.75 (range, 0-1) was achieved with 4 assessors on 4 occasions. The desired level of entrustment by academic year end was met by all 240 students (100%) for EPAs 1, 6, and 7, 237 (98.8%), 236 (98.3%), and 218 (90.8%) students for EPAs 2, 5, and 9, respectively, 197 students (82.1%) for EPA 3, 178 students (74.2%) for EPA 4, and 145 students (60.4%) for EPA 12. The most rapid growth was for EPA 2 (β 0 = .286), followed by EPA 1 (β 0 = .240), EPA 4 (β 0 = .236), and EPA 10 (β 0 = .230). CONCLUSIONS:The study findings suggest that EPA ratings provide reliable and dependable data to make entrustment decisions about students' performance.
Evidence-based practice (EBP) is currently considered as the golden standard for patient care. Many universities offer EBP courses to their healthcare professions students. However, no quantitative evidence synthesis has been conducted to compare EBP e-learning instructional methods to traditional methods, to better inform health education policymakers. Eight randomised studies reporting the effectiveness of e-learning methods compared to “no intervention” or to any other educational methods and including 1,243 learners met the inclusion criteria. The meta-analytical results revealed that e-learning was significantly better than “no intervention” [d = 1.4, 95% confidence interval (CI) = 1.060 to 1.776, I² = 99.5%, p < 0.0001] and as effective as other traditional methods such as lectures (d = 0.30, 95% CI = –0.348 to 0.952, I² = 90.5%, p = 0.3). The same conclusions were found when using the adjusted exam scores in relation to confounding variables such as the baseline characteristics and prior EBP knowledge of participants. The present meta-analysis demonstrates that teaching EBP via e-learning is an effective instructional method in times when lecture hours and face-to-face didactics are reduced or not possible such as during this COVID-19 pandemic and the likely-to-happen future outbreaks.
Phenomenon: Existing literature, as well as anecdotal evidence, suggests that tiered clinical grading systems may display systematic demographic biases. This study aimed to investigate these potential inequities in-depth. Specifically, this study attempted to address the following gaps in the literature: (1) studying grades actually assigned to students (as opposed to self-reported ones), (2) using longitudinal data over an 8-year period, providing stability of data, (3) analyzing three important, potentially confounding covariates, (4) using a comprehensive multivariate statistical design, and (5) investigating not just the main effects of gender and race, but also their potential interaction. Approach: Participants included 1,905 graduates (985 women, 51.7%) who received the Doctor of Medicine degree between 2014 and 2021. Most of the participants were white (n = 1,310, 68.8%) and about one-fifth were nonwhite (n = 397, 20.8%). There were no reported race data for 10.4% (n = 198). To explore potential differential grading, a two-way multivariate analysis of covariance was employed to examine the impact of race and gender on grades in eight required clerkships, adjusting for prior academic performance. Findings: There were two significant main effects, race and gender, but no interaction effect between gender and race. Women received higher grades on average on all eight clerkships, and white students received higher grades on average on four of the eight clerkships (Medicine, Pediatrics, Surgery, Obstetrics/Gynecology). These relationships held even when accounting for prior performance covariates. Insights: These findings provide additional evidence that tiered grading systems may be subject to systematic demographic biases. It is difficult to tease apart the contributions of various factors to the observed differences in gender and race on clerkship grades, and the interactions that produce these biases may be quite complex. The simplest solution to cut through the tangled web of grading biases may be to move away from a tiered grading system altogether.
Education in Doctor of Medicine programs has moved towards an emphasis on clinical competency, with entrustable professional activities providing a framework of learning objectives and outcomes to be assessed within the clinical environment. While the identification and structured definition of objectives and outcomes have evolved, many methods employed to assess clerkship students’ clinical skills remain relatively unchanged. There is a paucity of medical education research applying advanced statistical design and analytic techniques to investigate the validity of clinical skills assessment. One robust statistical method, multitrait-multimethod matrix analysis, can be applied to investigate construct validity across multiple assessment instruments and settings. Four traits were operationalized to represent the construct of critical clinical skills (professionalism, data gathering, data synthesis, and data delivery). The traits were assessed using three methods (direct observations by faculty coaches, clinical workplace-based evaluations, and objective structured clinical examination type clinical practice examinations). The four traits and three methods were intercorrelated for the multitrait-multimethod matrix analysis. The results indicated reliability values in the adequate to good range across the three methods with the majority of the validity coefficients demonstrating statistical significance. The clearest evidence for convergent and divergent validity was with the professionalism trait. The correlations on the same method/different traits analyses indicated substantial method effect; particularly on clinical workplace-based assessments. The multitrait-multimethod matrix approach, currently underutilized in medical education, could be employed to explore validity evidence of complex constructs such as clinical skills. These results can inform faculty development programs to improve the reliability and validity of assessments within the clinical environment.
Abstract Implementation of Competency based medical education (CBME) requires an organized and structured set of interrelated competencies known as a competency framework. Integration of competencies across residency educational programmes and meaningful competency-based clinical supervision is found to be lacking. Study conducted at Aga Khan University tested a five-dimensional model which can be used for competency based clinical supervision in health professionals at postgraduate medical education level. It investigated various factors, including faculty development through clinical supervisor self-assessment of competencies and resident evaluation to propose a Competency/Outcome-based Model of Clinical Supervision along with its working model.
Purpose: Faculty well-being impacts student learning and is a growing concern among medical schools.1 Burnout contributes negatively to faculty well-being and has been exacerbated by the demands of adapting basic and clinical teaching during the COVID-19 pandemic. Organizational drivers of burnout also include the workplace not meeting the wellness hierarchy of needs,2 work environment/culture, and lack of control over workload. By contrast, greater well-being has many benefits, including for student learning1 and future adoption of inclusive medical practice, about which little is currently known.3 Studies of faculty well-being in the United States focus primarily on clinicians and highlight critical differences in levels of burnout by role, race, and gender.4,5 However, researchers have not fully assessed well-being in all faculty in specific domains actionable through changes in policy and practice. We hypothesize that departmental type of work and faculty disability status may be crucial factors impacting drivers of burnout, faculty well-being and ultimately students’ learning. Method: Inspired by Maslow’s hierarchy of needs, Shapiro and colleagues5 proposed a wellness hierarchy of needs with 5 domains: (1) basics (e.g., adequate personal time, access to food/water), (2) safety, (3) respect, (4) appreciation, and (5) contribution. We developed a comprehensive survey based on Shapiro’s work, adding a domain on work environment, 1 question on perceived control over workload, and 1 question on frequency of burnout symptoms. Most survey items used a 5-point rating scale ranging from strongly disagree to strongly agree. The survey was administered to 3,570 medical school faculty members with 748 responses (approximately 20% response rate), of which 701 were complete. We conducted a 2-way analysis of variance (ANOVA) followed by a post hoc analysis to test for the differences in faculty well-being subdomains due to their self-reported departmental affiliation (basic, surgical, and nonsurgical) and disability status (yes and no). We also fitted a 2-way ordinal regression model with a cumulative link because burnout frequency and control over workload were assessed by only 1 ordinal item each. Results: The 2-way ANOVA showed a statistically significant effect of disability status on having basic well-being needs satisfied, (F(1, 697) = 4.56, P = .03, η2 = 0.01) and work environment (F(1, 690) = 7.74, P < .01, η2 = 0.0), indicating a lower percentage of faculty reporting disability responding that basic needs were met and their environment was favorable for well-being. There was also a statistically significant effect of department type on 3 domains: basics (F(2, 697) = 9.26, P < .01, η2 = 0.03); respect (F(2, 697) = 4.15, P = .02, η2 = 0.01); and work environment (F(2, 690) = 3.63, P = .03, η2 = 0.01). Multiple comparisons revealed that faculty in basic science departments had higher scores on the wellness hierarchy of needs, work environment, and control over workload than faculty in non–basic science departments. Similarly, nonsurgical and basic science departments had statistically significant higher scores than surgical departments in respect and work environment domains. Department type and disability status affected frequency of burnout and perceived control over workload, but were not statistically significant given the chi squares of 2 deviance tests. Descriptively, faculty reporting a disability working in surgical departments had lower medians, particularly for perceived control over workload. There were no statistically significant differences in safety and contribution due to department or disability status. Discussion: Overall, faculty who reported a disability scored significantly lower than other faculty in the basics domain of the wellness hierarchy of needs and the work environment domain. Results indicate a need to engage and explore with faculty with disability interventions in these 2 domains. Given a lower percentage of surgical faculty report favorably in the respect domain of the wellness hierarchy of needs and the work environment domain, interventions should consider addressing those specific areas. Significance: Tailoring well-being efforts aimed at domains shown to affect faculty’s well-being, particularly those who report a disability or in surgical departments, could ultimately impact faculty’s direct contributions to the health care system and their students’ learning.
Purpose To explore validity evidence for the use of entrustable professional activities (EPAs) as an assessment framework in medical education. Method Formative assessments on the 13 Core EPAs for entering residency were collected for 4 cohorts of students over a 9- to 12-month longitudinal integrated clerkship as part of the Education in Pediatrics Across the Continuum pilot at the University of Minnesota Medical School. The students requested assessments from clinical supervisors based on direct observation while engaging in patient care together. Based on each observation, the faculty member rated the student on a 9-point scale corresponding to levels of supervision required. Six EPAs were included in the present analyses. Student ratings were depicted as curves describing their performance over time; regression models were employed to fit the curves. The unit of analyses for the learning curves was observations rather than individual students. Results (1) Frequent assessments on EPAs provided a developmental picture of competence consistent with the negative exponential learning curve theory; (2) This finding was true across a variety of EPAs and across students; and (3) The time to attain the threshold level of performance on the EPA for entrustment varied by student and EPA. Conclusions The results provide validity evidence for an EPA-based program of assessment. Students assessed using multiple observations performing the Core EPAs for entering residency demonstrate classic developmental progression toward the desired level of competence resulting in entrustment decisions. Future work with larger data samples will allow further psychometric analyses of assessment of EPAs.
Background Physician professionalism, including anaesthesiologists and intensive care doctors, should be continuously assessed during training and subsequent clinical practice. Multi-source feedback (MSF) is an assessment system in which healthcare professionals are assessed on several constructs (e.g., communication, professionalism, etc.) by multiple people (medical colleagues, coworkers, patients, self) in their sphere of influence. MSF has gained widespread acceptance for both formative and summative assessment of professionalism for reflecting on how to improve clinical practice. Methods Instrument development and psychometric analysis (feasibility, reliability, construct validity via exploratory factor analysis) for MSF questionnaires in a postgraduate specialty training in Anaesthesiology and intensive care in Italy. Sixty-four residents at the Università del Piemonte Orientale (Italy) Anesthesiology Residency Program. Main outcomes assessed were: development and psychometric testing of 4 questionnaires: self, medical colleague, coworker and patient assessment. Results Overall 605 medical colleague questionnaires (mean of 9.3 ±1.9) and 543 coworker surveys (mean 8.4 ±1.4) were collected providing high mean ratings for all items (> 4.0 /5.0). The self-assessment item mean score ranged from 3.1 to 4.3. Patient questionnaires (n = 308) were returned from 31 residents (40%; mean 9.9 ± 6.2). Three items had high percentages of “unable to assess” (> 15%) in coworker questionnaires. Factor analyses resulted in a two-factor solution: clinical management with leadership and accountability accounting for at least 75% of the total variance for the medical colleague and coworker’s survey with high internal consistency reliability (Cronbach’s α > 0.9). Patient’s questionnaires had a low return rate, a limited exploratory analysis was performed. Conclusions We provide a feasible and reliable Italian language MSF instrument with evidence of construct validity for the self, coworkers and medical colleague. Patient feedback was difficult to collect in our setting.
Empathy is central to the physician–patient relationship, and affects clinical outcomes. There is uncertainty about the stability of empathy in medical students over the course of medical school, as well as differences in empathy between men and women. A panel study design was used to follow first year through fourth year medical students (MS1–4) during the 2018–2019 school year. Empathy was measured using the interpersonal reactivity index (IRI), a self-report scale that separates empathy into a cognitive perspective taking (PT) and affective empathic concern (EC) component. A total of 631 (359 women and 272 men) from 970 students (65% response rate) responded to a baseline survey, and a total of 536 students (300 women and 236 men) from 970 students (55% response rate) responded to surveys throughout the year. At baseline, women had significantly higher EC scores than men (p < 0.0001), with no significant PT difference between men and women (p > 0.05). These differences were stable for all MS cohorts. Women had self-reported higher affective empathy (EC component) than men, while there were no differences in cognitive empathy (PT component). We discuss these data in the context of defining gender vs. sex, socialized gender stereotypes, and implications for future research.
BACKGROUND:How effective have lockdowns been at reducing the covid-19 infection and mortality rates? Lockdowns influence contact among persons within or between populations including restricting travel, closing schools, prohibiting public gatherings, requiring workplace closures, all designed to slow the contagion of the virus. The purpose of the present study was to assess the impact of lockdown measures on the spread of covid-19 and test a theoretical model of the covid-19 pandemic employing structural equation modelling.METHODS:Lockdown variables, population demographics, mortality rates, infection rates, and health were obtained for eight countries: Austria, Belgium, France, Germany, Italy, Netherlands, Spain, and the United Kingdom. The dataset, owid-covid-data.csv, was downloaded on 06/01/2020 from: https://github.com/owid/covid-19-data/tree/master/public/data. Infection spread and mortality data were depicted as logistic growth and analyzed with stepwise multiple regression. The overall structure of the covid-19 data was explored through factor analyses leading to a theoretical model that was tested using latent variable path analysis.RESULTS:Multiple regression indicated that the time from lockdown had a small but significant effect (β = 0.112, p< 0.01) on reducing the number of cases per million. The stringency index produced the most important effect for mortality and infection rates (β = 0.588,β = 0.702, β = 0.518, β = 0.681; p< 0.01). Exploratory and confirmatory analyses resulted in meaningful and cohesive latent variables: 1) Mortality, 2) Infection Spread, 3) Pop Health Risk, and 4) Health Vulnerability (Comparative Fit Index = 0.91; Standardized Root Mean Square Residual = 0.08).DISCUSSION:The stringency index had a large impact on the growth of covid-19 infection and mortality rates as did percentage of population aged over 65, median age, per capita GDP, diabetes prevalence, cardiovascular death rates, and ICU hospital beds per 100K. The overall Latent Variable Path Analysis is theoretically meaningful and coherent with acceptable fit indices as a model of the covid-19 pandemic.
Mixed Methods Research (MMR) has been identified as a third research approach in psychological research mixing both quantitative and qualitative methods. MMR’s foundational basis is “an intuitive way of doing research that is constantly being displayed throughout our everyday lives”. Extensive work by cognitive psychologists has shown that intuition is laden with cognitive heuristics or biases that are misleading in interpreting the world around us. MMR rests on an ideology of transformative, emancipatory, advocacy, social justice, and post-modernist positions. Advocacy and politicizing, however, biases research, making the outcomes untenable, unreliable, and not valid, rather working as tools of indoctrination. This undermines scientific objectivity. These positions are critiqued as non-sustainable in the world of modern science. Nonetheless, the idea of mixing data formats is a good one and researchers should employ all possible data (words, numbers, perceptions, etc.) that help build theory, address research questions, test hypothesis, and evaluate theoretical models. MMR in psychological science can be improved with measurement as a foundation while eschewing intuition and transformative, emancipatory, and advocacy ideologies.
It is a pleasure to write on the “early history” of the Canadian Medical Education Journal on its tenth anniversary. In the editorial for the inaugural issue (March 2010), we wrote that we embarked on this adventure with some trepidation because of the many challenges of starting a new journal.1 These included establishing an editorial board, needing high quality submissions, seeking help from expert peer reviewers, and working hard for manuscript selection, preparation and distribution of our issues. At that time, while there was some interest and support for such an ambitious undertaking, it was hard to mobilize.
Purpose To conduct a study of the validity of the new Medical College Admission Test (MCAT). Method Deidentified data for first- and second-year medical students (185 women, 54.3%; 156 men, 45.7%) who matriculated in 2016 and 2017 to the University of Minnesota Medical School–Twin Cities were included. Of those students, 220 (64.5%) had taken the new MCAT exam and 182 (53.4%) had taken the old MCAT exam (61 [17.9%] had taken both). The authors calculated descriptive statistics and Pearson product moment correlations ( r ) between new and old MCAT section scores. They conducteda regression analysis of MCAT section scores with Step 1 scores and with preclerkship course performance. They also conducted an exploratory factor analysis (principal component analysis with varimax rotation) of MCAT scores, undergraduate grade point average, Step 1 scores, and course performance. Results The new MCAT exam section mean score percentiles ranged from 72 to 78 (mean composite score percentile of 80). The old MCAT exam section mean score percentiles ranged from 84 to 88 (mean composite score percentile of 83). The pattern of correlations among and between new and old MCAT exam section scores (range of r : 0.03–0.67; P < .01) provided evidence of both divergent and convergent validities. Backward multiple regression of new MCAT exam section scores and Step 1 scores resulted in a multiple R of .440; the same analysis with Human Behavior course performance as the dependent variable provided a similar solution with the expected sections of the new MCAT exam (multiple R = .502). The factor analysis resulted in 4 cohesive, theoretically meaningful factors: biomedical knowledge, basic science concepts, cognitive reasoning, and general achievement. Conclusions This study provided empirical evidence of multiple types of validity for the new MCAT exam.