Objective:The doctor-patient interaction is essential for successfull dental treatment. Although it is possible to consider social skills during student selection, these are rarely taken into account. The described project aims to identify and evaluate the social skills deemed necessary by various stakeholders and to assess whether these skills can be effectively measured using a Situational Judgment Test (SJT). Methods:The project involved conducting interviews with stakeholders (lecturers, students, patients, practicing dentists) to identify relevant social skills. This was followed by a Delphi survey to evaluate the importance of these skills. Additionally, the SJT was examined for its suitability in the context of dental medicine, and various methods for reliably measuring the identified skills were assessed. Results:Dental lecturers and students consider emotional resilience, particularly stress management, to be especially important during dental studies, while patient-related behaviors are of lesser priority - possibly due to the constraints of the academic environment. In contrast, patients and dentists emphasize the importance of helpfulness and caring conduct during treatment. Conclusion:Our research highlights the need to strengthen social skills in dental education. Although the SJT from general medicine is also suitable for dental studies, Multiple Mini Interviews (MMIs) are a more effective method for capturing complex skills, such as behavioral flexibility.
Educational attainment and admission tests have a longstanding history in the selection of medical students and are often used simultaneously in selection processes. Their value in the admission process is most frequently assessed by their ability to predict academic performance in medical school. However, their simultaneous use may overlook an overlap in their predictive validity. The present study aims to assess the predictive validity of both educational attainment and admission tests, as well as their incremental validities. In addition, subtest analyses are conducted to gain a more profound understanding of admission tests’ predictive power. A survey amongst test-takers of the German admission tests was conducted in 2022 and 2023. Self-reported preclinical performance was matched with admission test scores (i.e., TMS and HAM-Nat). Educational attainment was assessed by high-school grade point average (GPA). Based on n = 2113 medical students, hierarchical multiple regression analyses were conducted. Pearson’s correlations were used to assess the relationship of subtests with academic performance. For all analyses, the effects of range restriction were diminished using a multivariate correction formula. TMS and HAM-Nat as well as high-school GPA predicted academic performance separately. However, while both admission tests demonstrate substantial incremental validity over high-school GPA, the reverse is true to a far lesser extent. High-school GPA exhibits only small predictive power whilst controlling for admission test scores. Subtests containing elements of both crystallized and fluid intelligence proved to be of moderate effect size. The findings of this study suggest that both admission tests and high-school GPA are well-suited as selection criteria in the admission process. Given the growing concerns regarding high-school GPA, admission tests emerge as a compelling alternative, particularly because of their stronger predictive power. Within each examined admission test, content-rich subtests containing elements of both crystallized and fluid intelligence demonstrated the strongest association with academic performance in preclinical years, in line with the test-criterion content match hypothesis.
Introduction:Lack of diversity in the healthcare workforce harms patient care and outcomes, driving calls for more inclusive medical education. Physicians from disadvantaged backgrounds often serve underserved areas, and diversity improves cultural competence and trust. However, admissions still favour higher socioeconomic applicants, with barriers like standardized exams limiting access for underrepresented groups. This review examined barriers and enablers to medical school access for underrepresented groups, aiming to inform fairer admissions practices. Methods:A scoping review was conducted using Arksey and O'Malley framework to map literature on barriers and facilitators to medical school access for applicants from migration backgrounds and low socioeconomic status. A comprehensive search of PubMed, ERIC, and Google Scholar (Nov 2024-Apr 2025) included studies in five languages published in the past 10 years. Data were extracted in early 2025 and thematically analyzed using Braun and Clarke's method. Results:Underrepresented groups face structural, institutional, economic, social and psychological barriers to medical school entry. Key challenges included financial hardship, inadequate academic support, lack of social capital, exclusionary institutional practices and psychological factors. However, targeted interventions - such as pipeline and outreach programs emphasizing mentorship and support networks - can help mitigate these barriers. Conclusion:Despite ongoing efforts to widen participation, underrepresented groups continued to face complex, intersecting barriers to medical school admission. Addressing these challenges required more than general policy initiatives - it called for intentional, community-based approaches tailored to students' specific needs. This review highlighted the need for sustained, systemic change alongside targeted support strategies.
Situational Judgement Tests (SJTs) are popular to screen for social skills during undergraduate medical admission as they have been shown to predict relevant study outcomes. Two different types of SJTs can be distinguished: Traditional SJTs, which measure general effective behavior, and construct-driven SJTs which are designed to measure specific constructs. To date, there has been no comparison of the predictive validity of these two types of SJTs in medical admission. With the present research, we examine whether the HAM-SJT, a traditional SJT, and the CD-SJT, a construct-driven SJT with an agentic and a communal scale, administered during undergraduate medical admission can predict OSCE (i.e., objective structured clinical examination) performance in a low-stakes (nLS = 159) and a high-stakes (nHS = 160) sample of medical students. Results showed a moderate positive relation between the communal scale of the CD-SJT and performance in OSCE stations with trained patients in the high-stakes sample (r =.20, p =.009). This SJT had also an incremental value in predicting the OSCE performance above and beyond GPA (i.e., grade point average), a science test (i.e., HAM-Nat), and gender (ß = 0.18, 95
EDITORIAL article Front. Educ., 02 November 2023Sec. Assessment, Testing and Applied Measurement Volume 8 - 2023 | https://doi.org/10.3389/feduc.2023.1308436
Standardized ability tests that are associated with intelligence are often used for student selection. In Germany two different admission procedures to select students for medical studies are used simultaneously; the TMS and the HAM-Nat. Due to this simultaneous use of both a detailed analysis of the construct validity is mandatory. Therefore, the aim of the study is the construct validation of both selection procedures by using data of 4,528 participants ( M age = 20.42, SD = 2.74) who took part in a preparation study under low stakes conditions. This study compares different model specifications within the correlational structure of intelligence factors as well as analysis the g-factor consistency of the admission tests. Results reveal that all subtests are correlated substantially. Furthermore, confirmatory factor analyses demonstrate that both admission tests (and their subtests) are related to g as well as to a further test-specific-factor. Therefore, from a psychometric point of view, the simultaneous use of both student selection procedures appears to be legitimate.
As a component of many intelligence test batteries, figural matrices tests are an effective way to assess reasoning, which is considered a core ability of intelligence. Traditionally, the sum of correct items is used as a performance indicator (total solution procedure). However, recent advances in the development of computer-based figural matrices tests allow additional indicators to be considered for scoring. In two studies, we focused on the added value of a partial solution procedure employing log file analyses from a computer-based figural matrices test. In the first study (n = 198), we explored the internal validity of this procedure by applying both an exploratory bottom-up approach (using sequence analyses) and a complementary top-down approach (using rule jumps, an indicator taken from relevant studies). Both approaches confirmed that higher scores in the partial solution procedure were associated with higher structuredness in participants' response behavior. In the second study (n = 169), we examined the external validity by correlating the partial solution procedure in addition to the total solution procedure with a Grade Point Average (GPA) criterion. The partial solution procedure showed an advantage over the total solution procedure in predicting GPA, especially at lower ability levels. The implications of the results and their applicability to other tests are discussed.
Admission tests are among the most widespread and effective criteria for student selection in medicine in Germany. As such, the Test for Medical Studies (TMS) and the Hamburg Assessment Test for Medicine, Natural Sciences (HAM-Nat) are two major selection instruments assessing applicants’ discipline-specific knowledge and cognitive abilities. Both are currently administered in a paper-based format and taken by a majority of approximately 40,000 medicine applicants under high-stakes conditions yearly. Computer-based formats have not yet been used in the high-stakes setting, although this format may optimize student selection processes substantially. For an effective transition to computer-based testing, the test formats’ equivalence (i.e., measurement invariance) is an essential prerequisite. The present study examines measurement invariance across test formats for both the TMS and HAM-Nat. Results are derived from a large, representative sample of university applicants in Germany. Measurement invariance was examined via multiple-group confirmatory factor analysis. Analyses demonstrated partial scalar invariance for both admission tests indicating initial evidence of equivalence across test formats. Generalizability of the results is examined, and implications for the transition to computer-based testing are discussed.
Due to their high item difficulties and excellent psychometric properties, construction-based figural matrices tasks are of particular interest when it comes to high-stakes testing. An important prerequisite is that test preparation - which is likely to occur in this context - does not impair test fairness or item properties. The goal of this study was to provide initial evidence concerning the influence of test preparation. We administered test items to a sample of N = 882 participants divided into two groups, but only one group was given information about the rules employed in the test items. The probability of solving the items was significantly higher in the test preparation group than in the control group (M = 0.61, SD = 0.19 vs. M = 0.41, SD = 0.25; t(54) = 3.42, p = .001; d = .92). Nevertheless, a multigroup confirmatory factor analysis, as well as a differential item functioning analysis, indicated no differences between the item properties in the two groups. The results suggest that construction-based figural matrices are suitable in the context of high-stakes testing when all participants are provided with test preparation material so that test fairness is ensured.
Student selection at Hamburg medical school is based on the combination of a natural science knowledge test (HAM-Nat) and pre-university educational attainment.
The German Constitutional Court is currently reviewing whether the actual study admission process in medicine is compatible with the constitutional right of freedom of profession, since applicants without an excellent GPA usually have to wait for seven years. If the admission system is changed, politicians would like to increase the influence of psychosocial criteria on selection as specified by the Masterplan Medizinstudium 2020. What experiences have been made with the actual selection procedures? How could Situational Judgement Tests contribute to the validity of future selection procedures to German medical schools? High school GPA is the best predictor of study performance, but is more and more under discussion due to the lack of comparability between states and schools and the growing number of applicants with top grades. Aptitude and knowledge tests, especially in the natural sciences, show incremental validity in predicting study performance. The measurement of psychosocial competencies with traditional interviews shows rather low reliability and validity. The more reliable multiple mini-interviews are superior in predicting practical study performance. Situational judgement tests (SJTs) used abroad are regarded as reliable and valid; the correlation of a German SJT piloted in Hamburg with the multiple mini-interview is cautiously encouraging. A model proposed by the Medizinischer Fakultatentag and the Bundesvertretung der Medizinstudierenden considers these results. Student selection is proposed to be based on a combination of high school GPA (40%) and a cognitive test (40%) as well as an SJT (10%) and job experience (10%). Furthermore, the faculties still have the option to carry out specific selection procedures.
Das Bundesverfassungsgericht überprüft aktuell, ob das Vergabeverfahren der Medizinstudienplätze mit dem Grundrecht auf freie Berufswahl vereinbar ist, da BewerberInnen ohne sehr gute Abiturnoten meist sieben Jahre warten müssen. Bei einer Umstellung des Zulassungssystems möchte die Politik, dem Masterplan Medizinstudium 2020 folgend, psychosoziale Auswahlkriterien stärker gewichten.
In their article, Kesternich et al. [1] address the important topic of the shortage of general practitioners (GPs) and country doctors, which is considered in the “Masterplan Medizinstudium 2020” (master plan medical studies 2020) in the form of an admission quota for prospective country doctors. The study presents the results of a survey among medical students from Munich (Germany) in the clinical section of their studies. In a multivariate model Kesternich et al. determine the influence of socio-demographic factors, typical parameters considered in student selection procedures, and risk aversion on the intention to become a country doctor or a GP, respectively. Only
OBJECTIVES:Increasing numbers of educational institutions in the medical field choose to replace their conventional admissions interviews with a multiple mini-interview (MMI) format because the latter has superior reliability values and reduces interviewer bias. As the MMI format can be adapted to the conditions of each institution, the question of under which circumstances an MMI is most expedient remains unresolved. This article systematically reviews the existing MMI literature to identify the aspects of MMI design that have impact on the reliability, validity and cost-efficiency of the format.METHODS:Three electronic databases (OVID, PubMed, Web of Science) were searched for any publications in which MMIs and related approaches were discussed. Sixty-six publications were included in the analysis.RESULTS:Forty studies reported reliability values. Generally, raising the number of stations has more impact on reliability than raising the number of raters per station. Other factors with positive influence include the exclusion of stations that are too easy, and the use of normative anchored rating scales or skills-based rater training. Data on criterion-related validities and analyses of dimensionality were found in 31 studies. Irrespective of design differences, the relationship between MMI results and academic measures is small to zero. The McMaster University MMI predicts in-programme and licensing examination performance. Construct validity analyses are mostly exploratory and their results are inconclusive. Seven publications gave information on required resources or provided suggestions on how to save costs. The most relevant cost factors that are additional to those of conventional interviews are the costs of station development and actor payments.CONCLUSIONS:The MMI literature provides useful recommendations for reliable and cost-efficient MMI designs, but some important aspects have not yet been fully explored. More theory-driven research is needed concerning dimensionality and construct validity, the predictive validity of MMIs other than those of McMaster University, the comparison of station types, and a cost-efficient station development process.
Background: Multiple mini-interviews (MMIs) are a valuable tool in medical school selection due to their broad acceptance and promising psychometric properties. With respect to the high expenses associated with this procedure, the discussion about its feasibility should be extended to cost-effectiveness issues.Methods: Following a pilot test of MMIs for medical school admission at Hamburg University in 2009 (HAM-Int), we took several actions to improve reliability and to reduce costs of the subsequent procedure in 2010. For both years, we assessed overall and inter-rater reliabilities based on multilevel analyses. Moreover, we provide a detailed specification of costs, as well as an extrapolation of the interrelation of costs, reliability, and the setup of the procedure.Results: The overall reliability of the initial 2009 HAM-Int procedure with twelve stations and an average of 2.33 raters per station was ICC=0.75. Following the improvement actions, in 2010 the ICC remained stable at 0.76, despite the reduction of the process to nine stations and 2.17 raters per station. Moreover, costs were cut down from $915 to $495 per candidate. With the 2010 modalities, we could have reached an ICC of 0.80 with 16 single rater stations ($570 per candidate).Conclusions: With respect to reliability and cost-efficiency, it is generally worthwhile to invest in scoring, rater training and scenario development. Moreover, it is more beneficial to increase the number of stations instead of raters within stations. However, if we want to achieve more than 80 % reliability, a minor improvement is paid with skyrocketing costs.
Although some recent studies concluded that dexterity is not a reliable predictor of performance in preclinical laboratory courses in dentistry, they could not disprove earlier findings which confirmed the worth of manual dexterity tests in dental admission. We developed a wire bending test (HAM-Man) which was administered during dental freshmen's first week in 2008, 2009, and 2010. The purpose of our study was to evaluate if the HAM-Man is a useful selection criterion additional to the high school grade point average (GPA) in dental admission. Regression analysis revealed that GPA only accounted for a maximum of 9% of students' performance in preclinical laboratory courses, in six out of eight models the explained variance was below 2%. The HAM-Man incrementally explained up to 20.5% of preclinical practical performance over GPA. In line with findings from earlier studies the HAM-Man test of manual dexterity showed satisfactory incremental validity. While GPA has a focus on cognitive abilities, the HAM-Man reflects learning of unfamiliar psychomotor skills, spatial relationships, and dental techniques needed in preclinical laboratory courses. The wire bending test HAM-Man is a valuable additional selection instrument for applicants of dental schools.
Introduction: The present study examines the question whether the selection of dental students should be based solely on average school-leaving grades (GPA) or whether it could be improved by using a subject-specific aptitude test. Methods: The HAM-Nat Natural Sciences Test was piloted with freshmen during their first study week in 2006 and 2007. In 2009 and 2010 it was used in the dental student selection process. The sample size in the regression models varies between 32 and 55 students. Results: Used as a supplement to the German GPA, the HAM-Nat test explained up to 12% of the variance in preclinical examination performance. We confirmed the prognostic validity of GPA reported in earlier studies in some, but not all of the individual preclinical examination results. Conclusion: The HAM-Nat test is a reliable selection tool for dental students. Use of the HAM-Nat yielded a significant improvement in prediction of preclinical academic success in dentistry.
[english] Introduction: The present study examines the question whether the selection of dental students should be based solely on average school-leaving grades (GPA) or whether it could be improved by using a subject-specific aptitude test.Methods: The HAM-Nat Natural Sciences Test was piloted with freshmen during their first study week in 2006 and 2007. In 2009 and 2010 it was used in the dental student selection process. The sample size in the regression models varies between 32 and 55 students. Results: Used as a supplement to the German GPA, the HAM-Nat test explained up to 12% of the variance in preclinical examination performance. We confirmed the prognostic validity of GPA reported in earlier studies in some, but not all of the individual preclinical examination results. Conclusion: The HAM-Nat test is a reliable selection tool for dental students. Use of the HAM-Nat yielded a significant improvement in prediction of preclinical academic success in dentistry.[german] Einleitung: In der vorliegenden Untersuchung wird der Frage nachgegangen, ob die Auswahl der Studierenden in der Zahnmedizin alleine durch die Abiturdurchschnittsnote erfolgen sollte oder ob sie durch den Einsatz eines fachspezifischen Studierfähigkeitstest verbessert werden kann. Methoden: Der Naturwissenschaftstest HAM-Nat wurde in den Jahren 2006 und 2007 in der Erstsemesterwoche an den Studienanfängerinnen und -anfängern* erprobt sowie 2009 und 2010 im Auswahlverfahren eingesetzt. Die Stichprobengrößen der Regressionsmodelle variieren in allen Jahrgängen zwischen 32 und 55 Teilnehmern. Ergebnisse: Der HAM-Nat erklärte zusätzlich zur Abiturdurchschnittsnote bis zu 12 % der Leistungsvarianz in den vorklinischen Prüfungsleistungen. Die in anderen Studien gefundene prognostische Güte der Abiturdurchschnittsnote konnte für einige, aber nicht für alle Einzelprüfungen bestätigt werden. Schlussfolgerung: Der HAM-Nat erwies sich als zuverlässiges Auswahlinstrument in der Zahnmedizin. Durch den Einsatz des HAM-Nat wird die Vorhersage des vorklinischen, akademischen Studienerfolgs in der Zahnmedizin deutlich verbessert.
Aims: Tests with natural-scientific content are predictive of the success in the first semesters of medical studies. Some universities in the German speaking countries use the ‘Test for medical studies’ (TMS) for student selection. One of its test modules, namely “medical and scientific comprehension”, measures the ability for deductive reasoning. In contrast, the Hamburg Assessment Test for Medicine, Natural Sciences (HAM-Nat) evaluates knowledge in natural sciences. In this study the predictive power of the HAM-Nat test will be compared to that of the NatDenk test, which is similar to the TMS module “medical and scientific comprehension” in content and structure. Methods: 162 medical school beginners volunteered to complete either the HAM-Nat (N=77) or the NatDenk test (N=85) in 2007. Until spring 2011, 84.2% of these successfully completed the first part of the medical state examination in Hamburg. Via different logistic regression models we tested the predictive power of high school grade point average (GPA or “Abiturnote”) and the test results (HAM-Nat and NatDenk) with regard to the study success criterion “first part of the medical state examination passed successfully up to the end of the 7th semester” (Success7Sem). The Odds Ratios (OR) for study success are reported. Results: For both test groups a significant correlation existed between test results and study success (HAM-Nat: OR=2.07; NatDenk: OR=2.58). If both admission criteria are estimated in one model, the main effects (GPA: OR=2.45; test: OR=2.32) and their interaction effect (OR=1.80) are significant in the HAM-Nat test group, whereas in the NatDenk test group only the test result (OR=2.21) significantly contributes to the variance explained. Conclusions: On their own both HAM-Nat and NatDenk have predictive power for study success, but only the HAM-Nat explains additional variance if combined with GPA. The selection according to HAM-Nat and GPA has under the current circumstances of medical school selection (many good applicants and only a limited number of available spaces) the highest predictive power of all models.