
This study aims to adapt and validate the hyper-independence scale among Indonesian university students. The adaptation process followed the guidelines of the International Test Commission, which included forward translation, expert review, and pilot testing. Data were collected from 200 university students across various study programs. Psychometric evaluation was conducted using the Rasch model to examine reliability, unidimensionality, and item functioning. The results demonstrated high reliability at both the person (0.93) and item (0.96) levels, supported by strong separation indices, satisfactory item fit. In the Rasch model, item fit statistics provide evidence of construct validity, as misfitting items indicate response patterns inconsistent with the expected measurement structure. Most items fell within the acceptable MNSQ range (0.7–1.3), with minor deviations that remained tolerable and did not compromise overall validity. The findings suggest that a proportion of students exhibit moderate to high levels of hyper-independence, which may hinder help-seeking behavior, reduce academic engagement, and contribute to mental health risks. Within Indonesia’s collectivistic cultural context, these tendencies may reflect shifting attitudes toward autonomy in higher education. Overall, the validated instrument can be used to identify hyper-independence tendencies in university settings and supports the development of educational interventions that balance independence with adaptive support utilization.
Social Emotional Competence (SEC) was one of the essential competencies in 21st-century education, playing a crucial role in students’ personal and social development. However, valid and reliable instruments for assessing SEC for high school students in Indonesia were still limited. The instrument used in this study was adapted from the CASEL framework, which encompassed five core competencies: self-awareness, self-management, social awareness, relationship skills, and responsible decision-making. The adaptation process involved contextual and linguistic adjustments to align with the characteristics of Indonesian students. This study aimed to examine the validity and reliability of the adapted SEC assessment instrument. The research subjects consisted of 220 high school students in Demak who responded to a questionnaire. This study was categorized as descriptive quantitative. Construct validity was tested using CFA, while reliability was estimated using Cronbach’s alpha. The construct validity test produced model feasibility indices with CFI = 0.928, TLI = 0.9109, SRMR = 0.0534, and RMSEA = 0.0601. These results indicated that the measurement model demonstrated good feasibility. Based on the CFA analysis, 25 items were declared valid. Of these, 22 items had loading factor values greater than 0.5, while 3 items had loading factor values below 0.5. Despite the lower factor loadings, these 3 items were retained because they represented essential indicators. However, the item statements were revised to improve clarity and better represent the intended construct. The internal consistency reliability test showed a Cronbach’s alpha coefficient of 0.942. Since the coefficient value exceeded 0.70, the instrument could be considered reliable.
An ideal assessment should offer constructive feedback and insights into students' strengths and weaknesses in learning. This study aims to develop an assessment model integrating formative assessment with Differentiated Instruction (DIFAM) to assess learning achievements proportionally. The DIFAM model was developed using the ADDIE development framework. The research sample consisted of 99 students from four high schools in Bandung Regency. Student learning profiles were analyzed using N-Gain and paired sample t-test. Data analysis was conducted with R Studio and JASP software. Data analysis using the N-Gain formula revealed an average improvement in student learning achievements of 25% with the implementation of DIFAM. The formative tests conducted over eight sessions showed that students grasped the material more effectively compared to conventional teaching methods. Feedback from students and teachers indicated that DIFAM facilitated more structured learning and provided constructive feedback, contributing significantly to enhanced student performance. The DIFAM model demonstrates its ability to cater to diverse student needs, achieve significant learning improvements, and has the potential for broader application to ensure more inclusive and equitable learning outcomes.
This study aimed to examine the effectiveness of technology-enhanced learning (TEL) in improving students’ statistical graph interpretation skills through a rigorous Item Response Theory (IRT) analysis. Employing a quasi-experimental pretest–posttest control group design, the research involved 120 undergraduate students from four classes, equally divided into experimental and control groups. The experimental groups received TEL-based instruction featuring interactive graph visualizations and automated feedback, while the control groups followed conventional lectures and exercises over seven sessions. Data were collected using a 60-item multiple-choice test covering bar charts, histograms, boxplots, and scatterplots, which was content-validated by experts and trialed for clarity, yielding high reliability (Cronbach’s α = 0.833). Construct validity was ensured through unidimensionality and invariance testing, confirmed by eigenvalue and DETECT analysis. Data analysis applied IRT to calibrate item parameters discrimination (a), difficulty (b), and guessing (c) and to estimate students’ latent abilities (θ). Model comparison identified the 3PL model as the best fit, capturing both difficulty variation and guessing behavior. Calibration results showed that most items exhibited satisfactory psychometric quality, supporting the robustness of the instrument. Findings revealed that TEL groups achieved a nearly one-logit gain in ability from pretest to posttest, significantly higher than the minimal improvement observed in the control groups, as confirmed by independent t-tests and normalized gain analysis. These results indicate that TEL substantially strengthens students’ ability to interpret statistical graphs while demonstrating the diagnostic value of IRT in evaluating both item quality and learning effectiveness.
The concerns of this study were the literature used in English classrooms and students’ views on how it assists them in developing academically and professionally. Thus, its purpose was to evaluate the importance attached to literature courses in the curriculum of English Education programs. The study employed a sequential explanatory mixed-methods design. Quantitatively, a questionnaire instrument was used, which was completed by 107 students of the English Education Study Program at Universitas Bengkulu, regarding their perceptions of the implementation of literary courses, administered through Google Forms. Qualitatively, unstructured interviews provided more profound insights into the role of literature in students’ language proficiency. The results indicate that the majority of students (86%) are interested in studying literature, and they are highly engaged in their work. Those students felt that their grammar competence, vocabulary, and language skills had increased. Furthermore, 91% specifically claimed that literature enhanced their ability for critical thinking and intellectual enrichment, while 77% derived confidence in engaging with literary texts, thereby fostering further collaboration, empathy, and cultural sensitivity in their day-to-day offline routine. Thirty-five per cent of students were also encouraged to read literature. Yet, students encountered constraints such as insufficient time for studying, linguistic complexity, and exposure to unfamiliar cultural scenes. Students value literary education and the use of literature in preparing for future demands. Pedagogically, the curriculum development literature should be systematically integrated with odd in English Education, particularly instructional routines that value active learning, situated interpretation, and imaginative interaction with texts.
This study aimed to evaluate the effectiveness of a community health nursing internship course for final-year nursing students in comprehensive health centers in Tehran, Iran, using Kirkpatrick’s four-level model. The evaluation was performed in terms of the reaction, learning, behavior and outcomes dimensions based on the Kirkpatrick 4-level model. Fifty-six nursing students and 180 clients were randomly selected. Data were collected through researcher-made questionnaires and checklists and analyzed using descriptive and analytical statistical tests. At Level 1, students’ overall satisfaction averaged 61, with the highest satisfaction in clinical instructor performance (77.53%). At Level 2, the mean self-evaluation learning score was 66.29; the highest learning occurred in vaccination (89.28%) and growth monitoring/supplementary nutrition (69.64%). The overall performance evaluation averaged 70.14, with vaccination scoring the highest (91.07%). Clients reported high satisfaction with the care provided by students (Level 4 mean: 72.99). No significant association was found between students’ demographic characteristics and the first three levels of the model. The internship demonstrated effectiveness at all four Kirkpatrick levels. The findings support the value of structured community health internships and highlight the need for educational authorities to develop a standardized, evidence-based program that addresses the identified strengths and areas for improvement.
The one-parameter logistic (1-PL) model is widely used in Item Response Theory (IRT) to estimate student ability; however, ability-based scoring disregards item difficulty and guessing behavior, which can bias proficiency interpretations. This study evaluates three scoring alternatives derived from IRT: an ability-based conversion, a difficulty-weighted conversion, and a proposed guessing-justice method. Dichotomous responses from 400 students were analyzed using the Rasch (1-PL) model in the R environment with the ltm package. The 1-PL specification was retained to support a parsimonious and interpretable calibration framework consistent with the comparative scoring purpose of the study. Rasch estimation produced item difficulty values ranging from −1.03 to 0.18 and identified 268 unique response patterns. Ability-based scoring yielded only eight score distinctions, demonstrating limited discriminatory capacity. In contrast, the guessing-justice method produced a substantially more differentiated distribution, with approximately 70 percent of patterns consistent with knowledge-based responding and 30 percent indicative of guessing. The findings indicate that scoring models incorporating item difficulty and guessing behaviour provide a more equitable and accurate representation of student proficiency than traditional ability-based conversions. The proposed approach offers a practical and implementable alternative for classroom assessment and can be applied using widely accessible spreadsheet software such as Microsoft Excel.
Inequality is a condition characterized by an unbalanced assessment process. Physical and psychological factors, both those measuring and those being measured, may impact assessment inequality. The purpose of this research was to highlight the potential inequity in the performance assessment of madrasas in underdeveloped areas. A quantitative research design was employed. The data were collected using a questionnaire instrument that had been proven valid and reliable. Path analysis was used to determine both direct and indirect effects. The findings showed that measurement errors related to the instruments used have a direct positive effect on inequality in the performance assessment of madrasas in underdeveloped areas, as well as an indirect effect mediated through teacher quality. One alternative solution to reducing the imbalance in assessing the performance of madrasas in underdeveloped areas can be implemented through policy dimensions, including macro, meso, and micro dimensions.
Misconceptions in biology can prevent students from gaining a deeper understanding of biological concepts. There is a five-tier diagnostic test that can explore the misconceptions experienced by students. This research aims to develop a five-tier diagnostic test that is feasible to use to identify student misconceptions with Rasch model analysis. The research method used was quantitative descriptive, and the sample was 103 people with a purposive sampling technique, which is a purposeful sampling technique, namely, the school that is the research location, experiencing misconceptions related to the concept of cells. Based on the results of the study, it was found that the five-tier diagnostic test developed was very feasible to use as an instrument to identify students' misconceptions on the concept of cells. Each indicator is represented by several items that have been tested for validity, reliability, difficulty level and differentiation using Rasch model analysis with the help of the Winsteps program. Based on the analysis with the Rasch model, out of 36 items that were externally validated, 23 items were obtained that met the eligibility criteria and were declared valid for implementation.
This study aims to examine the fairness of Arabic language assessment instruments used in Muhammadiyah senior high schools by detecting the presence of Differential Item Functioning (DIF) in the Final Semester Summative Test (UAS) for 12th-grade students in the Special Region of Yogyakarta during the 2023/2024 academic year. Using a descriptive quantitative design, the research analyzed student response data from 1,157 participants across 25 schools. Data collection was conducted through documentation of test blueprints, item sheets, answer keys, and student responses. Analysis was performed using the Lord and Generalized Lord methods within the framework of Item Response Theory (IRT), focusing on three demographic variables: gender, study specialization (science vs. social studies), and school region (Yogyakarta City, Sleman, Bantul, and Kulon Progo). The Rasch model was identified as the most optimal model due to its superior fit and fulfillment of key psychometric assumptions, including unidimensionality and parameter invariance. The findings indicate that several items exhibit significant DIF across all examined variables. Eleven items showed gender-based DIF, with a higher number favoring male students. Twenty-three items demonstrated DIF by study specialization, and thirty-seven items displayed DIF based on school region, with students from Yogyakarta City benefiting the most. These results suggest that the test is not fully equitable and highlight the need for item revision to ensure fairness. The study contributes theoretically to the field of educational measurement and practically to the development of fairer evaluation practices in Islamic and language education settings.
This research is related to Item Response Theory (IRT), which is essential for determining the best method for estimating participants' abilities on a test measuring English listening ability. This study aims to (1) determine the characteristics of the test device measuring English listening ability, (2) determine the effect of the length of the test on the stability of the ability estimation using the maximum likelihood (ML) method, (3) determine the effect of test length on the stability of the ability estimation using the Bayes method, and (4) compare the stability of the ability estimate between ML and Bayes. This research is an exploratory descriptive study using a simulation approach. The best model is selected to generate data. The result of the generation is the actual ability (θ) and the participant's response, which is estimated with the maximum likelihood and Bayes, which produces the estimated ability with 10 replications, and is compared with calculating the MSE (mean square error). The method with a smaller MSE is stable and has a better estimation method. The results show that (1) the 2PL model is the best, (2) the length of the test affects the stability of the ability estimation in the ML method and the most stable case when the test contains 46 items, (3) the length of the test affects the stability of the ability estimate in the Bayes method and it is most stable when the test contains 46 items, and (4) the Bayes method is better and more accurate for estimating ability.
This study aims to develop instruments to measure digital literacy and critical thinking skills within the Civics Education course, addressing the challenge of assessing these essential competencies in the context of modern education. As digital literacy and critical thinking become increasingly crucial for active citizenship, there is a lack of comprehensive and reliable tools to evaluate these skills effectively. The research employs a development method using the 4D model, consisting of four phases: Define, Design, Develop, and Disseminate. In the Define phase, the competencies to be assessed were clearly identified. In the Design phase, the instruments were crafted based on specific indicators of digital literacy and critical thinking. The Develop phase involved testing the reliability and validity of the instruments, while the Disseminate phase prepared the instruments for broader use. The critical thinking instrument was found to have excellent internal consistency, with a Cronbach’s Alpha of 0.908. However, certain items exhibited low item-total correlations, indicating that revisions were necessary. This study contributes to filling the gap in Civics Education by providing a reliable and valid tool for evaluating digital literacy and critical thinking, ultimately supporting the enhancement of students' competencies in these crucial areas for active and informed citizenship.
Critical thinking is widely recognized as an essential competency in mathematics education, yet assessments often fail to capture its multidimensional nature. This study applied a Bayesian Cognitive Diagnostic Modeling (G-DINA) approach to identify the mastery profiles of tenth-grade students in Indonesia across four attributes: interpretation, analysis, evaluation, and inference. Data from 60 students revealed that most learners demonstrated partial rather than full mastery, with consistent challenges in evaluative reasoning and inference. These diagnostic profiles provide actionable insights for teachers, enabling more targeted instructional strategies that go beyond total test scores. The findings highlight the potential of Bayesian CDMs to enhance classroom assessment by offering fine-grained evidence of students’ reasoning patterns. This study contributes novelty by being among the first to implement Bayesian cognitive diagnosis in mathematics education within the Indonesian context, bridging methodological innovation with practical implications for teaching and assessment.
Literacy skills in reading and numeracy in Indonesia are classified as low, so the government has made new policies, one of which is the application of questions based on the Minimum Competency Assessment (AKM) in the National Assessment. Observation results from several high schools in Pekanbaru, Riau, have not shown the application of AKM questions to biology learning. This research aimed to produce AKM-based reading and numeracy literacy instruments on high school plant and animal bioprocess materials used in Research and Development (R&D) design, where the subjects were grade XII students at three high schools in Pekanbaru city. Data collection instruments were the test instruments that had been developed. Data were analyzed using Rasch modelling assisted by Winstep software, including Wright map analysis, person capability analysis, item capability analysis, scalogram analysis, and question item analysis. The results show that most students have a logic score below 0.0, meaning that reading and numeracy literacy skills and understanding concepts are still low. Four out of 63 students have inappropriate answer patterns, indicating that students did not work on the questions seriously. Then, the analysis results show 14 fit questions; having a Cronbach alpha value of 0.84 with a very high interpretation, a person reability value of 0.77 with sufficient interpretation, and an item reability value of 0.92 in the very high category; the difficulty level of the question items is in line with the rules of the test instrument development, since there is a spread of difficulty of the question items starting from very difficult, hard, medium, easy, and very easy. Thus, it is concluded that the instrument test can be used to measure reading literacy and numeracy skills based on AKM on plant and animal bioprocess materials.
The necessity of problem-solving skills has become a core competency that university students must possess, particularly through the appropriate and accurate use of the Indonesian language. This study aims to construct a theoretical framework of problem-solving abilities by analyzing the composition of opinion texts in Indonesian language learning. The research employs Polya’s theoretical approach, integrated with recent studies, and utilizes a quantitative methodology through Exploratory Factor Analysis (EFA). This method is used to examine the validity of the theoretical construction of problem-solving skills within the context of writing opinion texts in Indonesian language learning. The problem-solving theory derived from opinion-based learning was developed to produce a valid measurement instrument. The study began with the development of indicators drawn from various studies on problem-solving competencies. The resulting instrument consists of 19 items administered to students from both science and social studies tracks. A total of 298 first-semester students from Central Java participated in this study. The test reliability estimation yields a standardized alpha of 0.71. The findings include: (1) the adequacy of the sample was confirmed with a KMO-MSA value > 0.5, specifically 0.71, and a significance level of 0.001 on the Bartlett’s test; (2) all items were found to measure problem-solving skills, indicated by anti-image correlation values > 0.5; and (3) the study identified four dimensions of problem-solving skills based on opinion text analysis: initial problem identification, problem resolution, taking tangible action, and evaluation of implemented solutions solution reflection.
This study aims to develop, validate, and analyze test items for assessing the understanding of mechanical wave concepts among high school students. The test development process followed the Mardapi instrument development model, which includes: (1) constructing test specifications, (2) writing test items, (3) reviewing test items, (4) piloting the test, and (5) analyzing the items. The developed instrument consists of 12 multiple-choice items, covering three aspects of conceptual understanding: translation, interpretation, and interpolation. Content validity was assessed by three validators, and the results were analyzed using the Aiken V method. The instrument was then administered to 257 high school students in South Sulawesi Province. The results were analyzed using Item Response Theory (IRT) with the Rasch model through the Quest program. Item analysis included item fit estimation, reliability, and item difficulty. The content validity test results indicate that the instrument is valid. All items fit the Rasch model, with a reliability coefficient of 0.95, categorized as high reliability. Item difficulty analysis revealed that 8.3% of items were categorized as easy, 8.3% as difficult, and 83.3% as moderate. Overall, the results indicate that the test instrument is of good quality and can be used to assess high school students’ understanding of mechanical wave concepts.
Higher education institutions in Indonesia are currently facing significant challenges in maintaining public trust, requiring them to uphold autonomy, transparency, accountability, and continuously meet standards for quality assurance and improvement. This research aims to identify and address the unmet quality aspects in education to elevate the accreditation of private universities from C/Good to B/Very Good or even A/Excellent. By integrating the EduQual model with certification requirements from the Board of National Accreditation for Higher Education (Badan Akreditasi Nasional Perguruan Tinggi, BAN-PT), the research is supported by comprehensive studies of university accreditation reports, questionnaires, focus group discussions (FGD), and expert judgment through interviews and consultations. The data analysis employs gap analysis, importance-performance analysis, and quality function deployment. The study identified twenty-one priority improvement points to enhance accreditation, with six key improvements prioritized from the ninety indicators examined based on customer needs analysis. In the accreditation of private universities, the unmet aspects of education quality lie in the dimension of physical facilities, particularly human resources, as well as in the dimension of personal development, especially in the criteria for research and community service. The study recommends prioritizing investments in human resource development and strengthening research and community service initiatives, as these are critical areas where private universities fall short in meeting accreditation standards and fulfilling educational quality expectations.
This study aims to compare the analysis model of the characteristics of reading literacy items with the polytomous item response theory, which uses the Graded Response Model (GRM), Partial Credit Model (PCM), Generalized Partial Credit Model (GPCM), and Nominal Reasons Model (NRM). This research is quantitative research in nature, and secondary data were used from about 1000 test takers’ responses to reading literacy items in the 2018 reading literacy study analyzed with the R program. This model comparison was carried out so that the analysis results obtained were more accurate in representing the level of reading literacy skills in Indonesia. The results show that the GPCM model is the fit model with an AIC value of 23753.89 and a BIC value of 24042.45, and the number of suitable testlets is 7 out of a total of 7 testlets. Based on the relationship between information function scores and SEM, reading literacy items provide higher information when participants’ abilities range between -2.3 and +2.
In the field of English language education, the crucial role of assessment for learning (AfL) requires teachers to possess robust assessment literacy. This study explores AfL literacy among high school English teachers in Yogyakarta, Indonesia by utilizing Alonzo's validated AfL survey that establishes a comprehensive six-factor model, delineating teachers as assessors, pedagogists, student partners, motivators, learners, and stakeholder partners. Exploiting confirmatory factor analysis and examining demographic variations, this quantitative research invited 202 English teachers in the Special Region of Yogyakarta, Indonesia selected purposively based on the geographical service area. Data were collected through an online questionnaire adopting Alonzo's 42-questions AfL and were analyzed quantitatively via Confirmatory Factor Analysis (CFA) with four indices, namely comparative fit index (CFI), Tucker-Lewis index (TLI), root mean square error of approximation (RMSEA), and standard root mean square residual (SRMR).The findings substantiate the efficacy of the six-factor AfL model, underscoring educators' roles extending beyond traditional frameworks. The investigation also introduces a tool featuring detailed performance descriptors, addressing deficiencies, and harmonizing with AfL principles. It deduces that heightened foundational comprehension among English educators cultivates enhancements in AfL literacy and propels the refinement of professional evaluative competencies, thereby enriching the nuanced discourse surrounding AfL within language pedagogy. While the study's scope is confined to a specific geographical area and a limited cohort of participating instructors, it significantly enriches our comprehension of AfL literacy among English pedagogies. This research, therefore, provides a foundation for professional growth initiatives and facilitates enhancements in pedagogical approaches and academic achievement.
Teacher innovative behavior is a key capability for maximizing teacher performance as well as the student learning process, especially when teachers face new changes in their education system such as curriculum reform. However, insufficient attention has been given to understanding the interaction of predictor factors to encourage teacher innovative behavior. This study aims to assess the role of cognitive flexibility, positive affect, and negative affect in predicting teacher innovative behavior. A cross-sectional and quantitative design was used for this study. The data collection procedure used convenience sampling, with questionnaires distributed online via social media through several teacher communities. Three instruments were used: the Teacher Innovative Behavior Scale, the Cognitive Flexibility Inventory, and the Positive Affect and Negative Affect Schedule. Data were collected from 322 teachers from three educational levels. Descriptive analysis, correlation analysis, and hierarchical multiple regression analysis were conducted. The result showed that cognitive flexibility and positive affect positively predict teacher innovative behavior. On the other hand, negative affect negatively predicts teacher innovative behavior. Regarding the model, the result indicates that cognitive flexibility plays a more crucial role in predicting teacher innovative behavior, explaining 28,1% of the variance in the model. Researchers and policymakers could use the outcome to create future research, policies, and programs to enhance the capabilities of teachers to perform innovative behavior, especially during the educational system’s changes.