
Artificial Intelligence (AI) literacy is crucial for informed and responsible engagement with AI technologies. While various questionnaires exist to assess AI literacy, they differ in how they conceptualize it. A key component-users' ability to interact effectively with AI systems-is often overlooked or inconsistently addressed. This systematic literature review analyzes existing AI literacy questionnaires using the ABCE framework, which defines AI literacy across four dimensions: affective, behavioral, cognitive, and ethical. The review reveals a predominant focus on cognitive aspects, with affective and ethical dimensions receiving moderate attention, while behavioral aspects-critical for meaningful AI interaction-are frequently underrepresented. This imbalance points to a theoretical gap in current AI literacy assessments. Addressing this gap is essential for developing more comprehensive tools that reflect the full scope of AI literacy. The study's findings contribute to ongoing discussions in AI education and offer guidance for designing more balanced evaluation instruments that better support individuals' ability to navigate and critically engage with AI in diverse settings.
This paper analyzes the impact of recent curricular reforms in Austria aimed at improving students’ computational thinking and problem-solving skills. Using data from the Bebras Challenge, we compare identical tasks from 2016 (N = 11,186) and 2023 (N = 1,505) across grades 3 to 12, focusing on the introduction of mandatory Basic Digital Education and cross-disciplinary programs combining informatics, media literacy, and digital education. Results show a performance decline in most grades, except for grades 3 and 6, with sixth graders showing a 2.2
Large Language Models (LLMs) are emerging as promising tools for generating content, such as educational content for computer science education. When using LLMs, it is necessary to evaluate their outputs to ensure that the quality meets the required standards. This study explores the performance of two LLMs in evaluating the quality of existing Bebras tasks by comparing their evaluations to those made by human experts from the Bebras community. The evaluation is based on predefined criteria such as relevance to informatics, clarity, and learning experience. The study analyzes three tasks of varying difficulty levels aimed at 12–14-year-olds and examines differences in evaluations based on the roles of human experts (teachers, researchers, and organizers) and the corresponding roles prompted to the LLMs. The results indicate varying degrees of alignment between LLMs and human experts, with LLMs struggling to evaluate the learning experience and the time needed to solve the tasks. Additionally, an interesting phenomenon was observed among human expert participants: there was a lack of consensus in estimating the difficulty levels of the tasks. This highlights the inherent subjectivity in difficulty assessment and poses challenges for both human evaluators and AI models. This approach provides insights into the potential of LLMs to support the Bebras community in evaluating tasks and potentially generating new Bebras tasks in the future.
Various forms of unplugged computer science activities have been proposed and employed at the student level in the past. However, there have been few studies of its effectiveness at the teacher level. Active Computational Thinking (CT) Games are unplugged CT activities that encourage physical movement. In school systems where computer science is not part of the curriculum, questions arise regarding how well teachers can learn the fundamentals of CT from Active CT Games, and how effective an Active CT Games approach could be in the classroom. In this paper, we present our findings from Active CT Games workshops with 60 pre-service primary school teachers (41 responses) who had no prior knowledge of CT. We adopted the Approximation of Practice method to evaluate and improve CT lesson plans during the workshops. Our findings show that, among teachers with no CT experience, and who already appreciate the value of learning through play, Active CT Games were positively received. Survey responses indicated that active approaches to teaching CT are appreciated for making learning fun and engaging, while also fostering problem-solving, communication, and critical thinking skills. The Approximation of Practice method proved to be effective, although improvements were suggested. We discuss teacher concerns such as the value of lesson plan scaffolding, and significant setup and implementation times. Overall, these teachers, who had no prior experience in CT, reported high levels of confidence and enthusiasm in employing Active CT Games in the classroom after completion of the workshops.
Programming misconceptions, the incomplete or erroneous understanding of concepts and behaviors of computing agents, are a common and well documented hurdle to novices. A typical manifestation of a misconception is the erroneous application of constructs, e.g. iteration, in programming artifacts. Mere observation of these artifacts may only allow instructors to conjecture the presence of specific misconceptions. Confirmation thereof may be sought by directly asking students about their reasoning. This approach, however, does not scale well to larger populations of students or online learning situations. This study explores the possibility to identify misconceptions indirectly by using program-tracing tasks. Tracing is the ability to read, understand and symbolically execute a computer program. Our approach uses tracing questions where students predict the outputs of programs from multiple-choice answers carefully designed to reflect selected misconceptions. Specifically, we focus on misconceptions about iteration: a topic often reported as complex for complete novices. We evaluated our approach in a pre-test, post-test pilot study involving 35 fifth-grade students. The participants received an intervention based on block-based programming languages and exercises in the style of Code.org. We sought evidence of the presence of 9 predefined misconceptions about iteration at both the pre-test and post-test, identifying the changes that occurred after the intervention. Our results show that tracing evaluations may be useful to instructors as indicators of strengths and weaknesses of groups of learners that may help steer their learning activities.
Despite the growing focus on social issues in CE and their incorporation into curricula worldwide, little is known about teachers’ beliefs and self-efficacy in teaching them. We hypothesis teachers’ self-efficacy regarding teaching social issues to be lower than their overall confidence in teaching computing. To test this, we developed a survey instrument on teachers’ beliefs, perceived self-efficacy and implemented practices regarding social issues in CE in Germany. We adapted two instruments from computing and science education to assess their applicability in this context. The adapted version shows good reliability (Cronbach’s α > .82 ). Furthermore, we hypothesise that social issues in CE is a multi-dimensional area and that teachers’ self-efficacy may vary across sub-areas. To test this, we operationalised the domain via a qualitative content analysis of 32 secondary German CE curricula and identified 14 representative competence statements. Teachers are asked about their beliefs, self-efficacy and practices for each statement. Confirmatory factor analysis of an expert-derived six-dimensional model shows a better fit than a one-dimensional model. In this paper, we present the development process of the instrument, describe its structure in detail, and discuss its quality based on the current sample size (N = 809).
The 2021 revision of the national computing curriculum in Czechia reaffirmed the importance of teaching fundamental computing principles, including principles of computer hardware, at the ISCED level 2 (ages ∼ 11–15). Although these topics were present in the previous Czech curriculum, schools continue to face a shortage of structured, constructivist and evidence-based teaching materials that actively engage students in meaningful learning. To address this need, we developed three model lessons designed to help students grasp principles of core hardware concepts by building on their existing preconceptions. These lessons follow the constructivist Evocation – Realisation of Meaning – Reflection (ERR) framework, encouraging active exploration and reflection. The lessons were implemented and refined through design-based research in six schools across seven classes, involving approximately 160 students. Pre-post testing revealed large immediate improvements following the intervention ( n=45 , Cohen’s d = 0.79 ). This paper presents model lessons and impact of these lessons, contributing to the advancement of constructivist approaches in computing education at the lower secondary level.
In the school year 2023/24, the German federal state of Bavaria started teaching artificial intelligence (AI) as a new compulsory subject in the 11th grade of grammar school computer science education. The curriculum focuses on computer science related topics of machine learning and addresses the assessment of the opportunities and risks of artificial intelligence for individuals and society. The introduction of this new subject was prepared with a professional development course for all computer science teachers at Bavarian universities. We report on the structure of the professional development course at the authors' universities and on a study, which recorded the development of teachers' attitudes towards AI using the "General Attitudes towards Artificial Intelligence Scale" (GAAIS). As part of the four questionnaires of our longitudinal study, we also measured the development of the teachers' self-assessment of their competences with regard to the content to be taught in the future. As is desirable for a teacher training, teachers' self-assessment of AI competences develops positively over the course of the training. Teachers' attitudes towards artificial intelligence were only influenced by the training in the short term, but remained unchanged at the end of the three training days.
Over the last years, the situation of informatics in schools was significantly strengthened in Germany. Furthermore, the scientific discipline of informatics and its influence on our daily lives are constantly evolving. In light of this, in 2025 the German Informatics Society published new standards for lower secondary education in informatics in Germany, replacing the first version published in 2008. In this country report, we describe the development process and the resulting revision of the standards. In particular, we present overarching themes and illustrate changes to the preceding document. This way, we demonstrate how the developments and trends in informatics (education) were addressed within the German context, contributing to the international discourse on the curricular evolution of informatics in schools.
The topic of informatics has been adopted more systematically in many curricula around the world to prepare school students for the ever-changing digital world. Additionally, the annual Bebras Challenge on Informatics and Computational Thinking was created in Lithuania in 2004. Since then, the Challenge has grown into an international phenomenon, and in 2021, an international research consortium BeLLE was created by adopting the digital learning environment ViLLE for the Challenge. Prior to the Challenge, the tasks are assigned difficulty ratings (easy, medium, hard) by the Bebras Community and the organizers in each country. The Community difficulty ratings guide the organizers and the country difficulty ratings affect the points given or deducted depending on the students' answers. It is therefore vital to get the difficulty ratings as accurate as possible. This study investigates the accuracy of these difficulty ratings by comparing them with student performance, with a particular focus on determining whether the Bebras Community or the country ratings are more accurate. Data from two years (2022 and 2023) of the Challenge is used, including 99 tasks from more than 200,000 primary and secondary school students in eleven countries that are part of the BeLLE Consortium. Item Response Theory (IRT) results show that the Challenge is generally quite difficult and there is a lot of overlap between the Community difficulty categories with only the easy category differing from the other two categories statistically. Concerning the country difficulty ratings, all three difficulty categories differ from each other, and the results from a regression model indicate that the country difficulty ratings explain the IRT difficulty estimates more accurately which implies that the organizers in each country should evaluate the difficulty of the tasks again themselves rather than use the Community ratings as they are.
Beginners in programming solve tasks by creating programs which correctness could be evaluated by an automatic system. To do so, it must recognize all correct solutions and be able to decide whether the pupil's solution corresponds to them. This paper introduces an alternative method of assessing the gradation of programming tasks-Variance of Program Builds count (VPB). This is based on the fact that building a program takes the form of its successive iteration. In that case, the number of times such a program was built by the solver can be monitored. The variance of the program builds count can be considered as a criterion of the difficulty of the task. This variance is the highest in the age group for which the task is most suitable. If a series of tasks has a slow gradation in difficulty, all the tasks should be most suitable for the same age group. If the gradation is faster, each task should be most suitable for slightly older pupils than the preceding task. The VPB method was applied to a series of graded programming tasks in order to ascertain whether these tasks satisfied the requirement of increasing difficulty. The tasks were completed by 24,806 pupils from grades 4 to 9. The method identified an issue with the gradation of the series. This was in line with the opinions of two independent experts. Using the VPB method can be regarded as beneficial for the refinement of a series of programming tasks, with the objective of optimising its gradation.
The rapid digitalization of society and the economy necessitates a corresponding transformation in education. In response to this challenge, the pilot school subject "Digital World" was introduced in Hesse to provide grade 5 and 6 students with foundational competencies in digital technologies and to foster their interest in computing-related topics. This study investigates gender-specific differences in the development of self-efficacy and subject interest over the course of a school year. Using a pre-post design, data from over 1,200 students were collected and analyzed statistically. The results indicate a significant increase in self-efficacy, particularly among girls, while subject interest declines in both genders over time. Despite these developments, gender itself does not emerge as a dominant factor in explaining changes in self-efficacy or subject interest. These findings provide valuable insights into how digital education initiatives can be designed to be more inclusive and gender-sensitive. The observed increase in self-efficacy among girls suggests that structured exposure to digital topics may help mitigate traditional gender gaps in perceived competence. However, the decline in subject interest highlights the need for more engaging pedagogical approaches to sustain long-term motivation for digital subjects. The study's implications extend beyond the "Digital World" subject and contribute to broader discussions on equitable STEM education.
Debugging is a natural part of the programming process. It comes into play as soon as novices make their first mistakes in creating programming artifacts. It is also consistently reported to be a skill that is difficult to learn and teach effectively. Research in Computer Science Education has often focused on breaking down debugging in steps connected by temporal and causal dependencies. We refer to this as decomposition of the debugging process. In this work, we look at debugging from the standpoint of Cognitive Load theory, and break it down into a tree-shaped model of subskills that enable and are prerequisite to one another. Our decomposition of debugging in subskills complements the work done on the debugging process and suggests viable learning trajectories that take into account the Cognitive Load of learners.
Providing individualized support to students during debugging is a huge challenge for teachers in K-12 computing education. In these everyday assessment situations, they often have little time to gather relevant information to diagnose the student's problem and respond with an appropriate intervention. Thus, diagnostic and intervention processes in debugging are essential for teachers. Despite the importance, there is a lack of research on this topic and its possible implications for the classroom. Therefore, this paper aims to provide insights into teachers' diagnostic and intervention processes in debugging. In this qualitative study, we investigate situation-specific aspects teachers consider for diagnosing error situations and interventions they apply in a specific debugging-related situation using video vignettes. To this end, scripted video vignettes depicting a typical classroom debugging situation were presented to experienced teachers, who reported their observations in open-ended questionnaires. The data were then analyzed using qualitative content analysis. The results show a wide range of different aspects used in diagnostic processes and in proposed interventions. Furthermore, our results indicate that teachers rarely address motivational and emotional aspects of debugging in their interventions. These findings contribute to a better understanding of teachers' diagnostic and intervention processes and how they can be fostered in teacher education.
The international Bebras challenge (BC) aims to raise the students' interest in computer science (CS) and contribute to their problem-solving and computational thinking (CT) skills development. Through solving short conceptual tasks that are close to real-life context, students obtain the essential literacy needed to be successful in the modern world. Although the initiative has been running for about 20 years, teachers' motivation to engage with students in this challenge remains underexplored. This study employs self-determination theory as a theoretical framework and the work tasks motivation scale for teachers as an instrument. Data collected from 334 teachers across Estonia, Hungary, and Lithuania were analyzed. The results showed teachers' motivation most significantly stemmed from intrinsic motivation and identified regulation. There were no significant differences in all types of motivation between male and female teachers or teachers with different years of teaching experience. However, teachers who were more experienced in engaging in the BC demonstrated higher levels of almost all types of motivation. Mature teachers were more externally motivated than their younger colleagues. Teachers who used the Bebras tasks in their lessons had a higher level of intrinsic motivation and identified regulation than those who did not use these tasks. Teachers teaching various subjects showed different levels of identified regulation. The study provides implications for teachers, school administration, and the BC organization.
In introductory programming courses in higher education, systematic problem-solving consists of three interdependent steps: 1) the specification of the task, 2) the algorithm of the solving program, and 3) the implementation of the algorithm. During the specification phase, numerous decisions are made that help to describe subsequent steps, making specification a crucial stage in the design process, mainly when we use stricter analogous programming approaches, like derivation. However, the formal language used for the specification requires a considerable level of abstraction, so in practice this part of the task solution is often incomplete or incorrect from students. In this article, we aim to explore the challenges associated with the specification step in introductory programming courses in higher education. Additionally, we introduce a tool designed to alleviate these issues and elevate the specification process to the same level of experience as writing algorithms and code during problem solving.
Problem-solving strategies have been investigated in various informatics education contexts. However, no substantial research has yet been conducted on the problem-solving behavior of students in the field of Machine Learning (ML).This study aims to bridge this gap by analyzing the self-directed problem-solving processes employed by students in grades 8 to 10 (n = 93) when developing decision trees as classification models. A digital multi-touch puzzle game and a custom-developed toolchain were utilized to visually capture and subsequently analyze students’ gameplay behaviors using quantitative content analysis techniques.The results of this study indicate that learners within the examined age group predominantly employ exploratory problem-solving strategies in the self-directed construction of decision trees. In contrast, structured approaches are employed much less frequently and demonstrate lower persistence, yet they are significantly more correlated with successful game completion. These findings underscore the necessity of developing learning environments that promote the application and facilitate the persistence of structured problem-solving strategies, enabling learners to engage with the functioning and development of decision trees in a systematic and purposeful manner.
The fast-paced developments in GenAI technology have begun to change school realities. The aim of this study is to identify what competences teachers of technical subjects are relying on when managing AI use in the classroom and to what degree these competences are covered in the DigComp framework. This study used a qualitative research approach utilizing semi-structured interviews. Nine educators from two vocational upper secondary schools in Austria were interviewed and a thematic analysis of the interviews was conducted to identify competences which were then mapped to DigComp2.2. Results show that several of the AI-related competences identified, such as information literacy and privacy can smoothly be mapped to DigComp2.2. However, other competences such as prompting, and the reflected goal-oriented use of GenAI tools appear not to fit in well and hence point to the necessity of adapting DigComp2.2. Furthermore, the importance of interpersonal skills and a well-founded subject knowledge were highlighted as crucial. This research is targeted at curriculum designers, educators, educational researchers, and administrators. It will also speak to teachers who wish to better deal with and prosper from the emergence of GenAI in schools. Future work will address key factors in developing educators' and students' AI literacy.
In the Czech version of the Bebras Challenge, programming tasks in which a programming code is built from blocks are used alongside traditional contest tasks. This paper deals with the comparison of these programming tasks with other informatics tasks in terms of theworkload of their solvers. Programming tasks can strongly attract the contestants' attention. The aim of the paper is to find out to what extent pupils pay attention to programming tasks at the expense of other tasks and how this differs for successful and unsuccessful contestants. We prepared national round tests in which there were typically two programming tasks and ten other tasks in each age category. We conducted quantitative research with a total of 184,949 respondents from the 2023 contest. We discovered that pupils spend more time on programming tasks than on the other tasks. When designing tests that include programming tasks, it is necessary to take this into account, and thus either limit the number of tasks, increase the test time, or include less challenging tasks. Our research on programming tasks revealed that, for most age categories, successful pupils ran different versions of their solution fewer times than unsuccessful ones. We also discovered that successful solvers spent more time on the task than unsuccessful solvers. The youngest pupils were the outliers in these findings. Our paper contributes to the understanding of the extent to which it is a good idea to include programming tasks in informatics contests like the Bebras Challenge.
To teach programming effectively, instructors must possess diagnostic skills and the ability to provide individualized interventions. This requires a deep understanding of students' mental models and common misconceptions in programming, along with the capacity to assist students in refining their mental models where necessary. In our approach, we utilize videos depicting secondary school students' programming processes to train pre-service teachers in diagnostic skills. While these videos facilitate initial diagnoses, they alone cannot confirm them. Therefore, a complementary method is essential to teach diagnostic conversations and interventions effectively. To address this need, we developed tasks that aim at these skills by incorporating role-playing of students with misconceptions. Our research focused on evaluating pre-service teachers' engagement with these tasks and their reported outcomes. Our findings reveal that it is hard for them to get into the mind of a student who holds misconceptions. They also report that trying to do so is useful for understanding students' thought processes. The results suggest that role-playing tasks can foster the transition from theoretical knowledge about mental models to the practical ability to simulate their effects. In this way, our tasks contribute to bridging the gap between theory and practice in teacher training.