ABSTRACT The present study aimed to investigate (1) the development of probabilistic reasoning before formal instruction, measured yearly from 3rd grade preschool to 3rd grade elementary school and (2) the role of early numerical abilities measured in 2nd year preschool and vocabulary knowledge, measured at the beginning of 3rd year preschool for probabilistic reasoning. For the present study, 292 children (150 girls) were longitudinally followed from preschool ( M age = 4y10m) until the 3rd grade of elementary school. Probabilistic reasoning abilities were assessed by evaluating children's performance on three tasks: (a) distinguish certain from uncertain events, (b) compare probabilities, and (c) create equal probabilities. Children grew steadily regarding all probabilistic reasoning tasks (i.e., distinguish certain from uncertain events, p < 0.001, η 2 G = 0.31; compare probabilities, p < 0.001, η 2 G = 0.42; create equal probabilities, p < 0.001, η 2 G = 0.14), but the growth occurred later on the third task. Furthermore, early numerical abilities predicted children's ability to compare probabilities, and early vocabulary knowledge predicted the performance across all three probabilistic reasoning tasks and the growth in the ability to create equal probabilities.
Regression-based studies concluded that children’s general language abilities, mathematical language abilities, and mathematical abilities are positively related to each other. However, additional approaches are needed to better understand these relationships. A k-means cluster analysis based on 692 first, second, and third grader’s general language abilities, mathematical language abilities, and their mathematical abilities identified four distinct profiles of children. In a second step, group membership in these clusters in terms of maternal education level and home language use was examined. In Cluster 1, children attained a low performance on all three abilities, had low maternal education levels, and often spoke a different language than the instructional language. In Cluster 4, children attained a high performance on all three abilities, had high maternal education levels, and often only spoke the instructional language at home. Clusters 2 and 3 are situated between these extremes. These clusters exhibit similar mathematical language abilities, both representing intermediate performance levels. Children in Cluster 2 had intermediate general language abilities and low mathematical abilities. Children in Cluster 3 had low general language abilities and high mathematical abilities. Children in both clusters differed in their home language use but showed comparable maternal educational levels (i.e., medium to highly educated mothers). Our results indicate that the previously reported positive associations between these three abilities are more nuanced. In addition, children with lower maternal education levels (and especially those growing up with a different home language as well) may require more targeted support.
This paper is the second in a two-part series that investigates students’ reasoning on astronomical phenomena. In this work, we examined how (if at all) students rely on spatial scale assumptions when reasoning about Moon phases. To that end, we conducted N=25 semistructured interviews with last-year high school students. A personalized, physical scale model of the Sun-Earth-Moon system was constructed for every participant individually and placed on the interview table. Students were encouraged to clarify their thoughts by interacting with the scale model or by making a drawing. The scale estimates underpinning the scale model on the table were assessed through an interactive online survey, which was administered several days to weeks prior to the interviews. Only students with severely underestimated values for the relevant sizes and distances were selected for an interview. This selection procedure was motivated by both pragmatic and empirical considerations. We observed a wide variety of explanations for Moon phases, some of which were not found in the existing literature. Several students used their (inaccurate) scale ideas to reason about Moon phases. However, occasionally an interviewed student expressed arguments inconsistent with the scale model on the table. In these cases, scale assumptions seemed to be not taken into account when reasoning about Moon phases. The implications of these findings are discussed.
Spontaneous focusing on numerical order (SFONO) has been suggested as a relevant construct for the development of ordinality knowledge, as children who more often notice numerical order in everyday situations tend to exhibit better ordinality knowledge. However, earlier SFONO measures risked conflating SFONO with the skills needed in them and focused only on numerical sequences with small, consecutive numbers. This study addressed this gap by developing a revised SFONO measure. The construct validity of the measure was examined through three approaches: (1) assessing its ability to replicate individual differences in SFONO, (2) evaluating the influence of various task contexts and numerical sequences on SFONO scores, and (3) confirming its divergent validity from the requisite skills needed to perform the tasks. Fifty-one children (Mage = 5.75 years) completed four SFONO tasks featuring varied contexts and a wider range of numerical sequences. Results indicated that consistent individual differences in SFONO could be observed across diverse situations, providing evidence for the construct validity of its measurement. In addition, the SFONO measure showed divergent validity from the necessary skills, supporting the interpretation that SFONO reflects a distinct construct. Interestingly, SFONO responses appeared more affected by the numerical sequences used in the task than by the task context. Put together, the study highlights the need to carefully consider a wider range of task features when attempting to measure spontaneous mathematical focusing tendencies.
Students often encounter difficulties when learning and processing fractions. In fraction comparison tasks, they tend to rely on a mixture of successful strategies, e.g., fraction magnitude processing or benchmarking, and erroneous strategies, e.g., natural number-based reasoning or gap thinking. Reinhold et al. (2023) used a theory-driven approach to classify students into distinct profiles depending on their strategy choice based on their performance on a comparison task with 24 single-digit fractions. The authors identified single strategy profiles (e.g., typical natural number bias) and composite profiles (e.g., benchmarking or typical bias). The current study aimed to replicate this study (RO1) and extend its design by incorporating multi-digit fractions (i.e., the denominator has at least two digits; RO2) and additional biased comparison strategies, specifically gap thinking (RO3). A set of 101 fraction comparison tasks was administered to 285 fifth- and sixth-grade students in Flanders, Belgium, controlling for benchmarking to 1/2, numerical distance, natural number-based reasoning and gap thinking. A Bayesian classification approach based on students’ performance (accuracy, reaction time, and individual distance effect) replicated the distinct profiles identified by Reinhold et al. (2023), not only for single- (RO1) but also for multi-digit fractions (RO2). Furthermore, we identified additional single strategy and composite profiles (e.g., applying benchmarking where possible, otherwise gap thinking), both based on accuracy and reaction time data (RO3). This study enhances our understanding of individual differences in fraction processing and emphasizes the importance of controlling for gap thinking.
This paper is the first in a two-part series that investigates students’ reasoning on astronomical phenomena. We interviewed N=25 last-year high school students to uncover their reasoning on three observable phenomena, two of which will be discussed in this paper. The first phenomenon was the apparent sizes of the Moon and the Sun (as seen from Earth) and potential causes for their variation. The second phenomenon regarded the event of a solar eclipse and the difference between total and annular eclipses. The interviews were held in light of the students’ understanding of spatial scales in the Earth-Moon-Sun system. Based on their answers on a prior online survey assessing students’ estimates of these scales, a personal to-scale model of the Earth-Moon-Sun system was provided for every student. The students selected for interviews were those who showed a substantial underestimation of these scales. Students were repetitively encouraged to illustrate their reasoning on both phenomena by using the scale model or by providing a drawing. For both discussed astronomical phenomena, we encountered several alternative explanations that—as far as we know—were not previously documented in the research literature. Many of these relied on an inaccurate comprehension of the involved scales. Although for some students these explanations were in line with the scale model on the interview table representing their substantial underestimations, others did not seem to take spatial scales into account when making suggestions on the causes for these astronomical phenomena. The implications of these findings are discussed.
Many students persistently misinterpret histograms. This calls for closer inspection of students’ strategies when interpreting histograms and case-value plots (which look similar but are different). Using students’ gaze data, we ask: How and how well do upper secondary pre-university school students estimate and compare arithmetic means of histograms and case-value plots? We designed four item types: two requiring mean estimation and two requiring means comparison. Analysis of gaze data of 50 students (15–19 years old) solving these items was triangulated with data from cued recall. We found five strategies. Two hypothesized most common strategies for estimating means were confirmed: a strategy associated with horizontal gazes and a strategy associated with vertical gazes. A third, new, count-and-compute strategy was found. Two more strategies emerged for comparing means that take specific features of the distribution into account. In about half of the histogram tasks, students used correct strategies. Surprisingly, when comparing two case-value plots, some students used distribution features that are only relevant for histograms, such as symmetry. As several incorrect strategies related to how and where the data and the distribution of these data are depicted in histograms, future interventions should aim at supporting students in understanding these concepts in histograms. A methodological advantage of eye-tracking data collection is that it reveals more details about students’ problem-solving processes than thinking-aloud protocols. We speculate that spatial gaze data can be re-used to substantiate ideas about the sensorimotor origin of learning mathematics.
Fraction thinking poses a challenge for students, and several flawed comparison strategies have been identified: whole-number-strategy (choosing larger numerals), reverse-strategy (choosing smaller numerals), and gap-strategy (choosing the smaller difference between numerator and denominator). The prevalence of these strategies among college students is unknown. Here, we used cluster analysis to identify strategy use among 90 college students. Three cognitive factors were also assessed: general math achievement, inhibitory control, and working memory. The results revealed three clusters: Whole-Number-Strategy (14%), who used whole-number-strategy consistently; Partial-Reverse-Strategy (30%), who used reverse-strategy for more challenging fractions; and Gap-Tendency (56%), who performed well except when gap strategy fails. For cognitive factors, Gap-Tendency performed better than Whole-Number-Strategy, but did not differ from Partial-Reverse-Strategy, on any measure. These findings extended prior research on strategy choices to show that the majority of college students have overcome whole bias, but not yet achieved a complete understanding of fraction magnitude.
In the boxplot, the box always represents – regardless of its area – the middle half of the data and thus a measure of variability (interquartile range). However, when students first learn about boxplots, they are usual already familiar with other forms of statistical representations (e.g., bar or circle graphs) in which a larger area represents a higher frequency of observations. If students erroneously apply this well-established area-represents-frequency schema to boxplots, it results in a systematic error which we describe as the consequence of an incomplete conceptual change. We empirically validated difficulty-generating characteristics that allow the differentiation between item types with varying complexity (item level) and aimed to identify profiles (person level) that differ depending on which schema was used in which item type. For this purpose, we conducted two cross-sectional studies with N = 100 university students (study 1) and N = 297 participants who finished secondary school or higher (study 2) and used generalized linear mixed models (item level) and k-means clustering with predefined cluster centers (person level) to test our hypotheses. We could replicate the systematic error that was described in previous research and found new difficulty-generating characteristics in boxplot items. Our results support the notion of different profiles potentially emerging based on varying degrees of conceptual change. From an instructional perspective, information about individual progress in conceptual change could be considered for tailoring individualized interventions.
Estimating astronomical scales requires multiple complex mental processes, such as spatial thinking and interpreting large numbers. As such, it is a nontrivial question how these estimates can be most efficiently assessed. There is reason to believe that results from previous studies probing astronomical scale estimates are possibly susceptible to effects inherent in the questioning methods used. In this study, an interactive online survey was constructed and administered to 201 students in their last year of high school. To probe their estimates of spatial scales in the Solar neighborhood, we formulated five questions, two of those probing estimates on the relative sizes of astronomical bodies and three on the relative distances between those bodies. These questions were formulated in two different ways, and the effect of these formulations was studied. In one formulation, students were asked to numerically compare the magnitude of two sizes or distances, while in the other, these estimates were made in terms of the travel time of an imaginary spacecraft. After every answer, students were confronted with a customized visualization, which they could either agree with and move on to the next question or disagree and reconsider their previous answer until they agreed. Studying the effect of these visualizations on the students’ answers was another objective of this work. Most students had difficulties estimating both the relative sizes of the considered celestial bodies and the distances between them. The range of estimates covered many orders of magnitude for all questions, and for the distance-related questions, there was a clear trend of underestimation. We found a significant impact of question formulation on the magnitudes of student estimates. However, we found no indication that one question formulation led to more reliable results than the other. The effect of the visualizations was smaller than anticipated but noticeably larger for the size-related questions than for the distance-related questions. Self-assessments of certainty were made by the students after every answer, and those were found not to correlate with the accuracy of the answers. The implications of these findings are discussed.
Already in the early grades of primary school children develop notions of advanced mathematical concepts such as patterning, proportional reasoning, and probabilistic reasoning. However, there are large differences among children. We examined to which extent these differences could be explained by children's general language abilities, mathematical language abilities, and their home environment (i.e., maternal education level and home language). Data were collected in 717 first, second, and third graders who all engaged in a general language task, an advanced mathematical language task (addressing the mathematical language present in the domains of patterning, proportionality, and probability), and an advanced mathematical abilities task (addressing their reasoning in these domains). Path analysis revealed that both general language and advanced mathematical language abilities contributed to children's advanced mathematical abilities, although advanced mathematical language abilities were more impactful than general language abilities. Children's advanced mathematical language abilities partly mediated the relationship between general language abilities and advanced mathematical abilities. Advanced mathematical language abilities were in turn influenced by maternal education level and general language abilities. More precisely, children with more highly educated mothers and children with better general language abilities tend to have a better understanding of advanced mathematical language. Children who only spoke the instructional language at home did not perform better on the advanced mathematical language task than children with a different home language.
Boxplots are widely used in descriptive statistics but they are challenging to learn. Particularly the interpretation of the box area often leads to a systematic error, the area bias, when students assume a proportional relation between area and frequency, while the box area represents the interquartile range and thus a measure of variability. By analyzing eye-tracking data, research was able to identify systematic differences in gaze patterns depending on the assumed solution strategy that was used by an individual. While previous studies suggested—based on initial results—that extrafoveal perception and simultaneous processing might occur when boxplots are displayed in close spatial proximity, this aspect has not been the focus of systematic investigation so far. In the present study, we therefore test whether the application of a schema leads to fewer transitions between and fewer, but on average longer fixations on schema-relevant areas when the boxplots are in closer proximity. N= 27 university students participated in the study. Eye movement data were collected and evaluated in a generalized linear mixed model. Differences in line with the hypotheses regarding the number of transitions and the number of fixations were found for some but not all schemas. The findings provide initial indications that some, but not all, boxplot components may be perceived extrafoveally and processed simultaneously.
In the context of the European Erasmus+ project Teaching ASTronomy at the Educational level (TASTE), we investigated the extent to which a learning module at school and a set of activities during a planetarium visit help students to gain insight in the Apparent Motion of the Sun and Stars. Therefore, we have set up a two treatment study with a pretest posttest design. In the four participating countries (Belgium, Germany, Greece and Italy), secondary school students studied the concept of the celestial globe at school using newly designed learning materials. By using a latent class analysis, we identified different classes of student answers on the AMoSS test. We show how students evolve from one class to another between pretest and posttest. Overall the results of the pretest and posttest show that a good understanding of the different aspects of the apparent motion of celestial bodies is difficult to achieve.
The progress of Large Language Models (LLMs) like ChatGPT raises the question of how they can be integrated into education. One hope is that they can support mathematics learning, including word-problem solving. Since LLMs can handle textual input with ease, they appear well-suited for solving mathematical word problems. Yet their real competence, whether they can make sense of the real-world context, and the implications for classrooms remain unclear. We conducted a scoping review from a mathematics-education perspective, including three parts: a technical overview, a systematic review of word problems used in research, and a state-of-the-art empirical evaluation of LLMs on mathematical word problems. First, in the technical overview, we contrast the conceptualization of word problems and their solution processes between LLMs and students. In computer-science research this is typically labeled mathematical reasoning, a term that does not align with usage in mathematics education. Second, our literature review of 213 studies shows that the most popular word-problem corpora are dominated by s-problems, which do not require a consideration of realities of their real-world context. Finally, our evaluation of GPT-3.5-turbo, GPT-4o-mini, GPT-4.1, and o3 on 287 word problems shows that most recent LLMs solve these s-problems with near-perfect accuracy, including a perfect score on 20 problems from PISA. LLMs still showed weaknesses in tackling problems where the real-world context is problematic or non-sensical. In sum, we argue based on all three aspects that LLMs have mastered a superficial solution process but do not make sense of word problems, which potentially limits their value as instructional tools in mathematics classrooms.
Rational numbers, such as fractions and decimals, are harder to understand than natural numbers. Moreover, individuals struggle with fractions more than with decimals. The present study sought to disentangle the extent to which two potential sources of difficulty affect secondary-school students’ numerical magnitude understanding: number type (natural vs. rational) and structure of the notation system (place-value-based vs. non-place-value-based). To do so, a 2 (number type) × 2 (structure of the notation system) within-subjects design was created in which 61 secondary-school students estimated the position of four notations on a number line: natural numbers (e.g., 214 on a 0–1000 number line), decimals (e.g., 0.214 on a 0–1 number line), fractions (e.g., 3/14 on a 0–1 number line), and separated fractions (3 on a 0–14 number line). In addition to response times and error rates, eye tracking captured students’ on-line solution process. Students had slower response times and higher error rates for fractions than the other notations. Eye tracking revealed that participants encoded fractions longer than the other notations. Also, the structure of the notation system influenced participants’ eye movement behavior in the endpoint of the number line more than number type. Overall, our findings suggest that when a notation contains both sources of difficulty (i.e., rational and non-place-value-based, like fractions), this contributes to a worse understanding of its numerical magnitude than when it contains only one (i.e., natural but non-place-value-based, like separated fractions, or place-value-based but rational, like decimals) or neither (i.e., natural and place-value-based, like natural numbers) of these sources of difficulty.
We present two studies to investigate the extent to which attending a planetarium presentation increases secondary school students ' understanding of the apparent motion of the Sun and stars. In the first study, we used the Apparent Motion of Sun and Stars (AMoSS) test in a pretest/post-test/retention test setting to measure learning gains and improved insight of 404 students (16- to 17 -year -olds) after attending a classical planetarium presentation at the Brussels Planetarium. The AMoSS test is a questionnaire on the daily and yearly apparent motion and the observer 's position. It consists of six multiple-choice questions about the Sun and six similar multiple-choice questions about the stars. We asked the students to explain their choices. The learning gains are rather small and the scores improve more on the Sun questions than on the star questions. This difference is largest for questions about the yearly apparent motion. We found that this is due to the fact that many students copy their knowledge about the Sun to the stars. Based on the results of this survey, we developed a new planetarium presentation with particular attention to the use of the celestial sphere model. We also developed a learning module that prepares students at school for this planetarium presentation. In a second study, we measured the learning gains after attending this new planetarium presentation among 339 students, also 16- to 17 -year -olds. Some school groups had worked through the preparatory learning module at school and others had not. We find that the learning gains on the star questions are significantly higher than in the first study, due to better scores on the yearly apparent motion questions. In this regard, it is notable that we do not see significant differences between those students who prepared the presentation at school and those who did not. In the second study, the number of students who answer all questions correctly after attending the planetarium presentation or working through the learning module increases, but only significantly for those students who worked through the learning module at school.
Mathematical language (or the content-specific words in mathematics) has repeatedly shown to be related to mathematical abilities in preschool and in the early grades of primary education. Research in this field has predominantly focused on young children’s quantitative and spatial language. At the same time, recent research has discovered that children in the early years of primary education already possess advanced mathematical skills such as reasoning about patterns, proportions and probabilities. In this study we developed an advanced mathematical language task to measure children’s understanding of mathematical language proper of the domains of patterning, proportionality, and probability, specifically aimed at first, second, and third graders (ages 5 -8). To develop a valid test instrument, our task was substantiated by a review of relevant literature and previously conducted studies. Further, the task was optimized based on the feedback received from expert professionals and a pilot study with 36 children. After administering the advanced mathematical language task to 236 first, second, and third graders, the results suggest that the task provides reliable scores in these grades. The internal consistency is excellent (α = .993). In addition, the test–retest reliability was checked in 49 children (r = .704; p < .001), and shows that the test also produces stable results over time.
Tasks in which learners are asked to compare two data sets using box plots and decide which distribution contains more observations above a given threshold have already been investigated in research. There are indications that these tasks are solved schema-based and that different (correct and erroneous) schemas are used depending on the arrangement of the quartiles around the threshold. Erroneous schemas can cause systematic errors and are often based on typical misconceptions. For example, if learners did not complete the conceptual change and assume that in box plots – like in most other statistical representations (e.g., bar or circle diagrams) - more (box) area also represents more observations, they decide the task according to which box plot shows more box area above the threshold. However, this can lead to incorrect answers, as the box area does not represent frequency but the range of the middle half of the data (interquartile range) and thus a measure of variability. So far, these schema-based reasoning processes have mainly been investigated via differences in solution rates of congruent and incongruent items. The present study investigates whether eye-tracking data can help to better understand which information is processed in the different schemas. Our research interest is based on hypotheses specifying which box plot components are significantly involved in the different schemas. We assume that the gaze patterns of learners using different schemas differ both regarding the number and duration of fixations on the relevant box plot components (areas of interest) and in terms of the number of transitions between them. We asked N = 14 participants to solve congruent and incongruent items and simultaneously collected eye movement data. In the analysis, we first used the solution rates to assign the schemas most likely used. Subsequently, the eye-tracking data were analyzed regarding differences in line with our hypotheses. We found hypothesis-compliant effects in all schemas regarding the number of fixations and transitions, but not regarding fixation duration. These results not only validate the schemas identified in previous studies, but also indicate that the schemas differ primarily in terms of which quartile is focused.