Automated short answer grading with feedback (ASAG-F) systems currently face challenges in transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment. We introduce GraphRAG, a knowledge graph-based retrieval-augmented generation framework that addresses these limitations by grounding all large language model (LLM)-generated feedback and scores in instructor-curated atomic facts, ensuring traceability and verifiability. Using the Short Answer Feedback (SAF) dataset with 31 topics, we evaluate GraphRAG on unseen-question and unseen-answer splits. Our systematic evaluation demonstrates that GraphRAG achieves grading accuracy comparable to vector-based RAG and generally superior to a fine-tuned LLM baseline model while providing more transparent source attribution. Additional findings include: (1) Instructing the LLM to discretize continuous scores to match pedagogical rubrics, such as the 0.25 increments common in SAF, improves grading accuracy; (2) LLM-generated feedback exhibits length-dependent quality variations when unconstrained; prompt-based length control substantially enhances feedback quality and its stability, achieving optimal balance of instructional richness and conciseness; (3) Performance scaling analysis reveals that basic models like GPT-4o-mini offer cost-effective performance, while premium models like Claude-Opus-4 show diminishing returns. These results demonstrate that GraphRAG offers a robust, explainable, pedagogy-aligned, and cost-effective solution for large-scale educational applications, enabling transparent automated grading with effective pedagogical feedback and practical deployment costs.
We discuss the use of Bayesian networks as a general framework for diagnostic classification in educational assessments, showing how they can accommodate sophisticated capabilities useful in diagnostic assessment, including modeling hierarchical structure among latent attributes, diagnostic use of information from incorrect alternatives in multiple-choice items, and simultaneous diagnosis of both subskills and misconceptions. These capabilities are illustrated with an application reported by (Lee 2003; Lee & Corter, 2003, 2011), who proposed using Bayesian networks as the inference engine to learn from test data and diagnose individuals' misconceptions or bugs in the domain of multicolumn subtraction. Lee and Corter demonstrated that diagnosis of misconceptions or bugs is most effective when information from incorrect alternatives in multiple-choice items is used and when both bugs and skills are assessed simultaneously, with a hierarchical structure assumed for subskills and misconceptions. More recently, these innovations and issues have been investigated in the context of traditional CDM models. In this paper, we describe the approach taken by Lee and Corter and discuss some advantages and disadvantages of using Bayesian networks for diagnostic assessment.
The effects of induced incidental moods on patterns of information search and decision outcomes were investigated in a risky choice task with mixed-domain problems. Viewing of short videos was used to induce either happy or sad mood in participants, who then made choices between pairs of options consisting of a probabilistic gain coupled with a probabilistic loss. Eyetracking measures of information search, specifically frequencies of transitions between key aspects of the decision alternatives, were analyzed and related to use of heuristic or analytic compensatory strategies. Data were also gathered in a control condition, where participants were instructed to use an EV-calculation strategy, a prototypical integrative compensatory strategy. Results showed significant differences in choices and attention transitions between the EV-instruction and the induced mood conditions, but minimal differences between the happy and sad induced mood conditions. Participants in the induced mood conditions showed relatively more evidence of heuristic strategy use, but analytic strategies remained the modal strategy in all conditions. Importantly, key types of attention transitions were shown to reliably predict the frequency of observed choices consistent with optimal (EV- maximizing) and certain heuristic strategies.
Cluster analysis has been used widely in educational research, even more so recently due to the surge of interest in educational data mining. Cluster analysis refers to a set of exploratory data analysis (EDA) methods to find structure in multivariate data by sorting instances into distinct groups of relatively similar cases. Clustering methods can themselves be grouped according to the type of underlying model used: partitioning, hierarchical clustering, or overlapping clustering. Related methods include Latent Class Analysis and Latent Profile Analysis. Educational applications have included clusterings of learners, classrooms, teachers, countries, educational institutions, curricula, test items and content concepts.
In this chapter, two influential kinds of purported group knowledge that pose challenges to my account of justified group belief are examined. The first is often referred to as “social knowledge,” a paradigmatic instance of which is the so-called knowledge possessed by the scientific community, where no single individual knows a proposition, but the information plays a functional role in the community. The second is “collective knowledge,” where knowledge may be imputed to a group by aggregating bits of information had by its individual members. It is shown that both social knowledge and collective knowledge sever the crucial connection between knowledge and action, and open the door to serious abuses, not only epistemically, but morally and legally as well. Bits of information that are merely accessible to group members, or individual instances of knowledge that are aggregated with no communication, do not amount to group knowledge in any robust sense.
In a recent paper [1], Christián Carman advanced a tentative explanation for “overspecification” in medieval mathematical diagrams. Carman argues that the original (“correct”) diagrams were corrupted, presumably through incompetent copyists, while preparing the initial copies—often before the tenth consecutive copy. The diagrams then stabilized in an overspecified form and resisted further changes, sometimes for centuries of copies thereafter. I feel hesitant about this hypothesis for several reasons: (1) it assumes that the first Greek diagrams were essentially identical to modern diagrams; (2) pre-modern overspecification is ubiquitous and is rarely reversed; (3) the hypothesis ignores differing traditions of perspective; (4) the informal tests used to support the hypothesis do not precisely mirror the medieval copyist’s activity.
Allocation of Attention in Neural Network Models of Categorization Toshihiko Matsuka (tm249@columbia.edu) James E. Corter (jec34@columbia.edu) Department of Human Development, Teachers College, Columbia University, 525 W. 120 th St., New York, NY 10027 USA Arthur B. Markman (markman@psy.utexas.edu) Department of Psychology, University of Texas, Austin, TX 78712 USA We compared ALCOVE (Kruschke, 1992), RASHNL (Kruschke & Johansen, 1999), SUSTAIN (Love & Medin, 1998), and the Cortico-Hippocampal Model (CHM) (Gluck & Myers, 1993) to see how they account for selective attention in category learning. Such comparisons may usefully augment comparisons of the models’ classification accuracy. Method We simulated the results of studies of classification learning by Medin and Schaffer (1978) and Medin, Altom, Edelson & Freko (1982). The parameter values used for each model were adjusted to minimize the SSE in reproducing the training classification responses by human subjects. Attention allocation predictions for the models were derived as follows. ALCOVE and RASHNL have explicit attention weight parameters, which are reported below. For SUSTAIN, the dimension-specific tuning parameters, λ, are reported. In the CHM there are no explicitly defined dimension attention parameters. We defined implicit measures of a dimension’s attentional salience, by summing the absolute values of weights from all input nodes associated with a given dimension to the hidden node layer in the hippocampal net component of the CHM. To enhance comparability among the models, we computed and report relative attention weights for all the models. Summary of Results For Experiment 2 of Medin and Schaffer (1978), all the models fit the training set classification probabilities roughly equally well, but RASHNL and SUSTAIN were somewhat more accurate in predicting classification responses for the transfer stimuli. In this stimulus structure Dimensions 1 and 3 are highly predictive of the binary classification task, and Dimension 4 is moderately predictive. Somewhat surprisingly, ALCOVE, RASHNL, and the CHM gave as much or more attention weight to Dimension 4 as to the more diagnostic dimensions. For Experiment 4 of Medin, Altom, Edelson & Freko (1982), RASHNL and the CHM fit the training set classification probabilities best, but RASHNL was the best and the CHM worst in predicting the transfer classifications. In this stimulus structure, Dimensions 1 and 2 are diagnostic in the sense that each is highly correlated with the criterion classification response, but Dimensions 3 and 4 have a simple XOR pattern in regards to the criterion classification. ALCOVE, RASHNL, and SUSTAIN all learn to allocate more attention to Dimensions 3 and 4, that together define the classification in terms of a simple XOR relationship. In contrast, the CHM pays more attention to the individually, but merely probabilistically, diagnostic Dimensions 1 and 2. Conclusions The four models give different predictions about attention weights for some stimulus structures. Examining and comparing these predictions may shed light on how the models learn. A promising line for future research is to gather direct data on how humans allocate attention in category learning (Matsuka, 2002). References Gluck, M. A., & Myers, C. E. (1993). Hippocampal mediation of stimulus representation: A computational theory. Hippocampus, 3, 491-516. Kruschke, J. K. (1992). ALCOVE: An exemplar-based connectionist model of category learning. Psychological Review, 99, 22-44. Kruschke, J. K., & Johansen, M. K. (1999). A model of probabilistic category learning. Journal of Experimental Psychology: Learning, Memory, & Cognition, 25, 1083- Love, B. C., & Medin, D. L. (1998). SUSTAIN: A model of human category learning. Proceeding of the Fifteenth National Conference on AI (AAAI-98), 671-676. Matsuka, T. (2002). Attention processes in category learning. Unpublished doctoral dissertation (draft), Teachers College, Columbia University. Medin, D. L., Altom, M. W., Edelson, S. M., & Freko, D. Correlated symptoms and simulated medical classification. Journal of Experimental Psychology: Learning, Memory, and Cognition, 8, 37-50. Medin, D. L., & Schaffer, M. M. (1978). Context theory of classification learning. Psychological Review, 85, 207-238
In 2010, the province of Ontario introduced a new universal two-year play-based full-day kindergarten program. The authors exploited the phasing-in of this program over five years, allowing a natural experiment in which children from full-day kindergarten could be compared with those from half-day kindergarten in matched neighborhoods. Children (N=592) were followed from kindergarten to Grade 2 with direct learning and self-regulation measures. Grade 3 wide-scale achievement test scores were available for 269 of the children. Results showed lasting benefits of full-day kindergarten on children's self-regulation, reading, writing, and number knowledge to the end of Grade 2, including some benefits for vocabulary. Full-day kindergarten children were significantly more likely to meet provincial expectations for reading in Grade 3. The study points to the benefits of a play-based full-day kindergarten program and brings evidence to bear on the mixed findings in the research literature about the fade-out effects of full-day kindergarten.
Team climate is thought to play a vital role in shaping the process and outcomes of collaborative problem-solving. However, prior evidence for this proposition in the research literature is largely observational and correlational in nature. This report describes a field experiment that manipulates two key dimensions of group climate (West, 1990): group task orientation and support for innovation, and assesses the effects of this manipulation on the creativity of group solutions to a set of open-ended problems. Participants were 295 high school students in China working in 60 collaborative groups to solve four tasks used in prior creativity research. In order to establish a causal role for team climate in promoting creativity, rewards were used to manipulate the team climate factors of task orientation and support for innovation between groups in a 2 x 2 design. Task orientation, defined as a shared concern with excellence of group performance, was manipulated by informing some groups (but not others) in advance about both monetary and social performance rewards. Support for innovation, the expectation of and support for unique ideas, was manipulated by describing to some groups in advance an aspect of the scoring rule that rewarded originality (defined as unique answers to the tasks). The group responses to the tasks were scored on three dimensions often used to define creativity: fluency, flexibility, and originality. As hypothesized, the manipulations designed to incentivize a concern for excellence and support for innovation each showed a main effect in the form of higher creativity scores, but the effects were limited to the originality dimension. The findings suggest ways for educators and parents to plan and implement effective pedagogical strategies aimed at improving students' collaborative creativity.
Venn and Euler diagrams are valuable tools for representing the logical set relationships among events. Proportional Euler diagrams add the constraint that the areas of diagram regions denoting various compound and simple events must be proportional to the actual probabilities of these events. Such proportional Euler diagrams allow human users to visually estimate and reason about the probabilistic dependencies among the depicted events. The present paper focuses on the use of proportional Euler diagrams composed of rectangular regions and proposes an enhanced display format for such diagrams, dubbed “Euler boxes”, that facilitates quick visual determination of the independence or non-independence of two events and their complements. It is suggested to have useful applications in exploratory data analysis and in statistics education, where it may facilitate intuitive understanding of the notion of independence.
A problem sorting task was used to examine how the semantic content of probability word problems affects problem understanding and categorization, for students with various levels of statistical training. In the task, undergraduate and graduate students were asked to sort probability problems into groups by similarity of solution. The problems varied by relevant probability principle, by type of semantic schema, and by cover-story surface content. Results showed that both less-trained students and more-trained students tended to sort problems by relevant probability principle, but students with more statistics training did this more consistently. Both groups of students tended to be affected in the sorting task by semantic schema, defined here as intermediate-level abstractions of the problem structure. For example, when a permutation problem described assignment of people to people, students showed a strong tendency to group it with independent-events problems with a people-to-people matching schema.
Human studies of sleep and cognition have established thatdifferent sleep stages contribute to distinct aspects of cognitive and emotional processing. However, since the majority of these findings are based on single-night studies, it is difficult to determine whether such effects arise due to individual, between-subject differences in sleep patterns, or from within-subject variations in sleep over time. In the current study, weinvestigated the longitudinal relationship between sleep patterns and cognitive performance by monitoring both in parallel, daily, for a week. Using two cognitive tasks - one assessing emotional reactivity to facial expressions and the other evaluating learning abilities in a probabilistic categorization task - we found that between-subjectdifferences in the average time spent in particular sleep stages predicted performance in these tasks far more than within-subject daily variations. Specifically, the typical time individualsspent in Rapid-Eye Movement (REM) sleep and Slow-Wave Sleep (SWS) was correlated to their characteristic measures of emotional reactivity, whereas the typical time spent in SWS and non-REM stages 1 and 2 was correlated to their success in category learning. These effects were maintained even when sleep properties werebased onbaseline measures taken prior to the experimental week. In contrast, within-subject daily variations in sleep patterns only contributed to overnight difference in one particular measure of emotional reactivity. Thus, we conclude that the effects of natural sleep onemotional cognition and categorylearning are more trait-dependent than state-dependent, and suggest ways to reconcile these results with previous findings in the literature.