Widespread student adoption of large language models (LLMs) has prompted many CS instructors to assign greater weight to handwritten, proctored assessments. However, this approach struggles to scale as class sizes outpace course staff resources. To address this challenge, our study explores LLM-assisted grading to reduce required grading time. While prior work has emphasized tool accuracy, we evaluate both time and accuracy by comparing outcomes when course staff use an LLM-assisted grader versus Gradescope. We also incorporate a mixed-methods analysis of student and staff perceptions. In a CS1 course of 166 students supported by four teaching assistants (TAs), we observed that LLM-assisted grading reduced overall grading time by 40% compared to Gradescope, with time savings of 48% for exams and 25% for quizzes. Across all assessments, short answer questions showed a 46% time improvement, and free response questions showed a 37% time improvement. In terms of accuracy, accepted regrade requests increased negligibly from 0.1% to 0.5% across three exams and six quizzes. Students were generally neutral about LLM-assisted grading, but stressed the value of TA feedback and oversight. Meanwhile, TAs expressed positive sentiments towards the tool, tempered by concerns of skewed perceptions of students caused by the tool. Overall, these findings indicate that LLM-assisted grading can greatly reduce grading time, with only minor accuracy trade-offs that can be mitigated. As a result, LLM-assisted grading emerges as a promising approach for enhancing grading efficiency in CS courses, meriting further exploration for broader adoption.
Despite recent advances in fairness-aware machine learning, predictive models often exhibit discriminatory behavior towards marginalized groups. Such unfairness might arise from biased training data, model design, or representational disparities across groups, posing significant challenges in high-stakes decision-making domains such as college admissions. While existing fair learning models aim to mitigate bias, achieving an optimal trade-off between fairness and accuracy remains a challenge. Moreover, the reliance on black-box models hinders interpretability, limiting their applicability in socially sensitive domains. To circumvent these issues, we propose integrating Kolmogorov-Arnold Networks (KANs) within a fair adversarial learning framework. Leveraging the adversarial robustness and interpretability of KANs, our approach facilitates stable adversarial learning. We derive theoretical insights into the spline-based KAN architecture that ensure stability during adversarial optimization. Additionally, an adaptive fairness penalty update mechanism is proposed to strike a balance between fairness and accuracy. We back these findings with empirical evidence on two real-world admissions datasets, demonstrating the proposed framework's efficiency in achieving fairness across sensitive attributes while preserving predictive performance.
This work investigates the impact of Large Language Models (LLMs) and the COVID-19 pandemic on student behavior with autograder systems in three programming-heavy courses. We examine whether the release of LLMs like ChatGPT and GitHub Copilot, along with post-pandemic effects, has modified student interactions with autograders. Using data from student submissions over five years, totalling over 4,500 students across over 420,000 submissions, we analyze trends in submission behaviors before and after these events. Our methodology involves tracking submission patterns, focusing on timing, frequency, and score. Contrary to expectations, our findings reveal that metrics remain relatively consistent in the post-ChatGPT and post-pandemic era. Despite yearly fluctuations, no significant shift in student behaviors is attributable to these changes. Students continue to rely on a combination of manual debugging and autograder feedback without noticeable changes in their problem-solving approach. These findings highlight the resilience of the educational practices in these courses and suggest that integrating LLMs into mid-level CS curriculum may not necessitate the significant paradigm shift previously envisioned. Future work should extend these analyses to courses with different structures to determine if these results are generalizable. If not, the specific course aspects contributing to our observed ChatGPT and pandemic resilience should be identified.
This study examines disparities in self-reported HRQoL among English-speaking non-Latinx White, English-speaking Latinx, and Spanish-speaking Latinx children ages 4–12 years undergoing surgery. A total of 357 children completed the Child Health Rating Inventories, an animated, computer-administered method, to measure overall, physical, and mental health, as well as pre-operative anxiety. A multivariate general linear model was used to analyze the main effects of race/ethnicity and language on self-reported HRQoL. Results demonstrated differences in child self-reported overall [F(2,311) = 3.11, p = 0.05)] and mental health F(2,311) = 3.56, p = 0.03)], and preoperative anxiety F(2,311) = 5.70, p = 0.004)] by race/ethnicity and language. Post hoc comparisons using the Bonferroni test indicated that English-speaking Latinx children reported significantly poorer overall (p = 0.04) and mental health (p = 0.04) compared to English-speaking non-Latinx children. English-speaking and Spanish-speaking Latinx children reported significantly higher preoperative anxiety (p = 0.004 and p = 0.02, respectively) compared to English-speaking non-Latinx White children. Latinx children from English-speaking households as young as 4 years old reported their overall and mental health to be poorer compared to Non-Latinx White children from English-speaking households. Latinx children, regardless of spoken language, reported higher preoperative anxiety compared to non-Latinx White children. These findings highlight the need to consider early childhood experiences in understanding health disparities. Factors such as family dynamics, acculturative stress, and access to healthcare resources could potentially account for disparities in young children’s health experiences.
ABSTRACT Background A total of 80% of children experience postoperative pain following discharge. Effective postoperative pain management involves reliable caregiver pain assessment and/or child self‐report of pain. Unfortunately, caregiver and child ratings of postoperative pain are not always consistent (i.e., concordant). This study aimed to identify postoperative pain concordance among caregiver–child dyads and predictors for postoperative pain discordance. Methods Children and their caregivers completed preoperative baseline demographic, anxiety, and distress measures. Postoperatively, children and caregivers completed pain severity ratings using the Child Health Rating Inventories (CHRIS 2.0). On the basis of postoperative pain scores, caregiver–child dyads were classified as overestimators (i.e., caregivers rated pain as higher than children), in agreement, or underestimators (i.e., caregivers rated pain as lower than children). Results A large proportion of dyads disagreed on pain ratings ( n = 104; 44%), with 64 (27%) caregivers classified as overestimators and 40 (17%) caregivers classified as underestimators. Caregivers were more likely to underestimate male children's pain, β = 1.238, OR = 3.35 (95% CI: 1.26, 9.43), p = 0.16, and Spanish–speaking Latinx caregivers were more likely to underestimate children's pain, β = 2.27, OR = 9.63 (95% CI: 2.35, 39.37), p = 0.002. Conclusion Although most caregiver–child dyads agreed with pain ratings, 44% of the dyads disagreed. Among those who disagreed, males from Spanish–speaking Latinx households were at greatest risk of having their pain underestimated by their caregiver, which could be explained by the influence of intersecting social identities on pain beliefs, expression, and behaviors. Future studies should explore how pain discrepancies influence postoperative recovery outcomes for Latinx children.
Objectives:There has been a growing emphasis on holistic approaches to assessing postoperative recovery by using self-reported health-related quality of life (HRQoL). Identifying groups of children at higher risk of poor recovery has become important. The aim of the study is to identify predictors of paediatric postoperative recovery assessed by self-reported HRQoL. Methods:One hundred forty-eight children ages 4 to 12 years completed the Child Health Rating Inventories (CHRIS2.0) to measure overall, physical and mental health, preoperative anxiety, and postoperative pain. Four linear regressions were used to identify predictors of overall, physical and mental health and postoperative pain. Predictors included child gender, race/ethnicity and language, surgical severity, child and caregiver preoperative anxiety, and caregiver distress. Results:Child male gender (p = 0.03, 95% confidence interval [CI] [-10.15, -0.65]) and identifying as English-speaking Latinx (p = 0.03, 95% CI [0.58, 13.25]) predicted poorer postoperative overall health. Higher child preoperative anxiety (p < 0.001, 95% CI [0.39, 1.50]) and higher caregiver preoperative distress (p = 0.003, 95% CI [-1.09, 0.28]) predicted poorer postoperative overall health. Conclusions:The results of this self-reported study validated previously established predictors of recovery (preoperative anxiety and caregiver distress). Novel predictors, including child male gender and race/ethnicity and language, were identified, providing new insights into factors influencing recovery outcomes.
This research category full paper presents a lightweight, deployable Large Language Model (LLM) pipeline to automatically answer student forum questions about course logistics. By leveraging Retrieval-Augmented Generation (RAG) over course documents like the syllabus and calendar, our Llama 3.1-based model offers an accurate, low-cost alternative to commercial tools. In addition, many countries have strict student-privacy laws, making online third-party solutions difficult or impossible to gain approval for. Unlike large proprietary models, our pipeline runs on consumer hardware and avoids expensive API costs. We evaluated 159 authentic student forum questions across eleven Computer Science (CS) courses, showing that the model accurately answered 74% of logistical queries. Performance improved significantly when filtered by retrieval confidence scores above 0.5, allowing staff to easily identify questions needing review. To contextualize its effectiveness, we compared our pipeline to several state-of-the-art commercial models. While commercial tools slightly outperformed in raw accuracy, our model provided comparable results at a fraction of the computational and financial cost. This tradeoff makes our approach particularly well-suited for educational institutions seeking low-cost, scalable solutions compliant with student-privacy laws such as FERPA. Our findings support the use of lightweight LLMs for automating routine academic support, offering a viable and accessible path forward for institutions with limited infrastructure or privacy restrictions.
Stress is a pervasive issue in modern society, and a plethora of technological approaches have been developed for its detection. In this study, we propose an end-to-end framework for detecting chronic stress using only data collected from a smartwatch and a lightweight machine learning model that runs within a mobile application. Previous work achieved accurate results but required additional technology, such as cloud servers and custom sensors not readily available to the public. We tested lightweight models and found that a fine-tuned LightGBM achieved an $81.6 \% \mathrm{~F} 1$-score, while a hyperdimensional computing (HDC) model, optimizing efficiency with a slight accuracy tradeoff, reached 73%.
Excess alcohol consumption leads to serious health risks and severe consequences for both individuals and their communities. To advocate for healthier drinking habits, we introduce a groundbreaking mobile smartwatch application approach to just-in-time interventions for intoxication warnings. In this work, we have created a dataset gathering TAC, accelerometer, gyroscope, and heart rate data from the participants during a period of three weeks. This is the first study to combine accelerometer, gyroscope, and heart rate smartwatch data collected over an extended monitoring period to classify intoxication levels. Previous research had used limited smartphone motion data and conventional machine learning (ML) algorithms to classify heavy drinking episodes; in this work, we use smartwatch data and perform a thorough evaluation of different state-of-the-art classifiers such as the Transformer, Bidirectional Long Short-Term Memory (bi-LSTM), Gated Recurrent Unit (GRU), One-Dimensional Convolutional Neural Networks (1D-CNN), and Hyperdimensional Computing (HDC). We have compared performance metrics for the algorithms and assessed their efficiency on resource-constrained environments like mobile hardware. The HDC model achieved the best balance between accuracy and efficiency, demonstrating its practicality for smartwatch-based applications.
Excessive alcohol consumption was responsible for 6% of global deaths in 2023. To encourage healthier drinking habits and enhance user awareness of their current condition, just-in-time interventions prove to be a suitable approach for informing users about their current state of intoxication. Current methods for determining blood alcohol content are intrusive and many also invasive, requiring users to use breathalizers or actively engage in urine or blood tests. In this study, we introduce an application utilizing Hyperdimensional Computing to predict if a user is under the influence of alcohol, achieving an accuracy of 93.5% on average. Furthermore, this application is designed to run on both smartphones and smartwatches, enabling full on device computation and online learning through a C implementation utilizing vectorial operations. The application has shown to be very efficient, having a training time per instance of 13.2 and 1.25ms on smartwatch and smartphone respectively and inference time of 6.8 and 1.1ms. Moreover the energy consumption of the running application is negligible compared to the energy usage of the idle device.
With the mainstream adoption of Large Language Models (LLMs), members of both academia and the media have raised concerns around their impact on student learning and pedagogy. Many students and educators wonder about the pedagogical fit of this emerging technology. We aim to measure the adoption of and attitudes toward LLMs among the CS student population at an R1 University to determine how students are using these new tools. To this end, we conducted a large survey study targeting two populations participating in computing courses at the university: intro-sequence students (ISS) and experienced students (ES). In our preliminary results from Spring 2023, we've found several significant differences among the views of over 700 respondents across the two groups. Most students reported LLMs' unparalleled potential for quick information access, yet many harbor concerns about the reliability of the LLM responses, and the impact on academic integrity. Additionally, while ES have rapidly integrated LLMs into their learning, ISS remain cautious of the tools, highlighting a stark contrast in adoption rates between the groups. LLMs are clearly going to reshape pedagogical approaches and student engagement. Our study hopes to provide insight on the nuanced student attitudes toward LLMs. For example, the notable reservations expressed by ISS illustrate an imperative for careful, informed, and ethical integration to ensure these tools enhance rather than compromise the educational experience. In the future, we plan to continue tracking student attitudes in order to gain further understanding of the changing perceptions of LLMs and their impact.
Optimal group formation and project matching are critical and challenging tasks for instructors. We developed the Student-Project Matching Tool to optimize these processes and piloted it in a Computer Science Capstone course at the University of California, Irvine. The tool ensures that the team formation process balances individual preferences, project compatibility, and the overall performance potential of each team by considering students' skills and interests and sponsor projects' needs to maximize teams' success. Student perspectives and feedback showed an increase in student satisfaction with their team and the project they were matched to. Similarly, positive sponsor evaluations of the teams demonstrated that sponsors were pleased with the teams they were matched to. This tool provides the basis for effective team formation and project matching in Capstone courses, with a focus on maximizing student learning outcomes, real-world experiences, student-project ownership, and the number of fulfilled skills that each project requires for completion.
More than 50% of Computer Science (CS) students at the University of California, Irvine (UCI) experienced academic probation over a 10-year span. Particularly concerning, underrepresented groups (URG) faced an even higher probation rate, and probation students were twice as likely to leave the CS program. We conducted a comprehensive survey involving 308 CS1 students at UCI to delve further into their past and present academic experiences. Specifically, we studied (1) the role of academic preparedness and computing participation in CS students' success, and (2) to what extent these factors impact the academic success of URG students. Our findings reveal significant correlations between underperformance in CS1 and inadequate pre-college math preparation, consistent with existing literature. Moreover, our results show that most URG students also lack prior programming exposure and peer communities within their field (study groups or clubs). These insights highlight the urgency to consider curriculum and academic preparation modifications to establish a strong math and programming foundation for all students, which can guide them toward greater success in CS.
This full innovative practice paper describes a computational tool designed to optimally match students to industry-sponsored capstone projects in a software engineering capstone course for Computer Science undergraduates at an R1 University (R1U). In the context of these capstone courses, where students stand at the culmination of their academic journey, aligning students' personal learning goals and existing computing skills with team formations becomes critical. This paper presents the Student-Project Matching Tool (SPMT), created to help students find the best available industry-sponsored projects based on their desired learning outcomes, project requirements, and their interests in each project. To choose the learning outcomes they aim to achieve, students can select from a list of predefined software engineering categories and the skills needed to achieve proficiency in each category. The initial list of technical skills for each category was recorded from job postings on a variety of well-known job-search websites, and was further refined by the capstone program's industry partners. Allowing students to select the skills they will work on ensures that they have opportunities and exposure to the skill sets required for employment while still working on one of their most appealing projects. We have developed and piloted the SPMT, which utilizes student vectors to represent their interests and experiences across various software engineering skill sets. Similarly, this tool uses vectors to represent the skills required by each available project, aligning with the exact dimensions as those of the student vectors. The SPMT calculated Euclidean distances between the student interest and project requirement vectors. Next, the resulting Euclidean distances were multiplied with weights associated with students' level of interest in each industry-sponsored project. Subsequently, we framed the student-project matching process as a linear sum assignment problem, aiming to minimize the total sum of Euclidean distances between each student-project pair. The output of the SPMT process consistently matched students with teams that met their software engineering interests and project priorities. Our results reveal increased engagement and growth toward students' desired learning outcomes and computing skills. Specifically, after the first term of the capstone sequence, most students self-reported higher levels of proficiency growth in the skills within their desired software engineering category. This suggests that the SPMT effectively provides students with valuable learning experiences relevant to their career interests and representative of real-world settings.
Alcohol consumption has a significant impact on individuals' health, with even more pronounced consequences when consumption becomes excessive. One approach to promoting healthier drinking habits is implementing just-in-time interventions, where timely notifications indicating intoxication are sent during heavy drinking episodes. However, the complexity or invasiveness of an intervention mechanism may deter an individual from using it in practice. Previous research tackled this challenge using collected motion data and conventional Machine Learning (ML) algorithms to classify heavy drinking episodes, but with impractical accuracy and computational efficiency for mobile devices. Consequently, we have elected to use Hyperdimensional Computing (HDC) to design a just-in-time intervention approach that is practical for smartphones, smart wearables, and IoT deployment. HDC is a framework that has proven results in processing real-time sensor data efficiently. This approach offers several advantages, including low latency, minimal power consumption, and high parallelism. We explore various HDC encoding designs and combine them with various HDC learning models to create an optimal and feasible approach for mobile devices. Our findings indicate an accuracy rate of 89%, which represents a substantial 12% improvement over the current state-of-the-art.
This work considers the problem of enhancing the authenticity and fairness of undergraduate student admission decision-making process by employing state-of-the-art Deep Learning (DL) models with advanced bias-mitigation techniques. Traditional admission processes often introduce biases that can disadvantage underrepresented or marginalized groups, highlighting the need for more equitable and efficient methods. Although the DL models have emerged as a promising alternative offering superior performance and higher scalability than the classical Machine Learning approaches, fairness concerns remain a significant issue.We propose an adversarial debiasing-based DL framework that integrates the Optimistic Adam (OAdam) optimizer, ensuring consistent and stable model training crucial for achieving reliable and unbiased outcomes. Our framework leverages data from applicants to the Computer Science Department at the University of California, Irvine. To ensure holistic evaluation of applicants’ profile we utilize a dataset that encompasses a wide range of features showcasing demographics, academic records, high school information, and essay responses. By prioritizing the recall score alongside the fairness metrics, our approach effectively handles the fairness-accuracy trade-off, considerably minimizing the false negatives and ensuring equitable consideration for marginalized groups in admission decisions. Through rigorous experimentation and analysis, our comprehensive study demonstrates that the proposed fairness-aware Input Convex Neural Network model using OAdam optimizer, achieves high fairness metrics while ensuring a balanced predictive performance. The proposed model improves the p-% rule scores by an average of 39.989% across sensitive attributes and achieves recall scores 0.97% higher than those of unfair baseline models.
Computer Science (CS) students at the University of California, Irvine (UCI) have experienced academic probation rates higher than 50%. Particularly concerning, statistical analysis showed that students who self-identified as belonging to an underrepresented group (URG) experienced an even higher probation rate. Moreover, students who entered academic probation were twice as likely to leave the CS program. We designed and conducted a comprehensive survey involving 757 CS1 students at UCI to delve further into their past experiences, challenges, and perspectives to gain further insights into the factors contributing to these trends. Specifically, we studied (1) the role of socioeconomic factors such as mental health, academic preparedness, and computing participation in CS students' success, (2) to what extent these factors affect underrepresented group, first-generation, and female students, and (3) experiences that distinguish the most impacted minority groups. Our findings reveal significant correlations between underperformance in CS1 and socioeconomic factors, including satisfaction with course completion regardless of grade, mental health challenges, and insufficient pre-college math preparation. Many of these factors had strong associations with all minority groups. Moreover, our data shows that most URG students enter the program with weaker math preparation than their peers, often don't have prior programming experience, and once enrolled, they have limited interactions within the CS community. These insights highlight the urgency of redesigning academic support practices to support students with diverse backgrounds and experiences. There is a growing need to implement tailored interventions and support mechanisms for CS students, focusing on addressing the disparities in preparation, perspectives, and experiences. Our findings highlight the pressing need to reevaluate current academic support practices and provide a foundation for developing targeted support programs to guide struggling students toward greater success in CS.
With the mainstream adoption of Large Language Models (LLMs) over the last year, members of both academia and the media have raised concerns around the potential impact on student learning and pedagogy. Many students and educators wonder about the pedagogical fit of this emerging technology. We aim to measure the adoption and perception of LLMs among the CS education community in an R1 University to distinguish reality from hype. To this end, we conduct a large survey study targeting three populations participating in computing courses at the university: intro-sequence students (ISS), experienced students (ES), and faculty. Our survey seeks to gather insight around the different populations' perceptions of LLMs in education, as well as how these perceptions may be changing as LLMs improve. Our results show several significant differences across the views of 760 respondents. Most students report LLMs' un-paralleled potential for quick information access, yet many harbor concerns about their reliability and impact on academic integrity. Additionally, while ES rapidly integrate LLMs into their learning, ISS and faculty remain cautious, highlighting a stark contrast in adoption rates. Faculty are unconvinced of LLMs' educational benefits and are concerned about potential challenges in evaluating students' learning outcomes. LLMs are reshaping pedagogical approaches and student engagement. However, with the notable reservations expressed by certain segments, particularly by faculty and ISS, there is an imperative for careful, informed, and ethical integration to ensure that these tools enhance rather than compromise the educational experience.
Although the prevention of AI vulnerabilities is critical to preserve the safety and privacy of users and businesses, educational tools for robust AI are still underdeveloped worldwide. We present the design, implementation, and assessment of Maestro. Maestro is an effective open-source game-based platform that contributes to the advancement of robust AI education. Maestro provides goal-based scenarios where college students are exposed to challenging life-inspired assignments in a competitive programming environment. We assessed Maestro's influence on students' engagement, motivation, and learning success in robust AI . This work also provides insights into the design features of online learning tools that promote active learning opportunities in the robust AI domain. We analyzed the reflection responses (measured with Likert scales) of 147 undergraduate students using Maestro in two quarterly college courses in AI. According to the results, students who felt the acquisition of new skills in robust AI tended to appreciate highly Maestro and scored highly on material consolidation, curiosity, and maestry in robust AI . Moreover, the leaderboard, our key gamification element in Maestro , has effectively contributed to students' engagement and learning. Results also indicate that Maestro can be effectively adapted to any course length and depth without losing its educational quality.