Context: Many students now use generative AI (genAI) in their coursework, yet its effects ontheir intellectual development remain poorly understood. While prior work has investigatedstudents’ cognitive offloading during episodic interactions, it remains unclear whether usinggenAI routinely is tied to more fundamental shifts in students’ thinking habits.Objectives: To explore this possibility, we investigate (RQ1-How): how students’ trust in androutine use of genAI affect their cognitive engagement—specifically, reflection, the need forunderstanding, and critical thinking in STEM coursework. Further, we investigate (RQ2-Who):which students are particularly vulnerable to these cognitive disengagement effects.Methods: We drew on dual-process theory, cognitive offloading, and the automation biasliterature to develop a statistical model explaining how and to what extent students’ trust-driven routine use of genAI affected their cognitive engagement habits in STEM coursework, and how these effects differed across students’ diverse cognitive styles. We empirically evaluated this model using Partial Least Squares Structural Equation Modeling on survey data from 299 STEM students across five North American universities.Results: Students who trusted and routinely used genAI reported significantly lower cognitiveengagement. Unexpectedly, students with higher technophilic motivations, risk tolerance, andcomputer self-efficacy—traits often celebrated in STEM—were more prone to these effects.Interestingly, students’ prior experience with genAI or academia did not protect them fromcognitively disengaging.Conclusion: Our findings suggest a potential cognitive debt cycle in which routine genAIuse progressively weakens students’ intellectual habits, potentially driving over-reliance andescalating usage.This poses critical challenges for curricula and genAI system design, requiring interventions that actively support cognitive engagement.
Recruiting, retaining, and educating students in computing is a frequent research topic in CHI. However, students’ sociotechnical experiences of registering for classes are understudied—especially those of socioeconomic-diverse students. These experiences matter: research shows that registration problems bring long-term consequences to student successes. We investigate students’ socioeconomic status (SES) impact on registration experiences through three studies: a case study with education professionals using an emerging analytic method, SocioeconomicMag (SESMag); interviews with faculty/staff/students from 8 universities; and observations of 14 SES-diverse students registering for classes. Results showed: (1) 5 SES-inclusivity bugs which arose 30 times, 72% more often by lower-SES students than by higher-SES students. (2) 6/7 lower-SES students (but only 2/7 higher-SES students) expected downstream problems from the registration issues. (3) The risk-to-negative-outcomes rate was 3 times higher for lower-SES students. (4) The issues generalized across 8 universities and potentially to >700 other universities who use the same registration portal.
Objectives: When students use generative AI in coursework, what are its persistent effects on their intellectual development? We investigate (RQ1-How) how students' trust in and routine use of genAI affect their cognitive engagement habits in STEM coursework, and (RQ2-Who) which students are particularly vulnerable to cognitive disengagement. Method: Drawing on dual-process, cognitive offloading, and automation bias theories, we developed a statistical model explaining how and to what extent students' trust-driven routine genAI use affected their cognitive engagement – specifically, reflection, the need for understanding, and critical thinking in coursework, and how these effects differed across students' cognitive styles. We empirically evaluated this model using Partial Least Squares Structural Equation Modeling on survey data from 299 STEM students across five North American universities. Results: Students who trusted and routinely used genAI reported significantly lower cognitive engagement. Unexpectedly, students with higher technophilic motivations, risk tolerance, and computer self-efficacy – traits often celebrated in STEM – were more prone to these effects. Interestingly, students' prior experience with genAI or academia did not protect them from cognitively disengaging. Implications: Our findings suggest a potential cognitive debt cycle where routine genAI use weakens students' intellectual habits, potentially driving and escalating over-reliance. This poses challenges for curricula and genAI system design, requiring interventions that actively support cognitive engagement.
While much research has shown the presence of AI's "under-the-hood" biases (e.g., algorithmic, training data, etc.), what about "over-the-hood" inclusivity biases: barriers in user-facing AI products that disproportionately exclude users with certain problem-solving approaches? Recent research has begun to report the existence of such biases -- but what do they look like, how prevalent are they, and how can developers find and fix them? To find out, we conducted a field study with 3 AI product teams, to investigate what kinds of AI inclusivity bugs exist uniquely in user-facing AI products, and whether/how AI product teams might harness an existing (non-AI-oriented) inclusive design method to find and fix them. The teams' work resulted in identifying 6 types of AI inclusivity bugs arising 83 times, fixes covering 47 of these bug instances, and a new variation of the GenderMag inclusive design method, GenderMag-for-AI, that is especially effective at detecting certain kinds of AI inclusivity bugs.
Motivations. Explainable AI (XAI) systems aim to improve users' understanding of AI, but XAI research has shown that many XAI explanations serve some users well while failing others. In non-AI systems, software practitioners have used inclusive design approaches to address similar problems, sometimes creating "curb-cut" improvements that benefit both underserved users and everyone else. This raises the possibility that inclusive design approaches can bring similar curb-cut improvements to AI explanations. Objectives. Our objective was to investigate possible curb-cut effects of inclusivity-driven fixes an AI product team made using an inclusive design approach (GenderMag) to improve their XAI prototype. Methods. We ran a between-subject study with 69 participants who had no formal AI background. 34 participants used the original version of the XAI prototype and the rest used the version with the AI team's inclusivity fixes. We then compared the two groups' mental model concepts scores and prediction accuracy, and the two prototypes' inclusivity. Results. Our investigation produced four main results. First, the AI team's inclusivity fixes were overall effective, resulting in overall better conceptual mental models with the new prototype. Further (second), the AI team's inclusivity fixes were particularly beneficial to the underserved population's conceptual mental models-which, together with the first result, constitutes a curb-cut effect. However (third), the inclusivity fixes did not improve participants' prediction accuracy scores. Instead, it appears to have harmed them overall-a "curb-fence" effect (opposite of a curb-cut effect). Finally (fourth), the AI team's fixes improved equity, reducing the gender gap by 45%.
Large language model (LLM) code explanations can support people in solving code-related problems, yet prior work has shown that people have diverse problem-solving styles. If explanations fail to meet people's problem-solving needs, they may be less productive in their occupations and miss opportunities to learn and grow. Although some research has examined how LLMs can adapt their outputs to a user's age or expertise, no prior work has examined how LLMs can adapt their code explanations to people's problem-solving styles. To address this gap, we developed prompts from an established inclusive design method that considers 5 types of problem-solving styles, and we generated 1,072 code explanations from six open-weight LLMs. Using natural language processing techniques, we uncovered a taxonomy of 13 linguistic adaptations, with each adaptation supported by evidence from the literature, the prompts, or the LLMs' outputs. They also show which LLMs adapted their code explanations more frequently than others. This paper is the first to investigate problem-solving style adaptations in LLM code explanation, contributing two problem-solving adaptation approaches: declarative statements for each adaptation and 10 problem-solving style prompts.
Intersectional HCI recognizes that humans' interconnected social identities shape their experiences with technology. However, intersectional HCI requires extensive resources, such as access to intersectional populations, which many HCI practitioners may lack. For these practitioners, we present an analytical approach to bring intersectional lenses to HCI practices. The approach uses types-not at the level of identities, but at the level of personal traits drawn from foundational research. We first formally prove that certain analytical methods for detecting inclusivity issues can be meaningfully composed to provide equitable consideration of typically overlooked populations; then present four design use-cases to illustrate what the approach brings to HCI practices; and then empirically investigated one of the four use-cases with 24 HCI participants. Results show that practitioners using the compositional approach detected even more intersectional inclusivity problems than those using a complementary intersectional approach.
Software producers are now recognizing the importance of improving their products’ suitability for diverse populations, but little attention has been given to measurements to shed light on products’ suitability to individuals below the median socioeconomic status (SES)—who, by definition, make up half the population. To enable software practitioners to attend to both lower- and higher-SES individuals, this paper provides two new surveys that together can facilitate measuring how well a software product serves socioeconomically diverse populations. The first survey (SES-Subjective) is who-oriented: it measures who their potential or current users are in terms of their subjective SES (perceptions of their SES). The second survey (SES-Facets) is why-oriented: it collects individuals’ values for an evidence-based set of facet values (individual traits) that (1) statistically differ by SES and (2) affect how an individual works and problem-solves with software products. The surveys’ design goal is worldwide applicability, but as a first step, here we empirically validated both these surveys with deployments at University A and University B (464 and 522 responses, respectively), which showed reliability of both the surveys in a US context. Our results also statistically agree with both ground truth data on respondents’ socioeconomic statuses and with predictions from foundational literature. Finally, we explain how the pair of surveys can be uniquely actionable by software practitioners, such as in requirements gathering, debugging, quality assurance activities, maintenance activities, and fulfilling legal reporting requirements such as those being drafted by various governments for AI-powered software.
The GenderMag method has been successfully used by software teams to improve inclusivity in their software products across various domains. Given the success of this method, here we investigate how GenderMag can be systematically adopted in an organization. It is a conceptual replication of our prior work that identified a set of practices and pitfalls synthesized across different USA-based teams. Through Action Research, we trace the 3+ years long journey of GenderMag adoption in the MOSIP organization; starting from the initial ‘unfreeze’ stage to the institutionalization (‘re-freeze’) of GenderMag in the organization's processes. Our findings identify: (1) which practices from the prior work could be generalized and how some of them had to be modified to fit MOSIP organization's context (Digital Public Goods, open-source product, and fully remote work environment), and (2) the pitfalls that occurred.
Inclusive design appears rarely, if at all, in most undergraduate computer science (CS) curricula. As a result, many CS students graduate without knowing how to apply inclusive design to the software they build, and go on to careers that perpetuate the prolif- eration of software that excludes communities of users. Our panel of CS faculty will explain how we have been working to address this problem. For the past several years, we have been integrating bits of inclusive design in multiple courses in CS undergraduate programs, which has had very positive impacts on students' ratings of their instructors, students' ratings of the education climate, and students' retention. The panel's content will be mostly concrete examples of how we are doing this, so that attendees can leave with an in-the-trenches understanding of what this looks like for CS faculty across specialization areas and classes. We also show how it can be used in a department's BPC Plan and point to resources on the CRA's BPCnet Activity Library and on OERcommons, to enable interested faculty to go forward with this approach in their own classes and departments.
Software engineers use code-fluent large language models (LLMs) to help explain unfamiliar code, yet LLM explanations are not adapted to engineers' diverse problem-solving needs. We prompted an LLM to adapt to five problem-solving style types from an inclusive design method, the Gender Inclusiveness Magnifier (GenderMag). We ran a user study with software engineers to examine the impact of explanation adaptations on software engineers' perceptions, both for explanations which matched and mismatched engineers' problem-solving styles. We found that explanations were more frequently beneficial when they matched problem-solving style, but not every matching adaptation was equally beneficial; in some instances, diverse engineers found as much (or more) benefit from mismatched adaptations. Through an equity and inclusivity lens, our work highlights the benefits of having an LLM adapt its explanations to match engineers' diverse problem-solving style values, the potential harms when matched adaptations were not perceived well by engineers, and a comparison of how matching and mismatching LLM adaptations impacted diverse engineers.
Motivation: Many university CS programs have begun teaching various types of CS-related societal issues using approaches such as ethics, Responsible CS, inclusive design, and more. However, some recent research suggests that, although these programs have been able to teach awareness, students often fail to act upon this awareness. To address this problem, University X's CS program tried an unusual approach—integrating hands-on inclusive design skills in small ways across all four years of the CS major. But did it work? That is, did the students who experienced this change across the major actually build more inclusive technology than the students who did not experience it? Objectives: This paper aims to answer this through addressing two research questions: (RQ1): Did students who learned inclusive design across the curriculum act to create more inclusive software? (RQ2): How did inclusivity (or lack thereof) manifest in students’ projects? Method: To investigate these RQs, we conducted a case study of 22 term-long CS projects built by 22 teams consisting of a total of 92 3rd- and 4th-year CS students. Half of the student teams had experienced courses that had integrated inclusive design and the other half had not. The inclusive design elements University X taught were those of the GenderMag inclusive design method, so evaluating the students’ term-long projects was done by GenderMag experts—industry-experienced UX and Software professionals with real-world GenderMag experience. Results: The inclusiveness of students’ projects was higher Post-GenderMag, with fewer reports of inclusivity bugs and higher inclusivity ratings. Experts’ evaluations also revealed the ways in which bias (e.g. bias against risk-averse users) and inclusion (e.g. inclusion of users with diverse information processing styles) appeared in students’ projects. Implications: We believe this to be the first published evidence that compares student-built technology's inclusiveness before vs. after they have been taught inclusive design. These positive results suggest that teaching inclusive design across the curriculum can impact students beyond simply heightening awareness—moving them to act upon this new understanding by building technology that more inclusively serves a wider spectrum of society.
Computer science (CS) and information technology (IT) curricula are grounded in theoretical and technical skills. Topics like equity and inclusive design are rarely found in mainstream student studies. This results in graduates with outdated practices and limitations in software development. A research project was conducted to educate the faculty to integrate inclusive software design into the CS undergraduate curriculum. The objective is to produce graduates with the ability to develop inclusive software. This experience report presents the results of teaching inclusive design throughout the four-year CS and IT curriculum, focusing on the impact on faculty. This easy-to-adopt, high-impact approach improved student retention and classroom climate, broadening participation. Research questions address faculty understanding of inclusive software design, the approach's feasibility, improvement in students' ability to design equitable software, and assessment of the inclusiveness culture for students in computing programs. Faculty attended a summer workshop to learn about inclusive design and update their teaching materials to include the GenderMag method. Beginning in CS0 and CS1 and continuing through Senior Capstone, faculty used updated course assignments to include inclusive design in 10 courses for 44 sections taught. Faculty outcomes are positive, with the planning to include inclusive design and working with other department faculty most engaging. Faculty were impressed by student ownership and adoption of inclusive design methods, particularly in the culminating capstone senior project.
Online computer science (CS) courses have broadened access to CS education, yet inclusivity barriers persist for minoritized groups in these courses. One problem that recent research has shown is that often inclusivity biases (“inclusivity bugs”) lurk within the course materials themselves, disproportionately disadvantaging minoritized students. To address this issue, we investigated how a faculty member can use AID—an Automated Inclusivity Detector tool—to remove such inclusivity bugs from a large online CS1 (Intro CS) course and what is the impact of the resulting inclusivity fixes on the students’ experiences. To enable this evaluation, we first needed to (Bugs): investigate inclusivity challenges students face in 5 online CS courses; (Build): build decision rules to capture these challenges in courseware (“inclusivity bugs”) and implement them in the AID tool; (Faculty): investigate how the faculty member followed up on the inclusivity bugs that AID reported; and (Students): investigate how the faculty member’s changes impacted students’ experiences via a before-vs-after qualitative study with CS students. Our results from (Bugs) revealed 39 inclusivity challenges spanning courseware components from the syllabus to assignments. After implementing the rules in the tool (Build), our results from (Faculty) revealed how the faculty member treated AID more as a “peer” than an authority in deciding whether and how to fix the bugs. Finally, the study results with (Students) revealed that students found the after-fix courseware more approachable - feeling less overwhelmed and more in control in contrast to the before-fix version where they constantly felt overwhelmed, often seeking external assistance to understand course content.
Despite efforts to raise awareness of societal and ethical issues in CS education, research shows students often do not act upon their new awareness (Problem 1). One such issue, well-established by HCI research, is that much of technology contains barriers impacting numerous populations—such as minoritized genders, races, ethnicities, and more. HCI has inclusive design methods that help—but these skills are rarely taught, even in HCI classes (Problem 2). To address Problems 1 and 2, we created the Matchmaker Curriculum to pair CS faculty—including non-HCI faculty—with inclusive design elements to allow for inclusive design skill-building throughout their CS program. We present the curriculum and a field study, in which we followed 18 faculty along their journey. The results show how the Matchmaker Curriculum equipped 88% of these faculty with enough inclusive design teaching knowledge to successfully embed actionable inclusive design skill-building into 13 CS courses.
Artificial Intelligence (AI) is becoming more pervasive through all levels of society, trying to help us be more productive. Research like Amershi et al.'s 18 guidelines for human-AI interaction aim to provide high-level design advice, yet little remains known about how people react to Applications or Violations of the guidelines. This leaves a gap for designers of human-AI systems applying such guidelines, where AI-powered systems might be working better for certain sets of users than for others, inadvertently introducing inclusiveness issues. To address this, we performed a secondary analysis of 1,016 participants across 16 experiments, disaggregating their data by their 5 cognitive problem-solving styles from the Gender-inclusiveness Magnifier (GenderMag) method and illustrate different situations that participants found themselves in. We found that across all 5 cogniive style spectra, although there were instances where applying the guidelines closed inclusiveness issues, there were also stubborn inclusiveness issues and inadvertent introductions of inclusiveness issues. Lastly, we found that participants' cognitive styles not only clustered by their gender, but they also clustered across different age groups.
Assessing an AI system's behavior-particularly in Explainable AI Systems-is sometimes done empirically, by measuring people's abilities to predict the agent's next move-but how to perform such measurements? In empirical studies with humans, an obvious approach is to frame the task as binary (i.e., prediction is either right or wrong), but this does not scale. As output spaces increase, so do floor effects, because the ratio of right answers to wrong answers quickly becomes very small. The crux of the problem is that the binary framing is failing to capture the nuances of the different degrees of "wrongness." To address this, we begin by proposing three mathematical bases upon which to measure "partial wrongness." We then uses these bases to perform two analyses on sequential decision-making domains: the first is an in-lab study with 86 participants on a size-36 action space; the second is a re-analysis of a prior study on a size-4 action space. Other researchers adopting our operationalization of the prediction task and analysis methodology will improve the rigor of user studies conducted with that task, which is particularly important when the domain features a large output space.