Context: Many students now use generative AI (genAI) in their coursework, yet its effects ontheir intellectual development remain poorly understood. While prior work has investigatedstudents’ cognitive offloading during episodic interactions, it remains unclear whether usinggenAI routinely is tied to more fundamental shifts in students’ thinking habits.Objectives: To explore this possibility, we investigate (RQ1-How): how students’ trust in androutine use of genAI affect their cognitive engagement—specifically, reflection, the need forunderstanding, and critical thinking in STEM coursework. Further, we investigate (RQ2-Who):which students are particularly vulnerable to these cognitive disengagement effects.Methods: We drew on dual-process theory, cognitive offloading, and the automation biasliterature to develop a statistical model explaining how and to what extent students’ trust-driven routine use of genAI affected their cognitive engagement habits in STEM coursework, and how these effects differed across students’ diverse cognitive styles. We empirically evaluated this model using Partial Least Squares Structural Equation Modeling on survey data from 299 STEM students across five North American universities.Results: Students who trusted and routinely used genAI reported significantly lower cognitiveengagement. Unexpectedly, students with higher technophilic motivations, risk tolerance, andcomputer self-efficacy—traits often celebrated in STEM—were more prone to these effects.Interestingly, students’ prior experience with genAI or academia did not protect them fromcognitively disengaging.Conclusion: Our findings suggest a potential cognitive debt cycle in which routine genAIuse progressively weakens students’ intellectual habits, potentially driving over-reliance andescalating usage.This poses critical challenges for curricula and genAI system design, requiring interventions that actively support cognitive engagement.
Generative AI (GenAI) is rapidly reshaping software development workflows. While prior studies emphasize productivity gains, the adoption of GenAI also introduces new pressures that may harm developers' well-being. In this paper, we investigate the relationship between the adoption of GenAI and developers' burnout. We utilized the Job Demands–Resources (JD–R) model as the analytic lens in our empirical study. We employed a concurrent embedded mixed-methods research design, integrating quantitative and qualitative evidence. We first surveyed 442 developers across diverse organizations, roles, and levels of experience. We then employed Partial Least Squares–Structural Equation Modeling (PLS-SEM) and regression to model the relationships among job demands, job resources, and burnout, complemented by a qualitative analysis of open-ended responses to contextualize the quantitative findings. Our results show that GenAI adoption heightens burnout by increasing job demands, while job resources and positive perceptions of GenAI mitigate these effects, reframing adoption as an opportunity.
Claims that generative AI will soon write all of the code have led to predictions that programming is nearing its end. In this vision paper, we argue against this assumption that broader access to code generation necessarily democratizes software development, i.e., everyone can code but we have to distinguish between access and control: by access, we mean the ability of more people, including non-experts and less-experienced developers, to generate code-like artifacts with AI; by control, we mean the capacity to inspect, evaluate, integrate, maintain, and govern those artifacts as dependable software. While AI may broaden access to code production, control may become more concentrated among those who own or understand the code, software practices, infrastructure, evaluation practices, and deployment pipelines. Grounded in an expert panel, our vision paper argues that AI does not eliminate software engineering expertise but shifts where that expertise becomes most critical. The locus of software engineering expertise is shifting toward intent specification: orchestrating and governing AI behavior, evaluating software behavior, and integrating software systems. We conclude this paper by identifying research opportunities for education, tools, and policy that can help the software engineering community respond to the AI era with greater agency, accountability, and adaptability.
Recruiting, retaining, and educating students in computing is a frequent research topic in CHI. However, students’ sociotechnical experiences of registering for classes are understudied—especially those of socioeconomic-diverse students. These experiences matter: research shows that registration problems bring long-term consequences to student successes. We investigate students’ socioeconomic status (SES) impact on registration experiences through three studies: a case study with education professionals using an emerging analytic method, SocioeconomicMag (SESMag); interviews with faculty/staff/students from 8 universities; and observations of 14 SES-diverse students registering for classes. Results showed: (1) 5 SES-inclusivity bugs which arose 30 times, 72% more often by lower-SES students than by higher-SES students. (2) 6/7 lower-SES students (but only 2/7 higher-SES students) expected downstream problems from the registration issues. (3) The risk-to-negative-outcomes rate was 3 times higher for lower-SES students. (4) The issues generalized across 8 universities and potentially to >700 other universities who use the same registration portal.
Objectives: When students use generative AI in coursework, what are its persistent effects on their intellectual development? We investigate (RQ1-How) how students' trust in and routine use of genAI affect their cognitive engagement habits in STEM coursework, and (RQ2-Who) which students are particularly vulnerable to cognitive disengagement. Method: Drawing on dual-process, cognitive offloading, and automation bias theories, we developed a statistical model explaining how and to what extent students' trust-driven routine genAI use affected their cognitive engagement – specifically, reflection, the need for understanding, and critical thinking in coursework, and how these effects differed across students' cognitive styles. We empirically evaluated this model using Partial Least Squares Structural Equation Modeling on survey data from 299 STEM students across five North American universities. Results: Students who trusted and routinely used genAI reported significantly lower cognitive engagement. Unexpectedly, students with higher technophilic motivations, risk tolerance, and computer self-efficacy – traits often celebrated in STEM – were more prone to these effects. Interestingly, students' prior experience with genAI or academia did not protect them from cognitively disengaging. Implications: Our findings suggest a potential cognitive debt cycle where routine genAI use weakens students' intellectual habits, potentially driving and escalating over-reliance. This poses challenges for curricula and genAI system design, requiring interventions that actively support cognitive engagement.
Developers spend roughly one-tenth of their workday writing code, yet most AI tooling targets that fraction. This paper asks what should be built for the rest. We surveyed 860 Microsoft developers to understand where they want AI support, and where they want it to stay out. Using a human-in-the-loop, multi-model council-based thematic analysis, we identify 22 AI systems that developers want built across five task categories. For each, we describe the problem it solves, what makes it hard to build, and the constraints developers place on its behavior. Our findings point to a growing right-shift burden in AI-assisted development: developers wanted systems that embed quality signals earlier in their workflow to keep pace with accelerating code generation, while enforcing explicit authority scoping, provenance, uncertainty signaling, and least-privilege access throughout. This tension reveals a pattern we call "bounded delegation": developers wanted AI to absorb the assembly work surrounding their craft, never the craft itself. That boundary tracks where they locate professional identity, suggesting that the value of AI tooling may lie as much in where and how precisely it stops as in what it does.
Open source software (OSS) communities are facing increasing pressure from Generative AI (GenAI) tools. We call it AI-DDoS: a denial-of-service effect in which plausible but low-quality AI-generated contributions overwhelm OSS community capacity. Using a phenomenon-based mixed-methods approach, we first analyze practitioner accounts from Reddit, OSS mentor mailing lists, and blogs to identify six recurring themes and derive hypotheses. We then evaluate these hypotheses using Bayesian Structural Time Series analysis across 294 repositories with over 2 million pull requests and issues. Our results show that while PR volume increased in 2025, merge rates declined, with one-time contributors experiencing an 18.18
Context: The growing adoption of AI-assisted development tools is changing how software teams collaborate, share knowledge, and coordinate, yet its consequences for team social dynamics remain largely unexplored. Gap: It is unclear whether AI adoption is associated with an increase or reduction in community smells,socio-technical anti-patterns reflecting coordination and communication breakdowns,and through which mechanisms. Method: Grounded in Transactive Memory Systems (TMS) theory, we validate instruments for HumanAI and HumanHuman interaction along two TMS dimensions, Specialization and Coordination, and test five PLS-SEM models on survey data from 152 software professionals using AI tools. Community smell constructs were derived from the literature and validated through expert surveys and factor analysis. Results: AI adoption relates to community smells not in a single way, but through mechanisms depending on the work. In specialization work, AI is associated with higher knowledge-sharing peer interaction, which is in turn associated with fewer smells. In coordination work, AI is directly associated with higher communication quality, complementing rather than replacing human interaction. Contributions: We provide an empirically validated, TMS-grounded model showing that the AIcommunity-smell relationship is contingent on the type of collaboration, with a reusable instrument and evidence-based implications for research and practice.
Generative AI (genAI) tools promise productivity gains, yet developers still struggle to determine when to trust and effectively integrate these tools into their everyday work. Moreover, genAI can be exclusionary, failing to adequately support developers across individual differences. One such difference is cognitive style , which can shape how developers engage with genAI (e.g., risk-averse developers may gate outputs behind tests, whereas risk-tolerant ones may prototype directly and address issues post hoc). When tools fail to accommodate these differences, they can create additional usability barriers. Thus, to design tools that developers intend to use, we must understand which factors shape developers’ trust in and adoption of genAI at work . We developed a theoretical model of developers’ trust and adoption of genAI using a large-scale survey ( N = 238) conducted at GitHub and Microsoft. Using Partial Least Squares Structural Equation Modeling (PLS-SEM), we found aspects related to genAI’s system and output quality (e.g., presentation, safety/security, and performance), functional value (e.g., educational and practical benefits), and goal maintenance (ability to sustain alignment with task goals) to be significantly associated with trust. Trust, alongside developers’ cognitive styles (i.e., risk tolerance, technophilic motivations, and computer self-efficacy), was in turn significantly associated with adoption intentions, which ultimately were associated with reported usage. An Importance-Performance Matrix Analysis (IPMA) identified high-importance factors for which existing genAI tools provided insufficient support, highlighting actionable targets for design improvements. We bolstered these findings through a qualitative analysis of developers’ reported challenges and risks of genAI use, uncovering why these gaps persisted in development contexts. Our study offers practical guidance for designing genAI tools that support trustworthy and inclusive developer–AI interactions.
Context: The success and sustainability of open source software (OSS) projects depend not only on visible technical contributions but also on glue work, the often invisible and underappreciated efforts that hold projects and communities together. Despite their importance, glue work remains inconsistently defined, difficult to trace, and rarely acknowledged. Objectives: This study aims to develop an empirically grounded and theoretically informed understanding of glue work in OSS: what it is, where it appears, and how it can be traced and recognized. Methods: We conducted a multiple-case study across diverse OSS contexts using interviews, focus groups, and surveys. Drawing on invisible labor theory, we developed a taxonomy of glue work structured along sociocultural (what it is), sociospatial (where it appears), and sociolegal (how it is acknowledged) dimensions. We operationalized the taxonomy through glue work Tracker, a modular, extensible, multi-agent GitHub bot prototype, and conducted an expert evaluation with OSS practitioners. Results: The taxonomy identifies four categories and 12 types of sustaining OSS work and specifies how such contributions can be traced and recognized across development and community channels. The expert evaluation shows that the taxonomy and prototype align with real recognition needs and are feasible to integrate into project workflows. Conclusion: This work provides both a theoretical foundation and a practical infrastructure for making sustaining labor visible and supporting OSS sustainability.
Onboarding documentation is critical for attracting and retaining newcomers in open source software (OSS). However, it is often presented as dense, inconsistently structured, and fragmented presentations that are difficult to understand, which creates cognitive overload leading to frustration, errors, and abandonment. Here, we investigate how Cognitive Theory of Multimedia Learning (CTML) strategies can be used to restructure OSS documentation. We use a GenAI-based pipeline to operationalize these strategies to restructure OSS documentation through our prototype VisDoc. VisDoc segments documentation into task-based units, infers workflows, removes redundancy, and generates multimodal explanations. An expert evaluation (N=4) affirmed VisDoc's completeness, accuracy, and adoptability; A between-subjects evaluation (N=14) with newcomers found that VisDoc participants achieved higher task success, had significantly lower cognitive load, and perceived higher usability. The contributions of this work include a CTML-grounded analysis of onboarding challenges, a GenAI-based documentation restructuring pipeline, and empirical evidence that cognitively informed documentation restructuring reduces cognitive load and improves usability and task performance in OSS.
Open-source software (OSS) community managers face significant challenges in retaining contributors, as they must monitor activity and engagement while navigating complex dynamics of collaboration. Current tools designed for managing contributor retention (e.g., dashboards) fall short by providing retrospective rather than predictive insights to identify potential disengagement early. Without understanding how to anticipate and prevent disengagement, new solutions risk burdening community managers rather than supporting retention management. Following the Design Science Research paradigm, we employed a mixed-methods approach for problem identification and solution design to address contributor retention. To identify the challenges hindering retention management in OSS, we conducted semi-structured interviews, a multi-vocal literature review, and community surveys. Then through an iterative build-evaluate cycle, we developed and refined strategies for diagnosing retention risks and informing engagement efforts. We operationalized these strategies into a web-based prototype, incorporating feedback from 100+ OSS practitioners, and conducted an in situ evaluation across two OSS communities. Our study offers (1) empirical insights into the challenges of contributor retention management in OSS, (2) actionable strategies that support OSS community managers' retention efforts, and (3) a practical framework for future research in developing or validating theories about OSS sustainability.
Generative AI (GenAI) tools are increasingly being adopted in software development as productivity aids, since there is evidence that GenAI tools can improve individual aspects of productivity. However, productivity is multidimensional; accelerating one aspect of work may simply shift effort to another. In this paper, we investigate how GenAI adoption affects different dimensions of developer productivity. We surveyed 415 software practitioners to understand how they perceive productivity changes associated with AI adoption, using the SPACE framework (Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow). Our results reveal systematic redistribution of effort across SPACE dimensions. While frequent GenAI users reported faster task completion and higher output volume, these gains were offset by increased code review burden, persistent cognitive load from output verification, and unchanged collaboration patterns. We further provide an empirical mapping between the challenges perceived by developers and potential strategies to mitigate them. Overall, our findings suggest that, at the current stage of GenAI adoption, perceived productivity gains may be spurious – surface-level acceleration, often accompanied by redistributed effort and hidden costs.
As AI takes on more software work, the line between human and AI effort is shifting. Where developers draw that line around AI autonomy bears on how we design tools and roles that preserve meaningful work. Drawing on cognitive appraisal theory, work design, and automation research, we conducted a mixed-methods study of 448 professional developers at Microsoft to investigate their accepted levels of AI autonomy across software engineering work. Most developers accepted AI producing work under their oversight, although accepted autonomy varied substantively across tasks and individuals. Acceptance was lowest for identity-defining, human-facing, and design-oriented work, and higher among developers with more AI experience and risk tolerance. Task accountability was associated with lower odds of allowing AI to act on developers' behalf, whereas task identity was associated with lower odds of granting AI decision-making autonomy. Task demands had the opposite effect, increasing willingness to delegate decision-making to AI. Our findings suggest that preferences for AI autonomy reflect how developers cognitively experience their work, highlighting important considerations for designing meaningful work.
Our previous research showed that lack of belonging is associated with higher levels of burnout in software developers. We revisit this topic to offer guidelines on promoting a sense of belonging in software development teams, ultimately improving developers' well-being.
The sustainability of open source software (OSS) projects hinges on contributor retention. Interpersonal challenges can inhibit a feeling of welcomeness among contributors, particularly from underrepresented groups, which impacts their decision to continue with the project. How much this impact is, varies among individuals, underlining the importance of a thorough understanding of their effects. Here, we investigate the effects of interpersonal challenges on the sense of welcomeness among diverse populations within OSS, through the diversity lenses of gender, race, and (dis)ability. We analyzed the large-scale Linux Foundation Diversity and Inclusion survey (n = 706) to model a theoretical framework linking interpersonal challenges with the sense of welcomeness through Structural Equation Models Partial Least Squares (PLS-SEM). We then examine the model to identify the impact of these challenges on different demographics through Multi-Group Analysis (MGA). Finally, we conducted a regression analysis to investigate how differently people from different demographics experience different types of interpersonal challenges. Our findings confirm the negative association between interpersonal challenges and the feeling of welcomeness in OSS, with this relationship being more pronounced among gender minorities and people with disabilities. We found that different challenges have unique impacts on how people feel welcomed, with variations across gender, race, and disability groups. We also provide evidence that people from gender minorities and with disabilities are more likely to experience interpersonal challenges than their counterparts, especially when we analyze stalking, sexual harassment, and doxxing. Our insights benefit OSS communities, informing potential strategies to improve the landscape of interpersonal relationships, ultimately fostering more inclusive and welcoming communities.
Generative AI (genAI) tools, such as ChatGPT or Copilot, are advertised to improve developer productivity and are being integrated into software development. However, misaligned trust, skepticism, and usability concerns can impede the adoption of such tools. Research also indicates that AI can be exclusionary, failing to support diverse users adequately. One such aspect of diversity is cognitive diversity -- variations in users' cognitive styles -- that leads to divergence in perspectives and interaction styles. When an individual's cognitive style is unsupported, it creates barriers to technology adoption. Therefore, to understand how to effectively integrate genAI tools into software development, it is first important to model what factors affect developers' trust and intentions to adopt genAI tools in practice? We developed a theoretically grounded statistical model to (1) identify factors that influence developers' trust in genAI tools and (2) examine the relationship between developers' trust, cognitive styles, and their intentions to use these tools in their work. We surveyed software developers (N=238) at two major global tech organizations: GitHub Inc. and Microsoft; and employed Partial Least Squares-Structural Equation Modeling (PLS-SEM) to evaluate our model. Our findings reveal that genAI's system/output quality, functional value, and goal maintenance significantly influence developers' trust in these tools. Furthermore, developers' trust and cognitive styles influence their intentions to use these tools in their work. We offer practical suggestions for designing genAI tools for effective use and inclusive user experience.
Novice programming students frequently engage in help-seeking to find information and learn about programming concepts. Among the available resources, generative AI (GenAI) chatbots appear resourceful, widely accessible, and less intimidating than human tutors. Programming instructors are actively integrating these tools into classrooms. However, our understanding of how novice programming students trust GenAI chatbots-and the factors influencing their usage-remains limited. To address this gap, we investigated the learning resource selection process of 20 novice programming students tasked with studying a programming topic. We split our participants into two groups: one using ChatGPT (n=10) and the other using a human tutor via Discord (n=10). We found that participants held strong positive perceptions of ChatGPT's speed and convenience but were wary of its inconsistent accuracy, making them reluctant to rely on it for learning entirely new topics. Accordingly, they generally preferred more trustworthy resources for learning (e.g., instructors, tutors), preferring ChatGPT for low-stakes situations or more introductory and common topics. We conclude by offering guidance to instructors on integrating LLM-based chatbots into their curricula-emphasizing verification and situational use-and to developers on designing chatbots that better address novices' trust and reliability concerns.
Generative AI (genAI) tools (e.g., ChatGPT, Copilot) have become ubiquitous in software engineering (SE). As SE educators, it behooves us to understand the consequences of genAI usage among SE students and to create a holistic view of where these tools can be successfully used. Through 16 reflective interviews with SE students, we explored their academic experiences of using genAI tools to complement SE learning and implementations. We uncover the contexts where these tools are helpful and where they pose challenges, along with examining why these challenges arise and how they impact students. We validated our findings through member checking and triangulation with instructors. Our findings provide practical considerations of where and why genAI should (not) be used in the context of supporting SE students.