
User involvement is essential for aligning software systems with organizational needs in agile development; however, achieving meaningful engagement becomes challenging in large-scale settings. Large organizations may contain distributed teams and diverse user groups whose limited availability and unfamiliarity with agile practices hinder sustained involvement. Although participatory and user-centered approaches work in smaller contexts, they are challenging to scale, leaving a gap in how to organize user involvement across dispersed organizations and multiple development teams. To address this gap, the present study examines pilot testing in a large-scale agile environment. Drawing on a qualitative case study of a Norwegian public organization, based on interviews with 35 participants from development and user groups, the study defines iterative pilot testing as a scalable practice in which unfinished yet operational solutions are refined gradually through user involvement when functionality is novel or complex. The study demonstrates how this practice can make development work more visible, enable development teams to be closer to their users and their work practices, and facilitate teams in navigating cultural and functional differences, thereby strengthening user involvement in large-scale agile software development.
Software development in critical healthcare domains presents unique challenges due to the non-deterministic nature of Artificial Intelligence (AI). Traditional agile frameworks often face friction when aligning the exploratory cycles of model training with fixed sprint cadences. This paper presents a case study on the development of a mobile application for diabetes patient support, powered by a specialized Large Language Model (LLM). The project adapted Scrum ceremonies to manage the full AI lifecycle, from corpus preparation to clinical validation. We identify four key challenges, specifically the insufficiency of automated metrics (e.g., BERTScore) to ensure medical safety and the limitations of standard “Definition of Done” (DoD) in experimental contexts. As a solution, we implement “hybrid sprints” incorporating “question stories” and multi-level acceptance criteria that prioritize clinical expert validation over technical scores. The results demonstrate that while standard agile methods provide a foundation, they require pragmatic decoupling of data and model lifecycles to ensure safety-critical value delivery.
While Generative AI (GenAI) has rapidly transformed software development and GenAI tools receive widespread adoption, how humans and AI should collaborate has become a focal point for Agile practitioners, whose values emphasise teamwork and collaboration. Yet, despite growing interest, research on human-AI collaboration in software engineering, especially within Agile contexts, remains limited. Our study investigates Agile practitioners’ perceptions of collaborating with GenAI in software development activities. We conducted a survey with 73 Agile practitioners, revealing how they currently collaborate with GenAI across various activities and how they expect this collaboration to evolve. Key findings include that GenAI is currently not a substantial part of real-world workflows, being either unused or limited to an assistant role, but that there is a clear tendency towards greater AI involvement in the future, with practitioners increasingly viewing AI as a collaborator rather than merely a tool. These findings advance our understanding of human-AI collaboration in Agile settings and can guide both future research and the practical adoption of GenAI in Agile environments.
Agile Requirements Engineering (Agile RE) has been widely studied, but gaps remain regarding how user-centered research practices, specifically UX Research, are integrated into Agile RE. This paper investigates how UX Research is applied within Agile RE activities and identifies the positive and negative aspects of this integration. We conducted a multiple-case study with four distinct organizations, collecting data through interviews with 12 agile and UX professionals. Our results revealed that UX Research practices and tools are most frequently employed to discover requirements, both in organizations with and without UX professionals. However, the findings also indicate that although UX Research practices provide valuable input for requirements specification, there remains a need to better structure and communicate UX-related aspects so they become more visible to the agile team in all Agile RE activities. Our study contributes by discussing how UX Research practices and tools are employed across the stages of the Agile RE cycle, as well as the roles of the different stakeholders involved in these stages in different companies.
Agile coaches play a key role in supporting software teams and organisations during agile transformations, yet little is known about how they understand and design for the continuation of agility after their involvement ends. This study fills that research gap by examining how agile coaches perceive post-exit outcomes and describing the strategies they use to support endurance beyond their engagement. Drawing on sixteen semi-structured interviews with professional agile coaches across multiple organisations, we identify six interconnected themes that capture coaches’ accounts of embedding routines, transferring ownership, and aligning agile values with organisational contexts. We synthesise these themes into the concept of sustainment work, referring to the professional and organisational efforts that coaches describe undertaking to support the continuation of agile ways of working beyond their direct involvement. This synthesis is represented in an empirically derived model comprising three interrelated domains: Embedding Practices, Enabling Ownership, and Aligning Culture. The model reflects how coaches conceptualise the relationship between their interventions, anticipated post-exit trajectories, and the endurance of agile practices, rather than providing evidence of long-term organisational outcomes. The paper contributes theoretically grounded insights into how agility is understood and approached by highlighting coaches’ perceptions and intended sustainment strategies, as an ongoing organisational concern rather than a time-bound coaching intervention.
Team autonomy—the extent to which teams can make decisions—is a central tenet of agile development. When scaling agile, technical and coordination challenges necessitate restricting autonomy. While greater autonomy is often considered better, the reasons for, and the extent to which, autonomy is, or should be, restricted in large-scale agile development remain poorly understood. In this paper, we address this gap by analyzing large-scale agile frameworks using decision-making theory. Applying inductive thematic analysis to the primary documents of scaling frameworks, we identify and categorize their decisions and related decision-making elements by decision-maker, decision type, decision-making style, decision timing, and affected artifacts. This research contributes to decision-making knowledge by providing a systematic categorization of decision points in large-scale agile development. For practitioners, the resulting classification helps identify key decision points, define roles and responsibilities, and plan team autonomy in large-scale agile.
Architectural uncertainties arising from incomplete or unclear information pose significant challenges when making architectural decisions in Agile teams. Based on a limited number of case studies that employed a technique called ArchHypo, four patterns were identified that propose small adjustments in the development process to handle architectural uncertainties: Protective Guideline, Bring the Specialist, Plan for Preparation, and Quality Checkpoint. Although the patterns derived from these experiences can be useful in real projects, their applicability and consequences were based on limited evidence and specific scenarios. To address this issue, this paper presents an interview study with experienced software architects and engineers to gather further information on the application of these patterns. The research method employed semi-structured interviews to gather the experiences of professionals with the target practices, and thematic analysis was used to assess their recurrence, applicability, and consequences. The findings confirmed that most professionals recognized those practices in real projects and their suitability as actions in uncertainty management. Moreover, new positive and negative consequences, not previously documented in the patterns, were identified. As a result, this work contributes to the field by providing guidance to professionals on how to better evaluate the trade-offs of those patterns when applied to architecture uncertainty management.
Hybrid work has become the norm post pandemic, yet organizing it effectively in agile software engineering requires understanding how individual and contextual factors shape preferences and behaviors. This study surveyed 65 agile practitioners to explore these influences. Respondents generally preferred flexible office attendance policies, but not full-time remote work. Larger teams expressed interest in agreed on presence at the office, while small or medium size teams leaned slightly more toward individual choice. Social interaction was a strong reason to come to the office, and most respondents viewed in-person meetings as highly beneficial for social and problem-solving activities. In virtual meetings, camera use was guided mainly by social presence and visibility, while multitasking activities were mostly meeting-supportive and work related, with managers and agile leaders reporting slightly higher levels of multitasking and significantly higher levels of camera use.
Agile at scale introduces persistent tensions as organizations attempt to balance autonomy, coordination, and control across teams and layers. Some tensions are paradoxes, which are persistent contradictions in which both sides matter. Drawing on a qualitative case study of MarketCorp, a multinational marketplace company, we explore how paradoxes manifest in a large-scale agile context and how leaders respond. Interviews with 15 top managers and reflective conversations with 60 employees reveal three paradoxes: centralization vs. autonomy, internal vs. external focus, and product vs. project approaches. We explored two approaches for designing facilitated dialogue around paradoxes in leadership teams. The study highlights the importance of recognizing, normalizing, and collaboratively managing paradoxes and offers practical guidance for large-scale agile organizations, demonstrating that paradoxes can provide a valuable focus for continuous improvement and strengthen agile learning practices.
Although agile teams aim to maximize value, few studies have examined how teams collectively develop an understanding of value beyond the Product Owner’s perspective. While Product Owners often define value priorities, team members’ understanding of how daily work connects to business goals and broader organizational context often remains limited. To address this gap, this study introduces and evaluates a value retrospective process, a structured reflection on cost and value. Informed by Design Science Research principles, the process was jointly designed by the first author, acting as Agile Lead, and the company’s finance function. The process was implemented across 61 cross-functional teams in a large Finnish conglomerate, and iteratively refined over two years. This study examines process design, organizational implementation, and participant experiences using semi-structured interviews. The findings show that participants initially perceived retrospectives as potentially judgmental or audit-like, but over time, these perceptions shifted toward psychologically safe, constructive dialogue. Value retrospectives supported increased awareness of costs and value, strengthened team-sponsor alignment, and prompted concrete improvement actions. This study contributes an empirically grounded process model and insights into how structured value reflection can support learning and alignment in agile teams.
Experimentation has become a cornerstone of agile, evidence-driven software development. To explore how it may evolve in the coming decade, we collected qualitative data from 58 experts across the global experimentation community. Using an inductive thematic analysis inspired by the Gioia methodology, we identified six interrelated trends that capture how practitioners envision the future of experimentation: AI-augmented workflows, segment-level personalization, platformization and warehouse-native architectures, expansion beyond web contexts, rigor at scale, and cultural capability building. These trends highlight experimentation’s evolution from a technical testing practice toward a socio-technical learning system. Building on these insights, the paper outlines a practice-inspired research agenda for the next decade of evidence-based product development.
Agile software development often faces challenges related to architectural uncertainty. ArchHypo addresses this by providing a hypothesis-driven architecture technique that helps teams formulate, test, and learn from architectural assumptions iteratively. However, studies have shown that the lack of tooling integrated into everyday development workflows is a significant barrier to its adoption. In this paper, we present ArchHypo.AI, a Trello plugin that integrates hypothesis-driven architectural reasoning directly into agile boards and validates the ArchHypo technique in real projects. The plugin uses an LLM and a RAG mechanism to help teams generate and classify architectural hypotheses, develop technical plans, and link actions to architecture decision patterns. By operating on top of an existing project management tool, ArchHypo.AI aims to lower the adoption cost of hypothesis engineering and to make architectural decision processes more observable. We evaluated ArchHypo.AI in a controlled study with software professionals working on a realistic architecture scenario. The results indicate that the plugin helps structure architectural discussions, reduces manual effort in documenting hypotheses and plans, clarifies procedural steps, and surfaces differences in risk perception within teams. Qualitative feedback suggests that AI-assisted support facilitates collaborative reasoning about architecture. Our findings show that LLM-based tools can effectively support hypothesis-driven architecture in agile settings and highlight design considerations for integrating such tools into existing workflows.
To cope with ongoing change, an increasing number of companies view agile approaches as fundamental for remaining competitive in dynamic environments. However, realizing the outcomes associated with agility requires not only the implementation of agile practices but also employees who possess an agile mindset. Although scholarly understanding of the agile mindset has increased in recent years, a validated measurement instrument for individuals is still lacking. This paper addresses this research gap by following the validation procedure proposed by MacKenzie, Podsakoff, Podsakoff [1]. Based on a systematic review of existing conceptualizations of the agile mindset, we empirically developed a scale and specified a corresponding measurement model. To quantitatively evaluate and validate the model, we analyzed data from four independent samples (content validity: n = 184; pretest: n = 500; validation: n = 202; cross-validation: n = 218). As a result, we validate a four-dimensional conceptualization of the agile mindset comprising the attitude toward learning, collaborative exchange, empowered self-guidance, and customer co-creation. Furthermore, we demonstrate that the agile mindset is correlated with key individual outcomes: job satisfaction, goal orientation, and reduced resistance to change and thereby reinforcing its relevance. In addition to its theoretical contribution, this study provides practical value by offering an economically efficient instrument for measuring the agile mindset.
In agile software development, regression testing is an ongoing process activity. However, real-world time constraints necessitate selecting a subset of tests to run. Current regression test selection algorithms primarily focus on technical metrics such as requirement coverage while overlooking the business value each test validates. This study reframes regression test selection as a multi-objective optimization problem in which tests are selected to maximize business value within a constrained testing time while maintaining adequate requirement coverage. We apply an artificial intelligence-based search algorithm, and our results show that the proposed method consistently selects tests with higher business value than baseline approaches when time is limited, while maintaining comparable requirement coverage. These findings suggest that a driven value-aware selector can be incorporated into agile teams for decisions on allocating limited regression testing resources.
Background: Ensuring the quality of user stories is vital to Agile Software Development. Rule-based tools like AQUSA, based on the Quality User Story (QUS) framework, offer reliable structural checks but struggle with context-sensitive or pragmatic issues. Large Language Models (LLMs) have emerged as potential alternatives, yet prior studies often rely on small datasets, older models, or lack direct comparison with rule-based baselines. Objective: This study aims to assess the effectiveness of modern LLMs relative to a rule-based tool (AQUSA) for detecting defects in user stories, considering both structural and contextual dimensions. Method: We conduct a large-scale comparative evaluation involving AQUSA and three GPT-family LLMs (GPT-5, GPT-5-mini, and GPT-4), using 182 user stories drawn from three industrial datasets. We apply both quantitative metrics (precision, recall, F1-score) and qualitative analysis of feedback clarity and defect relevance. Results: GPT-5-mini achieved the highest recall (0.81) and overall F1-score (0.62), while AQUSA attained the highest precision (0.61) with significantly fewer false positives. GPT-5 showed high hallucination rates and instability; GPT-4 was overly conservative, leading to under-detection of defects. Conclusion: Neither rule-based nor GPT-family LLM-based approaches suffice in isolation. Rule-based tools enforce structural rigor, while LLMs capture nuanced linguistic and pragmatic flaws. We advocate a hybrid “Dual-gate” strategy—using AQUSA for structural validation followed by lightweight LLMs for contextual refinement—to improve the reliability and scalability of user story quality assessment in agile environments.
Scrum is widely adopted in software project management due to its adaptability and collaborative nature. The recent emergence of Large Language Models (LLMs) has created new opportunities to support knowledge-intensive Scrum practices. However, existing research has largely focused on technical activities such as coding and testing, with limited evidence on the use of LLMs in management-related Scrum activities. In this study, we investigate the use of LLMs in Scrum management activities through a survey of 70 Brazilian professionals. Among them, 49 actively use Scrum, and 33 reported using LLM-based assistants in their Scrum practices. The results indicate a high level of proficiency and frequent use of LLMs, with 85
Context: The rapid emergence of generative AI (GenAI) tools has begun to reshape software engineering practice. Yet, their adoption within agile environments remains underexplored. Objective: This study investigates how agile practitioners adopt GenAI tools in real-world organizational contexts, focusing on regulatory conditions, role-specific use cases, benefits, and barriers. Method: An exploratory multiple case study was conducted in three German organizations, involving 17 semi-structured interviews and document analysis. A cross-case thematic analysis was applied to identify GenAI adoption patterns. Results: Findings reveal that GenAI is primarily used for creative tasks, documentation, and code assistance. Benefits include efficiency gains and enhanced creativity, while barriers relate to data privacy, validation effort, and lack of governance. Using the Technology-Organization-Environment (TOE) framework, we find that these barriers stem from misalignments across the three dimensions. Regulatory pressures are often translated into policies without accounting for actual usage patterns or organizational constraints, leading to systematic policy-practice gaps and shadow IT behavior. Conclusion: GenAI offers significant potential to augment agile roles but requires alignment across TOE dimensions, including pragmatic policies, data protection measures, and user training to ensure responsible and effective integration.
In Agile software development, maintaining velocity requires the continuous management of Technical Debt (TD). However, the rapid iteration cycles inherent to Agile often obscure debt accumulation, making manual identification in issue trackers prohibitively expensive. To address this, we present TD-Suite, a comprehensive framework engineered to automate the classification of technical debt. It leverages state-of-the-art transformer models to analyze textual artifacts, such as developer discussions in issue reports, where subtle indicators of debt often lie hidden. TD-Suite provides a seamless end-to-end pipeline suitable for Agile ML Engineering, managing everything from initial data ingestion and rigorous preprocessing to model training, thorough evaluation, and final inference. It supports both binary classification (debt or no debt) and granular categorization—identifying code, architecture, design, or documentation debt—enabling Agile teams to formulate targeted refactoring strategies. To ensure robustness on real-world, imbalanced datasets, TD-Suite incorporates k-fold cross-validation, early stopping, and class weighting strategies. The framework explicitly integrates the tracking and reporting of carbon emissions associated with model training. Furthermore, it features a user-friendly Gradio web interface within a Docker container, simplifying integration into DevOps pipelines and democratizing access for practitioners without deep ML expertise.
This paper proposes the adaptation of Retrieval Augmented Generation (RAG) based systems into Agile Software Development workflows, addressing the need for specialized evaluation methods that align with Agile's iterative, feedback driven nature. Through a systematic mapping study of 20 papers, we identify existing evaluation frameworks for RAG systems and explore how they can be applied within Agile settings. The research highlights the role of Evaluator-Agents in aligning the responsiveness of RAG systems, with a focus on Multi-agent AI systems that are more efficient at handling complex, distributed tasks than Single-agent systems. The findings contribute to identify the potential of RAG and its evaluation in facilitating real time feedback and continuous improvement for Agile development.
In the wake of the COVID-19 pandemic, remote working seemed here to stay. However, attempts to return to working in the office are currently being made, mainly by large multinationals. We therefore set out to investigate software-industry workers' opinions of their productivity in their current work mode. This paper presents the results of surveys applied to software engineers from Mexico, Spain, Colombia and Argentina with the purpose of analysing how they perceive their productivity and that of their team when working either partially or totally in a remote manner. The results indicate that most of the respondents prefer to work remotely. Furthermore, according to their perception, their productivity and that of their team is, in general, equal or superior to that achieved in the office.