
Context: Despite the popularity of Agile software development, achieving consistent project success remains challenging. Furthermore, identifying what contributes to project success is inconsistent across studies. In addition measuring success is incoherent, and existing SLR's are not systematic, explicit, and comprehensive making it hard to replicate investigations. Objective: This systematic literature review (SLR) identifies critical success factors (CSFs) in Agile projects to improve outcomes. Method: We analyzed 53 primary studies published between 2001 and 2024, employing thematic synthesis with content analysis to uncover recurring themes and patterns. Results: Our analysis yielded 21 CSFs categorized into five themes: organizational, people, technical, process, and project. These factors range from management dedication and organizational environment to team effectiveness, emotional mindset, and Agile practices. Team effectiveness and project management emerged as the most frequently cited CSFs, highlighting the importance of people and process factors to Agile project outcomes. Conclusions: These interpreted and integrated themes and factors contributed to developing a theoretical framework to identify how these factors contribute to project success. This study offers valuable insights for researchers and practitioners, guiding future research to validate these findings and develop targeted interventions to enhance Agile project success.
Context: LLM-based Text-to-SQL has advanced quickly, but benchmark and training datasets may contain defects that distort evaluation and fine-tuning. Prior audits remain fragmented, addressing dimensions in isolation. Objective: This paper proposes text2sql-dataset-analyzer, an open-source framework that audits Text-to-SQL datasets across five complementary quality dimensions within a single reproducible pipeline. Method: The framework covers five dimensions: database schema integrity, SQL syntactic structure and complexity, execution testing, antipattern detection, and semantic correspondence. Semantic correspondence is evaluated via an LLM-as-a-judge committee with majority voting. An analytical database stores the resulting metrics for direct querying and Markdown report generation, while structured JSONL output records per-item annotations. Results: Auditing all 11,840 Spider 1.0 examples reveals quality issues despite 99.97% of queries passing execution checks. Schema and data checks identify 51 structural foreign-key errors and 41,927 row-level referential-integrity violations, 41,913 of them in just three of 206 databases. At the item level, the audit flags unanimously Incorrect (8-10%), disputed (23-27%), and Unanswerable NL-SQL pairs, together with SQL antipatterns associated with correctness, robustness, and portability concerns. Manual review of 367 flagged items confirms genuine defects in 263, including 26 consensus Unanswerable questions and all 17 Cartesian-product join bugs. Conclusions: Multi-dimensional validation exposes dataset defects missed by executability-based checks. We release the open-source framework and a prioritized remediation roadmap for a popular Text-to-SQL benchmark. The downstream impact of these defects on model training and benchmark scores remains future work.
Context: The post-ChatGPT surge has rapidly reframed Information Systems (IS) research and practice. As organizations and society grapple with Generative AI (GenAI) adoption, a body of secondary studies and research agendas has emerged to synthesize early evidence and chart directions for future inquiry. Objective: This study conducts a systematic literature review of secondary studies and research agenda/roadmap papers to synthesize the state of knowledge on GenAI's benefits and challenges in IS, and to identify future research directions. Method: We performed a systematic search across Scopus, Web of Science, and the AIS eLibrary for publications from 2023 onwards. Following a rigorous, multi-stage screening process, we selected a final set of 28 papers (18 secondary studies and 10 research agendas) for analysis using bibliometric mapping and thematic analysis. We also conducted a quality assessment of all sources to gauge confidence in each source's contribution to the findings. Results: GenAI offers transformative potential to drive productivity, accelerate innovation, personalize services, and democratize access to expertise. However, its adoption is constrained by interrelated challenges: technical unreliability (hallucinations, performance drift), societal-ethical risks (bias, malicious misuse, skill erosion), and a governance vacuum (privacy, accountability, intellectual property). Conclusions: Interpreted through a socio-technical lens, our findings reveal a persistent misalignment between GenAI's fast-evolving technical subsystem and the slower-adapting social subsystem, positioning IS research as critical for achieving joint optimization. To bridge this gap, we propose a research agenda that reorients IS scholarship from analyzing impacts toward actively shaping the co-evolution of technical capabilities with organizational routines, societal values, and regulatory institutions -- emphasizing hybrid human-AI ensembles, situated validation, design principles for probabilistic systems, and adaptive governance. For practitioners and policymakers, responsible adoption requires balancing automation with human augmentation alongside transparent governance and adaptive regulations to ensure broadly shared benefits.
Context: Large-scale agile (LSA) is inherently characterized by socio-technical complexities (e.g., system dependencies, extensive cross-team coordination, and distributed organizational structures). Onboarding of newcomers is a critical challenge in LSA environments and it remains underexplored compared to small-scale agile settings. Objective: Our study aims to investigate the onboarding processes and challenges within LSA projects. Method: In this exploratory qualitative study, we conducted 41 semi-structured interviews across four Swedish software companies. Results: We identified onboarding issues that result in negative outcomes, such as cognitive overload, psychological safety constraints, and cross-team collaboration hurdles. These factors create steep learning curves and knowledge deficiencies that contribute to long-term socio-technical debt. Conclusions: We provide actionable recommendations for practitioners to improve their onboarding processes. In particular, we recommend designing competency-based rather than time-based onboarding plans, expanding social integration through structured cross-team exposure, and addressing the gap between textbook agile and real-world LSA practices through explicit expectation setting.
Context: Software requirements outline customer expectations for their software and are critical for successful project outcomes. As software becomes increasingly complex due to its size and diverse features, it is vital to prioritize these requirements to utilize development resources effectively. To address this challenge, researchers are exploring new strategies and improved solutions using artificial intelligence (AI) tools. In our systematic literature review conducted in 2023, we found that existing requirements prioritization techniques predominantly rely on human input and have several limitations. These include inaccuracies, overlapping results, scalability issues, and excessive time consumption. These challenges can be mitigated by incorporating key features from AI-based software requirements prioritization techniques. Objective: The primary goal of this study is to develop a Hybrid AI-based requirements prioritization technique named as “KPSO-Fuzzy that effectively balances user preferences and technical dependencies. Method: We propose a Hybrid prioritization method called Hybrid KPSO-Fuzzy. First, we use K-Means clustering to group technically dependent functional requirements into three distinct clusters. In the next phase, we apply the Particle Swarm Optimization (PSO) algorithm along with Fuzzy Logic to prioritize each cluster simultaneously. Clusters that achieve higher accuracy will be selected to create the most effective prioritized lists of requirements. We conducted extensive experiments to validate our proposed approach, focusing on accuracy, scalability, and computation time. Results: The comparative analysis shows that our proposed Hybrid KPSO-Fuzzy method outperforms PSO, Fuzzy Logic, Multi-objective Artificial Bee Colony optimization and Hybrid Genetic Algorithms in terms of accuracy, scalability, and efficiency. Conclusions: Overall, this study clarifies the application domains for various AI-based techniques, enabling users to maximize their benefits.
Context: Recent advances in LLM-based diagram generation increasingly rely on coordinated agent systems rather than single-model prompts. Objective: This work highlights how modular multi-agent architectures improve reliability, semantic grounding, and iterative refinement in text-to-diagram workflows. Method: We analyze a pipeline composed of specialized agents for interpretation, synthesis, validation, and correction, each contributing a bounded and inspectable transformation. Results: The agent system provides deterministic validation, structured reasoning, and controlled refinement loops that outperform monolithic LLM generation. Conclusions: Multi-agent LLM pipelines represent a robust foundation for precise, verifiable diagram generation and serve as a reproducible alternative to single-pass text-to-diagram models.
Context : In the fast-growing API Economy scenario, API Gateways became an essential tool for microservice reliability and scalability, yet lack standardized quality models and measures to assess these tools in an objective fashion. Objective : This study quantitatively and qualitatively assessed the accuracy and relevance of a set of 59 metrics, gathered from prior analysis of 68 mainstream API Gateway software offerings. Method : An expert judgment evaluation study was conducted with seven domain specialists (N=7) from both academia and industry, who assessed each metric's accuracy and relevance using 5-point Likert scales. A multi-statistical method was performed to assess ratings, inter-observer consistency and reliability. Results : Although ratings were generally neutral to positive, statistical analysis revealed low consensus across experts. A subset of 24 core metrics with strong agreement and high scores was identified, alongside moderate- and low-agreement sets. These items cover key API management concerns, including traffic monitoring, error handling, latency, and resource usage. Conclusions : By combining practical definitions, expert judgment, and a tiered classification mechanism, this work established a broad and comprehensive set of metrics that will help to bridge the gap between metric adoption in real-world platforms and their theoretical grounding in software quality models.
Context: Sustainable tourism platforms require methodologies balancing stakeholder inclusivity with technical rigor. Existing approaches either emphasize participatory design at the cost of formalization or concentrate on technical precision while neglecting community voices, creating a critical gap in developing Web-3D frameworks that honor local knowledge while maintaining engineering excellence.Objective: We present a replicable methodology integrating Design Thinking, Software Domain Architecture (SDA), and Model-Driven Engineering (MDE) to bridge participatory engagement with formal software specification, validated through DIVEXPLORE-3D, a Web-3D platform for sustainable marine tourism with 54 stakeholders in Gorontalo, Indonesia.Method: The methodology comprises five integrated phases: (1) participatory workshops with 54 stakeholders across three SetupFloatingEnvironmentoups (tourists, community members, operators) using empathy mapping; (2) requirements formalization through SDA into four architectural layers; (3) model transformation from PIMs to PSMs via MDE; (4) bidirectional validation through a RTM; and (5) pilot deployment with technical and educational evaluation.Results: The validation of the methodology resulted in four evidence streams: (1) the four-layer architecture achieved 96% coverage of requirements across 68 specifications; (2) an expert evaluation (n=8) rated the sustainability adequacy of the methodology's specifications at 4.6/5.0; (3) a pilot deployment confirmed technical feasibility (mean latency 73ms, 97.2% uptime); and (4) educational testing (n=67) showed significant gains in conservation knowledge, with a Cohen's d of 1.52 (p<0.001). Additionally, community content moderation successfully published 78% of content without requiring expert modification.Conclusions: This study presents a replicable methodology integrating participatory co-design with formal software engineering. Instead of being opposing paradigms, they are complementary processes when mediated by structured traceability mechanisms. The DIVEXPLORE-3D framework promotes sustainable digital tourism through a modular, community-driven platform balancing ecological conservation, cultural preservation, and technical scalability.
Context: Clone detection is a common task in software engineering. Type-3 clones are fragments of code that can be slightly different in structure. Objective: The article presents a~new algorithm for Type-3 clone detection, its open-source implementation called DrDupLex3, and novel open-source tools that can be used for the automated assessment of Type-3 clones and to prepare training sets for machine learning-based clone detectors.Method: The algorithm for Type-3 clone detection builds upon the index of source code used in DrDupLex, the most accurate Type-2 clone detector to date.Results: A comparison with three state-of-the-art clone detectors (NiCad, CloneWorks, and SourcererCC) shows that DrDupLex3 is able to outperform them in precision, recall, and running time. It reported no false positives and found all clones reported by NiCad, CloneWorks, and SourcererCC.Conclusions: The presented clone detector outperforms three state-of-the-art competitors in a~scenario that can be easily repeated because it is based on tools for automated assessment of Type-3 clones.
Context: While conceptual research on AI in project management is advancing, empirical evidence of actual usage among IT project managers practising is limited. Objective: We investigate how Finnish IT project and programme managers use generative AI in daily work, identifying practices, organisational constraints, and future visions through individual interviews. Method: Using qualitative descriptive design, we conducted semi-structured interviews with 12 experienced Finnish IT project management consultants who work in multiple client organisations. Data were analysed through hybrid inductive-deductive thematic analysis to identify usage practices and future usage. Results: GenAI adoption is fragmented and peripheral, mainly constrained by organisational policies rather than individual resistance. Participants use GenAI mainly for support tasks rather than core project management functions. Despite varied usage, participants converge on envisioning AI as an 'assistant not replacement,' reflecting professional boundary work that preserves human authority. Conclusions: The adoption of GenAI in IT project management is limited by organisational constraints such as security policies and governance structures, which means that organisations should prioritise integration and data protection over bottom-up experimentation. More advanced capabilities remain aspirational, making incremental adoption through low-risk use cases more realistic than broad automation of core project management functions.
Context: Organizations adopting Artificial Intelligence (AI) face challenges in eliciting and analyzing requirements that align with strategic objectives, especially when human oversight and iterative refinement are needed. Large Language Models (LLMs)-based Multi-agent systems provide a~potential solution by supporting structured and collaborative Requirements Engineering (RE) processes for AI adoption planning. Objective: The objective of this study is to investigate whether a multi-agent system, built on LLMs and supported by human input, can assist in requirements analysis for AI adoption. Method: We used a mixed-method approach: (i) designed and developed a~multi-agent system to support the generation and prioritization of requirements for AI adoption, (ii) conducted multiple case studies with four companies to evaluate the system, and (iii) collected data through post-session questionnaires from nine participants and follow-up interviews, one per company. Results: Questionnaire and interview findings together indicate that the system may assist in identifying relevant and goal-aligned requirements. Seven participants considered the generated requirements relevant, and six found them aligned with organizational goals. Participants noted that iterative feedback improved completeness and feasibility, often within two feedback rounds. Both data sources show that human input was essential to clarify technical details, ensure contextual accuracy, and validate prioritization results. Participants from all companies also identified usability, transparency, and scalability as areas requiring further refinement for broader organizational use. Conclusions: LLM-based multi-agent systems can support strategic AI planning by enabling iterative refinement with human experts. Future work will include more interviews with stakeholders and adjustments to system features to improve transparency, usability, and scalability.
Context: The state-of-the-art and practice on software quality is growing constantly and presents several challenges. The ever-growing body of knowledge on the topic, obfuscates the situation further, the lack of explicit structure makes it difficult to identify which properties exist today and how they can be evaluated, two critical aspects to increase software quality. Objective: A step to disambiguate software properties descriptions is made via a Property Model Ontology (PMO) and steps towards the evaluation of the structure are carried out together with experts. The objectives of this paper are: 1) present in detail the PMO, 2) describe the research process used to develop and evaluate the PMO, and, 3) exemplify the usage of the PMO through real instantiations obtained from practitioners and researchers. Method: Expert interviews and qualitative research methods are used to evaluate the PMO and develop a proof of concept. Results: The PMO consists of concepts describing extra-functional properties (EFPs) and their evaluation methods, i.e., how to measure the properties. The PMO is instantiated in a modelling environment through a metamodel and an online web-content management system. Conclusions: Consensus on the definition and structure of EFPs is achieved and a common understanding on how they can be reused in practice.
Context: Artificial Intelligence (AI) is increasingly integrated into critical domains, making defect analysis essential to ensure system quality and reliability. Current defect classification frameworks do not adequately address the unique properties of AI systems. Objective: This paper proposes AIODC, a defect classification framework inspired by the Orthogonal Defect Classification (ODC) that incorporates AI-specific characteristics. Method: The framework extends ODC by introducing three new attributes – Data, Learning, and Thinking – and adds a “Catastrophic” severity level to account for risks associated with AI. Additionally, it modifies impact mapping utilizing AI/AIP quality models. The methodology was validated through a case study that examined 42 actual Keras defects. Results: This study demonstrated the feasibility of modifying ODC for AI systems to classify its defects. The case study indicated that defects occurring during the Learning phase are the most prevalent and were significantly linked to high severity, whereas defects in the Thinking phase primarily impact trustworthiness and accuracy. Conclusions: The results affirm the practicality and significance of AIODC in identifying high-risk defect categories, thus facilitating more focused and effective quality assurance strategies in AI-driven software systems.
Context: In software engineering, the presence of code smells is closely associated with increased maintenance costs and complexities, making their detection and remediation an important concern. Objective: Despite numerous deep learning approaches for code smell detection, many still heavily rely on feature engineering processes (metrics) and exhibit limited performance. To address these shortcomings, this paper introduces CSDXR, a novel approach for enhancing code smell detection based on Random Convolutional Kernel Transform---a state-of-the-art technique for time series classification. The proposed approach does not rely on a manual feature engineering process and follows a three-step process: first, it converts code snippets into numerical sequences through tokenization; second, it applies Random Convolutional Kernel Transform to generate pooled models from these sequences; and third, it constructs a classifier from the pooled models to identify code smells. Method: The proposed approach was evaluated on four real-world datasets and compared against four state-of-the-art methods---DeepSmells, AE-Dense, AE-CNN, and AE-LSTM---in detecting Complex Method, Multifaceted Abstraction, Feature Envy, and Complex Conditional smells. Results: Empirical results demonstrate that CSDXR outperformed the four state-of-the-art methods---DeepSmells, AE-Dense, AE-CNN, and AE-LSTM---in detecting Complex Method and Multifaceted Abstraction smells. Specifically, the enhancement rates in terms of F1 score were 1.99% and 6.09% for Complex Method and Multifaceted Abstraction smells, respectively. In terms of MCC, the improvement rates were 0.82% and 35.64% for these two smells, respectively. The results also show that while DeepSmells achieves superior overall performance on Feature Envy and Complex Conditional smells, CSDXR surpasses AE-Dense, AE-CNN, and AE-LSTM in detecting these two types of smells. Conclusions: The paper concludes that the proposed approach, CSDXR, demonstrates significant potential for effectively detecting various types of code smells.
Background: Poor communication of requirements between clients and suppliers contributes to project overruns,in both software and infrastructure projects. Existing literature offers limited insights into the communication challenges at this interface. Aim: Our research aim to explore the processes and associated challenges with requirements activities that include client-supplier interaction and communication. Method: we study requirements validation, communication, and digital asset verification processes through two case studies in the road and railway sectors, involving interviews with ten experts across three companies. Results: We identify 13 challenges, along with their causes and consequences, and suggest solution areas from existing literature. Conclusion: Interestingly, the challenges in infrastructure projects mirror those found in software engineering, highlighting a need for further research to validate potential solutions.
Context: Action research is popular in software engineering due to its industrial nature and promises of effective technology transfers. Yet, the methodology is still gaining popularity, and guidelines for conducting quality action research studies are needed. Objective: This paper aims to collect, summarize, and discuss guidelines for conducting action research in academia-industry collaborations. The guidelines are designed for researchers and practitioners alike. Method: I use existing guidelines for empirical studies and my own experiences to define guidelines for researchers and host organizations for conducting action research. Results: I identified 22 guidelines for conducting action research studies. They provide actionable recommendations on identifying the relevant context, planning and executing interventions (actions), reporting them, and reasoning around the ethics of action research. Conclusions: The paper concludes that the best way of engaging with action research is when we can be embedded in the host organization and when the collaboration leads to tangible change in the host organization and the generation of new scientific results.
Background: Early identification of software vulnerabilities is an intrinsic step in achieving software security. In the era of artificial intelligence, software vulnerability prediction models (VPMs) are created using machine learning and deep learning approaches. The effectiveness of these models aids in increasing the quality of the software. The handling of imbalanced datasets and dimensionality reduction are important aspects that affect the performance of VPMs. Aim: The current study applies novel metaheuristic approaches for feature subset selection. Method: This paper performs a comparative analysis of forty-eight combinations of eight machine learning techniques and six metaheuristic feature selection methods on four public datasets. Results: The experimental results reveal that VPMs productivity is upgraded after the application of the feature selection methods for both metrics-based and text-mining-based datasets. Additionally, the study has applied Wilcoxon signed-rank test to the results of metrics-based and text-features-based VPMs to evaluate which outperformed the other. Furthermore, it discovers the best-performing feature selection algorithm based on AUC for each dataset. Finally, this paper has performed better than the benchmark studies in terms of F1-Score. Conclusion: The results conclude that GWO has performed satisfactorily for all the datasets.
Background: With the rapid proliferation of question-and-answer websites for software developers like Stack Overflow, there is an increasing need to discern developers’ emotions from their posts to assess the influence of these emotions on their productivity such as efficiency in bug fixing. Aim: We aimed to develop a reliable emotion classification tool capable of accurately categorizing emotions in Software Engineering (SE) websites using data augmentation techniques to address the data scarcity problem because previous research has shown that tools trained on other domains can perform poorly when applied to SE domain directly. Method: We utilized four machine learning techniques, namely BERT, CodeBERT, RFC (Random Forest Classifier), and LSTM. Taking an innovative approach to dataset augmentation, we employed word substitution, back translation, and easy data augmentation methods. Using these we developed sixteen unique emotion classification models: EmoClassBERT- -Original, EmoClassRFC-Original, EmoClassLSTMOriginal, EmoClass- CodeBERT-Original EmoClassLSTM-Substitution, EmoClassBERT-Substitution, EmoClassRFC-Substitution, EmoClassCodeBERT-Substitution, Emo- ClassBERT-Translation, EmoClassLSTM-Translation, EmoClassRFC-Translation, EmoClassCodeBERT-Translation, EmoClassBERT-EDA, EmoClass- LSTM-EDA, EmoClassCodeBERT-EDA, and EmoClassRFC-EDA. We compared the performance of this model on a gold standard state-of-the-art database and techniques (Multi-label SO BERT and EmoTxt). Results: An initial investigation of models trained on the augmented datasets demonstrated superior performance to those trained on the original dataset. EmoClassLSTM-Substitution, EmoClassBERT-Substitution, EmoClassCodeBERT-Substitution, and EmoClassRFC-Substitution models show improvements of 13%, 5%, 5%, and 10% as compared to EmoClass- LSTM-Original, EmoClassBERT-Original, EmoClassCodeBERT-Original, and EmoClassRFC-Original, respectively, in average F1 score. The Emo- ClassCodeBERT-Substitution performed the best and outperformed the Multi-label SO BERT and Emotxt by 2.37% and 21.17%, respectively, in average F1-score. A detailed investigation of the models on 100 runs of the dataset shows that BERT-based and CodeBERT-based models gave the best performance. This detailed investigation reveals no significant differences in the performance of models trained on augmented datasets and the original dataset on multiple runs of the dataset. Conclusion: This research not only underlines the strengths and weaknesses of each architecture but also highlights the pivotal role of data augmentation in refining model performance, especially in the software engineering domain.