
A software requirement indicates a capability or characteristic that a software system must possess to provide value to its stakeholders. It is essential to ensure that the description of the requirements is unambiguous to allow for proper understanding and facilitate its evolution. However, since most software requirements are described in natural language, they may contain subjectivity and inconsistencies in their descriptions, which are conventionally referred to as "Software Requirements Anomalies". Several studies propose tools to aid in the detection of requirements anomalies. However, it can be observed that few of these studies evaluate the effectiveness (recall and precision) of the proposed tools. Therefore, this work presents a comparative study of three anomaly detection tools (RETA, Tactile Check, and Tiger Pro), as well as the ChatGPT model, analyzed based on requirements documents from different domains containing over 85 anomalies. The results show that the Tactile Check tool produced the best performance. Although ChatGPT offers advantages in terms of information visualization and flexibility of interaction, its performance was not satisfactory compared to tools specifically designed for anomaly detection in software requirements. All analyzed tools, including ChatGPT, demonstrated unsatisfactory levels of recall and precision, averaging below 66% and 57%, respectively. These results highlight the need for further contributions in this research area.
This study explores the state of the art of Technical Debt (TD) in software startups, seeking to understand the approaches, methods, and techniques used to manage this issue. TD is a concept that refers to suboptimal technical decisions made to accelerate development, such as implementing code or design that later requires revision to avoid future problems. In startups, where speed and innovation are essential for growth, the pressure to quickly release products often leads to the accumulation of TD. Although this offers short-term benefits, such as faster feature releases, in the long term, it can compromise product quality, increase maintenance costs, and reduce the ability to scale the business. We conducted the study through a systematic literature mapping and analyzed articles from the Scopus, IEEE Xplore, and ACM Digital Library databases. The authors selected fourteen studies published between 2017 and 2024. The analysis revealed that practitioners and researchers still lack standardized practices and tools for efficiently managing TD. Furthermore, the research highlights the need for more empirical studies that consider the specific context of startups, where limited resources and the need for accelerated innovation create unique challenges for TD management. The mapping carried out highlights the fragmentation of current approaches and the lack of a unified framework for managing TD in startups. The authors therefore conclude that there is a need to develop and empirically validate strategic models that integrate technical, human, and business aspects to prove more effective solutions for this specific context.
Accessibility is essential for inclusive digital experiences, yet it is often overlooked in early stages of software development. This paper presents two empirical studies focused on improving how accessibility requirements are specified. In the first study, undergraduate Software Engineering students explored accessibility principles through guided activities and questionnaires. The second study evaluated a structured method combining Personas, User Stories, Behavior-Driven Development (BDD), and WCAG guidelines. Participants applied the method to specify accessibility features in real scenarios. The findings show that although initial knowledge was limited, structured interventions led to more precise, WCAG-aligned, and testable requirements. These results highlight the value of embedding accessibility into software engineering education and demonstrate the effectiveness of combining user-centered design with formal specification techniques.
Context: Micro Frontend (MFE) is an architectural style that extends microservices principles to the frontend. Despite its growing adoption, misunderstandings about MFE foundations can create significant challenges during development. Preparing in-training software engineers to address these challenges and incorporating MFE into software architecture curricula is essential. Goal: We aim to address the gap in MFE education by presenting an experience report on teaching MFE in an undergraduate course. We compare two supporting materials to aid students in architectural decision-making: practitioner-provided guidelines and a catalog of MFE anti-patterns. Through two controlled experiments, we evaluate their effectiveness, with particular emphasis on understanding the role and benefits of using anti-patterns as a learning tool. Method: We taught MFE across five sessions and conducted a controlled experiment with two assessments, each one using one of the supporting materials. We compared them by analyzing differences in assessment scores and evaluated whether the catalog improved students' perceived learning. Additionally, we investigated how students used the catalog by applying the Technology Acceptance Model and collecting qualitative feedback regarding its use. Finally, we extend our previous work by introducing a new version of the catalog and comparing it with the original through a second controlled experiment. Results: Both supporting materials are equally helpful for solving MFE architectural problems. Students reported an increased perception of learning after engaging with the catalog. Their feedback indicated that the catalog was used to identify problems and solutions, promote efficient search for issues, and reinforce MFE knowledge. The results of the second experiment show no statistically significant differences between the versions. However, qualitative feedback indicates that preferences depend on individual reading styles and information needs. It also revealed additional ways students use the catalog to learn about MFE, such as using anti-patterns as a third learning phase. Conclusion: This paper provides insights into teaching MFE, introduces two supporting instructional materials, and highlights the value of anti-pattern catalogs for both education and practice. Our findings show that the catalog can support learning while also helping developers analyze architectural problems and make more informed decisions.
Background: Task interdependence is a core mechanism underlying coordination, performance, and collaboration in team-based work. Kiggundu’s task design theory has shaped how interdependence is conceptualized across knowledge-work settings, including Software Engineering (SE). However, there has been no consolidated synthesis of how Kiggundu’s theory has been empirically applied, operationalized, and adapted in SE and related domains. Objectives: This review aims to (1) systematically synthesize empirical studies that explicitly cite or build on Kiggundu’s task interdependence theory; (2) examine how task interdependence is conceptualized and measured in software development contexts; and (3) compare findings across other knowledge-work domains to identify convergent patterns, boundary conditions, and implications for SE research and practice. Method: We conducted a systematic literature review using forward snowballing from Kiggundu’s foundational publications, identifying 23 eligible empirical studies published between 2013 and 2024. A structured extraction protocol was used to code studies for theoretical framing, conceptualization and measurement of interdependence, analytical role, outcomes, domain, unit of analysis, and methodological characteristics. Open and axial coding supported thematic development, complemented by cross-tabulations and visual mappings to support integrative synthesis. Results: Most studies model task interdependence as a predictor or moderator of outcomes such as team performance, learning, coordination, and relational or affective states. In software development, interdependence is frequently conceptualized as a structural and directional feature of work design, often examined in relation to autonomy, coordination mechanisms, and distributed collaboration. Cross-domain evidence reveals both convergent patterns, such as positive associations with effectiveness under supportive conditions, and important boundary conditions shaped by factors including autonomy, social support, task complexity, and role clarity. A smaller but growing set of studies emphasizes perceived (psychological) interdependence and socio-cognitive or affective pathways, particularly in agile and distributed teams. Conclusions: The findings indicate that Kiggundu’s theory remains a relevant and adaptable framework for analyzing task interdependence in contemporary knowledge work, including software development. The synthesis highlights increasing attention to directionality, perceived interdependence, and emotional–cognitive dimensions, alongside persistent reliance on structural measures. By integrating evidence across domains, this review clarifies how structural task features interact with contextual and psychological factors, and outlines implications for software development practice as well as directions for future empirical research on software teams.
Context: Software companies that adopt agile methods face numerous challenges in sustaining the long-term evolution of software systems. Technical debt is a key contributor to poor maintainability, often leading to failures in agile software projects. This situation becomes even more problematic when managers do not adequately address technical debt items. Objective: This paper proposes the TechDebt Tracker, a method for supporting the documentation and monitoring of technical debt items within technical debt management activities in the context of agile projects. Method: We employed Design Science Research to develop and evaluate the proposal, following the summarized steps of related work review, problem definition, design and development, demonstration, and evaluation. In the related work review, we examined studies with similar research questions across multiple digital libraries. During the design and development phase, we used Design Thinking, the Business Model Canvas, and the Value Proposition Canvas to identify vulnerabilities and opportunities for improvement in the emerging solution. The proposal was demonstrated in a small software company, from which feedback was gathered to refine the method. The evaluation phase consisted of a small-scale study assessed through questionnaires. Results: The proposal comprises three components: two formulas, a kanban board, and a flow. The formulas are used to measure the impact of technical debt on the project and incorporate data such as the developer’s hourly rate, severity, and penalty associated with each technical debt item. The kanban board includes several columns—such as monitoring, technical debt backlog, and testing—as well as a card template used to register and prioritize each technical debt item. The flow consists of states and actions that, when used together with the kanban board, define a method for monitoring technical debt. The company evaluated the proposal positively, highlighting its technical adequacy. Conclusions: We recommend evaluating the proposal in additional contexts and hope that this proposal will be adopted by software companies that employ agile methodologies in diverse scenarios. Our goal is to take a step toward developing a method that can be used in the daily operations of companies to simplify technical debt management.
In recent years, LLM-based AI development platforms have gained widespread adoption, enabling both IT professionals and citizen developers to create AI-powered applications. However, the landscape remains fragmented, with a variety of API-based platforms, AI development frameworks, Low Code/No Code (LCNC) platforms, and Domain-Specific AI as a Service (AIaaS) solution, each offering varying levels of accessibility and customization. Due to the recency of the interest in LLM-based AI development platforms, there is limited systematic research categorizing these tools based on their functionalities and intended user groups. This paper addresses this gap by proposing a structured, feature-based categorization framework, distinguishing between platforms based on criteria such as primary target group, degree of customization, and level of abstraction. Methodologically, we apply a feature-driven analysis grounded in documented capabilities and design affordances across a representative set of tools, and we operationalize the two core dimensions (customization and abstraction) through an anchored ordinal scoring rubric to produce a visual map of categories and overlaps. However, further empirical research is needed to validate the attitude of users towards the different tools in the categories. By providing a clearer understanding of AI development tools, this research supports more informed decision-making and contributes to the democratization of AI adoption across industries.
Code smells are suboptimal structures that undermine software quality. While refactoring is the standard technique to address them, its manual application can degrade code if done without discipline. Despite its importance, refactoring is rarely explored in depth in undergraduate computing courses, creating a gap between academia and industry. Simultaneously, Open Source Software (OSS) projects offer authentic, hands-on learning environments for software maintenance. To address the academic gap and leverage this opportunity, this paper presents and evaluates a hands-on pedagogical approach for teaching code smell refactoring through student contributions to OSS projects. We implemented this approach in two undergraduate Software Quality and Maintenance courses. Our analysis of students' learning experiences reveals that they recognized quality improvements and the connection between refactoring and testing. However, they faced challenges with code complexity and cross-file changes, which sometimes inadvertently introduced new code smells. Regarding the OSS experience, students reported professional growth but struggled with contribution workflows and receiving feedback from maintainers. Our findings offer valuable insights and propose actionable pedagogical recommendations for educators seeking to integrate advanced software maintenance practices into their curricula by leveraging the real-world environment of OSS.
The introduction of the GitHub Discussions feature on GitHub provides a new, dedicated space for collaborative communication within open-source software (OSS) projects. This paper investigates the impact of GitHub Discussions on community engagement, focusing on newcomer onboarding and the activity surrounding issues and pull requests. Through a comprehensive empirical analysis of 285 OSS projects, we observe a significant shift in participation patterns, with GitHub Discussions emerging as a more attractive entry point for newcomers compared to traditional mechanisms like issues and pull requests. Our findings suggest that while GitHub Discussions lower the barrier for newcomer participation, it leads to a decrease in traditional contributions such as new issues and pull requests. Additionally, we explore the engagement levels of different contributor roles, demonstrating how GitHub Discussions foster diverse interactions and community involvement. The results provide critical insights into the evolving nature of collaboration on GitHub and highlight the role of GitHub Discussions in shaping the dynamics of OSS project ecosystems.
Developing the necessary skills in Software Engineering students to conduct effective requirements elicitation interviews is a complex challenge. Immersive role playing has emerged as a promising educational strategy, enabling students to simulate realistic interviews, receive real-time feedback, and improve their performance under pressure. This approach blends traditional role playing with immersive learning environments, providing engaging and authentic experiences that better prepare students for industry demands. This article presents an experience report on the use of an immersive role playing to teach the interview technique for requirements elicitation. Conducted with 86 undergraduate students enrolled in a Requirements Engineering course, the study offers a broader perspective on the effectiveness and challenges of this approach. The findings suggest that immersive activities foster reflection, help identify areas for improvement, and emphasize the importance of emotional regulation in real-world interactions. These insights reinforce that mastering Requirements Engineering requires not only technical proficiency, but also strong interpersonal and emotional skills.
[Context] In Brazil, 41% of companies use machine learning (ML) to some extent. However, several challenges have been reported when engineering ML-enabled systems, including unrealistic customer expectations and vagueness in ML problem specifications. Literature suggests that Requirements Engineering (RE) practices and tools may help to alleviate these issues, yet there is insufficient understanding of RE’s practical application and its perception among practitioners. [Goal] This study aims to investigate the application of RE in developing ML-enabled systems in Brazil, creating an overview of current practices, perceptions, and problems in the Brazilian industry. [Method] To this end, we extracted and analyzed data from an international survey focused on ML-enabled systems, concentrating specifically on responses from practitioners based in Brazil. We analyzed the cluster of RE-related answers gathered from 72 practitioners involved in data-driven projects. We conducted quantitative statistical analyses on contemporary practices using bootstrapping with confidence intervals and qualitative studies on the reported problems involving open and axial coding procedures. [Results] Our findings highlight distinct RE implementation aspects in Brazil’s ML projects. For instance, (i) RE-related tasks are predominantly conducted by data scientists; (ii) the most common techniques for eliciting requirements are interviews and workshop meetings; (iii) there is a prevalence of interactive notebooks in requirements documentation; (iv) practitioners report problems that include a poor understanding of the problem to solve and the business domain, low customer engagement, and difficulties managing stakeholders expectations. Our analysis suggests that development methodology plays a role in these challenges. Agile methods appear to facilitate the management of customer expectations compared to traditional approaches; however, they also appear to introduce greater difficulties in problem understanding and customer involvement. [Conclusion] These results provide an understanding of RE-related practices and challenges in the Brazilian ML industry, helping to guide research and initiatives toward improving the maturity of RE for ML-enabled system projects.
Software development is a complex and knowledge-intensive process that involves multiple participants across various stages of the development lifecycle. Managing knowledge effectively in this context is challenging, especially when it comes to capturing, organizing, and reusing information in software engineering projects. Software development artifacts and techniques play a crucial role in addressing these challenges by facilitating the storage and sharing of knowledge. This study examines how model-based engineering facilitates various aspects of knowledge management. A survey of 62 Brazilian companies was conducted, providing a comprehensive roadmap for the future. Using qualitative and quantitative analyses, the findings were compared with those of other studies. The results indicate that different artifacts effectively support various knowledge management concepts across different phases of software development. Furthermore, while companies predominantly adopt artifact models in the early stages of development and recognize their benefits, they do not fully utilize their potential.
Context: Software engineering (SE) artifacts and documents, such as requirements specifications, user stories, test cases, and concepts of operations (ConOps), are typically written in natural language, making their manipulation challenging. Natural Language Processing (NLP) is a viable solution for managing these tasks. Objective: To conduct a systematic literature review to explore the current use of NLP in SE artifacts and tasks, supplemented by a tertiary study focusing on the emerging role of Large Language Models (LLMs) in software engineering research. Method: We searched digital libraries for relevant papers and applied inclusion and exclusion criteria to filter the primary studies. We then analyzed NLP techniques applied to SE documents and examined their usage in this context. Our research methodology followed Kitchenham and Charters' guidelines. Additionally, we conducted a tertiary study to synthesize findings from existing systematic literature reviews and surveys specifically addressing LLMs in software engineering. Results: We selected 60 primary studies to identify the most common methods for NLP pipelines, feature extraction, language models, and machine learning algorithms used in SE. We also assessed the purposes of these methods, their benefits for SE, their difficulty, and their contribution to SE advancement. The tertiary study revealed a rapid proliferation of LLM-focused research, with comprehensive reviews documenting exponential growth in publications and widespread adoption across diverse SE tasks. Conclusion: Requirements are the most frequently addressed artifacts using NLP techniques, with preprocessing and part-of-speech (POS) tagging being widely used. There is a notable increase in the use of large language models for various SE tasks, such as requirements elicitation, source code generation, bug fixing, and software testing. The tertiary study confirms that LLMs represent a pivotal shift in the research landscape, warranting dedicated investigation to understand their transformative impact on NLP applications in software engineering.
Organizational Change (OC) in the software industry is essential for maintaining competitiveness in dynamic environments. However, OC initiatives often face inefficiencies due to the lack of structured methodologies and reliance on ad hoc practices. This tertiary study analyzes 17 secondary studies related to the OC theme and its instances affecting key organizational pillars, such as technology, processes, and people. The results reveal limited adoption of OC models in software organizations. In response, it synthesizes 41 OC models, 14 critical success factors, and 6 key characteristics, offering a synthesized overview of structured OC frameworks and their relevance to the software industry context, supporting future adaptation efforts by practitioners.
The extensive adoption of microservices architecture by technology companies is driven by its expected advantages, such as scalability, simplicity of development, and resilience, likely due to its cloud-native nature. However, the increasing complexity associated with this architecture can lead to the emergence of microservice smells, analogous to code smells, indicating potential architectural design issues. Despite the identification of numerous microservice smells in the literature, cohesive documentation to support architects detecting and refactoring them. This study conducts a Systematic Literature Review (SLR) to deepen the understanding of these microservice smells, explore detection tools and identify refactoring strategies to mitigate them. We conducted searches across six popular digital libraries, analyzing 27 relevant papers. As a result, we cataloged 104 distinct microservice bad smells, identified 7 detection tools, and compiled refactoring strategies for the most prevalent smells. This documentation aims to assist engineers and architects in identifying and effectively addressing microservice bad smells, thereby enhancing the quality and maintainability of microservice-based systems.
Technical Debt (TD) represents the effort required to address quality issues that affect a software system and progressively hinder code evolution over time. A pull request (PR) is a discrete unit of work that must meet specific quality standards to be integrated into the main codebase. PRs offer a valuable opportunity to assess how developers handle TD and how codebase quality evolves. In this work, we conducted two empirical analyses to understand how developers address TD within PRs and whether TD is effectively managed during PR reviews by both developers and reviewers. We examined 12 Java projects from Apache. The first study employed the SonarQube tool on 2,035 merged PRs to evaluate TD variation, identify the most frequently neglected and resolved types of TD issues, and analyze how TD evolves over time. The second study involved a qualitative analysis of review threads of 250 PRs, focusing on the types of PRs that frequently discuss TD, the characteristics of TD fix suggestions, and the reasons some suggestions are rejected. Our findings reveal that TD issues are prevalent in PRs, following a ratio of 1:2:1 (reduced: unchanged: increased). Among all TD issues, those related to code duplication and cognitive complexity are most frequently overlooked, while code duplication and obsolete code are the most commonly resolved. Regarding PR code review, we found that around 76% of review threads address TD, with code, design, and documentation being the most frequently discussed areas. Additionally, 96% of discussions include a fix suggestion, and over 80% of the discussed issues are resolved. These insights can help practitioners become more aware of TD management and may inspire the development of new tools to facilitate TD handling during PRs.
Context: In the last decade, machine learning (ML) components have become more and more present in contemporary software systems. A number of secondary literature studies reports challenges impacting on the development of ML-based systems, including those for requirements engineering (RE) activities. Motivation/Problem: Synthesizing secondary literature contributes to building knowledge and reaching conclusions about the existing RE approaches for ML-based systems (RE4ML), besides the novelty of a tertiary study on that subject. Objective: Through a tertiary study protocol we elaborated on, this paper synthesizes the body of evidence present in secondary studies on RE4ML systems. Method: We followed well-accepted guidelines about tertiary study protocols, including automatic search, the snowballing technique, selection and quality criteria, and data extraction and synthesis. Results: Nine secondary studies on RE4ML systems were aligned to our tertiary study's goal. We extracted and summarized the requirements elicitation, analysis, specification, validation, and management techniques for ML-based systems as well as the great challenges identified. Finally, we contribute with a nine-item research agenda to direct current and future searches to fill the gaps found. Conclusions: We conclude that RE has not been left aside in ML research, however, there are still challenges to be overcome, such as dealing with non-functional requirements, collaboration between stakeholders, and research in an industrial environment.
Formal verification of smart contracts is widely regarded as an effective method for ensuring correctness and security properties across all possible executions. Its practical relevance has been driven by the availability of automatic verification tools that discharge intricate proofs. Another area of growing interest is the integration of specification paradigms - for example, combining Hoare-logic–style specifications (pre/postconditions and invariants) with SMT and symbolic reasoning - so that each technique can precisely capture complementary aspects of contract behavior. In this article we present a comparative analysis of four leading Solidity verification tools - solc-verify, SMTChecker, VeriSmart and the Certora Prover - and define what is meant here by a formal verification tool: a system that provides mathematically rigorous proofs that stated properties hold for every possible execution of a contract. We also describe a consistent evaluation framework that considers the Solidity version support, the preservation of the original contract structure, the local execution capability, the verification time, and the modeling-language requirements, among other criteria. We used the ERC-20 token standard as a benchmark and applied this framework to obtain empirical evidence of each tool’s capabilities and limitations. Our results expose substantial variability in the tools performances that undermines their trustworthiness in practice and highlights a gap between an academic tool capabilities and the industrial requirements. Finally, we discuss how these findings can inform developers and researchers in selecting appropriate verification tools, thereby contributing to improved smart contract security and reliability.
Software development is an activity characterized by the continuous search for information by programmers. General-purpose search engines, such as Google, Bing, and Yahoo, are widely used by programmers to find solutions to their problems. However, the best solutions are not always among the initial pages of search results that are ranked according to the search engine's algorithms. This study aims to investigate the influence that the order of the pages returned by the ranking of search engines exerts on programmers' performance when performing programming tasks. We designed an empirical within-subject study with programmers to understand and evaluate their performance when solving programming tasks using a ranked list of pages returned by the Google search engine and artificially modified with two different ranking quality levels (Higher Quality Ranking and Lower Quality Ranking). Moreover, for each entry in the ranking of pages, the most frequent methods mentioned on the respective page were listed in the ranking visualization. Analysis of participants' recorded videos was conducted through a mixed-methods research approach to provide insights into the results. We found that programmers spent approximately 8 minutes longer resolving tasks associated with a Lower Quality Ranking, spending more time on irrelevant pages compared to relevant ones, due to efforts to fix problematic code or new searches for another page. The addition of a list of frequent methods in the ranking visualization could help programmers to skip irrelevant pages and reduce time wastage. The ranking quality influences the programmers' performance during the development of programming tasks. Therefore, we suggest the development of filters aimed at improving the quality of results delivered by search engines. Moreover, the results may encourage the adaptation of this study for other approaches that require information foraging, such as chatting with LLMs.
Developers use code comments for various reasons, such as explaining code, documenting specifications, communicating with other developers, and highlighting upcoming tasks. Software projects with minimal documentation often have a significant number of comments. In this sense, the code comment analysis technique can be used to examine more complex aspects of software projects, such as TD generated by merge conflicts. The TD resulting from the resolution of merge conflicts occurs when the resulting code contains comments that indicate tasks to be performed in the future. There are no studies directly linking merge conflicts and TD. The purpose of this study is to identify and analyze code comments generated when resolving merge conflicts from this perspective. This process can lead to improvements in software quality and help in managing TD. To this end, an exploratory analysis was carried out in 100 software projects, with a specific focus on task annotations originating from the resolution of merge conflicts. The results revealed that 60.61% of the analyzed projects have at least one code comment indicating the creation or maintenance of TD. In addition, metrics such as accuracy, precision, recall, and F1-score were applied across different software projects to enhance the effectiveness and reliability of the built data model. The metric results suggest that the tool was able to correctly classify in most cases, but was not particularly precise in this classification due to the size of the analyzed projects and the variety of task annotations, including the presence of some and the absence of others.