Diabetic retinopathy (DR) grading from fundus images remains challenging because severity labels are ordinal, class distributions are imbalanced, and neighboring stages are often visually ambiguous. This study presents a unified pipeline for five-class DR grading that integrates controlled CNN/Transformer backbone benchmarking, ensemble evaluation, computational-cost analysis, external validation, and proof-of-concept attribution reporting. APTOS 2019 dataset was used to benchmark six backbone architectures under stratified five-fold cross-validation with fixed fold assignments, identical preprocessing, imbalance-aware training, and model selection based on quadratic weighted kappa (QWK). Hard voting, weighted soft voting, stacking, and hybrid class-level fusion were investigated to combine complementary models, while trainable parameters, model size, inference latency, and throughput were reported to quantify deployment cost. Model generalization was assessed by directly evaluating APTOS-trained models on Messidor-2 without retraining, fine-tuning, or recalibration. Safety-filtered summaries generated with a vision-language model were restricted to non-diagnostic model-attribution descriptions, and fundus-masked Grad-CAM++ maps were paired with these summaries for attribution reporting. Among individual backbones, ResNet-50 and ConvNeXt-Tiny achieved the strongest performance, while weighted soft voting provided the most stable ensemble performance. External validation confirmed the difficulty of cross-dataset five-class DR grading, with lower performance than that observed during internal cross-validation. The generated summaries adhered to predefined safety constraints but were not intended for lesion localization, diagnosis, or clinical validation. Overall, the study demonstrates that controlled benchmarking and weighted ensembling can provide stable performance for ordinal DR grading, while computational overhead, domain shift, uncertainty calibration, and expert clinical validation remain important considerations for real-world deployment.
The quality of diabetic retinopathy (DR) screening relies on the ability to correctly grade severity; however, many deep-learning (DL) classifiers cannot be easily interpreted in the clinical context. This study presents a methodology that combines strong discriminative models with multimodal explanations, converting retinal pixels into clinically interpretable outputs. Using the APTOS 2019 benchmark, we evaluated six representative CNN- and transformer-based backbones under a controlled protocol with stratified five-fold cross-validation. We then compared ensembling strategies (hard voting, weighted soft voting, stacking) and investigated a hybrid class-level fusion variant to exploit grade-specific advantages. For interpretability, we produced Grad-CAM++ visual attribution maps and short textual rationales using vision-language models (VLMs) conditioned on the fundus image and classifier outputs under conservative prompting constraints. Modern CNN backbones (ResNet-50 and ConvNeXt-Tiny) provided the strongest single-model baselines, with cross-validated QWK up to 0.919 and 0.914, respectively. Ensembling improved ordinal agreement, and weighted soft voting was the most consistent across folds (QWK 0.934 +/- 0.017). Hybrid class-level fusion was competitive but did not yield a statistically reliable improvement over standard fusion in paired fold comparisons (Holm-adjusted p >= 1.000). For explanation quality, Grad-CAM++ offered plausible but coarse localization, and VLM rationales were generally grade-consistent. Quantitatively, VLM variants showed a trade-off between clinical completeness and template-level semantic similarity (coverage 0.700 vs. BERTScore 0.072), while image-text alignment was comparable (CLIPScore approximately 0.34).
Diabetic retinopathy (DR) screening requires artificial intelligence (AI) models that are not only highly accurate in grading five clinical stages but are also capable of generating quantitatively evaluated lesion-aware explanations to earn the trust of clinicians. We propose RobustDRNet, a hybrid ensemble model that combines local convolutional features from Residual Network-34 (ResNet-34) and ConvNeXt-Tiny with global transformer embeddings from Vision Transformer Base/16 (ViT-B16) via two-stage feature fusion and a disentangled multilayer perceptron (MLP), followed by a logistic regression stacking meta-learner for prediction aggregation. To address severe class imbalance, our training pipeline employs stratified sampling, contrast-limited adaptive histogram equalization (CLAHE) for contrast enhancement, strong data augmentation, and class-weighted focal loss. Evaluated on the Asia Pacific Tele-Ophthalmology Society (APTOS) 2019 dataset, RobustDRNet achieved 88.4% validation accuracy, a 0.967 macro-averaged area under the receiver operating characteristic curve (macro-AUC), and Cohen's kappa of 0.823, outperforming individual backbones and simple voting ensembles. In addition to classification performance, we integrated six complementary explainable AI (XAI) techniques: Gradient-weighted Class Activation Mapping++ (Grad-CAM++), Integrated Gradients, attention rollout, SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), and Testing with Concept Activation Vectors (TCAV). Each technique was quantitatively benchmarked against expert-annotated lesion maps from the Indian Diabetic Retinopathy Image Dataset (IDRiD). Saliency maps achieved mean Intersection over Union (IoU) scores of 0.06 for Grad-CAM++ and approximately 0.10 for Integrated Gradients; SHapley Additive exPlanations (SHAP) perturbations showed a deletion drop of 0.25 and an insertion gain of 0.22; and TCAV achieved complete classifier-level TCAV alignment (score = 1.0) with clinically coherent, grade-wise importance trajectories. By combining competitive grading performance with multi-perspective and quantitatively evaluated interpretability, RobustDRNet provides a promising DR screening framework whose decisions are supported by lesion-aware explanatory evidence.
Nowadays, virtual worlds are evolving into immersive metaverses, i.e., collective, shared virtual spaces that arise from the convergence of virtual reality (VR), augmented reality (AR), the internet, and additional digital technologies. These environments allow users to interact through avatars. While research has explored the technologies and social dynamics of the metaverse, two key limitations remain. First, there is no systematic overview of the expertise required to build metaverses and the key publication venues. Second, the sociotechnical issues have not been fully synthesized or made instrumental for developing holistic design approaches that address both social and technical constraints. A deeper investigation is needed to guide future research, highlight challenges, and foster collaboration across disciplines. In this article, we propose a systematic mapping study that addresses the limitations identified. From an initial pool of 2,323 sources, we identify 63 primary resources to (1) characterize the research community around the metaverse and (2) elicit, synthesize, and categorize 19 social and 18 technical issues affecting the development of an effective metaverse. Based on our results, we contextualize the catalog of issues with respect to the current body of knowledge, providing insights and a research roadmap that transforms issues into actionable challenges in the scope of a novel, unified research asset coined "metaverse engineering", i.e., the multidisciplinary discipline to identify processes and instruments to design socio-technical metaverses.
The metaverse has transitioned from a science fiction term to a rapidly growing area of research and application, with potential uses in education, professional training, social events, and the virtual economy. However, despite this progress, a fully realized and functional metaverse is not yet available, and its development still requires a clear understanding and definition of the research directions to follow. Nonetheless, the metaverse topic does not start from scratch; it shares its foundations with Virtual Environments (VEs), which represent the Virtual Reality applications' core. In this paper, we built on top of the knowledge given by the history of the two topics to present a scoping review of the historical development of research in VEs and the metaverse from the 1990s to early 2024, analyzing 352 papers from the Scopus database. We aimed to offer a comprehensive understanding of how past research informs the present and future directions of the metaverse. Our findings revealed that the metaverse, while emerging as a distinct research area in recent years, is deeply rooted in the history of VEs, with many of its concepts and technologies deriving from earlier work in the field. We also identified new, underexplored trends within the metaverse research, proposing a future research agenda informed by the shared history of the two topics.
This work implements a blockchain-based framework for educational metaverses within SENEM, an immersive learning environment. Ethereum and IPFS were integrated to support decentralized uploading, storage, and retrieval of lecture materials. The resulting prototype formalizes requirements for secure content management and possibly improve integrity, traceability, and privacy. Overall, this work offers a first step toward transparent immersive learning systems and a proof-of-concept for future usability research.
Instructor effectiveness is fundamental to student learning, with the ability to manage student inquiries serving as a critical component of effective teaching. Student questions represent a valuable training resource for instructors to strengthen their teaching strategies, yet interactions with students are often constrained by several factors. In this paper, we investigate how instructors perceive machine- and studentgenerated questions, considering the potential for the former to complement the latter in a cost-effective manner. Our study involved 121 undergraduate students and an equivalent number of simulated students modeled using a state-of-the-art large language model, generating over 360 questions in total based on video lectures given by seven university instructors. We assessed whether instructors could distinguish between human- and machine-generated questions and how they evaluated their relevance, clarity, answerability, challenge level, and cognitive depth. Results show that instructors struggle to differentiate between the two sets of questions, with accuracy close to random chance. Instructors tended to (i) rate machine-generated questions slightly higher in relevance, clarity, answerability, and challenge-though only relevance and answerability showed significant differences-and (ii) associate them marginally more often with higher-order cognitive skills. This confirms the potential of machine-generated questions as tools for instructor training. Repository: https://github.com/tail- unica/realistic-ai-generatedquestions.
Diabetes mellitus (DM) is a global health issue of significance that must be diagnosed as early as possible and managed well. This study presents a framework for diabetes prediction using Machine Learning (ML) models, complemented with eXplainable Artificial Intelligence (XAI) tools, to investigate both the predictive accuracy and interpretability of the predictions from ML models. Data Preprocessing is based on the Synthetic Minority Oversampling Technique (SMOTE) and feature scaling used on the Diabetes Binary Health Indicators dataset to deal with class imbalance and variability of clinical features. The ensemble model provided high accuracy, with a test accuracy of 92.50 Health, Income, and Physical Activity were the most influential predictors obtained from the model explanations. The results of this study suggest that ML combined with XAI is a promising means of developing accurate and computationally transparent tools for use in healthcare systems.
Diabetes mellitus (DM), a prevalent metabolic disorder, has significant global health implications. The advent of machine learning (ML) has revolutionized the ability to predict and manage diabetes early, offering new avenues to mitigate its impact. This systematic review examined 53 articles on ML applications for diabetes prediction, focusing on datasets, algorithms, training methods, and evaluation metrics. Various datasets, such as the Singapore National Diabetic Retinopathy Screening Program, REPLACE-BG, National Health and Nutrition Examination Survey (NHANES), and Pima Indians Diabetes Database (PIDD), have been explored, highlighting their unique features and challenges, such as class imbalance. This review assesses the performance of various ML algorithms, such as Convolutional Neural Networks (CNN), Support Vector Machines (SVM), Logistic Regression, and XGBoost, for the prediction of diabetes outcomes from multiple datasets. In addition, it explores explainable AI (XAI) methods such as Grad-CAM, SHAP, and LIME, which improve the transparency and clinical interpretability of AI models in assessing diabetes risk and detecting diabetic retinopathy. Techniques such as cross-validation, data augmentation, and feature selection are discussed in terms of their influence on the versatility and robustness of the model. Some evaluation techniques involving k-fold cross-validation, external validation, and performance indicators such as accuracy, area under curve, sensitivity, and specificity are presented. The findings highlight the usefulness of ML in addressing the challenges of diabetes prediction, the value of sourcing different data types, the need to make models explainable, and the need to keep models clinically relevant. This study highlights significant implications for healthcare professionals, policymakers, technology developers, patients, and researchers, advocating interdisciplinary collaboration and ethical considerations when implementing ML-based diabetes prediction models. By consolidating existing knowledge, this SLR outlines future research directions aimed at improving diagnostic accuracy, patient care, and healthcare efficiency through advanced ML applications. This comprehensive review contributes to the ongoing efforts to utilize artificial intelligence technology for a better prediction of diabetes, ultimately aiming to reduce the global burden of this widespread disease.
Stakeholders' conversations requirements elicitation meetings hold valuable insights into system and client needs. However, manually extracting requirements is time-consuming, labor-intensive, and prone to errors and biases. While current state-of-the-art methods assist in summarizing stakeholder conversations and classifying requirements based on their nature, there is a noticeable lack of approaches capable of both identifying requirements within these conversations and generating corresponding system requirements. These approaches would assist requirement identification, reducing engineers' workload, time, and effort. They would also enhance accuracy and consistency in documentation, providing a reliable foundation for further analysis. To address this gap, this paper introduces RECOVER (Requirements EliCitation frOm conVERsations), a novel conversational requirements engineering approach that leverages natural language processing and large language models (LLMs) to support practitioners in automatically extracting system requirements from stakeholder interactions by analyzing individual conversation turns. The approach is evaluated using a mixed-method research design that combines statistical performance analysis with a user study involving requirements engineers, targeting two levels of granularity. First, at the conversation turn level, the evaluation measures RECOVER's accuracy in identifying requirements-relevant dialogue and the quality of generated requirements in terms of correctness, completeness, and actionability. Second, at the entire conversation level, the evaluation assesses the overall usefulness and effectiveness of RECOVER in synthesizing comprehensive system requirements from full stakeholder discussions. Empirical evaluation of RECOVER shows promising performance, with generated requirements demonstrating satisfactory correctness, completeness, and actionability. The results also highlight the potential of automating requirements elicitation from conversations as an aid that enhances efficiency while maintaining human oversight.
SENEM-AI is a 3D virtual environment-based tool designed to enhance teaching and presentation skills by leveraging immersive simulations. Built upon the SENEM platform, it integrates virtual students powered by LLaMA with distinct personalities that simulate realistic classroom interactions. Educators can refine their communication strategies by reacting to dynamically generated questions. Preliminary evaluations highlighted its usability and potential impact, with participants valuing the immersive experience and engagement. SENEMAI represents a novel approach to supporting educators through accessible technology, paving the way for further research into AI-driven teaching aids and training environments in virtual settings. Tool video: https://www.youtube.com/watch?v=uHg25Gooi58
Diabetes mellitus (DM) is a global health issue of significance that must be diagnosed as early as possible and managed well. This study presents a framework for diabetes prediction using Machine Learning (ML) models, complemented with eXplainable Artificial Intelligence (XAI) tools, to investigate both the predictive accuracy and interpretability of the predictions from ML models. Data Preprocessing is based on the Synthetic Minority Oversampling Technique (SMOTE) and feature scaling used on the Diabetes Binary Health Indicators dataset to deal with class imbalance and variability of clinical features. The ensemble model provided high accuracy, with a test accuracy of 92.50% and an ROC-AUC of 0.975. BMI, Age, General Health, Income, and Physical Activity were the most influential predictors obtained from the model explanations. The results of this study suggest that ML combined with XAI is a promising means of developing accurate and computationally transparent tools for use in healthcare systems.
Commit messages are essential to understand changes in software projects, providing a way for developers to communicate code evolution. Generating effective commit messages that explain the rationale behind changes is a challenging and timeconsuming task. While previous research has shown success in automating straightforward commit messages (e.g., “add README”), our study explores a more complex task: generating rationale explanations for code changes. We developed a method to identify rationale sentences in commit messages and compiled a dataset of 45,945 commits with their corresponding rationales. A pre-trained model was trained on this dataset to generate rationale explanations. While the approach we engineered for the extraction of rationale from commit messages exhibited a 75 % precision, the model trained to generate the rationale only worked in a minority of cases. Our findings highlight the difficulty of the tackled task and the need for additional research in the area. We release our dataset and code to foster the investigation of this problem.
COmmon Software Measurement International Consortium (COSMIC) Functional Size Measurement is a method widely used in the software industry to quantify user functionality and measure software size, which is crucial for estimating development effort, cost, and resource allocation. COSMIC measurement is a manual task that requires qualified professionals and effort. To support professionals in COSMIC measurement, we propose an automatic approach, CosMet, that leverages Large Language Models to measure software size starting from use cases specified in natural language. To evaluate the proposed approach, we developed a web tool that implements CosMet using GPT-4 and conducted two studies to assess the approach quantitatively and qualitatively. Initially, we experimented with CosMet on seven software systems, encompassing 123 use cases, and compared the generated results with the ground truth created by two certified professionals. Then, seven professional measurers evaluated the analysis achieved by CosMet and the extent to which the approach reduces the measurement time. The first study’s results revealed that CosMet is highly effective in analyzing and measuring use cases. The second study highlighted that CosMet offers a transparent and interpretable analysis, allowing practitioners to understand how the measurement is derived and make necessary adjustments. Additionally, it reduces the manual measurement time by 60-80%.
The utmost importance of privacy and security requirements in software development calls for adopting methods that enable the identification and proactive mitigation of these issues during the system development. Our survey of 45 primary studies provides an overview of the methods, document types, and datasets employed in tackling this challenge, along with an analysis of approaches demonstrating superior performance based on document types and specific identification problems. Analysis reveals a wide adoption of ML-based systems on diverse datasets, showcasing the effectiveness of leveraging various sources of information to identify privacy and security requirements in software development.
Context: Early security requirements identification is crucial in software development, facilitating the integration of security measures into IT networks and reducing time and costs throughout software life-cycle. Objectives: This paper addresses the limitations of existing methods that leverage Natural Language Processing (NLP) and machine learning techniques for detecting security requirements. These methods often fall short in capturing syntactic and semantic relationships, face challenges in adapting across domains, and rely heavily on extensive domain-specific data. In this paper we focus on identifying the most effective approaches for this task, highlighting both domain-specific and domain-independent strategies. Method: Our methodology encompasses two primary streams of investigation. First, we explore shallow machine learning techniques, leveraging word embeddings. We test ensemble methods and grid search within and across domains, evaluating on three industrial datasets. Next, we develop several domain-independent models based on BERT, tailored to better detect security requirements by incorporating data on software weaknesses and vulnerabilities. Results: Our findings reveal that ensemble and grid search methods prove effective in domain-specific and domain-independent experiments, respectively. However, our custom BERT models showcase domain independence and adaptability. Notably, the CweCveCodeBERT model excels in Precision and F1-score, outperforming existing approaches significantly. It improves F1-score by similar to 3% and Precision by similar to 14% over the best approach currently in the literature. Conclusion: BERT-based models, especially with specialized pre-training, show promise for automating security requirement detection. This establishes a foundation for software engineering researchers and practitioners to utilize advanced NLP to improve security in early development phases, fostering the adoption of these state-of-the-art methods in real-world scenarios.
Context: Software development is a complex socio-technical process requiring a deep understanding of various aspects. In order to support practitioners in understanding such a complex activity, repository process metrics, like number of pull requests and issues, emerged as crucial for evaluating CI/CD workflows and guiding informed decision-making. The research community proposed different ways to visualize these metrics to increase their impact on developers' process comprehension: VR is a promising one. Nevertheless, despite such promising results, the role of VR, especially in educational settings, has received limited research attention. Objective: This study aims to address this gap by exploring how VR-based repository metrics visualization can support the teaching of process comprehension. Method: The registered report proposes the execution of a controlled experiment where VR and non-VR approaches will be compared, with the final aim to assess whether repository metrics in VR's impact on learning experience and software process comprehension. By immersing students in an intuitive environment, this research hypothesizes that VR can foster essential analytical skills, thus preparing software engineering students more effectively for industry requirements and equipping them to navigate complex software development tasks with enhanced comprehension and critical thinking abilities.
Machine learning (ML) is increasingly being used as a key component of most software systems, yet serious concerns have been raised about the fairness of ML predictions. Researchers have been proposing novel methods to support the development of fair machine learning solutions. Nonetheless, most of them can only be used in late development stages, e.g., during model training, while there is a lack of methods that may provide practitioners with early fairness analytics enabling the treatment of fairness throughout the development lifecycle. This paper proposes ReFair, a novel context-aware requirements engineering framework that allows to classify sensitive features from User Stories. By exploiting natural language processing and word embedding techniques, our framework first identifies both the use case domain and the machine learning task to be performed in the system being developed; afterward, it recommends which are the context-specific sensitive features to be considered during the implementation. We assess the capabilities of ReFair by experimenting it against a synthetic dataset---which we built as part of our research---composed of 12,401 User Stories related to 34 application domains. Our findings showcase the high accuracy of ReFair, other than highlighting its current limitations.
The metaverse represents a persistent, online 3D universe where people can interact, socialize, and work toward common goals. Education represents a key application domain, as it has the potential to enhance experiential learning and collaboration between learners and between learners and educators. However, challenges to the widespread adoption of educational metaverses persist. This paper focuses on emotional isolation, i.e., the feeling of emotional disconnection or loneliness, which can hinder learners’ motivation and participation. Machine learning-enabled emotional recognition systems have the potential to address this challenge, offering educators with feedback on the emotional states of learners within the metaverse. Yet, the integration of emotion recognition systems raises ethical concerns regarding consent, privacy, and algorithmic bias. In this short paper, we conduct a first step toward extracting ethical considerations from the literature on the use of emotion recognition in the educational metaverse. Then, we report these guidelines and finally implement one of the most critical —i.e., protection of privacy— within SENEM, an educational metaverse platform available in the literature. Through this research, we aim to raise awareness within the research community and promote responsible deployment of emotion recognition technology in educational metaverses, aiming to create a supportive and inclusive learning environment for all students.
Marcela Genero合作论文数Departamento de Tecnologias y Sistemas de Informacion
Escuela Superior de Informatica.
Universidad de Castilla-La Mancha7