Diabetic retinopathy (DR) grading from fundus images remains challenging because severity labels are ordinal, class distributions are imbalanced, and neighboring stages are often visually ambiguous. This study presents a unified pipeline for five-class DR grading that integrates controlled CNN/Transformer backbone benchmarking, ensemble evaluation, computational-cost analysis, external validation, and proof-of-concept attribution reporting. APTOS 2019 dataset was used to benchmark six backbone architectures under stratified five-fold cross-validation with fixed fold assignments, identical preprocessing, imbalance-aware training, and model selection based on quadratic weighted kappa (QWK). Hard voting, weighted soft voting, stacking, and hybrid class-level fusion were investigated to combine complementary models, while trainable parameters, model size, inference latency, and throughput were reported to quantify deployment cost. Model generalization was assessed by directly evaluating APTOS-trained models on Messidor-2 without retraining, fine-tuning, or recalibration. Safety-filtered summaries generated with a vision-language model were restricted to non-diagnostic model-attribution descriptions, and fundus-masked Grad-CAM++ maps were paired with these summaries for attribution reporting. Among individual backbones, ResNet-50 and ConvNeXt-Tiny achieved the strongest performance, while weighted soft voting provided the most stable ensemble performance. External validation confirmed the difficulty of cross-dataset five-class DR grading, with lower performance than that observed during internal cross-validation. The generated summaries adhered to predefined safety constraints but were not intended for lesion localization, diagnosis, or clinical validation. Overall, the study demonstrates that controlled benchmarking and weighted ensembling can provide stable performance for ordinal DR grading, while computational overhead, domain shift, uncertainty calibration, and expert clinical validation remain important considerations for real-world deployment.
The quality of diabetic retinopathy (DR) screening relies on the ability to correctly grade severity; however, many deep-learning (DL) classifiers cannot be easily interpreted in the clinical context. This study presents a methodology that combines strong discriminative models with multimodal explanations, converting retinal pixels into clinically interpretable outputs. Using the APTOS 2019 benchmark, we evaluated six representative CNN- and transformer-based backbones under a controlled protocol with stratified five-fold cross-validation. We then compared ensembling strategies (hard voting, weighted soft voting, stacking) and investigated a hybrid class-level fusion variant to exploit grade-specific advantages. For interpretability, we produced Grad-CAM++ visual attribution maps and short textual rationales using vision-language models (VLMs) conditioned on the fundus image and classifier outputs under conservative prompting constraints. Modern CNN backbones (ResNet-50 and ConvNeXt-Tiny) provided the strongest single-model baselines, with cross-validated QWK up to 0.919 and 0.914, respectively. Ensembling improved ordinal agreement, and weighted soft voting was the most consistent across folds (QWK 0.934 +/- 0.017). Hybrid class-level fusion was competitive but did not yield a statistically reliable improvement over standard fusion in paired fold comparisons (Holm-adjusted p >= 1.000). For explanation quality, Grad-CAM++ offered plausible but coarse localization, and VLM rationales were generally grade-consistent. Quantitatively, VLM variants showed a trade-off between clinical completeness and template-level semantic similarity (coverage 0.700 vs. BERTScore 0.072), while image-text alignment was comparable (CLIPScore approximately 0.34).
Diabetic retinopathy (DR) screening requires artificial intelligence (AI) models that are not only highly accurate in grading five clinical stages but are also capable of generating quantitatively evaluated lesion-aware explanations to earn the trust of clinicians. We propose RobustDRNet, a hybrid ensemble model that combines local convolutional features from Residual Network-34 (ResNet-34) and ConvNeXt-Tiny with global transformer embeddings from Vision Transformer Base/16 (ViT-B16) via two-stage feature fusion and a disentangled multilayer perceptron (MLP), followed by a logistic regression stacking meta-learner for prediction aggregation. To address severe class imbalance, our training pipeline employs stratified sampling, contrast-limited adaptive histogram equalization (CLAHE) for contrast enhancement, strong data augmentation, and class-weighted focal loss. Evaluated on the Asia Pacific Tele-Ophthalmology Society (APTOS) 2019 dataset, RobustDRNet achieved 88.4% validation accuracy, a 0.967 macro-averaged area under the receiver operating characteristic curve (macro-AUC), and Cohen's kappa of 0.823, outperforming individual backbones and simple voting ensembles. In addition to classification performance, we integrated six complementary explainable AI (XAI) techniques: Gradient-weighted Class Activation Mapping++ (Grad-CAM++), Integrated Gradients, attention rollout, SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), and Testing with Concept Activation Vectors (TCAV). Each technique was quantitatively benchmarked against expert-annotated lesion maps from the Indian Diabetic Retinopathy Image Dataset (IDRiD). Saliency maps achieved mean Intersection over Union (IoU) scores of 0.06 for Grad-CAM++ and approximately 0.10 for Integrated Gradients; SHapley Additive exPlanations (SHAP) perturbations showed a deletion drop of 0.25 and an insertion gain of 0.22; and TCAV achieved complete classifier-level TCAV alignment (score = 1.0) with clinically coherent, grade-wise importance trajectories. By combining competitive grading performance with multi-perspective and quantitatively evaluated interpretability, RobustDRNet provides a promising DR screening framework whose decisions are supported by lesion-aware explanatory evidence.
Nowadays, virtual worlds are evolving into immersive metaverses, i.e., collective, shared virtual spaces that arise from the convergence of virtual reality (VR), augmented reality (AR), the internet, and additional digital technologies. These environments allow users to interact through avatars. While research has explored the technologies and social dynamics of the metaverse, two key limitations remain. First, there is no systematic overview of the expertise required to build metaverses and the key publication venues. Second, the sociotechnical issues have not been fully synthesized or made instrumental for developing holistic design approaches that address both social and technical constraints. A deeper investigation is needed to guide future research, highlight challenges, and foster collaboration across disciplines. In this article, we propose a systematic mapping study that addresses the limitations identified. From an initial pool of 2,323 sources, we identify 63 primary resources to (1) characterize the research community around the metaverse and (2) elicit, synthesize, and categorize 19 social and 18 technical issues affecting the development of an effective metaverse. Based on our results, we contextualize the catalog of issues with respect to the current body of knowledge, providing insights and a research roadmap that transforms issues into actionable challenges in the scope of a novel, unified research asset coined "metaverse engineering", i.e., the multidisciplinary discipline to identify processes and instruments to design socio-technical metaverses.
The metaverse has transitioned from a science fiction term to a rapidly growing area of research and application, with potential uses in education, professional training, social events, and the virtual economy. However, despite this progress, a fully realized and functional metaverse is not yet available, and its development still requires a clear understanding and definition of the research directions to follow. Nonetheless, the metaverse topic does not start from scratch; it shares its foundations with Virtual Environments (VEs), which represent the Virtual Reality applications' core. In this paper, we built on top of the knowledge given by the history of the two topics to present a scoping review of the historical development of research in VEs and the metaverse from the 1990s to early 2024, analyzing 352 papers from the Scopus database. We aimed to offer a comprehensive understanding of how past research informs the present and future directions of the metaverse. Our findings revealed that the metaverse, while emerging as a distinct research area in recent years, is deeply rooted in the history of VEs, with many of its concepts and technologies deriving from earlier work in the field. We also identified new, underexplored trends within the metaverse research, proposing a future research agenda informed by the shared history of the two topics.
This work implements a blockchain-based framework for educational metaverses within SENEM, an immersive learning environment. Ethereum and IPFS were integrated to support decentralized uploading, storage, and retrieval of lecture materials. The resulting prototype formalizes requirements for secure content management and possibly improve integrity, traceability, and privacy. Overall, this work offers a first step toward transparent immersive learning systems and a proof-of-concept for future usability research.
Instructor effectiveness is fundamental to student learning, with the ability to manage student inquiries serving as a critical component of effective teaching. Student questions represent a valuable training resource for instructors to strengthen their teaching strategies, yet interactions with students are often constrained by several factors. In this paper, we investigate how instructors perceive machine- and studentgenerated questions, considering the potential for the former to complement the latter in a cost-effective manner. Our study involved 121 undergraduate students and an equivalent number of simulated students modeled using a state-of-the-art large language model, generating over 360 questions in total based on video lectures given by seven university instructors. We assessed whether instructors could distinguish between human- and machine-generated questions and how they evaluated their relevance, clarity, answerability, challenge level, and cognitive depth. Results show that instructors struggle to differentiate between the two sets of questions, with accuracy close to random chance. Instructors tended to (i) rate machine-generated questions slightly higher in relevance, clarity, answerability, and challenge-though only relevance and answerability showed significant differences-and (ii) associate them marginally more often with higher-order cognitive skills. This confirms the potential of machine-generated questions as tools for instructor training. Repository: https://github.com/tail- unica/realistic-ai-generatedquestions.
A few years after their release, Large Language Models (LLMs)-based tools are becoming an essential component of software education, as calculators are used in math courses. When learning software engineering (SE), the challenge is the extent to which LLMs are suitable and easy to use for different software development tasks. In this paper, we report the findings and lessons learned from using LLM-based tools-ChatGPT in particular-in five SE courses from four universities. After instructing students on the LLM potentials in SE and about prompting strategies, we ask participants to complete a survey and be involved in semi-structured interviews. The collected results report (i) indications about the usefulness of the LLM for different tasks, (ii) challenges to prompt the LLM, i.e., interact with it, (iii) challenges to adapt the generated artifacts to their own needs, and (iv) wishes about some valuable features students would like to see in LLM-based tools. Although results vary among different courses, also because of students' seniority and course goals, the perceived usefulness is greater for lowlevel phases (e.g., coding or debugging/fault localization) than for analysis and design phases. Interaction and code adaptation challenges vary among tasks and are mostly related to the need for task-specific prompts, as well as better specification of the development context.
Diabetes mellitus (DM) is a global health issue of significance that must be diagnosed as early as possible and managed well. This study presents a framework for diabetes prediction using Machine Learning (ML) models, complemented with eXplainable Artificial Intelligence (XAI) tools, to investigate both the predictive accuracy and interpretability of the predictions from ML models. Data Preprocessing is based on the Synthetic Minority Oversampling Technique (SMOTE) and feature scaling used on the Diabetes Binary Health Indicators dataset to deal with class imbalance and variability of clinical features. The ensemble model provided high accuracy, with a test accuracy of 92.50 Health, Income, and Physical Activity were the most influential predictors obtained from the model explanations. The results of this study suggest that ML combined with XAI is a promising means of developing accurate and computationally transparent tools for use in healthcare systems.
Diabetes mellitus (DM), a prevalent metabolic disorder, has significant global health implications. The advent of machine learning (ML) has revolutionized the ability to predict and manage diabetes early, offering new avenues to mitigate its impact. This systematic review examined 53 articles on ML applications for diabetes prediction, focusing on datasets, algorithms, training methods, and evaluation metrics. Various datasets, such as the Singapore National Diabetic Retinopathy Screening Program, REPLACE-BG, National Health and Nutrition Examination Survey (NHANES), and Pima Indians Diabetes Database (PIDD), have been explored, highlighting their unique features and challenges, such as class imbalance. This review assesses the performance of various ML algorithms, such as Convolutional Neural Networks (CNN), Support Vector Machines (SVM), Logistic Regression, and XGBoost, for the prediction of diabetes outcomes from multiple datasets. In addition, it explores explainable AI (XAI) methods such as Grad-CAM, SHAP, and LIME, which improve the transparency and clinical interpretability of AI models in assessing diabetes risk and detecting diabetic retinopathy. Techniques such as cross-validation, data augmentation, and feature selection are discussed in terms of their influence on the versatility and robustness of the model. Some evaluation techniques involving k-fold cross-validation, external validation, and performance indicators such as accuracy, area under curve, sensitivity, and specificity are presented. The findings highlight the usefulness of ML in addressing the challenges of diabetes prediction, the value of sourcing different data types, the need to make models explainable, and the need to keep models clinically relevant. This study highlights significant implications for healthcare professionals, policymakers, technology developers, patients, and researchers, advocating interdisciplinary collaboration and ethical considerations when implementing ML-based diabetes prediction models. By consolidating existing knowledge, this SLR outlines future research directions aimed at improving diagnostic accuracy, patient care, and healthcare efficiency through advanced ML applications. This comprehensive review contributes to the ongoing efforts to utilize artificial intelligence technology for a better prediction of diabetes, ultimately aiming to reduce the global burden of this widespread disease.
Stakeholders' conversations requirements elicitation meetings hold valuable insights into system and client needs. However, manually extracting requirements is time-consuming, labor-intensive, and prone to errors and biases. While current state-of-the-art methods assist in summarizing stakeholder conversations and classifying requirements based on their nature, there is a noticeable lack of approaches capable of both identifying requirements within these conversations and generating corresponding system requirements. These approaches would assist requirement identification, reducing engineers' workload, time, and effort. They would also enhance accuracy and consistency in documentation, providing a reliable foundation for further analysis. To address this gap, this paper introduces RECOVER (Requirements EliCitation frOm conVERsations), a novel conversational requirements engineering approach that leverages natural language processing and large language models (LLMs) to support practitioners in automatically extracting system requirements from stakeholder interactions by analyzing individual conversation turns. The approach is evaluated using a mixed-method research design that combines statistical performance analysis with a user study involving requirements engineers, targeting two levels of granularity. First, at the conversation turn level, the evaluation measures RECOVER's accuracy in identifying requirements-relevant dialogue and the quality of generated requirements in terms of correctness, completeness, and actionability. Second, at the entire conversation level, the evaluation assesses the overall usefulness and effectiveness of RECOVER in synthesizing comprehensive system requirements from full stakeholder discussions. Empirical evaluation of RECOVER shows promising performance, with generated requirements demonstrating satisfactory correctness, completeness, and actionability. The results also highlight the potential of automating requirements elicitation from conversations as an aid that enhances efficiency while maintaining human oversight.
SENEM-AI is a 3D virtual environment-based tool designed to enhance teaching and presentation skills by leveraging immersive simulations. Built upon the SENEM platform, it integrates virtual students powered by LLaMA with distinct personalities that simulate realistic classroom interactions. Educators can refine their communication strategies by reacting to dynamically generated questions. Preliminary evaluations highlighted its usability and potential impact, with participants valuing the immersive experience and engagement. SENEMAI represents a novel approach to supporting educators through accessible technology, paving the way for further research into AI-driven teaching aids and training environments in virtual settings. Tool video: https://www.youtube.com/watch?v=uHg25Gooi58
Diabetes mellitus (DM) is a global health issue of significance that must be diagnosed as early as possible and managed well. This study presents a framework for diabetes prediction using Machine Learning (ML) models, complemented with eXplainable Artificial Intelligence (XAI) tools, to investigate both the predictive accuracy and interpretability of the predictions from ML models. Data Preprocessing is based on the Synthetic Minority Oversampling Technique (SMOTE) and feature scaling used on the Diabetes Binary Health Indicators dataset to deal with class imbalance and variability of clinical features. The ensemble model provided high accuracy, with a test accuracy of 92.50% and an ROC-AUC of 0.975. BMI, Age, General Health, Income, and Physical Activity were the most influential predictors obtained from the model explanations. The results of this study suggest that ML combined with XAI is a promising means of developing accurate and computationally transparent tools for use in healthcare systems.
COmmon Software Measurement International Consortium (COSMIC) Functional Size Measurement is a method widely used in the software industry to quantify user functionality and measure software size, which is crucial for estimating development effort, cost, and resource allocation. COSMIC measurement is a manual task that requires qualified professionals and effort. To support professionals in COSMIC measurement, we propose an automatic approach, CosMet, that leverages Large Language Models to measure software size starting from use cases specified in natural language. To evaluate the proposed approach, we developed a web tool that implements CosMet using GPT-4 and conducted two studies to assess the approach quantitatively and qualitatively. Initially, we experimented with CosMet on seven software systems, encompassing 123 use cases, and compared the generated results with the ground truth created by two certified professionals. Then, seven professional measurers evaluated the analysis achieved by CosMet and the extent to which the approach reduces the measurement time. The first study’s results revealed that CosMet is highly effective in analyzing and measuring use cases. The second study highlighted that CosMet offers a transparent and interpretable analysis, allowing practitioners to understand how the measurement is derived and make necessary adjustments. Additionally, it reduces the manual measurement time by 60-80%.
Machine learning (ML) is increasingly being used as a key component of most software systems, yet serious concerns have been raised about the fairness of ML predictions. Researchers have been proposing novel methods to support the development of fair machine learning solutions. Nonetheless, most of them can only be used in late development stages, e.g., during model training, while there is a lack of methods that may provide practitioners with early fairness analytics enabling the treatment of fairness throughout the development lifecycle. This paper proposes ReFair, a novel context-aware requirements engineering framework that allows to classify sensitive features from User Stories. By exploiting natural language processing and word embedding techniques, our framework first identifies both the use case domain and the machine learning task to be performed in the system being developed; afterward, it recommends which are the context-specific sensitive features to be considered during the implementation. We assess the capabilities of ReFair by experimenting it against a synthetic dataset---which we built as part of our research---composed of 12,401 User Stories related to 34 application domains. Our findings showcase the high accuracy of ReFair, other than highlighting its current limitations.
Source code reuse is considered one of the holy grails of modern software development. Indeed, it has been widely demonstrated that this activity decreases software development and maintenance costs while increasing its overall trustwor-thiness. The Object-Oriented (OO) paradigm provides different internal mechanisms to favor code reuse, i.e., specification inheritance, implementation inheritance, and delegation. While previous studies investigated how inheritance relations impact source code quality, there is still a lack of understanding of their evolutionary aspects and, more particular, of how these mechanisms may impact source code quality over time. To bridge this gap of knowledge, this paper proposes an empirical investigation into the evolution of specification inheritance, implementation inheritance, and delegation and their impact on the variability of source code quality attributes. First, we assess how the implementation of those mechanisms varies over 15 releases of three software systems. Second, we devise a statistical approach with the aim of understanding how inheritance and delegation let source code quality—as indicated by the severity of code smells—vary in either positive or negative manner. The key results of the study indicate that inheritance and delegation evolve over time, but not in a statistically significant manner. At the same time, their evolution often leads code smell severity to be reduced, hence possibly contributing to improve code maintainability.
Tiago Massoni合作论文数Informatics Center - UFPE1