Breath analysis is emerging as a non-invasive and promising diagnostic approach capable of assessing a patient’s metabolic state by detecting volatile organic compounds in exhaled breath. This study investigates the potential of breath analysis for the early detection of lung cancer, respiratory and gastrointestinal diseases using open-access data from three distinct datasets. An artificial intelligence methodology is implemented to predict diagnostic labels while addressing class imbalance, an inherent challenge in medical datasets. After evaluating model performance and stability, the most relevant volatile organic compounds identified by the best-performing model for each dataset are analyzed. Using eXplainable Artificial Intelligence, the influence of volatile organic compound abundances on predictions is examined, enabling the identification of key variables and improving model interpretability. The proposed methodology provides a robust framework for breath-based diagnostics, emphasizing the potential of integrating breath analysis with machine learning to advance clinical decision-making despite ongoing challenges related to sampling variability, detection sensitivity, and standardization across studies.
The clinical adenoma - carcinoma progression represents a well-established framework for understanding colorectal cancer (CRC) development, although the molecular mechanisms underlying this transition remain only partially understood. Increasing evidence suggests the gut microbiome (GM) as a key modulator of colorectal carcinogenesis, positioning microbial profiling as a promising avenue for noninvasive risk stratification and early detection. In this study, Machine Learning (ML) classifiers integrated with eXplainable Artificial Intelligence (XAI) techniques were employed to identify microbiome-derived biomarkers predictive of CRC and adenomatous lesions. The models were trained on 16S rRNA sequencing data from 453 patients and evaluated through cross-validation, achieving AU-ROC and AU-PRC scores of 0.71 and 0.67, respectively. External validation on an independent Italian cohort (n=43) yielded AU-ROC and AU-PRC scores of 0.70 and 0.89, respectively. XAI-based interpretation revealed consistent microbial signatures across datasets. In detail, taxa belonging to the Fusobacterium and Peptostreptococcus genera were associated with increased CRC risk, whereas the Eubacterium eligens group was identified as a robust negative predictor. Beyond classification, patient-level explanations enabled by XAI facilitated the identification of adenoma subgroups exhibiting microbiome profiles converging toward those of CRC, suggesting the presence of transitional microbial states. Moreover, SHAP-based interaction networks uncovered microbial hubs and inter-species dependencies characterizing high-risk configurations, providing insights into the ecological dynamics of colorectal tumorigenesis. These findings demonstrate the added XAI value in elucidating microbiome interactions, enhancing model interpretability, and supporting biologically informed hypotheses. This integrative, explainable framework highlights the potential of AI-driven microbiome analysis in precision oncology and advances the development of interpretable, noninvasive tools for CRC risk prediction and management.
This study investigated the predictive performance of three regression models—Gradient Boosting (GB), Random Forest (RF), and XGBoost—in forecasting mortality due to endocrine, nutritional, and metabolic diseases across Italian provinces. Utilizing a dataset encompassing air pollution metrics and socio-economic indices, the models were trained and tested to evaluate their accuracy and robustness. Performance was assessed using metrics such as coefficient of determination (r2), mean absolute error (MAE), and root mean squared error (RMSE), revealing that GB outperformed both RF and XGB, offering superior predictive accuracy and model stability (r2 = 0.55, MAE = 0.17, and RMSE = 0.05). To further interpret the results, SHAP (SHapley Additive exPlanations) analysis was applied to the best-performing model to identify the most influential features driving mortality predictions. The analysis highlighted the critical roles of specific pollutants, including benzene and socio-economic factors such as life quality and instruction, in influencing mortality rates. These findings underscore the interplay between environmental and socio-economic determinants in health outcomes and provide actionable insights for policymakers aiming to reduce health disparities and mitigate risk factors. By combining advanced machine learning techniques with explainability tools, this research demonstrates the potential for data-driven approaches to inform public health strategies and promote targeted interventions in the context of complex environmental and social determinants of health.
Predicting phenotypes from genomic data can significantly advance agriculture. Genomic selection, which uses genome-wide DNA markers to identify individuals with high genetic value, enhances the accuracy of breeding programs. While linear models are routinely used for genomic selection (GS), machine learning (ML) models offer complementary potential. In this study, robust ML-based models were developed to predict five phenotypic traits—three related to flowering time and two to leaf number—in Arabidopsis thaliana, a model plant with a fully sequenced genome. Using explainable artificial intelligence (XAI), specifically SHapley Additive exPlanations (SHAP) values, we identified SNPs that contributed most to trait prediction. Many of these SNPs were located in or near genes known to regulate flowering and stem elongation, such as DOG1 and VIN3, supporting the biological plausibility of the model. SHAP also enabled local interpretability at the single-plant level, revealing the genotypic basis of individual predictions. Our results indicate that integrating ML with XAI improves model interpretability and provides predictive performance comparable to traditional methods. This approach confirms known genotype–phenotype relationships and highlights new candidate loci, paving the way for functional validation. The proposed methodology offers promising applications in precision breeding and translation of insights from Arabidopsis to crop species.
This systematic review explores the use of digital twins (DT) for sustainable agricultural water management. DTs simulate real-time agricultural environments, enabling precise resource allocation, predictive maintenance, and scenario planning. AI enhances DT performance through machine learning (ML) and data-driven insights, optimizing water usage. In this study, from an initial pool of 48 papers retrieved from well-known databases such as Scopus and Web of Science, etc., a rigorous eligibility criterion was applied, narrowing the focus to 11 pertinent studies. This review highlights major disciplines where DT technology is being applied: hydroponics, aquaponics, vertical farming, and irrigation. Additionally, the literature identifies two key sub-applications within these disciplines: the simulation and prediction of water quality and soil water. This review also explores the types and maturity levels of DT technology and key concepts within these applications. Based on their current implementation, DTs in agriculture can be categorized into two functional types: monitoring DTs, which emphasize real-time response and environmental control, and predictive DTs, which enable proactive irrigation management through environmental forecasting. AI techniques used within the DT framework were also identified based on their applications. These findings underscore the transformative role that DT technology can play in enhancing efficiency and sustainability in agricultural water management. Despite technological advancements, challenges remain, including data integration, scalability, and cost barriers. Further studies should be conducted to explore these issues within practical farming environments.
Background: The Yb-176(n, gamma)Yb-177 -> 177Lu reaction is of interest in nuclear medicine as it is the preferred production route for 177Lu. This radioisotope has seen a very fast growth of usage in nuclear medicine in recent years due to its outstanding properties. New data on this reaction could provide useful information for production at new facilities. Purpose: We aim to resolve resonances in the Yb-176(n, gamma)Yb-177 reaction for the first time. Previous capture measurement provided data at thermal point and encompassed integral measurements in the range from 3 keV to 1 MeV, where three time-of-flight measurements are available, but with low resolution to resolve the resonances. Transmission measurements from the 1970s resolved and analyzed some resonances. Method: We measure the neutron capture cross section of Yb-176(n, gamma)Yb-177 by means of the time-of-flight technique at the Experimental Area 1 of the n_TOF facility at CERN using an enriched (Yb2O3)-Yb-176 sample and an array of four C6D6 liquid scintillation detectors. Results: We have resolved 164 resonances up to 21 keV, including 96 new ones. We also provide new capture experimental data from 90 eV to 3 keV, and we extend the resolved resonance region up to 21 keV. In addition, resonance decay widths, Gamma(gamma) and Gamma(n), are provided for all resonances together with resonance energies. Conclusions: The Yb-176(n, gamma) Yb-177 reaction has been measured, providing resonance parameters for the first time from a few eV to 21 keV. The analysis of the resonances has been carried out and compared with previous works and existing libraries, revealing discrepancies due to the new information on Gamma(gamma) parameters. Our results are consistent with the Gamma(n) parameters obtained in transmission measurements.
BackgroundAdvances in DNA sequencing revolutionized plant genomics and significantly contributed to the study of genetic diversity. However, predicting phenotypes from genomic data remains a challenge, particularly in the context of plant breeding. Despite significant progress, accurately predicting phenotypes from high-dimensional genomic data remains a challenge, particularly in identifying the key genetic factors influencing these predictions. This study aims to bridge this gap by integrating explainable artificial intelligence (XAI) techniques with advanced machine learning models. This approach is intended to enhance both the predictive accuracy and interpretability of genotype-to-phenotype models, thereby improving their reliability and supporting more informed breeding decisions.ResultsThis study compares several ML methods for genotype-to-phenotype prediction, using data available from an almond germplasm collection. After preprocessing and feature selection, regression models are employed to predict almond shelling fraction. Best predictions were obtained by the Random Forest method (correlation = 0.727 ± 0.020, an R2 = 0.511 ± 0.025, and an RMSE = 7.746 ± 0.199). Notably, the application of the SHAP (SHapley Additive exPlanations) values algorithm to explain the results highlighted several genomic regions associated with the trait, including one, having the highest feature importance, located in a gene potentially involved in seed development.ConclusionsEmploying explainable artificial intelligence algorithms enhances model interpretability, identifying genetic polymorphisms associated with the shelling percentage. These findings underscore XAI’s efficacy in predicting phenotypic traits from genomic data, highlighting its significance in optimizing crop production for sustainable agriculture.
Identifying the origin of a food product holds paramount importance in ensuring food safety, quality, and authenticity. Knowing where a food item comes from provides crucial information about its production methods, handling practices, and potential exposure to contaminants. Machine learning techniques play a pivotal role in this process by enabling the analysis of complex data sets to uncover patterns and associations that can reveal the geographical source of a food item. This study aims to investigate the potential use of explainable artificial intelligence for identifying the food origin. The case of study of Mozzarella di Bufala Campana PDO has been considered by examining the composition of the microbiota in each samples. Three different supervised machine learning algorithms have been compared and the best classifier model is represented by Random Forest with an Area Under the Curve (AUC) value of 0.93 and the top accuracy of 0.87. Machine learning models effectively classify origin, offering innovative ways to authenticate regional products and support local economies. Further research can explore microbiota analysis and extend applicability to diverse food products and contexts for enhanced accuracy and broader impact.
METROFOOD-IT enhances the Italian Node of the ESFRI METROFOOD-RI infrastructure and promotes research and innovation in the agri-food sector through integrated services, with an emphasis on digitization, efficiency, traceability, and sustainability. The project fosters a research paradigm focused on improving metrological data flows to enhance, quality, traceability, security in the food and nutritional domain. This paper introduces the architectural model utilized for the integration of both physical and electronic facilities. A service-oriented architecture (SOA) has been designed to facilitate efficient and scalable data sharing. The METROFOOD-IT architectural model unfolds across three levels: services, data infrastructure, and linking infrastructure. The data architecture plays a central role in managing backend systems for services, ensuring continuous data availability to each ecosystem service. To guarantee a stable and recognized data model by the scientific and industrial community, the Smart Data Model Agrifood has been adopted. This model addresses standardized data sharing and standardization issues, ensuring uniformity in data syntax and semantics, implementing access restrictions, and providing essential technical aspects in data platforms.
This article explores the significant impact that artificial intelligence (AI) could have on food safety and nutrition, with a specific focus on the use of machine learning and neural networks for disease risk prediction, diet personalization, and food product development. Specific AI techniques and explainable AI (XAI) are highlighted for their potential in personalizing diet recommendations, predicting models for disease prevention, and enhancing data-driven approaches to food production. The article also underlines the importance of high-performance computing infrastructures and data management strategies, including data operations (DataOps) for efficient data pipelines and findable, accessible, interoperable, and reusable (FAIR) principles for open and standardized data sharing. Additionally, it explores the concept of open data sharing and the integration of machine learning algorithms in the food industry to enhance food safety and product development. It highlights the METROFOOD-IT project as a best practice example of implementing advancements in the agri-food sector, demonstrating successful interdisciplinary collaboration. The project fosters both data security and transparency within a decentralized data space model, ensuring reliable and efficient data sharing. However, challenges such as data privacy, model interoperability, and ethical considerations remain key obstacles. The article also discusses the need for ongoing interdisciplinary collaboration between data scientists, nutritionists, and food technologists to effectively address these challenges. Future research should focus on refining AI models to improve their reliability and exploring how to integrate these technologies into everyday nutritional practices for better health outcomes.
Respiratory malignancies, encompassing cancers affecting the lungs, the trachea, and the bronchi, pose a significant and dynamic public health challenge. Given that air pollution stands as a significant contributor to the onset of these ailments, discerning the most detrimental agents becomes imperative for crafting policies aimed at mitigating exposure. This study advocates for the utilization of explainable artificial intelligence (XAI) methodologies, leveraging remote sensing data, to ascertain the primary influencers on the prediction of standard mortality rates (SMRs) attributable to respiratory cancer across Italian provinces, utilizing both environmental and socioeconomic data. By scrutinizing thirteen distinct machine learning algorithms, we endeavor to pinpoint the most accurate model for categorizing Italian provinces as either above or below the national average SMR value for respiratory cancer. Furthermore, employing XAI techniques, we delineate the salient factors crucial in predicting the two classes of SMR. Through our machine learning scrutiny, we illuminate the environmental and socioeconomic factors pertinent to mortality in this disease category, thereby offering a roadmap for prioritizing interventions aimed at mitigating risk factors.
BackgroundColorectal cancer (CRC) is a type of tumor caused by the uncontrolled growth of cells in the mucosa lining the last part of the intestine. Emerging evidence underscores an association between CRC and gut microbiome dysbiosis. The high mortality rate of this cancer has made it necessary to develop new early diagnostic methods. Machine learning (ML) techniques can represent a solution to evaluate the interaction between intestinal microbiota and host physiology. Through explained artificial intelligence (XAI) it is possible to evaluate the individual contributions of microbial taxonomic markers for each subject. Our work also implements the Shapley Method Additive Explanations (SHAP) algorithm to identify for each subject which parameters are important in the context of CRC.ResultsThe proposed study aimed to implement an explainable artificial intelligence framework using both gut microbiota data and demographic information from subjects to classify a cohort of control subjects from those with CRC. Our analysis revealed an association between gut microbiota and this disease. We compared three machine learning algorithms, and the Random Forest (RF) algorithm emerged as the best classifier, with a precision of 0.729 ± 0.038 and an area under the Precision-Recall curve of 0.668 ± 0.016. Additionally, SHAP analysis highlighted the most crucial variables in the model's decision-making, facilitating the identification of specific bacteria linked to CRC. Our results confirmed the role of certain bacteria, such as Fusobacterium, Peptostreptococcus, and Parvimonas, whose abundance appears notably associated with the disease, as well as bacteria whose presence is linked to a non-diseased state.DiscussionThese findings emphasizes the potential of leveraging gut microbiota data within an explainable AI framework for CRC classification. The significant association observed aligns with existing knowledge. The precision exhibited by the RF algorithm reinforces its suitability for such classification tasks. The SHAP analysis not only enhanced interpretability but identified specific bacteria crucial in CRC determination. This approach opens avenues for targeted interventions based on microbial signatures. Further exploration is warranted to deepen our understanding of the intricate interplay between microbiota and health, providing insights for refined diagnostic and therapeutic strategies.
The neutron Time-of-Flight facility (n_TOF) is an innovative facility operative since 2001 at CERN, with three experimental areas. In this paper the n_TOF facility will be described, together with the upgrade of the facility during the Long Shutdown 2 at CERN. The main features of the detectors used for capture fission cross section measurements will be presented with perspectives for the future measurements.
Background Autism spectrum disorder (ASD) constitutes a pervasive developmental condition impacting social interaction and communication proficiency. Emerging evidence underscores a plausible association between ASD and alterations within the gut microbiome—an intricate assembly of microorganisms inhabiting the gastrointestinal tract. While machine learning (ML) techniques have emerged as a valuable tool for unraveling the intricate interactions between the gut microbiome and host physiology, their application faces limitations in assessing the individual contributions of microbial species for each subject. Addressing this constraint, explainable artificial intelligence (XAI) emerges as a solution. This paper delves into the potential of the Shapley Method Additive Explanations (SHAP) algorithm for personalized identification of microbiome biomarkers in the context of ASD. Results The study demonstrates the efficacy of the SHAP algorithm in overcoming conventional ML limitations. SHAP enables a personalized assessment of microbiome contributions, facilitating the identification of specific bacteria associated with ASD. Moreover, leveraging local explanation embeddings and an unsupervised clustering method successfully clusters ASD subjects into subgroups. Notably, a cluster with lower ASD probability is identified, uncovering false negatives in ASD classification. The recognition of false negatives holds clinical significance, prompting an exploration of contributing factors and insights for refining ASD classification accuracy. Conclusions In conclusion, XAI provides personalized insights into ASD-associated microbiome biomarkers. Its ability to address ML limitations enhances understanding of individualized microbial environment in ASD. The identification of ASD subgroups through clustering analysis emphasizes disorder heterogeneity. Additionally, recognizing false negatives within ASD classification introduces complexity to patient care considerations. These findings imply potential for tailored interventions based on individual microbiome profiles, advancing precision in ASD management and classification.
The origin of food products is an important factor to consider for consumers seeking authentic, high-quality, and safe products. Information about the origin provides valuable guidance for making informed decisions about food purchases and can help promote sustainable agricultural practices. Indeed, machine learning (ML) techniques offer a promising solution for evaluating the geographical origin of food products, including Mozzarella di Bufala Campana PDO. By analyzing various data sources, ML algorithms can discern patterns and associations that correlate with specific geographic regions. For instance, ML models can be trained on datasets containing information about the microbiota composition of Mozzarella di Bufala Campana PDO samples collected from different regions. By leveraging advanced classification or clustering algorithms, these models can learn to differentiate between microbiota profiles associated with distinct geographical origins. The proposed study aimed to implement an explainable artificial intelligence framework using microbiota data from Mozzarella PDO samples from Salerno and Caserta. Our analysis aimed to classify each sample into one of the two origin areas. We employed the XGB classifier, a machine learning algorithm, which achieved an accuracy of 0.825 ± 0.032, a F1-score of 0.849 ± 0.028 and an average AUC of 0.880 ± 0.026. The application of machine learning methods to classify the geographical origin of products, coupled with advanced techniques of chemical and biological analysis supported by artificial intelligence, promises to distinguish between authentic and adulterated products. It is important to ensure that the data used to train the models are representative and reliable. Machine learning could play a vital role in this context by enabling the analysis of vast amounts of data to identify patterns and characteristics unique to specific geographic regions.
Weed control is a critical challenge in agriculture, impacting crop yields and necessitating various management strategies, from manual to chemical methods. This paper explores the integration of advanced sensor technologies with machine learning for weed detection and management. Our study discusses the application of multiple sensors RGB, multispectral, hyperspectral, thermal, LiDAR, fluorescence, and ultrasonic each providing unique advantages in detecting and differentiating weeds from crops based on spectral, thermal, and structural characteristics. We delve into the capabilities of each sensor type, underlining their individual and combined utility in addressing the complexities of agricultural environments, such as varying lighting conditions, soil types, and crop stages. This review highlights the potential of ML algorithms to refine the data processing and enhance the accuracy of weed identification systems. Our findings indicate that while these technologies offer significant improvements in detecting and managing weeds, challenges remain in their adoption, particularly among small-scale farmers due to system complexity and cost.
A powder diffraction pattern is mainly affected by peak overlaps, difficulty in the correct background estimation, presence of preferred orientation effects, and limited experimental resolution.All this makes the structure solution process non-trivial.Most importantly, it is difficult to establish the critical initial steps such as pattern indexation and space group determination, especially if more than one chemical phase is present in the compound.On the other hand, an incorrectly defined unit cell does not lead to the structural solution.It happens despite the progress, availability, and variety (in terms of strategies and methods implemented) of automatic indexing software such as DICVOL [1], N-TREOR09 [2] and ITO [3].In the past few years, extraordinary advances in data-driven models and the availability of large amounts of experimental data from many different sources have enabled the development and application of Artificial Intelligence in materials science [4], especially machine learning (ML) algorithms for diffraction data analysis.A new machine learning (ML) based web platform, named CrystalMELA (Crystallographic MachinE LeArning) [5], for crystal system classification has been developed.The aim is to try to overcome the difficulties posed by the structure solution process from powder diffraction data, and to complement traditional indexing approaches.The tool is currently designed for the classification of the seven crystal classes.In the current version, the platform can run three different and complementary ML models: a Convolutional Neural Network (CNN), a Random Forest (RF) and an Extremely randomized trees (ExRT).The models have been trained on theoretical powder diffraction patterns of more than 280,000 crystal structures of inorganic, organic, organo-metallic compounds and minerals as collected in the POW_COD database [6].A 70% of classification accuracy was achieved, improved to 90% if the top-2 accuracy is considered.CrystalMELA is free available for the scientific community and its home web page is shown in Fig. 1.All the classification options in CrystalMELA platform are designed to be powerful and easy to use, supported by a user-friendly graphic interface.Their main aspects and some examples of applications to real cases will be presented.Figure 1.The Home web page of CrystalMELA platform
The approach based on atomic pair distribution function (PDF) has revolutionized structural investigations by X-ray/electron diffraction of nano or quasi-amorphous materials, opening up the possibility of exploring short-range order. However, the ab initio crystal structural solution by the PDF is far from being achieved due to the difficulty in determining the crystallographic properties of the unit cell. A method for estimating the crystal cell parameters directly from a PDF profile is presented, which is composed of two steps: first, the type of crystal cell is inferred using machine-learning approaches applied to the PDF profile; second, the crystal cell parameters are extracted by means of multivariate analysis combined with vector superposition techniques. The procedure has been validated on a large number of PDF profiles calculated from known crystal structures and on a small number of measured PDF profiles. The lattice determination step has been benchmarked by a comprehensive exploration of different classifiers and different input data. The highest performance is obtained using the k-nearest neighbours classifier applied to whole PDF profiles. Descriptors calculated from the PDF profiles by recurrence quantitative analysis produce results that can be interpreted in terms of PDF properties, and the significance of each descriptor in determining the prediction is evaluated. The cell parameter extraction step depends on the cell metric rather than its type. Monometric, dimetric and trimetric cells have top-1 estimates that are correct 40, 20 and 5% of the time, respectively. Promising results were obtained when analysing real nanocrystals, where unit cells close to the true ones are found within the top-1 ranked solution in the case of monometric cells and within the top-6 ranked solutions in the case of dimetric cells, even in the presence of a crystalline impurity with a weight fraction up to 40%.
Several international agencies recommend the study of new routes and new facilities for producing radioisotopes with application to nuclear medicine. 177Lu is a versatile radioisotope used for therapy and diagnosis (theranostics) of cancer with good success in neuroendocrine tumours that is being studied to be applied to a wider range of tumours. 177Lu is produced in few nuclear reactors mainly by the neutron capture on 176Lu. However, it could be produced at high-intensity celeratorbased neutron facilities. The energy of the neutrons in accelerator-based neutron facilities is higher than in thermal reactors.Thus, experimental data on the 176Yb(n,γ) cross-section in the eV and keV region are mandatory to calculate accurately the production of 177Yb, which beta decays to 177Lu. At present, there are not experimental data available from thermal to 3 keV of the 176Yb(n,γ) cross-section. In addition, there is no data in the resolved resonance region (RRR). This contribution shows the first results of the 176Yb capture measurement performed at the n_TOF facility at CERN.