Artificial intelligence systems for prostate cancer are increasingly moving beyond single-source image analysis toward models that combine radiology, clinical variables, pathology, molecular measurements, and radiomics. We conducted a PRISMA-ScR-guided evidence map of peer-reviewed English-language studies published from January 2021 to December 2025 and indexed in IEEE Xplore, Scopus, Web of Science, PubMed, or SpringerLink. Eligible studies used AI or machine learning to integrate at least two information sources for prostate cancer detection, grading, staging, prognosis, recurrence prediction, or treatment selection. Two reviewers screened records and extracted study design, clinical task, input modalities, fusion approach, validation strategy, performance, reproducibility, and implementation-relevant reporting. Twenty-six studies met eligibility criteria. Clinical variables were the most frequent non-imaging input (24/26), mpMRI was the dominant imaging source (18/26), and six studies incorporated whole-slide histopathology. Reported paired comparisons usually favoured multimodal approaches, with examples including PI-CAI performance above the median radiologist AUROC (0.91 vs. 0.86) and PET/MRI/clinical models reporting AUC values up to 0.955. However, the evidence remains mostly retrospective, geographically concentrated, weakly reproducible, and rarely calibrated; none of the included studies tested a multimodal AI system prospectively inside a clinical pathway. Current findings support continued development and rigorous evaluation rather than routine deployment. The next phase of research should prioritise prospective pathway studies, diverse external validation, transparent reporting, missing-modality robustness, and patient-centred outcomes.
Background: Glioblastoma (GBM) is a highly heterogeneous and vascularized malignancy in which the mesenchymal (MES) subtype is associated with poor prognosis, extensive macrophage infiltration and resistance to therapy. However, the signaling mechanisms integrating vascular remodeling with inflammatory tumor-macrophage crosstalk remain incompletely understood. Methods: We integrated magnetic resonance imaging-derived vascular phenotyping with transcriptomic analyses of human glioblastoma cohorts to identify molecular pathways associated with highly vascular tumors. Functional studies using glioblastoma cell lines, THP-1-derived macrophages and co-culture systems were performed to investigate the role of PI3K signaling in tumor-macrophage communication. Finally, an independent single-cell transcriptomic cohort of primary human glioblastoma was interrogated to determine whether the identified inflammatory programs were conserved in malignant cells from patient tumors. Results: Integrated imaging-transcriptomic analyses identified highly vascular glioblastomas as tumors enriched for the MES subtype, increased macrophage infiltration and activation of PI3K-associated signaling. Pharmacological inhibition of PI3K reduced the expression of macrophage-recruiting cytokines and impaired the ability of glioblastoma cells to educate macrophages toward an immunosuppressive phenotype. Reciprocally, tumor-educated macrophages enhanced inflammatory signaling, immune checkpoint expression and migratory capacity in glioblastoma cells, whereas IL-6 blockade attenuated these effects, identifying IL-6 as a key mediator of this bidirectional communication. To determine whether these inflammatory programs were conserved in human disease, we analyzed an independent single-cell transcriptomic dataset of primary glioblastomas. MES-like malignant cells exhibited the strongest inflammatory transcriptional programs among the four malignant transcriptional states, including higher NF-κB activation program scores and tumor-macrophage communication signature scores. At the tumor level, MES-like enrichment was positively associated with higher inflammatory program activity, supporting the clinical relevance of the proposed signaling axis. Conclusions: Together, our findings identify PI3K signaling as a central regulator integrating vascular remodeling with inflammatory tumor-macrophage communication in mesenchymal glioblastoma. These results provide a mechanistic framework linking PI3K signaling, macrophage education and the MES phenotype, and provide a rationale for therapeutic strategies aimed at disrupting inflammatory signaling within the glioblastoma microenvironment.
BACKGROUND:Glioblastoma (GBM) growth can alter surrounding brain tissue through location-dependent physiological changes. Two main growth phenotypes-(I) infiltrative, characterized by diffuse invasion with minimal mass effect, and (II) proliferative, characterized by pronounced tissue compression-are recognized, but their quantitative characterization and prognostic impact remain poorly explored. PURPOSE:To develop and validate a novel MRI-based biomarker, the Dynamic Infiltration Rate (DIR), that quantitatively assesses the balance between tumor volume expansion and peritumoral compression, and to evaluate its prognostic ability for stratifying patients based on overall survival (OS). METHODS:The DIR was defined as the ratio between tumor-volume enlargement and mass-effect-induced peritumoral compression. Technical validation was conducted using synthetic datasets with known ground truth spanning realistic infiltrative-proliferative spectra. Clinically, patients were dichotomized into high- and low-infiltration groups using a threshold optimized by maximizing the log-rank statistic for OS. Prognostic evaluation included multivariate Cox regression adjusted for age, sex, and MGMT methylation status. RESULTS:The synthetic dataset validation demonstrated high concordance with ground truth ( R 2 = 0.89 $R^2 = 0.89$ ). Clinical evaluation indicated significantly improved OS in the low DIR group (median = 35.2 weeks) compared to the high DIR group (median = 16.0 weeks; p = 0.0001 $p = 0.0001$ ). DIR effectively stratified patients based on survival (log-rank p < 0.001 $p < 0.001$ , HR = 2.49) and remained an independent prognostic factor on multivariate analysis (HR = 1.45, 95% CI 1.07-1.85; p = 0.0159 $p = 0.0159$ ). CONCLUSIONS:The DIR is a novel and robust quantitative MRI biomarker capable of distinguishing between proliferative and infiltrative GBM phenotypes, independently predicting OS. Early phenotype identification could facilitate personalized treatment strategies and individualized follow-up scheduling.
The European Health Data Space (EHDS) promotes health data sharing and secondary use across Europe. The QUANTUM project focuses on labelling data quality (DQ), utility, and maturity to support EU-wide standards within this context. This study examines current DQ assessment practices in European Data Holder institutions to inform the design of a labelling tool for the EHDS. The study explored institutional practices for DQ assurance and assessment through a survey of EU-wide Data Holders and a literature search on open-source health DQ tools with potential for labelling. The survey targeted QUANTUM partners and external institutions, addressing DQ practices and tools. The literature review followed PRISMA guidelines and combined PubMed, AI-based queries, and known sources. We obtained survey responses from 27 Institutions across 13 European countries. The results showed a high variety and heterogeneity in DQ practices and tools used, including in-house developed tools, open-source tools, commercial products, and manual procedures. Most practices allowed customization of DQ dimensions and export of DQ analyses. The systematic review identified 66 DQ tools, 53% specific to data types such as electronic health records, omics or imaging, and 47% general-purpose. The diverse DQ institutional practices and tools emphasize the need for an interoperable self-assessment DQ labelling tool that guides the measurement and completion of consolidated metrics aligned with the EHDS regulations. Based on these findings, the QUANTUM project is developing this tool to support DQ labelling of Data Holders' datasets candidate to publication at EU Health Data Access Bodies.
Recent advancements in blood-brain barrier permeability (BBBP) prediction of drug compounds have highlighted the growing role of machine learning, particularly deep learning. While considerable attention has been given to feature engineering and model design, their evaluation often receives insufficient attention despite its fundamental role in model credibility. In this work, we study a phenomenon we term overrepresentation bias, susceptible to be found in drug property databases, characterized by the presence of near-identical compounds with the same or nearly identical property values. Our findings reveal that overrepresentation bias leads to overly optimistic performance estimates in BBBP prediction models by significantly inflating test evaluation metrics─13.3% in average for the area under curve and 16.44% in average for the macro F1-score. To address this bias, we propose (i) an automatic detection algorithm and (ii) a bias-aware data handling procedure. We recommend adopting this approach to ensure more reliable model evaluations. Given that overrepresentation bias can affect performance estimation more than feature selection, model architecture, or even training data, we urge both academic and industrial communities to acknowledge its significance and take proactive measures to identify and address this bias in future studies.
While acute brain dysfunction (ABD, i.e., delirium and coma) is associated with significantly increased morbidity in critically ill patients, it presents with great heterogeneity that poses a challenge for management and prognostication. While machine learning may be promising for subgroup identification, this approach has not yet been applied to COVID-19 patients with ABD. The aim of our study was to identify distinct clusters among critically ill patients with COVID-19 based on ICU admission data and evaluate their association with clinical outcomes. We retrospectively analyzed an international multicenter database (COVID-D study) of critically ill adult patients with COVID-19 during the first pandemic wave and ABD using clinical features on day 1 of admission as input variables. We applied unsupervised machine learning in a pilot attempt to discover clusters of ABD patients. Hierarchical clustering was performed with a bootstrap-based robustness assessment after dimensionality reduction. Clusters were analyzed for differences in neurological outcomes, mechanical ventilation, and survival. We analyzed 1,631 critically ill COVID-19 patients with ABD, identifying four reproducible clusters with distinct clinical and neurological profiles. Cluster 1 ("mild respiratory failure,” n = 335) had the most favorable outcomes, with the shortest duration of delirium (4.13 days) and mechanical ventilation. Cluster 2 ("moderate ARDS," n = 508) showed a comparable delirium incidence but the longest duration (5.18 days). Cluster 3 ("early severe ARDS," n = 161) included patients who underwent prone positioning and mechanical ventilation early from the day of admission, with higher rates of coma (100
The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI development and use. Dataset shifts are defined as changes between train and test data distributions. Whether occurring over time (temporal) or across different sites (multi-source), they can severely degrade model performance and compromise data quality. This is particularly important in health AI, where the safety and fundamental rights of patients can be severely affected by uncontrolled shifts both at training and operational stages. While the theoretical foundations of covariate, prior, and concept shifts are well established, there is a lack of accessible and comprehensive software tools to perform their analysis. We introduce dashi, an open-source Python library designed for the exploration, quantification, and characterization of dataset shifts. dashi provides a dual approach: an unsupervised approach that leverages information geometry and non-parametric statistical manifolds to data variability characterization and analysis (e.g., Information Geometric Temporal plots and Multi-Source Variability metrics like Global Probabilistic Deviation and Source Probabilistic Outlyingness), and a supervised approach that quantifies and characterizes model performance degradation. Both unsupervised and supervised approaches work across user-defined temporal and domain/source batches. We demonstrate the utility of dashi on three simulated and real-world health AI case studies on gestational diabetes mellitus, COVID-19 and emergency medical dispatch. By providing interactive visual analytics and variability metrics, dashi supports trustworthiness of AI life cycle stages enabling robust and safe machine learning pipelines through the assessment of data coherence and AI performance.
Glioblastoma multiforme (GBM) is the most common and aggressive primary brain tumour in adults, with poor prognosis despite standard treatment with the Stupp protocol (surgical resection followed by concomitant temozolomide-based chemoradiotherapy and adjuvant temozolomide). Given the substantial clinical, functional and psychosocial burden of both the disease and its therapies, monitoring health-related quality of life (HRQoL) is essential. Smartphone-based approaches may enable continuous, low-burden monitoring; however, most existing tools rely mainly on self-reported outcomes and make limited use of passive digital phenotyping. To evaluate the feasibility and usability of Lalaby-Glio for monitoring HRQoL in patients with GBM, and to explore associations between subjective global QoL derived from patient-reported outcomes and objective sensor data. The study comprised four phases: (1) adaptation of the Lalaby platform for GBM-specific monitoring by combining passive smartphone sensor acquisitions (movement, step count, location, sound level/frequency, light exposure, internet use and call activity) with a clinician-designed daily 5-item questionnaire and weekly EORTC QLQ-C30 and QLQ-BN20 instruments; (2) a 6-week feasibility study in four adults (≥18 years) initiating Stupp treatment; (3) usability testing in GBM patients and a complementary non-patient sample of 103 healthy adults using either screenshot mock-ups or the installed app; and (4) exploratory analyses of associations between passive sensor data and subjective global QoL using Spearman correlations with patient-level bootstrap. Four patients (3 men, 1 woman; mean age 54.3 years) were enrolled; three completed all six weeks and one died after two weeks. Engagement was high, with 168 questionnaires completed (22 QLQ-C30, 19 QLQ-BN20 and 127 Daily Status entries). Mean app use was 1.8 (SD 2.7) minutes/day, and mean completion time for the weekly questionnaires was 7.2 (SD 3.3) minutes. Weekly PROMs revealed substantial inter- and intra-individual variability in functioning, symptoms and global QoL during treatment. Usability among GBM patients was favorable (overall mean 3.7/5). Among healthy participants (mean age 20.9 years; 69.9% women), those who installed the app (n=46) achieved a mean System Usability Scale score of 72.9, above the standard benchmark, and 91.3% judged the nature-inspired design appropriate for oncology (n=103). Passive sensing yielded 91.7 MB of data; global QoL correlated negatively with call activity (ρ=−0.598, P=.012) and movement (ρ=−0.604, P<.001), while steps showed a positive but non-significant association (ρ=0.392, P=.63). Lalaby-Glio appears feasible and usable for monitoring HRQoL in patients with GBM and can capture clinically meaningful fluctuations over time. Passive smartphone sensing shows promise for supporting digital phenotyping and identifying objective behavioral correlates of QoL. Future work should evaluate these findings in larger, multicenter cohorts and refine sensor data collection and processing to enable robust digital biomarkers for neuro-oncology monitoring and potential clinical integration.
BACKGROUND:Precise delineation of non-contrast-enhancing tumor (nCET) in glioblastoma (GB) is critical for maximal safe resection, yet routine imaging cannot reliably separate infiltrative tumor from vasogenic edema. The aim of this study was to develop and validate an automated method to identify peritumoral subregions compatible with nCET and assess its prognostic value. METHODS:Pre-operative T2-weighted and FLAIR MRI from 940 patients with newly diagnosed GB in four multicenter cohorts were analyzed. A deep-learning model segmented enhancing tumor, edema and necrosis; a non-local, spatially varying finite mixture model was applied to identify edema subregions characterized by relatively lower FLAIR hyperintensity, hypothesized to reflect nCET-related tissue. The ratio of these subregions to total edema volume defined the T2/FLAIR Heterogeneity Index (TFHI). Associations between TFHI and overall survival (OS) were examined with Kaplan-Meier curves and multivariable Cox regression. RESULTS:Higher TFHI values stratified patients with shorter OS. In the NCT03439332, TFHI above the optimal threshold was associated with a twofold increased hazard of death (hazard ratio [HR] 2.07, 95 % confidence interval 1.33-3.21; P = .0013) and a reduction in median survival of 98 days. Significant, though smaller, prognostic effects were confirmed in GLIOCAT & BraTS (HR= 1.37; P = .047), OUS (HR = 1.37; P = .0032) and pooled analysis (HR= 1.26; P = .0008). TFHI remained an independent predictor after adjustment for age, extent of resection and MGMT methylation. CONCLUSIONS:We present a reproducible, server-hosted tool for automated identification of imaging-defined, putative nCET-related peritumoral subregions and TFHI biomarker extraction that enables independent prognostic stratification. This approach provides a quantitative framework for studying peritumoral heterogeneity in GB.
Accurate early detection of clinically significant prostate cancer is crucial for improving patient outcomes. However, traditional diagnostic methods such as Digital Rectal Exam and Prostate-Specific Antigen (PSA) tests often lack the sensitivity and specificity needed for effective diagnosis. This study presents an AI-based approach for csPCa classification using MRI data, incorporating both the PI-CAI Challenge dataset and a newly compiled, diverse BIMCV Prostate dataset comprising over 9000 MRI sessions from 16 healthcare centers in the Valencian Region. The methodology includes a robust preprocessing pipeline, featuring prostate segmentation with a custom-trained nnUNet model, and utilizes a 3D variant of EfficientNet-B7. To ensure robustness, we employed a transfer learning strategy where five models pretrained on PI-CAI were fine-tuned on the BIMCV dataset and aggregated using a stacked meta-learner. This ensemble approach yielded a Receiver Operating Characteristic Area Under the Curve of 0.816 on the independent hold-out set, significantly outperforming a non-pretrained baseline (AUC 0.71). Furthermore, we demonstrated that synthesizing missing ADC maps using a mono-exponential model serves as an effective data augmentation strategy, preventing data loss without introducing domain shift. Interpretability techniques such as occlusion sensitivity and guided backpropagation were employed to provide insights into the model’s decision-making process, enhancing transparency. This research highlights the potential of AI-enhanced MRI techniques in advancing csPCa detection and diagnosis.
When developing machine learning models to support emergency medical triage, it is important to consider how changes over time in the input features can negatively affect the models’ performance. The objective of this study was to assess the effectiveness of novel deep continual learning pipelines in maximizing model performance when input features change over time, including the emergence of new features and the disappearance of existing ones. The model is designed to identify life-threatening situations, predict their admissible response delay, and determine their institutional jurisdiction. We analyzed a total of 1 414 575 events spanning from 2009 to 2019. We provided empirical evidence for covariate shifts and measured their negative effects on model performance. Then, we proposed a set of continual learning strategies to deal with 1) changing feature domains and 2) parameter updating over time. For changing feature domains, we proposed and assessed a static domain approach, a dynamic domain approach, and a predefined approach. For parameter updating over time, we designed and evaluated five approaches: from-scratch, fine-tuning, a cumulative, rehearsal and Elastic Weight Consolidation (EWC). Our findings demonstrate performance improvements, with the best strategy combination—dynamic feature domain coupled with EWC—yielding gains of up to 5.9
Ensuring trustworthy use of Artificial Intelligence (AI)-based Clinical Decision Support Systems (CDSSs) requires continuous evaluation of their performance and fairness, given the potential impact on patient safety and individual rights as high-risk AI systems. However, the practical implementation of health AI performance and fairness monitoring dashboards presents several challenges. Confusion-matrix-derived performance and fairness metrics are non-additive and cannot be reliably aggregated or disaggregated across time or population subgroups. Furthermore, acquiring ground-truth labels or sensitive variable information, and controlling dataset shifts-changes in data statistical distributions-may require additional interoperability with the electronic health records. We present the design of ShinAI-Agent, a modular system that enables continuous, interpretable, and privacy-aware monitoring of health AI and CDSS performance and fairness. An exploratory dashboard combines time series navigation for multiple performance and fairness metrics, model calibration and decision cutoff exploration, and dataset shift monitoring. The system adopts a two-layer database. First, a proxy database, mapping AI outcomes and essential case-level data such as the ground-truth and sensitive variables. And second, an OLAP architecture with aggregable primitives, including case-based confusion matrices and binned probability distributions for flexible computation of performance and fairness metrics across time or sensitive subgroups. The ShinAI-Agent approach supports compliance with the ethical and robustness requirements of the EU AI Act, enables advisory for model retraining and promotes the operationalisation of Trustworthy AI.
BackgroundReusing long-term data from electronic health records is essential for training reliable and effective health artificial intelligence (AI). However, intrinsic changes in health data distributions over time—known as dataset shifts, which include concept, covariate, and prior shifts—can compromise model performance, leading to model obsolescence and inaccurate decisions. ObjectiveIn this study, we investigate whether unsupervised, model-agnostic characterization of temporal dataset shifts using data distribution analyses through Information Geometric Temporal (IGT) projections is an early indicator of potential AI performance variations before model development. MethodsUsing the real-world Medical Information Mart for Intensive Care-IV (MIMIC-IV) electronic health record database, encompassing data from over 40,000 patients from 2008 to 2019, we characterized its inherent dataset shift patterns through an unsupervised approach using IGT projections and data temporal heatmaps. We trained and evaluated annually a set of random forests and gradient boosting models to predict in-hospital mortality. To assess the impact of shifts on model performance, we checked the association between the temporal clusters found in both IGT projections and the intertime embedding of model performances using the Fisher exact test. ResultsOur results demonstrate a significant relationship between the unsupervised temporal shift patterns, specifically covariate and concept shifts, identified using the IGT projection method and the performance of the random forest and gradient boosting models (P<.05). We identified 2 primary temporal clusters that correspond to the periods before and after ICD-10 (International Statistical Classification of Diseases, Tenth Revision) implementation. The transition from ICD-9 (International Classification of Diseases, Ninth Revision) to ICD-10 was a major source of dataset shift, associated with a performance degradation. ConclusionsUnsupervised, model-agnostic characterization of temporal shifts via IGT projections can serve as a proactive monitoring tool to anticipate performance shifts in clinical AI models. By incorporating early shift detection into the development pipeline, we can enhance decision-making during the training and maintenance of these models. This approach paves the way for more robust, trustworthy, and self-adapting AI systems in health care.
Glioblastoma (GBM) exhibits two principal growth phenotypes: infiltrative, characterized by diffuse invasion with minimal mass effect, and proliferative, characterized by pronounced tissue compression. Their quantitative delineation and prognostic implications remain uncertain. We introduce an MRI-derived biomarker, the dynamic infiltration rate (DIR), defined as the ratio of tumor-volume expansion to mass-effect–induced peritumoral compression, and evaluate it in silico and clinically. In a synthetic dataset spanning realistic infiltrative-proliferative spectra, DIR correlates strongly with ground truth (R^2=0.85). Applied to patient data, a data-driven threshold separates high- and low-infiltration groups with markedly different overall survival (median 16.0 versus 35.2 weeks; log-rank p<0.001; hazard ratio 2.49). Multivariate Cox analysis adjusted for age, sex, and MGMT status confirms DIR as an independent prognostic factor (HR = 1.38, 95 DIR therefore differentiates proliferative from infiltrative GBM phenotypes and provides prognostic information that could inform personalized therapy and follow-up.
INTRODUCTION:The management of HIV patients is inherently complex, requiring the integration of evolving clinical guidelines such as GESIDA 2023. These guidelines are extensive and demand coordinated implementation by healthcare professionals with diverse expertise. These challenges highlight the need for strategies that facilitate their adoption and improve patient care. In this study, we evaluated the impact of digitalizing clinical guidelines for AIDS patients on the efficiency of the clinical workflow in the internal medical services at a secondary public hospital. For this, we developed HIV-matic, a software to automate the implementation of the Spanish clinical guideline (GESIDA) for HIV management. MATERIALS & METHODS:The software codifies the guideline into a comprehensive set of 69 "if-then" rules within the CLIPS environment, including AIDS data, antiretroviral treatment recommendations, biochemistry data, comorbidities, and neoplasm screenings. These rules rely on 103 input variables, categorized into laboratory data from the Laboratory Information System (Gestlab), diagnostic data from the Hospital Information System (OrionClinic), and specific information compiled by the doctors through a structured questionnaire in the HIV-matic web interface. The clinical guideline digitalization was evaluated on a generated dataset of 191 prospective unique patient visits attended by clinical authors of the manuscript between 24/09/2024 and 20/12/2024. Evaluation metrics were calculated for rules and patients. RESULTS:When evaluating the rules, we observed a precision of 99,82%, a recall of 99,36%, a specificity of 99,96%, and an F1 score of 99,59% across the test dataset. Moreover, 87.43% of the evaluated patients received 100% correct messages. CONCLUSIONS:The GESIDA, implemented as a computerized clinical guideline in the HIV-matic software, supports clinicians in decision-making with a high level of performance. Its modular design and seamless integration with existing systems in the hospital make it scalable and adaptable for broader deployments in the health service for HIV patients.
BackgroundCoherence across sites in multicenter datasets is one substantial data quality dimension for reliable health data reuse, as unexpected heterogeneity in data can lead to biases in data analyses and suboptimal generalization of results. ObjectiveThis work aims to characterize and label the data coherence across sites in the first European multicenter dataset for cancer prevention in people and early detection among the homeless population in Europe: coadapting and implementing the health navigator model. This dataset emerged to enable research to address disparities in health challenges and health care access due to barriers such as unstable housing, limited resources, and social stigma in people experiencing homelessness. MethodsThe dataset comprises 652 cases: 142 from Austria, 158 from Greece, 197 from Spain, and 155 from the United Kingdom. All participants fit classifications from the European Typology of Homelessness and Housing Exclusion. This longitudinal study collected questionnaires at baseline, after 4 weeks, and at the end of the intervention. The 180-question survey covered sociodemographic data, overall health, mental health, empowerment, and interpersonal communication. Data variability was assessed using information theory and geometric methods to analyze discrepancies in distributions and completeness across the dataset. ResultsSubstantial variability was observed among the 4 pilot countries, both in the overall analysis and within specific domains. In particular, measures of health care empowerment, quality of life, and interpersonal communication demonstrated the greatest discrepancies among pilot sites, with the exception of the health domain. Notably, Spain exhibited the most pronounced differences, characterized by a high number of missing values related to interpersonal communication and the use of health care services. ConclusionsHealth data may be comparable across the 4 countries; however, substantial differences were observed in the other questionnaires, requiring independent, country-specific analyses. This study underscores the heterogeneity among people experiencing homelessness and the critical need for data quality assessments to inform future research and policymaking in this field.
The objective of this study was to build a multimodal, multitask predictive model-named E2eDeepEMC2-to improve out-of-hospital emergency incident severity assessments while coping with shifts in data distributions over time. We drew on 2054694 independent incidents recorded by the Valencian emergency medical dispatch service between 2009 and 2019 (excluding 2013), combining demographic, temporal, clinical and free-text inputs. To handle temporal drift, our model integrates continual learning strategies and comprises three encoder modules (for context, clinical data and text), whose outputs are merged to predict the life-threatening level, admissible response delay and emergency system jurisdiction. Compared with the Valencian Region's existing in-house triage protocol, E2eDeepEMC2 achieved absolute F1-score gains of 18.46% for life-threatening level, 25.96% for response delay and 3.63% for jurisdiction. Compared to non-continual learning baselines, it also outperformed them by 3.04%, 9.66% and 0.58%, respectively. Deployment of E2eDeepEMC2 is currently underway in the Valencian Region, underscoring its practical impact on real-world emergency dispatch decision-making.
Scientific software development in research institutions often focuses on demonstrating technology capabilities, with limited attention to formal documentation. This limits its use beyond academic projects, particularly in pharmaceutical industry collaborations, where Good Manufacturing Practices (GMP) are essential for operational deployment. We propose a methodology based on GMP guidelines tailored for research institutions to bridge this gap. Divided into four stages-project premise, prototyping, GAMP 5 V-Model development, and software transference-this approach accelerates development, ensures compliance, and facilitates the transfer of R&D results to the pharmaceutical industry.
Trustworthy health Artificial Intelligence (AI) must respect human rights and ethical standards, while ensuring AI robustness and safety. Despite the availability of general good practices, health AI developers lack a practical guide to address the construction of trustworthy AI (TAI). We introduce a TAI development framework (TAIDEV) as a reference guideline for the creation of TAI health systems. The framework core is a TAI matrix that classifies technical methods addressing the EU guideline for Trustworthy AI requirements (privacy and data governance; diversity, non-discrimination and fairness; transparency; and technical robustness and safety) across the different AI lifecycle stages (data preparation; model development, deployment and use, and model management). TAIDEV is complemented with generic, customizable example code pipelines for the different requirements with state-of-the-art AI techniques using Python. A related checklist is provided to help validate the application of different methods on new problems. The framework is validated using two open datasets, the UCI Heart Disease and the Diabetes 130-US Hospitals, with four code pipelines adapting TAIDEV for each dataset. The TAI framework and its example tutorials are provided as Open Source in the GitHub repository: https://github.com/bdslab-upv/trustworthy-ai. The TAIDEV framework provides health AI developers with an extensible theoretical development guideline with practical examples, aiming to ensure the development of ethical, robust and safe health AI and Clinical Decision Support Systems.
Good pharmacy and manufacturing practices identify data integrity as a critical factor for patient safety. However, many of the processes in hospital pharmacies are still documented manually, which can compromise the integrity of the information. This work presents a software customization prototype for data integrity assurance in hospital pharmacy formulation and dispensing processes. The proposed solution combines the implementation of ALCOA+ principles by design, GAMP5 V-model and human-center design methodologies to adapt a validated data integrity software for a hospital pharmacy use case. Moreover, our proposal is being evaluated in the hospital pharmacy environment to achieve a higher readiness level. After complete validation, the software customization can support the digitalization of drug formulation and dispensing processes and improve patient safety.
Alfons Juan合作论文数Departament of Computer Systems and Computation, Polytechnic University of Valencia5