Abstract Anthropomorphism—the human propensity to attribute human characteristics to nonhuman entities—has long preoccupied cognitive psychology, philosophy, and human–computer interaction. Yet the term is frequently mobilized as a catch-all label, potentially obscuring crucial differences among design strategies, user inferences, and social norms. Drawing on classical philosophy, Roman legal theory, psychology, and contemporary human–computer interaction research, this article offers a systematic framework that disaggregates anthropomorphism into seven interlocking constructs: anthropomimesis, ethopoiesis, impersonation, identification, theory of mind, theory of machine, and personification. This article argues that personification—understood as role-based recognition—constitutes the pivotal bridge between technical design and social expectations. The resulting model clarifies how humans design, interpret, and govern intelligent artifacts, and it suggests practical guidelines for ethically responsible AI and robot design.
Machine learning (ML) models are increasingly proposed to support clinical decision-making, yet their evidentiary basis remains weaker than their publication volume suggests. This editorial argues that the problem is not only translational, regulatory, or infrastructural, but methodological. Many medical ML pipelines rely on uncertain ground truths, optimize performance around clinically irrelevant thresholds, report unstable or prevalence-dependent metrics, neglect calibration and uncertainty, and lack rigorous external and temporal validation. These weaknesses produce optimistic estimates that do not reliably anticipate performance in heterogeneous clinical settings. We call for an evidence-based medical AI grounded in more reliable annotation practices, explicit modeling of uncertainty, clinically meaningful threshold selection, calibration and decision-utility analyses, robustness testing, external validation on independent datasets, and post-deployment monitoring. The editorial also invites authors, reviewers, users, and vendors to adopt stricter standards so that predictive models can become credible, accountable, and clinically useful tools in everyday practice, rather than merely publishable artifacts.
Flexible bronchoscopy is an essential tool for airway management and both diagnostic and therapeutic interventions, particularly in critical care. Accurate identification of tracheobronchial structures is crucial but challenging for less experienced clinicians, often leading to prolonged procedures and increased complication risks. Simulation-based training using virtual reality or manikins has shown promise, and recent studies suggest that artificial intelligence (AI)-based training outperforms self-directed learning. Limited data exist comparing AI-based bronchoscopy training to expert-led instruction. This study aimed to develop and evaluate a custom-made AI-based software for identifying key tracheobronchial structures and assessing its effectiveness as a training tool for anesthesia and intensive care residents. An AI-based software using YOLOv8 artificial neural networks was developed to recognize key tracheobronchial structures from bronchoscopy videos of a high-fidelity manikin. In a randomized trial, 22 second-year anesthesia residents with limited bronchoscopy experience were assigned to either AI-based unsupervised training (n=11) or traditional human-led training (n=11). Bronchoscopy skills were assessed using the modified Bronchoscopy Skill and Task Assessment Tool (BSTAT) before and after training. The AI model demonstrated high accuracy, with an average precision-recall AUC of 0.98 and a mean average precision of 0.98. Both groups of residents showed significant improvement in their BSTAT scores (from 30±4 to 53±2, p<0.001) and reduced procedural time (from 217±44 to 101±23 seconds, p<0.001). No significant differences were observed between the AI-based and expert-led training groups. We developed an AI-based software capable of real-time guidance during flexible bronchoscopy. The AI-based training demonstrated comparable efficacy to expert-led instruction, suggesting its potential as a viable tool for unsupervised medical training in flexible bronchoscopy.
Conversational agents are transforming digital interactions across various domains, including healthcare, education, and customer service, thanks to advances in large language models (LLMs). As these systems become more autonomous and ubiquitous, understanding what constitutes high-quality interaction from a user perspective is increasingly critical. Despite growing empirical research, the field lacks a unified framework for defining, measuring, and designing user-perceived interaction quality in human–artificial intelligence (AI) dialogue. Here, we present an integrative review of 125 empirical studies published between 2017 and 2025, spanning text-, voice-, and LLM-powered systems. Our synthesis identifies three consistent layers of user judgment: a pragmatic core (usability, task effectiveness, and conversational competence), a social–affective layer (social presence, warmth, and synchronicity), and an accountability and inclusion layer (transparency, accessibility, and fairness). These insights are formalised into a four-layer interpretive framework—Capacity, Alignment, Levers, and Outcomes—operationalised via a Capacity × Alignment matrix that maps distinct success and failure regimes. It also identifies design levers such as anthropomorphism, role framing, and onboarding strategies. The framework consolidates constructs, positions inclusion and accountability as central to quality, and offers actionable guidance for evaluation and design. This research redefines interaction quality as a dialogic construct, shifting the focus from system performance to co-orchestrated, user-centred dialogue quality.
Artificial intelligence (AI) systems are increasingly integrated into decision-making across high-stakes domains, influencing not only task performance but potentially human learning. While prior research has focused on AI’s impact on accuracy, its role in supporting long-term knowledge acquisition remains underexplored. This study investigates whether AI-based decision support systems can serve as implicit mentors, enabling users to internalize novel decision strategies through repeated interaction, a phenomenon we term machine mentoring. In a simulated diagnostic task, 289 medical students received different forms of AI support. Participants in the Feedback condition, who received trial-by-trial feedback on the correctness of both their own and the AI’s decisions, significantly improved their accuracy on cases involving a hidden diagnostic criterion and retained these gains in subsequent unaided trials. By contrast, AI advice alone, or advice paired with confidence indicators, did not lead to significant learning outcomes. These findings indicate that evaluative, trial-by-trial feedback is the key driver of AI-supported knowledge transfer: advice alone, even when paired with calibrated confidence cues, did not yield durable learning. Rather than attributing the effect to AI per se, our results indicate that what drives durable learning is access to trial-level, ground-truth corrective feedback—something competent AI systems can help deliver at scale. This is especially relevant in medical education and simulation contexts, where feedback can be reliably validated.
Conventional performance metrics in clinical decision support systems, such as accuracy or sensitivity, fail to reflect the reliability of individual predictions-an essential concern for clinicians operating in high-stakes environments. We introduce a calibration-informed framework featuring two novel metrics: the Local Predictive Value (LPV) and the Credible Predictive Value (CPV). LPV estimates the empirical reliability of a prediction by assessing the observed correctness frequency in the neighborhood of its confidence score. CPV refines this estimate using a Bayesian approach, integrating global predictive values as priors to produce a posterior distribution over correctness probabilities. LPV offers a descriptive, data-driven view of local reliability, while CPV provides a belief-adjusted estimate that mitigates overfitting to sparse local data. Applied to benchmark medical imaging datasets, these metrics yielded locally adaptive, interpretable reliability estimates. Divergences between LPV and CPV identified cases where local evidence was insufficient or misleading, highlighting how Bayesian smoothing improves stability against sparse or misleading local evidence. By combining local calibration with Bayesian inference, LPV and CPV advance the development of medical AI systems that are not only accurate but also interpretable and trustworthy at the individual case level.
Artificial intelligence (AI) is expanding in gastroenterology, particularly in endoscopy and imaging, where models support detection, classification, and risk stratification. However, translation into routine practice is hindered by inconsistent evaluation practices and limited clinical interpretability of reported metrics. We do not propose a generic reporting checklist; rather, we address a specific operational gap: insufficient support for clinicians in translating model outputs into use-case-specific decision thresholds, calibration assessment, and decision-analytic evaluation of clinical utility. These recommendations are based on narrative synthesis and expert opinion. We argue that evaluation should be anchored to pre-specified, clinically justified decision thresholds — or to a transparent interval of plausible thresholds elicited for the intended use case and prospective user community — rather than relying on global measures such as AUROC. We highlight how prevalence, calibration, and asymmetric error costs shape sensitivity, specificity, and predictive values. We recommend visualizations —ROC curves with marked operating points, calibration plots, and decision curve analysis — to communicate threshold-dependent behavior, probability reliability, and expected clinical utility. Beyond summary metrics, we emphasize the need for external validation, structured error analysis, and robustness assessment under distribution shift. Finally, we discuss transparency requirements under emerging European regulatory frameworks. These recommendations aim to support safer and more interpretable adoption of AI tools in gastroenterology
Artificial intelligence (AI) is increasingly being integrated into Clinical Decision-Support Systems (CDSSs), shifting attention from algorithmic performance alone to the broader sociotechnical conditions that shape effective human–AI collaboration. In this study, we investigated whether nine displacement-based structured coordination protocols can improve the collective diagnostic decision-making of hybrid human–AI teams (16 board-certified radiologists and a simulated AI model) in a radiological double-reading task for vertebral fracture detection from X-ray images. Among the protocols tested, the Accuracy-Oriented, Confidence-Oriented, and Presumptuous strategies achieved the highest (balanced) accuracy overall, with up to 97% among strong clinicians and 92% among weak ones, significantly outperforming simpler methods like majority voting. Conversely, approaches optimized for a single metric (e.g., sensitivity or specificity) introduced performance trade-offs. Benefits were strongest among less proficient clinicians, which exhibited substantial and consistent improvements, while proficient clinicians showed limited gains and occasional declines. Critically, Kasparov’s Law emerged as a comparative framework for empirically evaluating coordination quality relative to the diagnostic task, clinical objective, and clinician proficiency by identifying situations in which less proficient clinicians supported by superior coordination protocols outperformed more proficient clinicians operating under inferior ones. These findings demonstrate that coordination design is a critical determinant of hybrid human–AI decision-making, highlighting that a well-structured process can be more relevant than individual components’ performance and support process-centered approaches to the development and evaluation of CDSSs.
The article explores the shift from Asimov's laws, centered on machine obedience, to Kasparov's laws of "hybrid intelligence". While Asimov focused on preventing harm through autonomous constraints, Kasparov emphasizes that the best performance arises from the optimal orchestration of human, machine, and process. This perspective suggests that a "weak human + machine + superior process" can outperform a "strong human + machine + inferior process". Empirical studies in radiology are consistent with this socio-technical conjecture. Studies in radiology indicate that specific collaboration protocols allow human-AI teams to surpass isolated models. Notably, research confirms that less proficient clinicians embedded in effective protocols can achieve higher accuracy than clinicians with higher baseline accuracy operating under less effective protocols. This framework views AI as a component of "superminds" - collective cognitive architectures that enhance plural decision-making. Ultimately, the value of AI is an emergent property of the organizational system. Rather than focusing solely on model accuracy, designers must create interaction protocols that calibrate trust and prevent professional deskilling. The goal is to move toward a synergy where machines help human collectives become more intelligent.
AI systems are widely proposed as second-opinion advisors in clinical diagnosis, offering the promise of enhancing decision accuracy and clinician confidence while preserving human oversight. However, successful deployment in real-world practice faces a critical barrier: clinicians' reliance on AI is often miscalibrated, manifesting as misuse (over-reliance driven by automation bias) and disuse (under-utilization driven by self-anchoring bias). This paper addresses these deployment challenges by systematically analyzing how such reliance patterns affect diagnostic accuracy, confidence, and decision-making across diverse medical specialties. We report results from controlled simulations involving over 300 medical professionals across six diagnostic settings—including knee MRI analysis, spinal X-rays, cardiac ECG evaluation, and gastrointestinal endoscopy—using a human-first, AI-second workflow. Although AI advice improved average diagnostic accuracy (+2 percentage points) and clinician confidence (+3 points on a normalized scale), overall levels of appropriate reliance remained well below 50%, with disuse emerging as the more prevalent and consequential barrier. We introduce and validate Appropriate Reliance as an actionable metric for assessing and improving human-AI collaboration, providing practical guidance for developers, healthcare institutions, and policymakers seeking to deploy second-opinion AI systems safely and effectively. By identifying the sociotechnical barriers and offering evidence-based design insights, this work supports the emerging application of AI as a collaborative advisor in clinical workflows, charting a clear path toward deployment that enhances diagnostic safety, accountability, and patient care. Specifically, we propose integrating the Appropriate Reliance metric into system development workflows, clinician training, and regulatory evaluations to enable safe and effective deployment of second-opinion AI systems.
Spinal surgery carries substantial perioperative risks. Early identification of high-risk patients is critical for improving outcomes. Machine learning (ML) can enhance predictive accuracy over traditional risk scores by modeling complex clinical data, but many existing models lack large, heterogeneous cohorts. To develop and validate ML models for predicting perioperative complications in spine surgery, and assess fairness across patient subgroups. We conducted a retrospective cohort study of 5,060 adult patients from the SpineReg registry (2015–2023), each with 160 preoperative demographic, clinical, imaging, and patient-reported outcome features. Six ML algorithms—logistic regression (LR), decision tree, k-nearest neighbors, naïve Bayes, random forest (RF), and eXtreme gradient boosting—were trained with 80/20 train-test split, cross-validation, and hyperparameter optimization. Class imbalance was addressed via repeated undersampling. Performance was assessed using area under the ROC curve (AUC), balanced accuracy, positive predictive value (PPV), standardized net benefit (sNB), calibration, and fairness metrics. RF achieved the best performance, obtaining an AUC of 0.87, balanced accuracy of 0.80, PPV of 0.71, and sNB of 0.44. RF maintained robust performance across demographic and surgical subgroups. Key predictors of increased complication risk included sagittal imbalance, and multilevel surgery, whereas degenerative pathology and monosegmental fusion were protective. Sensitivity declined for rare complications but exceeded 75
Artificial intelligence (AI) is increasingly applied in clinical research to improve diagnostic accuracy, prognostic modeling, and disease monitoring, particularly in complex, data-constrained areas such as rare diseases. In this context, Myasthenia Gravis (MG)—a rare neuromuscular disorder—has seen a growing number of AI-driven studies aimed at addressing its diagnostic and clinical heterogeneity. However, the methodological quality and reporting standards of these studies remain largely unexamined. This study presents the first longitudinal evaluation of AI research in MG using the CLARITY AI framework—a modular tool designed to assess both structural rigor and macro-topical completeness. An analysis of 20 peer-reviewed studies published between 2020 and 2024 revealed that average total scores improved from 235.0 in 2021 to 256.15 in 2024, with notable gains in evaluation metrics (+2.33), data availability (+1.32), and study type and objectives (+1.55). A post hoc robustness check confirmed the stability of these temporal trends. Despite these improvements, critical deficiencies persist, particularly in usability testing, user engagement, and ethical reporting, where scores often remained below 1. These findings indicate that while technical sophistication is advancing, translational readiness continues to be limited by the underreporting of human-centered and ethical dimensions. This work provides actionable insights and establishes a benchmark for improving transparency, methodological rigor, and clinical relevance in AI applications for rare diseases.
Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human–AIpractices.
Pixel Attribution Methods (PAMs) are essential techniques in Explainable Artificial Intelligence (XAI) for improving the interpretability of black-box models in diagnostic imaging, particularly through saliency maps that highlight regions of interest in medical images. Despite their potential, existing PAMs often fail to align with clinical reasoning and radiological practices, limiting their applicability in real-world settings. This study addresses this gap by evaluating the semantic significance, defined as relevance and pertinence, of saliency maps generated by different PAMs (Grad-CAM, XGrad-CAM, Grad-CAM++, and Smooth Grad-CAM++) at layers 3 and 4 of a ResNeXt-50 neural network. Using a dataset of 12 thoraco-lumbar X-rays, we quantified the alignment between AI-generated saliency maps and human-defined maps created by six specialists, which were then aggregated using various criteria. Among the evaluated methods, Smooth Grad-CAM++ exhibited the lowest performance, while Grad-CAM, XGrad-CAM, and Grad-CAM++ yielded comparable results. Furthermore, saliency maps from layer 4 demonstrated significantly higher relevance than those from layer 3, although the difference in pertinence was less pronounced. However, saliency maps consistently underperform compared to expert ground truth annotations. These findings underscore the limited clinical significance of current saliency map approaches, highlighting the need for more human-centered AI models that better integrate into real-world diagnostic decision-making processes.
Achieving appropriate human reliance on Artificial Intelligence (AI) systems remains a central challenge in Human-Computer Interaction. Confidence scores—indicators of an AI system’s certainty in its recommendations—have been proposed as a means to help users calibrate their trust and reliance on AI Decision Support Systems (DSS). However, limited research has explored how well-calibrated versus miscalibrated confidence scores affect human decision-making. We report a study examining the effects of confidence calibration on user reliance, decision accuracy, and perceived utility of an AI DSS. In a within-subjects experiment involving 184 participants solving logic puzzles, we found that well-calibrated confidence scores significantly improved decision accuracy (+20%, 95% CI: [0.18, 0.23]), whereas miscalibrated scores yielded minimal accuracy gains (+2%, 95% CI: [-0.00, 0.04]) and increased vulnerability to automation bias and conservatism bias. Participants were more likely to accept AI recommendations when high confidence was expressed, even when those recommendations were incorrect, resulting in errors. Conversely, miscalibrated and low-confidence recommendations increased conservatism bias, leading users to reject even accurate AI suggestions. Perceived utility of the AI system was higher when confidence levels were high (p < 0.001) and when confidence was well-calibrated (p = 0.002). These findings underscore the importance of designing AI systems with properly calibrated confidence cues to improve human-AI collaboration and mitigate reliance-related biases.
The increasing integration of artificial intelligence (AI) in decision-making processes has amplified discussions surrounding algorithmic authority—the perceived epistemic legitimacy of AI systems over human judgment. This study investigates how individuals attribute epistemic authority to AI, focusing on psychological, contextual, and sociotechnical factors. Existing research highlights the importance of trust in automation, perceived performance, and moral frameworks in shaping such attributions. Unlike prior conceptual or philosophical accounts of algorithmic authority, our study adopts a relational and empirically grounded perspective by operationalizing algority through psychometric measures and contextual assessments. To address knowledge gaps in the micro-level dynamics of this phenomenon, we conducted an empirical study using psychometric tools and scenario-based assessments. Here, we report key findings from a survey of 610 participants, revealing significant correlations between trust in automation (TiA), perceptions of automated performance (PAS), and the propensity to defer to AI, particularly in high-stakes scenarios like criminal justice and job-matching. Trust in automation emerged as a primary factor, while moral attitudes moderated deference in ethically sensitive contexts. Our findings highlight the practical relevance of transparency and explainability for supporting critical engagement with AI outputs and for informing the design of contextually appropriate decision support. This study contributes to understanding algorithmic authority as a multidimensional construct, offering empirically grounded insights for designing AI systems that are trustworthy and context-sensitive.
BACKGROUND:Artificial intelligence (AI) is increasingly used in clinical workflows, but its psychological effects on diagnostic confidence remain unclear. OBJECTIVE:To examine how agreement or disagreement with AI recommendations affects clinicians' confidence, and whether effects vary by skill level. METHODS:Across three diagnostic domains, clinicians (N = 292) rated their diagnostic confidence before and after reviewing fixed-accuracy AI advice. Confidence change was analyzed by agreement/disagreement and user skills. RESULTS:Agreement with AI increased confidence (M = 0.046), even when both human and AI were wrong (M = 0.048). Disagreement slightly reduced confidence (M = -0.008). These effects were consistent across skill levels. CONCLUSION:Agreement with AI acts as a cognitive reinforcer, increasing confidence regardless of correctness or user skill, revealing risks of overreliance in AI-supported diagnosis.
Background: Machine learning (ML) is increasingly applied in medicine, underscoring the need for transparent and clinically relevant models. In gastrointestinal oncology, most ML studies rely on raw imaging data, which limits clinical adoption due to poor interpretability and the difficulty of collecting high-quality, large-scale video and image datasets in routine practice. Endoscopic ultrasound (EUS) plays a central role in the evaluation of pancreatic cancer; however, structured EUS features remain underused in predictive modeling. Objective: To assess the performance and interpretability of ML models for diagnosing pancreatic ductal adenocarcinoma (PDAC) using routinely collected EUS variables. Methods: We conducted a retrospective multicenter study using data from two Italian hospitals (n = 641) for model training and internal validation and from a third hospital (n = 120) for external validation, collected from 2015 to 2023. Decision trees, random forests, naïve Bayes and other classifiers were developed and evaluated. Model performance was assessed in terms of discriminative ability, calibration, and selective prediction. Results: All models demonstrated high discriminative performance (AUC ≥ 0.90). Decision trees provided the most favorable balance between interpretability and accuracy (balanced accuracy = 0.87; sensitivity = 0.89). Calibration and selective prediction analyses confirmed the robustness of the models. Conclusions: These findings demonstrate the feasibility of implementing interpretable yet high-performing ML models for PDAC diagnosis in real-life endoscopic settings.
Artificial Intelligence (AI) has become a pivotal tool in augmenting human decision-making across various do mains, yet its influence on user decisions often lacks comprehensive evaluation. While technical performance metrics such as accuracy and efficiency dominate AI design, integrating human-centered approaches that con sider trust and reliance remains underexplored. This study addresses the knowledge gap in understanding how AI systems influence decision-making quality, calibrated to user profiles, including their expertise, skills, professional role, confidence, and reliance tendencies. We present a novel and comprehensive metric framework for evaluating AI influence, emphasizing behavioral patterns and measurable improvements in decision outcomes beyond simple alignment with AI recommendations. The framework is applied to four medical domain case studies-MRI, ECG, X-ray, and ENDO-with user groups spanning specialists, sub-specialists, and trainees. Results reveal that while human and AI systems achieve high agreement rates (up to 81%), AI influence on decision quality varies significantly. Notably, X-ray decision-making showed the highest influence index (0.27), while MRI decisions exhibited substantial self-anchoring bias (6.94), undermining the potential positive impact of AI. Influence metrics unveiled nuances missed by agreement scores, highlighting domain-specific biases and opportunities to optimize AI-human interaction. This research underscores the necessity of adapting the type of AI system and affordances to user charac teristics and attitudes of reliance to foster calibrated trust and improve decision outcomes. Our findings inform the design of AI systems that better support diverse user needs and align with human decisions, driving progress toward human-centered AI integration in high-stakes domains.
Hugo Gamboa合作论文数Departamento de Sistemas e Informatica (Gabinete F269)
Escola Superior de Tecnologia de Setubal do I.P.S.6
Rosa Lanzilotti合作论文数Department of Computer Science, University of Bari4