
IntroductionRadiologists in resource-limited settings often face high workloads, especially in chest X-ray interpretation. Manual annotation of large-scale imaging datasets remains costly and time-consuming. This study aims to explore the feasibility of using large language models (LLMs), specifically GPT-4o, to generate binary disease presence labels from free-text radiology reports, and to use these labels to train deep learning models for automated chest X-ray classification.MethodsA two-stage supervised learning pipeline was developed using the publicly available MIMIC-CXR v2.1.0 dataset. First, GPT-4o was prompted with a structured clinical protocol to classify each radiology report as either “diseased” or “no disease.” Second, the generated labels were used to supervise the training of four convolutional neural networks: ResNet-18, DenseNet-121, EfficientNet-B1, and ConvNeXt-Tiny. A patient-level 70/10/20 split was employed to prevent data leakage across sets. Each model was trained across five random seeds (42–46), and 95% confidence intervals were computed using the t-distribution. Label quality was evaluated by comparing 210 generated labels against radiologist annotations from a board-certified radiologist.ResultsGPT-4o achieved an overall accuracy of 92.9% with expert labels on the 210-report validation set. For the “diseased” class, the precision was 97.4% and recall was 90.5%; for “no disease,” precision was 87.1% and recall was 96.4%. Among the CNN models evaluated on the held-out test set, ConvNeXt-Tiny achieved the highest area under the curve (AUC=0.832, 95% CI [0.801, 0.863]) and balanced accuracy (0.739), significantly outperforming EfficientNet-B1 (AUC=0.797; paired t-test, p=0.014). ResNet-18 (AUC=0.822) and DenseNet-121 (AUC=0.808) showed intermediate performance. All models demonstrated AUC values above 0.79, confirming the viability of LLM-derived weak supervision.DiscussionThis study demonstrates that LLMs can be effectively employed to generate supervision labels for medical imaging tasks. The proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability. The multi-seed evaluation with confidence intervals provides a rigorous assessment of model stability. Further work is needed to improve label reliability and expand to multi-label classification.
IntroductionTeleophthalmology depends on secure exchange of retinal fundus images between primary care screening locations and referral centers. Current image-sharing systems remain vulnerable to quantum attacks on public-key cryptography and to centralized storage failures with limited auditability.MethodsThis paper presents PQ-FundusChain, a consortium blockchain framework for quantum-resistant fundus image sharing. Each image is encrypted with an AES-256-GCM session key encapsulated using ML-KEM (FIPS 203). ML-DSA (FIPS 204) authenticates blockchain transactions. Encrypted images are stored in decentralized off-chain storage, while content hashes, metadata, and access policies are recorded on-chain. Access requests are evaluated through a risk-aware policy combining trust, access-history, and location scores with a fundus sensitivity score based on retinal anatomy and diabetic retinopathy severity.ResultsExperimental evaluation on the public Messidor-2 dataset shows a share-side cryptographic and hashing cost of approximately 4.02 ms per image, an access-side cost of approximately 3.48 ms, smart-contract execution of approximately 0.25 ms for both AccessGrant and PolicyUpdate, and a per-image on-chain record of approximately 4.83 KB independent of image resolution. The additional latency introduced by postquantum mechanisms remains below 1 ms for image sharing.DiscussionThe results demonstrate that postquantum protection can be incorporated into teleophthalmology workflows with modest computational overhead, with security analysis confirming confidentiality, integrity, authentication, non-repudiation, fine-grained access control, and resistance to harvest-now, decrypt-later attacks.
IntroductionOpioid Use Disorder (OUD) is an ongoing and pervasive public health problem. Despite the availability of efficacious Medication treatments for OUD (MOUD), many patients require additional support to address OUD's wide-ranging consequences. Mobile Health Interventions (MHIs), or smartphone-based software applications (“apps”), could serve as a scalable and flexible medium to increase adjunctive care access within MOUD.MethodAs an initial step in potential MHI creation, we conducted semi-structured interviews with U.S. military veterans receiving MOUD (N = 21) via rapid qualitative methods. Our team developed a start-list of interview questions. Interview transcripts were independently coded by interviewers, who then convened to discuss and refine code definitions, develop summary templates, organize codes into hierarchical structures, and refine domain categories through consensus.ResultsTwenty-one veterans participated [M[SD]Age = 59.95[10.50]; 90.48% male sex-assigned-at-birth; 80.95% White; 85.71% Non-Hispanic/Latinx; 52.38% on buprenorphine, 47.62% on methadone at time of interview]. Veteran attitudes were typically either neutral- positively valenced and characterized by willingness (Open Prudency) or neutral-negatively valenced and characterized by concern and contingencies for use, such as assurance for continued access to in-person care (Hesitancy). A small but notable proportion of veterans expressed strongly favorable views towards MHIs, often grounded in their own technological fluency (Enthusiasm). Outright Opposition to MHIs was rare. Desired MHI functions were variable and included their use for enhancing care accessibility and convenience (e.g., Resource Centralization, Patient-Provider Communication), as well as supporting autonomy, motivation and individualization (e.g., Self-Monitoring, Positive Reinforcement, Adjunctive Intervention).ConclusionsResults from this qualitative investigation of U.S. military veterans provide foundational data to guide future development of MHIs for supporting MOUD engagement. Findings suggest that veterans receiving MOUD are generally open towards and interested in using MHIs as part of their care regimens, which were broadly described as potential means to enhance treatment accessibility, convenience, and individualization.
BackgroundRadiology reports contain critical information for monitoring patients' clinical progress and evaluating treatment outcomes. Sequential radiological images are particularly valuable, providing not only current findings but also comparative assessments against previous reports. Automatically detecting temporal changes between free-text reports remains a significant challenge in Natural Language Processing.MethodsThis study evaluated the performance of Large Language Models in detecting temporal novelty using the LUNGUAGE dataset, which contains sequential chest radiography reports. The comparative multi-model and multi-prompt evaluation was carried out on a fixed, class-stratified subset of 1,500 of these examples. Seven language models from OpenAI, Anthropic, and Google Gemini were tested across five prompt strategies: zero-shot, few-shot, structured reasoning, memory-augmented, and temporal graph prompting. Each model classified target findings as new, improved, worsened, stable, or resolved by evaluating current and previous reports together. Performance was measured using accuracy, macro-F1, weighted-F1, and class-based metrics. The “new” class was analyzed separately due to its clinical importance. Model errors were categorized into clinically interpretable types, including missed new findings, missed resolved findings, temporal errors, and change direction errors. Bootstrap confidence intervals and paired McNemar significance tests assessed performance uncertainty.ResultsGPT-4.1 emerged as the top-performing model. Its few-shot prompt strategy achieved the best results: 0.8369 accuracy, 0.7888 macro-F1, 0.8322 weighted-F1, and 0.5385 recall for the “new” class.ConclusionThese findings demonstrate that Large Language Models hold substantial promise for temporal clinical reasoning in sequential radiology reports, and that performance depends not only on model architecture but also on prompt strategy and the class being evaluated.
BackgroundImproved methods of cognitive testing are urgently needed for primary care settings faced with a growing older adult population and higher rates of dementia. We tested the feasibility and acceptability of a novel online cognitive test that can be completed at home prior to an annual exam.Methods32 older adults completed testing at home on personal devices 1–4 weeks prior to an annual exam with their primary care provider. Additional cognitive screening was completed in-clinic to provide preliminary validation evidence. Completion rates were examined to assess feasibility, and patient acceptability was assessed via survey.ResultsCompletion rate was 78.5% for at-home testing online. Participants reported that they generally preferred at-home over in-clinic testing. At-home test performance was moderately correlated with the standard Montreal Cognitive Assessment completed in clinic.ConclusionFindings provide initial support for the acceptability of self-administered online cognitive screening for older adults in primary care. Completion rates were slightly below a preset benchmark, suggesting that additional support may be needed to facilitate at-home testing in routine care. Further validation research is needed in larger samples including examination of digital measure convergence with clinical diagnoses and additional cognitive testing.
BackgroundArtificial intelligence (AI) and digital health technologies are increasingly shaping how public healthcare services are planned, delivered, and governed. In Egypt, national digital health reforms have created an urgent need to assess whether public healthcare institutions have the administrative readiness, institutional capacity, and governance arrangements required for AI-enabled digital health transformation.MethodsA sequential explanatory mixed-methods design was employed across 24 Egyptian governorates. Phase 1 consisted of a cross-sectional survey of 387 healthcare administrators and policymakers from Ministry of Health and Population hospitals, university hospitals, primary healthcare units, and Health Insurance Organization facilities. The survey measured eight operationalized constructs: institutional capacity, administrative readiness, organizational governance, leadership support, IT infrastructure, staff training and skills, regulatory framework, and digital health transformation success. Phase 2 comprised semi-structured interviews with 18 senior policymakers, hospital leaders, digital health consultants, health information system managers, and health policy researchers. Quantitative findings were analyzed using descriptive statistics, confirmatory factor analysis, and structural equation modeling, while qualitative data were analyzed thematically and integrated with the survey results through explanatory joint display logic.ResultsAmong survey respondents, 62.3% were male, 71.6% held a master's degree or higher, and 58.1% had more than 10 years of professional experience. The structural model showed acceptable fit (CMIN/DF = 2.14, CFI = .96, TLI = .95, RMSEA = .045, SRMR = .038) and explained 68% of the variance in perceived digital health transformation success. Institutional capacity, administrative readiness, and organizational governance were positively associated with transformation success. The specified indirect pathways through leadership support, IT infrastructure, staff training and skills, and regulatory framework were also significant and are interpreted as hypothesis-generating associations because of the cross-sectional design. Interview findings helped explain the quantitative patterns by identifying fragmented governance, weak infrastructure, workforce digital literacy gaps, regulatory ambiguity, and resource allocation constraints.ConclusionsEgypt's public healthcare system faces interrelated institutional, administrative, workforce, infrastructure, and regulatory barriers to AI-enabled digital health transformation. Strengthening governance coordination, infrastructure equity, workforce capability, and AI-specific regulatory safeguards is essential before large-scale AI deployment can be reliably translated into public value.
Coronary artery disease (CAD) is one of the leading causes of global mortality, necessitating accurate assessment of coronary artery stenosis and plaque-associated angiographic findings from coronary angiography images. Although deep learning-based diagnostic models have demonstrated high predictive capability, their deployment on resource-constrained embedded platforms remains challenging because of computational complexity and memory requirements. To address these limitations, this study proposes a knowledge-distillation-based hardware-aware framework that transfers discriminative knowledge from a high-capacity teacher network to a lightweight convolutional neural network (CNN). The proposed framework integrates software-level localization of plaque-associated angiographic findings with FPGA-oriented optimization. Using a 90% training and 10% testing split, the proposed model achieved an accuracy of 98.64%, precision of 99.06%, recall of 97.82%, PPV of 0.99, NPV of 0.98, MCC of 0.96, and an AUC of 0.9940. And under 10-fold cross-validation, the model achieved a mean accuracy of 97.53% ± 0.27%, precision of 97.82%, recall of 96.94%, PPV of 0.98, NPV of 0.97, MCC of 0.95, and a mean AUC of 0.9897, which confirms that our proposed model exhibits high robust performance. For hardware realization, the optimized student CNN was deployed on a Xilinx Zynq UltraScale + MPSoC FPGA using the Vitis High-Level Synthesis (HLS) toolchain. The implemented accelerator achieved a kernel-level inference latency of 10.6 µs and a Peak On-Chip Kernel Throughput of approximately 94,339 inferences/s, while maintaining balanced hardware resource utilization and low power consumption. Overall, the proposed framework demonstrates accurate classification of plaque-associated angiographic findings with efficient FPGA deployment.
BackgroundWomen with HIV have increased risk of cardiovascular disease compared to women without HIV. Adopting healthy lifestyle habits in early adulthood, including adequate sleep and physical activity, can mitigate this risk. Fitness watches provide objective measures of sleep and physical activity, and may enable proactive interventions to optimize long-term cardiovascular health. However, the acceptability and feasibility of wearable devices for continuous monitoring of these two measures among women living with HIV has not been adequately explored.MethodsA pilot acceptability and feasibility study was conducted with thirty women with HIV participating in a US-based prospective observational study. Participants were assigned to one of three fitness watches, two requiring data transfer every 3–5 days, and one without a transfer requirement. Acceptability of using the fitness watch and feasibility of collecting sleep and physical activity data were assessed.ResultsAmong the 30 women, median participant age was 36.0 years (Interquartile Range: 30.1, 40.3) with 8 (27%) having > a high school education, and 19 (63%) never having owned a fitness watch. All participants agreed or strongly agreed that the device was easy to use, regardless of device type. Over 80% found it acceptable to collect and share sleep and physical activity data. Only three (10%) participants reported data transfer to be somewhat difficult. Sleep and physical activity data collection was feasible for 90% of devices requiring interim data transfer and 100% of devices without a transfer requirement.ConclusionsAll three fitness watches demonstrated high acceptability and feasibility for continuous data collection among women living with HIV, including participants without prior experience using wearable devices, with minimal concerns regarding privacy or confidentiality. Participants expressed a preference for devices that provided real-time access to personal data, suggesting that user visibility of health metrics may enhance engagement with sleep and physical activity interventions. These findings may inform the design and implementation of wearable device–based interventions to support sleep health and physical activity among women living with HIV of reproductive age.
IntroductionObjective assessment of infant suckling biomechanics remains limited despite decades of pressure-based measurement research, and tongue-tie (ankyloglossia) is still largely evaluated using semi-quantitative anatomical scoring. We developed a hydraulic differential-pressure device — the electronic baby bottle (EBB) — designed to record real-time net intraoral mechanical load during nutritive suckling through a fluid-filled sensing cavity with a controlled air inclusion coupled to a differential pressure transducer. This study aimed to: (1) describe the EBB measurement principle, (2) develop automated machine learning–based classification of effective suckling activity, and (3) evaluate whether quantitative biomechanical parameters derived from automated analysis demonstrate responsiveness one week after frenotomy.MethodsTime-series recordings from iterative prototype testing were segmented into overlapping windows and manually annotated to train a supervised classifier using the ROCKET transform implemented in Python (sktime) with ridge classification and cross-validation. Performance was evaluated on an internal held-out test set and an independent external validation set of full-length recordings excluded from model development. Quantitative parameters were computed from automatically identified effective suckling bursts in 25 infants with restrictive tongue-tie recorded before and one week after frenotomy and compared with 10 control infants with normal tongue mobility.ResultsAutomated classification achieved >99% accuracy on the internal test set and 92% segment-level classification accuracy on external validation recordings. Mean negative intraoral pressure during effective suckling bursts increased in 19/25 (76%) treated infants one week after frenotomy (paired t-test, p = 0.010), shifting toward control values.DiscussionHydraulic differential-pressure recording combined with time-series machine learning enables objective and reproducible quantification of nutritive suckling activity and detects short-term functional change following frenotomy in most infants with restrictive tongue-tie. This framework supports future work on normative reference ranges, standardized digital metrics, and data-driven diagnostic thresholds for infant feeding dysfunction.
IntroductionAlthough remote patient management (RPM) shows promise for improving chronic heart failure (HF) management, its effectiveness depends on technology acceptance, usability, and patient engagement. However, patients are often involved only after systems are implemented, limiting opportunities to align RPM with their needs. This study explored patients' expectations, wishes, and concerns regarding future RPM using participatory design methods to inform more patient-centred RPM.MethodsPatients with HF using RPM were recruited from a Dutch hospital for individual, semi-structured interviews. We employed a co-constructing stories approach, inviting participants to imagine future scenarios of remote care. Guided by narrative prompts about potential developments in RPM, participants were encouraged to discuss how such technologies might shape their experiences. During the interviews, one researcher concurrently created sketches of the participants' answers, creating an evolving trace of the dialogue. The visual data were subsequently analysed using Annotated Visual Analysis to identify themes.ResultsSix patients were interviewed, producing 7 h of audio and 18 A3 pages of sketch notes. Four key themes emerged: Data Feedback, reflecting patients' desire to receive clear and constructive feedback that supports self-management beyond alert-based monitoring, to be able to take more responsibility for their health through RPM; Invasiveness, including camera-based monitoring or collection of non-medical data, was perceived as a barrier to adoption, although concerns were reduced when monitoring was clearly explained and considered proportional to disease severity; Adaptability and Integration highlighted the need for RPM to extend support beyond monitoring and to remain useful in daily life; Healthcare Provider contact emphasised that human interaction is irreplaceable, with empathy from healthcare professionals fostering trust in the system and patients' confidence in managing their health.ConclusionInviting patients to reflect on remote care provided valuable insights into their needs and expectations for future RPM. Participatory design methodologies facilitated rich discussions about patients' experiences, needs, and expectations regarding future RPM. Patients emphasised RPM should be actionable, minimally invasive, adaptable to daily life, and complemented by human contact. This highlights the importance of socio-technical systems that empower self-management while maintaining professional support. These findings can guide more human-centred RPM designs for HF care.
BackgroundCervical cancer remains a leading cause of cancer death among women in sub-Saharan Africa, with Tanzania bearing a disproportionate burden. The critical shortage of trained pathologists, coupled with the unprecedented disease burden in low-resource settings, underscores the urgent need for point-of-care screening and diagnosis to enable timely decision-making. We assessed the diagnostic accuracy of AI-driven cytopathological tools to improve diagnostic efficiency and accessibility.MethodsThis retrospective secondary data analysis evaluated five convolutional neural network (CNN) architectures: EfficientNetB7, MobileNet, ResNet50, ResNet152, and InceptionNetV3, for semi-automated classification of cervical cell abnormalities. A total of 11,955 Pap smear cytological images were used from the publicly available Center for Recognition and Inspection of Cells (CRIC) Searchable Image Database, spanning six cellular classes: Normal, ASC-US, LSIL, ASC-H, HSIL, and carcinoma. Performance was assessed on a hold-out test set (n = 961) using macro-averaged metrics.ResultsEfficientNetB7 achieved the highest overall performance, with a macro F1 score of 0.9324 (95% CI: 0.920–0.945), an accuracy of 0.9775 (95% CI: 0.968–0.987), and a sensitivity of 0.9324 (95% CI: 0.920–0.945). ResNet50 ranked second (F1 score: 0.9282 [95% CI: 0.916–0.941]) and ResNet152 third (F1 score: 0.9240 [95% CI: 0.911–0.937]), showing minimal performance gaps. InceptionNetV3 followed closely (F1 score: 0.9220 [95% CI: 0.909–0.935]). MobileNet achieved the lowest F1 score (0.8918 [95% CI: 0.875–0.908]), but its lightweight architecture is suitable for edge deployment. The Carcinoma (CA) class achieved near-perfect recall across all models (> = 0.978). Notable interclass confusion was observed between ASC-US and LSIL, attributable to cytomorphological overlap.ConclusionEfficientNetB7 offers promising diagnostic accuracy for automated cervical cancer screening using Pap smear images and shows potential for integration into point-of-care workflows in low-resource settings. Future work should focus on training models on locally representative datasets and exploring whole smear analysis to reduce interclass misclassification.
BackgroundThe role of artificial intelligence (AI) in supporting clinical decision-making across the perioperative continuum remains incompletely defined. Although the presence of many AI models that perform well in terms of their predictive performance has been established, their role in the actual surgical decision-making process in the real-world in terms of specialties and perioperative phases has not been fully mapped.ObjectiveTo map the existing literature on the application of AI in surgical decision-making across the perioperative continuum, describe its use in preoperative, intraoperative, and postoperative phases, and identify key barriers and gaps affecting clinical implementation.MethodsA scoping review was conducted in accordance with the PRISMA-ScR guidelines. PubMed, Scopus, and Google Scholar were searched for studies evaluating AI applications in surgical decision-making. Eligible studies included primary research and evidence syntheses applying machine learning, deep learning, radiomics, or computer vision to diagnosis, risk stratification, surgical planning, intraoperative guidance, or postoperative outcome prediction. Study characteristics, perioperative phase, clinical application, AI methodology, and reported implementation barriers were extracted and charted using a standardized data form.ResultsFifty- five sources of evidence were included. Sources addressing multiple or cross-phase perioperative applications constituted the largest category, while among phase-specific applications, preoperative applications were the most frequent and primarily focused on diagnosis, risk stratification, and surgical planning. Intraoperative applications were less common and were limited by data availability, workflow integration, and real-time implementation challenges. Postoperative applications mainly addressed complication prediction, survival estimation, and recovery monitoring.ConclusionAI in surgical decision-making is expanding rapidly, with preoperative applications showing comparatively greater evidence maturity than intraoperative and postoperative applications. However, prospective validation and real-world implementation remain limited across the perioperative continuum.
Digital health innovations hold enormous potential to expand access to care. Yet individuals with disabilities often face inequities due to systemic, environmental, and design-related barriers that limit full participation in digital health ecosystems. These challenges often arise from the ways digital tools and systems are designed rather than the impairment itself. This perspective discusses the intersection of disability, digital health, and equitable access through key disability frameworks including the International Classification of Functioning (ICF), the Health Equity Framework for People with Disabilities, and other prominent medical and biopsychosocial models. We assert that equitable digital health for people with disabilities requires a shift toward accessibility-as-standard, grounded in Universal Design, participatory co-design, and intersectionality-informed approaches. We offer applied guidance on integrating Universal Design and Web Content Accessibility Guidelines (WCAG) principles from the earliest design stages; incorporating assistive technology compatibility; developing accessibility informed user profiles; establishing digital literacy and support pathways; adapting telehealth workflows for diverse needs; and building trust through culturally aligned content. We also outline strategies for equitable design such as embedding disability representation in co-design processes and developing measurement plans that capture accessibility, engagement, and long-term outcomes for people with disabilities. Achieving equitable reach and impact in digital health requires centering disability as a core design and implementation domain. By translating conceptual disability models into concrete design practices, and workflow adaptations, clinicians, researchers, and technology developers can create digital health environments that are more usable, inclusive, and sustainable improving access, experience, and health outcomes for disabled individuals across diverse settings.
IntroductionVirtual hospital systems support remote triage and diagnosis across distributed healthcare environments by integrating radiographic imaging with structured clinical information for clinical decision making. However, many existing diagnostic frameworks do not fully address real world challenges commonly encountered in scalable virtual healthcare settings. These challenges include severe class imbalance, incomplete clinical attributes, heterogeneous multimodal inputs, and variability in data quality across institutions.MethodsThis study introduces VHealth-MFusion, a fusion-based hierarchical multimodal deep learning framework that integrates chest X-ray (CXR) imaging and structured clinical data within a unified convolutional neural network (CNN) and multilayer perceptron (MLP) architecture. Using pneumonia Tele-Diagnosis as a use case, the proposed framework combines radiographic feature extraction through a CNN branch with structured clinical representation learning through an MLP branch, followed by late multimodal fusion for multiclass respiratory disease classification. The framework was evaluated using a hierarchical multimodal dataset containing CXR images across multiple diagnostic categories, including Normal, Bacterial, Adenovirus, Influenza, MERS-CoV, and SARS-CoV-2, together with 632 structured clinical variables comprising demographic information, vital signs, symptoms, and laboratory findings.ResultsUnder a controlled dataset-consistent comparison against an established hierarchical multimodal CNN baseline, VHealth-MFusion achieved 97.2% classification accuracy, outperforming the baseline accuracy of 95.9%.DiscussionThe findings demonstrate that multimodal integration of radiographic imaging and structured clinical information can improve diagnostic robustness and support more reliable Tele-Diagnosis workflows within virtual hospital environments. Overall, the proposed framework contributes a practically oriented multimodal diagnostic architecture designed to facilitate scalable remote clinical services and informed decision making within digitally enabled care systems where imaging and structured clinical information are jointly available.
Digital health interventions are increasingly used to support cardiovascular prevention, but their clinical maturity differs across risk factors and types of technology. This narrative review summarizes current evidence and implementation challenges related to digital therapeutics, mobile health applications, telemonitoring systems, and integrated digital health platforms in hypertension, dyslipidemia, and broader cardiometabolic prevention. The strongest evidence currently supports digitally enabled blood pressure management, particularly when home blood pressure monitoring, telemonitoring, behavioral support, medication adherence tools, and clinician-guided treatment adjustment are integrated into routine care. In contrast, digital interventions for dyslipidemia remain less established and are mainly focused on education, lifestyle modification, adherence support, shared decision-making, and long-term risk reduction. Therefore, most lipid-related digital tools should currently be interpreted as supportive mHealth or prevention interventions rather than fully established digital therapeutics. Successful implementation requires more than patient-facing applications. It depends on clinical workflow integration, professional responsibility for data review and telemonitoring alerts, regulatory and reimbursement pathways, data protection, interoperability, equity, and evidence of clinical and economic value. Future studies should distinguish between digital therapeutics, mHealth tools, telemonitoring systems, and digital ecosystems, and should evaluate clinically meaningful endpoints, safety, cost-effectiveness, and long-term sustainability.
Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.
BackgroundDuring the COVID-19 pandemic, the use of mental health chatbots received potential attention as a means of addressing the increase in demand for mental healthcare, which has become inaccessible due to a shortage of healthcare workers and lockdown. Even after the pandemic, chatbot versatility holds the potential to address the ongoing limited accessibility of mental healthcare due to a shortage of professionals, geographical challenges, cost, or stigma.ObjectiveThis review examines the scope and nature of the systematic review and meta-analysis evidence on chatbot interventions for mental health.MethodThis scoping review follows Arksey and O'Malley's five-step guide and was reported using the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for scoping reviews (PRISMA-ScR). The following databases were used to conduct all systematic literature searches: PubMed, Scopus, Medline, and Web of Science. Inclusion criteria include systematic reviews and meta-analyses, written in English, that focus on the use of chatbots as an intervention for mental conditions, with all primary articles being relevant. A total of 14 articles were relevant and included in the data extraction and analysis.ResultsIt was found that although current review-level evidence suggests potential short-term benefits, there was a discrepancy regarding the specific conditions under which chatbots are effective. Moreover, the use of chatbots for symptom improvement is still inconclusive, since there are issues with important aspects such as unclear risk of bias, lack of longitudinal studies, and lack of comparison with traditional therapy.ConclusionThere is potential for chatbot implementation as an intervention for mental conditions. However, this is still at its early stages of development and requires more research to improve communication quality and ensure safety, which is paramount when communicating with vulnerable patients. The findings of the review provide a useful foundation for informing clinical practice, policy, and future research.
Medicine is a predominantly physical profession, yet most medical artificial intelligence (AI) remains screen bound. Physical artificial intelligence (PAI) extends AI's capabilities and can directly address many of the healthcare system challenges and unmet needs such as workforce shortages, unsustainable healthcare costs, heightened expectations, aging population and need for pandemic preparedness. PAI relies on systems that autonomously perceive, decide, and actuate in real-time by synthesizing diverse environmental data from multiple sensory sources. These data undergo rapid processing and interpretation, facilitating immediate decision-making and responsive physical actions, including movement, object manipulation, and direct human interaction. This paper first identifies critical healthcare system needs that robotic agents can effectively address. It further examines core actions of PAI systems, emphasizing perceptual pathways, clinical decision-making processes, and human-robot interactions that translate sensory inputs into tailored, patient-specific responses. We explore essential data utilization aspects, including edge-device advances, PAI datasets, digital twins, and edge-to-cloud infrastructures that support real-time inference and are crucial for reducing barriers to PAI implementation. Finally, we analyze the cultural factors accelerating the adoption of PAI in healthcare and review PAI advancements in other sectors. We argue that the convergence of pressing healthcare demands, technological advancements, and cultural readiness signals a tipping point for PAI in medicine.
IntroductionChatbots are accessible and cost-effective tools that may be able to provide navigation support to people who lack access to conventional patient navigation services. To date, little is known about people's interest in using chatbots for patient navigation. It is also unclear what specific types of support people would like to see incorporated into patient navigation chatbots. The purpose of this study was to gain insight into these topics.MethodsParticipants were recruited through the Prolific research platform. They completed a cross-sectional survey that captured information on respondent characteristics, general interest in using a patient navigation chatbot, and interest in various navigation support functions. Data were analyzed using descriptive statistics and ANOVA and hierarchical regression statistical tests.ResultsA total of 276 participants (141 women; M age = 38.99 years, SD = 12.90) were included in the analysis. In general, participants were moderately interested in using a patient navigation chatbot. They had greater interest in functions related to care coordination, linkage to services and resources, and general education than functions related to needs assessment and treatment support. Interest varied by certain respondent characteristics, particularly demographic and personality factors.DiscussionThe findings from this study will help inform the design and marketing of patient navigation chatbots. The effective implementation of this technology may fill some of the accessibility gaps that characterize conventional patient navigation programs.