
Background: Generative Pre-trained Transformer (ChatGPT) is a generative artificial intelligence (AI) model developed by Open AI (San Francisco, CA, USA), which generates responses based on input received and can solve problems and complete tasks by using reinforcement techniques and machine learning from sources online. Despite the growing use of AI in healthcare, its application in vascular surgery remains limited. This study aimed to evaluate ChatGPT's ability to generate patient information related to digital subtraction angiography (DSA). Fifteen commonly asked patient questions regarding DSA were identified, and ChatGPT-3.5 was used to produce responses. Additionally, the model was tasked with generating a complete patient information leaflet for DSA. The outputs were assessed for readability, informational quality, and appropriateness. Methods: Fifteen questions were entered into ChatGPT-3.5, which also generated a complete patient information leaflet for DSA. The readability of the outputs was evaluated using two standardized scoring systems: the Flesch-Kincaid Reading Ease Score (FRES) and the Gunning Fog Index (GFI). The quality of the responses was assessed using the DISCERN tool, and their appropriateness was rated on a Likert scale. Results: The readability analysis using the FRES yielded an average score of 31.21 (range, 16.29-53.57), corresponding to a college-level reading ability. The mean Flesch-Kincaid Grade Level (FKGL) was 13.54 (range, 9.37-15.30). The patient information leaflet generated by ChatGPT-3.5 scored 41.30 on the Reading Ease scale, also indicating a college-level reading age. Using the GFI, the average score for the responses was 15.91 (range, 12.24-20.84), equivalent to the reading level of a college junior or senior, while the patient information leaflet scored 14.43. These findings suggest that the content is written at a significantly higher level than the recommended 6th-8th grade reading level for patient education materials. The quality of the responses, assessed using the DISCERN tool, averaged 41, which is considered "fair", whereas the patient information leaflet scored 34, classified as "poor". Despite this, the majority of the content was rated as factually accurate and generally appropriate for the clinical context. Conclusions: Using ChatGPT-3.5 to generate responses to patient questions and create a patient information leaflet resulted in content written at a significantly higher reading level than recommended for patient education. While individual responses demonstrated generally good appropriateness and achieved fair DISCERN scores, the overall quality and readability decreased when generating a complete leaflet. These findings suggest that ChatGPT-3.5 performs better with discrete questions than with comprehensive tasks. Care and caution should be exercised when considering the use of such tools for patient-facing materials.
Background: Atopic dermatitis (AD) significantly affects quality of life (QoL) with cutaneous manifestations and associated symptoms. Patients utilize online forums to discuss treatment options for many diseases, including AD. Unlike questionnaire-based evaluations, anonymous online discussions encourage extensive dialogue. The unstructured online narrative posts may serve as a new avenue for analysis of the overall sentiment of a treatment. This observational study aims to assess perception of different treatments for AD by developing a novel sentiment positivity index (SPI) of treatment-specific comments made to a popular online AD forum using natural language processing (NLP). Methods: All posts were extracted from a well-known online forum dedicated to AD made in a 5-year period and mentions of specific AD treatments were identified in 28,159 posts. These posts were analyzed with a pretrained language model to determine the overall sentiment, whether positive or negative. A SPI was developed, calculated as the ratio of posts expressing positive sentiment towards a particular treatment to the total number of posts. Treatment perceived more positively by participants have a higher SPI, and treatments perceived less positively have a lower SPI. Results: For AD treatments, there were 12,439 unique post authors with an average post length of 92 words. The SPI for all AD treatments varied from 0.116 to 0.365, with a mean of 0.221. The treatments with the highest SPI were Janus kinase (JAK) inhibitors [upadacitinib 0.365, 95% confidence interval (CI): 0.264-0.479; topical ruxolitinib 0.324, 95% CI: 0.243-0.417; and baricitinib 0.290, 95% CI: 0.161-0.466]. The treatments with the lowest SPI were topical corticosteroids (TCS) (Class V corticosteroids 0.116, 95% CI: 0.066-0.196; Class VI corticosteroids 0.128, 95% CI: 0.087-0.185). Conclusions: For all AD therapies, the mean SPI was 0.221, consistent with lower positivity regardless of therapy. A spectrum of SPI scores was demonstrated for the treatments studied, with TCS having the lowest positivity scores. Developing the SPI using natural-language processing of online forum commentary has the potential to give insight to the patient perspective and may be applicable to other disease states. Further validation and exploration of clinical relevance are warranted.