Background: In drug development clinical trials, there is a need for balance between restricting variables by setting eligibility criteria and representing the broader patient population that may use a product once it is approved. Similarly, although recent policy initiatives focusing on the inclusion of historically underrepresented groups are being implemented, barriers still remain. These limitations of clinical trials may mask potential product benefits and side effects. To bridge these gaps, online communication in health communities may serve as an additional population signal for drug side effects. Objective: The aim of this study was to employ a nontraditional dataset to identify drug side-effect signals. The study was designed to apply both natural language processing (NLP) technology and hands-on linguistic analysis to a set of online posts from known statin users to (1) identify any underlying crossover between the use of statins and impairment of memory or cognition and (2) obtain patient lexicon in their descriptions of experiences with statin medications and memory changes. Methods: Researchers utilized user-generated content on Inspire, looking at over 11 million posts across Inspire. Posts were written by patients and caregivers belonging to a variety of communities on Inspire. After identifying these posts, researchers used NLP and hands-on linguistic analysis to draw and expand upon correlations among statin use, memory, and cognition. Results: NLP analysis of posts identified statistical correlations between statin users and the discussion of memory impairment, which were not observed in control groups. NLP found that, out of all members on Inspire, 3.1% had posted about memory or cognition. In a control group of those who had posted about TNF inhibitors, 6.2% had also posted about memory and cognition. In comparison, of all those who had posted about a statin medication, 22.6% (P<.001) also posted about memory and cognition. Furthermore, linguistic analysis of a sample of posts provided themes and context to these statistical findings. By looking at posts from statin users about memory, four key themes were found and described in detail in the data: memory loss, aphasia, cognitive impairment, and emotional change. Conclusions: Correlations from this study point to a need for further research on the impact of statins on memory and cognition. Furthermore, when using nontraditional datasets, such as online communities, NLP and linguistic methodologies broaden the population for identifying side-effect signals. For side effects such as those on memory and cognition, where self-reporting may be unreliable, these methods can provide another avenue to inform patients, providers, and the Food and Drug Administration.
Background: Adverse drug reactions (ADRs) occur in nearly all patients on chemotherapy, causing morbidity and therapy disruptions. Detection of such ADRs is limited in clinical trials, which are underpowered to detect rare events. Early recognition of ADRs in the postmarketing phase could substantially reduce morbidity and decrease societal costs. Internet community health forums provide a mechanism for individuals to discuss real-time health concerns and can enable computational detection of ADRs. Objective: The goal of this study is to identify cutaneous ADR signals in social health networks and compare the frequency and timing of these ADRs to clinical reports in the literature. Methods: We present a natural language processing-based, ADR signal-generation pipeline based on patient posts on Internet social health networks. We identified user posts from the Inspire health forums related to two chemotherapy classes: erlotinib, an epidermal growth factor receptor inhibitor, and nivolumab and pembrolizumab, immune checkpoint inhibitors. We extracted mentions of ADRs from unstructured content of patient posts. We then performed population-level association analyses and time-to-detection analyses. Results: Our system detected cutaneous ADRs from patient reports with high precision (0.90) and at frequencies comparable to those documented in the literature but an average of 7 months ahead of their literature reporting. Known ADRs were associated with higher proportional reporting ratios compared to negative controls, demonstrating the robustness of our analyses. Our named entity recognition system achieved a 0.738 microaveraged F-measure in detecting ADR entities, not limited to cutaneous ADRs, in health forum posts. Additionally, we discovered the novel ADR of hypohidrosis reported by 23 patients in erlotinib-related posts; this ADR was absent from 15 years of literature on this medication and we recently reported the finding in a clinical oncology journal. Conclusions: Several hundred million patients report health concerns in social health networks, yet this information is markedly underutilized for pharmacosurveillance. We demonstrated the ability of a natural language processing-based signal-generation pipeline to accurately detect patient reports of ADRs months in advance of literature reporting and the robustness of statistical analyses to validate system detections. Our findings suggest the important contributions that social health network data can play in contributing to more comprehensive and timely pharmacovigilance.
Getting a new camera, or any other gadget, can be exciting. I found that learning how to use my camera was similar to learning how to gain the most from the Bible.
This study reports proof-of-principle early detection of chemotherapeutic-associated skin adverse drug reactions from social health networks using a deep learning–based signal generation pipeline to capture how patients describe cutaneous eruptions.