The Weill Cornell Heart to Heart Community Outreach Campaign (H2H) is a free outreach program that provides mobile health screenings. The program brings medical and nursing faculty and students to the underserved, uninsured communities of New York City. Participants are screened for diabetes and heart disease risk factors through onsite exams, including point of care blood tests. If an abnormality is found, they receive a medical consultation to offer personalized advice and referrals to free/low-cost clinics when needed. The goal is to help underserved individuals understand their cardiometabolic health and to promote early intervention. This article describes the development of the program, including factors that were essential to the collaboration, challenges faced, barriers to implementation, and its evolution throughout the first 12 years. The program has benefited from strong foundational program leadership, effective inter-institutional collaboration, and maintaining community trust.
In underserved communities across New York City, uninsured adults encounter a greater risk of cardiovascular disease (CVD) and diabetes. The Heart-to-Heart Community Outreach Program (H2H) addresses these disparities by screening for CVD risk factors, identifying healthcare access barriers, and fostering community engagement in translational research at the Weill Cornell Medicine Clinical and Translational Science Award (CTSA) hub. Screening events are hosted in partnership with faith-based institutions. Participants provide a medical history, complete a survey, and receive counseling by clinicians with referrals for follow-up care. This study aims to quantify H2H screening participant health status; identify socioeconomic, health access, and health-related barriers disproportionately promoting the onset of CVD and diabetes; and develop long-term community partnerships to enable underserved communities to influence activities across the translational research spectrum at our CTSA hub. The population served is disproportionately non-white, and uninsured, with many low-income and underserved individuals. The program was developed in partnership with our Community Advisory Board to empower this cohort to make beneficial lifestyle changes. Leveraging partnerships with faith-based institutions and community centers in at-risk New York City neighborhoods, H2H addresses the increasing burden of diabetes and CVD risk factors in vulnerable individuals while promoting community involvement in CTSA activities, serving as a model for similar initiatives.
Background While scientific knowledge of post–COVID-19 condition (PCC) is growing, there remains significant uncertainty in the definition of the disease, its expected clinical course, and its impact on daily functioning. Social media platforms can generate valuable insights into patient-reported health outcomes as the content is produced at high resolution by patients and caregivers, representing experiences that may be unavailable to most clinicians. Objective In this study, we aimed to determine the validity and effectiveness of advanced natural language processing approaches built to derive insight into PCC-related patient-reported health outcomes from social media platforms Twitter and Reddit. We extracted PCC-related terms, including symptoms and conditions, and measured their occurrence frequency. We compared the outputs with human annotations and clinical outcomes and tracked symptom and condition term occurrences over time and locations to explore the pipeline’s potential as a surveillance tool. Methods We used bidirectional encoder representations from transformers (BERT) models to extract and normalize PCC symptom and condition terms from English posts on Twitter and Reddit. We compared 2 named entity recognition models and implemented a 2-step normalization task to map extracted terms to unique concepts in standardized terminology. The normalization steps were done using a semantic search approach with BERT biencoders. We evaluated the effectiveness of BERT models in extracting the terms using a human-annotated corpus and a proximity-based score. We also compared the validity and reliability of the extracted and normalized terms to a web-based survey with more than 3000 participants from several countries. Results UmlsBERT-Clinical had the highest accuracy in predicting entities closest to those extracted by human annotators. Based on our findings, the top 3 most commonly occurring groups of PCC symptom and condition terms were systemic (such as fatigue), neuropsychiatric (such as anxiety and brain fog), and respiratory (such as shortness of breath). In addition, we also found novel symptom and condition terms that had not been categorized in previous studies, such as infection and pain. Regarding the co-occurring symptoms, the pair of fatigue and headaches was among the most co-occurring term pairs across both platforms. Based on the temporal analysis, the neuropsychiatric terms were the most prevalent, followed by the systemic category, on both social media platforms. Our spatial analysis concluded that 42% (10,938/26,247) of the analyzed terms included location information, with the majority coming from the United States, United Kingdom, and Canada. Conclusions The outcome of our social media–derived pipeline is comparable with the results of peer-reviewed articles relevant to PCC symptoms. Overall, this study provides unique insights into patient-reported health outcomes of PCC and valuable information about the patient’s journey that can help health care providers anticipate future needs. International Registered Report Identifier (IRRID) RR2-10.1101/2022.12.14.22283419
Background There remains significant uncertainty in the definition of the long COVID disease, its expected clinical course, and its impact on daily functioning. Social media platforms can generate valuable insights into patient-reported health outcomes as the content is produced at high resolution by patients and caregivers, representing experiences that may be unavailable to most clinicians. Objective We aim to determine the validity and effectiveness of advanced NLP approaches built to derive insight into Long COVID-related patient-reported health outcomes from social media platforms. Methodology We use Transformer-based BERT models to extract and normalize long COVID Symptoms and Conditions (SyCo) from English posts on Twitter and Reddit. Furthermore, we estimate the occurrence and co-occurrence of SyCo terms at any point or across time and locations. Finally, we compare the extracted health outcomes with human annotations and highly utilized clinical outcomes grounded in the medical literature. Result Based on our findings, the top three most commonly occurring groups of long COVID symptoms are systemic (such as “fatigue”), neuropsychiatric (such as “anxiety” and “brain fog”), and respiratory (such as “shortness of breath”). Regarding the co-occurring symptoms, the pair of ‘fatigue & headaches’ is most common. In addition, we show that other conditions, such as infection, hair loss, and weight loss, as well as mentions of other diseases, such as flu, cancer, or Lyme disease, are among the top reported terms by social media users. Conclusion The outcome of our social media-derived pipeline is comparable with the outcomes of peer-reviewed articles relevant to long COVID symptoms. Overall, this study provides unique insights into patient-reported health outcomes from long COVID and valuable information about the patient’s journey that can help healthcare providers anticipate future needs.
From the outset of the COVID-19 pandemic, social media has provided a platform for sharing and discussing experiences in real time. This rich source of information may also prove useful to researchers for uncovering evolving insights into post-acute sequelae of SARS-CoV-2 (PACS), commonly referred to as Long COVID. In order to leverage social media data, we propose using entity-extraction methods for providing clinical insights prior to defining subsequent downstream tasks. In this work, we address the gap between state-of-the-art entity recognition models and the extraction of clinically relevant entities which may be useful to provide explanations for gaining relevant insights from Twitter data. We then propose an approach to bridge the gap by utilizing existing configurable tools, and datasets to enhance the capabilities of these models. Code for this work is available at: https://github.com/VectorInstitute/ProjectLongCovid-NER .
Academic institutions need to maintain publication lists for thousands of faculty and other scholars. Automated tools are essential to minimize the need for direct feedback from the scholars themselves who are practically unable to commit necessary effort to keep the data accurate. In relying exclusively on clustering techniques, author disambiguation applications fail to satisfy key use cases of academic institutions. Algorithms can perfectly group together a set of publications authored by a common individual, but, for them to be useful to an academic institution, they need to programmatically and recurrently map articles to thousands of scholars of interest en masse. Consistent with a savvy librarian’s approach for generating a scholar’s list of publications, identity-driven authorship prediction is the process of using information about a scholar to quantify the likelihood that person wrote certain articles. ReCiter is an application that attempts to do exactly that. ReCiter uses institutionally-maintained identity data such as name of department and year of terminal degree to predict which articles a given scholar has authored. To compute the overall score for a given candidate article from PubMed (and, optionally, Scopus), ReCiter uses: up to 12 types of commonly available, identity data; whether other members of a cluster have been accepted or rejected by a user; and the average score of a cluster. In addition, ReCiter provides scoring and qualitative evidence supporting why particular articles are suggested. This context and confidence scoring allows curators to more accurately provide feedback on behalf of scholars. To help users to more efficiently curate publication lists, we used a support vector machine analysis to optimize the scoring of the ReCiter algorithm. In our analysis of a diverse test group of 500 scholars at an academic private medical center, ReCiter correctly predicted 98% of their publications in PubMed.
ABSTRACT IMPACT: Leveraging partnerships with faith-based institutions and community centers in at-risk NYC neighborhoods, the H2H Program breaks down barriers to engaging with the medical establishment and addresses the increasing burden of diabetes and CVD risk factors in the most vulnerable individuals. OBJECTIVES/GOALS: Screening for modifiable risk factors is critical for cardiovascular disease (CVD) risk reduction. Low-income, urban communities often encounter barriers to care. Community-academic outreach partnerships are vital in addressing such disparities and promoting health equity and culturally targeted interventions among high-risk populations. METHODS/STUDY POPULATION: In 2010, the Weill Cornell Clinical and Translational Science Center along with Weill Cornell Medicine (WCM) and Hunter-Bellevue School of Nursing (HBSON) launched Heart to Heart (H2H), a community outreach program partnering with faith-based centers to offer free health screenings and education to some of New York City’s (NYC) most vulnerable communities. Participants work with undergraduate, nursing, medical and dietician students to complete a demographics and health questionnaire followed by vital signs and point-of-care blood testing. Participants then receive personalized health education, nutrition and lifestyle counseling by student volunteers, precepted by WCM Primary Care and HBSON faculty. Participants are provided information on local free or low-cost clinics as necessary for follow-up. RESULTS/ANTICIPATED RESULTS: To date H2H held 125 events and 5,952 screenings. Mean age of the participants was 54.3 (SD 39.6) and 3,682 (63.1%) were female. 74.2% identified as non-white. 42.1% were uninsured. 32.3% reported annual income of less than $20k. 18.3% of participants reported not having seen a doctor in the past year. 40.7% reported preexisting hypertension, of which 74.5% were on medication and 78% with sub-optimal control. 15.7% had been previously diagnosed with diabetes, of which 75.8% were on medication and 41.4% with sub-optimal control (HbA1c <7). 37.7% had been diagnosed with dyslipidemia previously, of which 47.4% were on medication and 62.1% with sub-optimal control. Screenings revealed, 56.9% had undiagnosed hypertensive blood pressures, 4.7% had an elevated HbA1c >6.5, and 49.2% had dyslipidemia. DISCUSSION/SIGNIFICANCE OF FINDINGS: H2H screening revealed significant cardiovascular health disparities, many of which were poorly controlled or newly discovered. Cross-institutional academic partnerships can empower communities with knowledge of their health status and help facilitate access to medical care to further address health risk factors.
46,XY individuals with 5α-reductase-2 deficiency syndrome have ambiguous genitalia with a clitoral-like phallus, severely bifid scrotum, and pseudovaginal perineoscrotal hypospadias. As a result of this genital ambiguity, many are thought to be girls at birth and are treated as such from the time they are born. However, at puberty virilization occurs, whereupon affected individuals, with few exceptions, reveal their male gender identity and assume a male gender role. Much of the research on gender identity of individuals with 5α-reductase-2 deficiency was derived from a study of 38 males in three rural villages in the Dominican Republic. This chapter describes the clinical, biochemical, and genetic features of the syndrome; considers the impact of androgens on gender identity and gender role; and outlines treatment considerations.
We applied social network analysis (SNA) to Tweets mentioning cannabis or opioid-related terms to publicly available COVID-19 related Tweets collected from Jan 21st to May 3rd, 2020 (n= 2,558,474 Tweets). We randomly extracted 16,154 Tweets mentioning cannabis and 4,670 Tweets mentioning opioids from the COVID-19 Tweet corpora for our analysis. The cannabis related Tweets created by 6,144 users were disseminated to 280,042,783 users and retweeted 11 times the number of original messages while opioid-related Tweets created by 3,412 users were disseminated to smaller number of users. The opioids Twitter network showed more cohesive online group activities and a cleaner online environment with less disinformation. The cannabis Twitter network showed a less desirable online environment with more disinformation (false information to mislead the public) and stakeholders lacking strong science knowledge. Application of SNA to Tweets provides insights for future online-based drug abuse research during the outbreak.
We randomly extracted publicly available Tweets mentioning COVID-19 related terms (n=2,558,474 Tweets) from Tweet corpora collected daily using an API from Jan 21st to May 3rd, 2020. We applied a clustering algorithm to publicly available Tweets authored by African Americans (n=1,763) to detect topics and sentiment applying natural language processing (NLP). We visualized fifteen topics (four themes) using network diagrams (Newman modularity 0.74). Compared to the COVID-19 related Tweets authored by others, positive sentiments, cohesively encouraging online discussions (e.g., Black strong 27.1%, growing up Blacks 22.8%, support Black business 17.0%, how to build resilience 7.8%), and COVID-19 prevention behaviors (e.g., masks 4.7%, encouraging social distancing 9.4%) were uniquely observed in African American Twitter communities. Application of topic modeling techniques to streaming social media Twitter provides the foundation for research team insights regarding information and future virtual based intervention and social media based health disparity research for COVID-19.
We applied artificial intelligence techniques to build correlate models that predict general poor health in a national sample of caregivers with mild cognitive impairment (MCI). Our application of deep learning identified age, duration of caregiving, amount of alcohol intake, weight, myocardial infarction (MI) and frequency of MCI symptoms for Blacks and Hispanics whereas frequency of MCI symptoms, income, weight, coronary heart disease (CHD), age, and use of e-cigarette for the others as the strongest correlates of poor health among 81 variables entered. The application of artificial intelligence efficiently provided intervention strategies for Black and Hispanic caregivers with MCI.
Despite the demonstrated value of visualization-based modalities for measuring and mapping science, it remains common practice to search and explore the literature via databases that present lists of articles with little, if any, supplementary visual information. Identifying the desired item in a list is a familiar information retrieval paradigm with a low cognitive load. However, given the rapid emergence of the field of visual text analytics, it is time to challenge the notion that article lists should remain the dominant method to search and organize the scientific literature. One reason that visualization methods are applied relatively rarely in information retrieval may be that it is difficult to develop useful and user-friendly science mapping systems. This article summarizes key workflows for bibliometric mapping, a technique for visually representing information from scientific publications, including citation data, bibliographic metadata, and article content. It describes methods and challenges in extracting, processing, and normalizing data, reducing dimensionality, modeling topics, assigning labels, and visualizing data. It also describes software tools available to support bibliometric analysis and science mapping workflows, outlines methods from other domains that have not been widely applied in bibliometric mapping, and considers opportunities for next generation bibliometric analysis and mapping software systems.
Objective The paper provides a review of current practices related to evaluation support services reported by seven biomedical and research libraries. Methods A group of seven libraries from the United States and Canada described their experiences with establishing evaluation support services at their libraries. A questionnaire was distributed among the libraries to elicit information as to program development, service and staffing models, campus partnerships, training, products such as tools and reports, and resources used for evaluation support services. The libraries also reported interesting projects, lessons learned, and future plans. Results The seven libraries profiled in this paper report a variety of service models in providing evaluation support services to meet the needs of campus stakeholders. The service models range from research center cores, partnerships with research groups, and library programs with staff dedicated to evaluation support services. A variety of products and services were described such as an automated tool to develop rank-based metrics, consultation on appropriate metrics to use for evaluation, customized publication and citation reports, resource guides, classes and training, and others. Implementing these services has allowed the libraries to expand their roles on campus and to contribute more directly to the research missions of their institutions. Conclusions Libraries can leverage a variety of evaluation support services as an opportunity to successfully meet an array of challenges confronting the biomedical research community, including robust efforts to report and demonstrate tangible and meaningful outcomes of biomedical research and clinical care. These services represent a transformative direction that can be emulated by other biomedical and research libraries.
OBJECTIVE To collaborate with community members to develop tailored infographics that support comprehension of health information, engage the viewer, and may have the potential to motivate health-promoting behaviors. METHODS The authors conducted participatory design sessions with community members, who were purposively sampled and grouped by preferred language (English, Spanish), age group (18-30, 31-60, >60 years), and level of health literacy (adequate, marginal, inadequate). Research staff elicited perceived meaning of each infographic, preferences between infographics, suggestions for improvement, and whether or not the infographics would motivate health-promoting behavior. Analysis and infographic refinement were iterative and concurrent with data collection. RESULTS Successful designs were information-rich, supported comparison, provided context, and/or employed familiar color and symbolic analogies. Infographics that employed repeated icons to represent multiple instances of a more general class of things (e.g., apple icons to represent fruit servings) were interpreted in a rigidly literal fashion and thus were unsuitable for this community. Preliminary findings suggest that infographics may motivate health-promoting behaviors. DISCUSSION Infographics should be information-rich, contextualize the information for the viewer, and yield an accurate meaning even if interpreted literally. CONCLUSION Carefully designed infographics can be useful tools to support comprehension and thus help patients engage with their own health data. Infographics may contribute to patients' ability to participate in the Learning Health System through participation in the development of a robust data utility, use of clinical communication tools for health self-management, and involvement in building knowledge through patient-reported outcomes.
Author name disambiguation is a challenging problem in computer science. The problem arises from the fact that many authors share similar or identical names. Although some scholarly databases assign unique author identifiers, levels of accuracy are often unacceptable—especially for authors with common names. Existing algorithms have largely not leveraged institutional data on individual researchers. We are extending ReCiter, an agglomerative clustering algorithm for author name disambiguation, for use in publication management at our institution. The system uses available institutional data on researchers, including primary and secondary departments, history of co-investigatorships on grants and co-authorships, favored journals, and years of authors' terminal academic degrees. We are investigating the use of machine learning approaches to optimize system performance, and are planning to make the system available as a suite of freely available, open-source tools.
Repository Citation Bouquin, D. R., & Bales, M. E. (2015). Integrating External Resources into Health Informatics and Computing Instruction: Emerging Roles for Librarians and Information Professionals. University of Massachusetts and New England Area Librarian e-Science Symposium. https://doi.org/10.13028/ tk38-sp95. Retrieved from https://escholarship.umassmed.edu/escience_symposium/2015/posters/6
OBJECTIVE:Publications are a key data source for investigator profiles and research networking systems. We developed ReCiter, an algorithm that automatically extracts bibliographies from PubMed using institutional information about the target investigators. METHODS:ReCiter executes a broad query against PubMed, groups the results into clusters that appear to constitute distinct author identities and selects the cluster that best matches the target investigator. Using information about investigators from one of our institutions, we compared ReCiter results to queries based on author name and institution and to citations extracted manually from the Scopus database. Five judges created a gold standard using citations of a random sample of 200 investigators. RESULTS:About half of the 10,471 potential investigators had no matching citations in PubMed, and about 45% had fewer than 70 citations. Interrater agreement (Fleiss' kappa) for the gold standard was 0.81. Scopus achieved the best recall (sensitivity) of 0.81, while name-based queries had 0.78 and ReCiter had 0.69. ReCiter attained the best precision (positive predictive value) of 0.93 while Scopus had 0.85 and name-based queries had 0.31. DISCUSSION:ReCiter accesses the most current citation data, uses limited computational resources and minimizes manual entry by investigators. Generation of bibliographies using named-based queries will not yield high accuracy. Proprietary databases can perform well but requite manual effort. Automated generation with higher recall is possible but requires additional knowledge about investigators.
OBJECTIVES:To develop a method for investigating co-authorship patterns and author team characteristics associated with the publications in high-impact journals through the integration of public MEDLINE data and institutional scientific profile data. METHODS:For all current researchers at Columbia University Medical Center, we extracted their publications from MEDLINE authored between years 2007 and 2011 and associated journal impact factors, along with author academic ranks and departmental affiliations obtained from Columbia University Scientific Profiles (CUSP). Chi-square tests were performed on co-authorship patterns, with Bonferroni correction for multiple comparisons, to identify team composition characteristics associated with publication impact factors. We also developed co-authorship networks for the 25 most prolific departments between years 2002 and 2011 and counted the internal and external authors, inter-connectivity, and centrality of each department. RESULTS:Papers with at least one author from a basic science department are significantly more likely to appear in high-impact journals than papers authored by those from clinical departments alone. Inclusion of at least one professor on the author list is strongly associated with publication in high-impact journals, as is inclusion of at least one research scientist. Departmental and disciplinary differences in the ratios of within- to outside-department collaboration and overall network cohesion are also observed. CONCLUSIONS:Enrichment of co-authorship patterns with author scientific profiles helps uncover associations between author team characteristics and appearance in high-impact journals. These results may offer implications for mentoring junior biomedical researchers to publish on high-impact journals, as well as for evaluating academic progress across disciplines in modern academic medical centers.