During the twentieth century, inflammatory bowel disease (IBD) was considered a disease of early industrialized regions in North America, Europe and Oceania1. At the turn of the twenty-first century, IBD incidence increased in newly industrialized and emerging regions in Africa, Asia and Latin America, while the prevalence in early industrialized regions continued to grow steadily2-4. Changes in the incidence and prevalence denote the evolution of IBD across four epidemiologic stages: stage 1 (emergence), characterized by low incidence and prevalence; stage 2 (acceleration in incidence), marked by rapidly rising incidence and low prevalence; and stage 3 (compounding prevalence), where the incidence decelerates, plateaus or declines while the prevalence steadily increases. A fourth stage (prevalence equilibrium) has been proposed in which the prevalence slope plateaus due to demographic shifts in an ageing IBD population, but it has not yet been evidenced. To date, these stages have remained theoretical, lacking specific numerical indicators to define transition points. Here, using real-world data from 522 population-based studies encompassing 82 global regions and spanning more than a century (1920-2024), we show spatiotemporal transitions across stages 1-3 and model stage 4 progression. Understanding the evolution of IBD across epidemiologic stages enables healthcare systems to better anticipate the future worldwide burden of IBD.
Background Menopause is a normal transition in a woman’s life. For some women, it is a stage without significant difficulties; for others, menopause symptoms can severely affect their quality of life. This study developed and validated a case definition for problematic menopause using Canadian primary care electronic medical records, which is an essential step in examining the condition and improving quality of care. Methods We used data from the Canadian Primary Care Sentinel Surveillance Network including billing and diagnostic codes, diagnostic free-text, problem list entries, medications, and referrals. These data formed the basis of an expert-reviewed reference standard data set and contained the features that were used to train a machine learning model based on classification and regression trees. An ad hoc feature importance measure coupled with recursive feature elimination and clustering were applied to reduce our initial 86,000 element feature set to a few tens of the most relevant features in the data, while class balancing was accomplished with random under- and over-sampling. The final case definition was generated from the tree-based machine learning model output combined with a feature importance algorithm. Two independent samples were used: one for training / testing the machine learning algorithm and the other for case definition validation. Results We randomly selected 2,776 women aged 45–60 for this analysis and created a case definition, consisting of two occurrences within 24 months of International Classification of Diseases, Ninth Revision, Clinical Modification code 627 (or any sub-codes) OR one occurrence of Anatomical Therapeutic Chemical classification code G03CA (or any sub-codes) within the patient chart, that was highly effective at detecting problematic menopause cases. This definition produced a sensitivity of 81.5% (95% CI: 76.3-85.9%), specificity of 93.5% (91.9-94.8%), positive predictive value of 73.8% (68.3-78.6%), and negative predictive value of 95.7% (94.4-96.8%). Conclusion Our case definition for problematic menopause demonstrated high validity metrics and so is expected to be useful for epidemiological study and surveillance. This case definition will enable future studies exploring the management of menopause in primary care settings.
Abstract Background Menopause is a normal transition in a women’s life. For some women, it is a stage without significant difficulties; for others, menopause symptoms can severely affect their quality of life. Identifying problematic menopause is essential to study the condition and to improve quality of care. This study developed and validated a case definition for problem menopause using Canadian primary care electronic medical records. Methods We used data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN). A case definition was developed using a reference set created by expert reviewers and a machine learning approach was applied to produce a case definition. Methods to select the most appropriate features and to re-balance our cohort were also applied. Results We randomly selected 2,776 women aged 45–60 for this analysis. An algorithm of two occurrences of ICD-9-CM code 627 in diagnosis fields within 24 months OR one occurrence of ATC code G03CA in medication fields detected problem menopause. This definition produced sensitivity 81.5% (95%CI 76.3%-85.9%), specificity of 93.5% (95%CI 91.9%-94.8%), positive predicted value 73.8% (95%CI 68.3%-78.6%), and negative predicted value 95.7% (95%CI 94.4%-96.8%). Conclusion Our case definition for problem menopause is useful for epidemiological study and demonstrated strong validity metrics. This case definition will help inform future studies exploring management of menopause in primary care settings.
Background Primary care electronic medical record (EMR) data are emerging as a useful source for secondary uses, such as disease surveillance, health outcomes research, and practice improvement. These data capture clinical details about patients’ health status, as well as behavioural risk factors, such as smoking. While the importance of documenting smoking status in a healthcare setting is recognized, the quality of smoking data captured in EMRs is variable. This study was designed to test methods aimed at improving the quality of patient smoking information in a primary care EMR database. Methods EMR data from community primary care settings extracted by two regional practice-based research networks in Alberta, Canada were used. Patients with at least one encounter in the previous 2 years (2016–2018) and having hypertension according to a validated definition were included ( n = 48,377). Multiple imputation was tested under two different assumptions for missing data (smoking status is missing at random and missing not-at-random). A third method tested a novel pattern matching algorithm developed to augment smoking information in the primary care EMR database. External validity was examined by comparing the proportions of smoking categories generated in each method with a general population survey. Results Among those with hypertension, 40.8% ( n = 19,743) had either no smoking information recorded or it was not interpretable and considered missing. Those with missing smoking data differed statistically by demographics, clinical features, and type of EMR system used in the clinic. Both multiple imputation methods produced fully complete smoking status information, with the proportion of current smokers estimated at 25.3% (data missing at random) and 12.5% (data missing not-at-random). The pattern-matching algorithm classified 18.2% of patients as current smokers, similar to the population-based survey (18.9%), but still resulted in missing smoking information for 23.6% of patients. The algorithm was estimated to be 93.8% accurate overall, but varied by smoking status category. Conclusion Multiple imputation and algorithmic pattern-matching can be used to improve EMR data post-extraction but the recommended method depends on the purpose of secondary use (e.g. practice improvement or epidemiological analyses).
Background: Electronic medical records (EMR) are commonly used in primary care to document patient measurements including height and weight that are then used to produce body mass index (BMI) scores. However, little is known about the proportion of waist circumference (WC) documentation compared to BMI and the characteristics of patients with these measures. This study used a pan-Canadian research database, sourced from primary care EMRs, to describe BMI and WC documentation in primary care. Methods: A retrospective cohort design of primary care providers participating in the Canadian Primary Care Sentinel Surveillance Network (CPCSSN), this study presented descriptive, observational findings of EMR inputs. Frequencies and percentages of median BMI and WC documentation in CPCSSN EMRs and patient demographic characteristics are compared. Results: Of 707,819 Canadian patients aged of 40 or older, at least one BMI input was recorded for 58.6% and 11.5% had WC notations. The majority of patients (98.1%) with at least one WC measurement also had a BMI measurement while conversely 19.2% of patients with at least one BMI measurement also had a WC measurement. The most common median BMI category was overweight (36.9%) and median WC was 95.0 centimetres (IQR = 21.5). Conclusions: This study reports the documentation of obesity and overweight in select Canadian primary care EMRs infrequently recorded WC when compared to BMI. Future studies should examine the frequency and categories of anthropometric measurements in people with commonly managed chronic conditions and whether BMI and WC inputs are missing at random. Trial registration: Not applicable for this study.
IntroductionUse of administrative health data and primary care electronic medical record data are both ubiquitous in Alberta, but linkage between them at patient level and implementation of the linked data into primary care practice are rare. This demonstration project sought to achieve this for a sample of patients with diabetes. Objectives and ApproachAcademic family physicians in the Department of Family Medicine at the University of Calgary who participate in the Canadian Primary Care Sentinel Surveillance Network (CPCSSN) identified diabetes–related variables, either in their EMRs or in administrative data, that they wished to obtain in a linked dataset. Secure data linkage was obtained through Alberta Health Services (the provincial health authority) following transmission of patient mapping files direct from the clinics. The de-identified, linked, patient data was then transferred to CPCSSN-Alberta data managers for processing and displayed to users through an interactive Diabetes Dashboard. Results2598 patients with diabetes were identified using a validated CPCSSN case definition from 47 family physicians in three clinics. CPCSSN EMR data included primary care encounters, date of diagnosis, deprivation index, BMI, blood pressure, comorbidity, diabetes medications prescribed, risk factors, etc. Administrative data included laboratory results (HbA1c, fasting blood glucose, cholesterol, triglycerides, creatinine), medication dispensed, emergency room visits, inpatient admissions and costs. Integrated, interactive provider reports were created and sent to participating physicians. The reports presented the information about diabetes patients at individual provider level, bench-marked at clinic, primary care network and provincial levels. Follow-up with providers led to further dashboard development . We propose to scale up implementation of the integrated diabetes database and dashboard to include all 23,000 CPCSSN-identified diabetes patients in Alberta. Conclusion/ImplicationsIntegration of EMR and administrative data and its application to clinical care, panel management, and quality improvement in primary care, as well as to surveillance and research, was feasible and acceptable to the family physicians participating in this project.