Swiss cohort studies provide high-quality longitudinal data, but finding and comparing relevant studies across cohorts has historically been challenging. The Swiss Personalized Health Network Cohort Consortium (SPHN-CC) was established to address these limitations by creating the first coordinated network of Swiss cohort studies within the internationally recognized Maelstrom Research catalogue. Participating cohorts were invited in 2021-2022, including longitudinal and cross-sectional studies with 1010-21,993 participants. Data collected include questionnaires, physical and cognitive assessments, administrative records, and biological samples. Variables were classified into 18 domains and 134 subdomains, and an online metadata catalogue was implemented to document study designs, explore variable content, and assess harmonization potential. The catalogue enables researchers to identify study-specific and harmonized variables for co-analysis. Core variables, such as age, sex/gender, anthropometrics, and medication use, are widely available, while other variables vary across cohorts. Harmonization assessments demonstrate that several key variables can be co-analyzed across multiple studies, supporting collaborative research with over 37,000 participants. A use case illustrates the potential for harmonizing and co-analyzing data across studies. The SPHN-CC strengthens Swiss cohort research by enhancing data discoverability, supporting harmonization, and facilitating cross-cohort and international research, providing a model for more efficient use of high-value longitudinal data.
Introduction While Canada is rich in databases useful to support healthcare research, they are widely distributed, often poorly documented, and it is challenging to identify relevant databases, apply for access, and eventually use, link or harmonise the data. Even if the databases needed to address specific questions are known, it is difficult and time-consuming to find the metadata, the ``data about the data'' required to understand the characteristics and data content of these resources. A solution to these challenges is creation of metadata catalogues, which detail metadata for multiple databases, not the actual data. Objectives Describe a new catalogue including metadata about Canadian medical and non-medical databases' characteristics and variables, and information to assist catalogue users in seeking data access. Methods Starting with a list of 385 national, provincial and regional databases, a group of physician-investigators, epidemiologists, data scientists and patient partners prioritised databases for inclusion. Metadata cataloguing occurred in steps: (i) description of the database with listing of its characteristics, and when available, (ii) addition of information about collected variables. Results 83 individual databases are documented in the Metadata Catalogue of the Sepsis Canada Network (https://www.maelstrom-research.org/network/sepsis). 57 are registries, 13 are cohort and 13 cross-sectional databases. 16 cover all of Canada, while another 13 cover most of the country; 45 focus on a single province. For 33 databases (38\%) the catalogue includes detailed information about variables collected. Conclusions This metadata catalogue includes databases collecting information spanning the continuum of medical care, non-medical data, and determinants of health. It is freely available online and extensively searchable. It can facilitate implementation of a wide range of research initiatives into medical conditions, medical care, and outcomes.
Objectives The Healthy Life Trajectories Initiative (HeLTI) is an international multistudy consortium that supports the development and integration of four randomised controlled trials (RCTs) conducted in South Africa, India, China and Canada. HeLTI aims to evaluate interventions to improve the health and well-being of mothers and children, starting from preconception through pregnancy and early childhood until age 5 years. This paper describes the process by which we prospectively harmonised the participating studies and provides a descriptive analysis of the study-specific harmonisation potential.Design Prospective harmonisation of four international RCTs.Methods A list of core variables to be collected across ten waves of data collection was defined. Taking this list into consideration, investigators developed country-specific questionnaires that were then assessed and adjusted to optimise the harmonisation potential across countries. As questionnaires were not identical, where required, processing scripts were generated to help transform the collected data into the core variable format.Setting The four RCTs are conducted in Canada, China, India and South Africa. The prospective harmonisation was led by the Maelstrom Research team in Canada.Participants Between 4500 and 6000 women planning to get pregnant are recruited in each RCT. Women remain in the study if they become pregnant inside the planned interval of 1–3 years, depending on the country.Results A total of 1962 variables from questionnaires, physical measurements and biospecimen analyses were defined across 10 timepoints of data collection and 3 subpopulations (mothers, partners and children). These variables cover 47 different domains of information. For the preconception phase, following the development of questionnaires and their implementation in the data collection software, 77.2% of the core variables defined can be created across the four studies.Conclusion The HeLTI harmonisation process was successful, and the datasets generated represent a valuable resource allowing researchers to address a wide range of research questions on the impact of behaviour change interventions on maternal and child health indicators in different populations.
The significance of Findable, Accessible, Interoperable, and Reusable (FAIR) data is increasing, particularly in the context of enhancing data reuse in research. The National Research Data Infrastructure for Personal Health Data (NFDI4Health) aims to enhance the findability, reusability, and interoperability of health data derived from epidemiological, clinical, and public health studies. NFDI4Health has established the German Central Health Study Hub to improve health data findability through rich metadata. The Maelstrom Catalog, provided by Maelstrom Research, offers a comprehensive dataset of labeled and harmonized study variables, thereby enhancing the findability and reusability of epidemiological data. Both platforms rely on standardized categorization to optimize data reuse. To facilitate this process, NFDI4Health developed the Metadata Annotation Workbench, which supports metadata annotation with standardized vocabulary. This paper presents an AI solution for automatic classification and annotation integrated into this service, using a BioBERT-based text classifier. The model achieved a weighted F1-score of over 92% and demonstrated improved annotation performance, particularly for non-experts. It accelerates variable categorization, thereby enhancing data findability and re-use. As a result, the categorization of study variables can be accelerated and we are confident that the further development of such AI approaches will reduce curatorial workload and promote semantically annotated interoperable data catalogs.
Adopting universal data standards such as OMOP and SNOMED is crucial for ensuring the integrity, interoperability, and utility of data in global health research, setting the foundation for more effective studies.
While its etiology is not fully elucidated, preterm birth represents a major public health concern as it is the leading cause of child mortality and morbidity. Stress is one of the most common perinatal conditions and may increase the risk of preterm birth. In this paper we aimed to investigate the association of maternal perceived stress and anxiety with length of gestation. We used harmonized data from five birth cohorts from Canada, France, and Norway. A total of 5297 pregnancies of singletons were included in the analysis of perceived stress and gestational duration, and 55,775 pregnancies for anxiety. Federated analyses were performed through the DataSHIELD platform using Cox regression models within intervals of gestational age. The models were fit for each cohort separately, and the cohort-specific results were combined using random effects study-level meta-analysis. Moderate and high levels of perceived stress during pregnancy were associated with a shorter length of gestation in the very/moderately preterm interval [moderate: hazard ratio (HR) 1.92 (95%CI 0.83, 4.48); high: 2.04 (95%CI 0.77, 5.37)], albeit not statistically significant. No association was found for the other intervals. Anxiety was associated with gestational duration in the very/moderately preterm interval [1.66 (95%CI 1.32, 2.08)], and in the early term interval [1.15 (95%CI 1.08, 1.23)]. Our findings suggest that perceived stress and anxiety are associated with an increased risk of earlier birth, but only in the earliest gestational ages. We also found an association in the early term period for anxiety, but the result was only driven by the largest cohort, which collected information the latest in pregnancy. This raised a potential issue of reverse causality as anxiety later in pregnancy could be due to concerns about early signs of a possible preterm birth.
Chapter 13 Maelstrom Research Approaches to Retrospective Harmonization of Cohort Data for Epidemiological Research Tina W. Wey, Tina W. Wey 1The Maelstrom Research Platform, Research Institute of the Health Centre, McGill University, Montreal, CanadaSearch for more papers by this authorIsabel Fortier, Isabel Fortier 2Department of Medicine, Division of Experimental Medicine, The Maelstrom Research Platform, Health Centre, McGill University, Montreal, CanadaSearch for more papers by this author Tina W. Wey, Tina W. Wey 1The Maelstrom Research Platform, Research Institute of the Health Centre, McGill University, Montreal, CanadaSearch for more papers by this authorIsabel Fortier, Isabel Fortier 2Department of Medicine, Division of Experimental Medicine, The Maelstrom Research Platform, Health Centre, McGill University, Montreal, CanadaSearch for more papers by this author Book Editor(s):Irina Tomescu-Dubrow, Irina Tomescu-Dubrow The Ohio State University, Department of Sociology, Columbus, 43210 United StatesSearch for more papers by this authorChristof Wolf, Christof Wolf GESIS, Leibniz Institute for the Social Sciences, Mannheim, 68159 GermanySearch for more papers by this authorKazimierz M. Slomczynski, Kazimierz M. Slomczynski The Ohio State University, Department of Sociology, Columbus, 43210 United StatesSearch for more papers by this authorJ. Craig Jenkins, J. Craig Jenkins The Ohio State University, Department of Sociology, Columbus, 43210 United StatesSearch for more papers by this author First published: 10 November 2023 https://doi.org/10.1002/9781119712206.ch13 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary Maelstrom Research brings together an interdisciplinary team to address the challenges faced in epidemiological research collaborations and enhance the use of study data. Our activities include developing and implementing methods and tools to facilitate data discovery, documentation, harmonization, and integrated analysis. These resources include generic guidelines for retrospective harmonization, a study cataloging toolkit, and an open-source software suite, which aim to provide rigorous and standardized resources that can be adapted to different project needs. To date, these tools have been applied in multiple research initiatives to harmonize data collected by different population-based cohort studies and generate valuable research datasets. In this chapter, we describe the application of the Maelstrom harmonization approach and tools in two large research initiatives that harmonized individual participant data across cohort studies: CanPath (the Canadian Partnership for Tomorrow's Health) and MINDMAP Promoting mental well-being and healthy aging in cities. We use this opportunity to discuss the applied process and challenges, solutions, and limitations with illustrative examples from the two projects. Key aspects of all harmonization projects include rigor in the harmonization process, ensuring high-quality data input and outputs, comprehensive and transparent documentation, and close communication and collaboration among data harmonizers, data producers, and data users. Data harmonization is an important and challenging part of modern epidemiological research, and the continued development and application of standardized methods and tools, along with sharing of expertise and experience, remain important for promoting rigorous and effective harmonization to support the needs of the research community. References Beenackers , M.A. , Doiron , D. , Fortier , I. et al. ( 2018 ). MINDMAP: establishing an integrated database infrastructure for research in ageing, mental well-being, and the urban environment . BMC Public Health 18 ( 1 ): 158 . https://doi.org/10.1186/s12889-018-5031-7 . 10.1186/s12889-018-5031-7 PubMedWeb of Science®Google Scholar Bergeron , J. , Doiron , D. , Marcon , Y. et al. ( 2018 ). Fostering population-based cohort data discovery: the Maelstrom Research cataloguing toolkit . PLoS One 13 ( 7 ): e0200926 . https://doi.org/10.1371/journal.pone.0200926 . 10.1371/journal.pone.0200926 Web of Science®Google Scholar Bergeron , J. , Massicotte , R. , Atkinson , S. et al., & on behalf of the ReACH member cohorts' principal investigators. ( 2020 ). Cohort profile: research advancement through cohort cataloguing and harmonization (ReACH) . International Journal of Epidemiology , dyaa207. https://doi.org/10.1093/ije/dyaa207 . 10.1093/ije/dyaa207 Web of Science®Google Scholar Borugian , M.J. , Robson , P. , Fortier , I. et al. ( 2010 ). The Canadian Partnership for Tomorrow Project: building a pan-Canadian research platform for disease prevention . CMAJ 182 ( 11 ): 1197 – 1201 . https://doi.org/10.1503/cmaj.091540 . 10.1503/cmaj.091540 PubMedWeb of Science®Google Scholar Burton , P.R. , Hansell , A.L. , Fortier , I. et al. ( 2009 ). Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology . International Journal of Epidemiology 38 ( 1 ): 263 – 273 . https://doi.org/10.1093/ije/dyn147 . 10.1093/ije/dyn147 PubMedWeb of Science®Google Scholar Burton , P.R. , Banner , N. , Elliot , M.J. et al. ( 2017 ). Policies and strategies to facilitate secondary use of research data in the health sciences . International Journal of Epidemiology https://doi.org/10.1093/ije/dyx195 . 10.1093/ije/dyx195 Web of Science®Google Scholar Dalziel , M. , Roswell , J. , Tahmina , T.N. , and Xiao , Z. ( 2012 ). Impact of government investments in research & innovation: review of academic investigations . Optimum Online: The Journal of Public Sector Management 42 ( 2 ): http://www.optimumonline.ca/print.phtml?e=mesokurj&id=413 . Google Scholar Doiron , D. , Burton , P. , Marcon , Y. et al. ( 2013 ). Data harmonization and federated analysis of population-based studies: the BioSHaRE project . Emerging Themes in Epidemiology 10 ( 1 ): 12 . https://doi.org/10.1186/1742-7622-10-12 . 10.1186/1742-7622-10-12 PubMedGoogle Scholar Doiron , D. , Marcon , Y. , Fortier , I. et al. ( 2017 ). Software application profile: opal and mica: open-source software solutions for epidemiological data management, harmonization and dissemination . International Journal of Epidemiology 46 ( 5 ): 1372 – 1378 . https://doi.org/10.1093/ije/dyx180 . 10.1093/ije/dyx180 PubMedWeb of Science®Google Scholar Dummer , T.J.B. , Awadalla , P. , Boileau , C. et al. with the CPTP Regional Cohort Consortium. ( 2018 ). The Canadian Partnership for Tomorrow Project: a pan-Canadian platform for research on chronic disease prevention . CMAJ 190 ( 23 ): E710 – E717 . https://doi.org/10.1503/cmaj.170292 . 10.1503/cmaj.170292 PubMedWeb of Science®Google Scholar Fortier , I. , Raina , P. , Van den Heuvel , E.R. et al. ( 2017 ). Maelstrom research guidelines for rigorous retrospective data harmonization . International Journal of Epidemiology 46 ( 1 ): 103 – 105 . https://doi.org/10.1093/ije/dyw075 . 10.1093/ije/dyw075 PubMedWeb of Science®Google Scholar Fortier , I. , Dragieva , N. , Saliba , M. et al., & with the Canadian Partnership for Tomorrow Project's scientific directors and the Harmonization Standing Committee. ( 2019 ). Harmonization of the health and risk factor questionnaire data of the Canadian Partnership for Tomorrow Project: a descriptive analysis . CMAJ Open 7 ( 2 ): E272 – E282 . https://doi.org/10.9778/cmajo.20180062 . 10.9778/cmajo.20180062 PubMedGoogle Scholar Gallacher , J. and Hofer , S.M. ( 2011 ). Generating large-scale longitudinal data resources for aging research . The Journals of Gerontology: Series B 66B ( suppl_1 ): i172 – i179 . https://doi.org/10.1093/geronb/gbr047 . 10.1093/geronb/gbr047 Google Scholar Gaye , A. , Marcon , Y. , Isaeva , J. et al. ( 2014 ). DataSHIELD: taking the analysis to the data, not the data to the analysis . International Journal of Epidemiology 43 ( 6 ): 1929 – 1944 . https://doi.org/10.1093/ije/dyu188 . 10.1093/ije/dyu188 PubMedWeb of Science®Google Scholar Gaziano , J.M. ( 2010 ). The evolution of population science: advent of the mega cohort . JAMA 304 ( 20 ): 2288 – 2289 . https://doi.org/10.1001/jama.2010.1691 . 10.1001/jama.2010.1691 CASPubMedWeb of Science®Google Scholar Graham , E.K. , Weston , S.J. , Turiano , N.A. et al. ( 2020 ). Is healthy neuroticism associated with health behaviors? A coordinated integrative data analysis . Collabra: Psychology 6 ( 1 ): 32 . https://doi.org/10.1525/collabra.266 . 10.1525/collabra.266 PubMedGoogle Scholar Griffith , L.E. , van den Heuvel , E. , Raina , P. et al. ( 2016 ). Comparison of standardization methods for the harmonization of phenotype data: an application to cognitive measures . American Journal of Epidemiology 184 ( 10 ): 770 – 778 . https://doi.org/10.1093/aje/kww098 . 10.1093/aje/kww098 PubMedWeb of Science®Google Scholar Jaddoe , V.W.V. , Felix , J.F. , Andersen , A.-M.N. et al. ( 2020 ). The LifeCycle project-EU child cohort network: a federated analysis infrastructure and harmonized data of more than 250,000 children and parents . European Journal of Epidemiology 35 ( 7 ): 709 – 724 . https://doi.org/10.1007/s10654-020-00662-z . 10.1007/s10654-020-00662-z PubMedWeb of Science®Google Scholar Lesko , C.R. , Jacobson , L.P. , Althoff , K.N. et al. ( 2018 ). Collaborative, pooled and harmonized study designs for epidemiologic research: challenges and opportunities . International Journal of Epidemiology 47 ( 2 ): 654 – 668 . https://doi.org/10.1093/ije/dyx283 . 10.1093/ije/dyx283 PubMedWeb of Science®Google Scholar Little , J. , Higgins , J.P. , Ioannidis , J.P. et al. Studies, STrengthening the REporting of Genetic Association ( 2009 ). STrengthening the REporting of Genetic Association studies (STREGA): an extension of the STROBE statement . PLoS Medicine 6 ( 2 ): e22 . https://doi.org/10.1371/journal.pmed.1000022 . 10.1371/journal.pmed.1000022 Web of Science®Google Scholar Manolio , T.A. , Bailey-Wilson , J.E. , and Collins , F.S. ( 2006 ). Genes, environment and the value of prospective cohort studies . Nature Reviews Genetics 7 ( 10 ): 812 – 820 . https://doi.org/10.1038/nrg1919 . 10.1038/nrg1919 CASPubMedWeb of Science®Google Scholar Marcon , Y. , Bishop , T. , Avraam , D. et al. ( 2021 ). Orchestrating privacy-protected big data analyses of data from different resources with R and DataSHIELD . PLoS Computational Biology 17 ( 3 ): e1008880 . https://doi.org/10.1371/journal.pcbi.1008880 . 10.1371/journal.pcbi.1008880 PubMedWeb of Science®Google Scholar Moher , D. , Liberati , A. , Tetzlaff , J. et al. ( 2009 ). Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement . PLoS Medicine 6 ( 7 ): e1000097 . https://doi.org/10.1371/journal.pmed.1000097 . 10.1371/journal.pmed.1000097 Web of Science®Google Scholar Pastorino , S. , Bishop , T. , Crozier , S. et al. ( 2019 ). Associations between maternal physical activity in early and late pregnancy and offspring birth size: remote federated individual level meta-analysis from eight cohort studies . BJOG: An International Journal of Obstetrics & Gynaecology 126 ( 4 ): 459 – 470 . https://doi.org/10.1111/1471-0528.15476 . 10.1111/1471-0528.15476 CASPubMedWeb of Science®Google Scholar R Core Team ( 2019 ). R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing https://www.R-project.org . Google Scholar Richmond , R.C. , Al-Amin , A. , Davey Smith , G. , and Relton , C.L. ( 2014 ). Approaches for drawing causal inferences from epidemiological birth cohorts: a review . Early Human Development 90 ( 11 ): 769 – 780 . https://doi.org/10.1016/j.earlhumdev.2014.08.023 . 10.1016/j.earlhumdev.2014.08.023 PubMedWeb of Science®Google Scholar Roger , V.L. , Boerwinkle , E. , Crapo , J.D. et al. ( 2015 ). Strategic transformation of population studies: recommendations of the working group on epidemiology and population sciences from the National Heart, Lung, and Blood Advisory Council and Board of External Experts . American Journal of Epidemiology 363 – 368 . 10.1093/aje/kwv011 PubMedWeb of Science®Google Scholar RStudio Team ( 2016 ). RStudio: Integrated Development Environment for R . RStudio, Inc. http://www.rstudio.com . Google Scholar Sanchez-Niubo , A. , Egea-Cortés , L. , Olaya , B. et al. ( 2019 ). Cohort profile: the ageing trajectories of health – longitudinal opportunities and synergies (ATHLOS) project . International Journal of Epidemiology 48 ( 4 ): 1052 – 1053i . https://doi.org/10.1093/ije/dyz077 . 10.1093/ije/dyz077 PubMedWeb of Science®Google Scholar Thompson , A. ( 2009 ). Thinking big: large-scale collaborative research in observational epidemiology . European Journal of Epidemiology 24 ( 12 ): 727 . https://doi.org/10.1007/s10654-009-9412-1 . 10.1007/s10654-009-9412-1 PubMedWeb of Science®Google Scholar Turiano , N.A. , Graham , E.K. , Weston , S.J. et al. ( 2020 ). Is healthy neuroticism associated with longevity? A coordinated integrative data analysis . Collabra: Psychology 6 ( 1 ): 33 . https://doi.org/10.1525/collabra.268 . 10.1525/collabra.268 PubMedGoogle Scholar UNESCO Institute for Statistics ( 2012 ). International Standard Classification of Education: ISCED 2011 . Montreal, QC : UNESCO Institute for Statistics (UIS) http://dx.doi.org/10.15220/978-92-9189-123-8-en. 10.15220/978-92-9189-123-8-en Google Scholar Van den Heuvel , E.R. and Griffith , L.E. ( 2016 ). Statistical harmonization methods in individual participants data meta-analysis are highly needed . Biometrics & Biostatistics International Journal 3 ( 3 ): 70 – 72 . https://doi.org/10.15406/bbij.2016.03.00064 . 10.15406/bbij.2016.03.00064 Google Scholar Vandenbroucke , J.P. , Elm , E. von , Altman , D.G. et al., for the STROBE Initiative. ( 2007 ). Strengthening the reporting of observational studies in epidemiology (STROBE): explanation and elaboration . PLoS Medicine 4 ( 10 ): e297 . https://doi.org/10.1371/journal.pmed.0040297 . 10.1371/journal.pmed.0040297 Web of Science®Google Scholar Wey , T.W. , Doiron , D. , Wissa , R. et al. ( 2021 ). Overview of retrospective data harmonisation in the MINDMAP project: process and results . Journal of Epidemiology and Community Health 75 ( 5 ): 433 – 441 . https://doi.org/10.1136/jech-2020-214259 . 10.1136/jech-2020-214259 PubMedWeb of Science®Google Scholar Wilkinson , M.D. , Dumontier , M. , Aalbersberg , I.J.J. et al. ( 2016 ). The FAIR guiding principles for scientific data management and stewardship . Scientific Data 3 : https://doi.org/10.1038/sdata.2016.18 . 10.1038/sdata.2016.18 PubMedWeb of Science®Google Scholar Survey Data Harmonization in the Social Sciences ReferencesRelatedInformation
Background As a teratogen, alcohol exposure during pregnancy can impact fetal development and result in adverse birth outcomes. Despite the clinical and social importance of prenatal alcohol use, limited routinely collected information or epidemiological data exists in Canada. The aim of this study was to pool data from multiple Canadian cohort studies to identify sociodemographic characteristics before and during pregnancy that were associated with alcohol consumption during pregnancy and to assess the impact of different patterns of alcohol use on birth outcomes. Methods We harmonized information collected (e.g., pregnant women’s alcohol intake, infants' gestational age and birth weight) from five Canadian pregnancy cohort studies to consolidate a large sample ( n = 11,448). Risk factors for any alcohol use during pregnancy, including any alcohol use prior to pregnancy recognition, and binge drinking, were estimated using binomial regressions including fixed effects of pregnancy cohort membership and multiple maternal risk factors. Impacts of alcohol use during pregnancy on birth outcomes (preterm birth and low birth weight for gestational) were also estimated using binomial regression models. Results In analyses adjusting for multiple risk factors, women’s alcohol use during pregnancy, both any use and any binge drinking, was associated with drinking prior to pregnancy, smoking during pregnancy, and white ethnicity. Higher income level was associated with any drinking during pregnancy. Neither drinking during pregnancy nor binge drinking during pregnancy was significantly associated with preterm delivery or low birth weight for gestational age in our sample. Conclusions Pooling data across pregnancy cohort studies allowed us to create a large sample of Canadian women and investigate the risk factors for alcohol consumption during pregnancy. We suggest that future pregnancy and birth cohorts should always include questions related to the frequency and amount of alcohol consumed before and during pregnancy that are prospectively harmonized to support data reusability and collaborative research.
Introduction Severe bronchopulmonary dysplasia (BPD) is a well-known factor consistently associated with impaired cognitive outcomes. Regarding reported benefits on long-term neurodevelopmental outcomes, the potential adverse effects of high-dose docosahexaenoic acid (DHA) supplementation on this short-term neonatal morbidity need further investigations in infants born very preterm. This study will determine whether high-dose DHA enteral supplementation during the neonatal period is associated with the risk of severe BPD at 36 weeks’ postmenstrual age (PMA) compared with control, in contemporary cohorts of preterm infants born at less than 29 weeks of gestation. Methods and analysis As part of an Australian–Canadian collaboration, we will conduct an individual participant data (IPD) meta-analysis of randomised controlled trials targeting infants born at less than 29 weeks of gestation and evaluating the effect of high-dose DHA enteral supplementation in the neonatal period compared with a control. Primary outcome will be severe grades of BPD (yes/no) at 36 weeks’ PMA harmonised according to a recent definition that predicts early childhood morbidities. Other outcomes will be survival without severe BPD, death, BPD severity grades, serious brain injury, severe retinopathy of prematurity, patent ductus arteriosus and necrotising enterocolitis requiring surgery, sepsis, combined neonatal morbidities and growth. Severe BPD will be compared between groups using a multivariate generalised estimating equations log-binomial regression model. Subgroup analyses are planned for gestational age, sex, small-for-gestational age, presence of maternal chorioamnionitis and mode of delivery. Ethics and dissemination The conduct of each trial was approved by institutional research ethics boards and written informed consent was obtained from participating parents. A collaboration and data sharing agreement will be signed between participating authors and institutions. This IPD meta-analysis will document the role of DHA in nutritional management of BPD. Findings will be disseminated through conferences, media interviews and publications to peer-reviewed journals. PROSPERO registration number CRD42023431063. Trial registration number NCT05915806 .
Optimizing research on the developmental origins of health and disease (DOHaD) involves implementing initiatives maximizing the use of the available cohort study data; achieving sufficient statistical power to support subgroup analysis; and using participant data presenting adequate follow-up and exposure heterogeneity. It also involves being able to undertake comparison, cross-validation, or replication across data sets. To answer these requirements, cohort study data need to be findable, accessible, interoperable, and reusable (FAIR), and more particularly, it often needs to be harmonized. Harmonization is required to achieve or improve comparability of the putatively equivalent measures collected by different studies on different individuals. Although the characteristics of the research initiatives generating and using harmonized data vary extensively, all are confronted by similar issues. Having to collate, understand, process, host, and co-analyze data from individual cohort studies is particularly challenging. The scientific success and timely management of projects can be facilitated by an ensemble of factors. The current document provides an overview of the ‘life course’ of research projects requiring harmonization of existing data and highlights key elements to be considered from the inception to the end of the project.
OBJECTIVES:Existing individual-level human data cover large populations on many dimensions such as lifestyle, demography, laboratory measures, clinical parameters, etc. Recent years have seen large investments in data catalogues to FAIRify data descriptions to capitalise on this great promise, i.e. make catalogue contents more Findable, Accessible, Interoperable and Reusable. However, their valuable diversity also created heterogeneity, which poses challenges to optimally exploit their richness.METHODS:In this opinion review, we analyse catalogues for human subject research ranging from cohort studies to surveillance, administrative and healthcare records.RESULTS:We observe that while these catalogues are heterogeneous, have various scopes, and use different terminologies, still the underlying concepts seem potentially harmonizable. We propose a unified framework to enable catalogue data sharing, with catalogues of multi-center cohorts nested as a special case in catalogues of real-world data sources. Moreover, we list recommendations to create an integrated community of metadata catalogues and an open catalogue ecosystem to sustain these efforts and maximise impact.CONCLUSIONS:We propose to embrace the autonomy of motivated catalogue teams and invest in their collaboration via minimal standardisation efforts such as clear data licensing, persistent identifiers for linking same records between catalogues, minimal metadata 'common data elements' using shared ontologies, symmetric architectures for data sharing (push/pull) with clear provenance tracks to process updates and acknowledge original contributors. And most importantly, we encourage the creation of environments for collaboration and resource sharing between catalogue developers, building on international networks such as OpenAIRE and research data alliance, as well as domain specific ESFRIs such as BBMRI and ELIXIR.
Introduction Type 2 diabetes mellitus (T2DM) onset before 40 years of age has a magnified lifetime risk of cardiovascular disease. Diastolic dysfunction is its earliest cardiac manifestation. Low energy diets incorporating meal replacement products can induce diabetes remission, but do not lead to improved diastolic function, unlike supervised exercise interventions. We are examining the impact of a combined low energy diet and supervised exercise intervention on T2DM remission, with peak early diastolic strain rate, a sensitive MRI-based measure, as a key secondary outcome. Methods and analysis This prospective, randomised, two-arm, open-label, blinded-endpoint efficacy trial is being conducted in Montreal, Edmonton and Leicester. We are enrolling 100 persons 18–45 years of age within 6 years’ T2DM diagnosis, not on insulin therapy, and with obesity. During the intensive phase (12 weeks), active intervention participants adopt an 800–900 kcal/day low energy diet combining meal replacement products with some food, and receive supervised exercise training (aerobic and resistance), three times weekly. The maintenance phase (12 weeks) focuses on sustaining any weight loss and exercise practices achieved during the intensive phase; products and exercise supervision are tapered but reinstituted, as applicable, with weight regain and/or exercise reduction. The control arm receives standard care. The primary outcome is T2DM remission, (haemoglobin A1c of less than 6.5% at 24 weeks, without use of glucose-lowering medications during maintenance). Analysis of remission will be by intention to treat with stratified Fisher’s exact test statistics. Ethics and dissemination The trial is approved in Leicester (East Midlands – Nottingham Research Ethics Committee (21/EM/0026)), Montreal (McGill University Health Centre Research Ethics Board (RESET for remission/2021-7148)) and Edmonton (University of Alberta Health Research Ethics Board (Pro00101088). Findings will be shared widely (publications, presentations, press releases, social media platforms) and will inform an effectiveness trial. Trial registration number ISRCTN15487120 .
BACKGROUND:Preterm birth is one of the most important contributors to neonatal mortality and morbidity. Experiencing stress during pregnancy may increase the risk of adverse birth outcomes, including preterm birth. This association has been observed in previous studies, but differences in measures used limit comparability. OBJECTIVE:The objective of the study was to investigate the association between two measures of maternal stress during pregnancy, life stress and emotional distress, and gestation duration. METHODS:Women recruited in the Danish National Birth Cohort from 1996 to 2002, who provided information on their stress level during pregnancy and expecting a singleton baby, were included in the study. We assessed the associations between the level of life stress and emotional distress in quartiles, both collected at 31 weeks of pregnancy on average, and the rate of giving birth using Cox regression within intervals of the gestational period. RESULTS:A total of 80,991 pregnancies were included. Women reporting moderate or high levels of life stress vs no stress had a higher rate of giving birth earlier within all intervals of gestational age (e.g. high level: 27-33 weeks: hazard ratio (HR) 1.38, 95% confidence interval (CI) 1.04, 1.84; 34-36 weeks: 1.10, 95% CI 0.97, 1.25; 37-38 weeks: 1.21, 95% CI 1.15, 1.28). These associations between life stress and preterm birth were mainly driven by pregnancy worries. For emotional distress, a high level of distress was associated with shorter length of gestation in the preterm (27-33 weeks: 1.38, 95% CI 1.02, 1.86; 34-36 weeks: 1.05, 95% CI 0.91, 1.19) and early term (1.11, 95% CI 1.04, 1.17) intervals. CONCLUSIONS:Emotional distress and life stress were shown to be associated with gestational age at birth, with pregnancy-related stress being the single stressor driving the association. This suggests that reverse causality may, at least in parts, explain the earlier findings of stress as a risk factor for preterm birth.
Objectives: The BETTER (BEhaviors, Therapies, TEchnologies and hypoglycemic Risk in Type 1 diabetes) registry is a type 1 diabetes population surveillance system codeveloped with patient partners to address the burden of hypoglycemia and assess the impact of new therapies and technologies. The aim of this report was to describe the baseline characteristics of the BETTER registry cohort.Methods: A cross-sectional baseline evaluation was performed of a Canadian clinical cohort established after distribution of an online questionnaire. Participants were recruited through clinics, public foundations, advertising and social media. As of February 2021,1,430 persons >14 years of age and living with type 1 diabetes or latent-autoimmune diabetes (LADA) were enrolled. The trial was registered on ClinicalTrials.gov (NCT03720197).Results: Participants were (mean +/- standard deviation) 41.2 +/- 15.7 years old with a diabetes duration of 22.0 +/- 14.7 years, 62.0% female, 92.1% Caucasian and 7.8% self-reporting as LADA, with 40.9% using a continuous subcutaneous insulin infusion (CSII) system and 78.0% using a continuous glucose monitoring (CGM) system. The most recent glycated hemoglobin <7% was reported by 29.7% of participants. At least 1 episode of hypoglycemia <3.0 mmol/L (level 2-H) in the last month was reported by 78.4% of participants, with a median (interquartile range) of 5 (3, 10) episodes. The occurrence of severe hypoglycemia (level 3-H) in the last 12 months was reported by 13.3% of participants. Among these, the median number of episodes was 2 (1, 3).Conclusions: We have established the first surveillance registry for people living with type 1 diabetes in Canada relying on patient-reported outcomes and experiences. Hypoglycemia is a highly prevalent burden despite a relatively wide adoption of CSII and CGM use.(c) 2022 Canadian Diabetes Association.
Introduction Childhood overweight and obesity (OWO) is a primary global health challenge. Childhood OWO prevention is now a public health priority in China. The Sino-Canadian Healthy Life Trajectories Initiative (SCHeLTI), one of four trials being undertaken by the international HeLTI consortium, aims to evaluate the effectiveness of a multifaceted, community-family-mother-child intervention on childhood OWO and non-communicable diseases risk.Methods and analysis This is a multicentre, cluster-randomised, controlled trial conducted in Shanghai, China. The unit of randomisation is the service area of Maternal Child Health Units (N=36). We will recruit 4500 women/partners/families in maternity and district level hospitals. Participants in the intervention group will receive a multifaceted, integrated package of health promotion interventions beginning in preconception or in the first trimester of pregnancy, continuing into infancy and early childhood. The intervention, which is centred on a modified motivational interviewing approach, will target early-life maternal and child risk factors for adiposity. Through the development of a biological specimen bank, we will study potential mechanisms underlying the effects of the intervention. The primary outcome for the trial is childhood OWO (body mass index for age ≥85th percentile) at 5 years of age, based on WHO sex-specific standards. The study has a power of 0.8 (α=0.05) to detect a 30% risk reduction in the proportion of children with OWO at 5 years of age, from 24.4% in the control group to 17% in the intervention group. Recruitment was launched on 30 August 2018 for the pilot study and 10 January 2019 for the formal study.Ethics and dissemination The study has been approved by the Medical Research Ethics Committee of the International Peace Maternity and Child Health Hospital in Shanghai, China, and the Research Ethics Board of the Centre Intégré Universitaire de Santé et Services Sociaux de l’Estrie–CHUS in Sherbrooke, Canada. Data sharing policies are consistent with the governance policy of the HeLTI consortium and government legislation.Trial registration number ChiCTR1800017773.Protocol version November 11, 2020 (Version #5).
Background The MINDMAP project implemented a multinational data infrastructure to investigate the direct and interactive effects of urban environments and individual determinants of mental well-being and cognitive function in ageing populations. Using a rigorous process involving multiple teams of experts, longitudinal data from six cohort studies were harmonised to serve MINDMAP objectives. This article documents the retrospective data harmonisation process achieved based on the Maelstrom Research approach and provides a descriptive analysis of the harmonised data generated. Methods A list of core variables (the DataSchema) to be generated across cohorts was first defined, and the potential for cohort-specific data sets to generate the DataSchema variables was assessed. Where relevant, algorithms were developed to process cohort-specific data into DataSchema format, and information to be provided to data users was documented. Procedures and harmonisation decisions were thoroughly documented. Results The MINDMAP DataSchema (v2.0, April 2020) comprised a total of 2841 variables (993 on individual determinants and outcomes, 1848 on environmental exposures) distributed across up to seven data collection events. The harmonised data set included 220 621 participants from six cohorts (10 subpopulations). Harmonisation potential, participant distributions and missing values varied across data sets and variable domains. Conclusion The MINDMAP project implemented a collaborative and transparent process to generate a rich integrated data set for research in ageing, mental well-being and the urban environment. The harmonised data set supports a range of research activities and will continue to be updated to serve ongoing and future MINDMAP research needs.
Cohort Profile: Research Advancement through Cohort Cataloguing and Harmonization (ReACH) Julie Bergeron,* Rachel Massicotte, Stephanie Atkinson, Alan Bocking, William Fraser and Isabel Fortier, on behalf of the ReACH member cohorts’ principal investigators Child Health and Human Development, Research Institute of the McGill University Health Centre, Montreal, Canada, Department of Pediatrics, McMaster University, Hamilton, Canada, Department of Obstetrics and Gynecology, University of Toronto, Toronto, Canada and Department of Obstetrics and Gynecology, Université de Sherbrooke, Sherbrooke, Canada
The Developmental Origins of Health and Disease (DOHaD) framework aims to understand how environmental exposures in early life shape lifecycle health. Our understanding and the ability to prevent poor health outcomes and enrich for resiliency remain limited, in part, because exposure-outcome relationships are complex and poorly defined. We, therefore, aimed to determine the major DOHaD risk and resilience factors. A systematic approach with a 3-level screening process was used to conduct our Rapid Evidence Review following the established guidelines. Scientific databases using DOHaD-related keywords were searched to capture articles between January 1, 2009 and April 19, 2019. A final total of 56 systematic reviews/meta-analyses were obtained. Studies were categorized into domains based on primary exposures and outcomes investigated. Primary summary statistics and extracted data from the studies are presented in Graphical Overview for Evidence Reviews diagrams. There was substantial heterogeneity within and between studies. While global trends showed an increase in DOHaD publications over the last decade, the majority of data reported were from high-income countries. Articles were categorized under six exposure domains: Early Life Nutrition, Maternal/Paternal Health, Maternal/Paternal Psychological Exposure, Toxicants/Environment, Social Determinants, and Others. Studies examining social determinants of health and paternal influences were underrepresented. Only 23% of the articles explored resiliency factors. We synthesized major evidence on relationships between early life exposures and developmental and health outcomes, identifying risk and resiliency factors that influence later life health. Our findings provide insight into important trends and gaps in knowledge within many exposures and outcome domains.
Ambient air pollution increases the risk of respiratory mortality, but evidence for impacts on lung function and chronic obstructive pulmonary disease (COPD) is less well established. The aim was to evaluate whether ambient air pollution is associated with lung function and COPD, and explore potential vulnerability factors. We used UK Biobank data on 303 887 individuals aged 40-69 years, with complete covariate data and valid lung function measures. Cross-sectional analyses examined associations of land use regression-based estimates of particulate matter (particles with a 50% cut-off aerodynamic diameter of 2.5 and 10 mu m: PM2.5 and PM10, respectively; and coarse particles with diameter between 2.5 mu m and 10 mu m: PMcoarse) and nitrogen dioxide (NO2) concentrations with forced expiratory volume in 1 s (FEV1), forced vital capacity (FVC), the FEV1/FVC ratio and COPD (FEV1/FVC <lower limit of normal). Effect modification was investigated for sex, age, obesity, smoking status, household income, asthma status and occupations previously linked to COPD. Higher exposures to each pollutant were significantly associated with lower lung function. A 5 mu g.m(-3) increase in PM2.5 concentration was associated with lower FEV1 (-83.13 mL, -95% CI -92.50- -73.75 mL) and FVC (-62.62 mL, 95% CI -73.91--51.32 mL). COPD prevalence was associated with higher concentrations of PM2.5 (OR 1.52, 95% CI 1.42-1.62, per 5 mu g.m(-3)), PM10 (OR 1.08, 95% CI 1.00-1.16, per 5 mu g.m(-3)) and NO2 (OR 1.12, 95% CI 1.10-1.14, per 10 mu g.m(-3)), but not with PMcoarse. Stronger lung function associations were seen for males, individuals from lower income households, and "at-risk" occupations, and higher COPD associations were seen for obese, lower income, and non-asthmatic participants. Ambient air pollution was associated with lower lung function and increased COPD prevalence in this large study.
Combining data from different studies has a long tradition within the scientific community. It requires that the same information is collected from each study to be able to pool individual data. When studies have implemented different methods or used different instruments (e.g., questionnaires) for measuring the same characteristics or constructs, the observed variables need to be harmonized in some way to obtain equivalent content information across studies. This paper formulates the main concepts for harmonizing test scores from different observational studies in terms of latent variable models. The concepts are formulated in terms of calibration, invariance, and exchangeability. Although similar ideas are present in measurement reliability and test equating, harmonization is different from measurement invariance and generalizes test equating. In addition, if a test score needs to be transformed to another test score, harmonization of variables is only possible under specific conditions. Observed test scores that connect all of the different studies, are necessary to be able to test the underlying assumptions of harmonization. The concepts of harmonization are illustrated on multiple memory test scores from three different Canadian studies.