
Fragmentation across early care and education (ECE) systems in the United States obscures how many children access services and how these experiences shape school readiness. This study leverages Iowa's Integrated Data System for Decision-Making (I2D2), which links health, social determinants of health (SDOH), and education records, to provide one of the first unduplicated counts of ECE participation statewide. Drawing on a matched cohort of 27,321 children eligible for kindergarten in 2017--2018, we examine three aims: (1) describe patterns of centre-based ECE participation; (2) assess associations between child/family characteristics and participation; and (3) evaluate links between ECE and kindergarten suspensions and attendance. Results indicate that 73\% of children participated in at least one ECE program prior to school entry, with substantial overlap across public and private preschool, and subsidised or unsubsidised child care. ECE participation varied as a function of poverty, race/ethnicity, maternal education, and cumulative early-life risks. ECE participation was associated with better kindergarten attendance but not suspension. Findings underscore both the promise of integrated data for advancing equity-driven research and the need for policies that target families with the greatest barriers to ECE access.
Introduction:Adult social care in the UK faces increasing demand and persistent inequalities in terms of access and care quality, yet national-level understanding remains limited. The CARE Lab study aims to address these gaps using newly available individual-level routine administrative data for the whole of Wales from the Adults Receiving Care and Support (ARCS) census. This study will explore patterns of care provision, transitions from children's to adult services, and socio-demographic disparities, using linked data to inform service planning and policy. Methods:and analysis This quantitatively-led mixed-methods study comprises five research questions. Quantitative analysis will use ARCS census data, both standalone and linked to health, education, and social care datasets within the Secure Anonymised Information Linkage (SAIL) Databank. Qualitative interviews with people receiving care and support, carers, and professionals will contextualise findings. Key research questions address care patterns, demographic comparisons, regional variation, transitions from child to adult care, and the feasibility of evaluating care models using linked data. Statistical analyses will include descriptive and inferential statistics, propensity score matching, and there will be thematic analysis of qualitative data. Ethics:Ethical approval has been obtained from Cardiff University. Data access approvals will be sought from Welsh Government, SAIL, and the Office for National Statistics. Dissemination will occur through peer-reviewed publications, policy briefings, accessible multimedia outputs, and stakeholder engagement via an action group. The study will also produce a research-ready data asset and recommendations for future data infrastructure development across the UK.
Inclusion, Diversity, Equity, and Accessibility (IDEA) are increasingly recognised as essential to advancing population health research and addressing structural inequities. Yet, few publications describe how to develop IDEA strategies, leaving organisations with limited guidance on replicable processes. Here, Health Data Research Network Canada (HDRN Canada) details the steps it took to establish its own IDEA Strategy. The strategy was developed through an iterative five-phase process. Key steps included creating a project charter, defining shared governance and consensus-based decision making, and using professional facilitation to foster broad participation across member organisations. Visible executive sponsorship was critical to strategy development and implementation. The resulting strategy identifies four interconnected action areas: Learning and Unlearning, Facilitating IDEA in Research, Cultivating Trust and Reciprocity, and Providing Leadership and Advocacy to embed IDEA in organisational operations and research practices. HDRN Canada's experience demonstrates how a national distributed research network of organisations that work together to support multi-jurisdictional research can use best practices to transform IDEA principles into a concrete, actionable framework. This work offers a transferable model for population health research organisations seeking to integrate IDEA within their organisations and across the research ecosystem.
Introduction & BackgroundLiving conditions vary widely across England, and several indices have been developed to quantify spatial inequalities and support evidence-based decision-making. The Index of Multiple Deprivation (IMD) is an established measure of relative area-level deprivation, combining administrative data across seven domains including income, employment and education. The Priority Places for Food Index (PPFI) is a more recent index that also incorporates seven domains but focuses on food-related vulnerability and access. Although both indices aim to identify areas of need based on multiple domains, they capture distinct dimensions of disadvantage. As newer metrics like PPFI gain influence in policy and research, there is a growing need for approaches that enable clear comparison across indices. Objectives & ApproachThis project introduces an interactive visualisation tool designed to support systematic comparison between the newly released IMD 2025 and PPFI across England. The tool enables users to identify areas where the two indices align or differ and examine how these patterns vary by geography across domains. Developed using Python, the dashboard provides interactive choropleth maps at Lower Super Output Area (LSOA) and Local Authority District (LAD) level. Users can explore individual domains, conduct side-by-side comparisons and examine ranked mismatches using a dedicated explorer. Consistent visual design and scale-appropriate representations ensure interpretability and comparability. Relevance to Digital FootprintsPPFI incorporates indicators derived from administrative, web-scraped and crowdsourced data on spatial food access and service availability, reflecting digital footprints of everyday life. By comparing PPFI with IMD, the tool highlights how food-related and access-based vulnerabilities can exist even in areas not identified as highly deprived by traditional administrative measures alone. For example, Druridge Bay scores in the top 10% for food vulnerability under PPFI, yet falls outside the most deprived areas under IMD, a discrepancy that would be invisible without comparison. This demonstrates the value of integrating digital footprint data with established indices to provide a more nuanced understanding of local disadvantage. Conclusions & ImplicationsThis tool demonstrates that while IMD and PPFI frequently overlap, approximately 50% of high-priority areas differ depending on which index is used as they reflect different forms of disadvantage. It enables transparent, accessible comparison between the two indices, supporting more informed selection based on user needs and context. It is designed for use by policymakers, local authorities, food banks, and third-sector organisations to identify and prioritise areas of need without specialist data expertise. Integrating digital footprint data with established indices can strengthen evidence-based decision-making, ensuring hidden vulnerabilities do not go unaddressed.
Introduction & BackgroundCoffee shops function as contemporary third places where consumption, repeated presence, and informal social contact overlap. Their relevance extends beyond the service encounter, supporting everyday interactions and urban social value. Visitors experience these places through spatial, social, and symbolic qualities, which makes the coffee-shop servicescape central to third-place research. Prior research treats this servicescape as multidimensional, but it remains unclear which dimensions are associated with specific customer responses at scale. Objectives & ApproachThis study examines how servicescape dimensions appear in expressed third-place experience in New York City coffee-shop reviews. It uses the Stimulus-Organism-Response framework and 105,021 Google Maps reviews to connect environmental cues, expressed affective evaluation, and customer responses. BERTopic extracts review-derived experience aspects, which are manually classified into physical, social, socially symbolic, and restorative servicescape dimensions. Product, price, and operational aspects remain in the model as controls. Sentiment toward these cues represents the organism component as an expressed affective evaluation. Responses are measured through star ratings and review-based textual signals of recommendation and revisit intention. Regression models estimate how positive and negative sentiment toward each servicescape dimension is associated with these outcomes, controlling for other stimuli. Relevance to Digital FootprintsGoogle Maps reviews are used as digital footprints because they contain cues, evaluations, and response signals within the same user-generated text. ResultsThe results show that servicescape dimensions relate to customer responses in different ways. Positive social and socially symbolic sentiment is most closely associated with ratings. Recommendation sentiment is more strongly linked to social and restorative cues, while revisit sentiment is linked to physical and socially symbolic cues. Negative physical sentiment shows the most consistent decline across all three outcomes. Conclusions & ImplicationsThe study contributes to a digital-footprint operationalisation of the S-O-R framework for third-place research. It shows that customer responses are linked to specific servicescape dimensions, not to the coffee-shop environment as a whole. These findings can inform the design of coffee shops as commercial venues with social value in urban life. The results are limited by self-selected, platform-mediated Google Maps reviews, possible inauthentic or bot-generated content, and the focus on New York City coffee shops.
Introduction & BackgroundEnergy poverty affects nearly 20% of UK households, carrying substantial economic, health, and social costs. Yet, the metrics policymakers rely on most, such as the Low Income Low Energy Efficiency (LILEE) indicator, reduce a deeply lived experience to a static income threshold. They cannot capture what happens inside a household week to week: turning off the heating to afford food, letting a prepayment meter run out rather than accumulate debt, or enduring cold and damp conditions because no affordable alternative exists. These are behavioural trade-offs, and they are precisely what existing frameworks fail to see. Thus, there is a critical need to integrate high-resolution behavioural data with social frameworks to address this multifaceted challenge effectively. Objectives & ApproachWe operationalise a multi-stakeholder perspective on energy poverty by analysing longitudinal data from the Smart Energy Research Lab (SERL). Our methodological innovation is the Trade-off Proxy Variable (TPV): a composite behavioural indicator constructed from three flag signals directly observable in consumption records. The first, energy rationing, identifies households whose energy consumption significantly falls below a baseline adjusted for weather and property characteristics. This suggests that these households are intentionally restricting their energy use rather than simply being more efficient. The second, the inefficiency-consumption mismatch, flags households in poorly rated properties that display unusually low energy consumption. This indicates suppressed demand, as residents are unable to afford adequate heating in homes that are structurally expensive to heat. The third, extreme heating self-disconnection, captures instances of zero gas supply during freezing weather, highlighting households without heating during the coldest periods. A household is classified as acutely vulnerable when two or more signals co-occur. A competitive supervised machine learning framework, interpreted through SHapley Additive exPlanations (SHAP), is then applied to identify the structural drivers of these trade-offs. Relevance to Digital FootprintsBy moving away from static metrics, this research proposes that high-frequency energy logs serve as real-time digital proxies for financial vulnerability. AI modelling applied to these transactional records can uncover hidden patterns of self-disconnection and suppressed demand that demographic surveys routinely miss, offering a novel lens for behavioural data science with direct policy relevance. ResultsAs a primary outcome, we present the engineered TPV architecture, a methodological framework designed to overcome the limitations of subjective survey data and static models. Critically, our technical evaluation of the SERL dataset, informed by existing literature, reveals that conventional economic indicators structurally underrepresent hidden vulnerability, particularly among 'house-rich but cash-poor' demographics who objectively engage in extreme rationing but underreport subjective hardship. This insight informs the behavioural design of the TPV: by grounding vulnerability classification in observable consumption patterns rather than self-reported income, we can highlight household profiles in acute distress that may be overlooked by conventional poverty classifications. At the conference, we will present the complete mathematical formulation of the three behavioural flags, our data-engineering strategy for addressing SERL sample biases, and the design of the machine learning pipeline. In the next phase of the project, we will leverage the extracted SHAP personas to quantify measurable gaps in 'know-what' and 'know-how.' This analysis will lay the empirical groundwork for targeted Marketplace Literacy interventions. Conclusions & ImplicationsTackling energy poverty requires shifting from top-down, income-based definitions to bottom-up, behaviourally grounded insights. The TPV framework, SHAP-driven persona methodology, and Marketplace Literacy pipeline together offer a scalable and replicable blueprint for policymakers and charities, enabling a transition from reactive welfare support to proactive identification of at-risk households before they reach a crisis point.
Introduction & BackgroundSupermarket shopping data provides researchers with a novel source of information regarding health-related behaviours. Women’s health is an underexplored area, including how individuals manage menstrual symptoms. Whilst several tools exist to assess menstrual symptom severity in clinical and research settings, these are outdated and fail to consider relevant management behaviours, which may be traced through digital shopping records. Objectives & ApproachThis study aims to develop a modern survey tool to assess the use of self-care strategies and related shopping habits in managing menstrual symptoms. Pilot survey findings were used to identify relevant products, management strategies, and purchasing patterns which will guide future research harnessing shopping data. In October 2025, 150 female participants aged 18-55 completed an online survey. Recruitment was managed through Prolific with pre-set screening criteria to target participants who were paid for their participation. A range of statistical methods were deployed in data analysis. Relevance to Digital FootprintsThis study contributes to improving our understanding of applying novel data linkage to study an important public health issue. It provides important insights into how purchasing behaviours reflect management strategies, laying the laying the groundwork for using shopping data to track management of menstrual symptoms. ResultsThe mean participant age was 37 (SD = 8.77), most (80.00%) were in paid employment and 99.33% identified as cisgender. Use of self-care strategies was common with 72.67% of participants reporting this. The most common strategies for managing pain were paracetamol (68.00%), ibuprofen (53.33%), hot water bottle (50.00%) and herbal tea (26.67%). Most participants reported purchasing period products for themselves only (81.33%) or for themselves and another person (16.67%). Supermarkets were cited as the main location where participants purchased period products (83.33%) and pain relief for period pain (71.33%). Most participants reported purchasing period products regularly with 52.67% purchasing monthly or more frequently and 36.00% purchasing every 2-3 months. The mean reported monthly spending on managing menstruation was £11.49±£8.60. Such self-reported behaviours can indicate likely shopping patterns, helping to track and study management strategies. Conclusions & ImplicationsOur findings help to improve our understanding of how individuals self-manage menstrual symptoms, as well as inform approaches to future shopping data research to study female health-related behaviours at a national level.
Introduction & BackgroundPerfume purchasing is a highly multisensory process, relying on olfactory, tactile and visual cues. Although online retail already restricts direct olfactory engagement, the COVID-19 pandemic heightened these constraints by temporarily removing in-store sampling and therefore limiting physical interaction with products. Despite the growing interest in sensory marketing, little is known about how consumers' real-world purchasing behaviour changes when key sensory inputs are restricted. Objectives & ApproachThis research examines how consumers’ perfume purchasing behaviour shifts in limited multisensory retail environments by adopting a sensory deprivation perspective. Specifically, it investigates whether there are changes in purchasing behaviour across pre-COVID, during-COVID and post-COVID periods, including potential shifts in reliance on brand, price, product descriptions and visual cues. Relevance to Digital FootprintsThis research contributes to digital footprint literature by showing how purchasing behaviour may change under sensory restrictions. By analysing behavioural patterns captured in large-scale transactional data alongside product-level sensory cues, the study explores whether consumers adapt their purchase decisions when direct sensory evaluation is restricted. Overall, the study highlights the value of digital footprint data in examining purchasing behaviour in sensory-restricted retail environments. Conclusions & ImplicationsConceptually, this study advances sensory marketing research by shifting the focus from sensory optimisation to sensory constraint, positioning sensory deprivation as a meaningful condition that shapes consumer behavioural patterns. Empirically, it offers preliminary insights into how consumers’ perfume purchasing behaviour may shift when multisensory access is disrupted. The research demonstrates how natural disruptions can be used to deepen understanding of consumer psychology in digital and physical retail marketplaces.
Introduction & BackgroundResearch can inform strategies to improve population diets by understanding how to harness digital food retail environments to influence consumer purchase choices. This project aimed to use data, including eye tracking and purchases, from a pilot study to better understand “how” people shop within online supermarkets to buy food in real-life. Objectives & ApproachTo develop an interactive dashboard to visualise and explore eye-tracking, purchase and participant data from individual shoppers during their real-life online supermarket shops. Data was collected from ten participants of varying backgrounds, and prior online shopping experience, while they performed and paid for their own food shops (and home delivery) using the Ocado.com website on a University Tobii Pro desktop eye-tracker computer. Receipt data collected included cost, number and nature of all foods purchased. Analysis of eye-tracking data included classification of supermarket website navigation, and types of products (i.e. Fresh etc.) and webpages (i.e. product listings, basket, extra information pages etc.) viewed. Participants’ visual engagement with individual webpages was quantified using eye gaze (fixations) and movements (saccades) over time. Data from eye-tracking and purchases were aggregated and linked with participant characteristics and visualised using a quarto dashboard with the objective of enabling interactive exploration of trends. Relevance to Digital FootprintsWe show how participants’ eye-tracking data can be combined with digital food purchase footprints to characterise real-life shopping trips and explore online decision-making behaviours, including supermarket website and webpage use as well as food choice. ResultsA Quarto dashboard was created which enabled interactive exploration of visualised data on individual participant’s attention (fixation count and duration) across their online shop, and by specific supermarket webpage types. Across shops, participants took between 7-61mins, viewed between 21-99 distinct supermarket webpages and purchased 6-68 products (£42-£119). Webpages, including product-listing pages where most fixations were captured, were accessed by participants via “search” and “navigation” paths, with no clear relationship between types of products viewed and purchased. The dashboard allows users’ further exploration of the data interactively, filtering by participant, webpage, and product types. Conclusions & ImplicationsOur approach, data, and dashboard offer new insights into real-world consumers’ online food purchasing behaviours to inform future research which can underpin supermarket website design. Future work includes scale-up to increase sample size, and quantification of participants’ attention to specific within-webpage features such as pictures. Together these enable evaluation of marketing and product presentation within supermarket webpages to support sustainability and public health.
Introduction & BackgroundThe current approach to measuring fuel poverty in England, Low Income Low Energy Efficiency (LILEE), systematically underestimates the condition in low- and middle-income homes. This is due to its inability to accurately respond to macroeconomic shocks (e.g., energy price inflation) and its overstatement of energy inefficiency as the principal driver of fuel poverty. This inaccuracy poses a critical policy challenge: inaccurate measurement precludes meaningful alleviation. In response, the current study develops an area-level fuel poverty indicator that better captures the realities of fuel poverty by harnessing digital footprints: smart meter data and financial records, acquired from the Smart Energy Data Service (SENSE) and the Financial Data Service (FINDS), respectively. Objectives & ApproachThe aim is to develop a granular, area-level fuel poverty indicator (based on Lower Super Output Areas—LSOAs) through the linkage of smart meter and financial data. This approach facilitates the examination of the spatial and temporal overlap between low absolute energy consumption (SENSE) and low energy expenditure (FINDS). The proposed approach aims to complement or supplant the government’s official fuel poverty statistics, which are typically modelled using relatively small, static survey samples, by providing dynamic and responsive insights into energy affordability at regular intervals (e.g., monthly or seasonal). Relevance to Digital FootprintsEnergy and financial data represent rich digital footprints of sensitive social factors, which are key to a better understanding of fuel poverty in the UK. The SENSE and FINDS data are disparate digital footprints encoding complex consumer behaviours; yet their linkage could possess significant utility as an alternative resource for generating area-level fuel poverty insights. Conclusions & ImplicationsThe SENSE and FINDS collaboration is ongoing; we present a methodological framework for linking two independent datasets and new analytical approaches enabled by the linkage of energy and financial data. The expected outcomes are more accurate and responsive area-level fuel poverty indicators. This is particularly pertinent to England, where competing fuel poverty statistics are needed to overcome the weaknesses of the LILEE approach and to enable targeted measures for a fair energy transition. Further, the proposed indicator could harmonise diverging measurement standards across England and the devolved nations (Scotland, Wales and Northern Ireland), facilitating essential comparative statistics.
In the U.S., funding and staffing structures for data sharing and integration vary widely across state and local governments. Since 2009, we have conducted a biennial survey of a large network of integrated data systems (IDS), including questions to deepen our understanding of how IDS develop and are governed, funded, and staffed over time to provide guidance for practitioners and policymakers. This session presents findings from the 2023 and 2025 Network Surveys, as well as 30 qualitative interviews conducted in 2024. These data provide insights and guidance for the field regarding budgets, funding streams, staffing, and capacity development. Findings reveal three broad categories of IDS. Efforts with the smallest budgets have a local and domain specific focus; medium efforts are county or state efforts based at universities; and large efforts tend to be county or state efforts operated within government. Funding comes from a range of sources-federal, local, state, philanthropic, and fee for service. Sites spend the majority of their budgets on staffing and little, comparatively, on technology and infrastructure. Regardless of budget and funding mechanisms, sites spend more money on personnel than any other expense. It costs around $350,000 for the data system itself, including tools for storage and linkage. The major distinction between sites lies in their capacity to afford personnel to do research and design data products internally. These findings fill a critical gap in understanding what resources are needed to stand up IDS in diverse contexts to develop actionable insights for local policymakers.
Introduction While Canada is rich in databases useful to support healthcare research, they are widely distributed, often poorly documented, and it is challenging to identify relevant databases, apply for access, and eventually use, link or harmonise the data. Even if the databases needed to address specific questions are known, it is difficult and time-consuming to find the metadata, the ``data about the data'' required to understand the characteristics and data content of these resources. A solution to these challenges is creation of metadata catalogues, which detail metadata for multiple databases, not the actual data. Objectives Describe a new catalogue including metadata about Canadian medical and non-medical databases' characteristics and variables, and information to assist catalogue users in seeking data access. Methods Starting with a list of 385 national, provincial and regional databases, a group of physician-investigators, epidemiologists, data scientists and patient partners prioritised databases for inclusion. Metadata cataloguing occurred in steps: (i) description of the database with listing of its characteristics, and when available, (ii) addition of information about collected variables. Results 83 individual databases are documented in the Metadata Catalogue of the Sepsis Canada Network (https://www.maelstrom-research.org/network/sepsis). 57 are registries, 13 are cohort and 13 cross-sectional databases. 16 cover all of Canada, while another 13 cover most of the country; 45 focus on a single province. For 33 databases (38\%) the catalogue includes detailed information about variables collected. Conclusions This metadata catalogue includes databases collecting information spanning the continuum of medical care, non-medical data, and determinants of health. It is freely available online and extensively searchable. It can facilitate implementation of a wide range of research initiatives into medical conditions, medical care, and outcomes.
IntroductionLinked administrative data integrating health and non-health information can support population-based research about biological and contextual environmental factors that influence child health. Database linkage studies leverage existing data to provide more comprehensive information than would be available from any single source. However, it is unknown the extent by which child health studies capitalise on linked multi-domain Canadian administrative data. ObjectiveThis scoping review aims to describe Canadian population-based child health studies that used linked multi-domain (i.e., health and non-health) administrative data. MethodsA systematic search was conducted of MEDLINE, Embase, Scopus and Global Health from inception until March 12, 2025. Articles were included if they focused on children (birth to 18 years), used Canadian administrative data, and linked health with non-health data. Two reviewers independently screened titles/abstracts and full texts; a pilot test ensured consistency. Article characteristics, province/territory, parental linkage, and non-health variables, were collected using an extraction form. ResultsThe search yielded 4,437 articles, of which 42 met inclusion criteria. Most articles were conducted in Manitoba (45%) and Ontario (36%). Maternal linkage was common, whereas paternal linkage was limited to Manitoba and British Columbia. Immigration status was the most common non-health variable. Health service use, particularly preventive care, such as screening and vaccination coverage, was a common research theme. No multi-jurisdictional studies were identified. ConclusionsMulti-domain administrative data linkage studies remain concentrated in a few provinces. Expanding parental linkage, integrating non-health variables, and strengthening multi-jurisdictional studies are crucial for improving population-based understanding of child health influences across Canada
The integration of administrative data for population-based studies and research faces many challenges, e.g., integrity, quality, privacy, security, and availability of this data. The governance of integrated administrative data for research purposes is context-dependent and involves the articulation of methodological, technical, ethical, legal, and social issues. Brazil, the largest country in Latin America, has an estimated population size of approximately 213 million inhabitants. Despite the country’s diverse array of national information systems to support public administration, among them health data from the universal health system, the Sistema Único de Saúde (SUS), these systems are presented and made available in an isolated manner. Present and discuss CIDACS’ approach to provide requirements and recommendations for a national health data policy for scientific research and studies aimed at producing knowledge to support evidence-based public health in Brazil. Reporting experience on a collaboration between the Center for Data and Knowledge Integration for Health (CIDACS/Fiocruz Bahia) and the Information Technology Department of the Unified Health System (DATASUS) to situate proposed strategies to support requirements and recommendations for safe and secure access to, linkage and analysis of existing health data produced by the SUS. Results and Considerations We delineate an approach encompassing a literature review on technical, ethical, legal, and societal issues related to health data usage and reuse for scientific and public health research purposes, articulated with the mapping and analysis of the national and international regulatory landscape, legal consultancy, technical visits, as well as interviews and workshops with data stakeholders.
We started a family-based genetic epidemiology study in 2006-11 which recruited 24,000 adult volunteers from 7000 families across Scotland with consent for follow-up through medical record linkage and re-contact. In 2022-25 we have recruited a further 16,000 volunteers, with consent extended to administrative records, and age range now 12+. Original volunteers completed demographic, health and lifestyle questionnaires, provided biological samples, and underwent detailed clinical assessment. The samples, phenotype and genotype data form a resource for research on the genetics of conditions of public health importance. This has become a longitudinal dataset by linkage to routine NHS records: hospital, maternity, lab test, prescriptions, dentistry, mortality, imaging, cancer screening, GP data, Covid-19 testing and vaccinations, as well as follow-up questionnaires. The new wave of recruitment is all online with DNA from saliva collected by post. Teenagers aged 12-15 can join with parental consent. Researchers can find prevalent and incident disease cases and controls to test research hypotheses on a stratified population. They can also do targeted recruitment of participants to new studies, including recall by genotype. We have established and validated E-HR linkage with the NHS Scotland CHI Register, overcoming technical and governance issues in the process. We contribute to major international consortia, with collaborators from institutions worldwide, both academic and commercial. The Research Tissue Bank resources are available to academic and commercial researchers through a managed access process.
Background and objectives Timely monitoring of cancer cases is important for identifying emerging patterns, targeting and evaluating prevention strategies, and supporting resource allocation. There is a demonstrated lack of timely tracking of Métis-specific cancer incidence and prevalence, despite studies suggesting a higher incidence of cancer, lower rate of cancer screening uptake, and higher prevalence of modifiable risk factors for Métis people across Canada. The Métis are one of three constitutionally recognized Indigenous Peoples in Canada and are represented by the Métis Nation of Ontario (MNO) in the province of Ontario. This study examines cancer incidence and prevalence in MNO citizens across 9 types of cancers. Methods The MNO holds a registry of citizens that is shared annually with ICES. MNO Registry data was deterministically linked to the Ontario Cancer Registry (OCR) which has information about all Ontario residents diagnosed with cancer. The cancers examined included breast, cervix, colon, gallbladder, larynx, liver, lung, prostate, and skin. Cancer case counts between 2019 and 2024 were used to calculate incidence and prevalence. Results Analyses are ongoing. In total, 29,997 of the 30,885 registered MNO citizens were linked with the OCR data. There were 365 MNO citizens (52% male) diagnosed with one of the nine cancers between 2019 and 2024. Colon, breast, lung, and colon cancers were most common types of cancers. Implications Ongoing and timely monitoring of Métis-specific cancer rates in Ontario is important to inform programs and services targeted to MNO citizens for their health and wellness.
To advance First Nations (FN) data sovereignty in Manitoba by developing and implementing FN-specific training initiatives that build community capacity in epidemiology, biostatistics, data analysis, and data utilization. The First Nations Health and Social Secretariat of Manitoba (FNHSSM) established a regional training institute, youth training programs, and a mobile training lab designed to support FN leaders and community members in developing data skills grounded in FN methods and analytic frameworks. All training is FN-specific, recognizing the distinct policy and data contexts of First Nations in Canada. Development and facilitation have been guided by FNHSSM’s Grandmother’s Circle, with curriculum tailored through collaboration with Knowledge Keepers and local guest speakers to ensure cultural relevance and community alignment. In 2026, the mobile training lab will be piloted in FN communities to further enhance access and adaptability. These training initiatives strengthen community-level capacity, reduce reliance on external entities to conduct FN analyses, and support communities in asserting control over their data through the use of FN tools and frameworks. Equipped with these skills, trained individuals can participate meaningfully in data interpretation and decision-making and collaborate with external partners on their own terms. By building robust, community-driven data capacity, FNHSSM’s training initiatives promote FN data sovereignty and support communities in accessing and applying their own data to inform resource planning, respond to outbreaks and pandemics, and advocate for social justice where data already exists. This work reflects a transformative pathway for FN self-determination in data governance and public health.
The Scottish Longitudinal Study (SLS) is a largescale research ready record-linkage study created and supported by the SLS Development and Support Unit (SLS-DSU). The SLS is accessed via the SLS-DSU safe haven/trusted research environment (TRE). The SLS links Census through time 1991-2022 to administrative data on major life events, maps changing residential location and for children, their progress through the educational system. This paper will introduce the SLS as a data resource for researchers, the datasets held as part of it, the application process, the SLS-DSU, along with two innovations developed for making research using the SLS, and its sister studies, easer to do within the separate TREs. Census data are the building blocks of the SLS from 1991 onwards, for a 5% representative sample of the Scottish population (about 270,000 sample members each Census). The SLS links together a wealth of information from routinely collected administrative data, including vital events registrations (births, deaths and marriages), migration data, Scottish education data, and with appropriate additional permissions can be linked to NHS health data including cancer registry and hospital admission data. With the SLS-DSU two innovations were developed SYNTHPOP (an R package for creating bespoke synthetic project extracts) and eDataShield (an R package for doing combined federated type analysis with the SLS and one of the sister studies ONS LS for England and Wales, or the Norther Ireland LS). An overview of SYNTHPOP and eDataShield will be provided to demonstrate how they make working within TREs easier, given data restrictions.
UK Longitudinal Linkage Collaboration (UKLLC) is the national Trusted Research Environment (TRE) for record linkage in longitudinal research, partnering with SeRP and >20 Longitudinal Population Studies (LPS). Participating in LPS is rewarding, with many participants enrolling into multiple studies (e.g., ∼8% of ALSPAC mothers are also enrolled into UK Biobank). It's scientifically important to account for this in pooled and meta LPS analysis as most statistical assessments assume independence of sample membership. Currently, LPS records are treated separately, leading to potential duplication and over-counting of exposures/outcomes in research, where individuals participate in multiple cohorts. We used probabilistic record linkage to identify individuals across multiple cohorts. We engaged study Data Managers through a consensus-building workshop to reconcile governance issues. Of ∼570,000 participants from 22 partner LPS in UKLLC, we've currently identified 4,785 individuals in two cohorts, 155 in three, and <10 in four. These numbers are likely to increase as UKLLC scales to support larger studies (with an anticipated 2m participants hosted by 2027). Mappings of individuals belonging to multiple cohorts were delivered to end-users via our data provisioning pipeline in a manner compatible with the dynamic nature of the pooled UK LLC hosted sample. This linkage will support accurate pooled- and meta-analysis within UKLLC. Ensuring the governance challenges are accounted for is essential for the acceptability of our community LPS governance framework. Extending this linkage of individuals across different LPS and platforms (e.g., to UK Biobank) will be necessary for robust federated analysis across TREs.
This study examines the relationship between household financial circumstances and children's social care (CSC) involvement using newly linked local administrative data. Household benefits data (SHBE and UCDS) were securely linked and anonymised with Children in Need (CIN) records from six local authorities in London and South England, covering 2019-2021. Financial precarity is defined as households living below the relative poverty line or experiencing a cash shortfall. Children living in financial precarity were not more likely to be initially referred to CSC. However, once referred, they were significantly more likely to experience higher-intensity statutory interventions, including having a Child Protection Plan (12% vs 9%) and being re-referred (32% vs 29%). We estimate that an additional 270 Child Protection Plans were made over the study period among children referred from households below the poverty line. The 2020-21 Universal Credit (UC) uplift provides a natural experiment to examine the role of income support. Households receiving the uplift were substantially less likely to be in financial precarity, with a 17.5 percentage-point relative improvement. Children in uplift-eligible households were more likely to be referred to CSC but less likely to receive subsequent statutory protective interventions, suggesting that improved financial stability reduces the need for high-intensity CSC involvement. These findings demonstrate the value of ethically governed administrative data linkage for evaluating social policy. They highlight the importance of policies that strengthen families' financial circumstances as part of a preventative approach to improving child wellbeing and reducing statutory demand.