Introduction: Adolescents and young adults (AYAs) with cancer constitute a distinct population with unique biological and survivorship challenges. Long-term, population-based analyses of incidence and mortality trends in AYAs remain limited. Methods: We conducted a retrospective, longitudinal population-based cohort study using data from the Childhood, Adolescent, and Young Adult Cancer Survivor program in British Columbia, Canada, from 1971 to 2020. We analyzed age-standardized incidence rates (ASIR) for the 16 cancer types with the highest incidence and age-standardized mortality rates (ASMR) for the 8 cancer types with the highest mortality, overall and by sex. Trends were evaluated using annual percentage change (APC) and average APC from joinpoint regression analyses. Results: The cohort included 43,588 patients with a first cancer diagnosis and 6825 deaths. From 1971 to 2020, cancer incidence rose overall (AAPC 0.1; 95% confidence interval [CI]: −1.0, 1.3), with a brief decline in the mid-1990s. Significant increases in ASIR were observed for thyroid (APC 1983–2020 : 3.0, 95% CI: 2.4, 3.6) and gastrointestinal cancers (APC 1988–2020 : 1.9, 95% CI: 1.3, 2.6). Female AYAs showed rising trends in Hodgkin and non-Hodgkin lymphoma, while testicular cancer incidence increased among males. Male AYAs had higher all-cause mortality (51.6%). Cancer mortality declined beginning in the early 1990s; however, gastrointestinal cancer mortality increased in males (APC 1990–2020 : 1.7, 95% CI: 0.3, 3.1). Conclusions: Cancer incidence among AYAs has risen over five decades, particularly for thyroid and gastrointestinal cancers, with concerning increases in gastrointestinal cancer mortality among males. These findings emphasize the need for targeted prevention, screening, and intervention strategies.
In attempt to advance the current practice for assessing and predicting the primary ovarian insufficiency (POI) risk in female childhood cancer survivors, we propose two estimating function based approaches for age-specific logistic regression. Both approaches adapt the inverse probability of censoring weighting (IPCW) strategy and yield consistent estimators with asymptotic normality. The first approach modifies the IPCW weights used by Im et al. (2023) to account for doubly censoring. The second approach extends the outcome weighted IPCW approach to use the information of the subjects censored before the analysis time. We consider variance estimation for the estimators and explore by simulation the two approaches implemented in the situations where the conditional right-censoring time distribution required in the IPCW weighs is unknown and approximated using the survival random forest approaches, stratified empirical distribution functions, or the estimator under the Cox proportional hazards model. The numerical studies indicate that the second approach is more efficient when right-censoring is relatively heavy, whereas the first approach is preferable when the right-censoring is light. We also observe that the performance of the two approaches heavily relies on the estimation of censoring distribution in our simulation settings. The POI data from a childhood cancer survivor study are employed throughout the paper for motivation and illustration. Our data analysis provides new insight into understanding the POI risk among cancer survivors.
Administrative health data contain rich information for investigating public health issues; however, many restrictions and regulations apply to their use. Such data are usually not in the conventional format for statistical analysis since the databases are created and maintained to serve non-research purposes and only information for people who seek health services is recorded. Analysis of administrative health data is thus challenging in general. We aim to develop a tool for understanding how the mental health of youth aged younger than 18 years evolves over time through administrative records of mental health related emergency department (MHED) visits. The MHED records, a set of zero-truncated recurrent events data, are integrated with relevant population census information, and framed into a set of doubly censored recurrent events data. We present innovative strategies for overcoming the zero-truncation induced by the data collection and for processing doubly censored recurrent events data. We compare the dynamic patterns and exposure impacts of the MHED visits in two decades with a loosely structured model. The findings are verified empirically via simulation. The asymptotic properties of the proposed estimator are established. Through exploring the paediatric MHED visit records, we provide new insights into children/youths mental health changes over time.
Wastewater and environmental monitoring (WEM) was a critical public health surveillance tool for SARS-CoV-2 surveillance during the COVID-19 Pandemic. Yet, persistent methodological heterogeneity across laboratories continues to challenge the interpretation and compromises the actionability of WEM measurements. This study quantifies interlaboratory concordance in SARS-CoV-2 WEM measurements using influent wastewater samples (Sep 2021-Jan 2024) from a single Ontario Wastewater Surveillance Initiative (WSI) facility, independently analyzed by 12 laboratories with routine methods. Lacking known true viral concentrations, interlaboratory measurements were benchmarked against facility-specific longitudinal benchmark derived from the routine surveillance at the Ontario WSI facility. Concordance was assessed across four standard units: SARS-CoV-2 copies/mL, copies/copies of pepper mild mottle virus (PMMoV), and their Wastewater Viral Activity Level (WVAL)-standardized counterparts (WVAL-standardized SARS-CoV-2 copies/mL and copies/copies of PMMoV). Measurements in each unit were analyzed using complementary analytical frameworks, including categorical concordance metrics, principal component analysis, and linear mixed-effects modelling. Interlaboratory measurements consistently captured benchmark temporal dynamics, particularly major peaks and low-activity periods, but exhibited substantial variation in magnitude and public-health interpretation across methods. Concordance was strongest during epidemiological extremes but declined in transitional periods, elevating misclassification risks with potential implications for public health decision-making. Associations between benchmark concordance and the method-specific steps (concentration, extraction, and RT-qPCR) were assessed using Fisher's exact tests, alongside extracted settled-solids mass threshold analyses. No single methodological factor showed a statistically significant concordance association; however, several parameters, including RNA template volume, total RT-qPCR reaction volume, and extracted settled-solids mass, may warrant further investigation.
In the post-pandemic era of COVID-19, hospitalization remains a primary public health concern and wastewater surveillance has become an important tool for monitoring its dynamics at the level of community. However, there is usually no sufficient information to know the infection process that results in both wastewater viral signals and hospital admissions. That key challenge has motived a statistical framework proposed in this paper. We formulate the connection of overtime wastewater viral signals and hospitalization counts through a latent process of infection at the level of individual subject. We provide a strategy for accommodating aggregated data, a typical form of surveillance data. Moreover, we ease the conventional procedure of the statistical learning with the joint modeling using available information on the infection process, which can be under-reporting. A simulation study demonstrates that the proposed approach yields stable inference under different degrees of under-ascertainment. The COVID-19 surveillance data from Ottawa, Canada shows that the framework recovers coherent temporal patterns in infection prevalence and variant-specific hospitalization risk under several reporting assumptions.
Abstract This article presents a strategy for conducting regression analysis of zero‐truncated recurrent event data. The research is partly motivated by a pediatric mental health care (PMHC) program based on administrative data. We are particularly interested in how the occurrence of an event depends on its past occurrences and the associated covariates over time. We propose a stratified Cox regression model with time‐varying coefficients. An easy‐to‐implement procedure is provided for estimating the model parameters using the zero‐truncated data integrated with readily available population census information. We examine the proposed estimator's finite‐sample performance through simulation and establish its asymptotic properties. The PMHC program data are used throughout the article to motivate and illustrate the proposed approach.
This article presents a statistical learning framework for studying the evolution of pediatric mental health-related emergency department (MHED) visit patterns across the pre-, during-, and post-COVID-19 pandemic periods using population-based administrative health records. The MHED records are formulated as zero-truncated recurrent event data, partitioned into three successive time periods. We develop the modeling framework in a stepwise manner, guided by model fit using a collection of MHED records. The resulting framework progresses from nonparametric marginal rate models to more structured Cox-type regression models for characterizing visit patterns. We ultimately apply stratified regression analysis to investigate changes in visit frequencies and covariate effects across pandemic periods, accounting for prespecified period cut-off points and coarsened individual follow-up information. The proposed framework is motivated by and illustrated using pediatric MHED data throughout the article, providing a practical approach for analyzing recurrent healthcare utilization data with evolving temporal patterns.
This paper is motivated by an ongoing pediatric mental health care (PMHC) program in which records of mental health-related emergency department (MHED) visits are extracted from population-based administrative databases. A particular interest of this paper is to understand how the visit occurrence depends on the occurrences in the past in a general population. Only information on subjects experiencing MHED visits is available within a subject-specific time window. Thus, the MHED visits may be viewed as zero-truncated recurrent events. Some population census information can be utilized as supplementary information on the covariates of subjects without MHED visits during the study period. We consider an innovative stratified Cox regression model, which is an intensity-based model but requiring only a summary of the event history. We propose an estimation procedure with zero-truncated data integrated with some supplementary information. We establish the consistency and asymptotic normality of the proposed estimator. The finite-sample properties of the estimator are evaluated by simulation, which demonstrates improved performance of the proposed estimator over the maximum likelihood estimator based on zero-truncated data only. We use the PMHC program to illustrate the proposed approach throughout the paper.
Event history data from sports competitions have recently drawn increasing attention in sports analytics to generate data-driven strategies. Such data often exhibit self-excitation in the event occurrence and dependence within event clusters. The conventional event models based on gap times may struggle to capture those features. In particular, while consecutive events may occur within a short timeframe, the self-excitation effect caused by previous events is often transient and continues for a period of uncertain time. This paper introduces an extended Hawkes process model with random self-excitation duration to formulate the dynamics of event occurrence. We present examples of the proposed model and procedures for estimating the associated model parameters. We employ the collection of the corner kicks in the games of the 2019 regular season of the Chinese Super League to motivate and illustrate the modeling and its usefulness. We also design algorithms for simulating the event process under proposed models. The proposed approach can be adapted with little modification in many other research fields such as Criminology and Infectious Disease.
Corner kicks are an important event in soccer because they are often the result of strong attacking play and can be of keen interest to sports fans and bettors. Peng, Hu, and Swartz (2024, Computational Statistics) frame the commonly available corner kick data as right-censored event times, formulate the mixture feature of corner kick times caused by previous corner kicks, and explore patterns of corner kicks associated with several factors. This paper extends their modeling to accommodate the potential correlations between corner kicks by the same teams within the same games. We consider a frailty model for event times and apply the Monte Carlo Expectation Maximization (MCEM) algorithm to obtain the maximum likelihood estimates for the model parameters. We compare the proposed model with the model in Peng, Hu, and Swartz (2024) using likelihood ratio tests. The 2019 Chinese Super League (CSL) data are employed throughout the paper for motivation and illustration.
Recent research highlights a strong correlation between COVID-19 hospitalizations and wastewater viral signals. Increases in wastewater viral signals may be early warnings of increases in hospital admissions. That indicates a promising opportunity to assess and predict the burden of infectious diseases and has driven the widespread adoption and development of wastewater monitoring tools by public health organizations. Previous studies utilize distributed lag models to explore associations of COVID-19 hospitalizations with lagged SARS-CoV-2 wastewater viral signals. However, the conventional distributed lag models assume the duration time of the lag to be fixed, which is not always plausible. This paper presents Markov-modulated models with distributed lasting time, treating the duration of the lag as a random variable defined by a hidden process. We evaluate exposure effects over the duration time and estimate the distribution of the lasting time using the wastewater data and COVID-19 hospitalization records from Ottawa, Canada during June 2020 to November 2022. The different COVID-19 pandemic waves are accommodated in the statistical learning. Moreover, two strategies for comparing the associations over different time intervals are exemplified using the Ottawa data. Of note, the proposed Markov modulated models, an extension of distributed lag models, are potentially applicable to many different problems where the lag time is not fixed.
This paper presents a strategy for analyzing zero-truncated recurrent events data. Motivated by a pediatric mental health care (PMHC) program, we are particularly concerned with how the event occurrence depends on the occurrences in the past. We consider a stratified Cox regression model with time-varying coefficients and propose a procedure for estimating the model parameters using the zero-truncated data integrated with population census information. We evaluate the finite-sample performance of the proposed estimator through simulation and establish its asymptotic properties. Data from the PMHC program are used throughout the paper to motivate and to illustrate the proposed approach.
BackgroundCoronavirus disease (COVID-19) quickly spread around the world after its initial identification in Wuhan, China in 2019 and became a global public health crisis. COVID-19 related hospitalizations and deaths as important disease outcomes have been investigated by many studies while less attention has been given to the relationship between these two outcomes at a public health unit level. In this study, we aim to establish the relationship of counts of deaths and hospitalizations caused by COVID-19 over time across 34 public health units in Ontario, Canada, taking demographic, geographic, socio-economic, and vaccination variables into account.MethodsWe analyzed daily data of the 34 health units in Ontario between March 1, 2020 and June 30, 2022. Associations between numbers of COVID-19 related deaths and hospitalizations were explored over three subperiods according to the availability of vaccines and the dominance of the Omicron variant in Ontario. A generalized additive model (GAM) was fit in each subperiod. Heterogeneity across public health units was formulated via a random intercept in each of the models.ResultsMean daily COVID-19 deaths increased quickly as daily hospitalizations increased, particularly when daily hospitalizations were less than 20. In all the subperiods, mean daily deaths of a public health unit was significantly associated with its population size and the proportion of confirmed cases in subjects over 60 years old. The proportion of fully vaccinated (2 doses of primary series) people in the 60 + age group was a significant factor after the availability of the COVID-19 vaccines. The deprivation index, a measure of poverty, had a significantly positive effect on COVID-19 mortality after the dominance of the Omicron variant in Ontario. Quantification of these effects was provided, including effects related to public health units.ConclusionsThe differences in COVID-19 mortality across health units decreased over time, after adjustment for other covariates. In the last subperiod when most public health protections were released and the Omicron variant dominated, the least advantaged group might suffer higher COVID-19 mortality. Interventions such as paid sick days and cleaner indoor air should be made available to counter lifting of health protections.
We extended the Wikle's Bayesian hierarchical model based on a diffusion-reaction equation [Wikle, 2003] to investigate the COVID-19 spatio-temporal spread events across the USA from Mar 2020 to Feb 2022. Our model incorporated an advection term to account for the intra-state spread trend. We applied a Markov chain Monte Carlo (MCMC) method to obtain samples from the posterior distribution of the parameters. We implemented the approach via the collection of the COVID-19 infections across the states overtime from the New York Times. Our analysis shows that our approach can be robust to model misspecification to a certain extent and outperforms a few other approaches in the simulation settings. Our analysis results confirm that the diffusion rate is heterogeneous across the USA, and both the growth rate and the advection velocity are time-varying.
To understand the patterns of times to corner kicks in soccer and how they are associated with a few important factors, we analyze the corner kick records from the 2019 regular season of the Chinese Super League. This paper is particularly concerned with the elapsed time to a corner kick from a natural starting point. We overcome 2 challenges arising from such time-to-event analyses, which have not been discussed in the sports analytics literature. The first is that observations of times to corner kicks are subject to right-censoring. A given soccer starting point rarely ends with a corner kick but the occurrence of a different terminal event. The second issue is the mixture feature of short and typical gap times to the next corner kick from a particular one. There is often a subsequent corner kick quickly following a corner kick. The conventional event time models are thus inappropriate for formulating distributions of corner kick times. Our analysis reveals how the timing of corner kicks is associated with the factors of first versus second half of the game, home versus away team, score differential, betting odds prior to the game, and red card differential. We present applications of the developed statistical model for prediction to support tactics and sports betting.
We aim to develop a tool for understanding how the mental health of youth aged less than 18 years evolve over time through administrative records of mental health related emergency department (MHED) visits in two decades. Administrative health data usually contain rich information for investigating public health issues; however, many restrictions and regulations apply to their use. Moreover, the data are usually not in a conventional format since administrative databases are created and maintained to serve non-research purposes and only information for people who seek health services is accessible. Analysis of administrative health data is thus challenging in general. In the MHED data analyses, we are particularly concerned with (i) evaluating dynamic patterns and impacts with doubly-censored recurrent event data, and (ii) re-calibrating estimators developed based on truncated data by leveraging summary statistics from the population. The findings are verified empirically via simulation. We have established the asymptotic properties of the inference procedures. The contributions of this paper are twofold. We present innovative strategies for processing doubly-censored recurrent event data, and overcoming the truncation induced by the data collection. In addition, through exploring the pediatric MHED visit records, we provide new insights into children/youths mental health changes over time.
Administrative databases have become an increasingly popular data source for population-based health research. We explore how mortality risk is associated with some health service utilization process via linked administrative data. A generalized Cox regression model is proposed using a time-dependent stratification variable to summarize lifetime service utilization. Recognizing the service utilization over time as an internal covariate in the survival analysis, conventional likelihood methods are inapplicable. We present an estimating function based procedure for estimating model parameters, and provide a testing procedure for updating the stratification levels. The proposed approach is examined both asymptotically and numerically via simulation. We motivate and illustrate the proposed approach using an on-going program pertaining to opioid agonist treatment (OAT) management for individuals identified with opioid use disorders. Our analysis of the OAT data indicates that the OAT effect on mortality risk decreases in successive OAT attempts, in which two risk classes based on an individual's treatment episode number are established: one with 1-3 OAT episodes, and the other with 4+ OAT episodes.
Monitoring of viral signal in wastewater is considered a useful tool for monitoring the burden of COVID-19, especially during times of limited availability in testing. Studies have shown that COVID-19 hospitalizations are highly correlated with wastewater viral signals and the increases in wastewater viral signals can provide an early warning for increasing hospital admissions. The association is likely nonlinear and time-varying. This project employs a distributed lag nonlinear model (DLNM) (Gasparrini et al., 2010) to study the nonlinear exposure-response delayed association of the COVID-19 hospitalizations and SARS-CoV-2 wastewater viral signals using relevant data from Ottawa, Canada. We consider up to a 15-day time lag from the average of SARS-CoV N1 and N2 gene concentrations to COVID-19 hospitalizations. The expected reduction in hospitalization is adjusted for vaccination efforts. A correlation analysis of the data verifies that COVID-19 hospitalizations are highly correlated with wastewater viral signals with a time-varying relationship. Our DLNM based analysis yields a reasonable estimate of COVID-19 hospitalizations and enhances our understanding of the association of COVID-19 hospitalizations with wastewater viral signals.
INTRODUCTION:Urine drug tests (UDTs) are commonly used for monitoring opioid agonist treatment (OAT) responses, supporting the clinical decision for take-home doses and monitoring potential diversion. However, there is limited evidence supporting the utility of mandatory UDTs-particularly the impact of UDT frequency on OAT retention. Real-world evidence can inform patient-centred approaches to OAT and improve current strategies to address the ongoing opioid public health emergency. Our objective is to determine the safety and comparative effectiveness of alternative UDT monitoring strategies as observed in clinical practice among OAT clients in British Columbia, Canada from 2010 to 2020. METHODS AND ANALYSIS:We propose a population-level retrospective cohort study of all individuals 18 years of age or older who initiated OAT from 1 January 2010 to 17 March 2020. The study will draw on eight linked health administrative databases from British Columbia. Our primary outcomes include OAT discontinuation and all-cause mortality. To determine the effectiveness of the intervention, we will emulate a 'per-protocol' target trial using a clone censoring approach to compare fixed and dynamic UDT monitoring strategies. A range of sensitivity analyses will be executed to determine the robustness of our results. ETHICS AND DISSEMINATION:The protocol, cohort creation and analysis plan have been classified and approved as a quality improvement initiative by Providence Health Care Research Ethics Board and the Simon Fraser University Office of Research Ethics. Results will be disseminated to local advocacy groups and decision-makers, national and international clinical guideline developers, presented at international conferences and published in peer-reviewed journals electronically and in print.
Since the beginning of the global pandemic of Coronavirus (SARS-COV-2), there has been many studies devoted to predicting the COVID-19 related deaths/hospitalizations. The aim of our work is to (1) explore the lagged dependence between the time series of case counts and the time series of death counts; and (2) utilize such a relationship for prediction. The proposed approach can also be applied to other infectious diseases or wherever dynamics in lagged dependence are of primary interest. Different from the previous studies, we focus on time-varying coefficient models to account for the evolution of the coronavirus. Using two different types of time-varying coefficient models, local polynomial regression models and piecewise linear regression models, we analyze the province-level data in Canada as well as country-level data using cumulative counts. We use out-of-sample prediction to evaluate the model performance. Based on our data analyses, both time-varying coefficient modeling strategies work well. Local polynomial regression models generally work better than piecewise linear regression models, especially when the pattern of the relationship between the two time series of counts gets more complicated (e.g., more segments are needed to portray the pattern). Our proposed methods can be easily and quickly implemented via existing R packages.