Supplementary Figure S4 shows comparison of yearly counts: modeled vs observed counts for four rare cancers.
Supplementary Figure S7 shows observed vs. modeled age-adjusted incidence rates for all cancer sites combined (female) at three diagnosis periods (2005-2007, 2011-2013, and 2017-2019).
Age-standardization is a key statistical method used in health statistics to adjust rates such as mortality or incidence, enabling comparisons across populations or time points with different age structures. This review traces its historical development, global and country-specific practices, and future directions. The method dates back to the 19th century, with major adoption in the 20th century through the Segi and Doll's World Standard Population. While the World Health Organization (WHO) introduced an updated standard in 2000, the International Agency for Research on Cancer (IARC) continues to use the Segi and Doll's standard in the Cancer Incidence in Five Continents series, prioritizing consistency and comparability in long-term cancer surveillance. Case studies from the IARC, the United States (U.S.), Japan, and the Republic of Korea (Korea) illustrate different responses to changing demographics. The U.S. adopted the 2000 standard with expanded age detail for the elderly population. Japan introduced the 2015 Japan Standard Population to account for its rapidly aging society, though regional data limitations presented challenges. Korea, experiencing one of the fastest aging transitions globally, updated to a 2020 standard for more accurate national and sub-national reporting. The review also emphasizes that age-standardization can obscure important age-specific trends. Methods like Joinpoint clustering help detect divergent trends by age groups. Looking forward, age-standardization remains essential amid global demographic shifts. However, updates of standard populations must balance improved relevance with the need for continuity and robust data. International coordination and digital tools will support more flexible and transparent health statistics in the future.
Supplementary Table S2 lists the pool of covariates from years 2005 to 2019 and data sources.
Supplementary Figure S5 shows ratios of county-level observed to modeled age-adjusted incidence rates for five selected cancer sites.
Supplementary Figure S1 shows comparison of the yearly counts: modeled vs observed counts for all cancer sites combined (female).
Supplemental Figure S2 shows comparison of yearly counts: modeled vs observed counts for multiple common cancers.
Supplementary Table S3 shows results from the exploratory model comparison for prior selection: DIC & WAIC from the standard Poisson model (Model 1) with different spatial random effect assumptions (Besag2 and Leroux) for the 16 selected cancer sites.
Joinpoint regression can model trends in time-specific aggregated estimates. These methods have been developed mainly for non-survey data such as cancer registry data, and only recently have been extended to utilize survey data that accounts for complex sample designs resulting in non-zero correlation between the time-specific estimates. This correlation can occur for surveys with data from the same sampled units used across time points, for example, the annual National Health Interview Survey with multistage cluster samples using the same first-stage sampled clusters over consecutive time points. Another issue when modeling aggregated data is that the degrees of freedom for joinpoint analyses of multistage cluster samples are based on the number of time points, not the number of first-stage sampled clusters as used in survey methods. To address this, we propose models of individual-level data that incorporate both the correlation between time points and correct the degrees of freedom due to the sampling design that is needed for accurate inferences. Also, a modified design-based Akaike Information Criterion (M-dAIC) for model selection is proposed to account for complex sample designs. These new methods are empirically compared to existing methods using simulation studies and health survey data examples. The simulation studies indicated that this new individual-level model identified the true number of joinpoints more accurately than the established aggregate-level models for data collected using complex survey designs with moderate to large interclass correlation coefficients (ICC).
Supplementary Model Descriptions provides additional details on the models assumed for the random effects and hyperparameters.
Supplementary Figure S6 shows ratios of county-level observed to modeled age-adjusted incidence rates for five selected cancer sites.
Supplementary Table S1 includes the pool of cancer sites where the 16 sex-specific cancer sites were selected from.
BACKGROUND:Mapping cancer incidence is crucial for analyzing and visualizing patterns across geographic areas. Although many studies map cancer incidence at subnational levels (e.g., state, county), publicly available county-level data, especially for less common cancers, are limited. METHODS:Using data from the North American Association of Central Cancer Registries CiNA research database (2005-2019), we developed spatiotemporal hierarchical models to smooth/predict annual age group-specific case counts for all US counties. We compared Poisson and zero-truncated Poisson likelihoods and various priors. Model performance was assessed using the deviance information criterion, weighted Akaike information criterion, and average absolute relative deviation (AARD). Modeled age-adjusted rates were mapped to visualize spatial and temporal patterns. RESULTS:Modeled age-adjusted rates were produced for 16 selected sex-specific cancer sites across 3,109 counties from 2005 to 2019. AARD values varied by site and context, being lowest for common cancers and populous counties and highest for rare cancers and sparsely populated areas. Compared with maps of observed rates, modeled maps were smoother and more coherent, filling gaps and reducing extreme values driven by small case counts while preserving large-scale geographic gradients and temporal trends. CONCLUSIONS:The standard Poisson hierarchical mixed-effects model showed superior accuracy and computational efficiency and was selected for final estimation. As expected, the most accurate predictions are for more common cancer sites in more populous areas, and the least accurate predictions are for rarer cancers in areas with lower populations. IMPACT:The resulting estimates and maps could support surveillance, trend analysis, disparity identification, targeted interventions, and broader research efforts.
Supplementary Figure S3 shows comparison of yearly counts: modeled vs observed counts for five relatively less common cancers.
Supplementary Figure S8 shows observed vs. modeled age-adjusted incidence rates for prostate cancer at three diagnosis periods (2005-2007, 2011-2013, and 2017-2019).
Supplementary Figure S10 shows observed vs. modeled age-adjusted incidence rates for bones & joints (male) cancer at three diagnosis periods (2005-2007, 2011-2013, and 2017-2019).
Supplementary Figure S9 shows observed vs. modeled age-adjusted incidence rates for brain (female) cancer at three diagnosis periods (2005-2007, 2011-2013, and 2017-2019).