In disease surveillance, capture-recapture methods are commonly used to estimate the number of diseased cases in a defined target population. Since the number of cases never identified by any surveillance system cannot be observed, estimation of the case count typically requires at least one crucial assumption about the dependency between surveillance systems. However, such assumptions are generally unverifiable based on the observed data alone. In this paper, we advocate a modeling framework hinging on the choice of a key population-level parameter that reflects dependencies among surveillance streams. With the key dependency parameter as the focus, the proposed method offers the benefits of (a) incorporating expert opinion in the spirit of prior information to guide estimation; (b) providing accessible bias corrections, and (c) leveraging an adapted credible interval approach to facilitate inference. We apply the proposed framework to two real human immunodeficiency virus surveillance datasets exhibiting three-stream and four-stream capture-recapture-based case count estimation. Our approach enables estimation of the number of human immunodeficiency virus positive cases for both examples, under realistic assumptions that are under the investigator's control and can be readily interpreted. The proposed framework also permits principled uncertainty analyses through which a user can acknowledge their level of confidence in assumptions made about the key non-identifiable dependency parameter.
In this paper, we advocate and expand upon a previously described monitoring strategy for efficient and robust estimation of disease prevalence and case numbers within closed and enumerated populations such as schools, workplaces, or retirement communities. The proposed design relies largely on voluntary testing, which is notoriously biased (e.g., in the case of coronavirus disease 2019) due to nonrepresentative sampling. The approach yields unbiased and comparatively precise estimates with no assumptions about factors underlying selection of individuals for voluntary testing, building on the strength of what can be a small random sampling component. This component enables the use of a recently proposed “anchor stream” estimator, a well-calibrated alternative to classical capture-recapture (CRC) estimators based on 2 data streams. We show that this estimator is equivalent to a direct standardization based on “capture,” that is, selection (or not) by the voluntary testing program, made possible by means of a key parameter identified by design. This equivalency simultaneously allows for novel 2-stream CRC-like estimation of general mean values (e.g., means of continuous variables like antibody levels or biomarkers). For inference, we propose adaptations of Bayesian credible intervals when estimating case counts and bootstrapping when estimating means of continuous variables. We use simulations to demonstrate significant precision benefits relative to random sampling alone.
Monitoring key elements of disease dynamics (e.g., prevalence, case counts) is of great importance in infectious disease prevention and control, as emphasized during the COVID-19 pandemic. To facilitate this effort, we propose a new capture-recapture (CRC) analysis strategy that takes misclassification into account from easily-administered, imperfect diagnostic test kits, such as the Rapid Antigen Test-kits or saliva tests. Our method is based on a recently proposed "anchor stream" design, whereby an existing voluntary surveillance data stream is augmented by a smaller and judiciously drawn random sample. It incorporates manufacturer-specified sensitivity and specificity parameters to account for imperfect diagnostic results in one or both data streams. For inference to accompany case count estimation, we improve upon traditional Wald-type confidence intervals by developing an adapted Bayesian credible interval for the CRC estimator that yields favorable frequentist coverage properties. When feasible, the proposed design and analytic strategy provides a more efficient solution than traditional CRC methods or random sampling-based biased-corrected estimation to monitor disease prevalence while accounting for misclassification. We demonstrate the benefits of this approach through simulation studies that underscore its potential utility in practice for economical disease monitoring among a registered closed population.
Capture–recapture methods are widely applied in estimating the number ( ) of prevalent or cumulatively incident cases in disease surveillance. Here, we focus the bulk of our attention on the common case in which there are 2 data streams. We propose a sensitivity and uncertainty analysis framework grounded in multinomial distribution-based maximum likelihood, hinging on a key dependence parameter that is typically nonidentifiable but is epidemiologically interpretable. Focusing on the epidemiologically meaningful parameter unlocks appealing data visualizations for sensitivity analysis and provides an intuitively accessible framework for uncertainty analysis designed to leverage the practicing epidemiologist’s understanding of the implementation of the surveillance streams as the basis for assumptions driving estimation of . By illustrating the proposed sensitivity analysis using publicly available HIV surveillance data, we emphasize both the need to admit the lack of information in the observed data and the appeal of incorporating expert opinion about the key dependence parameter. The proposed uncertainty analysis is a simulation-based approach designed to more realistically acknowledge variability in the estimated associated with uncertainty in an expert’s opinion about the nonidentifiable parameter, together with the statistical uncertainty. We demonstrate how such an approach can also facilitate an appealing general interval estimation procedure to accompany capture–recapture methods. Simulation studies illustrate the reliable performance of the proposed approach for quantifying uncertainties in estimating in various contexts. Finally, we demonstrate how the recommended paradigm has the potential to be directly extended for application to data from >2 surveillance streams.
Epidemiologic screening programs often make use of tests with small, but non-zero probabilities of misdiagnosis. In this article, we assume the target population is finite with a fixed number of true cases, and that we apply an imperfect test with known sensitivity and specificity to a sample of individuals from the population. In this setting, we propose an enhanced inferential approach for use in conjunction with sampling-based bias-corrected prevalence estimation. While ignoring the finite nature of the population can yield markedly conservative estimates, direct application of a standard finite population correction (FPC) conversely leads to underestimation of variance. We uncover a way to leverage the typical FPC indirectly toward valid statistical inference. In particular, we derive a readily estimable extra variance component induced by misclassification in this specific but arguably common diagnostic testing scenario. Our approach yields a standard error estimate that properly captures the sampling variability of the usual bias-corrected maximum likelihood estimator of disease prevalence. Finally, we develop an adapted Bayesian credible interval for the true prevalence that offers improved frequentist properties (i.e., coverage and width) relative to a Wald-type confidence interval. We report the simulation results to demonstrate the enhanced performance of the proposed inferential methods.
In epidemiological studies, the capture-recapture (CRC) method is a powerful tool that can be used to estimate the number of diseased cases or potentially disease prevalence based on data from overlapping surveillance systems. Estimators derived from log-linear models are widely applied by epidemiologists when analyzing CRC data. The popularity of the log-linear model framework is largely associated with its accessibility and the fact that interaction terms can allow for certain types of dependency among data streams. In this work, we shed new light on significant pitfalls associated with the log-linear model framework in the context of CRC using real data examples and simulation studies. First, we demonstrate that the log-linear model paradigm is highly exclusionary. That is, it can exclude, by design, many possible estimates that are potentially consistent with the observed data. Second, we clarify the ways in which regularly used model selection metrics (e.g., information criteria) are fundamentally deceiving in the effort to select a best model in this setting. By focusing attention on these important cautionary points and on the fundamental untestable dependency assumption made when fitting a log-linear model to CRC data, we hope to improve the quality of and transparency associated with subsequent surveillance-based CRC estimates of case counts.
Surveillance research is of great importance for effective and efficient epidemiological monitoring of case counts and disease prevalence. Taking specific motivation from ongoing efforts to identify recurrent cases based on the Georgia Cancer Registry, we extend recently proposed "anchor stream" sampling design and estimation methodology. Our approach offers a more efficient and defensible alternative to traditional capture-recapture (CRC) methods by leveraging a relatively small random sample of participants whose recurrence status is obtained through a principled application of medical records abstraction. This sample is combined with one or more existing signaling data streams, which may yield data based on arbitrarily non-representative subsets of the full registry population. The key extension developed here accounts for the common problem of false positive or negative diagnostic signals from the existing data stream(s). In particular, we show that the design only requires documentation of positive signals in these non-anchor surveillance streams, and permits valid estimation of the true case count based on an estimable positive predictive value (PPV) parameter. We borrow ideas from the multiple imputation paradigm to provide accompanying standard errors, and develop an adapted Bayesian credible interval approach that yields favorable frequentist coverage properties. We demonstrate the benefits of the proposed methods through simulation studies, and provide a data example targeting estimation of the breast cancer recurrence case count among Metro Atlanta area patients from the Georgia Cancer Registry-based Cancer Recurrence Information and Surveillance Program (CRISP) database.
The application of serial principled sampling designs for diagnostic testing is often viewed as an ideal approach to monitoring prevalence and case counts of infectious or chronic diseases. Considering logistics and the need for timeliness and conservation of resources, surveillance efforts can generally benefit from creative designs and accompanying statistical methods to improve the precision of sampling-based estimates and reduce the size of the necessary sample. One option is to augment the analysis with available data from other surveillance streams that identify cases from the population of interest over the same timeframe, but may do so in a highly nonrepresentative manner. We consider monitoring a closed population (e.g., a long-term care facility, patient registry, or community), and encourage the use of capture-recapture methodology to produce an alternative case total estimate to the one obtained by principled sampling. With care in its implementation, even a relatively small simple or stratified random sample not only provides its own valid estimate, but provides the only fully defensible means of justifying a second estimate based on classical capture-recapture methods. We initially propose weighted averaging of the two estimators to achieve greater precision than can be obtained using either alone, and then show how a novel single capture-recapture estimator provides a unified and preferable alternative. We develop a variant on a Dirichlet-multinomial-based credible interval to accompany our hybrid design-based case count estimates, with a view toward improved coverage properties. Finally, we demonstrate the benefits of the approach through simulations designed to mimic an acute infectious disease daily monitoring program or an annual surveillance program to quantify new cases within a fixed patient registry.
We propose a monitoring strategy for efficient and robust estimation of disease prevalence and case numbers within closed and enumerated populations such as schools, workplaces, or retirement communities. The proposed design relies largely on voluntary testing, notoriously biased (e.g., in the case of COVID-19) due to non-representative sampling. The approach yields unbiased and comparatively precise estimates with no assumptions about factors underlying selection of individuals for voluntary testing, building on the strength of what can be a small random sampling component. This component unlocks a previously proposed "anchor stream" estimator, a well-calibrated alternative to classical capture-recapture (CRC) estimators based on two data streams. We show here that this estimator is equivalent to a direct standardization based on "capture", i.e., selection (or not) by the voluntary testing program, made possible by means of a key parameter identified by design. This equivalency simultaneously allows for novel two-stream CRC-like estimation of general means (e.g., of continuous variables such as antibody levels or biomarkers). For inference, we propose adaptations of a Bayesian credible interval when estimating case counts and bootstrapping when estimating means of continuous variables. We use simulations to demonstrate significant precision benefits relative to random sampling alone.
Decreased cognitive function is related to undesirable psychological outcomes such as greater emotional distress and lower quality of life, particularly among women living with HIV who experience cognitive impairment (WLWH-CI). Yet, few studies have examined the psychosocial resources that may attenuate these negative emotional outcomes. The current study sought to identify the interrelated contributions of social relationships and psychological resources in 399 WLWH-CI by applying Socio-Emotional Adaptation (SEA) theory using data from the Women's Interagency HIV Study (WIHS). Cognitive impairment (CI) was defined as impairment on two or more cognitive domains. Logistic regression models were used to estimate the odds of experiencing specific emotions due to a combination of four psychosocial resources. Emotions (i.e., depression, apathy, fear, anger, and acceptance) were related to a combination of binary (positive/negative) psychosocial resources including relationship with an informal support partner, relationship with a formal caregiver, coping, and perceived control. Understanding the conditions that may influence emotions in WLWH-CI is important for identifying and appropriately addressing the needs of this population. As CI increases, these individuals experience increasing challenges with articulating their care needs and having their needs met. As such, it becomes increasingly important to identify possible triggers for emotional responses to best address these underlying challenges.
BACKGROUND:Pain management approaches during uterine aspiration vary, which include local anesthetic, oral analgesics, moderate sedation, deep sedation, or a combination of approaches. For local anesthetic approaches specifically, we continue to have suboptimal pain control. Gabapentin as an adjunct to pain management has proven to be beneficial in gynecologic surgery. We sought to evaluate the impact of gabapentin on perioperative pain during surgical management of first-trimester abortion or early pregnancy loss with uterine aspiration under local anesthesia. OBJECTIVE:We hypothesized that adding gabapentin to local anesthesia will reduce perioperative and postoperative pain associated with uterine aspiration. Secondary outcomes included tolerability of gabapentin and postoperative pain, nausea, vomiting, and anxiety. STUDY DESIGN:We conducted a randomized double-blinded placebo-controlled trial of gabapentin 600 mg given 1 to 2 hours preoperatively among subjects receiving a first-trimester uterine aspiration under paracervical block in an outpatient ambulatory surgery center. There were 111 subjects randomized. The primary outcome was pain at time of uterine aspiration as measured on a 100-mm visual analog scale. Secondary outcomes included pain at other perioperative time points. To assess changes in pain measures, an intention to treat mixed effects model was fit with treatment groups (gabapentin vs control) as a between-subjects factor and time point as a within-subjects factor plus their interaction term. Because of a non-normal distribution of pain scores, the area under the curve was calculated for secondary outcomes with comparison of groups utilizing Mann-Whitney U tests. RESULTS:Among the 111 randomized, most subjects were Black or African American (69.4%), mean age was 26 years (±5.5), and mean gestational age was 61.3 days (standard deviation, 14.10). Mean pain scores at time of uterine aspiration were 66.77 (gabapentin) vs 71.06 (placebo), with a mean difference of -3.38 (P=.51). There were no significant changes in pain score preoperatively or intraoperatively. Subjects who received gabapentin had significantly lower levels of pain at 10 minutes after surgery (mean difference [standard error (SE)]=-13.0 [-5.0]; P=.01) and 30 minutes after surgery (mean difference [SE]=-10.8 [-5.1]; P=.03) compared with subjects who received placebo. Median nausea scores and incidence of emesis pre- and postoperatively did not differ between groups. Similarly, anxiety scores did not differ between groups, before or after the procedure. At 10 and 30 minutes after the procedure, most participants reported no side effects or mild side effects, and this did not differ between groups. CONCLUSION:Preoperative gabapentin did not reduce pain during uterine aspiration. However, it did reduce postoperative pain, which may prove to be a desired attribute of its use, particularly in cases where postoperative pain may be a greater challenge.
Penalized methods for variable selection such as the Smoothly Clipped Absolute Deviation penalty have been increasingly applied to aid variable section in regression analysis. Much of the literature has focused on parametric models, while a few recent studies have shifted the focus and developed their applications for the popular semi-parametric, or distribution-free, generalized estimating equations (GEEs) and weighted GEE (WGEE). However, although the WGEE is composed of one main and one missing-data module, available methods only focus on the main module, with no variable selection for the missing-data module. In this paper, we develop a new approach to further extend the existing methods to enable variable selection for both modules. The approach is illustrated by both real and simulated study data.
The visualization of tomato growth can be used in 3D computer games and virtual gardens. Based on the growth theory involving the respiration theory, the photosynthesis, and dry matter partition, a visual system is developed. The tomato growth visual simulation system is light-and-temperature-dependent and shows plausible visual effects in consideration of the continuous growth, texture map, gravity influence, and collision detection. In addition, the virtual tomato plant information, such as the plant height, leaf area index, fruit weight, and dry matter, can be updated and output in real time.
Longitudinal studies are used in mental health research and services studies. The dominant approaches for longitudinal data analysis are the generalized linear mixed-effects models (GLMM) and the weighted generalized estimating equations (WGEE). Although both classes of models have been extensively published and widely applied, differences between and limitations about these methods are not clearly delineated and well documented. Unfortunately, some of the differences and limitations carry significant implications for reporting, comparing and interpreting research findings. In this report, we review both major approaches for longitudinal data analysis and highlight their similarities and major differences. We focus on comparison of the two classes of models in terms of model assumptions, model parameter interpretation, applicability and limitations, using both real and simulated data. We discuss caveats and cautions when applying the two different approaches to real study data.