Causal mediation analysis is an effective method for understanding the mechanism between the exposure and the outcome, often assuming that the mediation model is consistent for each individual in the target population. In practice, however, the natural indirect effect (NIE) may vary across individuals due to their distinct characteristics. As a result, the population can be partitioned into subgroups according to the varying sizes of the NIEs. Distinguishing subgroups within the study population enables the development of more precise and targeted treatment strategies. In this paper, we propose an identifiable mixture mediation model with latent subgroups for the survival data, where the outcome follows an accelerated failure time model and the mediator is Gaussian distributed. We further employ three information criteria including the AIC, BIC, and singular BIC (sBIC) to select the number of subgroups, followed by the expectation-maximization (EM) algorithm to estimate the model parameters and NIEs. Simulation study shows that the sBIC is the most robust and efficient criterion for selecting the number of subgroups; therefore, we recommend the sBIC-EM algorithm for practical use. Lastly, we apply our algorithm to the lung cancer data and discover two latent groups with opposing NIEs.
Causal mediation analysis is a popular approach for investigating whether the effect of an exposure on an outcome is through a mediator to better understand the underlying causal mechanism. In recent literature, mediation analysis with multiple mediators has been proposed for continuous and dichotomous outcomes. In contrast, methods for mediation analysis for an ordinal outcome are still underdeveloped. In this paper, we first review mediation analysis methods with a continuous mediator for an ordinal outcome and then develop mediation analysis with a binary mediator for an ordinal outcome. We further consider multiple mediators for an ordinal outcome in the counterfactual framework and provide identification assumptions for identifying the mediation effects. Under the identification assumptions, we propose a regression-based method to estimate the mediation effects through multiple mediators while allowing the presence of exposure-mediator interactions. The closed-form expressions of mediation effects are also obtained for three scenarios: multiple continuous mediators, multiple binary mediators, and multiple mixed mediators. We conduct simulation studies to assess the finite sample performance of our new methods and present the biases, standard errors, and confidence intervals to demonstrate that our proposed estimators perform well in a wide range of practical settings. Finally, we apply our proposed methods to assess the mediation effects of candidate DNA methylation CpG sites in the causal pathway from socioeconomic index to body mass index.
传统的卷积神经网络(Convolutional Neural Network,CNN)的预测结果错误率不可控并且不具备置信度衡量,当面对高风险低容错率的问题时,这种方法的预测结果可靠性低.针对这个缺点,基于归纳一致性预测提出一种错误率可控的图像分类算法.该算法通过训练更小的训练集有效缩短了卷积神经网络的训练时间,并根据奇异值函数衡量测试样本与各个类别之间的一致性程度,使得输出结果具有极高的可靠性.实验结果表明,提出的算法能够有效地控制神经网络的预测准确率,并且能够提供具有可靠置信度的预测集合.
因果中介分析用于研究自变量是否通过中介变量对因变量产生影响,阐明潜在因果机制.本文研究当中介变量具有测量误差时,利用优势比方法定义中介效应,对无交互作用和有交互作用分别建立逻辑回归和线性回归模型的结构方程,获得自然直接效应和间接效应的解析表达式.
针对一致性预测支持向量机的多分类问题,提出了两种多分类算法,分别是基于一致性预测一对多支持向量机算法(One-Vs-Rest Support Vector Machine Algorithm Via Conformal Predictors,OVR SVM CP)和基于一致性预测一对一支持向量机算法(One-Vs-One Support Vector Machine Algorithm Via Conformal Predictors,OVO SVM CP).首先,将多分类问题转化为二分类问题,利用决策函数定义奇异值函数.然后,对这两种算法进行数值模拟实验,并与OVO SVM、OVO LSSVM、OVO TWSVM、HSVM算法相比较.最后,将两种算法应用于6组真实数据集测试其分类预测效果.仿真实验和真实数据应用结果表明,提出的两种算法预测效果较好,相比于其他3种的支持向量机算法有更高的预测准确率.
为有效落实高等教育"立德树人"根本任务,文章立足于对"多元统计分析"课程传统教学模式存在问题的分析,构建了"多元统计分析"课程"翻转课堂+课程思政"教学模式,并采用实证分析法对该教学模式的教学效果进行了探讨.实证分析结果表明:相较于传统教学模式,"多元统计分析"课程"翻转课堂+课程思政"教学模式不仅有效地提升了学生学习积极性,明显地提高了教学效果,而且专业课程与思想政治理论课同向同行,形成了协同效应,更有利于高等教育立德树人根本任务的落实.
In social and behavioral sciences, the mediation test based on the indirect effect is an important topic. There are many methods to assess intervening variable effects. In this paper, we focus on the difference method and the product method in mediation models. Firstly, we analyze the regression functions in the simple mediation model, and provide an expectation-consistent condition. We further show that the difference estimator and the product estimator are numerically equivalent based on the least-squares regression regardless of the error distribution. Secondly, we generalize the equivalence result to the three-path model and the multiple mediators model, and prove a general equivalence result in a class of restricted linear mediation models. Thirdly, we investigate the empirical distributions of the indirect effect estimators in the simple mediation model by simulations, and show that the indirect effect estimators are normally distributed as long as one multiplicand of the product estimator is large. Finally, we introduce some popular R packages for mediation analysis and also provide some useful suggestions on how to correctly conduct mediation analysis.
因果中介分析是通过中介变量识别解释自变量和因变量之间关系的机制.在单中介变量的因果模型下,通过中介变量建立起了自变量与因变量之间的结构方程,利用极大似然方法获得结构方程中参数估计以及间接效应估计,并由delta方法获得间接效应估计量的渐近正态分布.在有限样本下处理了模拟研究,模拟研究的结果表明我们提出的估计表现良好.最后,应用提出的方法通过DNA 甲基化位点分析社会经济指数对体重指数的影响.
Hypothesis testing for normal distributions is one important problem in statistics and related fields including management science, engineering science and medical science. In this paper, from a very unique perspective, we propose a unified framework to comprehensively review the existing literature on the one- and two-sample testing problems of normal distributions. The unified framework has integrated the literature in a way that it includes most commonly used tests as special cases, including the one-sample mean test, the one-sample variance test, the two-sample mean test, the two-sample variance test, and the Behrens-Fisher test. The unified framework has also put forward two new hypothesis tests that are rarely studied in the literature. To complete the puzzle, we propose two likelihood ratio test statistics to solve those new testing problems. Simulation studies and real data examples are also provided to demonstrate that our proposed test statistics are appropriate for practical implementation.
因果中介分析通过中介变量识别和解释自变量与因变量之间的因果机制.基于存在混杂变量的因果中介模型,运用线性回归建立结构方程,运用最小二乘获得参数估计量.由正态分布性质与delta方法获得自然直接效应与自然间接效应估计量服从正态分布.在有限样本下处理模拟研究,所提出的估计表现良好.并应用所提出的方法来评估DNA甲基化CpG位点在从社会经济指数到体重指数的因果路径中的中介作用.
因果中介分析用于研究自变量是否通过中介变量对因变量产生效应,阐明潜在因果机制.针对含有混杂变量与交互作用的因果中介模型,运用线性回归建立结构方程,给出基于反事实框架的间接效应和直接效应表达式,获得间接效应和直接效应极大似然估计量,由delta方法获得效应值估计量分布.在有限样本下处理了模拟研究,模拟研究结果表明,本文提出的估计表现良好,可将该方法用于体质肥胖数据研究.
因果中介分析是调查自变量通过中介变量对因变量的影响.近年来,中介分析研究成果越来越多,但主要针对一个中介变量模型.在本文中,我们构造了具有多个中介变量的二变量中介模型,提出了中介效应的估计,并对估计结果进行了模拟研究.模拟结果表明我们提出的估计效果良好.
随着互联网技术蓬勃发展,近年来各种网络资源及网络教学方式纷纷涌现,这极大推动了网络教学的发展.基于安徽省某高校学生的调查数据,本文首先从专业、网络课程学习兴趣、学习时间、学习因素、学习效果等方面对大学生网络课程学习情况进行统计分析,分析影响学生参与网络课程学习的主要影响因素,并结合调查数据建立相应的统计模型.实证分析表明:现阶段大学生对网络课程的兴趣程度很高,网络课程与传统课堂学习具有很强的互补性.同时,网络课程也需要在扩大宣传、提高质量、改进学习模式、增强教学管理与监督等方面持续改进.
双一流建设和传统行业面临产业升级,这对地方高校具有工程创新能力的引领性复合型人才培养提出了更高要求.文章以安徽理工大学研究生《工程数学》课程建设为例,就高水平地方理工类院校如何强化数学理论与工程背景相结合,进一步夯实研究生数学基础,助力高素质、创新型工程人才的培养进行了探讨.
Difference-based methods have attracted increasing attention for analyzing partially linear models in the recent literature. In this paper, we first propose to solve the optimal sequence selection problem in difference-based estimation for the linear component. To achieve the goal, a family of new sequences and a cross-validation method for selecting the adaptive sequence are proposed. We demonstrate that the existing sequences are only extreme cases in the proposed family. Secondly, we propose a new estimator for the residual variance by fitting a linear regression method to some difference-based estimators. Our proposed estimator achieves the asymptotic optimal rate of mean squared error. Simulation studies also demonstrate that our proposed estimator performs better than the existing estimator, especially when the sample size is small and the nonparametric function is rough.
In nonparametric regression, it is often needed to detect whether there are jump discontinuities in the mean function. In this paper, we revisit the difference-based method in [13] and propose to further improve it. To achieve the goal, we first reveal that their method is less efficient due to the inappropriate choice of the response variable in their linear regression model. We then propose a new regression model for estimating the residual variance and the total amount of discontinuities simultaneously. In both theory and simulation, we show that the proposed variance estimator has a smaller mean-squared error compared to the existing estimator, whereas the estimation efficiency for the total amount of discontinuities remains unchanged. Finally, we construct a new test procedure for detection of discontinuities using the proposed method; and via simulation studies, we demonstrate that our new test procedure outperforms the existing one in most settings.
统计学在经济学科、管理学科等其他学科有着广泛的应用,因此统计人才的培养相当重要.本文针对工科院校统计学专业人才的培养定位和存在的问题,以统计学人才培养定位和目标为出发点,从专业建设、课程教学和实践教学等方面对统计学本科教学提出了几点思考和想法,从而提高统计学本科教学质量,以达到培养合格的统计人才目的.
In this paper,to test the equality of two normal populations with one variance known,a new likelihood ratio test statistic is proposed,which is different from the one proposed by Neyman and Pearson.Then the exact null distribution of the new test statistic is derived.To compare the new test statistic with the one proposed by Neyman and Pearson,two simulation studies are conducted based on Monte Carlo method:one is the type I error and the other is the power.The simulation results show that the method proposed in this paper performs better.
In this paper, the likelihood ratio test for the equality of two normal populations with one mean known is studied.A new likelihood ratio test statistic that is different from the method proposed by Neyman and Pearson is constructed and its exact null distribution is also derived.We conduct the simulation analysis for the type I er-ror and the power test of the test statistic.Simulation studies show that our proposed test statistic performs better than Neyman and Pearson ' s in controlling the type I error and the power.
Interest in variance estimation in nonparametric regression has grown greatly in the past several decades. Among the existing methods, the least squares estimator in Tong and Wang (2005) is shown to have nice statistical properties and is also easy to implement. Nevertheless, their method only applies to regression models with homoscedastic errors. In this paper, we propose two least squares estimators for the error variance in heteroscedastic nonparametric regression: the intercept estimator and the slope estimator. Both estimators are shown to be consistent and their asymptotic properties are investigated. Finally, we demonstrate through simulation studies that the proposed estimators perform better than the existing competitor in various settings.