Correspondence analysis is generally a data analysis technique that expresses a two-dimensional or higher contingency table in a low-dimensional space to see the combination of rows and columns. In general, a statistical analysis method applied to contingency table analysis uses a chi-square test, but the chi-square test has a limitation in that it cannot show the coupling pattern for rows and columns. As an alternative, correspondence analysis is used. In general, the row profiles of correspondence analysis sum to 1. For data evaluated on the Likert scale for various attributes of each entity, the sum of rows is different for each row, so general correspondence analysis cannot be performed. Correspondence analysis can be applied to such data by doubling the number of column categories by adding the positive and negative values of each attribute so that the sum is equal. This is called the doubling technique. As a result of exploratory multivariate data analysis that does not rely on normal distribution and statistical models for Likert data, a correspondence analysis study using a doubling technique was proposed (Han, 2019). By quantifying the examiners and subjects, and expressing the result using a graphic technique, it was easy to visually recognize and interpret geometrically clear meaning. However, for the method proposed by Han (2019), the stability evaluation for quantification analysis results has not yet been developed. Therefore, in order to develop a methodology for the stability evaluation of the materialistic method for Likert data analysis, which is often seen in real life such as public opinion surveys or consumer preference surveys, but lacks development of analysis methods, and to show its usefulness through case analysis do.
한부모가족은 혼인 여부와 상관없이 남성가구주 또는 여성가구주 한명이 미혼자녀를 양육하는 가구이다. 한부모가족의 부모는 경제적, 심리적으로 많은 어려움을 겪고 있으며, 이로 인해 양부모 가정보다 우울증에 걸리는 비중이 더 높은 것으로 알려져 있다. 최근 우리나라는 한부모 가족이 점점 더 증가하는 추세에 있어, 한부모가족의 어려움을 해결하기 위한 방안이 필요로 되어지고 있다. 본 논문에서는 우리나라의 대표적 패널자료인 한국복지패널자료를 이용하여 한부모가족 부모들의 우울증에 영향을 주는 사회적, 경제적 요인을 파악하고자 한다. 본 논문에서는 패널자료분석을 위해 주변모형을 이용하였으며, 주변모형의 설정, 모수추정 방법 및 가설검정 방법 등에 대해 소개하였다. 또한 모형적합 후, 모형검진을 위한 방법을 소개하여 향후 주변모형을 이용하고자하는 연구자들에게 도움이 되고자 하였다. 분석결과, 기존에 알려진 경제적 요인뿐만 아니라, 자녀관계, 부모의 건강 등 비경제적 요인이 경제적 요인보다 더 우울증에 영향을 주는 것으로 밝혀졌다. 또한 모형검진 방법인 잔차누적합 방법을 이용한 결과 적합된 모형이 타당함을 알 수 있었다.
대응분석(correspondence analysis)은 일반적으로 2차원 이상의 분할표(contingency table)를 저차원 공간에 표현하여 행과 열의 결합 양상(pattern)을 볼 수 있는 다변량 자료기술(data description) 기법이다. 분할표에 흔히 적용되는 통계적 방법은 카이제곱 검정(chi-square test)인데, 이는 열범주의 상대적 빈도에 대한 행 표본들의 동일성(homogeneity) 혹은 행과 열의 독립성(independence)가설을 검정할 수 있다. 그러나 카이제곱 검정은 행과 열의 결합양상을 보여주지는 못한다. 이에 대한 문제를 해결할 수 있는 기법이 대응분석이다. 대응분석에서 각 행 프로파일 요소들은 합이 1이다. 그러나 각 개체에 대하여 여러 속성이 리커트 척도로 평가된 자료는 행 합계가 각 행마다 다르기 때문에 그대로는 대응분석을 할 수가 없다. 이런 자료에 대해서는 각 속성을 긍정성 척도와 부정성 척도를 합해서 합이 동일하게 열의 범주수를 2배로 늘여서 대응분석을 적용할 수 있다. 이런 기법을 더블링 기법이라 하는데, 본 연구에서는 행 합계가 동일하지 않은 리커트 척도로 평가된 자료에 대해 기존의 경우 더블링 기법을 통해 대응분석을 수행하기는 하였지만 그 수리적 체계가 미흡하였던 부분을 체계적으로 정리하였고 사례를 통해 이에 대한 활용성을 보이고자 하였다.Correspondence analysis is a multivariate data description technique that allows a two-dimensional or more contingency table to be expressed in a low-dimensional space to see the pattern of combining rows and columns. The statistical method commonly applied to partition tables is the chi-square test, which can verify the homogeneity of row samples or the independence of rows and columns with respect to the relative frequency of column categories. However, chi-square verification does not show the combination of rows and columns. In the corresponding analysis, each row profile element has a sum of one. However, for each object, the data for which various attributes are evaluated on the basis of the Likert scale can not be directly analyzed because the row sum is different for each row. For these data, a corresponding analysis can be applied by doubling the number of categories of columns with the same sum of positive and negative values for each attribute. In this study, we used the doubling technique to perform the correspondence analysis on the data that were evaluated as the Likert scale in which the row totals are not the same. And to demonstrate their usefulness through case study.
유권자들의 올바른 투표행위에 기여하기 위하여 또는 후보나 정당의 적절한 선거전략 수립을 위하여, 선거여론조사를 통하여 신뢰성 있고 객관적인 정보를 확보하는 것은 매우 중요한 문제이다. 따라서 정당, 언론기관, 조사회사 등 관련 기관에서는 여론조사의 결과와 선거예측의 정확도 향상을 위해 지속적으로 노력해 왔다. Kim et al.(2017)에서는 선거여론조사에서 지지후보가 없다고 응답한 무응답층을 분류하여 득표율 예측의 정확도를 높일 수 있는지를 분석하였는데, 결과적으로 무응답층에 대하여 적절한 분류를 수행함으로써 득표율 추정의 정확도를 상당히 높일 수 있음을 확인한 바 있다. 본 연구에서는 특정 선거구(지역)에 대하여 전체 투표율이 주어져 있다는 조건 하에서 각 층(성, 연령대)별 투표율을 추정하는 방안을 제안하고, 투표율을 반영하여 득표율을 예측하는 절차를 제시하였다. 또한 2016년 20대 국회의원선거에 대한 여론조사에서 전화면접조사를 통해 얻어진 자료를 사용하여 사례 분석을 수행하였다.It is very important to obtain objective and credible information through election polls in order to contribute to the correct voting behavior of the voters or to establish appropriate election strategies for candidates or political parties. Therefore, many related organizations such as political parties, media organizations, and research institutions have been making efforts to improve the accuracy of the results of the polls and the election prediction. Kim et al. (2017) analyzed whether the non-response group responded that there is no support candidate in the election survey to increase the accuracy of the estimation of the vote rate. As a result, it has been confirmed that the accuracy of the estimation of the vote rate can be significantly improved by performing an appropriate classification on the non-response layer. In this study, we propose a method to estimate the turnout by each strata (sex, age group) under the condition that the total turnout rate is given for a specific district (region) and propose a procedure to predict the vote rate by reflecting the turnout. In addition, case studies were conducted using data gathered through telephone interviews for the 20th National Assembly elections in 2016.
It is one of the hardest challenges to predict the movement of the stock price. We propose the modified bootstrap method in random forests to predict the direction of movement of the stock index price. The training set generated by the modified bootstrapping considers the impact of response variable simultaneously and is applied in random forests. The real KOSPI data are used for the experiments and the result shows that the proposed method performs better than the original method in various situations.
PURPOSE:The purpose of this study was to introduce the main concepts of statistical testing and effect size and to provide researchers in nursing science with guidance on how to calculate the effect size for the statistical analysis methods mainly used in nursing.METHODS:For t-test, analysis of variance, correlation analysis, regression analysis which are used frequently in nursing research, the generally accepted definitions of the effect size were explained.RESULTS:Some formulae for calculating the effect size are described with several examples in nursing research. Furthermore, the authors present the required minimum sample size for each example utilizing G*Power 3 software that is the most widely used program for calculating sample size.CONCLUSION:It is noted that statistical significance testing and effect size measurement serve different purposes, and the reliance on only one side may be misleading. Some practical guidelines are recommended for combining statistical significance testing and effect size measure in order to make more balanced decisions in quantitative analyses.
Sensitivity analysis is to study the influence of a small change in the input data on the output of the analysis. Han and Huh (1995) developed a quantification method for the ranked data. However, the question of stability in the analysis of ranked data has not been considered. Here, we propose a method of sensitivity analysis for ranked data. Our aim is to evaluate perturbations by using a graphical approach suggested by Han and Huh (1995). It extends the results obtained by Tanaka (1984) and Huh (1989) for the sensitivity analysis in Hayashi’s third methodof quantification and those byHuh and Park (1990) for the principal component reduction of the case influence derivatives in regression. A numerical example is provided to explain how to conduct sensitivity analysis basedon the proposed approach.