PurposeThe purpose of this paper is to create an automatic interpretation of the results of the method of multiple correspondence analysis (MCA) for categorical variables, so that the nonexpert user can immediately and safely interpret the results, which concern, as the authors know, the categories of variables that strongly interact and determine the trends of the subject under investigation.Design/methodology/approachThis study is a novel theoretical approach to interpreting the results of the MCA method. The classical interpretation of MCA results is based on three indicators: the projection (F) of the category points of the variables in factorial axes, the point contribution to axis creation (CTR) and the correlation (COR) of a point with an axis. The synthetic use of the aforementioned indicators is arduous, particularly for nonexpert users, and frequently results in misinterpretations. The current study has achieved a synthesis of the aforementioned indicators, so that the interpretation of the results is based on a new indicator, as correspondingly on an index, the well-known method principal component analysis (PCA) for continuous variables is based.FindingsTwo (2) concepts were proposed in the new theoretical approach. The interpretative axis corresponding to the classical factorial axis and the interpretative plane corresponding to the factorial plane that as it will be seen offer clear and safe interpretative results in MCA.Research limitations/implicationsIt is obvious that in the development of the proposed automatic interpretation of the MCA results, the authors do not have in the interpretative axes the actual projections of the points as is the case in the original factorial axes, but this is not of interest to the simple user who is only interested in being able to distinguish the categories of variables that determine the interpretation of the most pronounced trends of the phenomenon being examined.Practical implicationsThe results of this research can have positive implications for the dissemination of MCA as a method and its use as an integrated exploratory data analysis approach.Originality/valueInterpreting the MCA results presents difficulties for the nonexpert user and sometimes lead to misinterpretations. The interpretative difficulty persists in the MCA's other interpretative proposals. The proposed method of interpreting the MCA results clearly and accurately allows for the interpretation of its results and thus contributes to the dissemination of the MCA as an integrated method of categorical data analysis and exploration.
Defining a distance in a mixed setting requires the quantification of observed differences of variables of different types and of variables that are measured on different scales. There exist several proposals for mixed variable distances, however, such distances tend to be biased towards specific variable types and measurement units. That is, the variable types and scales influence the contribution of individual variables to the overall distance. In this paper, we define unbiased mixed variable distances for which the contributions of individual variables to the overall distance are not influenced by measurement types or scales. We define the relevant concepts to quantify such biases and we provide a general formulation that can be used to construct unbiased mixed variable distances.
Cluster analysis relates to the task of assigning objects into groups which ideally present some desirable characteristics. When a cluster structure is confined to a subset of the feature space, traditional clustering techniques face unprecedented challenges. We present an information-theoretic framework that overcomes the problems associated with sparse data, allowing for joint feature weighting and clustering. Our proposal constitutes a competitive alternative to existing clustering algorithms for sparse data, as demonstrated through simulations on synthetic data. The effectiveness of our method is established by an application on a real-world genomics data set.
Multidimensional and multivariate datasets encompassing diverse data types provide researchers a platform to apply a range of dimensionality reduction methods. This study assessed principal components analysis, factor analysis, multiple correspondence analysis, categorical principal components analysis and factor analysis for mixed data. We examined different strategies based on input variable measurement scale selection and variable value coding. The objectives were to highlight the importance of applying different analysis strategies, ascertain the applicability of these methods to multidimensional mixed-type data, compare outcomes, and evaluate execution times from three different statistical software to identify notable computational and interpretive drawbacks. Significant issues included the 'curse of dimensionality' concerning the determination of crucial dimensions, the need for increased computing power, the absence of software code for some methods and criteria, discrepancies in results' calculations across software packages, and the inability of some software packages to handle numerous variables or binary-coded variables and perform parallel analysis.
In this paper, we present an information-theoretic method for clustering mixed-type data, that is, data consisting of both continuous and categorical variables. The proposed approach extends the Information Bottleneck principle to heterogeneous data through generalised product kernels, integrating continuous, nominal, and ordinal variables within a unified optimisation framework. We address the following challenges: developing a systematic bandwidth selection strategy that equalises contributions across variable types, and proposing an adaptive hyperparameter updating scheme that ensures a valid solution into a predetermined number of potentially imbalanced clusters. Through simulations on 28,800 synthetic data sets and ten publicly available benchmarks, we demonstrate that the proposed method, named DIBmix, achieves superior performance compared to four established methods (KAMILA, K-Prototypes, FAMD with K-Means, and PAM with Gower’s dissimilarity). Results show DIBmix particularly excels when clusters exhibit size imbalances, data contain low or moderate cluster overlap, and categorical and continuous variables are equally represented. The method presents a significant advantage over traditional centroid-based algorithms, establishing DIBmix as a competitive and theoretically grounded alternative for mixed-type data clustering.
Understanding genotype × environment (G × E) interaction is essential for the improvement of aromatic crops such as basil (Ocimum basilicum), where yield is strongly influenced by environmental variability. In this study, five basil varieties (Burns Lemon, Cinnamon, Sweet, Red Rubin, and Thai) were evaluated across two years (2015–2016, 2016–2017) and three irrigation levels (40%, 70%, and 100% of the full water requirement) to assess dry biomass yield. ANOVA and mean performance plots confirmed significant varietal differences (F = 33.972, p < 0.001) and substantial Y × V interaction (F = 23.578, p < 0.001), motivating a deeper exploration of association patterns. To this end, we proposed the use of a modified version of Simple Correspondence Analysis (CA) combined with three variations of bi-plot analyses in order to explore the (G × E) interaction. In addition, another modification of CA is proposed and used, CA of raw data (CA-raw), for the same reason. For the purpose of the study, the combinations of the two cultivation periods (years) by the three irrigation levels were considered as six environments. Results showed that the proposed modification of CA of raw data serves as a faithful baseline for the study of (G × E) interaction. On the other hand, the proposed modified version of simple CA, after proper normalization (row, column, symmetrical, principal) of the factorial scores of the five basil varieties and the six environments, provide insights depending on whether the research focus lies on varieties, environments, or their joint associations (interaction). Overall, the combined use of ANOVA, mean plots, and CA under multiple normalizations and modifications demonstrated the robustness of the primary varietal–environment contrast, while also showing how methodological choices shape interpretation. The proposed methods are “model free” and can be used also with secondary published data.
In this study, we examined the potential of integrating multivariate data analysis methods as a preliminary stage for machine learning techniques to augment their predictive power. These methods encompass principal component analysis, multiple correspondence analysis, and non-linear categorical principal component analysis with optimal scaling. The machine learning approaches evaluated include Support Vector Machines, Stochastic Gradient Descent, Naïve Bayes, K-Nearest Neighbor, Decision Trees, Random Forests, Adaptive Boosting, and Multinomial Logistic Regression. We conducted experiments using data from a nationwide survey, comprising a total sample of 42,593 adolescents who answered more than 155 questions related to their eating habits. The dependent variable, body mass index (BMI), was measured and employed in the analysis as both a quantitative and qualitative variable. The index values were initially classified based on the World Health Organization’s recommendations. The results indicated that predictions are more reliable when utilizing the BMI as a qualitative variable within a four-class structure. Implementing a multivariate data analysis strategy before applying machine learning algorithms not only conserves time but also facilitates the selection of the most effective predictive model. Although dimensionality reduction may not consistently enhance the models’ predictive abilities, it contributes to the "interpretability" of the results. Received: 22 July 2024 | Revised: 23 September 2024 | Accepted: 2 January 2025 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement Data available on request from the corresponding author upon reasonable request. Author Contribution Statement Nikolaos Papafilippou: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – original draft, Writing – review & editing, Visualization. Zacharenia Kyrana: Resources. Emmanouil Pratsinakis: Conceptualization. Efstratios Kiranas: Conceptualization, Resources. Alexandra-Maria Michailidou: Resources. Angelos Markos: Conceptualization, Methodology, Software, Formal analysis, George Menexes: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – review & editing, Supervision, Project administration.
In the social sciences, accurately identifying the dimensionality of measurement scales is crucial for understanding latent constructs such as anxiety, happiness, and self-efficacy. This study presents a rigorous comparison between Parallel Analysis (PA) and Exploratory Graph Analysis (EGA) for assessing the dimensionality of scales, particularly focusing on ordinal data. Through an extensive simulation study, we evaluated the effectiveness of these methods under various conditions, including varying sample size, number of factors and their association, patterns of loading magnitudes, and symmetrical or skewed item distributions with assumed underlying normality or non-normality. Results show that the performance of each method varies across different scenarios, depending on the context. EGA consistently outperforms PA in correctly identifying the number of factors, particularly in complex scenarios characterized by more than a single factor, high inter-factor correlations and low to medium primary loadings. However, for datasets with simpler and stronger factor structures, specifically those with a single factor, high primary loadings, low cross-loadings, and low to moderate interfactor correlations, PA is suggested as the method of choice. Skewed item distributions with assumed underlying normality or non-normality were found to noticeably impact the performance of both methods, particularly in complex scenarios. The results provide valuable insights for researchers utilizing these methods in scale development and validation, ensuring that measurement instruments accurately reflect theoretical constructs.
PurposeThe purpose of this paper is to develop a software-library in the R programming language that implements the concepts of the interpretive coordinate, interpretive axis and interpretive plane. This allows for the automatic and reliable interpretation of results from the multiple correspondence analysis (MCA) as previously proposed and published. Consequently, the users can seamlessly apply these concepts to their data, both via R commands and a corresponding graphical interface.Design/methodology/approachWithin the context of this study, and through extensive literature review, the advantages of developing software using the Shiny library were examined. This library allows for the development of full-stack applications for R users without the need for knowledge of the corresponding technologies required for the development of complex applications. Additionally, the structural components of a Shiny application were presented, leading ultimately to the proposed software application.FindingsSoftware utilizing the Shiny library enables nonexpert developers to rapidly develop specialized applications, either to present or to assist in the understanding of objects or concepts that are scientifically intriguing and complex. Specifically, with this proposed application, the users can promptly and effectively apply the scientific concepts addressed in this study to their data. Additionally, they can dynamically generate charts and reports that are readily available for download and sharing.Research limitations/implicationsThe proposed package is an implementation of the fundamental concepts of the exploratory MCA method. In the next step, discoveries from the geometric data analysis will be added as features to provide more comprehensive information to the users.Practical implicationsThe practical implications of this work include the dissemination of the method's use to a broader audience. Additionally, the decision to implement it with open-source code will result in the integration of the package's functions by other third-party user packages.Originality/valueThe proposed software introduces the initial implementation of concepts such as interpretive coordination, the interpretive axis and the interpretive plane. This package aims to broaden and simplify the application of these concepts to benefit stakeholders in scientific research. The software can be accessed for free in a code repository, the link to which is provided in the full text of the study.
This study investigates the impact of an explicit and integrated dictionary awareness program on primary school pupils' dictionary use strategies. The survey involved a total of 150 participants, aged 10–12 years old, from mainstream and intercultural schools. Data was collected before and after the implementation of the program using the Strategy Inventory for Dictionary Use (SIDU), a reliable and validated self-report tool that accurately profiles paper dictionary users' reported use in real-life contexts (Gavriilidou 2013). The dictionary awareness program consisted of targeted activities and was implemented to a group of 75 students, including 50 from mainstream schools and 25 from an intercultural school. The findings suggest that there is a lack of dictionary culture among students attending Greek schools, as evidenced by the moderate strategic use of dictionaries and the incomplete integration of dictionaries as reference tools in the educational process. Additionally, the comparison of the percentage of each strategy category before and after the implementation of the program showed a significant effect of the program on all categories of Dictionary Use Strategies (DUS) employed by the experimental group. This study contributes to the discussion of the "teachability" of dictionary use strategies by highlighting the effectiveness of dictionary awareness programs in promoting a dictionary culture. Keywords: dictionary use strategies, dictionary awareness program, explicit and integrated strategy instruction, dictionary culture, CALLA, strategy based instruction, look up strategies, lemmatisation strategies
The degree to which objects differ from each other with respect to observations on a set of variables, plays an important role in many statistical methods. Many data analysis methods require a quantification of differences in the observed values which we can call distances. An appropriate definition of a distance depends on the nature of the data and the problem at hand. For distances between numerical variables, there exist many definitions that depend on the size of the observed differences. For categorical data, the definition of a distance is more complex as there is no straightforward quantification of the size of the observed differences. In this paper, we introduce a flexible framework for efficiently computing distances between categorical variables, supporting existing and new formulations tailored to specific contexts. In supervised classification, it enhances performance by integrating relationships between response and predictor variables. This framework allows measuring differences among objects across diverse data types and domains.
Analyzing data about different aspects of human behavior is a valuable and challenging task and it represents the general aim of behavioral data science.Both behavioral science and data science are interdisciplinary fields.In particular, behavioral science ranges from economics and finance to psychology and sociology, up to health-related behaviors; similarly, under the umbrella of data science are statistics, machine learning, computer science, to name a few.Therefore, behavioral data science requires a faceted approach, where specific knowledge domains must drive the choice and the development of methodological tools: while this is an important driver in any scientific field, in behavioral data science it becomes mandatory.When the data analysis goal is to understand human behavior, obtaining explanations is key, no black boxes can be used or trusted.Domain specific knowledge is of course essential throughout the learning pipeline, from pre-processing and feature engineering to the interpretation of the results.However, it is at the same time important to develop tools that support domain experts and ease their interpretation of the results: visualization tools can be crucial in this respect.This special issue features a range of papers that fit the above description: the contributions explore complex behaviors from multiple angles and show how behavioral data science blends elements from finance, education, healthcare, sociology, and text analysis, all through the lens of sophisticated data science methods, with a special focus on the interpretability of the results.A rough taxonomy of the articles in the special issue is: (i) contributions that tailor methods to specific application fields; (ii) contributions that present methodological enhancements to extend applicability and explainability of specific data science methods.
The primary objective of this study is to contribute to the conservation and sustainable use of seas by promoting Ocean Literacy. It investigates the impact of an educational program on Greek primary and secondary public school students' knowledge about coastal lagoons and attitudes towards marine environment conservation. An educational resource titled “Exploring the Coastal Lagoons” was developed to facilitate the non-formal educational intervention. The program involved classroom, fieldwork/outdoor and laboratory activities, focusing on enhancing understanding of coastal lagoons' abiotic and biotic characteristics and human interconnection. Results showed improved knowledge and slightly more positive attitudes after the didactic intervention. The study underlines the effectiveness of targeted educational interventions in marine sciences, suggesting that non-formal educational settings influence student outcomes more than family or informal sources. Younger students appeared more adaptable and responsive to educational stimuli. The study advocates for refined educational strategies integrating cognitive and emotional elements, emphasizing real nature experience.
The purpose of this paper is to investigate linguistic factors affecting the evaluation of the argumentative essays in written tests taken by junior and senior students, aged 16 to 18, attending high schools in Greece. To achieve this, we analyzed textual characteristics and scoring of 265 juniors and seniors, graded by 15 different raters. To examine the contribution of linguistic parameters to the assessment, we developed an automated tool to record and evaluate students' lexical and syntactic features in the Greek language. The results revealed that the extensive use of nominal groups including an adjective and a noun and the utilization of both impersonal and passive syntax, as well as adverbs to a lesser extent, contribute the most to positive grading in language tests. Furthermore, we identified a correlation between language and the other criteria of the evaluation rubric, namely content and organization. The paper contributes to the discussion about objectivity in writing evaluation in the Greek setting and to the creation of a rubric that ensures a more effective assessment of writing tasks.
ChatGPT (GPT-3.5), an intelligent Web-based tool capable of conducting text-based conversations akin to human interaction across various subjects, has recently gained significant popularity. This surge in interest has led researchers to examine its impact on numerous fields, including education. The aim of this paper is to investigate the perceptions of undergraduate students regarding ChatGPT’s utility in academic environments, focusing on its strengths, weaknesses, opportunities, and threats. It responds to emerging challenges in educational technology, such as the integration of artificial intelligence in teaching and learning processes. The study involved 257 students from two university departments in Greece—namely primary and early childhood education pre-service teachers. Data were collected using a structured questionnaire. Various methods were employed for data analysis, including descriptive statistics, inferential analysis, K-means clustering, and decision trees. Additional insights were obtained from a subset of students who undertook a project in an elective course, detailing the types of inquiries made to ChatGPT and their reasons for recommending (or not recommending) it to their peers. The findings offer valuable insights for tutors, researchers, educational policymakers, and ChatGPT developers. To the best of the authors’ knowledge, these issues have not been dealt with by other researchers.
Cognitive diagnostic models (CDMs) are psychometric models developed to categorize individuals based on their latent attributes. These attributes, often discrete, possess scientific interpretations, such as skill mastery in educational assessments, mental disorders in psychiatric diagnoses, or the presence of disease pathogens in biological samples. Nonparametric CDMs directly classify subjects into latent profiles by minimizing the distance between observed item responses and the centroids of the latent profiles. Two widely used nonparametric methods are nonparametric classification (NPC) and general nonparametric classification (GNPC). However, existing nonparametric algorithms do not offer information regarding the variability of estimates of latent profile membership. This chapter introduces a resampling scheme employing out-of-bag observations to quantify uncertainty in estimates of attributes, latent profiles, and individuals. The proposed approach is illustrated using simulated data.
The study investigates secondary students’ understanding of “orbital” and “electron cloud” concepts in different quantum contexts (for values of the ʽprincipal quantum number n = 1 and n = 2) on the basis of their verbal and pictorial representations, evaluating also their consistency. Participants, which were 192 12th-grade students from six urban secondary schools of Northern Greece, represented these two concepts through two corresponding tasks of a paper-and-pencil assessment tool, each of which comprised two parts for verbal and pictorial representations, respectively. Results provide evidence that although students struggle to express verbally the orbital and electron cloud concepts, their competences in the corresponding pictorial representations are relatively better, exhibiting inconsistencies between verbal and pictorial representations. Inconsistencies also exist between representations of the orbital and electron cloud concepts, since students appear to have verbally a better understanding of the electron cloud than the orbital, whereas the opposite holds true for their pictorial representations. Comparing verbal and pictorial representations, the pictorial ones appear to be more consistent tools, whereas a quantum context defined by n = 2 seems to be more challenging for students compared to that of n = 1. Furthermore, an analysis of student profiles leads to their categorization in four classes, providing additional relevant information. Implications for science education are also discussed.