
Overconfidence refers to the overestimation of one’s own abilities, and it is generally measured as the difference between the expected (or subjective) and actual results in a task. We propose an alternative way to measure overconfidence using psychometric tools. Participants are asked to answer a Multiple Choice Test (MCT), and the data obtained are used to fit an Item Response Theory (IRT) model which allows us to take into account test features such as the degree of difficulty and the discrimination power of each item in the test. These features, which are not even considered in the classical measures of overconfidence, open up the opportunity to provide a more reliable measure of overconfidence.
The Basic Local Independence Model (BLIM) is the most frequently used probabilistic model in knowledge structure theory. One of its fundamental assumptions is that the item-specific probabilities of a lucky guess and a careless error are constant across individual knowledge states. This assumption has been challenged on theoretical and empirical grounds. The development of a generalized yet parsimonious model can be founded upon measuring how far an item is from a knowledge state. The paper examines such a previously suggested item-state discrepancy based on the notions of inner and outer layer, and introduces a novel kind of discrepancy induced by a canonical metric on the knowledge states. This so-called minimum discrepancy reflects the minimal length of a possible learning path. The discrepancies are compared theoretically, and conditions for their equality are provided. The advantages that minimum discrepancy offers for generalizing the BLIM are discussed.
We consider the widely used Mazur hyperbolic delay discounting model. Our scientific contributions are two-fold. First, we provide a novel linearized version of the Mazur model that leads to an analytic estimator of the delay discount rate. Our novel estimator bypasses an outstanding problem: the instability in the numerical optimization to compute the nonlinear estimator. As a result, a simulation study shows that when compared to the conventional nonlinear estimator, our novel estimator more accurately estimates the discount rate. Second, we propose a linearized Mazur random effects model and develop a related F-test for the comparison of the impact of different experimental conditions on the delay discount rate. Because our approach does not suffer from numerical instability, all observations can be used in the computation of the F-statistic. As a result, when compared to the conventional approach, our F-test provides much higher statistical power to detect nonnull effect sizes. Analysis of a previously published dataset examining the effects of a brief intervention to reduce delay discounting provides empirical support for the utility of our proposed linearized methods.
Artificial neural networks are widely used in modeling sentence processing but often exhibit overconfident point estimates in their output behavior, even in the presence of ambiguous or conflicting linguistic cues. This limitation is illustrated by reversal anomalies—sentences with unexpected role reversals inducing a conflict between syntactic and semantic information, and has been observed in established models such as the Sentence Gestalt (SG) model. To address this, we introduce a Bayesian formulation of the SG model by applying an extension of the ensemble Kalman filter for Bayesian inference at the level of model parameters. Framing sentence comprehension as a Bayesian inverse problem allows us to characterize posterior predictive uncertainty in the model’s output representations, rather than relying on point estimates. Through numerical experiments and comparisons with a standard maximum likelihood-trained SG model, we show that the Bayesian approach yields systematically less overconfident output activations under cue conflict, reflecting increased uncertainty when processing linguistic ambiguities.
Networks with nodes (variables) that are either on or off (Ising model) have been of great value to psychopathology. However, the dynamics of such networks (evolution over time) has been less investigated. We assert that one of the issues with the dynamics on networks is the difficulty in obtaining qualitative features, like what ratio of on and off nodes will the network end up with after some time. These questions are relevant to psychopathology since the sum of variables (often a proxy for symptoms) provides an indication of the severity of the disorder. Here, we propose a framework to establish long term features of networks. The framework allows for the translation of some phenomena of multiple variables to dynamical systems, which has been a very fruitful area for psychology. We show that the approximations of the framework (i.e., Markov chains and dynamical systems) are accurate. Furthermore, we illustrate with a time series from a patient diagnosed with depression some of the properties that can be investigated by reducing networks to dynamical systems. One of those properties, which we borrow from weather predictions, is statistical regularity for different initial conditions of the process to see where it ends up.
This work presents a comprehensive theoretical framework for polytomous knowledge structures by establishing the fundamental equivalence of three core constructs: polytomous learning spaces, well-graded polytomous knowledge spaces, and polytomous antimatroids. We introduce axioms for learning smoothness and consistency, formulate a polytomous metric based on the rank function on nontrivial finite modular lattices, and generalize accessibility properties to item-specific response scales to rigorously prove this tripartite equivalence under the assumptions that Q is a nonempty finite set of items and each response scale Vp (p is an element of Q) is a nontrivial finite modular lattice. Thus, our work provides a coherent theoretical basis for advancing polytomous cognitive assessment and adaptive learning.
Cultural consensus theory (CCT), developed by Batchelder and colleagues in the mid-1980s, is a cognitively driven methodology to assess informants’ consensus in which the culturally correct answers are unknown to researchers a priori. The primary goal of CCT is to uncover the cultural knowledge, preferences or beliefs shared by group members. One of the CCT models, calledthe general Condorcet model (GCM), deals with dichotomous (e.g., true/false) response data which are collected from a group of informants who share the same cultural knowledge. We propose a new model, called the general Condorcet-Luce-Krantz (GCLK) model, which incorporates the GCMwith the Luce-Krantz threshold theory. The GCLK accounts for ordinal categorical data (including Likert-type questionnaires) in which informants can express confidence levels when answering the items/questions. In addition to finding out the consensus truth to the items, the GCLK also estimates other response characteristics, including the item-difficulty levels, informants’ competency levels, and guessing biases. We introduce the multicultural version of the GCLK that can help researchers detect the number of cultures for a given dataset. We use the hierarchical Bayesian modeling approach andthe Markov chain Monte Carlo sampling method for estimation. A posterior predictive check is established to test the central assumptions of the model. Through a series of simulations, we evaluate the model’s applicability and find that the GCLK performs well on parameter recovery.
Common hypotheses in psychiatric research regard relationships between latent variables such as risk-taking and ambiguity attitudes, alongside psychiatric symptoms. The prevailing methodological approach is to estimate the latent variables in a first step, and in a second step, use these estimates in the statistical analysis. Using estimates instead of their true values can lead to bias and reduced power. This study aims to develop a new hierarchical method, based on a Laplace-based variational approximation, to mitigate these issues. The developed method outperforms the prevailing method in terms of power, accuracy and out-of-sample prediction. An empirical example demonstrates the usefulness of the method. The developed approach is computationally efficient and suitable to apply to many data types.
In the past years, several theories for assessment have been developed within the overlapping fields of Psychometrics and Mathematical Psychology. The most notable are Item Response Theory (IRT), Cognitive Diagnostic Assessment (CDA), and Knowledge Structure Theory (KST). In spite of their common goals, these theories are largely independent and focus on different aspects. In Part I of this work, a general framework was introduced that, by means of two primitives (structure and process) and two operations (factorization and reparametrization), allows to derive the models of these theories. A two-processes approach provided a taxonomy encompassing IRT, CDA, and KST models. In Part II, the framework was used to derive KST and CDA models based on dichotomous latent variables. In this third contribution, IRT, CDA, and KST models based on polytomous and continuous latent variables are derived. Implications for IRT models are discussed. A special focus is on the implicit assumptions required to move from the discrete case to the continuous one.
Competence-based extensions of polytomous knowledge space theory have recently seen significant development, primarily through specialized conjunctive or disjunctive models. The conjunctive attribute map proposed by Stefanutti et al. (2023) assumes that a single, conjunctive set of attributes is required for each item-response, while other models, such as that of Wang et al. (2025), have explored purely disjunctive mechanisms. However, these frameworks do not readily account for scenarios where a given response can be justified by one of several distinct sets of attributes—a structure analogous to the competency model in dichotomous theory. This paper introduces a generalized attribute function which maps each item-response pair to a collection of attribute sets. This framework naturally subsumes the conjunctive model as the special case where this collection is a singleton. We establish a rigorous axiomatic foundation for this function, specified by several general conditions. We then derive the corresponding item-response function and the maximal polytomous assessment structure. We show that while our framework generalizes the conjunctive model, its disjunctive specialization is structurally distinct from other disjunctive approaches. The utility of this generalized model for capturing complex reasoning paths is demonstrated with a detailed example.
Evidence accumulation models (EAMs) have dominated theoretical accounts of speeded decision-making for decades, with a proliferation of different variants that each provide different theoretical accounts of how the decision-making process operates. However, these variants often show strong mimicry in their predictions about empirical data, hindering the ability to contrast these accounts and make further theoretical progress. Evans et al. (2020a) proposed double responding (DR) - where participants make a second response after their initial response - as an additional constraint to help distinguish between theoretical accounts, and found that models with lateral inhibition provided a superior account of the data to models with feed-forward inhibition or no inhibition. However, the study of Evans et al. (2020a) had two key limitations: (1) the focus on "implicit'' double responding behaviour, where participants were not told that they could change their response, and (2) the inability of any model to explain the positively skewed DR response time (DRRT) distributions. Here, we provide a conceptual replication of Evans et al. (2020a), where we explicitly instructed participants that they could change their mind, and included a leaky integrating threshold (LIT) framework to attempt to explain the positively skewed DRRT distributions. As in Evans et al. (2020a), models with lateral inhibition consistently provided a superior account of the data relative to models with feed-forward inhibition and without inhibition, mostly in terms of predictions about DR proportions. However, in contrast to Evans et al. (2020a), we found that the models - even without the LIT framework - were able to predict positively skewed DRRT distributions, with these contrasting findings likely influenced by the shorter DR deadline in Evans et al. (2020a). Our findings provide further evidence that lateral inhibition is crucial to explaining changes of mind in decision-making, and that DRs can serve as an important extra constraint in distinguishing between competing decision-making models.
Successful human-AI collaboration requires accurate uncertainty communication. While prior research emphasizes calibration, we demonstrate that metacognitive sensitivity-the ability of confidence scores to discriminate between correct and wrong decisions-is equally important. We derive the Bayes-optimal decision rule for combining human and AI predictions and use signal detection theory to analytically express combined accuracy based on human and AI metacognitive sensitivities. We prove that above-chance metacognitive sensitivity in either agent guarantees complementarity, where joint accuracy exceeds both individual accuracies. Monte Carlo simulations and empirical validation on human-AI image classification demonstrate that our analytic solution remains robust to non-Gaussian confidence distributions. These findings establish metacognitive sensitivity as an important determinant of human-AI collaboration.
Response time is probably the most important metric in cognitive psychology. Many studies focus on average response time, but even more information can be obtained considering the shape of the response time distributions. In their seminal work, Townsend and Nozawa (1995) investigated the interaction contrast of response time distributions in two-factorial experiments for different cognitive architectures, including serial and parallel processing, with exhaustive and self-terminating stopping rules. They derived distinct, nonparametric predictions for the interaction contrast under fairly weak assumptions: selective influence of the factorial manipulations, and stochastic ordering of the processing times for the different factor levels. Their original theory is limited to tasks with ceiling accuracy. Extensions to more difficult tasks either conditioned the response time distributions on accuracy, which interferes with the assumption of stochastic ordering. Other approaches focused on a special case of the paradigm (i.e., redundant signals tasks). Here we show that with a slight extension of the stochastic dominance assumption, Townsend and Nozawa’s theorems can be generalized to difficult tasks that entail non-negligible error rates. We demonstrate statistical tests and illustrate methods to assess their power. We investigate additional generalizations of the theory to higher-order experimental manipulations and larger networks of serial and parallel processes. Interesting special cases such as redundant signals tasks and parametric experimental variations, as well as applications of the theory in areas outside cognitive psychology are proposed and discussed. More than 150 years after Donders, we can finally design and analyze interesting reaction time experiments with non-negligible error rates.
This article presents a unified information-theoretic model of decision-making in the Shafir-Tversky Prisoner's Dilemma experiment. By combining the Principle of Maximum Entropy for prior beliefs and the Free Energy Principle for action selection, the model reproduces the observed pattern of cooperation and high willingness to pay for information reported by Shafir and Tversky (1992). The analysis demonstrates how epistemic value, captured as expected information gain, drives behaviour beyond classical utility theory. Explicit definitions and parameter choices ensure reproducibility and clarity. Results highlight the explanatory power of information-theoretic approaches in social dilemmas under uncertainty.
This paper introduces the BF-Poly-H algorithm for computing Bayes factors for binomial models that are characterized by arbitrary linear constraints. BF-Poly-H combines the simulated annealing-based integration technique of Lov & aacute;sz and Vempala (2006a) with Constrained Riemannian Hamiltonian Monte Carlo (CRHMC) sampling. This approach estimates the Bayes factor via a sequence of distributions that gradually anneals from a uniform distribution to the target posterior. CRHMC efficiently navigates the geometry of the constrained space at each step in the annealing process. We provide analytical results on scalability and convergence of the algorithm, alongside a suite of simulations demonstrating its performance across a broad range of scenarios. BF-Poly-H delivered accurate and precise results in all scenarios we tested. The most challenging scenario is a complicated combination of inequality and equality constraints among 100 binomials. By offering a robust solution for assessing the performance of high-dimensional and elaborately constrained binomial models, BF-Poly-H substantially expands the scope of feasible Bayesian hypothesis testing.
Nonlinear relationships among variables play an important role in psychological modeling and understanding changes over time from intensive longitudinal data (ILD). Most methods focus on linear relationships, with a few exceptions developed for specific nonlinear interactions or general system dynamics. Methods considering multiple possible nonlinear relationships among all variables remain rare, hindered by challenges like overfitting and interpretation difficulties. This article examines the feasibility of applying a quadratic vector autoregression method to psychological ILD, using the Regularization Algorithm under Marginality Principle (RAMP) alongside a local linearization method for interpretation. We evaluated its performance with simulated and empirical datasets using classification metrics, information criteria, and cross-validation. Results show this method often requires a long time series for satisfactory performance. While information criteria favor the quadratic model in empirical datasets, cross-validation favors simpler AR models. Nonetheless, these two challenges are comparable to those in idiographic linear VAR models. Clear evidence of nonlinear relationships among variables supports the value of this method for exploratory studies. We developed an R package, quadVAR, as an implementation of this method.
We provide a correction to Davis-Stober et al. (2015) “The algebraic structure of individual preferences”. In the original paper, the Bayes factors for the Weak Order Mixture Model (WOMM) against an unconstrained, encompassing model were calculated using an under-estimate of the Euclidean volume occupied by WOMM. The correct estimate of the volume used in the Bayes factor calculations should be .000453, not .0000453, as was used in the original analysis. This computational error caused the distribution of WOMM classifications to be slightly too high in the original paper. We provide updated model comparison tables. The main empirical results of the original paper still hold.
We evaluate the performance of the Bayes factor under model misspecification, with a focus on violations of compound symmetry and deviations from the normality assumption in repeated-measures designs. Simulation studies reveal that a paradox—discrepancies between inferences drawn from classical repeated-measures analysis of variance (rANOVA) and Bayesian analyses using Bayes factors with random intercept-only models—can arise under certain model misspecifications. We argue that the underlying misspecifications are best viewed as manifestations of covariance heterogeneity, rather than the existence of individual differences (random slopes) in the effects of independent variables. Our results demonstrate that models with both random intercepts and slopes effectively mitigate such discrepancies. We elucidate the rationale for adopting random-slope models as a means of accommodating a more complex marginal covariance structure than is assumed by a random intercept-only model. This study supports recent recommendations advocating the use of random intercept and slope models in analyzing repeated-measures data and explains the theoretical foundations of their effectiveness. We additionally substantiate that Bayes factors are generally robust to misspecifications in the distributions of random effects and residuals in linear mixed-effects models. We present a theorem demonstrating model selection consistency of the Bayes factor under model misspecification for the simple case of a one-sample t-test.
Associative learning of events’ co-occurrence rates across time can generate useful predictive representations (e.g., the successor representation), but temporal contiguities alone are not enough to infer causal relations. Recent work suggests that neural substrates long thought to implement temporal difference learning may perform causal inference by tracking temporal contingences — coincidences between events corrected by background co-occurrence rate. We show that changing the activation function enables simple associative updates to directly compute temporal contingencies. Temporal contiguities can be learned as the cross-correlation between a stimulus and a memory trace, and temporal contingencies can be learned as the cross-covariance. An implication is that neurally plausible causal learning algorithms can be implemented through simple associative updates. These results highlight a family of learning rules for incremental computation of forward, backward, and joint temporal contiguities and contingencies.
Falmagne's representation problem is revisited by maximizing Shannon entropy applied to ranking probabilities, under the linear constraints imposed by choice probabilities. Unlike Falmagne's recursive construction, our method leads directly to an explicit solution, obtained after transforming the initial system into an equivalent one via alternating sums, in the spirit of Block-Marschak polynomials. We compute this solution for the Luce model and the generalized extreme value model, and show that, as soon as there are at least four alternatives, the construction based on Shannon entropy is only one among infinitely many possible representations. Other solutions could be obtained by maximizing alternative entropy functions, further highlighting the potential role of information theory in enriching the analysis of stochastic choice.