This paper introduces the BF-Poly-H algorithm for computing Bayes factors for binomial models that are characterized by arbitrary linear constraints. BF-Poly-H combines the simulated annealing-based integration technique of Lov & aacute;sz and Vempala (2006a) with Constrained Riemannian Hamiltonian Monte Carlo (CRHMC) sampling. This approach estimates the Bayes factor via a sequence of distributions that gradually anneals from a uniform distribution to the target posterior. CRHMC efficiently navigates the geometry of the constrained space at each step in the annealing process. We provide analytical results on scalability and convergence of the algorithm, alongside a suite of simulations demonstrating its performance across a broad range of scenarios. BF-Poly-H delivered accurate and precise results in all scenarios we tested. The most challenging scenario is a complicated combination of inequality and equality constraints among 100 binomials. By offering a robust solution for assessing the performance of high-dimensional and elaborately constrained binomial models, BF-Poly-H substantially expands the scope of feasible Bayesian hypothesis testing.
Chaotic responses to COVID-19, political polarization, and pervasive misinformation raise the question of whether some or many individuals exercise irrational moral judgment. We provide the first mathematically correct test for transitivity of moral preferences. Transitivity is the most prominent rationality criterion of the behavioral, biological, and economic sciences. However, transitivity is conceptually, mathematically, and statistically difficult to evaluate empirically. We tested three parsimonious, order-constrained, probabilistic characterizations: First, the weak utility model treats an individual's choices as noisy reflections of a single, deterministic, underlying transitive preference; second, a variant severely limits the allowable response noise; and third, by the general random utility hypothesis, individuals' choices reveal uncertain, but transitive, moral preferences. Among 28 individuals, everyone's data were consistent with the weak utility model and general random utility model, thus supporting both operationalizations. Tightening the bounds on error rates in noisy responses yielded a poorly performing model, thus rejecting the model according to which choices are highly consistent with a single transitive preference. Bayesian model selection favored probabilistic transitive preferences and hence the equivalent random utility hypothesis. This suggests that there is some order underlying the apparent chaos: Rather than presume widespread disregard for moral principles, policymakers may build on navigating and reconciling extreme heterogeneity compounded with individual uncertainty.
Classical random utility models imply a consistency property called regularity. Decision makers who satisfy regularity are at least as likely to choose an option x from a set X of available options as from any larger set Y that contains X. In light of ample empirical evidence for context-dependent choice that violates regularity, some researchers have questioned the descriptive validity of all random utility models. In this article, we show that not all random utility models imply regularity. We propose a general framework for random utility models that accommodate context dependence and may violate regularity. Mathematically, like the classical models, context-dependent random utility models form convex polytopes. They yield behavioral predictions for those choice sets from which choices are made, by specifying combinations of preference rankings across two or more contexts. We discuss how context-dependent models can be less or more parsimonious than the classical models. Random utility models with or without regularity can be tested with contemporary methods of order-constrained inference.
Transitivity of preference (ToP) is a central axiom of rational choice theory. While violating ToP is rare and subject to debate, there have been reports of such violations (Tsetsos et al., 2016a; Tversky, 1969, but see Iverson & Falmagne, 1985; Regenwetter et al., 2011). If humans indeed violate ToP, an important challenge is understanding the conditions and mechanisms that promote either transitive or intransitive preferences. Here, we report the presence of ToP violations using a data analysis method that was prescribed as statistically adequate (Regenwetter et al., 2011), and an experimental design where each choice is presented on a single visual display to avoid artifacts that can be associated with sequential presentation and aggregation across choice stimuli. We introduce two cognitive heuristics that predict certain violations of ToP and we translate them into probabilistic choice models. Then, in three experiments (one of which is a preregistered replication), we evaluate violations of ToP and we assess the models that predict such violations. We find that, despite pervasive individual differences, the ToP adherence rate is much enhanced when the task was presented in a fashion that facilitates within-alternative integration. We also find that the proposed heuristic models successfully explain those ToP violations that do occur. These findings shed light on the conditions and cognitive mechanisms that support ToP. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The present study examines the effect of social distance on choice behavior through the lens of a probabilistic modeling framework. In an experiment, participants made incentive-compatible choices between lotteries in three different social distance conditions: self, friend, and stranger. We conduct a layered, within-subjects analysis that considers four properties of preferential choice. These properties vary in their granularity. At the coarsest level, we test whether choices are consistent with transitive underlying preferences. At a finer level of granularity, we evaluate whether each participant is best described as having fixed preferences with random errors or probabilistic preferences with error-free choices. In the latter case, we further distinguish three different bounds on response error rates. At the finest level, we identify the specific transitive preference ranking of the choice options that best describes a person’s choices. At each level of the analysis, we find that the stability between the self and friend conditions exceeds that between the self and stranger conditions. Stability increases with the coarseness of the analysis: Nearly all people are consistent with transitive preferences regardless of the social distance condition, but only for very few do we infer the same preference ranking in every social distance condition. Overall, while it matters whether one makes a choice on behalf of a friend versus for a stranger, the differences are most apparent when analyzing the data at a detailed level of granularity.
Just as we formulate detailed theories of utility or preference, so too should we theorize carefully about strength of preference. Likewise, because behavior is inherently uncertain, we need a theoretical framework for understanding choice probabilities. This paper fleshes out the simple premise that more strongly preferred options are more likely to be chosen. The resulting distribution-free Fechnerian models (DFMs) eschew convenience assumptions underlying popular models like the logit and probit, revealing which aspects of a core decision theory do or do not remain invariant across different ways of constructing strengths of preference, as well as across different monotonic links between those strengths of preference and choice probabilities. We formulate DFMs in a unifying polyhedral geometric space that allows for direct comparisons of theories that can be as categorically different as, say, regret theory, expected utility theory, and lexicographic semiorders. The geometric representation also provides a nuanced perspective on theoretical parsimony beyond parameter counting. Through a series of examples, we demonstrate the derivation and mathematical characterization of DFMs for decision theories with and without utilities and the inferences one can draw from data. We show how DFMs provide a multi-layered quantitative approach to the identifiability of hypothetical constructs. We highlight specific cases where DFMs protect the researcher against mistaken conclusions caused by overspecified models.
A common approach to theory testing in behavioral and experimental economics relies on null hypothesis significance testing via (generalized) linear regression models. Here, we showcase order-constrained inference as an alternative route to theory testing. Order-constrained inference can improve the precision and nuance of behavioral decision analytics. For example, the method can be leveraged to quantify the evidence in support of, or against, a given hypothesis. It also offers advanced model selection tools for quantitative competition among multiple theories. To illustrate our case for order-constrained methods, we re-analyze data from a pre-registered experiment on incentives, cognitive reflection, and dishonest behavior. Building on this publicly available dataset, we further highlight the advantages of Bayesian order-constrained inference. We discuss how the method can deliver more convincing and more nuanced evidence than frequentist null hypothesis significance testing, pointing to new research avenues for supplementing and expanding on experimental designs in behavioral economics.
Mediation analysis investigates the covariation of variables in a population of interest. In contrast, the resolution level of psychological theory, at its core, aims to reach all the way to the behaviors, mental processes, and relationships of individual persons . It would be a logical error to presume that the population-level pattern of behavior revealed by a mediation analysis directly describes all, or even many, individual members of the population. Instead, to reconcile collective covariation with theoretical claims about individual behavior, one needs to look beyond abstract aggregate trends. Taking data quality as a given and a mediation model’s estimated parameters as accurate population-level depictions, what can one say about the number of people properly described by the linkages in that mediation analysis? How many individuals are exceptions to that pattern or pathway? How can we bridge the gap between psychological theory and analytic method? We provide a simple framework for understanding how many people actually align with the pattern of relationships revealed by a population-level mediation. Additionally, for those individuals who are exceptions to that pattern, we tabulate how many people mismatch which features of the mediation pattern. Consistent with the person-oriented research paradigm, understanding the distribution of alignment and mismatches goes beyond the realm of traditional variable-level mediation analysis. Yet, such a tabulation is key to designing potential interventions. It provides the basis for predicting how many people stand to either benefit from, or be disadvantaged by, which type of intervention.
Stylized characteristics, such as loss aversion, risk aversion for gains, risk seeking for losses, overweighting of small probabilities, and diminishing sensitivity permeate both popular science and scholarly treatises about how 'people' make decisions. This note highlights that behavioral decision research is, in effect, a large-scale Linda problem: The likelihood that a given individual satisfies the conjunction of many such stylized characteristics may be vanishingly small. We concentrate on a case study, namely the pervasive oversimplifications surrounding Amos Tversky and Daniel Kahneman's Prospect Theory and Cumulative Prospect Theory (CPT). Focussing entirely on evidence from within the original papers, we show that each and every person may be an exception to (Cumulative) Prospect Theory as advertised. Similar problems afflict many other behavioral research paradigms. We call on scholars to relinquish overly simplified characterizations of choice behavior. Telling practitioners and laypersons in stylized fashion how 'people' think promotes conjunction fallacies on a huge scale. Rather than conceptualize individual differences as a mere add-on to a schematic decision theory of central tendencies, decision scholars should recognize heterogeneity as a major theoretical primitive when proposing new theories.
We agree with Erev and Feigin (2022) that one should model heterogeneity at different levels. We do not promote either distribution-first or individual-first approaches over the others because population-level heterogeneity compounds sources of heterogeneity. We qualify Erev and Feigin's proposals in that neither approach is immune to scientific reasoning errors. We agree with Scheibehenne (2022) that aggregate statistics can appropriately summarize behavior, provided that the "effects" are robust across individuals. In contrast to this idealized scenario, decision researchers often deal with the joint occurrence of important qualitative differences on numerous attributes. Misconstruing individual differences as error variance carries a cost and violates the definition of overfitting. We agree with Kellen (2022) that the literature is often vague enough not to state verbatum that CPT MED is more descriptive of behavior than CPT with free parameters, but scientific conjunction errors are not so limited in scope. Regenwetter et al. (2022) intentionally glossed over potential limitations of studies, such as response errors, sample quality, reliability of measures, and diagnosticity of stimuli, to make a conceptual point. Speculating about the joint influence of these factors would render us agnostic about the lower and upper bounds on the number of people who satisfy a stylized theory. "Recipes" in study design are not exempt from conjunction errors: Fallacious reasoning is pernicious at any stage of scientific theorizing. We agree with Kellen that efforts to lead decision research beyond stylized theory deserve much further attention and future work.
Davis-Stober and Regenwetter (2019; Psychological Review) discussed the ‘paradox’ of converging evidence, whereby, with more and more positive Cohen's d values across multiple studies, support for a theory does not accumulate. Instead, more and more people may be exceptions to the theory. Using a psychometric framework, Heck (2021; Psychological Review) argued that Davis-Stober and Regenwetter's worst case scenarios are too pessimistic. His upwards-adjusted lower bounds on the number of people who satisfy multiple predictions of a theory jointly only occur when true score distributions are equally, and maximally, negative correlated across conditions. We show that Heck's conclusions hinge on untestable auxiliary assumptions. If one drops those assumptions, then the lower bounds of Davis-Stober and Regenwetter (2019) are attainable for any combination of effect sizes and number of predictions - and can still occur even when correlations across predictions are no longer all negative. Our arguments point to larger issues in quantitative psychology where seemingly innocuous modeling assumptions can unintentionally rule out empirically possible, as well as entirely plausible, outcomes. Even with generously large Cohen's d values and generously high correlations among individuals across effects, the proportion of the population that satisfies a theory quickly becomes smaller as of the number of predictions increases, under both frameworks. Said simply, the 'paradox' does not dissipate in either framework.
Scholars heavily rely on theoretical scope as a tool to challenge existing theory. We advocate that scientific discovery could be accelerated if far more effort were invested into also overtly specifying and painstakingly delineating the intended purview of any proposed new theory at the time of its inception. As a case study, we consider Tversky and Kahneman (1992). They motivated their Nobel-Prize-winning cumulative prospect theory with evidence that in each of two studies, roughly half of the participants violated independence, a property required by expected utility theory (EUT). Yet even at the time of inception, new theories may reveal signs of their own limited scope. For example, we show that Tversky and Kahneman’s findings in their own test of loss aversion provide evidence that at least half of their participants violated their theory, in turn, in that study. We highlight a combination of conflicting findings in the original article that make it ambiguous to evaluate both cumulative prospect theory’s scope and its parsimony on the authors’ own evidence. The Tversky and Kahneman article is illustrative of a social and behavioral research culture in which theoretical scope plays an extremely asymmetric role: to call existing theory into question and motivate surrogate proposals.
Testing rationality of decision-making and choice by evaluating the mathematical property of transitivity has a long tradition in biology, economics, psychology, and zoology. This paradigm is fraught with conceptual, mathematical, and statistical pitfalls. In this overview, we tackle five major obstacles. One challenge lies in spelling out what transitivity of latent preferences really says and what it actually implies about observable choice behavior. Most notably, this step is fraught with aggregation artifacts, in that aggregated behavior can be profoundly misleading about individual behavior. Another hurdle comes from hard mathematical problems associated with characterizing the properties of heterogeneous transitive populations. A third challenge is the prevalence of straw man hypotheses in this area of research. The fourth difficulty is associated with adopting appropriate statistical inference tools that correctly accommodate the idiosyncratic mathematical properties of order-constrained statistical hypotheses. The fifth hurdle arises with the role of scientific parsimony in rationality research. We walk readers through key concepts, mathematical models, and statistical techniques for testing rationality. Throughout, we provide examples using the methods and data of two prominent published papers on animal choice behavior as our case studies. We explain how these papers tackled the five hurdles to varying degrees of success.
As has been known for over a century, aggregated preferences of a group may bear little or no similarity to the preference of any single individual, regardless of the aggregation method. Yet, it remains routine to fit or test theories of individual decision making on pooled data, and it remains routine to cast theories of individual decision making at the aggregate level. This mindset may have disastrous policy and business implications. A population of individuals who all satisfy one theory may behave collectively as though they satisfied a competing theory. A collection of individuals satisfying a given theory may collectively satisfy a version of the same theory with qualitatively different scientific or decision analytic implications. Because the resulting artifacts apply at the population level, replications, large samples, and high-quality data can do nothing to detect or repair them.
Testing rationality of decision-making and choice by evaluating the mathematical property of transitivity has a long tradition in biology, economics, psychology, and zoology. This paradigm is fraught with conceptual, mathematical, and statistical pitfalls. In this overview, we tackle five major obstacles. One challenge lies in spelling out what transitivity of latent preferences really says and what it actually implies about observable choice behavior. Most notably, this step is fraught with aggregation artifacts, in that aggregated behavior can be profoundly misleading about individual behavior. Another hurdle comes from hard mathematical problems associated with characterizing the properties of heterogeneous transitive populations. A third challenge is the prevalence of straw man hypotheses in this area of research. The fourth difficulty is associated with adopting appropriate statistical inference tools that correctly accommodate the idiosyncratic mathematical properties of order-constrained statistical hypotheses. The fifth hurdle arises with the role of scientific parsimony in rationality research. We walk readers through key concepts, mathematical models, and statistical techniques for testing rationality. Throughout, we provide examples using the methods and data of two prominent published papers on animal choice behavior as our case studies. We explain how these papers tackled the five hurdles to varying degrees of success.
Psychological theory should guide the method. A method should not dictate theory. Extraneous assumptions entering psychological theories through the backdoor of a method may differentially affect the analysis of different data sets. This introduces noise and jeopardizes successful replication of valid theoretical claims. Auxiliary theoretical assumptions can also bias substantive conclusions (including across replications). It is therefore becoming ever more crucial that theoretical claims genuinely represent the given theory, no more, no less. Recent work has highlighted a disconnect between some theories and their ‘predictions,’ questioned the scope of theories in the presence of heterogeneity in hypothetical constructs, and developed methods to avoid extraneous assumptions. This tutorial merges these strands of research using a simple, illustrated case study on formulating and testing order-constrained theories. The tutorial applies to empirical paradigms in which scholars can state ordinal constraints on the outcome probabilities for several binary variables such as binary responses or the presence/absence of symptoms, and where the collection of binary variables is associated with a finite set of distinct conditions, such as group membership, treatment condition, or discrete levels of an independent variable. The goal is to let scholars spell out very precise hypotheses that (1) areunadulterated reflections of their theory, (2) provide exceptional theoretical nuance, (3) formally accommodate substantive heterogeneity and (4) offer rigorous and strong quantitative diagnosticity.
This stand-alone tutorial gives an introduction to the QTESIR 2.1 public domain software package for the specification and statistical analysis of certain order-constrained probabilistic choice models. Like its predecessors, QTEST 2.1 allows a user to specify a variety of probabilistic models of binary responses and to carry out state-of-the-art frequentist order-constrained hypothesis tests within a Graphical User Interface (GUI). QTEST 2.1 automatizes the mathematical characterization of so-called "random preference models", adds some parallel computing capabilities, and, most importantly, adds tools for Bayesian inference and model selection. In this tutorial, we provide an in-depth introduction to the Bayesian features: We review order-constrained Bayesian p-values, DIC and Bayes factors, building on the data, models, and prior QTEST based frequentist data analyses of an earlier (frequentist) tutorial by Regenwetter et al. (2014). (C) 2019 Elsevier Inc. All rights reserved.