Large language models (LLMs) have been shown to acquire sequence-level planning abilities during training, yet their planning behavior exhibited at inference time often appears short-sighted and inconsistent with these capabilities. We propose a Bayesian account for this gap by grounding planning behavior in the evolving generative context: given the subtle differences between natural language and the language internalized by LLMs, accumulated self-generated context drives a planning-shift during inference and thereby creates the appearance of compromised planning behavior. We further validate the proposed model through two controlled experiments: a random-generation task demonstrating constrained planning under human prompts and increasing planning strength as self-generated context accumulates, and a Gaussian-sampling task showing reduced initial bias when conditioning on self-generated sequences. These findings provide a theoretical explanation along with empirical evidence for characterizing how LLMs plan ahead during inference.
Incorporating individual-level cognitive priors offers an important route to personalizing neural networks, yet accurately eliciting such priors remains challenging: existing methods either fail to uniquely identify them or introduce systematic biases. Here, we introduce PriorProbe, a novel elicitation approach grounded in Markov Chain Monte Carlo with People that recovers fine-grained, individual-specific priors. Focusing on a facial expression recognition task, we apply PriorProbe to individual participants and test whether integrating the recovered priors with a state-of-the-art neural network improves its ability to predict an individual's classification on ambiguous stimuli. The PriorProbe-derived priors yield substantial performance gains, outperforming both the neural network alone and alternative sources of priors, while preserving the network's inference on ground-truth labels. Together, these results demonstrate that PriorProbe provides a general and interpretable framework for personalizing deep neural networks.
Recent decision-making models have explained behaviour using mental sampling mechanisms, but there is still little agreement on the specific sampling process, such as whether sampling rates match true probabilities. Here, we seek to trace the sampling process using generation tasks: in two experiments using general online samples (Ns = 52, 51), participants repeatedly produced potential outcomes from pairs of monetary gambles before choosing between them. Results found over-generation of rarer outcomes and under-generation of common outcomes overall, but not in initial responses, as well as avoidance of direct repetitions. Participants also tended to select options with higher average utility across their responses, implying generations guided choice. These findings suggest systematic biases in the information people may consider before a choice, and the influence that this can have on subsequent decisions, carrying implications for mental sampling models of this behaviour. We thus suggest explicit generation is a valuable method to access underlying choice processes, offering new assessments of existing theories of decision making.
Being unpredictable is useful for creativity, exploration, and good decision-making. Decades of research asking people to generate random sequences of numbers have concluded that people systematically deviate from randomness, but individuals do so in very idiosyncratic ways. However, there is no consensus on the cause of these rich individual differences: most theories postulate that people achieve this by spending cognitive resources monitoring their own output and changing the way they say items accordingly, whereas a Local Sampling account postulates that people draw into a general-purpose ability to produce samples, which they would use to make judgments and choices, and which are inherently somewhat unpredictable. Here we distinguish between these possibilities by asking people to generate sequences both at random and as they come to mind, across two experiments at different production speeds, using a non-uniform distribution. We employ several measures of how sequences deviate from randomness. We find that, consistent with the Local Sampling explanation, people deviate from randomness in virtually identical ways in both sequences, with high individual differences pointing to a common cognitive process. We follow up these findings by computationally modelling human performance using several Local Sampling models. We find many participants had the same model as best-fitting in both sequences; with estimated parameters correlating strongly across tasks showing individual differences in how dependent on previous items sampling is. Overall, we conclude that random generation is better understood as employing a general-purpose faculty, Local Sampling, which is stable across time and tasks, with differences in the sampler’s features resulting in differences in random generation performance.
Societal expectations have been found to determine which social roles people should occupy. However, so far, these beliefs have been mainly explored with self-report and response conflict measures where expectation-confirming (vs. violating) judgments elicit faster responding. The present lab study (N = 57) applied a novel approach – the random generation paradigm – to understand how pre-existing social assumptions determine which information is retrieved from memory when prompted by different social categories. Specifically, we asked participants to imagine (hypothetical) people working in certain professions and to say their names out loud. We found that the statistics of the uttered names reflected societal gender stereotypes and environmental statistics of actual people working in these occupations. Importantly, the proportion of female and male names generated for each profession by each participant predicted their performance in a sequential priming task (prime = stereotyped professions, target = female and male faces) better than the environmental statistics or participants’ estimates of gender proportions. Together, these findings offer a new, and widely applicable, method for exploring cultural beliefs and help clarify how social information is sampled from memory when making social judgments.
Noise in behavior is often considered a nuisance: Although the mind aims for the best possible action, it is let down by unreliability in the sensory and response systems. Researchers often represent noise as additive, Gaussian, and independent. Yet a careful look at behavioral noise reveals a rich structure that defies easy explanation. First, in both perceptual and preferential judgments sensory and response noise may potentially play only minor roles, with most noise arising in the cognitive computations. Second, the functional form of the noise is both non-Gaussian and nonindependent, with the distribution of noise being better characterized as heavy-tailed and as having substantial long-range autocorrelations. It is possible that this structure results from brains that are, for some reason, bedeviled by a fundamental design flaw, albeit one with intriguingly distinctive characteristics. Alternatively, noise might not be a bug but a feature. Specifically, we propose that the brain approximates probabilistic inference with a local sampling algorithm, one using randomness to drive its exploration of alternative hypotheses. Reframing cognition in this way explains the rich structure of noise and leads to the surprising conclusion that noise is not a symptom of cognitive malfunction but plays a central role in underpinning human intelligence.
Behaving randomly can be advantageous: it prevents others from capitalizing on patterns in our behavior. Unfortunately, the consensus from sixty years of psychological research is that people cannot do so: when attempting to be random, people’s responses exhibit systematic patterns. Random phenomena, however, are not instantaneously random. They require sufficient time between observations (e.g. the weather) or for enough iterations of a randomizing process to have occurred (e.g. card shuffling). It is unknown whether human sequences can be random if afforded such delays. Here we show that a modest temporal separation between items can make human sequences indistinguishable from random ones. We carried out our own experiment (N = 54) and analyzed ten existing datasets with different production rates, response sets, and response modalities. We found that when the delay between items was between two and four seconds, differences between human and random sequences disappeared. Furthermore, by comparing sequences produced by the same participants at different speeds, we confirmed that when participants make an effort, the needed delay is independent of production rate, akin to the weather. Our results show that people are able to generate randomness, and within a few seconds, giving us an accessible protection against potential exploits.
Does the utility of an outcome influence people’s assessment of risk and uncertainty? Growing evidence suggests that people often rely on mental simulations to evaluate probability and risky events. However, prior experimental findings offer conflicting predictions about how utility biases this mental sampling process. Across four experiments (total N=206, with Experiment 4 pre-registered), we investigated the influence of utility using a random generation paradigm. These responses were then compared to probability judgments and predictions. While we identified individual differences, the majority of participants exhibited neutrality, with no systematic impact of utility on their sampling distributions. Nevertheless, biases emerged under specific conditions, including a preference for smaller or more probable outcomes as the starting point of simulations and optimism in single-response predictions. Additionally, we found evidence suggesting that probability judgments, predictions, and random generation tasks may rely on a shared underlying mental process. Our findings suggest that models of judgment and decision-making should account for individual differences in utility influences, particularly distinguishing between unbiased sampling and optimistic sampling—the selective over-representation of high-utility outcomes.
Individuals make biased and variable probability judgements. Recent models such as the Bayesian Sampler and Probability Theory Plus Noise capture these effects by assuming people randomly sample events but are biased towards indifference (i.e., 0.5). However there is a bias they do not capture: systematic violations of binary complementarity, i.e., violations of the simple constraint that judgments of P(A) and P(not A) should sum to 1. Until now, this bias was only captured by the sampling process of the Quantum Sequential Sampler. Here we develop straightforward generalisations of the Bayesian Sampler, by introducing an asymmetric prior, and Probability Theory Plus Noise, by introducing asymmetric noise, that can generate violations of binary complementarity. We next show that these three models make distinct predictions for the mean-variance relationship in repeated judgments. Finally, we investigate violations of binary complementarity in five experiments, where participants judged the probabilities of dice rolls. Participants consistently violated binary complementarity, independent of whether they were in a high or low probability environment or how the alternative options are partitioned. Crucially, participants showed the highest variability for probability judgements below 0.5, an effect captured by an asymmetric prior in the generalised Bayesian Sampler, but not by the biasing mechanisms in the other models.
In many real-life settings, feedback is only available for cases that decision makers accept and so may be biased toward positive events. How do people learn to distinguish good from bad alternatives from such selective feedback, and can they correct for this bias? We describe the computational problems of classification learning from biased samples and examine how exemplar and model-based methods can deal with this challenge: Model-based methods can adjust their representation of the task based on what information is available while exemplar models can impute fictive negative outcomes in missing cases to avoid positivistic biases. Importantly, these methods imply distinct assumptions about the task and reactions to missing feedback, which can be assessed empirically. In three experiments, we test whether participants rely on imputation or use a Bayesian model of the task to correct for selection bias. We find that many participants were best described by an exemplar model, most with imputation, but an almost equal proportion was best described by a Bayesian model. People best described by different models reacted somewhat differently to missing feedback. We also observe substantial stability in whether individuals were best described by model-based or exemplar models across tasks, though participants were more likely to use exemplar models when there was greater uncertainty about the task structure. Overall, our findings show that people deal with missing feedback in an adaptive manner by adopting diverse approaches that are partially stable and partially reflect assumptions made about the experimental context. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Choices made in risky scenarios are considered fundamentally noisy because decisions have often been found to be inconsistent when repeated. Past measures of noise may, however, be confounded by the use of randomized contextual factors that are known to influence choice, in particular, the order of trials. In two experiments, we control trial order to test the extent to which inconsistent choice is attributable to changes in experimental context. Both tasks find strong evidence that trial order has no effect on choice consistency, indicating such experimental factors have little influence on behavior compared with internal noise. Choices also showed an increase in consistency across multiple repetitions, suggesting a fall in noise with experience, but this increase was not associated with any improvement in performance, with choices showing no greater adherence to either expected value or expected utility across repetitions. Instead, choices increasingly adhered to simplistic heuristic decision rules, possibly indicating greater reliance on such strategies as the tasks progressed. These results carry implications for a number of decision-making theories, including true-and-error models, rank-based methods, and strategy shift approaches.
Repeated forecasts of changing values are a key aspect of many everyday tasks, from predicting the weather to financial markets. A particularly simple and informative instance of such fluctuating values are random walks: sequences in which each point is a random movement from only its preceding value, unaffected by any previous points. Moreover, random walks often yield basic rational forecasting solutions in which predictions of new values should repeat the most recent value, and hence replicate the properties of the original series. In previous experiments, however, we have found that human forecasters do not adhere to this standard, showing systematic deviations from the properties of a random walk such as excessive volatility and extreme movements between subsequent predictions. We suggest that such deviations reflect general statistical signatures of human cognition displayed across multiple tasks, offering a window into underlying cognitive mechanisms. Using these deviations as new criteria, we here explore several cognitive models of forecasting drawn from various approaches developed in the existing literature, including Bayesian, error-based learning, autoregressive and sampling mechanisms. These models are contrasted with human data from two experiments to determine which best accounts for the particular statistical features displayed by participants. We find support for sampling models in both aggregate and individual fits, suggesting that these variations are attributable to the use of inherently stochastic prediction systems. We thus argue that variability in predictions is primarily driven by computational noise within the decision making process, rather than "late" noise at the output stage.
Elicitation methods, such as asking people to produce the deciles of a distribution, are standard practices in policy or applied statistics. Similarly, much of cognitive science and psychology focuses on determining people's people's beliefs or latent traits through questionnaires or judgment tasks. However, these approaches often only capture a rough outline of what people know and are usually limited to point estimates of people's beliefs. Here, we present a novel experimental paradigm that allows us to access people's beliefs and how variable these beliefs are. Our task is based on an established random generation paradigm in which participants produce quantities from a particular domain as randomly as possible. We hypothesize that due to the minds' general-purpose mechanisms for probabilistic inferences, these random sequences represent the participants' underlying prior beliefs. We show that our method can infer participants' beliefs for a wide range of numeric quantities at comparable accuracy as an established elicitation method.Moreover, these inferred beliefs are consistent with individual participants' generalization and inference patterns in a subsequent conditional prediction task. We then extend our approach to non-numeric belief elicitation, highlighting how our method can go beyond numeric elicitation and provide insight into complex beliefs that are challenging to assess experimentally. Empirically, our results highlight that people know the rough shapes of environmental distributions, and these beliefs guide inference and generalization. Moreover, using our novel approach, we also show that people know the fine details of environmental distributions. Finally, our experimental results show that random generation paradigms can be a useful tool for cognitive scientists, psychologists, and applied statisticians.
The categorization of complex real-world stimuli, such as facial expressions, appears to vary greatly between people. This raises a crucial methodological challenge: how is it possible to elicit the mental representation of a complex category for a specific individual? Comprehensive category-elicitation methods such as Markov Chain Monte Carlo with People (MCMCP) work well across populations, but converge too slowly to be usable with individual participants. Here, we address the problem of slow convergence with a new method: combining MCMCP with an adapted Variational Auto-Encoder (VAE) acting as a “gatekeeper”. We tested this approach in a new experiment (N=90) on facial affect comparing MCMCP with the “gatekeeper” (MCMCPG) against baseline MCMCP and other variants. MCMCPG converged substantially faster than the other methods, in about 10 minutes for our task, with showing more representative recovered faces than its competitors. Further analyses captured participants’ substantial individual differences in a categorization task at an individual level. And, uniquely, the resulting model generalized these individual differences to real-world faces outside of our training set. Our study demonstrates the potential of MCMCPG for investigating generalizable human representations of complex stimuli at the individual level and illustrates the power of integrating Artificial Intelligence into psychological experiments.
Confirmation bias is defined as searching for and assimilating information in a way that favours existing beliefs. We show that confirmation bias emerges as a natural consequence of boundedly rational belief updating by presenting the BIASR model (Bayesian updating with an Independence Approximation and Source Reliability). In this model, an individual's beliefs about a hypothesis and the source reliability form a Bayesian network. Upon receiving information, an individual simultaneously updates beliefs about the hypothesis in question and the reliability of the information source. If the individual updates rationally then this introduces numerous dependencies between beliefs, the tracking of which represents an unrealistic demand on memory. We propose that human cognition overcomes this memory limitation by assuming independence between beliefs, evidence for which is provided in prior research. We show how a Bayesian belief updating model incorporating this independence approximation generates many types of confirmation bias, including biased evaluation, biased assimilation, attitude polarisation, belief perseverance and confirmation bias in the selection of sources.
Societal expectations have been found to determine which social roles people should occupy. However, so far, these beliefs have been mainly explored using implicit measures where expectation-confirming (vs. violating) judgments tend to be more efficient. The present study (N = 44) applied a novel approach – the random generation paradigm – to explore how pre-existing social assumptions determine which information is retrieved from memory when prompted by different social categories. Specifically, we asked participants to imagine people working in certain professions and to say their names out loud. We found that the statistics of the uttered names reflected societal gender stereotypes and environmental statistics of actual people working in these occupations. Importantly, the proportion of female and male names generated for each profession by each participant predicted their performance in a sequential priming task (prime = stereotyped professions, target = female and male faces) better than the environmental statistics or participants’ explicit estimates of gender proportions. Together, these findings offer a new, and widely applicable, method for exploring cultural beliefs and help clarify how social information is sampled from memory when making social judgments.