de Varda et al (2025) reported that human response times and the number of tokens generated by a range of large reasoning models (LRMs) are correlated across 7 different tasks. The authors took the findings to reflect a “strong alignment” between reasoning effort in LRMs and humans. In response to a critique of Vankov et al. (2026), de Varda et al (2026a) carried out a new series of simulation studies taken to support their original claims and to clarify what they mean by strong alignment. Here we reinterpret their findings and conclude that their findings have no implications for human reasoning. We also take their clarification of strong alignment as a “motte and bailey” argument that is common in NeuroAI.
The presence of outliers in response times can affect statistical analyses and lead to incorrect interpretation of the outcome of a study. Therefore, it is a widely accepted practice to try to minimize the effect of outliers by preprocessing the raw data. There exist numerous methods for handling outliers and researchers are free to choose among them. In this article, we use computer simulations to show that serious problems arise from this flexibility. Choosing between alternative ways for handling outliers can result in the inflation of p-values and the distortion of confidence intervals and measures of effect size. Using Bayesian parameter estimation and probability distributions with heavier tails eliminates the need to deal with response times outliers, but at the expense of opening another source of flexibility.
Visual translation tolerance refers to our capacity to recognize objects over a wide range of different retinal locations. Although translation is perhaps the simplest spatial transform that the visual system needs to cope with, the extent to which the human visual system can identify objects at previously unseen locations is unclear, with some studies reporting near complete invariance over 10 degrees and other reporting zero invariance at 4 degrees of visual angle. Similarly, there is confusion regarding the extent of translation tolerance in computational models of vision, as well as the degree of match between human and model performance. Here, we report a series of eye-tracking studies (total N = 70) demonstrating that novel objects trained at one retinal location can be recognized at high accuracy rates following translations up to 18 degrees. We also show that standard deep convolutional neural networks (DCNNs) support our findings when pretrained to classify another set of stimuli across a range of locations, or when a global average pooling (GAP) layer is added to produce larger receptive fields. Our findings provide a strong constraint for theories of human vision and help explain inconsistent findings previously reported with convolutional neural networks (CNNs).
Visual translation tolerance refers to our capacity to recognize objects over a wide range of different retinal locations. Although translation is perhaps the simplest spatial transform that the visual system needs to cope with, the extent to which the human visual system can identify objects at previously unseen locations is unclear, with some studies reporting near complete invariance over 10° and other reporting zero invariance at 4° of visual angle. Similarly, there is confusion regarding the extent of translation tolerance in computational models of vision, as well as the degree of match between human and model performance. Here we report a series of eye-tracking studies (total N=70) demonstrating that novel objects trained at one retinal location can be recognized at high accuracy rates following translations up to 18°. We also show that standard deep convolutional networks (DCNNs) support our findings when pretrained to classify another set of stimuli across a range of locations, or when a Global Average Pooling (GAP) layer is added to produce larger receptive fields. Our findings provide a strong constraint for theories of human vision and help explain inconsistent findings previously reported with CNNs.
Combinatorial generalization—the ability to understand and produce novel combinations of already familiar elements—is considered to be a core capacity of the human mind and a major challenge to neural network models. A significant body of research suggests that conventional neural networks cannot solve this problem unless they are endowed with mechanisms specifically engineered for the purpose of representing symbols. In this paper, we introduce a novel way of representing symbolic structures in connectionist terms—the vectors approach to representing symbols (VARS), which allows training standard neural architectures to encode symbolic knowledge explicitly at their output layers. In two simulations, we show that neural networks not only can learn to produce VARS representations, but in doing so they achieve combinatorial generalization in their symbolic and non-symbolic output. This adds to other recent work that has shown improved combinatorial generalization under some training conditions, and raises the question of whether specific mechanisms or training routines are needed to support symbolic processing.This article is part of the theme issue ‘Towards mechanistic models of meaning composition’.
Combinatorial generalization - the ability to understand and produce novel combinations of already familiar elements - is considered to be a core capacity of the human mind and a major challenge to neural network models. A significant body of research suggests that conventional neural networks can't solve this problem unless they are endowed with mechanisms specifically engineered for the purpose of representing symbols. In this paper we introduce a novel way of representing symbolic structures in connectionist terms - the vectors approach to representing symbols (VARS), which allows training standard neural architectures to encode symbolic knowledge explicitly at their output layers. In two simulations , we show that out-of-the-box neural networks not only can learn to produce VARS representations, but in doing so they achieve combinatorial generalization. This adds to other recent work that has shown improved combinatorial generalization under specific training conditions, and raises the question of whether special mechanisms are indeed needed to support symbolic processing.
What mechanism supports our ability to recognize objects over a wide range of different retinal locations? Most research in psychology and neuroscience suggests that learning to identify a novel object at one retinal location only supports the ability to identify that object at nearby retinal locations, and to date, neural network models of object identification show a similar restriction in generalization. As a consequence, it is widely assumed that objects need to be learned at multiple locations. We challenge this view and show the capacity to generalize across retinal locations (what we call on-line translation tolerance) has been underestimated in humans and artificial neural networks. Two eye tracking studies demonstrate that novel objects can be recognized following translations of 9° and even 18°. Additionally, computational studies showed that convolutional neural networks can achieve similarly robust generalization when a mechanism (Global Average Pooling) was built in to generate larger receptive fields.
We argue that making accept/reject decisions on scientific hypotheses, including a recent call for changing the canonical alpha level from p = .05 to .005, is deleterious for the finding of new discoveries and the progress of science. Given that blanket and variable alpha levels both are problematic, it is sensible to dispense with significance testing altogether. There are alternatives that address study design and sample size much more directly than significance testing does; but none of the statistical tools should be taken as the new magic method giving clear-cut mechanical answers. Inference should not be based on single studies at all, but on cumulative evidence from multiple independent studies. When evaluating the strength of the evidence, we should consider, for example, auxiliary assumptions, the strength of the experimental design, and implications for applications. To boil all this down to a binary decision based on a p -value threshold of .05, .01, .005, or anything else, is not acceptable.
The Parallel Distributed Processing (PDP) approach to cognitive modelling assumes that knowledge is distributed across multiple processing units. This view is typically justified on the basis of the computational advantages and biological plausibility of distributed representations. However, both these assumptions have been challenged. First, there is growing evidence that some neurons respond to information in a highly selective manner. Second, it has been demonstrated that localist representations are better suited for certain computational tasks. In this paper, we continue this line of research by investigating whether localist representations are learned in tasks involving arbitrary input-output mappings. The results imply that the pressure to learn local codes in such tasks is weak, but still there are conditions under which feed-forward PDP networks learn localist representation. Our findings further challenge the assumption that PDP modelling always goes hand in hand with distributed representations and provide directions for future research.
Previous studies show that eye movement trajectory curves away from a remembered visual location if a saccade needs to be made in the same direction as the location. Data suggest that part of the process of maintaining the location in working memory is the mental simulation of that location, so that the oculomotor system treats the remembered location as a real one. Other research suggests that word meaning may also behave like a ‘real object’ in space. The current study aimed to combine the two streams of research examining the effect of word meaning on the memory of a dot location. The results of two experiments showed that word meaning for ‘up’ (but not ‘down’) modulated both eye movement trajectory and location recognition time. Thus, mental simulation of task-irrelevant space-related word meaning affected both earlier stages of memory processes (maintenance of the location in the working memory) and later ones (location recognition).
The ability to recognize the same image projected to different retinal locations is critical for visual object recognition in natural contexts. According to many theories, the translation invariance for objects extends only to trained retinal locations, so that a familiar object projected to a nontrained location should not be identified. In another approach, invariance is achieved “online,” such that learning to identify an object in one location immediately affords generalization to other locations. We trained participants to name novel objects at one retinal location using eyetracking technology and then tested their ability to name the same images presented at novel retinal locations. Across three experiments, we found robust generalization. These findings provide a strong constraint for theories of vision.
Why do some neurons in hippocampus and cortex respond to information in a highly selective manner? It has been hypothesized that neurons in hippocampus encode information in a highly selective manner in order to support fast learning without catastrophic interference, and that neurons in cortex encode information in a highly selective manner in order to co-activate multiple items in short-term memory (STM) without suffering a superposition catastrophe. However, the latter hypothesis is at odds with the widespread view that neural coding in the cortex is highly distributed in order to support generalization. We report a series of simulations that characterize the conditions in which recurrent Parallel Distributed Processing (PDP) models of immediate serial can recall novel words. We found that these models learned localist codes when they succeeded in generalizing to novel words. That is, just as fast learning may explain selective coding in hippocampus, STM and generalization may help explain the existence of selective codes in cortex.
A key insight from 50 years of neurophysiology is that some neurons in cortex respond to information in a highly selective manner. Why is this? We argue that selective representations support the coactivation of multiple "things" (e.g., words, objects, faces) in short-term memory, whereas nonselective codes are often unsuitable for this purpose. That is, the coactivation of nonselective codes often results in a blend pattern that is ambiguous; the so-called superposition catastrophe. We show that a recurrent parallel distributed processing network trained to code for multiple words at the same time over the same set of units learns localist letter and word codes, and the number of localist codes scales with the level of the superposition. Given that many cortical systems are required to coactivate multiple things in short-term memory, we suggest that the superposition constraint plays a role in explaining the existence of selective codes in cortex.
Cohen’s classic study on statistical power (Cohen, 1962) showed that studies in the 1960 volume of the Journal of Abnormal and Social Psychology lacked sufficient power to detect anything other than large effects (r ~ 0.60). Sedlmeier and Gigerenzer (Sedlmeier & Gigerenzer, 1989) conducted a similar analysis on studies in the 1984 volume and found that, if anything, the situation had worsened. Recently, Button and colleagues showed that the average power of neuroscience studies is probably around 20% (Button et al., 2013b). Clearly repeated exhortations that researchers should “pay attention to the power of their tests rather than … focus exclusively on the level of significance” (Sedlmeier & Gigerenzer, 1989) have failed. Here we consider why this might be so. One reason might be a lack of appreciation of the importance of statistical power within a null hypothesis significance testing (NHST) framework. NHST grew out of the distinct statistical theories of Fisher (Fisher, 1955), and Neyman and Pearson (Rucci & Tweney, 1980). From Fisher we take the concept of null hypothesis testing, and from Neyman-Pearson the concepts of Type I (α) and Type II error (β). Power is a concept arising from Neyman-Pearson theory, and reflects the likelihood of correctly rejecting the null hypothesis (i.e., 1- β). However, the hybrid statistical theory typically used leans most heavily on Fisher’s concept of null hypothesis testing. Sedlmeier and Gigerenzer (1989) argued that a lack of understanding of these distinctions partly explained the lack of consideration of statistical power; while we (nominally, at least) adhere to a 5% Type I error rate, we pay little attention to the Type II error rate, despite the need to consider both when evaluating whether a research finding is likely to be true (Button et al., 2013b). Another reason might be the incentive structures within which scientists operate. Scientists are human and therefore will respond (consciously or unconsciously) to incentives; when personal success (e.g., promotion) is associated with the quality and (critically) the quantity of publications produced, it makes more sense to use finite resources to generate as many publications as possible. A single transformative study in a highly-regarded journal might confer the most prestige, but this is a high-risk strategy – the experiment may not produce the desired (i.e., publishable) results, or the journal may not accept it for publication (Sekercioglu, 2013). A safer strategy might be to “salami-slice” one’s resources to generate more studies which, with sufficient analytical flexibility (Simmons, Nelson, & Simonsohn, 2011), will almost certainly produce a number of publishable studies (Sullivan, 2007). There is some support for the second reason. Studies published in some countries may over-estimate true effects more than those published in other countries (Fanelli & Ioannidis, 2013; Munafo, Attwood, & Flint, 2008). This may be because, in certain countries, publication in even medium-rank journals confers substantial direct financial rewards on the authors (Shao & Shen, 2011), which may in turn be related to over-estimates of true effects (Pan, Trikalinos, Kavvoura, Lau, & Ioannidis, 2005). Authors may therefore (consciously or unconsciously) conduct a larger number of smaller studies, which are still likely to generate publishable findings, rather than risk investing their limited resources in a smaller number of larger studies. However, to the best of our knowledge, the first possible reason has not been systematically explored. We therefore surveyed studies published recently in a high-ranking psychology journal, and contacted authors to establish the rationale used for deciding sample size (see Supplementary Material). This indicated that approximately one third held beliefs that would serve, on average, to reduce statistical power (see Table 1). In particular, they used accepted norms within their area of research to decide on sample size, in the belief that this would be sufficient to replicate previous results (and therefore, presumably, to identify new findings). Given empirical evidence for a high prevalence of findings close to the p = 0.05 threshold (Masicampo & Lalande, 2012), this belief is likely to be unwarranted. If an experiment finds an effect with p ~ 0.05, and we assume the effect size observed is accurate, then if we repeat the experiment with the same sample size we will on average replicate that finding only 50% of the time (see Supplementary Material). In reality, power will be much lower than 50% because the effect size estimate observed in the original estimate is probably an over-estimate (Simonsohn, 2013). However, in our survey, over one third of respondents inaccurately believed that in this scenario the finding would replicate over 80% of the time (see Supplementary Material). Table 1 Beliefs about sample size and statistical power. There are unlikely to be simple solutions to the continued lack of appreciation of statistical power. One reason for pessimism, as we have shown, is that these concerns are not new; occasional discussion of these issues has not led to any lasting change. Structural change may be required, including more rigorous enforcement by journals and editors of guidelines which are often found in instructions for authors, but not always followed. Recently, Nature introduced a submission checklist for life sciences articles, which includes a requirement that sample size be justified (http://www.nature.com/authors/policies/checklist.pdf). Other journals are introducing novel submission formats which place greater emphasis on study design (including statistical power), rather than results, including Registered Reports at Cortex, and Registered Replication Reports at Perspectives on Psychological Science. The poor reproducibility of scientific findings continues to be a cause of major concern. Small studies with low statistical power contribute to this problem (Bertamini & Munafo, 2012; Button et al., 2013b), and arguments in defence of “small-scale science” (Quinlan, 2013) overlook the fact that larger studies protect against inferences from trivial effect sizes by allowing a better estimation of the magnitude of true effects (Button et al., 2013a). Reasons to resist NHST, and in particular the dichotomous interpretation of p-values, have been well-rehearsed (Sterne & Davey Smith, 2001), and alternative approaches, such as focusing on effect size estimation or implementing Bayesian approaches do exist. However, while NHST remains the dominant model for statistical inference, we should ensure that it is appropriately used.
Cohen’s classic study on statistical power (Cohen, 1962) showed that studies in the 1960 volume of the Journal of Abnormal and Social Psychology lacked sufficient power to detect anything other than large effects (r ~ 0.60). Sedlmeier and Gigerenzer (Sedlmeier & Gigerenzer, 1989) conducted a similar analysis on studies in the 1984 volume and found that, if anything, the situation had worsened. Recently, Button and colleagues showed that the average power of neuroscience studies is probably around 20% (Button et al., 2013b). Clearly repeated exhortations that researchers should “pay attention to the power of their tests rather than … focus exclusively on the level of significance” (Sedlmeier & Gigerenzer, 1989) have failed. Here we consider why this might be so. One reason might be a lack of appreciation of the importance of statistical power within a null hypothesis significance testing (NHST) framework. NHST grew out of the distinct statistical theories of Fisher (Fisher, 1955), and Neyman and Pearson (Rucci & Tweney, 1980). From Fisher we take the concept of null hypothesis testing, and from Neyman-Pearson the concepts of Type I (α) and Type II error (β). Power is a concept arising from Neyman-Pearson theory, and reflects the likelihood of correctly rejecting the null hypothesis (i.e., 1- β). However, the hybrid statistical theory typically used leans most heavily on Fisher’s concept of null hypothesis testing. Sedlmeier and Gigerenzer (1989) argued that a lack of understanding of these distinctions partly explained the lack of consideration of statistical power; while we (nominally, at least) adhere to a 5% Type I error rate, we pay little attention to the Type II error rate, despite the need to consider both when evaluating whether a research finding is likely to be true (Button et al., 2013b). Another reason might be the incentive structures within which scientists operate. Scientists are human and therefore will respond (consciously or unconsciously) to incentives; when personal success (e.g., promotion) is associated with the quality and (critically) the quantity of publications produced, it makes more sense to use finite resources to generate as many publications as possible. A single transformative study in a highly-regarded journal might confer the most prestige, but this is a high-risk strategy – the experiment may not produce the desired (i.e., publishable) results, or the journal may not accept it for publication (Sekercioglu, 2013). A safer strategy might be to “salami-slice” one’s resources to generate more studies which, with sufficient analytical flexibility (Simmons, Nelson, & Simonsohn, 2011), will almost certainly produce a number of publishable studies (Sullivan, 2007). There is some support for the second reason. Studies published in some countries may over-estimate true effects more than those published in other countries (Fanelli & Ioannidis, 2013; Munafo, Attwood, & Flint, 2008). This may be because, in certain countries, publication in even medium-rank journals confers substantial direct financial rewards on the authors (Shao & Shen, 2011), which may in turn be related to over-estimates of true effects (Pan, Trikalinos, Kavvoura, Lau, & Ioannidis, 2005). Authors may therefore (consciously or unconsciously) conduct a larger number of smaller studies, which are still likely to generate publishable findings, rather than risk investing their limited resources in a smaller number of larger studies. However, to the best of our knowledge, the first possible reason has not been systematically explored. We therefore surveyed studies published recently in a high-ranking psychology journal, and contacted authors to establish the rationale used for deciding sample size (see Supplementary Material). This indicated that approximately one third held beliefs that would serve, on average, to reduce statistical power (see Table 1). In particular, they used accepted norms within their area of research to decide on sample size, in the belief that this would be sufficient to replicate previous results (and therefore, presumably, to identify new findings). Given empirical evidence for a high prevalence of findings close to the p = 0.05 threshold (Masicampo & Lalande, 2012), this belief is likely to be unwarranted. If an experiment finds an effect with p ~ 0.05, and we assume the effect size observed is accurate, then if we repeat the experiment with the same sample size we will on average replicate that finding only 50% of the time (see Supplementary Material). In reality, power will be much lower than 50% because the effect size estimate observed in the original estimate is probably an over-estimate (Simonsohn, 2013). However, in our survey, over one third of respondents inaccurately believed that in this scenario the finding would replicate over 80% of the time (see Supplementary Material). Table 1 Beliefs about sample size and statistical power. There are unlikely to be simple solutions to the continued lack of appreciation of statistical power. One reason for pessimism, as we have shown, is that these concerns are not new; occasional discussion of these issues has not led to any lasting change. Structural change may be required, including more rigorous enforcement by journals and editors of guidelines which are often found in instructions for authors, but not always followed. Recently, Nature introduced a submission checklist for life sciences articles, which includes a requirement that sample size be justified (http://www.nature.com/authors/policies/checklist.pdf). Other journals are introducing novel submission formats which place greater emphasis on study design (including statistical power), rather than results, including Registered Reports at Cortex, and Registered Replication Reports at Perspectives on Psychological Science. The poor reproducibility of scientific findings continues to be a cause of major concern. Small studies with low statistical power contribute to this problem (Bertamini & Munafo, 2012; Button et al., 2013b), and arguments in defence of “small-scale science” (Quinlan, 2013) overlook the fact that larger studies protect against inferences from trivial effect sizes by allowing a better estimation of the magnitude of true effects (Button et al., 2013a). Reasons to resist NHST, and in particular the dichotomous interpretation of p-values, have been well-rehearsed (Sterne & Davey Smith, 2001), and alternative approaches, such as focusing on effect size estimation or implementing Bayesian approaches do exist. However, while NHST remains the dominant model for statistical inference, we should ensure that it is appropriately used.