Large language models (LLMs) are being increasingly incorporated into scientific workflows. However, we have yet to fully grasp the implications of this integration. How should the advent of large language models affect the practice of science? For this opinion piece, we have invited four diverse groups of scientists to reflect on this query, sharing their perspectives and engaging in debate. Schulz et al. make the argument that working with LLMs is not fundamentally different from working with human collaborators, while Bender et al. argue that LLMs are often misused and over-hyped, and that their limitations warrant a focus on more specialized, easily interpretable tools. Marelli et al. emphasize the importance of transparent attribution and responsible use of LLMs. Finally, Botvinick and Gershman advocate that humans should retain responsibility for determining the scientific roadmap. To facilitate the discussion, the four perspectives are complemented with a response from each group. By putting these different perspectives in conversation, we aim to bring attention to important considerations within the academic community regarding the adoption of LLMs and their impact on both current and future scientific practices.
Graphical perception is an important part of the scientific endeavour, and the interpretation of graphical information is increasingly important among educated consumers of popular media, who are often presented with graphs of data in support of different policy positions. However, graphs are multidimensional and data in graphs are comprised not only of overall global trends but also local perturbations. We presented a novel function estimation task in which scatterplots of noisy data that varied in the number of data points, the scale of the data, and the true generating function were shown to observers. 170 psychology undergraduates with mixed experience of mathematical functions were asked to draw the function that they believe generated the data. Our results indicated not only a general influence of various aspects of the presented graph (e.g., increasing the number of data points results in smoother generated functions) but also clear individual differences, with some observers tending to generate functions that track the local changes in the data and others following global trends in the data.
In the present research, we produce a coherent account of the storage and retrieval processes in short- and long-term event memory, and long-term knowledge, that produce response accuracy and response time in a wide variety of conditions in our studies of recognition memory. Two to nine pictures are studied sequentially followed by a target or foil test picture in four conditions used in Nosofsky et al. Journal of Experimental Psychology: Learning, Memory, and Cognition, 47, 316–342, (2021) and in our new paradigm: VM: target and foil responses to a given stimulus change from trial to trial; CM: the responses do not change from trial to trial; AN: every trial uses new stimuli; MIXED: combinations of VM, CN, and AN occur on each trial. In the new paradigm a given picture is equally often tested as old or new, but only in CM is the response key the same and learnable. Our model has components that have appeared in a variety of prior accounts, including learning and familiarity, but are given support by our demonstration that accuracy and response time data from a large variety of conditions can be predicted by these processes acting together, with parameter values that largely are unchanged. A longer version of this article, containing information not found here due to space, is available online https://doi.org/10.31234/osf.io/h8msp . The avalibility of the data (supplement materials), info and link is attached at the end section ( https://psyarxiv.com/h8msp .).
Vivid episodic memories in humans have been described as the replay of the flow of past events in sequential order. Recently, Panoz-Brown et al. Current Biology, 28, 1628-1634, (2018) developed an olfactory memory task in which rats were presented with a list of trial-unique odors in an encoding context; next, in a distinctive memory assessment context, the rats were rewarded for choosing the second to last item from the list while avoiding other items from the list. In a different memory assessment context, the fourth to last item was rewarded. According to the episodic memory replay hypothesis, the rat remembers the list items and searches these items to find the item at the targeted locations in the list. However, events presented sequentially differ in memory trace strength, allowing a rat to use the relative familiarity of the memory traces, instead of episodic memory replay, to solve the task. Here, we directly manipulated memory trace strength by manipulating the odor intensity of target odors in both the list presentation and memory assessment. The rats relied on episodic memory replay to solve the memory assessment in conditions in which reliance on memory trace strength is ruled out. We conclude that rats are able to replay episodic memories.
This commentary argues against the indictment of current experimental practices such as piecemeal testing, and the proposed integrated experiment design (IED) approach, which we see as yet another attempt at automating scientific thinking. We identify a number of undesirable features of IED that lead us to believe that its broad application will hinder scientific progress.
The binary distinction De Neys questions has been put forward many times since the beginnings of psychology, in slightly different forms and under different names. It has proved enormously useful and has received detailed empirical support and careful modeling. At heart the distinction is that between knowledge in long-term memory and control processes in short-term memory.
In 1998/1999, three participants trained for up to 74-h-long sessions to find a target present on half the trials in visual displays of 1, 2, or 4 initially novel objects. There were four targets and four foils that never changed. Displays occurred simultaneously, or the objects occurred successively, or the four features of each object occurred successively. When successive, the SOAs were short (17, 33, or 50 ms), so the displays appeared simultaneous, making it likely that the search strategy was the same in all conditions. A 2004 publication examined only the simultaneous condition and found evidence suggesting serial search as well as some small amount of automatic attention to targets and occasional early or late termination of search. A 2021 publication examined only the displays with single objects, obtaining evidence for dynamic perception of features. These studies drew conclusions from modeling subtle aspects of the response time distributions; extending such modeling to all conditions would have been complex making it difficult to understand the main processes at work. Here, we present a simple way to extend the 2021 model to the conditions with multiple item displays. It is a hybrid model with parallel automatic processing of features from all display items, processing that finishes during the first comparison, combined with serial comparisons that terminate when a target is found, or when none is found. When objects occur sequentially, there is a tendency to compare first the first object presented that probability rising with SOA. This model gives a good qualitative account of the accuracy and median response times from all the conditions. This success suggests that a more complex model incorporating the dynamic processes of the 2021 model would provide an excellent quantitative account for the accuracy and response time distributions for all the conditions of this visual search study.
Scientists studying decision-making often provide a set of choices, each specified with values or distributions of values, and probabilities or distributions of probabilities. For example, "Would you prefer $100 with probability 1.0 or $1 with probability .9 and $1,000 with probability 0.1?" Other decision research examines choices made in the absence of most quantitative information; for example, "Would you prefer a Ford now or a Porsche a year from now?," "Which food would you prefer," but models the findings with precise quantitative assumptions. Yet other research does neither; for example, modeling verbally stated choices with verbally stated heuristics. This article asks about the relevance of the first two research approaches for much of the decision-making made in life. The use of quantitative research and modeling is unsurprising, given that this approach underlies most of science. In life, values and probabilities are almost always partly or wholly vague and qualitative rather than quantitative. For example, when deciding which house to buy, there are relevant features such as size, color, neighborhood schools, construction materials, attractiveness, and many more, but the decision-maker finds it difficult and of little use to assign these precise values or weights. Nonetheless, humans have evolved to make decisions in such vaguely specified settings. I provide an example showing how a very high degree of uncertainty can defeat the application of quantitative decision-making, but such a demonstration is not critical if quantitative research and modeling produce a good understanding of and a good approximation to decision-making in the natural environment. This perspective addresses these issues.
Author(s): Harding, Samuel; Shiffrin, Richard | Abstract: We explore the dynamic coordination of perception, decision, and action underpinning perceptual choices by recording cursor movements during a binary response task. Stimuli were presented sequentially to control the time-course of perception, and we utilized a Hidden Markov Model (HMM) to relate measured movements of the mouse cursor to latent cognitive processes. Stimuli were simple perceptual objects comprised of two features, one of which was fully diagnostic of the correct response, while the other provided a probabilistic cue. The order of their arrival varied across trials, allowing us to manipulate the order of feature processing. The model builds upon response time methods and makes predictions about when individual features were perceived and the accumulation of evidence towards a response, every 10 milliseconds of each trial.
Participants gave recognition judgments for short lists of pictures of everyday objects. Pictures in a given list were an equal mixture of three types that varied according to the way they were used as targets and foils earlier in the same session. Under consistent-mapping (CM), targets and foils never switch roles; under varied-mapping (VM), targets and foils switch roles randomly across trials; whereas all-new (AN) items are novel on each trial of the experiment. Past research has shown that markedly enhanced performance occurs in CM conditions, leading to conclusions that item-response learning takes place in CM, perhaps automatically. However, almost all past research has compared CM, VM, and AN performance in between-blocks designs in which participants may adopt different cognitive strategies and criterion settings across the conditions. The present mixed-list design holds constant the strategy and criterion settings that are used for CM, VM, and AN items, and produced patterns of performance dramatically different than those observed in pure-list control conditions. We develop an extended version of an exemplar-based random-walk model of probe recognition to account for the major qualitative effects in the data. The data and the modeling provide evidence for strong item-response learning for CM foils but weak item-response learning for CM targets. We consider possible explanations for these effects in our General Discussion. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Eight initially novel objects with four features were learned by three participants over about 70 sessions in a variety of present-absent search tasks. This article analyzes and models trials with a single object presented for test. The features of the object were presented simultaneously, or successively at rates fast enough that the objects appeared to be simultaneous (inter-stimulus intervals were 16, 33, or 50 ms). Classification of a test object as target or foil required a conjunction of two features. When successively presented, features diagnostic for target presence could arrive first or last, and vice versa for features diagnostic for foil presence. Two results were particularly important: (1) the order in which target-diagnostic or foil-diagnostic features appeared produced large changes in accuracy and response times; (2) simultaneous feature presentation produced lower accuracy than sequential presentation with target-diagnostic features arriving first, despite the delay in such features arriving. The results required a dynamic model for perception and decision. The model has features perceived at independent times. It accumulates evidence at each moment based on the features perceived up to that time, and the diagnosticity of those features for classifying the test object as target or foil. The model also assumes that configurations of features provide evidence as processing continues: when all four features of an object are perceived the evidence points without error to the correct response. The results and modeling support the view that perceptual and decision processes operate concurrently and interactively during identification, recognition, and classification of well-learned objects, rather than in successive stages.
The connection of brain and mind has been a source of intense speculation at least since humanity became aware that the brain was the source of our behavior. Brain refers to the neurons, cells, and chemicals that govern activities of the organism. Mind is often considered consciously aware perceptions and thoughts. However, there is a gradient from unconscious to conscious, demonstrated by enormous amounts of research, such as the effects upon behavior of subliminal primes, so that mind is best considered to be the conscious and unconscious processes that act as an intermediate stage between the organism’s biology and its behavior, or a translation from one to the other. The National Academy of Sciences Colloquium “Brain Produces Mind by Modeling” was held May 1–3, 2019 at the Arnold and Mabel Beckman Center of the National Academy of Sciences in Irvine, CA. It was organized by Richard M. Shiffrin, Danielle S. Bassett, Nikolaus Kriegeskorte, and Joshua B. Tenenbaum. The theme of the colloquium and the foundation for the set of articles in this issue of PNAS is that the “mind” consists of a model formed by the “brain”: This would be a model of the entire environment, including the self, the body, the physical environment, other agents, and the social environment. Furthermore, the model would be a best guess about the most likely state of this environment. It uses this model to learn, decide, attend, remember, perceive, predict, and produce action. This model develops as the brain matures, rapidly during infancy and more slowly later. It has structural components that remain stable over long times. It has labile elements that change at multiple time scales, adapting to the current environment and goals. The mind’s formation through modeling of the world might be likened to the way scientists build models: through a … [↵][1]1To whom correspondence may be addressed. Email: shiffrin{at}indiana.edu. [1]: #xref-corresp-1-1
Over the past decade, a growing number of studies have investigated the relationship between the structure and function of human brain by predicting the resting-state functional connectivity (rsFC) from structural connectivity (SC). Yet how the whole-brain patterns of FC emerge from SC still remains incompletely understood. Unlike previous studies, here we propose an alternative approach for addressing this issue by predicting SC from rsFC. We first hypothesize that the functional couplings among brain areas at rest are shaped at least in three phases temporally: the initial direct interplay between brain areas, the communications within and between network modules, and followed by the indirect interactions ascribed to indirect structural pathways. We then introduce a network deconvolution (ND) algorithm inspired from the mechanism of cell differentiation, named CDA, to distinguish the direct dependencies from the functional network followed by a weight trimming algorithm based on Euclidean distance kernel function for shrinking the modular effects. Finally, we keep those region pairs with shorter shortest path length (SPL) together with shorter Euclidean distance as the structural connections. We apply the model and the algorithms to three intensively studied group averaged empirical connectome datasets with different parcellation resolutions and the results demonstrate that the predicted intrahemispheric structural connections and the weights distribution are highly consistent with the empirical SC derived from diffusion magnetic resonance imaging (dMRI) and probabilistic tractography, thus strongly supporting the model and algorithms proposed.
Proponents of preregistration argue that, among other benefits, it improves the diagnosticity of statistical tests [1]. In the strong version of this argument, preregistration does this by solving statistical problems, such as family-wise error rates. In the weak version, it nudges people to think more deeply about their theories, methods, and analyses. We argue against both: the diagnosticity of statistical tests depend entirely on how well statistical models map onto underlying theories, and so improving statistical techniques does little to improve theories when the mapping is weak. There is also little reason to expect that preregistration will spontaneously help researchers to develop better theories (and, hence, better methods and analyses).
In this article we review the framework proposed in 1968 by Atkinson and Shiffrin. We discuss the prior context that led to its production, including the advent of cognitive and mathematical modeling, its principal concepts, the subsequent refinements and elaborations that followed, and the way that the framework influenced other researchers to test the ideas and, in some cases, propose alternatives. The article illustrates the large amount of research and the large number of memory models that were directly influenced by this chapter over the past 50 years.
The three examples Gronau and Wagenmakers (Computational Brain and Behavior, 2018; hereafter denoted G&W) use to demonstrate the limitations of Bayesian forms of leave-one-out cross validation (let us term this LOOCV) for model selection have several important properties: The true model instance is among the model classes being compared; the smaller, simpler model is a point hypothesis that in fact generates the data; the larger class contains the smaller. As G&W admit, there is a good deal of prior history pointing to the limitations of cross validation and LOOCV when used in such situations (e.g., Bernardo and Smith 1994). We do not wish to rehash this literature trail, but rather give a conceptual overview of methodology that allows discussion of the ways that various methods of model selection align with scientific practice and scientific inference, and give our recommendation for the simplest approach that matches statistical inference to the needs of science. The methods include minimum description length (MDL) as reported by Grünwald (2007); Bayesian model selection (BMS) as reported by Kass and Raftery (Journal of the American Statistical Association, 90, 773–795, 1995); and LOOCV as reported by Browne (Journal of Mathematical Psychology, 44, 108–132, 2000) and Gelman et al. (Statistics and Computing, 24, 997–1016, 2014). In this commentary, we shall restrict the focus to forms of BMS and LOOCV. In addition, in these days of “Big Data,” one wants inference procedures that will give reasonable answers as the amount of data grows large, one focus of the article by G&W. We discuss how the various inference procedures fare when the data grow large.
The article “Robust Modeling in Cognitive Science” ( 2019 ) by Lee et al. makes several recommendations about best practices for cognitive science modelers. Many of these are reasonable and will not be discussed in this commentary. I believe several other critically important recommendations either put too much emphasis on less important components of good practice, or are somewhat misguided, and suggest that these are distorted in part because they are based on a misunderstanding of, or a failure to take into account, the goals of modeling. This commentary will highlight those areas where I believe the recommendations for good modeling practice deserve refinement and change.
Reproducibility has been one of the major tools science has used to help establish the validity and importance of scientific findings since Philosophical Transactions of the Royal Society was established in 1665 (1). Since that time the process of discovery has evolved to make use of new technologies and methods in a changing regulatory and social environment. The Sackler Colloquium “Reproducibility of Research: Issues and Proposed Remedies,” which took place on March 8–10, 2017, convened a wide range of research community stakeholders to address our understanding of transparency and reproducibility in the modern research context with two related questions: what does reproducibility mean in different research contexts, and what remedies increase reproducibility and transparency? We approach the topic of reproducibility with sensitivity to its complexity, spanning a wide range of issues from data collection and reporting to communication of scientific findings by scientists and nonscientists alike. The Colloquium was organized by David Allison, Richard Shiffrin, Victoria Stodden, and Stephen Fienberg. Before the Colloquium our esteemed and respected friend and colleague, Stephen Fienberg, unfortunately died and could not witness the outcome of his vision. A PNAS retrospective by Larry Wasserman describes Stephen’s extraordinary career (2). The 12 articles in this special issue fall into three categories that shaped the Colloquium. The first … [↵][1]1To whom correspondence should be addressed. Email: shiffrin{at}indiana.edu. [1]: #xref-corresp-1-1