Lexique 4, an updated French lexical database, expands upon its predecessor, Lexique 3, by incorporating several significant improvements to enhance its utility in psycholinguistics, computational linguistics, and education. The new version is based on a larger corpus of 316 million words derived from 65,317 documents, including movie, TV show, and documentary subtitles, which offers more accurate frequency estimates and includes contemporary neologisms. Lexique 4 introduces new variables, such as orthographic surface frequency, contextual diversity (CD), and detailed morphological structure, which provide a more comprehensive view of lexical properties. We find that contextual diversity is a slightly better predictor than word frequency, in line with previous work. Moreover, the integration of lexical decision times from the French Lexicon Project into Lexique 4 facilitates more in-depth linguistic research. Enhancements to the user interface, including a redesigned web platform, enable dynamic searches and sorting capabilities, increasing accessibility and usability for researchers. Statistical analyses indicate that the updated frequency measures in Lexique 4 are better predictors of lexical decision times compared to Lexique 3, supporting the value of these enhancements. Overall, Lexique 4 represents a comprehensive and flexible tool for analyzing French lexical properties, making it an essential asset for a broad range of users.
Experimental psychologists routinely need to randomise the order of trials across participants while avoiding undesirable sequential patterns, for example, long runs of trials belonging to the same stimulus category, or target items appearing too close together. We present shuffle, a freely available, cross-platform tool for generating quasi-randomised sequences that satisfy user-defined sequential constraints. Input is a plain-text or CSV table whose rows represent trials and whose columns represent categorical variables (\textit{e.g.}, stimulus type, response mapping, condition). Two types of constraint can be imposed on any column: a *maximum-repetition* constraint limits the number of consecutive rows that may share the same label, and a *minimum-gap* constraint requires a minimum number of intervening rows between any two rows sharing the same label. Two algorithms are provided. The *constructive* algorithm builds a valid permutation row by row and is very fast. The *equiprobable* algorithm draws random permutations repeatedly until one satisfies the constraints, guaranteeing an unbiased sample from the set of valid permutations; it is slower but preferable when distributional properties of the randomisation matter. Shuffle-go is implemented in Go and distributed as self-contained binaries for Windows, macOS, and Linux (both x86-64 and ARM64), requiring no runtime installation. It offers a command-line interface, two graphical desktop interfaces, an embeddable Go library, and a Python library. Source code is available at https://github.com/chrplr/shuffle-go under the GPLv3 licence.
Abstract Although the brain areas for language processing are well delimited, whether lexical-semantic and syntactic processes are spatially segregated remains debated. To clarify this issue, we conducted two experiments using 7-Tesla functional MRI in 20 participants performing: a functional localizer involving reading sequences of words of increasing linguistic complexity; and a presentation of short, semantically impoverished three-word mini-sentences, flashed in a single glance (e.g., “he does it”), whose grammaticality and syntactic complexity was manipulated through syntactic movement. Our results reveal two functionally dissociable sets of cortical patches within the language system: one sensitive to syntactic structure even in the absence of meaning, and the other involved in semantic composition. This dual-network architecture was consistently observed in the majority of participants, although its precise anatomical localization varied. The two types of voxels coexisted even within a given brain region of the Glasser atlas. Results were confirmed using subject-specific analyses and region-by-condition interactions, as voxels in those two systems displayed markedly different responses to mini-sentences. Thus, high-resolution functional imaging reveals a division of labor between syntactic and semantic composition within the classical language network.
We introduce `Goxpyriment', a new open-source software framework for programming behavioral and cognitive experiments using the Go programming language. The library is designed to address some limitations of existing Python-based experiment tools, particularly the runtime environment complexity that frequently complicates deployment across laboratories. Because Go is a compiled language that can natively embed assets (e.g., graphics, audio files, and stimulus lists), Goxpyriment compiles entire experiments into single, self-contained executable binaries with zero runtime dependencies. This drastically simplifies distribution to collaborators and testing computers. The programming interface, inspired by Expyriment (Krause Lindemann, 2014), was designed to be human friendly. The library includes an array of visual stimuli (text, shapes, images, Gabor patches, motion clouds, ...) and audio capabilities (WAV playback and tone generation). While developing Goxpyriment, we focused on timing reliability. Input events are timestamped by the operating system at hardware-interrupt time, so reaction times are computed by subtracting two OS-level timestamps rather than relying on continuous polling. Go's garbage collector can be disabled, greatly reducing the probability of unpredictable pauses that could corrupt stimulus timing. Finally, a set of over forty psychology experiments implemented in Goxpyriment are provided that promote not only learning by humans but also improve the ability of modern AI-assisted coding tools to help program experiments. The framework is released under the GNU General Public License v3 and is freely available at https://github.com/chrplr/goxpyriment.
Abstract How the brain encodes abstract concepts remains poorly understood. Current theories propose that, in brains and computers alike, word meanings are represented by vectors of neural activation whose similarities reflect semantic relationships. Here, we tested whether this hypothesis also applies to abstract concepts of elementary mathematics. We collected behavioral, 7 Tesla functional MRI and magneto-encephalography (MEG) data and used representational similarity analysis to ask where, when and how fifteen concepts of integers, fractions, and geometric shapes are encoded in the adult brain. Behavioral similarity ratings revealed a rich conceptual structure characterized by both categorical distinctions (numbers vs shapes, integers vs fractions), a numerical distance effect for integers, and systematic correspondences between items involving the same number (e.g. three, third, triangle). Functional MRI identified a bilateral cortical network whose neural encodings of concepts correlated with their semantic similarity, overlapping with classic math-responsive regions and encompassing IPS and ITG as well as dorsolateral prefrontal cortex (dlPFC). A double dissociation was observed, with a preference for arithmetic in the right anterior intraparietal sulcus (IPS), and for geometry in left inferior temporal gyrus (ITG) and bilateral posterior IPS. MEG revealed that a semantic neural code common to written words and symbols is activated by about 230 ms, again primarily distinguishing integers, fractions and geometry concepts. Together, these findings suggest that mathematical concepts are organized in the brain along both categorical and numerical dimensions, with overlapping but partially distinct sites supporting arithmetic and geometry domains.
When humans and large language models (LLMs) process the same text, activations in the LLMs correlate with brain activity measured, e.g., with functional magnetic resonance imaging (fMRI). Moreover, it has been shown that, as the training of an LLM progresses, the performance in predicting brain activity from its internal activations improves more in the left hemisphere than in the right one. The aim of the present work is to understand which kind of competence acquired by the LLMs underlies the emergence of this left-right asymmetry. Using the OLMo-2 7B language model at various training checkpoints and fMRI data from English participants, we compare the evolution of the left-right asymmetry in brain scores alongside performance on several benchmarks. We observe that the asymmetry co-emerges with the formal linguistic abilities of the LLM. These abilities are demonstrated in two ways: by the model's capacity to assign a higher probability to an acceptable sentence than to a grammatically unacceptable one within a minimal contrasting pair, or its ability to produce well-formed text. On the opposite, the left-right asymmetry does not correlate with the performance on arithmetic or Dyck language tasks; nor with text-based tasks involving world knowledge and reasoning. We generalize these results to another family of LLMs (Pythia) and another language, namely French. Our observations indicate that the left-right asymmetry in brain predictivity matches the progress in formal linguistic competence (knowledge of linguistic patterns).
How do the language areas of the human brain combine multiple words into meaningful phrases and sentences remains ill-understood. Here, to address this question, we determined the response profile of temporal and inferior frontal language areas to the composition of up to four words into phrases. We tested whether brain activity increases with the number of merged words, and whether this profile differs for noun and verb phrases. To this aim, we used fMRI to quantify the brain responses to individual noun and verb phrases of varying length and to tightly matched word lists. Increasing phrase length was associated to an increase in activation in all regions of the temporo-frontal language network. The effect was more pronounced for phrases built around verbs than for phrases built around nouns, suggesting that verbs involve a more complex syntactic tree structure than nouns. Even with word lists, several regions, notably the inferior frontal gyrus (IFG) pars triangularis and opercularis and the posterior superior temporal sulcus showed clear increases in activity with the length of sequences, although the words could not be merged into phrases. By contrast, other regions (IFG pars orbitalis, anterior temporal lobe, temporo-parietal junction) did not react to scrambled word lists. Those different functional response profiles inform theories of how composition is implemented in the human brain.
While deep learning has enabled the decoding of language from intracranial brain recordings, achieving this with non-invasive recordings remains an open challenge. We introduce a deep learning pipeline to decode individual words from electro- (EEG) and magneto-encephalography (MEG) signals. We evaluate our approach on seven public datasets and two datasets which we collect ourselves, amounting to a total of 723 participants reading or listening to five million words in three languages. Our model outperforms existing methods consistently across participants, devices, languages, and tasks, and can decode words absent from the training set. Our analyses highlight the importance of the recording device and experimental protocol: MEG and reading are easier to decode than EEG and listening, and decoding performance consistently increases with the amount of data used for training and for averaging during testing. Overall, our findings delineate the path and remaining challenges towards building non-invasive brain decoders for natural language.
We investigate optimal strategies for decoding perceived natural speech from fMRI data acquired from a limited number of participants. Leveraging Lebel et al. (2023)'s dataset of 8 participants, we first demonstrate the effectiveness of training deep neural networks to predict LLM-derived text representations from fMRI activity. Then, in this data regime, we observe that multi-subject training does not improve decoding accuracy compared to single-subject approach. Furthermore, training on similar or different stimuli across subjects has a negligible effect on decoding accuracy. Finally, we find that our decoders better model syntactic than semantic features, and that stories containing sentences with complex syntax or rich semantic content are more challenging to decode. While our results demonstrate the benefits of having extensive data per participant (deep phenotyping), they suggest that leveraging multi-subject for natural speech decoding likely requires deeper phenotyping or a substantially larger cohort.
The ability to compose complex mental representations by recombining simpler primitives is a characteristic of the human brain that manifests itself in a variety of domains such as spoken or written language, mathematics, or social reasoning. While it is known that these domains rest on partially distinct brain networks, they are rarely investigated together in the same paradigm. Here, we present a systematic functional MRI study of written and spoken word and sentence processing in each of these domains. In a true-false judgment task, adult participants were presented with a hierarchy of meaningless and meaningful stimuli bearing on various domains of general semantic knowledge and mathematics. The results show that distinct brain regions are activated by (1) syntax and semantics; (2) math versus non-math knowledge; (3) different domains of mathematics, such as geometry versus arithmetic; (4) social knowledge versus other knowledge. All of these regions are activated essentially identically in the written and spoken modalities. The results were replicated in a group of adolescents. Our experiments show that the human brain comprises distinct amodal networks for various domains of linguistic and semantic knowledge, and provide a simple paradigm to dissect them within a short fMRI session.
Deep learning has recently enabled the decoding of language from the neural activity of a few participants with electrodes implanted inside their brain. However, reliably decoding words from non-invasive recordings remains an open challenge. To tackle this issue, we introduce a novel deep learning pipeline to decode individual words from non-invasive electro- (EEG) and magneto-encephalography (MEG) signals. We train and evaluate our approach on an unprecedentedly large number of participants (723) exposed to five million words either written or spoken in English, French or Dutch. Our model outperforms existing methods consistently across participants, devices, languages, and tasks, and can decode words absent from the training set. Our analyses highlight the importance of the recording device and experimental protocol: MEG and reading are easier to decode than EEG and listening, respectively, and it is preferable to collect a large amount of data per participant than to repeat stimuli across a large number of participants. Furthermore, decoding performance consistently increases with the amount of (i) data used for training and (ii) data used for averaging during testing. Finally, single-word predictions show that our model effectively relies on word semantics but also captures syntactic and surface properties such as part-of-speech, word length and even individual letters, especially in the reading condition. Overall, our findings delineate the path and remaining challenges towards building non-invasive brain decoders for natural language.
A central feature in music is the hierarchical organization of its components. Musical pieces are not a simple concatenation of chords, but are characterized by rhythmic and harmonic structures. Here, we explore if sensitivity to music structure might emerge in the absence of any experience with musical stimuli. For this, we tested if rats detect the difference between structured and unstructured musical excerpts and compared their performance with that of humans. Structured melodies were excerpts of Mozart's sonatas. Unstructured melodies were created by the recombination of fragments of different sonatas. We trained listeners (both human participants and Long-Evans rats) with a set of structured and unstructured excerpts, and tested them with completely novel excerpts they had not heard before. After hundreds of training trials, rats were able to tell apart novel structured from unstructured melodies. Human listeners required only a few trials to reach better performance than rats. Interestingly, such performance was increased in humans when tonality changes were included, while it decreased to chance in rats. Our results suggest that, with enough training, rats might learn to discriminate acoustic differences differentiating hierarchical music structures from unstructured excerpts. More importantly, the results point toward species-specific adaptations on how tonality is processed.
The understanding of the human brain is one of the main scientific challenges of the twenty-first century. In the early 2000s, the French Atomic Energy Commission launched a program to conceive and build a human magnetic resonance imaging scanner operating at 11.7 T. We have now acquired human brain images in vivo at such a magnetic field. We deployed parallel transmission tools to mitigate the radiofrequency field inhomogeneity problem and tame the specific absorption rate. The safety of human imaging at such high field strength was demonstrated using physiological, vestibular, behavioral and genotoxicity measurements on the imaged volunteers. Our technology yields T2 and T2*-weighted images reaching mesoscale resolutions within short acquisition times and with a high signal and contrast-to-noise ratio. In a technological tour de force, a whole-body 11.7-T MRI scanner has been developed. Here images of the human brain are presented while safety for the imaged human volunteers has been ascertained.
Pseudowords are letter strings that look like words but are not words. They are used in psycholinguistic research, particularly in tasks such as lexical decision. In this context, it is essential that the pseudowords respect the orthographic statistics of the target language. Pseudowords that violate them would be too easy to reject in a lexical decision and would not enforce word recognition on real words. We propose a new pseudoword generator, UniPseudo, using an algorithm based on Markov chains of orthographic n-grams. It generates pseudowords from a customizable database, which allows one to control the characteristics of the items. It can produce pseudowords in any language, in orthographic or phonological form. It is possible to generate pseudowords with specific characteristics, such as frequency of letters, bigrams, trigrams, or quadrigrams, number of syllables, frequency of biphones, and number of morphemes. Thus, from a list of words composed of verbs, nouns, adjectives, or adverbs, UniPseudo can create pseudowords resembling verbs, nouns, adjectives, or adverbs in any language using an alphabetic or syllabic system.
Over the past decade, studies of naturalistic language processing where participants are scanned while listening to continuous text have flourished. Using word embeddings at first, then large language models, researchers have created encoding models to analyze the brain signals. Presenting these models with the same text as the participants allows to identify brain areas where there is a significant correlation between the functional magnetic resonance imaging (fMRI) time series and the ones predicted by the models' artificial neurons. One intriguing finding from these studies is that they have revealed highly symmetric bilateral activation patterns, somewhat at odds with the well-known left lateralization of language processing. Here, we report analyses of an fMRI dataset where we manipulate the complexity of large language models, testing 28 pretrained models from 8 different families, ranging from 124M to 14.2B parameters. First, we observe that the performance of models in predicting brain responses follows a scaling law, where the fit with brain activity increases linearly with the logarithm of the number of parameters of the model (and its performance on natural language processing tasks). Second, although this effect is present in both hemispheres, it is stronger in the left than in the right hemisphere. Specifically, the left-right difference in brain correlation follows a scaling law with the number of parameters. This finding reconciles computational analyses of brain activity using large language models with the classic observation from aphasic patients showing left hemisphere dominance for language.
Abstract The understanding of the human brain is one of the main scientific challenges of the 21st century. In this context, in the early 2000s the French Atomic Energy Commission (CEA) launched a program to conceive and build the first human MRI scanner operating at 11.7T. More than a decade of developments followed to deliver the magnet while six more years were necessary to complete the commissioning and finally obtain approval from the regulatory agencies to acquire the first ever human brain images in vivo at such magnetic field. We deployed parallel transmission tools to mitigate the radiofrequency field inhomogeneity problem and tame the specific absorption rate. To assure the safety of human imaging at such high field strength, we performed physiological, vestibular, behavioral and genotoxicity measurements on the volunteers. The data shows no evidence of adverse effects. The unprecedented field strength combined with the acquisition techniques deployed yield T2 and T2*-weighted images reaching mesoscale resolutions within short acquisition times and with a high signal and contrast to noise ratio.
Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether two models share a similar representational geometry, but they cannot explain why. Here, we propose a new, simple, approach to address this question. This approach maps neural activity in each model layer onto a set of interpretable linguistic features and quantifies how much each of them drives similarities and differences between models. We use this approach to compare 43 language models across 10 families, including decoder Transformers, State-Space Models, and Recurrent Neural Networks. We find that model-level similarity is driven most strongly by release date, a proxy for general LLM development, and model family, suggesting that linguistic signatures are not primarily shaped by scale or architecture class. Overall, our approach provides a way to link theoretically-motivated symbolic descriptions to neural representations and can readily be extended to other domains such as speech and vision, and to other neural systems such as biological brains.
We often express our thoughts through words, but thinking goes well beyond language. Here we focus on an elementary but basic thinking process, disjunction elimination, elicited by elementary visual scenes deprived of linguistic content, describing its neural and oculomotor correlates. We track two main components of a nonverbal deductive process: the construction of a logical representation (A or B), and its simplification by deduction (not A, therefore B). We identify the network active in the two phases and show that in the latter, but not in the former, it overlaps with areas known to respond to verbal logical reasoning. Oculomotor markers consistently differentiate logical processing induced by the construction of a representation, its simplification by deductive inference, and its maintenance when inferences cannot be drawn. Our results reveal how integrative logical processes incorporate novel experience in the flow of thoughts induced by visual scenes.
Two fundamental questions in neurolinguistics concerns the brain regions that integrate information beyond the lexical level, and the size of their window of integration. To address these questions we introduce a new approach named masked-attention generation. It uses GPT-2 transformers to generate word embeddings that capture a fixed amount of contextual information. We then tested whether these embeddings could predict fMRI brain activity in humans listening to naturalistic text. The results showed that most of the cortex within the language network is sensitive to contextual information, and that the right hemisphere is more sensitive to longer contexts than the left. Masked-attention generation supports previous analyses of context-sensitivity in the brain, and complements them by quantifying the window size of context integration per voxel.