Abstract This study explores the potential of Generative AI (GenAI), specifically ChatGPT, for the acquisition and semantic organization of collocations in support of collocation-resource construction. We examine ChatGPT’s ability to generate English collocation candidates for eleven semantic collocation categories formalized as Lexical Functions (LFs). LFs are treated as a fine-grained intermediate analytical layer from which broader, user-oriented semantic groupings can be derived. Two prompting strategies are compared: a few-shot approach, in which the model receives twenty illustrative examples of a given LF and generates fifty further candidates, and an iterative profile-refinement approach, in which a lexicographer guides the model in constructing a semantic and syntactic LF profile subsequently used for sample generation. The outputs were manually evaluated by specialists in Explanatory Combinatorial Lexicology. The results show that the few-shot strategy captures some central semantic properties of several LFs, while profile refinement generally yields higher accuracy and greater semantic coherence, especially for semantically complex LFs. The acquired LF profiles were then used to prompt ChatGPT to construct a collocation-dictionary entry for Fear, modelled on the format of the Oxford Collocations Dictionary (OCD). The comparison shows that LF-guided prompting favours semantic differentiation and explicit structural organization by sense and semantic function, whereas the OCD entry displays greater idiomatic selectivity and lexicographic compactness. Overall, LF-guided interaction with GenAI appears to be a promising semi-automatic instrument for compiling semantically structured collocation resources that can support dictionary construction, although expert validation, generalization, and post-editing remain necessary.
Frequent English verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'make a decision' vs. 'make a cake'. Whether language models represent this distinction remains unclear. We introduce a large-scale controlled dataset of minimally varying English sentence series in which the same context contains the same verb in light-verb and full-verb uses. Two probing experiments show that language models differentiate between these uses even in minimal contexts and exhibit separable patterns across object types. We release the dataset, generation code, and materials as a reusable resource. The framework supports extensions to broader contexts, additional verbs, and other languages.
This paper investigates to what extent the integration of morphological information can improve subword tokenization and thus also language modeling performance. We focus on Spanish, a language with fusional morphology, where subword segmentation can benefit from linguistic structure. Instead of relying on purely data-driven strategies like Byte Pair Encoding (BPE), we explore a linguistically grounded approach: training a tokenizer on morphologically segmented data. To do so, we develop a semi-supervised segmentation model for Spanish, building gold-standard datasets to guide and evaluate it. We then use this tokenizer to pre-train a masked language model and assess its performance on several downstream tasks. Our results show improvements over a baseline with a standard tokenizer, supporting our hypothesis that morphology-aware tokenization offers a viable and principled alternative for improving language modeling.
Graded readers (GRs) are a popular language-learning resource, as they provide contextualized input adapted to any level. Still, their creation process is non-systematic and their quantity is limited. This article investigates, firstly, the progression of linguistic complexity in a series of Spanish GRs of consecutive levels, and secondly, whether literary works (LWs) targeted at specific age groups of L1 speakers exhibit a similar gradation, thus representing suitable complementary material. For this purpose, we (1) fitted two random forests on 40 complexity features computed by processing 50 GRs, 50 LWs, and 8,585 graded lexical items; (2) performed intergroup comparisons with a reference corpus through permutation tests on the four features most informative to the random forests; and (3) further explored vocabulary using distributional techniques. Our findings indicate that complexity does not progress uniformly across levels or linguistic dimensions for both GRs and LWs. Moreover, LWs only differ substantially from GRs in their lowest level. This demonstrates the importance of considering quantitative measures such as the ones presented in this study to develop balanced language teaching materials.
Hate speech (HS) classifiers do not perform equally well in detecting hateful expressions towards different target identities. They also demonstrate systematic biases in predicted hatefulness scores. Tapping on two recently proposed functionality test datasets for HS detection, we quantitatively analyze the impact of different factors on HS prediction. Experiments on popular industrial and academic models demonstrate that HS detectors assign a higher hatefulness score merely based on the mention of specific target identities. Besides, models often confuse hatefulness and the polarity of emotions. This result is worrisome as the effort to build HS detectors might harm the vulnerable identity groups we wish to protect: posts expressing anger or disapproval of hate expressions might be flagged as hateful themselves. We also carry out a study inspired by social psychology theory, which reveals that the accuracy of hatefulness prediction correlates strongly with the intensity of the stereotype.
While the competence of LLMs to cope with agreement constraints has been widely tested for English, only a very limited number of works deals with morphologically rich(er) languages. In this work, we experiment with 25 mono- and multilingual LLMs, applying them to a collection of more than 5,000 test examples that cover the main agreement phenomena in three Romance languages (Italian, Portuguese, and Spanish) and one Slavic Language (Russian). We identify which of the agreement phenomena are most difficult for which models and challenge some common assumptions of what makes a good model. The test suites into which the test examples are organized are openly available and can be easily adapted to other agreement phenomena and other languages for further research.
In the modern labor market, accurate matching of job vacancies with suitable candidate CVs is critical. We present a novel multilingual knowledge graph-based framework designed to enhance the matching by accurately extracting the skills requested by a job and provided by a job seeker in a multilingual setting and aligning them via the standardized skill labels of the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy. The proposed framework employs a combination of state-ofthe-art techniques to extract relevant skills from job postings and candidate experiences. These extracted skills are then filtered and mapped to the ESCO taxonomy and integrated into a multilingual knowledge graph that incorporates hierarchical relationships and cross-linguistic variations through embeddings. Our experiments demonstrate a significant improvement of the matching quality compared to the state of the art.
Advances in cognitive science, sensing technologies, the arts and creative industries are paving the way for a deeper understanding of the behaviour of individuals regarding the land/soundscape they live in. Through a symbiotic relationship between artists, scientists and technology experts ReSilence explores the borders between sound and silence in a changing world by producing sound awareness in urban spaces (not only reducing the intensity of noise, but also considering it as energy producer and designing positive sounds, sounds we want to preserve and multiply). More specifically ReSilence focuses in musical experience design centred on the active participation of citizens, in the new silence of mobility, in the acoustic perception of outdoor urban soundscapes and in enhancing experiences for people with hearing and vision impairments.
Apart from being an economic struggle, migration is first of all a societal challenge; most migrants come from different cultural and social contexts, do not speak the language of the host country, and are not familiar with its societal, administrative, and labour market infrastructure. This leaves them in need of dedicated personal assistance during their reception and integration. However, due to the continuously high number of people in need of attendance, public administrations and non-governmental organizations are often overstrained by this task. The objective of the Welcome Platform is to address the most pressing needs of migrants. The Platform incorporates advanced Embodied Conversational Agent and Virtual Reality technologies to support migrants in the context of reception, integration, and social inclusion in the host country. It has been successfully evaluated in trials with migrants in three European countries in view of potentially deviating needs at the municipal, regional, and national levels, respectively: the City of Hamm in Germany, Catalonia in Spain, and Greece. The results show that intelligent technologies can be a valuable supplementary tool for reducing the workload of personnel involved in migrant reception, integration, and inclusion.
In the context of the increasingly globalised economy and labour market, recruitment agencies face the challenge to deal with a magnitude of job offers and job applications written in a variety of languages, formats, and styles. Quite often, this leads to a suboptimal evaluation of the CVs of job seekers with respect to their relevance to a job offer. To address this challenge, we propose an interactive system that follows the ``human-in-the-loop'' approach, actively involving recruiters in the job offer -- applicant CV matching. The system uses a fine-tuned state-of-the-art classification model that aligns job seeker CVs with labels of the {\it European Skills, Competences, Qualifications and Occupations} taxonomy to propose an initial match between job offers with the CVs of job candidates. This match is refined in sequential LLM driven-interaction with the recruiter, which culminates in CV relevance scores and reports that justify them.
We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results and findings, which include that just 13% of papers had (i) sufficiently low barriers to reproduction, and (ii) enough obtainable information, to be considered for reproduction, and that all but one of the experiments we selected for reproduction was discovered to have flaws that made the meaningfulness of conducting a reproduction questionable. As a result, we had to change our coordinated study design from a reproduce approach to a standardise-then-reproduce-twice approach. Our overall (negative) finding that the great majority of human evaluations in NLP is not repeatable and/or not reproducible and/or too flawed to justify reproduction, paints a dire picture, but presents an opportunity for a rethink about how to design and report human evaluations in NLP.
As pointed out by several scholars, current research on hate speech (HS) recognition is characterized by unsystematic data creation strategies and diverging annotation schemata. Subsequently, supervised-learning models tend to generalize poorly to datasets they were not trained on, and the performance of the models trained on datasets labeled using different HS taxonomies cannot be compared. To ease this problem, we propose applying extremely weak supervision that only relies on the class name rather than on class samples from the annotated data. We demonstrate the effectiveness of a state-of-the-art weakly-supervised text classification model in various in-dataset and cross-dataset settings. Furthermore, we conduct an in-depth quantitative and qualitative analysis of the source of poor generalizability of HS classification models.
Stress can be considered a mental/physiological reaction in conditions of high discomfort and challenging situations. The levels of stress can be reflected in both the physiological responses and speech signals of a person. Therefore the study of the fusion of the two modalities is of great interest. For this cause, public datasets are necessary so that the different proposed solutions can be comparable. In this work, a publicly available multimodal dataset for stress detection is introduced, including physiological signals and speech cues data. The physiological signals include electrocardiograph (ECG), respiration (RSP), and inertial measurement unit (IMU) sensors equipped in a smart vest. A data collection protocol was introduced to receive physiological and audio data based on alterations between well-known stressors and relaxation moments. Five subjects participated in the data collection, where both their physiological and audio signals were recorded by utilizing the developed smart vest and audio recording application. In addition, an analysis of the data and a decision-level fusion scheme is proposed. The analysis of physiological signals includes a massive feature extraction along with various fusion and feature selection methods. The audio analysis comprises a state-of-the-art feature extraction fed to a classifier to predict stress levels. Results from the analysis of audio and physiological signals are fused at a decision level for the final stress level detection, utilizing a machine learning algorithm. The whole framework was also tested in a real-life pilot scenario of disaster management, where users were acting as first responders while their stress was monitored in real time.
Nowadays, vast amounts of multimedia content are being produced, archived, and digitized, resulting in great troves of data of interest. Examples include user-generated content, such as images, videos, text, and audio posted by users on social media and wikis, or content provided through official publishers and distributors, such as digital libraries, organizations, and online museums. This digital content can serve as a valuable source of inspiration to the creative industries, such as architecture and gaming, to produce new innovative assets or to enhance and (re-)use existing ones. However, in its current form, this content is difficult to be reused and repurposed due to the lack of appropriate solutions for its retrieval, analysis, and integration into the design process. In this article, we present V4Design, a novel framework for the automatic content analysis, linking, and seamless transformation of heterogeneous multimedia content to help architects and virtual reality game designers establish innovative value chains and end-user applications. By integrating and intelligently combining state-of-the-art technologies in computer vision, 3-D generation, text analysis, generation and semantic integration, and interlinking, V4Design provides architects and video game designers with innovative tools to draw inspiration from archive footage and documentaries, inspiring and eventually supporting the design process.
Recognizing and categorizing lexical collocations in context is useful for language learning, dictionary compilation and downstream NLP. However, it is a challenging task due to the varying degrees of frozenness lexical collocations exhibit. In this paper, we put forward a sequence tagging BERT-based model enhanced with a graph-aware transformer architecture, which we evaluate on the task of collocation recognition in context. Our results suggest that explicitly encoding syntactic dependencies in the model architecture is helpful, and provide insights on differences in collocation typification in English, Spanish and French.
According to urban planner Kevin Lynch, imageability is the ability of a physical object to evoke a strong image in any viewer, making it memorable. The concept of imageability is important for architects and urban designers, so that their creations meet the needs of the citizens and improve the aesthetics of the place. Recently, computer vision and textual analysis techniques have been investigated for calculating the imageability of a place. In this paper, we propose a novel multi-modal system that utilises both visual and textual analysis methods to estimate the imageability score of a place. In addition, an image sentiment analysis deep learning model had been developed to provide supplementary information about the sentiment that is evoked to citizens by urban locations. Finally, a text generation algorithm is used to provide an explanation of the information extracted by the data analysis in a form of text to facilitate the works of architects and urban designers.
Addressing hate speech in online spaces has been conceptualized as a classification task that uses Natural Language Processing (NLP) techniques. Through this conceptualization, the hate speech detection task has relied on common conventions and practices from NLP. For instance, inter-annotator agreement is conceptualized as a way to measure dataset quality and certain metrics and benchmarks are used to assure model generalization. However, hate speech is a deeply complex and situated concept that eludes such static and disembodied practices. In this position paper, we critically reflect on these methodologies for hate speech detection, we argue that many conventions in NLP are poorly suited for the problem and encourage researchers to develop methods that are more appropriate for the task.
Abstract The correspondence between the communicative intention of a speaker in terms of Information Structure and the way this speaker reflects communicative aspects by means of prosody have been a fruitful field of study in Linguistics. However, text-to-speech applications still lack the variability and richness found in human speech in terms of how humans display their communication skills. Some attempts were made in the past to model one aspect of Information Structure, namely thematicity for its application to intonation generation in text-to-speech technologies. Yet, these applications suffer from two limitations: (i) they draw upon a small number of made-up simple question-answer pairs rather than on real (spoken or written) corpus material; and (ii) they do not explore whether any other interpretation would better suit a wider range of textual genres beyond dialogs. In this paper, two different interpretations of thematicity in the field of speech technologies are examined: the state-of-art binary (and flat) theme-rheme, and the hierarchical thematicity defined by Igor Mel’čuk within the Meaning-Text Theory. The outcome of the experiments on a corpus of native speakers of US English suggests that the latter interpretation of thematicity has a versatile implementation potential for text-to-speech applications of the Information Structure–prosody interface.
Jens Grivolla合作论文数Barcelona Media Innovation Centre5
Wolfgang Minker合作论文数Faculty of Engineering and Computer Science,University of Ulm
Institute of Information Technology5