This study investigates emotion recognition from spontaneous Tunisian Dialect (TD) speech, focusing on the challenges posed by natural, varied vocal expressions. We evaluate three acoustic feature sets—13 MFCCs, eGeMAPSv02, and emobase—across classical machine learning classifiers and deep learning models like LSTM and Transformer Encoder. To address class imbalance in the SERTUS dataset, especially for Joy and Neutral emotions, we applied data augmentation techniques such as white noise addition, pitch shifting, and time stretching, and enriched the corpus with manually validated speech segments from other TD sources. Our results highlight that random forest combined with emobase features yields the highest overall accuracy (67
Online misinformation increasingly exploits images that contain deceptive text, such as memes, screenshots of headlines, or edited posters. In such cases, the visual content may look plausible while the overlaid writing or associated title deliberately misrepresents the depicted scene or event. Existing fake news detection systems largely focus on textual articles or coarse image-article alignment, and are ill-equipped to capture this fine-grained, within-image incongruity. In this paper, we address the task of detecting fake news through deceptive text in images, defined as the lack of semantic conformity between an image and the textual content presented on it or in its title. We propose a multimodal consistency methodology that combines optical character recognition (OCR), deep language modeling, and visual representation learning to assess whether an image and its in-image or associated text are mutually consistent. Our approach integrates (i) a text encoder for the extracted writing and accompanying caption, (ii) a visual encoder for the image, and (iii) a cross-modal interaction module that explicitly models semantic alignment. We experiment on the widely used Weibo multimodal rumor detection corpus introduced by Wang et al. [4], and derive an OCR-filtered subset in which images contain non-trivial textual content. On this subset, a BERT-based text-only baseline achieves 0.79 F 1 and outperforms an imageonly ResNet baseline (0.70 F1), while both clearly improve over a random classifier. These initial results confirm the importance of textual information—including text extracted from images—for detecting fake news in multimodal posts.
Large Language Models (LLMs) have demonstrated strong performance in Natural Language Processing (NLP), yet their evaluation in Arabic—particularly low-resource varieties like the Tunisian Dialect (TD)—remains a pressing need. This study investigates data augmentation using GPT-4o, Arabic-pretrained models, and LLMs with few-shot learning for spontaneous Speech Emotion Recognition (SER). We employ both open-source and proprietary Automatic Speech Recognition (ASR) systems to generate transcriptions of TD speech, which are then used as input features for SER models. Our findings show that the best performance was achieved with AraBERT, reaching an accuracy of 77
The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current approaches rely on unstructured text corpora, this study explores two complementary strategies for leveraging structured knowledge from the UMLS Metathesaurus: (i) Continual pretraining that embeds knowledge into model parameters, and (ii) Graph Retrieval-Augmented Generation (GraphRAG) that consults a knowledge graph at inference time. We first construct a large-scale biomedical knowledge graph from UMLS (3.4 million concepts and 34.2 million relations), stored in Neo4j for efficient querying. We then derive a 100-million-token textual corpus from this graph to continually pretrain two models: BERTUMLS (from BERT) and BioBERTUMLS (from BioBERT). We evaluate these models on six BLURB (Biomedical Language Understanding and Reasoning Benchmark) datasets spanning five task types and evaluate GraphRAG on the two QA (Question Answering) datasets (PubMedQA, BioASQ). On BLURB tasks, BERTUMLS improves over BERT, with the largest gains on knowledge-intensive QA. Effects on BioBERT are more nuanced, suggesting diminishing returns when the base model already encodes substantial biomedical text knowledge. Finally, augmenting LLaMA 3-8B with our GraphRAG pipeline yields over than 3 points accuracy on PubMedQA and 5 points on BioASQ without any retraining, delivering transparent, multi-hop, and easily updated knowledge access. We release the processed UMLS Neo4j graph to support reproducibility.
Effective detection of cyberbullying requires understanding both textual and visual signals, including images with embedded text and user generated comments. This need is even more evident in low resource and multilingual environments such as Tunisia. In this context, this paper establishes CyberDTD (Cyberbullying Detection in Tunisian Dialect), a multimodal dataset designed to support research on cyberbullying detection in the Tunisian Dialect (TD). With 10,802 images across five categories, humor, sarcasm, hate, violence, and neutral. We present, to the best of our knowledge, the first cyberbullying dataset in TD. We provide a comprehensive description covering a wide range of online harassment, while also including neutral examples for balanced analysis. Key challenges such as class imbalance, multimodality, and cultural specificity are highlighted. CyberDTD represents an important resource for building and evaluating machine learning models in low-resource settings, supporting the development of more robust and culturally aware cyberbullying detection systems.
In this paper, we are interested in developing a sentiment analysis (SA) system on the basis of machine learning (ML) techniques. We propose a Libyan dialect twitter dataset which we have built. It contains 6,000 comments and tweets cleaned, pre-processed, and annotated. To evaluate the performance of sentiment classification of the tweets and comments, four machine learning classifiers have been applied. These results show that all classifiers achieved good results. Furthermore, we have created a basic sentiment lexicon which includes words and phrases with sentiment polarity (positive, negative, and neutral). This sentiment lexicon is a valuable resource for the Libyan dialect.
The effects of psychological crises are evolving at an astounding rate nowadays, presenting a significant challenge for everyone involved in tracking these disorders. Therefore, we propose in this paper a hybrid approach based on linguistic processing and numerical techniques allowing to: (i) identify the presence of psychological emergencies among social network users by analyzing their textual production, (ii) determine the specific type of emergency case, (iii) elaborate a graph for each type of emergency, reflecting the different dimensions linked to the psychological emergency, allowing for a better diagnosis of the situation and providing an overall view of the crisis type, (iv) combine the separate graphs for each emergency to address the various semantic aspects. The work was accomplished using advanced language model techniques, knowledge graphs and neural network graphs. The combination of these techniques ensures that their advantages are leveraged while overcoming their limitations in terms of result generalization. The evaluation of different parts related to detecting the presence of psychological problems, predicting specific type of emergency cases, and detecting links between knowledge graphs was measured using the F-measure metric. The values derived from this measure, corresponding to the evaluation of these three tasks, are, respectively, 83%, 87% and 80%. For the evaluation of the elaboration of each graph related to specific type of emergency cases, this was accomplished using qualitative metric standards. The results obtained can be considered encouraging given the significant scale of our approach.
As the volume of information exchanged on the Internet has increased, Aspect-Based Sentiment Analysis (ABSA) has become an indispensable task for addressing the needs of various areas that rely on the exploitation of public opinion. This task first involves extracting the different aspects (e.g., price, quality) of an entity (e.g., laptops) and then assigning sentiment polarity (e.g., positive, negative) to them. In this paper, we focus exclusively on the Aspect Extraction (AE) task, which is the most critical component of ABSA. First, we provide a comprehensive overview of the AE task, which involves identifying aspects from text. Subsequently, due to the significant impact of datasets on the performance of aspect extraction models, a statistical analysis is performed on various datasets created for the AE task, considering factors such as availability, sources, and more. then, we thoroughly review the studies in this field, classifying them into four main approaches: linguistic knowledge-based approach, machine learning-based approach, deep learning-based approach, and hybrid approach. Finally, we evaluate the strengths and weaknesses of the proposed approaches, providing insights for researchers to stay informed about the latest advancements in this area. For future work, we explore the challenges associated with AE and provide potential directions for future research to further advance the field.
Social networks have become a major source for the expression and dissemination of information and news in various languages. Users can express their opinions and sentiments towards different topics through written comments in dialectal languages. The objective of this work is to build a multi-class sentiment analysis model for Tunisian dialect comments scraped from social media websites, based on a lexical ontology and deep learning models. First, we created the TDCOR_Train, a multi-domain Tunisian Arabic dialect dataset. Subsequently, we developed the TDSO lexical ontology, which will be projected to automatically annotate each comment as positive, negative, or neutral. Further-more, we trained four deep learning models, namely Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Bi-LSTM, and RCNN (RNN + CNN) on the annotated dataset. Both RCNN and Bi-LSTM produced the best results, achieving accuracy rates of 81.27 and 82.01. To assess the performance of our sentiment analysis method, we fine-tuned the bert-base-arabertv02 transfer transformer which returned an accuracy rate of 81.17. Moreover, all models exceeded the accuracy rate of 85.00 when evaluated on two test corpora, the first is TSAC_New, and the second is TSAC_New_TDSO, which is TSAC_New dataset auto-annotated by the TDSO lexical ontology.
Aiming to support researchers by answering their questions, this paper presents the first module of a virtual assistant for Tunisian Arabic (TA): Automatic Speech Recognition (ASR). Given the specialized lexicon used, which is characterized by a high rate of Code-Switching (CS), our primary objective is to create a new dataset, Tunacad, tailored to this context. Tunacad comprises 12.56 hours of CS spontaneous speech in the academic domain. Additionally, we experiment with our proposed CS corpus on predefined ASR models, Whisper and Wav2vec2-XLR-S. To enhance ASR performance on our corpus, we apply several techniques, such as automatic orthographic correction, normalization, and prompt-based correction. Additionally, we incorporate human evaluation to assess transcription quality beyond surface-level accuracy. Through these experiments and CS errors analysis statistics, we highlight the persistent challenges in processing TA speech.
With the advent of complex language models and the massive amount of data available on the web, students have had an easier time committing plagiarism. This research describes a web-based system for identifying plagiarism in student reports using intrinsic analysis. To detect plagiarism, we use a combination of stylistic and semantic features as well as a similarity matching technique. We experimented with a dataset of scientific papers mostly published in French, the predominant language in our institutions. Our plagiarism detection method examines the writing style of suspect documents, locates relevant sources on the internet, and compares them to the suspicious documents using external text matching. The preliminary results are promising, with our intrinsic and extrinsic methods reaching an F-score of 40.3% and 89% accuracy, respectively.
In contrast to written texts and prepared speeches, conversational/spontaneous speech has a very high degree of freedom and includes a huge number of disfluencies. Detecting disfluencies using transformer-based models has advanced state-of-the-art performance. In this work, we aim to process disfluencies in the spontaneous tunisian dialect speech by generating fluent utterances from disfluent transcripts. We propose a transformer-based model by fine-tuning the pre-trained T5 language model. Using this model, we achieved an F-Measure score of 74,71
This paper introduces SERTUS (Speech Emotion Recognition TUnisian Spontaneous), an extensive dataset collection intended to propel research in Speech Emotion Recognition (SER), particularly within the realm of Tunisian Dialect (TD). SERTUS encompasses both registers of the Tunisian Dialect: the Popular (familiar) register and the intellectual register, capturing a diverse range of emotions in spontaneous environments and natural interactions across different regions of Tunisia. This work delineates the methodology utilized in crafting SERTUS, highlighting the challenges and strategies involved in capturing spontaneous interactions. Moreover, we underscore the importance of including TD and the multidomain nature of the dataset, illustrating its potential applications in various domains such as sports, politics, and culture. It's imperative to note that this corpus collection adhered to a rigorous protocol to ensure corpus quality, as it will undergo annotation in future research endeavors.
The adaptation of Large Language Models (LLMs) to specialized domains such as biomedicine holds significant promise for advancing Natural Language Processing (NLP) applications in healthcare and medical research. In this survey, we investigate various techniques and approaches for adapting LLMs to better understand and process biomedical text. We explore the challenges faced by LLMs in biomedical tasks, including the presence of specialized terminology and the scarcity of annotated datasets. The survey discusses adaptation techniques such as training LLMs from scratch, fine-tuning on biomedical data, and injecting domain-specific knowledge. Furthermore, we compare the efficacy of these techniques and highlight their advantages and challenges. Our examination reveals that while each approach offers distinct benefits, the choice depends on factors such as available resources, task requirements, and desired performance outcomes. Finally, we outline future research directions focusing on knowledge integration in the biomedical domain to enhance the capabilities of LLMs and their potential impact on healthcare and medical research.
As smartphones and social media usage grow among young people, particularly in Arab communities, the risk of encountering cyberbullying and harmful content increases. However, existing cyberbullying detection solutions are primarily tailored for English, leaving Arabic users underserved. This study addresses this gap by developing improved detection models for Arabic content. Through training various classifiers with annotated Arabic datasets, including traditional machine learning and deep learning techniques, this research aims to enhance the effectiveness of cyberbullying detection in Arabic online spaces, collected from three different platforms (Facebook, Twitter, and YouTube). Additionally, we conducted domain-specific data extraction from our existing datasets, focusing solely on political discourse. This was followed by testing the extracted data with deep learning algorithms. The results indicate that the proposed model outperforms other classifiers examined in the study. The overall enhancement achieved by the proposed model reaches an accuracy of 89
This paper provides an overview of the application of big data in psychology, exploring its relevance, potential applications, and associated challenges. It begins with a presentation of the concept of big data, including its architectures and technologies, emphasizing their importance in psychological research. After that, the paper delves into a set of existing data in the literature, taking into account high-level criteria such as volume, variety, and velocity. These criteria are expanded to include additional standards like ethical considerations, acquisition velocity, and value. Following that, the paper explores various studies about the analysis and processing of massive datasets in psychology, with an emphasis on study objectives and techniques employed, such as data warehouses and parallel processing. Finally, the paper concludes with a discussion on the criticisms of the use of big data in psychology and the significance of integrating other techniques such as artificial intelligence (AI) and natural language processing (NLP). Overall, the present study highlights the potential of big data in advancing psychological research, while acknowledging the need to address associated challenges and adopt complementary methodologies.
Sentiment analysis (SA) has emerged as a crucial computational method for extracting subjective information from text, facilitating organizations to transform unstructured opinions through actionable insights that drive strategic decision-making across domains covering from business intelligence to public policy formation [46]. Pre-training models for SA have gained significant attention for improving opinion extraction from text. In recent years, social media has become a crucial platform for customer engagement, with SA playing a key role in maintaining client loyalty. Extracting sentiments from comments and reviews is particularly challenging for under-resourced languages like the Tunisian Dialect (TD), which is written in both Arabizi and Arabic scripts. Despite advancements in SA, processing TD remains complex. In this study, BERT and CNN-Bidirectional LSTM models are employed to perform SA on unstructured data collected from Facebook. The dataset, TUNisian TElecom Sentiment Analysis (TUNTESA), consists of 27,080 Arabizi and 17,816 Arabic comments sourced from official telecommunications operators' Facebook pages. The comments are labeled as positive, negative, or neutral. The results demonstrate high accuracy (Acc), with the BERT Arabic model achieving 0.99 and the BERT Arabizi model reaching 0.94, outperforming existing studies. These findings highlight the practical applications of SA for businesses leveraging social media interactions. By effectively analyzing sentiments, telecom operators can enhance customer satisfaction, manage relationships, and extract valuable feedback, ultimately maintaining a competitive edge.
Data Augmentation (DA) is an important technique for limited resources languages to address the deficiency in data. It’s based on the generation of new instances to tackle imbalanced data. In this work, we implemented a data augmentation technique by using a dictionary to replace synonym or antonym words within the original text to generate new text. This must keep the semantic meaning of the text. In this paper, we present the Libyan Dialect Dictionary LDD which we have created in order to generate new comments/tweets from Arabic Sentiment Analysis dataset for Libyan Dialect in the Airline domain ASALDA. To evaluate the performance of sentiment classification of new augmented data, three deep learning (DL) models LSTM, CNN, and RNN were used. The result confirms that the LSTM model achieved the highest accuracy of 94% when using original data in the test set and augmented data in the train and dev set. The LDD proposed in this paper can be used for all researchers who work on Libyan dialects and data augmentation.
In the digital age, the proliferation of social media platforms such as Facebook, YouTube, and Twitter has led to an unprecedented surge in content produced in both standard languages and dialects. Translating dialects, such as Tunisian Dialect (TD), presents unique challenges for automatic translation systems due to their informal and regionally specific nature. This study aims to address these challenges by developing a robust translation model for Tunisian Dialect into English, a globally recognized formal language. Tunisian Dialect is often written using a mixture of Latin script and Arabic, while frequently incorporating foreign words, adding complexity to the translation task. Moreover, the scarcity of parallel Tunisian dialect to English corpora has hindered the development of accurate translation systems. To overcome these limitations, we created a parallel corpus between Tunisian Dialect and English by leveraging existing data between Tunisian Dialect and Modern Standard Arabic, applying data augmentation techniques to build a substantial trilingual dataset. Our research explores two translation approaches: one using Modern Standard Arabic as a pivot language, and the other translating directly from Tunisian dialect to English. The best model achieved a result of 67.89%, demonstrating significant improvements in translation accuracy and offering valuable insights for tackling dialectal translation challenges.
Abdelmajid Ben Hamadou合作论文数Higher Institute of Computer Science and Multimedia, Sfax University20
Rim Faiz合作论文数University of Carthage - IHEC
9