Purpose Research into harmful language detection is hindered by fragmented data sets and incompatible label schemas, which significantly limit evaluation. This paper aims to introduce evaluation of harmful language detection (EHLD): a general, extensible data framework supporting the integration, enrichment and evaluation of diverse harmful language resources, with a particular focus on Italian. Design/Methodology/Approach EHLD defines a unified schema with essential detection attributes (e.g. harmful_or_not and harm_subcategory), optional contextual dimensions (e.g. target, category and intersectionality), evaluation-oriented text complexity features and metadata for provenance. The authors’ instantiate EHLD by integrating three data sources: curated data sets from the literature, large-scale large language model (LLM)-based annotation from the Mappa dell’Intolleranza and LLM-generated data for inclusive language detection. They further introduce label_confidence to explicitly encode label reliability. Model selection follows an iterative cycle combining distributional analysis, linguistic diversity assessment and performance evaluation. Findings The resulting data set contains 237,956 instances with 18 features and exhibits substantial linguistic variety (Self-bilingual evaluation understudy drops from 0.759 at 1 gram to 0.107 at 4 grams). After iterative, data-centric refinement, generative pretrained transformer-5-nano is selected as the reference model and achieves an F1-score of 0.813 on binary harmful language detection, outperforming reported baselines such as bidirectional encoder representations from transformers (BERT)-based multitask models and other LLMs on comparable Italian settings. Research limitations/implications Due to differences in data sets, domains, label definitions and evaluation protocols, comparisons across studies are not fully controlled. The framework partly relies on automatically generated labels, which, despite confidence-aware handling, may introduce residual noise. Practical implications By combining standardized labels, provenance tracking and complexity-oriented evaluation features, EHLD supports the construction of reproducible data sets and richer model auditing, enabling a more transparent deployment of harmful language detection systems. Originality/value EHLD contributes a reusable, confidence-aware integration framework that connects different harmful language resources and allows iterative, data-driven model selection and evaluation, with a focus on Italian.
This paper introduces the HH4AI Methodology, a structured approach to assessing the impact of AI systems on human rights, focusing on compliance with the EU AI Act and addressing technical, ethical, and regulatory challenges. The paper highlights AIs transformative nature, driven by autonomy, data, and goal-oriented design, and how the EU AI Act promotes transparency, accountability, and safety. A key challenge is defining and assessing "high-risk" AI systems across industries, complicated by the lack of universally accepted standards and AIs rapid evolution. To address these challenges, the paper explores the relevance of ISO/IEC and IEEE standards, focusing on risk management, data quality, bias mitigation, and governance. It proposes a Fundamental Rights Impact Assessment (FRIA) methodology, a gate-based framework designed to isolate and assess risks through phases including an AI system overview, a human rights checklist, an impact assessment, and a final output phase. A filtering mechanism tailors the assessment to the system's characteristics, targeting areas like accountability, AI literacy, data governance, and transparency. The paper illustrates the FRIA methodology through a fictional case study of an automated healthcare triage service. The structured approach enables systematic filtering, comprehensive risk assessment, and mitigation planning, effectively prioritizing critical risks and providing clear remediation strategies. This promotes better alignment with human rights principles and enhances regulatory compliance.
This paper presents the Agent for Anti-Discriminatory Language (AAL), a legally-informed large language model designed to detect discriminatory, stereotypical, and intersectional language in Italian social media. By embedding anti-discrimination legal principles from European and Italian law into the model's behavior—via instruction tuning and expert-guided active learning—we investigate the feasibility of normatively aligned classification in high-stakes digital discourse. Using an interdisciplinary approach, we integrate structured legal knowledge, expert-annotated examples, and staged model feedback. Evaluation results demonstrate high precision in identifying overt hate speech and stereotypes, as well as the ability to generate legal justifications. However, challenges remain in identifying implicit and intersectional bias, particularly when lexical cues are weak or the social context is complex. We discuss the implications for trustworthy AI, calibrating model confidence, and integrating legal reasoning into multilingual generative systems.
Manual data annotation is often slow, expensive, and difficult to scale-especially for tasks that are subjective and context-dependent, like detecting hate speech and stereotypes. With recent progress in Large Language Models (LLMs), there is growing potential to automate this process. In this study, first, we explore the use of a committee of LLMs (GPT-4o-mini, Gemini-1.5-flash, and DeepSeek-R1) to generate annotations for Italian social media content. Then, the quality of LLM-generated labels is evaluated by following a teacher-student approach, where we trained a smaller student model (phi3.5-mini-instruct) on them and testing its performance against a human-labeled dataset. The results indicate that fine-tuning the student model over LLM-generated labels, with careful dataset balancing and hyperparameter tuning, can yield a significant improvement over the baseline (approximately 17% improvement), suggesting that committee-based LLM annotation can provide high-quality labels. Therefore, this work shows the proposed approach can be a reliable and scalable alternative to manual labeling by properly addressing challenges like class imbalance to be sure about the fairness and accuracy of the finetuned models.
The LLM-DPM Workshop investigates the transformative impact of Large Language Models (LLMs) and Explainable AI (XAI) on Data and Process Management. With organizations facing growing dependence on intricate data-driven workflows, there is an urgent demand for systems that prioritize not only efficiency but also transparency, reliability, and equity. This workshop serves as a dedicated platform to explore the convergence of LLMs, process mining, and database technologies, tackling critical issues such as enhancing data integrity, refining query understanding, forecasting process outcomes, and optimizing system performance. Featuring technical discussions, keynote speeches, and collaborative panels, LLM-DPM seeks to bridge the gap between academic research and industry applications, inspiring innovative approaches and advancing the creation of accountable, AI-powered solutions for process management challenges.
Concerning the subjectivity of genres and potential users’ preferences in streaming platforms, our study examines the capabilities of audio features integrated with user-attributed tags in improving the quality of a Recommender System. With this paper, we propose a novel approach to music genre prediction that exploits the interconnected nature of sub-genres. While we have implemented various supervised, unsupervised and semi-supervised algorithms in search of an efficient solution, the Agent-based Vector Label Propagation Algorithm (AVPRA) has yielded promising preliminary results, allowing for a more nuanced characterization of song genres and supporting the identification of sub-genres rather than or single-genre assignment.
Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers have begun to explore the use of LLMs to generate synthetic datasets as a viable alternative. This study proposes a pipeline for generating synthetic data and a comprehensive approach for investigating the factors that influence the validity of synthetic data generated by LLMs by examining how model performance is affected by metrics such as prompt strategy, text length and target position in a specific task, i.e. inclusive language detection in Italian job advertisements. Our results show that, in most cases and across different metrics, the fine-tuned models trained on synthetic data consistently outperformed other models on both real and synthetic test datasets. The study discusses the practical implications and limitations of using synthetic data for language detection tasks with LLMs.
This study introduces a structured and rigorous methodology for evaluating AI systems in line with the EU AI Act, based on the seven Trustworthy AI Principles. Unlike many existing frameworks that rely on simplistic additive scoring - where high scores in one area can mask serious deficiencies in others - our approach uses a more precise and discriminating scoring scheme. This ensures that critical weaknesses are not masked by strong performance elsewhere, thereby supporting more responsible and informed AI system design decisions. The approach combines qualitative self-assessments with quantitative indicators, allowing stakeholders to identify flaws in a system’s design, trace their origins, and monitor changes throughout the AI lifecycle.
This paper is a collaborative effort between Linguistics, Law, and Computer Science to evaluate stereotypes and biases in automated translation systems. We advocate gender-neutral translation as a means to promote gender inclusion and improve the objectivity of machine translation. Our approach focuses on identifying gender bias in English-to-Italian translations. First, we define gender bias following human rights law and linguistics literature. Then we proceed by identifying gender-specific terms such as she/lei and he/lui as key elements. We then evaluate the cosine similarity between these target terms and others in the dataset to reveal the model's perception of semantic relations. Using numerical features, we effectively evaluate the intensity and direction of the bias. Our findings provide tangible insights for developing and training gender-neutral translation algorithms.
A Language Model is a term that encompasses various types of models designed to understand and generate human communication. Large Language Models (LLMs) have gained significant attention due to their ability to process text with human-like fluency and coherence, making them valuable for a wide range of data-related tasks fashioned as pipelines. The capabilities of LLMs in natural language understanding and generation, combined with their scalability, versatility, and state-of-the-art performance, enable innovative applications across various AI-related fields, including eXplainable Artificial Intelligence (XAI), Automated Machine Learning (AutoML), and Knowledge Graphs (KG). Furthermore, we believe these models can extract valuable insights and make data-driven decisions at scale, a practice commonly referred to as Big Data Analytics (BDA). In this position paper, we provide some discussions in the direction of unlocking synergies among these technologies, which can lead to more powerful and intelligent AI solutions, driving improvements in data pipelines across a wide range of applications and domains integrating humans, computers, and knowledge.
In this paper, we propose an innovative approach to thoroughly explore dataset features that introduce bias in downstream machine-learning tasks. Depending on the data format, we use different techniques to map instances into a similarity feature space. Our method's ability to adjust the resolution of pairwise similarity provides clear insights into the relationship between the dataset classification complexity and model fairness. Experimental results confirm the promising applicability of the similarity network in promoting fair models. Moreover, leveraging our methodology not only seems promising in providing a fair downstream task such as classification, it also performs well in imputation and augmentation of the dataset satisfying the fairness criteria such as demographic parity and imbalanced classes.
Machine Learning is a powerful tool for uncovering relationships and patterns within datasets. However, applying it to a large datasets can lead to biased outcomes and quality issues, due to confounder variables indirectly related to the outcome of interest. Achieving fairness often alters training data, like balancing imbalanced groups (privileged/unprivileged) or excluding sensitive features, impacting accuracy. To address this, we propose a solution inspired by similarity network fusion , preserving dataset structure by integrating global and local similarities. We evaluate our method, considering data set complexity, fairness, and accuracy. Experimental results show the similarity network’s effectiveness in balancing fairness and accuracy . We discuss implications and future directions.
To gain a comprehensive understanding of a patient’s health, advanced analytics must be applied to the data collected by electronic health record (EHR) systems. However, managing and curating this data requires carefully designed workflows. While digitalization and standardization enable continuous health monitoring, missing data values and technical issues can compromise the consistency and timeliness of the data. In this paper, we propose a workflow for developing prognostic models that leverages the SMART BEAR infrastructure and the capabilities of the Big Data Analytics (BDA) engine to homogenize and harmonize data points. Our workflow improves the quality of the data by evaluating different imputation algorithms and selecting one that maintains the distribution and correlation of features similar to the raw data. We applied this workflow to a subset of the data stored in the SMART BEAR repository and examined its impact on the prediction of emerging health states such as cardiovascular disease and mild depression. We also discussed the possibility of model validation by clinicians in the SMART BEAR project, the transmission of subsequent actions in the decision support system, and the estimation of the required number of data points.
Natural Language Processing (NLP) algorithms have significantly advanced the capabilities of understanding, processing and generating human language. However, one persistent challenge in NLP is the problem of uncertainty, due e.g. to the inherent complexity of human language, variations in language usage across different contexts and domains, and the presence of noisy or incomplete data. In this work we take in consideration this problem for statistics derived from court documents by NLP systems.
In network models of propagation processes, the individual, microscopic level perspective is the norm, with aggregations studied as possible outcomes. On the contrary, we adopted a mesoscale perspective with groups as the core element and in this sense we present a novel agent-group dynamic model of propagation in networks. In particular, we focus on ephemeral groups that dynamically form, create new links, and dissolve. The experiments simulated 160 model configurations and produced results describing cases of consecutive and non-consecutive dynamic grouping, bounded or unbounded in the number of repetitions. Results revealed the existence of complex dynamics and multiple behaviors. An efficiency metric is introduced to compare the different cases. A Null Model analysis disclosed a pattern in the difference between the group and random models, varying with the size of groups. Our findings indicate that a mesoscopic construct like the ephemeral group, based on assumptions about social behavior and absent any microscopic level change, could produce and describe complex propagation dynamics. A conclusion is that agent-group dynamic models may represent a powerful approach for modelers and a promising new direction for future research in models of coevolution between propagation and behavior in society.
Criminal investigation adopts Artificial Intelligence to enhance the volume of the facts that can be investigated and documented in trials. However, the abstract reasoning implied in legal justification and argumentation requests to adopt solutions providing high precision, low generalization error, and retrospective transparency. Three requirements that hardly coexist in today’s Artificial Intelligence solutions. In a controlled experiment, we then investigated the use of graph embeddings procedures to retrieve potential criminal actions based on patterns defined in enquiry protocols. We observed that a significant level of accuracy can be achieved but different graph reformation procedures imply different levels of precision, generalization, and transparency.
Structural network analysis retrieves the holistic patterns of interactions among network instances. Due to the unprecedented growth of data availability, it is time to take advantage of Machine Learning to integrate the outcome of the structural analysis with better predictions on the upcoming states of large networks. Concerning the existing challenges of adopting methods embracing multi-dimensional, multi-task, transparent representations within incremental procedures, in our recent study, we proposed the AVPRA algorithm. It works as an embedder of both the network structure and domain-specific features making the aforementioned challenges feasible to address. In this paper, we elaborate on the validation of AVPRA by adopting it in multiple downstream Machine Learning tasks on the Twitter network of the Italian Parliament. Comparing the outcome with state-of-the-art algorithms of graph embedding, the capability of AVPRA in retaining either network structure properties or domain-specific features of the nodes is promising. In addition, the method is incremental and transparent.
Anisio Azzini合作论文数Department of Information Technology
University of Milan1
Maurice Van Keulen合作论文数University of Twente;Mathematics and Computer Science (EEMCS);Faculty of Electrical Engineering1