Legal text summarization is increasingly relevant due to the growing availability of digital legal documents. In this study, we address the challenges of summarizing judgments from the Portuguese Supreme Court of Justice (STJ) by leveraging domain-adapted embedding models tailored to the Portuguese legal domain, which lacks a robust ecosystem of specialized embeddings. This paper presents the development of domain-specific embeddings for Portuguese legal language, designed to capture the semantics of jurisprudential statements with higher precision. Beyond proposing this resource, the study includes an empirical validation through expert (judge) assessment, demonstrating the embeddings’ applicability to real judicial reasoning. Our findings also demonstrate that domain-specific embeddings significantly enhance the quality of extractive summaries, even when employing simple algorithms. Additionally, we propose a segmentation model that identifies the inherent structure and roles of different sections within judgments, a critical step given their extensive length and varying purposes. This model further improves the quality and relevance of generated summaries. Our approach achieved a ROUGE-1 score of 57.23 and a ROUGE-2 score of 25.44. We evaluated summaries for coherence, completeness, and accuracy using state-of-the-art large language models (Llama and GPT-4o). The LLM-based ratings show limited agreement with ROUGE and with each other, and a cross-evaluation analysis combining these signals with feedback from STJ judges highlights the limitations of any single metric for evaluating legal summaries. Finally, a human evaluation conducted by STJ judges provided valuable insights into the strengths and limitations of our approach.
Journalistic manual fact-checking is the usual way to address fake news; however, this labor-intensive task regularly is not a match for the scale of the problem. The literature introduced automated fact-checking (AFC) as a potential solution; however, there is still missing functionality in the AFC pipeline, a lack of research benchmarking data, and a disconnect between their design and human factors crucial for adoption. We present a fully explainable AFC framework designed to augment professional journalists in the wild. A novel human annotation-free approach surpasses state-of-the-art multi-label classification by 12%. It is the first to demonstrate strong generalization across different claim subjects without retraining and to generate complete verdict explanation articles and their summaries. A focused user study of 103 professional journalists, with 93% having dedicated experience with fact-checking, validates the framework's level of explainability, transparency, and quality of generated fact-checking artifacts. The importance of establishing clear source selection and bias evaluation criteria reinforced the need for human augmentation, not replacement, by AFC systems.
Extreme Multi-label Classification (XML) involves predicting multiple labels for a given input, a fundamental problem in domains such as text categorization, recommendation systems, and image tagging. This task presents significant challenges for machine learning and information retrieval, particularly given the exponential growth of online data and the concomitant need for algorithms capable of handling large-scale datasets with numerous labels. Traditional classification methods are inadequate for this task due to the vast number of possible label combinations and the sparsity of label assignments. This paper reports the results of a project with the Supreme Court of Justice of Portugal ("Supremo Tribunal de Justica Portugues") to address the problem using Sparse Local Embeddings for Extreme Multi-label Classification (SLEEC), an embedding-based approach that showed promising results in legal datasets. Our goal was to associate descriptors, which categorize court judgments, with the judgments themselves. This work tackled various challenges, including a large number of descriptors, an unbalanced dataset, numerous tail labels, and extensive document lengths. Our experimental results demonstrate that our approach achieved a precision/recall variation ranging between 0.57 and 0.68, indicating promising performance in this complex task.
The rapid advancements in legal text summarization have not been matched by equivalent progress in evaluation metrics capable of assessing the quality of legal summaries. Traditional evaluation approaches, such as ROUGE, remain widely used despite their inability to capture semantic fidelity. While more recent metrics focus on semantic evaluation, their applicability to legal summarization has not been thoroughly tested, and their performance is highly dependent on embedding models and computational resources, particularly for long and complex legal texts. Furthermore, the absence of publicly available datasets with expert annotations hinders the development and validation of domain-specific evaluation methods. In this paper, we address these challenges by introducing the first publicly available dataset of Portuguese legal summaries, annotated by legal experts across multiple dimensions such as Coherence and Relevance. We use this dataset to systematically evaluate several recent evaluation metrics, comparing their performance against ROUGE, the standard metric for summarization tasks. Our analysis, based on Spearman correlation with human judgments, reveals that ROUGE-2 maintains the highest correlation across almost every evaluated dimension, outperforming more recent metrics, including semantic-based approaches. These results emphasize the challenges of adapting new evaluation frameworks to the legal domain and underscore the need for further research into metrics that can better capture domain-specific requirements.
A Classificação Extrema Multi-etiqueta (XML) consiste na predição de múltiplas etiquetas para um determinado input, sendo um problema fundamental em domínios como categorização de texto, sistemas de recomendação e marcação de imagens. Esta tarefa apresenta desafios significativos para a aprendizagem automática e a recuperação de informação, especialmente devido ao crescimento exponencial de dados online e à consequente necessidade de algoritmos capazes de lidar com conjuntos de dados de grande escala e com um elevado número de etiquetas. Os métodos tradicionais de classificação são inadequados para esta tarefa devido ao vasto número de possíveis combinações de etiquetas e à dispersão das atribuições. Este artigo apresenta os resultados de um projeto realizado com o Supremo Tribunal de Justiça de Portugal, onde abordámos este problema utilizando Sparse Local Embeddings for Extreme Multi-label Classification (SLEEC), uma abordagem baseada em embeddings que demonstrou resultados promissores no domínio legal. O nosso objetivo foi associar descritores, que categorizam os acórdãos do tribunal Português, aos respetivos acórdãos. Este trabalho enfrentou diversos desafios, nos quais se incluem um elevado número de descritores, um conjunto de dados desbalanceado, a presença de muitas etiquetas raras (tail labels) e a extensão considerável dos documentos. Os resultados experimentais demonstram que a nossa abordagem alcançou uma variação de precisão/cobertura entre 0,57 e 0,68, indicando um desempenho promissor nesta tarefa complexa.
Legal document segmentation is a critical task in the field of natural language processing (NLP), enabling efficient analysis, retrieval, and understanding of legal content. Despite its importance, research in this area for European Portuguese has been limited. To address this gap, we present a novel approach to automate the segmentation of legal judgments from the Portuguese Supreme Court of Justice into distinct sections. Leveraging a Bi-LSTM-CRF model, we developed a dataset and achieved significant results, including an accuracy of 0.9997, precision of 0.9986, recall of 0.996, and F1-Score of 0.9973. Our methodology and experimental results demonstrate the effectiveness and potential applications of our approach for the European Portuguese language.
Fake news has been linked to the rise of psychological disorders, the increased disbelief in science, and the erosion of democracy and freedom of speech. Online social networks are arguably the main vehicle of fake news spread. Educating online users with explanations is one way of preventing this spread. Understanding how online belief is formed and changed may offer a roadmap for such education. The literature includes surveys addressing online opinion formation and polarization; however, they usually address a single domain, such as politics, online marketing, health, and education, and do not make online belief change their primary focus. Unlike other studies, this work is the first to present a cross-domain systematic literature review of user studies, methodologies, and opinion model dimensions. It also includes the orthogonal polarization dimension, focusing on online belief change. We include peer-reviewed works published in 2020 and later found in four relevant scientific databases, excluding theoretical publications that did not offer validation through dataset experimentation or simulation. Bibliometric networks were constructed for better visualization, leading to the organization of the papers that passed the review criteria into a comprehensive taxonomy. Our findings show that a person’s individuality is the most significant influential force in online belief change. We show that online arguments that balance facts with emotionally evoking content are more efficient in changing their beliefs. Polarization was shown to be cross-correlated among multiple subjects, with politics being the central polarization pole. Polarized online networks start as networks with high opinion segregation, evolve into subnetworks of consensus, and achieve polarization around social network influencers. Trust in the information source was demonstrated to be the chief psychological construct that drives online users to polarization. This shows that changing the beliefs of influencers may create a positive snowball effect in changing the beliefs of polarized online social network users. These findings lay the groundwork for further research on using personalized explanations to reduce the harmful effects of online fake news on social networks.
Legal documents are commonly known for being lengthy and having a specific vocabulary. For professionals and non-jurists, having a summary of each document is crucial so they can use it as a reference for other cases without spending too much time reading the entire document. In the Portuguese Supreme Court of Justice, summaries are done manually, by its Judges which is very time-consuming because of the length of the legal documents. Aiming to support the Judges in this task, the goal of this work is to investigate how different techniques and methods of automated text summarization can achieve good performance on Portuguese legal documents.
Anxiety disorders are among the most prevalent mental health disorders. In various settings, patients are using apps for the relief of their symptoms. Digital applications are recommended by international guidelines, but most apps available are not scientifically validated. Little is known about the profile of adult users searching for mental health apps, in the general population. In this study, we propose to identify the user profile of real-world users searching for mental health apps to reduce anxiety. We used the app TUYB Ansiedade® for this study and divided the users in three groups: health professional referral, social network and search engine. We got 453 completed first assessments, indicating that the users that found the app through Social Networks have higher values of anxiety, depression and lower perceived quality of life and resilience. Our results support the conclusion that users searching and engaging with apps in social networks have a lower mental well-being profile and this should inform app developers, guidelines on app use for mental health and policy makers. The policies regarding Social Networks and app developers should be improved to account for the mental health of their users.
Argumentation Mining (AM) is a growing sub-field of Natural Language Processing (NLP) which aims at extracting argumentative structures from text. In this work, neural learning and symbolic reasoning are combined in a system named N-SAUR, that extracts the argumentative structures present in a collection of texts, and then assesses each argument’s strength. The extraction is based on Toulmin’s model and the result quality surpasses previous approaches over an existing benchmark. Complementary scores are also extracted and combined with a set of rules that produce the final calculation of argument strength. The performance of the system was evaluated through human assessments. Users can also interact with the system in various ways, allowing for the strength calculation to change through user-cooperative reasoning.
AbstractIn the current digital era, language technologies are playing an increasingly vital role in the legal domain, assisting users, lawyers, judges, and legal professionals to solve many real-world problems. While open datasets and innovative deep learning methodologies have led to recent breakthroughs in the area, significant efforts are still being made to transfer the theoretical/algorithmic developments, associated with general text and speech processing, into real applications in the legal-domain. This chapter presents a brief survey on language technologies for addressing legal tasks, covering studies and applications related to both text and speech processing (Manuscript submitted in May 2022).
Fake news has become a serious and destabilizing problem in our increasingly polarized society. The core tasks of detecting and characterizing untrue or misleading online content are quite challenging. News consumers intervention has been identified as a crucial addition to Fake News Detection Systems (FNDS). However, even though detection explanation is starting to gain research momentum, not as much attention has been given to personalization of explanations. As humans are the obvious targets of fake news, explanations that can evoke emotional responses and/or are aligned with individual personality traits and cognitive styles can be leveraged to nudging of the news consumer into a reflective state about the subject, which has been shown to be more effective than the crude presentation of facts in changing pre-conceived beliefs. This paper adds to the main goal of misinformation detection systems, aiming to expand them onto personalized FNDS. It proposes a metric to be used in their evaluation. It offers a definition to help research on the implementation of personalized fake news explanations. And finally, it proposes a personalized fake news taxonomy, discussing its components centered around emotion-based and personality-based explanations. This taxonomy highlights several opportunities for those researching in the area of personalized fake news explanation systems.
Psychosis is a brain condition that affects the subject and the way it perceives the world around, impairing its cognitive and speech capabilities, and creating a disconnection from reality in which the subject is inserted. Psychosis lacks formal and pre-cise diagnostic tools, relying on self-reports from patients, their families, and specialized clinicians. Previous studies have fo-cused on the identification and prediction of psychosis through surface-level analysis of diagnosed patients targeting audio, time, and paucity features to predict or identify psychosis. More recent studies have started focusing on high-level and complex language analysis such as semantics, structure, and pragmatics. Only a reduced number of studies have targeted the Portuguese language. Currently, no study has targeted structural or semantic features in European Portuguese, thus this is our objective. The results obtained through our work suggest that the use of structural and semantic features, particularly for European Portuguese, holds some power in classifying subjects as diagnosed with psychosis or not. However, further research is required to identify possible improvements to the techniques employed and to concretely identify which particular features hold the most power during the classification tasks.
In this paper, several methodologies were applied to automatically detect feeding activity in the Oceanário de Lisboa Main Aquarium, using videos acquired on the outside of the aquarium by a static camera. We propose three methods. The first one is based on Convolutional Neural Networks (CNN) and learns to detect the feeding patterns at each frame based on training data. The second one is based on motion variability, computed either from frame difference or optical flow. The third one uses the analysis of spatial patterns (aggregation) formed by fish present in each frame, assuming their prior detection. For the development of the methods and quantitative evaluation, several videos were filmed at Oceanário de Lisboa. To evaluate each of the approaches, several metrics are extracted such as accuracy, precision, and recall. We analysed videos of the feeding patterns of rays (bottom feeding) and sharks (surface feeding). These different patterns are quite distinct in terms of motion and aggregation of fish. For bottom feeding, we concluded that the frame difference approach was the best performing, but the CNN and aggregation methods showed competitive results. For surface feeding, the CNN was the only method able to perform with reasonable accuracy.
Psychosis is a clinical syndrome characterized by the presence of symptoms such as hallucinations, thought disorder and disorganized speech. Several studies have used machine learning, combined with speech and natural language processing methods to aid in the diagnosis process of this disease. This paper describes the creation of the first European Portuguese corpus for the identification of the presence of speech characteristics of psychosis, which contains samples of 92 participants, 56 controls and 36 individuals diagnosed with psychosis and medicated. The corpus was used in a set of experiments that allowed identifying the most promising feature set to perform the classification: the combination of acoustic and speech metric features. Several classifiers were implemented to study which ones entailed the best performance depending on the task and feature set. The most promising results obtained for the entire corpus were achieved when identifying individuals with a Multi-Layer Perceptron classifier and reached an 87.5% accuracy. Focusing on the gender dependent results, the overall best results were 90.9% and 82.9% accuracy, for female and male subjects respectively. Lastly, the experiments performed lead us to conjecture that spontaneous speech presents more identifiable characteristics than read speech to differentiate healthy and patients diagnosed with psychosis.
In this article, a system that takes a 3D model of a sculpture as starting point to compose music is presented. We raised the hypothesis that cross-domain mapping can be an approach to model inspiration. The semantic meaning of the sculpture is not used directly but rather a more abstract approach was used. A Genetic Algorithm was used to obtain results with more musical interest. The results were promising: the majority of the participants gave a classification of 4 out of 5 to the preferred interpretations of the compositions and related them to the respective sculpture. This is a step toward a possible model for inspiration.
In this paper, we present an inspirational system that takes a 3D model of a sculpture as starting point to compose music. It is considered that cross-domain mapping can be an approach to model inspiration. Our approach does not consider the interpretation of the sculpture but rather looks at it abstractly. The results were promising: the majority of the participants gave a classification of 4 out of 5 to the preferred interpretations of the compositions and related them to the respective sculpture. This is a step to a possible model for inspiration.
Christoph Tempich合作论文数Universitat Karlsruhe14