Diversity, Equity, and Inclusion (DEI) policies have recently become extremely controversial, with many companies vowing to end their support. This has led to mixed reactions online. This was intensified by the ongoing “woke” vs “anti-woke” culture war. Both groups defend and consume content that aligns with their ideologies. In this context, understanding the discourse surrounding these issues online is essential, as such movements have the potential to lead to real-world harm. For this reason, we conduct a large-scale study around the DEI and “woke” discussion on the Reddit platform from 2020–2024, finding that it has grown significantly during the studied period, spreading across a large variety of seemingly unrelated topics. Finally, we note that the discourse has become increasingly polarized, with a growing trend of toxicity and negative sentiments, coupled with changes in the meaning of the terms “woke” and DEI on the platform. These findings have important implications for public policy related to social issues.
This paper investigates how Responsible AI principles can be systematically integrated into public health predictive modeling using data from the Brazilian Unified Health System. We operationalize a four-layer architecture embedding governance, leakage-controlled temporal validation, fairness auditing, explainability, and structured documentation into the modeling lifecycle. In predicting monthly respiratory hospitalizations, LightGBM achieved an RMSE of 13.81 but exhibited a 44.8 percent sMAPE disparity in the highest disparity state. A resampling adjustment reduced this gap by 8.51 percentage points. The results demonstrate how structured Responsible AI controls reshape evaluation beyond aggregate predictive accuracy.
O racismo se manifesta de maneiras complexas nas mídias sociais, exigindo abordagens eficazes para detecção automatizada. Este estudo contribui construindo um novo conjunto de dados de racismo anotado por pesquisadores negros, garantindo representatividade na rotulagem de postagens racistas direcionadas à população negra na plataforma X. Avaliamos o desempenho de modelos de aprendizado de máquina tradicionais (Naive Bayes, Regressão Logística, Random Forest e XGBoost) e modelos baseados em Transformer, como BERTimbau, voltado para a língua Portuguesa. Embora BERTimbau tenha alcançado uma pontuação F1 de 0,83, indicando razoável eficácia, não superou modelos mais simples, como Regressão Logística e Naive Bayes. Os resultados evidenciam desafios na detecção automatizada de racismo online, como a escassez de dados anotados e as complexidades linguísticas, incluindo ironia, sarcasmo e ambiguidades. Análises de erros revelam que esses fatores de fato impactam a eficácia dos classificadores, sugerindo a necessidade de métodos mais robustos para identificar o racismo em português. Aviso de conteúdo: Este artigo contém exemplos de frases racistas. As postagens incluídas exemplificam os desafios encontrados no processo de classificação dos dados.
Social media networks have amplified the reach of social and political movements, but most research focuses on mainstream platforms such as X, Reddit, and Facebook, overlooking Discord. As a rapidly growing, community-driven platform with optional decentralized moderation, Discord offers unique opportunities to study political discourse. This study analyzes over 30 million messages from political servers on Discord discussing the 2024 U.S. elections. Servers were classified as Republican-aligned, Democratic-aligned, or unaligned based on their descriptions. We tracked changes in political conversation during key campaign events and identified distinct political valence and implicit biases in semantic association through embedding analysis. We observed that Republican servers emphasized economic policies, while Democratic servers focused on equality-related and progressive causes. Furthermore, we detected an increase in toxic language, such as sexism, in Republican-aligned servers after Kamala Harris's nomination. These findings provide a first look at political behavior on Discord, highlighting its growing role in shaping and understanding online political engagement.
Large Language Models (LLMs) are increasingly used in applications that shape public discourse, yet little is known about whether they reflect distinct opinions on global issues like climate change. This study compares climate change-related responses from multiple LLMs with human opinions collected through the People's Climate Vote 2024 survey (UNDP - United Nations Development Programme and Oxford, 2024). We compare country and LLM's answer probability distributions and apply Exploratory Factor Analysis (EFA) to identify latent opinion dimensions. Our findings reveal that while LLM responses do not exhibit significant biases toward specific demographic groups, they encompass a wide range of opinions, sometimes diverging markedly from the majority human perspective.
Since 2022 we have been exploring application areas and technologies in which Artificial Intelligence (AI) and modern Natural Language Processing (NLP), such as Large Language Models (LLMs), can be employed to foster the usage and facilitate the documentation of Indigenous languages which are in danger of disappearing. We start by discussing the decreasing diversity of languages in the world and how working with Indigenous languages poses unique ethical challenges for AI and NLP. To address those challenges, we propose an alternative development AI cycle based on community engagement and usage. Then, we report encouraging results in the development of high-quality machine learning translators for Indigenous languages by fine-tuning state-of-the-art (SOTA) translators with tiny amounts of data and discuss how to avoid some common pitfalls in the process. We also present prototypes we have built in projects done in 2023 and 2024 with Indigenous communities in Brazil, aimed at facilitating writing, and discuss the development of Indigenous Language Models (ILMs) as a replicable and scalable way to create spell-checkers, next-word predictors, and similar tools. Finally, we discuss how we envision a future for language documentation where dying languages are preserved as interactive language models.
Estimativas automáticas do tom de pele enfrentam desafios devido a vieses raciais e de gênero em abordagens de aprendizado de máquina. Neste trabalho, exploramos um conjunto de dados rotulado e avaliamos duas abordagens computacionais amplamente exploradas na literatura, ITA e CASCo, a fim de investigar a robustez e limitações nessa tarefa. Nossos resultados mostram que essas abordagens ainda apresentam falhas significativas, comprometendo sua aplicação em contextos reais onde a precisão é essencial.
This work analyzes the differences in how Oscar-nominated films in the ‘Best Picture’ category are addressed in the English, Spanish and Portuguese versions on Wikipedia. Using text and graph analysis techniques, the connectivity and semantics of the articles in these languages were investigated, revealing that each language highlights different aspects of the films, reflecting the cultural and linguistic priorities of their respective communities.
Nearly half of Brazil's 180 Indigenous languages face extinction within the next 20 years. What's more concerning is that most of these languages lack a single scientific article describing them, which means they could disappear without leaving any documented evidence of their existence. This work investigates the state of articles about those languages in Wikipedia, both in the English and Portuguese versions, regarded here as indicative of the minimum world-level trace of the previous existence of these languages. Our study shows that over 30% of these languages do not have a single Wikipedia article describing them. It also highlights that the Portuguese and English editing communities are not only distinct, but have different practices, achieving similar levels of quality through different temporal dynamics. These results, although encouraging, suggest that any effort to enhance coverage comprehensiveness in both Wikipedias should consider different strategies for engaging each editing community.
In this paper we discuss how AI can contribute to support the documentation and vitalization of Indigenous languages and how that involves a delicate balancing of ensuring social impact, exploring technical opportunities, and dealing with ethical constraints. We start by surveying previous work on using AI and NLP to support critical activities of strengthening Indigenous and endangered languages and discussing key limitations of current technologies. After presenting basic ethical constraints of working with Indigenous languages and communities, we propose that creating and deploying language technology ethically with and for Indigenous communities forces AI researchers and engineers to address some of the main shortcomings and criticisms of current technologies. Those ideas are also explored in the discussion of a real case of development of large language models for Brazilian Indigenous languages.
Este artigo apresenta um conjunto de dados desenvolvido a partir das informações disponibilizadas nas revistas de registro de patentes do Instituto Nacional de Propriedade Industrial (INPI). Devido à fragmentação das informações em múltiplas revistas, torna-se inviável realizar buscas ou fazer análises sobre o cenário de registro de patentes no Brasil. O conjunto de dados apresentado neste artigo, agrega as informações fragmentadas dos processos em um repositório no GitHub e apresenta um fluxo automatizado para obtenção de novas revistas, utilizando GitHub Actions com o objetivo de manter os dados atualizados e relevantes.
The slow increasing rate of representation of women and other minorities, particularly in Information Technology (IT) companies, suggests that the current recruiting strategies for attracting underrepresented minorities (URMs) may not be effective. While much can be done to improve hiring strategies, little attention has been paid to how job seekers’ perceptions (positive or otherwise) of a company may affect its attractiveness. For instance, newer generations of professionals are more likely to avoid companies that are not committed to diversity causes. Studies have in fact shown that job seekers increasingly rely on social media to inform themselves about targeted companies. In this paper, we investigate the “social-mediascape” of Black Tech communities and IT companies, namely, the spaces of interactions and communications and their influences on individuals’ perceptions of particular companies. To this end, we look into the Twitter Black and Tech communities in Brazil and in the United States. We rely on “Twitter lists” that are curated by the users of the platform and effective in capturing topical homophily. A research challenge in itself, we provide the first large-scale compositions of Black Tech communities and their connection to IT companies based on followership data on Twitter. After analyzing these compositions, we can then create perceptual maps between communities and IT companies. Our results suggest that there is a stronger correlation between Black activism and technology in the US context. For the Brazilian context, we found a stronger correlation between the Black Tech and general Software Developer users communities suggesting that both racial activism and technology are important topics to attract the Brazilian Black Tech community’s interests.
How can we best address the dangerous impact that deep learning-generated fake audios, photographs, and videos (a.k.a. deepfakes) may have in personal and societal life? We foresee that the availability of cheap deepfake technology will create a second wave of disinformation where people will receive specific, personalized disinformation through different channels, making the current approaches to fight disinformation obsolete. We argue that fake media has to be seen as an upcoming cybersecurity problem, and we have to shift from combating its spread to a prevention and cure framework where users have available ways to verify, challenge, and argue against the veracity of each piece of media they are exposed to. To create the technologies behind this framework, we propose that a new Science of Disinformation is needed, one which creates a theoretical framework both for the processes of communication and consumption of false content. Key scientific and technological challenges facing this research agenda are listed and discussed in the light of state-of-art technologies for fake media generation and detection, argument finding and construction, and how to effectively engage users in the prevention and cure processes.
Developing and middle-income countries increasingly emphasize higher education and entrepreneurship in their long-term development strategy. Thus, our work focuses on the influence of higher education institutions (HEIs) on startup ecosystems in Brazil, an emerging economy. As traditional data to perform this type of study, such as surveys, are challenging to get, we propose an alternative approach. Given the growing capability of social media databases such as Crunchbase and LinkedIn to provide startup and individual-level data, we draw on computational methods to mine data for social network analysis. Our approach enables different types of analysis. First, we describe regional variability in entrepreneurial network characteristics. Second, we examine the influence of elite HEIs in economic hubs on entrepreneur networks. Third, we investigate the influence of the academic trajectories of startup founders, including their courses of study and HEIs of origin, on the fundraising capacity of startups. We find that HEI quality and the maturity of the ecosystem influence startup success. We also observe that elite HEIs have a powerful influence on local entrepreneur ecosystems. Surprisingly, while the most nationally prestigious HEIs in the South and Southeast have the longest geographical reach, their network influence remains local. This means that investments in entrepreneurship, in the Brazilian context, tend to remain concentrated in wealthier cities, and may actually reinforce or increase regional inequalities. We also find that the startup ecosystem in the wealthier South and Southeast is more diverse in terms of sectors, which is more advantageous to economic development. Our approach can be helpful, especially in countries with limited studies of the interaction between startups and institutional factors supporting them. In terms of policy recommendations, we would recommend more investment at the regional level in terms of cultivating entrepreneurship, given the limited spillover from wealthier regions.
Chatbots have received significant attention in the last years. These systems have improved operational efficiency by reducing the cost of customer service and more and more are used in customer service channels. More recently, with the addition of speech capabilities added to those systems, bot language modifications have to be done in order to improve user experience. In this paper, we analyze the language used in those systems using datasets collected from real chatbots. For that, we first propose models that are able to identify the language style (writing, speech, and computer-mediated) using as linguist features syntactic (Biber's dimensions) and sentence embeddings as well as state-of-art datasets. Our results show that our models were able to distinguish among the three classes with an accuracy of up to 66%. Finally, we evaluate real chatbots systems, either speech- and text-based ones, using our proposed models. We found that these real chatbots are generally using a computer-mediated style, but there is indication that their developers tend to adapt the language style to be more speech-like when text-to-speech is used.
O Twitter é uma das redes sociais online mais utilizadas pelo público em geral para o debate de diferentes assuntos, sendo a pandemia da COVID-19 um dos temas mais debatidos nos últimos meses. A partir de dezembro de 2020, o foco do debate passou a ser a vacinação. Neste artigo, investigamos a percepção do público sobre o tema, analisando mais de 9 milhões de tweets em português, em um período de dois meses correspondentes aos estágios iniciais da vacinação no Brasil e no mundo. Nossos resultados fornecem um entendimento inicial sobre a dinâmica do debate online sobre a vacinação contra a COVID-19, evidenciando como as pessoas usam o mundo online para compartilhar suas impressões e preocupações sobre o assunto.
Workforce diversification is essential to increase productivity in any world economy. In the context of the Fourth Industrial Revolution, that need is even more urgent since technological sectors are men-dominated. Despite the significant progress made towards gender inequality in the last decades, we are far from the ideal scenario. Changes towards equality are too slow and uneven across different world regions. Monitoring gender parity is essential to understand priorities and specificities in each world region. However, it is challenging because of the scarcity and the cost to obtain data, especially in less developed countries. In this paper we study how the Facebook Advertising Platform (Facebook Ads) can be used to assess gender imbalance in education, focusing on STEM (Science, Technology, Engineering, and Mathematics) areas, which are the main focus of the Fourth Revolution. As a case study, we apply our methodology to characterize Brazil in terms of gender balance in STEM as well as to correlate the results using Facebook Ads data with official Brazilian government numbers. Our results suggest that even considering a biased population where the majority is female, the proportion of men interested in some majors is higher than the proportion of women. Within STEM areas, we can identify two different patterns. Life Science and Math/Physical Sciences have female dominance, Environmental Science, Technology, and Engineering majors are still concentrated towards men. We also assess the impact of educational level and age on the interest in majors. The gender gap in STEM increases with the women's educational level and age, as confirmed by official data in Brazil.